Prosecution Insights
Last updated: October 01, 2026
Application No. 19/008,192

METHOD AND SYSTEM FOR GENERATING INDEXED CORPUS FOR DOMAIN-DRIVEN KNOWLEDGE AUGMENTED QUESTION ANSWERING

Non-Final OA §101
Filed
Jan 02, 2025
Priority
Jan 08, 2024 — IN 202421001329
Examiner
KIM, ETHAN DANIEL
Art Unit
Tech Center
Assignee
Tata Group
OA Round
1 (Non-Final)
77%
Grant Probability
Favorable
1-2
OA Rounds
1y 1m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 77% — above average
77%
Career Allowance Rate
88 granted / 114 resolved
+17.2% vs TC avg
Strong +25% interview lift
Without
With
+25.4%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
10 currently pending
Career history
131
Total Applications
across all art units

Statute-Specific Performance

§101
10.0%
-30.0% vs TC avg
§103
49.6%
+9.6% vs TC avg
§102
35.2%
-4.8% vs TC avg
§112
1.4%
-38.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 114 resolved cases

Office Action

§101
Notice of Pre-AIA or AIA Status 1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement 2. The information disclosure statement (IDS) submitted on January 02, 2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Rejections - 35 USC § 101 3. 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent thereof, subject to the conditions and requirements of this title. 4. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The independent claim 1 recites “A processor implemented method, comprising: receiving, via one or more hardware processors, a raw corpus data as input, wherein the raw corpus data is a domain specific data; and generating, via the one or more hardware processors, an indexed corpus from the raw corpus data, comprising: transforming the raw corpus data into an associated plain text document, wherein a file type of the associated plain text document is determined based on a determined file type of the raw corpus data; normalizing the associated plain text document, comprising standardizing and cleaning the plain text document; generating a plurality of document chunks associated with the normalized associated plain text document, comprising: determining a document type of the plain text document based on a calculated keyword score and a large language model (LLM) score; and performing a section-based chunking, a paragraph-based chunking and a sentence-based chunking on the plain text document to obtain the plurality of document chunks, based on the determined document type; generating a plurality of domain centric chunks from the plurality of document chunks using a domain model graph and the LLM, wherein the domain model graph comprises a set of associations and a root node; aggregating the plurality of domain centric chunks by grouping similar chunks from the plurality of domain centric chunks; and generating the indexed corpus by performing indexing on the aggregated plurality of domain centric chunks using vector embeddings”. The limitations “receiving, via one or more hardware processors, a raw corpus data as input, wherein the raw corpus data is a domain specific data; and generating, via the one or more hardware processors, an indexed corpus from the raw corpus data, comprising: transforming the raw corpus data into an associated plain text document, wherein a file type of the associated plain text document is determined based on a determined file type of the raw corpus data; normalizing the associated plain text document, comprising standardizing and cleaning the plain text document; generating a plurality of document chunks associated with the normalized associated plain text document, comprising: determining a document type of the plain text document based on a calculated keyword score and a large language model (LLM) score; and performing a section-based chunking, a paragraph-based chunking and a sentence-based chunking on the plain text document to obtain the plurality of document chunks, based on the determined document type; generating a plurality of domain centric chunks from the plurality of document chunks using a domain model graph and the LLM, wherein the domain model graph comprises a set of associations and a root node; aggregating the plurality of domain centric chunks by grouping similar chunks from the plurality of domain centric chunks; and generating the indexed corpus by performing indexing on the aggregated plurality of domain centric chunks using vector embeddings” as drafted, covers a mental process, as this could be done by mentally or by hand with pen and paper. This judicial exception is not integrated into a practical application. Claim 1 recites “A processor implemented method, comprising:…” this limitation directs towards using a computer for the method, and does not impose any meaningful limits on practicing the abstract idea. The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception. The addition of the generic computer components recited above with regards to claim 1 do not amount to more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Claim 1 does not recite any additional limitations. The claim as drafted, is not patent eligible. The independent claim 8 recites “A system, comprising: one or more hardware processors; a communication interface; and a memory storing a plurality of instructions, wherein the plurality of instructions cause the one or more hardware processors to: receive a raw corpus data as input, wherein the raw corpus data is a domain specific data; and generate an indexed corpus from the raw corpus data, by: transforming the raw corpus data into an associated plain text document, wherein a file type of the associated plain text document is determined based on a determined file type of the raw corpus data; normalizing the associated plain text document, comprising standardizing and cleaning the plain text document; generating a plurality of document chunks associated with the normalized associated plain text document, by: determining a document type of the plain text document based on a calculated keyword score and a large language model (LLM) score; and performing a section-based chunking, a paragraph-based chunking and a sentence-based chunking on the plain text document to obtain the plurality of document chunks, based on the determined document type; generating a plurality of domain centric chunks from the plurality of document chunks using a domain model graph and the LLM, wherein the domain model graph comprises a set of associations and a root node; aggregating the plurality of domain centric chunks by grouping similar chunks from the plurality of domain centric chunks; and generating the indexed corpus by performing indexing on the aggregated plurality of domain centric chunks using vector embeddings”. The limitations “receive a raw corpus data as input, wherein the raw corpus data is a domain specific data; and generate an indexed corpus from the raw corpus data, by: transforming the raw corpus data into an associated plain text document, wherein a file type of the associated plain text document is determined based on a determined file type of the raw corpus data; normalizing the associated plain text document, comprising standardizing and cleaning the plain text document; generating a plurality of document chunks associated with the normalized associated plain text document, by: determining a document type of the plain text document based on a calculated keyword score and a large language model (LLM) score; and performing a section-based chunking, a paragraph-based chunking and a sentence-based chunking on the plain text document to obtain the plurality of document chunks, based on the determined document type; generating a plurality of domain centric chunks from the plurality of document chunks using a domain model graph and the LLM, wherein the domain model graph comprises a set of associations and a root node; aggregating the plurality of domain centric chunks by grouping similar chunks from the plurality of domain centric chunks; and generating the indexed corpus by performing indexing on the aggregated plurality of domain centric chunks using vector embeddings” as drafted, covers a mental process, as this could be done by mentally or by hand with pen and paper. This judicial exception is not integrated into a practical application. Claim 8 recites “A system, comprising: one or more hardware processors; a communication interface; and a memory storing a plurality of instructions, wherein the plurality of instructions cause the one or more hardware processors to:…” this limitation directs towards using a computer for the method, and does not impose any meaningful limits on practicing the abstract idea. The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception. The addition of the generic computer components recited above with regards to claim 8 do not amount to more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Claim 8 does not recite any additional limitations. The claim as drafted, is not patent eligible. The independent claim 15 recites “One or more non-transitory machine-readable information storage mediums comprising one or more instructions which when executed by one or more hardware processors cause: receiving a raw corpus data as input, wherein the raw corpus data is a domain specific data; and generating an indexed corpus from the raw corpus data, comprising: transforming the raw corpus data into an associated plain text document, wherein a file type of the associated plain text document is determined based on a determined file type of the raw corpus data; normalizing the associated plain text document, comprising standardizing and cleaning the plain text document; generating a plurality of document chunks associated with the normalized associated plain text document, comprising: determining a document type of the plain text document based on a calculated keyword score and a large language model (LLM) score; and performing a section-based chunking, a paragraph-based chunking and a sentence-based chunking on the plain text document to obtain the plurality of document chunks, based on the determined document type; generating a plurality of domain centric chunks from the plurality of document chunks using a domain model graph and the LLM, wherein the domain model graph comprises a set of associations and a root node; aggregating the plurality of domain centric chunks by grouping similar chunks from the plurality of domain centric chunks; and generating the indexed corpus by performing indexing on the aggregated plurality of domain centric chunks using vector embeddings”. The limitations “receiving a raw corpus data as input, wherein the raw corpus data is a domain specific data; and generating an indexed corpus from the raw corpus data, comprising: transforming the raw corpus data into an associated plain text document, wherein a file type of the associated plain text document is determined based on a determined file type of the raw corpus data; normalizing the associated plain text document, comprising standardizing and cleaning the plain text document; generating a plurality of document chunks associated with the normalized associated plain text document, comprising: determining a document type of the plain text document based on a calculated keyword score and a large language model (LLM) score; and performing a section-based chunking, a paragraph-based chunking and a sentence-based chunking on the plain text document to obtain the plurality of document chunks, based on the determined document type; generating a plurality of domain centric chunks from the plurality of document chunks using a domain model graph and the LLM, wherein the domain model graph comprises a set of associations and a root node; aggregating the plurality of domain centric chunks by grouping similar chunks from the plurality of domain centric chunks; and generating the indexed corpus by performing indexing on the aggregated plurality of domain centric chunks using vector embeddings” as drafted, covers a mental process, as this could be done by mentally or by hand with pen and paper. This judicial exception is not integrated into a practical application. Claim 15 recites “One or more non-transitory machine-readable information storage mediums comprising one or more instructions which when executed by one or more hardware processors cause:…” this limitation directs towards using a computer for the method, and does not impose any meaningful limits on practicing the abstract idea. The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception. The addition of the generic computer components recited above with regards to claim 15 do not amount to more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Claim 15 does not recite any additional limitations. The claim as drafted, is not patent eligible. Allowable Subject Matter 5. Claims 1-20 have been rejected under 35 U.S.C. 101, but would be allowable if the rejection is overcome. The following is a statement of reasons for the indication of allowable subject matter: The prior art could not overcome or render obvious the limitations of “determining a document type of the plain text document based on a calculated keyword score and a large language model (LLM) score; and performing a section-based chunking, a paragraph-based chunking and a sentence-based chunking on the plain text document to obtain the plurality of document chunks, based on the determined document type; generating a plurality of domain centric chunks from the plurality of document chunks using a domain model graph and the LLM, wherein the domain model graph comprises a set of associations and a root node; aggregating the plurality of domain centric chunks by grouping similar chunks from the plurality of domain centric chunks; and generating the indexed corpus by performing indexing on the aggregated plurality of domain centric chunks using vector embeddings” as found in the independent claims. Conclusion 6. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Bargeron (U.S. Publication No. 20090083257) discloses the method and subsystem for information acquisition and aggregation to facilitate ontology and language-model generation within a content-search-service system. Cunningham (U.S. Publication No. 20250094708) discloses generating large language model outputs from stored content items. Hasan (U.S. Publication No. 20230350929) discloses the method and system for generating intent responses through virtual agents. Any inquiry concerning this communication or earlier communications from the examiner should be directed to ETHAN DANIEL KIM whose telephone number is (571) 272-1405. The examiner can normally be reached on Monday - Friday 9:00 - 5:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached on (571) 272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see https://ppair-my.uspto.gov/pair/PrivatePair. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ETHAN DANIEL KIM/ Examiner, Art Unit 2658 /RICHEMOND DORVIL/Supervisory Patent Examiner, Art Unit 2658
Read full office action

Prosecution Timeline

Jan 02, 2025
Application Filed
Sep 10, 2026
Non-Final Rejection mailed — §101 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12731573
ENHANCED SPOKEN DIALOGUE MODIFICATION
3y 3m to grant Granted Sep 08, 2026
Patent 12706101
ROTATION OF SOUND COMPONENTS FOR ORIENTATION-DEPENDENT CODING SCHEMES
3y 2m to grant Granted Aug 11, 2026
Patent 12688220
MODEL PERFORMANCE THROUGH TEXT-TO-TEXT TRANSFORMATION VIA DISTANT SUPERVISION FROM TARGET AND AUXILIARY TASKS
5y 2m to grant Granted Jul 21, 2026
Patent 12682894
SYSTEM AND METHOD FOR SIMULTANEOUSLY IDENTIFYING INTENT AND SLOTS IN VOICE ASSISTANT COMMANDS
3y 11m to grant Granted Jul 14, 2026
Patent 12664993
INTEGRATION OF HIGH FREQUENCY RECONSTRUCTION TECHNIQUES WITH REDUCED POST-PROCESSING DELAY
1y 6m to grant Granted Jun 23, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
77%
Grant Probability
99%
With Interview (+25.4%)
2y 10m (~1y 1m remaining)
Median Time to Grant
Low
PTA Risk
Based on 114 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month