Notice of Pre-AIA or AIA Status
1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
2. The information disclosure statement (IDS) submitted on January 02, 2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 101
3. 35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent thereof, subject to the conditions and requirements of this title.
4. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
The independent claim 1 recites “A processor implemented method, comprising: receiving, via one or more hardware processors, a raw corpus data as input, wherein the raw corpus data is a domain specific data; and generating, via the one or more hardware processors, an indexed corpus from the raw corpus data, comprising: transforming the raw corpus data into an associated plain text document, wherein a file type of the associated plain text document is determined based on a determined file type of the raw corpus data; normalizing the associated plain text document, comprising standardizing and cleaning the plain text document; generating a plurality of document chunks associated with the normalized associated plain text document, comprising: determining a document type of the plain text document based on a calculated keyword score and a large language model (LLM) score; and performing a section-based chunking, a paragraph-based chunking and a sentence-based chunking on the plain text document to obtain the plurality of document chunks, based on the determined document type; generating a plurality of domain centric chunks from the plurality of document chunks using a domain model graph and the LLM, wherein the domain model graph comprises a set of associations and a root node; aggregating the plurality of domain centric chunks by grouping similar chunks from the plurality of domain centric chunks; and generating the indexed corpus by performing indexing on the aggregated plurality of domain centric chunks using vector embeddings”.
The limitations “receiving, via one or more hardware processors, a raw corpus data as input, wherein the raw corpus data is a domain specific data; and generating, via the one or more hardware processors, an indexed corpus from the raw corpus data, comprising: transforming the raw corpus data into an associated plain text document, wherein a file type of the associated plain text document is determined based on a determined file type of the raw corpus data; normalizing the associated plain text document, comprising standardizing and cleaning the plain text document; generating a plurality of document chunks associated with the normalized associated plain text document, comprising: determining a document type of the plain text document based on a calculated keyword score and a large language model (LLM) score; and performing a section-based chunking, a paragraph-based chunking and a sentence-based chunking on the plain text document to obtain the plurality of document chunks, based on the determined document type; generating a plurality of domain centric chunks from the plurality of document chunks using a domain model graph and the LLM, wherein the domain model graph comprises a set of associations and a root node; aggregating the plurality of domain centric chunks by grouping similar chunks from the plurality of domain centric chunks; and generating the indexed corpus by performing indexing on the aggregated plurality of domain centric chunks using vector embeddings” as drafted, covers a mental process, as this could be done by mentally or by hand with pen and paper.
This judicial exception is not integrated into a practical application. Claim 1 recites “A processor implemented method, comprising:…” this limitation directs towards using a computer for the method, and does not impose any meaningful limits on practicing the abstract idea.
The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception. The addition of the generic computer components recited above with regards to claim 1 do not amount to more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Claim 1 does not recite any additional limitations. The claim as drafted, is not patent eligible.
The independent claim 8 recites “A system, comprising: one or more hardware processors; a communication interface; and a memory storing a plurality of instructions, wherein the plurality of instructions cause the one or more hardware processors to: receive a raw corpus data as input, wherein the raw corpus data is a domain specific data; and generate an indexed corpus from the raw corpus data, by: transforming the raw corpus data into an associated plain text document, wherein a file type of the associated plain text document is determined based on a determined file type of the raw corpus data; normalizing the associated plain text document, comprising standardizing and cleaning the plain text document; generating a plurality of document chunks associated with the normalized associated plain text document, by: determining a document type of the plain text document based on a calculated keyword score and a large language model (LLM) score; and performing a section-based chunking, a paragraph-based chunking and a sentence-based chunking on the plain text document to obtain the plurality of document chunks, based on the determined document type; generating a plurality of domain centric chunks from the plurality of document chunks using a domain model graph and the LLM, wherein the domain model graph comprises a set of associations and a root node; aggregating the plurality of domain centric chunks by grouping similar chunks from the plurality of domain centric chunks; and generating the indexed corpus by performing indexing on the aggregated plurality of domain centric chunks using vector embeddings”.
The limitations “receive a raw corpus data as input, wherein the raw corpus data is a domain specific data; and generate an indexed corpus from the raw corpus data, by: transforming the raw corpus data into an associated plain text document, wherein a file type of the associated plain text document is determined based on a determined file type of the raw corpus data; normalizing the associated plain text document, comprising standardizing and cleaning the plain text document; generating a plurality of document chunks associated with the normalized associated plain text document, by: determining a document type of the plain text document based on a calculated keyword score and a large language model (LLM) score; and performing a section-based chunking, a paragraph-based chunking and a sentence-based chunking on the plain text document to obtain the plurality of document chunks, based on the determined document type; generating a plurality of domain centric chunks from the plurality of document chunks using a domain model graph and the LLM, wherein the domain model graph comprises a set of associations and a root node; aggregating the plurality of domain centric chunks by grouping similar chunks from the plurality of domain centric chunks; and generating the indexed corpus by performing indexing on the aggregated plurality of domain centric chunks using vector embeddings” as drafted, covers a mental process, as this could be done by mentally or by hand with pen and paper.
This judicial exception is not integrated into a practical application. Claim 8 recites “A system, comprising: one or more hardware processors; a communication interface; and a memory storing a plurality of instructions, wherein the plurality of instructions cause the one or more hardware processors to:…” this limitation directs towards using a computer for the method, and does not impose any meaningful limits on practicing the abstract idea.
The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception. The addition of the generic computer components recited above with regards to claim 8 do not amount to more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Claim 8 does not recite any additional limitations. The claim as drafted, is not patent eligible.
The independent claim 15 recites “One or more non-transitory machine-readable information storage mediums comprising one or more instructions which when executed by one or more hardware processors cause: receiving a raw corpus data as input, wherein the raw corpus data is a domain specific data; and generating an indexed corpus from the raw corpus data, comprising: transforming the raw corpus data into an associated plain text document, wherein a file type of the associated plain text document is determined based on a determined file type of the raw corpus data; normalizing the associated plain text document, comprising standardizing and cleaning the plain text document; generating a plurality of document chunks associated with the normalized associated plain text document, comprising: determining a document type of the plain text document based on a calculated keyword score and a large language model (LLM) score; and performing a section-based chunking, a paragraph-based chunking and a sentence-based chunking on the plain text document to obtain the plurality of document chunks, based on the determined document type; generating a plurality of domain centric chunks from the plurality of document chunks using a domain model graph and the LLM, wherein the domain model graph comprises a set of associations and a root node; aggregating the plurality of domain centric chunks by grouping similar chunks from the plurality of domain centric chunks; and generating the indexed corpus by performing indexing on the aggregated plurality of domain centric chunks using vector embeddings”.
The limitations “receiving a raw corpus data as input, wherein the raw corpus data is a domain specific data; and generating an indexed corpus from the raw corpus data, comprising: transforming the raw corpus data into an associated plain text document, wherein a file type of the associated plain text document is determined based on a determined file type of the raw corpus data; normalizing the associated plain text document, comprising standardizing and cleaning the plain text document; generating a plurality of document chunks associated with the normalized associated plain text document, comprising: determining a document type of the plain text document based on a calculated keyword score and a large language model (LLM) score; and performing a section-based chunking, a paragraph-based chunking and a sentence-based chunking on the plain text document to obtain the plurality of document chunks, based on the determined document type; generating a plurality of domain centric chunks from the plurality of document chunks using a domain model graph and the LLM, wherein the domain model graph comprises a set of associations and a root node; aggregating the plurality of domain centric chunks by grouping similar chunks from the plurality of domain centric chunks; and generating the indexed corpus by performing indexing on the aggregated plurality of domain centric chunks using vector embeddings” as drafted, covers a mental process, as this could be done by mentally or by hand with pen and paper.
This judicial exception is not integrated into a practical application. Claim 15 recites “One or more non-transitory machine-readable information storage mediums comprising one or more instructions which when executed by one or more hardware processors cause:…” this limitation directs towards using a computer for the method, and does not impose any meaningful limits on practicing the abstract idea.
The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception. The addition of the generic computer components recited above with regards to claim 15 do not amount to more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Claim 15 does not recite any additional limitations. The claim as drafted, is not patent eligible.
Allowable Subject Matter
5. Claims 1-20 have been rejected under 35 U.S.C. 101, but would be allowable if the rejection is overcome.
The following is a statement of reasons for the indication of allowable subject matter: The prior art could not overcome or render obvious the limitations of “determining a document type of the plain text document based on a calculated keyword score and a large language model (LLM) score; and performing a section-based chunking, a paragraph-based chunking and a sentence-based chunking on the plain text document to obtain the plurality of document chunks, based on the determined document type; generating a plurality of domain centric chunks from the plurality of document chunks using a domain model graph and the LLM, wherein the domain model graph comprises a set of associations and a root node; aggregating the plurality of domain centric chunks by grouping similar chunks from the plurality of domain centric chunks; and generating the indexed corpus by performing indexing on the aggregated plurality of domain centric chunks using vector embeddings” as found in the independent claims.
Conclusion
6. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Bargeron (U.S. Publication No. 20090083257) discloses the method and subsystem for information acquisition and aggregation to facilitate ontology and language-model generation within a content-search-service system. Cunningham (U.S. Publication No. 20250094708) discloses generating large language model outputs from stored content items. Hasan (U.S. Publication No. 20230350929) discloses the method and system for generating intent responses through virtual agents.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ETHAN DANIEL KIM whose telephone number is (571) 272-1405. The examiner can normally be reached on Monday - Friday 9:00 - 5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached on (571) 272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see https://ppair-my.uspto.gov/pair/PrivatePair. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ETHAN DANIEL KIM/
Examiner, Art Unit 2658
/RICHEMOND DORVIL/Supervisory Patent Examiner, Art Unit 2658