DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on July 27, 2026 has been entered.
Response to Amendment
The Amendment filed July 27, 2026, has been entered. Claims 1, 12 and 16 have been amended. Claims 8 and 17 have been cancelled. Claims 1 to 7, 9 to 16 and 18 to 20 are pending in the application.
Response to Arguments
Applicant’s arguments, see Remarks Part III. Section A to D, filed July 27, 2026, with respect to the rejection(s) of claim(s) 1, 12 and 16 under 35 USC § 10 have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of Poirier; Louis et al.( US 20240202539 A1).
Applicant’s arguments, see Remarks Part IV., filed July 27, 2026, with respect to the rejection(s) of claim(s) 2 to 7, 9 to 11, 13 to 15 and 18 to 20 under 35 USC § 10 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1 to 20 are rejected under 35 U.S.C. 103 as being unpatentable over Mondlock (US 12020140 B1) hereinafter Mondlock in view of Arunachalam; Elavarasi et al. (US 20250211549 A1) hereinafter Arunachalam in further view of Poirier; Louis et al.( US 20240202539 A1) hereinafter Poirier.
Regarding claim 1, Mondlock teaches:
A non-transitory computer-readable medium configured to store a computer program having logical instructions for enabling one or more processing devices to perform the steps of: in response to receiving a user query directed to subject information retrievable from documentation stored in a private database, using a section-based chunking procedure to obtain, from the private database, a relevant section of the documentation as context,
Shown in “FIG. 7B illustrates aspects of the computer-implemented method 700 involving splitting documents into text chunks, selecting relevant text chunks, and concurrently sending augmented text chunks to an LLM to extract relevant information. In some aspects, the computer-implemented method 700 may include at block 722 causing the plurality of documents to be split into a text chunks set. The plurality of documents may be split by the document/asset/expert module 122 or the LLM service 170. Splitting the plurality of documents may be performed at block 380A of the RAG pipeline 300.” (Mondlock Column 23 lines 10 to 20).
Where the private database is discussed in “The internal data store 140 may be owned or operated by the same organization that owns or operates the generative AI pipeline. The internal data store 140 may include a relational database (e.g., a PostgreSQL database), a non-relational datastore (e.g., a NoSQL database), a vector database (e.g., Pinecone), a web server, file server, and/or application server. In some aspects, the internal data store 140 may be located remotely from the server 110, such as in a public cloud environment. The internal data store 140 may store one or more data sources, such as a chat history 142, document collections 144, asset collections 146, and/or expert collections 148. Chat history 142 may include one or more records of the queries submitted by users and the responses output by server 110. Document collections 144 may include one or more sets of documents, such as web pages, PDFs, Word documents, text files, or any other suitable file containing text. Asset collections 146 may include one or more sets of databases, data sets, applications, models, knowledge graphs, or any other suitable sources of data. Expert collections 148 may include or more sets of identifying information for experts, corpora of experts' works, and/or experts' biographies for one or more subject matter experts. For example, any of the document collections 144, asset collections 146, or expert collections 148 may include data from a catalogue of documents, assets, or experts, such as a data set of assets and descriptions of the respective assets (e.g., applications or models for generating predictions or other data).” (Mondlock Column 4 line 65 to column 5 line 25).
and feeding the user query and the relevant section as context to a Large Language Model (LLM) as part of a prompt configured to cause the LLM to generate a response based on the retrieved section of the documentation.
“The computer-implemented method 600 may continue at block 622 by sending an augmented user query to an LLM to cause the LLM to obtain an answer from the LLM, such as an LLM service 170. The augmented user query may be sent by the LLM interface module 132. Sending the augmented user query may occur at block 388 of the RAG pipeline 300. The augmented user query may include the relevant information responses, the user query, and a prompt to cause the LLM to generate an answer.
“In some aspects, the LLM interface module 132 may transmit prompts to and receive answers from the LLM service 170. The LLM interface module 132 may transmit, for each relevant text chunk and relevant data chunk, a prompt that includes the user query and the relevant text chunk and/or relevant data chunk to the LLM service 170 and may receive relevant information from the relevant text chunk and/or relevant data chunk. The LLM interface module 132 may concurrently transmit a plurality of prompts to the LLM service 170. The LLM interface module 132 may transmit a prompt that includes the user query and each relevant information and may receive an answer. ” (Mondlock Column 9 line 20 to 31).
Mondlock does not teach, but Arunachalam teaches:
wherein the section-based chunking procedure obtains the relevant section based on a section header associated with semantically closest embedded content vectors stored in the private database,
“Systems, methods, and computer program products for using a generative artificial intelligence system to generate answers or summaries is provided. During the ingestion stage, the system receives documents and transcripts that include data associated with a theme. The data is converted into a common format and is divided into chunks. The chunks are associated with metadata tags that include chunk and data information. From the chunks, the system generates embedding vectors. During the inference stage, the system receives an information request. If the information request is a question, the system generates a vector from the question, and uses a similarity search to identify similar vectors. From the similar vectors, the system identifies chunks. If the information request includes a summary request, the system uses the metadata tags to identify chunks with summary information. The system generates an answer or a summary from the identified chunks.” (Arunachalam [Abstract]).
“Data processing engine 114 may also generate a metadata tags 406A, and associate each metadata tag in metadata tags 406A with each chunk in chunks 404A. Each metadata tag may include a project identifier associated with project data 202A. The metadata tag may also include a chunk identifier, a size of the chunk, a title of document 204A or transcript 206A that is included in the chunk, a subtitle corresponding to a section of document 204A or transcript 206A, a hierarchy of the chunk compared to other chunks, and the like. ” (Arunachalam [0055]).
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of Mondlock the capability to receive semantically close vectors related to the relevant sections. The benefit and motivation of such modification is discussed by Arunachalam in the following portion: “One of LLM models, such as LLM 110C may receive one or more chunks in the subset of chunks to generate an answer to the question in the information request. In some embodiments, LLM 110C may receive the chunk that corresponds to the most similar vector first, and then refine the answer with each subsequent chunk. For example, LLM 110C may receive a first chunk in the subset of chunks to generate an answer. Next, LLM 110C may receive a second chunk and the answer generated from the first chunk to generate a second answer. Next, LLM 110C may receive a third chunk and the answer generated using the first and second chunks. The process may continue until LLM 110C uses all chunks in the subset of chunks or until LLM 110C determines that the content of the answer no longer changes.” (Arunachalam [0068]).
Wherein the main benefit of receiving those chunks is to refine the output of an LLM using them as context for the answer.
Mondlock in view of Arunachalam does not explicitly teach, but Poirier teaches:
and wherein the section-based chunking procedure returns content associated with the section header as a complete section of the documentation rather than an arbitrary sliding-window portion of the documentation;
“In step 1508, the computing system generates a respective segment embedding for each data record segment based on the respective segment and the respective contextual metadata. In some embodiments, an embeddings generator module (e.g., embeddings generator module 512) generates the segment embeddings. The respective segment embeddings may be vector values. The respective segment embeddings may comprise respective vector embeddings and the embeddings datastore may comprise one or more vector datastores. In some embodiments, the contextual metadata is stored in respective headers of the respective segments and/or respective data records.” (Poirier [0203])
“In some embodiments, the chunking is performed by an agent, the segment embeddings are generated by another agent, and information is retrieved from the embeddings store based on the respective embeddings by an additional agent, and wherein the agents are supervised by an orchestrator. For example, the addition agent may use a machine learning model to implement a similarity machine learning process that determines a similarity between two of more of segments and/or data records based on the embeddings (e.g., vector embeddings generated based on the data segments, and the contextual metadata). Thus, for example, similarity of segments can be determined based on a similarity of the segments (e.g., similarity of passages of the segments) and similarity of the contextual metadata (or information indicated by the contextual metadata). In some embodiments, the plurality of data records may be processed by a type system prior to being any of scanned and chunked.” (Poirier [0206])
“Another step can include separating text and code (step 1605). The goal of this step is to identify and separate code and pure text. This then allows the system to further process these modalities separately. An additional step can include chunking and parsing text and code (step 1606). Having separated the text and code modalities, the purpose of this step is to identify, locate and extract contiguous pieces of code (step 1607), and conduct chunking of the text contents (step 1606) in reasonable and as contiguous fashion (e.g., that is no cutting mid-sentence or mid-paragraph especially due to breakages between pages and if possible, having chunks with contiguous topics).” (Poirier [0212])
Wherein Poirier [0203] introduce the concept of contextual metadata as the information related to the chunk that is saved in a header of the chunk, Poirier [0206] describes the agent in charge of retrieving the data based on the embeddings, that includes the context metadata saved in the header, in order to find similar segments. Poirier [0212] further describes chunking of a document in a reasonable fashion, avoiding cutting sections like paragraphs and sentences, the steps of locating, identifying contiguous chunks.
And wherein obtaining the relevant section further includes a) detecting a header of the vectors semantically closest to the query vector
“An enterprise generative artificial intelligence architecture is disclosed herein which can intelligently and efficiently crawl and index disparate data records (e.g., data records of one or more enterprise systems) across a variety of different domains using contextual information (e.g., contextual metadata) to provide improved data record identification and retrieval, access control (e.g., role-based access), and map relationships between data records. In one example, contextual information may prevent some users from accessing (e.g., viewing, retrieving) certain data records, and improve similarity evaluations used in retrieval operations (e.g., of a generative artificial intelligence process). Accordingly, the systems described herein can provide more accurate and reliable results that are also faster and more secure than existing techniques.” (Poirier [0023])
“In some embodiments, an enterprise generative artificial intelligence system can crawl, chunk, and index a corpus of data records. Data records can include documents (e.g., PDF, text, html, markdown source code or other source code, etc.), database tables, information generated by applications (e.g., artificial intelligence application insights), images, audiovisual files, executables, models (e.g., data models, machine learning models, large language models, multimodal models), and the like. More specifically, the enterprise generative artificial intelligence system preprocesses and chunks data records across different domains (e.g., data domains, industry-specific domains) of an enterprise. The chunking process partitions data records into segments (or, chunks) and can insert and/or append contextual information for the segments (e.g., as a header of the segment).” (Poirier [0024])
“The contextual information may include, for example, one or more attributes describing the segment (e.g., type of data record, size of chunk, semantic or contextual description of the chunk, access control restrictions or permissions, etc.). A segment may include a passage of a text document, a portion of database table, a sub-model of model, and so forth. The segments and/or contextual information can be stored for efficient retrieval (e.g., as part of a generative artificial intelligence process). For example, the segments and/or contextual information can be stored as embeddings (e.g., vector embeddings) which can allow efficient retrieval. In one example, contextual information can include explicit and/or inferred references between segments and/or data records. For example, the references may indicate relationships that can be used (e.g., traversed) when performing similarity evaluations or other aspects of retrieval operations.” (Poirier [0025])
“In some embodiments, embeddings generator module 128 may generate segment embeddings based on the respective segment and the respective contextual metadata. For example, the segment embedding may include a vector embedding that can used as part of a similarity machine learning process that determines similarities between data records 105-106 and 109-110 and/or segments. Accordingly, similarities can easily be determined (e.g., as part of a generative artificial intelligence retrieval operation) based on various segments and corresponding contextual information.” (Poirier [0042])
“FIG. 8A depicts a flowchart 800 of an example iterative generative artificial intelligence process using unstructured data according to some embodiments. This example process may be implemented by an enterprise generative artificial intelligence system (e.g., enterprise generative artificial intelligence system 402). In step 802, a user query is provided to a retrieval model (e.g., a retriever module of a retrieval agent module). In step 804, the retrieval model receives the query and performs a similarly search (e.g., an ANN-based search) of the vector store 806. The retrieved information is returned the retriever model and provided to a large language model in step 810. …”(Poirier [0163])
Wherein Poirier [0023] to Poirier [0025] describe an invention where the contextual information is described and utilized for an efficient retrieval and the embeddings of the segments are created (including the contextual information, that could include a header). In Poirier [0042] similarities are calculated for the retrieval, taking in account the segment and the metadata and the relationship between the queries and the similarity search is further described in Poirier [0163].
And b) retrieving from the private database, all paragraphs and table entries having the matched header to reconstruct the entire section of the documentation for use as the context;
“FIG. 15 depicts a flowchart 1500 of an example method of intelligent crawling and chunking according to some embodiments. In step 1502, a computing system (e.g., enterprise generative artificial intelligence system and/or intelligent crawling and chunking subsystem 120) scans a plurality of different data domains of an enterprise information environment. In some embodiments, a crawling module (e.g., crawling module 514 and/or crawling module 122) scans the different data domains of the enterprise information environment. In step 1504, the computing system chunks a plurality of data records of multiple enterprise data sources of the plurality of different data domains of the enterprise information environment. The data records can include any of documents, database tables, models, text, images, video, audio, artificial intelligence insights, application outputs, applications, source code, scripts, and/or compiled source code. The chunking can generate one or respective data record segments for each of the plurality of data records. In some embodiments, a chunking module (e.g., chunking module 510 and/or chunking module 124) chunks the data records.” (Poirier [0201])
“In some embodiments, the chunking is performed by an agent, the segment embeddings are generated by another agent, and information is retrieved from the embeddings store based on the respective embeddings by an additional agent, and wherein the agents are supervised by an orchestrator. For example, the addition agent may use a machine learning model to implement a similarity machine learning process that determines a similarity between two of more of segments and/or data records based on the embeddings (e.g., vector embeddings generated based on the data segments, and the contextual metadata). Thus, for example, similarity of segments can be determined based on a similarity of the segments (e.g., similarity of passages of the segments) and similarity of the contextual metadata (or information indicated by the contextual metadata). In some embodiments, the plurality of data records may be processed by a type system prior to being any of scanned and chunked.” (Poirier [0206])
“At the end of the first stage, the system can have several instances of the text, code, table and image classes, outlining different modalities in each data record. Having done this, the system can continue to the second stage, namely, building an information graph for each data record.” (Poirier [0216])
“In the second stage, in order to facilitate an effective information retrieval, the system can represent the information in each data record 1601 using an information graph. The nodes 1630, 1641, 1643 and 1645 of this graph correspond to different instances of the four classes of modality from stage one. The depiction is a (directed) bipartite graph, with edges going out from the text nodes to all other modality nodes. Establishing these edges are the primary goal of this stage.” (Poirier [0217])
”In the third stage, given the information graph, the system can outline a process for information retrieval. One approach for this starts with first embedding the content of the text nodes (and/or other textual metadata associated with other modalities) (step 1629), and storing the embeddings in a vector store 1626. Given a query 1622, the system can embed it (step 1628) and find the most relevant text chunks or text nodes 1630 associated with it. This would be the entry to the graph. At that point, the system can follow the out-going edges to other modality nodes. The classes associated with these modalities, (e.g., code, image, and table) may have a method that enables generation of relevant insights given the query. This method can be powered by different approaches, including multimodal models or other tools for understanding and querying the specific modalities. These insights, together with the text chunks and the user query can then be combined in an aggregator 1650 to a body of text or a prompt to be used for querying a model (e.g., multimodal model, large language model, etc.).” (Poirier [0219])
Wherein Poirier [0201] describes the information that could be carried in the segments, in portions, Poirier [0206] describe retrieving relevant portions that are similar and Poirier [0216], Poirier [0217] and Poirier [0219] describe retrieving those portions and organizing them in a graph, later using an aggregator to combine those portions to include them in a body of text or a prompt for an LLM.
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of Mondlock in view of Arunachalam the capability to detect and manage operations of a retrieval augmented generation based on the information present on a header of a chunk (contextual information in contextual metadata stored in the header) and to retrieve multiple sections related to each other and feed them as context of a prompt to an LLM. The benefit and motivation of such modification is discussed by Poirier in the following portion: “In some implementations, the chunking module 124 can preprocess the data records and/or segments to generate corresponding contextual information. In some embodiments, the contextual information may be included and/or represented in contextual metadata, and/or the contextual metadata may be generated from the contextual information. The contextual information may improve security, as well as accuracy and reliability of associated retrieval operations. In one example, contextual information includes contextual metadata. The contextual information can include references between segments and/or data records 105-106 and 109-110. For example, the references may indicate relationships that can be used (e.g., traversed) when performing similarity evaluations or other aspects of retrieval operations (e.g., by one or more of the agents 506). Contextual information may also include information that can assist a large language model in generating a plan and/or answers. For example, the chunking module 124 may generate contextual information for structured data chunks (or passages) that include natural language descriptions of the data records 105-106 and 109-110, locations of related data records 105-106 and 109-110, and the like.” (Poirier [0038])
Regarding Claim 2, the rejection of claim 1 is incorporated, furthermore Mondlock teaches:
The non-transitory computer-readable medium of claim 1, wherein the section-based chunking procedure uses Retrieval-Augmented Generation (RAG) to parse the user query and retrieve the relevant section. In the following sections:
“FIG. 3B illustrates an aspect of the RAG pipeline 300 involving receiving and processing a user query. In some aspects, the RAG pipeline 300 may include at block 320 receiving a user query. The user query may be received by the input/output module 120 or any other suitable program. The user query may be received from the user device 180 or from a generative AI pipeline, such as RAG pipeline 300. User queries may be distributed among a plurality of servers 110 by the load balancer 190. The user query may comprise a question or a request. The user query may comprise a selection or deselection of one or more document collections, expert collections, or asset collections. The user query may comprise a selection of whether to provide relevant information to the LLM to assist in answering the query (i.e., use RAG) or submit the query without providing relevant information (i.e., do not use RAG).” (Mondlock Column 11 lines 23 to 38).
“FIG. 3E illustrates an aspect of the RAG pipeline 300 involving splitting documents or assets into text or data chunks and determining the most relevant text or data chunks. In some aspects, the RAG pipeline 300 may include at block 380A splitting each selected document into a plurality of text chunks and/or splitting each selected asset into a plurality of data chunks at block 380B. The selected documents and/or selected assets may be split by the document/asset/expert module 122 or any other suitable program.” (Mondlock Column 15 line 63 to column 16 line 4).
Regarding Claim 3, the rejection of claim 1 is incorporated, furthermore Mondlock teaches:
The non-transitory computer-readable medium of claim 1, wherein the section-based chunking procedure uses an inherent structure of the documentation to select, for the relevant section, one or more of subsections, paragraphs, bullet point lists, and tables.
“In some aspects, the document/asset/expert module 122 may include instructions for splitting documents or assets into chunks and generating embeddings of those chunks. The document/asset/expert module 122 may split each document of document collections 144 and document collections 162 into a plurality of text chunks and split each asset of asset collections 146 and asset collections 164 into a plurality of text chunks and/or data chunks. The text chunks may be paragraph-sized, sentence-sized, fixed-sized (e.g., 50 words) or any other appropriate size. The document/asset/expert module 122 may use a tool, such as Natural Language Toolkit (NLTK) or Sentence Splitter, to perform the splitting. In some aspects, document/asset/expert module 122 may transmit the documents and/or assets, via the LLM interface module 132, to the LLM service 170 and receive text chunks and/or asset chunks from the LLM service 170.” (Mondlock Column 7 lines 4 to 18).
Regarding claim 4, the rejection of claim 1 is incorporated, furthermore Mondlock teaches:
The non-transitory computer-readable medium of claim 1, wherein, before receiving the user query, the logical instructions further enable the one or more processing devices to perform a data preparation procedure to separate the documentation into sections,
“In some aspects, the document/asset/expert module 122 may include instructions for splitting documents or assets into chunks and generating embeddings of those chunks. The document/asset/expert module 122 may split each document of document collections 144 and document collections 162 into a plurality of text chunks and split each asset of asset collections 146 and asset collections 164 into a plurality of text chunks and/or data chunks. The text chunks may be paragraph-sized, sentence-sized, fixed-sized (e.g., 50 words) or any other appropriate size. The document/asset/expert module 122 may use a tool, such as Natural Language Toolkit (NLTK) or Sentence Splitter, to perform the splitting. In some aspects, document/asset/expert module 122 may transmit the documents and/or assets, via the LLM interface module 132, to the LLM service 170 and receive text chunks and/or asset chunks from the LLM service 170.” (Mondlock Column 7 lines 4 to 18).
Mondlock does not teach: each section including content under a respective section header. On the other hand, Arunachalam teaches: “More specifically, during an ingestion stage, a generative AI system may receive data from various sources. The data may correspond to a project and may be included in documents, transcripts, and the like. The generative AI system may convert the data into a uniform format and may also generate an identifier, such as a project identifier, for the data. Next, the generative AI system may divide the data in the uniform format into chunks having a predefined length. The generative AI system may also associate metadata tags with the chunks. There may be one metadata tag for one chunk in some embodiments. The metadata tag may include the identifier, the title of the chunk, the hierarchy of the chunk compared to other chunks, the data source (e.g., the document or transcript from where the chunk came from), and the like. From the chunks, an embedding large language model (LLM) may generate embedding vectors (or simply vectors). The vectors may include numeric embeddings that represent information in the chunks. The generative AI system may generate a dictionary for the data, where the dictionary includes or points to an identifier, chunks, metadata tags, and vectors associated with the data. The generative AI system may store the dictionary, including the identifier, chunks, metadata tags, and vectors in a vector storage, or a combination of various storage devices.” (Arunachalam [0015]).
It would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Mondlock to incorporate the teachings of Arunachalam to include the header of the text chunk in the embeddings of the text chunk. The motivation to include the respective header of the text chunk to the text chunk embedding is discussed by Arunachalam and can be found in “At operation 608, a dictionary is generated. For example, generative AI system 108 may generate dictionary 410A that links the project identifier, vectors 408A, chunks 404A, and metadata tags 406A associated with project data 202A.” (Arunachalam [0081]). Wherein the collection of headers and sections could be used to create a dictionary that contains the project data.
Regarding claim 5, the rejection of claim 4 is incorporated, furthermore Mondlock teaches:
The non-transitory computer-readable medium of claim 4, wherein the data preparation procedure further includes dividing the content of each section into one or more of paragraphs, table entries, and subsections.
“In some aspects, the document/asset/expert module 122 may include instructions for splitting documents or assets into chunks and generating embeddings of those chunks. The document/asset/expert module 122 may split each document of document collections 144 and document collections 162 into a plurality of text chunks and split each asset of asset collections 146 and asset collections 164 into a plurality of text chunks and/or data chunks. The text chunks may be paragraph-sized, sentence-sized, fixed-sized (e.g., 50 words) or any other appropriate size. The document/asset/expert module 122 may use a tool, such as Natural Language Toolkit (NLTK) or Sentence Splitter, to perform the splitting. In some aspects, document/asset/expert module 122 may transmit the documents and/or assets, via the LLM interface module 132, to the LLM service 170 and receive text chunks and/or asset chunks from the LLM service 170.” (Mondlock Column 7 lines 4 to 18).
Regarding claim 6, the rejection of claim 4 is incorporated, furthermore Mondlock teaches:
The non-transitory computer-readable medium of claim 4, wherein the data preparation procedure further includes embedding a content value of each paragraph or table entry within a section as a vector and storing the vector together with metadata including (i) a header identifying the section,
“Data processing engine 114 may also generate a metadata tags 406A, and associate each metadata tag in metadata tags 406A with each chunk in chunks 404A. Each metadata tag may include a project identifier associated with project data 202A. The metadata tag may also include a chunk identifier, a size of the chunk, a title of document 204A or transcript 206A that is included in the chunk, a subtitle corresponding to a section of document 204A or transcript 206A, a hierarchy of the chunk compared to other chunks, and the like. ” (Arunachalam [0055]).
“At operation 606, vectors are generated. For example, an embedding LLM in LLMs 110 may generate vectors 408A from chunks 404A. Vectors 408A may include embeddings in the embedding space that correspond to the data stored in chunks 404A. There may be one chunk in chunks 404A for one vector in vectors 408A.” (Arunachalam [0080]).
“In some instances, data processing engine 114 may assign metadata tags to chunks. Typically there may be one metadata tag for each chunk. The metadata tag may include a project identifier, a title of the document, a subtitle of the document, a hierarchy of the chunk as compared to other chunks in the document or in the section of the document, a hierarchy of the chunk in the project, etc. ” (Arunachalam [0038]).
(ii) a content type identifying paragraph or table entry,
“As discussed above, data may include text, images, tables, audio and video files, etc., that may be collected from multiple data sources. Data sources may include documents, transcripts, audio and video recordings from various meetings, handwritten notes, source code, and the like. …” (Arunachalam [0027]).
“Data processing engine 114 may also generate a metadata tags 406A, and associate each metadata tag in metadata tags 406A with each chunk in chunks 404A. Each metadata tag may include a project identifier associated with project data 202A. The metadata tag may also include a chunk identifier, a size of the chunk, a title of document 204A or transcript 206A that is included in the chunk, a subtitle corresponding to a section of document 204A or transcript 206A, a hierarchy of the chunk compared to other chunks, and the like. ” (Arunachalam [0055]).
and (iii) a table path identifying a stored table representation, in the private database to enable the documentation to be searched by section.
“In some embodiments, generative AI system 108 may generate a dictionary 410A. Dictionary 410A may include or be associated with a project identifier, document chunks 404A, metadata tags 406A, and vectors 408A. Alternatively dictionary 410A may include pointers to locations in data storage 120 that stores chunks 404A, metadata tags 406A, and vectors 408A. In this way, dictionary 410 may be used to identify locations of one or more document chunks 404A, metadata tags 406A, and vectors 408A associated with project data 202A. ” (Arunachalam [0057]).
“Generative AI system 108 may also include vector storage 412. Vector storage 412 may be data storage 120 discussed in FIG. 1. Vector storage 412 may store a project identifier, document chunks 404A, metadata tags 406A, and vectors 408A associated with project data 202A. ” (Arunachalam [0058]).
“Based on the similarity search, vector retrieving module 506 may retrieve the top K vectors from vectors 408A. Vector retrieving module 506 may rank the top K vectors according to similarity from the most similar vector to the least similar vector. Next, vector retrieving module 506 may pass the ranked top K vectors to a chunk retrieving module 508. The chunk retrieving module 508 may convert the ranked top K vectors into chunks. Alternatively, chunk retrieving module 508 may use dictionary 410A to access the subset of chunks from chunks 404A that correspond to the ranked top K vectors. In some embodiments, the subset of chunks may be in the same order as the ranked top K vectors. ” (Arunachalam [0067]).
It would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Mondlock to incorporate the teachings of Arunachalam to include information about the header, identification and location of the paragraphs and tables in the embeddings of the text chunk. The motivation to include the respective header of the text chunk to the text chunk embedding is discussed by Arunachalam and can be found in “… Additionally, the chunks may be divided according to one or more rules, such as including complete words, sentences, and paragraphs in one chunk, keeping data from one or more tables within one chunk, not dividing images between chunks, and the like. Data processing engine 114 may recursively generate chunks by dividing the available data from each document 204A or transcript 206A in half, until data processing engine 114 generates chunks that are less than a predefined chunk size.” (Arunachalam [0054]).
Regarding Claim 7, the rejection of claim 1 is incorporated, furthermore Mondlock teaches:
The non-transitory computer-readable medium of claim 1, wherein the logical instructions further enable the one or more processing devices to embed the user query as a query vector, wherein obtaining the relevant section of the documentation as context includes searching the private database for vectors semantically closest to the query vector. In the following sections:
“In some aspects, the document/asset/expert module 122 may include instructions for generating embeddings from each text chunk and/or data chunk. The embeddings represent the text chunks and data chunks as multi-dimensional (e.g., 768 or 1,536 dimension) vectors of numerical values. The document/asset/expert module 122 may use Word2Vec, Bidirectional Encoder Representations from Transformers (BERT), or other suitable algorithms to generate the embeddings. Alternatively, the document/asset/expert module 122 may transmit the text chunks and/or data chunks, via the LLM interface module 132, to the LLM service 170 (e.g., using the text-embedding-ada-002 model) and receive embeddings from the LLM service 170. The document/asset/expert module 122 may save the embeddings into embeddings 168 in the external data sources 160. The embeddings 168 may comprise a vector database, such as ChromaDB, Pinecone, or Milvus.” (Mondlock Column 7 lines 19 to 35).
“In some aspects, the relevant information identification module 130 may include instructions for identifying relevant text chunks and/or data chunks from the relevant documents and/or relevant assets. The relevant information identification module 130 may use a semantic search to compare the embedding of the user query to embeddings of the text chunks and/or data chunks to identify relevant text chunks and/or relevant data chunks” (Mondlock Column 9 lines 7 to 14).
“The computer-implemented method 600 may continue at block 614 by causing chunk similarity scores to be calculated. The chunk similarity scores may be calculated by the relevant information identification module 130 or the LLM service 170. Chunk similarity scores may be calculated at blocks 382A and/or 384A of the RAG pipeline 300. Chunk similarity scores may indicate semantic similarity of the user query to each text chunk of the plurality of documents. The chunk similarity scores may be calculated using various techniques, such as cosine similarity between the vectors representing the chunks.” (Mondlock Column 20 lines 5 to 12).
Regarding Claim 9, the rejection of claim 1 is incorporated, furthermore Mondlock teaches:
The non-transitory computer-readable medium of claim 1, wherein the section-based chunking procedure obtains the relevant section of the documentation in a manner unrelated to a sliding window procedure.
“In some aspects, the RAG pipeline 300 may include at block 384A identifying the top relevant text chunks and/or identifying the top relevant data chunks at block 384B. The top relevant text chunks and/or data chunks may be identified by the relevant information identification module 130 or any other suitable program. The top relevant text and data chunks may be identified by a semantic search. The semantic search may comprise performing a KNN search of the user query embedding and text chunk and/or data chunk embeddings.” (Mondlock Column 16 lines 14 to 23).
Regarding Claim 10, the rejection of claim 1 is incorporated, furthermore Mondlock teaches:
The non-transitory computer-readable medium of claim 1, wherein a size of the user query and relevant section is configured to fall within an input token limit of the LLM.
“In some aspects, the query module 128 may include instructions for generating an augmented user query from the rephrased user query. The query module 128 may supplement the rephrased user query with information obtained (via the document/asset/expert module 122 and the relevant information identification module 130) from document collections 144, document collections 162, asset collections 146, asset collections 164, and/or other suitable sources to generate a prompt. For example, a rephrased user query may ask a question regarding Acme Corp.'s most recent earnings report. The query module 128 may append the contents of Acme Corp.'s earnings report when generating the augmented user query. The query module 128 may summarize an augmented user query in order to satisfy a maximum word or token limit of the LLM service 170. For example, the query module 128 may implement map reduce functionality to split Acme Corp.'s earnings report document into a plurality of text chunks and summarize each text chunk to generate a summarized output text suitable for submission to the LLM service 170.” (Mondlock Column 8 lines 19 to 38).
Regarding claim 11, the rejection of claim 1 is incorporated, furthermore Mondlock teaches:
The non-transitory computer-readable medium of claim 1, wherein the private database is a vector store.
“The internal data store 140 may be owned or operated by the same organization that owns or operates the generative AI pipeline. The internal data store 140 may include a relational database (e.g., a PostgreSQL database), a non-relational datastore (e.g., a NoSQL database), a vector database (e.g., Pinecone), a web server, file server, and/or application server.” (Mondlock Column 4 line 65 to column 5 line 4).
“In some aspects, the computer-implemented method 700 may include at block 726 saving the text chunks set and text embeddings into a data store. The text chunks set and text embeddings may be saved by the document/asset/expert module 122. The text chunks set and text embeddings may be saved at block 382A of the RAG pipeline 300. The text chunks set and text embeddings may be saved into internal data store 140.” (Mondlock Column 23 lines 28 to 35).
Regarding claim 12, Mondlock teaches:
A method comprising the steps of: in response to receiving a user query directed to subject information retrievable from documentation stored in a private database, using a section-based chunking procedure to obtain, from the private database, a relevant section of the documentation as context,
“FIG. 7B illustrates aspects of the computer-implemented method 700 involving splitting documents into text chunks, selecting relevant text chunks, and concurrently sending augmented text chunks to an LLM to extract relevant information. In some aspects, the computer-implemented method 700 may include at block 722 causing the plurality of documents to be split into a text chunks set. The plurality of documents may be split by the document/asset/expert module 122 or the LLM service 170. Splitting the plurality of documents may be performed at block 380A of the RAG pipeline 300.” (Mondlock Column 23 lines 10 to 20).
“The internal data store 140 may be owned or operated by the same organization that owns or operates the generative AI pipeline.” (Mondlock Column 4 lines 65 to 67).
and feeding the user query and the relevant section as context to a Large Language Model (LLM) as part of a prompt configured to cause the LLM to generate a response based on the retrieved section of the documentation.
“The computer-implemented method 600 may continue at block 622 by sending an augmented user query to an LLM to cause the LLM to obtain an answer from the LLM, such as an LLM service 170. The augmented user query may be sent by the LLM interface module 132. Sending the augmented user query may occur at block 388 of the RAG pipeline 300. The augmented user query may include the relevant information responses, the user query, and a prompt to cause the LLM to generate an answer.” (Mondlock Column 20 line 13 to 21).
“In some aspects, the LLM interface module 132 may transmit prompts to and receive answers from the LLM service 170. The LLM interface module 132 may transmit, for each relevant text chunk and relevant data chunk, a prompt that includes the user query and the relevant text chunk and/or relevant data chunk to the LLM service 170 and may receive relevant information from the relevant text chunk and/or relevant data chunk. The LLM interface module 132 may concurrently transmit a plurality of prompts to the LLM service 170. The LLM interface module 132 may transmit a prompt that includes the user query and each relevant information and may receive an answer. ” (Mondlock Column 9 line 20 to 31).
Mondlock does not teach, but Arunachalam teaches:
wherein the section-based chunking procedure obtains the relevant section based on a section header associated with semantically closest embedded content vectors stored in the private database,
“Systems, methods, and computer program products for using a generative artificial intelligence system to generate answers or summaries is provided. During the ingestion stage, the system receives documents and transcripts that include data associated with a theme. The data is converted into a common format and is divided into chunks. The chunks are associated with metadata tags that include chunk and data information. From the chunks, the system generates embedding vectors. During the inference stage, the system receives an information request. If the information request is a question, the system generates a vector from the question, and uses a similarity search to identify similar vectors. From the similar vectors, the system identifies chunks. If the information request includes a summary request, the system uses the metadata tags to identify chunks with summary information. The system generates an answer or a summary from the identified chunks.” (Arunachalam [Abstract]).
“Data processing engine 114 may also generate a metadata tags 406A, and associate each metadata tag in metadata tags 406A with each chunk in chunks 404A. Each metadata tag may include a project identifier associated with project data 202A. The metadata tag may also include a chunk identifier, a size of the chunk, a title of document 204A or transcript 206A that is included in the chunk, a subtitle corresponding to a section of document 204A or transcript 206A, a hierarchy of the chunk compared to other chunks, and the like. ” (Arunachalam [0055]).
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of Mondlock the capability to receive semantically close vectors related to the relevant sections. The benefit and motivation of such modification is discussed by Arunachalam in the following portion: “One of LLM models, such as LLM 110C may receive one or more chunks in the subset of chunks to generate an answer to the question in the information request. In some embodiments, LLM 110C may receive the chunk that corresponds to the most similar vector first, and then refine the answer with each subsequent chunk. For example, LLM 110C may receive a first chunk in the subset of chunks to generate an answer. Next, LLM 110C may receive a second chunk and the answer generated from the first chunk to generate a second answer. Next, LLM 110C may receive a third chunk and the answer generated using the first and second chunks. The process may continue until LLM 110C uses all chunks in the subset of chunks or until LLM 110C determines that the content of the answer no longer changes.” (Arunachalam [0068]).
Wherein the main benefit of receiving those chunks is to refine the output of an LLM using them as context for the answer.
Mondlock in view of Arunachalam does not explicitly teach, but Poirier teaches:
and wherein the section-based chunking procedure returns content associated with the section header as a complete section of the documentation rather than an arbitrary sliding-window portion of the documentation;
“In step 1508, the computing system generates a respective segment embedding for each data record segment based on the respective segment and the respective contextual metadata. In some embodiments, an embeddings generator module (e.g., embeddings generator module 512) generates the segment embeddings. The respective segment embeddings may be vector values. The respective segment embeddings may comprise respective vector embeddings and the embeddings datastore may comprise one or more vector datastores. In some embodiments, the contextual metadata is stored in respective headers of the respective segments and/or respective data records.” (Poirier [0203])
“In some embodiments, the chunking is performed by an agent, the segment embeddings are generated by another agent, and information is retrieved from the embeddings store based on the respective embeddings by an additional agent, and wherein the agents are supervised by an orchestrator. For example, the addition agent may use a machine learning model to implement a similarity machine learning process that determines a similarity between two of more of segments and/or data records based on the embeddings (e.g., vector embeddings generated based on the data segments, and the contextual metadata). Thus, for example, similarity of segments can be determined based on a similarity of the segments (e.g., similarity of passages of the segments) and similarity of the contextual metadata (or information indicated by the contextual metadata). In some embodiments, the plurality of data records may be processed by a type system prior to being any of scanned and chunked.” (Poirier [0206])
“Another step can include separating text and code (step 1605). The goal of this step is to identify and separate code and pure text. This then allows the system to further process these modalities separately. An additional step can include chunking and parsing text and code (step 1606). Having separated the text and code modalities, the purpose of this step is to identify, locate and extract contiguous pieces of code (step 1607), and conduct chunking of the text contents (step 1606) in reasonable and as contiguous fashion (e.g., that is no cutting mid-sentence or mid-paragraph especially due to breakages between pages and if possible, having chunks with contiguous topics).” (Poirier [0212])
Wherein Poirier [0203] introduce the concept of contextual metadata as the information related to the chunk that is saved in a header of the chunk, Poirier [0206] describes the agent in charge of retrieving the data based on the embeddings, that includes the context metadata saved in the header, in order to find similar segments. Poirier [0212] further describes chunking of a document in a reasonable fashion, avoiding cutting sections like paragraphs and sentences, the steps of locating, identifying contiguous chunks.
And wherein obtaining the relevant section further includes a) detecting a header of the vectors semantically closest to the query vector
“An enterprise generative artificial intelligence architecture is disclosed herein which can intelligently and efficiently crawl and index disparate data records (e.g., data records of one or more enterprise systems) across a variety of different domains using contextual information (e.g., contextual metadata) to provide improved data record identification and retrieval, access control (e.g., role-based access), and map relationships between data records. In one example, contextual information may prevent some users from accessing (e.g., viewing, retrieving) certain data records, and improve similarity evaluations used in retrieval operations (e.g., of a generative artificial intelligence process). Accordingly, the systems described herein can provide more accurate and reliable results that are also faster and more secure than existing techniques.” (Poirier [0023])
“In some embodiments, an enterprise generative artificial intelligence system can crawl, chunk, and index a corpus of data records. Data records can include documents (e.g., PDF, text, html, markdown source code or other source code, etc.), database tables, information generated by applications (e.g., artificial intelligence application insights), images, audiovisual files, executables, models (e.g., data models, machine learning models, large language models, multimodal models), and the like. More specifically, the enterprise generative artificial intelligence system preprocesses and chunks data records across different domains (e.g., data domains, industry-specific domains) of an enterprise. The chunking process partitions data records into segments (or, chunks) and can insert and/or append contextual information for the segments (e.g., as a header of the segment).” (Poirier [0024])
“The contextual information may include, for example, one or more attributes describing the segment (e.g., type of data record, size of chunk, semantic or contextual description of the chunk, access control restrictions or permissions, etc.). A segment may include a passage of a text document, a portion of database table, a sub-model of model, and so forth. The segments and/or contextual information can be stored for efficient retrieval (e.g., as part of a generative artificial intelligence process). For example, the segments and/or contextual information can be stored as embeddings (e.g., vector embeddings) which can allow efficient retrieval. In one example, contextual information can include explicit and/or inferred references between segments and/or data records. For example, the references may indicate relationships that can be used (e.g., traversed) when performing similarity evaluations or other aspects of retrieval operations.” (Poirier [0025])
“In some embodiments, embeddings generator module 128 may generate segment embeddings based on the respective segment and the respective contextual metadata. For example, the segment embedding may include a vector embedding that can used as part of a similarity machine learning process that determines similarities between data records 105-106 and 109-110 and/or segments. Accordingly, similarities can easily be determined (e.g., as part of a generative artificial intelligence retrieval operation) based on various segments and corresponding contextual information.” (Poirier [0042])
“FIG. 8A depicts a flowchart 800 of an example iterative generative artificial intelligence process using unstructured data according to some embodiments. This example process may be implemented by an enterprise generative artificial intelligence system (e.g., enterprise generative artificial intelligence system 402). In step 802, a user query is provided to a retrieval model (e.g., a retriever module of a retrieval agent module). In step 804, the retrieval model receives the query and performs a similarly search (e.g., an ANN-based search) of the vector store 806. The retrieved information is returned the retriever model and provided to a large language model in step 810. …”(Poirier [0163])
Wherein Poirier [0023] to Poirier [0025] describe an invention where the contextual information is described and utilized for an efficient retrieval and the embeddings of the segments are created (including the contextual information, that could include a header). In Poirier [0042] similarities are calculated for the retrieval, taking in account the segment and the metadata and the relationship between the queries and the similarity search is further described in Poirier [0163].
And b) retrieving from the private database, all paragraphs and table entries having the matched header to reconstruct the entire section of the documentation for use as the context;
“FIG. 15 depicts a flowchart 1500 of an example method of intelligent crawling and chunking according to some embodiments. In step 1502, a computing system (e.g., enterprise generative artificial intelligence system and/or intelligent crawling and chunking subsystem 120) scans a plurality of different data domains of an enterprise information environment. In some embodiments, a crawling module (e.g., crawling module 514 and/or crawling module 122) scans the different data domains of the enterprise information environment. In step 1504, the computing system chunks a plurality of data records of multiple enterprise data sources of the plurality of different data domains of the enterprise information environment. The data records can include any of documents, database tables, models, text, images, video, audio, artificial intelligence insights, application outputs, applications, source code, scripts, and/or compiled source code. The chunking can generate one or respective data record segments for each of the plurality of data records. In some embodiments, a chunking module (e.g., chunking module 510 and/or chunking module 124) chunks the data records.” (Poirier [0201])
“In some embodiments, the chunking is performed by an agent, the segment embeddings are generated by another agent, and information is retrieved from the embeddings store based on the respective embeddings by an additional agent, and wherein the agents are supervised by an orchestrator. For example, the addition agent may use a machine learning model to implement a similarity machine learning process that determines a similarity between two of more of segments and/or data records based on the embeddings (e.g., vector embeddings generated based on the data segments, and the contextual metadata). Thus, for example, similarity of segments can be determined based on a similarity of the segments (e.g., similarity of passages of the segments) and similarity of the contextual metadata (or information indicated by the contextual metadata). In some embodiments, the plurality of data records may be processed by a type system prior to being any of scanned and chunked.” (Poirier [0206])
“At the end of the first stage, the system can have several instances of the text, code, table and image classes, outlining different modalities in each data record. Having done this, the system can continue to the second stage, namely, building an information graph for each data record.” (Poirier [0216])
“In the second stage, in order to facilitate an effective information retrieval, the system can represent the information in each data record 1601 using an information graph. The nodes 1630, 1641, 1643 and 1645 of this graph correspond to different instances of the four classes of modality from stage one. The depiction is a (directed) bipartite graph, with edges going out from the text nodes to all other modality nodes. Establishing these edges are the primary goal of this stage.” (Poirier [0217])
”In the third stage, given the information graph, the system can outline a process for information retrieval. One approach for this starts with first embedding the content of the text nodes (and/or other textual metadata associated with other modalities) (step 1629), and storing the embeddings in a vector store 1626. Given a query 1622, the system can embed it (step 1628) and find the most relevant text chunks or text nodes 1630 associated with it. This would be the entry to the graph. At that point, the system can follow the out-going edges to other modality nodes. The classes associated with these modalities, (e.g., code, image, and table) may have a method that enables generation of relevant insights given the query. This method can be powered by different approaches, including multimodal models or other tools for understanding and querying the specific modalities. These insights, together with the text chunks and the user query can then be combined in an aggregator 1650 to a body of text or a prompt to be used for querying a model (e.g., multimodal model, large language model, etc.).” (Poirier [0219])
Wherein Poirier [0201] describes the information that could be carried in the segments, in portions, Poirier [0206] describe retrieving relevant portions that are similar and Poirier [0216], Poirier [0217] and Poirier [0219] describe retrieving those portions and organizing them in a graph, later using an aggregator to combine those portions to include them in a body of text or a prompt for an LLM.
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of Mondlock in view of Arunachalam the capability to detect and manage operations of a retrieval augmented generation based on the information present on a header of a chunk (contextual information in contextual metadata stored in the header) and to retrieve multiple sections related to each other and feed them as context of a prompt to an LLM. The benefit and motivation of such modification is discussed by Poirier in the following portion: “In some implementations, the chunking module 124 can preprocess the data records and/or segments to generate corresponding contextual information. In some embodiments, the contextual information may be included and/or represented in contextual metadata, and/or the contextual metadata may be generated from the contextual information. The contextual information may improve security, as well as accuracy and reliability of associated retrieval operations. In one example, contextual information includes contextual metadata. The contextual information can include references between segments and/or data records 105-106 and 109-110. For example, the references may indicate relationships that can be used (e.g., traversed) when performing similarity evaluations or other aspects of retrieval operations (e.g., by one or more of the agents 506). Contextual information may also include information that can assist a large language model in generating a plan and/or answers. For example, the chunking module 124 may generate contextual information for structured data chunks (or passages) that include natural language descriptions of the data records 105-106 and 109-110, locations of related data records 105-106 and 109-110, and the like.” (Poirier [0038])
Regarding claim 13, the rejection of claim 12 is incorporated, furthermore Mondlock teaches:
The method of claim 12, wherein the section-based chunking procedure uses Retrieval-Augmented Generation (RAG) to parse the user query and retrieve the relevant section.
“FIG. 3B illustrates an aspect of the RAG pipeline 300 involving receiving and processing a user query. In some aspects, the RAG pipeline 300 may include at block 320 receiving a user query.” (Mondlock Column 11 lines 23 to 26).
“FIG. 3E illustrates an aspect of the RAG pipeline 300 involving splitting documents or assets into text or data chunks and determining the most relevant text or data chunks.” (Mondlock Column 15 lines 63 to 66).
Regarding claim 14, the rejection of claim 12 is incorporated, furthermore Mondlock teaches:
The method of claim 12, wherein the section-based chunking procedure uses an inherent structure of the documentation including, for the relevant section, one or more subsections, paragraphs, bullet point lists, and tables.
“In some aspects, the document/asset/expert module 122 may include instructions for splitting documents or assets into chunks and generating embeddings of those chunks. The document/asset/expert module 122 may split each document of document collections 144 and document collections 162 into a plurality of text chunks and split each asset of asset collections 146 and asset collections 164 into a plurality of text chunks and/or data chunks. The text chunks may be paragraph-sized, sentence-sized, fixed-sized (e.g., 50 words) or any other appropriate size. The document/asset/expert module 122 may use a tool, such as Natural Language Toolkit (NLTK) or Sentence Splitter, to perform the splitting. In some aspects, document/asset/expert module 122 may transmit the documents and/or assets, via the LLM interface module 132, to the LLM service 170 and receive text chunks and/or asset chunks from the LLM service 170.” (Mondlock Column 7 lines 4 to 18).
Regarding claim 15, the rejection of claim 12 is incorporated, furthermore Mondlock teaches:
dividing the content of each section into one or more of paragraphs, table entries, and subsections;
“In some aspects, the document/asset/expert module 122 may include instructions for splitting documents or assets into chunks and generating embeddings of those chunks. The document/asset/expert module 122 may split each document of document collections 144 and document collections 162 into a plurality of text chunks and split each asset of asset collections 146 and asset collections 164 into a plurality of text chunks and/or data chunks. The text chunks may be paragraph-sized, sentence-sized, fixed-sized (e.g., 50 words) or any other appropriate size. The document/asset/expert module 122 may use a tool, such as Natural Language Toolkit (NLTK) or Sentence Splitter, to perform the splitting. In some aspects, document/asset/expert module 122 may transmit the documents and/or assets, via the LLM interface module 132, to the LLM service 170 and receive text chunks and/or asset chunks from the LLM service 170.” (Mondlock Column 7 lines 4 to 18).
and embedding a content value of each section as vectors in the private database to enable the documentation to be searched by section.
“In some aspects, the document/asset/expert module 122 may include instructions for generating embeddings from each text chunk and/or data chunk. The embeddings represent the text chunks and data chunks as multi-dimensional (e.g., 768 or 1,536 dimension) vectors of numerical values. The document/asset/expert module 122 may use Word2Vec, Bidirectional Encoder Representations from Transformers (BERT), or other suitable algorithms to generate the embeddings. Alternatively, the document/asset/expert module 122 may transmit the text chunks and/or data chunks, via the LLM interface module 132, to the LLM service 170 (e.g., using the text-embedding-ada-002 model) and receive embeddings from the LLM service 170. The document/asset/expert module 122 may save the embeddings into embeddings 168 in the external data sources 160. The embeddings 168 may comprise a vector database, such as ChromaDB, Pinecone, or Milvus.” (Mondlock Column 7 lines 19 to 35).
“In some aspects, the relevant information identification module 130 may include instructions for identifying relevant text chunks and/or data chunks from the relevant documents and/or relevant assets. The relevant information identification module 130 may use a semantic search to compare the embedding of the user query to embeddings of the text chunks and/or data chunks to identify relevant text chunks and/or relevant data chunks” (Mondlock Column 9 lines 7 to 14).
Mondlock does not teach: The method of claim 12, wherein, before receiving the user query, the process further comprises the steps of: performing a data preparation procedure to separate the documentation into sections, each section including content under a respective section header;
In Mondlock there is no specific mention that the header of the text chunk is embedded with the text chunk as a searchable tag of the text chunk.
On the other hand, Arunachalam define sections of metadata that could be embedded in a vector with a text chunk that is being prepared for an LLM in here: “The generative AI system may also associate metadata tags with the chunks. There may be one metadata tag for one chunk in some embodiments. The metadata tag may include the identifier, the title of the chunk, the hierarchy of the chunk compared to other chunks, the data source (e.g., the document or transcript from where the chunk came from), and the like.” (Arunachalam [0015] lines 9 to 15).
Such metadata portions in the embeddings of the text chunks could modify the data preparation procedure disclosed by Mondlock in: “In some aspects, the document/asset/expert module 122 may include instructions for splitting documents or assets into chunks and generating embeddings of those chunks. The document/asset/expert module 122 may split each document of document collections 144 and document collections 162 into a plurality of text chunks and split each asset of asset collections 146 and asset collections 164 into a plurality of text chunks and/or data chunks. The text chunks may be paragraph-sized, sentence-sized, fixed-sized (e.g., 50 words) or any other appropriate size. The document/asset/expert module 122 may use a tool, such as Natural Language Toolkit (NLTK) or Sentence Splitter, to perform the splitting. In some aspects, document/asset/expert module 122 may transmit the documents and/or assets, via the LLM interface module 132, to the LLM service 170 and receive text chunks and/or asset chunks from the LLM service 170.” (Mondlock Column 7 lines 4 to 18).
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of Mondlock the capability to prepare and separate document into sections for the LLM. The benefit and motivation of such modification is discussed by Arunachalam in the following portion: “As discussed above, generative AI system 108 may operate during the ingestion stage and the inference stage. The various embodiments that use the above components during the ingestion stage and the inference stage are discussed below. In particular, FIGS. 2-3 are directed to one embodiment for training and using generative AI system 108, and FIGS. 4-8 are directed to a different embodiment for training and using the generative AI system 108.” (Arunachalam [0040]). Wherein the motivation to do that preparation of the chunks is so an LLM can ingest the chunks generated.
Regarding claim 16, Mondlock teaches:
A system comprising: a processing device; and memory configured to store computer logic having instructions enabling the processing device to perform the steps of: in response to receiving a user query directed to subject information retrievable from documentation stored in a private database, using a section-based chunking procedure to obtain, from the private database, a relevant section of the documentation as context,
“FIG. 7B illustrates aspects of the computer-implemented method 700 involving splitting documents into text chunks, selecting relevant text chunks, and concurrently sending augmented text chunks to an LLM to extract relevant information. In some aspects, the computer-implemented method 700 may include at block 722 causing the plurality of documents to be split into a text chunks set. The plurality of documents may be split by the document/asset/expert module 122 or the LLM service 170. Splitting the plurality of documents may be performed at block 380A of the RAG pipeline 300.” (Mondlock Column 23 lines 10 to 20).
“The internal data store 140 may be owned or operated by the same organization that owns or operates the generative AI pipeline.” (Mondlock Column 4 lines 65 to 67).
and feeding the user query and the relevant section as context to a Large Language Model (LLM) as part of a prompt configured to cause the LLM to generate a response based on the retrieved section of the documentation.
“The computer-implemented method 600 may continue at block 622 by sending an augmented user query to an LLM to cause the LLM to obtain an answer from the LLM, such as an LLM service 170. The augmented user query may be sent by the LLM interface module 132. Sending the augmented user query may occur at block 388 of the RAG pipeline 300. The augmented user query may include the relevant information responses, the user query, and a prompt to cause the LLM to generate an answer.” (Mondlock Column 20 line 13 to 21).
“In some aspects, the LLM interface module 132 may transmit prompts to and receive answers from the LLM service 170. The LLM interface module 132 may transmit, for each relevant text chunk and relevant data chunk, a prompt that includes the user query and the relevant text chunk and/or relevant data chunk to the LLM service 170 and may receive relevant information from the relevant text chunk and/or relevant data chunk. The LLM interface module 132 may concurrently transmit a plurality of prompts to the LLM service 170. The LLM interface module 132 may transmit a prompt that includes the user query and each relevant information and may receive an answer. ” (Mondlock Column 9 line 20 to 31).
Mondlock does not teach, but Arunachalam teaches:
wherein the section-based chunking procedure obtains the relevant section based on a section header associated with semantically closest embedded content vectors stored in the private database,
“Systems, methods, and computer program products for using a generative artificial intelligence system to generate answers or summaries is provided. During the ingestion stage, the system receives documents and transcripts that include data associated with a theme. The data is converted into a common format and is divided into chunks. The chunks are associated with metadata tags that include chunk and data information. From the chunks, the system generates embedding vectors. During the inference stage, the system receives an information request. If the information request is a question, the system generates a vector from the question, and uses a similarity search to identify similar vectors. From the similar vectors, the system identifies chunks. If the information request includes a summary request, the system uses the metadata tags to identify chunks with summary information. The system generates an answer or a summary from the identified chunks.” (Arunachalam [Abstract]).
“Data processing engine 114 may also generate a metadata tags 406A, and associate each metadata tag in metadata tags 406A with each chunk in chunks 404A. Each metadata tag may include a project identifier associated with project data 202A. The metadata tag may also include a chunk identifier, a size of the chunk, a title of document 204A or transcript 206A that is included in the chunk, a subtitle corresponding to a section of document 204A or transcript 206A, a hierarchy of the chunk compared to other chunks, and the like. ” (Arunachalam [0055]).
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of Mondlock the capability to receive semantically close vectors related to the relevant sections. The benefit and motivation of such modification is discussed by Arunachalam in the following portion: “One of LLM models, such as LLM 110C may receive one or more chunks in the subset of chunks to generate an answer to the question in the information request. In some embodiments, LLM 110C may receive the chunk that corresponds to the most similar vector first, and then refine the answer with each subsequent chunk. For example, LLM 110C may receive a first chunk in the subset of chunks to generate an answer. Next, LLM 110C may receive a second chunk and the answer generated from the first chunk to generate a second answer. Next, LLM 110C may receive a third chunk and the answer generated using the first and second chunks. The process may continue until LLM 110C uses all chunks in the subset of chunks or until LLM 110C determines that the content of the answer no longer changes.” (Arunachalam [0068]).
Wherein the main benefit of receiving those chunks is to refine the output of an LLM using them as context for the answer.
Mondlock in view of Arunachalam does not explicitly teach, but Poirier teaches:
and wherein the section-based chunking procedure returns content associated with the section header as a complete section of the documentation rather than an arbitrary sliding-window portion of the documentation;
“In step 1508, the computing system generates a respective segment embedding for each data record segment based on the respective segment and the respective contextual metadata. In some embodiments, an embeddings generator module (e.g., embeddings generator module 512) generates the segment embeddings. The respective segment embeddings may be vector values. The respective segment embeddings may comprise respective vector embeddings and the embeddings datastore may comprise one or more vector datastores. In some embodiments, the contextual metadata is stored in respective headers of the respective segments and/or respective data records.” (Poirier [0203])
“In some embodiments, the chunking is performed by an agent, the segment embeddings are generated by another agent, and information is retrieved from the embeddings store based on the respective embeddings by an additional agent, and wherein the agents are supervised by an orchestrator. For example, the addition agent may use a machine learning model to implement a similarity machine learning process that determines a similarity between two of more of segments and/or data records based on the embeddings (e.g., vector embeddings generated based on the data segments, and the contextual metadata). Thus, for example, similarity of segments can be determined based on a similarity of the segments (e.g., similarity of passages of the segments) and similarity of the contextual metadata (or information indicated by the contextual metadata). In some embodiments, the plurality of data records may be processed by a type system prior to being any of scanned and chunked.” (Poirier [0206])
“Another step can include separating text and code (step 1605). The goal of this step is to identify and separate code and pure text. This then allows the system to further process these modalities separately. An additional step can include chunking and parsing text and code (step 1606). Having separated the text and code modalities, the purpose of this step is to identify, locate and extract contiguous pieces of code (step 1607), and conduct chunking of the text contents (step 1606) in reasonable and as contiguous fashion (e.g., that is no cutting mid-sentence or mid-paragraph especially due to breakages between pages and if possible, having chunks with contiguous topics).” (Poirier [0212])
Wherein Poirier [0203] introduce the concept of contextual metadata as the information related to the chunk that is saved in a header of the chunk, Poirier [0206] describes the agent in charge of retrieving the data based on the embeddings, that includes the context metadata saved in the header, in order to find similar segments. Poirier [0212] further describes chunking of a document in a reasonable fashion, avoiding cutting sections like paragraphs and sentences, the steps of locating, identifying contiguous chunks.
And wherein obtaining the relevant section further includes a) detecting a header of the vectors semantically closest to the query vector
“An enterprise generative artificial intelligence architecture is disclosed herein which can intelligently and efficiently crawl and index disparate data records (e.g., data records of one or more enterprise systems) across a variety of different domains using contextual information (e.g., contextual metadata) to provide improved data record identification and retrieval, access control (e.g., role-based access), and map relationships between data records. In one example, contextual information may prevent some users from accessing (e.g., viewing, retrieving) certain data records, and improve similarity evaluations used in retrieval operations (e.g., of a generative artificial intelligence process). Accordingly, the systems described herein can provide more accurate and reliable results that are also faster and more secure than existing techniques.” (Poirier [0023])
“In some embodiments, an enterprise generative artificial intelligence system can crawl, chunk, and index a corpus of data records. Data records can include documents (e.g., PDF, text, html, markdown source code or other source code, etc.), database tables, information generated by applications (e.g., artificial intelligence application insights), images, audiovisual files, executables, models (e.g., data models, machine learning models, large language models, multimodal models), and the like. More specifically, the enterprise generative artificial intelligence system preprocesses and chunks data records across different domains (e.g., data domains, industry-specific domains) of an enterprise. The chunking process partitions data records into segments (or, chunks) and can insert and/or append contextual information for the segments (e.g., as a header of the segment).” (Poirier [0024])
“The contextual information may include, for example, one or more attributes describing the segment (e.g., type of data record, size of chunk, semantic or contextual description of the chunk, access control restrictions or permissions, etc.). A segment may include a passage of a text document, a portion of database table, a sub-model of model, and so forth. The segments and/or contextual information can be stored for efficient retrieval (e.g., as part of a generative artificial intelligence process). For example, the segments and/or contextual information can be stored as embeddings (e.g., vector embeddings) which can allow efficient retrieval. In one example, contextual information can include explicit and/or inferred references between segments and/or data records. For example, the references may indicate relationships that can be used (e.g., traversed) when performing similarity evaluations or other aspects of retrieval operations.” (Poirier [0025])
“In some embodiments, embeddings generator module 128 may generate segment embeddings based on the respective segment and the respective contextual metadata. For example, the segment embedding may include a vector embedding that can used as part of a similarity machine learning process that determines similarities between data records 105-106 and 109-110 and/or segments. Accordingly, similarities can easily be determined (e.g., as part of a generative artificial intelligence retrieval operation) based on various segments and corresponding contextual information.” (Poirier [0042])
“FIG. 8A depicts a flowchart 800 of an example iterative generative artificial intelligence process using unstructured data according to some embodiments. This example process may be implemented by an enterprise generative artificial intelligence system (e.g., enterprise generative artificial intelligence system 402). In step 802, a user query is provided to a retrieval model (e.g., a retriever module of a retrieval agent module). In step 804, the retrieval model receives the query and performs a similarly search (e.g., an ANN-based search) of the vector store 806. The retrieved information is returned the retriever model and provided to a large language model in step 810. …”(Poirier [0163])
Wherein Poirier [0023] to Poirier [0025] describe an invention where the contextual information is described and utilized for an efficient retrieval and the embeddings of the segments are created (including the contextual information, that could include a header). In Poirier [0042] similarities are calculated for the retrieval, taking in account the segment and the metadata and the relationship between the queries and the similarity search is further described in Poirier [0163].
And b) retrieving from the private database, all paragraphs and table entries having the matched header to reconstruct the entire section of the documentation for use as the context;
“FIG. 15 depicts a flowchart 1500 of an example method of intelligent crawling and chunking according to some embodiments. In step 1502, a computing system (e.g., enterprise generative artificial intelligence system and/or intelligent crawling and chunking subsystem 120) scans a plurality of different data domains of an enterprise information environment. In some embodiments, a crawling module (e.g., crawling module 514 and/or crawling module 122) scans the different data domains of the enterprise information environment. In step 1504, the computing system chunks a plurality of data records of multiple enterprise data sources of the plurality of different data domains of the enterprise information environment. The data records can include any of documents, database tables, models, text, images, video, audio, artificial intelligence insights, application outputs, applications, source code, scripts, and/or compiled source code. The chunking can generate one or respective data record segments for each of the plurality of data records. In some embodiments, a chunking module (e.g., chunking module 510 and/or chunking module 124) chunks the data records.” (Poirier [0201])
“In some embodiments, the chunking is performed by an agent, the segment embeddings are generated by another agent, and information is retrieved from the embeddings store based on the respective embeddings by an additional agent, and wherein the agents are supervised by an orchestrator. For example, the addition agent may use a machine learning model to implement a similarity machine learning process that determines a similarity between two of more of segments and/or data records based on the embeddings (e.g., vector embeddings generated based on the data segments, and the contextual metadata). Thus, for example, similarity of segments can be determined based on a similarity of the segments (e.g., similarity of passages of the segments) and similarity of the contextual metadata (or information indicated by the contextual metadata). In some embodiments, the plurality of data records may be processed by a type system prior to being any of scanned and chunked.” (Poirier [0206])
“At the end of the first stage, the system can have several instances of the text, code, table and image classes, outlining different modalities in each data record. Having done this, the system can continue to the second stage, namely, building an information graph for each data record.” (Poirier [0216])
“In the second stage, in order to facilitate an effective information retrieval, the system can represent the information in each data record 1601 using an information graph. The nodes 1630, 1641, 1643 and 1645 of this graph correspond to different instances of the four classes of modality from stage one. The depiction is a (directed) bipartite graph, with edges going out from the text nodes to all other modality nodes. Establishing these edges are the primary goal of this stage.” (Poirier [0217])
”In the third stage, given the information graph, the system can outline a process for information retrieval. One approach for this starts with first embedding the content of the text nodes (and/or other textual metadata associated with other modalities) (step 1629), and storing the embeddings in a vector store 1626. Given a query 1622, the system can embed it (step 1628) and find the most relevant text chunks or text nodes 1630 associated with it. This would be the entry to the graph. At that point, the system can follow the out-going edges to other modality nodes. The classes associated with these modalities, (e.g., code, image, and table) may have a method that enables generation of relevant insights given the query. This method can be powered by different approaches, including multimodal models or other tools for understanding and querying the specific modalities. These insights, together with the text chunks and the user query can then be combined in an aggregator 1650 to a body of text or a prompt to be used for querying a model (e.g., multimodal model, large language model, etc.).” (Poirier [0219])
Wherein Poirier [0201] describes the information that could be carried in the segments, in portions, Poirier [0206] describe retrieving relevant portions that are similar and Poirier [0216], Poirier [0217] and Poirier [0219] describe retrieving those portions and organizing them in a graph, later using an aggregator to combine those portions to include them in a body of text or a prompt for an LLM.
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of Mondlock in view of Arunachalam the capability to detect and manage operations of a retrieval augmented generation based on the information present on a header of a chunk (contextual information in contextual metadata stored in the header) and to retrieve multiple sections related to each other and feed them as context of a prompt to an LLM. The benefit and motivation of such modification is discussed by Poirier in the following portion: “In some implementations, the chunking module 124 can preprocess the data records and/or segments to generate corresponding contextual information. In some embodiments, the contextual information may be included and/or represented in contextual metadata, and/or the contextual metadata may be generated from the contextual information. The contextual information may improve security, as well as accuracy and reliability of associated retrieval operations. In one example, contextual information includes contextual metadata. The contextual information can include references between segments and/or data records 105-106 and 109-110. For example, the references may indicate relationships that can be used (e.g., traversed) when performing similarity evaluations or other aspects of retrieval operations (e.g., by one or more of the agents 506). Contextual information may also include information that can assist a large language model in generating a plan and/or answers. For example, the chunking module 124 may generate contextual information for structured data chunks (or passages) that include natural language descriptions of the data records 105-106 and 109-110, locations of related data records 105-106 and 109-110, and the like.” (Poirier [0038])
Regarding claim 18, the rejection of claim 16 is incorporated, furthermore Mondlock teaches:
The system of claim 16, wherein the section-based chunking procedure obtains the relevant section of the documentation in a manner unrelated to a sliding window procedure.
“In some aspects, the RAG pipeline 300 may include at block 384A identifying the top relevant text chunks and/or identifying the top relevant data chunks at block 384B. The top relevant text chunks and/or data chunks may be identified by the relevant information identification module 130 or any other suitable program. The top relevant text and data chunks may be identified by a semantic search. The semantic search may comprise performing a KNN search of the user query embedding and text chunk and/or data chunk embeddings.” (Mondlock Column 16 lines 14-23)
Regarding claim 19, the rejection of claim 16 is incorporated, furthermore Mondlock teaches:
The system of claim 16, wherein a size of the user query and relevant section is configured to fall within an input token limit of the LLM.
“In some aspects, the query module 128 may include instructions for generating an augmented user query from the rephrased user query. The query module 128 may supplement the rephrased user query with information obtained (via the document/asset/expert module 122 and the relevant information identification module 130) from document collections 144, document collections 162, asset collections 146, asset collections 164, and/or other suitable sources to generate a prompt. For example, a rephrased user query may ask a question regarding Acme Corp.'s most recent earnings report. The query module 128 may append the contents of Acme Corp.'s earnings report when generating the augmented user query. The query module 128 may summarize an augmented user query in order to satisfy a maximum word or token limit of the LLM service 170. For example, the query module 128 may implement map reduce functionality to split Acme Corp.'s earnings report document into a plurality of text chunks and summarize each text chunk to generate a summarized output text suitable for submission to the LLM service 170. The query module 128 may implement prompt engineering to supplement the augmented user query. For example, the query module 128 may add text instructing the LLM service 170 to answer the user query with the supplemental information instead of relying upon pre-trained data.” (Mondlock Column 8 lines 19-43)
Regarding claim 20, the rejection of claim 16 is incorporated, furthermore Mondlock teaches:
The system of claim 16, wherein the private database is a vector store, and wherein the system includes one or more of a server and a retriever.
“The internal data store 140 may be owned or operated by the same organization that owns or operates the generative AI pipeline. The internal data store 140 may include a relational database (e.g., a PostgreSQL database), a non-relational datastore (e.g., a NoSQL database), a vector database (e.g., Pinecone), a web server, file server, and/or application server.” (Mondlock Column 8 lines 19 to 38).
“In some aspects, the computer-implemented method 700 may include at block 726 saving the text chunks set and text embeddings into a data store. The text chunks set and text embeddings may be saved by the document/asset/expert module 122. The text chunks set and text embeddings may be saved at block 382A of the RAG pipeline 300. The text chunks set and text embeddings may be saved into internal data store 140.” (Mondlock Column 4 line 65 to column 5 line 4).
“The server 110 may be an individual server, a group (e.g., cluster) of multiple servers, or another suitable type of computing device or system (e.g., a collection of computing resources). The server 110 may be located within the enterprise network of an organization that owns or operates the generative AI pipeline or hosted by a third-party provider.” (Mondlock Column 4 lines 31 to 36).
“In some aspects, if the intent is to query one or more document collections, then the RAG pipeline 300 may include at block 354A retrieving the one or more documents from the one or more selected document collections. The documents may be retrieved by the document/asset/expert module 122 or any other suitable program. The retrieved documents may be stored, short term or long term, in document collections 144.” (Mondlock Column 14 lines 1 to 8).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HECTOR J. CRESPO FEBLES whose telephone number is (571)272-4512 or email hcrespofebles@uspto.gov. The examiner can normally be reached Mon - Fri 7:30 - 5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at (571) 272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HECTOR J. CRESPO FEBLES/Examiner, Art Unit 2657
/DANIEL C WASHBURN/Supervisory Patent Examiner, Art Unit 2657