DETAILED ACTION
This Office Action is responsive to the Applicant’s submission, filed on March 31, 2026, amending claims 2, 22, 23 and 32. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Because new grounds of rejection presented below have not been necessitated by the Applicant’s amendments, this Office Action is non-final.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on March 17, 2026 has been considered by the Examiner.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-6, 22-25, 27-30 and 32-34 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
As described in MPEP § 2106, the analysis as to whether a claim qualifies as eligible subject matter under 35 U.S.C. § 101 includes the following determinations:
(1) Whether the claim is to a statutory category, i.e. to a process, machine, manufacture or composition of matter (“Step 1”) – see MPEP §§ 2106, subsection III, and 2106.03
(2) If the claim is to a statutory category, whether the claim recites any judicial exceptions, including certain groupings of abstract ideas (i.e., mathematical concepts, certain methods of organizing human activity, or mental processes) (“Step 2A, Prong One”) – see MPEP §§ 2106, subsection III, and 2106.04
(3) If the claim recites a judicial exception, whether the claim recites additional elements that integrate the judicial exception into a practical application (“Step 2A, Prong Two”) – see MPEP §§ 2106, subsection III, and 2106.04
(4) If the claim does not recite additional elements that integrate the judicial exception into a practical application, whether the claim recites additional elements that amount to significantly more than the judicial exception (“Step 2B”) – see MPEP §§ 2106, subsection III, and 2106.05
Claim 1
Step 1. The claim recites a process.
Step 2a, prong 1. The claim recites a judicial exception, particularly a mental process. The recitations of “perform[ing] initial embedding-based retrieval of subsets of the chunks of information from the corpus” and “re-rank[ing] at least some of the retrieved chunks of information in the subsets” are considered mental processes. As broadly recited, a human can select data from a set of data and rank the selected data.
Step 2a, prong 2. This judicial exception is not integrated into a practical application. Additional elements include training a retrieval pipeline comprising a bi-encoding model and a cross-encoding model using the corpus and a plurality of seed queries to perform the above-noted mental processes, and outputting a trained bi-encoding model and a trained cross-encoding model. This claimed training and outputting of models to perform the judicial exception, wherein the models are recited at a high level of generality, represents no more than mere instructions to apply the judicial exception on a computer, and thus does not integrate the judicial exception into a practical application. See MPEP § 2106.05(f). Claim 1 also recites “obtaining (i) a corpus comprising chunks of information and (ii) a plurality of seed queries.” However, this is indicative of insignificant extra-solution activity, i.e. mere data gathering, and is also insufficient to integrate the abstract idea into a practical application. See MPEP § 2106.5(g).
Step 2b. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As noted above, the claimed training and outputting of models (i.e. a bi-encoding model and a cross-encoding model) to perform the judicial exception represents no more than mere instructions to apply the judicial exception on a computer. As such, these claimed features do not amount to significantly more than the judicial exception. See MPEP § 2106.05(f). As also noted above, the recitation of “obtaining (i) a corpus comprising chunks of information and (ii) a plurality of seed queries” is indicative of insignificant extra-solution activity, i.e. mere data gathering. Such data gathering is well-understood, routine and conventional. See, e.g., Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014).
Accordingly, claim 1 is not eligible.
Claims 27 and 34
These claims are similar in scope to claim 1 and are therefore rejected under a similar rationale. The additional element of “at least one processing device” in claim 27 is a generic computing component and represents no more than mere instructions to apply the judicial exception on a computer. Similarly, the additional elements of a “non-transitory computer readable medium” and “at least one a processor” in claim 34 are generic computer components that represent no more than mere instructions to apply the judicial exception on a computer.
Accordingly, claims 27 and 34 are also ineligible.
Dependent claims 2-6, 22-25, 28-30 and 32-33
In claim 2, identifying one or more evaluation metrics for comparing the trained retrieval pipeline and an untrained retrieval pipeline is indicative of a mental process. Claim 2 is thus ineligible.
In claims 3 and 30, providing the plurality of seed queries to at least one large language model, and receiving additional queries from the at least one large language model, is considered insignificant extra-solution activity, i.e. mere data gathering. As noted above, such data gathering is well-understood, routing and conventional. Accordingly, claims 3 and 30 are ineligible.
In claim 4, the recitations of “identify specified chunks of information relevant to the input query” and “create a response to the input query, the response based on the one or more specified chunks of information” are considered indicative of a mental process. The additional recitations of “obtaining at input query at a retriever model” and “providing one or more of the specified chunks of information from the retriever model to a generative model” are indicative of insignificant extra-solution activity, i.e. mere data gathering. As noted above, such data gathering is well-understood, routing and conventional. Lastly, the recitations of “the retriever model including the trained bi-encoding model and the trained cross-encoding model, the retriever model configured to” identify the specific chunks, and “using the generative model to create” the response to the input query represent no more than mere instructions to apply the judicial exception on a computer. Accordingly, claim 4 is ineligible.
Claims 5 and 28 recite a mental process. Using a corpus and a plurality of seed queries to generate positive and negative training examples for a bi-encoding model and a cross-encoding model can practically be performed in the human mind. The recited formats of the training examples do not change this analysis; the claimed formats can be generated in the human mind or by using pen and paper. Accordingly, claims 5 and 28 are ineligible.
Claims 6 and 29 recite at least a mental process. Performing an evaluation of the trained bi-encoding model and the trained cross-encoding model can practically be performed in the human mind, as can performing bi-clustering to generate clusters of data related to the corpus. Determining a probability distribution based on one or more metrics associated with the clusters is also indicative of a mental process, or alternatively, a mathematical concept. Sampling data from a testing corpus based on the probability distribution is also indicative of a mental process. Accordingly, claims 6 and 29 are ineligible.
In claim 22, “identifying subsets of the chunks of information” and “determin[ing] whether the subsets of the chunks of information actually are or are not relevant to the seed queries” is considered a mental process. That the determination is made “using at least one large language model” is indicative of mere instructions to apply the judicial exception on a computer. Accordingly, claim 22 is ineligible.
In claim 23, “wherein identifying the subsets of the chunks of information comprises identifying positive and negative training examples” is indicative of a mental process. Claim 23 is thus ineligible.
In claim 24, “ranks the positive training examples and the negative training examples…ranking the positive training examples as being more relevant to the seed queries and ranking the negative training examples as being less relevant or irrelevant to the seed queries” is considered a mental process. That “the at least one large language model” performs this ranking task, like claimed, is indicative of mere instructions to apply the judicial exception on a computer. Moreover, “ranking the positive training examples and the negative training examples…ranking the negative training examples higher than the positive training examples” is considered a mental process. The claimed recitation that the cross-encoding model and the bi-encoding model perform this ranking task is also indicative of mere instructions to apply the judicial exception on a computer. Accordingly, claim 24 is ineligible.
In claims 25 and 33, the recitation of “…judge the positive and negative training examples and determine which of the positive and negative training examples to use” is considered a mental process. That the claim recites using “the at least one large language model” to perform this judgment task is indicative of mere instructions to apply the judicial exception on a computer. Accordingly, claims 25 and 33 are ineligible.
In claim 32, “identify subsets of the chunks of information” and “identify positive and negative training examples using the subsets” is considered a mental process. That “at least one large language model” is used to identify the positive and negative training examples is indicative of mere instructions to apply the judicial exception on a computer. Accordingly, claim 32 is ineligible.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1, 4, 27 and 34 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by U.S. Patent Application Publication No. 2024/0070403 to Meng et al. (“Meng”).
Regarding claims 1, 27 and 34, Meng generally describes systems and methods for training and using information-seeking dialogue systems, wherein the training is performed in a pipelined process having components of passage retrieval, re-ranking, and generating a response to a query (see e.g. paragraph 0004). Meng particularly teaches that such training entails:
obtaining (i) a corpus comprising chunks of information and (ii) a plurality of seed queries (see e.g. paragraph 0016: Meng describes a training phase in which machine-learning models are trained to perform the tasks of retrieving passages relevant to a query, re-ranking the retrieved passages, and generating a response based at least in part on a passage and a query. Meng further discloses that such training entails obtaining a training set of training data, wherein the training set includes examples of queries, passages and responses – see e.g. paragraphs 0005 and 0017. The passages in the training set constitute a corpus comprising chunks of information, wherein each individual passage is considered a chunk. The queries in the training set are considered seed queries.);
training a retrieval pipeline comprising a bi-encoding model and a cross-encoding model using the corpus and the plurality of seed queries (see e.g. paragraphs 0004: like noted above, Meng discloses that the training is performed in a pipelined process having components of passage retrieval, re-ranking, and generating a response to a query. More particularly, the training entails using the training set to train: a retriever model to identify similarities between a query and a passage; a re-ranking model to rank a plurality of passages based on a distance metric defining a distance between a passage and a query; and a generator model to generate a response to a query based on a passage – see e.g. paragraphs 0005 and 0018. Once trained, the retriever model is used to retrieve a predetermined number of passages in response to a query, the re-ranking model is then used to re-rank the retrieved passages based on the query, and lastly the generator model is used to generate a response based on one or more of the highest ranked passages of the re-ranked passages – see e.g. paragraphs 0019, 0035 and 0043. The combination of the retriever model and re-ranking model is considered a “retrieval pipeline” like claimed. Meng particularly discloses that the retriever model is implemented with a bi-encoding model – see e.g. paragraph 0036. The re-ranking model is implemented with a cross-encoding model – see e.g. paragraph 0037. Accordingly, Meng teaches training a retrieval pipeline comprising a retriever model and a re-ranking model, i.e. a bi-encoding model and a cross-encoding model, respectively. As noted, this training is performed using the training set, i.e. the corpus and plurality of seed queries.); and
outputting a trained bi-encoding model and a trained cross-encoding model (see e.g. paragraphs 0019 and 0043: Meng discloses that, once trained, the bi-encoding model and cross-encoding model are output, e.g. for testing or inference.);
wherein training the retrieval pipeline comprises:
training the bi-encoding model to perform initial embedding-based retrieval of subsets of the chunks of information from the corpus (see e.g. paragraphs 0004, 0006, 0035-0036: Meng discloses that the retriever model is trained to retrieve passages by, in part, generating an encoding for the query and encodings for each of the passages. The encodings, which are considered embeddings, are particularly used to identify the similarity between the query and each of the passages – see e.g. paragraphs 0006 and 0036. During inference, the retriever model identifies a predetermined number of the passages most similar to the query via the encodings – see e.g. paragraphs 0019 and 0036. Like noted above, the retriever model is particularly implemented with a bi-encoding model – see e.g. paragraph 0036. Accordingly, Meng teaches training the retriever model, i.e. the bi-encoding model, to perform initial embedding-based retrieval of subsets of the chunks of information, i.e. a predetermined number of the passages, from the corpus.); and
training the cross-encoding model to re-rank at least some of the retrieved chunks of information in the subsets (see e.g. paragraphs 0005 and 0018: like noted above, Meng teaches training the re-ranking model to rank a plurality of passages based on a distance metric defining a distance between a passage and a query – see e.g. paragraphs 0005 and 0018. During inference, and like further noted above, the trained re-ranking model is then used to re-rank the predetermined number of passages retrieved by the retriever model – see e.g. paragraph 0019. Like also noted above, Meng discloses that the re-ranking model is implemented with a cross-encoding model – see e.g. paragraph 0037. Accordingly, Meng teaches training the re-ranking model, i.e. the cross-encoding model, to re-rank at least some of the retrieved chunks of information in the subsets, i.e. to re-rank the predetermined number of passages retrieved by the retriever model.).
Meng thus teaches a method like that of claim 1. Meng further discloses that the above-described tasks can be implemented with at least one processor (e.g. a gpu) of a computing device (see e.g. paragraphs 0044 and 0058). Such a computing device for implementing the above-described teachings of Meng is considered an apparatus like that of claim 27. Moreover, Meng also discloses that the above-described tasks can be implemented via program code stored on non-transitory computer-readable storage media for execution by a processor (see e.g. paragraph 0059). Such a non-transitory computer-readable storage medium comprising code to implement the above-described teachings of Meng is considered a non-transitory computer readable medium like that of claim 34.
As per claim 4, Meng further teaches: (i) obtaining an input query at a retriever model, the retriever model including the trained bi-encoding model and the trained cross-encoding model, the retriever model configured to identified specified chunks of information relevant to the input query; (ii) providing one or more of the specified chunks of information from the retriever model to a generative model; and (iii) using the generative model to create a response to the input query, the response based on the one or more specified chunks of information (see e.g. paragraphs 0005 and 0018: like noted above, Meng teaches training a retriever model to identify similarities between a query and a passage, a re-ranking model to rank a plurality of passages, and a generator model to generate a response to a query based on a passage. Once trained, and like further noted above, the retriever model is used to retrieve a predetermined number of passages in response to a query, the re-ranking model is then used to re-rank the retrieved passages based on the query, and lastly the generator model is used to generate a response based on one or more of the highest ranked passages of the re-ranked passages – see e.g. paragraphs 0019, 0035 and 0043. Like also noted above, Meng discloses that the retriever model is implemented with a bi-encoding model, and that the re-ranking model is implemented with a cross-encoding model – see e.g. paragraphs 0036 and 0037. The combination of the retriever model and re-ranking model taught by Meng is considered a “retriever model” like in claim 4. Accordingly, Meng teaches: (i) obtaining an input query at a retriever model, i.e. at the retriever model and re-ranking model, wherein the retriever model includes the trained bi-encoding model and the trained cross-encoding model, and is configured to identify specified chunks of information, i.e. one or more highest ranked passages, relevant to the input query; (ii) providing one or more of the specified chunks of information from the retriever model to a generative model, i.e. to the generator model; and (iii) using the generative model to create a response to the input query, wherein the response is based on the one or more specified chunks of information.). Consequently, Meng further teaches a method like that of claim 4.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claim 2 is rejected under 35 U.S.C. 103 as being unpatentable over the U.S. Patent Application Publication to Meng cited above, and also over the article entitled, “Enhancing Q&A with Domain-Specific Fine-Tuning and Iterative Reasoning: A Comparative Study” by Nguyen et al. (“Nguyen”).
Regarding claim 2, Meng teaches a method like in claim 1, as is described above, which comprises training a retrieval pipeline comprising a bi-encoding model and a cross-encoding model using a corpus and a plurality of seed queries. Meng, however, does not disclose that training the retrieval pipeline further comprises identifying and outputting one or more evaluation metrics for comparing the trained retrieval pipeline and an untrained retrieval pipeline, as is required by claim 2.
Nguyen nevertheless generally teaches identifying and outputting one or more evaluation metrics (e.g. relevancy, faithfulness) for comparing a trained (i.e. fine-tuned) retrieval pipeline (e.g. an embedding/retriever model) and an untrained retrieval pipeline:
Retrieval-augmented generation (RAG) techniques, which combine information retrieval and generative models (Lewis et al. 2021), have shown promise in boosting the quality of LLM output in Q&A tasks. RAG systems leverage the strengths of both retrieval and generation components to provide contextually relevant and informative responses. While there is a lack of established quantification of RAG accuracy, early findings suggest that generic RAG does not perform well in complex domains such as finance. In one instance, RAG based on generic LLMs such as GPT-4-Turbo fails to answer 81% of the questions derived from Securities and Exchange Commission (SEC) financial filings (Islam et al. 2023).
The underperformance of generic LLMs and RAG in domain-specific Q&A has motivated us to research into and create methods to adapt and extend such models and techniques. In this paper, we describe and quantify the gains in accuracy from two major methods: model fine-tuning and iterative reasoning.
Fine-tuning is a way to adapt language models to specific domains and tasks (Devlin et al. 2018; Liu et al. 2019) by training them on domain-specific data and having them capture the nuances and intricacies of a particular field. In a typical RAG workflow, there are two principal models that can be considered for fine-tuning: the Embedding Model, whose tasks are indexing the information in the corpus and retrieving information relevant to the posed question, and the Generative Model, whose task is synthesizing an answer. We have performed experiments and analyses comparing baseline generic-LLM-based RAG and RAG with one or both of fine-tuned retrieval and fine-tuned generation on the FinanceBench dataset, and observed promising gains from fine-tuning.
(Section 1 “Introduction.” Emphasis added.).
We use automated retrieval quality metrics from LlamaIndex’s evaluation suite, which employs an LLM to validate the retrieved documents as they relate to the query, response, and reference documents.
Relevancy. Relevancy measures whether or not the response for the query is in line with the context information. A response scores 0 if it is irrelevant and 1 otherwise.
Faithfulness. Faithfulness measures whether a response is supported by the retrieved context, scoring 0 when it is not and 1 when it is.
Context similarity. Context similarity measures the semantic difference between the retrieved contexts and the reference contexts. It is calculated by finding the cosine similarity between the contexts when mapped into an embedding space. By default, LlamaIndex rounds this score to 0 or 1 if the cosine is below or above 0.8 respectively.
(Section 3.3.1 “Retrieval Quality Metrics.” Emphasis added.)
We experimented with and evaluated the outputs of the following technical configurations:
Generic RAG: bge-large-en with gpt-3.5-turbo-0125.
RAG with Fine-Tuned Generator: bge-large-en with fine-tuned gpt-3.5-turbo-0125.
RAG with Fine-Tuned Retriever: fine-tuned bge-large-en with gpt-3.5-turbo-0125.
Fully Fine-Tuned RAG: fine-tuned bge-large-en with fine-tuned gpt-3.5-turbo-0125.
Generic RAG with OODA Reasoning: bge-large-en with gpt-3.5-turbo-0125 and OODA reasoning.
Each system uses the same VectorStoreIndex for the documents, a top-k of 10 documents, and default RAG prompts, all from base LlamaIndex. For the full reproducible experimental setup, see OpenSSA’s FinanceBench example. Other combinations such as Fine-Tuned RAG with OODA Reasoning shall be addressed in a future publication (see Section 6).
Our test set consists of the remaining 41 questions in the publicly-available FinanceBench test of valid questions that were not the 100 selected for training.
(Section 4 “Experiments & Results.” Footnote omitted and emphasis added.).
Table 1 summarizes the retrieval quality scores achieved by various configurations across the FinanceBench test set questions. As the OODA reasoning involves multiple iterative retrievals, it is not directly comparable to the one-step retrieval processes of these RAG systems, and is therefore not included in this table.
(Section 4.1 “Retrieval Quality Results”).
PNG
media_image1.png
154
476
media_image1.png
Greyscale
It would have been obvious to one of ordinary skill in the art, having the teachings of Meng and Nguyen before the effective filing date of the claimed invention, to modify the method taught by Meng so as to identify and output one or more evaluation metrics for comparing the trained retrieval pipeline and an untrained retrieval pipeline, as is taught by Nguyen. It would have been advantageous to one of ordinary skill to utilize such a combination because it would enable a user to identify whether the training actually improves the retrieval pipeline, as is evident from Nguyen (see e.g. the excerpts provided above). Accordingly, Meng and Nguyen are considered to teach, to one of ordinary skill in the art, a method like that of claim 2.
Claims 3 and 30 are rejected under 35 U.S.C. 103 as being unpatentable over the U.S. Patent Application Publication to Meng cited above, and also over the article entitled, “Context Aware Query Rewriting for Text Rankers using LLM” by Anand et al. (“Anand”).
Regarding claims 3 and 30, Meng teaches a method and apparatus like in claims 1 and 27, respectively, which comprises training a retrieval pipeline comprising a bi-encoding model and a cross-encoding model using a corpus and a plurality of seed queries, as is described above. Meng, however, does not disclose that training the retrieval pipeline comprises (i) providing the plurality of seed queries to at least one large language model, and (ii) receiving additional queries from the at least one large language model, as is required by claims 3 and 30.
Anand nevertheless generally teaches training a retrieval pipeline (and particularly, a ranker therein) by, in part, providing a plurality of seed queries to at least one large language model (LLM), and receiving additional queries (i.e. rewritten queries) from the at least one large language model:
In this section, we describe the proposed Context Aware Rewriter (CAR) framework for ambiguous query reformulation. Figure 1 shows the working of the proposed framework during the training and inferences stages. During training, the framework is composed of a query reformulation phase and a document ranking phase. In the query reformulation phase, we employ a context-aware prompting of a LLM (Section 3.1). During inference, the ranker which was fine-tuned on disambiguated queries is directly employed to rank documents for new queries without rewriting them, as shown in Figure 1.
(Section 3 “METHOD.” Emphasis added.).
Algorithm 1 details the query reformulation process. Given an ambiguous query
q
, our goal is to generate a rewrite
q
*
that helps disambiguate the intent of the original query. We generate query rewrites by conditioning the LLM (
R
E
) through few-shot prompting on the relevant document for the query
d
+
to avoid topic drift. An example of the few-shot prompting employed is shown in Figure 2.
q
*
=
R
E
(
q
,
d
+
)
We first instruct the LLM (Figure 2) to perform the task of query reformulation in the context of the given document. We then provide an ambiguous query (
q
) concatenated with the corresponding relevant document as a part of the prompt. This context-aware prompting of LLMs results in better rewrites without topic drifts, as the LLM is conditioned on intents conveyed in the document. We also introduce a constraint on the maximum length of the output sequence generated. This constraint coupled with grounding of the generation on the relevant document context through prompting to prevent hallucination and topic drift.
(Section 3.1 “Query Rewriter.” Emphasis added.).
Our goal is to train a model for re-ranking documents on the query rewrites for under-specified queries. Given a query-document pair (
q
*
,
d
) as input, the ranking models output a relevance score. This score can then be used to rank documents based on their relevance to the given query.
Formally, the training set comprises pairs
q
i
*
,
d
i
, where
q
i
*
is a rewritten (disambiguated) query using an LLM and
d
i
is a relevant or irrelevant document to the query. The aim is to fine-tune a ranker
R
that predicts a relevance score
y
^
∈
[
0
;
1
]
given a reformulated query
q
*
and a document
d
:
R
:
q
*
,
d
⟼
y
^
(1)
The fine-tuned ranking model (
R
) can be employed to re-rank a set of documents obtained from a first-stage lightweight frequency-based retriever. Recent studies have indicated that pre-trained language models which jointly model queries and documents tasks [11, 32, 39] have demonstrated significant performance on ranking tasks. In this work, we employ the BERT[13] model for ranking. The input to the ranker is of format:
C
L
S
q
S
E
P
d
[
S
E
P
]
. (2)
We employ the pointwise loss to train the ranker. Let us assume of a mini batch of
N
training examples
x
i
,
y
i
i
=
1
,
…
,
N
.
The ranking task is cast as a binary classification problem, where each training instance
x
i
=
q
i
*
,
d
i
is a query-document pair and
y
i
∈
0,1
is a relevance label. The predicted score of
x
i
is denoted as
y
^
i
. The cross-entropy loss function is defined as follows:
L
P
o
i
n
t
=
-
1
N
∑
i
=
1
N
y
i
∙
l
o
g
y
^
i
+
1
-
y
i
∙
l
o
g
(
1
-
y
^
i
)
(3)
(Section 3.2 “Ranking Phase.” Emphasis added.).
Anand discloses that training the retrieval pipeline (particularly, the ranker) with the additional (i.e. rewritten) queries provides significant improvement (see e.g. the Abstract, which recites “[i]n our extensive experiments, we find that fine-tuning a ranker using re-written queries offers a significant improvement of up to 33% on the passage ranking task and up to 28% on the document ranking task when compared to the baseline performance of using original queries.”).
Accordingly, it would have been obvious to one of ordinary skill in the art, having the teachings of Meng and Anand before the effective filing date of the claimed invention, to modify the method and apparatus taught by Meng so as to train the retrieval pipeline by, in part, (i) providing the plurality of seed queries to at least one large language model, and (ii) receiving additional queries from the at least one large language model, as is taught by Anand. It would have been advantageous to one of ordinary skill to utilize such a combination because training the retrieval pipeline with the additional (i.e. rewritten) queries can provide significant improvement, as is taught by Anand (see e.g. the Abstract). Accordingly, Meng and Anand are considered to teach, to one of ordinary skill in the art, a method like that of claim 3 and an apparatus like that of claim 30.
Claims 5 and 28 are rejected under 35 U.S.C. 103 as being unpatentable over the U.S. Patent Application Publication to Meng cited above, over the article entitled “Hybrid retrievers with generative re-rankers” by Marek Kozlowski (“Kozlowski”), and also over U.S. Patent Application Publication No. 2025/0328565 to Larson et al. (“Larson”).
Regarding claims 5 and 28, Meng teaches a method and apparatus like in claims 1 and 27, respectively, which entail training a retrieval pipeline comprising a bi-encoding model and a cross-encoding model using a corpus and a plurality of seed queries, as is described above. Meng further suggests using the corpus and the plurality of seed queries to generate positive and negative training examples for the bi-encoding model and the cross-encoding model (see e.g. paragraphs 0005 and 0017: like noted above, Meng teaches obtaining a training set of training data, wherein the training set includes examples of corresponding queries, passages and responses; the passages of the training set are considered a corpus and the queries are considered seed queries. Meng further teaches that the bi-encoding model is trained using a query and corresponding positive and negative passages – see e.g. paragraph 0036. The cross-encoder model is similarly trained with a query and corresponding positive and negative passages – see e.g. paragraph 0037. Training examples comprising such a query and corresponding positive and/or negative passages are considered positive and negative training examples.). However, Meng does not explicitly describe the particular format of the positive and negative training examples, and thus does not disclose that: (i) the positive and negative training examples for the bi-encoding model include triples having a form (query, more relevant document/chunk, less relevant document/chunk); and (ii) the positive and negative training examples for the cross-encoding model include triples having a form (query, passage/chunk, score), as is required by each of claims 5 and 28.
Kozlowski nevertheless generally teaches training a bi-encoding model with positive and negative training examples that include triples having the form (query, more relevant document/chunk, less relevant document/chunk):
The dense retrievers in the proposed approach are Sentence-BERT-type encoders. We trained them in the Bi-encoder architecture, and used them with FAISS index support. FAISS is a library for the efficient similarity search and clustering of dense vectors. Bi-encoder architectures are used because sentence embedding in a vector space is needed for efficient comparison during semantic search or clustering. The queries and passages are passed independently to the sentence transformer to produce fixed-size embeddings. These can then be compared using cosine similarity to identify matching passages for a given query. The training of the dense retrievers is performed by fine-tuning encoders—specifically, we fine-tuned the RoBERTa model used in our bi-encoder architecture. During training, we used a loss function called MultipleNegativesRankingLoss, number_of_epochs = 1−10 and batch_size = 32. We pass triplets in the format: (query, positive_passage, negative_passage), where negative passage is a hard negative example (not positive one, but lexically similar to the positive one) that is retrieved by lexical search in the whole passage corpora. We used Elasticsearch to obtain (max = 10) hard negative examples for given positive passages. We trained two types of dense retrievers: a) using RoBERTa-base-v2 as a transformer model, and training one epoch on 500 000 triplets (135 000 Poleval ones and 370 000 Polish MS Marco ones randomly selected); and b) using RoBERTa-large-v2 as a transformer model, and training 10 epochs on a few million triplets (several million Polish MS Marco ones randomly selected). We then fine-tuned for one epoch on all 135 000 PolEval triplets.
(Section IV.A “Candidate passage retrieval.” Footnotes omitted and emphasis added.).
It would have been obvious to one of ordinary skill in the art, having the teachings of Meng and Kozlowski before the effective filing date of the claimed invention, to modify the method and apparatus taught by Meng such that the positive and negative training examples for training the bi-encoding model includes triples having a form of (query, more relevant document/chunk, less relevant document/chunk), as is taught by Kozlowski. It would have been advantageous to one of ordinary skill to utilize such triples because they can effectively train the bi-encoding model, as is evident from Kozlowski (see e.g. the excerpt provided above).
Larson generally teaches training a cross-encoding model (i.e. a “re-ranker”) using training examples of the form (query, passage/chunk, score) (see e.g. paragraphs 0176-0182: Larson describes a re-ranker, which can be implemented with a cross-encoder, that receives a question and a potentially-relevant text block and generates a relevance score indicating how much relevance the text block has to the question. Larson further suggests that the re-ranker can be trained with training examples that include triples having a question, text block, and ground-truth annotation, i.e. relevance score – see e.g. paragraph 0214. The triples, that is, have the form of (query, passage/chunk, score) like claimed.).
It would have been obvious to one of ordinary skill in the art, having the teachings of Meng, Kozlowski and Larson before the effective filing date of the claimed invention, to modify the method and apparatus taught by Meng and Kozlowski such that the positive and negative training examples for training the cross-encoding model includes triples having a form of (query, passage/chunk, score), as is taught by Larson. It would have been advantageous to one of ordinary skill to utilize such triples because they can effectively train the cross-encoding model, as is evident from Larson (see e.g. paragraph 0214). Accordingly, Meng, Kozlowski and Larson are considered to teach, to one of ordinary skill in the art, a method like that of claim 5 and an apparatus like that of claim 28.
Claims 6 and 29 are rejected under 35 U.S.C. 103 as being unpatentable over the U.S. Patent Application Publication to Meng cited above, over the article entitled “Data splitting for artificial neural networks using SOM-based stratified sampling” by May et al. (“May”), and also over the article entitled “Biclustering in data mining” by Busygin et al. (“Busygin”).
Regarding claims 6 and 29, Meng teaches a method and apparatus like in claims 1 and 27, respectively, which comprises training a retrieval pipeline comprising a bi-encoding model and a cross-encoding model using a corpus and a plurality of seed queries, as is described above. Meng further teaches performing an evaluation of the trained bi-encoding model and the trained cross-encoding model, e.g. in a testing phase (see e.g. paragraphs 0019 and 0053-0054). Meng, however, does not disclose that the evaluation comprises: (i) performing bi-clustering to generate clusters of data related to the corpus; (ii) determining a probability distribution based on one or more metrics associated with the clusters; and (iii) sampling data from a testing corpus based on the probability distribution, as is required by claims 6 and 29.
May generally teaches that the selection of training, testing and validation data from a dataset is important to generate accurate test or validation performance of an artificial neural network (ANN) model:
Methods commonly used during the development of statistical models to ensure good generalization include hold-out cross validation, k-fold cross-validation, ensemble training, and Bayesian regularization (Sarle, 1997). In ANN applications, the hold-out is most commonly employed, and it is synonymous with stop training or early stopping. In this approach, a subset of data is reserved to periodically test the performance of the network during training. Training is stopped when the test error reaches an optimum value, as further training will result in over-fitting, and hence ensures a generalized fit. Furthermore, when model selection is employed to compare alternative models or optimize ANN architectures, the models can potentially be optimistically biased towards the test data. In order to avoid testing bias, a second hold-out is required for validating the optimal ANN model (Maier & Dandy, 2000).
Regardless of the number of data subsets required, the issue for modellers is that the hold-out of data itself can prove to be yet another source of bias and variance. If the data subsets are selected inappropriately, then the training, test and validation data may not be equally representative of the problem domain, and will generate inaccurate test or validation performance. Variation in test and validation error may be observed for repeated instances of sampling, which creates uncertainty regarding the model performance that is gauged based on a single instance of training, test and validation data. The uncertainty due to sampling variance may be significantly greater than other sources of model uncertainty such as network initialization, training and architecture (LeBaron & Weigend, 1998).
(Section 1. “Introduction.” Emphasis added).
May further discloses that the training, testing and validation data can be selected from the dataset according to a “stratified sampling” approach, which entails performing clustering to generate clusters of data related to the dataset:
Stratified sampling partitions the data into
H
homogeneous groups (or strata) of size
N
h
, and data are sampled from within each stratum (Cochran, 1977). The partitioning of the data forces sampling to be distributed throughout all regions of the input–output space, and ensures that adequate representation of input–output tuplets can be achieved. For multi-variate data it is convenient to use partitioning or clustering algorithms to generate the strata (Mulvey, 1983). Gill, Smith, and Bagnall (2004) refer to this as cluster-based stratified sampling (CBSS). Several examples of data splitting have been described using different clustering algorithms, including k-means clustering, the self-organizing map (SOM) (Kohonen, 1995) and fuzzy c-means clustering (Kaufman & Rousseeuw, 1990). Bowden et al. (2002), Daszykowski, Walczak, and Massart (2002) and Svozil, Pospichal, and Kvnasnicka (1995) applied a partitioning of data based on the self-organizing map prior to sampling. The methodology has since been adopted in several similar ANN applications to water resources modelling (Anctil & Lauzon, 2004; Kingston, 2006; Zhang, Stanley, & Smith, 2004). Shahin et al. (2004) used fuzzy c-means clustering to partition the data for ANN model development, where the membership values were used to guide sampling.
(Section 2. “Data splitting methods.” Emphasis added.)
Moreover, May suggests that the training, testing and validation data can then be selected by determining a probability distribution based on one or more metrics (e.g. the size and/or standard deviation) associated with the clusters, and by sampling data from the dataset (particularly, from each cluster) based on the probability distribution:
Sample allocation refers to the sampling within each stratum, and is an important aspect of SBSS in terms of selecting data for ANN development. In past applications of SBSS, the sample allocation has varied. Bowden et al. (2002), Daszykowski et al. (2002) and Svozil et al. (1995) draw single samples from each cell for training, test and validation, although these are based on different grid sizes. Alternatively, Kingston (2006) randomly samples all data within each partition in proportion to the desired sample sizes for training, testing and validation, so that all of the available data are used. It remains unclear which, if any, is the most appropriate approach to take.
In stratified random sampling, the sampling within strata is usually uniform random sampling, and an allocation rule identifies the number of samples drawn per stratum, which is referred to as the quota. Three basic rules can be considered for determining the sample quota (Cochran, 1977; Kpedekpo, 1973): equal allocation, proportional allocation, and Neyman allocation.
(Section 3.3. “Sample allocation.” Emphasis added.).
The allocation of samples from within strata can also be determined based on the size of individual strata,
N
h
, to yield proportional allocation. In this case, the number of samples
n
h
to be taken from stratum
h
can be determined according to
n
h
=
N
h
∑
j
=
1
H
N
j
n
N
, (3)
which results in the overall selection of
n
samples.
(Section 3.3.2. “Proportional allocation.” Emphasis added.).
Neyman allocation considers both the size of the stratum,
N
h
, and the intra-stratum standard deviation,
σ
h
. The Neyman allocation for stratum
h
is determined according to
n
h
=
N
h
σ
h
∑
j
=
1
H
N
j
σ
j
n
N
, (4)
where
σ
j
is the intra-stratum standard deviation. Here, the sample allocation is increased for strata that are either large, or have increased variance. Neyman allocation yields an optimal sample when used to draw a stratified sample for the estimation of conventional statistics (Cochran,1977). It should be noted that
σ
j
conventionally refers to the intra-stratum standard deviation of the variable for which statistical estimates are later generated, since this defines the optimality of the sample allocation rule. However, in this study the standard deviation is based on the multivariate form,
σ
=
σ
x
1
2
+
…
+
σ
x
p
2
+
σ
y
2
, (5)
where
σ
x
1
2
is the intra-stratum variance of component
x
i
and
σ
y
2
is the variance of the output variable,
y
. The multivariate standard deviation describes the intra-stratum variability with respect to input–output tuplets, which is for sampling data for regression. In this case the sampling rate is expected to be greater where there is greater variance in either the inputs, output, or both.
(Section 3.3.3. “Neyman allocation.” Emphasis added.).
It would have been obvious to one of ordinary skill in the art, having the teachings of Meng and May before the effective filing date of the claimed invention, to modify the method and apparatus taught by Meng such that performing the evaluation of the trained bi-encoding model and the trained cross-encoding model (i.e. generating the test and or validation dataset to perform the evaluation) comprises: (i) performing clustering to generate clusters of data related to the dataset (i.e. the corpus); (ii) determining a probability distribution based on one or more metrics associated with the clusters; and (iii) sampling data from the dataset (i.e. a testing corpus) based on the probability distribution, as is taught by May. It would have been advantageous to one of ordinary skill to utilize such a combination because it can generate more accurate test or validation performance, as is suggested by May (see e.g. the portion of section 1 “Introduction” excerpted above.). Accordingly, Meng and May are considered to teach a method similar to that of claim 6 and an apparatus similar to that of claim 29, but do not explicitly teach performing bi-clustering to generate the clusters of data related to the corpus, as is required by claims 6 and 29.
Busygin nevertheless generally teaches performing bi-clustering to generate clusters of data (see e.g. the Abstract, which recites: “Biclustering consists in simultaneous partitioning of the set of samples and the set of their attributes (features) into subsets (classes)…In this paper we review the most widely used and successful biclustering techniques and their related applications.”).
It would have been obvious to one of ordinary skill in the art, having the teachings of Meng, May and Busygin before the effective filing date of the claimed invention, to modify the method and apparatus taught by Meng and May such that bi-clustering like taught by Busygin is used to generate the clusters of data related to the corpus. It would have been advantageous to one of ordinary skill to utilize such bi-clustering because it enables the clustering of documents and words simultaneously, and also the discovery of important relations between document and word classes, as is taught by Busygin (see e.g. section 6.2. “Text mining”). Accordingly, Meng, May and Busygin are considered to teach, to one of ordinary skill in the art, a method like that of claim 6 and an apparatus like that of claim 29.
Claims 21 and 31 are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Meng and Anand, which is described above, over the article entitled “Balanced Data Sampling for Language Model Training with Clustering” by Shao et al. (“Shao”), and also over the article entitled “Sampling Techniques in Statistics for Machine Learning” by Kartik Chaudhary (“Chaudhary”).
Regarding claims 21 and 31, Meng and Anand teach a method like in claim 3 and an apparatus like in claim 30, as is described above, and which entail training a retrieval pipeline by providing a plurality of seed queries to at least one large language model, and retrieving additional queries from the at least one large language model. Anand suggests generating the additional (i.e. rewritten) queries by using a template to create a prompt from selected example queries, and providing the prompt to at least one large language model to generate one or more additional queries (see e.g. section 3.1 “Query Rewriter” and Figure 2: Anand demonstrates a prompt given to a large language model, wherein the prompt comprises a selected query and is provided to the large language model to generate an additional, i.e. rewritten, query based on the selected query. Anand teaches that each of a plurality of queries in a training batch is rewritten via the large language model – see e.g. “Algorithm 1” – and so it is apparent that a template would be used to create the prompt for each query to be rewritten.). Meng and Anand, however, do not teach or suggest that the selected example queries used in the prompt to create the additional queries are selected by: (i) converting the plurality of seed queries into dense vector representations; (ii) clustering the dense vector representations into multiple clusters; (iii) randomly selecting a subset of the clusters; and (iv) randomly selecting an example query from each of the selected subset of the clusters, as is required by claims 21 and 31.
Shao generally describes “ClusterClip Sampling,” which “balance[s] the text distribution of training data for better model training.” (Abstract). In particular, Shao discloses that ClusterClip Sampling entails: (i) converting training samples (e.g. text documents) into dense vector representations (i.e. data embeddings); (ii) clustering the dense vector representation into multiple clusters; (iii) selecting a subset of the clusters; and (iv) selecting one or more example training samples from each of the selected clusters:
To address the above challenges, we propose ClusterClip, a cluster-based sampling strategy with clip operation to mitigate overfitting. This sampling strategy has two steps. Firstly, we leverage data clustering to reflect the data distribution. Based on off-the-shelf NLP tools, the semantic related texts could be grouped into the same cluster. By calculating the size of each semantic cluster, we can evaluate the data rarity. Secondly, we perform the data sampling using the cluster information. The documents from different clusters are sampled evenly at the beginning of the training, thus encouraging the model to learn from rare documents instead of wasting computation on common texts. As the training progresses, a clip operation is applied. If certain documents are sampled too many times, these documents are clipped and no longer sampled from the dataset. Thus, ClusterClip rebalances the data distribution, which facilitates the model to learn rare documents but avoids severe overfitting of repeated texts by clipping.
(Section 1 “Introduction.” Emphasis added.).
Aiming to describe and manipulate the distribution of texts in the training set, we introduce data clustering to group samples into semantic clusters. We choose not to rely on metadata from the dataset itself, as such information is often absent or fuzzy (Gao et al., 2021; Azerbayev et al., 2023). Instead, data clustering can automatically discover semantic similar documents and group these data points into clusters. Specifically, we first utilize off-the-shelf transformer-based models to generate text representations for each data sample. Then we conduct a K-Means clustering on these generated data embeddings to group samples into clusters. By clustering, we classify data with similar topics into the same subset. Thus, we can analyze the data distribution and rebalance the data distribution when sampling the training data.
We choose out-of-the-box transformer-based embedding models and the K-Means method in the experiments as these methods are well-established, efficient at scale, and can produce semantic-related data clusters. Other embedding functions, including rule-based or model-based, could also be utilized for more accurate clustering. Comparing the impact of different embedding or clustering methods on the data sampling strategies would be a valuable topic and is our future work.
(Section 3.1 “Data Clustering.” Emphasis added.).
After data clustering, the clusters can describe the long tail distribution of the training set. Based on the cluster information, ClusterClip Sampling increases the sampling weights of rare documents and decreases the weights of common texts. Moreover, a clip operation is introduced to mitigate over fitting. Thus, ClusterClip Sampling balances the learning on both very common texts and extremely rare documents
Uniform Sampling At the beginning of training, ClusterClip Sampling performs a Uniform Sampling from the clusters, which aims to up-sample rare data points and down-sample common texts. We ensure that each cluster has the same probability of being sampled. After sampling the cluster, amount of tokens in each cluster. This also improves the data diversity within the batch as it balances the occurrence of samples in each cluster in a batch.
Clip Operation When uniformly sampling the data, documents from small clusters can be sampled a huge number of times. In this case, the model will suffer from overfitting on these small semantic clusters and not learn well on the whole training set. To solve this issue, we further propose a clip operation to add a maximum repetition of each sample. When different clusters are uniformly sampled, small clusters can be consumed multiple times. The ClusterClip will record the repeated times of each cluster. When one cluster has been consumed a certain number of times, the cluster will be knocked out and will not be sampled in further training. Thus, the model will see a sample at most a certain number of times, which mitigates the overfitting.
(Section 3.2 “ClusterClip Sampling.” Emphasis added.).
It would have been obvious to one of ordinary skill in the art, having the teachings of Meng, Anand and Shao before the effective filing date of the claimed invention, to modify the method and apparatus taught by Meng and Anand such that that the sample queries (i.e. the queries used in the prompt provided to the large language model to create the additional, rewritten queries) are selected by: (i) converting the plurality of training samples (i.e. the seed queries) into dense vector representations; (ii) clustering the dense vector representations into multiple clusters; (iii) selecting a subset of the clusters; and (iv) selecting an example training sample (i.e. example query) from each of the selected subset of the clusters, as is taught by Shao. It would have been advantageous to one of ordinary skill to utilize such a combination because it provides “the ability to improve model learning efficiency and generalization without relying on dataset-specific metadata or complicated optimizations,” as is taught by Shao (see Section 1 “Introduction”). Accordingly, Meng, Anand and Shao teach a method similar to that of claim 21 and an apparatus similar to that of claim 31, but do not explicitly disclose that the subset of clusters is randomly selected, or that the example query is randomly selected from each of the selected clusters, as is required by claims 21 and 31.
Chaudhary nevertheless generally teaches using cluster sampling to select a number of data points within a population, wherein the cluster sampling entails (i) clustering the data points into multiple clusters, (ii) randomly selecting a subset of the clusters, and (iii) randomly selecting a data point from each of the selected subset of the clusters:
Cluster Sampling
Cluster sampling is done when the total population can be grouped into several groups such that the groups are mutually homogenous and internally heterogeneous.
In this technique, the population is first divided into small groups as discussed above and these groups are known as clusters.
Once clusters are formed, simple random sampling is applied to select a few clusters for the study. Once some clusters are selected(sampled), there are two possibilities-
take all the elements from each selected cluster,
Choose samples from each cluster based on simple random sampling or stratified sampling technique and combine later.
In the second case, we are performing sampling in two stages. This kind of cluster sampling is called ‘two-stage’ cluster sampling.
Thus, cluster sampling could be ‘multi-stage’ depending upon the requirements.
(Section 2.1 “Probability Sampling.” Emphasis added.).
It would have been obvious to one of ordinary skill in the art, having the teachings of Meng, Anand, Shao and Chaudhary before the effective filing date of the claimed invention, to modify the method and apparatus taught by Meng, Anand and Shao such that that the subset of clusters is randomly selected, and the data point (i.e. the example query) is randomly selected from each of the selected clusters, as is taught by Chaudhary. It would have been advantageous to one of ordinary skill to utilize such a combination because the resulting set of selected samples would approximate the entire set, as is suggested by Chaudhary (see e.g. Section 3 “Qualities of Good Sampling Techniques,” which recites “[s]ampling techniques aim at selecting a small portion of the total population that is representative of that population.“). Accordingly, Meng, Anand, Shao and Chaudhary are considered to teach, to one of ordinary skill in the art, a method like that of claim 21 and an apparatus like that of claim 31.
Claims 22, 23 and 32 are rejected under 35 U.S.C. 103 as being unpatentable over the U.S. Patent Application Publication to Meng cited above, and also over U.S. Patent Application Publication No. 2026/0017496 to Gofman et al. (“Gofman”).
Regarding claim 22, Meng teaches a method like in claim 1, as is described above, which comprises obtaining a corpus comprising chunks of information and obtaining a plurality of seed queries, and training a retrieval pipeline using the corpus and the plurality of seed queries. Meng, however, does not disclose that training the retrieval pipeline comprises (i) identifying subsets of the chunks of information, and (ii) using at least one large language model to determine whether the subsets of the chunks of information actually are or are not relevant to the seed queries, as is required by claim 22.
Similar to Meng, Gofman describes a retrieval pipeline that comprises a retriever model and a reranker model (e.g. a cross-encoding model), wherein the retrieval model performs initial embedding-based retrieval of subsets of documents from a corpus, and the reranker model re-ranks at least some of the retrieved documents (see e.g. paragraphs 0002-0003 and 0033). Gofman further teaches training the retrieval pipeline by, in part, identifying subsets of chunks of information (e.g. documents) and using at least one large language model to determine whether the subsets of the chunks of information actually are or are not relevant to seed queries (see e.g. paragraphs 0005 and 0046-0055: Gofman teaches using a large language model (LLM) to generate, for each of a plurality of documents in a corpus, one or more synthetic queries. The synthetic queries are considered “seed queries” like claimed. Gofman further teaches using the LLM to rank, for each synthetic query, a plurality of documents associated with the synthetic query based on a relevance of the plurality of documents to the synthetic query – see e.g. paragraphs 0005 and 0078-0079. The plurality of documents relevant to the synthetic query are identified from the corpus using the LLM and/or a retriever model – see e.g. paragraphs 0006-0009 and 0075-0077. Gofman discloses that the generated synthetic queries and ranked plurality of documents are used as a training dataset to train the retrieval pipeline, particularly the reranker model therein – see e.g. paragraphs 0019 and 0095-0097. Accordingly, Gofman teaches training a retrieval pipeline by, in part, identifying subsets of the chunks of information, e.g. documents, and using at least one LLM to determine whether the subsets of the chunks of information actually are or are not relevant to the seed queries, i.e. synthetic queries.).
It would have been obvious to one of ordinary skill in the art, having the teachings of Meng and Gofman before the effective filing date of the claimed invention, to modify the method taught by Meng such that training the retrieval pipeline (i.e. generating a dataset to train the retrieval pipeline) comprises identifying subsets of the chunks of information, and using at least one large language model to determine whether the subsets of the chunks of information actually are or are not relevant to seed queries, as is taught by Gofman. It would have been advantageous to one of ordinary skill to utilize such a combination, because it can reduce the amount of human intervention needed to generate a training dataset to train the retrieval pipeline for a specialized domain, as is taught by Gofman (see e.g. paragraphs 0034-0035). Accordingly, Meng and Gofman are considered to teach, to one of ordinary skill in the art, a method like that of claim 22.
As per claim 23, it would have been obvious, as is described above, to modify the method taught by Meng such that training the retrieval pipeline comprises identifying subsets of the chunks of information, and using at least one large language model to determine whether the subsets of the chunks of information actually are or are not relevant to the seed queries, as is taught by Gofman. Gofman particularly teaches that identifying the subsets of the chunks of information comprises identifying positive and negative training examples (see e.g. paragraphs 0075 and 0095). Accordingly, the above-described combination of Meng and Gofman is further considered to teach a method like that of claim 23.
Regarding claim 32, Meng teaches an apparatus like that of claim 27, as is described above, which obtains a corpus comprising chunks of information and obtains a plurality of seed queries, and trains a retrieval pipeline using the corpus and the plurality of seed queries. Meng, however, does not disclose that training the retrieval pipeline comprises (i) identifying subsets of the chunks of information, and (ii) using at least one large language model to identify positive and negative training examples using the subsets, as is required by claim 32.
Like noted above, Gofman describes a retrieval pipeline that comprises a retriever model and a reranker model (e.g. a cross-encoding model), wherein the retrieval model performs initial embedding-based retrieval of subsets of documents from a corpus, and the reranker model re-ranks at least some of the retrieved documents (see e.g. paragraphs 0002-0003 and 0033). Gofman further teaches training the retrieval pipeline by, in part, identifying subsets of chunks of information (e.g. documents) and using at least one large language model to identify positive and negative training examples using the subsets (see e.g. paragraphs 0005 and 0046-0055: Gofman teaches using a large language model (LLM) to generate, for each of a plurality of documents in a corpus, one or more synthetic queries. The synthetic queries are considered “seed queries” like claimed. Gofman further teaches using the LLM to rank, for each synthetic query, a plurality of documents associated with the synthetic query based on a relevance of the plurality of documents to the synthetic query – see e.g. paragraphs 0005 and 0078-0079. The plurality of documents relevant to the synthetic query are identified from the corpus using the LLM and/or a retriever model – see e.g. paragraphs 0006-0009 and 0075-0077. Documents identified as relevant to the synthetic query are positive examples, and documents not relevant to the synthetic query are negative examples – see e.g. paragraphs 0075 and 0095. Gofman discloses that the generated synthetic queries and positive and negative examples are used as a training dataset to train the retrieval pipeline, particularly the reranker model therein – see e.g. paragraphs 0019 and 0095-0097. Accordingly, Gofman teaches training a retrieval pipeline by, in part, identifying subsets of the chunks of information, e.g. documents, and using at least one LLM to identify positive and negative training examples using the subsets.).
It would have been obvious to one of ordinary skill in the art, having the teachings of Meng and Gofman before the effective filing date of the claimed invention, to modify the apparatus taught by Meng such that training the retrieval pipeline (i.e. generating a dataset to train the retrieval pipeline) comprises identifying subsets of the chunks of information, and using at least one large language model to identify positive and negative training examples using the subsets, as is taught by Gofman. It would have been advantageous to one of ordinary skill to utilize such a combination, because it can reduce the amount of human intervention needed to generate a training dataset to train the retrieval pipeline for a specialized domain, as is taught by Gofman (see e.g. paragraphs 0034-0035). Accordingly, Meng and Gofman are considered to teach, to one of ordinary skill in the art, an apparatus like that of claim 32.
Claims 24, 25 and 33 are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Meng and Gofman described above, and also over the article entitled, “Gecko: Versatile Text Embeddings Distilled from Large Language Models” by Lee et al. (“Lee”).
Regarding claim 24, Meng and Gofman teach a method like in claim 23, as is described above, which comprises identifying subsets of chunks of information, including positive and negative training examples, and using at least one large language model to determine whether the subsets of the chunks of information are relevant to seed queries. Gofman particularly teaches that the at least one large language model ranks the positive training examples (see e.g. paragraph 0078). Meng and Gofman, however, do not teach that the at least one large language model ranks the positive training examples and the negative training examples, wherein the at least one large language model ranks the positive training examples as being more relevant to the seed queries and ranks the negative training examples as being less relevant or irrelevant to the seed queries, as is required by claim 24. Moreover, Meng and Gofman also do not teach that the cross-encoding model and the bi-encoding model rank the positive training examples and the negative training examples, wherein the cross-encoding model and the bi-encoding model rank the negative training examples higher than the positive training examples, as is further required by claim 24.
Similar to Meng and Gofman, Lee teaches training a retrieval pipeline (particularly, an embedding model) by, in part, identifying subsets of chunks of information (i.e. the top N passages for each query), and using at least one large language model to determine whether the subsets of the chunks of information actually are or are not relevant to seed queries (i.e. the respective queries):
Text embedding models represent natural language as dense vectors, positioning semantically similar text near each other within the embedding space (Gao et al., 2021; Le and Mikolov, 2014; Reimers and Gurevych, 2019). These embeddings are commonly used for a wide range of downstream tasks including document retrieval, sentence similarity, classification, and clustering (Muennighoff et al., 2023). Instead of building separate embedding models for each downstream task, recent efforts seek to create a single embedding model supporting many tasks.
(Section 1. “Introduction.” Emphasis added).
In this section, we introduce our two-stage approach that uses LLMs to generate FRet. Traditional approaches for training embedding models often rely on large, manually labeled datasets. However, creating such datasets is time-consuming, expensive, and often results in undesirable biases and lack of diversity. In this work, we present a novel method for generating synthetic data for training multi-task text embedding models, leveraging the power of LLMs through a two-step distillation process. The overall process of generating FRet is illustrated in Figure 2.
…
LLM-based Positive and Negative Mining Most models that utilize synthetic queries are trained with
(
q
,
p
s
e
e
d
)
pairs, which assumes that
p
s
e
e
d
is a good positive target for 𝑞 (Dai et al., 2022; Jeronymo et al., 2023). While this is likely true in most cases, we hypothesize that there could be a more relevant passage than
p
s
e
e
d
somewhere in our corpus of web passages. Essentially, in the previous section, we sampled
P
t
,
q
p
s
e
e
d
from the LLM, but this does not guarantee that
p
s
e
e
d
maximizes
P
p
q
,
t
over all the passages in the corpus. This intuition is supported by our observation that generated queries often focus on a particular aspect of a relatively long passage. Hence, we propose a method that leverages LLMs to discover more relevant positive passages along with a good hard negative for the generated query.
In particular, we use an existing embedding model to retrieve top 𝑁 neighbors
P
=
p
(
1
)
,
…
,
p
(
N
)
from the corpus given a generated query 𝑞. We then employ the same LLM used for the query generation to rank these retrieved passages based on their relevance to the query. Specifically, we use two well-known few-shot prompted LLM ranking functions: query likelihood and relevance classification. Query likelihood uses an LLM to measure the log-likelihood of a generated query 𝑞 given a passage 𝑝, i.e.,
Q
L
q
,
p
=
L
L
M
q
p
,
P
Q
L
(Sachan et al., 2022). Herein,
P
Q
L
is a prompt containing an instruction for judging query likelihood and several few-shot examples of relevant query and passage pairs (Drozdov et al., 2023). Relevance classification (Zhuang et al., 2023) uses an LLM to measure the log-likelihood of a specific relevance label given the query 𝑞 and a passage 𝑝, i.e.,
R
C
q
,
p
=
L
L
M
l
a
b
e
l
q
,
p
,
P
R
C
, where
P
R
C
is a prompt with few-shot examples for grading the relevance of each query-passage pair. The prompts
P
Q
L
and
P
R
C
are identical for every example. Our pilot study demonstrated that each prompting method (i.e. QL and RC) excels in different tasks, so we ensemble the rankings from two different prompting results with the standard Reciprocal Rank Fusion (RRF) approach (Cormack et al., 2009), obtaining a ranking function 𝑅(𝑞, 𝑝). As shown in Appendix A, the ensembling greatly improves the robustness of our model across diverse tasks.
Given the scores from LLMs after ensembling, we index the set of passages
P
according to their ranking, i.e.
P
=
p
1
,
…
,
p
N
where if
i
<
j
,
R
q
,
p
i
≥
R
q
,
p
j
. We then choose a new positive target:
p
+
=
arg max
p
∈
P
R
q
,
p
=
p
1
Importantly,
p
+
can be different from
p
s
e
e
d
and conveys an approximation to the global preference of the LLM over the entire corpus. Table 5 lists examples where the
p
+
differs from
p
s
e
e
d
, demonstrating that the pair (𝑞,
p
s
e
e
d
) can be sub-optimal and there can be more relevant passages for 𝑞 globally. We find that the relabeling of the positive passage (i.e.,
p
+
≠
p
s
e
e
d
) happens for about 15% in our dataset.
Similarly, the LLM scores can also be used to select hard negative passages. One straightforward option is to select the lowest scoring negative, i.e.
p
-
=
p
N
. Another is to sample from the remaining nearest neighbors, i.e.
p
-
~
P
\
{
p
+
}
. We explore both options in §4.3. Combining all of our generation results along with the positive and negative mining, we create the FRet dataset, comprised of 6.6M examples, each containing a task, a query, a positive passage, and a negative passage.
(Section 3.2. “Fret: Two-Step LLM Distillation.” Footnote omitted and emphasis added.).
As indicated above, Lee particularly teaches that identifying the subsets of the chunks of information comprises identifying positive and negative training examples, and wherein the at least one large language model ranks the positive training examples and the negative training examples, the at least one large language model ranking the positive training examples as being more relevant (i.e. most relevant) to the seed queries and ranking the negative training examples as being less relevant (e.g. the lowest or lower ranked passage/chunk in the subset of chunks) or irrelevant to the seed queries.
It would have been obvious to one of ordinary skill in the art, having the teachings of Meng, Gofman and Lee before the effective filing date of the claimed invention, to modify the method taught by Meng and Gofman such that the at least one large language model ranks the positive training examples and the negative training examples, wherein the at least one large language model ranks the positive training examples as being more relevant to the seed queries and ranks the negative training examples as being less relevant or irrelevant to the seed queries, as is taught by Lee. It would have been advantageous to one of ordinary skill to utilize such a combination, because it can produce “hard negative” training examples, which can enable the retrieval pipeline to distinguish positive passages from less relevant passages (i.e. the hard negatives) when given the queries, as is suggested by Lee (see e.g. section 3.3 “Unified Fine-tuning Mixture,” which recites “[f]or fine-tuning we optimize the in-batch cross-entropy loss, where
q
i
should distinguish
p
i
+
from the hard negative
p
i
-
, other passages in the batch
p
j
+
j
=
1
B
, and other queries in the batch
q
j
j
=
1
B
\
q
i
.”). Gofman generally teaches that, during training, the retrieval pipeline ranks the training examples; if the ranking differs from the ranking of the training examples provided by the LLM, the weights of the retrieval pipeline are updated so that the rankings provided by the retrieval pipeline align more with those of the LLM (see e.g. paragraphs 0096 and 0097). It thus follows that, during training, the retrieval pipeline (i.e. the cross-encoding model and the bi-encoding model) ranks the positive training examples and the negative training examples, wherein the ranking can differ from the ranking provided by the LLM (e.g. the cross-encoding model and the bi-encoding model rank the negative training examples higher than the positive training examples). Accordingly, Meng, Gofman and Lee are considered to teach, to one of ordinary skill in the art, a method like that of claim 24.
Regarding claim 25, Meng and Gofman teach a method like in claim 23, as is described above, which comprises identifying subsets of chunks of information, including positive and negative training examples, and using at least one large language model to determine whether the subsets of the chunks of information are relevant to seed queries. Meng and Gofman, however, do not explicitly disclose that using the at least one large language model to determine whether the subsets of the chunks of information are or are not relevant to the seed queries comprises using the at least one large language model to judge the positive and negative training examples and determine which of the positive and negative training examples to use, as is required by claim 25.
Lee nevertheless teaches using at least one large language model to judge (i.e. rank) positive and negative training examples and thereby determine which of the positive and negative training examples to use (see e.g. the portions of section 1. “Introduction” and section 3.2. “Fret: Two-Step LLM Distillation,” which are excerpted above).
It would have been obvious to one of ordinary skill in the art, having the teachings of Meng, Gofman and Lee before the effective filing date of the claimed invention, to modify the method taught by Meng and Gofman such that using the at least one large language model to determine whether the subsets of the chunks of information are or are not relevant to the seed queries comprises using the at least one large language model to judge the positive and negative training examples and determine which of the positive and negative training examples to use, as is taught by Lee. It would have been advantageous to one of ordinary skill to utilize such a combination, because it can produce “hard negative” training examples, which can enable the retrieval pipeline to distinguish positive passages from less relevant passages (i.e. the hard negatives) when given the queries, as is suggested by Lee (see e.g. section 3.3 “Unified Fine-tuning Mixture,” which recites “[f]or fine-tuning we optimize the in-batch cross-entropy loss, where
q
i
should distinguish
p
i
+
from the hard negative
p
i
-
, other passages in the batch
p
j
+
j
=
1
B
, and other queries in the batch
q
j
j
=
1
B
\
q
i
.”). Accordingly, Meng, Gofman and Lee are considered to teach, to one of ordinary skill in the art, a method like that of claim 25.
Regarding claim 33, Meng and Gofman teach an apparatus like that of claim 32, as is described above, which identifies subsets of chunks of information, and uses at least one large language model to identify positive and negative training examples using the subsets. Meng and Gofman, however, do not explicitly disclose using the at least one large language model to determine whether the subsets of the chunks of information actually are or are not relevant to the seed queries including using the at least one large language model to judge the positive and negative training examples and determine which of the positive and negative training examples to use, as is required by claim 33.
Lee nevertheless teaches using at least one large language model to judge (i.e. rank) positive and negative training examples and thereby determine which of the positive and negative training examples to use (see e.g. the portions of section 1. “Introduction” and section 3.2. “Fret: Two-Step LLM Distillation,” which are excerpted above).
It would have been obvious to one of ordinary skill in the art, having the teachings of Meng, Gofman and Lee before the effective filing date of the claimed invention, to modify the apparatus taught by Meng and Gofman so as to use the at least one large language model to determine whether the subsets of the chunks of information actually are or are not relevant to the seed queries, including by using the at least one large language model to judge the positive and negative training examples and determine which of the positive and negative training examples to use, as is taught by Lee. It would have been advantageous to one of ordinary skill to utilize such a combination, because it can produce “hard negative” training examples, which can enable the retrieval pipeline to distinguish positive passages from less relevant passages (i.e. the hard negatives) when given the queries, as is suggested by Lee (see e.g. section 3.3 “Unified Fine-tuning Mixture,” which recites “[f]or fine-tuning we optimize the in-batch cross-entropy loss, where
q
i
should distinguish
p
i
+
from the hard negative
p
i
-
, other passages in the batch
p
j
+
j
=
1
B
, and other queries in the batch
q
j
j
=
1
B
\
q
i
.”). Accordingly, Meng, Gofman and Lee are considered to teach, to one of ordinary skill in the art, an apparatus like that of claim 33.
Claim 26 is rejected under 35 U.S.C. 103 as being unpatentable over the combination of Meng, Gofman and Lee described above, and also over U.S. Patent Application Publication No. 2025/0077940 to Cui et al. (“Cui”).
Regarding claim 26, Meng, Gofman and Lee teach a method like in claim 25, as is described above, which comprises using at least one large language model to determine whether subsets of chunks of information are relevant to seed queries, including using the at least one large language model to judge positive and negative training examples and determine which of the positive and negative training examples to use. Gofman particularly teaches that a first large language model processes pairs of chunks of information to determine which of the chunks of information is more or less relevant to the seed queries (see e.g. paragraphs 0078-0081 and 0083-0093). The training samples are generated based on the chunks selected by the first large language model (see e.g. paragraph 0095). Lee provides a similar teaching (see e.g. “LLM-based Positive and Negative Mining” on pages 5-6). Meng, Gofman and Lee, however, do not disclose that a second large language model processes candidate chunks of information selected by the first large language model to generate training samples, as is further required by claim 26.
Cui nevertheless generally teaches using a second large language model to review the outputs of a first large language model (see e.g. paragraphs 0002-0003).
It would have been obvious to one of ordinary skill in the art, having the teachings of Meng, Gofman, Lee and Cui before the effective filing date of the claimed invention, to modify the method taught by Meng, Gofman and Lee so as to use a second large language model to process the outputs (i.e. the candidate chunks of information) selected by the first large language model like taught by Cui and thereby generate the training samples. It would have been advantageous to one of ordinary skill to utilize such a combination, because it can improve the accuracy of the outputs, as is suggested by Cui (see e.g. paragraphs 0002-0003). Accordingly, Meng, Gofman, Lee and Cui are considered to teach, to one of ordinary skill in the art, a method like that of claim 26.
Response to Arguments
The Examiner acknowledges the Applicant’s amendments to claims 2, 22, 23 and 32. In response to these amendments, the objections presented in the previous Office Action to claims 2, 22, 23 and 33 are respectfully withdrawn.
In response to the Applicant’s arguments, the 35 U.S.C. § 101 rejections presented in the previous Office Action to claims 27-33, for being directed to non-statutory subject matter, are respectfully withdrawn.
The 35 U.S.C. § 101 rejections presented in the previous Office Action to claims 1-6, 22-25, 27-30 and 32-34, for being directed to an abstract idea without significantly more, are however respectfully maintained. Regarding these rejections, the Applicant refers to Example 39 of the Office’s § 101 guidance to demonstrate that training a neural network (i) involves mathematical operations but is not directed to mathematical operations, (ii) does not involve a mental process, and (iii) is not directed to a method of organizing human activity.
In response, the Examiner respectfully notes that the rejection does not rely on training per se being judicial exception. Instead, taking claim 1 for example, the rejection indicates that the recited “perform[ing] initial embedding-based retrieval of subsets of the chunks of information from the corpus” and “re-rank[ing] at least some of the retrieved chunks of the information in the subsets,” and not the training per se, is considered a judicial exception (i.e. a mental process). The Applicant’s reliance on Example 39, which is found eligible because it does not recite any judicial exceptions, is thus misplaced.
The Applicant also refers to claim 3 of Example 47 of the Office’s § 101 guidance. The Examiner respectfully notes that, like the claims of the instant application, claim 3 does recite judicial exceptions. The training step in claim 3 is directed to a mathematical concept because the claim explicitly recites using a selected training algorithm that includes a backpropagation algorithm and a gradient descent algorithm. The claim also recites a mental process (e.g. “detecting one or more anomalies in network traffic”). Claim 3 of Example 47 is found eligible not because of the neural network training, but because it comprises additional elements (e.g. “detecting a source address associated with the one or more malicious network packets in real time”) that integrate the judicial exception into an abstract idea. The rejected claims of the instant application on the other hand have no such additional elements, as is indicated in the rejections above.
The Applicant argues that humans cannot mentally perform claim 1, e.g. mentally train machine learning models to perform initial embedding-based retrieval of subsets of chunks of information. The Examiner agrees. However, the Examiner respectfully submits that humans can perform elements of claim 1, e.g. perform initial embedding-based retrieval of subsets of chunks of information and re-rank some of the retrieved chunks of information in the subsets. These elements constitute an abstract idea (i.e. mental process), as is indicated above. Claim 1 additionally recites training models (i.e. a bi-encoding model and a cross-encoding model) to perform this recited abstract idea. However, training models to perform an abstract idea does not necessarily integrate the abstract idea into a practical application or amount to significantly more than the abstract idea.
For example, claim 2 of Example 47 of the Office’s § 101 guidance recites “detecting one or more anomalies in a data set using the trained ANN.” Here, “detecting one or more anomalies in a data set” is deemed a mental process. The additional element reciting “using the trained ANN” provides nothing more than mere instructions to implement the abstract idea on a computer, and thus does not integrate the mental process into a practical application or amount to significantly more than the mental process. The claims of the instant application similarly recite a mental process, e.g. performing embedding-based retrieval of subsets of chunks of information, as is noted above. Similar to claim 2 of Example 47, the claims of the instant application further recite using a neural network (i.e. training a bi-encoding model) to perform the mental process. Consequently, like in claim 2 of Example 47, training the bi-encoding model provides nothing more than mere instructions to implement the abstract idea on a computer, and thus does not integrate the mental process into a practical application or amount to significantly more than the mental process.
The Applicant further argues that the claimed features are directed to solving a technical problem by providing training for a retrieval pipeline that trains machine learning models for performing initial embedding-based retrieval of subsets of chunks of information and re-ranking retrieved chunks in the subsets. However, like noted above, performing initial embedding-based retrieval of subsets of chunks of information and re-ranking retrieved chunks in the subsets is a mental process. The Applicant’s arguments support this notion:
In addition, the Applicant's claims are directed to solving a specific technical problem. As noted in the Applicant's specification, applying large language models to real-world mission-critical applications remains challenging. Among other reasons, this can be due to the tendency of large language models to be trained for general-purpose usage. This makes it difficult to apply the large language models to specialized domains. To address these deficiencies, the Applicant's specification discloses a retrieval engine configured to search a domain-specific corpus for relevant documents or chunks to provide context to a large language model that is trained and fine-tuned for specific data. While searching for and retrieving these relevant documents one-by-one could theoretically be performed, it may not be effectively performed in a timely manner, such as when a universe of documents could include huge quantities of documents, websites, or other information across any number of fields that are irrelevant to the specialized field at issue. The sheer quantity of documents, websites, or other information could be as broad as all publicly-accessible data over the Internet. In such cases, individual users would be overwhelmed, and selecting individual documents, websites, or other information would be time- and resource-consuming and likely lead to highly-relevant information being missed.
The claimed features help to overcome these types of issues by providing training for a retrieval pipeline that trains machine learning models for performing initial embedding-based retrieval of subsets of chunks of information and re-ranking retrieved chunks in the subsets. This alleviates the need for users to manually sift through massive quantities of searchable information without becoming overwhelmed. Among other things, this can help eliminate the need for users to individually review and rank information on a document-by-document basis. The Applicant's claims are therefore (i) not directed to any abstract idea, (ii) are directed to a practical application, and (iii) are directed to significantly more than any abstract idea.
(Applicant’s Remarks, pages 12-13. Emphasis added.).
As indicated in the above excerpt, the retrieval pipeline training is intended to improve the mental process, e.g. it “alleviates the need for users to manually sift through massive quantities of searchable information without being overwhelmed.” However, an improvement in the judicial exception itself is not an improvement in technology. See MPEP 2106.05(a), subsection II. Mere automation of manual processes is also not sufficient to show an improvement in computer functionality. Credit Acceptance Corp. v. Westlake Services, 859 F.3d 1044, 1055, 123 USPQ2d 1100, 1108-09 (Fed. Cir. 2017). Accordingly, the Examiner respectfully maintains that the above-noted claims rejected under 35 U.S.C. § 101 are directed to an abstract idea without significantly more.
The Applicant’s arguments addressing the 35 U.S.C. § 103 rejections presented in the previous Office Action have been fully considered and are persuasive. Therefore, the rejections have been withdrawn. However, upon further search and consideration, new grounds of rejection are presented above.
Conclusion
The new grounds of rejection presented above are not necessitated by Applicant's amendments. Accordingly, this Office Action is non-final.
The prior art made of record on form PTO-892 and not relied upon is considered pertinent to applicant’s disclosure. The applicant is required under 37 C.F.R. §1.111(C) to consider these references fully when responding to this action. In particular, the U.S. Patent Application Publication to Ngan et al. cited therein demonstrates a retrieval pipeline that comprises a trained bi-encoding model that performs initial embedding-based retrieval of subsets of chunks of information from a corpus, and a trained cross-encoding model that re-ranks at least some of the retrieved chunks of information.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to BLAINE T BASOM whose telephone number is (571)272-4044. The examiner can normally be reached Monday-Friday, 9:00 am - 5:30 pm, EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Matt Ell can be reached at (571)270-3264. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/BTB/
7/24/2026
/MATTHEW ELL/Supervisory Patent Examiner, Art Unit 2141