DETAILED ACTION
This action is responsive to the application filed on 06/22/2026. Claims 1, 3-5, and 7-11 are pending and have been examined. This action is Non-final.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Applicant’s claim for the benefit of a prior-filed application under 35 U.S.C. 119(e) or under 35 U.S.C.
120, 121, 365(c), or 386(c) is acknowledged.
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 06/22/2026 has been entered.
Response to Arguments
Argument 1 (101 rejection): The applicant’s core position is that the amended claims are eligible because they are directed to “improvements in training the machine learning model itself” and to “an improvement to how the machine learning model itself operates,” which the applicant contends the updated MPEP 2106.05(a) and the precedential Ex Parte Desjardins now expressly recognize as improvements to computer functionality. The applicant seizes on the examiner’s own Response-to-Arguments language (“training another model using that generated data,”, “model training based on calculated similarity information”) as an admission that the claim recites the eligible improvement, argues that Desjardins draws no distinction between improving a “model-training mechanism itself” and “training another model using generated data,” and insists there is no legal requirement that a claim “explain how the computer itself is improved” (see MPEP 2106.04(d)(1)). The applicant points to specification [0068]-[0069] for the asserted benefits (better similarity accuracy, reduced compute and time, higher-accuracy matching), cites McRO, and runs the same theory through Step 2A Prong Two (a “specific transfer-learning architecture” that should not be assessed at a high level of generality) and Step 2B (“significantly more” solving insufficient training data for certain user segments). Claims 10 and 11 make no independent argument and simply ride on claim 1 (see Remarks, pages 9-18).
Examiner Response to Argument 1: The examiner has considered the applicant's arguments but does not find them persuasive. Ex Parte Desjardins is distinguishable: the claims held eligible there recited adjusting a model's parameters to optimize performance on a new task while protecting its performance on previously-learned tasks (mitigating “catastrophic forgetting”), which the panel found to be “an improvement to how the machine learning model itself operates,” and the new MPEP 2106.05(a) examples (xiii) and (xiv) reflect that same narrow holding of improvements based on adjustments to a model's parameters. Amended claim 1 recites no such improvement to how a model operates as it recites generating vector representations, computing and comparing similarity values, selecting a subset of combinations, a manual determination of similarity, and training a model on the selected data, which are mathematical concepts and mental processes carried out using generic bi-encoder, cross-encoder, and learning-device components, and using a cross-encoder to generate labels for training a bi-encoder is a conventional knowledge-distillation arrangement in which each model operates in its ordinary/generic manner. The examiner's prior statement that the claim recites “training another model using that generated data” is a characterization of the abstract idea, not an admission of eligibility, and Desjardins does not hold otherwise, because its eligible improvement was to the operation of the model itself rather than the mere fact that one model is trained on data generated by another. The asserted benefits at paragraphs [0068]-[0069] do not prove eligibility, because unlike the reduced storage in Desjardins, which flowed from an architectural change to the model, the reduction here flows from processing fewer product combinations, i.e., performing less of the abstract calculation, and a faster or less resource-intensive result from carrying out an abstract algorithm on less data is not an improvement to the computer. Contrary to the applicant's reliance on MPEP 2106.04(d))), the rejection does not require the claim to “explain” any improvement, rather, neither the claim nor the specification identifies an improvement to computer functionality, only an improved abstract result. “McRO” is also distinguishable, as the rules there improved an existing technological process and the claim was not directed to the abstract idea of the rules themselves, whereas here the claimed operations are the mathematical computations and selections themselves. Under Step 2A Prong Two the additional elements were evaluated individually and in combination and are generic model components that amount to instructions to apply the exception, the specificity residing in the abstract steps, not in any technology-improving element, and under Step 2B they are well-understood, routine, and conventional and supply no inventive concept, with the new “manual determination” and “consists only of” limitations adding only a mental process and a further mathematical selection step;. Accordingly, the rejection under 101 is maintained, and claims 10 and 11 are rejected for the same reasons.
Argument 2: The examiner has considered the applicant's arguments but does not find them persuasive. Regarding newly-added Feature [12], Thakur teaches a manually-annotated “gold” set of highly-similar pairs used with a cross-encoder that labels pairs and retains “only certain pairs,” while Qu teaches using a cross-encoder to confirm labels and retain only the pairs it scores as highly similar with high confidence, alongside manually labeled data, and to the extent neither reference expressly limits the second data to pairs confirmed by both a manual determination and the cross-encoder, the “consists only of” limitation is at least rendered obvious, as it would have been obvious, to maximize training-data quality and remove noisy or false labels (the stated purpose of Qu's confidence-based selection and Thakur's “keep only certain pairs” filtering), to retain only the combinations that both a manual determination and the cross-encoder confirm, which is a predictable use of known label-quality techniques with a reasonable expectation of success. Regarding the “product” elements, the recited “product information,” “product title,” and “product” terms describe the textual content input to the models and impose no limitation on how the bi-encoder, cross-encoder, or third model operate, each performing the identical operations whether the text is a product title, a sentence, or a passage; the difference is only in the content of the text and is not entitled to patentable weight where it does not change how the steps are performed (see MPEP 2111.05), and there is no inconsistency because Zhang supplies the “product” context (item title tokens as input, item embeddings as output) while Thakur, Lu, and Qu supply the multi-model training pipeline performed on that same textual input. Regarding Features [9]-[10], the “product title for one product as input” is not supplied by a generic rationale alone, as Zhang clearly discloses receiving a product title for a single item and outputting a corresponding embedding, and Thakur teaches the third model independently encoding each single input into a dense vector. Regarding the “silver dataset,” it is the set of pairs the cross-encoder has labeled as highly similar, and combined with Zhang, whose encoded items are products described by item titles, it reads on the recited second data indicating combinations of products determined highly similar based on the second vector representations. Finally, regarding motivation and reasonable expectation of success, the rejection sets forth a rationale because each reference operates on textual inputs and the encoding techniques are questionable to whether the text is an item title, a sentence, or a passage, so Zhang's “short” item titles are encoded in the same manner. In re Fine, In re Wilson, and Amgen don't encourage a different result because the rejection identifies the specific teachings relied upon and articulates reasoning with rational underpinning. The applicant's arguments regarding claims 10 and 11 and the dependent claims rely on the same contentions as claim 1 and are not persuasive for the same reasons, and the rejection under 103 is maintained.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition
of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the
conditions and requirements of this title.
Claims 1, 3-5, and 7-11 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding claim 1,
Step 1: This claim is directed to a method, which is one of the four statutory categories. Therefore, claim 1 satisfies Step 1.
Step 2A Prong 1:
“generating, using a first model comprising a bi-encoder, first vector representations based on product information data for the plurality of products, and generating first data indicating one or more combinations of products determined to be highly similar to each other from among the plurality of products based on similarities between the first vector representations” - This limitation is directed to a mathematical concept, as generating vector representations and determining similarity between those vector representations to identify highly similar combinations involves mathematical calculations and operations performed on data, and thus the limitation is directed to math.
“generating, using a second model comprising a cross-encoder, second vector representations from product information of two products included in the first data and generating second data indicating one or more combinations of products determined to be highly similar to each other based on the second vector representations” – This limitation is directed to a mathematical concept, as generating vector representations for paired inputs and determining similarity between products based on those representations involves mathematical calculation and mathematical relationships, and thus the limitation is directed to math.
“wherein the second model is configured to output a similarity value between products with a higher accuracy than the first model” – This limitation is directed to a mathematical concept, as it recites a comparative, quantified relationship between the similarity outputs of two models, which is an evaluation and comparison of mathematical results, and thus the limitation is directed to math.
“wherein the generating the second data further comprises determining, based on data related to a combination of products that are manually determined to be highly similar to each other according to a manual determination, whether the combination of products manually determined to be highly similar is also determined to be highly similar based on the second vector representations” - This limitation is directed to a mental process and a mathematical concept. The recited “manual determination” of whether products are highly similar is a process that can be performed in the human mind using evaluation, observation, and judgment, or with the aid of pen and paper, and is thus a mental process. Determining whether the manually-identified combination “is also determined to be highly similar based on the second vector representations” is a comparison against the model’s computed similarity output, and is thus a mathematical concept, and thus the limitation is considered both math and a mental process.
“wherein the second data consists only of combinations of products that are determined to be highly similar based on the second vector representations and determined to be highly similar according to the manual determination” – This limitation is directed to a mathematical concept and a mental process, as restricting the second data to only those combinations satisfying both a computed similarity condition based on the second vector representations and a manual determination of similarity involves a mathematical selection/comparison operation and a human evaluation and judgment, and thus the limitation is considered both math and a mental process.
Step 2A Prong 2 and Step 2B:
“A method for training a model for identifying a product in accordance with a predetermined search condition from among a plurality of products, performed by a learning device, the method comprising:” - The limitation recites a method of learning a model used for searching for a product with a predetermined search condition to be performed on a learning device. The limitation recites mere instructions to apply onto a computer, and thus the limitation does not integrate to a practical application, nor provide significantly more than the judicial exception (see MPEP 2106.05(f)).
“wherein the first model is configured to receive product information comprising a product title for one product as input, and to output a vector corresponding to the input product information…wherein the second model is configured to receive product information comprising product titles for two products, and to output a vector corresponding to the input product information of the two products…wherein the third model is configured to receive product information comprising a product title for one product as input, and to output a vector corresponding to the input product information” - This limitation recites receiving product information as input and outputting a vector corresponding to that input. This is directed to insignificant extra-solution activity in the nature of mere data gathering and data output, which cannot be integrated to a practical application (see MPEP 2106.05(g)). Furthermore, under Step 2B, receiving/sending data is a well-understood, routine, and conventional activity that cannot provide significantly more than the judicial exception (see MPEP 2106.05(d)(II)).
“training a third model to generate vector representations from the product information using the second data as training data” - This limitation recites training a third model using the generated data as training data. This amounts to no more than mere instructions to apply the exception on a computer, which cannot be integrated to a practical application, nor can it provide significantly more than the judicial exception (see MPEP 2106.05(f)).
“wherein the first model, the second model, and the third model are different models, wherein a number of product combinations of the plurality of products represented in the first data is less than a number of all combinations of the plurality of products” - This limitation recites that the models are different and that the number of combinations represented in the first data is less than all combinations. This amounts to no more than generally linking the use of the exception to a particular field of use and merely limiting the volume of data processed, which cannot be integrated to a practical application, nor can it provide significantly more than the judicial exception (see MPEP 2106.05(h)).
Thus, claim 1 is non-patent eligible. Claims 10 and 11 are analogous to claim 1 and thus will face the same rejection.
Regarding claim 3,
Step 1: The claim is directed to a method, which falls under the category of a process. The claim satisfies Step 1.
Step 2A Prong 1:
“The method according to claim 1, wherein similarity between the first vector representations is computed based on distance between two of the first vector representations.” - The limitation is directed to the computing a similarity based on distance between two vector representations. The limitation is directed to mathematical calculation/concept, and it is considered math.
There are no elements to be evaluated under Step 2A Prong 2 and Step 2B.
Thus, claim 3 is non-patent eligible.
Regarding claim 4,
Step 1: The claim is directed to a method, which falls under the category of a process. The claim satisfies Step 1.
Step 2A Prong 1:
“The method according to claim 1, wherein in the step of generating the second data, similarity is computed based on a score computed based on the second vector representations.” - The limitation is directed to computing similarity based on a score value based on the second vector representations. The limitation is directed to mathematical calculation/concept, and it is considered math.
There are no elements to be evaluated under Step 2A Prong 2 and Step 2B.
Thus, claim 4 is non-patent eligible.
Regarding claim 5,
Step 1: The claim is directed to a method, which falls under the category of a process. The claim satisfies Step 1.
There are no elements to be evaluated under Step 2A Prong 1.
Step 2A Prong 2 and Step 2B:
“The method according to claim 1, wherein the third model comprises a bi-encoder.” - The limitation recites that the third model comprises a bi-encoder. The limitation amounts to no more than merely limiting to a field of use/environment, and thus the limitation does not integrate to a practical application, nor provide significantly more than the judicial exception (see MPEP 2106.05(h)).
Thus, claim 5 is non-patent eligible.
Regarding claim 7,
Step 1: The claim is directed to a method, which falls under the category of a process. The claim satisfies Step 1.
Step 2A Prong 1:
“compute vector representations lower in dimension than the vector representations.” - The limitation is directed to computing vector representations that are lower in dimension. The limitation is directed to the use of mathematical calculations/concept, and thus the limitation is directed to math.
Step 2A Prong 2 and Step 2B:
“The method according to claim 1, further comprising a step of using a dimensionality reduction encoder to” - The limitation recites a step of using a dimensionality reduction encoder to compute the vector representations that are lower than in dimension. The limitation is directed to mere instructions to apply the encoder for executing the abstract idea, and thus it does not integrate to a practical application, nor provide significantly more than the judicial exception (see MPEP 2106.05(h)).
Thus, claim 7 is non-patent eligible.
Regarding claim 8,
Step 1: The claim is directed to a method, which falls under the category of a process. The claim satisfies Step 1.
Step 2A Prong 1:
“The method according to claim 1, wherein in the step of generating the second data, when there is data related to a combination of products manually determined to be highly similar to each other, the similarity is determined to be high based on the second vector representations, and second data including the combination of products manually determined to be highly similar to each other is generated.” - The limitation is directed to manually determining combination of products to be highly similar to one another for the data. The limitation is directed to a process that can be performed in the human mind using evaluation, observation, and judgement, with aid of pen and paper, and thus the limitation is directed to a mental process.
There are no elements to be evaluated under Step 2A Prong 2 and Step 2B.
Thus, claim 8 is non-patent eligible.
Regarding claim 9,
Step 1: The claim is directed to a method, which falls under the category of a process. The claim satisfies Step 1.
Step 2A Prong 1:
“generating third vector representations based on product information of the new product in order to generate third data including new combinations of products highly similar to the new product based on similarity between the third vector representations and the first vector representations with using the first model; generating fourth vector representations from combinations of product information of two products included in the new combinations included in the third data and generating fourth data including combinations of highly similar products based on the fourth vector representations with using the second model; - The limitation is directed to generating vector representations based on information of new product orders and generating new combinations of the products based on similarity. The limitation is directed to the use of mathematical calculations/concept, and thus the limitation is directed to math.
Step 2A Prong 2 and Step 2B:
“The method according to claim 1, further comprising: receiving a new product put up for sale;” - The limitation recites a step to receive products that are put up for sale. The limitation is directed to an insignificant, extra-solution activity that cannot be integrated to a practical application (see MPEP 2106.05(g)). Furthermore, under Step 2B, the act of sending/receiving information and data over a network is a well-understood, routine, and conventional activity (WURC) and cannot provide significantly more than the judicial exception (see MPEP 2106.05(d)(II)).
“executing learning of the third model using the second encoder annotation data as training data.” - The limitation is directed to executing the third model using another encoder’s annotation data as the training data. The limitation amounts to no more than mere further limiting to e field of use/environment, and it does not integrate to a practical application, nor provide significantly more than the judicial exception (see MPEP 2106.05(h)).
Thus, claim 9 is non-patent eligible.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this
Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not
identically disclosed as set forth in section 102, if the differences between the claimed invention and the
prior art are such that the claimed invention as a whole would have been obvious before the effective filing
date of the claimed invention to a person having ordinary skill in the art to which the claimed invention
pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are
summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1, 3-5, and 8-11 is/are rejected under 35 U.S.C. 103 as being unpatentable over NPL reference “Towards Personalized and Semantic Retrieval: An End-to-End Solution for E-commerce Search via Embedding Learning,” by Zhang et. al. (referred herein as Zhang) in view of NPL reference “Augmented SBERT: Data Augmentation Method for Improving Bi-Encoders for Pairwise Sentence Scoring Tasks,” by Thakur et. al. (referred herein as Thakur) in view of NPL reference “ERNIE-Search: Bridging Cross-Encoder with Dual-Encoder via Self On-the-fly Distillation for Dense Passage Retrieval,” by Lu et. al. (referred herein as Lu) further in view of NPL reference “RocketQA: An Optimized Training Approach to Dense Passage Retrieval for Open-Domain Question Answering,” by Qu et. al. (referred herein as Qu).
Regarding claim 1, Zhang teaches:
A method for training a model for identifying a product in accordance with a predetermined search condition from among a plurality of products, performed by a learning device, the method comprising: ([Zhang, page 2409, sec 3] “Offline Model Training module trains a two tower model…for the uses in online serving and offline indexing…enable fast online embedding retrieval…transform any user input query text to query embedding, which is then fed to the item embedding index to retrieve K similar items.”, wherein the examiner interprets “training a two tower model for online serving and transforming a user’s input query to retrieve K similar items” to be the same as “training a model used for identifying a product corresponding to a predetermined search condition from among a plurality of products,” because they are both describing a trained model deployed in a search system that retrieves products matching a given user condition.)
wherein the first model is configured to receive product information comprising a product title for one product as input, and to output a vector corresponding to the input product information, ([Zhang, Figure 3, page 2409] showing “Item Title Tokens” as an input feature to the item tower, and [Zhang, page 2409-2410, sec 4.1] “[The item tower] concatenates all item features as input layer, then goes through multi-layer perceptron (MLP)…to output a single item embedding”, wherein the examiner interprets Zhang’s item tower receiving item title tokens (product title information) for a single item and outputting a single item embedding vector to be the same as “the first model being configured to receive product information comprising a product title for one product as input and to output a vector corresponding to the input product information,” because they are both describing a model that ingests a product title for a product and produces a corresponding vector representation (embedding) for that product.)
Zhang does not teach generating, using a first model comprising a bi-encoder, first vector representations based on product information data for the plurality of products, and generating first data indicating one or more combinations of products determined to be highly similar to each other based on similarities between the first vector representations; generating, using a second model comprising a cross-encoder, second vector representations from product information of two products included in the first data and generating second data indicating one or more combinations of products determined to be highly similar to each other based on the second vector representations, wherein the second model is configured to output a similarity value between products with a higher accuracy than the first model; training a third model to generate vector representations from the product information using the second data as training data, wherein the third model is configured to receive product information comprising a product title for one product as input, and to output a vector corresponding to the input product information; and wherein the first model, the second model, and the third model are different models, and wherein a number of product combinations of the plurality of products represented in the first data is less than a number of all combinations of the plurality of products.
Thakur teaches:
generating, using a first model comprising a bi-encoder, first vector representations based on product information data for the plurality of products, and generating first data indicating one or more combinations of products determined to be highly similar to each other from among the plurality of products based on similarities between the first vector representations, ([Thakur, page 1, sec 1] “bi-encoders such as Sentence BERT (SBERT)…encode each sentence independently and map them to a dense vector space.” AND [Thakur, page 4, sec 3.1] “We train a bi-encoder (SBERT) on the gold training set…and use it to sample further, similar sentence pairs. We use cosine-similarity and retrieve for every sentence the top k most similar sentences in our collection. For large collections, approximate nearest neighbour search like Faiss4 could be used to quickly retrieve the k most similar sentences.”, wherein the examiner interprets training an initial SBERT bi-encoder to independently encode each input into a dense vector and using cosine similarity to retrieve the top-k most similar pairs to be the same as using a first model comprising a bi-encoder to generate first vector representations and generating first data indicating combinations of highly similar products, because they are both describing an independent-encoding architecture that maps each item into a dense vector and uses vector similarity to select a subset of similar pairs from the full collection.)
generating, using a second model comprising a cross-encoder, second vector representations from product information of two products included in the first data and generating second data indicating one or more combinations of products determined to be highly similar to each other based on the second vector representations, ([Thakur, page 3, sec 3.1] “Given a pre-trained, well-performing cross-encoder, we sample sentence pairs according to a certain sampling strategy (discussed later) and label these using the cross-encoder. We call these weakly labeled examples the silver dataset and they will be merged with the gold training dataset.” and [Thakur, page 1, Abstract] “Cross-encoders, which perform full-attention over the input pair”, wherein the examiner interprets the cross-encoder performing full-attention over pairs drawn from the set identified by the bi-encoder and producing a labeled silver dataset to be the same as using a second model comprising a cross-encoder to generate second vector representations from two products included in the first data and generating second data indicating combinations of highly similar products, because they are both describing a cross-encoder that takes two items jointly as input from a previously identified candidate set and produces a labeled dataset of highly similar pairs for downstream training.)
wherein the second model is configured to output a similarity value between products with a higher accuracy than the first model; ([Thakur, page 1, Abstract] “While cross-encoders often achieve higher performance, they are too slow for many practical use cases.” and [Thakur, page 1, sec 1] “A drawback of the SBERT bi-encoder is usually a lower performance in comparison with the BERT cross-encoder.” AND [Thakur, page 2, sec. 2] “This drawback was addressed by SBERT (Reimers and Gurevych, 2019), which applies BERT independently on the inputs followed by mean pooling on the output to create fixed-sized sentence embeddings… Our proposed data augmentation approach is based on semi-supervision”, wherein the examiner interprets the cross-encoder achieving higher performance/accuracy than the bi-encoder, and “pooling on the output to create fixed-sized sentence embeddings” to be the same as the second model outputting a similarity value with higher accuracy than the first model, because they are both describing the relative superiority of cross-encoders over bi-encoders in the accuracy of similarity scoring between two items.)
and training a third model to generate vector representations from the product information using the second data as training data, wherein the third model is configured to receive product information comprising a product title for one product as input, and to output a vector corresponding to the input product information, ([Thakur, page 3, sec 3.1] “We then train the bi-encoder on this extended training dataset. We refer to this model as Augmented SBERT (AugSBERT).” and [Thakur, page 1, sec 1] “bi-encoders…encode each sentence independently and map them to a dense vector space.”, wherein the examiner interprets training the new AugSBERT bi-encoder on the cross-encoder-labeled silver dataset, where that bi-encoder independently encodes each single input into a dense vector, to be the same as training the third model to generate vector representations using the second data as training data and being configured to receive a product title for one product and output a corresponding vector, because they are both describing a model trained last in the pipeline on cross-encoder-generated data that independently maps a single input item to a vector representation.)
and wherein the first model, the second model, and the third model are different models, ([Thakur, page 3, sec 3.1] “Given a pre-trained, well-performing cross encoder, we sample sentence pairs according to a certain sampling strategy (discussed later) and label these using the cross-encoder. We call these weakly labeled examples the silver dataset and they will be merged with the gold training dataset. We then train the bi-encoder on this extended training dataset. We refer to this model as Augmented SBERT (AugSBERT). The process is illustrated in Figure 2..” and [Thakur, page 4, sec 3.1] “We train a bi-encoder (SBERT) on the gold training set as described in section 5 and use it to sample further, similar sentence pairs.”, wherein the examiner interprets Thakur's explicit use of three structurally and parametrically distinct models; (1) an initial SBERT bi-encoder used for semantic search sampling to identify candidate similar pairs, (2) a separate BERT cross-encoder used to label those pairs, and (3) a newly trained AugSBERT bi-encoder trained on the cross-encoder-labeled silver data; to be the same as the first, second, and third models being different models, because they are both describing a pipeline with three distinct models each serving a different role, where no two models are the same model being reused.)
and wherein a number of product combinations of the plurality of products represented in the first data is less than a number of all combinations of the plurality of products, ([Thakur, page 3, sec 3.1] “there are n × (n - 1)/2 possible combinations for n sentences. Weakly labeling all possible combinations would create an extreme computational overhead, and, as our experiments show, would likely not lead to a performance improvement.”, wherein the examiner interprets the teaching that it is impractical to label all possible pair combinations and that only a sampled subset is used to be the same as the number of product combinations represented in the first data being less than the number of all combinations of the plurality of products, because they are both recognizing that processing every possible pair from the full item set is impractical and that only a selected subset of combinations, retrieved by the bi-encoder, is carried forward.)
wherein the generating the second data further comprises determining, based on data related to a combination of products that are manually determined to be highly similar to each other according to a manual determination, whether the combination of products manually determined to be highly similar is also determined to be highly similar based on the second vector representations, ([Thakur, page 2-3, sec 3.1] “In our in-domain experiments, we re-use the sentences from the gold training set.” and [Thakur, page 3, sec 3.1] “Given a pre-trained, well-performing crossencoder, we sample sentence pairs according to a certain sampling strategy (discussed later) and label these using the cross-encoder. We call these weakly labeled examples the silver dataset and they will be merged with the gold training dataset.…weakly label a large set of randomly sampled pairs and then keep only certain pairs…we keep all the positive pairs.”, wherein the examiner interprets Thakur’s gold training set, which is a set of human-annotated (manually determined) highly-similar pairs being re-used and re-labeled by the cross-encoder, and only certain pairs being kept based on that cross-encoder labeling, to be the same as determining, based on manually-determined highly-similar combinations, whether those combinations are also determined highly similar based on the second vector representations, because they are both using the cross-encoder’s second-stage similarity output to evaluate pairs that were manually/human-annotated as highly similar.)
Zhang and Thakur do not teach wherein the second model is configured to receive product information comprising product titles for two products, and to output a vector corresponding to the input product information of the two products.
Lu teaches:
wherein the second model is configured to receive product information comprising product titles for two products, and to output a vector corresponding to the input product information of the two products, ([Lu, page 3, sec 3.1] “cross-encoder computes the relevance score sce(q, p), where the input is the concatenation of q and p with a special token [SEP]. Subsequently, the [CLS] representation of the output is fed into a linear function to compute the relevance score.”, wherein the examiner interprets Lu’s cross-encoder receiving a paired input (the concatenation of two inputs separated by a special token [SEP]) and producing a [CLS] representation corresponding to that paired input to be the same as the second model being configured to receive product information comprising product titles for two products and to output a vector corresponding to the input product information of the two products, because they are both describing a model that jointly ingests two inputs and produces a vector representation corresponding to the combined paired input.)
Zhang, Thakur, and Lu do not teach wherein the second data consists only of combinations of products that are determined to be highly similar based on the second vector representations and determined to be highly similar according to the manual determination.
Qu teaches:
wherein the second data consists only of combinations of products that are determined to be highly similar based on the second vector representations and determined to be highly similar according to the manual determination, ([Qu, page 5, sec 3.3] “we utilize [the cross-encoder] to annotate unlabeled questions for data augmentation…To ensure the quality of the automatically labeled data, we only select the predicted positive and negative passages with high confidence scores estimated by the cross-encoder.” and [Qu, page 7, sec 4.1.3] “Specifically, we select the top retrieved passages with a score higher than 0.9 as positive examples and those with a score less than 0.1 as negative examples.” and [Qu, page 6, STEP 4] “and then train a dual-encoder M(2)D on both the manually labeled training data DL and the automatically augmented training data DU.”, wherein the examiner interprets Qu’s retaining only the pairs that the cross-encoder confirms as highly similar with high confidence, together with the manually labeled training data D_L, to be the same as the second data consisting only of combinations that are determined to be highly similar based on the second vector representations and determined to be highly similar according to the manual determination, because they are both restricting the retained training data to only those pairs endorsed by both a manual/human label and the cross-encoder’s high-confidence similarity determination.)
Zhang, Thakur, Lu, Qu, and the instant application are analogous art because they are all directed to neural retrieval systems that train encoder-based models and compute similarity between vector representations of textual inputs.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the offline training module of Zhang to include the bi-encoder/cross-encoder augmentation pipeline disclosed by Thakur. One would be motivated to do so to effectively generate high-quality dense vector representations for large collections of textual items using a cross-encoder that produces labeled similarity data used to train an efficient bi-encoder for retrieval, as suggested by Thakur ([Thakur, page 1, sec 1] “bi-encoders such as Sentence BERT (SBERT) encode each sentence independently and map them to a dense vector space.”).
It would have further been obvious to include the paired-input cross-encoder of Lu. One would be motivated to do so to effectively compute more accurate similarity or relevance scores between two jointly-encoded inputs, as suggested by Lu ([Lu, page 3, sec 3.1] “cross-encoder computes the relevance score sce(q, p), where the input is the concatenation of q and p with a special token [SEP].”).
It would have further been obvious to include the cross-encoder-based label confirmation and denoising of Qu. One would be motivated to do so to ensure the quality of the training data by retaining only the pairs confirmed as highly similar by both the cross-encoder and the manual labels, thereby removing false labels and improving retrieval accuracy, as suggested by Qu ([Qu, page 5, sec 3.3] “To ensure the quality of the automatically labeled data, we only select the predicted positive and negative passages with high confidence scores estimated by the cross-encoder.”). Claims 10 and 11 are analogous to claim 1, aside from claim type and minute differences, thus the same rejection applies.
Regarding claim 3,
Zhang, Thakur, Lu, and Qu teach The method according to claim 1, (see rejection of claim 1).
Zhang further teaches wherein similarity between the first vector representations is computed based on distance between two of the first vector representations; ([Zhang, page 2409, sec 3] “we employ one of state-of-the-art algorithms [15] for efficient nearest-neighbor search of dense vectors.” and [Zhang, 2410, sec 4.1] “simple dot product interaction between query and item towers, the query and item embeddings are still theoretically in the same geometric space. Thus finding K nearest items for a given query embedding is equivalent to minimizing the loss for K query item pairs where the query is given…G(Q(q), S(s)) = ∑ wᵢ eᵢᵀ g”, wherein the examiner interprets nearest-neighbor search of dense vectors and finding K nearest items to be the same as computing similarity based on distance between two vector representations, because they are both procedures that compare embeddings in a shared space and identify the pairs with the smallest distance (i.e., highest similarity)).
Regarding claim 4,
Zhang, Thakur, Lu, and Qu teach The method according to claim 1, (see rejection of claim 1).
Zhang further teaches wherein in the step of generating the second data, similarity is computed based on a score computed based on the second vector representations. ([Zhang, page 2410, sec 4.3] “the soft dot product interaction between query and item can be defined as follows, G(Q(q), S(s)) = ∑ wᵢ eᵢᵀ g” and [Zhang, page 2410, sec 4.3] “This scoring function is basically a weighted sum of all inner products between m query embeddings and one item embedding.”, wherein the examiner interprets Zhang’s discussion of a scoring function formed by weighted inner products of the query and item embeddings to be the same as computing similarity based on a score derived from the second vector representations, because they are both describing how a similarity measure is produced by applying a mathematical function (dot-product weighting) to the vector outputs of the second model.)
Regarding claim 5,
Zhang, Thakur, Lu, and Qu teach The method according to claim 1, (see rejection of claim 1).
Thakur further teaches wherein the third model comprises a bi-encoder. ([Thakur, page 3, sec 3.1] “We then train the bi-encoder on this extended training dataset. We refer to this model as Augmented SBERT (AugSBERT).”, wherein the examiner interprets the newly trained AugSBERT model (which is a bi-encoder trained on the cross-encoder-labeled silver dataset) to be the same as the “third model comprising a bi-encoder,” because they are both describing a model that is trained last in the pipeline, using data generated by the cross-encoder as training data, and that independently encodes each single input into a dense vector representation, which is the defining characteristic of a bi-encoder)
Zhang, Thakur, Lu, Qu, and the instant application are analogous art because they are all directed to neural retrieval systems that train encoder-based models.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of claim 1 disclosed by Zhang, Thakur, Lu, and Qu to include the bi-encoder process disclosed by Thakur. One would be motivated to do so to effectively improve the quality of vector representations used for retrieval by retraining a bi-encoder model on additional labeled similarity data generated during the training pipeline, as suggested by Thakur ([Thakur, page 3, sec 3.1] “We then train the bi-encoder on this extended training dataset.”).
Regarding claim 8,
Zhang, Thakur, Lu, and Qu teach The method according to claim 1, (see rejection of claim 1).
Thakur further teaches wherein in the step of generating the second data, when there is data related to a combination of products manually determined to be highly similar to each other, the similarity is determined to be high based on the second vector representations, and second data including the combination of products manually determined to be highly similar to each other is generated. ([Thakur, page 3, sec 3.1] “Given a pre-trained, well-performing cross-encoder, we sample sentence pairs…and label these using the cross-encoder. We call these weakly labeled examples the silver dataset and they will be merged with the gold training dataset. We then train the bi-encoder on this extended training dataset.” and “we can re-use the sentences from the gold training set [human-annotated]”, wherein the examiner interprets “merged with the gold training dataset” (gold = human-labeled pairs already judged highly similar) and “weakly labeled…silver dataset” (pairs that the cross-encoder deems highly similar using second-stage vector representations) to be the same as “data related to a combination of products manually determined to be highly similar to each other” and “similarity is determined to be high based on the second vector representations…second data including the combination of products manually determined to be highly similar,” because they are both describing a process in which previously human-verified similar pairs (gold) are carried forward into a new dataset only after the model’s second-stage vectors (cross-encoder) confirm high similarity, thereby forming the updated second data.)
Zhang, Thakur, Lu, Qu, and the instant application are analogous art because they are all directed to enhancing the quality of product-pair training data.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method according to claim 1 disclosed by Zhang, Thakur, Lu, and Qu, to include the gold training dataset disclosed by Thakur. One would be motivated to do so to efficiently train the bi-encoder as suggested by Thakur ([Thakur, page 4, sec 3.1] “We train a bi-encoder (SBERT) on the gold training set as described in section 5 and use it to sample further, similar sentence pairs.”).
Regarding claim 9,
Zhang, Thakur, Lu, and Qu teach The method according to claim 1, (see rejection of claim 1).
Zhang further teaches further comprising receiving a new product put up for sale; generating third vector representations based on product information of the new product in order to generate third data including new combinations of products highly similar to the new product based on similarity between the third vector representations and the first vector representations with using the first model; ([Zhang, 2409, sec 3] “Offline Indexing module loads the item embedding model (i.e., the item tower) to compute all the item embeddings from the item”, and [Zhang, page 2409, sec 3] “transform any user input query text to query embedding…retrieve K similar items”, wherein the examiner interprets computing and embedding for a query item and retrieving K similar items by nearest-neighbor search to be the same as generating a vector for the new product with the first model and forming new combinations of highly similar products, because they are both embedding the new item and selecting its closest neighbors in the existing product-vector space.)
Thakur further teaches:
generating fourth vector representations from combinations of product information of two products included in the new combinations included in the third data and generating fourth data including combinations of highly similar products based on the fourth vector representations with using the second model; ([Thakur, page 3, sec 3.1] “we sample sentence pairs according to a certain sampling strategy (discussed later) and label these using the cross-encoder”, wherein the examiner interprets labeling these using the cross-encoder (which jointly encodes each item pair) to be the same as generating fourth vector representations with the second model to score the new pairs, because they are both re-embedding each candidate pair with a stronger cross-encoder to assess similarity.)
generating second encoder annotation data by annotating each of the combinations of highly similar products included in the fourth data to be positive; and executing learning of the third model using the second encoder annotation data as training data. ([Thakur, page 3, sec 3.1] “we call these weakly labeled examples the silver dataset and they will be merged with the gold training dataset. We then train the bi-encoder on this extended training dataset”, wherein the examiner interprets the silver dataset of cross-encoder-approved pairs to be the same as second-encoder annotation data marked positive, because they are both collections of pairs that the second model has confirmed as highly similar; the examiner further interprets “train the bi-encoder on this extended training dataset” to be the same as executing learning of the third model with the second-encoder annotation data, because they are both retraining the serving bi-encoder using the positives produced by the cross-encoder.)
Zhang, Thakur, Lu, Qu, and the instant application are analogous art because they are all directed to automated pipelines that ingest newly-arriving items.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method according to claim 1 disclosed by Zhang, Thakur, Lu, and Qu, to include the training data set fine-tuning process disclosed by Thakur. One would be motivated to do so to effectively improve the accuracy of the serving bi-encoder without costly manual labeling, as suggested by Thakur ([Thakur, page 1] “We use the cross-encoder to label new input pairs, which are added to the training set for the bi-encoder. The SBERT bi-encoder is then fine-tuned on this larger augmented training set, which yields a significant performance increase”).
Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Zhang in view of Thakur in view of Lu in view of Qu further in view of NPL reference “Billion-Scale Similarity Search with GPUs”, by Johnson et. al. (referred herein as Johnson).
Regarding claim 7, Zhang, Thakur, Lu, and Qu teaches The method according to claim 1, (see rejection of claim 1).
Zhang, Thakur, Lu, and Qu do not teach further comprising a step of using a dimensionality reduction encoder to compute vector representations lower in dimension than the vector representations.
Johnson teaches further comprising a step of using a dimensionality reduction encoder to compute vector representations lower in dimension than the vector representations. ([Johnson, page 535, sec 1] “several approaches employ compressed representations of the vectors using an encoding. This is especially convenient for memory-limited devices like GPUs. It turns out that accepting a minimal accuracy loss can result in orders of magnitude of compression” and [Johnson, page 535, sec 1] “the optimized product quantization or OPQ is a linear transformation on the input vectors that improves the accuracy of the product quantization; it can be applied as a pre-processing.”, wherein the examiner interprets “internal compressed representation…using an encoding” and “a linear transformation…applied as a pre-processing” to be the same as employing a dimensionality-reduction encoder that outputs lower-dimensional vectors, because they are both transforming higher-dimensional embeddings into more compact representations to reduce storage and accelerate subsequent similarity search.)
Zhang, Thakur, Lu, Qu, Johnson, and the instant application are analogous art because they are all directed to methods of product search that employ vector embeddings and similarity search.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method according to claim 1 disclosed by Zhang, Thakur, Lu and Qu to include the compressed representation technique disclosed by Johnson. One would be motivated to do so to efficiently increase the amount of compression achieved, as suggested by Johnson ([Johnson, page 535] “compressed representation of the vectors using an encoding…accepting a minimal accuracy loss results in orders of magnitude of compression”).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DEVAN KAPOOR whose telephone number is (703)756-1434. The examiner can normally be reached Monday - Friday: 9:00AM - 5:00 PM EST (times may vary).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Yi can be reached at (571) 270-7519. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DEVAN KAPOOR/Examiner, Art Unit 2126
/DAVID YI/Supervisory Patent Examiner, Art Unit 2126