Prosecution Insights
Last updated: October 01, 2026
Application No. 18/641,257

INFORMATION RETRIEVAL USING MULTIVARIATE DISTRIBUTIONS

Non-Final OA §101§103§112
Filed
Apr 19, 2024
Priority
Apr 19, 2023 — provisional 63/460,599
Examiner
PHUNG, QUOC LY PHU
Art Unit
Tech Center
Assignee
Google LLC
OA Round
1 (Non-Final)
45%
Grant Probability
Moderate
1-2
OA Rounds
1y 10m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 45% of resolved cases
45%
Career Allowance Rate
14 granted / 31 resolved
-14.8% vs TC avg
Strong +94% interview lift
Without
With
+94.4%
Interview Lift
resolved cases with interview
Typical timeline
4y 3m
Avg Prosecution
16 currently pending
Career history
47
Total Applications
across all art units

Statute-Specific Performance

§101
29.0%
-11.0% vs TC avg
§103
47.2%
+7.2% vs TC avg
§102
4.0%
-36.0% vs TC avg
§112
18.8%
-21.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 31 resolved cases

Office Action

§101 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims 1-20 are presented for examination. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. With respect to claim 1 [line 5], claim 19 [line 8] and claim 20 [line 7], it is confusing what the limitation “the parameters of the probability distributions over the space of multivariate representations for the query” refers to. In machine learning, multidimensional and multivariate are related but not exactly the same. Multidimensional refers to multiple input features and multivariate refers to multiple output variables. The limitation recites the space of multivariate representations which has never been defined in independent claims. For the purposes of examination, Examiner will interpret the limitation as either “the parameters of the probability distributions over the space of multidimensional representations for the query” OR “the parameters of the probability distributions over a space of multivariate representations for the query.” With respect to claims 2-18, they are rejected based on their virtual dependency on claim 1. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefore, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Independent claims Step 1 Claim 1 is drawn to a method, claim 19 is drawn to a system and claim 20 is drawn to a non-transitory computer-readable storage media storing instructions that when executed to perform the method of claim 1. Therefore, each of these claim groups falls under one of four categories of statutory subject matter (process/method, machines/product/apparatus, manufactures, and composition of matter). Step 2A – Prong 1 Claims 1 and 19 and 20 are directed to a judicially recognized exception of an abstract idea without significantly more. Claims 1 and 19 and 20 recite a method of processing the query using a query encoder neural network to generate parameters of a probability distribution over a space of multi-dimensional latent representations for the query that under its broadest reasonable interpretation enumerates a mathematical concept. A human can perform the calculation using words or using mathematical symbols to generate parameters of probability distribution in the space of latent representations. Therefore, the step of processing the query using query encoder neural network to generate parameters is nothing more than a mathematical concept (MPEP 2106.04(a)(2)(I)). Claims 1 and 19 and 20 recite further a method of identifying, using the parameters of the probability distribution over the space of multivariate representations for the query, a subset of a plurality of content items that under its broadest reasonable interpretation enumerates a mathematical concept. A human can perform the calculation using words or using mathematical symbols to identify a subset of content items using parameters of the probability distribution. Therefore, the step of identifying, using the parameters, a subset of a plurality of content items is nothing more than a mathematical concept (MPEP 2106.04(a)(2)(I)). Step 2A – Prong 2 Claims 1, 19 and 20 recite further a method of receiving a query that fails to integrate the abstract idea into a practical application. The step of receiving a query is a form of insignificant input and output solution activities, where receiving a query is necessary for all uses of the judicial exception. This additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (MPEP 2106.05(g)). Claims 1, 19 and 20 recite further a method of generating a response to the query that identifies at least one of the content items in the subset that fails to integrate the abstract idea into a practical application. The step of generating a response is a form of insignificant input and output solution activities, where generating a response to the query is necessary for all uses of the judicial exception. This additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (MPEP 2106.05(g)). Step 2B The additional elements in step 2A-Prong 2 those are forms of insignificant extra-solution activities, do not amount to significantly more than an abstract idea because the court decision have determined that these additional elements of receiving a query; and generating a response to the query to be well-understood, routine, and conventional when claimed in a merely generic manner (MPEP 2106.05(d)(II)). As such, claims 1, 19 and 20 are not patent eligible. Dependent claims Claims 2-18 merely narrow the previously recited abstract idea limitations. For the reasons described above with respect to claim 1, this judicial exception is not meaningfully integrated into a practical application, or significantly more than the abstract idea. The claims disclose similar limitations described for the independent claims above and do not provide anything more than the mental processes that are practically capable of being performed in the human mind with the assistance of pen. Therefore, claims 2-18 also recite abstract ideas that do not integrate into a practical application or amount to significantly more than the judicial exception, and are rejected under U.S.C. 101. Step 1 Claims 2-18 are drawn to a method. Therefore, this claim group falls under one of four categories of statutory subject matter (process/method, machines/product/apparatus, manufactures, and composition of matter). Step 2A – Prong 1 Dependent claim 4 recites further the mathematical process by processing a sequence that includes a first token, a second token, and a plurality of tokens representing the query using the query encoder neural network to generate a respective embedding of each of the tokens; processing the embedding of the first token using a first output neural network head to generate the respective means for each of the k dimensions; and processing the embedding of the second token using a second output neural network head to generate the respective variances for each of the k dimensions those are based on one or more features of the ML project (MPEP 2106.04(a)(2)(I)). Dependent claim 5 recites further the mathematical process by maintaining, for each of the plurality of content items, a respective content vector; generating, from the parameters of the probability distribution, a query vector for the query; and identifying the subset of the plurality of content items using the content vectors and the query vector those are based on one or more features of the ML project (MPEP 2106.04(a)(2)(I)). Dependent claim 6 recites further the mathematical process by searching the plurality of content items using a search technique that outputs a subset of content items that have content vectors that are most similar to the query vector according to a similarity measure that is based on one or more features of the ML project (MPEP 2106.04(a)(2)(I)). Dependent claim 9 recites further the mathematical process by processing the content item using a content item encoder neural network to generate parameters of a probability distribution over the space of multi-dimensional latent representations for the content item; and generating the content vector for the content item from the parameters of a probability distribution over the space of multi-dimensional latent representations for the content item those are based on one or more features of the ML project (MPEP 2106.04(a)(2)(I)). Dependent claim 10 recites further the mathematical process by processing a sequence that includes a first token, a second token, and a plurality of tokens representing the content item using the content item encoder neural network to generate a respective embedding of each of the tokens; processing the embedding of the first token using a first output neural network head to generate the respective means for each of the k dimensions; and processing the embedding of the second token using a second output neural network head to generate the respective variances for each of the k dimensions those are based on one or more features of the ML project (MPEP 2106.04(a)(2)(I)). Dependent claim 15 recites further the mathematical process by ranking the subset of content items according to the respective similarity scores for the content items in the subset that is based on one or more features of the ML project (MPEP 2106.04(a)(2)(I)). Dependent claim 16 recites further the mathematical process by generating a respective search result that identifies each of the one or more content items; and ordering the respective search results according to the respective similarity scores for the content items those are based on one or more features of the ML project (MPEP 2106.04(a)(2)(I)). Dependent claim 17 recites further the mathematical process by providing the response for presentation on the user device that is based on one or more features of the ML project (MPEP 2106.04(a)(2)(I)). Dependent claim 18 recites further the mathematical process by wherein maintaining, for each of the plurality of content items, a respective content vector comprises indexing the respective content vectors in an index database; and wherein searching the plurality of content items using the search technique comprises searching the indexed content vectors in the index database using the search technique those are based on one or more features of the ML project (MPEP 2106.04(a)(2)(I)). Step 2A – Prong 2 Dependent claim 2 recites further the insignificant extra solution activities by wherein each multi-dimensional latent representation in the space has a fixed number k of dimensions, wherein k is greater than one. This additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (MPEP 2106.05(g)). Dependent claim 3 recites further the insignificant extra solution activities by wherein the probability distribution over the space is a multi-variate normal distribution with a diagonal covariance matrix and wherein the parameters comprise a respective mean for each of the k dimensions and a respective variance for each of the k dimensions. This additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (MPEP 2106.05(g)). Dependent claim 7 recites further the insignificant extra solution activities by wherein the similarity measure is a dot product. This additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (MPEP 2106.05(g)). Dependent claim 8 recites further the insignificant extra solution activities by wherein the search technique is an approximate nearest neighbor technique. This additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (MPEP 2106.05(g)). Dependent claim 11 recites further the insignificant extra solution activities by wherein the query and content item encoder neural networks are the same neural network. This additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (MPEP 2106.05(g)). Dependent claim 12 recites further the insignificant extra solution activities by wherein, for each content item, the similarity measure between the query vector and the content vector for the content item approximates a negative KL divergence between (i) the probability distribution over the space of multi-dimensional latent representations defined by the parameters generated for the query and (ii) the probability distribution over the space of multi-dimensional latent representations defined by the parameters generated for the content item. This additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (MPEP 2106.05(g)). Dependent claim 13 recites further the insignificant extra solution activities by wherein the query encoder neural network is an encoder-only self-attention neural network. This additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (MPEP 2106.05(g)). Dependent claim 14 recites further the insignificant extra solution activities by wherein the query encoder neural network has been trained through distillation from a pre-trained teacher neural network. This additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (MPEP 2106.05(g)). As such, dependent claims 2-18 are not patent eligible. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Clement et al (US 20230042051 A1) hereafter Clement, and further in view of Grgicak et al (US 20230162044 A1) hereafter Grgicak. With respect to claim 1, Clement teaches a method performed by one or more computers (a method of a distillation system disclosed that extracts knowledge from a pre-trained sequence-to-sequence neural transformer model into a smaller bi-encoder. A pre-trained sequence-to-sequences neural transformer model is referred to as the teacher model and a bi-encoder referred to as the student model [par. 0011, 0016-0020 and FIG. 4]), the method comprising: receiving a query (bi-encoder consists of a query encoder to receive a query embedding of a target query [par. 0020-0022]); processing the query using a query encoder neural network to generate parameters of query (the bi-encoder consists of a query encoder and a source code encoder for each programming language where the query encoder and the source code encoders share a joint embedding space. The distillation system extracts knowledge with attention based on a conditional probability distribution into the bi-encoder [par. 0004, 0020-0022]); identifying, using the parameters of (content items are retrieved in response to the received queries. A retrieval system is able to perform a cross-domain search to find docstrings or natural language text that closely match a source code snippet. Bi-encoder embeds natural language queries and the corresponding source code snippet so they are mapped to a continuous vector space in which the queries are close to the corresponding source code snippet. Bi-encoder is used to search for natural language text associated with a source code snippet given the source code snippet as the query [par. 0006, 0020-0022, 0067-0071]); and generating a response to the query that identifies at least one of the content items in the subset (Approximate Nearest Neighbor search is used to perform the above search. The top-k search results are obtained and ranked in order of closest distance to the query embedding [par. 0068]). However, Clement does not particularly disclose probability distribution over a space of multi-dimensional latent representations; and probability distribution over the space of multivariate representations. In the same field of endeavor, Grgicak teaches probability distribution over a space of multi-dimensional latent representations (a clustering engine may utilize any cluster model or algorithm to group signal profile vectors. Any suitable algorithm for determining similarity or probability may be employed, such as similarity-based clustering models (centroid models, density models), distribution models (multivariate distribution models like Gaussian model), or any other model for clustering multidimensional vectors according to commonalities. Signal profiles are represented in vector form as multidimensional vectors in a multidimensional space. Mean and variance are statistics to describe the probability distribution of the log-normalized data [par. 0090-0102, 0132, 0154-0161, 0167-0169]); and probability distribution over the space of multivariate representations (any suitable algorithm for determining similarity or probability may be employed, such as similarity-based clustering models (centroid models, density models), distribution models (multivariate distribution models like Gaussian model), or any other model for clustering multidimensional vectors according to commonalities. Mean and variance are statistics to describe the probability distribution of the log-normalized data [par. 0090-0102, 0154-0161]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have incorporated the concept of generating cell vectors by concatenating allele vectors derived from signal profiles of each cell as suggested by Grgicak into the concept of extracting knowledge by a distillation system from a large sequence-to-sequence neural transformer model into a smaller bi-encoder as suggested by Clement because both of these systems addressing the process of employing a suitable neural network model for clustering vectors based on a certain probability distribution. Doing so would be desirable because the concept of Clement would be more efficient by employing a multidimensional model or a multivariate model for clustering vectors to determine the probability distribution in a multidimensional space to normalize the data and to visualize the data in a meaningful way (Grgicak, [par. 0092, 0098-0102, 0167-0171]). With respect to claim 2, the combination of Clement and Grgicak teaches wherein each multi-dimensional latent representation in the space has a fixed number k of dimensions, wherein k is greater than one (Grgicak, allele vectors may be concatenated together to form a high dimensional space. A dimensionality engine may output a data plot which represents multidimensional data in a low dimension space, such as two-dimensional space or three-dimensional space (number of dimensions is 2 or above) [par. 0090, 0159-161, 0169]). With respect to claim 3, the combination of Clement and Grgicak teaches wherein the probability distribution over the space is a multi-variate normal distribution with a diagonal covariance matrix and wherein the parameters comprise a respective mean for each of the k dimensions and a respective variance for each of the k dimensions (Grgicak, any suitable algorithm for determining similarity or probability may be employed, such as Gaussian distributions or multivariate normal distribution models. A distribution-based cluster model may model each component k by the Gaussian distribution, characterized by a mean vector, a covariance matrix and an associated probability. The Principle Component Analysis (PCA) may assume the mean and the variance are sufficient to describe the probability distribution of the log-normalized data [par. 0092, 0133, 0159-0161]). With respect to claim 4, the combination of Clement and Grgicak teaches wherein processing the query using a query encoder neural network to generate parameters of a probability distribution over a space of multi-dimensional latent representations for the query comprises: processing a sequence that includes a first token, a second token, and a plurality of tokens representing the query using the query encoder neural network to generate a respective embedding of each of the tokens (Clement, a first encoder jointly learns embeddings for sequences of tokens in the first domain and a second encoder learns embeddings for sequences of tokens in the second domain. The embeddings are mapped with encoder models, such as a LTSM that encodes the tokens representing a query and a source code snippet [par. 0005, 0021-0022]); processing the embedding of the first token using a first output neural network head to generate the respective means for each of the k dimensions (Clement, an encoder block consists of two layers. The first layer includes a masked multi-head attention component followed by a layer normalization component. The output of the layer normalization component is input into the encoder-decoder multi-head attention component with a residual connection to layer normalization component. The second layer includes an encoder-decoder multi-head attention component followed by a layer normalization component [par. 0044-0053]; and Grgicak, the mean and variance are statistics to describe the probability distribution of the log-normalized data [par. 0160]); and processing the embedding of the second token using a second output neural network head to generate the respective variances for each of the k dimensions (Clement, an encoder block consists of two layers. The first layer includes a masked multi-head attention component followed by a layer normalization component. The output of the layer normalization component is input into the encoder-decoder multi-head attention component with a residual connection to layer normalization component. The second layer includes an encoder-decoder multi-head attention component followed by a layer normalization component [par. 0044-0053]; and Grgicak, the mean and variance are statistics to describe the probability distribution of the log-normalized data [par. 0160]). With respect to claim 5, the combination of Clement and Grgicak teaches further comprising: maintaining, for each of the plurality of content items, a respective content vector (Clement, the embeddings are mapped with encoder models, such as a LTSM that encodes the tokens representing a query and a source code snippet and returns a corresponding vector embedding. The pre-trained sequence-to-sequence neural transformer model translate a natural language text into a source code snippet. A true translation pairs include query-source code snippet pairs [par. 0022, 0037]), and wherein identifying, using the parameters of the probability distribution over the space of multivariate representations for the query, a subset of a plurality of content items comprises: generating, from the parameters of the probability distribution, a query vector for the query (LTSM encodes the tokens representing a query ad a source code snippet and returns a corresponding vector embedding Eq (query) and Ec (code) [par. 0022]); and identifying the subset of the plurality of content items using the content vectors and the query vector (Clement, attention mechanisms gather information about relevant context of a given subtoken and then encode that context into a vector which represents the subtoken [par. 0045-0048]). With respect to claim 6, the combination of Clement and Grgicak teaches wherein identifying the subset of the plurality of content items using the content vectors and the query vector comprises: searching the plurality of content items using a search technique that outputs a subset of content items that have content vectors that are most similar to the query vector according to a similarity measure (Clement, bi-encoder consists of a query encoder and a source code encoder, and the bi-encoder is then used to search for source code snippets of a target programming language having a source code embedding that is close to the query embedding of a target query. The embeddings construct a measure of their similarity [par. 0020-0022]). With respect to claim 7, the combination of Clement and Grgicak teaches wherein the similarity measure is a dot product (Clement, a traditional bi-encoder code search process is the bi-encoder model of the log-likelihood of the similarity of the query and source code embeddings of a matched query-code pair. The attention function is scaled dot-product attention [par. 0025-0027, 0046, 0047]). With respect to claim 8, the combination of Clement and Grgicak teaches wherein the search technique is an approximate nearest neighbor technique (Clement, a natural language query, such as a docstring, is parsed by a tokenizer into a sequence of subtoken which are then transformed into a query embedding by a query encoder. An approximate nearest neighbor search, such as Approximate Nearest Neighbor is used to perform the search [par. 0068]). With respect to claim 9, the combination of Clement and Grgicak teaches further comprising: generating the respective content vectors for each of the content items, comprising, for each content item: processing the content item using a content item encoder neural network to generate parameters of a probability distribution over the space of multi-dimensional latent representations for the content item (Grgicak, signal profiles are represented in vector form as multidimensional vectors in a multidimensional space. Mean and variance are statistics to describe the probability distribution of the log-normalized data [par. 0090-0102, 0132, 0154-0161, 0167-0169]); and generating the content vector for the content item from the parameters of a probability distribution over the space of multi-dimensional latent representations for the content item (Clement, attention mechanisms gather information about relevant context of a given subtoken and then encode that context into a vector which represents the subtoken [par. 0045-0048]). With respect to claim 10, the combination of Clement and Grgicak teaches wherein each multi-dimensional latent representation in the space has a fixed number k of dimensions, wherein k is greater than one (Grgicak, allele vectors may be concatenated together to form a high dimensional space. A dimensionality engine may output a data plot which represents multidimensional data in a low dimension space, such as two-dimensional space or three-dimensional space (number of dimensions is 2 or above) [par. 0090, 0159-161, 0169]), wherein the probability distribution over the space is a multi-variate normal distribution with a diagonal covariance matrix, wherein the parameters comprise a respective mean for each of the k dimensions and a respective variance for each of the k dimensions (Grgicak, any suitable algorithm for determining similarity or probability may be employed, such as Gaussian distributions or multivariate normal distribution models. A distribution-based cluster model may model each component k by the Gaussian distribution, characterized by a mean vector, a covariance matrix and an associated probability. The Principle Component Analysis (PCA) may assume the mean and the variance are sufficient to describe the probability distribution of the log-normalized data [par. 0092, 0133, 0159-0161]), and wherein processing the content item using a content item encoder neural network to generate parameters of a probability distribution over the space of multi-dimensional latent representations for the content item comprises: processing a sequence that includes a first token, a second token, and a plurality of tokens representing the content item using the content item encoder neural network to generate a respective embedding of each of the tokens (Clement, a first encoder jointly learns embeddings for sequences of tokens in the first domain and a second encoder learns embeddings for sequences of tokens in the second domain. The embeddings are mapped with encoder models, such as a LTSM that encodes the tokens representing a query and a source code snippet [par. 0005, 0021-0022]); processing the embedding of the first token using a first output neural network head to generate the respective means for each of the k dimensions (Clement, an encoder block consists of two layers. The first layer includes a masked multi-head attention component followed by a layer normalization component. The output of the layer normalization component is input into the encoder-decoder multi-head attention component with a residual connection to layer normalization component. The second layer includes an encoder-decoder multi-head attention component followed by a layer normalization component [par. 0044-0053]; and Grgicak, the mean and variance are statistics to describe the probability distribution of the log-normalized data [par. 0160]); and processing the embedding of the second token using a second output neural network head to generate the respective variances for each of the k dimensions (Clement, an encoder block consists of two layers. The first layer includes a masked multi-head attention component followed by a layer normalization component. The output of the layer normalization component is input into the encoder-decoder multi-head attention component with a residual connection to layer normalization component. The second layer includes an encoder-decoder multi-head attention component followed by a layer normalization component [par. 0044-0053]; and Grgicak, the mean and variance are statistics to describe the probability distribution of the log-normalized data [par. 0160]). With respect to claim 11, the combination of Clement and Grgicak teaches wherein the query and content item encoder neural networks are the same neural network (Clement, the query encoder, Eq, and the source code encoder, Ec, may be the same model. The embeddings construct a measure of their similarity [par. 0022]). With respect to claim 12, the combination of Clement and Grgicak teaches wherein, for each content item, the similarity measure between the query vector and the content vector for the content item approximates a negative KL divergence between (i) the probability distribution over the space of multi-dimensional latent representations defined by the parameters generated for the query and (ii) the probability distribution over the space of multi-dimensional latent representations defined by the parameters generated for the content item (Clement, the natural way to measure the difference between the probability distributions between the teacher and the student model is the Kullback-Leibler (KL) divergence. Parameters θ of the student probability distribution q are modified which simplifies the KL divergence by omitting constant terms to a new learning objective as the average KL divergence per query. The distance computation component may utilize a cosine similarity function to compute the distance. The cosine similarity is the cosine of the angle between the vectors representing the query and source code snippet embeddings [par. 0028-0031, 0059]). With respect to claim 13, the combination of Clement and Grgicak teaches wherein the query encoder neural network is an encoder-only self-attention neural network (Clement, the teacher model and the pre-trained sequence-to-sequence model are neural transformer models. A neural transformer model is a type of deep learning model that utilizes an attention mechanism, wherein attention directs the neural network to focus on a subset of features or tokens. The attention mechanism provides the model with a better capability to learn the task at hand thereby generating more accurate predictions [par. 0004, 0040]). With respect to claim 14, the combination of Clement and Grgicak teaches wherein the query encoder neural network has been trained through distillation from a pre-trained teacher neural network (Clement, a teacher model is used to draw 2*samples of augmented data pairs using the true translation pairs. The teacher model is distilled into the bi-encoder through training of the true translation pairs and their associated data pairs. Pre-trained conditional-probability sequence-to-sequences neural transformer model is referred to as the teacher model and the bi-encoder is referred to as the student model [par. 0004, 0016-0018, 0021]). With respect to claim 15, the combination of Clement and Grgicak teaches wherein the similarity measure generates a respective similarity score for each of the content items in the subset (Clement, the query encoder, Eq, and the source code encoder, Ec, may be the same model. The embeddings construct a measure of their similarity [par. 0022]), and wherein generating a response to the query that identifies at least one of the content items in the subset comprises: ranking the subset of content items according to the respective similarity scores for the content items in the subset (Clement, an approximate nearest neighbor search, such as Approximate Nearest Neighbor is used to perform the search. The top-k search results are obtained and ranked in order of closest distance to the query embedding. The repository is indexed by the embedding table. The k closest-matching docstrings are ranked in order of the closest distance to the source code embedding [par. 0068-0070]). With respect to claim 16, the combination of Clement and Grgicak teaches wherein generating the response comprises: generating a respective search result that identifies each of the one or more content items (Clement, the top-k search results are obtained and ranked in order of closest distance to the query embedding [par. 0068]); and ordering the respective search results according to the respective similarity scores for the content items (Clement, the distance computation component may utilize a cosine similarity function to compute the distance. The cosine similarity is the cosine of the angle between the vectors representing the query and source code snippet embeddings. The search results are based on the similarity score between the query and the code embeddings [par. 0059-0060]). With respect to claim 17, the combination of Clement and Grgicak teaches wherein receiving a query comprises receiving the query from a user device (Clement, the client device interacts with the web service through the REST APIs. Receiving queries from the user devices [par. 0061, 0062, 0087, 0088]), and wherein the method further comprises: providing the response for presentation on the user device (Clement, the web service responds to the client device's request by transmitting a response including the ranked retrieved results [par. 0088]). With respect to claim 18, the combination of Clement and Grgicak teaches wherein maintaining, for each of the plurality of content items, a respective content vector comprises indexing the respective content vectors in an index database; and wherein searching the plurality of content items using the search technique comprises searching the indexed content vectors in the index database using the search technique (Clement, the source code repository is indexed by the source code embeddings produced by the bi-encoder which are stored in the embedding table. The repository is indexed by the embedding table. [par. 0066-0070]). With respect to claim 19, it is a system claim that is corresponding to the method of claim 1. Therefore, it is rejected for the same as claimed in claim 1 above. With respect to claim 20, it is a non-transitory computer-readable claim that is corresponding to the method of claim 1. Therefore, it is rejected for the same as claimed in claim 1 above. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Glass et al (US 12694236 B2) disclosed methods, systems, and computer program products for natural language data generation using automated knowledge distillation techniques are provided herein. A computer-implemented method includes retrieving, in response to an input query, a set of passages from at least one knowledge base by processing the input query using a first set of artificial intelligence techniques; ranking at least a portion of the set of passages by processing the set of passages using a second set of artificial intelligence techniques; generating at least one natural language answer, in response to the input query, by processing a subset of the set of passages in connection with automated knowledge distillation techniques based on the ranking of the at least a portion of the set of passages; and performing automated actions based on the ranking of the at least a portion of the set of passages and/or the at least one generated natural language answer. Auerbach et al (US 11763176 B1) disclosed techniques for improved searching and querying in computer-based reasoning systems are discussed and include receiving multiple new multidimensional data element to store in a computer-based reasoning data model; determining a feature bucket for each feature of each data element and storing a reference identifier in the feature bucket(s). A query on the computer-based reasoning system includes input data element (e.g., an actual data element, or a set of restrictions on features). Shotton et al (US 11710309 B2) disclosed a method to relocalize a mobile camera (such as on a smart phone) in a known environment or to compute the pose of an object moving relative to a fixed camera. The pose information is useful for robotics, augmented reality, navigation and other applications. In various embodiments where camera pose is calculated, a trained machine learning system associates image elements from an image of a scene, with points in the scene's 3D world coordinate frame. Hazard et al (US 10817750 B2) disclosed techniques for creating well-balanced computer-based reasoning systems and using those to control systems. The techniques include receiving a request to determine whether to use one or more particular data elements, features, cases, etc. in a computer-based reasoning model (e.g., as data elements, cases or features are being added, or as part of pruning existing features or cases). Conviction measures (such as targeted or untargeted conviction, contribution, surprisal, etc.) are determined and inclusivity conditions are tested. The result of comparing the conviction measure can be used to determine whether to include or exclude the feature, case, etc. in the computer-based reasoning model. A controllable system may then be controlled using the computer-based reasoning model. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Quoc Phung whose telephone number is (703) 756 1330. The examiner can normally be reached on Monday through Friday from 9am to 5pm PT. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) athttp://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Welch can be reached on 571-272-7212. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Q.L.P./Examiner, Art Unit 2143 /JENNIFER N WELCH/Supervisory Patent Examiner, Art Unit 2143
Read full office action

Prosecution Timeline

Apr 19, 2024
Application Filed
Sep 16, 2026
Non-Final Rejection mailed — §101, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743613
THERMODYNAMIC NEURAL NETWORK
5y 3m to grant Granted Sep 22, 2026
Patent 12737583
METHOD AND APPARATUS FOR CONSTRUCTING NETWORK STRUCTURE OPTIMIZER, AND COMPUTER-READABLE STORAGE MEDIUM
4y 11m to grant Granted Sep 15, 2026
Patent 12725055
CURATED MACHINE LEARNING WORKFLOW SUGGESTIONS AND CLUSTERING TECHNIQUES
5y 6m to grant Granted Sep 01, 2026
Patent 12725053
SYSTEMS AND METHODS OF OPTIMIZING RESOURCE ALLOCATION USING MACHINE LEARNING AND PREDICTIVE CONTROL
3y 10m to grant Granted Sep 01, 2026
Patent 12718115
SYSTEM AND METHOD FOR CALCULATING GENERALIZED UTILITIES AND CHOICE PREDICTIONS
4y 3m to grant Granted Aug 25, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
45%
Grant Probability
99%
With Interview (+94.4%)
4y 3m (~1y 10m remaining)
Median Time to Grant
Low
PTA Risk
Based on 31 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month