DETAILED ACTION
Notice of Pre-AIA or AIA Status
1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
2. This action is responsive to the Applicant’s amendment filed on January 05, 2026.
3. Claims 21-22, 25-31, and 34-40 are pending, of which claims 21, 30, and 39 are in independent form.
4. Claims 21, 30, and 40 are amended.5. Claims 1-20, 22-23, and 32-33 are cancelled by the applicant.
Terminal Disclaimer
6. The terminal disclaimer filed on January 05, 2026 disclaiming the terminal portion of any patent granted on this application which would extend beyond the expiration date of US Patent 12,135,701 B2 has been reviewed and is accepted.
Response to Arguments
7. Applicant's arguments, see “A. Non-Statutory Double Patenting Rejection”, filed on January 05, 2026 have been carefully considered. Based on the filing of Terminal Disclaimer, the non-statutory obviousness-type double patenting rejection has been withdrawn.8. Applicant's arguments, see “B. Rejection Under 35 U.S.C. §101”, filed on January 05, 2026 have been carefully considered and are not persuasive.9. Applicant recites MPEP §2106 and notes that Step 2A has two progs (judicial exception; practical application). Applicant then cites MPEP §2106.04(a) for the rule that if claim limitations do not fall within the abstract-idea groupings, the analysis ends at Prong One.
Response: Examiner has carefully considered the argument but respectfully
disagrees. This argument is not persuasive. Examiner agrees that MPEP §2106 and Step 2A control the eligibility analysis. However, as explained below, the limitations of claims 21, 30, and 39 do fall within the “mathematical concepts” grouping identified in MPEP §2106.04(a)(1), so the analysis does not end at Prong One. Instead, the claims recite specific numerical and vector-space operation (masking elements via dropout ratios, mapping to high-dimensional vectors, forming a correlation matrix, and applying a decorrelation loss) that constitute mathematical relationships and calculations, rather than merely “involving” an abstract idea at a high level.
10. Applicant contents that claims 21, 30, and 39 do not recite any of the abstract idea groupings (mathematical concepts, methods or organizing human activity, mental processes), and thus do not recite an abstract idea at all.
Response: Examiner has carefully considered the argument but respectfully
disagrees. This argument is not persuasive. Examiner maintains that the claims recite mathematical concepts under the USPTO’s abstract idea groupings. The core limitations of the independent claims 21, 30, and 39 require: - Generating first and second vector representation of sentences (purely numerical representation). - Applying dropout ratios that “modify, in a vector domain, the sentences by masking one or more … numerical elements” of the vector representations. - Mapping these vectors into high-dimensional vector representations. - Generating a correlation matrix from the high-dimensional vectors. - Applying a decorrelation operation based on a loss objective with explicit terms (augmentation invariance, feature redundancy management, hyper-parameter weighting).
Each of these is a mathematical operation on numerical data—specifically, linear algebra and statistical correlation/decorrelation in vector space. Under MPEP §2106.04(a)(1), such operations are squarely within the “mathematical concepts” category (e.g., mathematical relationships, formulas, and calculations), even if not expressed with symbols like “x2 + y2”. The claims therefore recite a judicial exception at Step 2A, Prong One.
11. Applicant cites the August 4, 2025 memo for the proposition that examiner must distinguish between claims that recite an exception (requiring further analysis) and those that merely involve an exception (which are eligible). Applicant argues the present claims only “involve” an abstract idea.
Response: Examiner has carefully considered the argument but respectfully
disagrees. This argument is not persuasive. Examiner considered the August 4, 2025 memo. In this case, the claims recite the mathematical operations themselves, rather than merely involving them implicitly. The claims do not just say “train a machine learning model” in an open-ended way. They expressly require: - Using specific dropout ratios to generate different masked versions of the same sentence vector. - Computing a correlation matrix between the high-dimensional representations. - Applying decorrelation operation using a particular loss objection decorrelation function comprising an “augmentation invariance term”, a “feature redundancy management term”, and a “hyper-parameter”, and combining these terms (multiplying the hyper-parameter by the redundancy term and summing with the augmentation term).
Those elements define the mathematical relationships and structure of the training objective. Because the claims explicitly recite these mathematical constructs, they do more than merely “involve” an abstract idea—they “recite” the abstract mathematical concept that performs the heart of the claimed method. That places them within the judicial—exception category under the August memo and MPEP §2106.04(a).
12. Applicant analogized the claims to Example 39, where the limitation “training the neural network in a first stage using the first training set” was held not recite a judicial exception because it did not set forth any specific mathematical relationships, formula, or equations. Applicant contrasts this with Example 47, where claim 2 recites backpropagation and gradient descent by name and thus recites a mathematical concept. Applicant argues that the present claims, like Example 39, do not specify particular mathematical calculations by name, and therefore do not recite a judicial exception.
Response: Examiner has carefully considered the argument but respectfully
disagrees. This argument is not persuasive. Examiner disagrees with Applicant’s analogy to Example 39. The claims here go well beyond a generic “train a model” limitation: - They fix the structure of the training process by requiring different dropout ratios applied to the same sentence in vector space, which produces two specific correlated views. - They define the training objective mathematically via the correlation matrix and the loss objective decorrelation function, including particular terms (augmentation invariance term, features redundancy management term) and a hyper-parameter that weights those terms.
In Example 39, the claim simply instructs “training the neural network” without specifying any particular training algorithm or internal mathematical structure; and number of training algorithms or loss functions could satisfy the limitation. Here, by contrast, the claims constrain how the training is performed via explicit mathematical constructs correlation matrix, decorrelation objective with specified component and weighting). That is analogous to Example 47, where the references to backpropagation and gradient descent was found to recite a specific mathematical algorithm. The fact that the present claims do not literally use the words “backpropagation” or “gradient descent” is not dispositive. What matters under the USPTO guidance is whether the claims set forth specific mathematical relationships or calculations. By requiring a particular correlation-based loss structure with named terms and a hyper-parameter relationship, the claims do exactly that. They, thus recite a mathematical concept, not merely a high-level “training” step.
13. Applicant argues that none of the limitations “set forth or describe any specific mathematical relationships, formulas, or equations using words or mathematical symbols”, so they cannot be mathematical concept.
Response: Examiner has carefully considered the argument but respectfully
disagrees. This argument is not persuasive. Examiner does not interpret MPEP §2106.04(a)(1) or the August memo as requiring explicit algebraic notation or named equation to find a mathematical concept. A claim can “set forth” a mathematical relationship in words by defining the structure of a loss function or algorithm without writing the exact equation. Here, the claims explicitly define: - A correlation matrix based on two high-dimensional vectors. - A loss objective decorrelation function comprising (i) an augmentation invariance term, (ii) a feature redundancy management term, and (iii) a hyper-parameter that multiples the redundancy term, with the result then summed with the augmentation term. That wording describes the mathematical relationship between these terms—multiplication by the hyper-parameter and summation with the other term—even if the claim does not user symbols like L = L + αLred.. Under the USPTO guidance, that is sufficient to constitute a mathematical relationship and thus a mathematical concept.
Given the above, Examiner maintains: - Step 2A, Prong One: The claims recite mathematical concept (vector-space masking under specified dropout ratios, mapping into high-dimensional spaces, correlation matrices, and a decorrelation loss function with defined terms and weighting).
- Step 2A, Prong Two: The additional computer recitations (generic processors, memory, non-transitory medium, query input/result output) do not integrate the exception into a practical application because they merely invokes generic computing to carry out the mathematical operations and display/use the results.
- Step 2B: There is no recited improvement to the functioning of the computer itself, nor any unconventional hardware or arrangement; the claims use known computing components to implement an abstract training objective. Thus, they do not amount to “significantly more” than the abstract idea.
Accordingly, Applicant’s argument under the August 4, 2025 memo and Examples 39/47 are not persuasive, and the rejection under 35 U.S.C. §101 is maintained.
14. Applicant’s arguments, see “C. Rejections under 35 U.S.C. §103”, filed on January 05, 2026 have been carefully considered and are not persuasive.
14. A) Applicant argue that “Independent claims 21, 30, and 39 are herein amended to incorporate the matter of claims 23, 24, 32, and 33 (now canceled). The Office Action acknowledges that WU, Zhiyanov, and Kalantidis fail to teach the subject matter of claims 23, 24, 32, and 33, and instead relies on Klein … to cure deficiencies”.
Response: Examiner has carefully considered the argument but respectfully
disagrees. This argument is not persuasive. Examiner acknowledges that independent claims 21, 30, and 39 have been amended to incorporate subject matter previously recited in claims 23, 24, 32, and 33. To the extent the prior Office Action relied on Klein to satisfy those particular limitations (e.g., different dropout ratios with masking, correlation-matrix-based decorrelation objectives), the current rejection has been reconsidered in view of the amendment and applicant’s grace-period argument. As explained below, Klein is not relies upon as explained below, Klein is not relied upon as prior art unless and until applicant fails to establish that the §102(b)(1)(A) exception applies.
14. B) Applicant argue that “Klein cannot cure these deficiencies as Klein does not qualify as prior art … Klein is deemed a ‘grace period inventor disclosure’ … and the publication is not prior art under 35 U.S.C. §102(a)(1), pursuant to MPEP §2153.01(a). Nor is the publication available as prior art under 35 U.S.C. §103”.
Response: Examiner has carefully considered the argument but respectfully
disagrees. This argument is not persuasive. Applicant asserts that Klein is an inventor disclosure within one year of the effective filing data and therefore qualifies for the §102(b)(1)(A) grace-period exception. The examiner has considered applicant’s argument. However, mere attorney argument is insufficient to establish that a reference qualifies for the §102(b)(1)(A) exception. The record must contain evidence that the disclosure was made by the inventor or by another who obtained the subject matter directly or indirectly from the inventor, such as a statement under 37 CFR 1.130(a) or other appropriate evidence.
Accordingly, unless and until applicant submits a timely and proper statement under 37 CFR 1.130 or equivalent evidence establishing that Klein is an inventor disclosure under §102(b)(1)(A), Klein remains available as prior art under §102(a)(1) and may be relied upon in an obviousness rejection under §103.
14. C) Applicant argue that “Therefore, because Wu, Zhiyanov, Kalantidis fail to disclose and would not have rendered obvious the subject matter of claims 21, 30, and 39 (as amended), claims 21, 30, and 39 are patentable over Wu, Zhiyanov, and Kalantidis”.
Response: Examiner has carefully considered the argument but respectfully
disagrees. This conclusion is not persuasive at this time. As discussed in the prior Office Action, Wu discloses training a contrastive learning framework for sentence representations with two augmented views of each sentence and corresponding vector embeddings. Zhiyanov teaches multi-stage deep neural network processing that maps input text descriptors to high-dimensional embeddings using a redundancy-reduction loss term in conjunction with similarity-preservation/augmentation term. Taken together, Wu, Zhiyanov, and Kalantidis provide explicit teachings and motivations to (i) generate multiple sentence embeddings via different augmentation/encoding conditions, (ii) map those embeddings into high-dimensional spaces, and (iii) decorrelate the dimensions via correlation-matrix-based or redundancy-reduction loss objectives.
14. D) Applicant argue that “claim 22 depends from claim 21; claim 31 depends from claim 30; and claim 40 depends from claim 39. Therefore, claims 22, 31, and 40 are also patentable … based at least on their dependencies … “.
Response: Examiner has carefully considered the argument but respectfully
disagrees. This argument is not persuasive. The patentability of dependent claims 22, 31, and 40 under §103 depends on the patentability of their respective independent claims and the additional subject matter they recite. To the extent the §103 rejection is maintained for independent claims 21, 30, and 39, the dependent claims likewise remain rejected under §103, subject to any additional limitations being separately considered. If the §103 rejection of the independent claims is withdrawn, the examiner will reconsider the dependent claims accordingly.
14. E) Applicant argue that “The rejection of claims 23, 24, 32, and 33 is rendered moot in view of cancelation of those claims”.
Response: The examiner agrees that the rejection of canceled claims 23, 24, 32, and 33 is moot. No further action is taken with respect to those claims.
14. F) Applicant argue that “As explained above, Klein does not qualify as prior art to the present application”.
Response: Applicant’s position that Klein is not prior art has been considered. As noted above, applicant must provide evidence, such as a 37 CFR 1.130(a) statement, to establish that Klein qualifies for the §102(b)(1)(A) grace-period exception. Absent such evidence, Klein remains available as prior art under §102(a)(1) and may be used in an obviousness rejection. If applicant provides sufficient evidence, the examiner will reconsider and, if appropriate, withdraw the Klein-based rejection.
14. G) Applicant argue that “Furthermore, claims 25-29 depend from claim 21 and 34-38 depend from claim 30. Claims 21 and 30 are patentable … Therefore, claims 25-29 and 34-38 are also patentable … based at least on their dependencies”.
Response: Applicant’s argument that dependent claims 25-29 and 34-38 are patentable solely based on the asserted patentability of independent claims 21 and 30 is not persuasive. The patentability analysis for each dependent claim must be consider both (i) whether the base claim is patentable, and (ii) whether the additional limitations recited in the dependent claims are themselves taught or suggested by the prior art. To the extent the §103 rejection is maintained for the independent claims and/or Klein remains usable as prior art, the Examiner continues to rely on the combination including Klein to teach the specific additional limitations in claims 25-29 and 34-38, such as:
First and second dropout ratios with different masking behavior;
Relative magnitude of the dropout ratios (second larger than first);
Correlation matrix constructed from two high-dimensional views;
Loss objective decorrelation function with augmentation invariance terms, feature redundancy management term, and hyper-parameter;
Specific combination of terms in the loss (hyper-parameter weighting and summation).
If Klein is ultimately removed as prior art and no other reference is cited for those added limitations, the §103 rejection of claims 25-29 and 34-38 would need to be withdrawn or replaced with a new art-based rejection.
15. The rejection under 35 U.S.C. §101 from the previous Office Action is maintained in the instant Office Action.
Claim Rejections - 35 USC § 103
16. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
17. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
18. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
19. Claims 21-22, 25-31, and 34-40 are rejected under 35 U.S.C. 103 as being unpatentable over Wu et al. “CLEAR: Contrastive Learning for Sentence Representation” (hereinafter Wu) in view of Zhiyanov US Patent 10,565,498 B1 (hereinafter Zhiyanov), further in view of Kalantidis et al. US 2023/0106141 A1 (hereinafter Kalantidis), further in view of Klein et al. “SCD: Self-Contrastive Decorrelation for Sentence Embeddings” (hereinafter Klein).
Regarding claim 21, Wu discloses a computer-implemented method comprising: training a machine learning model, the training comprising (Wu [Section: Introduction] e.g., “…we propose a new framework, CLEAR, combining word-level MLM objective with sentence-level CL objective to pre-train a language model”. CLEAR is explicitly about training a machine learning model): performing a first encoding operation on a sentence including a plurality of text, the first encoding operation comprising generating a first vector representation of the sentence, the first vector representation including first numerical elements representing the plurality of text of the sentence (Wu [Section – 2.1 Sentence Representation] e.g., “Later on, many pre-trained language models such as BERT (Devlin et al., 2019) propose to use the manually-inserted token (the [CLS] token) as the representation of the whole sentence”. The [CLS] embedding is a numerical vector representing the entire sentence), wherein the first vector representation is generated using a first dropout ratio that modifies, in a vector domain (Wu [Section – Introduction] e.g., “Later on, pre-trained models such as BERT (Devlin et al. , 2019) propose to insert a special token (i.e., [CLS] token) during the pre-training and take its embedding as the representation for the sentence.”. This shows Wu teaches that the embedding (vector) of the [CLS] token is used as sentence representation. See also [Section – 3.1 The Contrastive Learning Framework] e.g., “For each original sentence s, we generate two random augmentations se1 = … A transformer-based encoder f(·) that learns the representation of the input augmented sentences H1 = f(se1) and H2 = f(se2)”. Here H1 and H2 are vector outputs (“representations” of the sentence (augmented) embedding, meaning they ae numerical vectors). See also [Section – 3.1 The Contrastive Learning Framework] e.g., “), [the sentence by masking one or more of the first numerical elements of the first vector representation;] Wu does not explicitly teach performing a second encoding operation on the sentence, the second encoding operation comprising generating a second vector representation of the sentence, the second vector representation including second numerical elements representing the plurality of text of the sentence, wherein the second vector representation is generated using a second dropout ratio that modifies, in the vector domain, [the sentence by masking one or more of the second numerical elements of the second vector representation; mapping the first vector representation to a first high dimensional vector representation, and further mapping the second vector representation to a second high dimensional vector representation; and decorrelating the first high dimensional vector representation and the second high dimensional vector representation; wherein the trained machine learning model comprises the decorrelated first high dimensional vector representation and the decorrelated second high dimensional vector representation.
However, Zhiyanov discloses performing a second encoding operation on the sentence, the second encoding operation comprising generating a second vector representation of the sentence, [the second vector representation including second numerical elements representing the plurality of text of the sentence], wherein the second vector representation is generated using a second dropout ratio that modifies, in the vector domain, the sentence (Zhiyanov [col…] e.g., “…a relationship analysis system may use the neural network model to generate, for a given pair of entities for which at least some text information is available, relationship indicators of any of several types in different embodiments. For example, according to some embodiments, the relationship analysis system may generate, given a pair of item descriptors containing values for various attributes of a corresponding pair of items, a similarity score which represents how similar the two items of the pair are to each other…”. This is effectively a second encoding operation producing a second vector representation, analogous to the limitation’s “second vector representation the sentence/entity” or The section explicitly mentions generating vectors/representations for each entity descriptor to produce relationship indicator (similarity, inclusion, participation);
mapping the first vector representation to a first high dimensional vector representation (Zhiyanov [col. Lines..] e.g., “The raw text of the attributes may be processed and converted into a set of intermediate vectors by a token model layer”, see also [Zhiyanov [col. Line..] e.g., “ each token (such as the word “companya”, a normalized version of the word “CompanyA” of the original or raw attribute) being represented by a selected number of real values labeled X0, X1, and so on”, see also [col. Line…] e.g., “Individual ones of the attribute model subnetworks… may comprise a plurality of nodes, including for example LSTM units… Corresponding to a given attribute's value, an attribute model output vector (AMOV) may be generated in the depicted embodiment”, see also [col. Lines..] e.g., “In at least some embodiments, the AMOVs may be combined (e.g., by concatenation) and provided as input to a first dense or dully-connected layer 250A of the deep neural network 202.. The output of the first dense layer 250A may comprise another intermediate values vector 255 in the depicted embodiment…”), and further mapping the second vector representation to a second high dimensional vector representation (Zhiyanov [col. ..] e.g., “The output of the first dense layer 250A may comprise another intermediate values vector 255 in the depicted embodiment, which may in turn comprise the input to a second dense layer 250B with associated weight matrix 260B. The output of the second dense layer 250B may comprise the similarity score 270…”, see also [col. …] e.g., “During training of the deep neural network, a common objective function may be used… a respective relationship indicator of the desired type (such as a similarity score) for various pairs of entity descriptors”. This chain paragraphs shows toke [Wingdings font/0xE0] AMOV [Wingdings font/0xE0] first dense layer [Wingdings font/0xE0] second dense layer [Wingdings font/0xE0] final vector, which directly maps to the limitations’ s functional steps). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include the functionality deep neural network-based relationship analysis with multi-feature token model taught and disclosed by Zhiyanov, in the contrastive learning for sentence representation taught by Wu to yield the predictable results of deep neural network models that utilize an extensible multi-feature token model to generate relationship indicators with respect to various kinds of entity pairs may be extremely useful in a number of scenarios (Zhiyanov [col. 22, lines 22 – col. 23, line 3]). The proposed combination of Wu and Zhiyanov does not explicitly teach decorrelating the first high dimensional vector representation and the second high dimensional vector representation; and wherein a trained machine learning model comprises the decorrelated first high dimensional vector representation and the decorrelated second high dimensional vector representation.
However, Kalantidis discloses decorrelating the first high dimensional vector representation and the second high dimensional vector representation (Kalantidis [0084] e.g., “The loss is composed of two terms… The second term pushes off-diagonal elements towards zero, reducing the redundancy between output dimensions…”. Uses a cross-correlation matrix to decorrelate embeddings by minimizing redundancy); and
wherein the trained machine learning model comprises the decorrelated first high dimensional vector representation and the decorrelated second high dimensional vector representation (Kalantidis [0089] e.g., “…providing more dimensions to decorrelate leads to a bottleneck representation that is more informative, allowing the network to learn an encoder that also has more decorrelated outputs”, see also [0128] e.g., “…encoding the input vector using the trained dimensionality reduction model to generate an encoded output vector in the d-dimensional space; and outputting the encoded output vector.”. The trained model outputs decorrelated representation). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include the functionality of dimensionality reduction model and method for training same, taught and disclosed by Kalantidis, in the system taught by the proposed combination of Wu and Zhiyanov to yield the predictable results of training a dimensionality reduction model to minimize a total loss based on the computed similarity preservation loss and the computed redundancy reduction loss. The computational cost or training time for any uses and applications that follow is reduced. The training or task performance in some cases is improved (Kalantidis [0109]).
The combined teaching of Wu, Zhiyanov, and Kalantidis does not explicitly discloses: the sentence by masking one or more of the first numerical elements of the first vector representation; the sentence by masking one or more of the second numerical elements of the second vector representation. Klein discloses the sentence by masking one or more of the first numerical elements of the first vector representation (Klein et al. [Section 2 Method] e.g., “…one augmentation is generated with high dropout and one with low dropout. This entails employing different random masks during the encoding phase. The random masks are associated with different ratios, rA and rB, with rA < rB. Integrating the distinct dropout rates into the encoder, we yield h A i = fθ(xi , rA) and h B i = fθ(xi , rB)”. “Employing different random masks during the encoding phase” is exactly the masking of vector elements hA corresponds to the first vector representation, rA to the first dropout ration); and the second vector representation including a second numerical elements representing the plurality of text of the sentence (Klein [Section 2 Method] e.g., “…Integrating the distinct dropout rates into the encoder, we yield h A i = fθ(xi , rA) and h B i = fθ(xi , rB)”. hBis the second vector representation, rB is the second dropout ration which uses a random mask to modify the embedding); and the sentence by masking one or more of the second numerical elements of the second vector representation (Klein et al. [Section 2 Method] e.g., “…one augmentation is generated with high dropout and one with low dropout. This entails employing different random masks during the encoding phase. The random masks are associated with different ratios, rA and rB, with rA < rB. Integrating the distinct dropout rates into the encoder, we yield h A i = fθ(xi , rA) and h B i = fθ(xi , rB)”. “Employing different random masks during the encoding phase” is exactly the masking of vector elements hA corresponds to the first vector representation, rA to the first dropout ration). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include the functionality of scd: self-contrastive decorrelation for sentence embeddings taught and disclosed by Klein, in the proposed combination of Wu, Zhiyanov, and Kalantidis to yield the predictable results in a system performing contrastive learning of sentence representation (Klein [Section 1 Introduction]).
Claims 30 and 39 incorporate substantially all the limitations of claim 21 in a system form (see Wu, [Section 4 Setup] e.g., “All of the models are pre-trained on 256 NVIDIA Tesla V100 32GB GPUs”. This implies a computing system with: Processor: The GPUs themselves serve as parallel processing units for model training along with CPUs that manage GPU execution Memory: Each GPU has 32 GB of VRAM), and a non-transitory computer-readable storage medium (see Wu, [Section 4.1 Setup] e.g., “Pre-training data: We pre-train all the models on a combination of BookCorpus (Zhu et al., 2015) and English Wikipedia datasets, the data BERT used for pre-training”. The dataset must reside in non-transitory memory (disk or SSD) to be read during training) and are rejected under the rationale.
Regarding claim 22, the proposed combination of Wu, Zhiyanov, Kalantidis, and Klein teaches a computer-implemented method of claim 21, further comprising: receiving, by the trained machine learning model, a query that includes a target sentence (Kalantidis [0034] e.g., “The dimensionality reduction model in the system 100 is embodied in or a component of an encoder 102, which is configured to receive data such as can be represented by an input vector, in an input space that is a higher dimensional representation space…”); and outputting, using at least in part the trained machine learning model, a result sentence that satisfies a similarity metric relative to the target sentence (Kalantidis [0034] e.g., “…and generate an output vector in an output space that is a lower-dimensional representation space, e.g., a d-dimensional space, where D is greater than d. The input vectors can be, for instance, provided by one or more datapoints 104a, 104b, 104c, 104d,”. This directly supports a target sentence query and outputting a result sentence that satisfies a similarity metric).
23.(Canceled)
24.(Canceled)
Regarding claim 25, he proposed combination of Wu, Zhiyanov, Kalantidis, and Klein teaches a computer-implemented method of claim 21, wherein the second dropout ratio is larger than the first dropout ratio (Klein [Section 2 Method] e.g., “…The random masks are associated with different ratios, rA and rB, with rA < rB”).
Regarding claim 26, the proposed combination of Wu, Zhiyanov, Kalantidis, and Klein teaches a computer-implemented method of claim 21, further comprising: generating a correlation matrix using the first high dimensional vector representation and the second high dimensional vector representation, wherein the decorrelating is performed on the correlation matrix (Klein [Section 2 Method] e.g., “A projector maps embeddings to a high-dimensional feature space, where the features are decorrelated”, see also [Figure 1]).
Regarding claim 27, the proposed combination of Wu, Zhiyanov, Kalantidis, and Klein teaches a computer-implemented method of claim 26, wherein the decorrelating on the correlation matrix is based at least in part on a loss objective decorrelation function, wherein the loss objective decorrelation function comprises an augmentation invariance term, feature redundancy management term, and a hyper-parameter (Klein [Section 2 Method], e.g., “…Given the embeddings, we leverage a joint loss, consisting of two objectives: min θ1,θ2 LS(fθ1 ) + αLC(fθ1 , pθ2, …The objective of LS is to increase the contrast of the augmented embedding…The objective of LC is to reduce the redundancy and promote invariance w.r.t. augmentation in a high-dimensional space P. Here α ∈ R denotes a hyperparameter”).
Regarding claim 28, the proposed combination of Wu, Zhiyanov, Kalantidis, and Klein teaches a computer-implemented method of claim 27, wherein the decorrelating is based at least in part on the loss objective decorrelation function further comprises multiplying the hyper-parameter with the feature redundancy management term and summing the augmentation invariance term with the multiplying of the hyper-parameter with the feature redundancy management term (Klein [Section 2 Method] e.g., “Here α ∈ R denotes a hyperparameter and p : T → P is a projector (MLP) parameterized by θ2, which maps the embedding to P, with |P| |T |. The objective of LS is to increase the contrast of the augmented embedding, pushing apart the embeddings h A i and h B i . The objective of LC is to reduce the redundancy and promote invariance w.r.t. augmentation in a high-dimensional space P” The formula explicitly shows the weighting (multiplication) of the redundancy term by α and summing with LS).
Regarding claim 29, the proposed combination of Wu, Zhiyanov, Kalantidis, and Klein teaches a computer-implemented method of claim 21, wherein the first vector representation satisfies a first similarity threshold relative to a sentiment of the sentence including the plurality of text, and wherein the second vector representation satisfies a second similarity threshold relative to the sentiment of the sentence including the plurality of text, wherein the second similarity threshold is less than the first similarity threshold (Klein [Section 2 Method] e.g., “…one augmentation is generated with high dropout and one with low dropout … The objective of LS is to increase the contrast of the augmented embedding, pushing apart the embeddings h A i and h Bi”. hA retains more semantic information than hB (low dropout vs. high dropout), which supports the interpretation that hA satisfies a stricter similarity threshold than hB).
Regarding claim 31, the proposed combination of Wu, Zhiyanov, Kalantidis, and Klein teaches a system of claim 30, further comprising: receiving, by the trained machine learning model, a query that includes a target sentence (Kalantidis [0034] e.g., “The dimensionality reduction model in the system 100 is embodied in or a component of an encoder 102, which is configured to receive data such as can be represented by an input vector, in an input space that is a higher dimensional representation space…”); and outputting, using at least in part the trained machine learning model, a result sentence that satisfies a similarity metric relative to the target sentence (Kalantidis [0034] e.g., “…and generate an output vector in an output space that is a lower-dimensional representation space, e.g., a d-dimensional space, where D is greater than d. The input vectors can be, for instance, provided by one or more datapoints 104a, 104b, 104b, 104c, 104d,”. This directly supports a target sentence query and outputting a result sentence that satisfies a similarity metric).
32.(Canceled)
33.(Canceled)
Regarding claim 34, the proposed combination of Wu, Zhiyanov, Kalantidis, and Klein teaches a system of claim 30, wherein the second dropout ratio is larger than the first dropout ratio (Klein [Section 2 Method] e.g., “…The random masks are associated with different ratios, rA and rB, with rA < rB”).
Regarding claim 35, the proposed combination of Wu, Zhiyanov, Kalantidis, and Klein teaches a system of claim 30, further comprising: generating a correlation matrix using the first high dimensional vector representation and the second high dimensional vector representation, wherein the decorrelating is performed on the correlation matrix (Klein [Section 2 Method] e.g., “A projector maps embeddings to a high-dimensional feature space, where the features are decorrelated”, see also [Figure 1]).
Regarding claim 36, the proposed combination of Wu, Zhiyanov, Kalantidis, and Klein teaches a system of claim 35, wherein the decorrelating on the correlation matrix is based at least in part on a loss objective decorrelation function, wherein the loss objective decorrelation function comprises an augmentation invariance term, feature redundancy management term, and a hyper-parameter (Klein [Section 2 Method], e.g., “…Given the embeddings, we leverage a joint loss, consisting of two objectives: min θ1,θ2 LS(fθ1 ) + αLC(fθ1 , pθ2, …The objective of LS is to increase the contrast of the augmented embedding…The objective of LC is to reduce the redundancy and promote invariance w.r.t. augmentation in a high-dimensional space P. Here α ∈ R denotes a hyperparameter”).
Regarding claim 37, the proposed combination of Wu, Zhiyanov, Kalantidis, and Klein teaches a system of claim 36, wherein the decorrelating is based at least in part on the loss objective decorrelation function further comprises multiplying the hyper-parameter with the feature redundancy management term and summing the augmentation invariance term with the multiplying of the hyper-parameter with the feature redundancy management term (Klein [Section 2 Method], e.g., “…Given the embeddings, we leverage a joint loss, consisting of two objectives: min θ1,θ2 LS(fθ1 ) + αLC(fθ1 , pθ2, …The objective of LS is to increase the contrast of the augmented embedding…The objective of LC is to reduce the redundancy and promote invariance w.r.t. augmentation in a high-dimensional space P. Here α ∈ R denotes a hyperparameter”).
Regarding claim 38, the proposed combination of Wu, Zhiyanov, Kalantidis, and Klein teaches a system of claim 30, wherein the first vector representation satisfies a first similarity threshold relative to a sentiment of the sentence including the plurality of text, and wherein the second vector representation satisfies a second similarity threshold relative to the sentiment of the sentence including the plurality of text, wherein the second similarity threshold is less than the first similarity threshold (Klein [Section 2 Method] e.g., “…one augmentation is generated with high dropout and one with low dropout … The objective of LS is to increase the contrast of the augmented embedding, pushing apart the embeddings h A i and h Bi”. hA retains more semantic information than hB (low dropout vs. high dropout), which supports the interpretation that hA satisfies a stricter similarity threshold than hB).
Regarding claim 40, the proposed combination of Wu, Zhiyanov, Kalantidis, and Klein teaches a non-transitory computer-readable storage medium of claim 39, further comprising: receiving, by the trained machine learning model, a query that includes a target sentence (Kalantidis [0034] e.g., “The dimensionality reduction model in the system 100 is embodied in or a component of an encoder 102, which is configured to receive data such as can be represented by an input vector, in an input space that is a higher dimensional representation space…”); and outputting, using at least in part the trained machine learning model, a result sentence that satisfies a similarity metric relative to the target sentence (Kalantidis [0034] e.g., “…and generate an output vector in an output space that is a lower-dimensional representation space, e.g., a d-dimensional space, where D is greater than d. The input vectors can be, for instance, provided by one or more datapoints 104a, 104b, 104b, 104c, 104d,”. This directly supports a target sentence query and outputting a result sentence that satisfies a similarity metric).
Conclusion
20. THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
21. Any inquiry concerning this communication or earlier communications from the examiner should be directed to BERHANU MITIKU whose telephone number is (571)270-1983. The examiner can normally be reached Monday – Friday 8:30AM – 4:00PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ajay Bhatia can be reached at 571-272-3906. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/BERHANU MITIKU/Examiner, Art Unit 2156
/AJAY M BHATIA/Supervisory Patent Examiner, Art Unit 2156