DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 04/29/2026 was filed after the mailing date of the Non-Final Rejection on 04/02/2026. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Response to Arguments
1. Regarding the rejection under 101, Applicant has amended the independent claims to include additional elements which integrate the judicial exception into a practical application. Specifically, Applicant has amended the independent claims to recite the use of a first and second batch normalization layer in a first and second projection head, respectively, for generating first and second sentence representations, wherein the first and second batch normalization layers are applied in different modes simultaneously. This asymmetric batch normalization leads to a technical improvement via improved encoder training, as information leakage due to intra-batch communication, which is prone to occur in contrastive model training, is avoided with this method (see pg. 10, lines 5-21 of Applicant’s specification). Hence, the claim language reflects a technical improvement to the field of cross-lingual retrieval, and thus is subject matter eligible under Step 2A Prong 2. Accordingly, the rejection is withdrawn.
2. Regarding the rejections under 103, Applicant's arguments filed 04/29/2026 have been fully considered but they are not persuasive.
Applicant argues on pgs. 10-13 of the Remarks that the cited prior art fails to teach “generating, by a first projection head, a first sentence representation of a first sentence of the sentence pair, the first projection head applying a first batch normalization layer in a first mode; and generating, by a second projection head, a second sentence representation of a second sentence of the sentence pair, the second projection head applying a second batch normalization layer in a second mode different from the first mode simultaneously with the first mode…”. Specifically, Applicant argues that the applied He reference’s shuffle batch normalization does not read on the claimed “asymmetric” batch normalization. The Examiner respectfully disagrees with these arguments. While the Examiner agrees that shuffle BN taught in He is ultimately different from the batch normalization method described in Applicant’s specification, the broadest reasonable interpretation of the claimed invention of claim 1reads on the He reference. Under the broadest reasonable interpretation, “mode” is being interpreted as a particular way, manner, or method of performing the batch normalization. Thus, the limitation only requires that the cited art teach a first and second batch normalization layer which process in a way/manner/method different from each other at the same time. The limitation as currently written does not require that the first and second mode be, for example, a training mode and an evaluation mode, as described in Applicant’s specification. Thus, the He reference teaches this limitation. Specifically, He teaches a first and second BN layer (see pg. 4, section ‘Shuffling BN’, where fq and fk both have a respective BN layer) and teaches that the manner/way/method of performing the first BN is different and simultaneous to the second manner/way/method of performing the second BN (see pg. 4, section ‘Shuffling BN’ “We resolve this problem by shuffling BN. We train with multiple GPUs and perform BN on the samples independently for each GPU (as done in common practice). For the key encoder fk, we shuffle the sample order in the current mini-batch before distributing it among GPUs (and shuffle back after encoding); the sample order of the mini-batch for the query encoder fq is not altered.”; i.e. a first BN is performed (with a first sample order) and a second BN is performed (with a second sample order); see also Algorithm 1, pseudocode, where x_q and x_k are different augmentations/sample orders of a minibatch being respectively input to f_q and f_k for a same input minibatch). This concept reads on the claim language recited.
Hence, Applicant’s arguments are not persuasive.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
3. Claims 1-3, 7, 10, and 14-15 are rejected under 35 U.S.C. 103 as being unpatentable over Niu et al. (US 2023/0153542 A1, hereinafter Niu) in view of Fu et al. (NPL ABSent: Cross-Lingual Sentence Representation Mapping with Bidirectional GANs, hereinafter Fu) and further in view of Chen et al. (NPL “A Simple Framework for Contrastive Learning of Visual Representations”, hereinafter Chen) and He et al. (NPL “Momentum Contrast for Unsupervised Visual Representation Learning”, hereinafter He).
Regarding claim 1, Niu discloses A method for sentence representation generation for cross-lingual retrieval, the method comprising: pretraining an encoder, the pretraining comprising (para. 0029 “To address this challenge, a contrastive learning approach can be adopted and the aligner model 110 is trained on a classification task with in-batch negatives.”; para. 0030-0031 “Fig. 3 is a simplified logic flow diagram illustrating a method of training an aligner model…At step 302, a training dataset is received…”): obtaining a plurality of sentence pairs (para. 0032 “…OPUS-100 contains approximately 55 M sentence pairs. For example, 99 language pairs are chosen for training the aligner model, 44 of which are chosen from 1M sentence pairs of training data, 73 chosen from at least 100 k, and 95 chosen from at least 10 k. Following OPUS-100's choice, the training data for each language pair in New-Tatoeba is capped at 1 M to make it easier to compare with OPUS-trained models.”); for each sentence pair of the plurality of sentence pairs, generating a sub-contrastive prediction loss (para. 0029 “To address this challenge, a contrastive learning approach can be adopted and the aligner model 110 is trained on a classification task with in-batch negatives. For example, for the batch of sentences S… in a source language, and a batch of sentences T… in a target language, wither S.sub.i is aligned with T.sub.i for each I, a pairwise semantic similarity between S and T is computed to obtain N similarities for the positive alignments, and N.sup.2-N similarities for the negative ones (in total N.sup.2 similarities computed). During training, these similarity scores are used as logits and pair each positive log with all negative ones. These logits are then sed to compute the contrastive loss 120… ”) by: …optimizing, by an optimizer, the encoder by minimizing a contrastive prediction loss generated based on the sub-contrastive prediction loss corresponding to each respective sentence pair of the plurality of sentence pairs; obtaining a target sentence (para. 0029 “To address this challenge, a contrastive learning approach can be adopted and the aligner model 110 is trained on a classification task with in-batch negatives. For example, for the batch of sentences S… in a source language, and a batch of sentences T… in a target language, wither S.sub.i is aligned with T.sub.i for each I, a pairwise semantic similarity between S and T is computed to obtain N similarities for the positive alignments, and N.sup.2-N similarities for the negative ones (in total N.sup.2 similarities computed). During training, these similarity scores are used as logits and pair each positive log with all negative ones. These logits are then used to compute the contrastive loss 120… ”; para. 0042 “In some examples, the cross-lingual transfer module 430, may receive an input 440, e.g., such as an input text in a source language and/or target language, via a data interface 415.”); generating, by the encoder, an initial target sentence representation of the target sentence; generating …a target sentence representation of the target sentence for cross-lingual retrieval… (para. 0047 “In some embodiments, the aligner model may be tested via cross-lingual sentence retrieval tasks, which retrieve a matching sentence in the target language from a collection of sentences.”).
Niu does not specifically disclose:
generating…a first sentence representation of a first sentence of the sentence pair…
generating…a second sentence representation of a second sentence of the sentence pair…
and [generating], by a cross-lingual calibration unit, a target sentence representation of the target sentence for cross-lingual retrieval] based on the initial target sentence representation.
Fu teaches generating…a first sentence representation of a first sentence of the sentence pair… (Fu teaches generating a first sentence representation utilizing an initial first sentence representation: Fig. 2 caption: “It learns two generators Gx and Gy to approximate the joint distribution of vectors from both language. Gx projects sentence embeddings x from language X to Y…”; see Gx(X) in Fig. 2), generating…a second sentence representation of a second sentence of the sentence pair… (Fu teaches generating a second sentence representation utilizing an initial second sentence representation: Fig. 2 caption: “It learns two generators Gx and Gy to approximate the joint distribution of vectors from both language... while, conversely, Gy projects sentence embeddings y from language Y to X…”; see Gy(Y) in Fig. 2), and generating a target sentence representation by a cross-lingual calibration unit and based on the initial target sentence representation through cross-lingual calibration (Fig. 2; pg. 4 section 3.4 “With regard to obtaining the sentence represntations, we adopt pre-trained word vectors…we adopt simple (weighted) averages of word vectors, which are surprisingly powerful, although our method could also be applied to other sentence embedding methods. Given source sentence embeddings {x} and target sentence embeddings {y} acquired as above, we can train the generators Gx and Gy through the joint loss function in Equation 5. Subsequently, we evaluate the obtained transformation via standard sentence retrieval task. For each source sentence embedding, we compute its k neural neighbors in terms of the distance function fd among all target embeddings. The corresponding k target sentences are regarded as the candidate set of mapping results.”).
Niu and Fu are considered to be analogous to the claimed invention as
they both are in the same field of cross-lingual embedding. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Niu to incorporate the teachings of Fu in order to generate a target sentence representation of the target sentence for cross-lingual retrieval specifically based on the initial target sentence representation through cross-lingual calibration, and to generate a first sentence representation and a second sentence representation representing a first and second sentence in a sentence pair. Doing so would improve performance on retrieval tasks (pg. 8, Conclusion).
Niu in view of Fu does not specifically teach [generating], by a first projection head, [a first sentence representation…]; and [generating], by a second projection head, [a second sentence representation…] …
Chen teaches [generating], by a first projection head, [a first sentence representation…] (Fig. 2, hi passed through projection head g(.) to obtain zi; pg. 2 section 2.1, 3rd bullet pt: “A small neural network projection head g(.) that maps representations to the space where contrastive loss is applied.”; 4th bullet pt: “A contrastive loss function for a contrastive prediction task…Given a set {xk} including a positive pair of examples…”; batch normalization utilized ); and [generating], by a second projection head, [a second sentence representation…] … (Fu teaches generating a second sentence representation utilizing an initial second sentence representation: Fig. 2 caption: “It learns two generators Gx and Gy to approximate the joint distribution of vectors from both language... while, conversely, Gy projects sentence embeddings y from language Y to X…”; see Gy(Y) in Fig. 2).
Niu, Fu, and Chen are considered to be analogous to the claimed invention as
Niu and Fu are in the same field of natural language processing and Chen is in the same field of contrastive learning. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Niu in view of Fu to incorporate the teachings of Chen in order to specifically obtain a center sentence and context sentence representation through a first and second projection head respectively, and to use these representations to generate the sub-contrastive loss. Doing so would improve the representation quality of the hidden representation preceding the projection head (pg. 6, section 4.2).
Niu in view of Fu and Chen does not specifically disclose [the first projection head] applying a first batch normalization layer in a first mode; [and…the second projection head] applying a second batch normalization layer in a second mode different from the first mode simultaneously with the first mode…
He teaches applying a first batch normalization layer in a first mode (pg. 4, section 3.3 “Shuffling BN”: “Our encoders Fq and Fk both have Batch Normalization (BN) [37] as in the standard ResNet [33].”); and applying a second batch normalization layer in a second mode different from the first mode simultaneously with the first mode… (pg. 4, section 3.3 “Shuffling BN”: “Our encoders Fq and Fk both have Batch Normalization (BN) [37] as in the standard ResNet [33].”; pg. 4, section 3.3 “Shuffling BN”: “… We resolve this problem by shuffling BN. We train with multiple GPUs and perform BN on the samples independelty for each GPU (as done in common practice). For the key encoder Fk, we shuffle the sample order in the current mini-batch before distributing it among GPUs (and shuffle back after encoding); the sample order of the mini-batch for the query encoder Fq is not altered…”).
Niu, Fu, Chen, and He are considered to be analogous to the claimed invention as Niu and Fu are in the same field of natural language processing and Chen and He are in the same field of contrastive learning. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Niu in view of Fu and Chen in order to specifically have the first projection head include a first batch normalization layer, the second projection head include a second batch normalization layer, and the first and second batch normalization layers be in different batch normalization modes at the same time. Doing do would be beneficial, as this would allow the contrastive training to benefit from batch normalization while avoiding the model from “cheating” and learning poor representations (He, pg. 4, section 3.3 “Shuffling BN”).
Regarding claim 2, Niu in view of Fu, Chen, and He discloses wherein the target sentence is in a first language (Niu, para. 0042 “In some examples, the cross-lingual transfer module 430, may receive an input 440, e.g., such as an input text in a source language and/or target language, via a data interface 415.”), and the target sentence representation is suitable for performing a cross-lingual retrieval task across the first language and a second language (Fu, pg. 4 section 3.4 “Given source sentence embeddings {x} and target sentence embeddings {y} acquired as above, we can train the generators Gx and Gy through the joint loss function in Equation 5. Subsequently, we evaluate the obtained transformation via standard sentence retrieval task. For each source sentence embedding, we compute its k neural neighbors in terms of the distance function fd among all target embeddings. The corresponding k target sentences are regarded as the candidate set of mapping results.”; pg. 4, section 4.1 “We focus on German and English as well as Spanish and English translation retrieval.”).
Niu, Fu, Chen, and He are considered to be analogous to the claimed invention as Niu and Fu are in the same field of natural language processing and Chen and He are in the same field of contrastive learning. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated the teachings of Fu in order to specifically have the target sentence representation be suitable for performing a cross-lingual retrieval task across the first and second languages, for the same rationale as given in claim 1.
Regarding claim 3, Niu in view of Fu, Chen, and He discloses wherein each sentence pair of the plurality of sentence pairs includes two sentences located in a same context window (para. 0032 “In one embodiment, the training data set may be (1) an English-centered dataset such as OPUS-100; (2) a non-English-centered language dataset, e.g., the v2021-08-07 Tatoeba Challenge. OPUS-100 is English-centered, meaning that all training pairs include English on either the source or target side. The corpus covers 100 languages (including English). The languages for training are selected based on the volume of parallel data available in OPUS. The OPUS collection is comprised of multiple corpora, ranging from movie subtitles to GNOME documentation to the Bible. OPUS-100 contains approximately 55 M sentence pairs. For example, 99 language pairs are chosen for training the aligner model, 44 of which are chosen from 1M sentence pairs of training data, 73 chosen from at least 100 k, and 95 chosen from at least 10 k. Following OPUS-100's choice, the training data for each language pair in New-Tatoeba is capped at 1 M to make it easier to compare with OPUS-trained models.”), and the pretraining further comprises combining the plurality of sentence pairs into a training dataset (sentence pairs selected to make up training data: para. 0032 “In one embodiment, the training data set may be (1) an English-centered dataset such as OPUS-100; (2) a non-English-centered language dataset, e.g., the v2021-08-07 Tatoeba Challenge. OPUS-100 is English-centered, meaning that all training pairs include English on either the source or target side. The corpus covers 100 languages (including English). The languages for training are selected based on the volume of parallel data available in OPUS. The OPUS collection is comprised of multiple corpora, ranging from movie subtitles to GNOME documentation to the Bible. OPUS-100 contains approximately 55 M sentence pairs. For example, 99 language pairs are chosen for training the aligner model, 44 of which are chosen from 1M sentence pairs of training data, 73 chosen from at least 100 k, and 95 chosen from at least 10 k. Following OPUS-100's choice, the training data for each language pair in New-Tatoeba is capped at 1 M to make it easier to compare with OPUS-trained models.”).
Regarding claim 7, Niu in view of Fu, Chen, and He discloses wherein the generating a sub-contrastive prediction loss further comprises: predicting an initial center sentence representation of the center sentence through the encoder (Niu, para. 0029 “To address this challenge, a contrastive learning approach can be adopted and the aligner model 110 is trained on a classification task with in-batch negatives. For example, for the batch of sentences S… in a source language, and a batch of sentences T… in a target language, wither S.sub.i is aligned with T.sub.i for each I, a pairwise semantic similarity between S and T is computed to obtain N similarities for the positive alignments, and N.sup.2-N similarities for the negative ones (in total N.sup.2 similarities computed). During training, these similarity scores are used as logits and pair each positive log with all negative ones. These logits are then used to compute the contrastive loss 120… ”; para. 0034 “At step 306, a pretrained multi-lingual model may be used to compute a pairwise token-level similarity between the two sentences within each positive input pair or negative input pair. For example, the pairwise token-level similarity between two sentences may be computed as the BERT score described in relation to FIG. 2.”; this similarity computation requires a first sentence (Si, Fig. 2) to be embedded via the encoder 105 to obtain contextual embedding); and predicting an initial context sentence representation of the context sentence through the encoder (para. 0029 “To address this challenge, a contrastive learning approach can be adopted and the aligner model 110 is trained on a classification task with in-batch negatives. For example, for the batch of sentences S… in a source language, and a batch of sentences T… in a target language, wither S.sub.i is aligned with T.sub.i for each I, a pairwise semantic similarity between S and T is computed to obtain N similarities for the positive alignments, and N.sup.2-N similarities for the negative ones (in total N.sup.2 similarities computed). During training, these similarity scores are used as logits and pair each positive log with all negative ones. These logits are then used to compute the contrastive loss 120… ”; para. 0034 “At step 306, a pretrained multi-lingual model may be used to compute a pairwise token-level similarity between the two sentences within each positive input pair or negative input pair. For example, the pairwise token-level similarity between two sentences may be computed as the BERT score described in relation to FIG. 2.”; this similarity computation requires a second sentence (Tj, Fig. 2) to be embedded via the encoder 105 to obtain contextual embedding).
Regarding claim 10, Niu in view of Fu, Chen, and He discloses wherein the generating a target sentence representation comprises: generating the target sentence representation through performing, on the initial sentence representation, at least one of shifting, scaling, and rotating (Fu, projecting sentence embeddings from a language X/Y to a different language Y/X reads on the BRI of shifting that initial sentence representation to get a target representation).
Niu, Fu, Chen, and He are considered to be analogous to the claimed invention as Niu and Fu are in the same field of natural language processing and Chen and He are in the same field of contrastive learning. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated the teachings of Fu in order to specifically have the target sentence representation be obtained via performing on the initial sentence representation at least one of shifting, scaling, or rotating, for the same rationale as given in claim 1.
Regarding claim 14, claim 14 is a system claim with limitations similar to those recited in method claim 1, and is thus rejected under similar rationale.
Additionally, Niu teaches An apparatus for sentence representation generation for cross-lingual retrieval, comprising: at least one processor; and a memory storing computer-executable instructions that, when executed, cause the at least one processor to (para. 0006 “According to a third aspect of the disclosure, an electronic device is provided and includes: at least one processor; and a memory communicatively connected to the at least one processor; in which the memory is stored with instructions executable by the at least one processor, and when the instructions are performed by the at least one processor, the at least one processor is caused to perform the method for generating a cross-lingual textual semantic model according to the first aspect, or the method for determining a textual semantic according to the second aspect.”).
Regarding claim 15, claim 15 is a computer product claim with limitations similar to those recited in method claim 1, and is thus rejected under similar rationale.
Additionally, Niu discloses A computer program product for sentence representation generation for cross-lingual retrieval, comprising a computer program that is executed by at least one processor for (para. 0007 “According to a fourth aspect of disclosure, a non-transitory computer-readable storage medium stored with computer instructions is provided, in which the computer instructions are configured to cause a computer to perform the method for generating a cross-lingual textual semantic model according to the first aspect or perform the method for determining a textual semantic according to the second aspect.”).
4. Claims 4, 9, and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Niu in view of Fu, Chen, and He, and further in view of Goswami et al. (NPL Cross-Lingual Sentence Embedding using Multi-Task Learning, hereinafter Goswami).
Regarding claim 4, Niu in view of Fu, Chen, and He does not specifically disclose wherein the two sentences are two sentences in a same language.
Goswami teaches wherein the two sentences are two sentences in a same language (pg. 3 section 3. “In this section, we describe the components and working of the proposed Dual Encoder with Anchor Model (DuEAM) architecture for multilingual sentence embeddings, trained using an unsupervised multi-task joint loss function…”; pg. 5 section 5 “Weakly Supervised Data”: “This training data contains both monolingual and cross-lingual sentence pairs, where the monolingual sentence pairs are same as those of XNLI (without annotated labels)…”).
Niu, Fu, Chen, He, and Goswami are considered to be analogous to the claimed invention as Niu, Fu, and Goswami are in the same field of natural language processing and Chen and He are in the same field of contrastive learning. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Niu in view of Fu to incorporate the teachings of Goswami in order to have the two sentences be sentences in the same language. Doing so would be beneficial, as this would allow the model to learn semantic textual similarity between monolingual datasets which is a major task for sentence embedding (pg. 5, section 6.1, 1st para.).
Regarding claim 9, Niu in view of Fu, Chen, and He discloses a previous representation set corresponding to a previous training dataset is stored in a memory bank (He, teaches queue/memory bank (see Fig. 2, (b)(c), and caption “(c): MoCo encodes the new keys on-the-fly by a momentum-updated encoder, and maintains a queue (not illustrated in this figure) of keys…”); pg. 1, section 1, 4th para. “…We maintain the dictionary as a queue of data samples: the encoded representation of the current mini-batch are enqueued, and the oldest are dequeued…”; see pg. 3, ‘Dictionary as a queue’: “At the core of our approach is maintaining the dictionary as a queue of data samples. This allows us to reuse the encoded keys from the immediate preceding mini-batches…”), and the generating the sub-contrastive prediction losses comprises: extracting a language-specific representation set for the third language from the previous representation set (pg. 1, section 1, 4th para. “…We maintain the dictionary as a queue of data samples: the encoded representation of the current mini-batch are enqueued, and the oldest are dequeued…”; see also pg. 4, Algorithm 1, enqueue operation performed for current mini-batch); and generating the sub-contrastive prediction loss based further on the language-specific representation set (He teaches the language-specific representation set in a memory bank/queue; Niu teaches the generation of sub-contrastive prediction losses based on encoded representations (see above claim mapping for claim 1)).
Niu, Fu, Chen, and He are considered to be analogous to the claimed invention as Niu and Fu are in the same field of natural language processing and Chen and He are in the same field of contrastive learning. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated the teachings of He in order to specifically have a previous representation set corresponding to a previous training dataset is stored in a memory bank, and the generating the sub-contrastive prediction loss involve extracting a language-specific representation set for the third language from the previous representation set to be used for generating the loss. Doing so would be beneficial, as this would enable large-scale contrastive learning which requires large numbers of negative training samples (see NPL Sun et al., Contrastive Distillation on Intermediate Representations for Language Model Compression, pg. 5, section 3.3).
Niu in view of Fu, Chen, and He does not specifically disclose wherein the first sentence and the second sentence are sentences in a third language…
Goswami teaches wherein the first sentence and the second sentence are sentences in a third language… (pg. 3 section 3. “In this section, we describe the components and working of the proposed Dual Encoder with Anchor Model (DuEAM) architecture for multilingual sentence embeddings, trained using an unsupervised multi-task joint loss function…”; pg. 5 section 5 “Weakly Supervised Data”: “This training data contains both monolingual and cross-lingual sentence pairs, where the monolingual sentence pairs are same as those of XNLI (without annotated labels)…”)).
Niu, Fu, Chen, He, and Goswami are considered to be analogous to the claimed invention as Niu, Fu, and Goswami are in the same field of natural language processing and Chen and He are in the same field of contrastive learning. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Niu in view of Fu to incorporate the teachings of Goswami in order to have the two sentences be sentences in a third language. Doing so would be beneficial, given the same rationale as claim 4.
Regarding claim 16, Niu in view of Fu, Chen, He, and Goswami discloses wherein the memory bank is maintained in a first-in-first-out manner (He, see pg. 4, Algorithm 1, enqueue and dequeue operation; by enqueuing a current minibatch and dequeuing an earliest minibatch, a FIFO memory bank is realized), and the pretraining further comprises storing the first sentence representation and the second sentence representation in the memory bank for use as a previous representation set in a subsequent training dataset (He teaches storing encoded representations in a memory bank: pg. 1, section 1, 4th para. “…We maintain the dictionary as a queue of data samples: the encoded representation of the current mini-batch are enqueued, and the oldest are dequeued…”; see also pg. 4, Algorithm 1, enqueue operation performed for current mini-batch; Fu teaches the generating the first and second sentence representations (see claim mapping for claim 1)).
Niu, Fu, Chen, He, and Goswami are considered to be analogous to the claimed invention as Niu, Fu, and Goswami are in the same field of natural language processing and Chen and He are in the same field of contrastive learning. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated the teachings of He in order to have the memory bank be maintained in a first-in-first-out manner, and to store a first and second sentence representation in the memory bank for use as a previous representation set in a subsequent training dataset. Doing so would be beneficial, given the same rationale as claim 9.
5. Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Niu in view of Fu, Chen, and He, and further in view of Kim (NPL [Part 1] Predicting on Text Pairs with Transformers: Cross-Encoding with BERT).
Regarding claim 5, Niu in view of Fu, Chen, and He does not specifically disclose identifying a plurality of center sentences in at least one document; for each center sentence in the plurality of center sentences, determining a context window centered on the center sentence in the at least one document, extracting a context sentence form the context window, and combining the center sentence and the context sentence into a sentence pair corresponding to the center sentence, and obtaining the plurality of sentence pairs corresponding to the plurality of center sentences.
Kim teaches identifying a plurality of center sentences in at least one document (pg. 4 4th para. “…one of the main ways it trained itself was through looking at sentence pairs…From their massive text data-set”; see figure on pg. 4 “Sentence 1” reads on a center sentence in a document (data-set corpus)); for each center sentence in the plurality of center sentences, determining a context window centered on the center sentence in the at least one document (pg. 4, 5th para. “…they made pairs on sentences that were next to each other in the corpus…”; context window is sentence that is adjacent to a first sentence), extracting a context sentence form the context window (pg. 4, 5th para. “…they made pairs on sentences that were next to each other in the corpus…”; context window is sentence that is adjacent to a first sentence; see figure on pg. 4, “Sentence 2”), and combining the center sentence and the context sentence into a sentence pair corresponding to the center sentence (pg. 4, 5th para. “From their massive text data-set, they made these pairs of sentences that were next to each other in the corpus…”) and obtaining the plurality of sentence pairs corresponding to the plurality of center sentences (pg. 4, 5th para. “From their massive text data-set, they made these pairs of sentences that were next to each other in the corpus…”; see pseudo-code, sentence pairs used for training model).
Niu, Fu, Chen, He, and Kim are considered to be analogous to the claimed invention as Niu, Fu, and Kim are in the same field of natural language processing, and Chen and He are in the same field of contrastive learning. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Niu in view of Fu, Chen, and He to incorporate the teachings of Kim in order to identify a plurality of center sentences in at least one document, determine a context window centered on the center sentence, extract a context sentence from the context window, and combine the center and context sentences to obtain the plurality of sentence pairs. Doing so would enable the model to learn the notion of text sequences (Kim, pg. 4).
6. Claims 11 and 18-19 are rejected under 35 U.S.C. 103 as being unpatentable over Niu in view of Fu, Chen, and He, and further in view of Zheng et al. (US 2023/0214177 A1, hereinafter Zheng)
Regarding claim 11, Niu in view of Fu, Chen, and He discloses wherein the target sentence is a sentence in a first language (Niu, para. 0042 “In some examples, the cross-lingual transfer module 430, may receive an input 440, e.g., such as an input text in a source language and/or target language, via a data interface 415.”) and the shifting (see above claim mapping for claim 10). However, Niu in view of Fu, Chen, and He does not specifically disclose [the shifting] comprises: subtracting a predetermined mean from a current sentence representation, the predetermined mean computed based on a set of representations corresponding to a set of sentences in the first language.
Zheng teaches [the shifting] comprises: subtracting a predetermined mean from a current sentence representation, the predetermined mean computed based on a set of representations corresponding to a set of sentences in the first language (Zheng teaches shifting using a predetermined mean for a particular set of embeddings: para. 0124 “As shown in Fig. 3, at step 302, process 300 may include receiving at least one embedding set…”; para. 0128-0129 “As shown in Fig. 3, at step 304, process 300 may include applying mean centering…In some non-limiting embodiments or aspects, applying mean centering may include determining a mean based on all embedding vectors of the set of embedding vectors. Additionally or alternatively, the mean may be subtracted from each embedding vector of the set of embedding vectors.”).
Niu, Fu, Chen, He, and Zheng are considered to be analogous to the claimed invention as Niu, Fu, and Zheng are in the same field of natural language processing and Chen and He are in the same field of contrastive learning. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Niu in view of Fu, Chen, and He in order to incorporate the teachings of Zheng in order to specifically perform shifting by subtracting a predetermined mean from a current sentence representation, the predetermined mean computed based on a set of representations corresponding to a set of sentences in the first language. Doing so would be beneficial, as applying such preprocessing on embedding vectors would improve cross lingual word embedding performance (Zheng, para. 0003).
Regarding claim 18, Niu in view of Fu, Chen, and He discloses wherein generating the target sentence representation comprises performing at least one of shifting…(Fu, see claim mapping for claim 10).
Niu, Fu, Chen, and He are considered to be analogous to the claimed invention as Niu and Fu are in the same field of natural language processing and Chen and He are in the same field of contrastive learning. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated the teachings of Fu in order to specifically have the target sentence representation be obtained via performing on the initial sentence representation at least one of shifting, scaling, or rotating, for the same rationale as given in claim 1.
Niu in view of Fu, Chen, and He does not specifically disclose performing at least one of [shifting] by subtracting a predetermined mean from the initial target sentence representation, scaling by dividing the initial target sentence representation by a predetermined variance, or rotating the initial target sentence representation based on a predetermined rotation matrix between a first language and a second language.
Zheng teaches performing at least one of [shifting] by subtracting a predetermined mean from the initial target sentence representation, scaling by dividing the initial target sentence representation by a predetermined variance, or rotating the initial target sentence representation based on a predetermined rotation matrix between a first language and a second language (1st option taught: Zheng teaches shifting using a predetermined mean for a particular set of embeddings: para. 0124 “As shown in Fig. 3, at step 302, process 300 may include receiving at least one embedding set…”; para. 0128-0129 “As shown in Fig. 3, at step 304, process 300 may include applying mean centering…In some non-limiting embodiments or aspects, applying mean centering may include determining a mean based on all embedding vectors of the set of embedding vectors. Additionally or alternatively, the mean may be subtracted from each embedding vector of the set of embedding vectors.”).
Niu, Fu, Chen, He, and Zheng are considered to be analogous to the claimed invention as Niu, Fu, and Zheng are in the same field of natural language processing and Chen and He are in the same field of contrastive learning. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Niu in view of Fu, Chen, and He in order to incorporate the teachings of Zheng in order to specifically perform shifting by subtracting a predetermined mean from a current sentence representation. Doing so would be beneficial, given the same rationale as claim 11.
Regarding claim 19, Niu in view of Fu, Chen, He, and Zheng discloses wherein the predetermined mean and the predetermined variance are computed based on a set of representations corresponding to a set of sentences in a language of the target sentence (Zheng teaches 1st option of claim 18 of performing shifting by subtracting a predetermined mean; this predetermined mean is computed based on a set of representations corresponding to a set of sentences in a particular language: para. 0124 “As shown in Fig. 3, at step 302, process 300 may include receiving at least one embedding set…”; para. 0128-0129 “As shown in Fig. 3, at step 304, process 300 may include applying mean centering…In some non-limiting embodiments or aspects, applying mean centering may include determining a mean based on all embedding vectors of the set of embedding vectors. Additionally or alternatively, the mean may be subtracted from each embedding vector of the set of embedding vectors.”; Niu teaches a set of sentences in a language of a target language, see above claim mapping for claim 1).
Niu, Fu, Chen, He, and Zheng are considered to be analogous to the claimed invention as Niu, Fu, and Zheng are in the same field of natural language processing and Chen and He are in the same field of contrastive learning. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Niu in view of Fu, Chen, and He in order to incorporate the teachings of Zheng in order to specifically have the predetermined mean be computed based on a set of representations corresponding to a set of sentences in a language of the target sentence. Doing so would be beneficial, given the same rationale as claim 11.
7. Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Niu in view of Fu, Chen, and He, and further in view of Zhao et al. (NPL Inducing Language-Agnostic Multilingual Representations, hereinafter Zhao).
Regarding claim 12, Niu in view of Fu, Chen, and He discloses wherein the target sentence is a sentence in a first language (Niu, para. 0042 “In some examples, the cross-lingual transfer module 430, may receive an input 440, e.g., such as an input text in a source language and/or target language, via a data interface 415.”) and a current sentence representation (Niu, Fig. 4 “Cross-Lingual Transfer Module 230”; para. 0042 “The cross-lingual transfer module 430 may generate an output 450 such as an alignment with a sentence in the target language corresponding to the input 440) and a set of sentences in the first language (Niu para. 0032 “In one embodiment, the training data set may be (1) an English-centered dataset such as OPUS-100; (2) a non-English-centered language dataset, e.g., the v2021-08-07 Tatoeba Challenge. OPUS-100 is English-centered, meaning that all training pairs include English on either the source or target side. The corpus covers 100 languages (including English)…). However, Niu in view of Fu, Chen, and He does not specifically disclose and the scaling comprises: dividing a current sentence representation by a predetermined variance, the predetermined variance computed based on a set of representations corresponding to a set of sentences in the first language.
Zhao teaches the scaling comprises: dividing an embedding by a predetermined variance, the predetermined variance computed based on a set of representations corresponding to a set of embeddings (pg. 4, section 3.2 “Vector space normalization”: “We add a batch normalization layer that constrains all embeddings of different language into a distribution with zero mean and unit variance…where ε is a constant value for numerical stability, µβ and σβ are mean and variance, serving as per batch statistics for each time step in a sequence…”; see Eq. 4, embedding f(I,s) divided by variance σ ).
Niu, Fu, Chen, He, and Zhao are considered to be analogous to the claimed invention as Niu, Fu, and Zhao are in the same field of natural language processing and Chen and He are in the same field of contrastive learning. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Niu in view of Fu, Chen, and He in order to incorporate the teachings of Zhao in order to specifically perform scaling by dividing a current sentence representation by a predetermined variance, the predetermined variance computed based on a set of representations corresponding to a set of sentences in the first language. Doing so would be beneficial, as this would remove remove language identity signal such as variance, from multi-lingual embeddings and increasing the discriminativeness of the embeddings (Zhao, pg. 4, section 3.2).
8. Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over Niu in view of Fu, Chen, and He, and further in view of Liu et al. (US 2020/0372106 A1, hereinafter Liu).
Regarding claim 13, Niu in view of Fu, Chen, and He discloses wherein the target sentence is a sentence in a first language (Niu, para. 0042 “In some examples, the cross-lingual transfer module 430, may receive an input 440, e.g., such as an input text in a source language and/or target language, via a data interface 415.”), the target sentence representation is to be used for performing a cross-lingual retrieval task across the first language and a second language…(Fu, pg. 4 section 3.4 “Given source sentence embeddings {x} and target sentence embeddings {y} acquired as above, we can train the generators Gx and Gy through the joint loss function in Equation 5. Subsequently, we evaluate the obtained transformation via standard sentence retrieval task. For each source sentence embedding, we compute its k neural neighbors in terms of the distance function fd among all target embeddings. The corresponding k target sentences are regarded as the candidate set of mapping results.”; pg. 4, section 4.1 “We focus on German and English as well as Spanish and English translation retrieval.”).
Niu, Fu, Chen, and He are considered to be analogous to the claimed invention as Niu and Fu are in the same field of natural language processing and Chen and He are in the same field of contrastive learning. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention in order to use the teachings of Fu to specifically have the target sentence representation be suitable for performing a cross-lingual retrieval task across the first and second languages, for the same rationale as given in claim 1.
Niu in view of Fu, Chen, and He discloses a current sentence representation (Niu, Fig. 4 “Cross-Lingual Transfer Module 230”; para. 0042 “The cross-lingual transfer module 430 may generate an output 450 such as an alignment with a sentence in the target language corresponding to the input 440), but does not specifically disclose the rotating comprises: rotating a current sentence representation based on a predetermined rotation matrix between the first language and the second language.
Liu teaches the rotating comprises: rotating a representation based on a predetermined rotation matrix between the first language and the second language (para. 0043 “To provide additional details for an improved understanding of selected embodiments of the present disclosure, reference is now made to FIG. 4 which depicts a simplified illustration 400 of a sequence for aligning embeddings E1, E2 from different languages or domains into a shared embedding space. In FIG. 4A, there is shown two non- aligned sets of embeddings E1, E2 that are trained independently on monolingual data, including a first embedding E1 that includes English words and a second embedding E2 that includes Spanish words to be aligned/translated. Each dot represents a word in that space, and the size of the dot is proportional to the frequency of the words in the training corpus of that language. In FIG. 4B, the first embedding E1 is rotated into rough alignment with the second embedding E2, such as by using adversarial learning to learn a rotation matrix W for roughly aligning the two distributions. In FIG. 4C, the mapping rotation matrix W may be further refined using a geometric transformation, such as a Procrustes transformation, that involves only translation, rotation, uniform scaling, or a combination of these transformations whereby frequent words aligned by the previous step are used as anchor points to minimize an energy function that corresponds to a spring system between anchor points. The refined mapping rotation matrix W′ is then applied to the first embedding E1 to map all words in the dictionary. In FIG. 4D, the first embedding E1 is translated by using the mapping rotation matrix W′ and a distance metric that expands the space where there is high density of points (like the area around the word “cat”), so that “hubs” (like the word “cat”) become less close to other word vectors than they would otherwise (compared to the same region in FIG. 4A)…”).
Niu, Fu, Chen, He, and Liu are considered to be analogous to the claimed invention as Niu, Fu, and Liu are in the same field of natural language processing and Chen and He are in the same field of contrastive learning. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Niu in view of Fu, Chen, and He in order to incorporate the teachings of Liu to specifically perform rotating by rotating a current sentence representation based on a predetermined rotation matrix between the first language and the second language. Doing so would be beneficial, as this would align the monolingual embeddings from different language such that they are aligned in a shared space where words of high semantic similarity across language are close to each other (Liu, para. 0024, para. 0043).
9. Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Niu in view of Fu, Chen, and He, and further in view of Chung et al. (NPL “W2V-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training”, hereinafter Chung).
Regarding claim 20, Niu in view of Fu, Chen, and He discloses …minimizing…the contrastive prediction loss (Niu, see above claim mapping for claim 1). However, Niu in view of Fu, Chen, and He does not specifically disclose generating a masked language model loss and optimizing the encoder by further minimizing the masked language model loss in combination with the contrastive prediction loss.
Chung teaches generating a masked language model loss (Fig. 1, see ‘MLM Loss’; see pg. 3, ‘Masked prediction loss’) and optimizing the encoder by further minimizing the masked language model loss in combination with the contrastive prediction loss (Fig. 1, see ‘Contrastive loss’; see pg. 2, ‘Contrastive loss’; pg. 3, ‘Masked prediction loss’, 2nd para. “w2v-BERT is trained to solve the two self-supervised tasks at the same time. The final training loss to be minimized is… [Equation 2]”; Equation 2 minimizes for both contrastive loss Lc and masked language model loss Lm).
Niu, Fu, Chen, He, and Chung are considered to be analogous to the claimed invention as Niu, Fu, and Chung are in the same field of natural language processing and Chen and He are in the same field of contrastive learning. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Niu in view of Fu, Chen, and He in order to incorporate the teachings of Chung to specifically generate a masked language model loss and to optimize the encoder by further minimizing the masked language model loss in combination with the contrastive prediction loss. Doing so would be beneficial, as combining contrastive learning with masked language modeling was found to be more effective than contrastive learning along for pretraining of the encoder (Chung, pg. 4, section 5.1).
Allowable Subject Matter
10. Claim 17 is objected to as being dependent upon a rejected base claim but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CODY DOUGLAS HUTCHESON whose telephone number is (703)756-1601. The examiner can normally be reached M-F 8:00AM-5:00PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre-Louis Desir can be reached at (571)-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/CODY DOUGLAS HUTCHESON/Examiner, Art Unit 2659
/PIERRE LOUIS DESIR/Supervisory Patent Examiner, Art Unit 2659