Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This action is responding to application papers dated 7/19/2024.
Claims 1-20 are pending in the application.
Claim Objections
Claims 3, 6, 7, and 20 are objected to because of the following informalities:
Per claim 3, it appears that “wherein train …in a first …transform” needs to be “wherein training … in the first …transforming.” Similarly, claims 6, 7, and 20 need to be corrected in the same manner.
Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Per claim 1:
It is unclear whether the model being trained in the second training stage refers to the model trained in the first stage. Interpretation: train the pre-trained neural network model, trained in the first training stage, in a second training stage. For the limitation “the neural network model is produced” on line 21, there is insufficient antecedent basis for the limitation in the claim. Interpretation: the pre-trained neutral network model is trained. On line 16, it is not clear to which snippet in the context of “back translation of the source code snippet in a fourth programming language” it is referring as there are multiple code snippets recited. Interpretation: the source code snippet in the third programming language is back translated into the fourth programming language.
Per claim 5, on the last line, “a second programming language” is interpreted as the second programming language recited in claim 1.
Per claim 9, It is unclear whether the model being trained in the second training stage refers to the model trained in the first stage. Interpretation: train the pre-trained neural network model, trained in the first training stage, in a second training stage. “neural network” on line 3 is interpreted as “neural network model.” On the last line, “the neural network model” is interpreted as “the pre-trained neural network model.”
Per claims 10, 11, and 13-15, “the neural network” is interpreted as “the pre-trained neural network model.”
Per claim 12, “creating the second fine-tuning training dataset comprising a plurality of second training samples” is interpreted as: creating the data-augmented training dataset comprising the plurality of data-augmented training samples. “the pre-trained neural network” is interpreted as: the pre-trained neural network model.
Per claim 16, It is unclear whether the model being trained in the second training stage refers to the model trained in the first stage. Interpretation: train the pre-trained neural network model, trained in the first training stage, in a second training stage. The limitation “neural network” on line 4 is interpreted as “neural network model.” “the neural network” is interpreted as “the pre-trained neural network model.” On the last line, “the neural network model” is interpreted as “the pre-trained neural network model.”
Per claims 18 and 19, “the neural network” is interpreted as “the pre-trained neural network model.”
Per claim 20, “the pre-trained neural network” is interpreted as “the pre-trained neural network model.”
Per claims 2-8, 10-15 and 17-20, these claims are rejected because they depend on claims 1, 9 and 16 respectively.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1, 3-16 and 20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Lachaux et al. (“Unsupervised Translation of Programming Languages,” hereafter Lachaux, 9/2020).
Per claim 1:
Lachaux teaches: A system for training a neural network model for code translation, comprising: a processor; and a memory that stores a program that is configured to be executed by the processor, the program comprising instructions to perform acts that: obtain a pre-trained neural network model trained on an unsupervised set of source code snippets; create a first fine-tuning training dataset comprising a plurality of first training samples, wherein a first training sample of the plurality of first training samples comprises a source code snippet in a first programming language and a known translation of the source code snippet in a second programming language, wherein the first programming language and the second programming language differ (Lachaux, see at least abstract, A transcompiler; page 2, translate functions from a programming language to another, that is purely based on monolingual source code…. TransCoder successfully manages to grasp complex patterns specific to each language, and to translate them to other languages; page 3, For TransCoder, we consider a sequence-to-sequence (seq2seq) model with attention … composed of an encoder and a decoder with a transformer architecture … subsequent work showed that pretraining the entire model (and not only word representations) in a cross-lingual way could lead to significant improvements in unsupervised machine translation … we follow the pretraining strategy of Lample and Conneau … where a Cross-lingual Language Model (XLM) is pretrained with a masked language modeling objective … on monolingual source code datasets; page 5, The first symbol given as input to the decoder is a special token indicating the output programming language. At test time, a Python sequence can be encoded by the model, and decoded using the C++ start symbol to generate a C++ translation; page 7, Figure 2: Example of unsupervised Python to C++ translation; page 5, In the unsupervised setting, a source-to-target model is coupled with a backward target-to-source model trained in parallel. The target-to-source model is used to translate target sequences into the source language, producing noisy source sequences corresponding to the ground truth target sequences. The source-to-target model is then trained in a weakly supervised manner to reconstruct the target sequences from the noisy source sequences generated by the target-to-source model, and vice versa. The two models are trained in parallel until convergence. An example of back-translation is illustrated in Figure 1 … 4.2 Training data; Note that the two phase feedback loop with forward translation and sample filtering (first fine-tuning of the data with filtering));
train the pre-trained neural network model in a first training stage using the first fine-tuning training dataset; create a second fine-tuning training dataset comprising a plurality of second training samples, wherein a second training sample comprises a source code snippet in a third programming language and a back translation of the source code snippet in a fourth programming language, wherein the third programming language and the fourth programming language differ; and train the pre-trained neural network model in a second training stage using the second fine-tuning training dataset (Lachaux, page 5, In the unsupervised setting, a source-to-target model is coupled with a backward target-to-source model trained in parallel. The target-to-source model is used to translate target sequences into the source language, producing noisy source sequences corresponding to the ground truth target sequences. The source-to-target model is then trained in a weakly supervised manner to reconstruct the target sequences from the noisy source sequences generated by the target-to-source model, and vice versa. The two models are trained in parallel until convergence. An example of back-translation is illustrated in Figure 1 … 4.2 Training data, We filter projects whose license explicitly permits the re-distribution of parts of the project, and select the C++, Java, and Python files within those projects. Ideally, a transcompiler should be able to translate whole projects. In this work, we decide to translate at function level. Unlike files or classes, functions are short enough to fit into a single batch, and working at function level allows for a simpler evaluation of the model with unit tests (c.f. Section 4.4). We pretrain TransCoder on all source code available, and train the denoising auto-encoding and back-translation objectives on functions only… 4.2 Training data; Note that the two phase feedback loop with forward translation and sample filtering (first fine-tuning of the data with filtering) and the second phase with back translation and sample filtering by translating the source language back into the target language where the new round of the back translation is tested against the original inputs; for the back translation of C++ (third language) to Python as an example, Python is the fourth language in a back translation and first language in a known translation).
where upon completion of the second training stage, the neural network model is produced to translate source code in one programming language into a different programming language (Lachaux, see at least page 2, we propose to apply recent approaches in unsupervised machine translation, by leveraging large amount of monolingual source code from GitHub to train a model, TransCoder, to translate between three popular languages: C++, Java and Python; page 4, For the masked language modeling (MLM) objective, at each iteration we consider an input stream of source code sequences, randomly mask out some of the tokens, and train TransCoder to predict the tokens that have been masked out based on their contexts. We alternate between streams of batches of different languages. This allows the model to create high quality, cross-lingual sequence representations. An example of XLM pretraining is given on top of Figure 1.; page 3, We describe now some of these methods and how they can be instantiated in the setting of unsupervised transcompilation …we follow the pretraining strategy of Lample and Conneau [29], where a Cross-lingual Language Model (XLM) is pretrained with a masked language modeling objective [14] on monolingual source code datasets; abstract, We train our model on source code from open source GitHub projects, and show that it can translate functions between C++, Java, and Python with high accuracy. Our method relies exclusively on monolingual source code, requires no expertise in the source or target languages, and can easily be generalized to other programming languages).
3. The system of claim 1, wherein train the pre-trained neural network model in a first training stage using the first fine-tuning training dataset further comprises: transform the first training sample into an input sequence comprising a plurality of tokens and associated token types (Lachaux, page 4, For the masked language modeling (MLM) objective, at each iteration we consider an input stream of source code sequences, randomly mask out some of the tokens, and train TransCoder to predict the tokens that have been masked out based on their contexts. We alternate between streams of batches of different languages. This allows the model to create high quality, cross-lingual sequence representations. An example of XLM pretraining is given on top of Figure 1 … XLM pretraining allows the seq2seq model to generate high quality representations of input sequences. However, the decoder lacks the capacity to translate, as it has never been trained to decode a sequence based on a source representation. To address this issue, we train the model to encode and decode sequences with a Denoising Auto-Encoding (DAE) objective [46]. The DAE objective operates like a supervised machine translation algorithm, where the model is trained to predict a sequence of tokens given a corrupted version of that sequence. To corrupt a sequence, we use the same noise model as the one described in Lample et al. [30]. Namely, we randomly mask, remove and shuffle input tokens; Python input function SumOfKsubArray into C++. TransCoder infers the types of the arguments, of the variables, and the return type of the function; Note that TransCoder processes multiple languages within a model injecting special language token types into the sequence).
4. The system of claim 3, wherein the program comprises instructions to perform acts that: train the pre-training neural network model in the first training stage with the input sequence for the pre-trained neural network model to learn to predict a token from vocabulary of the pre-trained neural network model or to predict copying a token from the input sequence (Lachaux, page 6, reduces the overall vocabulary size, and maximizes the token overlap between languages, improving the cross-linguality of the model … The BPE codes are learned with fastBPE8 on the concatenation of tokenized C++, Java, and Python files; page 4, For the masked language modeling (MLM) objective, at each iteration we consider an input stream of source code sequences, randomly mask out some of the tokens, and train TransCoder to predict the tokens that have been masked out based on their contexts. … where the model is trained to predict a sequence of tokens given a corrupted version of that sequence).
5. The system of claim 1, wherein train the pre-trained neural network model in a first training stage using the first fine-tuning training dataset further comprises: generate an input sequence to the pre-trained neural network model comprising the source code snippet in the first programming language followed by the known translation of the source code snippet in a second programming language (Lachaux, page 2, translate functions from a programming language to another, that is purely based on monolingual source code…. TransCoder successfully manages to grasp complex patterns specific to each language, and to translate them to other languages; page 3, For TransCoder, we consider a sequence-to-sequence (seq2seq) model with attention … composed of an encoder and a decoder with a transformer architecture … subsequent work showed that pretraining the entire model (and not only word representations) in a cross-lingual way could lead to significant improvements in unsupervised machine translation … page 4, For the masked language modeling (MLM) objective, at each iteration we consider an input stream of source code sequences, randomly mask out some of the tokens, and train TransCoder to predict the tokens that have been masked out based on their contexts; page 5, In the unsupervised setting, a source-to-target model is coupled with a backward target-to-source model trained in parallel. The target-to-source model is used to translate target sequences into the source language, producing noisy source sequences corresponding to the ground truth target sequences. The source-to-target model is then trained in a weakly supervised manner to reconstruct the target sequences from the noisy source sequences generated by the target-to-source model, and vice versa).
6. The system of claim 1, wherein train the pre-trained neural network model in a first training stage using the first fine-tuning training dataset further comprises: generate an input sequence to the pre-trained neural network model comprising the known translation of the source code snippet in the second programming language followed by the source code snippet in the first programming language (Lachaux, page 2, translate functions from a programming language to another, that is purely based on monolingual source code…. TransCoder successfully manages to grasp complex patterns specific to each language, and to translate them to other languages; page 3, For TransCoder, we consider a sequence-to-sequence (seq2seq) model with attention … composed of an encoder and a decoder with a transformer architecture … subsequent work showed that pretraining the entire model (and not only word representations) in a cross-lingual way could lead to significant improvements in unsupervised machine translation … page 4, For the masked language modeling (MLM) objective, at each iteration we consider an input stream of source code sequences, randomly mask out some of the tokens, and train TransCoder to predict the tokens that have been masked out based on their contexts).
7. The system of claim 1, wherein create a second fine-tuning training dataset comprising a plurality of second training samples further comprises: generate, by the pre-trained neural network model of the first training stage, the back translation of the source code snippet in the fourth programming language, wherein the pre-trained neural network model of the first training stage is given the source code snippet in the third programming language dataset (Lachaux, page 5, In the unsupervised setting, a source-to-target model is coupled with a backward target-to-source model trained in parallel. The target-to-source model is used to translate target sequences into the source language, producing noisy source sequences corresponding to the ground truth target sequences. The source-to-target model is then trained in a weakly supervised manner to reconstruct the target sequences from the noisy source sequences generated by the target-to-source model, and vice versa. The two models are trained in parallel until convergence. An example of back-translation is illustrated in Figure 1 … 4.2 Training data, We filter projects whose license explicitly permits the re-distribution of parts of the project, and select the C++, Java, and Python files within those projects. Ideally, a transcompiler should be able to translate whole projects. In this work, we decide to translate at function level. Unlike files or classes, functions are short enough to fit into a single batch, and working at function level allows for a simpler evaluation of the model with unit tests (c.f. Section 4.4). We pretrain TransCoder on all source code available, and train the denoising auto-encoding and back-translation objectives on functions only … 4.2 Training data; Note that the two phase feedback loop with forward translation and sample filtering (first fine-tuning of the data with filtering) and the second phase with back translation and sample filtering by translating the source language back into the target language where the new round of the back translation is tested against the original inputs).
8. The system of claim 1, wherein the neural network model comprises a neural transformer model with attention (Lachaux, page 3, For TransCoder, we consider a sequence-to-sequence (seq2seq) model with attention …composed of an encoder and a decoder with a transformer architecture).
Per claim 9:
Lachaux teaches: A computer-implemented method for training a neural network for code translation, comprising: obtaining a neural network pre-trained on an unsupervised training dataset of source code; creating a first fine-tuning training dataset comprising a plurality of first training samples, wherein a first training sample of the plurality of first training samples comprises a source code snippet in a first programming language and a known translation of the source code snippet in a second programming language, wherein the first programming language and the second programming language differ (Lachaux, see at least abstract, A transcompiler; page 2, translate functions from a programming language to another, that is purely based on monolingual source code…. TransCoder successfully manages to grasp complex patterns specific to each language, and to translate them to other languages; page 3, For TransCoder, we consider a sequence-to-sequence (seq2seq) model with attention … composed of an encoder and a decoder with a transformer architecture … subsequent work showed that pretraining the entire model (and not only word representations) in a cross-lingual way could lead to significant improvements in unsupervised machine translation … we follow the pretraining strategy of Lample and Conneau … where a Cross-lingual Language Model (XLM) is pretrained with a masked language modeling objective … on monolingual source code datasets; page 5, The first symbol given as input to the decoder is a special token indicating the output programming language. At test time, a Python sequence can be encoded by the model, and decoded using the C++ start symbol to generate a C++ translation; page 7, Figure 2: Example of unsupervised Python to C++ translation; page 5, In the unsupervised setting, a source-to-target model is coupled with a backward target-to-source model trained in parallel. The target-to-source model is used to translate target sequences into the source language, producing noisy source sequences corresponding to the ground truth target sequences. The source-to-target model is then trained in a weakly supervised manner to reconstruct the target sequences from the noisy source sequences generated by the target-to-source model, and vice versa. The two models are trained in parallel until convergence. An example of back-translation is illustrated in Figure 1 … 4.2 Training data; Note that the two phase feedback loop with forward translation and sample filtering (first fine-tuning of the data with filtering));
transforming each first training sample into a token sequence comprising a plurality of tokens, wherein each token comprises a corresponding token type; training the pre-trained neural network model in a first training stage with the first fine-tuning training dataset; creating a data-augmented training dataset comprising a plurality of data-augmented training samples, wherein a data-augmented training sample of the plurality of data-augmented training samples comprises a source code snippet in a third programming language and a back translation of the source code snippet in a fourth programming language, wherein the third programming language and the fourth programming language differ; (Lachaux, page 5, In the unsupervised setting, a source-to-target model is coupled with a backward target-to-source model trained in parallel. The target-to-source model is used to translate target sequences into the source language, producing noisy source sequences corresponding to the ground truth target sequences. The source-to-target model is then trained in a weakly supervised manner to reconstruct the target sequences from the noisy source sequences generated by the target-to-source model, and vice versa. The two models are trained in parallel until convergence. An example of back-translation is illustrated in Figure 1 … 4.2 Training data, We filter projects whose license explicitly permits the re-distribution of parts of the project, and select the C++, Java, and Python files within those projects. Ideally, a transcompiler should be able to translate whole projects. In this work, we decide to translate at function level. Unlike files or classes, functions are short enough to fit into a single batch, and working at function level allows for a simpler evaluation of the model with unit tests (c.f. Section 4.4). We pretrain TransCoder on all source code available, and train the denoising auto-encoding and back-translation objectives on functions only… 4.2 Training data; Note that the two phase feedback loop with forward translation and sample filtering (first fine-tuning of the data with filtering) and the second phase with back translation and sample filtering by translating the source language back into the target language where the new round of the back translation is tested against the original inputs; for the back translation of C++ (third language) to Python as an example, Python is the fourth language in a back translation and first language in a known translation).
transforming each data-augmented training sample into a token sequence a token sequence comprising a plurality of tokens, wherein each token comprises a corresponding token type; and training the pre-trained neural network model in a second training stage with the data-augmented training dataset, wherein upon completion of the second training stage, the neural network model is trained for code translation (Lachaux, see at least page 2, we propose to apply recent approaches in unsupervised machine translation, by leveraging large amount of monolingual source code from GitHub to train a model, TransCoder, to translate between three popular languages: C++, Java and Python; page 4, For the masked language modeling (MLM) objective, at each iteration we consider an input stream of source code sequences, randomly mask out some of the tokens, and train TransCoder to predict the tokens that have been masked out based on their contexts. We alternate between streams of batches of different languages. This allows the model to create high quality, cross-lingual sequence representations. An example of XLM pretraining is given on top of Figure 1.; page 3, We describe now some of these methods and how they can be instantiated in the setting of unsupervised transcompilation …we follow the pretraining strategy of Lample and Conneau [29], where a Cross-lingual Language Model (XLM) is pretrained with a masked language modeling objective [14] on monolingual source code datasets; abstract, We train our model on source code from open source GitHub projects, and show that it can translate functions between C++, Java, and Python with high accuracy. Our method relies exclusively on monolingual source code, requires no expertise in the source or target languages, and can easily be generalized to other programming languages; page 4, For the masked language modeling (MLM) objective, at each iteration we consider an input stream of source code sequences, randomly mask out some of the tokens, and train TransCoder to predict the tokens that have been masked out based on their contexts. We alternate between streams of batches of different languages. This allows the model to create high quality, cross-lingual sequence representations. An example of XLM pretraining is given on top of Figure 1 … XLM pretraining allows the seq2seq model to generate high quality representations of input sequences. However, the decoder lacks the capacity to translate, as it has never been trained to decode a sequence based on a source representation. To address this issue, we train the model to encode and decode sequences with a Denoising Auto-Encoding (DAE) objective [46]. The DAE objective operates like a supervised machine translation algorithm, where the model is trained to predict a sequence of tokens given a corrupted version of that sequence. To corrupt a sequence, we use the same noise model as the one described in Lample et al. [30]. Namely, we randomly mask, remove and shuffle input tokens; Python input function SumOfKsubArray into C++. TransCoder infers the types of the arguments, of the variables, and the return type of the function; Note that TransCoder processes multiple languages within a model injecting special language token types into the sequence).
10. The computer-implemented method of claim 9, wherein training the pre-trained neural network model in the first training stage with the first fine-tuning training dataset trains the neural network to learn to predict a token from a vocabulary of the neural network or from the token sequence (Lachaux, page 6, reduces the overall vocabulary size, and maximizes the token overlap between languages, improving the cross-linguality of the model … The BPE codes are learned with fastBPE8 on the concatenation of tokenized C++, Java, and Python files; page 4, For the masked language modeling (MLM) objective, at each iteration we consider an input stream of source code sequences, randomly mask out some of the tokens, and train TransCoder to predict the tokens that have been masked out based on their contexts. … where the model is trained to predict a sequence of tokens given a corrupted version of that sequence).
11. The computer-implemented method of claim 9, wherein training the pre-trained neural network model in the second training stage with the data-augmented training dataset trains the neural network to learn to predict a token from a vocabulary of the neural network or from the token sequence (Lachaux, page 6, reduces the overall vocabulary size, and maximizes the token overlap between languages, improving the cross-linguality of the model … The BPE codes are learned with fastBPE8 on the concatenation of tokenized C++, Java, and Python files; page 4, For the masked language modeling (MLM) objective, at each iteration we consider an input stream of source code sequences, randomly mask out some of the tokens, and train TransCoder to predict the tokens that have been masked out based on their contexts. … where the model is trained to predict a sequence of tokens given a corrupted version of that sequence).
12. The computer-implemented method of claim 9, wherein creating the second fine-tuning training dataset comprising a plurality of second training samples further comprises: generating, by the pre-trained neural network of the first training stage, the back translation of the source code snippet in the fourth programming language, wherein the pre-trained neural network of the first training stage is given the source code snippet in the third programming language and generates the back translation of the source code snippet in the fourth programming language (Lachaux, page 5, In the unsupervised setting, a source-to-target model is coupled with a backward target-to-source model trained in parallel. The target-to-source model is used to translate target sequences into the source language, producing noisy source sequences corresponding to the ground truth target sequences. The source-to-target model is then trained in a weakly supervised manner to reconstruct the target sequences from the noisy source sequences generated by the target-to-source model, and vice versa. The two models are trained in parallel until convergence. An example of back-translation is illustrated in Figure 1 … 4.2 Training data, We filter projects whose license explicitly permits the re-distribution of parts of the project, and select the C++, Java, and Python files within those projects. Ideally, a transcompiler should be able to translate whole projects. In this work, we decide to translate at function level. Unlike files or classes, functions are short enough to fit into a single batch, and working at function level allows for a simpler evaluation of the model with unit tests (c.f. Section 4.4). We pretrain TransCoder on all source code available, and train the denoising auto-encoding and back-translation objectives on functions only … 4.2 Training data; Note that the two phase feedback loop with forward translation and sample filtering (first fine-tuning of the data with filtering) and the second phase with back translation and sample filtering by translating the source language back into the target language where the new round of the back translation is tested against the original inputs).
13. The computer-implemented method of claim 9, wherein the neural network is a neural transformer model with attention (Lachaux, page 3, For TransCoder, we consider a sequence-to-sequence (seq2seq) model with attention …composed of an encoder and a decoder with a transformer architecture).
14. The computer-implemented method of claim 13, wherein the neural network comprises at least one encoder block coupled to at least one decoder block (Lachaux, page 3, For TransCoder, we consider a sequence-to-sequence (seq2seq) model with attention …composed of an encoder and a decoder with a transformer architecture).
15. The computer-implemented method of claim 14, wherein the at least one encoder block comprises a token-type head that produces a token type for each token of a training sample, wherein the at least one decoder block comprises a copy segment prediction head that outputs an output probability indicating whether to select a token from vocabulary of the neural network or select a token from the training sample (Lachaux, page 5, We use a transformer with 6 layers, 8 attention heads, and set the dimensionality of the model to 1024. We use a single encoder and a single decoder for all programming languages. … auto-encoding and back-translation objectives; page 4, The first symbol given as input to the decoder is a special token indicating the output programming language. At test time, a Python sequence can be encoded by the model, and decoded using the C++ start symbol to generate a C++ translation. The quality of the C++ translation will depend on the “cross-linguality” of the model: if the Python function and a valid C++ translation are mapped to the same latent representation by the encoder, the decoder will successfully generate this C++ translation).
Per claims 16 and 20, these claims are the hardware storage device versions of claims 9, 12, and 14, and are rejected for the same reasons set forth in connection with the rejection of claims 9, 12 and 14 above.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 2 is rejected under 35 U.S.C. 103 as being unpatentable over Lachaux In view of Singh et al. (US20220206785, hereafter Singh).
Per claim 2:
Lachaux further teaches: The system of claim 1, wherein the pre-trained neural network model is jointly pre-trained with a masked language model objective and an objective on an unsupervised source code corpus (Lachaux, see at least Page 4, we train the model to encode and decode sequences with a Denoising Auto-Encoding (DAE) objective [46]. The DAE objective operates like a supervised machine translation algorithm, where the model is trained to predict a sequence of tokens given a corrupted version of that sequence. To corrupt a sequence, we use the same noise model as the one described in Lample et al. [30]. Namely, we randomly mask, remove and shuffle input tokens; page 5, we use back-translation, which is one of the most effective methods to leverage monolingual data in a weakly-supervised scenario. Initially introduced to improve the performance of machine translation in the supervised setting [41]; page 2, we propose to apply recent approaches in unsupervised machine translation, by leveraging large amount of monolingual source code from GitHub to train a model, TransCoder, to translate between three popular languages: C++, Java and Python; page 3, section 3 Model, We train it using the three principles of unsupervised machine translation identified in Lample et al. [32], namely initialization, language modeling, and back-translation; page 3, section 3.1, Pretraining is a key ingredient of unsupervised machine translation …Cross-lingual Language Model (XLM) is pretrained with a masked language modeling objective [14] on monolingual source code datasets).
Lachaux does not explicitly teach jointly pre-training with an autoregressive objective. However, such an autoregressive objective is known in the industry. Singh teaches an autoregressive language model to improve code migration (Singh, see at least [0005] In various implementations, the machine learning model may be an autoregressive language model (e.g., trained to perform natural language processing, or “NLP”) such as a bidirectional encoder representations from transformers (BERT)-based model (also referred to herein as a “transformer” model). By conditioning such an autoregressive language model with demonstration(s), the autoregressive language model is effectively “primed” to perform a task established by the demonstration(s), e.g., by being more likely to select output candidates that are aligned with the demonstrated task; [0007], Training the autoregressive language model specifically using computer-programming-related corpuses enables the model, upon conditioning with demonstrations pertinent to a source code migration, to more accurately generate post-migration source code; [0031]). It would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to have combined Singh’s autoregressive objective with Lachaux’s code translation to modify Lachaux’s system to combine the autoregressive objective as taught by Singh, with a reasonable expectation of success, since they are analogous art because they are from the same field of endeavor related to machine training. Combining Singh’s functionality with that of Lachaux results in a system that allows incorporating an autoregressive objective for the trainings. The modification would be obvious because one having ordinary skill in the art would be motivated to make this combination to training the model in an autoregressive manner and “more accurately generate post-migration source code (Singh, see at least [0005] the machine learning model may be an autoregressive language model (e.g., trained to perform natural language processing, or “NLP”) such as a bidirectional encoder representations from transformers (BERT)-based model (also referred to herein as a “transformer” model). By conditioning such an autoregressive language model with demonstration(s), the autoregressive language model is effectively “primed” to perform a task established by the demonstration(s), e.g., by being more likely to select output candidates that are aligned with the demonstrated task; [0007], Training the autoregressive language model specifically using computer-programming-related corpuses enables the model, upon conditioning with demonstrations pertinent to a source code migration, to more accurately generate post-migration source code; [0031]).”
Claims 17-19 are rejected under 35 U.S.C. 103 as being unpatentable over Lachaux In view of See et al. (“Get To The Point: Summarization with Pointer-Generator Networks,” 2017).
Per claim 17:
Lachaux further teaches wherein the at least one encoder block comprises a token-type head (Lachaux, page 5, We use a transformer with 6 layers, 8 attention heads, and set the dimensionality of the model to 1024. We use a single encoder and a single decoder for all programming languages. During XLM pretraining, we alternate between batches of C++, Java, and Python, composed of 32 sequences of source code of 512 tokens. At training time, we alternate between the denoising auto-encoding and back-translation objectives, and use batches of around 6000 tokens. … the decoder is always trained to generate a valid function, even when the encoder output is noisy).
Lachaux does not explicitly teach wherein the at least one decoder block comprises a pointer-generator network. However, See teaches such a pointer-generator network (See, see, see at least abstract, we use a hybrid pointer-generator network that can copy words from the source text via pointing, which aids accurate repro duction of information, while retaining the ability to produce novel words through the generator). It would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to have combined See’s pointer-generator network with Lachaux’s code translation to modify Lachaux’s system to combine the network as taught by See, with a reasonable expectation of success, since they are analogous art because they are from the same field of endeavor related to machine training. Combining See’s functionality with that of Lachaux results in a system that allows the pointer-generator network to be used. The modification would be obvious because one having ordinary skill in the art would be motivated to make this combination to aid “accurate reproduction of information, while retaining the ability to produce novel words through the generator (See, see at least abstract, we use a hybrid pointer-generator network that can copy words from the source text via pointing, which aids accurate reproduction of information, while retaining the ability to produce novel words through the generator).”
18. The hardware storage device of claim 17, wherein train the pre-trained neural network model in the first training stage using the first fine-tuning training dataset trains the pointer-generator network to learn to predict a token from a vocabulary of the neural network or from the token sequence (Lachaux, page 6, reduces the overall vocabulary size, and maximizes the token overlap between languages, improving the cross-linguality of the model … The BPE codes are learned with fastBPE8 on the concatenation of tokenized C++, Java, and Python files; page 4, For the masked language modeling (MLM) objective, at each iteration we consider an input stream of source code sequences, randomly mask out some of the tokens, and train TransCoder to predict the tokens that have been masked out based on their contexts. … where the model is trained to predict a sequence of tokens given a corrupted version of that sequence).
19. The hardware storage device of claim 17, wherein train the pre-trained neural network model in the second training stage using the data-augmented training dataset trains the pointer-generator network to learn to predict a token from a vocabulary of the neural network or from the token sequence. (Lachaux, page 6, reduces the overall vocabulary size, and maximizes the token overlap between languages, improving the cross-linguality of the model … The BPE codes are learned with fastBPE8 on the concatenation of tokenized C++, Java, and Python files; page 4, For the masked language modeling (MLM) objective, at each iteration we consider an input stream of source code sequences, randomly mask out some of the tokens, and train TransCoder to predict the tokens that have been masked out based on their contexts. … where the model is trained to predict a sequence of tokens given a corrupted version of that sequence).
Examiner’s Note
The Examiner has pointed out particular references contained in the prior art of record within the body of this action for the convenience of the Applicant. Although the specified citations are representative of the teachings in the art and are applied to the specific limitations within the individual claim, other passages and figures may apply. Applicant, in preparing the response, should consider fully the entire reference as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior art or disclosed by the Examiner.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
CN 110956045 is related to multi-language translation;
Guo et al. is related to Jointly Masked Sequence-to-Sequence Model for Non-Autoregressive Neural Machine Translation.
US20210027025 is related to utilizing an autoregressive transformer using multiple layers of masked multi-head self-attention to map a sequence of input tokens to a sequence of output tokens;
US20200034436 is related to machine translation from one language to another language.
US20220108688 is related to a multilingual language model mBERT converted into an autoregressive transformer decoder.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to INSUN KANG whose telephone number is (571)272-3724. The examiner can normally be reached M-TR 9am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chat Do can be reached at 571-272-3721. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/INSUN KANG/ Primary Examiner, Art Unit 2193