Prosecution Insights
Last updated: August 17, 2026
Application No. 17/968,653

TRANSLATION MODEL WITH LEARNED POSITION AND CORRECTIVE LOSS

Final Rejection §101§102§103
Filed
Oct 18, 2022
Priority
Oct 20, 2021 — provisional 63/257,916
Examiner
KIM, SEHWAN
Art Unit
2100
Tech Center
2100 — Computer Architecture & Software
Assignee
The Toronto-dominion Bank
OA Round
2 (Final)
60%
Grant Probability
Moderate
3-4
OA Rounds
2m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 60% of resolved cases
60%
Career Allowance Rate
90 granted / 150 resolved
+5.0% vs TC avg
Strong +67% interview lift
Without
With
+66.7%
Interview Lift
resolved cases with interview
Typical timeline
4y 0m
Avg Prosecution
28 currently pending
Career history
185
Total Applications
across all art units

Statute-Specific Performance

§101
19.9%
-20.1% vs TC avg
§103
46.2%
+6.2% vs TC avg
§102
7.3%
-32.7% vs TC avg
§112
24.2%
-15.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 150 resolved cases

Office Action

§101 §102 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Examiner’s Note Providing supporting paragraph(s) for each limitation of amended/new claim(s) in Remarks is strongly requested for clear and definite claim interpretations by Examiner (e.g., to avoid rejections under 35 U.S.C § 112(a) “Lack of written description”) Applicant can schedule interviews (via Automated Interview Request (AIR)) at any stage of the prosecution (e.g., Non-Final, Final, and After-Final) to discuss any issues related to, for example, rejections under 35 U.S.C § 101 and § 102/103, for moving toward allowance. Priority Acknowledgment is made of applicant’s claim for domestic priority based on provisional application 63/257,916 filed on October 20, 2021. Response to Arguments Applicant's arguments filed on 02/27/2026 have been fully considered but they are not persuasive. In Remarks, regarding 35 USC § 101, Applicant contends: “The learned positional combination layer thus learns the particular parameters for effectively combining the token and the positional information. The resulting token-position encodings more effectively distinguish nearby or adjacent positions and may discourage repetition of the same token in the resulting output." Specification at 8” “That is, using the learned positional combination layer reduced the measurable similarity of positional information, providing a more discriminatory signal for the decoder to effectively generate tokens. In applications that use the positional information in parallel translation, "one and two token repetitions are reduced by over 30% and 35% respectively," providing a significant improvement to a problem in parallel translation in which tokens are repeated at adjacent positions in early iterations of applying the decoder”. “Similarly, the computer model architecture in this invention enables improved representation of positional information with tokens for language models that, in turn, improves effective output token generation, such as reducing repeating tokens in parallel translation models.”. Examiner’s response: The examiner understands the applicant’s assertion. However, it appears that each processing step is just applying the abstract idea to a general field of endeavor with additional elements. In addition, improvements to technology or technical field are not necessarily reflected in the claims. Thus, the claim does not integrate the judicial exception into a practical application, and the claim does not amount to significantly more than the judicial exception. The examiner understands the applicant’s assertion “The learned positional combination layer thus learns the particular parameters for effectively combining the token and the positional information. The resulting token-position encodings more effectively distinguish nearby or adjacent positions and may discourage repetition of the same token in the resulting output." Specification at 8” and “That is, using the learned positional combination layer reduced the measurable similarity of positional information, providing a more discriminatory signal for the decoder to effectively generate tokens. In applications that use the positional information in parallel translation, "one and two token repetitions are reduced by over 30% and 35% respectively," providing a significant improvement to a problem in parallel translation in which tokens are repeated at adjacent positions in early iterations of applying the decoder”. However, as rejected under Claim Rejections - 35 USC § 101, the claims are recited in a high-level, and it is not clear how the recited claims reflect the asserted improvements. In addition, note that increasing accuracy and/or reducing some token repetitions do not always provide improvements. Providing more details and/or explaining how the claims reflect the asserted technical improvements may help overcome the existing rejections. The examiner understands the applicant’s assertion “Similarly, the computer model architecture in this invention enables improved representation of positional information with tokens for language models that, in turn, improves effective output token generation, such as reducing repeating tokens in parallel translation models.” As mentioned in the Remarks, Desjardins showed improvements clearly by explaining how the machine learning model is trained to learn new tasks while protecting knowledge about previous tasks. However, it is not clear how the recited claims reflect the asserted improvements of effective output token generation. Providing more details and/or explaining how the claims reflect the asserted technical improvements may help overcome the existing rejections. For now, the limitations do not clearly show e.g., improvements in computer technology and improvements to other technical fields. Rather, the improvements in Remarks are about just improving the abstract ideas of the independent claims. It doesn’t seem that the specification and/or the independent claims clearly show how the inventive concept of the claims enables improvements and how they are tied together. The applicant may need to amend the claims to show how the claim languages and improvements are tied together. To find a valid improvement to a technology, MPEP 2106.04(d)(1) says the specification must explain the improvement and that the claim must reflect the disclosed improvement. Furthermore, the improvement should not be merely a consequence of the abstract idea. See MPEP 2106.05(a). An improvement in the abstract idea itself is not an improvement to technology. For at least these reasons, Applicant's arguments are not convincing. Applicant’s arguments regarding 35 USC § 102/103 with respect to the independent claims have been considered but are moot because the arguments are directed to amended limitation(s) that has/have not been previously examined. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-4, 6-12 and 14-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The analysis of the claims will follow the 2016 Revised Patent Subject Matter Eligibility Guidance, 84 Fed. Reg. 50 (“2019 PEG”). Claim 1 Step 1: The claim recites [a] system; therefore, it is directed to the statutory category of a machine. Step 2A Prong 1: The claim recites, inter alia: identifying an input sequence representation of an encoded sequence of input tokens: This limitation encompasses the mental process of identifying an encoded input sequence, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper. identifying an output estimate including a sequence of estimated output tokens: This limitation encompasses the mental process of identifying an estimated output sequence, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper. identifying a set of positional encodings corresponding to each position in the sequence of estimated output tokens: This limitation encompasses the mental process of identifying positional encodings for each position in an output sequence, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper. determining a sequence of output token-position encodings: This limitation encompasses the mental process of generating token-position encodings, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper. determining a sequence of output token probabilities: This limitation encompasses the mental process of generating a sequence of output probabilities, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper. Step 2A Prong 2: The abstract ideas listed above are not integrated into a practical application. Specifically, the additional elements, a processor that executes instructions; and a non-transitory computer-readable medium having instructions executable by the processor for, amount to invoking computers or other machinery merely as tools to perform an existing process. Thus, these additional elements represent no more than mere instructions to apply the abstract idea on a computer (see MPEP § 2106.05(f)). The additional elements, by applying a learned position combination layer position-wise to each estimated output token in the sequence of estimated output tokens with the corresponding positional encoding and by applying a decoder block to the sequence of output token-position encodings and the input sequence representation, amount to invoking computers or other machinery merely as tools to perform an existing process. Thus, these additional elements represent no more than mere instructions to apply the abstract idea on a computer (see MPEP § 2106.05(f)). Nothing in the claim integrates the abstract ideas into a practical application, and the claim is thus directed to the abstract ideas. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. The claim recites additional hardware and neural network elements that amount to invocations of computer elements as tools to perform existing processes. The additional elements listed above do not amount to significantly more than the abstract ideas. Therefore, the claim is subject-matter ineligible. Claim 2 Step 1: A machine, as above. Step 2A Prong 1: The claim inherits the abstract ideas of claim 1, from which it depends. Step 2A Prong 2: The abstract ideas inherited from claim 1 are not integrated into a practical application. Specifically, the additional element, wherein the learned position combination layer is a fully-connected layer, merely limits the architecture of the neural network layer outlined in claim 1. This feature does not affect the analysis presented in the rejection of claim 1 outlining the neural network as a computer element applying an abstract idea (see MPEP § 2106.05(f)). Nothing in the claim integrates the abstract ideas into a practical application, and the claim is thus directed to the abstract ideas. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. The claim merely recites an architectural feature of the neural network recited in claim 1. The additional element presented above does not amount to significantly more than the abstract ideas. Therefore, the claim is subject-matter ineligible. Claim 3 Step 1: A machine, as above. Step 2A Prong 1: The claim inherits the abstract ideas of claim 1, from which it depends. Step 2A Prong 2: The abstract ideas inherited from claim 1 are not integrated into a practical application. Specifically, the additional element, wherein the decoder block includes a full self-attention layer and a masked self-attention layer applied to the sequence of output token-position encodings, merely limits the architecture of the neural network layer outlined in claim 1. This feature does not affect the analysis presented in the rejection of claim 1 outlining the neural network as a computer element applying an abstract idea (see MPEP § 2106.05(f)). Nothing in the claim integrates the abstract ideas into a practical application, and the claim is thus directed to the abstract ideas. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. The claim merely recites an architectural feature of the neural network recited in claim 1. The additional element presented above does not amount to significantly more than the abstract ideas. Therefore, the claim is subject-matter ineligible. Claim 4 Step 1: A machine, as above. Step 2A Prong 1: The claim inherits the abstract ideas of claim 1, from which it depends. Step 2A Prong 2: The abstract ideas inherited from claim 1 via claim 3 are not integrated into a practical application. Specifically, the additional element, wherein the decoder block includes an attention layer for the input sequence representation after the masked self-attention layer, merely limits the architecture of the neural network layer outlined in claim 1. This feature does not affect the analysis presented in the rejection of claim 1 outlining the neural network as a computer element applying an abstract idea (see MPEP § 2106.05(f)). Nothing in the claim integrates the abstract ideas into a practical application, and the claim is thus directed to the abstract ideas. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. The claim merely recites an architectural feature of the neural network recited in claim 1. The additional element presented above does not amount to significantly more than the abstract ideas. Therefore, the claim is subject-matter ineligible. Claim 6 Step 1: A machine, as above. Step 2A Prong 1: The claim inherits the abstract ideas of claim 1, from which it depends. Step 2A Prong 2: The abstract idea inherited from claim 1 are not integrated into a practical application. Specifically, the additional element, wherein the instructions are further executable for, amounts to invoking computers or other machinery merely as a tool to perform an existing process. Specifically, that the method of training the decoder with a masked loss function takes the form of executable instructions amounts simply to applying the abstract idea on a computer (see MPEP § 2106.05(f)). The additional element, training parameters of the decoder block without distillation from another trained model, merely recites the idea of a solution or outcome (see MPEP § 2106.05(f)). Specifically, it is unclear from the limitation how the training of the parameters takes place, only that it is done without distillation. Nothing in the claim integrates the abstract ideas into a practical application, and the claim is thus directed to the abstract ideas. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. The claim merely applies the abstract idea using computer implemented-instructions. The additional element presented above does not amount to significantly more than the abstract ideas. Therefore, the claim is subject-matter ineligible. Claim 7 Step 1: A machine, as above. Step 2A Prong 1: The claim recites, inter alia: training parameters of the decoder block with a masked loss based on a masked output estimate and a corrective loss based on a predicted output of the model when the sequence of output tokens is masked: This limitation encompasses the mathematical concept of training a decoder with a masked loss function, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper. Step 2A Prong 2: The abstract idea presented above is not integrated into a practical application. Specifically, the additional element, wherein the instructions are further executable for, amounts to invoking computers or other machinery merely as a tool to perform an existing process. Specifically, that the method of training the decoder with a masked loss function takes the form of executable instructions amounts simply to applying the abstract idea on a computer (see MPEP § 2106.05(f)). Nothing in the claim integrates the abstract ideas into a practical application, and the claim is thus directed to the abstract ideas. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. The claim merely applies the abstract idea using computer implemented-instructions. The additional element presented above does not amount to significantly more than the abstract ideas. Therefore, the claim is subject-matter ineligible. Claim 8 Step 1: A machine, as above. Step 2A Prong 1: The claim inherits the abstract ideas of claim 1, from which it depends. Step 2A Prong 2: The abstract ideas inherited from claim 1 are not integrated into a practical application. Specifically, the additional element, wherein the input sequence representation is generated by an encoder that includes another learned position combination layer for a sequence of input tokens and another set of positional encodings for the sequence of input tokens, amounts to invoking computers or other machinery merely as a tool to perform an existing process. Thus, these additional elements represent no more than mere instructions to apply the abstract idea on a computer (see MPEP § 2106.05(f)). Nothing in the claim integrates the abstract ideas into a practical application, and the claim is thus directed to the abstract ideas. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. The claim recites additional neural network elements that amount to invocations of computer elements as tools to perform existing processes. The additional element presented above does not amount to significantly more than the abstract ideas. Therefore, the claim is subject-matter ineligible. Claim 9 Step 1: The claim recites [a] method; therefore, it is directed to the statutory category of a process. Step 2A Prong 1: The claim recites, inter alia: identifying an input sequence representation of an encoded sequence of input tokens: This limitation encompasses the mental process of identifying an encoded input sequence, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper. identifying an output estimate including a sequence of estimated output tokens: This limitation encompasses the mental process of identifying an estimated output sequence, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper. identifying a set of positional encodings corresponding to each position in the sequence of estimated output tokens: This limitation encompasses the mental process of identifying positional encodings for each position in an output sequence, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper. determining a sequence of output token-position encodings: This limitation encompasses the mental process of generating token-position encodings, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper. determining a sequence of output token probabilities: This limitation encompasses the mental process of generating a sequence of output probabilities, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper. Step 2A Prong 2: The abstract ideas listed above are not integrated into a practical application. Specifically, the additional elements, by applying a learned position combination layer position-wise to each estimated output token in the sequence of estimated output tokens with the corresponding positional encoding and by applying a decoder block to the sequence of output token-position encodings and the input sequence representation, amount to invoking computers or other machinery merely as tools to perform an existing process. Thus, these additional elements represent no more than mere instructions to apply the abstract idea on a computer (see MPEP § 2106.05(f)). Nothing in the claim integrates the abstract ideas into a practical application, and the claim is thus directed to the abstract ideas. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. The claim recites additional neural network elements that amount to invocations of computer elements as tools to perform existing processes. The additional elements listed above do not amount to significantly more than the abstract ideas. Therefore, the claim is subject-matter ineligible. Claim 10 Step 1: A process, as above. Step 2A Prong 1: The claim inherits the abstract ideas of claim 9, from which it depends. Step 2A Prong 2: The abstract ideas inherited from claim 9 are not integrated into a practical application. Specifically, the additional element, wherein the learned position combination layer is a fully-connected layer, merely limits the architecture of the neural network layer outlined in claim 9. This feature does not affect the analysis presented in the rejection of claim 9 outlining the neural network as a computer element applying an abstract idea (see MPEP § 2106.05(f)). Nothing in the claim integrates the abstract ideas into a practical application, and the claim is thus directed to the abstract ideas. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. The claim merely recites an architectural feature of the neural network recited in claim 9. The additional element presented above does not amount to significantly more than the abstract ideas. Therefore, the claim is subject-matter ineligible. Claim 11 Step 1: A process, as above. Step 2A Prong 1: The claim inherits the abstract ideas of claim 9, from which it depends. Step 2A Prong 2: The abstract ideas inherited from claim 9 are not integrated into a practical application. Specifically, the additional element, wherein the decoder block includes a full self-attention layer and a masked self-attention layer applied to the sequence of output token-position encodings, merely limits the architecture of the neural network layer outlined in claim 9. This feature does not affect the analysis presented in the rejection of claim 9 outlining the neural network as a computer element applying an abstract idea (see MPEP § 2106.05(f)). Nothing in the claim integrates the abstract ideas into a practical application, and the claim is thus directed to the abstract ideas. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. The claim merely recites an architectural feature of the neural network recited in claim 9. The additional element presented above does not amount to significantly more than the abstract ideas. Therefore, the claim is subject-matter ineligible. Claim 12 Step 1: A process, as above. Step 2A Prong 1: The claim inherits the abstract ideas of claim 9, from which it depends. Step 2A Prong 2: The abstract ideas inherited from claim 9 via claim 11 are not integrated into a practical application. Specifically, the additional element, wherein the decoder block includes an attention layer for the input sequence representation after the masked self-attention layer, merely limits the architecture of the neural network layer outlined in claim 9. This feature does not affect the analysis presented in the rejection of claim 9 outlining the neural network as a computer element applying an abstract idea (see MPEP § 2106.05(f)). Nothing in the claim integrates the abstract ideas into a practical application, and the claim is thus directed to the abstract ideas. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. The claim merely recites an architectural feature of the neural network recited in claim 9. The additional element presented above does not amount to significantly more than the abstract ideas. Therefore, the claim is subject-matter ineligible. Claim 14 Step 1: A process, as above. Step 2A Prong 1: The claim inherits the abstract ideas of claim 9, from which it depends. Step 2A Prong 2: The abstract idea inherited from claim 9 are not integrated into a practical application. Specifically, the additional element, wherein the instructions are further executable for, amounts to invoking computers or other machinery merely as a tool to perform an existing process. Specifically, that the method of training the decoder with a masked loss function takes the form of executable instructions amounts simply to applying the abstract idea on a computer (see MPEP § 2106.05(f)). The additional element, training parameters of the decoder block without distillation from another trained model, merely recites the idea of a solution or outcome (see MPEP § 2106.05(f)). Specifically, it is unclear from the limitation how the training of the parameters takes place, only that it is done without distillation. Nothing in the claim integrates the abstract ideas into a practical application, and the claim is thus directed to the abstract ideas. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. The claim merely applies the abstract idea using computer implemented-instructions. The additional element presented above does not amount to significantly more than the abstract ideas. Therefore, the claim is subject-matter ineligible. Claim 15 Step 1: A process, as above. Step 2A Prong 1: The claim recites, inter alia: training parameters of the decoder block with a masked loss based on a masked output estimate and a corrective loss based on a predicted output of the model when the sequence of output tokens is masked: This limitation encompasses the mathematical concept of training a decoder with a masked loss function, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper. Step 2A Prong 2: There are no additional elements in the claim that integrate the abstract idea into a practical application, and the claim is thus directed to the abstract idea. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. Therefore, the claim is subject- matter ineligible. Claim 16 Step 1: A process, as above. Step 2A Prong 1: The claim inherits the abstract ideas of claim 9, from which it depends. Step 2A Prong 2: The abstract ideas inherited from claim 9 are not integrated into a practical application. Specifically, the additional element, wherein the input sequence representation is generated by an encoder that includes another learned position combination layer for a sequence of input tokens and another set of positional encodings for the sequence of input tokens, amounts to invoking computers or other machinery merely as a tool to perform an existing process. Thus, these additional elements represent no more than mere instructions to apply the abstract idea on a computer (see MPEP § 2106.05(f)). Nothing in the claim integrates the abstract ideas into a practical application, and the claim is thus directed to the abstract ideas. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. The claim recites additional neural network elements that amount to invocations of computer elements as tools to perform existing processes. The additional element presented above does not amount to significantly more than the abstract ideas. Therefore, the claim is subject-matter ineligible. Claim 17 Step 1: The claim recites [a] non-transitory computer-readable storage medium; therefore, it is directed to the statutory category of an article of manufacture. Step 2A Prong 1: The claim recites, inter alia: identifying an input sequence representation of an encoded sequence of input tokens: This limitation encompasses the mental process of identifying an encoded input sequence, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper. identifying an output estimate including a sequence of estimated output tokens: This limitation encompasses the mental process of identifying an estimated output sequence, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper. identifying a set of positional encodings corresponding to each position in the sequence of estimated output tokens: This limitation encompasses the mental process of identifying positional encodings for each position in an output sequence, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper. determining a sequence of output token-position encodings: This limitation encompasses the mental process of generating token-position encodings, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper. determining a sequence of output token probabilities: This limitation encompasses the mental process of generating a sequence of output probabilities, which is an evaluation practically capable of being performed in the human mind with the assistance of pen and paper. Step 2A Prong 2: The abstract ideas listed above are not integrated into a practical application. Specifically, the additional element, the non-transitory computer-readable medium comprising instructions executable by a processor for, amounts to invoking computers or other machinery merely as a tool to perform an existing process. Thus, this additional element represents no more than mere instructions to apply the abstract idea on a computer (see MPEP § 2106.05(f)). The additional elements, by applying a learned position combination layer position-wise to each estimated output token in the sequence of estimated output tokens with the corresponding positional encoding and by applying a decoder block to the sequence of output token-position encodings and the input sequence representation, amount to invoking computers or other machinery merely as tools to perform an existing process. Thus, these additional elements represent no more than mere instructions to apply the abstract idea on a computer (see MPEP § 2106.05(f)). Nothing in the claim integrates the abstract ideas into a practical application, and the claim is thus directed to the abstract ideas. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. The claim recites additional hardware and neural network elements that amount to invocations of computer elements as tools to perform existing processes. The additional elements listed above do not amount to significantly more than the abstract ideas. Therefore, the claim is subject-matter ineligible. Claim 18 Step 1: An article of manufacture, as above. Step 2A Prong 1: The claim inherits the abstract ideas of claim 17, from which it depends. Step 2A Prong 2: The abstract ideas inherited from claim 17 are not integrated into a practical application. Specifically, the additional element, wherein the learned position combination layer is a fully-connected layer, merely limits the architecture of the neural network layer outlined in claim 1. This feature does not affect the analysis presented in the rejection of claim 1 outlining the neural network as a computer element applying an abstract idea (see MPEP § 2106.05(f)). Nothing in the claim integrates the abstract ideas into a practical application, and the claim is thus directed to the abstract ideas. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. The claim merely recites an architectural feature of the neural network recited in claim 17. The additional element presented above does not amount to significantly more than the abstract ideas. Therefore, the claim is subject-matter ineligible. Claim 19 Step 1: An article of manufacture, as above. Step 2A Prong 1: The claim inherits the abstract ideas of claim 17, from which it depends. Step 2A Prong 2: The abstract ideas inherited from claim 17 are not integrated into a practical application. Specifically, the additional element, wherein the decoder block includes a full self-attention layer and a masked self-attention layer applied to the sequence of output token-position encodings, merely limits the architecture of the neural network layer outlined in claim 1. This feature does not affect the analysis presented in the rejection of claim 1 outlining the neural network as a computer element applying an abstract idea (see MPEP § 2106.05(f)). Nothing in the claim integrates the abstract ideas into a practical application, and the claim is thus directed to the abstract ideas. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. The claim merely recites an architectural feature of the neural network recited in claim 17. The additional element presented above does not amount to significantly more than the abstract ideas. Therefore, the claim is subject-matter ineligible. Claim 20 Step 1: An article of manufacture, as above. Step 2A Prong 1: The claim inherits the abstract ideas of claim 17, from which it depends. Step 2A Prong 2: The abstract ideas inherited from claim 17 via claim 19 are not integrated into a practical application. Specifically, the additional element, wherein the decoder block includes an attention layer for the input sequence representation after the masked self-attention layer, merely limits the architecture of the neural network layer outlined in claim 17. This feature does not affect the analysis presented in the rejection of claim 17 outlining the neural network as a computer element applying an abstract idea (see MPEP § 2106.05(f)). Nothing in the claim integrates the abstract ideas into a practical application, and the claim is thus directed to the abstract ideas. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because when considered separately or in combination, they do not constitute an inventive concept. The claim merely recites an architectural feature of the neural network recited in claim 17. The additional element presented above does not amount to significantly more than the abstract ideas. Therefore, the claim is subject-matter ineligible. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 1, 3-6, 8, 9, 11-14, 16, 17, 19 and 20 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Chen et al. (US 2022/0237380 A1, hereinafter Chen). Regarding claim 1, Chen teaches [a] system comprising: a processor that executes instructions; and (Chen, [0111]; “CPU 502 and/or GPU 504 are processors that perform processing necessary to implement Transformer 100A according to the present embodiment.”) a non-transitory computer-readable medium having instructions executable by the processor for: (Chen, [0117]; “Although the optical recording medium such as optical disk 526 is illustrated as the exemplary non-transitory recording medium in FIG. 4, it is not limited thereto, and there can be used a semiconductor recording medium such as a flash memory, a magnetic recording medium such as a hard disk or a storage tape, or a magneto-optical recording medium such as an MO (magneto-optical disk).”) identifying an input sequence representation of an encoded sequence of input tokens; (Chen, Fig. 1; PNG media_image1.png 809 665 media_image1.png Greyscale Chen, [0040]; “An input token string generated by an input embedding layer 4,” wherein generating “an input token string” using “an input embedding layer” is equivalent to identifying an input sequence representation of an encoded sequence of input tokens.) identifying an output estimate including a sequence of estimated output tokens; (Chen, [0050]; “An output token string generated by an output embedding layer 14,” wherein generating “an output token string” using “an output embedding layer” is equivalent to identifying an output estimate including a sequence of estimated output tokens.) identifying a set of positional encodings corresponding to each position in the sequence of estimated output tokens; (Chen, [0052]; “Positional embedding layer 16 outputs a positional embedding, which is a value indicating a position at which the token is present in already output sequence 12,” wherein “a position at which the token is present” encompasses each position in the sequence of estimated output tokens.) determining a sequence of output token-position encodings by applying a learned position combination layer position-wise to each estimated output token in the sequence of estimated output tokens with the corresponding positional encoding; and (Chen, [0053]; “Adder 18 adds the positional embedding from positional embedding layer 16, to the token string from output embedding layer 14. As a result, adder 18 outputs an output token string (vector) obtained by adding, to the vector indicating the value of the token included in the sentence, the value (relative or absolute position in existing output sequence 12) indicating the position at which the token is present in the sentence,” wherein the resulting vector is a sequence of output token-position encodings. In addition, e.g., “Adder 18 adds the positional embedding from positional embedding layer 16, to the token string from output embedding layer 14” read(s) on “position-wise”. Furthermore, e.g., under a broadest reasonable interpretation (BRI), Adder 18 along with positional embedding layer 16 read(s) on “a learned position combination layer” since the whole model of fig 1 including the positional embedding layer 16 has been trained.) determining a sequence of output token probabilities by applying a decoder block to the sequence of output token-position encodings and the input sequence representation (Fig 1; As illustrated in the figure, a decoder block “40” produces a set of output token probabilities “64” from the sequence of output-token position encodings produced by adder “18” and the input sequence representation “2.” Chen, [0062]; “Output sequence 64 indicates a probability of a translation target sentence (target sentence) corresponding to input sequence 2 (source sentence).”). Regarding claim 3, Chen teaches [t]he system of claim 1 (and thus the rejection of claim 1 is incorporated). Chen further teaches wherein the decoder block includes a full self-attention layer and a masked self-attention layer applied to the sequence of output token-position encodings (Chen, [0054-55]; “Each of decoder blocks 40 includes a MMHA (Masked Multi-head Attention) layer 42, an MHA (Multi-head Attention) layer 46,” wherein a “MMHA (Masked Multi-head Attention) layer” encompasses a masked-self attention layer and “an MHA (Multi-head Attention) layer” encompasses a full self-attention layer.). Regarding claim 4, Chen teaches [t]he system of claim 3 (and thus the rejection of claim 3 is incorporated). Chen further teaches wherein the decoder block includes an attention layer for the input sequence representation after the masked self-attention layer (Chen, Fig. 1; As can be seen from the image, “MHA layer 46” corresponding to an attention layer is include[d] after the “MMHA layer 42” corresponding to the masked self-attention layer.). Regarding claim 5, Chen teaches [t]he system of claim 1 (and thus the rejection of claim 1 is incorporated). Chen further teaches wherein the decoder block estimates the sequence of output token probabilities in parallel (Chen, [0105]; “The original positional embeddings of the sentence are used to prevent the Transformer from recursively obtaining the word order dependency between the words. This ensures that the stacked SANs (Self-Attention Networks) learn the sentence representation completely in parallel,” wherein to “learn the sentence representation completely in parallel” is equivalent to estimat[ing] the sequence of output token probabilities in parallel.). Regarding claim 6, Chen teaches [t]he system of claim 1 (and thus the rejection of claim 1 is incorporated). Chen further teaches wherein the instructions are further executable for training parameters of the decoder block without distillation from another trained model (Chen, [0120-21]; “Training program 514 is executed by the processor (CPU 502 and/or GPU 504) to implement the training processing for determining parameter set 518. That is, training program 514 causes the computer to execute a training method for training Transformer 100A… Each of the parameters included in parameter set 518 is optimized by executing training program 514. Training data set 90 includes a combination of pieces of data as shown in FIG. 4,” thereby indicating that “Transformer 100A” and its decoder block are trained directly or without distillation from another trained model. More specifically, Chen does not disclose a separate transformer model whose knowledge is distilled to an updated transformer model. Therefore, the training is done without distillation.). Regarding claim 8, Chen teaches [t]he system of claim 1 (and thus the rejection of claim 1 is incorporated). Chen further teaches wherein the input sequence representation is generated by an encoder that includes another learned position combination layer for a sequence of input tokens and another set of positional encodings for the sequence of input tokens (Chen Fig. 1; As can be seen from the image, the encoder block “20” includes another “adder 8” along with “positional embedding layer 6” corresponding to the learned position combination layer and “positional embedding[s] 6” corresponding to another set of positional encodings for the sequence of input tokens. Chen, [0040]; “An input token string generated by an input embedding layer 4.” Chen, [0042-43]; “Positional embedding layer 6 outputs a positional embedding, which is a value indicating a position at which the token is present in input sequence 2. Adder 8 adds the positional embedding from positional embedding layer 6, to the sequence from input embedding layer 4. As a result, adder 8 outputs an input token string (vector) obtained by adding, to the vector indicating the value of the token (for example, word) included in the sentence, the value (relative or absolute position in input sequence 2) indicating the position at which the token is present in the sentence.”). Claims 9, 11-14 and 16 are method claims corresponding to the steps of claims 1, 3-6 and 8 and are therefore rejected for the same reasons. Claims 17, 19 and 20 are non-transitory computer-readable medium claims corresponding to the steps of claims 1, 3 and 4 and are therefore rejected for the same reasons. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 2, 10 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Chen in view of Yang et al. (US 2022/0129638 A1, hereinafter Yang). Regarding claim 2, Chen teaches [t]he system of claim 2 (and thus the rejection of claim 2 is incorporated). Chen does not explicitly teach wherein the learned position combination layer is a fully-connected layer. However, Yang, in the area of predicting semantic similarity using transformers, teaches this limitation (Yang, [0091]; “More particularly, the block encoding portion 454 can process the token embeddings 452 and the position embeddings 450 to obtain a sentence token representation for each sentence of the textual blocks 452. At least one of these sentence token representations can be processed using one or more dense layer(s) 458 to obtain textual block representations 460 for each of the textual blocks 452,” wherein “one or more dense layer(s)” encompasses a fully-connected layer.). Yang is analogous to the claimed invention as both are from the same field of endeavor, that is, applying transformer architectures to natural language processing tasks. Chen teaches a dedicated adder layer that combines token and positional embeddings but does not explicitly specify that said layer is fully-connected. Yang teaches this limitation. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to substitute the adder of Chen with the one or more dense layers of Yang. The motivation to do so is to utilize the benefits of fully-connected layers in evaluating positional locality which is an important consideration in assessing sentence-level context from textual embeddings (Yang, [0031]; “The first contextual block representation of the plurality of contextual block representations can be selected to represent the entire document. After selecting the first contextual block representation, a dense layer can be utilized to transform the first contextual block representation with L2 normalization.”). Claim 10 is a method claim corresponding to the steps of claim 2 and is therefore rejected for the same reasons. Claim 18 is a non-transitory computer-readable medium claim corresponding to the steps of claim 2 and is therefore rejected for the same reasons. Claims 7 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Chen in view of Zhang et al. (US 2024/0104352 A1, hereinafter Zhang). Regarding claim 7, Chen teaches [t]he system of claim 1 (and thus the rejection of claim 1 is incorporated). Chen further teaches wherein the instructions are further executable for training parameters of the decoder block (Chen, [0120-21]; “Training program 514 is executed by the processor (CPU 502 and/or GPU 504) to implement the training processing for determining parameter set 518. That is, training program 514 causes the computer to execute a training method for training Transformer 100A… Each of the parameters included in parameter set 518 is optimized by executing training program 514. Training data set 90 includes a combination of pieces of data as shown in FIG. 4,” wherein a “[t]raining program” for “determining [a] parameter set” is equivalent to training parameters of the decoder block.). Chen does not explicitly teach with a masked loss based on a masked output estimate and a corrective loss based on a predicted output of the model when the sequence of output tokens is masked. However, Zhang, in the area of masked modeling-based pre-training, teaches this limitation (Zhang, [0006]; “The method includes evaluating, by the computing system, a loss function comprising a contrastive loss term and a masked modeling term, wherein the contrastive loss term evaluates a contrastive pre-training output generated based on the first set of context vectors and the plurality of target quantized vectors, and wherein the masked modeling loss term evaluates a masked modeling pre-training output generated based on the second set of context vectors and the plurality of discretized identifiers,” wherein “a masked modeling loss term” corresponds to a masked loss based on a masked output estimate, and “a contrastive loss term” corresponds to a corrective loss. Zhang, [0040-41]; “Specifically, the contrastive loss term 38 can evaluate a contrastive pre-training output generated based on the first set of context vectors 34 and the plurality of target quantized vectors 28. In particular, in one example, for a context vector c t , corresponding to a masked time step (position) t, the model 14 (inclusive of supplemental pre-training prediction components) is asked to identify its true quantized vector q, from a set of K distractors { q ~ 1 , q ~ 2 , … , q ~ k } ,” thereby indicating that “the contrastive loss term” corresponding to a corrective loss [is] based on a predicted output of the model when the sequence of output tokens is masked.). Zhang is analogous to the claimed invention as both are from the same field of endeavor, that is, training transformer models on masked input sequences. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to train the decoder block of Chen with the combined loss function of Zhang. The motivation to do so is to construct a loss function that can train the model to simultaneously optimize predictive accuracy of individual masked positions and sentence-wide context (Zhang, [0040]; “Specifically, the contrastive loss term 38 can evaluate a contrastive pre-training output generated based on the first set of context vectors 34 and the plurality of target quantized vectors 28.” Zhang, [0049-50]; “The masked modeling loss term 40 can evaluate whether the predicted identifier corresponds to a true discretized identifier of the plurality of discretized identifiers 30 that corresponds to the masked position. The machine learning model 14 can be trained end-to-end based on the loss function that includes both the contrastive loss term 38 and the masked modeling term 40. Thus, the model 14 can be trained to solve the two self-supervised tasks at the same time.”). Claim 15 is a method claim corresponding to the steps of claim 7 and is therefore rejected for the same reasons. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Guo et al. (“Jointly Masked Sequence-to-Sequence Model for Non-Autoregressive Neural Machine Translation”) discloses a masked, non-autoregressive transformer model trained on separate masked and n-gram count loss functions (Cited by Applicant in IDS). Shazeer et al. (US 2020/0342316 A1) discloses a transformer architecture with a dedicated embedding layer that combines input token embeddings and position embeddings. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SEHWAN KIM whose telephone number is (571)270-7409. The examiner can normally be reached Mon - Thu 7:00 AM - 5:00 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michael J Huntley can be reached at (303) 297-4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SEHWAN KIM/Examiner, Art Unit 2129
Read full office action

Prosecution Timeline

Oct 18, 2022
Application Filed
Sep 03, 2025
Non-Final Rejection mailed — §101, §102, §103
Feb 27, 2026
Response Filed
Aug 05, 2026
Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12619853
DECISION-MAKING DEVICE, UNMANNED SYSTEM, DECISION-MAKING METHOD, AND PROGRAM
5y 6m to grant Granted May 05, 2026
Patent 12619921
PREDICTIVE FOG DATA CENTER MIGRATION
3y 8m to grant Granted May 05, 2026
Patent 12608592
AUTOMATED ELECTRIC SUBMERSIBLE PUMP (ESP) FAILURE ANALYSIS
3y 4m to grant Granted Apr 21, 2026
Patent 12602595
SYSTEM AND METHOD OF USING A KNOWLEDGE REPRESENTATION FOR FEATURES IN A MACHINE LEARNING CLASSIFIER
9y 4m to grant Granted Apr 14, 2026
Patent 12602580
Dataset Dependent Low Rank Decomposition Of Neural Networks
6y 9m to grant Granted Apr 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
60%
Grant Probability
99%
With Interview (+66.7%)
4y 0m (~2m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 150 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month