CTNF 18/240,954 CTNF 100952 Notice of Pre-AIA or AIA Status This Non-Final communication is in response to application No. 18/240,954 filed on 8/31/2023 which claims priority to 63/402,706 filed on 8/32/2022. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA. Claim Rejections - 35 USC § 101 07-04-01 AIA 07-04 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e., a law of nature, a natural phenomenon, or an abstract idea) without significantly more. 101 Subject Matter Eligibility Analysis Step 1: Claims 1-20 are within the four statutory (a process, machine, manufacture or composition of matter.) Claims 1-11 describe a process and 12-20 describes a machine. With respect to claim 1: Step 2A Prong 1: The claim recites an abstract idea enumerated in the 2019 PEG. determining, by the computing system, a first reward value for the first candidate output sequence according to a reward function; (This is an abstract idea of a "Mental Process." The " determining " step under its broadest reasonable interpretation, covers concepts that can be practically performed in the human mind. The determination could be made manually by an individual.) determining, by the computing system, a second reward value for the corrected candidate output sequence according to the reward function; (This is an abstract idea of a "Mental Process." The " determining " step under its broadest reasonable interpretation, covers concepts that can be practically performed in the human mind. The determination could be made manually by an individual.) evaluating, by the computing system, a loss function that contrasts the first candidate output sequence with the corrected output sequence based on the first reward value and the second reward value; and (This is an abstract idea of a "Mental Process." The " evaluating " step under its broadest reasonable interpretation, covers concepts that can be practically performed in the human mind. The evaluation could be made manually by an individual.) Step 2A Prong 2 : The judicial exception is not integrated into a practical application Additional elements: processing, by the computing system, a training input sequence with a machine- learned sequential labeling model to generate one or more candidate output sequences; (this limitation amounts to adding insignificant extra-solution activity to the judicial exception). applying, by the computing system, a correction function to at least a first candidate output sequence of the one or more candidate output sequences to generate a corrected output sequence; (This amounts to no more than mere instructions to “apply” the exception using a generic computer component.) modifying, by the computing system, one or more parameter values of the machine-learned sequential labeling model based at least in part on the loss function. (this limitation amounts to adding insignificant extra-solution activity to the judicial exception). Step 2B: the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception The additional elements “processing…” and “modifying…” add insignificant extra-solution activity to the judicial exception and cannot provide an inventive concept. Storing and retrieving information in memory is directed to a well understood routine conventional activity of data transmission (MPEP 2106.05(d)(II)(iv)). The additional element “applying…” is recited in a generic level and they represent generic computer components to apply the abstract idea. Mere instructions to apply an exception cannot provide an inventive concept (MPEP 2106.05(f)). When considered in combination, these additional elements represent insignificant extra-solution activity and mere instructions to apply an expectation, which do not provide an inventive concept. Therefore, claim 1 is ineligible. With respect to claim 2: Step 2A Prong 1: claim 2, which incorporates the rejection of claim 1, recites an additional abstract idea: the method further comprises selecting, by the computing system, the first candidate output sequence from the plurality of candidate output sequences according to one or more criteria. (This is an abstract idea of a "Mental Process." The " selecting " step under its broadest reasonable interpretation, covers concepts that can be practically performed in the human mind. The selection could be made manually by an individual.) Step 2A Prong 2: The judicial exception is not integrated into a practical application. processing, by the computing system, the training input sequence with the machine- learned sequential labeling model to generate the one or more candidate output sequences comprises processing, by the computing system, the training input sequence with the machine- learned sequential labeling model to generate a plurality of candidate output sequences; (This amounts to no more than mere instructions to “apply” the exception using a generic computer component.) Step 2B: the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception The additional element is recited in a generic level and they represent generic computer components to apply the abstract idea. Mere instructions to apply an exception cannot provide an inventive concept (MPEP 2106.05(f)). Therefore, claim 2 is ineligible. With respect to claim 3: Step 2A Prong 1: claim 3, which incorporates the rejection of claim 2, does not recite an abstract idea. Step 2A Prong 2: The judicial exception is not integrated into a practical application. processing, by the computing system, the training input sequence with the machine-learned sequential labeling model to generate the plurality of candidate output sequences comprises performing a beam search over an output space of the machine-learned sequential labeling model to generate the plurality of candidate output sequences. (This amounts to no more than mere instructions to “apply” the exception using a generic computer component.) Step 2B: the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception The additional element is recited in a generic level and they represent generic computer components to apply the abstract idea. Mere instructions to apply an exception cannot provide an inventive concept (MPEP 2106.05(f)). Therefore, claim 3 is ineligible. With respect to claim 4: Step 2A Prong 1: claim 4, which incorporates the rejection of claim 2, recites an additional abstract idea: determining, according to the reward function, a respective reward value for each of the plurality of candidate output sequences; and (This is an abstract idea of a "Mental Process." The " determining " step under its broadest reasonable interpretation, covers concepts that can be practically performed in the human mind. The determination could be made manually by an individual.) selecting as the first candidate output sequence the candidate output sequence that receives the largest respective reward value. (This is an abstract idea of a "Mental Process." The " selecting " step under its broadest reasonable interpretation, covers concepts that can be practically performed in the human mind. The selection could be made manually by an individual.) Step 2A Prong 2: claim 4 does not recite any additional elements and thus cannot be integrated into a practical application. Step 2B: claim 4 does not recite an additional element. Therefore, claim 4 is ineligible. With respect to claim 5: Step 2A Prong 1: claim 5, which incorporates the rejection of claim 1, does not recite an abstract idea. Step 2A Prong 2: The judicial exception is not integrated into a practical application. the correction function comprises a pre-defined function that, when applied to a function input, with positive probability relative to a training dataset, results in a function output that receives a larger reward value than the function input according to the reward function. (This amounts to no more than mere instructions to “apply” the exception using a generic computer component.) Step 2B: the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception The additional element is recited in a generic level and they represent generic computer components to apply the abstract idea. Mere instructions to apply an exception cannot provide an inventive concept (MPEP 2106.05(f)). Therefore, claim 5 is ineligible. With respect to claim 6: Step 2A Prong 1: claim 6, which incorporates the rejection of claim 1, does not recite an abstract idea. Step 2A Prong 2: The judicial exception is not integrated into a practical application. drops any predicted spans in the first candidate output sequence that do not match with spans in a ground truth output sequence associated with the training input sequence; and (This amounts to no more than mere instructions to “apply” the exception using a generic computer component.) replaces one or more incorrect tags in the first candidate output sequence with one or more ground truth tags included in the ground truth output sequence associated with the training input sequence. (This amounts to no more than mere instructions to “apply” the exception using a generic computer component.) Step 2B: the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception The additional element is recited in a generic level and they represent generic computer components to apply the abstract idea. Mere instructions to apply an exception cannot provide an inventive concept (MPEP 2106.05(f)). Therefore, claim 6 is ineligible. With respect to claim 7: Step 2A Prong 1: claim 7, which incorporates the rejection of claim 1, recites an additional abstract idea: the reward function evaluates both precision and recall. (this is an abstract idea of a “mathematical concept”. The recited “reward function” represents a mathematical function that would fall under the “mathematical concepts” grouping.) Step 2A Prong 2: claim 7 does not recite any additional elements and thus cannot be integrated into a practical application. Step 2B: claim 7 does not recite an additional element. Therefore, claim 7 is ineligible. With respect to claim 8: Step 2A Prong 1: claim 8, which incorporates the rejection of claim 1, recites an additional abstract idea: the loss function comprises a corrective margin loss term that contrasts the first candidate output sequence with the corrected output sequence based on the first reward value and the second reward value, wherein the corrective margin loss term provides a loss value that is equal to a maximum of zero or a margin loss value, the margin loss value equal to a margin value minus a log likelihood of the corrected output sequence plus a log likelihood of the first candidate output sequence. (this is an abstract idea of a “mathematical concept”. The recited “loss function” represents a mathematical function that would fall under the “mathematical concepts” grouping.) Step 2A Prong 2: claim 8 does not recite any additional elements and thus cannot be integrated into a practical application. Step 2B: claim 8 does not recite an additional element. Therefore, claim 8 is ineligible. With respect to claim 9: Step 2A Prong 1: claim 9, which incorporates the rejection of claim 8, recites an additional abstract idea: the margin value is equal to a scaling coefficient times a difference between the second reward value and the first reward value. (this is an abstract idea of a “mathematical concept”. The recited “margin value” represents a mathematical calculation that would fall under the “mathematical concepts” grouping.) Step 2A Prong 2: claim 9 does not recite any additional elements and thus cannot be integrated into a practical application. Step 2B: claim 9 does not recite an additional element. Therefore, claim 9 is ineligible. With respect to claim 10: Step 2A Prong 1: claim 10, which incorporates the rejection of claim 1, recites an additional abstract idea: determining, by the computing system, a third reward value for a ground truth output sequence according to the reward function; (This is an abstract idea of a "Mental Process." The " determining " step under its broadest reasonable interpretation, covers concepts that can be practically performed in the human mind. The determination could be made manually by an individual.) wherein loss function further contrasts the corrected output sequence with the ground truth output sequence based on the second reward value and the third reward value. (this is an abstract idea of a “mathematical concept”. The recited “loss function” represents a mathematical function that would fall under the “mathematical concepts” grouping.) Step 2A Prong 2: claim 10 does not recite any additional elements and thus cannot be integrated into a practical application. Step 2B: claim 10 does not recite an additional element. Therefore, claim 10 is ineligible. With respect to claim 11: Step 2A Prong 1: claim 11, which incorporates the rejection of claim 1, does not recite an abstract idea. Step 2A Prong 2: The judicial exception is not integrated into a practical application. the training input sequence comprises a sequence of natural language tokens; and (this limitation amounts to adding insignificant extra-solution activity to the judicial exception). the first candidate output sequence comprises: a translation of the sequence of natural language tokens; one or more part of speech classifications of the sequence of natural language tokens; one or more chunking classifications of the sequence of natural language tokens; one or more named entity recognitions for the sequence of natural language tokens; one or more segmentations for the sequence of natural language tokens; one or more named entity recognitions for the sequence of natural language tokens; one or more natural language answers to the sequence of natural language tokens; or one or more dialog responses for the sequence of natural language tokens. (this limitation amounts to adding insignificant extra-solution activity to the judicial exception). Step 2B: the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception The additional elements add insignificant extra-solution activity to the judicial exception and cannot provide an inventive concept. Storing and retrieving information in memory is directed to a well understood routine conventional activity of data transmission (MPEP 2106.05(d)(II)(iv)). Therefore, claim 11 is ineligible. With respect to claim 12: The claim recites similar limitations as corresponding to claim 1. Therefore, the same subject matter analysis that was utilized for claim 1, as described above, is equally applicable to claim 12. Therefore, claim 12 is ineligible. With respect to claim 13: The claim recites similar limitations as corresponding to claim 2. Therefore, the same subject matter analysis that was utilized for claim 2, as described above, is equally applicable to claim 13. Therefore, claim 13 is ineligible. With respect to claim 14: The claim recites similar limitations as corresponding to claim 3. Therefore, the same subject matter analysis that was utilized for claim 3, as described above, is equally applicable to claim 14. Therefore, claim 14 is ineligible. With respect to claim 15: The claim recites similar limitations as corresponding to claim 4. Therefore, the same subject matter analysis that was utilized for claim 4, as described above, is equally applicable to claim 15. Therefore, claim 15 is ineligible. With respect to claim 16: The claim recites similar limitations as corresponding to claim 5. Therefore, the same subject matter analysis that was utilized for claim 5, as described above, is equally applicable to claim 16. Therefore, claim 16 is ineligible. With respect to claim 17: The claim recites similar limitations as corresponding to claim 6. Therefore, the same subject matter analysis that was utilized for claim 6, as described above, is equally applicable to claim 17. Therefore, claim 17 is ineligible. With respect to claim 18: The claim recites similar limitations as corresponding to claim 7. Therefore, the same subject matter analysis that was utilized for claim 7, as described above, is equally applicable to claim 18. Therefore, claim 18 is ineligible. With respect to claim 19: The claim recites similar limitations as corresponding to claim 8. Therefore, the same subject matter analysis that was utilized for claim 8, as described above, is equally applicable to claim 19. Therefore, claim 19 is ineligible. With respect to claim 20: The claim recites similar limitations as corresponding to claim 1. Therefore, the same subject matter analysis that was utilized for claim 1, as described above, is equally applicable to claim 20. Therefore, claim 20 is ineligible. Claim Rejections - 35 USC § 103 07-06 AIA 15-10-15 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. 07-20-aia AIA The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 07-21-aia AIA Claim s 1-9, 11-20 are rejected under 35 U.S.C. 103 as being unpatentable over Shu (NPL ‘Reward Optimization for Neural Machine translation with learned Metrics’ (from applicants IDS)) in view of Gui (NPL: ‘Uncertainty-Aware Label Refinement for Sequence Labeling’ (2020)) . Regarding claim 1, Shu teaches: A computer-implemented method to perform corrective reward optimization for sequential labeling, the method comprising: (Page 4 Section 3 Contrastive-Margin Loss for Reward Optimization “As the policy gradient method optimizes the expected reward without the candidate set constraint, it may diverge from best-response reward maximization.”) processing, by the computing system, a training input sequence with a machine- learned sequential labeling model to generate one or more candidate output sequences; (Page 2 Section 2 “Given a source-target sequence pair (X,Y ∗ ) from the training set, let S(X,θ,N) be a set of N candidate sequences generated by model pθ(Y |X), which is reachable by a certain decoding strategy (e.g. beam search).”)) determining, by the computing system, a first reward value for the first candidate output sequence according to a reward function; (Page 3 Expected Reward Maximization “The first category of solutions approximates the best-response reward using the expected reward.”) determining, by the computing system, a second reward value for the corrected candidate output sequence according to the reward function; (Shu does not teach corrected candidate sequence but does teach calculating a second reward as outlined in Ranking Optimization on page 3) evaluating, by the computing system, a loss function that contrasts the first candidate output sequence with the corrected output sequence based on the first reward value and the second reward value; and (Page 4 Section 3 Contrastive-Margin Loss for Reward Optimization “In this paper, we propose to use a more conservative variation of the max-margin loss, which we refer to as contrastive-margin loss.” And equation 9 shows the loss equation.) modifying, by the computing system, one or more parameter values of the machine-learned sequential labeling model based at least in part on the loss function. (Page 5 Training and Hyperparameters “We start with the baseline model parameters and use the contrastive-margin loss to fine tune the parameters.”) Shu does not teach: applying, by the computing system, a correction function to at least a first candidate output sequence of the one or more candidate output sequences to generate a corrected output sequence; However, Gui does: applying, by the computing system, a correction function to at least a first candidate output sequence of the one or more candidate output sequences to generate a corrected output sequence; (Page 4 Draft Label and Uncertainty Estimation “We find when the epistemic uncertainty ui is larger than some threshold value Γ, then the draft label y ∗ i has a high probability of being wrong. Hence, we utilize a novel two-stream self-attention model to refine those uncertain labels using long-term label dependencies and word-label interactions”) Shu and Gui are considered analogous art to the claimed invention because they are in the same field of endeavor being neural machine translation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the reward system of Shu with the label refinement of Gui. One would want to do this to improve label accuracy (Gui Introduction). Regarding claim 2, Shu in view of Gui teaches claim 1 as outlined above. Shu further teaches: processing, by the computing system, the training input sequence with the machine- learned sequential labeling model to generate the one or more candidate output sequences comprises processing, by the computing system, the training input sequence with the machine- learned sequential labeling model to generate a plurality of candidate output sequences; and (Page 2 Section 2 “Given a source-target sequence pair (X,Y ∗ ) from the training set, let S(X,θ,N) be a set of N candidate sequences generated by model pθ(Y |X), which is reachable by a certain decoding strategy (e.g. beam search).”)) the method further comprises selecting, by the computing system, the first candidate output sequence from the plurality of candidate output sequences according to one or more criteria. (Page 3 Ranking Optimization “Assume the candidates in S(X,θ,N) are sorted by their model scores in descending order.”) Regarding claim 3, Shu in view of Gui teaches claim 2 as outlined above. Shu further teaches: processing, by the computing system, the training input sequence with the machine-learned sequential labeling model to generate the plurality of candidate output sequences comprises performing a beam search over an output space of the machine-learned sequential labeling model to generate the plurality of candidate output sequences. (Page 2 Section 2 “Given a source-target sequence pair (X,Y ∗ ) from the training set, let S(X,θ,N) be a set of N candidate sequences generated by model pθ(Y |X), which is reachable by a certain decoding strategy (e.g. beam search).”)) Regarding claim 4, Shu in view of Gui teaches claim 2 as outlined above. Shu further teaches: determining, according to the reward function, a respective reward value for each of the plurality of candidate output sequences; and (Page 3 Expected Reward Maximization “The first category of solutions approximates the best-response reward using the expected reward.”) selecting as the first candidate output sequence the candidate output sequence that receives the largest respective reward value. (Page 3 Ranking Optimization “Various methods are explored in this category. Here, we denote Y ∗ R = argmaxY ∈ SR(Y) as the candidate with the highest reward.”) Regarding claim 5, Shu in view of Gui teaches claim 1 as outlined above. Gui further teaches: the correction function comprises a pre-defined function that, when applied to a function input, with positive probability relative to a training dataset, results in a function output that receives a larger reward value than the function input according to the reward function . (Page 5 Section 3.3 Training and Decoding “When training is complete, we can obtain the draft labels Y ∗ = {y ∗ 1,y ∗ 2,...,y ∗ n} and corresponding uncertainties U = {u1,u2,...,un} from variational LSTM, and refined labels ˆY = {ˆy1, ˆ y2, . .., ˆyn} from two-stream self-attention model. To avoid the correct labels being incorrectly modified, we set an uncertainty threshold Γ to distinguish which labels should be used, i.e., we use refined labels when ui > Γ and vice versa (as an example, given u1 > Γ, u2 ≤ Γ, and un > Γ, decoding labels will become {ˆy1,y ∗ 2,..., ˆyn}).”) Regarding claim 6, Shu in view of Gui teaches claim 1 as outlined above. Gui further teaches: the correction function one or both of: drops any predicted spans in the first candidate output sequence that do not match with spans in a ground truth output sequence associated with the training input sequence; and replaces one or more incorrect tags in the first candidate output sequence with one or more ground truth tags included in the ground truth output sequence associated with the training input sequence. (Page 4 “We find when the epistemic uncertainty ui is larger than some threshold value Γ, then the draft label y ∗ i has a high probability of being wrong. Hence, we utilize a novel two-stream self-attention model to refine those uncertain labels using long-term label dependencies and word-label interactions.” And page 5 “When training is complete, we can obtain the draft labels Y ∗ = {y ∗ 1,y ∗ 2,...,y ∗ n} and corresponding uncertainties U = {u1,u2,...,un} from variational LSTM, and refined labels ˆY = {ˆy1, ˆ y2, . .., ˆyn} from two-stream self-attention model. To avoid the correct labels being incorrectly modified, we set an uncertainty threshold Γ to distinguish which labels should be used, i.e., we use refined labels when ui > Γ and vice versa (as an example, given u1 > Γ, u2 ≤ Γ, and un > Γ, decoding labels will become {ˆy1,y ∗ 2,..., ˆyn}).”) Regarding claim 7, Shu in view of Gui teaches claim 1 as outlined above. Shu further teaches: the reward function evaluates both precision and recall. (Page 5 “In our experiments, we test with two rewards: Smoothed BLEU (SLBEU) and BLEURT” Both of these encompass prevision and recall.) Regarding claim 8, Shu in view of Gui teaches claim 1 as outlined above. Shu further teaches: the loss function comprises a corrective margin loss term that contrasts the first candidate output sequence with the corrected output sequence based on the first reward value and the second reward value, wherein the corrective margin loss term provides a loss value that is equal to a maximum of zero or a margin loss value, the margin loss value equal to a margin value minus a log likelihood of the corrected output sequence plus a log likelihood of the first candidate output sequence. (Page 4 Section 3 Contrastive-Margin Loss for Reward Optimization “In this paper, we propose to use a more conservative variation of the max-margin loss, which we refer to as contrastive-margin loss. Denote Y ∼ R as the candidate with the worst reward in S(X,θ,N).” and equation 9 shows the loss formula) Regarding claim 9, Shu in view of Gui teaches claim 8 as outlined above. Shu further teaches: the margin value is equal to a scaling coefficient times a difference between the second reward value and the first reward value. (Page 4 “The margin m is computed with the reward difference, scaled by a hyperparameter α.”) Regarding claim 11, Shu in view of Gui teaches claim 1 as outlined above. Shu further teaches: the training input sequence comprises a sequence of natural language tokens; and (page 4 & 5 Section 5 Experiments subsection dataset explains the datasets used.) the first candidate output sequence comprises: a translation of the sequence of natural language tokens; (page 5 “the translations are generated using baseline transformers”) one or more part of speech classifications of the sequence of natural language tokens; one or more chunking classifications of the sequence of natural language tokens; one or more named entity recognitions for the sequence of natural language tokens; one or more segmentations for the sequence of natural language tokens; one or more named entity recognitions for the sequence of natural language tokens; one or more natural language answers to the sequence of natural language tokens; or one or more dialog responses for the sequence of natural language tokens. (page 5 “the translations are generated using baseline transformers”) Regarding claim 12, Shu teaches: A computing system configured to perform corrective reward optimization for sequential labeling, the computing system comprising one or more processors and one or more non-transitory computer-readable storing instructions that when executed by the one or more processors cause the computing system to perform operations (Abstract. They are using a computing system to perform the operations and thus would need some sort of computer-readable medium to store instructions.) processing, by the computing system, a training input sequence with a machine- learned sequential labeling model to generate one or more candidate output sequences; (Page 2 Section 2 “Given a source-target sequence pair (X,Y ∗ ) from the training set, let S(X,θ,N) be a set of N candidate sequences generated by model pθ(Y |X), which is reachable by a certain decoding strategy (e.g. beam search).”)) determining, by the computing system, a first reward value for the first candidate output sequence according to a reward function; (Page 3 Expected Reward Maximization “The first category of solutions approximates the best-response reward using the expected reward.”) determining, by the computing system, a second reward value for the corrected candidate output sequence according to the reward function; (Shu does not teach corrected candidate sequence but does teach calculating a second reward as outlined in Ranking Optimization on page 3) evaluating, by the computing system, a loss function that contrasts the first candidate output sequence with the corrected output sequence based on the first reward value and the second reward value; and (Page 4 Section 3 Contrastive-Margin Loss for Reward Optimization “In this paper, we propose to use a more conservative variation of the max-margin loss, which we refer to as contrastive-margin loss.” And equation 9 shows the loss equation.) modifying, by the computing system, one or more parameter values of the machine-learned sequential labeling model based at least in part on the loss function. (Page 5 Training and Hyperparameters “We start with the baseline model parameters and use the contrastive-margin loss to fine tune the parameters.”) Shu does not teach: applying, by the computing system, a correction function to at least a first candidate output sequence of the one or more candidate output sequences to generate a corrected output sequence; However, Gui does: applying, by the computing system, a correction function to at least a first candidate output sequence of the one or more candidate output sequences to generate a corrected output sequence; (Page 4 Draft Label and Uncertainty Estimation “We find when the epistemic uncertainty ui is larger than some threshold value Γ, then the draft label y ∗ i has a high probability of being wrong. Hence, we utilize a novel two-stream self-attention model to refine those uncertain labels using long-term label dependencies and word-label interactions”) Shu and Gui are considered analogous art to the claimed invention because they are in the same field of endeavor being neural machine translation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the reward system of Shu with the label refinement of Gui. One would want to do this to improve label accuracy (Gui Introduction). Regarding claim 13, Shu in view of Gui teaches claim 12 as outlined above. Claim 13 recites similar limitations corresponding to claim 2 and is rejected for similar reasons as claim 2 using similar teachings and rationale. Regarding claim 14, Shu in view of Gui teaches claim 13 as outlined above. Claim 14 recites similar limitations corresponding to claim 3 and is rejected for similar reasons as claim 3 using similar teachings and rationale. Regarding claim 15, Shu in view of Gui teaches claim 13 as outlined above. Claim 15 recites similar limitations corresponding to claim 4 and is rejected for similar reasons as claim 4 using similar teachings and rationale. Regarding claim 16, Shu in view of Gui teaches claim 12 as outlined above. Claim 16 recites similar limitations corresponding to claim 5 and is rejected for similar reasons as claim 5 using similar teachings and rationale. Regarding claim 17, Shu in view of Gui teaches claim 12 as outlined above. Claim 17 recites similar limitations corresponding to claim 6 and is rejected for similar reasons as claim 6 using similar teachings and rationale. Regarding claim 18, Shu in view of Gui teaches claim 12 as outlined above. Claim 18 recites similar limitations corresponding to claim 7 and is rejected for similar reasons as claim 7 using similar teachings and rationale. Regarding claim 19, Shu in view of Gui teaches claim 12 as outlined above. Claim 19 recites similar limitations corresponding to claim 8 and is rejected for similar reasons as claim 8 using similar teachings and rationale. Regarding claim 20, Shu teaches: One or more non-transitory computer readable media that store computer-executable instructions for performing operations (Abstract. They are using a computing system to perform the operations and thus would need some sort of computer-readable medium to store instructions.) processing, by the computing system, a training input sequence with a machine- learned sequential labeling model to generate one or more candidate output sequences; (Page 2 Section 2 “Given a source-target sequence pair (X,Y ∗ ) from the training set, let S(X,θ,N) be a set of N candidate sequences generated by model pθ(Y |X), which is reachable by a certain decoding strategy (e.g. beam search).”)) determining, by the computing system, a first reward value for the first candidate output sequence according to a reward function; (Page 3 Expected Reward Maximization “The first category of solutions approximates the best-response reward using the expected reward.”) determining, by the computing system, a second reward value for the corrected candidate output sequence according to the reward function; (Shu does not teach corrected candidate sequence but does teach calculating a second reward as outlined in Ranking Optimization on page 3) evaluating, by the computing system, a loss function that contrasts the first candidate output sequence with the corrected output sequence based on the first reward value and the second reward value; and (Page 4 Section 3 Contrastive-Margin Loss for Reward Optimization “In this paper, we propose to use a more conservative variation of the max-margin loss, which we refer to as contrastive-margin loss.” And equation 9 shows the loss equation.) modifying, by the computing system, one or more parameter values of the machine-learned sequential labeling model based at least in part on the loss function. (Page 5 Training and Hyperparameters “We start with the baseline model parameters and use the contrastive-margin loss to fine tune the parameters.”) Shu does not teach: applying, by the computing system, a correction function to at least a first candidate output sequence of the one or more candidate output sequences to generate a corrected output sequence; However, Gui does: applying, by the computing system, a correction function to at least a first candidate output sequence of the one or more candidate output sequences to generate a corrected output sequence; (Page 4 Draft Label and Uncertainty Estimation “We find when the epistemic uncertainty ui is larger than some threshold value Γ, then the draft label y ∗ i has a high probability of being wrong. Hence, we utilize a novel two-stream self-attention model to refine those uncertain labels using long-term label dependencies and word-label interactions”) Shu and Gui are considered analogous art to the claimed invention because they are in the same field of endeavor being neural machine translation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the reward system of Shu with the label refinement of Gui. One would want to do this to improve label accuracy (Gui Introduction) . 07-21-aia AIA Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Shu in view of Gui and Shen (NPL: ‘Minimum Risk Training for Neural Machine Translation’ (from applicants IDS)) Regarding claim 10, Shu in view of Gui teaches claim 1 as outlined above. Neither Shu nor Gui teaches: determining, by the computing system, a third reward value for a ground truth output sequence according to the reward function; wherein loss function further contrasts the corrected output sequence with the ground truth output sequence based on the second reward value and the third reward value. However, Shen does: determining, by the computing system, a third reward value for a ground truth output sequence according to the reward function; wherein loss function further contrasts the corrected output sequence with the ground truth output sequence based on the second reward value and the third reward value. (Page 2 Section 3 Minimum Risk Training for Neural Machine Translation “We use a loss function ∆(y,y(s)) to measure the discrepancy between the model prediction y and the gold standard translation y(s).”) Shu, Gui and Shen are considered analogous art to the claimed invention because they are in the same field of endeavor being neural machine translation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the reward system of Shu with the label refinement of Gui and incorporating the ground truth into the loss function of Shen. One would want to do this to maximize prediction accuracy (Shen conclusion). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to DANIEL P GRUSZKA whose telephone number is (571)272-5259. The examiner can normally be reached M-F 9:00 AM - 6:00 PM ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Li Zhen can be reached at (571) 272-3768. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DANIEL GRUSZKA/Examiner, Art Unit 2121 /Li B. Zhen/Supervisory Patent Examiner, Art Unit 2121 Application/Control Number: 18/240,954 Page 2 Art Unit: 2121 Application/Control Number: 18/240,954 Page 3 Art Unit: 2121 Application/Control Number: 18/240,954 Page 4 Art Unit: 2121 Application/Control Number: 18/240,954 Page 5 Art Unit: 2121 Application/Control Number: 18/240,954 Page 6 Art Unit: 2121 Application/Control Number: 18/240,954 Page 7 Art Unit: 2121 Application/Control Number: 18/240,954 Page 8 Art Unit: 2121 Application/Control Number: 18/240,954 Page 9 Art Unit: 2121 Application/Control Number: 18/240,954 Page 10 Art Unit: 2121 Application/Control Number: 18/240,954 Page 11 Art Unit: 2121 Application/Control Number: 18/240,954 Page 12 Art Unit: 2121 Application/Control Number: 18/240,954 Page 13 Art Unit: 2121 Application/Control Number: 18/240,954 Page 14 Art Unit: 2121 Application/Control Number: 18/240,954 Page 15 Art Unit: 2121 Application/Control Number: 18/240,954 Page 16 Art Unit: 2121 Application/Control Number: 18/240,954 Page 17 Art Unit: 2121 Application/Control Number: 18/240,954 Page 18 Art Unit: 2121 Application/Control Number: 18/240,954 Page 19 Art Unit: 2121 Application/Control Number: 18/240,954 Page 20 Art Unit: 2121 Application/Control Number: 18/240,954 Page 21 Art Unit: 2121 Application/Control Number: 18/240,954 Page 22 Art Unit: 2121 Application/Control Number: 18/240,954 Page 23 Art Unit: 2121 Application/Control Number: 18/240,954 Page 24 Art Unit: 2121