Prosecution Insights
Last updated: August 15, 2026
Application No. 17/576,724

SUBCOMPONENT MODEL TRAINING

Non-Final OA §101§103
Filed
Jan 14, 2022
Examiner
TRAN, DAVID HOANG
Art Unit
2147
Tech Center
2100 — Computer Architecture & Software
Assignee
Salesforce Inc.
OA Round
3 (Non-Final)
21%
Grant Probability
At Risk
3-4
OA Rounds
0m
Est. Remaining
35%
With Interview

Examiner Intelligence

Grants only 21% of cases
21%
Career Allowance Rate
4 granted / 19 resolved
-33.9% vs TC avg
Moderate +14% lift
Without
With
+14.1%
Interview Lift
resolved cases with interview
Typical timeline
4y 5m
Avg Prosecution
22 currently pending
Career history
57
Total Applications
across all art units

Statute-Specific Performance

§101
29.4%
-10.6% vs TC avg
§103
48.3%
+8.3% vs TC avg
§102
8.5%
-31.5% vs TC avg
§112
12.7%
-27.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 19 resolved cases

Office Action

§101 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 02/20/2026 has been entered. Response to Arguments Applicant’s arguments filed 01/21/2026 on pages 8-13 of Remarks regarding the rejection under 35 U.S.C. 101 with respect to claims 1-20 have been fully considered but they are not persuasive. Beginning on page 9, Applicant asserts that under 101 Step 2A Prong One the claims are not directed to an abstract idea because the claims recite “process[] the user dialog input with the machine learning model and the plurality subcomponent models” and “training the plurality of subcomponent models” cannot be performed in the human mind. However, Examiner respectfully disagrees. MPEP 2106.04(a)(2)(III)(c) talks about mental processes on a generic computer which include machine learning models. Also, see MPEP 2106.04(d) and 2106.05(f). Beginning on page 10, Applicant asserts that under 101 Step 2A Prong Two the claims are not directed to an abstract idea because the claims recite features that integrate any alleged judicial exception into a practical application as the claims recite “computing one or more weights for data points of the one or more subcomponent training datasets, wherein the one or more weights are based at least in part on a contribution, associated with performing the plurality of sequential subtasks, of the data points to an end-to-end error loss measurement that corresponds collectively to the plurality of sequential subtasks and the final task;”. MPEP 2106.04(d)(1) talks about a claim reciting a judicial exception is not directed to the judicial exception if it also recites additional elements demonstrating that the claim as a whole integrates the exception into a practical application. MPEP 2106.05(a) talks about the judicial exception alone cannot provide the improvement. The improvement can be provided by one or more additional elements. Improving an abstract element is not sufficient to integrate into a practical application. See updated rejection below. Applicant’s arguments on pages 13-16 regarding the rejection under 35 U.S.C. 103 with respect to claims 1-20 have been fully considered but are moot. New reference Arik has been incorporated below to teach the newly presented limitations. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Regarding Claim 1, Claim 1 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1 Analysis: Claim 1 is directed to a method, i.e., a process, one of the statutory categories. Step 2A Prong One Analysis: The limitations: “computing one or more weights for data points of the one or more subcomponent training datasets, wherein the one or more weights are based at least in part on a contribution, associated with performing the plurality of sequential subtasks, of the data points to an end-to-end error loss measurement that corresponds collectively to the plurality of sequential subtasks and the final task;” As drafted, under their broadest reasonable interpretation, cover mathematical concepts (including mathematical relationships, mathematical formulas or equations, or mathematical calculations, e.g., computing). The above limitations in the context of this claim encompass, inter alia, computing weights based at least in part on a loss measurement (corresponding to mathematical concepts). “[transmitting one or more responses to the user dialog input based at least in part on] processing the user dialog input [with the machine learning model and the plurality of subcomponent models.]” As drafted, under their broadest reasonable interpretation, cover concepts performed in human mind (including an observation, evaluation, judgement, or opinion, e.g., processing). The above limitations in the context of this claim encompass, inter alia, processing user dialog input (corresponding to mental processes which can be done mentally or by pen and paper). Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. The limitations: “wherein the machine learning model is configured to perform a final task comprising a plurality of sequential subtasks and the plurality of subcomponent models are configured to perform the plurality of sequential subtasks; “training the plurality of subcomponent models based at least in part on the one or more weights for the data points of the one or more subcomponent training datasets:” “[transmitting one or more responses to the user dialog input based at least in part on processing the user dialog input] with the machine learning model and the plurality of subcomponent models.” As drafted, are additional elements that amount to no more than mere instructions to apply the exception for the abstract ideas. See MPEP 2106.05(f). Specifically, they amount to mere instructions to apply the exception using a machine-learning based model (e.g., by using these elements as tools). The limitations: “inputting one or more subcomponent training datasets into the plurality of subcomponent models of the machine learning model,” “receiving, at the task-oriented dialog system, user dialog input; and” “transmitting one or more responses to the user dialog input based at least in part on [processing the user dialog input with the machine learning model and the plurality of subcomponent models.]” As drafted, amount to insignificant extra-solution activities, which do not integrate a judicial exception into a practical application. For example, the additional elements of "inputting subcomponent training datasets", “receiving user dialog input”, and “transmitting responses” amount to mere data gathering and data storage, respectively, which are insignificant extra-solution activities that do not integrate a judicial exception into a practical application. See MPEP 2106.05(g). Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, all of the additional elements are insignificant extra-solution activities or mere instructions to apply an exception. (i.e., the additional element describes a unit for applying the abstract ideas). Insignificant extra-solution activities and mere instructions to apply an exception cannot provide an inventive concept. Moreover, receiving, communicating, and storing data are insignificant extra-solution activities that are well-understood, routine, and conventional. See MPEP 2106.05(d)(II) ("The courts have recognized the following computer functions as well-understood, routine, and conventional functions ... i. Receiving or transmitting data over a network") (citing OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015)). The limitations: “wherein the machine learning model is configured to perform a final task comprising a plurality of sequential subtasks and the plurality of subcomponent models are configured to perform the plurality of sequential subtasks; “training the plurality of subcomponent models based at least in part on the one or more weights for the data points of the one or more subcomponent training datasets:” “[transmitting one or more responses to the user dialog input based at least in part on processing the user dialog input] with the machine learning model and the plurality of subcomponent models.” As drafted, are additional elements that amount to no more than mere instructions to apply the exception for the abstract ideas. See MPEP 2106.05(f). Specifically, they amount to mere instructions to apply the exception using a machine-learning based model (e.g., by using these elements as tools). The claim is not patent eligible. Regarding Claim 2, Claim 2 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1 Analysis: Claim 2 is directed to a method, i.e., a process, one of the statutory categories. Step 2A Prong One Analysis: The limitations: “calculating a first end-to-end error loss gradient based at least in part on the baseline end-to-end error loss measurement and the end-to-end error loss measurement.” As drafted, under their broadest reasonable interpretation, cover mathematical concepts (including mathematical relationships, mathematical formulas or equations, or mathematical calculations, e.g., calculating). The above limitations in the context of this claim encompass, inter alia, calculating a first end-to-end error loss gradient (corresponding to mathematical concepts). Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. The limitations: “obtaining a baseline end-to-end error loss measurement of the machine learning model in a non-updated state;” “obtaining the end-to-end error loss measurement based at least in part on inputting a first subcomponent training dataset of the one or more subcomponent training datasets into a first subcomponent model of the plurality of subcomponent models; and” As drafted, amount to insignificant extra-solution activities, which do not integrate a judicial exception into a practical application. For example, the additional elements of "obtaining loss measurements" amount to mere data gathering and data storage, respectively, which are insignificant extra-solution activities that do not integrate a judicial exception into a practical application. See MPEP 2106.05(g). Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, all of the additional elements are insignificant extra-solution activities or mere instructions to apply an exception. (i.e., the additional element describes a unit for applying the abstract ideas). Insignificant extra-solution activities and mere instructions to apply an exception cannot provide an inventive concept. Moreover, receiving, communicating, and storing data are insignificant extra-solution activities that are well-understood, routine, and conventional. See MPEP 2106.05(d)(II) ("The courts have recognized the following computer functions as well-understood, routine, and conventional functions ... i. Receiving or transmitting data over a network") (citing OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015)). The claim is not patent eligible. Regarding Claim 3, Claim 3 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1 Analysis: Claim 3 is directed to a method, i.e., a process, one of the statutory categories. Step 2A Prong One Analysis: The limitations: “wherein calculating the first end-to-end error loss gradient comprises calculating a finite difference approximation based at least in part on the baseline end-to-end error loss measurement and the end-to-end error loss measurement.” As drafted, under their broadest reasonable interpretation, cover mathematical concepts (including mathematical relationships, mathematical formulas or equations, or mathematical calculations, e.g., calculating). The above limitations in the context of this claim encompass, inter alia, calculating a finite difference approximation (corresponding to mathematical concepts). Step 2A Prong Two Analysis: Please see the corresponding analysis of Claim 1. Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claim is not patent eligible. Regarding Claim 4, Claim 4 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1 Analysis: Claim 4 is directed to a method, i.e., a process, one of the statutory categories. Step 2A Prong One Analysis: The limitations: “[training a critic model for the first subcomponent training dataset [based at least in part on the first end-to-end error loss gradient and a predicted future end-to-end error loss gradient for the first subcomponent training dataset;” “wherein computing the one or more weights is based at least in part on the critic model.” As drafted, under their broadest reasonable interpretation, cover mathematical concepts (including mathematical relationships, mathematical formulas or equations, or mathematical calculations, e.g., calculating, computing). The above limitations in the context of this claim encompass, inter alia, calculating a first end-to-end error loss gradient and computing weights (corresponding to mathematical concepts). Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. The limitations: “training a critic model for the first subcomponent training dataset [based at least in part on the first end-to-end error loss gradient and a predicted future end-to-end error loss gradient for the first subcomponent training dataset;]” As drafted, are additional elements that amount to no more than mere instructions to apply the exception for the abstract ideas. See MPEP 2106.05(f). Specifically, they amount to mere instructions to apply the exception using a machine-learning based model (e.g., by using these elements as tools). Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The limitations: “training a critic model for the first subcomponent training dataset [based at least in part on the first end-to-end error loss gradient and a predicted future end-to-end error loss gradient for the first subcomponent training dataset;]” As drafted, are additional elements that amount to no more than mere instructions to apply the exception for the abstract ideas. See MPEP 2106.05(f). Specifically, they amount to mere instructions to apply the exception using a machine-learning based model (e.g., by using these elements as tools). The claim is not patent eligible. Regarding Claim 5, Claim 5 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1 Analysis: Claim 5 is directed to a method, i.e., a process, one of the statutory categories. Step 2A Prong One Analysis: The limitations: “[training a critic model for the first subcomponent training dataset] based at least in part on a ranking between the first end-to-end error loss gradient and a second end-to-end error loss gradient calculated based at least in part on the baseline end-to-end error loss measurement and a second end-to-end error loss measurement, wherein the second end-to-end error loss measurement is calculated based at least in part on inputting a second subcomponent training dataset of the one or more subcomponent training datasets into the first subcomponent model of the plurality of subcomponent models;” “wherein computing the one or more weights is based at least in part on the critic model.” As drafted, under their broadest reasonable interpretation, cover mathematical concepts (including mathematical relationships, mathematical formulas or equations, or mathematical calculations, e.g., ranking, calculated, computing). The above limitations in the context of this claim encompass, inter alia, ranking between the first end-to-end error loss gradient and a second end-to-end error loss gradient calculated and computing weights (corresponding to mathematical concepts). Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. The limitations: “training a critic model for the first subcomponent training dataset [based at least in part on a ranking between the first end-to-end error loss gradient and a second end-to-end error loss gradient calculated based at least in part on the baseline end-to-end error loss measurement and a second end-to-end error loss measurement, wherein the second end-to-end error loss measurement is calculated based at least in part on inputting a second subcomponent training dataset of the one or more subcomponent training datasets into the first subcomponent model of the plurality of subcomponent models;]” As drafted, are additional elements that amount to no more than mere instructions to apply the exception for the abstract ideas. See MPEP 2106.05(f). Specifically, they amount to mere instructions to apply the exception using a machine-learning based model (e.g., by using these elements as tools). Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The limitations: “training a critic model for the first subcomponent training dataset [based at least in part on a ranking between the first end-to-end error loss gradient and a second end-to-end error loss gradient calculated based at least in part on the baseline end-to-end error loss measurement and a second end-to-end error loss measurement, wherein the second end-to-end error loss measurement is calculated based at least in part on inputting a second subcomponent training dataset of the one or more subcomponent training datasets into the first subcomponent model of the plurality of subcomponent models;]” As drafted, are additional elements that amount to no more than mere instructions to apply the exception for the abstract ideas. See MPEP 2106.05(f). Specifically, they amount to mere instructions to apply the exception using a machine-learning based model (e.g., by using these elements as tools). The claim is not patent eligible. Regarding Claim 6, Claim 6 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1 Analysis: Claim 6 is directed to a method, i.e., a process, one of the statutory categories. Step 2A Prong One Analysis: The limitations: “wherein the second end-to-end error loss gradient is calculated based at least in part on a finite difference approximation based at least in part on the baseline end-to-end error loss measurement and the second end-to-end error loss measurement.” As drafted, under their broadest reasonable interpretation, cover mathematical concepts (including mathematical relationships, mathematical formulas or equations, or mathematical calculations, e.g., calculating). The above limitations in the context of this claim encompass, inter alia, calculating a finite difference approximation (corresponding to mathematical concepts). Step 2A Prong Two Analysis: Please see the corresponding analysis of Claim 1. Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claim is not patent eligible. Regarding Claim 7, Claim 7 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1 Analysis: Claim 7 is directed to a method, i.e., a process, one of the statutory categories. Step 2A Prong One Analysis: Please see the corresponding analysis of Claim 1. Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. The limitations: “training a critic model for a first subcomponent training dataset of the one or more subcomponent training datasets;” “updating the critic model based at least in part on the end-to-end error loss measurement;” “updating the one or more weights based at least in part on the updated critic model; and” “retraining the plurality of subcomponent models based at least in part on the updated one or more weights.” As drafted, are additional elements that amount to no more than mere instructions to apply the exception for the abstract ideas. See MPEP 2106.05(f). Specifically, they amount to mere instructions to apply the exception using a machine-learning based model (e.g., by using these elements as tools). Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The limitations: “training a critic model for a first subcomponent training dataset of the one or more subcomponent training datasets;” “updating the critic model based at least in part on the end-to-end error loss measurement;” “updating the one or more weights based at least in part on the updated critic model; and” “retraining the plurality of subcomponent models based at least in part on the updated one or more weights.” As drafted, are additional elements that amount to no more than mere instructions to apply the exception for the abstract ideas. See MPEP 2106.05(f). Specifically, they amount to mere instructions to apply the exception using a machine-learning based model (e.g., by using these elements as tools). The claim is not patent eligible. Regarding Claim 8, Claim 8 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1 Analysis: Claim 8 is directed to a method, i.e., a process, one of the statutory categories. Step 2A Prong One Analysis: Please see the corresponding analysis of Claim 1. Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. The limitations: “training the plurality of subcomponent models based at least in part on a Monte Carlo tree search.” As drafted, are additional elements that amount to no more than mere instructions to apply the exception for the abstract ideas. See MPEP 2106.05(f). Specifically, they amount to mere instructions to apply the exception using a machine-learning based model (e.g., by using these elements as tools). Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The limitations: “training the plurality of subcomponent models based at least in part on a Monte Carlo tree search.” As drafted, are additional elements that amount to no more than mere instructions to apply the exception for the abstract ideas. See MPEP 2106.05(f). Specifically, they amount to mere instructions to apply the exception using a machine-learning based model (e.g., by using these elements as tools). The claim is not patent eligible. Regarding Claim 9, Claim 9 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1 Analysis: Claim 9 is directed to a method, i.e., a process, one of the statutory categories. Step 2A Prong One Analysis: Please see the corresponding analysis of Claim 1. Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. The limitations: “wherein at least one of the one or more subcomponent training datasets comprises data points associated with a subtask that is not included in the plurality of sequential subtasks.” As drafted, amount to insignificant extra-solution activities, which do not integrate a judicial exception into a practical application. For example, the additional elements of "inputting subcomponent training datasets" amount to mere data gathering and data storage, respectively, which are insignificant extra-solution activities that do not integrate a judicial exception into a practical application. See MPEP 2106.05(g). Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, all of the additional elements are insignificant extra-solution activities or mere instructions to apply an exception. (i.e., the additional element describes a unit for applying the abstract ideas). Insignificant extra-solution activities and mere instructions to apply an exception cannot provide an inventive concept. Moreover, receiving, communicating, and storing data are insignificant extra-solution activities that are well-understood, routine, and conventional. See MPEP 2106.05(d)(II) ("The courts have recognized the following computer functions as well-understood, routine, and conventional functions ... i. Receiving or transmitting data over a network") (citing OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015)). The claim is not patent eligible. Regarding Claim 10, Claim 10 recites an apparatus for performing steps similar of claim 1 and is rejected with the same rationale, mutatis mutandis, in view of the following additional elements, considered individually and as an ordered combination with the additional elements identified above, failing to integrate the abstract idea into a practical application or amount to significantly more than the abstract idea: “a processor;” “memory coupled with the processor; and” “instructions stored in the memory and executable by the processor to cause the apparatus to:” This is a recitation of generic computing components to be used in performing the abstract idea, which does not integrate the abstract idea into a practical application or amount to significantly more than the abstract idea. See MPEP 2106.05(f). Regarding Claim 11, Claim 11 recites an apparatus for performing steps substantially similar to those of claim 2 and is rejected with the same rationale, mutatis mutandis. Regarding Claim 12, Claim 12 recites an apparatus for performing steps substantially similar to those of claim 3 and is rejected with the same rationale, mutatis mutandis. Regarding Claim 13, Claim 13 recites an apparatus for performing steps substantially similar to those of claim 4 and is rejected with the same rationale, mutatis mutandis. Regarding Claim 14, Claim 14 recites an apparatus for performing steps substantially similar to those of claim 5 and is rejected with the same rationale, mutatis mutandis. Regarding Claim 15, Claim 15 recites an apparatus for performing steps substantially similar to those of claim 6 and is rejected with the same rationale, mutatis mutandis. Regarding Claim 16, Claim 16 recites an apparatus for performing steps substantially similar to those of claim 7 and is rejected with the same rationale, mutatis mutandis. Regarding Claim 17, Claim 17 recites an apparatus for performing steps substantially similar to those of claim 8 and is rejected with the same rationale, mutatis mutandis. Regarding Claim 18, Claim 18 recites an apparatus for performing steps substantially similar to those of claim 9 and is rejected with the same rationale, mutatis mutandis Regarding Claim 19, Claim 19 recites an apparatus for performing steps similar of claim 1 and is rejected with the same rationale, mutatis mutandis, in view of the following additional elements, considered individually and as an ordered combination with the additional elements identified above, failing to integrate the abstract idea into a practical application or amount to significantly more than the abstract idea: “A non-transitory computer-readable medium storing code for training a plurality of subcomponent models of a machine learning model, the code comprising instructions executable by a processor to:” This is a recitation of generic computing components to be used in performing the abstract idea, which does not integrate the abstract idea into a practical application or amount to significantly more than the abstract idea. See MPEP 2106.05(f). Regarding Claim 20, Claim 20 recites a non-transitory computer-readable medium for performing steps substantially similar to those of claim 2 and is rejected with the same rationale, mutatis mutandis. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 2, 3, 4, 7, 9, 10, 11, 12, 13, 16, 18, 19 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Lin et al. (US 20200349464 A1); hereinafter Lin in view of Arik et al. (US 20210089870 A1); hereinafter Arik in further view of Maroengsit et al. (A Survey on Evaluation Methods for Chatbots); hereinafter Maroengsit Claim 1 is rejected over Lin, Arik and Maroengsit. Regarding claim 1, Lin teaches a method for training a plurality of subcomponent models of a machine learning model [of a task-oriented dialog system], the method comprising: inputting one or more subcomponent training datasets into the plurality of subcomponent models of the machine learning model, wherein the machine learning model is configured to perform a final task comprising a plurality of sequential subtasks and the plurality of subcomponent models are configured to perform the plurality of sequential subtasks; (Lin [0036]: “As described in more detail below, the multi-task learning is based on the data from the datasets 102 a, 102 b, through 102 n (related to the different tasks) (subcomponent training datasets) being combined to provide an ensemble of training data that can be used to strategically train certain sub-modules of the machine learning model 106 (subcomponent models) in a fine-tuned manner. In some cases, the machine learning model 106 can be trained to perform multiple tasks. For instance, the multi-task learning can be used to train the machine learning model 106 to perform well on multiple tasks, leading to high performance on multiple evaluation datasets during inference.”; and [0129]: “Process 1000 is illustrated as logical flow diagrams, the operation of which represent a sequence of operations that can be implemented in hardware, computer instructions, or a combination thereof.”; Note: See Figure 1 of Lin to see that Mth Sub-Module 108m can be the final task. See Figure 10 of Lin to see that the first and second modules are sequentially performing the tasks.) training the plurality of subcomponent models based at least in part on the one or more weights for the data points of the one or more subcomponent training datasets. (Lin [0044]: “In some implementations, the percentage of data used from the different datasets 102 a, 102 b, through 102 n can change automatically throughout the training of the machine learning model 106. For instance, easier datasets can be weighted higher (with higher percentages) at the beginning of the training, and as training progresses, more difficult datasets can be weighted higher (with higher percentages) and the easier datasets can be weighted lower (with lower percentages). The gradual adjustment of the percentages of the data used from the different datasets can be controlled manually or can be controlled automatically by the machine learning system 100.”; and [0045]: “As noted above, each sub-module 108 a, 108 b, through 108 m can be designed for a specific dataset, and each input dataset 102 a, 102 b, through 102 n can be used to train its corresponding sub-module. The sub-module determination engine 104 can determine which sub-module 108 a, 108 b, through 108 m to use for processing a given dataset 102 a, 102 b, through 102 n. The determination can be based on formats of the different datasets 102 a, 102 b, through 102 n. The format of each dataset 102 a, 102 b, through 102 n can include the name of the dataset, a task for which the dataset is applicable (e.g., instance detection, image classification, referring expression, phrase matching, among others), content of the dataset (e.g., images, phases, labels, and/or other content), a combination thereof, and/or other suitable information.”;) Lin does not appear to explicitly teach computing one or more weights for data points of the one or more subcomponent training datasets, wherein the one or more weights are based at least in part on a contribution, associated with performing the plurality of sequential subtasks, of the data points to an end-to-end error loss measurement that corresponds collectively to the plurality of sequential subtasks and the final task; However, Arik teaches computing one or more weights for data points of the one or more subcomponent training datasets, (Arik [0020]: “A data value estimator function, modeled by a deep neural network, outputs a likelihood a training sample will be used in training of the predictor model.”; and [0023]: “the data value estimator model 120 is a neural network The data value estimator model 120, for each training sample 102 in the batch of training samples 102, determines a selection probability 106 based on estimator parameter values 122 of the data value estimator model 120. The selection probability 106 represents a prediction of how valuable each training sample 102 in the batch of the training samples 102 will be to the predictor model 142.”) wherein the one or more weights are based at least in part on a contribution of the data points to an end-to-end error loss measurement that corresponds collectively to the plurality of sequential subtasks and the final task; and (Arik [0027]: “The DVLR framework 110 adjusts and/or updates the model parameter values 143 of the predictor model 142 and the estimator parameter values 122 of the data value estimator model 120 based on the performance measurements 144.”; and [0029]: “the DVRL framework 110 updates the estimator parameter values 122 of the data value estimator model 120 based an the reinforcement signal 260. The reinforcement signal 260 may also include reward data 220. The performance evaluator 150 may determine the reward data 220 by quantifying the performance measurements 144. For example, when the performance measurements 144 indicate low loss data 144 (i.e., minimal error or an accurate prediction) from the subset of training samples 102 received by the predictor model 142,”; Note: The data value estimator model 120 is the critic model that is updated based on the performance measurements (end-to-end loss measurement.) It would have been obvious before the effective filing date to combine the sub-module processing of multiple datasets of Lin with quantifying the value of training data of Arik to improve model performance (Arik, [0019]). Lin and Arik are analogous art because they both concern optimizing machine learning model training by selecting training samples. Lin does not appear to explicitly teach receiving, at the task-oriented dialog system, user dialog input; and transmitting one or more responses to the user dialog input based at least in part on processing the user dialog input with the machine learning model and the plurality of subcomponent models. However, Maroengsit teaches receiving, at the task-oriented dialog system, user dialog input; and (See page 113, Figure 1 of Maroengsit to see that the user inputs message into the user interface as part of the task-oriented dialog system.) transmitting one or more responses to the user dialog input based at least in part on processing the user dialog input with the machine learning model and the plurality of subcomponent models. (Maroengsit [page 112, 2. Architecture]: “In Figure 1, chatbots are categorized into two types by how it generates a response. Rule-based chatbots use algorithms either handcraft or already existing ones as a decision maker to identify both knowledge and response.”; Note: See Figure 1 of Maroengsit to see that subcomponent models include TF-IDF, Word2Vec, Intents Classification, Name Entity Recognitions and LSTM.) It would have been obvious before the effective filing date to combine the sub-module processing of multiple datasets of Lin with the task-oriented dialog system using subcomponent models of Maroengsit to effectively generate a response and keep on learning and improving based on previous learning models (Maroengsit, page 112, 2. Architecture). Lin and Maroengsit are analogous art because they both concern multi-tasking with multiple subprocesses. Claim 2 is rejected over Lin, Arik and Maroengsit with the incorporation of claim 1. Lin does not appear to explicitly teach obtaining a baseline end-to-end error loss measurement of the machine learning model in a non-updated state; obtaining the end-to-end error loss measurement based at least in part on inputting a first subcomponent training dataset of the one or more subcomponent training datasets into a first subcomponent model of the plurality of subcomponent models; and calculating a first end-to-end error loss gradient based at least in part on the baseline end-to-end error loss measurement and the end-to-end error loss measurement. However, Arik teaches obtaining a baseline end-to-end error loss measurement of the machine learning model in a non-updated state; (Arik [0030]: “the performance evaluator 150 calculates reward data 220 based on historical loss data. For example, the performance evaluator 150 determines, using a moving average calculator 146, a moving average of loss data based on N-most recent training iterations of the predictor model 142”; and [0033]: “the DVLR 110 accepts the set of training samples 102 (i.e., D), and initializes the estimator parameter values of the data value estimator model 120, the model parameter values of the predictor model 142, and resets the moving average loss in the moving average loss calculator 146.”) obtaining the end-to-end error loss measurement based at least in part on inputting a first subcomponent training dataset of the one or more subcomponent training datasets into a first subcomponent model of the plurality of subcomponent models; and (Arik [0025]: “The predictor model 142 determines performance measurements 144 based on the subset of training samples 102 sampled from the batch of input training samples 102 selected for the current training iteration”; [0028]: “the DVRL framework 110 trains the predictor model 142 using a stochastic gradient descent optimization algorithm with a loss function (e.g., mean squared error (MSE) for regression or cross entropy for classification). When the performance evaluator 150 determines the loss data 144 based on the loss function, the DVLR 110 updates the model values parameter 143 of the predictor model 142 with the performance measurements 144 (e.g., loss data 144) using the feedback loop 148.”) calculating a first end-to-end error loss gradient based at least in part on the baseline end-to-end error loss measurement and the end-to-end error loss measurement. (Arik [0029]: “After the DVRL framework 110 determines the loss data 144 for the training iteration, the DVLR 110 may generate a reinforcement signal 260… The performance evaluator 150 may determine the reward data 220 by quantifying the performance measurements 144.”; and [0030]: “the moving average calculator 146 may obtain the loss data 144 and determine the difference between the current training iteration loss data 144 and the average of the N-most recent training iterations of loss data… The reward value 230 may be based on the difference between the current training iteration loss data 144 and the average of the N-most recent training iterations of loss data.”) It would have been obvious before the effective filing date to combine the sub-module processing of multiple datasets of Lin with quantifying the value of training data of Arik to improve model performance (Arik, [0019]). Lin and Arik are analogous art because they both concern optimizing machine learning model training by selecting training samples. Claim 3 is rejected over Lin, Arik and Maroengsit with the incorporation of claim 1. Regarding claim 3, Lin does not appear to explicitly teach wherein calculating the first end-to-end error loss gradient comprises calculating a finite difference approximation based at least in part on the baseline end-to-end error loss measurement and the end-to-end error loss measurement. However, Arik teaches wherein calculating the first end-to-end error loss gradient comprises calculating a finite difference approximation based at least in part on the baseline end-to-end error loss measurement and the end-to-end error loss measurement. (Arik [0030]: “the moving average calculator 146 may obtain the loss data 144 and determine the difference between the current training iteration loss data 144 and the average of the N-most recent training iterations of loss data.”) It would have been obvious before the effective filing date to combine the sub-module processing of multiple datasets of Lin with quantifying the value of training data of Arik to improve model performance (Arik, [0019]). Lin and Arik are analogous art because they both concern optimizing machine learning model training by selecting training samples. Claim 4 is rejected over Lin, Arik and Maroengsit with the incorporation of claim 1. Regarding claim 4, Lin does not appear to explicitly teach training a critic model for the first subcomponent training dataset based at least in part on the first end-to-end error loss gradient and a predicted future end-to-end error loss gradient for the first subcomponent training dataset; wherein computing the one or more weights is based at least in part on the critic model. However, Arik teaches training a critic model for the first subcomponent training dataset (Arik [0020]: “A data value estimator function, modeled by a deep neural network, outputs a likelihood a training sample will be used in training of the predictor model.”; and [0023]: “the data value estimator model 120 is a neural network The data value estimator model 120, for each training sample 102 in the batch of training samples 102, determines a selection probability 106 based on estimator parameter values 122 of the data value estimator model 120. The selection probability 106 represents a prediction of how valuable each training sample 102 in the batch of the training samples 102 will be to the predictor model 142.”; Note: The data value estimator model is the critic model.) based at least in part on the first end-to-end error loss gradient and a predicted future end-to-end error loss gradient for the first subcomponent training dataset; (Arik [0029]: “when the performance measurements 144 indicate low loss data 144 (i.e., minimal error or an accurate prediction) from the subset of training samples 102 received by the predictor model 142, the reward data 220 may reinforce the estimator parameters values 122 of the data value estimator model 120. Conversely, when the performance measurements 144 indicate high loss data 144 (i.e., high error) from the subset of training samples 102 received by the predictor model 142, the reward data 220 may indicate that the estimator parameter values 122 of the data value estimator model 120 need further updating.”; and [0033]: “The DVLR 110 next updates the estimator parameter values 122 of the data value estimator model 120 based on the performance measurements 144 for the training iteration including the moving average loss front the moving average loss calculator 146.”) wherein computing the one or more weights is based at least in part on the critic model. (Arik [0024]: “The sampler 130 selects, based on the selection probabilities 106 of each training sample 102, a subset of training, samples 102 to provide to the predictor model 142.”; and [0008]: “for each training sample in the batch of training samples, determining, using a data value estimator, a selection probability. The selection probability for the training sample is based on estimator parameter values of the data value estimator. The operations also include selecting, based on the selection probabilities of each training sample, a subset of training samples from the batch of training samples, and determining, using a predictor model with the subset of training samples, performance measurements.”; Note: The selection probabilities output by the DVE are the weights for each data point.) It would have been obvious before the effective filing date to combine the sub-module processing of multiple datasets of Lin with quantifying the value of training data of Arik to improve model performance (Arik, [0019]). Lin and Arik are analogous art because they both concern optimizing machine learning model training by selecting training samples. Claim 7 is rejected over Lin, Arik and Maroengsit with the incorporation of claim 1. Regarding claim 7 Lin teaches retraining the plurality of subcomponent models based at least in part on the updated one or more weights. (Lin [0045]: “In some implementations, the percentage of data used from the different datasets 102 a, 102 b, through 102 n can change automatically throughout the training of the machine learning model 106. For instance, easier datasets can be weighted higher (with higher percentages) at the beginning of the training, and as training progresses, more difficult datasets can be weighted higher (with higher percentages) and the easier datasets can be weighted lower (with lower percentages). The gradual adjustment of the percentages of the data used from the different datasets can be controlled manually or can be controlled automatically by the machine learning system 100.”; [0044] and “As noted above, each sub-module 108 a, 108 b, through 108 m can be designed for a specific dataset, and each input dataset 102 a, 102 b, through 102 n can be used to train its corresponding sub-module. The sub-module determination engine 104 can determine which sub-module 108 a, 108 b, through 108 m to use for processing a given dataset 102 a, 102 b, through 102 n. The determination can be based on formats of the different datasets 102 a, 102 b, through 102 n. The format of each dataset 102 a, 102 b, through 102 n can include the name of the dataset, a task for which the dataset is applicable (e.g., instance detection, image classification, referring expression, phrase matching, among others), content of the dataset (e.g., images, phases, labels, and/or other content), a combination thereof, and/or other suitable information.”;) Lin does not appear to explicitly teach training a critic model for a first subcomponent training dataset of the one or more subcomponent training datasets; updating the critic model based at least in part on the end-to-end error loss measurement; updating the one or more weights based at least in part on the updated critic model; and However, Arik teaches training a critic model for a first subcomponent training dataset of the one or more subcomponent training datasets; (Arik [0020]: “A data value estimator function, modeled by a deep neural network, outputs a likelihood a training sample will be used in training of the predictor model.”) updating the critic model based at least in part on the end-to-end error loss measurement; (Arik [0027]: “The DVLR framework 110 adjusts and/or updates the model parameter values 143 of the predictor model 142 and the estimator parameter values 122 of the data value estimator model 120 based on the performance measurements 144.”; and [0029]: “the DVRL framework 110 updates the estimator parameter values 122 of the data value estimator model 120 based an the reinforcement signal 260.”) updating the one or more weights based at least in part on the updated critic model; and (Arik [0033]: “the DVLR 110 updates the model parameter values 143 of the predictor model 142 based on the performance measurements 144 for the training iteration… At the final step, the DVLR updates the moving average loss in the moving average loss calculator 146.”) It would have been obvious before the effective filing date to combine the sub-module processing of multiple datasets of Lin with quantifying the value of training data of Arik to improve model performance (Arik, [0019]). Lin and Arik are analogous art because they both concern optimizing machine learning model training by selecting training samples. Claim 9 is rejected over Lin, Arik and Maroengsit with the incorporation of claim 1. Regarding claim 9, Lin teaches wherein at least one of the one or more subcomponent training datasets comprises data points associated with a subtask that is not included in the plurality of sequential subtasks. (Lin [0044]: “In such an instance, the machine learning system 600 can start training with 100% data from RefCOCO, and can gradually add Visual Genome data to the training data up until the percentage reaches 20% RefCOCO and 80% Visual Genome, or other suitable ratio. In another illustrative example, the percentage of data used from the different datasets 102 a, 102 b, through 102 n can be adjusted according to the loss (or based on the decrease of the loss) from the different datasets 102 a, 102 b, through 102 n, which can lead to balanced performance over the various tasks related to the different datasets 102 a, 102 b, through 102 n.”;) Claim 10 is rejected over Lin, Arik and Maroengsit. Regarding claim 10, Lin teaches an apparatus for training a plurality of subcomponent models of a machine learning model of a task-oriented dialog system, comprising: (Lin [0128]: “the process 1000 may be performed by a computing device or apparatus,”;) a processor; (Lin [0128]: “the computing device or apparatus may include an input device, a sub-module determination engine, an output device, a processor, microprocessor,”;) memory coupled with the processor; and (Lin [0034]: “the machine learning system 100 may also include, in some instances, one or more memory devices (e.g., one or more random access memory (RAM) components”;) instructions stored in the memory and executable by the processor to cause the apparatus to: (Lin [0139]: “Processes and methods according to the above-described examples can be implemented using computer-executable instructions that are stored or otherwise available from computer-readable media. Such instructions can include, for example, instructions and data which cause or otherwise configure a general purpose computer, special purpose computer, or a processing device to perform a certain function or group of functions”;) The remainder of claim 10 is claim 1 in the form of an apparatus and is rejected for the same reasons as claim 1 stated above. Dependent claim 11 is claim 2 in the form of an apparatus and is rejected for the same reasons as claim 2 stated above. For the rejection of the limitations specifically pertaining to the apparatus of claim 10, see the rejection of claim 10 above. Dependent claim 12 is claim 3 in the form of an apparatus and is rejected for the same reasons as claim 3 stated above. For the rejection of the limitations specifically pertaining to the apparatus of claim 10, see the rejection of claim 10 above. Dependent claim 13 is claim 4 in the form of an apparatus and is rejected for the same reasons as claim 4 stated above. For the rejection of the limitations specifically pertaining to the apparatus of claim 10, see the rejection of claim 10 above. Dependent claim 16 is claim 7 in the form of an apparatus and is rejected for the same reasons as claim 7 stated above. For the rejection of the limitations specifically pertaining to the apparatus of claim 10, see the rejection of claim 10 above. Dependent claim 18 is claim 9 in the form of an apparatus and is rejected for the same reasons as claim 9 stated above. For the rejection of the limitations specifically pertaining to the apparatus of claim 10, see the rejection of claim 10 above. Claim 19 is rejected over Lin, Arik and Maroengsit. Regarding claim 19, Lin teaches a non-transitory computer-readable medium storing code for training a plurality of subcomponent models of a machine learning model, the code comprising instructions executable by a processor to: (Lin [0130]: “the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.”;) The remainder of claim 19 is claim 1 in the form of a non-transitory computer-readable medium and is rejected for the same reasons as claim 1 stated above. Dependent claim 20 is claim 2 in the form of a non-transitory computer-readable medium and is rejected for the same reasons as claim 2 stated above. For the rejection of the limitations specifically pertaining to a non-transitory computer-readable medium of claim 19, see the rejection of claim 19 above. Claims 5, 6, 14 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Lin, Arik and Maroengsit in further view of Du et al. (Sequential Scenario-Specific Meta Learner for Online Recommendation); hereinafter Du Claim 5 is rejected under Lin, Arik, Maroengsit and Du with the incorporation of claim 1. Regarding claim 5, Lin does not appear to explicitly teach training a critic model for the first subcomponent training dataset based at least in part on a ranking between the first end-to-end error loss gradient and a second end-to-end error loss gradient calculated based at least in part on the baseline end-to-end error loss measurement and a second end-to-end error loss measurement, wherein the second end-to-end error loss measurement is calculated based at least in part on inputting a second subcomponent training dataset of the one or more subcomponent training datasets into the first subcomponent model of the plurality of subcomponent models; PNG media_image1.png 54 356 media_image1.png Greyscale However, Du teaches training a critic model for the first subcomponent training dataset based at least in part on a ranking between the first end-to-end error loss gradient and a second end-to-end error loss gradient calculated based at least in part on the baseline end-to-end error loss measurement and (Du [page 2899]: “To overcome the drawback of hand-crafted stop rules, we propose to learn the stop policy with a neural network, which we call stop controller Ms (critic model).”; and “Then the accumulative reward at t is: [Equation (10)] which is the loss decrease from step t to the end of learning.”; Note: Equation 10 shows the ranking between two loss values, one at a previous time step (t-1) and one after the final training step (T).) a second end-to-end error loss measurement, wherein the second end-to-end error loss measurement is calculated based at least in part on inputting a second subcomponent training dataset of the one or more subcomponent training datasets into the first subcomponent model of the plurality of subcomponent models; (See Figure 1 of Du to see that all datasets including the second dataset Party are passed through the same subcomponent model structure.) It would have been obvious before the effective filing date to combine the sub-module processing of multiple datasets of Lin with the stop controller and reward equation of Du to effectively improve model performance (Du, page 2900, 5.2 Performance Comparison). Lin and Du are analogous art because they both concern optimizing machine learning model training by processing data sets. Lin does not appear to explicitly teach wherein computing the one or more weights is based at least in part on the critic model. However, Arik teaches wherein computing the one or more weights is based at least in part on the critic model. (Arik [0022]: “In an embodiment, the disclosure provides a computer-implemented process of building a predictive (ML) model (critic model) to predict the usefulness (weight) of a record (data point) in the context of the training process of a machine learning model.”;) It would have been obvious before the effective filing date to combine the sub-module processing of multiple datasets of Lin with quantifying the value of training data of Arik to improve model performance (Arik, [0019]). Lin and Arik are analogous art because they both concern optimizing machine learning model training by selecting training samples. Claim 6 is rejected under Lin, Arik, Maroengsit and Du with the incorporation of claim 1. Regarding claim 6, Lin does not appear to explicitly teach wherein the second end-to-end error loss gradient is calculated based at least in part on a finite difference approximation based at least in part on the baseline end-to-end error loss measurement and the second end-to-end error loss measurement. However, Du teaches wherein the second end-to-end error loss gradient is calculated based at least in part on a finite difference approximation based at least in part on the baseline end-to-end error loss measurement and the second end-to-end error loss measurement. (Du [page 2899]: “To overcome the drawback of hand-crafted stop rules, we propose to learn the stop policy with a neural network, which we call stop controller Ms (critic model).”; and “Then the accumulative reward at t is: [Equation (10)] which is the loss decrease PNG media_image2.png 56 344 media_image2.png Greyscale from step t to the end of learning.”; Note: Equation 10 shows the ranking between two loss values, one at a previous time step (t-1) and one after the final training step (T). And see Figure 1 of Du to see that all datasets including the second dataset Party are passed through the same subcomponent model structure.) It would have been obvious before the effective filing date to combine the sub-module processing of multiple datasets of Lin with the stop controller and reward equation of Du to effectively improve model performance (Du, page 2900, 5.2 Performance Comparison). Lin and Du are analogous art because they both concern optimizing machine learning model training by processing data sets. Dependent claim 14 is claim 5 in the form of an apparatus and is rejected for the same reasons as claim 5 stated above. For the rejection of the limitations specifically pertaining to the apparatus of claim 10, see the rejection of claim 10 above. Dependent claim 15 is claim 6 in the form of an apparatus and is rejected for the same reasons as claim 6 stated above. For the rejection of the limitations specifically pertaining to the apparatus of claim 10, see the rejection of claim 10 above. Claims 8 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Lin, Arik and Maroengsit in further view of Danihelka et al. (US 20220366246 A1); hereinafter Danihelka Claim 8 is rejected under Lin, Arik, Maroengsit and Danihelka with the incorporation of claim 1. Regarding claim 8, Lin does not appear to explicitly teach training the plurality of subcomponent models based at least in part on a Monte Carlo tree search. However, Danihelka teaches training the plurality of subcomponent models based at least in part on a Monte Carlo tree search. (Danihelka [0070]: “the MDP-based planning algorithm may be a Monte Carlo tree search (MCTS) algorithm. At each time step, the system 100 makes use of an action selection policy, the reward estimate, and, when relevant, the value estimate generated by the environment model in accordance with current model parameters. Each value estimate, when considered, specifies a value of the environment being in the predicted next environment state to performing the task.”; and [0084]: “The training engine 116 trains the policy neural network 110 and the environment model 150 (subcomponent models) by using a reinforcement learning technique to iteratively adjust the values of the set of parameters of the policy neural network 110 and the environment model 150 based on the interactions of the agent with the environment. An example of a suitable reinforcement learning technique is described in Espeholt, Lasse, e.sub.t al. “IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures.” ICML. 2018.”; Note: See Figure 1 of Danihelka to see that the Policy Neural Network 110 and Environment Model are subcomponent models.) It would have been obvious before the effective filing date to combine the sub-module processing of multiple datasets of Lin with Monte Carlo tree search algorithm of Danihelka to effectively improve agent performance (Danihelka, [0070]). Lin and Danihelka are analogous art because they both concern optimizing machine learning model training by processing data sets. Dependent claim 17 is claim 7 in the form of an apparatus and is rejected for the same reasons as claim 7 stated above. For the rejection of the limitations specifically pertaining to the apparatus of claim 10, see the rejection of claim 10 above. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to DAVID H TRAN whose telephone number is (703)756-1525. The examiner can normally be reached M-F 9:30 am - 5:30 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached at (571) 270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DAVID H TRAN/Examiner, Art Unit 2147 /VIKER A LAMARDO/Supervisory Patent Examiner, Art Unit 2147
Read full office action

Prosecution Timeline

Show 4 earlier events
Aug 14, 2025
Response Filed
Nov 21, 2025
Final Rejection mailed — §101, §103
Jan 21, 2026
Response after Non-Final Action
Feb 20, 2026
Request for Continued Examination
Mar 04, 2026
Response after Non-Final Action
May 05, 2026
Non-Final Rejection mailed — §101, §103
Jul 29, 2026
Examiner Interview Summary
Jul 29, 2026
Applicant Interview (Telephonic)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12632724
CANONICALIZATION OF DATA WITHIN OPEN KNOWLEDGE GRAPHS
4y 8m to grant Granted May 19, 2026
Patent 12579404
PROCESSOR FOR NEURAL NETWORK, PROCESSING METHOD FOR NEURAL NETWORK, AND NON-TRANSITORY COMPUTER READABLE STORAGE MEDIUM
4y 2m to grant Granted Mar 17, 2026
Study what changed to get past this examiner. Based on 2 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
21%
Grant Probability
35%
With Interview (+14.1%)
4y 5m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 19 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month