Prosecution Insights
Last updated: October 02, 2026
Application No. 18/037,149

INFORMATION PROCESSING DEVICE, INFORMATION PROCESSING METHOD, AND RECORDING MEDIUM

Final Rejection §102§103
Filed
May 16, 2023
Priority
Nov 30, 2020 — nonprovisional of PCTJP2020044486
Examiner
BALAKRISHNAN, VIJAY MURALI
Art Unit
2143
Tech Center
2100 — Computer Architecture & Software
Assignee
NEC Corporation
OA Round
2 (Final)
41%
Grant Probability
Moderate
3-4
OA Rounds
7m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 41% of resolved cases
41%
Career Allowance Rate
11 granted / 27 resolved
-14.3% vs TC avg
Strong +73% interview lift
Without
With
+73.3%
Interview Lift
resolved cases with interview
Typical timeline
3y 11m
Avg Prosecution
15 currently pending
Career history
44
Total Applications
across all art units

Statute-Specific Performance

§101
27.5%
-12.5% vs TC avg
§103
36.6%
-3.4% vs TC avg
§102
12.6%
-27.4% vs TC avg
§112
23.3%
-16.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 27 resolved cases

Office Action

§102 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Interpretation As recited in MPEP § 2111, during patent examination, “the pending claims must be given their broadest reasonable interpretation consistent with the specification”. Under a broadest reasonable interpretation (BRI), claim terms must be given their plain and ordinary meaning (i.e., the meaning that the term would have to a person of ordinary skill in the art), unless applicant sets forth a special definition of a claim term within the specification. The plain and ordinary meaning of a term “may be evidenced by a variety of sources, including the words of the claims themselves, the specification, drawings, and prior art”. Claims 1, 8, and 9, each recite the limitations “assign[ing] labels to the training examples using a teacher model” and “calculat[ing] errors between predictions of the one or more student models and predictions of the teacher model”. The specification further describes a teacher model as a model “which can be regarded as outputting absolutely correct predictions” [see ¶ 0013-0014], i.e., an oracle [see ¶ 0002] that always assigns correct labels to examples. It is well understood in the art that it is virtually impossible for a trained machine learning model to be “absolutely correct” in its predictions, i.e., 100% accurate, due to inherent probabilistic noise/randomness in real-world data and fundamental model limitations. The specification also does not appear to further explain how such a trained model would be generated or prepared. Based on the limited description of the specification and what would be understood by one of ordinary skill in the art, the examiner has thereby broadly interpreted a “teacher model” to be encompass any oracle/expert system with access to correct predictions, e.g., a human annotator that assigns ground truth labels to examples, or a system/component that receives, stores and accesses ground truth labels. Claims 1, 8, and 9 further recite the limitation “generating one or more student models using at least a part of the training examples to which the labels are assigned”. Although the specification describes the recited models in the context of machine learning, the claims do not expressly define the recited “student models” as being machine learning models, and the additional limitations of the claim do not particularly require the functionality of machine learning models. As such, the examiner has broadly interpreted “student models” to encompass any rule-based system or algorithm that is developed based on observed examples (i.e., training examples). Claims 1, 8, and 9 further recite the limitation “extract[ing] and output[ting] each example for which the error is to be significant based on the calculated errors”. While neither the specification or claims explicitly set forth a requisite degree for determining an error to be “significant”, applicant’s specification does describe errors “greater than a predetermined threshold value” [¶ 0041], or errors “in which a weighted sum of a degree of appearance and the error increases” [¶ 0042], to be examples of significance. The examiner has thereby broadly interpreted “significant” errors to be any errors that meet a predetermined condition and/or exceed a predetermined measure or threshold. Claim 3 recites the term “a degree of appearance”. While neither the specification or claims explicitly set forth a special definition, the specification does appear to describe “degrees of appearance” as synonymous to probability values with respect to examples within a distribution of predicted class outputs [¶ 0032, 0042]. The examiner has thereby broadly interpreted “a degree of appearance” to be any measure of likelihood or confidence with respect to a given example. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 3, 5-6, and 8-9 are rejected under 35 U.S.C. 103 as being unpatentable over Hady (“Semi-supervised Learning for Regression with Co-training by Committee”, available 2009), in view of Kee (“Query-by-committee improvement with diversity and density in batch active learning”, available online 3 May 2018) and Shi (“Knowledge Distillation for Recurrent Neural Network Language Modeling with Trust Regularization”, available online 17 Apr 2019). Regarding claim 1, Hady discloses An information processing device (“In this paper, a semi-supervised regression framework, denoted by CoBCReg is proposed, in which an ensemble of diverse regressors is used for semi-supervised learning that requires neither redundant independent views nor different base learning algorithms. Experimental results show that CoBCReg can effectively exploit unlabeled data to improve the regression estimates” [Hady Abstract]; “An experimental study is conducted to evaluate CoBCReg framework on six data sets described in Table 1…All algorithms are implemented using WEKA library [12]” [Hady pages 126-127 Methodology]; Evaluating the disclosed CoBCReg framework through processing datasets and utilizing open source machine learning libraries inherently relies upon conventional computer implementation (i.e., device comprising memory coupled to at least one processor) to perform necessary functions) comprising: a memory storing instructions; ([Hady pages 126-127 Methodology] as detailed above) and one or more processors configured to execute the instructions ([Hady pages 6-7 Methodology] as detailed above) to: receive training examples formed by features; (“An experimental study is conducted to evaluate CoBCReg framework on six data sets described in Table 1… The input features and the real-valued outputs are scaled to [0, 1]. For each experiment, 5 runs of 4-fold cross-validation have been performed. That is, for each data set, 25% are used as test set, while the remaining 75% are used as training examples where 10% of the training examples are randomly selected as the initial labeled data set L while the remaining 90% of the 75% of data are used as unlabeled data set U” [Hady pages 126-127 Methodology]) assign labels to the training examples using a teacher model; (“For each experiment, 5 runs of 4-fold cross-validation have been performed. That is, for each data set, 25% are used as test set, while the remaining 75% are used as training examples where 10% of the training examples are randomly selected as the initial labeled data set L while the remaining 90% of the 75% of data are used as unlabeled data set U” [Hady pages 126-127 Methodology]; Via execution of cross-validation, a component of the disclosed device (i.e., teacher model) accesses stored data to assign a set of labeled training examples L) generate one or more student models using at least a part of the training examples to which the labels are assigned, (“In the experiments, an initial ensemble of four RBF network regressors, N = 4, is constructed by Bagging… Table 2 present the average of the RMSEs of the four RBF Network regressors used in CoBCReg and the RMSE of CoBCReg on the test set at iteration 0 (initial) trained only on the 10% available labeled data L” [Hady page 127 Methodology and Results]; see lines 1-3 of Algorithm 1. COBC for Regression (wherein L is the set of m labeled training examples, N is the number of committee members (ensemble size), and hi is each ensemble member) – PNG media_image1.png 70 666 media_image1.png Greyscale [Hady page 123]; The disclosed algorithm initially generates each RBF network regressor (i.e., student model) from a respective set of labeled training examples Li) calculate, using error calculation examples different from the part of the training examples used to generate the one or more student models, errors between predictions by one or more student models and predictions by the teacher model; (“For each iteration t and for each ensemble member hi, a set U’ of u examples is drawn randomly from U. The SelectRelevantExamples method is applied such that the companion committee Hi (ensemble consists of all members except hi) estimates the output of each unlabeled example in U’” [Hady page 124 Co-training by Committee for Regression (CoBCReg)]; “Then, the root mean squared error (RMSE) of hj is evaluated first (∈ j )… It is worth mentioning that the RMSEs ∈ j and ∈’ j should be estimated accurately. If the training data of hj is used, this will under-estimate the RMSE. Fortunately, since the bootstrap sampling [5] is used to construct the committee, the out-of-bootstrap examples are considered for a more accurate estimate of ∈’j” [Hady pages 125-126 Confidence Measure]; see lines 8-9 of Algorithm 1 – PNG media_image2.png 51 586 media_image2.png Greyscale , and lines 2-6 of Algorithm 2. SelectRelevantExamples – PNG media_image3.png 156 712 media_image3.png Greyscale [Hady page 123]; The disclosed algorithm calculates error for out-of-bag examples from validation set Vj (i.e., error calculation examples) by calculating root mean squared error (RMSE) between example label (i.e., predictions of teacher model) and RBF network output (i.e., predictions of student model) – note that examples of Vj are out-of-bag, i.e., separate from bagging examples Li used to initially generate ensemble members, as shown in line 2 of Algorithm 1) retain examples formed by features in a data retention means; (That is, for each data set, 25% are used as test set, while the remaining 75% are used as training examples where 10% of the training examples are randomly selected as the initial labeled data set L while the remaining 90% of the 75% of data are used as unlabeled data set U” [Hady pages 126-127 Methodology]; All examples, including unlabeled examples, are accessed from the data set and retained by the disclosed device) extract, from among the examples formed by features retained in the dat retention means, each example for which the error is to be significant based on the calculated errors (“Thus, for each regressor hj , create a pool U_ of u unlabeled examples… Finally, the unlabeled example ˜xj which maximizes the relative improvement of the RMSE (Δxu ) is selected as the most relevant example labeled by companion committee Hj” [Hady page 125 Confidence Measure]; see lines 7-14 of Algorithm 2 – PNG media_image4.png 202 508 media_image4.png Greyscale [Hady page 123]; Based on measure Δxu, which is determined based on calculated errors (see PNG media_image5.png 25 192 media_image5.png Greyscale in line 5 of Algorithm 2), the disclosed algorithm returns each example that meets a condition with respect to its impact on the calculated errors (i.e., error is significant)) However, Hady does not expressly teach converting outputs of the predictions by the one or more student models and outputs of predictions by the teacher model into respective probability distributions, and taking the Kullback-Leibler divergence (KL divergence) of the converted probability distributions. In the same field of endeavor, Shi teaches a teacher-student knowledge distillation framework (“In this paper, we examine the effect of applying knowledge distillation in reducing the model size for RNNLMs. In addition, we propose a trust regularization method to improve the knowledge distillation training for RNNLMs. Using knowledge distillation with trust regularization, we reduce the parameter size to a third of that of the previously published best model while maintaining the state-of-the-art perplexity result on Penn Treebank data” [Shi Abstract]) that convert[s] outputs of the predictions by the one or more student models and outputs of predictions by the teacher model into respective probability distributions; and tak[es] the Kullback-Leibler divergence (KL divergence) of the converted probability distributions (“In knowledge distillation, the student model is learned to minimize the combination of cross-entropy loss based on training data labels and Kullback-Leibler (KL) divergence to teacher model distribution. In this paper, our experiments in the context of language modeling show that the knowledge distillation methods [8] that use cross-entropy loss and KL divergence with fixed interpolation weights get worse results than using KL divergence alone. Hence we propose a trust regularization (TR) method to dynamically adjust the combination weights for these two types of losses.” [Shi page 7230 Introduction]; “In language modeling, each training data label is represented as a degenerated data distribution which gives all probability mass to one class. So the degenerated data distribution is localized and can be over-confident, comparing with the teacher’s probability distribution learned over the whole training data. Different from previous observations using knowledge distillation for acoustic modeling [8] and image classification [15], our experiment results show that the student model learned by minimizing the interpolation of cross-entropy loss and KL divergence performs worse than the student model learned by minimizing the KL divergence alone.” [Shi page 7231 Trust Regularization]) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated converting outputs of the predictions by the one or more student models and outputs of predictions by the teacher model into respective probability distributions; and taking the Kullback-Leibler divergence (KL divergence) of the converted probability distributions as taught by Shi into Hady because they are both directed towards teacher-student knowledge distillation frameworks. Incorporating the trust regularization method taught by Shi into the knowledge distillation framework of Hady would improve performance of the learned student model while simultaneously reducing memory footprint and computational cost [Shi Abstract]. However, the combination of Hady and Shi does not expressly teach assign[ing], using the teacher model, labels to the extracted examples for which the error is significant; and us[ing] the labeled examples for which the error is significant for retraining the student model. In the same field of endeavor, Kee teaches a query-by-committee ensemble learning framework that utilizes sets of labeled and unlabeled data to answer queries (“In this study, we utilize query-by-committee (QBC) for uncertainty and demonstrate that its performance can be improved by introducing diversity and density in instance utility. Test results show that uncertainty sampling by QBC can be significantly improved with diversity and density incorporated in instance selection. Furthermore, we investigate several distance measures for use in diversity and density and show that random forest dissimilarity can be an effective distance measure in batch active learning” [Kee Abstract]; “We describe general BAL procedures discussed in Sections 3.1 to 3.3 , namely QBC only (QO), QBC and diversity (QD), and QBC, diversity, and density (QDD) settings, respectively. The overall procedures are similar in BAL scheme, yet show difference in the objective function and instance selection. Initially, a set of labeled instances, L , a set of unlabeled instances, U, and a constant batch size, q , are given” [Kee page 405 Batch active learning procedures]) that assign[s], using the teacher model, labels to the extracted examples for which the error is significant; and us[es] the labeled examples for which the error is significant for retraining the student model (see, e.g., Algorithm 2 for QD BAL procedure – while the termination condition is not satisfied, the most uncertain extracted examples are selected (see lines 11-12 – PNG media_image6.png 62 822 media_image6.png Greyscale ) and labeled (see line 15 – PNG media_image7.png 40 276 media_image7.png Greyscale ), and the model continues to retrain on the newly labeled examples (see line 7 – PNG media_image8.png 41 281 media_image8.png Greyscale ) [Kee page 407]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated assign[s], using the teacher model, labels to the extracted examples for which the error is significant; and us[es] the labeled examples for which the error is significant for retraining the student model as taught by Kee into the combination of Hady and Shi because both Hady and Kee are directed towards query-by-committee ensembles that utilize sets of labeled and unlabeled data to answer queries. It is noted that Hady expressly discusses adaptability of the disclosed CoBCReg semi-supervised learning framework to a query by committee active learning environment (“There are many interesting directions for future work…Finally, to enhance the performance of CoBCReg by interleaving it with Query by Committee [8]. Combining semi-supervised learning and active learning within the Co-Training setting has been applied effectively for classification” [Hady page 130 Conclusions and Future Work]), as well as the suitability of additional confidence measures, beyond that disclosed in Hady, for selecting relevant examples (“Third, to explore other confidence measures that are more efficient and effective” [Hady page 130 Conclusions and Future Work]). A person of ordinary skill in the art would thereby recognize the value of incorporating the uncertainty, diversity, and density functions [Kee pages 404-405 Incorporating density and diversity], as taught by Kee, into the CoBCReg framework of Hady by modifying the algorithmic selection of relevant examples to prioritize selection of uncertain / least confident unlabeled examples, as is typical for an active learning framework. Incorporating these teachings would thereby enable utilization of the CoBCReg framework for an active learning training objective as suggested by Hady, and further boost model performance through inclusion of diversity and density terms for example selection [Kee Abstract]. Regarding claim 3, the combination of Hady, Shi, and Kee teaches the limitations of parent claim 1, and Kee further teaches calculat[ing] a degree of appearance, and determin[ing], as an example for which error is significant, each error calculation example for which: a weighted sum of the degree of appearance is significant; and the error is significant (“Unlike diversity which is generally incorporated in BAL, density has been often overlooked, but it is necessary to take into account density to prevent outliers in queries and label more representative instances which may improve uncertainty sampling further. Some studies incorporate density consideration with uncertainty under serial AL setting. One approach is to utilize density estimation (DE) methods. … Both DE and ER approaches assume that instances from denser regions are more informative in estimating the decision boundary, yet how they compare denser regions are different. DE approaches evaluate the information of an instance based on the actual density at the instance, p ( x ), derived from the underlying probability distribution in the feature space...Introducing a density factor extends Eq. (6) to the following form PNG media_image9.png 31 377 media_image9.png Greyscale where f ( x ), d ( x ), and h ( x ) are the uncertainty, diversity, and density functions, respectively, and 0 ≤λ, β ≤1 such that λ + β ≤ 1 control the relative importance” [Kee page 404-405 Incorporating diversity and density]; see line 12 in Algorithm 3 QDD: Batch active learning with query-by-committee, diversity and density – PNG media_image10.png 22 592 media_image10.png Greyscale ” [Kee page 407]; A related unlabeled example s may be selected through adjusted uncertainty measure u(x), which is a sum (weighted through importance terms λ and B) of uncertainty f(x), diversity d(x) and density h(x), wherein density h(x) is a confidence measure (i.e., degree of appearance) drawn from the underlying distribution of the data) Regarding claim 5, the combination of Hady, Shi, and Kee discloses the limitations of parent claim 1, and Hady further teaches generat[ing] the one or more student models, (see lines 1-3 of Algorithm 1. COBC for Regression [Hady page 123] as detailed in claim 1 above; The disclosed algorithm initially generates each RBF network regressor (i.e., student model) from a respective set of labeled training examples Li) and calculat[ing], using a remaining part of the training examples as the error calculation examples, the errors between the predictions by the one or more student models and the predictions by the teacher models (see line 2 of Algorithm 1 and lines 2-6 of Algorithm 2 [Hady page 123] as detailed in claim 1 above; Examples of validation set Vj used to calculate errors are out-of-bag, i.e., separate from bagging examples Li used to initially generate ensemble members). Regarding claim 6, the combination of Hady, Shi, and Kee discloses the limitations of parent claim 1, and Hady further teaches generat[ing] a plurality of sample groups by random sampling with duplicates from the training examples (see line 2 of Algorithm 1 – PNG media_image11.png 38 885 media_image11.png Greyscale [Hady page 123]), generat[ing] the one or more student models using respective sampling groups (see line 3 of Algorithm 1 – PNG media_image12.png 26 337 media_image12.png Greyscale [Hady page 123]), calculate, for each of the one or more student models, using as the error calculation examples, samples included in the training examples but not included in the sample groups, the errors between the predictions by the one or more student models and the predictions by the teacher model, (see line 2 of Algorithm 1 and lines 2-6 of Algorithm 2 [Hady page 123] as detailed in claim 1 above; Examples of validation set Vj used to calculate errors are out-of-bag, i.e., separate from bagging examples Li used to initially generate ensemble members), and calculate as the errors between the predictions of the one or more student models and the predictions of the teacher model, an average of the errors calculated for the one or more student models (“Figure 1 shows the RMSE of CoBCReg (CoBCReg), and the average of the RMSEs of the four regressors used in CoBCReg (RBFNNs) at the different SSL iterations. The dash and solid horizontal lines show the average of the RMSEs of the four regressors and the RMSE of the ensemble trained using only the 10% labeled data, respectively, as a basline for the comparison” [Hady page 217 Results]; see Figure 1 [Hady page 218]). Regarding claims 8 and 9, they are method and product claims that largely correspond to the apparatus of claim 1, which is already disclosed by Hady as detailed above. Consequently, they are rejected for the same reasons as claim 1. Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over Hady, Shi, and Kee, as applied to claim 1 above, further in view of Nandi (“Sampling Based Methods for Class Imbalance in Datasets”, available online 15 May 2017). Regarding claim 4, the combination of Hady, Shi, and Kee teaches the limitations of parent claim 1. However, the combination does not expressly teach generat[ing] new error calculation examples by oversampling from the training examples. In the same field of endeavor, Nandi teaches a means of data re-sampling through bootstrapping techniques (“In practice, we can't travel to a parallel universe (...yet) and re-collect this data, but we can simulate this using the bootstrap method. The idea behind bootstrap is simple: If we resample points with replacement from our data, we can treat the re-sampled dataset as a new dataset we collected in a parallel universe. Using the bootstrap method, I can create 2,000 re-sampled datasets from our original data and compute the mean of each of these datasets” [Nandi pages 3-4]) that generates new examples by oversampling from existing examples (“We want to use the general principle of bootstrap to sample with replacement from our minority class, but we want to adjust each re-sampled value to avoid exact duplicates of our original data. This is where the Synthetic Minority Oversampling Technique (SMOTE) algorithm comes in. The SMOTE algorithm can be broken down into four steps: 1. Randomly pick a point from the minority class. 2. Compute the k-nearest neighbors (for some pre-specified k) for this point. 3. Add k new points somewhere between the chosen point and each of its neighbors” [Nandi page 7]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated generates new examples by oversampling from existing examples as taught by Nandi into the combination because both Hady and Nandi are directed towards data re-sampling through bootstrapping techniques. Hady expressly utilizes bootstrap sampling to generate sampling groups for each ensemble predictor, the sampling groups further including labeled out-of-bag samples for validation error calculation (“Fortunately, since the bootstrap sampling [5] is used to construct the committee, the out-of-bootstrap examples are considered for a more accurate estimate of ∈’ j” [Hady page 126]). Given that availability of labeled training data is commonly limited in real-life scenarios (“For regression tasks, labeling the examples for training is a time consuming, tedious and expensive process” [Hady page 129 Conclusions and Future Work]) incorporating the oversampling taught by Nandi into Hady would thereby address potential class imbalance in labeled examples and thereby improve representativeness of bootstrap sampled groups. Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Hady, Shi, and Kee, as applied to claim 1 above, further in view of Yang (“Active Learning Using Uncertainty Information”, available conference 2016). Regarding claim 7, Hady teaches the limitations of parent claim 1. However, Hady does not expressly teach calculating the errors using examples other than the training example as the error calculation examples. In the same field of endeavor, Yang discloses an ensemble learning framework that utilizes sets of labeled and unlabeled data to answer queries (“The second class, retraining free active learning, contains the remaining methods which not need repeatedly train the model for each unlabeled instance during one single selection. For example, uncertainty sampling and query-by-committee belong to this category…We concentrate on the pool-based active learning setting which assumes a large pool of unlabeled data along with a small set of labeled data already available [2].” [Yang page 2647]) that calculat[es] errors using examples other than the training example as the error calculation examples (“Firstly, let us introduce some preliminaries and notation. Let PNG media_image13.png 31 151 media_image13.png Greyscale represent the training data set that consists of m labeled instances and U be the pool of unlabeled instances” [Yang page 2647 Retraining-Based Active Learning]; “Expected error reduction has demonstrated its effectiveness on text classification domain [8]. There are also some followup work of EER contributed by other researchers [9] [10] [11]. EER aims to select the sample which will reduce the future generalization error. Since we can not see the test data, the unlabeled pool can be used as the validation set to predict the future test error” [Yang page 2647 Expected Error Reduction]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated calculat[ing] errors using examples other than the training example as the error calculation examples as taught by Yang into the combination because both Hady and Yang are both directed towards ensemble learning frameworks that utilizes sets of labeled and unlabeled data to answer queries. Given that availability of labeled training data is commonly limited in real-life scenarios (“For regression tasks, labeling the examples for training is a time consuming, tedious and expensive process” [Hady page 129 Conclusions and Future Work]) incorporating the teaching of Yang into Hady by modifying sampling of validation sets to consist of unlabeled examples would be beneficial in instances where all available labeled data is reserved for initial training of ensemble models. Response to Arguments The amendment filed 07/08/2026 has been entered. Applicant’s amendment with respect to resolving specification objections has been considered, and the objections are consequently withdrawn. Applicant’s amendment to the claims with respect to resolving objections and indefiniteness rejections under 35 U.S.C. 112(b) has been considered, and the objections and 112(b) rejections are consequently withdrawn. Applicant’s amendment to the claims with respect to resolving provisional non-statutory double patenting rejections has been considered, and the rejections are consequently withdrawn. Applicant’s amendment to the claims with respect to resolving non-eligible subject matter rejections under 35 U.S.C. 101 has been considered, and the rejections are consequently withdrawn. The remarks filed 07/08/2026 have been fully considered. Applicant’s remarks traversing the prior art rejections under 35 U.S.C. 102 and 35 U.S.C. 103 set forth in the office action mailed 04/08/2026, in view of claims 1 and 3-9 as amended, have been considered, but are moot because the new grounds of rejection set forth above does not rely on the reference(s) applied in the prior rejection of record for the subject matter being specifically challenged in applicant' s argument. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to VIJAY M BALAKRISHNAN whose telephone number is (571) 272-0455. The examiner can normally be reached 10am-5pm EST Mon-Thurs. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, JENNIFER WELCH can be reached on (571) 272-7212. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /V.M.B./ Examiner, Art Unit 2143 /JENNIFER N WELCH/Supervisory Patent Examiner, Art Unit 2143
Read full office action

Prosecution Timeline

May 16, 2023
Application Filed
Apr 08, 2026
Non-Final Rejection mailed — §102, §103
Jul 01, 2026
Applicant Interview (Telephonic)
Jul 01, 2026
Examiner Interview Summary
Jul 08, 2026
Response Filed
Sep 23, 2026
Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743623
INFORMATION PROCESSING DEVICE AND MACHINE LEARNING METHOD THAT OPTIMIZE A DECODING PROCESS USING BACK-PROPAGATION
4y 0m to grant Granted Sep 22, 2026
Patent 12731026
METHOD AND SYSTEM FOR PROGRAM SAMPLING USING NEURAL NETWORK
3y 10m to grant Granted Sep 08, 2026
Patent 12711407
REASONING METHOD BASED ON STRUCTURAL ATTENTION MECHANISM FOR KNOWLEDGE-BASED QUESTION ANSWERING AND COMPUTING APPARATUS FOR PERFORMING THE SAME
3y 8m to grant Granted Aug 18, 2026
Patent 12645933
Method and System for Training a Neural Network for Generating Universal Adversarial Perturbations
4y 7m to grant Granted Jun 02, 2026
Patent 12619871
INTERPRETABLE NEURAL NETWORK ARCHITECTURE USING CONTINUED FRACTIONS
3y 11m to grant Granted May 05, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
41%
Grant Probability
99%
With Interview (+73.3%)
3y 11m (~7m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 27 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month