Prosecution Insights
Last updated: October 02, 2026
Application No. 18/166,696

DIRECTIONAL DRIVERS OF DEEP LEARNING MODELS BASED ON MODEL GRADIENTS

Final Rejection §103
Filed
Feb 09, 2023
Examiner
WU, NICHOLAS S
Art Unit
2148
Tech Center
2100 — Computer Architecture & Software
Assignee
The Bank of New York Mellon
OA Round
2 (Final)
52%
Grant Probability
Moderate
3-4
OA Rounds
4m
Est. Remaining
84%
With Interview

Examiner Intelligence

Grants 52% of resolved cases
52%
Career Allowance Rate
33 granted / 63 resolved
-2.6% vs TC avg
Strong +31% interview lift
Without
With
+31.4%
Interview Lift
resolved cases with interview
Typical timeline
4y 0m
Avg Prosecution
17 currently pending
Career history
89
Total Applications
across all art units

Statute-Specific Performance

§101
25.0%
-15.0% vs TC avg
§103
53.5%
+13.5% vs TC avg
§102
3.8%
-36.2% vs TC avg
§112
16.7%
-23.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 63 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant's arguments filed 04/20/2026 have been fully considered but they are not fully persuasive. Regarding the 101 rejections, applicant’s arguments and amendments to the independent claims are persuasive and overcome the previous 101 rejections. Specifically, applicant’s amended limitation generate a data object storing the gradients, the data object comprising a plurality of dimensions including at least a temporal dimension and a feature dimension; collapse the temporal dimension of the data object to aggregate, based on the one or more groups of features defined by the group definition, the gradients in the data object to generate a reduced-dimension data object; and for each group of features from among the one or more groups of features: determine a directional driver based on the aggregated gradients in the reduced-dimension data object, the directional driver indicating an impact of the group of features on the model output provides a technical improvement because using a feature directional driver based on collapsed data objects improves feature-level analysis by reducing the complexity required for analyzing feature impact. See pg. 9-10 of “Remarks”: “As disclosed in the specification, conventional deep learning models frequently rely on large numbers of features, which creates a technical challenge because the resulting model outputs become difficult to interpret and the influence of underlying variables cannot easily be explained. In particular, when multiple derived features correspond to a single underlying variable, traditional feature-level analysis makes it difficult to determine the true impact of that variable on model predictions. Additionally, in time-series environments, dependencies across sequences of input data further complicate attempts to determine how variables affect predictions over time. To address these technical shortcomings, the disclosed system uses a structured gradient- aggregation architecture that groups features into variables and variable groups and aggregates gradients according to those predefined groupings. Specifically, gradients produced by the deep learning model are determined for individual features and stored in multi-dimensional tensor data structures that include temporal dimensions. The system then collapses the temporal dimension of the tensor and aggregates gradients according to hierarchical feature group definitions to determine directional drivers representing the direction and magnitude of each group's influence on model outputs. The resulting directional drivers provide both magnitude and direction of impact, enabling identification of how grouped features influence model outputs (see paras. [0006]-[0008]). This architecture improves the operation of machine-learning systems by enabling aggregate feature analysis that reduces feature complexity while preserving the underlying variable relationships represented in the model inputs. Instead of attempting to interpret large numbers of individual feature gradients, the system produces structured directional drivers associated with higher-level variables and variable groups. This significantly improves the interpretability and usability of deep learning predictions, particularly in high-dimensional and time-series environments.” Applicant’s amendments and corresponding arguments that the claimed invention provides a technical improvement to the field of feature analysis are persuasive. Therefore, the 101 rejections are withdrawn. Regarding the 103 rejections, applicant's arguments filed with respect to the prior art rejections have been fully considered but they are moot. Applicant has amended the claims to recite new combinations of limitations. Applicant's arguments are directed at the amendment. Please see below for new grounds of rejection, necessitated by Amendment. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-2, 5-7, 9-12, and 15-20 are rejected under 35 U.S.C. 103 as being unpatentable over Miller, et al., US Pre-Grant Publication US20250045439A1 (“Miller”) in view of Sikdar, et al., Non-Patent Literature “Integrated Directional Gradients: Feature Interaction Attribution for Neural NLP Models” (“Sikdar”) and further in view of Cerliani, Non-Patent Literature “Feature Importance with Time Series and Recurrent Neural Network” (“Cerliani”). Regarding claim 1, Miller discloses: A system, comprising: a processor programmed to: (Miller, ⁋94, “The computing device 400 can include a processor 402 that is communicatively coupled to a memory 404 [A system, comprising: a processor programmed to:].”). access a plurality of features and a group definition that specifies at least one or more groups of features; (Miller, ⁋48, “One or more families [and a group definition that specifies at least one or more groups of features;] of time-series transforms can be applied to the time-series data [access a plurality of features] for a predictor variable 124 to generate transformed time-series data instances 334.”). provide the plurality of features as input to a deep learning model trained to generate a model output based on a model function and the plurality of features, wherein the deep learning model, when executed, generates the model output; (Miller, ⁋48, “Each of the transformed time-series data instances 334 can be fed into one input node of the input layer 340 [provide the plurality of features as input to a deep learning model trained to generate a model output based on a model function and the plurality of features,]. Input nodes taking data instances for one family of transformations can be connected to one hidden node in the first hidden layer of the risk prediction model 120… Hidden nodes in the first hidden layer 350A are connected to the nodes in the second hidden layer 350B, which are further connected to the output layer 360 [wherein the deep learning model, when executed, generates the model output;].”). for each feature from among the plurality of features: obtain a gradient that represents a rate of change of the model function based on the feature; (Miller, ⁋49, “The explanatory data can indicate relationships between the time-series data instances [for each feature from among the plurality of features:] of the predictor variable and the output risk indicator or between the transformed time-series data instances and the output risk indicator…The explanatory data can be calculated using a points-below-max algorithm or an integrated gradients algorithm [obtain a gradient that represents a rate of change of the model function based on the feature;].”). and for each group of features from among the one or more groups of features: determine a directional driver based on the aggregated gradients,…the directional driver indicating an impact of the group of features on the model output. (Miller, ⁋85, “Alternatively, or additionally, an integrated gradients algorithm may be used to generate the explanatory data [and for each group of features from among the one or more groups of features: determine a directional driver based on the aggregated gradients,…]. The integrated gradients algorithm can involve a reference point consisting of an alternative set of input variable values (X′, Y′), which produce an alternative score F(X′, Y′).”, and Miller, ⁋49, “The explanatory data can indicate relationships between the time-series data instances of the predictor variable and the output risk indicator or between the transformed time-series data instances and the output risk indicator [the directional driver indicating an impact of the group of features on the model output.].”). While Miller teaches a time based machine learning system that determines which features impacts model performance, Miller does not explicitly teach: directional driver generate a data object storing the gradients, the data object comprising a plurality of dimensions including at least a temporal dimension and a feature dimension; collapse the temporal dimension of the data object to aggregate, based on the one or more groups of features defined by the group definition, the gradients in the data object to generate a reduced-dimension data object; Sikdar teaches directional driver (Sikdar, pg. 868 col. 2, “Thus we propose to use absolute value of IDG, which is the path integral of the directional gradient over the straight line path from the baseline b to the input x as the dividend of the feature group. Further, the sign of IDG may be used to signify the nature of contribution (positive or negative) to model output [directional driver].”). Miller and Sikdar are both in the same field of endeavor (i.e. model explainability). It would have been obvious for a person having ordinary skill in the art before the effective filing date of the claimed invention to combine Miller and Sikdar to teach the above limitation(s). The motivation for doing so is that knowing the positive or negative impact of a feature improves the performance of the model (cf. Sikdar, pg. 868 col. 2, “the sign of IDG may be used to signify the nature of contribution (positive or negative) to model output.”). While Miller in view of Sikdar teaches a time based machine learning system that determines which features positively or negatively impacts model performance, the combination does not explicitly teach: generate a data object storing the gradients, the data object comprising a plurality of dimensions including at least a temporal dimension and a feature dimension; collapse the temporal dimension of the data object to aggregate, based on the one or more groups of features defined by the group definition, the gradients in the data object to generate a reduced-dimension data object; Cerliani teaches: generate a data object storing the gradients, the data object comprising a plurality of dimensions including at least a temporal dimension and a feature dimension; (Cerliani, pg. 7-8, “Our idea is to check the contribution of each single input feature on the final prediction output. The contribution in our case is given by the value of the gradients obtained from the differentiation operation of the input sequences on the forecasts. With Tensorflow, the implementation of this method is only 3 steps: use the GradientTape object to capture the gradients on the input; [generate a data object storing the gradients,] get the gradients with tape.gradient: this operation produces gradients of the same shape of the single input sequence (time dimension x features) [the data object comprising a plurality of dimensions including at least a temporal dimension and a feature dimension;]; obtain the impact of each sequence feature as average over the time dimension.”). collapse the temporal dimension of the data object to aggregate, based on the one or more groups of features defined by the group definition, the gradients in the data object to generate a reduced-dimension data object; (Cerliani, pg. 7-8, “Our idea is to check the contribution of each single input feature on the final prediction output. The contribution in our case is given by the value of the gradients obtained from the differentiation operation of the input sequences on the forecasts. With Tensorflow, the implementation of this method is only 3 steps: use the GradientTape object to capture the gradients on the input; get the gradients with tape.gradient: this operation produces gradients of the same shape of the single input sequence (time dimension x features); obtain the impact of each sequence feature as average over the time dimension; averaging over time is interpreted as collapsing the temporal dimension (i.e. collapse the temporal dimension of the data object to aggregate, based on the one or more groups of features defined by the group definition, the gradients in the data object to generate a reduced-dimension data object;).”). Miller, Sikdar, and Cerliani are all in the same field of endeavor (i.e. feature importance). It would have been obvious for a person having ordinary skill in the art before the effective filing date of the claimed invention to combine Miller, Sikdar, and Cerliani to teach the above limitation(s). The motivation for doing so is that collapsing gradients along the time dimension provides the impact of each feature on the model (cf. Cerliani, pg. 8, “obtain the impact of each sequence feature as average over the time dimension.”). Regarding claim 2, Miller in view of Sikdar and Cerliani teaches the system of claim 1. Cerliani further teaches wherein the -plurality of dimensions further comprise batch size, the temporal dimension comprising input time steps corresponding to a time period, and the feature dimension comprising the plurality of features. (Cerliani, pg. 4, “To train our sequential neural network we rearrange the data properly as 3D sequences of dimension: sample x time dimension x features [wherein the -plurality of dimensions further comprise batch size, the temporal dimension comprising input time steps corresponding to a time period, and the feature dimension comprising the plurality of features.].”). Miller, in view of Sikdar, and Cerliani are both in the same field of endeavor (i.e. feature importance). It would have been obvious for a person having ordinary skill in the art before the effective filing date of the claimed invention to combine Miller, in view of Sikdar, and Cerliani to teach the above limitation(s). The motivation for doing is that considering time steps, batch size, and features aids in explaining time-constrained models (cf. Cerliani, pg. 3, “In this post, I investigate the decision taken by a neural network trained to forecast the future. I used a recurrent structure to automatically learn information also from the time dimension. With some simple steps, we extract all we needed to understand the output of our model.”). Regarding claim 5, Miller in view of Sikdar and Cerliani teaches the system of claim 1. Miller further teaches wherein the group definition specifies a hierarchical grouping of features comprising: a first level having one or more variable groups each comprising a plurality of variables; a second level having each variable from among the plurality of variables, each variable comprising a group of features; and a third level comprising the features. (Miller, ⁋48 and Figure 3, “One or more families of time-series transforms [a second level having each variable from among the plurality of variables, each variable comprising a group of features;] can be applied to the time-series data for a predictor variable 124 [wherein the group definition specifies a hierarchical grouping of features comprising: a first level having one or more variable groups each comprising a plurality of variables;] to generate transformed time-series data instances 334 [and a third level comprising the features.].”). Regarding claim 6, Miller in view of Sikdar and Cerliani teaches the system of claim 5. Sikdar also teaches the directional driver as seen in claim 1. Miller further teaches: wherein to aggregate, based on the one or more groups of features, the gradients obtained from the deep learning model, the processor is further programmed to: for each variable from among the plurality of variables, aggregate the gradients of the group of features pertaining to the variable; and for each variable group, aggregate the aggregate gradients of the plurality of variables pertaining to the variable group, (Miller, ⁋86, “Integrated gradients may be applied to a model with correlated input variables, including a model with multiple compound time-series transformations constructed as linear combinations of individual transformations [wherein to aggregate, based on the one or more groups of features, the gradients obtained from the deep learning model, the processor is further programmed to: for each variable from among the plurality of variables, aggregate the gradients of the group of features pertaining to the variable;]. Treating each of the compound transformations as an input variable in its own right, the integrated gradients algorithm can be applied to express the score difference as a sum of contributions from each of the input variables, including the compound time-series transformations [and for each variable group, aggregate the aggregate gradients of the plurality of variables pertaining to the variable group,].”). wherein each directional driver indicates an impact of each variable group on the model output. (Miller, ⁋85, “Alternatively, or additionally, an integrated gradients algorithm may be used to generate the explanatory data [wherein each directional driver]. The integrated gradients algorithm can involve a reference point consisting of an alternative set of input variable values (X′, Y′), which produce an alternative score F(X′, Y′).”, and Miller, ⁋49, “The explanatory data can indicate relationships between the time-series data instances of the predictor variable and the output risk indicator or between the transformed time-series data instances and the output risk indicator [indicates an impact of each variable group on the model output.].”). Regarding claim 7, Miller in view of Sikdar and Cerliani teaches the system of claim 6. Sikdar also teaches the directional driver as seen in claim 1. Miller further teaches: wherein the processor is further programmed to: generate an output report based on the directional drivers for the variable groups, (Miller, ⁋49, “The explanatory data may indicate an impact a predictor variable has or a group of predictor variables have on the value of the risk indicator, such as credit score (e.g., the relative impact of the predictor variable(s) on a risk indicator) [generate an output report based on the directional drivers for the variable groups,].”). the output report visually showing an impact of each variable group on the model output. (Miller, ⁋49, “The explanatory data may indicate an impact a predictor variable has or a group of predictor variables have on the value of the risk indicator, such as credit score (e.g., the relative impact of the predictor variable(s) on a risk indicator) [the output report…showing an impact of each variable group on the model output.].”, and Miller, ⁋51, “As discussed above with regard to FIG. 1 , the risk assessment computing system 130 can communicate with client computing systems 104, which may send risk assessment queries to the risk assessment server 118 to request risk assessment.”, and Miller, ⁋100, “Another example of an output device is the presentation device 412 depicted in FIG. 4 . A presentation device 412 can include any device or group of devices suitable for providing visual [visually], auditory, or other suitable sensory output…In some aspects, the presentation device 412 can include a remote client-computing device”). Regarding claim 9, Miller in view of Sikdar and Cerliani teaches the system of claim 1. Sikdar further teaches wherein each directional driver comprises a positive value or a negative value, and wherein the processor is further programmed to: determine, for each directional driver, whether the impact is positive or negative based on the positive value or the negative value. (Sikdar, pg. 868 col. 2, “Thus we propose to use absolute value of IDG, which is the path integral of the directional gradient over the straight line path from the baseline b to the input x as the dividend of the feature group. Further, the sign of IDG may be used to signify the nature of contribution (positive or negative) [wherein each directional driver comprises a positive value or a negative value, and wherein the processor is further programmed to:] to model output [determine, for each directional driver, whether the impact is positive or negative based on the positive value or the negative value.].”). It would have been obvious to one of ordinary skill in the art before the effective filling date of the present application to combine the teachings of Sikdar with the teachings of Miller and Cerliani for the same reasons disclosed in claim 1. Regarding claim 10, Miller in view of Sikdar and Cerliani teaches the system of claim 1. Sikdar further teaches wherein the processor is further programmed to: determine, for each directional driver, a magnitude of the impact based on a value of the directional driver. (Sikdar, pg. 868 col. 2, “The dividend of a group of features is distinct from its value and is the measure of the importance of the interaction of the features in the group…Thus we propose to use absolute value of IDG, which is the path integral of the directional gradient over the straight line path from the baseline b to the input x as the dividend of the feature group [wherein the processor is further programmed to: determine, for each directional driver, a magnitude of the impact based on a value of the directional driver.].”). It would have been obvious to one of ordinary skill in the art before the effective filling date of the present application to combine the teachings of Sikdar with the teachings of Miller and Cerliani for the same reasons disclosed in claim 1. Regarding claim 11, the claim is similar to claim 1 and rejected under the same rationales. Miller further teaches the additional limitations of by a processor (Miller, ⁋94, “The computing device 400 can include a processor 402 that is communicatively coupled to a memory 404 [by a processor].”). Regarding claim 12, the claim is similar to claim 2 and rejected under the same rationales. Regarding claims 15-19, the claims are similar to claims 5-7 and 9-10 and are rejected under the same rationales. Regarding claim 20, the claim is similar to claim 1 and is rejected under the same rationales. Miller further teaches the additional limitations of A non-transitory storage medium storing instructions that, when executed by a processor, programs the processor to: (Miller, 5, “a non-transitory computer-readable storage medium having program code that is executable by a processor device to cause a computing device to perform operations [A non-transitory storage medium storing instructions that, when executed by a processor, programs the processor to:]”). Claims 3-4 and 13-14 are rejected under 35 U.S.C. 103 as being unpatentable over Miller, et al., US Pre-Grant Publication US20250045439A1 (“Miller”) in view of Sikdar, et al., Non-Patent Literature “Integrated Directional Gradients: Feature Interaction Attribution for Neural NLP Models” (“Sikdar”) and further in view of Cerliani, Non-Patent Literature “Feature Importance with Time Series and Recurrent Neural Network” (“Cerliani”) and Brownlee, “A Gentle Introduction to Mini-Batch Gradient Descent and How to Configure Batch Size” (“Brownlee”). Regarding claim 3, Miller in view of Sikdar and Cerliani teaches the system of claim 2. Cerliani further teaches wherein to collapse the temporal dimension of the data object, the processor is further programmed to: collapse the data object storing the gradients from three dimensions into a two dimensional array shaped by the batch size and the plurality of features. (Cerliani, pg. 7-8, “Our idea is to check the contribution of each single input feature on the final prediction output. The contribution in our case is given by the value of the gradients obtained from the differentiation operation of the input sequences on the forecasts. With Tensorflow, the implementation of this method is only 3 steps: use the GradientTape object to capture the gradients on the input; get the gradients with tape.gradient: this operation produces gradients of the same shape of the single input sequence (time dimension x features); obtain the impact of each sequence feature as average over the time dimension [collapse the data object storing the gradients from three dimensions into a two dimensional array shaped by the batch size and the plurality of features.].”). It would have been obvious to one of ordinary skill in the art before the effective filling date of the present application to combine the teachings of Cerliani with the teachings of Miller and Sikdar for the same reasons disclosed in claim 1. While Cerliani teaches collapsing a gradient to a 2D matrix, the combination does not explicitly teach …shaped by the batch size…. Brownlee teaches …shaped by the batch size… (Brownlee, pg. 4, “Mini-batch gradient descent is a variation of the gradient descent algorithm that splits the training dataset into small batches that are used to calculate model error and update model coefficients […shaped by the batch size…].”). Miller, in view of Sikdar and Cerliani, and Brownlee are both in the same field of endeavor (i.e. machine learning). It would have been obvious for a person having ordinary skill in the art before the effective filing date of the claimed invention to combine Miller, in view of Sikdar and Cerliani, and Brownlee to teach the above limitation(s). The motivation for doing so is that including batch size in model update gradients can affect the amount of update steps seen thus improving visibility into model performance (cf. Brownlee, pg. 4, “The model update frequency is higher than batch gradient descent which allows for a more robust convergence, avoiding local minima.”). Regarding claim 4, Miller in view of Sikdar, Cerliani, and Brownlee teaches the system of claim 3. Cerliani further teaches wherein to collapse the data object, the processor is further programmed to: average the gradients across the time dimension. (Cerliani, pg. 7-8, “Our idea is to check the contribution of each single input feature on the final prediction output. The contribution in our case is given by the value of the gradients [gradients] obtained from the differentiation operation of the input sequences on the forecasts. With Tensorflow, the implementation of this method is only 3 steps: use the GradientTape object to capture the gradients on the input; get the gradients with tape.gradient: this operation produces gradients of the same shape of the single input sequence (time dimension x features); obtain the impact of each sequence feature as average over the time dimension [wherein to collapse the data object the processor is further programmed to: average the gradients across the time dimension.].”). It would have been obvious to one of ordinary skill in the art before the effective filling date of the present application to combine the teachings of Cerliani with the teachings of Miller, Sikdar, and Brownlee for the same reasons disclosed in claim 3. Regarding claims 13-14, the claims are similar to claims 3-4 and rejected under the same rationales. Allowable Subject Matter Claim 21 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for indication of allowable subject matter: Regarding claim 21, Below are the closest cited references, each of which disclose various aspects of the claimed invention: Srinivas, et al., “Full-gradient representation for neural network visualization” discloses teaching a system that shows the importance of having visual representations for input sensitivity for neural networks. The system teaches visualizing the importance of inputs based on a saliency map based on aggregated feature gradients. While Srinivas teaches using visualizations for gradient feature importance, Srinivas does not explicitly teach a visual representation generated from a time-series data structure that simultaneously displays (i) time-varying directional drivers for each variable group and (ii) relative magnitudes of the directional drivers across the plurality of variable groups Choo, et al., “Visual Analytics for Explainable Deep Learning” discloses the importance of having visualizations in the deep learning pipeline specifically for explainability and analysis portion of a model. Choo notes the importance that feature importance visualizations provide for identifying features that impact the model’s performance the most. While Choo teaches the benefits of including visualizations of features impact on the model’s performance, Choo does not explicitly teach gradient feature importance or a visual representation generated from a time-series data structure that simultaneously displays (i) time-varying directional drivers for each variable group and (ii) relative magnitudes of the directional drivers across the plurality of variable groups. Selvaraju, et al., “Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization” discloses a system that tries to visually explain the decision of an image classification CNN. The system does this by leveraging the gradient weighted class activation mapping to a class to highlight the important regions of an image when determining a class prediction. While Selvaraju teaches a system that considers gradient feature importance, Selvaraju is focused on image classification and not the temporal aspect of the gradient impact and therefore does not explicitly teach a visual representation generated from a time-series data structure that simultaneously displays (i) time-varying directional drivers for each variable group and (ii) relative magnitudes of the directional drivers across the plurality of variable groups. While the above prior arts disclose the aforementioned concepts, however, none of the prior arts, individually or in reasonable combination, discloses all the limitations in the manner recited in claim 21. Specifically, the claim recites: “a visual representation generated from the time-series data structure that simultaneously displays (i) time-varying directional drivers for each variable group and (ii) relative magnitudes of the directional drivers across the plurality of variable groups.” While the references cited above mention aspects of visually showing output reports for gradient feature importance, the output reports do not recite the specific visual representations of a visual representation generated from a time-series data structure that simultaneously displays (i) time-varying directional drivers for each variable group and (ii) relative magnitudes of the directional drivers across the plurality of variable groups. Therefore, claim 21 is allowable over the prior art. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Angelo, et al., US20220138532A1 discloses different techniques of feature attribution and highlights feature attribution using gradient based methods. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to NICHOLAS S WU whose telephone number is (571)270-0939. The examiner can normally be reached Monday - Friday 8:00 am - 4:00 pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michelle Bechtold can be reached at 571-431-0762. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /N.S.W./Examiner, Art Unit 2148 /MICHELLE T BECHTOLD/Supervisory Patent Examiner, Art Unit 2148
Read full office action

Prosecution Timeline

Feb 09, 2023
Application Filed
Dec 22, 2025
Non-Final Rejection mailed — §103
Mar 31, 2026
Examiner Interview Summary
Mar 31, 2026
Applicant Interview (Telephonic)
Apr 20, 2026
Response Filed
Jul 21, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737619
OPTIMIZING ALGORITHMS FOR HARDWARE DEVICES
3y 11m to grant Granted Sep 15, 2026
Patent 12725027
PROACTIVE ANOMALY DETECTION
5y 9m to grant Granted Sep 01, 2026
Patent 12725405
LEARNING APPARATUS, ESTIMATION APPARATUS, DATA GENERATION APPARATUS, LEARNING METHOD, AND COMPUTER-READABLE STORAGE MEDIUM STORING A LEARNING PROGRAM
5y 0m to grant Granted Sep 01, 2026
Patent 12645939
SPIKING NEURAL NETWORK
3y 5m to grant Granted Jun 02, 2026
Patent 12619880
METHODS, DEVICES AND MEDIA FOR RE-WEIGHTING TO IMPROVE KNOWLEDGE DISTILLATION
5y 0m to grant Granted May 05, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
52%
Grant Probability
84%
With Interview (+31.4%)
4y 0m (~4m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 63 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month