Prosecution Insights
Last updated: October 02, 2026
Application No. 18/621,773

REGULARIZING AND INTERPRETABILITY-ENHANCING LOSS FOR ATTENTION-BASED NEURAL NETWORKS

Final Rejection §102§103
Filed
Mar 29, 2024
Examiner
TRAN, VINCENT HUY
Art Unit
2115
Tech Center
2100 — Computer Architecture & Software
Assignee
Robert Bosch GmbH
OA Round
2 (Final)
87%
Grant Probability
Favorable
3-4
OA Rounds
1m
Est. Remaining
96%
With Interview

Examiner Intelligence

Grants 87% — above average
87%
Career Allowance Rate
970 granted / 1120 resolved
+31.6% vs TC avg
Moderate +10% lift
Without
With
+9.7%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
23 currently pending
Career history
1143
Total Applications
across all art units

Statute-Specific Performance

§101
8.4%
-31.6% vs TC avg
§103
44.5%
+4.5% vs TC avg
§102
26.5%
-13.5% vs TC avg
§112
10.5%
-29.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1120 resolved cases

Office Action

§102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims 1-20 are pending in the application. Examiner’s Note: The examiner has cited particular passages including column and line numbers, paragraphs as designated numerically and/or figures as designated numerically in the references as applied to the claims below for the convenience of the applicant. Although the specified citations are representative of the teachings in the art and are applied to the specific limitations within the individual claims, other passages, paragraphs and figures of any and all cited prior art references may apply as well. It is respectfully requested from the applicant, in preparing an eventual response, to fully consider the context of the passages, paragraphs and figures as taught by the prior art and/or cited by the examiner while including in such consideration the cited prior art references in their entirety as potentially teaching all or part of the claimed invention. MPEP 2141.02 VI: “PRIOR ART MUST BE CONSIDERED IN ITS ENTIRETY, INCLUDING DISCLOSURES THAT TEACH AWAY FROM THE CLAIMS." Information Disclosure Statement The information disclosure statement (IDS) submitted on 06/11/2026, 08/17/2026 was filed after the mailing date of the first office action. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 1 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Max-Heinrich Laves “Calibration of Model Uncertainty for Dropout Variational Inference” 20 Jun 2020 (“Laves”). Regarding claim 1, Laves discloses A non-transitory computer-readable medium having computer-readable instructions stored thereon, the computer-readable instructions operable by a processor to train a machine learning model, the instructions operable to perform the following functions: receive input data; given a training data set D of labeled images and an unseen test image x with class label y, we are interested in evaluating the predictive distribution [page 2 and read further Section 4. Experiments on page 6] process the input data through a machine learning layer that applies a softmax function to provide a softmax output [p]; Let fw(x) be the out-put (logits) of a neural network with weight matrices w, and with model likelihood p(y = c | fw(x)) for class c, which is sampled from a probability vector p = sSM(fw(x)), obtained by passing the model output through the softmax function sSM(.) [page 2-3] determine a training loss [equation 28] including a regularizing loss term derived from the softmax output; and Pereyra et al. link label smoothing to confidence penalty and propose a simple way to prevent overconfident networks (Pereyra et al., 2017). Low entropy output distributions are penalized by adding the negative entropy to the training objective. [page 2] Additionally, we compare temperature scaling to entropy regularization, where low entropy output distributions are penalized by adding the negative entropy H of the softmax output to the negative log-likelihood training objective, weighted by an additional hyperparameter. This leads to the following optimization function: [page 6] PNG media_image1.png 58 448 media_image1.png Greyscale train the machine learning model on the training loss. First, fw is trained with Gaussian dropout until convergence on the training set [page 5]. We additionally train all networks in the exact same manner with confidence penalty loss with fixed β = 0.1. [page 6] Calibration by confidence penalty must be performed during the training and cannot be done afterwards. [page 6] Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 2 is/are rejected under 35 U.S.C. 103 as being unpatentable over Laves as applied to claim 1 above, and further in view of Xiao. Regarding claim 2, Laves does not teach softmax function is characterized by formula PNG media_image2.png 41 209 media_image2.png Greyscale Xiao is directed to transformer-based language models and attention mechanisms. Xiao identified a problem resulting from the conventional SoftMax operation in attention and proposes an alternative SoftMax function [page 9]. Xiao specifically teaches PNG media_image2.png 41 209 media_image2.png Greyscale Xiao further explains that SaftMax1 does not require the attention scores on all contextual tokens to sum up to one and that SoftMax1 is equivalent to using a token having all-zero key and Value features in the attention computation [page 5]. Xiao further teaches actually using SoftMax1 in a machine-learning model. Xiao reports pre-training language modes in which a model replaced the regular attention mechanism in with SoftMax1 [page 10]. Thus, SoftMax1 is not merely a mathematical alternative disclosed in the abstract, it discloses as an operational attention function in a trained language model. Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to modify the neural-network training and confidence-penalty technique of Laves to employ the SoftMax1 attention function disclosed by Xiao. A person of ordinary skill in the art would have been motivated to make this modification because Xiao identifies a limitation of conventional SoftMax in attention mechanism: conventional Softmax forces attention scores over contextual tokens to sum to one, thereby causing the model to assign attention to token even when the tokens are not semantically relevant. Xiao proposes SoftMax1 specifically to address this problem by adding the “1” term to the denominator and thereby allowing the attention scores on the contextual tokens to sum to less than one. On ordinary skill in the art would recognize that the SoftMax1 function of Xiao is a known alternative parameterization of the SoftMax operation and could be substituted for the conventional softmax operation used by Laves while retaining the neural-network training framework and the confidence penalty regularization of Laves. Therefore, the combination of Laves and Xiao would have yielded the claimed method with a reasonable expectation of success. Claim(s) 3 is/are rejected under 35 U.S.C. 103 as being unpatentable over Laves as applied to claim 1 above, and further in view of Bondarenko et al. US Pub. No. 2024/0386239 (“Bondarenko”). Regarding claim 3, Laves teaches a method of processing input data using a neural network model and applying a softmax function to an output model. In particular, Laves explains that a probability vector is obtained by passing a model output through a softmax function. Laves further teaches penalizing the softmax output using a confidence penalty. However, Laves does not expressly teach wherein softmax function is applied by a plurality of heads of a multi-head attention layer where each head corresponds to a parallel linear layer respectively are represented by Q, K, V such that the softmax function is applied on QKT as represented by formula (6): Sof tmax(QKT) (6). Bondarenko teaches a transformer having a self-attention block. Bondarenko explains that input data is linearly projected into three matrices: a query matrix Q, a key matrix K, and a value Matrix V. Specifically, Bondarenko teaches softmax function is applied by a plurality of heads of a multi-head attention layer where each head corresponds to a parallel linear layer respectively are represented by Q, K, V such that the softmax function is applied on QKT as represented by formula (6): Sof tmax(QKT) (6) [See fig. 1]. [0022] Generally, the transformer 110 includes a self-attention block 120 (labeled “SA”) and a feedforward block 140 (labeled “FF”). In the self-attention block 120, the input data 105 may be linearly projected (e.g., multiplied using learned parameters) into three matrices: a query matrix Q 122 (also referred to in some aspects as a “query representation” or simply “queries”), a key matrix K 124 (also referred to in some aspects as a “key representation” or simply “keys”), and a value matrix V 126 (also referred to in some aspects as a “value representation” or simply “values”). For example, during training, one or more query weights, key weights, and value weights are learned based on training data, and the queries Q 122, the keys K 124, and the values V 126 can be generated by multiplying the input data by the learned weights. [0023] In some aspects, an attention matrix A (also referred to as an “attention map” or simply “attention” in some aspects) is then generated as an output of an attention block 130 based on the queries and keys. For example, the self-attention block 120 may, at a combiner 128, compute the dot product of the query matrix and the transposed key matrix (e.g., Q.Math.K.sup.T). In some aspects, the attention block 130 can apply one or more operations (e.g., a row-wise softmax operation) to the dot product generated by the combiner 128 to yield the attention matrix A. That is, the attention matrix A generated by the attention block 130 may be defined as A=σ(Q.Math.K.sup.T), where σ corresponds to a regularizing function usable in a transformer neural network, such as a softmax function or the like. Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to modify the neural networking processing of Laves using the transformer self-attention architecture of Bondarenko. Both references concern neural network processing in which an input is processed through learning neural network transformation and a softmax operation is used to generate probability-like output values. Laves expressly teaches softmax-based processing and penalization of resulting probability distribution. Bondarenko teaches a known transformer implementation in which neural network input is linearly projected into Q, K, and V representations, the Q and K representations are combined as QKT, a softmax operation is applied to produce an attention matrix. One of ordinary skill in the art would motivated to employ the self-attention architecture of Bondarenko in the neural network processing of Laves because doing so would provide the recognized advantages of transformer-based attention, including allowing the model to determine relationships among different portions of the input data through query-key attention. Such modification would have amounted to applying a well-known neural network architecture and its conventional Q/K/V attention mechanism to Laves’s softmax-based neural network processing, with a reasonable expectation of success. Claim(s) 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Laves in view of Guangxuan Xiao et al. “EFFICIENT STREAMING LANGUAGE MODELS WITH ATTENTION SINKS” submitted on 20 Sep 2023 (“Xiao”). Regarding claim 13, Laves teaches training a machine learning model, the method comprising: receiving input data; given a training data set D of labeled images and an unseen test image x with class label y, we are interested in evaluating the predictive distribution [page 2 and read further Section 4. Experiments on page 6] processing the input data through a machine learning layer that applies a softmax function to the input data to provide output data; Let fw(x) be the out-put (logits) of a neural network with weight matrices w, and with model likelihood p(y = c | fw(x)) for class c, which is sampled from a probability vector p = sSM(fw(x)), obtained by passing the model output through the softmax function sSM(.) [page 2-3] determining a training loss including a regularizing loss term derived from the output data; and Pereyra et al. link label smoothing to confidence penalty and propose a simple way to prevent overconfident networks (Pereyra et al., 2017). Low entropy output distributions are penalized by adding the negative entropy to the training objective. [page 2] Additionally, we compare temperature scaling to entropy regularization, where low entropy output distributions are penalized by adding the negative entropy H of the softmax output to the negative log-likelihood training objective, weighted by an additional hyperparameter. This leads to the following optimization function: [page 6] PNG media_image1.png 58 448 media_image1.png Greyscale training the machine learning model on the training loss. First, fw is trained with Gaussian dropout until convergence on the training set [page 5]. We additionally train all networks in the exact same manner with confidence penalty loss with fixed β = 0.1. [page 6] Calibration by confidence penalty must be performed during the training and cannot be done afterwards. [page 6] Laves does not teach softmax function represented by formula PNG media_image2.png 41 209 media_image2.png Greyscale Xiao is directed to transformer-based language models and attention mechanisms. Xiao identified a problem resulting from the conventional SoftMax operation in attention and proposes an alternative SoftMax function [page 9]. Xiao specifically teaches PNG media_image2.png 41 209 media_image2.png Greyscale Xiao further explains that SaftMax1 does not require the attention scores on all contextual tokens to sum up to one and that SoftMax1 is equivalent to using a token having all-zero key and Value features in the attention computation [page 5]. Xiao further teaches actually using SoftMax1 in a machine-learning model. Xiao reports pre-training language modes in which a model replaced the regular attention mechanism in with SoftMax1 [page 10]. Thus, SoftMax1 is not merely a mathematical alternative disclosed in the abstract, it discloses as an operational attention function in a trained language model. Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to modify the neural-network training and confidence-penalty technique of Laves to employ the SoftMax1 attention function disclosed by Xiao. A person of ordinary skill in the art would have been motivated to make this modification because Xiao identifies a limitation of conventional SoftMax in attention mechanism: conventional Softmax forces attention scores over contextual tokens to sum to one, thereby causing the model to assign attention to token even when the tokens are not semantically relevant. Xiao proposes SoftMax1 specifically to address this problem by adding the “1” term to the denominator and thereby allowing the attention scores on the contextual tokens to sum to less than one. On ordinary skill in the art would recognize that the SoftMax1 function of Xiao is a known alternative parameterization of the SoftMax operation and could be substituted for the conventional softmax operation used by Laves while retaining the neural-network training framework and the confidence penalty regularization of Laves. Therefore, the combination of Laves and Xiao would have yielded the claimed method with a reasonable expectation of success. Claim(s) 14-15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Laves/Xiao as applied to claim 13 above, and further in view of Bondarenko et al. US Pub. No. 2024/0386239 (“Bondarenko”). Regarding claim 14, Laves teaches with T = diag(t1,……, tC). Auxiliary scaling makes use of a more powerful auxiliary recalibration model R consisting of a two-layer fully-connected network with C hidden units and leaky ReLU activations after the hidden layer. Laves/Xiao does not teach passing the input data through a plurality of linear layers prior to applying the softmax1 function. Bondarenko teaches passing the input data through a plurality of linear layers prior to applying the softmax.sub.1 function [See par. 0016-0018, 0024]. Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to modify the neural network processing of Laves to include passing the input data through a plurality of linear layers prior to applying the softmax1 function of Bondarenko. One of ordinary skill in the art would motivated to make this modification because the transformer self-attention architecture of Bondarenko provides a known mechanism for processing relationships among portions of input data through multiple learned linear projections and attention weighting. Incorporating such linear projections into Laves’s neural network processing would therefore have provided the predictable benefit of allowing the softmax operation to operate on attention-derived representations of the input data. Regarding claim 15, Bondarenko discloses the softmax.sub.1 function is applied through multi-head attention layer [See fig. 1]. Claim(s) 7-8, 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Laves as applied to claim 1 or 13 above, and further in view of Cherian US Pub. No. 2024/0241508. Regarding claim 7, Laves teaches advances in deep learning have led to high accuracy predictions for classification tasks, making deep-learning classifiers an attractive choice for safety-critical applications like computer-aided diagnosis (Esteva et al., 2017). Laves does not teach the input data is tabular manufacturing data. Cherian teaches a system comprises one or multiple tools to perform one or multiple tasks. The anomaly detector collects a feedforward signal indicative of a sequence of control inputs to the plurality of actuators and a feedback signal indicative of a sequence of outputs of the system caused by the plurality of actuators operated based on the sequence of control inputs. Cherian further teaches a system includes an attention model [See fig. 2] for the plurality of control steps of the manufacturing process of the system 102. The block diagram 200D includes an expansion of the attention layer 256 of the attention model 222. The attention layer 256 includes an alignment layer 258, a softmax layer 260, and a context layer 262. Specifically, Cherian teaches the input data is tabular manufacturing data [SEE fig. 9; par. 0078-0083]. Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the teachings of the cited reference because they both directed to the method and system of applying neural network in order to identify a set of parameters that results in a model that achieves a target level of performance. Cherian teaches the input data is tabular manufacturing data would further help the system of Laves to apply in the field of automated control to control the performance of a manufacturing process to reduce waste material, cause downtimes, decrease output. Regarding claim 8, Cherian teaches the tabular manufacturing data includes a plurality of measurement entries, each column of the tabular manufacturing data corresponding to a manufacturing station and/or properties therefrom, and each row of the tabular manufacturing data corresponding to a different product of manufacture [par. 0002, 0074, 0103]. Regarding claim 20, Cherian teaches the tabular manufacturing data includes a plurality of measurement entries, each column of the tabular manufacturing data corresponding to a manufacturing station and/or properties therefrom, and each row of the tabular manufacturing data corresponding to a different product of manufacture [par. 0002, 0074, 0103]. Claim(s) 17-18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Laves/Xiao as applied to claim 13 above, and further in view of Tu, Xuan CN 111680788 A (“Tu”). Regarding claim 17, Laves discussed how model uncertainty is obtained by Monte Carlo Gaussian dropout and how it can be calibrated with logit scaling. Laves/Xiao does not teach a dropout is applied after the softmax function. Tu teaches a diagnosis method based on deep learning, aiming at the characteristic of the device data, adding parameter regularization and Dropout method to improve the generalization capability of the model on the original basic convolutional neural network model. Specifically, Tu teaches a dropout is applied after the softmax function. S3 constructing a basic convolutional neural network fault diagnosis model, and using the test sample to optimize the network structure layer, specifically comprising: the basic structure of the convolutional neural network is shown in FIG. 3; the network is composed of a convolutional layer, a pool layer and a full connection layer… Finally, it is a fully connected output layer, the activation function is softmax function, finally outputting the classification result. S4 using regularization, Dropout method to optimize the structure layer optimization of the convolutional neural network internal structure, specifically comprising: regularization by the regression loss function of the model plus a constraint form, effectively reducing the complexity of the model, inhibiting over-fitting, improving the adaptability of the model. L1, L2 regularization method is the most common parameter regularization method, isa norm constraint. L1 is generally also called Lasso regression, adding the absolute value of the parameter, L2 is ridge regression (Ridge regression), operation is the square of each parameter and then calculating square root, both are the regression loss function is added with a constraint form. the difference is that L1 produces little characteristic; other features are 0, L2 regularization retains more features, which makes them approach to 0, reduces the difference between the parameters, the parameter becomes more smooth, model can adapt more data set, so the application adopts L2 regularization method [page. 6-8] Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to modify the method of Bondarenko with the step of a dropout is applied after the softmax function of Tu. The motivation for doing so would have been, as suggested by Tu on page 7, to effectively solve the over-fitting problem caused by too much depth neural network parameter, no more overly dependent on local features, reinforcing model generalization ability. Regarding claim 18, Tu teaches the dropout is greater than 0.3 [page 7 - The Dropout size of the present application is set to 0.5]. Allowable Subject Matter Claims 9-12 allowed. Claims 4-6, 16, 19 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to VINCENT HUY TRAN whose telephone number is (571)272-7210. The examiner can normally be reached M-F 7:00-4:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kamini S Shah can be reached at 571-272-2279. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. VINCENT H TRAN Primary Examiner Art Unit 2115 /VINCENT H TRAN/Primary Examiner, Art Unit 2115
Read full office action

Prosecution Timeline

Mar 29, 2024
Application Filed
Apr 28, 2026
Non-Final Rejection mailed — §102, §103
Jul 22, 2026
Applicant Interview (Telephonic)
Jul 23, 2026
Examiner Interview Summary
Jul 28, 2026
Response Filed
Sep 08, 2026
Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12746722
DATA PROCESSING DEVICE FOR GENERATING MICROSTRUCTURES HAVING CONTROLLABLE DEFORMABLE PROPERTIES
3y 0m to grant Granted Sep 29, 2026
Patent 12749030
PREDICTING POWER GENERATION OF A RENEWABLE ENERGY INSTALLATION
3y 0m to grant Granted Sep 29, 2026
Patent 12741743
SYSTEM AND METHOD FOR CONTROLLING AN AIRCRAFT SEAT AND ITS ENVIRONMENT VIA A WIRELESS CONNECTION
4y 1m to grant Granted Sep 22, 2026
Patent 12735255
ARTICLE DELIVERY SYSTEM AND METHOD
3y 0m to grant Granted Sep 15, 2026
Patent 12729677
SENSORS, MULTIPLEXED COMMUNICATION TECHNIQUES, AND RELATED SYSTEMS
3y 3m to grant Granted Sep 08, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
87%
Grant Probability
96%
With Interview (+9.7%)
2y 7m (~1m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 1120 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month