Prosecution Insights
Last updated: August 01, 2026
Application No. 18/176,037

TRANSFORMER NETWORK WITH NORMALIZATION INCLUDING SCALING PARAMETER

Final Rejection §101§103
Filed
Feb 28, 2023
Examiner
SIPPEL, MOLLY CLARKE
Art Unit
2122
Tech Center
2100 — Computer Architecture & Software
Assignee
Microsoft Technology Licensing, LLC
OA Round
2 (Final)
52%
Grant Probability
Moderate
3-4
OA Rounds
4m
Est. Remaining
78%
With Interview

Examiner Intelligence

Grants 52% of resolved cases
52%
Career Allowance Rate
12 granted / 23 resolved
-2.8% vs TC avg
Strong +26% interview lift
Without
With
+26.1%
Interview Lift
resolved cases with interview
Typical timeline
3y 9m
Avg Prosecution
19 currently pending
Career history
41
Total Applications
across all art units

Statute-Specific Performance

§101
26.5%
-13.5% vs TC avg
§103
52.9%
+12.9% vs TC avg
§102
2.0%
-38.0% vs TC avg
§112
18.6%
-21.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 23 resolved cases

Office Action

§101 §103
DETAILED ACTION This action is responsive to the amendment filed on 03/02/2026. Claims 1-20 are currently pending in the case. Claims 1, 10, and 19 are currently amended. Claims 1, 10, and 19 are independent claims. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) submitted on 12/29/2025 is being considered by the examiner. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Regarding claim 1: Step 1 Statutory Category: Claim 1 is directed to a system, which falls under one of the four statutory categories. Step 2A Prong 1 Judicial Exception: Claim 1 recites, in part, “wherein each of the plurality of normalization sub-layers is configured to apply layer normalization to a sum of: a first scaling parameter multiplied by an input vector of the sub-layer, …; and an output vector of the sub-layer, such that the first scaling parameter weights respective contributions of the input vector and the output vector to the sum on which layer normalization is performed”. This limitation, under the broadest reasonable interpretation, covers the recitation of a mathematical calculation, as directed to “a claim that recites a mathematical calculation, when the claim is given its broadest reasonable interpretation in light of the specification, will be considered as falling within the "mathematical concepts" grouping. A mathematical calculation is a mathematical operation (such as multiplication) or an act of calculating using mathematical methods to determine a variable or number”. See MPEP §2106.04(a)(2)(I)(C). Step 2A Prong 2 Integration into a Practical Application: This judicial exception is not integrated into a practical application. In particular, the claim recites: “a computing system” and “a processor”. These limitations are additional elements that amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f). Further, the claim recites: “receive a training set”. This limitation amounts to mere data gathering. It is necessary to acquire the data in order to use the recited judicial exception. Therefore, this limitation is insignificant extra-solution activity to the judicial exception, see MPEP §2106.05(g). Further, the claim recites: “based at least in part on the training data set, train a transformer network that includes a plurality of layers”. This limitation is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f). Further, the claim recites: “wherein the plurality of layers each respectively include a plurality of sub-layers including: an attention sub-layer; a feed-forward sub-layer; and a plurality of normalization sub-layers downstream from corresponding sub-layers of the plurality of sub-layers”. This limitation is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Further, the claim recites: “wherein the input vector is passed to the normalization sub-layer in a residual stream, of the transformer network”. This limitation is an additional element that amounts to insignificant extra-solution activity to the judicial exception, see MPEP §2106.05(g). Step 2B Significantly More: The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements: “a computing system”, “a processor”, and “based at least in part on the training data set, train a transformer network that includes a plurality of layers” amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. Elements that merely amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process cannot provide an inventive concept. Further, the additional element “receive a training set” is insignificant extra-solution activity to the judicial exception and is directed to receiving or transmitting data over a network which courts have recognized as well-understood, routine, and conventional when they are claimed in a generic manner, see MPEP §2106.05(d)(II). Further, the additional element “wherein the plurality of layers each respectively include a plurality of sub-layers including: an attention sub-layer; a feed-forward sub-layer; and a plurality of normalization sub-layers downstream from corresponding sub-layers of the plurality of sub-layers” generally links the use of the judicial exception to a particular technological environment or field of use. Elements that merely generally link the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. Finally, the additional element “wherein the input vector is passed to the normalization sub-layer in a residual stream, of the transformer network” amounts to insignificant extra-solution activity to the judicial exception and further, the additional element is well‐understood, routine, and conventional as taught by activity is supported under Berkheimer Option 2, Dey et al., U.S. Patent Application Publication No. 20220292266, Paragraph 0002, Lines 21-28, “The original Transformer design uses L-layers in the encoder 110, where L=6 to perform sequential operations (e.g., 110_1, 110_2, . . . , 110_L, where L=6) and L layers for the decoder 102. Layers of encoder 110 are processed by Add and Norm components 112, 114 via skip feeds 115, 116 to perform residual connection followed by layer normalization, which are well known functions applied in deep architectures”; See also, Figure 1. The claim is not patent eligible. Regarding claim 2, the rejection of claim 1 is incorporated, and further, the claim recites: “wherein, at each of the plurality of layers, the processor is further configured to scale a plurality of value projection weights and a plurality of output projection weights of the attention sub-layer and a plurality of feed-forward weights of the feed-forward sub-layer by a second scaling parameter when training the transformer network”. This limitation recites mathematical concepts in addition to those identified in the rejection of the parent claim. Thus, the claim recites a judicial exception. The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible. Regarding claim 3, the rejection of claim 2 is incorporated, and further, the claim recites: “wherein the processor is further configured to determine the first scaling parameter and the second scaling parameter based at least in part on a number of the plurality of layers”. This limitation recites mathematical concepts in addition to those identified in the rejection of the parent claim. Thus, the claim recites a judicial exception. The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible. Regarding claim 4, the rejection of claim 3 is incorporated, and further, the claim recites: “the processor is further configured to determine the first scaling parameter and the second scaling parameter based at least in part on whether or not the transformer network includes both an encoder and a decoder”. This limitation is a continuation of the “wherein the processor is further configured to determine the first scaling parameter and the second scaling parameter based at least in part on a number of the plurality of layers” limitation of the parent claim. Thus, the claim recites a judicial exception. Further, the claim recites: “the transformer network includes an encoder and/or a decoder”. This limitation is an additional element that amounts to generally linking the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Elements that merely generally link the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. The claim is not patent eligible. Regarding claim 5, the rejection of claim 4 is incorporated, and further, the claim recites: “wherein: the transformer network includes the encoder without including the decoder or includes the decoder without including the encoder; the first scaling parameter is equal to ( 2 N ) 1 4 , where N is the number of the plurality of layers; and the second scaling parameter is equal to ( 8 N ) - 1 4 ”. This limitation recites mathematical concepts in addition to those identified in the rejection of the parent claim. Thus, the claim recites a judicial exception. The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible. Regarding claim 6, the rejection of claim 4 is incorporated, and further, the claim recites: “the first scaling parameter and the second scaling parameter differ between the encoder and the decoder”. This limitation is a continuation of the “the processor is further configured to determine the first scaling parameter and the second scaling parameter based at least in part on whether or not the transformer network includes both an encoder and a decoder” limitation of the parent claim. Thus, the claim recites a judicial exception. Further, the claim recites: “the transformer network includes both the encoder and the decoder”. This limitation is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Elements that merely generally link the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. The claim is not patent eligible. Regarding claim 7, the rejection of claim 6 is incorporated, and further, the claim recites: “at the encoder: the first scaling parameter is equal to 0.81 ( N 4 M ) 1 16 , where N is a number of encoder layers included in the encoder and M is a number of decoder layers included in the decoder; and the second scaling parameter is equal to 0.87 ( N 4 M ) - 1 16 ; and at the decoder: the first scaling parameter is equal to ( 3 M ) 1 4 ; and the second scaling parameter is equal to ( 12 M ) - 1 4 ”. This limitation recites mathematical concepts in addition to those identified in the rejection of the parent claim. Thus, the claim recites a judicial exception. The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible. Regarding claim 8, the rejection of claim 1 is incorporated, and further, the claim recites: “wherein the transformer network includes 100 or more layers”. This limitation is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Elements that merely generally link the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. The claim is not patent eligible. Regarding claim 9, the rejection of claim 1 is incorporated, and further, the claim recites: “wherein the transformer network is a machine translation model”. This limitation is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Elements that merely generally link the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. The claim is not patent eligible. Regarding claim 10: Step 1 Statutory Category: Claim 10 is directed to a method, which falls under one of the four statutory categories. Step 2A Prong 1 Judicial Exception: Claim 10 recites, in part, “wherein each of the plurality of normalization sub-layers is configured to apply layer normalization to a sum of: a first scaling parameter multiplied by an input vector of the sub-layer, …; and an output vector of the sub-layer, such that the first scaling parameter weights respective contributions of the input vector and the output vector to the sum on which layer normalization is performed”. This limitation, under the broadest reasonable interpretation, covers the recitation of a mathematical calculation, as directed to “a claim that recites a mathematical calculation, when the claim is given its broadest reasonable interpretation in light of the specification, will be considered as falling within the "mathematical concepts" grouping. A mathematical calculation is a mathematical operation (such as multiplication) or an act of calculating using mathematical methods to determine a variable or number”. See MPEP §2106.04(a)(2)(I)(C). Step 2A Prong 2 Integration into a Practical Application: This judicial exception is not integrated into a practical application. In particular, the claim recites: “a computing system”. This limitation is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f). Further, the claim recites: “receiving a training set”. This limitation amounts to mere data gathering. It is necessary to acquire the data in order to use the recited judicial exception. Therefore, this limitation is insignificant extra-solution activity to the judicial exception, see MPEP §2106.05(g). Further, the claim recites: “based at least in part on the training data set, training a transformer network that includes a plurality of layers”. This limitation is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f). Further, the claim recites: “wherein the plurality of layers each respectively include a plurality of sub-layers including: an attention sub-layer; a feed-forward sub-layer; and a plurality of normalization sub-layers downstream from corresponding sub-layers of the plurality of sub-layers”. This limitation is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Finally, the claim recites: “wherein the input vector is passed to the normalization sub-layer in a residual stream of the transformer network”. This limitation is an additional element that amounts to insignificant extra-solution activity to the judicial exception, see MPEP §2106.05(g). Step 2B Significantly More: The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements: “a computing system” and “based at least in part on the training data set, training a transformer network that includes a plurality of layers” amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. Elements that merely amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process cannot provide an inventive concept. Further, the additional element “receiving a training set” is insignificant extra-solution activity to the judicial exception and is directed to receiving or transmitting data over a network which courts have recognized as well-understood, routine, and conventional when they are claimed in a generic manner, see MPEP §2106.05(d)(II). Further, the additional element “wherein the plurality of layers each respectively include a plurality of sub-layers including: an attention sub-layer; a feed-forward sub-layer; and a plurality of normalization sub-layers downstream from corresponding sub-layers of the plurality of sub-layers” generally links the use of the judicial exception to a particular technological environment or field of use. Elements that merely generally link the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. Finally, the additional element: “wherein the input vector is passed to the normalization sub-layer in a residual stream of the transformer network” amounts to insignificant extra-solution activity to the judicial exception, and further, the additional element is well‐understood, routine, and conventional as taught by activity is supported under Berkheimer Option 2, Dey et al., U.S. Patent Application Publication No. 20220292266, Paragraph 0002, Lines 21-28, “The original Transformer design uses L-layers in the encoder 110, where L=6 to perform sequential operations (e.g., 110_1, 110_2, . . . , 110_L, where L=6) and L layers for the decoder 102. Layers of encoder 110 are processed by Add and Norm components 112, 114 via skip feeds 115, 116 to perform residual connection followed by layer normalization, which are well known functions applied in deep architectures”; See also, Figure 1. The claim is not patent eligible. Regarding claim 11, the rejection of claim 10 is incorporated, and further, claim 11 is substantially similar to claim 2 respectively, and is rejected in the same manner and reasoning applying. Regarding claim 12, the rejection of claim 11 is incorporated, and further, claim 12 is substantially similar to claim 3 respectively, and is rejected in the same manner and reasoning applying. Regarding claim 13, the rejection of claim 12 is incorporated, and further, claim 13 is substantially similar to claim 4 respectively, and is rejected in the same manner and reasoning applying. Regarding claim 14, the rejection of claim 13 is incorporated, and further, claim 14 is substantially similar to claim 5 respectively, and is rejected in the same manner and reasoning applying. Regarding claim 15, the rejection of claim 13 is incorporated, and further, claim 15 is substantially similar to claim 6 respectively, and is rejected in the same manner and reasoning applying. Regarding claim 16, the rejection of claim 15 is incorporated, and further, claim 16 is substantially similar to claim 7 respectively, and is rejected in the same manner and reasoning applying. Regarding claim 17, the rejection of claim 10 is incorporated, and further, claim 17 is substantially similar to claim 8 respectively, and is rejected in the same manner and reasoning applying. Regarding claim 18, the rejection of claim 10 is incorporated, and further, claim 18 is substantially similar to claim 9 respectively, and is rejected in the same manner and reasoning applying. Regarding claim 19: Step 1 Statutory Category: Claim 19 is directed to a system, which falls under one of the four statutory categories. Step 2A Prong 1 Judicial Exception: Claim 19 recites, in part, “process the inferencing input data … to generate inferencing output data”. This limitation, under the broadest reasonable interpretation, covers the recitation of a mathematical calculation, as directed to “a claim that recites a mathematical calculation, when the claim is given its broadest reasonable interpretation in light of the specification, will be considered as falling within the "mathematical concepts" grouping. A mathematical calculation is a mathematical operation (such as multiplication) or an act of calculating using mathematical methods to determine a variable or number”. See MPEP §2106.04(a)(2)(I)(C). Further, the claim recites: “wherein each of the plurality of normalization sub-layers is configured to apply layer normalization to a sum of: a first scaling parameter multiplied by an input vector of the sub-layer, …; and an output vector of the sub-layer, such that the first scaling parameter weights respective contributions of the input vector and the output vector to the sum on which layer normalization is performed”. This limitation, under the broadest reasonable interpretation, covers the recitation of a mathematical calculation, as directed to “a claim that recites a mathematical calculation, when the claim is given its broadest reasonable interpretation in light of the specification, will be considered as falling within the "mathematical concepts" grouping. A mathematical calculation is a mathematical operation (such as multiplication) or an act of calculating using mathematical methods to determine a variable or number”. See MPEP §2106.04(a)(2)(I)(C). Step 2A Prong 2 Integration into a Practical Application: This judicial exception is not integrated into a practical application. In particular, the claim recites: “a computing system” and “a processor”. These limitations are additional elements that amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f). Further, the claim recites: “receive inferencing input data”. This limitation amounts to mere data gathering. It is necessary to acquire the data in order to use the recited judicial exception. Therefore, this limitation is insignificant extra-solution activity to the judicial exception, see MPEP §2106.05(g). Further, the claim recites: “at a transformer network”. This limitation is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Further, the claim recites: “wherein the transformer network includes a plurality of layers that each respectively include a plurality of sub-layers including: an attention sub-layer; a feed-forward sub-layer; and a plurality of normalization sub-layers downstream from corresponding sub-layers of the plurality of sub-layers”. This limitation is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Further, the claim recites: “wherein the input vector is passed to the normalization sub-layer in a residual stream of the transformer network”. This limitation is an additional element that amounts to insignificant extra-solution activity to the judicial exception, see MPEP §2106.05(g). Finally, the claim recites: “output the inferencing output data”. This limitation is insignificant extra-solution activity to the judicial exception, see MPEP §2106.05(g). Step 2B Significantly More: The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements: “a computing system” and “a processor”, amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. Elements that merely amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process cannot provide an inventive concept. Further, the additional elements “receive inferencing input data” and “output the inferencing output data” are insignificant extra-solution activity to the judicial exception and are directed to receiving or transmitting data over a network which courts have recognized as well-understood, routine, and conventional when they are claimed in a generic manner, see MPEP §2106.05(d)(II). Further, the additional elements “at a transformer network” and “wherein the transformer network includes a plurality of layers that each respectively include a plurality of sub-layers including: an attention sub-layer; a feed-forward sub-layer; and a plurality of normalization sub-layers downstream from corresponding sub-layers of the plurality of sub-layers” generally link the use of the judicial exception to a particular technological environment or field of use. Elements that merely generally link the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. Finally, the additional element: “wherein the input vector is passed to the normalization sub-layer in a residual stream of the transformer network” amounts to insignificant extra-solution activity to the judicial exception, and further, the additional element is well‐understood, routine, and conventional as taught by activity is supported under Berkheimer Option 2, Dey et al., U.S. Patent Application Publication No. 20220292266, Paragraph 0002, Lines 21-28, “The original Transformer design uses L-layers in the encoder 110, where L=6 to perform sequential operations (e.g., 110_1, 110_2, . . . , 110_L, where L=6) and L layers for the decoder 102. Layers of encoder 110 are processed by Add and Norm components 112, 114 via skip feeds 115, 116 to perform residual connection followed by layer normalization, which are well known functions applied in deep architectures”; See also, Figure 1. The claim is not patent eligible. Regarding claim 20, the rejection of claim 19 is incorporated, and further, the claim recites: “wherein the transformer network is a machine translation model configured to: receive, as the inferencing input data, a text input in a first language; and output, as the inferencing output data, the text input translated into a second language”. This limitation is an additional element that generally links the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Elements that merely generally link the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. The claim is not patent eligible. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 9-10, and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Ott et al., Scaling Neural Machine Translation, 09/04/2018, http://arxiv.org/pdf/1806.00187, hereinafter referred to as “Ott” in view of Yin et al., Patent Application Publication No. 20240127000, hereinafter referred to as “Yin” further in view of Liu et al., Rethinking Skip Connection with Layer Normalization in Transformers and ResNets, 05/15/2021, https://arxiv.org/pdf/2105.07205, hereinafter referred to as “Liu”. Regarding claim 1, Ott teaches A computing system comprising: a processor (Ott, Page 3, Col 2, Paragraph 2, “All experiments are run on DGX-1 nodes with 8 NVIDIA c V100 GPUs interconnected by Infiniband. We use the NCCL2 library and torch.distributed for inter-GPU communication”) configured to: receive a training data set (Ott, Page 2, Section 3.1, Lines 3-5, “For En–De we replicate the setup of Vaswani et al. (2017) which relies on the WMT’16 training data with 4.5M sentence pairs”); and based at least in part on the training data set, train a transformer network that includes a plurality of layers (Ott, Page 3, Section 3.2, Lines 1-6, “We use the Transformer model (Vaswani et al., 2017) implemented in PyTorch in the fairseq-py toolkit (Edunov et al., 2017). All experiments are based on the “big” transformer model with 6 blocks in the encoder and decoder networks” The “blocks” are considered to be the “plurality of layers”), wherein the plurality of layers each respectively include a plurality of sub-layers including: an attention sub-layer; a feed-forward sub-layer; and a plurality of normalization sub-layers downstream from corresponding sub-layers of the plurality of sub-layers (Ott, Page 3.2, Lines 6-16, “Each encoder block contains a self-attention layer, followed by two fully connected feed-forward layers with a ReLU non-linearity between them. Each decoder block contains self-attention, followed by encoder-decoder attention, followed by two fully connected feed-forward layers with a ReLU between them. We include residual connections (He et al., 2015) after each attention layer and after the combined feedforward layers, and apply layer normalization (Ba et al., 2016) after each residual connection”). Ott does not explicitly teach wherein each of the plurality of normalization sub-layers is configured to apply layer normalization to a sum of: a first scaling parameter multiplied by an input vector of the sub-layer, wherein the input vector is passed to the normalization sub-layer in a residual stream of the transformer network; and an output vector of the sub-layer, such that the first scaling parameter weights respective contributions of the input vector and the output vector to the sum on which layer normalization is performed. Yin teaches wherein each of the plurality of normalization sub-layers is configured to apply layer normalization to a sum of: … an input vector of the sub-layer, wherein the input vector is passed to the normalization sub-layer in a residual stream of the transformer network; and an output vector of the sub-layer… (Yin, Paragraph 0060, Lines 1-4, “As shown in FIG. 2A, module 224 may further aggregate the output of MHA module 222 with the input of MHA module 222, by applying: H i M H A = L a y e r N o r m H i - 1 + M H A H i - 1 ; See also, Yin, Figure 2A, shown are residual connections passing input to “Add & Norm” layers which are considered to be equivalent to the “normalization sub-layer”). It would have been obvious to a person of ordinary skill in the art to modify the transformer training method of Ott to include the layer normalization method of Yin. The motivation for doing so would have been that Ott does use layer normalization (Ott, Page 3, Section 3.2, Lines 12-16) and Yin allows the model to aggregate the input and the output of the sub-layer, which is then input into the next sub-layer (Yin, Paragraph 0060), further, layer normalization would stabilize the dynamics of the hidden states of the transformer layers (Yin, Paragraph 0061, “Each of MHA module 222 and FFN 232 may implement layer normalization (“LN”), which is used to stabilize the dynamics of the hidden states in transformer layers 220”). Ott in view of Yin does not explicitly teach the input vector of the sub-layer being multiplied by a first scaling parameter nor applying the layer normalization …such that the first scaling parameter weights respective contributions of the input vector and the output vector to the sum on which layer normalization is performed. Liu teaches the input vector of the sub-layer being multiplied by a first scaling parameter and applying the layer normalization …such that the first scaling parameter weights respective contributions of the input vector and the output vector to the sum on which layer normalization is performed (Liu, Page 4, Section 3, Paragraph 3, Lines 1-3 and equation 3, “Motivated by Transformer, which combines skip connection with layer normalization, we further examine the effect of layer normalization on expanded skip connection, which takes the form of y = L N λ x + F x ,   W ”; “ λ ” is considered to be the “first scaling parameter”; Liu, Page 2, Lines 1-2, “λ denotes the modulating factor that controls the relative importance of the skip connection or the shortcut”). It would have been obvious to a person of ordinary skill before the effective filing date of the invention, to have modified the transformer network of Ott in view of Yin to include a scaling parameter during layer normalization as taught by Liu. The motivation to do so would have been to better incorporated the effect of the transformed input (Liu, Page 2, Final Bullet Point, “The proposed recursive skip connection with layer normalization further facilities the optimization by separating the expanded skip connection into multiple stages to better incorporate the effect of the transformed input”), further, feeding more input signals to the layer output is helpful to learning the overall model, and coupled with layer normalization do not raise optimization difficulty (Liu, Page 7, Paragraph 2, Lines 1-3, “In summary, our experiments show that feeding more input signals to the layer output should be helpful to learning of the overall model. Previous concerns on the optimization difficulty introduced by the expanded skip connection could be alleviated by the usage of layer normalization”). Regarding claim 9, the rejection of claim 1 is incorporated, and further, the proposed combination teaches wherein the transformer network is a machine translation model (Ott, Page 1, Abstract, Lines 4-17, “This paper shows that reduced precision and large batch training can speedup training by nearly 5x on a single 8- GPU machine with careful tuning and implementation.1 On WMT’14 English-German translation, we match the accuracy of Vaswani et al. (2017) in under 5 hours when training on 8 GPUs and we obtain a new state of the art of 29.3 BLEU after training for 85 minutes on 128 GPUs. We further improve these results to 29.8 BLEU by training on the much larger Paracrawl dataset. On the WMT’14 English-French task, we obtain a state-of-the-art BLEU of 43.2 in 8.5 hours on 128 GPUs”). Regarding claim 10, Ott teaches A method for use with a computing system, (Ott, Page 3, Col 2, Paragraph 2, “All experiments are run on DGX-1 nodes with 8 NVIDIA c V100 GPUs interconnected by Infiniband. We use the NCCL2 library and torch.distributed for inter-GPU communication”) the method comprising: receiving a training data set (Ott, Page 2, Section 3.1, Lines 3-5, “For En–De we replicate the setup of Vaswani et al. (2017) which relies on the WMT’16 training data with 4.5M sentence pairs”); and based at least in part on the training data set, training a transformer network that includes a plurality of layers (Ott, Page 3, Section 3.2, Lines 1-6, “We use the Transformer model (Vaswani et al., 2017) implemented in PyTorch in the fairseq-py toolkit (Edunov et al., 2017). All experiments are based on the “big” transformer model with 6 blocks in the encoder and decoder networks” The “blocks” are considered to be the “plurality of layers”), wherein the plurality of layers each respectively include a plurality of sub-layers including: an attention sub-layer; a feed-forward sub-layer; and a plurality of normalization sub-layers downstream from corresponding sub-layers of the plurality of sub-layers (Ott, Page 3.2, Lines 6-16, “Each encoder block contains a self-attention layer, followed by two fully connected feed-forward layers with a ReLU non-linearity between them. Each decoder block contains self-attention, followed by encoder-decoder attention, followed by two fully connected feed-forward layers with a ReLU between them. We include residual connections (He et al., 2015) after each attention layer and after the combined feedforward layers, and apply layer normalization (Ba et al., 2016) after each residual connection”). Ott does not explicitly teach wherein each of the plurality of normalization sub-layers is configured to apply layer normalization to a sum of: a first scaling parameter multiplied by an input vector of the sub-layer; and an output vector of the sub-layer. Yin teaches wherein each of the plurality of normalization sub-layers is configured to apply layer normalization to a sum of: … an input vector of the sub-layer; and an output vector of the sub-layer (Yin, Paragraph 0060, Lines 1-4, “As shown in FIG. 2A, module 224 may further aggregate the output of MHA module 222 with the input of MHA module 222, by applying: H i M H A = L a y e r N o r m ( H i - 1 + M H A H i - 1 ) ”. It would have been obvious to a person of ordinary skill in the art to modify the transformer training method of Ott to include the layer normalization method of Yin. The motivation for doing so would have been that Ott does use layer normalization (Ott, Page 3, Section 3.2, Lines 12-16) and Yin allows the model to aggregate the input and the output of the sub-layer, which is then input into the next sub-layer (Yin, Paragraph 0060). Ott in view of Yin does not explicitly teach the input vector of the sub-layer being multiplied by a first scaling parameter nor applying the layer normalization …such that the first scaling parameter weights respective contributions of the input vector and the output vector to the sum on which layer normalization is performed. Liu teaches the input vector of the sub-layer being multiplied by a first scaling parameter and applying the layer normalization …such that the first scaling parameter weights respective contributions of the input vector and the output vector to the sum on which layer normalization is performed (Liu, Page 4, Section 3, Paragraph 3, Lines 1-3 and equation 3, “Motivated by Transformer, which combines skip connection with layer normalization, we further examine the effect of layer normalization on expanded skip connection, which takes the form of y = L N λ x + F x ,   W ”; “ λ ” is considered to be the “first scaling parameter”; Liu, Page 2, Lines 1-2, “λ denotes the modulating factor that controls the relative importance of the skip connection or the shortcut”). It would have been obvious to a person of ordinary skill before the effective filing date of the invention, to have modified the transformer network of Ott in view of Yin to include a scaling parameter during layer normalization as taught by Liu. The motivation to do so would have been to better incorporated the effect of the transformed input (Liu, Page 2, Final Bullet Point, “The proposed recursive skip connection with layer normalization further facilities the optimization by separating the expanded skip connection into multiple stages to better incorporate the effect of the transformed input”), further, feeding more input signals to the layer output is helpful to learning the overall model, and coupled with layer normalization do not raise optimization difficulty (Liu, Page 7, Paragraph 2, Lines 1-3, “In summary, our experiments show that feeding more input signals to the layer output should be helpful to learning of the overall model. Previous concerns on the optimization difficulty introduced by the expanded skip connection could be alleviated by the usage of layer normalization”). Regarding claim 18, the rejection of claim 10 is incorporated, and further, the proposed combination teaches wherein the transformer network is a machine translation model (Ott, Page 1, Abstract, Lines 4-17, “This paper shows that reduced precision and large batch training can speedup training by nearly 5x on a single 8- GPU machine with careful tuning and implementation.1 On WMT’14 English-German translation, we match the accuracy of Vaswani et al. (2017) in under 5 hours when training on 8 GPUs and we obtain a new state of the art of 29.3 BLEU after training for 85 minutes on 128 GPUs. We further improve these results to 29.8 BLEU by training on the much larger Paracrawl dataset. On the WMT’14 English-French task, we obtain a state-of-the-art BLEU of 43.2 in 8.5 hours on 128 GPUs”). Regarding claim 19, Ott teaches A computing system comprising: a processor (Ott, Page 3, Col 2, Paragraph 2, “All experiments are run on DGX-1 nodes with 8 NVIDIA c V100 GPUs interconnected by Infiniband. We use the NCCL2 library and torch.distributed for inter-GPU communication”) configured to: receive inferencing input data (Ott, Page 2, Section 3.1, Lines 3-5, “For En–De we replicate the setup of Vaswani et al. (2017) which relies on the WMT’16 training data with 4.5M sentence pairs”); and process the inferencing input data at a transformer network to generate inferencing output data, wherein the transformer network includes a plurality of layers (Ott, Page 3, Section 3.2, Lines 1-6, “We use the Transformer model (Vaswani et al., 2017) implemented in PyTorch in the fairseq-py toolkit (Edunov et al., 2017). All experiments are based on the “big” transformer model with 6 blocks in the encoder and decoder networks”; Ott, Page 5, Section 4.4, Lines 1-2, “We report results on newstest14 for English-to-German (En-De) and English-to-French (En-Fr)”; Ott, Page 5, Section 4.4, Lines 8-10, “Table 2 reports 29.3 BLEU for En-De in 1h 25min and 43.2 BLEU for En-Fr in 8h 32min” The “blocks” are considered to be the “plurality of layers” and because the experiments were performed and the BLEU calculated, inferencing output data must have been generated in the form of a translated sentence), that each respectively include a plurality of sub-layers including: an attention sub-layer; a feed-forward sub-layer; and a plurality of normalization sub-layers downstream from corresponding sub-layers of the plurality of sub-layers (Ott, Page 3.2, Lines 6-16, “Each encoder block contains a self-attention layer, followed by two fully connected feed-forward layers with a ReLU non-linearity between them. Each decoder block contains self-attention, followed by encoder-decoder attention, followed by two fully connected feed-forward layers with a ReLU between them. We include residual connections (He et al., 2015) after each attention layer and after the combined feedforward layers, and apply layer normalization (Ba et al., 2016) after each residual connection”); and output the inferencing output data (Ott, Page 5, Section 4.4, Lines 1-2, “We report results on newstest14 for English-to-German (En-De) and English-to-French (En-Fr)”; Ott, Page 5, Section 4.4, Lines 8-10, “Table 2 reports 29.3 BLEU for En-De in 1h 25min and 43.2 BLEU for En-Fr in 8h 32min”; In order for the BLEU to be calculated, the “inferencing output data” must have been output by the transformer) Ott does not explicitly teach wherein each of the plurality of normalization sub-layers is configured to apply layer normalization to a sum of: a first scaling parameter multiplied by an input vector of the sub-layer; and an output vector of the sub-layer. Yin teaches wherein each of the plurality of normalization sub-layers is configured to apply layer normalization to a sum of: … an input vector of the sub-layer; and an output vector of the sub-layer (Yin, Paragraph 0060, Lines 1-4, “As shown in FIG. 2A, module 224 may further aggregate the output of MHA module 222 with the input of MHA module 222, by applying: H i M H A = L a y e r N o r m ( H i - 1 + M H A H i - 1 ) ”. It would have been obvious to a person of ordinary skill in the art to modify the transformer training method of Ott to include the layer normalization method of Yin. The motivation for doing so would have been that Ott does use layer normalization (Ott, Page 3, Section 3.2, Lines 12-16) and Yin allows the model to aggregate the input and the output of the sub-layer, which is then input into the next sub-layer (Yin, Paragraph 0060). Ott in view of Yin does not explicitly teach the input vector of the sub-layer being multiplied by a first scaling parameter nor applying the layer normalization …such that the first scaling parameter weights respective contributions of the input vector and the output vector to the sum on which layer normalization is performed. Liu teaches the input vector of the sub-layer being multiplied by a first scaling parameter and applying the layer normalization …such that the first scaling parameter weights respective contributions of the input vector and the output vector to the sum on which layer normalization is performed (Liu, Page 4, Section 3, Paragraph 3, Lines 1-3 and equation 3, “Motivated by Transformer, which combines skip connection with layer normalization, we further examine the effect of layer normalization on expanded skip connection, which takes the form of y = L N λ x + F x ,   W ”; “ λ ” is considered to be the “first scaling parameter”; Liu, Page 2, Lines 1-2, “λ denotes the modulating factor that controls the relative importance of the skip connection or the shortcut”). It would have been obvious to a person of ordinary skill before the effective filing date of the invention, to have modified the transformer network of Ott in view of Yin to include a scaling parameter during layer normalization as taught by Liu. The motivation to do so would have been to better incorporated the effect of the transformed input (Liu, Page 2, Final Bullet Point, “The proposed recursive skip connection with layer normalization further facilities the optimization by separating the expanded skip connection into multiple stages to better incorporate the effect of the transformed input”), further, feeding more input signals to the layer output is helpful to learning the overall model, and coupled with layer normalization do not raise optimization difficulty (Liu, Page 7, Paragraph 2, Lines 1-3, “In summary, our experiments show that feeding more input signals to the layer output should be helpful to learning of the overall model. Previous concerns on the optimization difficulty introduced by the expanded skip connection could be alleviated by the usage of layer normalization”). Regarding claim 20, the rejection of claim 19 is incorporated, and further, the proposed combination teaches wherein the transformer network is a machine translation model configured to: receive, as the inferencing input data, a text input in a first language; and output, as the inferencing output data, the text input translated into a second language (Ott, Page 1, Abstract, Lines 4-17, “This paper shows that reduced precision and large batch training can speedup training by nearly 5x on a single 8- GPU machine with careful tuning and implementation.1 On WMT’14 English-German translation, we match the accuracy of Vaswani et al. (2017) in under 5 hours when training on 8 GPUs and we obtain a new state of the art of 29.3 BLEU after training for 85 minutes on 128 GPUs. We further improve these results to 29.8 BLEU by training on the much larger Paracrawl dataset. On the WMT’14 English-French task, we obtain a state-of-the-art BLEU of 43.2 in 8.5 hours on 128 GPUs”). Claims 2 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Ott in view of Yin further in view of Liu in further view of Xu et al., Understanding and Improving Layer Normalization, 11/16/2019, https://arxiv.org/pdf/1911.07013, hereinafter referred to as “Xu”. Regarding claim 2, the rejection of claim 1 is incorporated. The proposed combination does not explicitly teach wherein, at each of the plurality of layers, the processor is further configured to scale a plurality of value projection weights and a plurality of output projection weights of the attention sub-layer and a plurality of feed-forward weights of the feed-forward sub-layer by a second scaling parameter when training the transformer network. Xu teaches wherein, at each of the plurality of layers, the processor is further configured to scale a plurality of value projection weights and a plurality of output projection weights of the attention sub-layer and a plurality of feed-forward weights of the feed-forward sub-layer by a second scaling parameter when training the transformer network (Xu, Page 2, Paragraph 4, Lines 1-3, “we propose a novel normalization method, Adaptive Normalization (AdaNorm). AdaNorm replaces the bias and gain with a new transformation function. This function adaptively adjusts scaling weights based on input values”; Xu, Page 7, Section 4.1). It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to modify the transformer training method of the proposed combination to include scaling the weights of the sub-layers as taught by Xu. The motivation for doing so would have been to reduce over-fitting, which results in better model performance (Xu, Page 1, Abstract, Lines 13-16). Regarding claim 11, the rejection of claim 10 is incorporated. The proposed combination does not explicitly teach wherein, at each of the plurality of layers, the processor is further configured to scale a plurality of value projection weights and a plurality of output projection weights of the attention sub-layer and a plurality of feed-forward weights of the feed-forward sub-layer by a second scaling parameter when training the transformer network. Xu teaches wherein, at each of the plurality of layers, the processor is further configured to scale a plurality of value projection weights and a plurality of output projection weights of the attention sub-layer and a plurality of feed-forward weights of the feed-forward sub-layer by a second scaling parameter when training the transformer network (Xu, Page 2, Paragraph 4, Lines 1-3, “we propose a novel normalization method, Adaptive Normalization (AdaNorm). AdaNorm replaces the bias and gain with a new transformation function. This function adaptively adjusts scaling weights based on input values”; Xu, Page 7, Section 4.1). It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to modify the transformer training method of the proposed combination to include scaling the weights of the sub-layers as taught by Xu. The motivation for doing so would have been to reduce over-fitting, which results in better model performance (Xu, Page 1, Abstract, Lines 13-16). Claims 8 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Ott in view of Yin further in view of Liu in further view of HUANG, et al., "GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism", In Proceedings of Advances in Neural Information Processing Systems, Volume 32, December 08, 2019, 10 Pages, hereinafter referred to as “Huang”. Regarding claim 8, the rejection of claim 1 is incorporated. The proposed combination does not explicitly teach wherein the transformer network includes 100 or more layers. Huang teaches wherein the transformer network includes 100 or more layers (Huang, Page 1, Abstract, Lines 16-18, “We train a single 6-billion-parameter, 128-layer Transformer model on a corpus spanning over 100 languages and achieve better quality than all bilingual models”). It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to modify the transformer training method of the proposed combination, to include the transformer network including 100 or more layers as taught by Huang. The motivation for doing so would have been that scaling up a network improves model quality (Huang, Page 1, Abstract, Lines 1-2, “Scaling up deep neural network capacity has been known as an effective approach to improving model quality for several different machine learning tasks”). Regarding claim 17, the rejection of claim 10 is incorporated. The proposed combination does not explicitly teach wherein the transformer network includes 100 or more layers. Huang teaches wherein the transformer network includes 100 or more layers (Huang, Page 1, Abstract, Lines 16-18, “We train a single 6-billion-parameter, 128-layer Transformer model on a corpus spanning over 100 languages and achieve better quality than all bilingual models”). It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to modify the transformer training method of the proposed combination, to include the transformer network including 100 or more layers as taught by Huang. The motivation for doing so would have been that scaling up a network improves model quality (Huang, Page 1, Abstract, Lines 1-2, “Scaling up deep neural network capacity has been known as an effective approach to improving model quality for several different machine learning tasks”). Allowable Subject Matter Claims 3-7 and 12-16 are objected to as being dependent upon a rejected base claim, but would be allowable over prior art if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Claims 3-7 and 12-16 have been rejected under 35 U.S.C. 101 only. A complete prior art search was performed for these claims; however no prior art was uncovered that disclose or fairly suggest the following claimed features: Reason for allowance after detailed search, the cited arts, neither alone nor in combination, teach the claimed subject matter of claims 3 and 12, “wherein the processor is further configured to determine the first scaling parameter and the second scaling parameter based at least in part on a number of the plurality of layers”. Pertinent art (Liu) discloses a scaling parameter to be used while determining the input to layer normalization (Liu, Page 4, Section 3, Expanded Skip Connection with Layer Normalization (xSkip+LN); Liu, Page 9, Learning λ), however the art does not disclose determining the scaling parameter based at least in part on a number of the plurality of layers as required by the claims. Pertinent art (Xu) discloses the second scaling parameter, but adaptively adjusts the scaling parameter based on input values (Xu, Page 2, Paragraph 4, Lines 1-3, “we propose a novel normalization method, Adaptive Normalization (AdaNorm). AdaNorm replaces the bias and gain with a new transformation function. This function adaptively adjusts scaling weights based on input values) and not based at least in part on a number of the plurality of layers, as required by the claims. Response to Arguments Applicant’s arguments regarding the 35 U.S.C. 101 rejections of the claims have been fully considered but are unpersuasive. Applicant first argues, on page 11, paragraph 2 of the response, that claim 1 is directed to eligible subject matter at Step 2A Prong 2 of the subject matter eligibility analysis because the claim provides an improvement in the functioning of the computing device itself, thus integrating the abstract idea into a practical application. Examiner respectfully disagrees. While the applicant argues the improvement lies in the architecture of the transformer network, the claim reflects an improvement in layer normalization. An improvement to layer normalization maybe be an improvement in an abstract idea, but not an improvement in the functioning of a computer, as a computer. Further, it is important to note, the judicial exception alone cannot provide the improvement, see MPEP 2106.05(a). Applicant next argues, on page 11, paragraph 3 of the response, that claim 1 is analogous to the claims at issue in Desjardins because it is directed to an improvement in the structure of a machine learning model that provides an improvement to that model. Examiner respectfully disagrees. Claim 1 does not reflect an improvement in the “structure of a machine learning model”, but rather to the mathematical concept of layer normalization. Thus, the fact pattern of the instant case is not identical to the fact pattern of Desjardins and therefore, the same logic cannot be applied. An improvement to layer normalization maybe be an improvement in an abstract idea, but not an improvement in the functioning of a computer, as a computer. Applicant's arguments regarding the remainder of the claims rely upon the arguments asserted with respect to the independent claims, and are thus unpersuasive. Applicant’s amendments to the claims with regard to the 35 U.S.C. 101 rejections do not overcome the rejection, for an analysis of the limitations added during amendment please see the updated 35 U.S.C. 101 rejections above. Applicant’s arguments regarding the 35 U.S.C. 103 rejections of the claims have been fully considered but are unpersuasive. Applicant’s arguments with respect to the claims have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Specifically, applicant argued, on pages 14-15 of the response, that Unity Technologies does not disclose scalar multiplication be applied to a term of an input to a LayerNorm function. The Unity Technologies reference is no longer relied upon for that or any other limitation of the claims. Please see updated 35 U.S.C. 103 rejection above. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MOLLY CLARKE SIPPEL whose telephone number is (571)272-3270. The examiner can normally be reached Monday - Friday, 7:30 a.m. - 4:30 p.m. ET.. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kakali Chaki can be reached at (571)272-3719. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /M.C.S./ Examiner, Art Unit 2122 /KAKALI CHAKI/ Supervisory Patent Examiner, Art Unit 2122
Read full office action

Prosecution Timeline

Feb 28, 2023
Application Filed
Dec 01, 2025
Non-Final Rejection mailed — §101, §103
Feb 25, 2026
Examiner Interview Summary
Feb 25, 2026
Applicant Interview (Telephonic)
Mar 02, 2026
Response Filed
May 05, 2026
Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12670387
SYSTEM, METHOD, AND COMPUTER-READABLE MEDIA FOR LEAKAGE CORRECTION IN GRAPH NEURAL NETWORK BASED RECOMMENDER SYSTEMS
4y 1m to grant Granted Jun 30, 2026
Patent 12664398
SYSTEM, METHOD AND NON-TRANSITORY COMPUTER READABLE MEDIUM
3y 9m to grant Granted Jun 23, 2026
Patent 12657427
Systems, Methods, and Computer Program Products for Determining Uncertainty from a Deep Learning Classification Model
4y 1m to grant Granted Jun 16, 2026
Patent 12632779
HYPERPARAMETER SELECTION USING BUDGET-AWARE BAYESIAN OPTIMIZATION
4y 5m to grant Granted May 19, 2026
Patent 12626098
METHOD AND SYSTEM FOR CREATING AN ENSEMBLE OF NEURAL NETWORK-BASED CLASSIFIERS THAT OPTIMIZES A DIVERSITY METRIC
3y 8m to grant Granted May 12, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
52%
Grant Probability
78%
With Interview (+26.1%)
3y 9m (~4m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 23 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month