Prosecution Insights
Last updated: October 02, 2026
Application No. 18/229,272

System, Method, Computer Program Product for Operating a Gated Multilayer Perceptron Machine Learning Model Architecture

Final Rejection §103
Filed
Aug 02, 2023
Examiner
KAPOOR, DEVAN
Art Unit
2126
Tech Center
2100 — Computer Architecture & Software
Assignee
Visa International Service Association
OA Round
2 (Final)
7%
Grant Probability
At Risk
3-4
OA Rounds
1y 2m
Est. Remaining
18%
With Interview

Examiner Intelligence

Grants only 7% of cases
7%
Career Allowance Rate
1 granted / 14 resolved
-47.9% vs TC avg
Moderate +11% lift
Without
With
+11.1%
Interview Lift
resolved cases with interview
Typical timeline
4y 4m
Avg Prosecution
29 currently pending
Career history
47
Total Applications
across all art units

Statute-Specific Performance

§101
34.0%
-6.0% vs TC avg
§103
57.4%
+17.4% vs TC avg
§102
5.8%
-34.2% vs TC avg
§112
2.2%
-37.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 14 resolved cases

Office Action

§103
DETAILED ACTION This action is responsive to the application filed on 07/28/2026. Claims 1, 3-4, 6-8, 10-11, 13-15, 17-18, and 20-21 are pending and have been examined. This action is Final. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Argument 1: The applicant argues that amended claim 1 is directed to an improved classification system that uses clustering information of data instances in each class of a dataset to provide a more accurate result, and therefore integrates any judicial exception into a practical application under Step 2A, Prong 2. The applicant relies on paragraphs [0003]-[0004] and [0045]-[0048] of the publication, which describe that a conventional multilayer perceptron cannot use class clustering information and that the disclosed gated architecture is more accurate and requires fewer computational resources. Citing MPEP 2106.05(a), the applicant contends that the specification makes this improvement apparent. The applicant identifies the claimed technical solution as determining center values of first and second classifications in a feature space, generating three intermediate embeddings with three neural network models, generating an intermediate classification with a gating model, multiplying that classification with the second and third embeddings, concatenating the result with the first embedding, and classifying the combined input with a head model. Under Step 2B, the applicant argues that the additional elements, considered individually and in combination (MPEP 2106.05(e)), provide meaningful limitations beyond routine or conventional computer use. The dependent claims and independent claims 8 and 15 are argued as eligible for the same reasons. Response to Argument 1: The applicant’s arguments with respect to the rejection of the claims under 35 U.S.C. 101 have been fully considered and are persuasive. Accordingly, the rejection under 35 U.S.C. 101 is withdrawn. Argument 2: The applicant argues that none of Mao, Friedrich, or Shazeer, alone or in combination, teaches or suggests the limitations added to amended claim 1. Specifically, the applicant asserts that the references do not teach determining a center value of a first classification and of a second classification in a feature space of the interaction data; generating the second and third intermediate embeddings from inputs comprising those center values; generating those embeddings with separate second and third machine learning models; or the first, second, and third machine learning models each being neural network models. The applicant does not address the specific passages or interpretations relied on in the prior action and does not challenge the rationale for combining the references. The applicant concludes that claim 1 is allowable, that independent claims 8 and 15 are allowable for similar reasons, and that the dependent claims are allowable based on their dependency. Response to Argument 2: The applicant’s arguments have been fully considered but are not persuasive. The amendments incorporate the subject matter of previously rejected claims 2, 3, and 5 into claim 1, of previously rejected claims 9, 10, and 12 into claim 8, and of previously rejected claims 16, 17, and 19 into claim 15, and each of those limitations was rejected on the merits in the previous action over Mao in view of Friedrich further in view of Shazeer. With respect to determining a center value of a first classification and of a second classification in a feature space of the interaction data, Friedrich determines an aggregation of embeddings by “an aggregation function such as average” ([Friedrich, [0059]]), which is a representative central value computed in the embedding, i.e. feature, space of the input data, for each of the classifications determined at the first classifier and at the second classifier ([Friedrich, [0003]] AND [Friedrich, [0005]]). With respect to generating the second and third intermediate embeddings from inputs comprising those center values, Friedrich discloses “providing the embedding and/or a hidden state of the first classifier resulting from the provided embedding as input to the second classifier” and that “The classification or the hidden state of the classifier representing a parent may serve as input for the classifier representing its child” ([Friedrich, [0005]]). With respect to the second and third machine learning models and to those models each being neural network machine learning models, Mao discloses “two independent MLP networks as two streams” ([Mao, page 3]) and Shazeer discloses “a plurality of expert neural networks”, each of which “can be a feed-forward neural network with its own parameters” ([Shazeer, col 1, lines 35-44] AND [Shazeer, col 4, lines 57-59]). The applicant’s remarks otherwise recite the amended limitations and assert that the references fail to teach them without pointing out how the claim language patentably distinguishes over the specific teachings relied upon, and thus the rejection over 35 U.S.C 103 is maintained. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: Determining the scope and contents of the prior art. Ascertaining the differences between the prior art and the claims at issue. Resolving the level of ordinary skill in the pertinent art. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1, 3-4, 6-8, 10-11, 13-15, 17-18, and 20-21 are rejected under 35 U.S.C. 103 as being unpatentable over the NPL reference “FinalMLP: An Enhanced Two-Stream MLP Model for CTR Prediction”, by Mao et. al. (referred herein as Mao) in view of US 20220092440 A1, by Friedrich et. al. (referred herein as Friedrich) further in view of US 10719761 B2, by Shazeer et. al. (referred herein as Shazeer). Regarding claim 1, Mao teaches: A system, comprising: at least one processor programmed or configured to: receive interaction data associated with a plurality of interactions, the interaction data comprising a plurality of features; ([Mao, page 2] “Feature Embedding: Embedding is a common way to map high-dimensional and sparse raw features into dense numeric representations. Specifically, suppose that the raw input feature is x = {x1, ..., xM} with M feature fields, where xi is the feature of the i-th field. In general, xi can be a categorical, multi-valued, or numerical feature. Each of them can be transformed into embedding vectors accordingly”, wherein the examiner interprets “suppose that the raw input feature is x = {x1, ..., xM} with M feature fields” and “Each of them can be transformed into embedding vectors accordingly” to be the same as “receive interaction data associated with a plurality of interactions, the interaction data comprising a plurality of features” because they are both directed to receiving multi-field feature input data for downstream machine-learning processing.) generate a first intermediate embedding based on providing a first input to at least one machine learning model, wherein the first input comprises the plurality of features; ([Mao, page 2] “Each of them can be transformed into embedding vectors accordingly. … Then, these feature embeddings will be concatenated and fed into the following layer” and [Mao, page 4] “... the feature inputs h1 and h2 are usually set as the same one, which is a concatenation of feature embeddings e (optionally with some pooling), i.e., h1 = h2 = e.”, wherein the examiner interprets “feature embeddings will be concatenated and fed into the following layer” to be the same as “generate a first intermediate embedding based on providing a first input to at least one machine learning model, wherein the first input comprises the plurality of features” because they are both directed to taking feature inputs and transforming them into an embedding representation for further model processing.) wherein, when generating the second intermediate embedding, the at least one processor is programmed or configured to: generate the second intermediate embedding based on providing the second input to a second machine learning model; ([Mao, page 3] “DualMLP, which simply combines two independent MLP networks as two streams. Specifically, the two-stream MLP model can be formulated as follows”, wherein the examiner interprets “two independent MLP networks as two streams” to be the same as a second machine learning model because they are both directed to a separate machine learning model used as another stream for processing. The examiner further interprets the second of the “two independent MLP networks” to be the same as generate the second intermediate embedding based on providing the second input to a second machine learning model because they are both directed to using a second model stream to generate an intermediate output from a corresponding input.) provide the first intermediate embedding as an input to a gating machine learning model to generate an intermediate classification of the first intermediate embedding, wherein the gating machine learning model is configured to provide a prediction of a classification label of an input as an output; ([Mao, page 4] “Specifically, we make stream-specific feature selection through the context-aware feature gating layer as follows: g1 = Gate1(x1), g2 = Gate2(x2), (Eq. 3), and h1 = 2σ(g1) ⊙ e, h2 = 2σ(g2) ⊙ e, (Eq. 4). where Gatei denotes an MLP-based gating network, which takes stream-specific conditional features xi as input and outputs element-wise gating weights gi. Note that it is flexible to either choose xi from a set of user/item features or set it as learnable parameters. The feature importance weights are obtained by using the sigmoid function σ and a multiplier of 2 to transform them to the range of [0, 2] with an average of 1.”, wherein the examiner interprets “Gatei denotes an MLP-based gating network, which takes stream-specific conditional features xi as input and outputs element-wise gating weights gi” to be the same as “provide the first intermediate embedding as an input to a gating machine learning model to generate an intermediate classification of the first intermediate embedding” because they are both directed to providing an embedding/input representation to a gating or classifier model that produces a classification-related output.) multiply the intermediate classification of the first intermediate embedding, the second intermediate embedding, and the third intermediate embedding to provide an intermediate product of outputs; ([Mao, page 4] “Given the concatenated feature embeddings e, we can then obtain weighted feature outputs h1 and h2 via element-wise product ⊙.”, wherein the examiner interprets “feature outputs h1 and h2 via element-wise product ⊙” to be the same as “multiply the intermediate classification of the first intermediate embedding, the second intermediate embedding, and the third intermediate embedding to provide an intermediate product of outputs” because they are both directed to multiplying learned representations / gating-related outputs to produce a fused intermediate result.) combine the first intermediate embedding and the intermediate product of outputs to provide a combined final input; ([Mao, page 2] “existing two-stream models often combine two streams via summation or concatenation … To fuse the stream outputs with stream-level feature interaction, we propose an interaction aggregation layer based on second-order bilinear fusion.” and [Mao, page 2] “Stream-level fusion is required to fuse the outputs of two streams … F denotes the fusion operation which is commonly set as summation or concatenation.”, wherein the examiner interprets “fuse the stream outputs” and “fusion operation which is commonly set as summation or concatenation” to be the same as combine the first intermediate embedding and the intermediate product of outputs to provide a combined final input because they are both directed to combining separate stream outputs into one final fused input representation). Mao does not teach determine a center value of a first classification in a feature space of the interaction data…determine a center value of a second classification in the feature space of the interaction data…generate a second intermediate embedding based on providing a second input to the at least one machine learning model, wherein the second input comprises the center value of the first classification in the feature space of the interaction data; or generate a third intermediate embedding based on providing a third input to the at least one machine learning model, wherein the third input comprises the center value of the second classification in the feature space of the interaction data. Friedrich teaches: determine a center value of a first classification in a feature space of the interaction data; ([Friedrich, [0003]] “determining a first classification for the embedding at a first classifier, determining if the first classification meets a first condition” and [Friedrich, [0059]] “The aggregation is instead determined from the token embeddings 306-1, 306-2, …, 306-n by a component 309 which may be a convolutional neural network layer, CNN layer, or an aggregation function such as average, with or without attention.”, wherein the examiner interprets “determining a first classification for the embedding at a first classifier” to be the same as a first classification of the interaction data because they are both directed to producing a first classification result from processed input data. The examiner further interprets the “aggregation function such as average” that is “determined from the token embeddings” to be the same as a center value of the first classification in a feature space of the interaction data because they are both directed to a representative central value computed from the embedded representation, i.e. the feature space, of the data being classified.) determine a center value of a second classification in the feature space of the interaction data; ([Friedrich, col 1, lines 35-44] “in accordance with an example embodiment of the present invention, the method comprises determining a second classification at a second classifier,” and [Friedrich, [0059]] “The aggregation is instead determined from the token embeddings 306-1, 306-2, …, 306-n by a component 309 which may be a convolutional neural network layer, CNN layer, or an aggregation function such as average, with or without attention.”, wherein the examiner interprets “determining a second classification at a second classifier” to be the same as a second classification of the interaction data because they are both directed to producing a second classification result from processed input data. The examiner further interprets the “aggregation function such as average” that is “determined from the token embeddings” to be the same as a center value of the second classification in the same feature space of the interaction data because they are both directed to a representative central value computed from the embedded representation, i.e. the feature space, of the data being classified.) generate a second intermediate embedding based on providing a second input to the at least one machine learning model, wherein the second input comprises the center value of the first classification in the feature space of the interaction data; ([Friedrich, [0003]] “determining a first classification for the embedding at a first classifier, determining if the first classification meets a first condition” and [Friedrich, [0005]] “providing the embedding and/or a hidden state of the first classifier resulting from the provided embedding as input to the second classifier. The classifiers of different levels in a hierarchy are assigned to different levels in a hierarchy”, wherein the examiner interprets “determining a first classification for the embedding at a first classifier” to be the same as a first classification of the interaction data because they are both directed to producing a classification result based on input data processed by a model. The examiner further interprets “providing the embedding and/or a hidden state of the first classifier resulting from the provided embedding as input to the second classifier” to be the same as providing a second input to the at least one machine learning model because they are both directed to supplying information derived from a first model into a subsequent model stage). generate a third intermediate embedding based on providing a third input to the at least one machine learning model, wherein the third input comprises the center value of the second classification in the feature space of the interaction data; ([Friedrich, col 1, lines 35-44] “in accordance with an example embodiment of the present invention, the method comprises determining a second classification at a second classifier,” and [Friedrich, [0005]] “The classification or the hidden state of the classifier representing a parent may serve as input for the classifier representing its child.”, wherein the examiner interprets “The classification or the hidden state of the classifier representing a parent may serve as input for the classifier representing its child” to be the same as providing a third input to the at least one machine learning model because they are both directed to supplying information from a prior classification stage as input to a subsequent machine learning model stage. The examiner further interprets “determining a second classification at a second classifier” and “The classification or the hidden state of the classifier representing a parent” to be the same as a center value of a second classification of the interaction data because they are both directed to a representative value derived from the second classification that is used for subsequent processing.) Mao and Friedrich do not teach wherein, when generating the first intermediate embedding, the at least one processor is programmed or configured to: generate the first intermediate embedding based on providing the first input to a first machine learning model…wherein, when generating the third intermediate embedding, the at least one processor is programmed or configured to: generate the third intermediate embedding based on providing a third input to a third machine learning model…wherein the first machine learning model, the second machine learning model, and the third machine learning model are each neural network machine learning models…generate an output classification label of the combined final input based on providing the combined final input to a head machine learning model, wherein the head machine learning model is configured to provide a prediction of a classification label of an input as an output. Shazeer teaches: wherein, when generating the first intermediate embedding, the at least one processor is programmed or configured to: generate the first intermediate embedding based on providing the first input to a first machine learning model; ([Shazeer, col 1, lines 35-44] “The neural network includes a Mixture of Experts (MoE) subnetwork between a first neural network layer and a second neural network layer in the neural network. The MoE subnetwork includes a plurality of expert neural networks, in which each expert neural network is configured to process a first layer output generated by the first neural network layer in accordance with a respective set of expert parameters of the expert neural network to generate a respective expert output.” AND [Shazeer, col 4, lines 35-39] “In particular, as shown in FIG. 1, the neural network 102 includes a MoE subnetwork 130 arranged between a first neural network layer 104 and a second neural network layer 108 in the neural network 102. The first neural network layer 104 and the second neural network layer 108”, wherein the examiner interprets “each expert neural network is configured to process a first layer output generated by the first neural network layer” to be the same as generate the first intermediate embedding based on providing the first input to a first machine learning model because they are both directed to a first machine learning model receiving a first input for processing. The examiner further interprets “generate a respective expert output” to be the same as the first intermediate embedding because they are both directed to an intermediate output representation generated by that first machine learning model.) wherein, when generating the third intermediate embedding, the at least one processor is programmed or configured to: generate the third intermediate embedding based on providing a third input to a third machine learning model; ([Shazeer, Abstract] “The MoE subnetwork includes a plurality of expert neural networks, in which each expert neural network is configured to process a first layer output generated by the first neural network layer in accordance with a respective set of expert parameters of the expert neural network to generate a respective expert output.”, wherein the examiner interprets “each expert neural network” and “a plurality of expert neural networks” to be the same as a third machine learning model because they are both directed to a model that performs machine learning processing, wherein the examiner further interprets “process a first layer output generated by the first neural network layer” to be the same as providing a third input because they are both directed to supplying an input representation to that model for processing, and wherein the examiner further interprets “generate a respective expert output” to be the same as the third intermediate embedding because they are both directed to an intermediate output representation generated by that model). wherein the first machine learning model, the second machine learning model, and the third machine learning model are each neural network machine learning models; ([Shazeer, col 4, lines 57-59] “Each expert neural network can be a feed-forward neural network with its own parameters”, wherein the examiner interprets “Each expert neural network can be a feed-forward neural network” to be the same as the first machine learning model, the second machine learning model, and the third machine learning model are each neural network machine learning models because they are both directed to machine learning models implemented as neural networks.) generate an output classification label of the combined final input based on providing the combined final input to a head machine learning model, wherein the head machine learning model is configured to provide a prediction of a classification label of an input as an output. ([Shazeer, Abstract] “ provide the MoE [Mixture of Experts] output as input to the second neural network layer” and [Shazeer, col 3, lines 38-46] “the system 100 includes a neural network 102 that can be configured to receive any kind of digital data input and to generate any kind of score, classification, or regression output based on the input. For example, if the inputs to the neural network 102 are images or features that have been extracted from images, the output generated by the neural network 102 for a given image may be scores for each of a set of object categories, with each score representing an estimated likelihood that the image contains an image of an object belonging to the category”, wherein the examiner interprets “provide the MoE output as input to the second neural network layer” to be the same as providing the combined final input to a head machine learning model because they are both directed to supplying a processed intermediate output to a subsequent neural network stage for further processing. The examiner further interprets “generate any kind of score, classification, or regression output based on the input” and “the output generated by the neural network 102 for a given image may be scores for each of a set of object categories, with each score representing an estimated likelihood that the image contains an image of an object belonging to the category” to be the same as generate an output classification label of the combined final input because they are both directed to producing a classification output from an input provided to a neural network model.) Mao, Friedrich, Shazeer, and the instant application are analogous art because they are all directed to machine learning architectures that generate embeddings from input data, process those embeddings through multiple model stages, and combine intermediate representations and generate a final classification output. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the feature embedding technique disclosed by Mao to include the technique which provides the second classifier input from the output of the first classifier disclosed by Friedrich. One would be motivated to do so to effectively enable hierarchical classification and propagation of learned representations across multiple classifier stages, as suggested by Friedrich ([Friedrich, [0005]] “The classification or the hidden state of the classifier representing a parent may serve as input for the classifier representing its child.” AND [Friedrich, [0059]] “an aggregation function such as average”). It would have also been obvious to a person of ordinary skill in the art before the effective filing date of the invention to include the output mixture technique disclosed by Shazeer. One would be motivated to do so to efficiently enable downstream neural network layers to operate on combined intermediate outputs for improved prediction performance and generate intermediate output representations from separate neural network model stages, as suggested by Shazeer ([Shazeer, Abstract] “provide the Mixture of Experts [MoE] output as input to the second neural network layer.” AND ([Shazeer, col 1, lines 35-44] “each expert neural network is configured to process a first layer output generated by the first neural network layer in accordance with a respective set of expert parameters of the expert neural network to generate a respective expert output”). Claims 8 and 15 are analogous to claim 1, aside from claim type and minute differences, and thus face the same rejection. Regarding claim 3, Mao, Friedrich, and Shazeer teach The system of claim 1, (see the rejection of claim 1). Shazeer further teaches and wherein an input layer of each of the first machine learning model, the second machine learning model, and the third machine learning model are the same size. ([Shazeer, col 4, lines 58-62] “The expert neural networks are configured to receive the same sized inputs and produce the same-sized outputs. In some implementations, the expert neural networks are feed-forward neural networks with identical architectures, but with different parameters”, wherein the examiner interprets “The expert neural networks are configured to receive the same sized inputs and produce the same-sized outputs.” to be the same as an input layer of each of the first machine learning model, the second machine learning model, and the third machine learning model are the same size because they are both directed to multiple machine learning models being configured to accept inputs of the same dimensional size.) Mao, Friedrich, Shazeer, and the instant application are analogous art because they are all directed to machine learning model architectures that generate intermediate representations using multiple machine learning models. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the system of claim disclosed by Mao, Friedrich, and Shazeer to include the mixture of expert (MoE) neural networks disclosed by Shazeer. One would be motivated to do so to effectively enable additional intermediate model stages using multiple neural network models having consistent input sizing for downstream processing, as suggested by Shazeer ([Shazeer, col 4, lines 58-62] “The expert neural networks are configured to receive the same sized inputs and produce the same-sized outputs.”). Claims 10 and 17 are analogous to claim 3, aside from claim type and minute differences, and thus face the same rejection. Regarding claim 4, Mao, Friedrich, and Shazeer teach The system of claim 1, (see the rejection of claim 1) Mao further teaches: wherein the at least one processor is further programmed or configured to: train the gating machine learning model and the head machine learning model ([Mao, page 2] “To address these problems, in this paper, we build an enhanced two-stream MLP model, namely FinalMLP, which integrates feature gating and interaction aggregation layers on top of two MLP module networks. More specifically, we propose a stream-specific feature gating layer that allows obtaining gating-based feature importance weights for soft feature selection.”, wherein the examiner interprets “feature gating layer” to be the same as the gating machine learning model because they are both directed to a model component that performs gating-based processing of learned representations. The examiner further interprets “two MLP module networks” and “interaction aggregation layers” to be the same as the head machine learning model because they are both directed to downstream machine learning model components that process and aggregate intermediate outputs.) using a binary cross-entropy loss function, ([Mao, page 5] “To train FinalMLP, we apply the widely used binary cross-entropy loss”, wherein the examiner further interprets “To train FinalMLP, we apply the widely used binary cross-entropy loss” to be the same as train the gating machine learning model and the head machine learning model using a binary cross-entropy loss function because they are both directed to training the recited machine learning architecture with binary cross-entropy as the training loss.) Claims 11 and 18 are analogous to claim 4, aside from claim type and minute differences, and thus face the same rejection. Regarding claim 6, Mao, Friedrich, and Shazeer teach The system of claim 1, (see the rejection of claim 1) Mao further teaches: wherein, when combining the first intermediate embedding and the intermediate product of outputs to provide the combined final input, the at least one processor is programmed or configured to: ([Mao, page 2] “Stream-Level Fusion is required to fuse the outputs of two streams to obtain the final predicted click probability ŷ.”, wherein the examiner interprets “fuse the outputs of two streams” to be the same as combining the first intermediate embedding and the intermediate product of outputs to provide the combined final input because they are both directed to merging separate intermediate representations into one fused representation for subsequent processing.). concatenate the first intermediate embedding and the intermediate product of outputs to provide the combined final input. ([Mao, page 4] “Bilinear Fusion As mentioned before, existing work mostly employs summation or concatenation as the fusion layer…Given the concatenated feature embeddings e, we can then obtain weighted feature outputs h1 and h2 via element-wise product ⊙.”, wherein the examiner interprets “concatenation as the fusion layer…Given the concatenated feature embeddings e,” to be the same as concatenate the first intermediate embedding and the intermediate product of outputs to provide the combined final input because they are both directed to using concatenation to combine separate intermediate representations into a single fused input representation). wherein the examiner interprets “concatenation as the fusion layer” to be the same as concatenate the first intermediate embedding and the intermediate product of outputs to provide the combined final input because they are both directed to using concatenation to combine separate intermediate representations into a single fused input representation). Claims 13 and 20 are analogous to claim 6, aside from claim type and minute differences, and thus face the same rejection. Regarding claim 7, Mao, Friedrich, and Shazeer teach The system of claim 1, (see the rejection of claim 1) Mao further teaches wherein the gating machine learning model ([Mao, page 4] “Gate_i denotes an MLP-based gating network, which takes stream-specific conditional features xi as input and outputs element-wise gating weights gi.”, wherein the examiner interprets “MLP-based gating network” to be the same as the gating machine learning model because they are both directed to a machine learning model that performs gating-related processing.). Friedrich further teaches and the head machine learning model each comprises a binary classification machine learning model. ([Friedrich, [0057]] “The task specific classifier in the example comprises a dense layer with a softmax output 324 of dimension 2.” AND ([Friedrich, [0008]] “Preferably the first classification and /or the second classification is a binary classification.” AND [Friedrich, [0050]] “The binary classification in the example is true if the instance is assigned to the label and false otherwise.” and [Friedrich, [0057]] “The task specific classifier in the example makes a binary classification to predict whether the given instance, e.g. the text or a word sequence from the text or the digital image, belongs to a particular class or not.”, wherein the examiner interprets “task specific classifier” to be the same as the head machine learning model because they are both directed to a downstream machine learning model that produces the classification output. The examiner further interprets “binary classification” and “task specific classifier” to be the same as binary classification machine learning model because they are both directed to a machine learning classifier configured to determine one of two classification outcomes.). Mao, Friedrich, Shazeer, and the instant application are analogous art because they are all directed to machine learning model architectures that use model components to produce classification outputs. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the system claim 1 disclosed by Mao, Friedrich, and Shazeer to include the task specific binary classifier disclosed by Friedrich. One would be motivated to do so to effectively configure the downstream model to produce one of two possible classification outcomes for improved classification decision-making, as suggested by Friedrich ([Friedrich, [0057]] “The task specific classifier in the example makes a binary classification to predict whether the given instance, e.g. the text or a word sequence from the text or the digital image, belongs to a particular class or not.”) Claims 14 and 21 are analogous to claim 7, aside from claim type and minute differences, and thus face the same rejection. Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to DEVAN KAPOOR whose telephone number is (703)756-1434. The examiner can normally be reached Monday - Friday: 9:00AM - 5:00 PM EST (times may vary). Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Yi can be reached at (571) 270-7519. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DEVAN KAPOOR/Examiner, Art Unit 2126 /DAVID YI/Supervisory Patent Examiner, Art Unit 2126
Read full office action

Prosecution Timeline

Aug 02, 2023
Application Filed
Apr 28, 2026
Non-Final Rejection mailed — §103
Jun 05, 2026
Interview Requested
Jul 07, 2026
Examiner Interview Summary
Jul 07, 2026
Applicant Interview (Telephonic)
Jul 28, 2026
Response Filed
Sep 22, 2026
Final Rejection mailed — §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
7%
Grant Probability
18%
With Interview (+11.1%)
4y 4m (~1y 2m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 14 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month