Prosecution Insights
Last updated: October 02, 2026
Application No. 18/605,901

System and Method for Transformation of Discrete Input for Adversarial Robustness

Non-Final OA §103
Filed
Mar 15, 2024
Examiner
GONZALES, VINCENT
Art Unit
Tech Center
Assignee
Mitsubishi Electric Corporation
OA Round
1 (Non-Final)
79%
Grant Probability
Favorable
1-2
OA Rounds
10m
Est. Remaining
90%
With Interview

Examiner Intelligence

Grants 79% — above average
79%
Career Allowance Rate
424 granted / 539 resolved
+18.7% vs TC avg
Moderate +11% lift
Without
With
+11.2%
Interview Lift
resolved cases with interview
Typical timeline
3y 5m
Avg Prosecution
14 currently pending
Career history
557
Total Applications
across all art units

Statute-Specific Performance

§101
21.0%
-19.0% vs TC avg
§103
41.7%
+1.7% vs TC avg
§102
13.8%
-26.2% vs TC avg
§112
14.3%
-25.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 539 resolved cases

Office Action

§103
Detailed Action This action is written in response to the application filed March 15, 2024. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Subject Matter Eligibility In determining whether the claims are subject matter eligible, the examiner has considered and applied guidance from MPEP § 2106. The examiner finds that the independent claims are directed to the practical application of using a neural network to embed and process discrete input data. Furthermore, the combination of steps performed in the recited method cannot be practically performed as a mental process. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103(a) which forms the basis for all obviousness rejections set forth in this Office action: (a) A patent may not be obtained though the invention is not identically disclosed or described as set forth in section 102 of this title, if the differences between the subject matter sought to be patented and the prior art are such that the subject matter as a whole would have been obvious at the time the invention was made to a person having ordinary skill in the art to which said subject matter pertains. Patentability shall not be negatived by the manner in which the invention was made. The examiner cites the following references in the rejections below: Goodfellow (Goodfellow, Ian J., Jonathon Shlens, and Christian Szegedy. "Explaining and harnessing adversarial examples." arXiv preprint arXiv:1412.6572 (2014).) Ikeda (US 2022/0129764 A1) Li (Li, Bai, et al. "Certified adversarial robustness with additive noise." Advances in neural information processing systems 32 (2019). arXiv preprint arXiv:1809.03113. Cited by Applicant in IDS dated 3/15/24.) Mikolov (Mikolov, Tomas, et al. "Efficient estimation of word representations in vector space." arXiv preprint arXiv:1301.3781 (2013).) Mittal (US 2023/0215427 A1) Claims 1-2, 7-9, 13-14, 16 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Mikolov and Goodfellow. Regarding claims 1, 16 and 20, Mikolov discloses a computer-implemented artificial intelligence (AI) method (and a related system and non-transitory computer-readable medium) for robust transformation of a discrete input with a neural network including an embedding subnetwork and a transformation subnetwork, comprising: embedding the discrete input into a continuous space using the embedding subnetwork to produce a continuous embedding; PP. 2-3, “2 Model Architectures Many different types of models were proposed for estimating continuous representations of words, including the well-known Latent Semantic Analysis (LSA) and Latent Dirichlet Allocation (LDA). In this paper, we focus on distributed representations of words learned by neural networks, as it was previously shown that they perform significantly better than LSA for preserving linear regularities among words [20, 31]; LDA moreover becomes computationally very expensive on large data sets.” Goodfellow discloses the following further limitation comprising: injecting a set of random noises of a predetermined magnitude into the continuous embedding to produce a set of perturbed embeddings; P. 6, “As control experiments, we trained training a maxout network with noise based on randomly adding ± ϵ to each pixel, or adding noise in U(-ϵ, ϵ ) to each pixel. These obtained an error rate of 86.2% with confidence 97.3% and an error rate of 90.4% with a confidence of 97.8% respectively on fast gradient sign adversarial examples.” processing each of the set of perturbed embeddings with the transformation subnetwork to produce a set of transformations; and P. 5, “Training on adversarial examples is somewhat different from other data augmentation schemes; usually, one augments the data with transformations such as translations that are expected to actually occur in the test set. This form of data augmentation instead uses inputs that are unlikely to occur naturally but that expose flaws in the ways that the model conceptualizes its decision function.” outputting a combination of the set of transformations as the robust transformation of the discrete input. Id. P. 3 “we train a single model to recognize labels y ∈ {-1, 1} with P(y = 1) = σ (wTx + b) where σ(z) is the logistic sigmoid function”. At the time of filing, it would have been obvious to a skilled machine learning engineer to combine the technique for continuous embedding of discrete inputs (as taught by Mikolov) with the technique for augmenting training examples with additive noise (as taught by Goodfellow). Both disclosures illustrate that the techniques disclosed therein improve neural network performance: “We observed that it is possible to train high quality word vectors using very simple model architectures, compared to the popular neural network models (both feedforward and recurrent). Because of the much lower computational complexity, it is possible to compute very accurate high dimensional word vectors from a much larger data set.” Mikolov, p. 10. “The adversarial training procedure can be seen as minimizing the worst case error when the data is perturbed by an adversary.” Goodfellow, p. 5. Regarding independent claims 16 and 20, the computer hardware components recited therein (ie a memory and a processor, and a non-transitory computer readable storage medium) are inherent throughout both Mikolov and Goodfellow. Regarding claim 2, Goodfellow discloses the further limitation wherein wherein the discrete input includes a tensor of one or more categorical values. P. 3 “we train a single model to recognize labels y ∈ {-1, 1} with P(y = 1) = σ (wTx + b) where σ(z) is the logistic sigmoid function”. Regarding claim 7, Mikilov discloses the further limitation wherein the transformation subnetwork is a deep neural network trained for one or a combination of: automatic speech recognition, language modelling, and log data modelling. P. 1, sec. 1, “automatic speech recognition”. Regarding claim 8, Mikolov discloses the further limitation wherein the combination of the set of transformations is an aggregation of the set of transformations. P. 10, “By using ten examples instead of one to form the relationship vector (we average the individual vectors together), we have observed improvement of accuracy of our best models by about 10% absolutely on the semantic-syntactic test.” Regarding claim 9, Mikolov discloses the further limitation wherein each of the set of transformations is a continuous tensor, such that the robust transformation is determined as an average of the set of transformations. P. 10, “By using ten examples instead of one to form the relationship vector (we average the individual vectors together), we have observed improvement of accuracy of our best models by about 10% absolutely on the semantic-syntactic test.” Regarding claim 13, Goodfellow discloses the further limitation wherein the neural network is trained with training samples of the discrete input. P. 3 “we train a single model to recognize labels y ∈ {-1, 1} with P(y = 1) = σ (wTx + b) where σ(z) is the logistic sigmoid function”. Regarding claim 14, Goodfellow discloses the further limitation wherein the neural network is trained with training samples of the discrete input modified with noise in the continuous space. P. 3 “we train a single model to recognize labels y ∈ {-1, 1} with P(y = 1) = σ (wTx + b) where σ(z) is the logistic sigmoid function”. P. 5, “That can be interpreted as learning to play an adversarial game, or as minimizing an upper bound on the expected cost over noisy samples with noise from U(-ϵ , ϵ ) added to the inputs.” Claims 3 are rejected under 35 U.S.C. 103 as being unpatentable over Mikolov, Goodfellow and Ikeda. Regarding claim 3, Ikeda discloses the further limitation which Mikolov/Goodfellow do not disclose wherein the discrete input is internet proxy log data. [0049] Hereinafter, a method of defining the node feature vector will be described using a specific example. FIG. 4 is a table showing an example of the proxy log data, and shows an example of an access log of a proxy server. The proxy log data shown in FIG. 4 includes information about time, domain, method, path, the number of transmission bytes, the number of reception bytes, and client IP. At the time of filing, it would have been obvious to a skilled machine learning engineer to apply the combined system of Mikolov/Goodfellow to the problem of anomaly detection within a computer network (as addressed by Ikeda) because the latter is a classification task, and neural networks are widely used to address such tasks. Proxy server log data can help identify anomalies within network traffic. Claims 4 are rejected under 35 U.S.C. 103 as being unpatentable over Mikolov, Goodfellow and Mittal. Regarding claim 4, Mittal discloses the further limitation which nether Mikolov/Goodfellow discloses wherein the continuous embedding includes a tensor of floating-point values. [0021] FIG. 4 depicts user 402 providing a spoken utterance to ASR encoder 404, which processes the spoken utterance and generates an output embedding (e.g., an advanced representation of speech input (e.g., in the form of a float vector)). At the time of filing, it would have been obvious to a skilled machine learning engineer to use floating point values to store values (as taught by Mittal) in combination with the Mikolov/Goodfellow system because they are one of the fundamental data types for computing, and the one most suitable for storing the decimal values commonly used in neural networks. Claims 5-6 are rejected under 35 U.S.C. 103 as being unpatentable over Mikolov, Goodfellow, Mittal and Li. Regarding claim 5, Li discloses the further limitation wherein 4, wherein the set of random noises includes a set of Gaussian noise tensors, wherein each of the Gaussian noise tensors has a shape of the tensor of floating-point values and includes independent Gaussian samples having a mean of zero and a standard deviation defined by the predetermined magnitude. P. 4, algorithm 1, “Add i.i.d. Gaussian noise N(0; σ2) to each pixel of x and apply the classifier f on it. Let the output be ci = f(x + N(0; σ2I)). At the time of filing, it would have been obvious to a skilled machine learning engineer to apply the technique disclosed by Li for adding gaussian noise with a mean of zero and a specified standard deviation to the combined system of Mikolov/Goodfellow/Mittal because a Gaussian distribution requires these two parameters (ie a mean and a standard deviation). A zero mean is a typical default parameter because it does not alter the mean of the distribution to which it is added. Regarding claim 6, Goodfellow discloses the further limitation wherein 5, wherein each of the perturbed embeddings is formed by adding the tensor of floating-point values to one of the Gaussian noise tensors. P. 11, “To solve the problem introduced by Nguyen et al. (2014) of generating a fooling image for a particular class, we propose adding rxp(y = i j x) to a Gaussian sample x as a fast method of generating a fooling image classified as class i.” Claim Objections and Allowable Subject Matter Dependent claims 10-12, 15 and 17-19 are allowable over the prior art, but are objected to as depending upon a rejected parent claim. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Vincent Gonzales whose telephone number is (571) 270-3837. The examiner can normally be reached on Monday-Friday 7 a.m. to 4 p.m. MT. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Miranda Huang, can be reached at (571) 270-7092. Information regarding the status of an application may be obtained from the USPTO Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. /Vincent Gonzales/Primary Examiner, Art Unit 2124
Read full office action

Prosecution Timeline

Mar 15, 2024
Application Filed
Aug 18, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737626
JOINT INPUT PERTUBATION AND TEMPERATURE SCALING FOR NEURAL NETWORK CALIBRATION
3y 3m to grant Granted Sep 15, 2026
Patent 12705472
Prefetching Weights For Use In A Neural Network Processor
2y 9m to grant Granted Aug 11, 2026
Patent 12675989
FUSION MODEL TRAINING USING DISTANCE METRICS
2y 3m to grant Granted Jul 07, 2026
Patent 12651182
IDENTIFYING TRAITS OF PARTITIONED GROUP FROM IMBALANCED DATASET
4y 11m to grant Granted Jun 09, 2026
Patent 12639623
FAIR SELECTIVE CLASSIFICATION VIA A VARIATIONAL MUTUAL INFORMATION UPPER BOUND FOR IMPOSING SUFFICIENCY
4y 4m to grant Granted May 26, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
79%
Grant Probability
90%
With Interview (+11.2%)
3y 5m (~10m remaining)
Median Time to Grant
Low
PTA Risk
Based on 539 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month