Prosecution Insights
Last updated: October 02, 2026
Application No. 18/278,473

CONTINUAL LEARNING NEURAL NETWORK SYSTEM TRAINING FOR CLASSIFICATION TYPE TASKS

Final Rejection §103
Filed
Aug 23, 2023
Priority
May 27, 2021 — provisional 63/194,056 +2 more
Examiner
LAHAM BAUZO, ALVARO SALIM
Art Unit
2146
Tech Center
2100 — Computer Architecture & Software
Assignee
DeepMind Technologies Limited
OA Round
2 (Final)
50%
Grant Probability
Moderate
3-4
OA Rounds
8m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 50% of resolved cases
50%
Career Allowance Rate
4 granted / 8 resolved
-5.0% vs TC avg
Strong +100% interview lift
Without
With
+100.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 9m
Avg Prosecution
23 currently pending
Career history
39
Total Applications
across all art units

Statute-Specific Performance

§101
15.6%
-24.4% vs TC avg
§103
62.4%
+22.4% vs TC avg
§102
6.9%
-33.1% vs TC avg
§112
15.0%
-25.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 8 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Amendments This Office Action is in response to the amendment filed on 07/02/2026. Claims 1-2, 6, 9, 14-15, 17, and 28-29 have been amended. Claims 12 and 13 have been cancelled. Claims 30 and 31 have been added. The objections and rejections from the prior correspondence that are not restated herein are withdrawn. Information Disclosure Statement The information disclosure statement(s) (IDS) submitted on 06/30/2026 is/are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement(s) is/are being considered by the examiner. Response to Arguments Applicant's arguments filed on 07/02/2026 have been fully considered. Applicant's arguments regarding the 35 U.S.C. 101 rejections of the previous office action have been fully considered and are persuasive. The claims reflect an improvement in how the neural network-based system is trained. Specifically, the claims update only the parameters of the selected subset of neural networks based upon the determined update, which mitigates catastrophic forgetting in continual learning settings. Therefore, the 35 U.S.C. 101 rejections of the previous office action are withdrawn. Applicant's arguments regarding the 35 U.S.C. 102/103 rejections of the previous office action have been fully considered but are not persuasive. Applicant argues: “Applicant respectfully submits that the cited references do not teach or suggest at least this combination of features. The Office Action asserts that Shazeer describes "selecting a subset of neural networks from a plurality of neural networks stored in a memory based upon the encoding" (Office Action, page 22). However, the cited portions of Shazeer do not describe "selecting the subset of neural networks based upon ... similarity measures", where the similarity measures are similarity measures between "(i) the encoding of the training data item in the latent embedding space, and (ii) the key embedding representing the neural network in the latent embedding space," as recited in amended claim 1. The Office Action asserts that Meidar describes "each of the plurality of neural networks are associated with a respective key", "determining a similarity between the encoding and each respective key", and "selecting a subset of neural networks ... based upon the determined similarity" (Office Action, pages 36-38). However, the cited portions of Meidar describe "a set of identified characteristics serves ... as a key to retrieve a set of classifiers" (Meidar, paragraph [0084]). A set of characteristics serving as a key does not disclose or suggest a "key embedding in the latent embedding space of the encoder", as recited in amended claim 1. Neither does a set of characteristics serving as a key disclose or suggest "determining ... a respective similarity measure between ... the encoding of the training data item in the latent embedding space, and ... the key embedding representing the neural network in the latent embedding space", as recited in amended claim 1.” Examiner respectfully disagrees. SHAZEER is not relied upon for teaching selecting the subset of neural networks based upon the similarity measures, which is taught by MEIDAR. MEIDAR teaches two processes. The Deep Features extracted from the product images by a DNN are established as the retrieval key of each classifier and stored in a characteristics database (MEIDAR [0042], [0076], [0085]). Then, the features extracted from input images of the inserted product are compared with the stored retrieval keys using a nearest neighbor algorithm (i.e., determining a respective similarity measure), and the set of identified characteristics serves as a key to retrieve a set of classifiers (i.e., selecting a subset of neural networks) from the classifier database (MEIDAR [0044] and [0084]). Thus, the set of identified characteristics in MEIDAR [0084] is the retrieval key established in the first process. Under BRI, key embedding in the latent space of the encoder can be interpreted as MEIDAR’s retrieval key, which is a numerical representation of the Deep Features produced by the DNN (MEIDAR [0054]). In the combination, SHAZEER’s first neural network layer 104 (i.e., encoder) generates the first layer output 124 (i.e., encoding), which is compared with MEIDAR’s retrieval keys. Therefore, SHAZEER in view of MEIDAR teaches the amended limitations of claims 1, 28, and 29, as shown in the 35 U.S.C. 103 rejections below. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 14, 18, and 28-29 are rejected under 35 U.S.C. 103 as being unpatentable over SHAZEER (US 20190251423 A1) in view of MEIDAR (US 20210312205 A1), hereafter SHAZEER and MEIDAR, respectively. Regarding Claim 1: SHAZEER teaches: A method performed by one or more computers, the method comprising: (SHAZEER [0008] teaches: "Another innovative aspect of the subject matter described in this specification can be embodied in a method including: receiving a network input; and processing the network input using the system described above to generate a network output for the network input." SHAZEER [0037] teaches: “For convenience, the process will be described as being performed by a system of one or more computers located in one or more locations. For example, a neural network system, e.g., the neural network system 100 of FIG. 1, appropriately programmed, can perform the process.”) training a neural network-based system by a machine learning training technique, the training comprising: (SHAZEER [0054] teaches: “During training of the neural network 102, the processes 200 and 300 can be used as part of generating a network output for a training input.” SHAZEER [0063] teaches: “Data processing apparatus for implementing machine learning models can also include, for example, special-purpose hardware accelerator units for processing common and compute-intensive parts of machine learning training or production, i.e., inference, workloads.” SHAZEER [0064] teaches: “Machine learning models can be implemented and deployed using a machine learning framework, e.g., a TensorFlow framework, a Microsoft Cognitive Toolkit framework, an Apache Singa framework, or an Apache MXNet framework.”) (a) receiving a training data item and target data associated with the training data item; (SHAZEER [0054] teaches: "During training of the neural network 102, the processes 200 and 300 can be used as part of generating a network output for a training input (i.e., a training data item). The gradient of an objective function can be backpropagated to adjust the values of the parameters of various components of the neural network 102 to improve the quality of the network output relative to a known output for the training input (i.e., target data associated with the training data item)." Examiner's note: The network generates outputs for the received training input (i.e., receiving a training data item).) (b) processing the training data item using an encoder to generate an encoding of the training data item in a latent embedding space of the encoder; (SHAZEER [0003] teaches: “Neural networks are machine learning models that employ one or more layers of nonlinear units to predict an output for a received input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to another layer in the network, i.e., another hidden layer or the output layer. Each layer of the network generates an output from a received input in accordance with current values of a respective set of parameters.” SHAZEER [0017] teaches: “For example, if the inputs to the neural network 102 are images (i.e., training data item) or features that have been extracted from images, the output generated by the neural network 102 for a given image may be scores for each of a set of object categories, with each score representing an estimated likelihood that the image contains an image of an object belonging to the category.” SHAZEER [0024] teaches: “the neural network 102 includes a MoE subnetwork 130 arranged between a first neural network layer 104 and a second neural network layer 108 in the neural network 102.” SHAZEER [0025] teaches: “Each expert neural network in the MoE subnetwork 130 can be configured to process a first layer output 124 (i.e., encoding) generated by the first neural network layer 104 (i.e., processing the training data item using an encoder to generate an encoding of the training data item) in accordance with a respective set of expert parameters of the expert neural network to generate a respective expert output.” SHAZEER [0024] teaches: “the neural network 102 includes a MoE subnetwork 130 arranged between a first neural network layer 104 and a second neural network layer 108 in the neural network 102. The first neural network layer 104 and the second neural network layer 108 can be any kind of neural network layer, for example, a LSTM neural network layer or other recurrent neural network layer, a convolutional neural network layer, or a fully-connected neural network layer." Examiner's note: LSTM and convolutional neural networks are types of encoders, and thus their outputs are encodings. Additionally, one of ordinary skill in the art would recognize that the output of the first neural network layer 104, such as an LSTM layer, is a hidden representation (i.e., encoding of the training data item in a latent embedding space of the encoder) of the received input image.) (c) selecting a subset of neural networks from a plurality of neural networks […] and stored in a memory, comprising: (SHAZEER [0027] teaches: "Although the MoE subnetwork includes a large number of expert neural networks, only a small number of them are selected during the processing of any given network input by the neural network 102, e.g., only a small number of the expert neural networks are selected (i.e., selecting a subset of neural networks from a plurality of neural networks) to process the first layer output 124. The neural network 102 includes a gating subsystem 110 that is configured to select, based on the first layer output 124, one or more of the expert neural networks, and determine a respective weight for each selected expert neural network." SHAZEER [0060] teaches: "Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data." Examiner's note: The neural networks are part of the data stored in memory used for executing the instructions.) wherein the plurality of neural networks are configured to process the encoding to generate output data indicative of a classification of an aspect of the training data item; (SHAZEER [0017] teaches: "For example, if the inputs to the neural network 102 are images or features that have been extracted from images (i.e., process the encoding), the output generated by the neural network 102 for a given image may be scores for each of a set of object categories, with each score representing an estimated likelihood that the image contains an image of an object belonging to the category (i.e., classification of an aspect of the training data item).” SHAZEER [0027] teaches: “The gating subsystem 110 then provides the first layer output as input to each of the selected expert neural networks (i.e., the plurality of neural networks are configured to process the encoding). The gating subsystem 110 combines the expert outputs generated by the selected expert neural networks (i.e., to generate output data) in accordance with the weights for the selected expert neural networks to generate a MoE output 132." SHAZEER [0016] teaches: “Generally, the system 100 includes a neural network 102 that can be configured to receive any kind of digital data input and to generate any kind of score, classification, or regression output based on the input.” Examiner’s note: Under BRI, configured to process the encoding to generate output data indicative of a classification can be interpreted as the mixture of experts (MoE) whose outputs are combined to generate a MoE output, which can be a classification score, as described in SHAZEER [0016].) (d) processing the encoding using the selected subset of neural networks to generate the output data; (SHAZEER [0041] teaches: “The system provides the first layer output as input to each of the selected expert neural networks (step 206). Each of the selected expert neural networks is configured to process the first layer output in accordance with the respective current values of parameters of the selected expert neural network to generate a respective expert output.”) (e) determining an update to the parameters of the selected subset of neural networks based upon a loss function comprising a relationship between the generated output data and the target data associated with the training data item; (SHAZEER [0054] teaches: "The gradient of an objective function (i.e., determining an update [...] based upon a loss function) can be backpropagated to adjust the values of the parameters of various components of the neural network 102 (i.e., to the parameters of the selected subset of neural networks) to improve the quality of the network output relative to a known output for the training input (i.e., comprising a relationship between the generated output data and the target data associated with the training data item).") (f) updating the parameters of the selected subset of neural networks based upon the determined update. (SHAZEER [0054] teaches: "The gradient of an objective function (i.e., based upon the determined update) can be backpropagated to adjust the values of the parameters of various components of the neural network 102 (i.e., updating the parameters of the selected subset of neural networks) to improve the quality of the network output relative to a known output for the training input.") SHAZEER is not relied upon for teaching, but MEIDAR teaches: […] neural networks each associated with a respective key embedding in the latent embedding space of the encoder […]: (MEIDAR [0042] teaches: “Another type of product characteristics can be provided from a Neural networks such as Deep Neural Network (DNN). A specific example for this can be utilizing a Convolutonal Neural Network (CNN), such as Inception, Xception or any other, to extract image related features. In literature, such features produced by a DNN are termed ‘Deep Features’. Such ‘Deep Features’ can be extracted a single or multiple network layers.” MEIDAR [0044] teaches: “The classifiers can be selected by a nearest-neighbors (NN) or approximated nearest neighbors (ANN) algorithm. There are multiple variants that can be considered for classifier selection form the database.” MEIDAR [0047] teaches: “Other variations of the formula utilizing one or more of the characteristics for classifier retrieval and selection can be interchangeably used without altering the scope described and claimed herein. In similarity-based retrieval, the various characteristics can be presented numerically.” MEIDAR [0054] teaches: “The classifiers in the classifier data base are also associated with similar presentation. The vector can be a numerical representation (i.e., key embedding in the latent space of the encoder) of the entire set of characteristics obtained, for example by concatenation of numerical representation of extracted features) and can serve as the basis for the retrieval method described.” MEIDAR [0055] teaches: “In certain examples, the models used for creating the classifiers and retrieving the models used to retrieve the classifiers are determined using a preliminary neural network (i.e., neural networks), […]. It is further noted, that the preliminary neural network used to select the model(s) or other CNN, RNN, used for classifiers creation (in other words, building the classifier) and/or retrieval, do not necessarily reside on the AIC, but are rather located on a backend management server, configured to communicate with a transceiver in the AIC through a communication network.” MEIDAR [0076] teaches: “FIG. 1A is a flow diagram showing how product images 100 are used for characteristics extraction 101-106 such as color palette 101; box/wrap type & shape 102; weight 103, key-words & logo 104 and scale and rotation invariant features 105 and Deep Feature 106. The deep feature 106 may represent a sub-field of machine learning involving learning of a product characteristic representations provided from a deep-neural-network (DNN) trained to identify that product. A DNN can be used to provide product characteristics by collecting its response to a given input, such as a product's image, from at least one of its layers. The response of a DNN to a product image is hierarchical by nature. The term ‘Deep Features’ typically corresponds to responses produced by deeper layers of a DNN model. Such ‘Deep Features’ are typically specific to the input product or to similar products, therefore can be used as a unique descriptors for the product, such descriptors can be used as characteristics for classifier retrieval and re-identification once the product is scanned.” MEIDAR [0084] teaches: "Extracted features are input to a processing module 420, which compares them with characteristics already stored in characteristic database 430 whereupon finding a corresponding characteristic set in characteristics database 430, the set of identified characteristics serves either together as a single concatenated vector or each individually, as a key to retrieve a set of classifiers from classifier database 440." MEIDAR [0085] teaches: “Images captured by acquisition and synchronization module 410 can be transmitted to product image database 470, and to classifier training module 460, where and discrepancy between retrieved classifiers and product characteristics extracted is identified and based on extracted features, a new classifier is established, and the characteristic associated with the classifier is established as a retrieval key (i.e., key embedding in the latent space of the encoder) for that particular classifier.” MEIDAR [0087] teaches: “[…] wherein (xii) the set of executable instructions, are configured, when executed, to cause the processor to assign a retrieval key to each classifier (i.e., neural networks each associated with a respective key), […].”) determining, for each of the plurality of neural networks, a respective similarity measure between: (i) the encoding of the training data item in the latent embedding space, and (ii) the key embedding representing the neural network in the latent embedding space; (MEIDAR [0015] teaches: "The extracted product characteristics are then utilized as keys to extract a set of relevant classifiers that are needed to produce a single correct recognition." MEIDAR [0084] teaches: "Extracted features are input to a processing module 420, which compares them with characteristics already stored in characteristic database 430 (i.e., determining, for each of the plurality of neural networks, a respective similarity measure), whereupon finding a corresponding characteristic set in characteristics database 430, the set of identified characteristics serves either together as a single concatenated vector or each individually, as a key to retrieve a set of classifiers from classifier database 440." MEIDAR [0085] teaches: “Images captured by acquisition and synchronization module 410 can be transmitted to product image database 470, and to classifier training module 460, where and discrepancy between retrieved classifiers and product characteristics extracted is identified and based on extracted features, a new classifier is established, and the characteristic associated with the classifier is established as a retrieval key for that particular classifier (i.e., the key embedding representing the neural network in the latent embedding space).” MEIDAR [0087] teaches: "wherein (xii) the set of executable instructions, are configured, when executed, to cause the processor to assign a retrieval key to each classifier (i.e., for each of the plurality of neural networks), wherein (xiii) the set of plurality of the classifiers identifying a single product are selected by applying a nearest-neighbors algorithm and approximated nearest neighbors (ANN) algorithm to the classifiers associated with the product independently; and selecting the classifiers selected by both algorithms, […]." MEIDAR [0047] teaches: “Other variations of the formula utilizing one or more of the characteristics for classifier retrieval and selection can be interchangeably used without altering the scope described and claimed herein. In similarity-based retrieval, the various characteristics can be presented numerically. The conversion to a numeric presentation can be made in a straight-forward fashion for hierarchical characteristics that have numerical values by nature, such as weight, volume, and color or color-histogram. Other characteristics, for example, categorical characteristics may need to be converted to a numeric presentation prior to using them in a classifier search and retrieval method.” MEIDAR [0053] teaches: “In another exemplary implementation, each characteristic is used separately to retrieve a larger set of relevant classifiers (i.e., plurality of neural networks) followed by an intersecting step to select the set of classifiers which match multiple characteristics. Methods for extracting nearest neighbors (NN) or k-Nearest neighbors (kNN) from database can be used. The classifier database can be very large (typically more than 100 k classifiers), which implies that retrieving the appropriate classifiers by comparing each character to the entire database may be time consuming and not appropriate for real-time product recognition as needed in the AIC.” Examiner’s note: MEIDAR teaches two processes. First, the characteristics associated with each classifier are established as a retrieval key for that specific classifier, and a retrieval key is assigned to each classifier (MEIDAR [0085] and [0087]). Then, the features extracted from the input images of the inserted product are compared with the characteristics already stored in characteristics database 430 by applying a nearest neighbor algorithm and an approximated nearest neighbor algorithm, where the characteristics are presented numerically (MEIDAR [0047], [0084], and [0087]). Under BRI, determining, for each of the plurality of neural networks, a respective similarity measure between: (i) the encoding of the training data item in the latent embedding space, and (ii) the key embedding representing the neural network in the latent embedding space can be interpreted as MEIDAR’s applying a nearest neighbor algorithm to compare the extracted characteristics stored in the characteristics database as numerical representations (i.e., key embedding representing the neural network in the latent embedding space) with the extracted features of the inserted product image.) and selecting the subset of neural networks based upon the similarity measures; (MEIDAR [0084] teaches: "Extracted features are input to a processing module 420, which compares them with characteristics already stored in characteristic database 430 whereupon finding a corresponding characteristic set (i.e., based upon the similarity measures) in characteristics database 430, the set of identified characteristics serves either together as a single concatenated vector or each individually, as a key to retrieve a set of classifiers (i.e., selecting the subset of neural networks) from classifier database 440." Examiner’s note: MEIDAR [0084] teaches that the set of identified characteristics found by the comparison (i.e., based upon the similarity measures) serves as a key to retrieve a set of classifiers (i.e., selecting the subset of neural networks) from the classifier database 440.) Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of SHAZEER and MEIDAR before them, to include MEIDAR’s retrieval keys in SHAZEER’s mixture of experts (MoE) method, where SHAZEER’s first layer output 124 is compared with MEIDAR’s retrieval keys to select the expert neural network. One would have been motivated to make such a combination in order to avoid the classifier capacity problem, where the increase in the number of items used for training leads to increased error in the test dataset (MEIDAR [0004]), by splitting the recognition task among multiple classifiers, each classifier trained to recognize groups of products sharing similar characteristics, and by using the extracted product characteristics to extract a set of relevant classifiers (MEIDAR [0015]). Regarding Claim 14: SHAZEER in view of MEIDAR teaches the elements of claim 1 as outlined above. MEIDAR further teaches: wherein the similarity measure is based upon a cosine distance between the encoding and the key embedding. (MEIDAR [0048] teaches: "Other similarity metrics, such as, for example Cosine similarity [...] may be used additionally or alternatively.") Regarding Claim 18: SHAZEER in view of MEIDAR teaches the elements of claim 1 as outlined above. SHAZEER further teaches: The method of claim 1, wherein processing the encoding of the training data item using the selected subset of neural networks comprises processing the encoding through each respective neural network of the subset of neural networks to generate intermediate data for each respective neural network; (SHAZEER [0041] teaches: "The system provides the first layer output as input to each of the selected expert neural networks (step 206). Each of the selected expert neural networks is configured to process the first layer output in accordance with the respective current values of parameters of the selected expert neural network to generate a respective expert output." SHAZEER [0042] teaches: "The system then combines the expert outputs generated (i.e., intermediate data) by the selected expert neural networks in accordance with the weights for the selected expert neural networks to generate the MoE output (step 208).” SHAZEER [0043] teaches: "In particular, the system weights the expert output generated by each of the selected expert neural networks by the weight for the selected expert neural network to generate a weighted expert output. The system then sums the weighted expert outputs to generate the MoE output.") and aggregating the intermediate data for each respective neural network to generate the output data indicative of the classification of an aspect of the training data item. (SHAZEER [0017] teaches: "For example, if the inputs to the neural network 102 are images or features that have been extracted from images, the output generated by the neural network 102 for a given image may be scores for each of a set of object categories, with each score representing an estimated likelihood that the image contains an image of an object belonging to the category (i.e., classification of an aspect of the training data item)." SHAZEER [0042] teaches: "The system then combines (i.e., aggregating) the expert outputs generated (i.e., intermediate data) by the selected expert neural networks in accordance with the weights for the selected expert neural networks to generate the MoE output (step 208).” SHAZEER [0043] teaches: "In particular, the system weights the expert output generated by each of the selected expert neural networks by the weight for the selected expert neural network to generate a weighted expert output. The system then sums the weighted expert outputs to generate the MoE output (i.e., output data indicative of the classification of an aspect of the training data item).") Regarding Claim 28: The claim recites similar limitations as corresponding claim 1 and is rejected for similar reasons as claim 1 using similar teachings and rationale. Additionally, SHAZEER teaches: A system comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising: (SHAZEER [0060] teaches: "Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.”) Regarding Claim 29: The claim recites similar limitations as corresponding claim 1 and is rejected for similar reasons as claim 1 using similar teachings and rationale. Additionally, SHAZEER teaches: One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising: (SHAZEER [0056] teaches: “Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, data processing apparatus.” SHAZEER [0060] teaches: "Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.”) Claims 2-3 and 30-31 are rejected under 35 U.S.C. 103 as being unpatentable over SHAZEER in view of MEIDAR, as applied to claims 1 and 29 respectively above, and further in view of MCCALL (GB 2289970 A), hereafter MCCALL. Regarding Claim 2: SHAZEER in view of MEIDAR teaches the elements of claim 1 as outlined above. SHAZEER further teaches: wherein training the neural network-based system further comprises repeating steps (a) to (f) for a plurality of training data items; (SHAZEER [0047] teaches: "The system adds a noise to the modified first layer output to generate an initial gating output (step 304). The noise helps with load balancing, i.e., to encourage expert neural networks to receive roughly equal numbers of training examples in a training dataset during training, or to receive roughly equal numbers of input examples in input data during testing." SHAZEER [0054] teaches: "During training of the neural network 102, the processes 200 and 300 can be used as part of generating a network output for a training input." SHAZEER [0063] teaches: “Data processing apparatus for implementing machine learning models can also include, for example, special-purpose hardware accelerator units for processing common and compute-intensive parts of machine learning training or production, i.e., inference, workloads.” SHAZEER [0064] teaches: “Machine learning models can be implemented and deployed using a machine learning framework, e.g., a TensorFlow framework, a Microsoft Cognitive Toolkit framework, an Apache Singa framework, or an Apache MXNet framework.” Examiner's note: SHAZEER [0047] teaches the neural networks receiving more than one training sample and performing processes 200 and 300 for a training input, and thus repeating the process for each training input.) However, SHAZEER in view of MEIDAR is not relied upon for teaching, but MCCALL teaches: wherein the plurality of training data items comprises a first training data item drawn from a first data distribution and a second training data item drawn from a second data distribution, and wherein the first and second data distributions are different. (MCCALL [page 3, lines 20-24] teaches: "This is achieved by interleaving patterns from different classes rather than presenting them in an arbitrary fashion as in the prior art. In interleaved training, according to the preferred embodiment, each class is treated as a separate sequence, and patterns are taken from each class in rotation.") Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of SHAZEER, MEIDAR, and MCCALL before them, to include MCCALL’s data interleaving in SHAZEER and MEIDAR’s mixture of experts (MoE) method. One would have been motivated to make such a combination in order to solve the problem of feed-forward neural networks converging to local minima in pattern classification tasks, wherein the order of pattern presentation is modified during training so that the minority classes get the same number of presentations as the majority class. Patterns from different classes are interleaved. This prevents the training process from finding a deep local minimum which classifies all of the minority class wrongly (MCCALL [page 1, (57)]). Regarding Claim 3: SHAZEER in view of MEIDAR and MCCALL teaches the elements of claim 2 as outlined above. MCCALL further teaches: wherein the plurality of training data items comprise training data items drawn from the first data distribution interspersed with training data items drawn from the second data distribution. (MCCALL [page 3, lines 20-24] teaches: "This is achieved by interleaving patterns from different classes rather than presenting them in an arbitrary fashion as in the prior art. In interleaved training, according to the preferred embodiment, each class is treated as a separate sequence, and patterns are taken from each class in rotation.") Regarding Claim 30: SHAZEER in view of MEIDAR teaches the elements of claim 29 as outlined above. Additionally, the claim recites similar limitations as corresponding claim 2 and is rejected for similar reasons as claim 2 using similar teachings and rationale. Regarding Claim 31: SHAZEER in view of MEIDAR and MCCALL teaches the elements of claim 30 as outlined above. Additionally, the claim recites similar limitations as corresponding claim 3 and is rejected for similar reasons as claim 3 using similar teachings and rationale. Claims 4 and 5 are rejected under 35 U.S.C. 103 as being unpatentable over SHAZEER in view of MEIDAR, as applied to claim 1 above, and further in view of BARZ ("Deep Learning on Small Datasets without Pre-Training using Cosine Loss"), hereafter BARZ. Regarding Claim 4: SHAZEER in view of MEIDAR teaches the elements of claim 1 as outlined above. SHAZEER in view of MEIDAR is not relied upon for teaching, but BARZ teaches: wherein the relationship between the generated output data and the target data for the training data item is based upon a dot product between the generated output data and the target data. (BARZ [page 1373, section 3.1 Cosine Loss] teaches: "We consider the class embeddings φ as fixed and aim at learning the parameters θ of a neural network f θ by maximizing the cosine similarity between the image features and the embeddings of their classes. To this end, we define the cosine loss function to be minimized by the neural network: L c o s x , y = 1 - σ c o s f θ x , φ y                                                         ( 3 ) In practice, this is implemented as a sequence of two operations. First, the features learned by the network (with d = n ) are L 2 -normalized: φ x = x x 2 . This restricts the prediction space to the unit hypersphere, where the cosine similarity is equivalent to the dot product: L c o s x , y = 1 - φ y , ψ f θ x                                                         ( 4 ) The class embeddings φ ( y ) need to lie on the unit hypersphere as well for this equation to hold. One-hot vectors, for example, have unit-norm by definition and hence do not need to be L2-normalized explicitly. When working with batches of multiple samples, we compute the average loss over all instances in the batch.”) Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of SHAZEER, MEIDAR, and BARZ before them, to include BARZ’s cosine loss in SHAZEER and MEIDAR’s mixture of experts (MoE) method. One would have been motivated to make such a combination because the cosine loss function provides substantially better performance than cross-entropy on datasets with only a handful of samples per class (BARZ [Abstract]). Regarding Claim 5: SHAZEER in view of MEIDAR teaches the elements of claim 1 as outlined above. SHAZEER in view of MEIDAR is not relied upon for teaching, but BARZ teaches: wherein the target data is in the form of a one-hot vector. (BARZ [page 1371, section 1. Introduction] teaches: "In this work, however, we propose an extremely simple but surprisingly effective loss function for learning from scratch on small datasets: the cosine loss, which maximizes the cosine similarity between the output of the neural network and one-hot vectors indicating the true class.") Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of SHAZEER, MEIDAR, and BARZ before them, to include BARZ’s one-hot vectors indicating the true class in SHAZEER and MEIDAR’s mixture of experts (MoE) method. One would have been motivated to make such a combination because the cosine loss function provides substantially better performance than cross-entropy on datasets with only a handful of samples per class (BARZ [page 1371, Abstract]). Claims 6-9 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over SHAZEER in view of MEIDAR, as applied above to claim 1, and further in view of GALLARDO ("Self-Supervised Training Enhances Online Continual Learning"), hereafter GALLARDO. Regarding Claim 6: SHAZEER in view of MEIDAR teaches the elements of claim 1 as outlined above. SHAZEER in view of MEIDAR is not relied upon for teaching, but GALLARDO teaches: wherein the encoder is pre-trained using a second dataset different than a dataset that the training data item belongs to. (GALLARDO [page 4, section 4. Algorithms] teaches: "For all experiments, we use ResNet-18 as the CNN." GALLARDO [page 7, section 6.2.1 Domain Transfer with Deep SLDA] teaches: "A natural next step is to study how well self-supervised pretraining of features on one dataset generalize to continual learning on another dataset, e.g., pre-train on ImageNet and continually learn on Places-365." GALLARDO [page 5, section 5. Experimental Setup] teaches: "In addition to doing continual learning on ImageNet itself, we also study how well the pre-trained features learned on ImageNet from various sizes of pre-train sets generalize to another dataset from an entirely different domain [...]" GALLARDO [page 4, section 4.2. Online Continual Learning Models] teaches: "A new input is classified by assigning it the label of the closest Gaussian in feature space." Examiner's note: A CNN is a type of encoder built using ResNet. A CNN is an encoder, and GALLARDO [page 7, section 6.2.1 Domain Transfer with Deep SLDA] explicitly teaches pre-training the CNN (i.e., encoder) using different datasets.) Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of SHAZEER, MEIDAR, and GALLARDO before them, to include GALLARDO’s continual learning on different datasets in SHAZEER and MEIDAR’s mixture of experts (MoE) method. One would have been motivated to make such a combination because self-supervised features are versatile in generalizing to datasets beyond what they were trained on, which is especially useful when generalizing to datasets that have very few classes and/or samples, as pre-training directly on a subset of the dataset would be difficult (GALLARDO [page 6, section 6.1. Offline Linear Evaluation Results]). Regarding Claim 7: SHAZEER in view of MEIDAR and GALLARDO teaches the elements of claim 6 as outlined above. GALLARDO further teaches: wherein the encoder is pre-trained using a self-supervised learning technique. (GALLARDO [page 7, section 6.2.1 Domain Transfer with Deep SLDA] teaches: "A natural next step is to study how well self-supervised pretraining of features on one dataset generalize to continual learning on another dataset, e.g., pre-train on ImageNet and continually learn on Places-365." GALLARDO [page 8, section 8. Conclusion] teaches: "In this paper, we demonstrated the discriminative power of using self-supervised CNN features produced by the MoCo-V2 and SwAV models to generalize to unseen classes/datasets in online continual learning settings.") Regarding Claim 8: SHAZEER in view of MEIDAR and GALLARDO teaches the elements of claim 7 as outlined above. GALLARDO further teaches: wherein the self-supervised learning technique comprises training based upon transformed views of training data items. (GALLARDO [page 4, section 4.1. Pre-Training Approaches] teaches: "MoCo-V2 [16] makes two additional improvements over MoCo: it uses an additional projection layer and blur augmentation. […] SwAV: The self-supervised SwAV [9] method learns to assign clusters to different augmentations or “views” of the same image.") Regarding Claim 9: SHAZEER in view of MEIDAR teaches the elements of claim 1 as outlined above. SHAZEER in view of MEIDAR is not relied upon for teaching, but GALLARDO teaches: wherein parameters of the encoder are held fixed. (GALLARDO [page 4, section 4.2. Online Continual Learning Models] teaches: "SLDA [29]: Deep Streaming Linear Discriminant Analysis (SLDA) keeps the pre-trained CNN features fixed. It solely learns the output layer of the network.") Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of SHAZEER, MEIDAR, and GALLARDO before them, to include GALLARDO’s encoder fixed parameters in SHAZEER and MEIDAR’s mixture of experts (MoE) method. One would have been motivated to make such a combination because self-supervised features are versatile in generalizing to datasets beyond what they were trained on, which is especially useful when generalizing to datasets that have very few classes and/or samples, as pre-training directly on a subset of the dataset would be difficult (GALLARDO [page 6, section 6.1. Offline Linear Evaluation Results]). Regarding Claim 11: SHAZEER in view of MEIDAR teaches the elements of claim 1 as outlined above. SHAZEER in view of MEIDAR is not relied upon for teaching, but GALLARDO teaches: wherein the encoder is based upon a ResNet architecture. (GALLARDO [page 4, section 4. Algorithms] teaches: "For all experiments, we use ResNet-18 as the CNN.") Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of SHAZEER, MEIDAR, and GALLARDO before them, to include GALLARDO’s ResNet as the CNN in SHAZEER’s mixture of experts (MoE) method. One would have been motivated to make such a combination because ResNet-18 has been adopted as the universal standard CNN architecture by the community (GALLARDO [page 4, section 4. Algorithms]). Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over SHAZEER in view of MEIDAR, as applied to claim 1 above, and further in view of KINGMA ("Auto-Encoding Variational Bayes"), hereafter KINGMA. Regarding Claim 10: SHAZEER in view of MEIDAR teaches the elements of claim 1 as outlined above. SHAZEER in view of MEIDAR is not relied upon for teaching, but KINGMA teaches: wherein the encoder is based upon a variational autoencoder. (KINGMA [page 4, section 2.3 Our estimator of the lower bound] teaches: "A connection with auto-encoders becomes clear when looking at the objective function given at eq. (6). The variational approximation q ϕ z | x i (the encoder) maps a datapoint x i to a distribution over latent variables z from which the datapoint could have been generated.") Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of SHAZEER, MEIDAR, and KINGMA before them, to include KINGMA’s auto-encoder in SHAZEER and MEIDAR’s mixture of experts (MoE) method. One would have been motivated to make such a combination because efficient approximate posterior inference of the latent variable z given an observed value x for a choice of parameters θ is useful for coding or data representation tasks (KINGMA [page 2, section 2.1 Problem scenario]). Claims 15-16 are rejected under 35 U.S.C. 103 as being unpatentable over SHAZEER in view of MEIDAR, as applied to claim 1 above, and further in view of CHOQUE (US 20210012182 A1), hereafter CHOQUE. Regarding Claim 15: SHAZEER in view of MEIDAR teaches the elements of claim 1 as outlined above. SHAZEER in view of MEIDAR is not relied upon for teaching, but CHOQUE teaches: wherein the respective key embeddings are generated by sampling a probability distribution based upon the latent embedding space represented by the encoder. (CHOQUE [0047] teaches: "At step 614, a density model may be calculated. The density model may be a probabilistic distribution of the positive neighborhood within the memory module based on the encoded data and/or class label. In a variety of embodiments, the density model may include a Gaussian mixture model calculated as described herein." CHOQUE [0048] teaches: "At step 616, embedded data may be calculated. The embedded data may be a synthetic data point that statistically represents the encoded data and/or class label within the density model generated based on the positive neighborhood. In a number of embodiments, the conditional distribution of embedded data k′ (e.g. a new key) given a particular Gaussian distribution may be defined as: […]" CHOQUE [0049] teaches: "Sampling the embedded data k′ may include first sampling an index i from a set of Gaussian mixture components under the distribution and generating a random variable from: […]") Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of SHAZEER, MEIDAR, and CHOQUE before them, to include CHOQUE’s Gaussian mixture model in SHAZEER and MEIDAR’s mixture of experts (MoE) method. One would have been motivated to make such a combination in order to improve the quality, efficiency, and speed of machine learning systems by offering improved model training and performance through the management of memories in memory augmented neural networks (CHOQUE [0004]). Regarding Claim 16: SHAZEER in view of MEIDAR and CHOQUE teaches the elements of claim 15 as outlined above. CHOQUE further teaches: wherein the probability distribution is determined based upon a sample of encoded data. (CHOQUE [0047] teaches: "The density model may be a probabilistic distribution of the positive neighborhood within the memory module based on the encoded data and/or class label. In a variety of embodiments, the density model may include a Gaussian mixture model calculated as described herein.") Claim 17 is rejected under 35 U.S.C. 103 as being unpatentable over SHAZEER in view of MEIDAR and CHOQUE, as applied above to claim 16, and further in view of GALLARDO. Regarding Claim 17: SHAZEER in view of MEIDAR and CHOQUE teaches the elements of claim 16 as outlined above. CHOQUE further teaches: wherein the sample of encoded data comprises encoded data generated by processing data items using the encoder […] (CHOQUE [0031] teaches: "At step 412, input data may be encoded. The input data may be encoded for processing by the memory augmented neural network. The encoded input data may include a vector representation of the input data.") SHAZEER in view of MEIDAR and CHOQUE is not relied upon for teaching, but GALLARDO teaches: […] and wherein the data items are drawn from a second dataset different than a dataset that the training data item belongs to. (GALLARDO [page 4, section 4. Algorithms] teaches: "For all experiments, we use ResNet-18 as the CNN." GALLARDO [page 7, section 6.2.1 Domain Transfer with Deep SLDA] teaches: "A natural next step is to study how well self-supervised pretraining of features on one dataset generalize to continual learning on another dataset, e.g., pre-train on ImageNet and continually learn on Places-365." GALLARDO [page 5, section 5. Experimental Setup] teaches: "In addition to doing continual learning on ImageNet itself, we also study how well the pre-trained features learned on ImageNet from various sizes of pre-train sets generalize to another dataset from an entirely different domain [...]" GALLARDO [page 4, section 4.2. Online Continual Learning Models] teaches: "A new input is classified by assigning it the label of the closest Gaussian in feature space." Examiner's note: A CNN is a type of encoder built using ResNet. A CNN is an encoder, and GALLARDO [page 7, section 6.2.1 Domain Transfer with Deep SLDA] explicitly teaches pre-training the CNN (i.e., encoder) using different datasets.) Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of SHAZEER, MEIDAR, CHOQUE, and GALLARDO before them, to include GALLARDO’s continual learning on different datasets in SHAZEER, MEIDAR, and CHOQUE’s mixture of experts (MoE) method. One would have been motivated to make such a combination because using features pre-trained on one dataset for continual learning on another dataset could prove especially useful when performing continual learning on small datasets that require rich and generalizable feature representations (GALLARDO [page 7, section 6.2.1 Domain Transfer with Deep SLDA]). Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Alvaro S Laham Bauzo whose telephone number is 571-272-5650. The examiner can normally be reached Mon-Fri 7:30 AM - 11:00 AM | 1:00 PM - 5:30 PM ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Usmaan Saeed, can be reached on 571-272-4046. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /A.S.L./Examiner, Art Unit 2146 /USMAAN SAEED/Supervisory Patent Examiner, Art Unit 2146
Read full office action

Prosecution Timeline

Aug 23, 2023
Application Filed
Apr 06, 2026
Non-Final Rejection mailed — §103
Jul 01, 2026
Examiner Interview Summary
Jul 01, 2026
Applicant Interview (Telephonic)
Jul 02, 2026
Response Filed
Sep 15, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12725047
SYSTEM AND METHOD FOR DEFECT CLASSIFICATION AND LOCALIZATION WITH SELF-SUPERVISED PRETRAINING
3y 6m to grant Granted Sep 01, 2026
Patent 12632705
ADVERSARIAL 3D DEFORMATIONS LEARNING
4y 4m to grant Granted May 19, 2026
Patent 12475388
MACHINE LEARNING MODEL SEARCH METHOD, RELATED APPARATUS, AND DEVICE
3y 4m to grant Granted Nov 18, 2025
Study what changed to get past this examiner. Based on 3 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
50%
Grant Probability
99%
With Interview (+100.0%)
3y 9m (~8m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 8 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month