Prosecution Insights
Last updated: October 02, 2026
Application No. 17/584,505

MACHINE LEARNING DEVELOPMENT USING SUFFICIENTLY-LABELED DATA

Final Rejection §103
Filed
Jan 26, 2022
Priority
Feb 16, 2021 — provisional 63/200,120
Examiner
DIEP, DUY T
Art Unit
2123
Tech Center
2100 — Computer Architecture & Software
Assignee
University of Florida Research Foundation Inc.
OA Round
4 (Final)
37%
Grant Probability
At Risk
5-6
OA Rounds
0m
Est. Remaining
61%
With Interview

Examiner Intelligence

Grants only 37% of cases
37%
Career Allowance Rate
13 granted / 35 resolved
-17.9% vs TC avg
Strong +24% interview lift
Without
With
+23.7%
Interview Lift
resolved cases with interview
Typical timeline
4y 4m
Avg Prosecution
18 currently pending
Career history
65
Total Applications
across all art units

Statute-Specific Performance

§101
29.0%
-11.0% vs TC avg
§103
60.5%
+20.5% vs TC avg
§102
2.8%
-37.2% vs TC avg
§112
7.7%
-32.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 35 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment The amendments and arguments filed 06/09/2026 have been entered. Claims 1-7, 9-16, 18-20 remain pending in the application. Applicant’s amendments and arguments, with respect to claim rejections of claims 1-7, 9-16, 18-20 under 35 U.S.C 103 filed 03/24/2026 have been considered and are not persuasive. Therefore, the previous rejections as set forth in the previous office action will be maintained. Applicant argues that the cited combination fails to teach or suggest the amended limitations of claim 1. In particular, Applicant contends that Tan’s freeze-out technique, which freezes selected hidden units during training, does not teach training an output module in a second training stage using a portion of fully-labeled examples while freezing the hidden representation generated from sufficiently-labeled data in a first training stage. Applicant further argues that Appalaraju’s disclosure of reducing the number of training image pairs merely concerns reducing the amount of training data generally and does not teach or suggest determining the portion of fully-labeled examples based on the quantity of sufficiently-labeled data used to train the hidden module. Applicant additionally contends that none of the cited references teaches or suggests that an increase in the quantity of sufficiently-labeled data provides a corresponding decrease in the quantity of fully-labeled examples required to achieve a target performance of the machine learning model. Accordingly, Applicant asserts that the cited references, individually or in combination, fail to teach or suggest the newly amended limitations The examiner respectfully disagrees. Applicants’ arguments are not persuasive with respect to the limitation of training the output module while freezing the hidden representation. Applicants focus on Tan’s exemplary disclosure of randomly freezing a percentage of hidden units; however, the rejection is based on the combined teachings of the references rather than Tan individually. Appalaraju teaches learning a hidden representation from sufficiently-labeled positive/negative image pairs, Hagen teaches subsequent training of a classifier corresponding to the claimed output module, and Hagen further teaches freezing previously trained feature-learning layers while leaving the classification layer available for training. Tan additionally teaches selecting and freezing neural-network units such that the associated weights are not updated during a training run. Thus, the combined teachings of Appalaraju, Hagen, and Tan teach or at least suggest maintaining the previously learned hidden representation in a frozen state while subsequently training the classification/output module. Moreover, the amended claim does not recite a particular freeze-out mechanism, percentage of units to be frozen, or random or non-random manner of performing the freezing. Accordingly, Applicant’s distinction based on Tan’s exemplary freeze-out implementation does not distinguish the claimed limitation. With respect to the newly amended data-quantity limitations, Appalaraju teaches training the hidden representation using sufficiently-labeled similarity-based image pairs, while Hagen teaches subsequent training using a reduced portion of training data and teaches that reduced-data training may provide improved classification performance. Thus, the previously applied combination already teaches the underlying concepts of first learning a representation from sufficiently-labeled data and subsequently training a classifier using a reduced amount of labeled data. The Examiner acknowledges only that the previously applied combination does not expressly teach or suggest the specific newly claimed cross-stage relationship wherein the portion of fully-labeled examples is determined based on the quantity of sufficiently-labeled data used to train the hidden module such that an increase in the quantity of sufficiently-labeled data provides a decrease in the quantity of fully-labeled examples required to achieve the target performance. This specific relationship was introduced by Applicant in the present amendment. Accordingly, upon further consideration, the amendment necessitates a new ground of rejection (as set forth below). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-5, 7, 9-14, 16, 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Appalaraju et.al (US 10467526 B1), in view of Hagen et.al (US 20240070554 A1), further in view of Tan et.al (US 20210232909 A1), further in view of Singh et.al (US 20210124993 A1) Regarding claim 1, Appalaraju teaches or at least suggest the limitation “generating, by one or more processors, sufficiently-labeled data from comprising a plurality of example-pairs by generating, for an example-pair of the plurality of example- pairs, a sufficient label based on a comparison of a first set of original labels that corresponds to a first example of the example-pair and a second set of original labels that corresponds to a second example of the example-pair, wherein the sufficient label indicates whether the first set of original labels and the second set of original labels correspond to one or more same or a plurality of different classes” (“Column 3 lines 24-33 “Individual ones of the training examples may comprise a pair of images I1 and I2 and a similarity-based label (e.g., the logical equivalent of a 0/1 indicating similarity/dissimilarity respectively) corresponding to the pair. A given pair of images (I1, I2) may be labeled either as a positive pair or a negative pair in some embodiments, with the two images of a positive pair being designated as similar to one another based on one or more criteria, and the two images of a negative pair being designated as dissimilar to one another”, Column 5 lines 8-13 “a given pair of images may be labeled, for the purposes of model training, as a positive or similar pair based on a variety of criteria. Visual similarity, in which for example the shapes, colors, patterns and the like of the objects are taken into account, may be used to classify images as similar in various embodiments. In addition, for the purposes of some applications, it may be useful to classify images as similar based on semantic criteria ... For example, with respect to logos or symbols representing an organization, some symbols may be graphics-oriented while others may be more text-oriented. For the purposes of identifying images that may both represent a given organization, an image pair consisting of one graphical logo and one image of a company's name may be considered a positive image pair in some embodiments” Appalaraju discloses an Artificial Intelligence system for image similarity analysis using optimized image pair and neural network. Within the disclosure, Appalaraju discloses using a pair of images, wherein each image may have their own classification of shapes, colors, patterns and the like of the objects. In other words, each image within the image pair may be provided as pre-labeled images based on their own classification and each image within the image pair is analogous to the first and second input example of the sufficiently-labeled example-pair. Appalaraju further discloses labeling the pair of images as a positive pair or a negative pair based on their similarity, wherein the similarity determination is based on the classified/pre-labeled images, which is analogous to the process of generating a sufficient label to indicate a same or different class between two input examples based on relative information of original label to obtain a sufficiently-labeled example-pair, wherein the positive/negative label corresponds to the sufficient label the indicate a same or different class, and the positive/negative pair resulted from the positive/negative label corresponds to the sufficiently-labeled example-pair within the claim.) Appalaraju teaches or at least suggest the limitation “training, by the one or more processors and in a first training stage, a hidden module of a machine learning model by generating a hidden representation based on the sufficiently-labeled data in accordance with a minimization of a loss function that corresponds to training the machine learning model with the sufficiently-labeled data relative to a plurality of fully-labeled examples, wherein the plurality of fully-labeled examples comprises a plurality of original labels that correspond to one or more actual classes” (Column 6 lines 17-22 “As shown, system 100 includes various components and artifacts of an image analytics service (IAS) 102, which may be logically subdivided into a training subsystem 104 (at which machine learning models for image processing may be trained”, Column 7 lines 38-65 “semantic information may be used to classify some pairs of images … as similar based on the meaning or semantics associated with the images in some embodiments, even if the images of the pair do not necessarily look like one another ... Some image sources may provide pre-labeled images from which positive and negative pairs can be generated”, Column 8 lines 56-62 “The training of a neural network-based model of the kind discussed above may comprise a plurality of iterations or epochs in some embodiments. In a given training iteration, a plurality of candidate training image pairs may be initially be identified at the image pair selection subsystem, including for example one or more positive pairs and one or more negative pairs”, Column 10 lines 47-50 “A loss function 270 may be calculated based on the output vectors and a label y indicating whether the input image pair was a positive (similar) pair or a negative (dissimilar) pair in some embodiment”, and Column 11 lines 28-39 “... the loss becomes the distance between the embeddings of two similar images. As such, the model learns to reduce the distance between similar images, which is a desirable result ... If the two images are very dissimilar, the maximum function returns zero, so no minimization may be required. However, depending on the selected value of the m hyperparameter, if the images are not very dissimilar, a non-zero error value may be computed as the loss” Appalaraju discloses training the machine learning model using similarity-based positive/negative image pairs, corresponding to the sufficiently-labeled data. During training, Appalaraju calculates a loss based on the output vectors or embeddings of the image pair and the corresponding similarity label and trains the model to minimize the loss. Accordingly, the resulting embedding corresponds to the claimed hidden representation generated based on the sufficiently-labeled data in accordance with minimization of the loss function. Appalaraju further discloses that the positive/negative image pairs may be generated from pre-labeled/classified images. Thus, the pre-labeled images correspond to the claimed fully-labeled examples, and their respective classifications correspond to the original labels associated with one or more actual classes.) Appalaraju teaches the limitation “providing, by the one or more processors, the machine learning model for use in one or more prediction tasks” (Column 6, lines 47-62“The run-time subsystem 105 may comprise, for example, one or more run-time coordinators 135, image pre-processors 128, as well as resources 155 that may be used to execute trained versions 124 of the models produced at the training subsystem 104 in the depicted embodiment ... Class information stored regarding the previously-processed images may be used to select a subset of the images with which the similarity analysis for a given source image is to be performed at run-time in some embodiments. In at least one embodiment, a given run-time request submitted by a client may include two images for which a similarity score is to be generated using the trained model 124” Appalaraju discloses the implementation of the machine learning model at the run-time subsystem, wherein the trained machine learning model produced by the training subsystem is provided to and executed by a run-time subsystem to perform similarity analysis on input images. Such output constitutes a prediction regarding the similarity or dissimilarity between the images. The run-time subsystem utilizes this predicted similarity to perform tasks such as identifying related images. Accordingly, Appalaraju teaches the claimed process of machine learning model for use in one or more prediction tasks, wherein generating an output of a similarity score corresponds performing a prediction.) Appalaraju does not teach a part of the limitation “training, by the one or more processors and in a second training stage, an output module of the machine learning model based on a portion of the plurality of fully-labeled examples ...” However, Hagen teaches or at least suggest this part of the limitation (paragraph 15 “In some embodiments, the training module 106 may be configured to train a machine learning tool in two stages. A first training stage may be conducted based on a full set of training data, and a second stage may be conducted on a reduced set of training data that is a subset of the full training data set, after identifying and eliminating certain data points from the full training data set”, and paragraph 27 “The method 200 may further include a step 210 that includes continuing to train the machine learning algorithm with the reduced training data set to create a trained prediction model or classifier”. Hagen discloses training a machine-learning model in two stages, including a second training stage using a reduced training data set that is a subset of the full training data set, and continuing training using the reduced training data set to generate a trained prediction model or classifier. Accordingly, Hagen’s second-stage classifier corresponds to the claimed output module trained based on a portion of the plurality of fully-labeled examples. When combined with Appalaraju, the learned output-vector/embedding representations generated from Appalaraju’s similarity-based positive/negative image pairs may be used as input features to Hagen’s second-stage classifier.) Before the effective filing date, it would have been obvious to a person ordinary skilled in the art to combine the teaching of an Artificial Intelligence system for image similarity analysis using optimized image pair and neural network by Appalaraju, with the teaching of training a classifier in a second stage using reduced training data by Hagen. The motivation to do so is recited in Hagen’s disclosure (paragraph 54 “Further experiments illustrated that the use of second-stage classifiers (e.g., a Siamese network), instead of only a single advanced neural network, can improve classification accuracy”, and paragraph 27 “The method 200 may further include a step 210 that includes continuing to train the machine learning algorithm with the reduced training data set to create a trained prediction model or classifier. Step 210 may include training the first state of the model with the reduced training data set generated at step 208 to generate a second model state. As a result of training on an improved, reduced training data set, the second model state may be a more accurate classifier than the first model state.” Hagen discloses the benefit of training a classifier using an improved reduced training data set such that the model may be more accurate. Given that Appalaraju teaches learning representations from similarity-based positive/negative image pairs, one of ordinary skill in the art would have recognized that incorporating Hagen’s second-stage classifier training would permit the learned representations of Appalaraju to be used as input features for a classifier trained in a second stage using a reduced portion of fully-labeled training examples, thereby improving classification accuracy while reducing the amount of fully-labeled data used for second-stage training. Thus, the combined teaching results in a training subsystem by Appalaraju that learns from positive/negative image pair (sufficient labeled data), which corresponds to the claimed hidden module, to obtain output vector representation of image similarity. These learned representations may be used as input features to the classifier disclosed by Hagen, which may be further trained in a second stage using a reduced training dataset, wherein the classifier by Hagen corresponds to the claimed output module that generates outputs based on the learned hidden representation as claimed.) Appalaraju does not teach a part of the limitation “training ... an output module of the machine learning model … while freezing the hidden representation” However, Hagen in view of Tan teachers or at least suggest this limitation (Tan at paragraph 6 “The assessment component identifies units of a neural network. The selection component selects a subset of units of the neural network. The freeze-out component freezes the selected subset of units of the neural network” and paragraph 30 “Freeze-out provides an improved regularization technique as it eliminates the need to update the weights of output connections... the freeze-out technique involves randomly freezing a certain percentage of hidden units... The output of frozen units is not included for a training run but the weights of output connections from the frozen units to units of the following layer or layers are not changed for the training run. Thus, there is no need to update the weights of output connections from the frozen units to units of the following layer or layers during each training run.” Tan discloses systems and techniques that facilitate freeze-out as a regularizer in training neural networks. Within the disclosure, Tan discloses freeze-out of selected neural-network units such that the weights of output connections associated with the frozen units are not updated during a training run. Tan further discloses applying the freeze-out technique to hidden units and maintaining the neural-network architecture while the selected units remain frozen. Accordingly, Tan teaches or at least suggests maintaining a previously learned hidden or feature portion of a neural network in a frozen state during subsequent training. When incorporated into the Appalaraju/Hagen combination, Tan’s freeze-out technique corresponds to maintaining the learned hidden representation generated from Appalaraju’s sufficiently-labeled data in a frozen state while the output/classification module is subsequently trained. Hagen further support the mapping by Tan by teaching fine-tuning a pre-trained neural network via freezing at paragraph 32 “In some embodiments, the fine-tuning step 306 may include freezing all copied layers of the neural network from epoch to epoch except the classification layer. In other embodiments, the fine-tuning step 306 may include freezing initial layers of the neural network that train lower level features from epoch to epoch and fine-tuning subsequent layers”. Hagen discloses freezing all copied layers except the classification layer, or alternatively freezing initial layers that learn lower-level features while fine-tuning subsequent layers. Thus, Hagen teaches an arrangement in which previously learned feature-generating layers remain frozen while a classification portion remains trainable. In the Appalaraju/Hagen combination, the frozen feature-learning portion corresponds to maintaining Appalaraju’s previously learned hidden representation, while Hagen’s trainable classification layer corresponds to the claimed output module. Therefore, Hagen in view of Tan teaches or at least suggests training the output module while freezing the hidden representation, as claimed.) Before the effective filing date, it would have been obvious to a person ordinary skilled in the art to combine the teaching of an Artificial Intelligence system for image similarity analysis using optimized image pair and neural network by Appalaraju, and the teaching of training by using a reduced training data by Hagen, with the teaching of systems and techniques that facilitate freeze-out as a regularizer in training neural networks by Tan. The motivation to do so is referred to in Tan’s disclosure (paragraph 30 “Freeze-out provides an improved regularization technique as it eliminates the need to update the weights of output connections... Additionally, the reduction in steps and elimination of the need to update weights of output connections can reduce amount of time and effort required in training a neural network, thus optimizing training. Furthermore, the reduction of steps and elimination of need to update weights of output connections can mitigate reduction of errors as well as improve accuracy prediction by the neural network as evidenced by results of experiments shown in this specification below.” Tan discloses that freeze-out eliminates unnecessary updating of weights associated with frozen units, thereby reducing the amount of time and effort required for training a neural network and improving training efficiency and prediction accuracy. Hagen also teaches freezing previously trained feature-learning layers while leaving the classification layer trainable. Since Appalaraju teaches learning a hidden representation from similarity-based image pairs and Hagen teaches subsequently training a classifier using reduced training data, a person of ordinary skill in the art would have been motivated to incorporate Tan’s freeze-out technique into the Appalaraju/Hagen combination to maintain the previously learned hidden representation in a frozen state while training the classification/output module. Such a modification would avoid unnecessary updating of the learned feature portion while obtaining the training-efficiency and accuracy benefits taught by Tan, with predictable results.) Appalaraju in view of Hagen teaches or at least suggest “(i) the portion of the plurality of fully-labeled examples is determined based on a quantity of the sufficiently-labeled data used in training the hidden module” (Appalaraju discloses at Column 6 lines 1-11 “In some embodiments, new images may be synthesized or created (e.g., using additional machine learning models) to help obtain a sufficient number of positive image pairs to train the model … In some embodiments, to identify an image to be included in a negative pair with a given source image, a random selection algorithm may be used” Appalaraju discloses that the positive/negative image pairs used to train the model are generated from underlying pre-labeled/classified images, and further teaches obtaining or generating enough underlying images to provide a sufficient number of positive image pairs for training. Thus, Appalaraju teaches or at least suggests a relationship between the quantity of sufficiently-labeled image-pair data used for training and the amount of underlying pre-labeled image data needed to provide that training data. As discussed above, Hagen teaches selecting a reduced portion of the labeled training data for second-stage classifier training. Accordingly, when Hagen’s reduced portion is applied to Appalaraju’s pre-labeled images, a person of ordinary skill in the art would have understood that the amount of the reduced fully-labeled portion may be determined based on the quantity of Appalaraju’s positive/negative image-pair data used to train the hidden representation, thereby teaching or at least suggesting the recited limitation.) Appalaraju/Hagen/Tan does not teach the limitation “(ii) an increase in the quantity of the sufficiently-labeled data provides a decrease in the quantity of the plurality of fully- labeled examples in achieving a target performance of the machine learning model” However, Singh teaches or at least suggest this limitation (paragraph 20 “In few-shot learning or classification, the digital image classification system can train a base neural network on a set of base classes with abundant examples in a fashion that facilitates the neural network to classify digital images into novel classes with few (or no) labeled instance”, paragraph 21 “In the first phase, the digital image classification system can train a base neural network (including a feature extractor and a first classifier) based on base classes to develop robust and general-purpose feature representations aimed to be useful for classifying digital images into novel classes. In the second phase, the digital image classification system can exploit the learning of the first phase in the form of a prior to perform classification over novel classes. For example, the digital image classification system can utilize a transfer learning approach”, paragraph 24 “The digital image classification system gains improvements in accuracy versus conventional systems in few-shot classification as N increases in N-way K-shot evaluation”, and paragraph 25 “many conventional systems require large numbers of supervisory examples to effectively train a neural network to classify digital images, especially for identifying novel classes from training on base classes. By utilizing manifold mixup together with self-supervised training, the digital image classification system reduces the number of labeled examples required for training a neural network to classify digital images. With these techniques, the digital image classification system further reduces the amount of training data as compared to semi-supervised systems that require additional unlabeled data on top of labeled example”. Singh discloses training a neural network using abundant examples to develop robust and general-purpose feature representations, and then exploits the learning obtained from that first training stage to classify novel classes using few labeled examples. Singh further explains that this approach reduces the number of labeled examples required while achieving improved classification accuracy. When applied to Appalaraju, Appalaraju’s positive and negative labeled image pairs correspond to the sufficiently labeled data used to train the hidden representation, while Appalaraju’s original pre-labeled images correspond to the fully labeled examples. Accordingly, a person of ordinary skill in the art would have understood that increasing the quantity of Appalaraju’s sufficiently labeled image pair data used to train the model would provide more learning and thereby permit a reduction in the quantity of fully labeled examples required for subsequent classification while still achieving the desired performance, as taught or suggested by Singh.) Before the effective filing date, it would have been obvious to a person ordinary skilled in the art to combine the teaching of an Artificial Intelligence system for image similarity analysis using optimized image pair and neural network by Appalaraju, the teaching of training by using a reduced training data by Hagen, and the teaching of systems and techniques that facilitate freeze-out as a regularizer in training neural networks by Tan with the teaching of … by Singh. The motivation to do so is referred to in Singh’s disclosure (paragraph 25 “By utilizing manifold mixup together with self-supervised training, the digital image classification system reduces the number of labeled examples required for training a neural network to classify digital images. With these techniques, the digital image classification system further reduces the amount of training data as compared to semi-supervised systems that require additional unlabeled data on top of labeled examples. Indeed, the digital image classification system does not require extra unlabeled digital images for training like many conventional semi-supervised systems.”, paragraph 26 “The digital image classification system, on the other hand, can flexibly scale for various deep learning models (e.g., neural networks) based on utilizing smaller amounts of labeled data to classify digital images into novel classes. For example, the digital image classification system can readily modify a neural network to adapt to different classes because the digital image classification system requires such smaller amounts of labeled digital images. In addition, the digital image classification system can flexibly adapt to classify digital images within different domains based”. Singh discloses that its training techniques reduce the number of labeled examples required to train a neural network and allow the system to use smaller amounts of labeled data when adapting the network to classify novel classes. Accordingly, a person of ordinary skill in the art would have been motivated to incorporate Singh’s teaching into the combined system of Appalaraju, Hagen, and Tan in order to reduce the amount of labeled training data required for subsequent classifier training while maintaining effective classification performance, thereby improving training efficiency and scalability.) Regarding claim 2 depends on claim 1, thus the rejection of claim 1 is incorporated. Appalaraju in view of Hagen teaches the limitation “The method of claim 1, wherein the plurality of fully-labeled examples is obtained from fully-labeled data used to generate the sufficiently-labeled data” (Appalaraju discloses at Column 7 lines 38-65 “Images that can be used to train the model(s) may be acquired from a variety of image sources 110 … semantic information may be used to classify some pairs of images … as similar based on the meaning or semantics associated with the images in some embodiments, even if the images of the pair do not necessarily look like one another ... Some image sources may provide pre-labeled images from which positive and negative pairs can be generated”, and Hagen discloses at paragraph 15 “In some embodiments, the training module 106 may be configured to train a machine learning tool in two stages. A first training stage may be conducted based on a full set of training data, and a second stage may be conducted on a reduced set of training data that is a subset of the full training data set, after identifying and eliminating certain data points from the full training data set”. Appalaraju discloses that some image sources may provide pre-labeled images from which positive and negative image pairs may be generated. Thus, Appalaraju’s collection of pre-labeled images corresponds to the claimed fully labeled data used to generate the sufficiently labeled data, while the positive and negative labeled image pairs generated from those images correspond to the sufficiently labeled data. As discussed with respect to claim 1, the pre-labeled images correspond to fully labeled examples. Hagen further teaches obtaining a reduced training data set that is a subset of a full training data set. Accordingly, when Hagen’s subset selection is applied to Appalaraju’s collection of pre-labeled images from images sources, the reduced subset corresponds to the claimed plurality of fully labeled examples obtained from the fully labeled data. Therefore, Appalaraju in view of Hagen teaches or at least suggests that the plurality of fully labeled examples is obtained from fully labeled data that is used to generate sufficiently-labeled data .) Regarding claim 3 depends on claim 1, thus the rejection of claim 1 is incorporated. Appalaraju teaches the limitation “The method of claim 1, wherein the machine learning model is configured to, for the one or more prediction tasks, identify an original label from the plurality of original labels for an unseen input provided to the machine learning model.” (Column 7 lines 38-65 “Images that can be used to train the model(s) may be acquired from a variety of image sources 110 ... Such semantic information may be used to classify some pairs of images (such as images of company logos and company names) as similar based on the meaning or semantics associated with the images in some embodiments, even if the images of the pair do not necessarily look like one another ... Some image sources may provide pre-labeled images from which positive and negative pairs can be generated” Appalaraju discloses obtaining images used to train the model, wherein the image comprises semantic information (pre-labeled images) that may be used to classify some pairs of images such as company names or logos. These images training data corresponds to the unseen input with identified original label for the one or more prediction tasks, as recited in the claim.) Regarding claim 5 depends on claim 1, thus the rejection of claim 1 is incorporated. Appalaraju teaches the limitation “obtaining fully-labeled data comprising a plurality of input examples each having one of the plurality of original labels” (Column 5 lines 8-13 “a given pair of images may be labeled, for the purposes of model training, as a positive or similar pair based on a variety of criteria. Visual similarity, in which for example the shapes, colors, patterns and the like of the objects are taken into account, may be used to classify images as similar in various embodiments.” Appalaraju discloses the training image pairs (positive/negative) may comprise pre-labeled images (e.g., shapes, colors, patterns) for the purposes of model training, which corresponds to fully-labeled data comprising a plurality of input examples each having one of the plurality of original labels) Appalaraju teaches the limitation “generating the plurality of example-pairs, each example-pair comprising a first input example selected from the fully-labeled data and a second input example selected from the fully-labeled data” (Page 16 column 5 “images to be used for training may be collected from a variety of image sources (such as web sites), at least some of which may also comprise semantic metadata”, Page 23 column 20 “metadata associated with the collected images may be used to classify images into categories with associated labels. In at least one embodiment”, and Page 17 column 8“In a given training iteration, a plurality of candidate training image pairs may be initially be identified at the image pair selection subsystem, ... A number of operations may be performed to optimize the selection of training image pairs from among the candidates in some embodiments” Appalaraju discloses a selection subsystem to select image pairs from among the candidates, wherein images in image pairs to be used for training may be collected from a variety of image sources, wherein images from image sources may comprise semantic metadata to classify images into categories with associated labels, which corresponds to a first and second input example as claimed.) Appalaraju teaches the limitation “generating one or more sufficient labels for the plurality of example-pairs based at least in part on summarizing a first original label of the plurality of original labels and a second original label of the plurality of original labels” (page 15 column 3 “individual ones of the training examples may comprise a pair of images I1 and I2 and a similarity-based label (e.g., the logical equivalent of a 0/1 indicating similarity/dissimilarity respectively) corresponding to the pair. A given pair of images (I1, I2) may be labeled either as a positive pair or a negative pair in some embodiments, with the two images of a positive pair being designated as similar to one another based on one or more criteria, and the two images of a negative pair being designated as dissimilar to one another”, and page 16 column 5 “it may be useful to classify images as similar based on semantic criteria ... images to be used for training may be collected from a variety of image sources (such as web sites), at least some of which may also comprise semantic metadata associated with the images ... In at least some embodiments, such semantic metadata may be analyzed to determine whether two images are to be designated as similar, even if they may not look alike from a purely visual perspective.”. Appalaraju discloses generating the similarity-based label for the image pair based on the criteria, wherein the criteria may be semantic criteria of semantic data associated with each image within the image pair, wherein the positive/negative label of each image pair corresponds to the generating of sufficient labels for the plurality of example-pairs based at least in part on summarizing a first original label of the plurality of original labels and a second original label of the plurality of original labels, as claimed.) Regarding claim 7 depends on claim 5, thus the rejection of claim 5 is incorporated. Appalaraju teaches the limitation “The method of claim 1, further comprising storing the sufficiently-labeled data in a storage medium as an encrypted representation of the first input example and the second input example” (Page 25 column 24 “Various embodiments may further include receiving, sending or storing instructions and/or data implemented in accordance with the foregoing description upon a computer-accessible medium.” Appalaraju discloses the embodiment may include storing data in a computer-accessible medium, wherein the data can be the similarity-based label representing the relationship between the image pair input data. A person ordinary skilled in the art would have been able to configure such label to be encrypted to be stored in the computer storage medium.) Regarding claim 9 depends on claim 1, thus the rejection of claim 1 is incorporated. Appalaraju teaches the limitation “The method of claim 1, wherein the machine learning model comprises a neural network configured as a classifier” (Page 24 column 22 “In at least some embodiments, a server that implements a portion or all of one or more of the technologies described herein, including the various components of the training and run-time subsystems of an image analytics service, including for example ... classifiers”, and Page 24 column 21 “In the depicted embodiment, the images available at the analytics service may be analyzed, e.g., with the help of models similar to the CNN models discussed earlier, and placed into various classes”. Appalaraju discloses the computer system of the method may comprise of classifier, wherein the classifier may be the classifier model as disclosed by Frandsen based on the teaching combination, and the images with semantic metadata may be placed into classes.) Regarding claim 10 depends on claim 1, thus the rejection of claim 1 is incorporated. Appalaraju teaches the limitation “The method of claim 1, wherein the hidden module is trained using one of a hinge loss function, a negative cosine similarity function, a contrastive function, or a mean squared error” (Page 18 column 10 “A loss function 270 may be calculated based on the output vectors and a label y indicating whether the input image pair was a positive (similar) pair or a negative (dissimilar) pair in some embodiments”, and Page 17 column 8 “A contrastive loss function may be used for the overall neural network model in at least some embodiments”. Appalaraju discloses calculating a loss function for the training of the machine learning model, wherein the loss function can be a contrastive loss function.) Regarding claim 11, Appalaraju teaches limitations “one or more processors”, and “at least one memory storing processors-executable instructions that, when executed by any of the one or more processors, cause the one or more processors to perform operations” (page 24 column 22 “In various embodiments, computing device 9000 may be a uniprocessor system including one processor 9010, or a multiprocessor system”, and “System memory 9020 may be configured to store instructions and data accessible by processor(s) 9010... In the illustrated embodiment, program instructions and data implementing one or more desired functions, such as those methods, techniques, and data described above, are shown stored within system memory 9020 as code 9025 and data 9026.” Appalaraju discloses the embodiment may be a computing device with one or more processors, and memory configured to store coded instructions to be executed by the one or more processor.) The applicant is further directed to the rejection of claim 1, because claim 11 recites similar limitation to claim 1, thus the claim is rejected under the same rationale. Regarding claim 12 depends on claim 11, thus the rejection of claim 11 is incorporated. The applicant is further directed to the rejection of claim 3, because claim 12 recites similar limitation to claim 3, thus the claim is rejected under the same rationale. Regarding claim 14 depends on claim 11, thus the rejection of claim 11 is incorporated. The applicant is further directed to the rejection of claim 5, because claim 14 recites similar limitation to claim 5, thus the claim is rejected under the same rationale. Regarding claim 16 depends on claim 14, thus the rejection of claim 14 is incorporated. The applicant is further directed to the rejection of claim 7, because claim 16 recites similar limitation to claim 7, thus the claim is rejected under the same rationale. Regarding claim 18 depends on claim 11, thus the rejection of claim 11 is incorporated. The applicant is further directed to the rejection of claim 9, because claim 18 recites similar limitation to claim 9, thus the claim is rejected under the same rationale. Regarding claim 19, which recites a machine, one of the four statutory categories of patentable subject matter. Appalaraju teaches the limitation: “One or more non-transitory computer storage media storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations” (page 25 column 23 “Generally speaking, a computer-accessible medium may include non-transitory storage media or memory media”, and page 24 column 22 “System memory 9020 may be configured to store instructions and data accessible by processor(s) 9010... In the illustrated embodiment, program instructions and data implementing one or more desired functions, such as those methods, techniques, and data described above, are shown stored within system memory 9020 as code 9025 and data 9026.” Appalaraju discloses a computer-accessible medium may include non-transitory storage media or memory media, wherein the storage media may comprise of memories which contain instructions to cause program operations executed by one or more processors.) The applicant is further directed to the rejection of claim 1, because claim 19 recites similar limitation to claim 1, thus the claim is rejected under the same rationale. Regarding claim 20 depends on claim 19 thus the rejection of claim 19 is incorporated. The applicant is further directed to the rejection of claim 5, because claim 20 recites similar limitation to claim 5, thus the claim is rejected under the same rationale. Claims 4, 13 are rejected under 35 U.S.C. 103 as being unpatentable over Appalaraju et.al (US 10467526 B1), in view of Hagen et.al (US 20240070554 A1), further in view of Tan et.al (US 20210232909 A1), further in view of Singh et.al (US 20210124993 A1), further in view of Frandsen et.al (NPL: Machine Learning for Disease Prediction). Regarding claim 4 depends on claim 1, thus the rejection of claim 1 is incorporated. Frandsen teaches the limitation “The method of claim 1, wherein the plurality of original labels comprises a first original label classifying an individual as contracting a disease and a second original label classifying the individual as not contracting the disease, and wherein the one or more prediction tasks includes identification of either the first original label or the second original label for an unseen individual to indicate a likelihood of the unseen individual having contracted the disease” (Chapter 2 Page 13 section 2.1 “The aim of classification is to assign labels to objects. More formally, assume there is a set X of objects and a finite set Y of labels”, Chapter 2 Page 15 section 2.1 “Example 2.5 (Disease Prediction). Suppose we have a soft classifier that takes values in the unit interval. This classifier has been created to predict if someone is at risk of developing 14 late stage CKD; the label 0 indicates no risk, and the label 1 indicates high risk”, and Page 24 section 2.4 “... Since the effectiveness of a classifier resides in its ability to make useful predictions even on unseen examples, it is important to investigate its generalization. We measure the generalization by gathering a new collection of labeled data, which we call the test set, and reporting possibly multiple numerical quantities that evaluate how well the classifier was able to predict the labels in the test set.” Frandsen discloses the label 1 indicates high risk of disease and label 0 indicates no risk of disease, which is analogous to the claimed fourth and fifth label within the claim respectively, wherein the classifier may obtain image data from Appalaraju with semantic metadata being these 0 and 1 labels to perform its training and after the training, the classifier may be able to perform prediction to assign either 0 or 1 label for respectively no risk of contracting disease and high risk of contracting disease for unseen input, wherein the unseen input can be data associated with unseen individuals.) Before the effective filing date, it would have been obvious to a person ordinary skilled in the art to combine the teaching of Appalaraju/Hagen/Tan/Singh with the teaching of machine learning for disease prediction by Frandsen. The motivation to do so is referred to in Frandsen’s disclosure (page 14 example 2.5 “This classifier has been created to predict if someone is at risk of developing late stage CKD; the label 0 indicates no risk, and the label 1 indicates high risk. Should an individual be flagged as high risk, the doctor may order potentially costly tests and treatments”, and page 24 section 2.4 “However, our ultimate goal is to have a classifier that performs well on all appropriate data, not just the particular examples in the training set. We use the term generalization to denote the performance of a classifier on new, unseen data. Since the effectiveness of a classifier resides in its ability to make useful predictions even on unseen examples, it is important to investigate its generalization.” Frandsen discloses the utilization of the trained machine learning model in predicting task such as predicting diseases within the medical field and the trained classifier model is evaluated on unseen labeled data to measure generalization performance and to produce prediction outputs. Frandsen provides examples such as in example 2.5, wherein the classifier machine learning model within the example is capable to provide label of 0 or 1, which is similar to the machine learning model by Appalaraju as disclosed above, wherein the model has been trained on similarity-based example pairs. Frandsen further discloses evaluating the performance of a trained classifier model on unseen labeled data to determine how well the model performed. This post training evaluation step is a standard component of supervised machine learning workflows. Once the model has been trained and stabilized as taught by Appalaraju/Hagen/Tan, one of ordinary skilled in the art would have found it obvious to apply the trained model to predict the similarity in unseen labeled data using the testing procedure in Frandsen in order to measure the model generalization and enable prediction. Therefore, the incorporation of the teaching by Frandsen into the teaching combination would have represented a predictable use of known machine learning techniques to perform their established functions.) Regarding claim 13 depends on claim 11, thus the rejection of claim 11 is incorporated. The applicant is further directed to the rejection of claim 4, because claim 13 recites similar limitation to claim 4, thus the claim is rejected under the same rationale. Claims 6, 15 are rejected under 35 U.S.C. 103 as being unpatentable over Appalaraju et.al (US 10467526 B1), in view of Hagen et.al (US 20240070554 A1), further in view of Tan et.al (US 20210232909 A1), further in view of Singh et.al (US 20210124993 A1), further in view of Frandsen et.al (NPL: Machine Learning for Disease Prediction), further in view of Esteva et.al (US 20220222484 A1). Regarding claim 6 depends on claim 5, thus the rejection of claim 5 is incorporated. Appalaraju/Hagen/Tan/Singh/Frandsen does not teach “The method of claim 5, wherein the one or more sufficient labels are generated using an annotation machine learning model”. However, Esteva teaches this limitation (paragraph 28 “The labeling interface module 220 generates information for displaying subsets of the segmented portions of the image data to an operator for annotation. In one embodiment, the labeling interface module 220 identifies a region within the image that includes one or more portions to be annotated”, and paragraph 29 “When a session begins, the portions to be annotated may be presented to the operator for labeling without recommendations. The labeling interface module 220 receives data indicating the labels assigned to portions by the operator. The classifier training module 230 trains a classifier using the assigned labels”. Esteva discloses an AI-enhanced data labeling tool with annotation. Within the disclosure, Esteva discloses the tool comprises of modules for annotation operation of images, such that the annotated image may then be labeled. A person ordinary skilled in the art would have been able to configure the labeling process to assign similarity-based label based on the annotation image based on the teaching combination as explained below.) Before the effective filing date, it would have been obvious to a person ordinary skilled in the art to combine the teaching of Appalaraju/Hagen/Tan/Singh/Frandsen with the teaching of the AI-enhanced data labeling tool with annotation by Esteva. The motivation to do so is referred to in Esteva’s disclosure (paragraph 20 “the AI-enhanced labeling tool may enable the operator to annotate more examples in a given time period, reducing the cost of data annotation. The tool may also increase accuracy of the labels assigned to image portions.”, paragraph 53 “FIG. 7 includes experimental data demonstrating efficiency improvements that may be realized using one embodiment of the AI-enhanced data labeling tool.”, and paragraph 56 “FIG. 8 is an example comparison plot of a dataset annotated with the AI-enhanced data labeling tool versus one annotated without... a model trained with 50, 75, and 100 training examples benefits from an 11%, 11%, and 5% boost in model validation accuracy, respectively. In benchmark machine learning competitions, top performing models typically win by fractions of a percent to single percentage points. Thus, this improvement is effectiveness is significant”. Esteva discloses the benefit of the AI-enhanced data labeling tool with annotation feature, which increases accuracy of the labels assigned to image portions. The tool provides labeled annotated image training data that help boosting model validation accuracy, which is demonstrated through experiment and example in fig. 7 and fig. 8. Given that Appalaraju discloses neural network model training based on similarity-based example pairs of images, thus the teaching combination by Appalaraju/Hagen/Tan/Frandsen may further incorporate the teaching of annotation image data for further improvement on obtaining training data for better training result and better machine learning model performance.) Regarding claim 15 depends on claim 14, thus the rejection of claim 14 is incorporated. The applicant is further directed to the rejection of claim 6, because claim 15 recites similar limitation to claim 6, thus the claim is rejected under the same rationale. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to DUY TU DIEP whose telephone number is (703)756-1738. The examiner can normally be reached M-F 8-4:30. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571) 270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DUY T DIEP/Examiner, Art Unit 2123 /ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123
Read full office action

Prosecution Timeline

Show 10 earlier events
Jan 15, 2026
Request for Continued Examination
Jan 26, 2026
Response after Non-Final Action
Mar 24, 2026
Non-Final Rejection mailed — §103
May 13, 2026
Interview Requested
May 28, 2026
Applicant Interview (Telephonic)
May 28, 2026
Examiner Interview Summary
Jun 09, 2026
Response Filed
Aug 27, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748949
A Central Node and Method Therein for Enabling an Aggregated Machine Learning Model from Local Machine Learnings Models in a Wireless Communications Newtork
4y 1m to grant Granted Sep 29, 2026
Patent 12743603
KNOWLEDGE GRAPH REASONING MODEL, SYSTEM, AND REASONING METHOD BASED ON BAYESIAN FEW-SHOT LEARNING
3y 11m to grant Granted Sep 22, 2026
Patent 12725004
NEURAL ARCHITECTURE SEARCH BASED OPTIMIZED DNN MODEL GENERATION FOR EXECUTION OF TASKS IN ELECTRONIC DEVICE
5y 5m to grant Granted Sep 01, 2026
Patent 12651158
NEURAL NETWORK TRAINING METHOD AND APPARATUS USING TREND
4y 1m to grant Granted Jun 09, 2026
Patent 12608642
MODEL PARAMETER LEARNING METHOD AND MOVEMENT MODE DETERMINATION METHOD
4y 7m to grant Granted Apr 21, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
37%
Grant Probability
61%
With Interview (+23.7%)
4y 4m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 35 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month