Prosecution Insights
Last updated: August 17, 2026
Application No. 17/477,370

TRAINING OBJECT DETECTION MODELS USING TRANSFER LEARNING

Non-Final OA §101§103
Filed
Sep 16, 2021
Examiner
PAULA, CESAR B
Art Unit
2145
Tech Center
2100 — Computer Architecture & Software
Assignee
NVIDIA Corporation
OA Round
3 (Non-Final)
34%
Grant Probability
At Risk
3-4
OA Rounds
0m
Est. Remaining
42%
With Interview

Examiner Intelligence

Grants only 34% of cases
34%
Career Allowance Rate
58 granted / 173 resolved
-21.5% vs TC avg
Moderate +8% lift
Without
With
+8.2%
Interview Lift
resolved cases with interview
Typical timeline
4y 6m
Avg Prosecution
8 currently pending
Career history
195
Total Applications
across all art units

Statute-Specific Performance

§101
12.4%
-27.6% vs TC avg
§103
50.7%
+10.7% vs TC avg
§102
18.7%
-21.3% vs TC avg
§112
14.2%
-25.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 173 resolved cases

Office Action

§101 §103
DETAILED ACTION This Final rejection is response to the amendment to the claims filed 5/27/2025. Claims 1-20 are pending. Claims 1, 9, and 16 are independent claims. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Drawings The objection to the drawings has been withdrawn as necessitated by the amendment. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claim 1-20 remain rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more as shown below. Step 1 According to the first part of the analysis, in the instant case, claims 1-8 are directed to a method, claims 9-15 are directed to a system, and claims 16-20 are directed to a non-transitory computer readable storage medium. Thus, each of the claims falls within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter). Regarding Claim 1 Step 2A Prong 1: identifying a first set of images comprising a plurality of objects of a plurality of classes; (This step for identifying and classifying objects involves observations and evaluations, which are understood to be a recitation of a mental process.) providing the first set of images as input to a first machine learning model trained to detect, for a given input image, a presence of one or more objects of at least one of the plurality of classes depicted in the given input image and to predict at least mask data comprising an indication of one or more pixels of the given input image depicting one or more of the detected objects associated with one or more of the detected objects; (These steps for detecting objects and predicting mask data involve mathematical calculations, algorithms, and relationships in machine learning, which are understood to be a recitation of a mental or mathematical process for indicating the location of the pixels within the image.) determining, from one or more first outputs of the first machine learning model, object data associated with each of the first set of images, wherein the object data for each respective image of the first set of images comprises mask data indicating the one or more pixels of each respective image depicting each object detected in the respective image; (This step for determining object data and mask data involves recognizing and identifying features within an image, which are understood to be a recitation of a mental or mathematical process for indicating the location of the pixels within the image.) and training a second machine learning model to detect objects of a target class in a second set of images, wherein the second machine learning model is trained using at least a subset of the first set of images, and a target output for the at least a subset of the first set of images, wherein the target output comprises the mask data associated with each object detected in the at least a subset of the first set of images and an indication of whether a class associated with each object detected in the at least a subset of the first set of images corresponds to the target class. (This step for generating a target output and detecting objects of a certain class involves recognizing and categorizing visual features, which is understood to be a recitation of a mental process.) Step 2A Prong 2: providing the first set of images as input to a first machine learning model trained (This step links the judicial exception to a particular technological environment, i.e., using a trained machine learning model, which is within a field of technology. See MPEP § 2106.05(h).) from one or more first outputs of the first machine learning model, (This step links the judicial exception to a particular technological environment, i.e., processing machine learning model outputs in the field of image detection. See MPEP § 2106.05(h).) training a second machine learning model (This step links the judicial exception to a particular technological environment, i.e., training a machine learning model. See MPEP § 2106.05(h).) wherein the second machine learning model is trained using at least a subset of the first set of images, (This step links the judicial exception to a particular technological environment, i.e., training a machine learning model using a dataset. See MPEP § 2106.05(h).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because, when considered individually and in combination, they do not add significantly more (also known as an inventive concept) to the exception. The claim recites machine learning models and their application in image processing, including processing model outputs and training machine learning models using a dataset, which link the judicial exception to a particular technological environment. Regarding Claim 2 Step 2A Prong 1: wherein the first machine learning model is further trained to predict, for each of the one or more detected objects, a particular class of the plurality of classes associated with a respective detected object. (This step for predicting an object’s class involves recognizing and categorizing visual features, which is understood to be a recitation of a mental process.) Step 2A Prong 2: wherein the first machine learning model is further trained (This step links the judicial exception to a particular technological environment, i.e., training a machine learning model for object classification. See MPEP § 2106.05(h).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because, when considered individually and in combination, they do not add significantly more (also known as an inventive concept) to the exception. The claim recites a machine learning model and its training for object classification, which links the judicial exception to a particular technological environment. Regarding Claim 3 Step 2A Prong 1: generating the target output, wherein generating the target output comprises: determining whether the particular class associated with the respective detected object corresponds to the target class. (This step for generating a target output and determining class correspondence involves recognizing, categorizing, and making classification decisions, which are understood to be a recitation of a mental process.) Step 2A Prong 2: The claim does not include additional elements, when considered separately and in combination, that integrate the judicial exception into a practical application. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because, when considered individually and in combination, they do not add significantly more (also known as an inventive concept) to the exception. The claim is directed to a mental process of generating a target output and determining class correspondence without any technological improvement or inventive concept. Regarding Claim 4 Step 2A Prong 1: identifying, using an indication of one or more bounding boxes associated with the image, ground truth data associated with the respective object depicted in the image. (This step for identifying ground truth data using bounding boxes involves recognizing and categorizing visual information, which is understood to be a recitation of a mental process.) Step 2A Prong 2: The claim does not include additional elements, when considered separately and in combination, that integrate the judicial exception into a practical application. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because, when considered individually and in combination, they do not add significantly more (also known as an inventive concept) to the exception. The claim is directed to a mental process of identifying ground truth data using bounding boxes without any technological improvement or inventive concept. Regarding Claim 5 Step 2A Prong 1: at least one bounding box of the one or more bounding boxes were provided by at least one of an accepted bounding box authority entity or a user of a platform. (This step for providing bounding box data involves human judgment and visual identification, which is understood to be a recitation of a mental process.) Step 2A Prong 2: The claim does not include additional elements, when considered separately and in combination, that integrate the judicial exception into a practical application. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because, when considered individually and in combination, they do not add significantly more (also known as an inventive concept) to the exception. The claim is directed to a mental process of providing bounding boxes without any technological improvement or inventive concept. Regarding Claim 6 Step 2A Prong 1: wherein the second machine learning model is a multi-head machine learning model, and wherein the method further comprises: upon training the second machine learning model using at least a subset of the first set of images and the target output, identifying one or more heads of the second machine learning model that correspond to predicting mask data for a given input image; and updating the second machine learning model to remove the one or more identified heads. (This step for identifying model heads involves recognizing and categorizing relationships within a machine learning model, which is understood to be a recitation of a mental process.) Step 2A Prong 2: wherein the second machine learning model is a multi-head machine learning model, (This step links the judicial exception to a particular technological environment, i.e., using a multi-head machine learning model. See MPEP § 2106.05(h).) and wherein the method further comprises: upon training the second machine learning model using at least a subset of the first set of images and the target output, (This step links the judicial exception to a particular technological environment, i.e., training a machine learning model. See MPEP § 2106.05(h).) updating the second machine learning model to remove the one or more identified heads. (This step links the judicial exception to a particular technological environment, i.e., updating the structure of a machine learning model. See MPEP § 2106.05(h).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because, when considered individually and in combination, they do not add significantly more (also known as an inventive concept) to the exception. The claim recites training and updating a multi-head machine learning model, which links the judicial exception to a particular technological environment. Regarding claim 7 Step 2A Prong 1: providing a third set of images as input to the second machine learning model; obtaining one or more second outputs of the second machine learning model; and determining, based on the one or more second outputs, additional object data associated with each of the third set of images, wherein the additional object data for each respective image of the second set of images comprises an indication of a region of the respective image that includes an object detected in the respective image and a class associated with the detected object. (This step for determining additional object data involves recognizing and categorizing visual information, which is understood to be a recitation of a mental process.) Step 2A Prong 2: providing a third set of images as input to the second machine learning model, (This step is directed to providing data, which is an insignificant extra-solution activity. See MPEP § 2106.05(g).) obtaining one or more second outputs of the second machine learning model; (This step is a routine operation of generating model outputs and is an insignificant extra-solution activity. See MPEP § 2106.05(g).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because, when considered individually and in combination, they do not add significantly more (also known as an inventive concept) to the exception. The claim recites providing input data, generating model outputs, and extracting object data, which are well-understood, routine, and conventional activities, as recognized by the court decisions listed in MPEP § 2106.05(d). Regarding Claim 8 Step 2A Prong 1: The claim does not recite any abstract steps. Step 2A Prong 2: transmitting the updated second machine learning model to at least one of an edge device or an endpoint device via a network. (This step is directed to transmitting data, which is an insignificant extra-solution activity under MPEP § 2106.05(g).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because, when considered individually and in combination, it does not add significantly more (also known as an inventive concept) to the exception. The claim recites transmitting data over a network, which is a well-understood, routine, and conventional activity, as recognized by the court decisions listed in MPEP § 2106.05(d). Regarding Claim 9 Step 2A Prong 1: a memory device; and a processing device coupled to the memory device, wherein the processing device is to perform operations comprising: generating training data for a machine learning model, wherein generating the training data comprises: generating a training input comprising an image depicting an object; and generating a target output for the training input, wherein the target output comprises a bounding box associated with the depicted object, mask data associated with the depicted object, and an indication of a class associated with the depicted object; providing the training data to train the machine learning model on (i) a set of training inputs comprising the generated training input and (ii) a set of target outputs comprising the generated target output; identifying one or more heads of the trained machine learning model that correspond to predicting mask data for a given input image; and updating the trained machine learning model to remove the one or more identified heads. (These steps for generating a target output and identifying model heads involve recognizing and categorizing visual information, which are understood to be recitations of a mental process.) Step 2A Prong 2: a memory device; and a processing device coupled to the memory device, wherein the processing device is to perform operations comprising: (This step describes generic computer components that do not add significantly more. See MPEP § 2106.05(f).) generating training data for a machine learning model, wherein generating the training data comprises: generating a training input comprising an image depicting an object; (This step links the judicial exception to a particular technological environment, i.e., generating training data for training a machine learning model. See MPEP § 2106.05(h).) providing the training data to train the machine learning model on (i) a set of training inputs comprising the generated training input and (ii) a set of target outputs comprising the generated target output; (This step links the judicial exception to a technological environment but does not integrate the exception into a practical application. See MPEP § 2106.05(h).) updating the trained machine learning model to remove the one or more identified heads. (This step links the judicial exception to a particular technological environment, i.e., optimizing a machine learning model. See MPEP § 2106.05(h).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because, when considered individually and in combination, they do not add significantly more (also known as an inventive concept) to the exception. The claim recites a memory device, a processing device, and machine learning models, which link the judicial exception to a particular technological environment. Regarding Claim 10 Step 2A Prong 1: providing a set of images as input to the updated trained machine learning model; obtaining one or more outputs of the updated trained machine learning model; and determining, from the one or more outputs, object data associated with each of the set of images, wherein the object data for each respective image of the second set of images comprises an indication of a region of the respective image that includes an object detected in the respective image and a class associated with the detected object. (This step for determining additional object data involves recognizing and categorizing visual information, which is understood to be a recitation of a mental process.) Step 2A Prong 2: providing a set of images as input to the updated trained machine learning model; (This step is directed to providing data, which is an insignificant extra-solution activity. See MPEP § 2106.05(g).) obtaining one or more outputs of the updated trained machine learning model; (This step is a routine operation of generating model outputs and is an insignificant extra-solution activity. See MPEP § 2106.05(g).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because, when considered individually and in combination, they do not add significantly more (also known as an inventive concept) to the exception. The claim recites additional limitations directed to providing input data, generating model outputs, and determining object data, which are insignificant extra-solution activities. These are well-understood, routine, and conventional activities in the field of technology, as recognized by the court decisions listed in MPEP § 2106.05(d). Regarding Claim 11 Step 2A Prong 1: The claim does not recite any abstract steps. Step 2A Prong 2: deploying the updated trained machine learning model using at least one of an edge device or an endpoint device. (The step links the claim to a technological environment, i.e., deploying a trained machine learning model to an edge or endpoint device. See MPEP § 2106.05(h)). Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because, when considered individually and in combination, they do not add significantly more (also known as an inventive concept) to the exception. The claim recites deploying a trained machine learning model, which links the judicial exception to a particular technological environment. Regarding Claim 12 Step 2A Prong 1: providing the image depicting the object as input to an additional machine learning model, wherein the additional machine learning model is trained to detect, for a given input image, a presence of one or more objects depicted in the given input image and to predict at least mask data associated with one or more of the detected objects; and determining, from one or more outputs of the additional machine learning model, object data associated with the image, wherein the object data for the image comprises mask data associated with the depicted object. (These steps for detecting objects, predicting mask data, and determining object data involve recognizing and categorizing visual information, which are understood to be recitations of a mental process.) Step 2A Prong 2: providing the image depicting the object as input to an additional machine learning model, wherein the additional machine learning model is trained (This step links the judicial exception to a particular technological environment, i.e., using and training a machine learning model for object detection. See MPEP § 2106.05(h).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because, when considered individually and in combination, they do not add significantly more (also known as an inventive concept) to the exception. The claim recites using and training a machine learning model for object detection, which links the judicial exception to a particular technological environment. Regarding Claim 13 Step 2A Prong 1: wherein the additional machine learning model is further trained to predict, for each of the one or more detected objects, a class associated with the respective detected object, and wherein object data for the image further comprises the indication of the class associated with the depicted object. (These steps for predicting an object's class and determining classification involve recognizing and categorizing visual information, which are understood to be recitations of a mental process.) Step 2A Prong 2: wherein the additional machine learning model is further trained (This step links the judicial exception to a particular technological environment, i.e., training a machine learning model for object classification. See MPEP § 2106.05(h).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because, when considered individually and in combination, they do not add significantly more (also known as an inventive concept) to the exception. The claim recites a machine learning model trained for object classification, which links the judicial exception to a particular technological environment. Regarding Claim 14 Step 2A Prong 1: obtaining ground truth data associated with the image, wherein the ground truth data comprises the bounding box associated with the depicted object. (This step for obtaining bounding box data involves recognizing and labeling visual information, which is understood to be a recitation of a mental process.) Step 2A Prong 2: obtaining ground truth data associated with the image, (This step is directed to data gathering, which is an insignificant extra-solution activity. See MPEP § 2106.05(g).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because, when considered individually and in combination, they do not add significantly more (also known as an inventive concept) to the exception. The claim recites insignificant extra-solution activity, such as obtaining ground truth data, which is a well-understood, routine, and conventional practice in the field of computer vision and machine learning, as recognized by the court decisions listed in MPEP § 2106.05(d). Regarding Claim 15 Step 2A Prong 1: wherein the ground truth data is obtained from a database comprising an indication of one or more bounding boxes associated with objects depicted in a set of images, wherein the image is included in the set of images, and wherein the one or more bounding boxes is provided by an accepted bounding box authority entity or a user of a platform. (These steps for identifying bounding boxes and specifying their source involve recognizing and categorizing visual information, which are understood to be recitations of a mental process.) Step 2A Prong 2: the ground truth data is obtained from a database comprising (This step is directed to data retrieval, which is an insignificant extra-solution activity. See MPEP § 2106.05(g).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because, when considered individually and in combination, they do not add significantly more (also known as an inventive concept) to the exception. The claim recites insignificant extra-solution activity, such as obtaining ground truth data from a database, which is a well-understood, routine, and conventional practice in the field of computer vision and machine learning, as recognized by the court decisions listed in MPEP § 2106.05(d). Regarding Claim 16 Step 2A Prong 1: providing a set of current images as input to a first machine learning model, wherein the first machine learning model is trained to detect objects of a target class in a given set of images using (i) a training input comprising a set of training images, and (ii) a target output for the training input, the target output comprising, for each respective training image of the set of training images, ground truth data associated with each object depicted in the respective training image, wherein the ground truth data indicates a region of the respective training image that includes a respective object, mask data associated with each object depicted in the respective training image, wherein the mask data is obtained based on one or more outputs of a second machine learning model, and an indication of whether a class associated with each object depicted in the respective training image corresponds to the target class; obtaining one or more outputs of the first machine learning model; and determining, based on the one or more outputs of the first machine learning model, object data associated with each of the set of current images, wherein the object data for each respective current image of the set of current images comprises an indication of a region of the respective current image that includes an object detected in the respective current image and an indication of whether the detected object corresponds to the target class. (These steps for recognizing objects, defining regions, applying mask data, and classifying detected objects involve cognitive evaluations and are recitations of a mental process.) Step 2A Prong 2: providing a set of current images as input to a first machine learning model, wherein the first machine learning model is trained (This step links the judicial exception to a particular technological environment, i.e., training and using a machine learning model for object detection. See MPEP § 2106.05(h).) wherein the mask data is obtained based on one or more outputs of a second machine learning model, (This step links the judicial exception to a particular technological environment, i.e., using a second machine learning model to generate mask data. See MPEP § 2106.05(h).) obtaining one or more outputs of the first machine learning model; (This step links the judicial exception to a particular technological environment, i.e., obtaining model outputs as part of an AI-based image processing system. See MPEP § 2106.05(h).) based on the one or more outputs of the first machine learning model, (This step links the judicial exception to a particular technological environment, i.e., processing machine learning model outputs. See MPEP § 2106.05(h).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because, when considered individually and in combination, they do not add significantly more (also known as an inventive concept) to the exception. The claim recites machine learning models for object detection, obtaining and processing model outputs, and generating mask data using a second machine learning model, which link the judicial exception to a particular technological environment. Regarding Claim 17 Step 2A Prong 1: wherein the object data further comprises mask data associated with the object detected in the respective current image. (This step for specifying mask data involves recognizing and labeling visual information, which is understood to be a recitation of a mental process.) Step 2A Prong 2: The claim does not include additional elements, when considered separately and in combination, that integrate the judicial exception into a practical application. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because, when considered individually and in combination, it does not add significantly more (also known as an inventive concept) to the exception. The claim is directed to a mental process of recognizing and specifying mask data, without any technological improvement or inventive concept. Regarding Claim 18 Step 2A Prong 1: extracting one or more sets of object data from the one or more outputs of the first machine learning model, wherein each of the one or more sets of object data is associated with a level of confidence that the object data corresponds to an object detected in the respective current image; determining whether the level of confidence associated with a respective set of object data satisfies a level of confidence criterion. (These steps involve recognizing and evaluating confidence scores, which are understood to be a recitation of a mental process.) Step 2A Prong 2: extracting one or more sets of object data from the one or more outputs of the first machine learning model, (This step links the judicial exception to a particular technological environment, i.e., extracting confidence scores from a machine learning model. See MPEP § 2106.05(h).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because, when considered individually and in combination, it does not add significantly more (also known as an inventive concept) to the exception. The claim recites extracting confidence scores from a machine learning model, which links the judicial exception to a particular technological environment. Regarding Claim 19 Step 2A Prong 1: providing the set of training images as input to the second machine learning model, wherein the second machine learning model is trained to detect, for a given input image, one or more objects of at least one of a plurality of classes depicted in the given input image and to predict, for each of the one or more detected objects, at least mask data associated with the respective detected object; (This step for detecting objects and predicting mask data involves recognizing and categorizing visual information, which is understood to be a recitation of a mental process.) determine, from one or more outputs of the second machine learning model, object data associated with each of the set of training images, wherein the object data for each respective training image of the set of training images comprises mask data associated with each object detected in the respective image. (This step for determining object data involves recognizing and categorizing visual information, which is understood to be a recitation of a mental process.) Step 2A Prong 2: providing the set of training images as input to the second machine learning model, wherein the second machine learning model is trained to (This step links the judicial exception to a particular technological environment, i.e., training a machine learning model for object detection. See MPEP § 2106.05(h).) from one or more outputs of the second machine learning model, (This step links the judicial exception to a particular technological environment, i.e., processing machine learning model outputs, See MPEP § 2106.05(h).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because, when considered individually and in combination, it does not add significantly more (also known as an inventive concept) to the exception. The claim recites training a machine learning model, obtaining model outputs, and processing them, which links the judicial exception to a particular technological environment. Regarding Claim 20 Step 2A Prong 1: wherein the ground truth data is obtained using a database comprising an indication of one or more bounding boxes associated with the set of training images, (This step for recognizing and associating bounding boxes with training images involves cognitive evaluation and classification, which is understood to be a recitation of a mental process.) wherein each of the one or more bounding boxes were provided by at least one of an accepted bounding box authority entity or a user of a platform. (This step for specifying the source of bounding boxes involves recognizing and categorizing visual information, which is understood to be a recitation of a mental process.) Step 2A Prong 2: wherein the ground truth data is obtained using a database (This step is directed to data retrieval, which is an insignificant extra-solution activity. See MPEP § 2106.05(g).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because, when considered individually and in combination, it does not add significantly more (also known as an inventive concept) to the exception. The claim recites obtaining data from a database, which is an insignificant extra-solution activity and a well-understood, routine, and conventional practice, as recognized by the court decisions listed in MPEP § 2106.05(d). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-4 are rejected under 35 U.S.C. 103 as being unpatentable over Uijlings et al., “Revisiting knowledge transfer for training object class detectors” (hereinafter Uijlings) in view of Li et al., (US20210407090) (hereinafter Li). Regarding claim 1, Uijlings teaches: A method comprising: identifying a first set of images comprising a plurality of objects of a plurality of classes; (Uijlings: “We use ILSVRC 2013 [35]... ILSVRC 2013 has 200 object classes: we use the first 100 as sources S and second 100 as targets T... As our source training set we use all images of the augmented val1 set which have bounding-box annotations for 100 source classes, resulting in 63k images with 81k bounding-boxes.” (Section 3, p.5) (Note: This describes the use of ILSVRC 2013, which contains a dataset with annotated bounding boxes across 100 object classes. This satisfies the limitation of identifying a first set of images with objects of multiple classes.) providing the first set of images as input to a first machine learning model trained to detect, for a given input image, a presence of one or more objects of at least one of the plurality of classes depicted in the given input image ... ; (Uijlings: “We train SSD on the source set S. For each anchor box, SSD regresses to a single output box, along with one confidence score for each source class.” (Section 2.2, p.4); “SSD starts from a dense grid of ‘anchor boxes’ covering the image, and then adjusts their coordinates to match objects using regression.” (Section 2.2, p.3-4)) determining, from one or more first outputs of the first machine learning model, object data associated with each of the first set of images, ... ; (Uijlings: “It produces a set of proposals B and assigns to each proposal b ∈ B scores Fs(b, I) at all levels of the hierarchy. More precisely, it assigns a score for each class s ∈ H, including scores for the original leaf classes S, the intermediate-level classes, and the top-level ‘entity’ class.” (Section 2.2, p.4)) training a second machine learning model to detect objects of a target class in a second set of images, ... and an indication of whether a class associated with each object detected in the at least a subset of the first set of images corresponds to the target class. (Uijlings: “We now train an object detector from the bounding boxes produced on the target training set by MIL. We train a Faster-RCNN detector [33] with Inception-ResNet [43] as base network. We apply it to the target test set and report mean Average Precision (mAP).” (Sec. 3.2, p.7) (Note: This second model is therefore learned specifically for the ‘target classes’ identified under weak supervision.) Uijlings does not teach but Li teaches: ... and to predict at least mask data comprising an indication of one or more pixels of the given input image depicting one or more of the detected objects; (Li: “visual object instance segmentation typically identifies a label for each specific object of interest in an image and the boundary of each specific object of interest at the detailed pixel level within the image.” – Li, Paragraph [0030]; “The foreground-specialized teacher model generates various outputs, which include the classifications of the foreground objects contained in the foreground input images” – Li, Paragraph [0057]; “The mask segmentation loss acts as a measure of errors associated with boundaries of individual instances of the foreground objects in the input images identified by the foreground - specialized teacher model.” – Li, Paragraph [0058]) ... wherein the object data for each respective image of the first set of images comprises mask data indicating the one or more pixels of each respective image with depicting each object detected in the respective image; (Li: “The foreground-specialized teacher model generates various outputs, which include the classifications of the foreground objects contained in the foreground input images” – Li, Paragraph [0057]; “The mask segmentation loss acts as a measure of errors associated with boundaries of individual instances of the foreground objects-- mask data indicating the one or more pixels-- in the input images identified by the foreground-specialized teacher model” – Li, Paragraph [0058, 30]) ... wherein the second machine learning model is trained using at least a subset of the first set of images, and a target output for the at least a subset of the first set of images, wherein the target output comprises the mask data associated with each object detected in the at least a subset of the first set of images ... (Li: “The student model is essentially trained here to imitate the behavior of the foreground-specialized teacher model, meaning the student model is (ideally) trained to produce the same outputs that the foreground-specialized teacher model would have produced using the same images.” – Li, Paragraph [0050]; “The mask segmentation loss acts as a measure of errors associated with boundaries of individual instances of the foreground objects in the input images.” – Li, Paragraph [0058]) Uijlings and Li are analogous art to the present invention because they both address object detection and segmentation in images, focusing on training object‐level models using bounding‐box annotations or mask data for improved detection performance. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Uijlings (for multi‐class bounding‐box detection and knowledge transfer for detecting objects of target classes) and Li (for incorporating explicit instance‐level mask outputs and teacher–student knowledge transfer) in order to achieve finer‐grained object delineation and higher detection accuracy. This combination reflects a teaching, suggestion, or motivation in the prior art, as Li emphasizes that pixel‐level masks further refine object boundaries, and Uijlings demonstrates the benefits of transferring teacher outputs to a student model. One of ordinary skill in the art would have been motivated to make such a combination because it would allow the generation of precise pixel‐level masks to further refine object boundaries, as suggested by Li (see, e.g., Li, Paragraph [0030] and [0058]). Regarding claim 2, Uijlings teaches: The method of claim 1, wherein the first machine learning model is further trained to predict, for each of the one or more detected objects, a particular class of the plurality of classes associated with a respective detected object. (Uijlings: “We train SSD on the source set S. For each anchor box, SSD regresses to a single output box, along with one confidence score for each source class." (Section 2.2, p.4); "Each object bounding-box has multiple class labels, including its original label from S (e.g., ‘bear’) and all its ancestors up to ‘entity.’" (Section 2.2, p.4)) (Note: These disclosures align with the claimed feature that the first machine learning model predicts, for each detected object, a class label from the plurality of classes associated with that object.) Regarding claim 3, Uijlings teaches: The method of claim 2, further comprising: generating the target output, (Uijlings: “It produces a set of proposals B and assigns to each proposal b ∈ B scores Fs (b, I) at all levels of the hierarchy.” (Section 2.2, p.4); “We train SSD on the source set S. For each anchor box, SSD regresses to a single output box, along with one confidence score for each source class.” (Section 2.2)) (Note: These references describe generating outputs, including scores and proposals, that correspond to objects detected in the input images.) wherein generating the target output comprises: determining whether the particular class associated with the respective detected object corresponds to the target class. (Uijlings: “it assigns a score for each class s ∈ H, including scores for the original leaf classes S, the intermediate-level classes, and the top-level ‘entity’ class.” (Section 2.2, p.4); “We use these scoring functions during the re-localization stage of MIL on T, which greatly helps localizing target objects correctly.” (Section 2.2, p.3)) (Note: These references disclose assigning scores for detected objects and evaluating these scores to determine if they correspond to the intended target class during re-localization, which satisfies the limitation.) Regarding claim 4, Uijlings teaches: The method of claim 1, further comprising: identifying, using an indication of one or more bounding boxes associated with the image, ground truth data associated with the respective object depicted in the image. (Uijlings: “We train SSD on the source set S. For each anchor box, SSD regresses to a single output box, along with one confidence score for each source class.” (Section 2.2, p.4); “As our source training set we use all images of the augmented val1 set which have bounding-box annotations for 100 source classes, resulting in 63k images with 81k bounding-boxes.” (Section 3, p.5); “It produces a set of proposals B and assigns to each proposal b ∈ B scores Fs(b, I) at all levels of the hierarchy.” (Section 2.2, p.4)) (Note: These references collectively disclose the process of using bounding-box annotations during training, where the SSD model generates object proposals and aligns them with ground truth data, associating bounding boxes with detected objects in the image. This satisfies the limitation of identifying ground truth data using bounding boxes.) Claims 5 is rejected under 35 U.S.C. 103 as being unpatentable over Uijlings in view of Li further in view of Ksenia Konyushkova et al., “Learning Intelligent Dialogs for Bounding Box Annotation” (hereinafter Konyushkova). Regarding claim 5, Uijlings in view of Li teaches the method of claim 4. Uijlings and Li do not teach but Konyushkova teaches: at least one bounding box of the one or more bounding boxes were provided by at least one of an accepted bounding box authority entity or a user of a platform. (Konyushkova: “The annotator is asked to verify whether the box produced by the algorithm covers an object tightly enough. If not, the process iterates.” – Konyushkova (Section 1, p.1); “The detector is re-trained on all accepted boxes, giving a new detector.” – Konyushkova (Section 5.3, p.7) (Note: The description in Konyushkova (Section 1, p.1) demonstrates that bounding boxes are provided by a user verifying or refining boxes. The explanation in Konyushkova (Section 5.3, p.7) illustrates that bounding boxes are generated by an authoritative entity, such as a trained detector, which is improved iteratively using accepted boxes. Together, these references directly support the claim limitation of bounding boxes being provided by either an accepted bounding box authority entity or a user of a platform.) Konyushkova is analogous art to the present invention because it addresses the improvement of bounding-box accuracy through a user-verified, iterative annotation process. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Konyushkova’s iterative box verification approach with the bounding-box detection of Uijlings as integrated with Li’s instance-segmentation framework in order to yield bounding boxes that are both accurately localized and efficiently verified. One of ordinary skill in the art would have been motivated to make such a combination because it enables the generation of bounding boxes that have been confirmed to tightly cover the target objects, as indicated by Konyushkova’s disclosure that “the annotator is asked to verify whether the box produced by the algorithm covers an object tightly enough… If not, the process iterates” (see Konyushkova, Section 1, p.1). Claims 6-14 and 16-19 are rejected under 35 U.S.C. 103 as being unpatentable over Uijlings in view of Li further in view of Kaiming He et al., “Mask R-CNN” (hereinafter He) further in view of Jiaoda Li et al., “Differentiable Subset Pruning of Transformer Heads” (hereinafter Jiaoda). Regarding claim 6, Uijlings in view of Li teaches the method of claim 1. Uijlings and Li do not teach but He teaches: the second machine learning model is a multi-head machine learning model, (He: “Mask R-CNN adopts the same two-stage procedure, with an identical first stage (which is RPN). In the second stage, in parallel to predicting the class and box offset, Mask R-ClNN also outputs a binary mask for each RoI.” – He (Section 3, p.3)) (Note: This reference clearly describes that Mask R-CNN uses multiple parallel branches (or "heads") to perform distinct tasks such as, classification, bounding box regression, and mask prediction. This aligns with the concept of a multi-head machine learning model as required by the claim.) wherein the method further comprises: upon training the second machine learning model using the at least a subset of the first set of images and the target output, identifying one or more heads of the second machine learning model that correspond to predicting mask data for a given input image; (He: “Mask R-CNN decouples mask and class prediction: as the existing box branch predicts the class label, we generate a mask for each class without competition among classes (by a per-pixel sigmoid and a binary loss.)” – He (Section 4.2, p.6); “The mask branch is a small FCN applied to each RoI, predicting a segmentation mask in a pixel-to-pixel manner.” – He (Section 1, p.1)) (Note: These references describe how Mask R-CNN isolates the "mask branch" to perform pixel-to-pixel segmentation for each Region of Interest (RoI). This decoupling of mask prediction from other tasks demonstrates the identification of specific heads responsible for generating mask data. This directly supports the claim of identifying heads corresponding to mask prediction.) He is analogous art to the present invention because it introduces a multi-head architecture within the Mask R-CNN framework that enhances instance segmentation by employing parallel branches for distinct tasks. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Uijlings and Li with He’s multi-head instance segmentation approach in order to integrate a dedicated mask branch alongside the existing classification and bounding-box regression branches. One of ordinary skill in the art would have been motivated to make such a combination because it enables the generation of precise pixel-level masks that further refine object boundaries and improve model flexibility, as disclosed by He (see, e.g., He, Section 3, p.3). Uijlings, Li and He do not teach, but Jiaoda teaches: updating the second machine learning model to remove the one or more identified heads. (Jiaoda: “Inserting glh into the multi-head attention enables our pruning approach: setting the gate variable to glh=0 means the head attlh is pruned away.” – Jiaoda (Section 2, p.2); “Our method learns per-head importance variables and then enforces a user-specified hard constraint on the number of unpruned heads.” – Jiaoda (Abstract, p.1); “To make our subset pruner differentiable, we apply the Gumbel–softmax trick (Maddison et al., 2017) and its extension to subset selection (Vieira, 2014; Xie and Ermon, 2019). This gives us a pruning scheme that always returns the specified number of heads.” – Jiaoda (Section 3, p.3)) (Note: Jiaoda’s work explicitly describes a systematic method for removing specific heads from a multi-head model. By introducing gating variables (glh), the model dynamically determines which heads to retain or prune during training. The Gumbel-top-K algorithm further ensures that the pruning process adheres to user-defined constraints on the number of retained heads. This demonstrates a clear mechanism for updating the model to remove identified heads, as required by the claim.) Jiaoda is analogous art to the present invention because it presents a method for pruning specific heads in multi-head architectures to optimize model efficiency without significantly compromising performance. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Uijlings, Li, and He with Jiaoda’s head-pruning methodology in order to reduce model complexity while maintaining essential segmentation performance. One of ordinary skill in the art would have been motivated to make such a combination because Jiaoda explicitly discloses the use of gating variables and differentiable subset pruning to systematically remove less important heads from a multi-head model (see, e.g., Jiaoda, Abstract, p.1, and Section 2, p.2). Regarding claim 7, Uijlings in view of Li further in view of He further in view of Jioda teaches the method of claim 6. Li further teaches: providing a third set of images as input to the second machine learning model; (Li: “In some embodiments of this disclosure, for example, the processor 120 may obtain at least one trained student model and use the trained student model(s) to perform visual object instance segmentation for one or more images captured by the electronic device 101.” – Li, Paragraph [0034]) (Note: This describes the deployment of a trained student model to process new images, corresponding to the third set of images in claim 7.) obtaining one or more second outputs of the second machine learning model; (Li: “The method further includes deploying the trained student model to perform visual object instance segmentation in an external device.” – Li, Abstract) (Note: Deploying the student model to perform segmentation implies generating outputs such as classifications and regions, satisfying this limitation.) determining, based on the one or more second outputs, additional object data associated with each of the third set of images, wherein the additional object data for each respective image of the second set of images comprises an indication of a region of the respective image that includes an object detected in the respective image and a class associated with the detected object. (Li: “The foreground-specialized teacher model generates various outputs, which include the classifications of the foreground objects contained in the foreground input images.” – Li, Paragraph [0057]; “The knowledge selection operation helps to train the student model to imitate the behavior of the teacher model.” – Li, Paragraph [0066]) (Note: Paragraph [0057] confirms the generation of classifications as part of the outputs from the teacher model, and Paragraph [0066] explains that the student model is trained to replicate these outputs, including regions and classifications, when processing new images.) Regarding claim 8, Uijlings in view of Li further in view of He further in view of Jioda teaches the method of claim 6. Li further teaches: transmitting the updated second machine learning model to at least one of an edge device or an endpoint device via a network; (Li: “The deployment stage includes transmitting the trained student model to external devices such as a smartphone, an autonomous vehicle, or a virtual-, augmented-, or mixed-reality headset” – Li, Paragraph [0052]; “Any suitable communication mechanisms may be used to deploy the trained student models to the external devices, such as wired communications, wireless communications, or physical transport via a Universal Serial Bus (USB) Flash drive or other portable memory” Li – Paragraph [0052]) (Note: These references explicitly describe the transmission of a trained model, which corresponds to the "updated second machine learning model" in claim 8. The references also mention devices, such as smartphones and autonomous vehicles, which qualify as edge or endpoint devices, and specify wired and wireless networks as communication mechanisms.) Claim 9 is a system claim corresponding to a combination of method claims 1, 4, and 6, and is rejected for the same reasons as given in the rejections of claims 1, 4, and 6. Claim 10 is a system claim corresponding to a combination of method claims 1, 4, 6, and 7 and is rejected for the same reasons as given in the rejection of claims 1, 4, 6, and 7. Claim 11 is a system claim corresponding to a combination of method claims 1, 4, 6, and 8 and is rejected for the same reasons as given in the rejection of claims 1, 4, 6, and 8. Regarding claim 12, Uijlings in view of Li further in view of He further in view of Jioda teaches the system of claim 9. Li further teaches: wherein generating the target output for the training input comprises: providing the image depicting the object as input to the additional machine learning model, (Li: “Note that while a single specialized teacher model and a single student model may be described below in relation to FIG. 2, the training stage 202 may be used to train and the deployment stage 204 may be used to deploy any suitable number of specialized teacher models and any suitable number of student models.” Li – Paragraph [0048]) (Note: The reference indicates the use of multiple teacher models, suggesting that an image could be provided to an additional teacher model distinct from the main student model for further processing.) wherein the additional machine learning model is trained to detect, for a given input image, a presence of one or more objects depicted in the given input image and to predict at least mask data associated with one or more of the detected objects; (Li: “The foreground- specialized teacher model generates various outputs, which include the classifications of the foreground objects contained in the foreground input images.” – Li, Paragraph [0057]; “The mask segmentation loss acts as a measure of errors associated with boundaries of individual instances of the foreground objects in the input images.” – Li, Paragraph [0058]) (Note: These references demonstrate that the teacher model is trained for object detection and mask prediction, which aligns with the functionality of the additional machine learning model described in claim 12.) determining, from one or more outputs of the additional machine learning model, object data associated with the image, wherein the object data for the image comprises mask data associated with the depicted object. (Li: “The mask segmentation loss acts as a measure of errors associated with boundaries of individual instances of the foreground objects in the input images.” – Li, Paragraph [0058]) (Note: The prior art explains that the teacher model generates outputs, including mask data, which corresponds to object data as required by the claim. The additional machine learning model performs similarly.) Regarding claim 13, Uijlings in view of Li further in view of He further in view of Jioda teaches the system of claim 12. Li further teaches: wherein the additional machine learning model is further trained to predict, for each of the one or more detected objects, a class associated with the respective detected object; (Li: “The foreground-specialized teacher model generates various outputs, which include the classifications of the foreground objects contained in the foreground input images.” – Li, Paragraph [0057]; “Training, using at least one processor, a student model to perform visual object instance segmentation in order to segment and classify objects in second training images… wherein training the student model comprises using selected outputs of the specialized teacher model.” – Li, Paragraph [0004]) (Note: These references demonstrate that the teacher model (analogous to the additional machine learning model in claim 13) predicts and associates each detected object with a class.) wherein object data for the image further comprises the indication of the class associated with the depicted object. (Li: “The knowledge selection operation helps to train the student model to imitate the behavior of the teacher model… The distillation of the knowledge learned by the teacher model can focus primarily or exclusively on how the teacher model classifies the foreground objects in the input images.” – Li, Paragraph [0065], [0066]) (Note: The teacher model transfers class-specific data to the student model, ensuring that object data includes class labels.) Claim 14 is a system claim corresponding to a combination of method claims 1, 4, and 6 and is rejected for the same reasons as given in the rejection of claims 1, 4, and 6. Claim 16 is a non-transitory computer-readable storage medium claim corresponding to a combination of method claims 1, 4, and 6 and is rejected for the same reasons as given in the rejections of claims 1, 4, and 6. Regarding claim 17, Uijlings in view of Li further in view of He further in view of Jioda teaches the non-transitory computer-readable storage medium of claim 16: Li further teaches: wherein the object data further comprises mask data associated with the object detected in the respective current image. (Li: “The student model 404 is associated with a soft mask 704, which represents a latent soft feature mask that is applied to the learned features in order to embed foreground awareness into the student model 404. The soft mask 704 can adaptively calibrate the pixel-wise feature responses of the student model 404 based on guidance from the teacher model 310.” – Li, Paragraph [0080]; “In a first approach, the student model 404 is trained to perform a foreground segmentation task in addition to the object classification task, which may be achieved in some embodiments by adding a branch for the foreground segmentation task to the student model's backbone.” – Li, Paragraph [0071]) (Note: These references illustrate that the student model incorporates mechanisms like latent soft feature masks and additional segmentation tasks to generate detailed foreground-aware object data. Paragraph [0080] explains how the soft mask calibrates pixel-level responses based on teacher model guidance, emphasizing that the student model's outputs include detailed mask data for each detected object. Paragraph [0071] further supports this by describing how a segmentation branch added to the student model enables precise identification of object boundaries, aligning with the requirement of producing mask data for detected objects in the claim.) Regarding claim 18, Uijlings in view of Li further in view of He further in view of Jioda teaches the non-transitory computer-readable storage medium of claim 16: Li further teaches: wherein determining object data associated with each of the set of current images comprises: extracting one or more sets of object data from the one or more outputs of the first machine learning model, wherein each of the one or more sets of object data is associated with a level of confidence that the object data corresponds to an object detected in the respective current image; (Li: “In some embodiments, the outputs 312 include softmax classification outputs of the teacher model” – Li, Paragraph [0058]) (Note: The softmax classification outputs of the teacher model provide a probability distribution over the potential classes for detected objects, effectively embedding a level of confidence in the classification. This aligns with the requirement to associate object data with a level of confidence.) determining whether the level of confidence associated with a respective set of object data satisfies a level of confidence criterion. (Li: “In some embodiments, the outputs 312 include softmax classification outputs of the teacher model” – Li, Paragraph [0058]) (Note: The reference to softmax outputs also supports this limitation, as the confidence threshold can be applied to the probability values provided by the softmax outputs, enabling decisions about whether an object meets the confidence criterion. This is a standard interpretation of softmax output in machine learning.) Claim 19, which recites a non-transitory computer-readable storage medium of Claim 16 taught by Uijlings, Li, He, and Jioda, is rejected on the basis that it is not patentably distinct from Claim 1, which has been rejected. Claim 1 and Claim 19 describe substantially the same subject matter, albeit with different framing. Specifically, the first machine learning model in Claim 1 corresponds to the second machine learning model in Claim 19, and vice versa. Both claims outline processes where one model's outputs are used to train the other, with Claim 1 focusing on training the second model using outputs from the first and Claim 19 focusing on training the first model using outputs from the second. Claim 15 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Uijlings in view of Li further in view of He further in view of Jiaoda further in view of Doke et al., (US12094203) (hereinafter Doke) further in view of Konyushkova. Regarding claim 15, Uijlings in view of Li further in view of He further in view of Jioda teaches the system of claim 14. Uijlings, Li, He, and Jioda do not teach but Doke teaches: wherein the ground truth data is obtained from a database comprising an indication of one or more bounding boxes associated with objects depicted in a set of images, (Doke: “After determining the ground truth result, the system(s) may store, in the one or more databases, data representing the ground truth result for the second image data” – Doke (Detailed Description, p.14); “After identifying the region of the video data that represents the identifier, the localization component may provide an indication of the region to a reader component. For example, the localization component may provide coordinates of the bounding box, the video data corresponding to the region itself, and/or other information to the reader component.” – Doke (Detailed Description, p.15) (Note: Doke explicitly describes obtaining ground truth data (such as bounding box coordinates) associated with images and storing this data within one or more databases. The mention of "coordinates of the bounding box" linked with image data demonstrates that the ground truth includes bounding boxes for objects depicted in images) Doke is analogous art to the present invention because it discloses the organization and storage of annotated training data in a database. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Uijlings, Li, He, and Jiaoda with Doke’s disclosure in order to achieve efficient and organized retrieval of ground truth data. One of ordinary skill in the art would have been motivated to make such a combination because Doke explicitly discloses that, after determining the ground truth result, a system stores data representing the ground truth (including bounding box coordinates) in one or more databases (see, e.g., Doke, Detailed Description, p.14 and p.15). Uijlings, Li, He, Jioda, and Doke do not teach but Konyushkova teaches: wherein the image is included in the set of images, and wherein the one or more bounding boxes is provided by an accepted bounding box authority entity or a user of a platform. (Konyushkova: “The annotator is asked to verify whether the box produced by the algorithm covers an object tightly enough. If not, the process iterates.” – Konyushkova (Section 1, p.1); “The detector is re-trained on all accepted boxes, giving a new detector.” – Konyushkova (Section 5.3, p.7) (Note: The description in Konyushkova (Section 1, p.1) demonstrates that bounding boxes are provided by a user verifying or refining boxes. The explanation in Konyushkova (Section 5.3, p.7) illustrates that bounding boxes are generated by an authoritative entity, such as a trained detector, which is improved iteratively using accepted boxes. Together, these references directly support the claim limitation of bounding boxes being provided by either an accepted bounding box authority entity or a user of a platform.) Konyushkova is analogous art to the present invention because it discloses the incorporation of human verification and iterative refinement to ensure accurate bounding box annotations. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Uijlings, Li, He, Jiaoda, and Doke with Konyushkova’s iterative box verification technique in order to improve annotation quality and efficiency. One of ordinary skill in the art would have been motivated to make such a combination because Konyushkova discloses that annotators are required to verify whether a box produced by an algorithm accurately covers an object and, if not, to repeat the process until an acceptable bounding box is obtained (see, e.g., Konyushkova, Section 1, p.1). Claim 20 is a non-transitory computer readable storage medium claim corresponding to a combination of method claims 1, 4, 6 and system claim 15 and is rejected for the same reasons as given in the rejections of claims 1, 4, 6, and 15. Response to Arguments Applicant's arguments filed on 5/27/2025 have been fully considered but they are not persuasive. The Applicant indicates that “Uijlings is silent regarding "a first machine learning model trained to...predict at least mask data comprising an indication of one or more pixels of the given input image depicting one or more of the detected objects…”. Li teaches “visual object instance segmentation typically identifies a label for each specific object of interest in an image and the boundary of each specific object of interest at the detailed pixel level within the image.” –Paragraph [0030]. Claims 2-19 are rejected at least based on the rationale found in Li as indicated above. Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Cesar Paula whose telephone number is (571)272-4128. The examiner can normally be reached Monday - Friday, 8:30 am - 5:00 pm ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached on (571) 272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /CESAR B PAULA/Supervisory Patent Examiner, Art Unit 2145
Read full office action

Prosecution Timeline

Show 2 earlier events
May 14, 2025
Examiner Interview Summary
May 14, 2025
Applicant Interview (Telephonic)
May 27, 2025
Response Filed
Sep 23, 2025
Final Rejection mailed — §101, §103
Dec 15, 2025
Examiner Interview Summary
Dec 23, 2025
Request for Continued Examination
Jan 16, 2026
Response after Non-Final Action
Aug 12, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12699894
THREE-DIMENSIONAL OBJECT DETECTION USING PSEUDO-LABELS
4y 7m to grant Granted Aug 04, 2026
Patent 12670367
APPARATUS AND METHOD WITH NEURAL NETWORK OPERATION
3y 4m to grant Granted Jun 30, 2026
Patent 12596934
PREDICTION-MODEL-BUILDING METHOD, STATE PREDICTION METHOD AND DEVICES THEREOF
4y 0m to grant Granted Apr 07, 2026
Patent 12585982
MODEL MANAGEMENT USING CONTAINERS
4y 5m to grant Granted Mar 24, 2026
Patent 12585859
SYSTEM AND METHOD FOR IMPROVING THE CLARITY OF OVERLAPPING OBJECTS
1y 2m to grant Granted Mar 24, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
34%
Grant Probability
42%
With Interview (+8.2%)
4y 6m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 173 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month