DETAILED ACTION
This Non-Final rejection is response to the RCE amendment to the claims filed 12/23/2025. Claim 6 has been canceled. Claims 1-5, and 7-20 are pending. Claims 1, 9, and 16 are independent claims.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Drawings
The objection to the drawings has been withdrawn as necessitated by the amendment.
Claim Rejections - 35 USC § 101
The rejections of claims 1-20 rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more, have been withdrawn as necessitated by the amendment.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-4, 7-14, and 16-19 are rejected under 35 U.S.C. 103 as being unpatentable over Uijlings et al., “Revisiting knowledge transfer for training object class detectors” (hereinafter Uijlings) in view of Li et al., (US20210407090) (hereinafter Li), in view of Kaiming He et al., “Mask R-CNN” (hereinafter He), and further in view of Jiaoda Li et al., “Differentiable Subset Pruning of Transformer Heads” (hereinafter Jiaoda).
Regarding claim 1, Uijlings teaches:
A method comprising: identifying a first set of images comprising a plurality of objects of a plurality of classes; (Uijlings: “We use ILSVRC 2013 [35]... ILSVRC 2013 has 200 object classes: we use the first 100 as sources S and second 100 as targets T... As our source training set we use all images of the augmented val1 set which have bounding-box annotations for 100 source classes, resulting in 63k images with 81k bounding-boxes.” (Section 3, p.5) (Note: This describes the use of ILSVRC 2013, which contains a dataset with annotated bounding boxes across 100 object classes. This satisfies the limitation of identifying a first set of images with objects of multiple classes.)
providing the first set of images as input to a first machine learning model trained to detect, for a given input image, a presence of one or more objects of at least one of the plurality of classes depicted in the given input image ... ; (Uijlings: “We train SSD on the source set S. For each anchor box, SSD regresses to a single output box, along with one confidence score for each source class.” (Section 2.2, p.4); “SSD starts from a dense grid of ‘anchor boxes’ covering the image, and then adjusts their coordinates to match objects using regression.” (Section 2.2, p.3-4))
determining, from one or more first outputs of the first machine learning model, object data associated with each of the first set of images, ... ; (Uijlings: “It produces a set of proposals B and assigns to each proposal b ∈ B scores Fs(b, I) at all levels of the hierarchy. More
precisely, it assigns a score for each class s ∈ H, including scores for the original leaf classes S, the intermediate-level classes, and the top-level ‘entity’ class.” (Section 2.2, p.4))
training a second machine learning model to detect objects of a target class in a second set of images, ... and an indication of whether a class associated with each object detected in the at least a subset of the first set of images corresponds to the target class. (Uijlings: “We now train an object detector from the bounding boxes produced on the target training set by MIL. We train a Faster-RCNN detector [33] with Inception-ResNet [43] as base network. We apply it to the target test set and report mean Average Precision (mAP).” (Sec. 3.2, p.7) (Note: This second model is therefore learned specifically for the ‘target classes’ identified under weak supervision.)
Uijlings does not teach but Li teaches:
... and to predict at least mask data comprising an indication of one or more pixels of the given input image depicting one or more of the detected objects; (Li: “visual object instance segmentation typically identifies a label for each specific object of interest in an image and the boundary of each specific object of interest at the detailed pixel level within the image.” – Li, Paragraph [0030]; “The foreground-specialized teacher model generates various outputs, which include the classifications of the foreground objects contained in the foreground input images” – Li, Paragraph [0057]; “The mask segmentation loss acts as a measure of errors associated with boundaries of individual instances of the foreground objects in the input images identified by the foreground - specialized teacher model.” – Li, Paragraph [0058])
... wherein the object data for each respective image of the first set of images comprises mask data indicating the one or more pixels of each respective image with depicting each object detected in the respective image; (Li:
“The foreground-specialized teacher model generates various outputs, which include the classifications of the foreground objects contained in the foreground input images” – Li, Paragraph [0057]; “The mask segmentation loss acts as a measure of errors associated with boundaries of individual instances of the foreground objects-- mask data indicating the one or more pixels-- in the input images identified by the foreground-specialized teacher model” – Li, Paragraph [0058, 30])
... wherein the second machine learning model is trained using at least a subset of the first set of images, and a target output for the at least a subset of the first set of images, wherein the target output comprises the mask data associated with each object detected in the at least a subset of the first set of images ... (Li: “The student model is essentially trained here to imitate the behavior of the foreground-specialized teacher model, meaning the student model is (ideally) trained to produce the same outputs that the foreground-specialized teacher model would have produced using the same images.” – Li, Paragraph [0050]; “The mask segmentation loss acts as a measure of errors associated with boundaries of individual instances of the foreground objects in the input images.” – Li, Paragraph [0058])
Uijlings and Li are analogous art to the present invention because they both address object detection and segmentation in images, focusing on training object‐level models using bounding‐box annotations or mask data for improved detection performance. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Uijlings (for multi‐class bounding‐box detection and knowledge transfer for detecting objects of target classes) and Li (for incorporating explicit instance‐level mask outputs and teacher–student knowledge transfer) in order to achieve finer‐grained object delineation and higher detection accuracy. This combination reflects a teaching, suggestion, or
motivation in the prior art, as Li emphasizes that pixel‐level masks further refine object boundaries, and Uijlings demonstrates the benefits of transferring teacher outputs to a student model. One of ordinary skill in the art would have been motivated to make such a combination because it would allow the generation of precise pixel‐level masks to further refine object boundaries, as suggested by Li (see, e.g., Li, Paragraph [0030] and [0058]).
Uijlings does not teach:
wherein the trained second machine learning comprises a plurality of model heads each associated with a distinct task, the plurality of model heads comprising at least one model head associated with predicting a bounding box associated with an object detected in a given image and at least one model head associated with predicting mask data associated with the object detected in a given image. However, He teaches “Mask R-CNN adopts the same two-stage procedure, with an identical first stage (which is RPN). In the second stage, in parallel to predicting the class and box offset, Mask R-ClNN also outputs a binary mask for each RoI.” – He (Section 3, p.3, fig.2)) (Note: This reference clearly describes that Mask R-CNN uses multiple parallel branches (or "heads") to perform distinct tasks. One stage proposes bounding boxes around the objects in the images. The other stage outputs binary masks for each region of interest. This aligns with the concept of a multi-head machine learning model as required by the claim.)
He is analogous art to the present invention because it introduces a multi-head architecture within the Mask R-CNN framework that enhances instance segmentation by employing parallel branches for distinct tasks. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Uijlings and Li with He’s multi-head instance segmentation approach in order to integrate a dedicated mask branch alongside the existing classification and bounding-box regression branches. One of ordinary skill in the art would have been motivated to make such a combination
because it enables the generation of precise pixel-level masks that further refine object boundaries and improve model flexibility, as disclosed by He (see, e.g., He, Section 3, p.3).
Uijlings does not teach, but Jiaoda teaches:
and upon training the second machine learning model, updating the second machine learning model to remove the at least one model head associated with predicting the mask data associated with the object detected in a given image. (Jiaoda teaches using a pipelined training where pruning or removing of the model head(s) is performed after training, “Inserting glh into the multi-head attention enables our pruning approach: setting the gate variable to glh=0 means the head attlh is pruned away.” (p.2, col.1, parag.2, col.2, Section 2). Jiaoda also shows that the head(s) that are removed are those with least importance scores (fig.,1b); “Our method learns per-head importance variables and then enforces a user-specified hard constraint on the number of unpruned heads.” – Jiaoda (Abstract, p.1); “To make our subset pruner differentiable, we apply the Gumbel–softmax trick (Maddison et al., 2017) and its extension to subset selection (Vieira, 2014; Xie and Ermon, 2019). This gives us a pruning scheme that always returns the specified number of heads.” – Jiaoda (Section 3, p.3)) (Note: Jiaoda’s work explicitly describes a systematic method for removing specific heads from a multi-head model. By introducing gating variables (glh), the model dynamically determines which heads to retain or prune during training. The Gumbel-top-K algorithm further ensures that the pruning process adheres to user-defined constraints on the number of retained heads. This demonstrates a clear mechanism for updating the model to remove identified heads, as required by the claim.).
Jiaoda is analogous art to the present invention because it presents a method for pruning specific heads in multi-head architectures to optimize model efficiency without significantly compromising performance. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Uijlings, Li, He, Jiaoda’s head-pruning methodology in order to reduce model complexity while still achieving
high performance. One of ordinary skill in the art would have been motivated to make such a combination because Jiaoda explicitly discloses the use of gating variables and differentiable subset pruning to systematically remove less important heads from a multi-head model (see, e.g., Jiaoda, Abstract, p.1, and Section 2, p.2).
Regarding claim 2, Uijlings teaches:
The method of claim 1, wherein the first machine learning model is further trained to predict, for each of the one or more detected objects, a particular class of the plurality of classes associated with a respective detected object. (Uijlings: “We train SSD on the source set S. For each anchor box, SSD regresses to a single output box, along with one confidence score for each source class." (Section 2.2, p.4); "Each object bounding-box has multiple class labels, including its original label from S (e.g., ‘bear’) and all its ancestors up to ‘entity.’" (Section 2.2, p.4)) (Note: These disclosures align with the claimed feature that the first machine learning model predicts, for each detected object, a class label from the plurality of classes associated with that object.)
Regarding claim 3, Uijlings teaches:
The method of claim 2, further comprising: generating the target output, (Uijlings: “It produces a set of proposals B and assigns to each proposal b ∈ B scores Fs (b, I) at all levels of the hierarchy.” (Section 2.2, p.4); “We train SSD on the source set S. For each anchor box, SSD regresses to a single output box, along with one confidence score for each source class.”
(Section 2.2)) (Note: These references describe generating outputs, including scores and proposals, that correspond to objects detected in the input images.)
wherein generating the target output comprises: determining whether the particular class associated with the respective detected object corresponds to the target class. (Uijlings: “it assigns a score for each class s ∈ H, including scores for the original leaf classes S, the intermediate-level classes, and the top-level ‘entity’ class.” (Section 2.2, p.4); “We use these scoring functions during the re-localization stage of MIL on T, which greatly helps localizing target objects correctly.” (Section 2.2, p.3)) (Note: These references disclose assigning scores for detected objects and evaluating these scores to determine if they correspond to the intended target class during re-localization, which satisfies the limitation.)
Regarding claim 4, Uijlings teaches:
The method of claim 1, further comprising: identifying, using an indication of one or more bounding boxes associated with the image, ground truth data associated with the respective object depicted in the image. (Uijlings: “We train SSD on the source set S. For each anchor box, SSD regresses to a single output box, along with one confidence score for each source class.” (Section 2.2, p.4); “As our source training set we use all images of the augmented val1 set which have bounding-box annotations for 100 source classes, resulting in 63k images with 81k bounding-boxes.” (Section 3, p.5); “It produces a set of proposals B and assigns to each proposal b ∈ B scores Fs(b, I) at all levels of the hierarchy.” (Section 2.2, p.4)) (Note: These references collectively disclose the process of using bounding-box annotations during training, where the SSD model generates object proposals and aligns them with ground truth data,
associating bounding boxes with detected objects in the image. This satisfies the limitation of identifying ground truth data using bounding boxes.)
Regarding claim 7, Uijlings in view of Li further in view of He further in view of Jioda teaches the method of claim 6.
Li further teaches:
providing a third set of images as input to the second machine learning model; (Li: “In some embodiments of this disclosure, for example, the processor 120 may obtain at least one trained student model and use the trained student model(s) to perform visual object instance segmentation for one or more images captured by the electronic device 101.” – Li, Paragraph [0034]) (Note: This describes the deployment of a trained student model to process new images, corresponding to the third set of images in claim 7.)
obtaining one or more second outputs of the second machine learning model; (Li: “The method further includes deploying the trained student model to perform visual object instance segmentation in an external device.” – Li, Abstract) (Note: Deploying the student model to perform segmentation implies generating outputs such as classifications and regions, satisfying this limitation.)
determining, based on the one or more second outputs, additional object data associated with each of the third set of images, wherein the additional object data for each respective image of the second set of images comprises an indication of a region of the respective image that includes an object detected in the respective image and a class
associated with the detected object. (Li: “The foreground-specialized teacher model generates various outputs, which include the classifications of the foreground objects contained in the foreground input images.” – Li, Paragraph [0057]; “The knowledge selection operation helps to train the student model to imitate the behavior of the teacher model.” – Li, Paragraph [0066]) (Note: Paragraph [0057] confirms the generation of classifications as part of the outputs from the teacher model, and Paragraph [0066] explains that the student model is trained to replicate these outputs, including regions and classifications, when processing new images.)
Regarding claim 8, Uijlings in view of Li further in view of He further in view of Jioda teaches the method of claim 6.
Li further teaches:
transmitting the updated second machine learning model to at least one of an edge device or an endpoint device via a network; (Li: “The deployment stage includes transmitting the trained student model to external devices such as a smartphone, an autonomous vehicle, or a virtual-, augmented-, or mixed-reality headset” – Li, Paragraph [0052]; “Any suitable communication mechanisms may be used to deploy the trained student models to the external devices, such as wired communications, wireless communications, or physical transport via a Universal Serial Bus (USB) Flash drive or other portable memory” Li – Paragraph [0052]) (Note: These references explicitly describe the transmission of a trained model, which corresponds to the "updated second machine learning model" in claim 8. The references also mention devices, such as smartphones and autonomous vehicles, which qualify as edge or endpoint devices, and specify wired and wireless networks as communication mechanisms.)
Claim 9 is a system claim corresponding to a combination of method claims 1, 4, and 6, and is rejected for the same reasons as given in the rejections of claims 1, 4, and 6.
Claim 10 is a system claim corresponding to a combination of method claims 1, 4, 6, and 7 and is rejected for the same reasons as given in the rejection of claims 1, 4, 6, and 7.
Claim 11 is a system claim corresponding to a combination of method claims 1, 4, 6, and 8 and is rejected for the same reasons as given in the rejection of claims 1, 4, 6, and 8.
Regarding claim 12, Uijlings in view of Li further in view of He further in view of Jioda teaches the system of claim 9. Li further teaches:
wherein generating the target output for the training input comprises: providing the image depicting the object as input to the additional machine learning model, (Li: “Note that while a single specialized teacher model and a single student model may be described below in relation to FIG. 2, the training stage 202 may be used to train and the deployment stage 204 may be used to deploy any suitable number of specialized teacher models and any suitable number of student models.” Li – Paragraph [0048]) (Note: The reference indicates the use of multiple teacher models, suggesting that an image could be provided to an additional teacher model distinct from the main student model for further processing.)
wherein the additional machine learning model is trained to detect, for a given input image, a presence of one or more objects depicted in the given input image and to predict at least mask data associated with one or more of the detected objects; (Li: “The foreground-
specialized teacher model generates various outputs, which include the classifications of the foreground objects contained in the foreground input images.” – Li, Paragraph [0057]; “The mask segmentation loss acts as a measure of errors associated with boundaries of individual instances of the foreground objects in the input images.” – Li, Paragraph [0058]) (Note: These references demonstrate that the teacher model is trained for object detection and mask prediction, which aligns with the functionality of the additional machine learning model described in claim 12.)
determining, from one or more outputs of the additional machine learning model, object data associated with the image, wherein the object data for the image comprises mask data associated with the depicted object. (Li: “The mask segmentation loss acts as a measure of errors associated with boundaries of individual instances of the foreground objects in the input images.” – Li, Paragraph [0058]) (Note: The prior art explains that the teacher model generates outputs, including mask data, which corresponds to object data as required by the claim. The additional machine learning model performs similarly.)
Regarding claim 13, Uijlings in view of Li further in view of He further in view of Jioda teaches the system of claim 12. Li further teaches:
wherein the additional machine learning model is further trained to predict, for each of the one or more detected objects, a class associated with the respective detected object; (Li: “The foreground-specialized teacher model generates various outputs, which include the classifications of the foreground objects contained in the foreground input images.” – Li, Paragraph [0057]; “Training, using at least one processor, a student model to perform visual
object instance segmentation in order to segment and classify objects in second training images… wherein training the student model comprises using selected outputs of the specialized teacher model.” – Li, Paragraph [0004]) (Note: These references demonstrate that the teacher model (analogous to the additional machine learning model in claim 13) predicts and associates each detected object with a class.)
wherein object data for the image further comprises the indication of the class associated with the depicted object. (Li: “The knowledge selection operation helps to train the student model to imitate the behavior of the teacher model… The distillation of the knowledge learned by the teacher model can focus primarily or exclusively on how the teacher model classifies the foreground objects in the input images.” – Li, Paragraph [0065], [0066]) (Note: The teacher model transfers class-specific data to the student model, ensuring that object data includes class labels.)
Claim 14 is a system claim corresponding to a combination of method claims 1, 4, and 6 and is rejected for the same reasons as given in the rejection of claims 1, 4, and 6.
Claim 16 is a non-transitory computer-readable storage medium claim corresponding to a combination of method claims 1, 4, and 6 and is rejected for the same reasons as given in the rejections of claims 1, 4, and 6.
Regarding claim 17, Uijlings in view of Li further in view of He further in view of Jioda teaches the non-transitory computer-readable storage medium of claim 16:
Li further teaches:
wherein the object data further comprises mask data associated with the object detected in the respective current image. (Li: “The student model 404 is associated with a soft mask 704, which represents a latent soft feature mask that is applied to the learned features in order to embed foreground awareness into the student model 404. The soft mask 704 can adaptively calibrate the pixel-wise feature responses of the student model 404 based on guidance from the teacher model 310.” – Li, Paragraph [0080]; “In a first approach, the student model 404 is trained to perform a foreground segmentation task in addition to the object classification task, which may be achieved in some embodiments by adding a branch for the foreground segmentation task to the student model's backbone.” – Li, Paragraph [0071]) (Note: These references illustrate that the student model incorporates mechanisms like latent soft feature masks and additional segmentation tasks to generate detailed foreground-aware object data. Paragraph [0080] explains how the soft mask calibrates pixel-level responses based on teacher model guidance, emphasizing that the student model's outputs include detailed mask data for each detected object. Paragraph [0071] further supports this by describing how a segmentation branch added to the student model enables precise identification of object boundaries, aligning with the requirement of producing mask data for detected objects in the claim.)
Regarding claim 18, Uijlings in view of Li further in view of He further in view of Jioda teaches the non-transitory computer-readable storage medium of claim 16:
Li further teaches:
wherein determining object data associated with each of the set of current images comprises: extracting one or more sets of object data from the one or more outputs of the first machine learning model, wherein each of the one or more sets of object data is
associated with a level of confidence that the object data corresponds to an object detected in the respective current image; (Li: “In some embodiments, the outputs 312 include softmax classification outputs of the teacher model” – Li, Paragraph [0058]) (Note: The softmax classification outputs of the teacher model provide a probability distribution over the potential classes for detected objects, effectively embedding a level of confidence in the classification. This aligns with the requirement to associate object data with a level of confidence.)
determining whether the level of confidence associated with a respective set of object data satisfies a level of confidence criterion. (Li: “In some embodiments, the outputs 312 include softmax classification outputs of the teacher model” – Li, Paragraph [0058]) (Note: The reference to softmax outputs also supports this limitation, as the confidence threshold can be applied to the probability values provided by the softmax outputs, enabling decisions about whether an object meets the confidence criterion. This is a standard interpretation of softmax output in machine learning.)
Claim 19, which recites a non-transitory computer-readable storage medium of Claim 16 taught by Uijlings, Li, He, and Jioda, is rejected on the basis that it is not patentably distinct from Claim 1, which has been rejected. Claim 1 and Claim 19 describe substantially the same subject matter, albeit with different framing. Specifically, the first machine learning model in Claim 1 corresponds to the second machine learning model in Claim 19, and vice versa. Both claims outline processes where one model's outputs are used to train the other, with Claim 1 focusing on training the second model using outputs from the first and Claim 19 focusing on training the first model using outputs from the second.
Claims 5 is rejected under 35 U.S.C. 103 as being unpatentable over Uijlings in view of Li, in view of He, in view of Jiaoda, and further in view of Ksenia Konyushkova et al., “Learning Intelligent Dialogs for Bounding Box Annotation” (hereinafter Konyushkova).
Regarding claim 5, Uijlings in view of Li teaches the method of claim 4.
Uijlings and Li do not teach but Konyushkova teaches:
at least one bounding box of the one or more bounding boxes were provided by at least one of an accepted bounding box authority entity or a user of a platform. (Konyushkova: “The annotator is asked to verify whether the box produced by the algorithm covers an object tightly enough. If not, the process iterates.” – Konyushkova (Section 1, p.1); “The detector is re-trained on all accepted boxes, giving a new detector.” – Konyushkova (Section 5.3, p.7) (Note: The description in Konyushkova (Section 1, p.1) demonstrates that bounding boxes are provided by a user verifying or refining boxes. The explanation in Konyushkova (Section 5.3, p.7) illustrates that bounding boxes are generated by an authoritative entity, such as a trained detector, which is improved iteratively using accepted boxes. Together, these references directly support the claim limitation of bounding boxes being provided by either an accepted bounding box authority entity or a user of a platform.)
Konyushkova is analogous art to the present invention because it addresses the improvement of bounding-box accuracy through a user-verified, iterative annotation process. It
would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Konyushkova’s iterative box verification approach with the bounding-box detection of Uijlings as integrated with Li’s instance-segmentation framework in order to yield bounding boxes that are both accurately localized and efficiently verified. One of ordinary skill in the art would have been motivated to make such a combination because it enables the generation of bounding boxes that have been confirmed to tightly cover the target objects, as indicated by Konyushkova’s disclosure that “the annotator is asked to verify whether the box produced by the algorithm covers an object tightly enough… If not, the process iterates” (see Konyushkova, Section 1, p.1).
Claim 15 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Uijlings in view of Li further in view of He, in view of Jiaoda, further in view of Doke et al., (US12094203) (hereinafter Doke) further in view of Konyushkova.
Regarding claim 15, Uijlings in view of Li further in view of He further in view of Jioda teaches the system of claim 14.
Uijlings, Li, He, and Jioda do not teach but Doke teaches:
wherein the ground truth data is obtained from a database comprising an indication of one or more bounding boxes associated with objects depicted in a set of images, (Doke: “After determining the ground truth result, the system(s) may store, in the one or more databases, data representing the ground truth result for the second image data” – Doke (Detailed Description, p.14); “After identifying the region of the video data that represents the identifier, the localization component may provide an indication of the region to a reader component. For example, the localization component may provide coordinates of the bounding box, the video data corresponding to the region itself, and/or other information to the reader component.” – Doke (Detailed Description, p.15) (Note: Doke explicitly describes obtaining ground truth data (such as bounding box coordinates) associated with images and storing this data within one or more databases. The mention of "coordinates of the bounding box" linked with image data demonstrates that the ground truth includes bounding boxes for objects depicted in images)
Doke is analogous art to the present invention because it discloses the organization and storage of annotated training data in a database. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings
of Uijlings, Li, He, and Jiaoda with Doke’s disclosure in order to achieve efficient and organized retrieval of ground truth data. One of ordinary skill in the art would have been motivated to make such a combination because Doke explicitly discloses that, after determining the ground truth result, a system stores data representing the ground truth (including bounding box coordinates) in one or more databases (see, e.g., Doke, Detailed Description, p.14 and p.15).
Uijlings, Li, He, Jioda, and Doke do not teach but Konyushkova teaches:
wherein the image is included in the set of images, and wherein the one or more bounding boxes is provided by an accepted bounding box authority entity or a user of a platform. (Konyushkova: “The annotator is asked to verify whether the box produced by the algorithm covers an object tightly enough. If not, the process iterates.” – Konyushkova (Section 1, p.1); “The detector is re-trained on all accepted boxes, giving a new detector.” – Konyushkova (Section 5.3, p.7) (Note: The description in Konyushkova (Section 1, p.1) demonstrates that bounding boxes are provided by a user verifying or refining boxes. The explanation in Konyushkova (Section 5.3, p.7) illustrates that bounding boxes are generated by an authoritative entity, such as a trained detector, which is improved iteratively using accepted boxes. Together, these references directly support the claim limitation of bounding boxes being provided by either an accepted bounding box authority entity or a user of a platform.)
Konyushkova is analogous art to the present invention because it discloses the incorporation of human verification and iterative refinement to ensure accurate bounding box annotations. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Uijlings, Li, He, Jiaoda, and
Doke with Konyushkova’s iterative box verification technique in order to improve annotation quality and efficiency. One of ordinary skill in the art would have been motivated to make such a combination because Konyushkova discloses that annotators are required to verify whether a box produced by an algorithm accurately covers an object and, if not, to repeat the process until an acceptable bounding box is obtained (see, e.g., Konyushkova, Section 1, p.1).
Claim 20 is a non-transitory computer readable storage medium claim corresponding to a combination of method claims 1, 4, 6 and system claim 15 and is rejected for the same reasons as given in the rejections of claims 1, 4, 6, and 15.
Response to Arguments
Applicant's arguments filed on 12/23/2025 have been fully considered but they are not persuasive.
Regarding claim 1, the Applicant indicates that “During the Examiner interview, the Examiner stated that it appears that claim 1, as amended, overcomes the current rejection subject to further search and consideration.…”. Upon further review, the newly introduced amendment is taught by a combination of Uijilings in view of Li, He, and Jiaoda as indicated above.
Claims 2-5, 7-20 are rejected at least based on the rejection of and their dependency on claim 1 as introduced above.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Cesar Paula whose telephone number is (571)272-4128. The examiner can normally be reached Monday - Friday, 8:30 am - 5:00 pm ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached on (571) 272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available
to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.