DETAILED ACTION
Response to Arguments
Applicant’s arguments, see Page 2-5, filed 5/26/2026, with respect to 101 rejections have been fully considered and are persuasive. The 101 rejections of claims 1-20 has been withdrawn.
Regarding the 103 rejections, applicant's arguments filed on 5/26/2026 have been fully considered but they are not persuasive for the following reasons:
Applicant argued that “In McKay, each constituent ML labeler is trained for a different product category [ Examiner’s response: applicant’s argument is not persuasive because none of the rejected independent claims expressly requires the multiple AI models to be trained to recognize the same object, or product category. Nor do the claims exclude the arrangement in which different AI models identify different objects, features or categories within the same image and the resulting outputs are subsequently used to determine a label for the image. For example, under the broad language of the claims, different ML labelers may respectively identify features such as lags, arms, a head, and a face within the first image, and output of those labelers may collectively (aggregated) be used to determine/label that the first image depicts a human being]. Each labeler is looking for different objects in the image, not providing independent assessments of the same labeling task [ Examiner’s response: this argument relies on a limitation that is not recited in the claims. The claims do not require each AI model to provide an independent assessment of the same labeling task, nor do they require each model to answer the same classification question. The fact that Mckay’s individual ML labelers may be configured to identify different objects or categories doesn’t preclude their respective output from being collectively used to determine an overall label for the first image; as stated previously, just because different ML labelers are trained on different objects recognition (legs, faces, arms etc.….), doesn’t mean the system, as a whole, could not recognize an object (a human being) in the first image, when each recognition results are aggregated]. The splitter in McKay routes the same image to different category-specific labelers to detect different types of objects, with each labeler returning a label for only its own category. The final result is an aggregation of different category detections, not a consensus determination of a single label based on multiple independent assessments of the same labeling question [Examiner’s response: The examiner disagrees. Applicant’s characterization again imports a “consensus determination” requirement and a requirement that the models independently answer the “same labeling question” neither of which is expressly recited in the claims. Mckay teaches “labeling result fan-in to assemble labeling results from each constituent labeler into an overall labeling result”. McKay further explains, Para. 69, “Multiple ML labelers, human labelers or other labelers can be composed together into directed graphs as needed, such that each individual labeler solves a portion of an overall classification problem, and the results are aggregated together to form the overall labeled output. The overall labeling graph for a use case [for each image (first image)] can be thought of abstractly as a single labeler, and each labeler may itself be implemented as a directed graph.”. Thus, Mckay expressly contemplates multiple individual labelers producing respective labeling results that are aggregated to form an overall labeled output. The claims do not require the multiple responses to constitute competing answers to the same classification question. Rather, the claims broadly require receiving multiple responses from multiple AI model systems and determining, based on those multiple responses, a label for the first image. Accordingly, Mckay’s use of the outputs from multiple constituent labelers to produce an overall labeled output reasonably reads on the claimed relationship. Moreover, nothing in the claims excludes an embodiment in which the multiple AI models detect different types of objects for features and each AI model returns a label corresponding to its respective category, with those response subsequently being used to determine an overall label of an image].
NOTE: Nowhere in the claim it is expressly stated that the multiple artificial intelligence (AI) models do not detect different types of objects (sub-objects) category, where each AI models returns a label for only its own category (which is then aggregated for overall labeling). Furthermore, the claim states “receiving multiple responses from the multiple artificial intelligence model systems”. The claim doesn’t expressly state that the received response is based on (in response to) the provided inputs and first image. In other words, the received response could also anything else other than the provided inputs and first image.
The claim also states “determining, based on the multiple responses, a label for the first image”. Since the claim doesn’t explicitly state what the responses and the labels are, labeling an object in the image could also be considered labeling the image itself. For example, it image depicts the human and the object is labeled “human”, it could also be considered labeling the “first” image.
Applicant also argued that the present disclosure generates "multiple inputs to multiple artificial intelligence model systems" and then determines "based on the multiple responses, a label for the first image." As disclosed in the specification at paragraphs [0233]-[0238], the claimed system submits the same image with prompts to multiple AI model systems [Applicant’s reliance on the disclosed embodiment doesn’t establish that the claims require the narrower features described in that embodiment. The claims recite proving “inputs” not specifically “prompts”, which is a lot narrower than mere inputs. Thus, applicant’s argument that the specification describes submitting the same image together with prompts to multiple AI model systems doesn’t distinguish the claimed subject matter from McKay where the broader claim language doesn’t require such prompts.] and each system independently provides its own label or annotation for the same image [which is also performed by McKay references in Para. 122 and 125. Also, Para. 59 states “Based on the labels output for the data item by one or more labelers 110, the workflow can output a final labeled result”]. The label is then determined based on comparing the multiple responses [ Examiner’s response: the claims, however, do not recite comparing the multiple responses. Rather, the claims recite “determining, based on the multiple responses, a label for the first image”. The phrase “based on” is broader than “comparing” and does not require a particular consensus, comparison or conflict resolution operation among the multiple response. Mckay’s aggregation of multiple labeling results into an overall labeled output therefore satisfies the claimed relationship absent further claim language requiring a particular manner of determining the label]. This multi-model consensus approach is fundamentally different from McKay's category-splitting architecture, in which each labeler answers a different question about what is in the image.” [Examiner’s response: Applicatn’s argument characterizes the claimed method as requiring a “multi-model consensus approach” but the claims do not expressly, as written now, require consensus among the AI models, do not require the AI models to answer the same question, and do not require the response to be compared with one another. Instead, the claims broadly require multiple AI model systems, multiple responses, and a determination of a label for the first image based on those responses. McKay’s disclosure of multiple ML labelers producing respective labeling results and aggregating those results to form an overall labeled output therefore reads on the claim limitations at issue. The fact that Mckayh’s constituent labelers may be configured to detect different categories, object, parts and features doesn’t distinguish McKay from the claims because the claims do not require the multiple AI models to perform identical labeling tasks].
Applicant also argued the examiner's rationale for combining McKay and Shekhar is that the modification enables the system to implement policy based active learning that expressly improves training efficiency by selecting the most informative images for labeling and retraining, thereby reducing human labeling effort while improving model accuracy. However, this rationale does not address how or why a person of ordinary skill in the art would have been motivated to modify McKay's category-splitting labeler architecture to arrive at the claimed multi-model consensus labeling system. The rationale speaks only to the motivation to use active learning techniques (selecting informative samples), which is Shekhar's contribution, but does not provide any reasoning for why a person of ordinary skill would modify McKay to generate multiple inputs to multiple AI model systems, provide the same image to multiple AI systems, and determine a label based on multiple responses from those systems. The Examiner's rationale amounts to no more than a conclusory statement regarding the benefit of active learning. [ Examiner’s response: Applicant’s argument is not persuasive because it mischaracterizes both the proposed combination and the respective teachings for which Mckay and Shekhar are relied upon. Rather, as explained above, McKay itself teaches the use of multiple labelers, receiving respective labeling results from the labelers, and aggregating those results to form an overall labeled output. Thus, the features concerning multiple labelers and the use of their respective outputs in determining an overall label are supplied by Mckay, not by the proposed medication based on Shekhar.
Applicant’s argument further relies on requirements that are not recited in the claims. In particular, applicant characterizes the claimed invention as requiring a “multi-model consensus labeling system” in which multiple independent AI models provide independent assessments of the same labeling question and a final label is determined through comparison or consensus among those assessments. However, as discussed above, the claims do not require “consensus” do not require each AI model to independently answer the same labeling question, and do not require comparison of the multiple responses. The claims instead broadly recite receiving multiple responses from multiple AI model systems and determining, based on the multiple responses, a label for the first image. Accordingly, the proposed combination need not provide a motivation for modifying McKay to satisfy limitations that are not actually recited in the claims.
Shekhar is relied upon for the feature missing from McKay, namely the active-learning feature mapped in the rejection. The relevant question is therefore whether one of ordinary skill in the art would have had reason to incorporate Shekhar’s active learning technique into McKay’s labeling System. As set forth in the rejection, Shekhar teaches policy based active learning that selects informative images for labeling and retraining. Incorporating such a technique into Mckay’s labeling workflow would have predictably permitted informative samples to be selected for labeling and subsequent model training, thereby reducing unnecessary labeling effort and improving the efficiency of training the machine learning models. Thus, the stated rationale is directly related to the particular teaching for which Shkhar is relied upon and provides an articulated reason for incorporating that teaching into Mckay.
Accordingly, applicant’s hindsight argument is not persuasive. The proposed combination doesn’t depend upon knowledge obtained from applicant’s disclosure to supply the claimed features. McKay provides the multiple-laber and aggregated output teaching relied upon in the rejection, while Shkhar provides the active learning teaching. One of ordinary skill in the art would have had reason to incorporate Shekhar’s active learning technique into McKay’s machine learning labeling workflow to improve the efficiency of selecting informative samples for labeling and model training, thereby reducing unnecessary labeling effort while improving effectiveness of the training process].
Therefore, the 103 rejection is maintained.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-5, 8-12, and 15-19 are rejected under 35 U.S.C. 103 as being unpatentable over McKay et al. (Pub. No. US 2021/0192394) in view of Shekhar et al. (Pub. No. US 20220253630).
Regarding claim 1 McKay teaches a non-transitory computer-readable medium comprising executable instructions, the executable instructions being executable by one or more processors to perform a method [Para. 14 and 233], the method comprising: receiving a set/plurality of first images [para. 5 “machine learning (ML) labeler that receives a plurality of labeling requests, each of which includes a data item to be labeled.”; Para. 59 “During execution of a graph, the same data item to be labeled (e.g., image, video, word document, or other discrete unit to be labeled) may be sent to one or more ML labeling platforms to be processed by one or more ML models 120 and to one or more human labeler computer systems 140 to be labeled by one or more human users”; for each first image in a subset of the set of first images: generating multiple inputs (requests) to multiple artificial intelligence model systems [Para. 122 “Here the splitter splits a request to label an image (or image training data) into requests to constituent ML labelers 1002a, 1002b, 1002c, 1002d, where each constituent ML labeler is trained for a particular product category”]; providing the multiple inputs (requests) and the first image to the multiple artificial intelligence model systems (ML Labelers) [ Para. 122 “here the splitter splits a request to label an image (or image training data) into requests to constituent ML labelers 1002a, 1002b, 1002c, 1002d, where each constituent ML labeler is trained for a particular product category” and “splitter 1004 routes the labeling request to i) labeler 1002a to label the image with any tools that labeler 1002a detects in the image, ii) labeler 1002b to label the image with any vehicles that labeler 1002b detects in the image, iii) labeler 1002c to label the image with any clothing items that labeler 1002c detects in the image, and iv) labeler 1002d to label the image with any food items that labeler 1002d detects in the image”]; receiving multiple responses (labeling results) from the multiple artificial intelligence model systems [Para. 125];
determining, based on the multiple responses (labels), a label (final) for the first image [Para. 59 “Based on the labels output for the data item by one or more labelers 110, the workflow can output a final labeled result”]; training a computer vision model (ML model) based on the model training data set [Para. 5 “the first portion of the augmented results are provided to an experiment coordinator, which iteratively trains the ML model using this portion of the augmented results”].
McKay doesn’t explicitly teach the rest of claim limitations.
However, Shekhar teaches adding (including) the first (sample) image and the label (annotations) to a model training data set [Para. “The training set includes the newly annotated or labeled sample images, in which the annotations have been verified or corrected at previous operation 215” it’s clear that the image and tags/annotations/labels are included in the training data]; training a computer vision model (object detection network) based on the model training data set [Para. 52 “The training set is used to retrain the object detection network.”. Para 2 “In the field of computer vision, object detection refers to the task of identifying objects or object instances from a digital photograph”]; receiving a second image [Para. 30 “the object detection apparatus 110 receives an image including one or more of instances of an object”; Para. 89 “the system receives an image including a set of instances of an object.”]; applying the computer vision model (object detection network) to the second image [Para. 49 and Para. 92 “the system generates annotation data for the image using an object detection network that is trained at least in part together with a policy network that selects predicted output from the object detection network for use in training the object detection network”]; and receiving an output (identify instances) from the computer vision model [Para. 95].
It would have been obvious to one of ordinary skill in the art before the effective filing date to modify McKay by incorporating the missing features as correctly mapped above, feature as taught by Shekhar; because the modification enables the system implement policy based active learning expressly improves training efficiency by selecting the most informative images for labeling and retraining, thereby reducing human labeling effort while improving model accuracy.
Claim 8 is rejected for the same reasons as claim 1. Furthermore, McKay teaches having a method to perform all claim limitations [Abstract and Para. 5].
Claim 15 is rejected for the same reason as claim 1. Furthermore, McKay teaches having a system comprising at least one processor and memory containing executable instructions, the executable instructions being executable by the at least one processor [Abstract, and Para. 4].
Regarding Claims 2, 9 and 16, McKay in view of Shekhar teaches all claim limitation above. Furthermore, Shekhar teaches wherein the computer vision model includes an image classification model (pre-trained classifier), and the output includes a class (instance) of an object in the second image [Para. 32, and fig. 6 step 610 and related description].
Regarding claims 3, 10 and 17, McKay teaches wherein the computer vision model includes an object detection model, and the output includes a location (bounding box) of an object in the second image [Para. 2].
Regarding claims 4, 11 and 18, McKay teaches wherein the multiple artificial intelligence model systems include the computer vision model (object detection network) [Para. 5 and 7].
Regarding claims 5, 12 and 19, McKay in view of Shekhar teaches all claim limitation above. Furthermore, Shekhar teaches wherein the set of first images is a first set of first images, and wherein the method further comprises: receiving a second set of first images, the second set of first images a superset of the first set of first images [fig. 1, 6, 7 and related description]; and selecting the first set of first images from the second set of first images [fig. 1, 6, 7 and related description].
Regarding claims 6, 7, 13, 14 and 20, McKay in view of Shekbar further in view of Anderson (Pub. No. US 2013/0124474) doesn’t explicitly teach the claim limitations.
Regarding claims 6, 13 and 20, McKay in view of Shekhar doesn’t explicitly
teach the rest of claim limitations.
However, Anderson teaches wherein for each first image in a subset of the set of first images, determining, based on the multiple responses, the label (match code) for the first image includes performing one or more of a strict comparison, a fuzzy comparison (fuzzy match), and a semantic comparison (semantic for quality) of the multiple responses and determining, based on the performance (similar score) of one or more of the strict comparisons, the fuzzy comparison, and the semantic comparison of the multiple responses, the label for the first image [Para. 226 “Short field-values consisting of one or two characters may often only be compared for equality as there may be no basis for distinguishing error from intent.”; Para. 96 “Identifying such false positives is useful as they represent tokens paired on the basis of similarity that should not be paired on the basis of semantic meaning”; Para. 106 “match quality states within a match code for a pair of compared field values might include "exact match" if the values were identical or "fuzzy match" if the similarity score were greater than a fuzzy match threshold”].
It would have been obvious to one of ordinary skill in the art before the effective filing date to modify McKay in view of Shekhar by incorporating the missing features as correctly mapped above, feature as taught by Anderson; because the modification enables the system to improve the scalability and speed of large-scale record clustering by reducing expensive comparisons using search-based candidate selection and enabling efficient parallel processing.
Regarding claims 7 and 14, McKay in view of Shekbar doesn’t explicitly teach the claim limitations.
However, Anderson teaches comprising determining that the performance (best score) one or more of the strict comparisons, the fuzzy comparison, and the semantic comparison of the multiple responses exceeds a threshold (match threshold) [Para. 236 “If the best score is above a match threshold, the query record is added to the corresponding cluster”].
It would have been obvious to one of ordinary skill in the art before the effective filing date to modify McKay in view of Shekhar by incorporating the missing features as correctly mapped above, feature as taught by Anderson; because the modification enables the system to improve the scalability and speed of large-scale record clustering by reducing expensive comparisons using search-based candidate selection and enabling efficient parallel processing.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SOLOMON G BEZUAYEHU whose telephone number is (571)270-7452. The examiner can normally be reached on Monday-Friday 10 AM-8 PM.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Oneal Mistry can be reached on 313-446-4912. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 888-786-0101 (IN USA OR CANADA) or 571-272-4000.
/SOLOMON G BEZUAYEHU/
Primary Examiner, Art Unit 2666