DETAILED ACTION
This action is written in response to the remarks and amendments dated April 23, 2026. This action is made final. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
All outstanding rejections under §112(b) are withdrawn. The Applicants argue that the previous art of record does not anticipate or render obvious the claims as currently amended. The Examiner provides updated prior art rejections below necessitated by the current amendments.
Subject Matter Eligibility
In determining whether the claims are subject matter eligible, the examiner has considered and applied the 2019 USPTO Patent Eligibility Guidelines, as well as guidance in the MPEP chapter 2106. The examiner finds that the independent claims are directed to the practical application of improving the labeling of a machine-learning perception system, ie object recognition.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103(a) which forms the basis for all obviousness rejections set forth in this Office action:
(a) A patent may not be obtained though the invention is not identically disclosed or described as set forth in section 102 of this title, if the differences between the subject matter sought to be patented and the prior art are such that the subject matter as a whole would have been obvious at the time the invention was made to a person having ordinary skill in the art to which said subject matter pertains. Patentability shall not be negatived by the manner in which the invention was made.
The following are the references relied upon in the rejections below:
Mishra (US 2017/0103267 A1)
Yin (Yin, Dong, et al. "A fourier perspective on model robustness in computer vision." Advances in Neural Information Processing Systems 32 (2019).)
Claims 1-3, 5-10 and 16-17 are rejected under 35 U.S.C. 103 as being unpatentable over Mishra and Yin.
Regarding claims 1, 16 and 17, Mishra discloses a computer-implemented method (and a related computer system and non-transitory computer-readable medium) for improving the labeling performance of a machine-learned (ML) perception system on an input dataset, the method comprising processing, by a computing device, the input dataset through the perception system; …
[0031] “It has been recognized that video imaging systems and platforms which analyze image and video content for detection and feature extraction can accumulate significant amounts of data suitable for training and learning analytics that can be leveraged to improve over time the classifiers used to perform the detection and feature extraction by employing a larger search space and generating additional and more complex classifiers through distributed processing.”
isolating, by a computing device, at least one scene from the input dataset that includes the one or more perception weaknesses;
[0034] “A human user can validate the output and provide negative feedback to the algorithm when the algorithm performs poorly. This feedback is stored in a dataset and the classifier is retrained using the learning platform 12. In supervised learning, the goal is typically to label scenic elements and perform object detection.”
relabeling the one or more perception weaknesses in the at least one scene to obtain at least one relabeled scene; and
[0034] “A human user can validate the output and provide negative feedback to the algorithm when the algorithm performs poorly.”
[0074] “ One method is to use an independent observer to classify objects independently from the trained classifier, or to use two trained classifiers with independent implementation and algorithms. The observer can generate validation data which can be compared against the output from the trained classifier.”
retraining, by a computing device, the perception system with the at least one relabeled scene;
[0034] “A human user can validate the output and provide negative feedback to the algorithm when the algorithm performs poorly. This feedback is stored in a dataset and the classifier is retrained using the learning platform 12. In supervised learning, the goal is typically to label scenic elements and perform object detection.” (Emphasis added.)
whereby the labeling performance of the ML perception system is improved.
[The Examiner notes that this limitation is an intended result.]
[0032] “FIG. 1(a) illustrates a video analysis environment 10 which includes a learning platform 12 to generate and/or improve a set of classifiers 14 used in analyzing video content, e.g., for detection and extraction of features within image frames of a video.“
Yin discloses the following further limitations which Mishra does not disclose:
creating, by the computing device, an augmented dataset augmented with Gaussian noise;
P. 2, sec. 2, “We define Gaussian data augmentation with parameter σ as the following operation: In each iteration, we add i.i.d. Gaussian noise N(0; σ~2) to every pixel in all the images in the training batch, where σ~ is chosen uniformly at random from [0; σ].”
processing, by the computing device, the augmented dataset through the perception system;
P. 4, sec. 4.1, “Ford et al. [10] investigated the robustness of three models on CIFAR-10-C: a naturally trained model, a model trained by Gaussian data augmentation, and an adversarially trained model. It was observed that Gaussian data augmentation and adversarial training improve robustness to all noise and many of the blurring corruptions, while degrading robustness to fog and contrast.”
comparing, by the computing device, results of the input dataset to results of the augmented database to identify one or more perception weaknesses;
P. 5, “First, we test model sensitivity to perturbations along each Fourier basis vector. Results on CIFAR-10 are shown in Figure 3. The difference between the three models is striking. The naturally trained model is highly sensitive to additive perturbations in all but the lowest frequencies, while Gaussian data augmentation and adversarial training both dramatically improve robustness in the higher frequencies.”
P. 6, fig. 5 (reproduced below).
PNG
media_image1.png
344
808
media_image1.png
Greyscale
At the time of filing, it would have been obvious to a person of ordinary skill to apply the technique disclosed by Yin for data augmentation using Gaussian noise to the image processing system of Mishra because—as illustrated by the results in fig. 5 above—this can lead to improved model performance for image classification tasks.
Regarding independent claims 16 and 17, Mishra also discloses the generic computer components recited therein. (See eg [0006] ‘processors’ and [0069] ‘memory’.)
Regarding claim 2, Mishra discloses the further limitation wherein the step of identifying one or more perception weaknesses is performed by a defect detection engine.
[0083] “video analysis module 342”, further discussed at [0085]-[0086].
Regarding claim 3, Mishra discloses the further limitation wherein the one or more perception weaknesses comprise labeled objects in bounding boxes.
See fig. 9 (reproduced below) and 10.
PNG
media_image2.png
346
984
media_image2.png
Greyscale
[0067] “FIG. 9 illustrates how a classifier determines if an object of interest is located within a context. For this case, measurements are extracted from an image frame extracted from a traffic intersection video. The sliding window (box in leftmost image) is centered at each pixel within the frame and the pixel intensities are extracted (top 2nd column).”
Regarding claim 5, Mishra discloses the further limitation wherein the step of relabeling the at least one scene comprises sending the at least one scene to be labeled by a human.
[0032] “The results of the feature analysis stage 22 are validated in a validation stage 24, which can be performed manually by analysts responsible for accepting or rejecting the results of the feature analyses (e.g., to identify misclassified objects).”
Regarding claim 6, Mishra discloses the further limitation wherein the input dataset is comprised of one or more of baseline data, augmented data, ground truth data, and simulated data.
[The Examiner notes that this is a Markush group.]
‘baseline data’ :: “training data and labelled ground truth in a supervised learning mode”.
Regarding claim 7, Yin discloses the further limitation comprising filtering, by a computing device, the results of the processing step to concentrate on specific perception weaknesses.
P. 5, “First, we test model sensitivity to perturbations along each Fourier basis vector. Results on CIFAR-10 are shown in Figure 3. The difference between the three models is striking. The naturally trained model is highly sensitive to additive perturbations in all but the lowest frequencies, while Gaussian data augmentation and adversarial training both dramatically improve robustness in the higher frequencies.” (Emphasis added.)
Regarding claim 8, Mishra discloses the further limitation comprising repeating the method and comparing the perception weaknesses with the perception weaknesses identified in the processing step that occurred prior to the relabeling to determine an amount of improvement in the performance of the perception system.
[0069] “From the distributed processing, a learning and training analysis stage 112 is performed using the results of each partially computed algorithm in aggregation to iteratively improve the classifier. For example, the aggregate results could be used to determine the best feature to distinguish two categories, then that feature could be suppressed to find the next best feature, and the process can continue iteratively until acceptable at least one termination criterion is achieved.”
Regarding claim 9, Mishra discloses the further limitation comprising repeating the steps of processing, identifying, isolating, relabeling and retraining until retraining the ML perception system with the newly labeled scenes does not meaningfully improve the labeling performance.
[0069] “From the distributed processing, a learning and training analysis stage 112 is performed using the results of each partially computed algorithm in aggregation to iteratively improve the classifier. For example, the aggregate results could be used to determine the best feature to distinguish two categories, then that feature could be suppressed to find the next best feature, and the process can continue iteratively until acceptable at least one termination criterion is achieved.” (Emphasis added.)
Regarding claim 10, Mishra discloses the further limitation comprising reengineering the ML perception system when the labeling performance does not meaningfully improve with the newly labeled scenes.
[0069] “From the distributed processing, a learning and training analysis stage 112 is performed using the results of each partially computed algorithm in aggregation to iteratively improve the classifier. For example, the aggregate results could be used to determine the best feature to distinguish two categories, then that feature could be suppressed to find the next best feature, and the process can continue iteratively until acceptable at least one termination criterion is achieved.”
Additional Relevant Prior Art
The following references were identified by the Examiner as being relevant to the disclosed invention, but are not relied upon in any particular prior art rejection:
Wang discloses a computer vision system comprising techniques for image area classification. (US 2020/0242424 A1)
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Vincent Gonzales whose telephone number is (571) 270-3837. The examiner can normally be reached on Monday-Friday 7 a.m. to 4 p.m. MT. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Miranda Huang, can be reached at (571) 270-7092.
/Vincent Gonzales/Primary Examiner, Art Unit 2124