DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Receipt is acknowledged of certified copies of papers submitted under 35 U.S.C. 119(a)-(d), which papers have been placed of record in the file.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 11/27/2024 and 06/24/2025 have been considered by the examiner.
Claim Objections
Claims 1, 3, 4, 6, 13, and 16 are objected to because of the following informalities:
In claim 1, lines 3-19, the lettering/alphabetical bullets (a-f) should be removed in order to have the claim written in a proper format with only indentations to indicate sub-clauses, and semi-colons used to separate the limitations.
In claim 2, line 1, “wherein c. comprises” should be corrected to “wherein providing an ignore mask comprises” in order to align with the corrections of claim 1.
In claim 3, line 1, “wherein c. comprises” should be corrected to “wherein providing an ignore mask comprises” in order to align with the corrections of claim 1.
In claim 4, line 1, “wherein c. comprises” should be corrected to “wherein providing an ignore mask comprises” in order to align with the corrections of claim 1.
In claim 12, line 1, “wherein b. comprises” should be corrected to “wherein assigning a groundtruth annotation to each pixel comprises” in order to align with the corrections of claim 1.
In claim 13, line 1, “wherein d.-f. are repeated” should be corrected to “wherein receiving a prediction value from the machine learning model for each pixel, determining a loss value based on the prediction value and ground truth annotation for each pixel which has no ignore flag and training the machine learning model on a basis of the loss value, and ignoring the prediction value for pixels which have ignore flags, are repeated” in order to remove reference to the alphabetical bullets and align with the corrections of claim 1.
Appropriate correction is required.
Double Patenting
The non-statutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A non-statutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on non-statutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a non-statutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claims 1 and 5-20 are rejected on the ground of non-statutory double patenting as being unpatentable over claims 1 and 8-18 of co-pending Application No. 19/091357 in view of HUANG et al. (US 20170147905 A1).
Although the claims at issue are not identical, they are not patentably distinct from each other because the instant application and the conflicting co-pending application are claiming common subject matter, as follows:
This Application No. 18/963,110
Co-pending Application No. 19/091357
Claim 1: A method of training a machine learning model to identify image features, the method comprising:
a. providing training image data, the training image data comprising a plurality of pixels;
c. providing an ignore mask comprising a set of ignore flags, wherein each ignore flag relates to a respective one of the pixels and each ignore flag provides an indication that the pixel should be ignored;
d. for each pixel, receiving a prediction value from the machine learning model, wherein each prediction value provides an indication of a probability of the pixel corresponding with an image feature;
e. for each pixel which has no ignore flag, determining a loss value based on the prediction value
and training the machine learning model on a basis of the loss value;
and f. for each pixel which has an ignore flag, ignoring the prediction value for that pixel so that it is not used to train the machine learning model.
Claim 1: the method comprising:
b. assigning a groundtruth annotation to each pixel, wherein each groundtruth annotation relates to a respective one of the pixels and each groundtruth annotation indicates whether or not that the pixel corresponds with an image feature;
e. for each pixel which has no ignore flag, determining a loss value based on the prediction value and groundtruth annotation for that pixel,
Claim 1: A method of training a machine learning model to identify image features, the method comprising:
a. providing training image data, the training image data comprising a plurality of pixels;
d. for each pixel, receiving a prediction value from the machine learning model, wherein each prediction value provides an indication of a probability of the pixel corresponding with an image feature;
e. for each pixel which coincides with the groundtruth mask and does not coincide with an ignore mask, determining a loss value based on the prediction value for that pixel
and training the machine learning model on a basis of the loss value;
and f. for each pixel which lies between the inner and outer edges of an ignore mask, ignoring the prediction value for that pixel so that it is not used to train the machine learning model.
Claim 8: further comprising assigning a groundtruth annotation to each pixel, wherein each groundtruth annotation relates to a respective one of the pixels and each groundtruth annotation indicates whether or not that the pixel corresponds with an image feature;
Claim 8: and for each pixel which does not coincide with an ignore mask, a loss value is determined based on the prediction value and groundtruth annotation for that pixel.
Although co-pending application 19/091357 claim 1 teaches wherein a method of training a machine learning model to identify image features, the method comprising: a. providing training image data, the training image data comprising a plurality of pixels; d. for each pixel, receiving a prediction value from the machine learning model, wherein each prediction value provides an indication of a probability of the pixel corresponding with an image feature; e. for each pixel which has no ignore flag, determining a loss value based on the prediction value and training the machine learning model on a basis of the loss value; and f. for each pixel which has an ignore flag, ignoring the prediction value for that pixel so that it is not used to train the machine learning model. Co-pending application 19/091357, claim 1, as stated in the table above with respect to claim 1 of the present application 18/963,110, fails to clearly disclose c. providing an ignore mask comprising a set of ignore flags, wherein each ignore flag relates to a respective one of the pixels and each ignore flag provides an indication that the pixel should be ignored.
However, HUANG et al. (US 20170147905 A1) explicitly teaches c. providing an ignore mask comprising a set of ignore flags (Fig. 5A, Paragraph [0057] – HUANG discloses for a gray zone, i.e., the area between positive and negative regions, the loss weight is set to 0, and for each pixel labeled non-positive in the output coordinate space, an ignore flag fign is set to 1 if a pixel with positive label within rnear=2 pixel length exists. See also Paragraph [0059].), wherein each ignore flag relates to a respective one of the pixels and each ignore flag provides an indication that the pixel should be ignored (Fig. 5A, Paragraph [0057] – HUANG discloses for a gray zone, i.e., the area between positive and negative regions, the loss weight is set to 0, and for each pixel labeled non-positive in the output coordinate space, an ignore flag fign is set to 1 if a pixel with positive label within rnear=2 pixel length exists.);
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of the co-pending application 19/091357, claim 1, of having wherein a method of training a machine learning model to identify image features, the method comprising: a. providing training image data, the training image data comprising a plurality of pixels; d. for each pixel, receiving a prediction value from the machine learning model, wherein each prediction value provides an indication of a probability of the pixel corresponding with an image feature; e. for each pixel which has no ignore flag, determining a loss value based on the prediction value and training the machine learning model on a basis of the loss value; and f. for each pixel which has an ignore flag, ignoring the prediction value for that pixel so that it is not used to train the machine learning model, with the teachings of HUANG et al. (US 20170147905 A1) of having c. providing an ignore mask comprising a set of ignore flags, wherein each ignore flag relates to a respective one of the pixels and each ignore flag provides an indication that the pixel should be ignored.
Wherein having co-pending application 19/091357, claim 1, having c. providing an ignore mask comprising a set of ignore flags, wherein each ignore flag relates to a respective one of the pixels and each ignore flag provides an indication that the pixel should be ignored.
The motivation behind the modification would have been to obtain an enhanced method of training a machine learning model to identify image features that can be optimized directly for detection and can be easily improved by incorporating landmark information.
The further limitations of the present application’s 18/963,110 dependent claims 5-20 are similar to the cited claims 1 and 8-18 of co-pending application 19/091357 as indicated below:
This Application No. 18/963,110
Co-pending Application No. 19/091357
Claim 5: A method according to claim 2, wherein the training image data comprises one or more images of the object.
Claim 6: A method according to claim 2, wherein the training image data comprises a series of images of the object which each contain the same feature viewed from a different viewing angle.
Claim 7: A method according to claim 6, further comprising generating the training image by imaging the object from a series of different viewing angles.
Claim 8: A method according to claim 7, wherein the object is imaged with light.
Claim 9: A method according to claim 1, wherein the training image data comprises a series of images of an object which each contain the same feature viewed from a different viewing angle.
Claim 10: A method according to claim 9, further comprising generating the training image by imaging the object from a series of different viewing angles.
Claim 11: A method according to claim 10, wherein the object is imaged with light.
Claim 12: A method according to claim 1, wherein b. comprises displaying the training image data to a human annotator; and receiving the groundtruth mask via inputs from the human annotator,
the groundtruth mask providing an indication that a region of the training image data contains an image feature.
Claim 13: A method according to claim 1, wherein d.-f. are repeated, each repeat comprising a respective training epoch.
Claim 14: A method according to claim 1, wherein the image feature comprises a surface defect.
Claim 15: A method according to claim 14, wherein the image feature comprises a surface defect of an aircraft.
Claim 16: A method according to claim 14, wherein the image feature comprises a dent.
Claim 17: A method according to claim 1, wherein the loss value is determined by the algorithm:
PNG
media_image1.png
93
466
media_image1.png
Greyscale
wherein yk is a groundtruth annotation for that pixel; pk is a prediction value for that pixel, a pixel which corresponds with an image feature has a groundtruth annotation of yk of 1, and a pixel which does not correspond with an image feature has a groundtruth annotation of yk of 0.
Claim 18: A method according to claim 1, wherein after the machine learning model has been trained, it is used to segment an image in an inference phase.
Claim 19: A computer system configured to train a machine learning model by the method of claim 1.
Claim 20: A computer software configured to train a machine learning model by the method of claim 1.
Claim 10: wherein the training image data comprises a series of images of an object [wherein a series of images is one or more images] which each contain the same feature viewed from a different viewing angle.
Claim 10: wherein the training image data comprises a series of images of an object which each contain the same feature viewed from a different viewing angle.
Claim 11: further comprising generating the training image data by imaging the object from a series of different viewing angles.
Claim 12: wherein the object is imaged with light.
Claim 10: wherein the training image data comprises a series of images of an object which each contain the same feature viewed from a different viewing angle.
Claim 11: further comprising generating the training image data by imaging the object from a series of different viewing angles.
Claim 12: wherein the object is imaged with light.
Claim 13: wherein b. comprises displaying the training image data to a human annotator and receiving the groundtruth mask via inputs from the human annotator.
Claim 1: providing a groundtruth mask which provides an indication that a region of the training image data contains an image feature;
Claim 14: wherein d.-f. are repeated, each repeat comprising a respective training epoch.
Claim 15: wherein the image feature comprises a surface defect, optionally wherein the image feature comprises a surface defect of an aircraft, and further optionally wherein the image feature comprises a dent.
Claim 15: wherein the image feature comprises a surface defect, optionally wherein the image feature comprises a surface defect of an aircraft, and further optionally wherein the image feature comprises a dent.
Claim 15: wherein the image feature comprises a surface defect, optionally wherein the image feature comprises a surface defect of an aircraft, and further optionally wherein the image feature comprises a dent.
Claim 9: wherein the loss value is determined by an algorithm of:
PNG
media_image2.png
124
426
media_image2.png
Greyscale
wherein yk is the groundtruth annotation for that pixel; pk is the prediction value for that pixel, a pixel which corresponds with an image feature has a groundtruth annotation of yk of 1, and a pixel which does not correspond with an image feature has a groundtruth annotation of yk of 0.
Claim 16: wherein after the machine learning model has been trained, it is used to segment an image in an inference phase.
Claim 17: A computer system configured to train a machine learning model by the method of claim 1.
Claim 18: Computer software configured to train a machine learning model by the method of claim 1.
Claims 5-20 of the present application no. 18/963,110 contain the same limitations as claims 1 and 8-18 of co-pending application no. 19/091357 as cited above, and are therefore rejected on the ground of non-statutory double patenting as being unpatentable over claims 1 and 8-18 of co-pending Application No. 19/091357.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or
composition of matter, or any new and useful improvement thereof, may obtain a patent
therefor, subject to the conditions and requirements of this title.
Claim 20 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter.
Claim 20 is drawn to a “computer software” per se, therefore, fails to fall within a statutory category of invention, since applicant’s specification does not define the term “computer software.”
A claim directed to a computer program itself is non-statutory because it is not:
A process occurring as a result of executing the program, or
A machine programmed to operate in accordance with the program, or
A manufacture structurally and functionally interconnected with the program in a manner which enable the program to act as a computer component and realize its functionality, or
A composition of matter.
See MPEP § 2106.01. Data structures not claimed as embodied in computer readable media are descriptive material per se and are not statutory because they are not capable of causing functional change in the computer. See, e.g., Warmerdam, 33 F.3d at 1361, 31 USPQ2d at 1760 (claim to a data structure per se held non-statutory). Such claimed data structures do not define any structural and functional interrelationships between the data structure and other claimed aspects of the invention, which permit the data structure's functionality to be realized. In contrast, a claimed computer readable medium encoded with a data structure defines structural and functional interrelationships between the data structure and the computer software and hardware components which permit the data structure's functionality to be realized, and is thus statutory. Similarly, computer programs claimed as computer listings per se, i.e., the descriptions or expressions of the programs are not physical “things.” They are neither computer components nor statutory processes, as they are not “acts” being performed. Such claimed computer programs do not define any structural and functional interrelationships between the computer program and other claimed elements of a computer, which permit the computer program's functionality to be realized.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed
invention is not identically disclosed as set forth in section 102 of this title, if the
differences between the claimed invention and the prior art are such that the claimed
invention as a whole would have been obvious before the effective filing date of the
claimed invention to a person having ordinary skill in the art to which the claimed
invention pertains. Patentability shall not be negated by the manner in which the
invention was made.
Claims 1, 2, 5, 12-14, and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over BA (US 20230307132 A1), hereinafter referenced as BA in view of HUANG (US 20170147905 A1), hereinafter referenced as HUANG.
Regarding claim 1, BA teaches a method of training a machine learning model to identify image features (Fig. 6, Paragraph [0054] – BA discloses apparatuses, methods, and systems described herein may be embodied in a variety of other forms. Paragraph [0055] – BA discloses various embodiments of the invention relate to training and/or using a deep learning model for interactive segmentation digital pathology images to identify different regions within the image.), the method comprising:
a. providing training image data (Fig. 6, Paragraph [0060] – BA discloses deep learning model can be trained using a multi-class training data set that has images and corresponding annotations from one or more domains.),
the training image data comprising a plurality of pixels (Fig. 6, Paragraph [0072] – BA discloses deep learning model can be trained using the training data described here (e.g., that includes multi-class images and that may include ground-truth masks and/or click annotations), such that it learns how to segment a particular number of classes (e.g., two target classes and one background class). The training data enforce network learning of effective representations to match the segmentation prediction that correspond to specific pixels indicated via click annotations (or to specific regions within a ground-truth mask).);
b. assigning a groundtruth annotation to each pixel (Fig. 6, Paragraph [0064] – BA discloses click annotations are automatically identified for training images. For example, some training images may be associated with a ground-truth mask that indicates (or that can be used to determine) to which label each pixel is to be assigned.),
wherein each groundtruth annotation relates to a respective one of the pixels and each groundtruth annotation indicates whether or not that the pixel corresponds with an image feature (Fig. 6, Paragraph [0072] – BA discloses training data enforce network learning of effective representations to match the segmentation prediction that correspond to specific pixels indicated via click annotations (or to specific regions within a ground-truth mask). For example, such a model learns to segment any targeting regions pointed to by these pixel locations. It is thus trained to group together the image pixels of unified labels (i.e., similar/identical network representations) as the “annotated” image pixels, no matter what exact underlying semantic meaning they have.);
Although BA further teaches d. for each pixel, receiving a prediction value from the machine learning model, wherein each prediction value provides an indication of a probability of the pixel corresponding with an image feature (Fig. 6, Paragraph [0074] – BA discloses click annotations may identify—for each of a set of classes—one or more pixels in the image that do, or that do not, correspond to the class. The deep-learning model may use the click annotations and the learned model parameters to predict which portions of the images correspond to each of the set of classes.);
BA fails to explicitly teach c. providing an ignore mask comprising a set of ignore flags, wherein each ignore flag relates to a respective one of the pixels and each ignore flag provides an indication that the pixel should be ignored; e. for each pixel which has no ignore flag, determining a loss value based on the prediction value and groundtruth annotation for that pixel, and training the machine learning model on a basis of the loss value; and f. for each pixel which has an ignore flag, ignoring the prediction value for that pixel so that it is not used to train the machine learning model.
However, HUANG explicitly teaches c. providing an ignore mask comprising a set of ignore flags (Fig. 5A, Paragraph [0057] – HUANG discloses for a gray zone, i.e., the area between positive and negative regions, the loss weight is set to 0, and for each pixel labeled non-positive in the output coordinate space, an ignore flag fign is set to 1 if a pixel with positive label within rnear=2 pixel length exists. See also Paragraph [0059].),
wherein each ignore flag relates to a respective one of the pixels and each ignore flag provides an indication that the pixel should be ignored (Fig. 5A, Paragraph [0057] – HUANG discloses for a gray zone, i.e., the area between positive and negative regions, the loss weight is set to 0, and for each pixel labeled non-positive in the output coordinate space, an ignore flag fign is set to 1 if a pixel with positive label within rnear=2 pixel length exists.);
e. for each pixel which has no ignore flag (Fig. 5A, Paragraph [0057] – HUANG discloses for a gray zone, i.e., the area between positive and negative regions, the loss weight is set to 0, and for each pixel labeled non-positive in the output coordinate space, an ignore flag fign is set to 1 if a pixel with positive label within rnear=2 pixel length exists.),
determining a loss value based on the prediction value and groundtruth annotation for that pixel (Fig. 5A, Paragraph [0085] – HUANG discloses the classification loss and a bounding box regression loss are combined to obtain a multi-task loss. Paragraph [0053] – HUANG discloses independent branch 360 outputs the confidence score ŷ (per pixel in the output map) of being a target object [See also Eq. 1]. Paragraph [0060-61] – HUANG discloses the multi-task loss may be represented as,
PNG
media_image3.png
82
645
media_image3.png
Greyscale
where θ is the set of parameters in the network, and the Iverson bracket function [y*i>0] is activated only if the ground truth score y*i is positive.),
and training the machine learning model on a basis of the loss value (Fig. 5A, Paragraph [0078] – HUANG discloses FIG. 5A is a flowchart illustrating a method to train an end-to-end multi-task object detection network using object landmarks, according to various embodiments of the present disclosure. The process for training the object detection network begins at step 502 by inputting an image into the end-to-end multi-task object detection network to obtain an output feature map that comprises a confidence score and a bounding box for each pixel in the output feature map.); and
f. for each pixel which has an ignore flag (Fig. 5A, Paragraph [0057] – HUANG discloses for a gray zone, i.e., the area between positive and negative regions, the loss weight is set to 0, and for each pixel labeled non-positive in the output coordinate space, an ignore flag fign is set to 1 if a pixel with positive label within rnear=2 pixel length exists.),
ignoring the prediction value for that pixel so that it is not used to train the machine learning model (Fig. 5A, Paragraph [0078] – HUANG discloses FIG. 5A is a flowchart illustrating a method to train an end-to-end multi-task object detection network using object landmarks, according to various embodiments of the present disclosure. Paragraph [0079] – HUANG discloses a first set of pixels in the output feature map that fall in a margin region between a positive region and a negative region is ignored.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of BA of having a method of training a machine learning model to identify image features, the method comprising: a. providing training image data, the training image data comprising a plurality of pixels; b. assigning a groundtruth annotation to each pixel, wherein each groundtruth annotation relates to a respective one of the pixels and each groundtruth annotation indicates whether or not that the pixel corresponds with an image feature; d. for each pixel, receiving a prediction value from the machine learning model, wherein each prediction value provides an indication of a probability of the pixel corresponding with an image feature; with the teachings of HUANG having c. providing an ignore mask comprising a set of ignore flags, wherein each ignore flag relates to a respective one of the pixels and each ignore flag provides an indication that the pixel should be ignored; e. for each pixel which has no ignore flag, determining a loss value based on the prediction value and groundtruth annotation for that pixel, and training the machine learning model on a basis of the loss value; and f. for each pixel which has an ignore flag, ignoring the prediction value for that pixel so that it is not used to train the machine learning model.
Wherein having BA’s method of training a machine learning model to identify image features wherein having c. providing an ignore mask comprising a set of ignore flags, wherein each ignore flag relates to a respective one of the pixels and each ignore flag provides an indication that the pixel should be ignored; e. for each pixel which has no ignore flag, determining a loss value based on the prediction value and groundtruth annotation for that pixel, and training the machine learning model on a basis of the loss value; and f. for each pixel which has an ignore flag, ignoring the prediction value for that pixel so that it is not used to train the machine learning model.
The motivation behind the modification would have been to obtain an enhanced method of training a machine learning model to identify image features with improved accuracy and optimized detection performance, since both BA and HUANG relate to arrangements for image or video recognition/understanding, wherein BA discloses relate to techniques that facilitate improving the accuracy of cell classifications by analyzing data generated by two separate machine-learning models, and HUANG relates to computer processing and, more particularly, to systems, devices, and methods for end-to-end object recognition in computer vision applications wherein embodiments of the present disclosure can be optimized directly for detection and can be easily improved by incorporating landmark information. Please see BA (US 20230307132 A1), Paragraph [0098], and HUANG (US 20170147905 A1), Paragraph [0003, 0074].
Regarding claim 2, BA in view of HUANG teach a method according to claim 1,
BA fails to explicitly teach wherein c. comprises inspecting an object and generating the ignore mask on a basis of the inspection.
However, HUANG explicitly teaches wherein c. comprises inspecting an object and generating the ignore mask on a basis of the inspection (Fig. 5A, Paragraph [0078] – HUANG discloses FIG. 5A is a flowchart illustrating a method to train an end-to-end multi-task object detection network using object landmarks, according to various embodiments of the present disclosure. Paragraph [0079] – HUANG discloses at step 504, a first set of pixels in the output feature map that fall in a margin region between a positive region and a negative region is ignored.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of BA in view of HUANG of having a method of training a machine learning model to identify image features, the method comprising: a. providing training image data, the training image data comprising a plurality of pixels; b. assigning a groundtruth annotation to each pixel, wherein each groundtruth annotation relates to a respective one of the pixels and each groundtruth annotation indicates whether or not that the pixel corresponds with an image feature; c. providing an ignore mask comprising a set of ignore flags, wherein each ignore flag relates to a respective one of the pixels and each ignore flag provides an indication that the pixel should be ignored; with the teachings of HUANG having wherein c. comprises inspecting an object and generating the ignore mask on a basis of the inspection.
Wherein having BA’s method of training a machine learning model to identify image features wherein c. comprises inspecting an object and generating the ignore mask on a basis of the inspection.
The motivation behind the modification would have been to obtain an enhanced method of training a machine learning model to identify image features with improved accuracy and optimized detection performance, since both BA and HUANG relate to arrangements for image or video recognition/understanding, wherein BA discloses relate to techniques that facilitate improving the accuracy of cell classifications by analyzing data generated by two separate machine-learning models, and HUANG relates to computer processing and, more particularly, to systems, devices, and methods for end-to-end object recognition in computer vision applications wherein embodiments of the present disclosure can be optimized directly for detection and can be easily improved by incorporating landmark information. Please see BA (US 20230307132 A1), Paragraph [0098], and HUANG (US 20170147905 A1), Paragraph [0003, 0074].
Regarding claim 5, BA in view of HUANG teach a method according to claim 2,
BA further teaches wherein the training image data (Fig. 1, #110a-n called multi-class images, Paragraph [0106] – BA discloses the network includes a model training system 105 that retrieves or retrieves multi-class images 110a-n from one or more data sources.)
comprises one or more images of the object (Fig. 1, Paragraph [0106] – BA discloses each of the multi-class images 110a-n can include annotation data that indicates borders of or areas of depictions of different segments in the image. The segments in each of the multi-class images 110a-n may correspond to different types of objects or things. See also Paragraphs [0060-0061].).
Regarding claim 12, BA in view of HUANG teach a method according to claim 1,
BA further teaches wherein b. comprises displaying the training image data to a human annotator (Fig. 2, Paragraph [0136] – BA discloses part or all of the first dataset 210 may have been received via an interactive GUI from a user that is facilitating training a deep-learning model (e.g., deep-learning model 115) that is connected with the interactive GUI. See also Paragraph [0108].);
and receiving a groundtruth mask via inputs from the human annotator (Fig. 2, Paragraph [0136] – BA discloses a first dataset 210 may include a plurality of images (e.g., multi-class images 110a-n) from one or more classes (e.g., DP images, natural scene images, immunohistochemistry images, H&E images, or any other relevant images) and may include corresponding annotations (e.g., that identify click annotations, distinct segments within the images, a ground-truth map, etc.).),
the groundtruth mask providing an indication that a region of the training image data contains an image feature (Fig. 2, Paragraph [0064] – BA discloses some training images may be associated with a ground-truth mask that indicates (or that can be used to determine) to which label each pixel is to be assigned. Such ground-truth masks may change with different click annotation targets. For example, two out of multiple image regions in an input image can be used as an segmentation target and the corresponding ground-truth mask can be generated to indicate the target image regions and ignore rest of the image regions.).
Regarding claim 13, BA in view of HUANG teach a method according to claim 1,
BA fails to explicitly teach wherein d.-f. are repeated, each repeat comprising a respective training epoch.
However, HUANG explicitly teaches wherein d.-f. (Fig. 5A, steps 502-518, Paragraphs [0078-0086], see also Paragraph [0053, 0057]) are repeated,
each repeat comprising a respective training epoch (Fig. 5A, Paragraph [0078] – HUANG discloses FIG. 5A is a flowchart illustrating a method to train an end-to-end multi-task object detection network using object landmarks, according to various embodiments of the present disclosure. Paragraph [0063] – HUANG discloses in embodiments, the global learning rate starts with 0.001 and is reduced by a factor of 10 every 100,000 iterations [wherein iterations are epochs].).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of BA in view of HUANG of having a method of training a machine learning model to identify image features, the method comprising: d. for each pixel, receiving a prediction value from the machine learning model, wherein each prediction value provides an indication of a probability of the pixel corresponding with an image feature; e. for each pixel which has no ignore flag, determining a loss value based on the prediction value and groundtruth annotation for that pixel, and training the machine learning model on a basis of the loss value; and f. for each pixel which has an ignore flag, ignoring the prediction value for that pixel so that it is not used to train the machine learning model, with the teachings of HUANG having wherein d.-f. are repeated, each repeat comprising a respective training epoch.
Wherein having BA’s method of training a machine learning model to identify image features wherein d.-f. are repeated, each repeat comprising a respective training epoch.
The motivation behind the modification would have been to obtain an enhanced method of training a machine learning model to identify image features with improved accuracy and optimized detection performance, since both BA and HUANG relate to arrangements for image or video recognition/understanding, wherein BA discloses relate to techniques that facilitate improving the accuracy of cell classifications by analyzing data generated by two separate machine-learning models, and HUANG relates to computer processing and, more particularly, to systems, devices, and methods for end-to-end object recognition in computer vision applications wherein embodiments of the present disclosure can be optimized directly for detection and can be easily improved by incorporating landmark information. Please see BA (US 20230307132 A1), Paragraph [0098], and HUANG (US 20170147905 A1), Paragraph [0003, 0074].
Regarding claim 14, BA in view of HUANG teach a method according to claim 1,
BA further teaches wherein the image feature comprises a surface defect (Fig. 2, Paragraph [0056] – BA discloses the region labels may include one or more labels that convey a biological meaning (e.g., a tumor region or a stromal region), and the region labels may also include a background label that indicates that the corresponding regions are predicted not to include cells pertinent to an analysis of interest. For example, it may be inconsistent to have a classification that indicates that a cell does or does not include a biomarker when a depiction of the cell is within a background region. This inconsistency may indicate that (for example) the background region label indicates that no cells are depicted, that any depicted cells do not pertain to the analysis of interest, and/or that an artifact or image defect sufficiently obstructs visualization of signals such that exclusion of corresponding data is preferred.).
Regarding claim 18, BA in view of HUANG teach a method according to claim 1,
BA further teaches wherein after the machine learning model has been trained, it is used to segment an image in an inference phase (Fig. 1, Paragraph [0110] – BA discloses after the deep-learning model is trained, a segmentation controller 120 can use the trained deep-learning model to process each of one or more digital-pathology (DP) images 125a-n to identify different (e.g., non-overlapping or overlapping) segments in the image. See also Paragraph [0066].).
Regarding claim 19, BA in view of HUANG teach the method of claim 1,
BA further teaches a computer system configured to train a machine learning model by the method of claim 1 (Fig. 1, Paragraph [0105] – BA discloses a single physical or virtual computing system includes one or more processors and a computer readable medium with instructions that, when executed, perform actions described as being performed by multiple components in FIG. 1. For example, a single computing system may perform actions described herein as being performed by model training system 105, segmentation controller 120, GUI controller 160, and cell classifier 170.).
Regarding claim 20, BA in view of HUANG teach the method of claim 1,
BA further teaches a computer software configured to train a machine learning model by the method of claim 1 (Fig. 1, Paragraph [0194] – BA discloses embodiments of the present disclosure include a computer-program product tangibly embodied in a non-transitory machine-readable storage medium, including instructions configured to cause one or more data processors to perform part or all of one or more methods and/or part or all of one or more processes disclosed herein.).
Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over BA (US 20230307132 A1), hereinafter referenced as BA in view of HUANG (US 20170147905 A1), hereinafter referenced as HUANG in further view of LIU (US 20230343082 A1), hereinafter referenced as LIU.
Regarding claim 3, BA in view of HUANG teach a method according to claim 1,
BA in view of HUANG fail to explicitly teach wherein c. comprises providing receiving inputs from a manual inspection of an object and generating the ignore mask on a basis of the inputs.
However, LIU explicitly teaches wherein c. comprises providing receiving inputs from a manual inspection of an object (Fig. 2B, Paragraph [0021] – LIU discloses annotated samples may be manually annotated or auto annotated using a model or algorithm. Paragraph [0060] – LIU further discloses images annotated with the different classes A, B, C, and D are provided to the neural network 200 which feeds back on its predictions and performs validation steps, or training steps, to improve its predictions.)
and generating the ignore mask on a basis of the inputs (Fig. 7, Paragraph [0095] – LIU discloses the further dataset is encoded to include an ignore attribute to object classes that are annotated only in the other ones of the multiple datasets and to background classes associated with the other ones of the multiple datasets in step S206.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of BA in view of HUANG of having a method of training a machine learning model to identify image features, the method comprising: a. providing training image data, the training image data comprising a plurality of pixels; b. assigning a groundtruth annotation to each pixel, wherein each groundtruth annotation relates to a respective one of the pixels and each groundtruth annotation indicates whether or not that the pixel corresponds with an image feature; c. providing an ignore mask comprising a set of ignore flags, wherein each ignore flag relates to a respective one of the pixels and each ignore flag provides an indication that the pixel should be ignored; with the teachings of LIU having wherein c. comprises providing receiving inputs from a manual inspection of an object and generating the ignore mask on a basis of the inputs.
Wherein having BA’s method of training a machine learning model to identify image features wherein c. comprises providing receiving inputs from a manual inspection of an object and generating the ignore mask on a basis of the inputs.
The motivation behind the modification would have been to obtain an enhanced method of training a machine learning model to identify image features with improved accuracy and reduced workload, since both BA and LIU relate to arrangements for image or video recognition/understanding, wherein BA discloses relate to techniques that facilitate improving the accuracy of cell classifications by analyzing data generated by two separate machine-learning models, and LIU relates to encoding training data for training of a neural network for detecting and classifying objects wherein the performance of the neural network model is improved in terms of detection accuracy by training using training data encoded by the herein disclosed method; embodiments disclosed herein provide for reducing the workload and time for preparing annotations because every obtained data set does not need to have annotations for every class. Please see BA (US 20230307132 A1), Paragraph [0098], and LIU (US 20230343082 A1), Paragraph [0013, 0060].
Claims 4 and 6-11 are rejected under 35 U.S.C. 103 as being unpatentable over BA (US 20230307132 A1), hereinafter referenced as BA in view of HUANG (US 20170147905 A1), hereinafter referenced as HUANG in further view of LI (US 20210046861 A1), hereinafter referenced as LI.
Regarding claim 4, BA in view of HUANG teach a method according to claim 1,
BA in view of HUANG fail to explicitly teach wherein c. comprises inspecting an object with a sensor to generate three-dimensional inspection data and generating the ignore mask on a basis of the three-dimensional inspection data.
However, LI explicitly teaches wherein c. comprises inspecting an object with a sensor to generate three-dimensional inspection data (Fig. 1, Paragraph [0034] – LI discloses sensor data or image data pre-processor may use data representative of one or more images (or other data representations, such as LiDAR depth maps) and load the sensor data into memory in the form of a multi-dimensional array/matrix (alternatively referred to as tensor, or more specifically an input tensor, in some examples). See also Paragraph [0032].)
and generating the ignore mask on a basis of the three-dimensional inspection data (Fig 5, Paragraph [0058] – LI discloses leading vehicles 518 may be identified via their tail lights and, within a certain proximity to the high beam configuration or sensor collecting the sensor data 102, may likewise be identified as active vehicles. Any inactive vehicles (such as parked vehicles 520) identified via the segmentation masks may be ignored by the systems described herein, such that these parked/inactive vehicles may remain in the lit portions 512 of the environment.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of BA in view of HUANG of having a method of training a machine learning model to identify image features, the method comprising: a. providing training image data, the training image data comprising a plurality of pixels; b. assigning a groundtruth annotation to each pixel, wherein each groundtruth annotation relates to a respective one of the pixels and each groundtruth annotation indicates whether or not that the pixel corresponds with an image feature; c. providing an ignore mask comprising a set of ignore flags, wherein each ignore flag relates to a respective one of the pixels and each ignore flag provides an indication that the pixel should be ignored; with the teachings of LI having wherein c. comprises inspecting an object with a sensor to generate three-dimensional inspection data and generating the ignore mask on a basis of the three-dimensional inspection data.
Wherein having BA’s method of training a machine learning model to identify image features wherein c. comprises inspecting an object with a sensor to generate three-dimensional inspection data and generating the ignore mask on a basis of the three-dimensional inspection data.
The motivation behind the modification would have been to obtain an enhanced method of training a machine learning model to identify image features with improved accuracy and performance, since both BA and LI relate to arrangements for image or video recognition/understanding, wherein BA discloses relate to techniques that facilitate improving the accuracy of cell classifications by analyzing data generated by two separate machine-learning models, and LI relates to DNN-based image processing using pixel-level semantic segmentation of images received from cameras and/or sensors on the vehicles, combined with one or more post-processing techniques, may be used to adjust control and/or activation parameters of high beams of a vehicle; the segmentation masks output by the DNN(s), when combined with one or more post processing steps may result in accurate identification and localization of actionable actors such that high beam on/off activations and/or dimming or shading during activation may be effectively automated—thereby providing additional illumination of the environment while also controlling the downstream effect of the high beam activations for actors in the environment. Please see BA (US 20230307132 A1), Paragraph [0098], and LI (US 20210046861 A1), Paragraph [0004, 0005].
Regarding claim 6, BA in view of HUANG teach a method according to claim 2,
BA in view of HUANG fail to explicitly teach wherein the training image data comprises a series of images of the object which each contain the same feature viewed from a different viewing angle.
However, LI explicitly teaches wherein the training image data comprises a series of images of the object which each contain the same feature viewed from a different viewing angle (Fig. 9B, Paragraph [0032] – LI discloses in some examples, the sensor data 102 may be captured by a single camera with a forward-facing, substantially centered field of view with respect to a horizontal axis (e.g., left to right) of the vehicle 900. In a non-limiting embodiment, one or more forward-facing cameras may be used (e.g., a center or near-center mounted camera(s)), such as a wide-view camera 970, a surround camera 974, a stereo camera 968, and/or a long-range or mid-range camera 998. The sensor data 102 captured from this perspective may be useful for predictions described herein because a forward-facing camera may include a field of view (e.g., the field of view of the forward-facing stereo camera 968 and/or the wide-view camera 970 of FIG. 9B) that includes both a current lane of travel of the vehicle 900, adjacent lane(s) of travel of the vehicle 900, lanes for oncoming traffic, and/or boundaries of the driving surface. In some examples, more than one camera or other sensor (e.g., LiDAR sensor, RADAR sensor, etc.) may be used to incorporate multiple fields of view or sensory fields (e.g., the fields of view of the long-range cameras 998, the forward-facing stereo camera 968, and/or the forward-facing wide-view camera 970 of FIG. 9B).).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of BA in view of HUANG of having a method of training a machine learning model to identify image features, the method comprising: a. providing training image data, the training image data comprising a plurality of pixels; b. assigning a groundtruth annotation to each pixel, wherein each groundtruth annotation relates to a respective one of the pixels and each groundtruth annotation indicates whether or not that the pixel corresponds with an image feature; d. for each pixel, receiving a prediction value from the machine learning model, wherein each prediction value provides an indication of a probability of the pixel corresponding with an image feature; c. providing an ignore mask comprising a set of ignore flags, wherein each ignore flag relates to a respective one of the pixels and each ignore flag provides an indication that the pixel should be ignored; with the teachings of LI having wherein the training image data comprises a series of images of the object which each contain the same feature viewed from a different viewing angle.
Wherein having BA’s method of training a machine learning model to identify image features wherein the training image data comprises a series of images of the object which each contain the same feature viewed from a different viewing angle.
The motivation behind the modification would have been to obtain an enhanced method of training a machine learning model to identify image features with improved accuracy and performance, since both BA and LI relate to arrangements for image or video recognition/understanding, wherein BA discloses relate to techniques that facilitate improving the accuracy of cell classifications by analyzing data generated by two separate machine-learning models, and LI relates to DNN-based image processing using pixel-level semantic segmentation of images received from cameras and/or sensors on the vehicles, combined with one or more post-processing techniques, may be used to adjust control and/or activation parameters of high beams of a vehicle; the segmentation masks output by the DNN(s), when combined with one or more post processing steps may result in accurate identification and localization of actionable actors such that high beam on/off activations and/or dimming or shading during activation may be effectively automated—thereby providing additional illumination of the environment while also controlling the downstream effect of the high beam activations for actors in the environment. Please see BA (US 20230307132 A1), Paragraph [0098], and LI (US 20210046861 A1), Paragraph [0004, 0005].
Regarding claim 7, BA in view of HUANG in further view of LI teach a method according to claim 6,
BA in view of HUANG fail to explicitly teach further comprising generating the training image data by imaging the object from a series of different viewing angles.
However, LI explicitly teaches further comprising generating the training image data by imaging the object from a series of different viewing angles (Fig. 9B, Paragraph [0032] – LI discloses in some examples, the sensor data 102 may be captured by a single camera with a forward-facing, substantially centered field of view with respect to a horizontal axis (e.g., left to right) of the vehicle 900. In a non-limiting embodiment, one or more forward-facing cameras may be used (e.g., a center or near-center mounted camera(s)), such as a wide-view camera 970, a surround camera 974, a stereo camera 968, and/or a long-range or mid-range camera 998. The sensor data 102 captured from this perspective may be useful for predictions described herein because a forward-facing camera may include a field of view (e.g., the field of view of the forward-facing stereo camera 968 and/or the wide-view camera 970 of FIG. 9B) that includes both a current lane of travel of the vehicle 900, adjacent lane(s) of travel of the vehicle 900, lanes for oncoming traffic, and/or boundaries of the driving surface. In some examples, more than one camera or other sensor (e.g., LiDAR sensor, RADAR sensor, etc.) may be used to incorporate multiple fields of view or sensory fields (e.g., the fields of view of the long-range cameras 998, the forward-facing stereo camera 968, and/or the forward-facing wide-view camera 970 of FIG. 9B).).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of BA in view of HUANG in further view of LI of having a method of training a machine learning model to identify image features, the method comprising: a. providing training image data, the training image data comprising a plurality of pixels; b. assigning a groundtruth annotation to each pixel, wherein each groundtruth annotation relates to a respective one of the pixels and each groundtruth annotation indicates whether or not that the pixel corresponds with an image feature; d. for each pixel, receiving a prediction value from the machine learning model, wherein each prediction value provides an indication of a probability of the pixel corresponding with an image feature; c. providing an ignore mask comprising a set of ignore flags, wherein each ignore flag relates to a respective one of the pixels and each ignore flag provides an indication that the pixel should be ignored; with the teachings of LI having further comprising generating the training image data by imaging the object from a series of different viewing angles.
Wherein having BA’s method of training a machine learning model to identify image features further comprising generating the training image data by imaging the object from a series of different viewing angles.
The motivation behind the modification would have been to obtain an enhanced method of training a machine learning model to identify image features with improved accuracy and performance, since both BA and LI relate to arrangements for image or video recognition/understanding, wherein BA discloses relate to techniques that facilitate improving the accuracy of cell classifications by analyzing data generated by two separate machine-learning models, and LI relates to DNN-based image processing using pixel-level semantic segmentation of images received from cameras and/or sensors on the vehicles, combined with one or more post-processing techniques, may be used to adjust control and/or activation parameters of high beams of a vehicle; the segmentation masks output by the DNN(s), when combined with one or more post processing steps may result in accurate identification and localization of actionable actors such that high beam on/off activations and/or dimming or shading during activation may be effectively automated—thereby providing additional illumination of the environment while also controlling the downstream effect of the high beam activations for actors in the environment. Please see BA (US 20230307132 A1), Paragraph [0098], and LI (US 20210046861 A1), Paragraph [0004, 0005].
Regarding claim 8, BA in view of HUANG in further view of LI teach a method according to claim 7,
BA in view of HUANG fail to explicitly teach wherein the object is imaged with light.
However, LI explicitly teaches Hhwherein the object is imaged with light (Fig. 9A-B, Paragraph [0032] – LI discloses more than one camera or other sensor (e.g., LiDAR sensor, RADAR sensor, etc.) may be used to incorporate multiple fields of view or sensory fields (e.g., the fields of view of the long-range cameras 998, the forward-facing stereo camera 968, and/or the forward-facing wide-view camera 970 of FIG. 9B). Paragraph [0156] – LI further discloses LIDAR technologies, such as 3D flash LIDAR, may also be used. 3D Flash LIDAR uses a flash of a laser as a transmission source, to illuminate vehicle surroundings up to approximately 200 m. A flash LIDAR unit includes a receptor, which records the laser pulse transit time and the reflected light on each pixel, which in turn corresponds to the range from the vehicle to the objects. Flash LIDAR may allow for highly accurate and distortion-free images of the surroundings to be generated with every laser flash.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of BA in view of HUANG in further view of LI of having a method of training a machine learning model to identify image features, the method comprising: a. providing training image data, the training image data comprising a plurality of pixels; b. assigning a groundtruth annotation to each pixel, wherein each groundtruth annotation relates to a respective one of the pixels and each groundtruth annotation indicates whether or not that the pixel corresponds with an image feature; d. for each pixel, receiving a prediction value from the machine learning model, wherein each prediction value provides an indication of a probability of the pixel corresponding with an image feature; c. providing an ignore mask comprising a set of ignore flags, wherein each ignore flag relates to a respective one of the pixels and each ignore flag provides an indication that the pixel should be ignored; with the teachings of LI having wherein the object is imaged with light.
Wherein having BA’s method of training a machine learning model to identify image features wherein the object is imaged with light.
The motivation behind the modification would have been to obtain an enhanced method of training a machine learning model to identify image features with improved accuracy and performance, since both BA and LI relate to arrangements for image or video recognition/understanding, wherein BA discloses relate to techniques that facilitate improving the accuracy of cell classifications by analyzing data generated by two separate machine-learning models, and LI relates to DNN-based image processing using pixel-level semantic segmentation of images received from cameras and/or sensors on the vehicles, combined with one or more post-processing techniques, may be used to adjust control and/or activation parameters of high beams of a vehicle; the segmentation masks output by the DNN(s), when combined with one or more post processing steps may result in accurate identification and localization of actionable actors such that high beam on/off activations and/or dimming or shading during activation may be effectively automated—thereby providing additional illumination of the environment while also controlling the downstream effect of the high beam activations for actors in the environment. Please see BA (US 20230307132 A1), Paragraph [0098], and LI (US 20210046861 A1), Paragraph [0004, 0005].
Regarding claim 9, BA in view of HUANG teach a method according to claim 1,
BA in view of HUANG fail to explicitly teach wherein the training image data comprises a series of images of an object which each contain the same feature viewed from a different viewing angle.
However, LI explicitly teaches wherein the training image data comprises a series of images of an object which each contain the same feature viewed from a different viewing angle (Fig. 9B, Paragraph [0032] – LI discloses in some examples, the sensor data 102 may be captured by a single camera with a forward-facing, substantially centered field of view with respect to a horizontal axis (e.g., left to right) of the vehicle 900. In a non-limiting embodiment, one or more forward-facing cameras may be used (e.g., a center or near-center mounted camera(s)), such as a wide-view camera 970, a surround camera 974, a stereo camera 968, and/or a long-range or mid-range camera 998. The sensor data 102 captured from this perspective may be useful for predictions described herein because a forward-facing camera may include a field of view (e.g., the field of view of the forward-facing stereo camera 968 and/or the wide-view camera 970 of FIG. 9B) that includes both a current lane of travel of the vehicle 900, adjacent lane(s) of travel of the vehicle 900, lanes for oncoming traffic, and/or boundaries of the driving surface. In some examples, more than one camera or other sensor (e.g., LiDAR sensor, RADAR sensor, etc.) may be used to incorporate multiple fields of view or sensory fields (e.g., the fields of view of the long-range cameras 998, the forward-facing stereo camera 968, and/or the forward-facing wide-view camera 970 of FIG. 9B).).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of BA in view of HUANG of having a method of training a machine learning model to identify image features, the method comprising: a. providing training image data, the training image data comprising a plurality of pixels; b. assigning a groundtruth annotation to each pixel, wherein each groundtruth annotation relates to a respective one of the pixels and each groundtruth annotation indicates whether or not that the pixel corresponds with an image feature; d. for each pixel, receiving a prediction value from the machine learning model, wherein each prediction value provides an indication of a probability of the pixel corresponding with an image feature; c. providing an ignore mask comprising a set of ignore flags, wherein each ignore flag relates to a respective one of the pixels and each ignore flag provides an indication that the pixel should be ignored; with the teachings of LI having wherein the training image data comprises a series of images of an object which each contain the same feature viewed from a different viewing angle.
Wherein having BA’s method of training a machine learning model to identify image features wherein the training image data comprises a series of images of an object which each contain the same feature viewed from a different viewing angle.
The motivation behind the modification would have been to obtain an enhanced method of training a machine learning model to identify image features with improved accuracy and performance, since both BA and LI relate to arrangements for image or video recognition/understanding, wherein BA discloses relate to techniques that facilitate improving the accuracy of cell classifications by analyzing data generated by two separate machine-learning models, and LI relates to DNN-based image processing using pixel-level semantic segmentation of images received from cameras and/or sensors on the vehicles, combined with one or more post-processing techniques, may be used to adjust control and/or activation parameters of high beams of a vehicle; the segmentation masks output by the DNN(s), when combined with one or more post processing steps may result in accurate identification and localization of actionable actors such that high beam on/off activations and/or dimming or shading during activation may be effectively automated—thereby providing additional illumination of the environment while also controlling the downstream effect of the high beam activations for actors in the environment. Please see BA (US 20230307132 A1), Paragraph [0098], and LI (US 20210046861 A1), Paragraph [0004, 0005].
Regarding claim 10, BA in view of HUANG in further view of LI teach a method according to claim 9,
BA in view of HUANG fail to explicitly teach further comprising generating the training image data by imaging the object from a series of different viewing angles.
However, LI explicitly teaches further comprising generating the training image data by imaging the object from a series of different viewing angles (Fig. 9B, Paragraph [0032] – LI discloses in some examples, the sensor data 102 may be captured by a single camera with a forward-facing, substantially centered field of view with respect to a horizontal axis (e.g., left to right) of the vehicle 900. In a non-limiting embodiment, one or more forward-facing cameras may be used (e.g., a center or near-center mounted camera(s)), such as a wide-view camera 970, a surround camera 974, a stereo camera 968, and/or a long-range or mid-range camera 998. The sensor data 102 captured from this perspective may be useful for predictions described herein because a forward-facing camera may include a field of view (e.g., the field of view of the forward-facing stereo camera 968 and/or the wide-view camera 970 of FIG. 9B) that includes both a current lane of travel of the vehicle 900, adjacent lane(s) of travel of the vehicle 900, lanes for oncoming traffic, and/or boundaries of the driving surface. In some examples, more than one camera or other sensor (e.g., LiDAR sensor, RADAR sensor, etc.) may be used to incorporate multiple fields of view or sensory fields (e.g., the fields of view of the long-range cameras 998, the forward-facing stereo camera 968, and/or the forward-facing wide-view camera 970 of FIG. 9B).).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of BA in view of HUANG in further view of LI of having a method of training a machine learning model to identify image features, the method comprising: a. providing training image data, the training image data comprising a plurality of pixels; b. assigning a groundtruth annotation to each pixel, wherein each groundtruth annotation relates to a respective one of the pixels and each groundtruth annotation indicates whether or not that the pixel corresponds with an image feature; d. for each pixel, receiving a prediction value from the machine learning model, wherein each prediction value provides an indication of a probability of the pixel corresponding with an image feature; c. providing an ignore mask comprising a set of ignore flags, wherein each ignore flag relates to a respective one of the pixels and each ignore flag provides an indication that the pixel should be ignored; with the teachings of LI having further comprising generating the training image data by imaging the object from a series of different viewing angles.
Wherein having BA’s method of training a machine learning model to identify image features further comprising generating the training image data by imaging the object from a series of different viewing angles.
The motivation behind the modification would have been to obtain an enhanced method of training a machine learning model to identify image features with improved accuracy and performance, since both BA and LI relate to arrangements for image or video recognition/understanding, wherein BA discloses relate to techniques that facilitate improving the accuracy of cell classifications by analyzing data generated by two separate machine-learning models, and LI relates to DNN-based image processing using pixel-level semantic segmentation of images received from cameras and/or sensors on the vehicles, combined with one or more post-processing techniques, may be used to adjust control and/or activation parameters of high beams of a vehicle; the segmentation masks output by the DNN(s), when combined with one or more post processing steps may result in accurate identification and localization of actionable actors such that high beam on/off activations and/or dimming or shading during activation may be effectively automated—thereby providing additional illumination of the environment while also controlling the downstream effect of the high beam activations for actors in the environment. Please see BA (US 20230307132 A1), Paragraph [0098], and LI (US 20210046861 A1), Paragraph [0004, 0005].
Regarding claim 11, BA in view of HUANG in further view of LI teach a method according to claim 10,
BA in view of HUANG fail to explicitly teach wherein the object is imaged with light.
However, LI explicitly teaches Hhwherein the object is imaged with light (Fig. 9A-B, Paragraph [0032] – LI discloses more than one camera or other sensor (e.g., LiDAR sensor, RADAR sensor, etc.) may be used to incorporate multiple fields of view or sensory fields (e.g., the fields of view of the long-range cameras 998, the forward-facing stereo camera 968, and/or the forward-facing wide-view camera 970 of FIG. 9B). Paragraph [0156] – LI further discloses LIDAR technologies, such as 3D flash LIDAR, may also be used. 3D Flash LIDAR uses a flash of a laser as a transmission source, to illuminate vehicle surroundings up to approximately 200 m. A flash LIDAR unit includes a receptor, which records the laser pulse transit time and the reflected light on each pixel, which in turn corresponds to the range from the vehicle to the objects. Flash LIDAR may allow for highly accurate and distortion-free images of the surroundings to be generated with every laser flash.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of BA in view of HUANG in further view of LI of having a method of training a machine learning model to identify image features, the method comprising: a. providing training image data, the training image data comprising a plurality of pixels; b. assigning a groundtruth annotation to each pixel, wherein each groundtruth annotation relates to a respective one of the pixels and each groundtruth annotation indicates whether or not that the pixel corresponds with an image feature; d. for each pixel, receiving a prediction value from the machine learning model, wherein each prediction value provides an indication of a probability of the pixel corresponding with an image feature; c. providing an ignore mask comprising a set of ignore flags, wherein each ignore flag relates to a respective one of the pixels and each ignore flag provides an indication that the pixel should be ignored; with the teachings of LI having wherein the object is imaged with light.
Wherein having BA’s method of training a machine learning model to identify image features wherein the object is imaged with light.
The motivation behind the modification would have been to obtain an enhanced method of training a machine learning model to identify image features with improved accuracy and performance, since both BA and LI relate to arrangements for image or video recognition/understanding, wherein BA discloses relate to techniques that facilitate improving the accuracy of cell classifications by analyzing data generated by two separate machine-learning models, and LI relates to DNN-based image processing using pixel-level semantic segmentation of images received from cameras and/or sensors on the vehicles, combined with one or more post-processing techniques, may be used to adjust control and/or activation parameters of high beams of a vehicle; the segmentation masks output by the DNN(s), when combined with one or more post processing steps may result in accurate identification and localization of actionable actors such that high beam on/off activations and/or dimming or shading during activation may be effectively automated—thereby providing additional illumination of the environment while also controlling the downstream effect of the high beam activations for actors in the environment. Please see BA (US 20230307132 A1), Paragraph [0098], and LI (US 20210046861 A1), Paragraph [0004, 0005].
Claims 15 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over BA (US 20230307132 A1), hereinafter referenced as BA in view of HUANG (US 20170147905 A1), hereinafter referenced as HUANG in further view of AFRASIABI (US 20240249500 A1), hereinafter referenced as AFRASIABI.
Regarding claim 15, BA in view of HUANG teach a method according to claim 14,
BA in view of HUANG fail to explicitly teach wherein the image feature comprises a surface defect of an aircraft.
However, AFRASIABI explicitly teaches wherein the image feature comprises a surface defect of an aircraft (Fig. 1, Paragraph [0038] – AFRASIABI discloses above-described computing system can be applied in quality inspections to detect features on an aircraft skin such as scratches, rivets, dents, cracks, etc. For example, an object detector 38 configured to detect scratches on aircraft skin can generate cropped images 40 of scratches on the aircraft skin. Filtered cropped images 44 of the scratches can be inputted into a clustering model 54, and features 48 of the filtered cropped images 44 can be clustered into feature clusters 56, which are further filtered to generate a training dataset 72, which is used to train a pre-training machine learning model 66 to improve its performance in detecting scratches on images of aircraft skin.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of BA in view of HUANG of having a method of training a machine learning model to identify image features, the method comprising: a. providing training image data, the training image data comprising a plurality of pixels; b. assigning a groundtruth annotation to each pixel, wherein each groundtruth annotation relates to a respective one of the pixels and each groundtruth annotation indicates whether or not that the pixel corresponds with an image feature; with the teachings of AFRASIABI having wherein the image feature comprises a surface defect of an aircraft.
Wherein having BA’s method of training a machine learning model to identify image features wherein the image feature comprises a surface defect of an aircraft.
The motivation behind the modification would have been to obtain an enhanced method of training a machine learning model to identify image features with improved accuracy and performance, since both BA and AFRASIABI relate to processing image/video features in feature spaces, wherein BA discloses relate to techniques that facilitate improving the accuracy of cell classifications by analyzing data generated by two separate machine-learning models, and AFRASIABI relates to a system and method for training a machine learning model by efficiently generating a diverse, specific, and accurate training dataset with reduced computing resources. Please see BA (US 20230307132 A1), Paragraph [0098], and AFRASIABI (US 20240249500 A1), Paragraph [0018, 0038].
Regarding claim 16, BA in view of HUANG teach a method according to claim 14,
BA in view of HUANG fail to explicitly teach wherein the image feature comprises a dent.
However, AFRASIABI explicitly teaches wherein the image feature comprises a dent (Fig. 1, Paragraph [0038] – AFRASIABI discloses above-described computing system can be applied in quality inspections to detect features on an aircraft skin such as scratches, rivets, dents, cracks, etc. For example, an object detector 38 configured to detect scratches on aircraft skin can generate cropped images 40 of scratches on the aircraft skin. Filtered cropped images 44 of the scratches can be inputted into a clustering model 54, and features 48 of the filtered cropped images 44 can be clustered into feature clusters 56, which are further filtered to generate a training dataset 72, which is used to train a pre-training machine learning model 66 to improve its performance in detecting scratches on images of aircraft skin.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date the claimed invention was made to combine the teachings of BA in view of HUANG of having a method of training a machine learning model to identify image features, the method comprising: a. providing training image data, the training image data comprising a plurality of pixels; b. assigning a groundtruth annotation to each pixel, wherein each groundtruth annotation relates to a respective one of the pixels and each groundtruth annotation indicates whether or not that the pixel corresponds with an image feature; with the teachings of AFRASIABI having wherein the image feature comprises a dent.
Wherein having BA’s method of training a machine learning model to identify image features wherein the image feature comprises a dent.
The motivation behind the modification would have been to obtain an enhanced method of training a machine learning model to identify image features with improved accuracy and performance, since both BA and AFRASIABI relate to processing image/video features in feature spaces, wherein BA discloses relate to techniques that facilitate improving the accuracy of cell classifications by analyzing data generated by two separate machine-learning models, and AFRASIABI relates to a system and method for training a machine learning model by efficiently generating a diverse, specific, and accurate training dataset with reduced computing resources. Please see BA (US 20230307132 A1), Paragraph [0098], and AFRASIABI (US 20240249500 A1), Paragraph [0018, 0038].
Allowable Subject Matter
Claim 17, a dependent claim, is therefrom objected to as being dependent upon rejected base claim 1, but would be allowable if rewritten in independent form including all of the limitations of the base claims and any intervening claims, once the claim objections along with the double patenting rejection have been overcome.
The following is a statement of reasons for the indication of allowable subject matter:
Regarding claim 17, the prior arts fail to explicitly teach wherein the loss value is determined by the algorithm:
PNG
media_image1.png
93
466
media_image1.png
Greyscale
wherein yk is a groundtruth annotation for that pixel; pk is a prediction value for that pixel, a pixel which corresponds with an image feature has a groundtruth annotation of yk of 1, and a pixel which does not correspond with an image feature has a groundtruth annotation of yk of 0.
Conclusion
Listed below are the prior arts made of record and not relied upon but are considered pertinent to applicant’s disclosure.
Tang et al. (US 20220351386 A1) - The present disclosure provides a computer-implemented method, a device, and a storage medium. The method includes inputting an image into an attention-enhanced high-resolution network (AHRNet) to extract feature maps for generating a first feature map; generating a first probability map which is concatenated with the first feature map to form a concatenated first feature map, and updating the AHRNet using the first segmentation loss; generating a second feature map, and scaling the second feature map to form a third feature map; generating a second probability map which is concatenated with the third feature map to form a concatenated third feature map, and updating the AHRNet using the second segmentation loss; generating a fourth feature map, and scaling the fourth feature map to form a fifth feature map; updating the AHRNet using the third segmentation loss and the regional level set loss; and outputting the third probability map.… Fig. 1, Abstract.
HSIEH et al. (US 20230132180 A1) - The present disclosure relates to systems, methods, and non-transitory computer-readable media that upsample and refine segmentation masks. Indeed, in one or more implementations, a segmentation mask refinement and upsampling system upsamples a preliminary segmentation mask utilizing a patch-based refinement process to generate a patch-based refined segmentation mask. The segmentation mask refinement and upsampling system then fuses the patch-based refined segmentation mask with an upsampled version of the preliminary segmentation mask. By fusing the patch-based refined segmentation mask with the upsampled preliminary segmentation mask, the segmentation mask refinement and upsampling system maintains a global perspective and helps avoid artifacts due to the local patch-based refinement process.… Fig. 1, Abstract.
WANG et al. (US 20230135234 A1) - In various examples, to support training a deep neural network (DNN) to predict a dense representation of a 3D surface structure of interest, a training dataset is generated from real-world data. For example, one or more vehicles may collect image data and LiDAR data while navigating through a real-world environment. To generate input training data, 3D surface structure estimation may be performed on captured image data to generate a sparse representation of a 3D surface structure of interest (e.g., a 3D road surface). To generate corresponding ground truth training data, captured LiDAR data may be smoothed, subject to outlier removal, subject to triangulation to filling missing values, accumulated from multiple LiDAR sensors, aligned with corresponding frames of image data, and/or annotated to identify 3D points on the 3D surface of interest, and the identified 3D points may be projected to generate a dense representation of the 3D surface structure.… Fig. 1, Abstract.
CHOI et al. (US 20220051017 A1) - Apparatuses, systems, and techniques to identify one or more objects in one or more images. In at least one embodiment, one or more objects are identified in one or more images based, at least in part, on a likelihood that one or more objects is different from other objects in one or more images..… Fig. 1, Abstract.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to BEZAWIT N SHIMELES whose telephone number is (571)272-7663. The examiner can normally be reached M-F 7:30am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chineyere Wills-Burns can be reached at (571) 272-9752. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/BEZAWIT NOLAWI SHIMELES/Examiner, Art Unit 2673
/CHINEYERE WILLS-BURNS/ Supervisory Patent Examiner, Art Unit 2673