Prosecution Insights
Last updated: August 17, 2026
Application No. 18/320,265

OBJECT SEGMENTATION USING MACHINE LEARNING FOR AUTONOMOUS SYSTEMS AND APPLICATIONS

Final Rejection §102§103
Filed
May 19, 2023
Examiner
RHIM, WOO CHUL
Art Unit
2676
Tech Center
2600 — Communications
Assignee
NVIDIA Corporation
OA Round
4 (Final)
78%
Grant Probability
Favorable
5-6
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 78% — above average
78%
Career Allowance Rate
121 granted / 155 resolved
+16.1% vs TC avg
Strong +22% interview lift
Without
With
+22.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 8m
Avg Prosecution
26 currently pending
Career history
182
Total Applications
across all art units

Statute-Specific Performance

§101
7.2%
-32.8% vs TC avg
§103
49.3%
+9.3% vs TC avg
§102
22.9%
-17.1% vs TC avg
§112
17.1%
-22.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 155 resolved cases

Office Action

§102 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendments Submission dated 06/18/2026 amends 1, 5, 9, 13, and 17. Claims 1-20 are pending. In view of the amendments to the claims, the previously set forth claim objections have been withdrawn. The examiner notes that he amendments to claims 1 and 17 include claim language that was added but not underlined. The phase “, using the one or more neural network,” in line 9 of claim 1 is new and should have been underlined. For the prior art purposes, the examiner has entered the phrase as part of the amendment because the omission of the underline appears unintentional. Extra caution is advised. The phrase “the one or more candidate shapes of masks” in line 10 of claim 17 is new and should have been underlined. For the prior art purposes, the examiner has entered the phrase as part of the amendment because the omission of the underline appears unintentional. Extra caution is advised. Response to Arguments Applicant's arguments filed with the submission have been fully considered but they are not persuasive. On pages 9-10 of the submission, the applicant argues that us patent application publication no. 2023/0099494 to Kocamaz et al. (hereinafter Kocamaz) does not teach the amended claim languages of the independent claims because the teaching of Kocamaz is “not the same as "using a mask identifier of one or more neural networks" to determine "one or more candidate shapes or masks corresponding to the object for generation of a segmentation mask" and then using the neural network to determine "target portions of the one or more candidate shapes or masks outputted by the mask identifier" as recited in claim 1, and similarly in claims 9 and 17” (see the second full par. of page 10 of the submission). The examiner does not find the argument persuasive because other than above argument itself, the submission does not provide any evidence or the rationale for supporting the argument, and in the absence of such evidence and/or rationale, the above argument is “just an attorney argument and not the kind of not the kind of factual evidence that is required to rebut a prima facie case of obviousness" (see MPEP 2145(I)). Moreover, as provided below in the body of the rejection, Kocamaz applied teaches the amended claim language (see the 102 rejection provided below). On pages 10-11, the applicant argues that the dependent claims are patentable for their dependencies to the independent claims. For the reasons provided above and the evidence provided in the body of the rejection, the examiner finds this argument unpersuasive. Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: Mask identifier in claims 1, 9 and 17. Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 1, 2, 4-6, 8-10, 12-14, and 16-20 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Kocamaz. For claims 1 and 9, Kocamaz as applied discloses at least one processor (see, e.g., pars. 190-203 and FIGS. 8-10 and claims 1 and 12) comprising: one or more circuits to: determine a bounding shape corresponding to an object depicted in an image (see, e.g., pars. 35-37, 52, 60, 65-67, 72, and 77 and FIGS. 2B-D, 3, 5A-B, 6, and 7, which teach determining a bounding shape corresponding to a vehicle in an input image); determine, using a mask identifier of one or more neural networks, one or more candidate shapes or masks corresponding to the object for generation of a segmentation mask corresponding to the object, the one or more candidate shapes or masks determined based on at least on features extracted from the image (see, e.g., pars. 20-23, 26, 31-33, 46-47, 60-63, 70-71 and 75-76 and FIGS. 6 and 7, which teach determining one or more output masks, i.e., sets of pixels of the image, corresponding to the object based on object classes and lane identifiers determined from the image using the machine learning models such as a CNN; the examiner interprets the output mask generating/computing part in the machine learning model as the claimed mask identifier and the object classes and lane identifiers and also the implicit image features extracted by convolutional layers in the machine learning model as the claimed features); determine, using the one or more neural networks, target portions of the one or more candidate shapes or masks outputted by the mask identifier by excluding at least a non-target portion of the one or more candidate shapes or masks that are outside of a region identified by the bounding shape (see, e.g., pars. 64-67, 72-74 and 77- 79 and FIGS. 5-7, which teach determining the bounding shape using the machine learning model and associating the pixel sets with the bounding shapes by determining portions of the first and second set of pixels that are within the bounding shape and portions that are outside of the bounding shape); determine the segmentation mask corresponding to the object depicted in the image based at least on the one or more neural networks processing the image, bounding shape information corresponding to the bounding shape, and the target portions of the one or more candidate shapes or masks (see, e.g., pars. 35-37, 40-41, 53-57, 60, 65-67, and 69-79 and FIGS. 4-7, which teach determining the output mask assigned to the object based on the machine learning model, the bounding shape and portions of the pixel sets that are within the bounding shape), the segmentation mask including less pixels of the image than the bounding shape (see, e.g., par. 64, which teaches removing pixels outside of a bounding shape, suggesting that the mask corresponding to the object would have less pixels of the image than the bounding shape); and cause a machine to perform one or more planning, navigation, or control operations based at least on the segmentation mask (see, e.g., pars. 61-64, 68 and 80 and FIGS. 4 and 7, which teach performing one or more operations, such as vehicle control and navigation based on the output masks). For claim 17, Kocamaz as applied discloses: performing one or more operations by a machine based at least on a segmentation mask corresponding to an object (see, e.g., pars. 61-64, 68, and 80 and FIGS. 4 and 7, which teach performing one or more operations based on the output masks), the segmentation mask generated using (i) an image of the object, (ii) a bounding shape corresponding to the object, and (iii) target portions of one or more candidate shapes or masks (see, e.g., pars. 35-37, 40-41, 53-57, 60, 65-67, and 69-79 and FIGS. 4-7, which teach determining the output mask assigned to the object based on the bounding shape and portions of the image pixel sets that are within the bounding shape), the segmentation mask including less pixels of the image than the bounding shape (see, e.g., par. 64, which teaches removing pixels outside of a bounding shape, suggesting that the mask corresponding to the object would have less pixels of the image than the bounding shape), the one or more candidate shapes or masks corresponding to the object and for generation of the segmentation mask corresponding to the object determined using a mask identifier of a neural network based at least on features extracted from the image (see, e.g., pars. 20-23, 26, 31-33, 46-47, 60-63, 70-71 and 75-76 and FIGS. 6 and 7, which teach determining one or more output masks, i.e., sets of pixels of the image, corresponding to the object based on object classes and lane identifiers determined from the image using the machine learning models such as a CNN; the examiner interprets the output mask generating/computing part in the machine learning model as the claimed mask identifier and the object classes and lane identifiers and also the implicit image features extracted by convolutional layers in the machine learning model as the claimed features), the target portions of the one or more candidate shapes or masks determined by the neural network from the one or more candidate shapes of masks outputted by the mask identifier by excluding at least a non-target portion of the one or more candidate shapes or masks that are outside of a region identified by the bounding shape (see, e.g., pars. 64-67, 72-74 and 77- 79 and FIGS. 5-7, which teach determining the bounding shape using the machine learning model and associating the pixel sets with the bounding shapes by determining portions of the first and second set of pixels that are within the bounding shape and portions that are outside of the bounding shape), and the neural network that was trained using ground truth data generated using bounding shape labels (see, e.g., pars. 30-40, 53-57, 60, 65-67, 70, and 75 and FIGS. 1, 2A-D, 4, 5A-B, 6, and 7, which teach a machine learning model trained using GT masks encoded with the bounding shape). For claims 2 and 10, Kocamaz as applied discloses that the one or more planning, navigation, or control operations include at least one of: operating a simulation using the segmentation mask; operating a machine perception system using the segmentation mask (see, e.g., pars. 68, 175 and 179 and claim 11); assigning the segmentation mask to a data structure comprising the image. For claims 4 and 12, Kocamaz as applied discloses that the one or more neural networks are configured using training data comprising a plurality of training instances, at least one individual training instance of the plurality of training instances having a corresponding bounding shape and class indication (see, e.g., pars. 31-34 and FIGS. 1, 2A-C, which teach training a machine learning model using ground truth masks encoded with object class labels that include annotations corresponding to bounding shapes). For claims 5 and 13, Kocamaz as applied teaches that the bounding shape surrounds the object in the image and a first additional portion of the image corresponding to first pixels of the image within the bounding shape and outside an outline of the object (see, e.g., pars. 34-37 and FIGS. 2A-B and 5A-B, which teach that the bounding shapes includes the vehicles and areas surrounding the vehicles), the segmentation mask forms an outline of the object and a second additional portion of the image corresponding to second pixels of the image within the segmentation mask and outside the outline of the object (see, e.g., pars. 34-37 and FIGS. 2C-D, which teach that the segmentations masks include the vehicles and areas surrounding the vehicles), the second additional portion including less pixels than the first additional portion (see, e.g., par. 64, which teach removing pixels outside of a bounding shape, suggesting that the surrounding areas in the mask would have less pixels than the surrounding areas in the bounding shape) For claims 6 and 14, Kocamaz as applied discloses that one or more parameters of the one or more neural networks are updated using the bounding shape and the segmentation mask (see, e.g., pars. 40 and 50, which teach updating parameters of the machine learning models using the output masks encoded with the bounding shapes). For claims 8, 16 and 20, Kocamaz as applied discloses that the processor is comprised in at least one of: a control system for an autonomous or semi-autonomous machine (see, e.g., pars. 19-20, 25, 59, and 81, FIGS. 8A-D and claim 11); a perception system for an autonomous or semi-autonomous machine ; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing conversational Al operations; a system implementing one or more large language models (LLMs); a system for generating synthetic data; a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources. For claim 18, Kocamaz as applied discloses that the ground truth data includes segmentation masks that are automatically generated from the bounding shape labels (see, e.g., pars. 30-40, which does not mention human intervention and that the annotations used for training may be machine-automated; the examiner interprets the absence of human intervention and the automated annotating to suggests that GT masks are automatically generated from the annotations). For claim 19, Kocamaz as applied discloses that the one or more operations include at least one of a planning operation, a control operation, a navigation operation, or an actuation operation (see, e.g., pars. 4, 61 and 68, which teach a navigation operation). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 3, 7, 11, and 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kocamaz in view of us patent application publication no. 2021/0063578 to Wekel et al. (hereinafter Wekel). For claims 3 and 11, while Kocamaz as applied does not explicitly teach, Wekel in the analogous art teaches that the one or more circuits are to receive the bounding shape as a three-dimensional shape, and convert the bounding shape to a two- dimensional shape corresponding to a frame of reference of the image (see, e.g., pars. 35-36 and 44 of Wekel, which teach receiving and projecting 3D LiDAR point cloud to 2D range image, which includes the instance segmentation mask corresponding to the bounding shapes of the image labels in the image). It would have been obvious to modify Kocamaz to convert the bounding shapes as taught by Wekel because doing so would allow the location of the bounding shapes to be tracked during the conversion/transformation of the input data such that accurate ground truth data in the LiDAR range image domain may be generated (see, e.g., pars. 5, 24, 32, and 41 of Wekel). For claims 7 and 15, while Kocamaz does not explicitly teach, Wekel in the analogous art teaches that the bounding shape information comprises (i) a data structure indicating a position of at least one of a corner or an edge of the bounding shape (see, e.g., pars. 35, 48 and 82 of Wekel, which teach that the bounding shape information includes pixel locations of one or more vertices of a bounding shape or offsets from an edge of a bonding shape) and (ii) an identifier of the object (see, e.g., par. 36, which teach that the bounding shape information includes the boundary contours encoded with unique instance information including an instance value for each separate object). It would have been obvious to modify Kocamaz to use the bounding shape information as taught by Wekel because doing so would allow the location of the bounding shapes to be tracked during the conversion/transformation of the input data such that accurate ground truth data in the LiDAR range image domain may be generated (see, e.g., pars. 5, 24, 32, and 41 of Wekel). Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to WOO RHIM whose telephone number is (571)272-6560. The examiner can normally be reached Mon - Fri 9:30 am - 6:00 pm et. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Henok Shiferaw can be reached at 571-272-4637. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /WOO C RHIM/Examiner, Art Unit 2676
Read full office action

Prosecution Timeline

Show 6 earlier events
Jan 20, 2026
Response after Non-Final Action
Jan 30, 2026
Request for Continued Examination
Feb 02, 2026
Response after Non-Final Action
Mar 18, 2026
Non-Final Rejection mailed — §102, §103
Jun 02, 2026
Examiner Interview Summary
Jun 02, 2026
Applicant Interview (Telephonic)
Jun 18, 2026
Response Filed
Jul 24, 2026
Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705931
SIGN-LANGUAGE TRANSLATION
2y 7m to grant Granted Aug 11, 2026
Patent 12700062
IMAGE PROCESSING METHOD AND DEVICE
2y 7m to grant Granted Aug 04, 2026
Patent 12688549
PASS THROUGH USING COMMON IMAGE SENSOR FOR COLOUR AND DEPTH
2y 6m to grant Granted Jul 21, 2026
Patent 12670597
SYSTEMS AND METHODS FOR EFFICENTLY SENSING COLLISON THREATS
4y 0m to grant Granted Jun 30, 2026
Patent 12664622
METHODS FOR REDUCING THE APPEARANCE OF BLOCK-RELATED ARTIFACTS
2y 9m to grant Granted Jun 23, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
78%
Grant Probability
99%
With Interview (+22.5%)
2y 8m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 155 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month