Prosecution Insights
Last updated: October 02, 2026
Application No. 18/860,072

AUTOMATED OBJECTS LABELING IN VIDEO DATA FOR MACHINE LEARNING AND OTHER CLASSIFIERS

Non-Final OA §102§103
Filed
Oct 25, 2024
Priority
Apr 25, 2022 — provisional 63/363,526 +2 more
Examiner
ZAK, JACQUELINE ROSE
Art Unit
Tech Center
Assignee
Virginia Polytechnic Institute and State University
OA Round
1 (Non-Final)
64%
Grant Probability
Moderate
1-2
OA Rounds
1y 3m
Est. Remaining
77%
With Interview

Examiner Intelligence

Grants 64% of resolved cases
64%
Career Allowance Rate
23 granted / 36 resolved
+3.9% vs TC avg
Moderate +13% lift
Without
With
+12.7%
Interview Lift
resolved cases with interview
Typical timeline
3y 2m
Avg Prosecution
29 currently pending
Career history
64
Total Applications
across all art units

Statute-Specific Performance

§101
5.0%
-35.0% vs TC avg
§103
60.6%
+20.6% vs TC avg
§102
17.5%
-22.5% vs TC avg
§112
13.4%
-26.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 36 resolved cases

Office Action

§102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Status Claims 1-20 are pending for examination in the application filed 10/25/2024. Priority Acknowledgement is made of Applicant’s claim to priority of provisional application 63/363,526, filing date 04/25/2022. Acknowledgement is additionally made of the present application as a national stage entry of PCT/US23/66205, international filing date: 04/25/2023. Information Disclosure Statement The information disclosure statement (IDS) submitted on 10/25/2024 has been considered by the examiner. Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier, as explained in MPEP §2181, subsection I (note that the list of generic placeholders below is not exhaustive, and other generic placeholders may invoke 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph): A. The Claim Limitation Uses the Term “Means” or “Step” or a Generic Placeholder (A Term That Is Simply A Substitute for “Means”) With respect to the first prong of this analysis, a claim element that does not include the term “means” or “step” triggers a rebuttable presumption that 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, does not apply. When the claim limitation does not use the term “means,” examiners should determine whether the presumption that 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, paragraph 6 does not apply is overcome. The presumption may be overcome if the claim limitation uses a generic placeholder (a term that is simply a substitute for the term “means”). The following is a list of non-structural generic placeholders that may invoke 35 U.S.C. 112(f) or pre- AIA 35 U.S.C. 112, paragraph 6: “mechanism for,” “module for,” “device for,” “unit for,” “component for,” “element for,” “member for,” “apparatus for,” “machine for,” or “system for.” Welker Bearing Co., v. PHD, Inc., 550 F.3d 1090, 1096, 89 USPQ2d 1289, 1293-94 (Fed. Cir. 2008); Massachusetts Inst. of Tech. v. Abacus Software, 462 F.3d 1344, 1354, 80 USPQ2d 1225, 1228 (Fed. Cir. 2006); Personalized Media,161 F.3d at 704, 48 USPQ2d at 1886–87; Mas- Hamilton Group v. LaGard, Inc., 156 F.3d 1206, 1214-1215, 48 USPQ2d 1010, 1017 (Fed. Cir.1998). This list is not exhaustive, and other generic placeholders may invoke 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, paragraph 6. Such claim limitations are: obtaining, by a computing device, scan data and video data that each depict a scene comprising an object in independent claim 1 and dependent claims 2-10 A computing device, comprising: a memory device to store computer-readable instructions thereon; and at least one processing device configured through execution of the computer-readable instructions to: obtain scan data and video data that each depict a scene comprising an object in independent claim 11 and dependent claims 12-17 when executed by at least one computing device, directs the at least one computing device to: obtain scan data and video data that each depict a scene comprising an object in independent claim 18 and dependent claims 19-20 [0028] The computing device 102 can be embodied or implemented as, for example, a client computing device, a peripheral computing device, or both. The computing device 102, while described in the singular, may include a collection of computing devices 102. Examples of the computing device 102 can include a computer, a general-purpose computer, a special-purpose computer, a laptop, a tablet, a smartphone, another client computing device, or any combination thereof. Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 1, 4, 6-7, 11-12, 14-15, and 18-19 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Wrenninge (US20220237410A1). Regarding claim 1, Wrenninge teaches a method to label one or more objects captured in visual data ([0004] Thus, there is a need in the field to create a new and useful method of generating synthetic image datasets depicting simulated real-world imagery, that are intrinsically labeled, and efficiently cover the parameter space underlying machine learning models. This invention provides such a new and useful method), comprising: obtaining, by a computing device, scan data and video data that each depict a scene comprising an object, the scan data being generated by a scanner and the video data being generated by a camera ([0100] The simulation system 701 obtains input data 703. In some implementations, the input data 703 can include images that represent physical objects and environments. For example, the input data 703 can include one or more 3D models associated with particular properties. As another example, the input data 703 can include digital images that are captured using a camera. As another example, the input data 703 can include images that are optically digitized from analog images. As another example, the input data 703 can include images that are obtained from one or more videos… In some implementations, the input data 703 can include sensor data that are obtained from sensors such as LIDAR or radar sensors. [0019] the computing system can include a vehicle computing system (e.g., a vehicle ECU, central vehicle computer, etc.). However, the method can be otherwise implemented at any suitable computing system and/or network of computing systems); generating, by the computing device, a virtual representation of the scene in a virtual environment based on the scan data, the virtual representation comprising a virtual representation subset corresponding to the object, and the virtual environment being associated with a virtual camera ([0099] As another example, the synthetic dataset can include synthetic sensor data that simulate point clouds obtained from sensors such as radar or LIDAR sensors. [0059] Block S200 preferably produces, as output, a plurality of synthesized virtual scenes. Scene synthesis in conjunction with the method 100 preferably includes a defined model of the 3D virtual scene (e.g., determined in accordance with one or more variations of Block S100) that contains the geometric description of objects in the scene, a set of materials describing the appearance of the objects, specifications of the light sources in the scene, and a virtual camera model); applying, by the computing device, label data to the virtual representation subset to create a labeled virtual representation subset corresponding to the object ([0066] Block S300 includes generating a synthetic image of the generated scene, which functions to create a realistic synthetic two-dimensional (2D) representation of objects in the scene, wherein the objects are intrinsically labeled with the parameter values used to generate the scene (e.g., the objects in the scene, object classifications, the layout of the objects, all other parametrized metadata, etc.)…Block S300 can also function to produce realistic synthetic images with pixel-perfect ground truth annotations and/or labels (e.g., of what each object should be classified as, to any suitable level of subclassification and/or including any suitable geometric parameter, such as pose, heading, position, orientation, etc.)); and applying, by the computing device, the labeled virtual representation subset to the object depicted in the video data based on a correlation of the scanner, the camera, and the virtual camera ([0077] Block S500 can include combining synthetic and real images into a combined image dataset. [0066] The viewpoint of the virtual camera is determined as a parameter value in a variation of Block S100 (e.g., as a rendering parameter). [0045] Block S100 preferably includes determining a probability density function (PDF) for each parameter of the set of parameters, and sampling the PDF to obtain the value of the parameter (e.g., the parameter value)…The PDFs of multiple parameters can be coupled together (e.g., to form a joint PDF as shown in FIG. 2). In a first variation, the PDF of each parameter can be selected by a user (e.g., via an explicit choice by a human operator, available variables that are selected manually for each scene, object, and/or set of objects, etc.). In a second variation, the PDF of each parameter (or of a subset of the set of parameters) can be a learned function that is based on the output of training a model on the synthetic dataset, a real dataset, and/or a combination of synthetic and real data (e.g., the PDF of one or more parameters can be tuned to improve the performance of the model). [0042] Positions of each object can be defined by parameters (e.g., relative to camera position and orientation, relative to global coordinates, by specifying coordinates, heading, orientation, drag and drop, etc.)). Regarding claim 4, Wrenninge teaches the method of claim 1. Wrenninge further teaches wherein generating the virtual representation of the scene in the virtual environment based on the scan data comprises: capturing, by the computing device, a light detection and ranging (LiDAR) scan of the scene, the LiDAR scan comprising the scan data; and rendering, by the computing device, the LiDAR scan in the virtual environment to generate the virtual representation in the virtual environment based on the LiDAR scan ([0100] The simulation system 701 obtains input data 703. In some implementations, the input data 703 can include images that represent physical objects and environments…In some implementations, the input data 703 can include sensor data that are obtained from sensors such as LIDAR or radar sensors. [0101] The simulation system 701 includes a parameter set generator 705, a scenario generator 710, and a renderer 715. The parameter set generator 705 can generate parameters to process the input data 703. [0102] In some implementations, the parameters can include a plurality of scenario parameters and a plurality of rendering parameters. The scenario parameters can represent attributes that can describe certain objects in certain environments in input data. In some implementations, the scenario parameters can include one or more object attributes associated with physical features of objects that are imaged in input data. [0113] FIG. 8A is a flowchart of a method in accordance with an embodiment. In block 805, a plurality of parameters is determined, including a plurality of scenario parameters and a plurality of rendering parameters. In block 810, parameter values are determined for the scenario parameters. In block 815, a plurality of scenarios is generated based on the scenario parameters and their respective determined parameter values. In block 820, the values of the rendering parameters are determined. In block 825, a plurality of synthetic images is rendered based on the generated scenarios and the determined rendering parameters). Regarding claim 6, Wrenninge teaches the method of claim 1. Wrenninge further teaches wherein applying the label data to the virtual representation subset to create the labeled virtual representation subset corresponding to the object comprises: applying, by the computing device, the label data to the virtual representation subset to create the labeled virtual representation subset corresponding to the object based on receipt of input data that is indicative of a selection of at least one of the virtual representation subset or the object for label data annotation ([0100] The simulation system 701 obtains input data 703. In some implementations, the input data 703 can include images that represent physical objects and environments. [0103] In some implementations, the scenario parameters can include one or more environment attributes that can describe physical features of a particular environment in input data or environmental conditions in input data. [0108] The renderer 715 receives scenarios from the and generates synthetic images that correspond to the scenarios. For example, where the scenarios represent 3D scenes, the renderer 715 generates synthetic images corresponding to 3D scenes by rendering the 3D scenes based on rendering parameters. [0066] Block S300 includes generating a synthetic image of the generated scene, which functions to create a realistic synthetic two-dimensional (2D) representation of objects in the scene, wherein the objects are intrinsically labeled with the parameter values used to generate the scene (e.g., the objects in the scene, object classifications, the layout of the objects, all other parametrized metadata, etc.)…Block S300 can also function to produce realistic synthetic images with pixel-perfect ground truth annotations and/or labels (e.g., of what each object should be classified as, to any suitable level of subclassification and/or including any suitable geometric parameter, such as pose, heading, position, orientation, etc.). Regarding claim 7, Wrenninge teaches the method of claim 1. Wrenninge further teaches wherein applying the label data to the virtual representation subset to create the labeled virtual representation subset corresponding to the object comprises: extracting, by the computing device, a subset of three-dimensional (3D) point cloud data that is indicative of the virtual representation subset and the object from a 3D point cloud dataset that is indicative of the virtual representation and the scene ([0100] The simulation system 701 obtains input data 703… In some implementations, the input data 703 can include sensor data that are obtained from sensors such as LIDAR or radar sensors. For example, the input data 703 can include point clouds provided by sensors. [0101] The simulation system 701 includes a parameter set generator 705, a scenario generator 710, and a renderer 715. The parameter set generator 705 can generate parameters to process the input data 703. [0105] The scenario generator 710 can generate scenarios based on parameter values for the scenario parameters. In some implementations, the scenarios can represent 3D scenes defined by scenario parameters…In some implementations, the scenarios can represent simulated sensor data defined by sensor parameters. For example, the simulated sensor data can include a point cloud that simulates a real point cloud provided from sensors such as LIDAR and radar sensors); and applying, by the computing device, the label data to the subset of 3D point cloud data to create the labeled virtual representation subset corresponding to the object ([0108] The renderer 715 receives scenarios from the and generates synthetic images that correspond to the scenarios. For example, where the scenarios represent 3D scenes, the renderer 715 generates synthetic images corresponding to 3D scenes by rendering the 3D scenes based on rendering parameters. [0092] rendering a synthetic image of the 3D scene (e.g., made up of a set of pixels, some of which depict the object) based on the rendering parameter value; automatically labelling each pixel depicting the object within the image with a label (e.g., labeling the object with its object class, other metadata, etc.). [0109] In some implementations, the renderer 715 performs image synthesis and simulation of camera, optics, and sensors. The images and animations can include synthetic optical images, but can additionally or alternatively include other suitable data that is represented as a projection of a three-dimensional (3D) space (e.g., point cloud from LIDAR or radar)). Regarding claim 11, Wrenninge teaches a computing device, comprising: a memory device to store computer-readable instructions thereon; and at least one processing device configured through execution of the computer-readable instructions to ([0116] FIG. 9 illustrates an implementation 900 of the simulation system of FIG. 7 in which components are communicative coupled by a communication bus 910. The system 900 may include, for example, a processor 904, memory 906, database 908, input device 912, output device 914, and network interface communication unit 902. A parameter set generator 922 may include computer program instructions stored on a storage medium and executable on a processor to determine parameters and parameter values. [0019] the computing system can include a vehicle computing system (e.g., a vehicle ECU, central vehicle computer, etc.). However, the method can be otherwise implemented at any suitable computing system and/or network of computing systems): obtain scan data and video data that each depict a scene comprising an object, the scan data being generated by a scanner and the video data being generated by a camera ([0100] The simulation system 701 obtains input data 703. In some implementations, the input data 703 can include images that represent physical objects and environments. For example, the input data 703 can include one or more 3D models associated with particular properties. As another example, the input data 703 can include digital images that are captured using a camera. As another example, the input data 703 can include images that are optically digitized from analog images. As another example, the input data 703 can include images that are obtained from one or more videos… In some implementations, the input data 703 can include sensor data that are obtained from sensors such as LIDAR or radar sensors); generate a virtual representation of the scene in a virtual environment based on the scan data, the virtual representation comprising a virtual representation subset corresponding to the object, and the virtual environment being associated with a virtual camera ([0099] As another example, the synthetic dataset can include synthetic sensor data that simulate point clouds obtained from sensors such as radar or LIDAR sensors. [0059] Block S200 preferably produces, as output, a plurality of synthesized virtual scenes. Scene synthesis in conjunction with the method 100 preferably includes a defined model of the 3D virtual scene (e.g., determined in accordance with one or more variations of Block S100) that contains the geometric description of objects in the scene, a set of materials describing the appearance of the objects, specifications of the light sources in the scene, and a virtual camera model); apply label data to the virtual representation subset to create a labeled virtual representation subset corresponding to the object ([0066] Block S300 includes generating a synthetic image of the generated scene, which functions to create a realistic synthetic two-dimensional (2D) representation of objects in the scene, wherein the objects are intrinsically labeled with the parameter values used to generate the scene (e.g., the objects in the scene, object classifications, the layout of the objects, all other parametrized metadata, etc.)…Block S300 can also function to produce realistic synthetic images with pixel-perfect ground truth annotations and/or labels (e.g., of what each object should be classified as, to any suitable level of subclassification and/or including any suitable geometric parameter, such as pose, heading, position, orientation, etc.)); and apply the labeled virtual representation subset to the object depicted in the video data based on a correlation of the scanner, the camera, and the virtual camera ([0077] Block S500 can include combining synthetic and real images into a combined image dataset. [0066] The viewpoint of the virtual camera is determined as a parameter value in a variation of Block S100 (e.g., as a rendering parameter). [0045] Block S100 preferably includes determining a probability density function (PDF) for each parameter of the set of parameters, and sampling the PDF to obtain the value of the parameter (e.g., the parameter value)…The PDFs of multiple parameters can be coupled together (e.g., to form a joint PDF as shown in FIG. 2). In a first variation, the PDF of each parameter can be selected by a user (e.g., via an explicit choice by a human operator, available variables that are selected manually for each scene, object, and/or set of objects, etc.). In a second variation, the PDF of each parameter (or of a subset of the set of parameters) can be a learned function that is based on the output of training a model on the synthetic dataset, a real dataset, and/or a combination of synthetic and real data (e.g., the PDF of one or more parameters can be tuned to improve the performance of the model). [0042] Positions of each object can be defined by parameters (e.g., relative to camera position and orientation, relative to global coordinates, by specifying coordinates, heading, orientation, drag and drop, etc.)). Regarding claim 12, Wrenninge teaches the device of claim 11. Wrenninge further teaches wherein, to generate the virtual representation of the scene in the virtual environment based on the scan data, the at least one processing device is further configured to: capture a light detection and ranging (LiDAR) scan of the scene, the LiDAR scan comprising the scan data; and render the LiDAR scan in the virtual environment to generate the virtual representation in the virtual environment based on the LiDAR scan ([0100] The simulation system 701 obtains input data 703. In some implementations, the input data 703 can include images that represent physical objects and environments…In some implementations, the input data 703 can include sensor data that are obtained from sensors such as LIDAR or radar sensors. [0101] The simulation system 701 includes a parameter set generator 705, a scenario generator 710, and a renderer 715. The parameter set generator 705 can generate parameters to process the input data 703. [0102] In some implementations, the parameters can include a plurality of scenario parameters and a plurality of rendering parameters. The scenario parameters can represent attributes that can describe certain objects in certain environments in input data. In some implementations, the scenario parameters can include one or more object attributes associated with physical features of objects that are imaged in input data. [0113] FIG. 8A is a flowchart of a method in accordance with an embodiment. In block 805, a plurality of parameters is determined, including a plurality of scenario parameters and a plurality of rendering parameters. In block 810, parameter values are determined for the scenario parameters. In block 815, a plurality of scenarios is generated based on the scenario parameters and their respective determined parameter values. In block 820, the values of the rendering parameters are determined. In block 825, a plurality of synthetic images is rendered based on the generated scenarios and the determined rendering parameters). Regarding claim 14, Wrenninge teaches the device of claim 11. Wrenninge further teaches wherein, to apply the label data to the virtual representation subset to create the labeled virtual representation subset corresponding to the object, the at least one processing device is further configured to: apply the label data to the virtual representation subset to create the labeled virtual representation subset corresponding to the object based on receipt of input data that is indicative of a selection of at least one of the virtual representation subset or the object for label data annotation ([0100] The simulation system 701 obtains input data 703. In some implementations, the input data 703 can include images that represent physical objects and environments. [0103] In some implementations, the scenario parameters can include one or more environment attributes that can describe physical features of a particular environment in input data or environmental conditions in input data. [0108] The renderer 715 receives scenarios from the and generates synthetic images that correspond to the scenarios. For example, where the scenarios represent 3D scenes, the renderer 715 generates synthetic images corresponding to 3D scenes by rendering the 3D scenes based on rendering parameters. [0066] Block S300 includes generating a synthetic image of the generated scene, which functions to create a realistic synthetic two-dimensional (2D) representation of objects in the scene, wherein the objects are intrinsically labeled with the parameter values used to generate the scene (e.g., the objects in the scene, object classifications, the layout of the objects, all other parametrized metadata, etc.)…Block S300 can also function to produce realistic synthetic images with pixel-perfect ground truth annotations and/or labels (e.g., of what each object should be classified as, to any suitable level of subclassification and/or including any suitable geometric parameter, such as pose, heading, position, orientation, etc.). Regarding claim 15, Wrenninge teaches the device of claim 11. Wrenninge further teaches wherein, to apply the label data to the virtual representation subset to create the labeled virtual representation subset corresponding to the object, the at least one processing device is further configured to: extract a subset of three-dimensional (3D) point cloud data that is indicative of the virtual representation subset and the object from a 3D point cloud dataset that is indicative of the virtual representation and the scene ([0100] The simulation system 701 obtains input data 703… In some implementations, the input data 703 can include sensor data that are obtained from sensors such as LIDAR or radar sensors. For example, the input data 703 can include point clouds provided by sensors. [0101] The simulation system 701 includes a parameter set generator 705, a scenario generator 710, and a renderer 715. The parameter set generator 705 can generate parameters to process the input data 703. [0105] The scenario generator 710 can generate scenarios based on parameter values for the scenario parameters. In some implementations, the scenarios can represent 3D scenes defined by scenario parameters…In some implementations, the scenarios can represent simulated sensor data defined by sensor parameters. For example, the simulated sensor data can include a point cloud that simulates a real point cloud provided from sensors such as LIDAR and radar sensors); and apply the label data to the subset of 3D point cloud data to create the labeled virtual representation subset corresponding to the object ([0108] The renderer 715 receives scenarios from the and generates synthetic images that correspond to the scenarios. For example, where the scenarios represent 3D scenes, the renderer 715 generates synthetic images corresponding to 3D scenes by rendering the 3D scenes based on rendering parameters. [0092] rendering a synthetic image of the 3D scene (e.g., made up of a set of pixels, some of which depict the object) based on the rendering parameter value; automatically labelling each pixel depicting the object within the image with a label (e.g., labeling the object with its object class, other metadata, etc.). [0109] In some implementations, the renderer 715 performs image synthesis and simulation of camera, optics, and sensors. The images and animations can include synthetic optical images, but can additionally or alternatively include other suitable data that is represented as a projection of a three-dimensional (3D) space (e.g., point cloud from LIDAR or radar)). Regarding claim 18, Wrenninge teaches a non-transitory computer-readable medium embodying at least one program that, when executed by at least one computing device, directs the at least one computing device to ([0116] FIG. 9 illustrates an implementation 900 of the simulation system of FIG. 7 in which components are communicative coupled by a communication bus 910. The system 900 may include, for example, a processor 904, memory 906, database 908, input device 912, output device 914, and network interface communication unit 902. A parameter set generator 922 may include computer program instructions stored on a storage medium and executable on a processor to determine parameters and parameter values. [0019] the computing system can include a vehicle computing system (e.g., a vehicle ECU, central vehicle computer, etc.). However, the method can be otherwise implemented at any suitable computing system and/or network of computing systems): obtain scan data and video data that each depict a scene comprising an object, the scan data being generated by a scanner and the video data being generated by a camera ([0100] The simulation system 701 obtains input data 703. In some implementations, the input data 703 can include images that represent physical objects and environments. For example, the input data 703 can include one or more 3D models associated with particular properties. As another example, the input data 703 can include digital images that are captured using a camera. As another example, the input data 703 can include images that are optically digitized from analog images. As another example, the input data 703 can include images that are obtained from one or more videos… In some implementations, the input data 703 can include sensor data that are obtained from sensors such as LIDAR or radar sensors); generate a virtual representation of the scene in a virtual environment based on the scan data, the virtual representation comprising a virtual representation subset corresponding to the object, and the virtual environment being associated with a virtual camera ([0099] As another example, the synthetic dataset can include synthetic sensor data that simulate point clouds obtained from sensors such as radar or LIDAR sensors. [0059] Block S200 preferably produces, as output, a plurality of synthesized virtual scenes. Scene synthesis in conjunction with the method 100 preferably includes a defined model of the 3D virtual scene (e.g., determined in accordance with one or more variations of Block S100) that contains the geometric description of objects in the scene, a set of materials describing the appearance of the objects, specifications of the light sources in the scene, and a virtual camera model); apply label data to the virtual representation subset to create a labeled virtual representation subset corresponding to the object ([0066] Block S300 includes generating a synthetic image of the generated scene, which functions to create a realistic synthetic two-dimensional (2D) representation of objects in the scene, wherein the objects are intrinsically labeled with the parameter values used to generate the scene (e.g., the objects in the scene, object classifications, the layout of the objects, all other parametrized metadata, etc.)…Block S300 can also function to produce realistic synthetic images with pixel-perfect ground truth annotations and/or labels (e.g., of what each object should be classified as, to any suitable level of subclassification and/or including any suitable geometric parameter, such as pose, heading, position, orientation, etc.)); and apply the labeled virtual representation subset to the object depicted in the video data based on a correlation of the scanner, the camera, and the virtual camera ([0077] Block S500 can include combining synthetic and real images into a combined image dataset. [0066] The viewpoint of the virtual camera is determined as a parameter value in a variation of Block S100 (e.g., as a rendering parameter). [0045] Block S100 preferably includes determining a probability density function (PDF) for each parameter of the set of parameters, and sampling the PDF to obtain the value of the parameter (e.g., the parameter value)…The PDFs of multiple parameters can be coupled together (e.g., to form a joint PDF as shown in FIG. 2). In a first variation, the PDF of each parameter can be selected by a user (e.g., via an explicit choice by a human operator, available variables that are selected manually for each scene, object, and/or set of objects, etc.). In a second variation, the PDF of each parameter (or of a subset of the set of parameters) can be a learned function that is based on the output of training a model on the synthetic dataset, a real dataset, and/or a combination of synthetic and real data (e.g., the PDF of one or more parameters can be tuned to improve the performance of the model). [0042] Positions of each object can be defined by parameters (e.g., relative to camera position and orientation, relative to global coordinates, by specifying coordinates, heading, orientation, drag and drop, etc.)). Regarding claim 19, Wrenninge teaches the medium of claim 18. Wrenninge further teaches wherein, to generate the virtual representation of the scene in the virtual environment based on the scan data, the at least one computing device is further configured to: capture a light detection and ranging (LiDAR) scan of the scene, the LiDAR scan comprising the scan data; and render the LiDAR scan in the virtual environment to generate the virtual representation in the virtual environment based on the LiDAR scan ([0100] The simulation system 701 obtains input data 703. In some implementations, the input data 703 can include images that represent physical objects and environments…In some implementations, the input data 703 can include sensor data that are obtained from sensors such as LIDAR or radar sensors. [0101] The simulation system 701 includes a parameter set generator 705, a scenario generator 710, and a renderer 715. The parameter set generator 705 can generate parameters to process the input data 703. [0102] In some implementations, the parameters can include a plurality of scenario parameters and a plurality of rendering parameters. The scenario parameters can represent attributes that can describe certain objects in certain environments in input data. In some implementations, the scenario parameters can include one or more object attributes associated with physical features of objects that are imaged in input data. [0113] FIG. 8A is a flowchart of a method in accordance with an embodiment. In block 805, a plurality of parameters is determined, including a plurality of scenario parameters and a plurality of rendering parameters. In block 810, parameter values are determined for the scenario parameters. In block 815, a plurality of scenarios is generated based on the scenario parameters and their respective determined parameter values. In block 820, the values of the rendering parameters are determined. In block 825, a plurality of synthetic images is rendered based on the generated scenarios and the determined rendering parameters). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 2-3 are rejected under 35 U.S.C. 103 as being unpatentable over Wrenninge in view of Yang (US11676351B1). Regarding claim 2, Wrenninge teaches the method of claim 1. Wrenninge further teaches generating, by the computing device, a labeled visual dataset based on the video data and the labeled virtual representation subset corresponding to the object ([0076] Block S500 includes generating a synthetic image dataset. Block S500 functions to output a synthetic image dataset, made up of intrinsically labelled images, as a result of the procedural generation, rendering, and/or augmentation Blocks of the method, for use in downstream applications (e.g., model training, model evaluation, image capture methodology validation, etc.). Block S500 can thus also include combining a plurality of images into dataset (e.g., made up of the plurality of images). [0077] Block S500 can include combining synthetic and real images into a combined image dataset. [0023] Second, variants of the method result in labels of objects (and poses of such objects) within synthetic images that are inherently “perfect” (e.g., as accurate as possible) without human or other manual intervention (e.g., human-derived object or pose labels), because object types, layouts, relative orientations and positions are deterministic and known due to programmatic (e.g., procedural) generation of the virtual scene containing the objects. For example, hand-annotated data in conventional, manually-generated labeled datasets can fail to train ML models (e.g., networks) to recognize objects that are not correctly annotated in the ground-truth datasets (e.g., the hand-annotated datasets), whereas an intrinsically-labeled, procedurally generated synthetic dataset generated in accordance with variants of the method 100 are programmatically prevented from containing annotation errors. FIG. 5 depicts an example of pixel-by-pixel segmentation (e.g., intrinsic labeling) of synthetic images based on the underlying objects and/or groups of objects). Wrenninge does not explicitly teach the labeled visual dataset comprising an annotation of the labeled virtual representation subset applied to the object in one or more video frames of the video data. Yang, in the same field of endeavor of image annotating, teaches the labeled visual dataset comprising an annotation of the labeled virtual representation subset applied to the object in one or more video frames of the video data ([col. 17 ln. 23-36] Once a virtual object model is generated for the real-world object, a subject matter expert (SME) may annotate selected points of the virtual object model for the particular purpose desired and provide annotation data that the AR applications can render in their AR representations of the real-world environment to facilitate understanding of the real-world environment, performance of tasks in the real-world environment, or otherwise augment the AR representation of the real-world environment with additional digital content. The annotation data may be linked to the selected annotation points such that the annotation data is made part of the virtual object model and associated with the points at coordinate locations corresponding to the annotation point coordinates. [col. 21 ln. 10-15] Assuming that the correlation between points in the VO model(s) M1 are accurate with the at least one second model of the currently viewed real-world environment M2, then the annotations and/or overlays of the virtual objects in the AR representation, will be accurately shown in the output of the AR application engine 220. [col. 12 ln. 7-10] The AR model M3 is used by the AR application to superimpose or otherwise overlay AR model M3 elements on captured images or video of the currently viewed real-world environment). Therefore, it would have been obvious to a person of ordinary skill in the art before the time of filing to modify the method of Wrenninge with the teachings of Yang for the labeled visual dataset to comprise an annotation of the labeled virtual representation subset applied to the object in one or more video frames of the video data because "Thus, the human operator may interact with the AR application's, such as via the GUI generated by the GUI engine 210, to follow the procedure or perform operations specified in the annotations with regard to real-world objects to thereby perform a desired task defined by the combination of annotations" [col. 21 ln. 15-20]. Regarding claim 3, Wrenninge and Yang teach the method of claim 2. Wrenninge further teaches training, by the computing device, a model to detect the object in different visual data based on the labeled visual dataset, the different visual data depicting a different scene comprising the object ([0080] The method can include Block S600, which includes training a model based on the synthetic image dataset. Block S600 functions to train a learning model (e.g., an ML model, a synthetic neural network, a computational network, etc.) using supervised learning, based on the intrinsically labeled synthetic image dataset. Block S600 can function to modify a model to recognize objects (e.g., classify objects, detect objects, etc.) depicted in images with an improved accuracy (e.g., as compared to an initial accuracy, a threshold accuracy, a baseline accuracy, etc.). [0135] In some implementations, the previously described methods may be used to generate a distribution of variations in the synthetic image dataset for specific training or evaluation purposes. For example, there are real-world scenarios in which accurate object detection by a perception system is more difficult. For example, at higher ego-vehicle speeds, there may be more motion blur for certain objects, such as fences by the side of a road. As another example, object detection may be more difficult in illumination conditions for which there is less optical contrast. However, more generally, there may be a variety of scenarios for which it may be useful to generate training data, such as generating training data for particular object classes (e.g., specific actor vehicles such as bicycles, cars, trucks, buses, motorcycles; environmental features such as buildings, walls, vegetation; road width and road surface), specific combinations of objects, specific illumination conditions, etc. As another example, curb height and sidewalk width may vary in different driving environment, such that generating synthetic data for images having differences in curb height and sidewalk width may be desirable. [0137] For example, the ego-vehicle speed and a selection of static and non-static objects in a scene may be varied to train a machine learning model based on scenarios in which motion blur of at least one class of objects is more likely to occur. For example, objects may be included in scenes that have attributes making their detection more susceptible to motion blur). Claims 5, 13, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Wrenninge in view of Eustice (US20160209846A1). Regarding claim 5, Wrenninge teaches the method of claim 1. Wrenninge does not explicitly teach wherein generating the virtual representation of the scene in the virtual environment based on the scan data comprises: applying, by the computing device, a smoothing and mapping (SAM) algorithm to the scan data; and tracking, by the computing device, three-dimensional (3D) location data corresponding to at least one of the object or the virtual representation subset with respect to at least one vantage point of the virtual camera based on applying the SAM algorithm to the scan data. Eustice, in the same field of endeavor of scanning data analysis, teaches wherein generating the virtual representation of the scene in the virtual environment based on the scan data comprises: applying, by the computing device, a smoothing and mapping (SAM) algorithm to the scan data ([0046] We use the state-of-the-art in nonlinear least-squares, pose-graph SLAM and measurements from our survey vehicle's 3D LIDAR scanners to produce a map of the 3D structure in a self-consistent frame. We construct a pose-graph to solve the full SLAM problem, as shown in FIG. 2, where nodes in the graph are poses (X) and edges are either odometry constraints (U), laser scan-matching constraints (Z), or GPS prior constraints (G). These constraints are modeled as Gaussian random variables; resulting in a nonlinear least-squares optimization problem that we solve with the incremental smoothing and mapping (iSAM) algorithm); and tracking, by the computing device, three-dimensional (3D) location data corresponding to at least one of the object or the virtual representation subset with respect to at least one vantage point of the virtual camera based on applying the SAM algorithm to the scan data ([0009] According to the present teachings, 3D prior maps (augmented with surface reflectivities) constructed by a survey vehicle equipped with 3D LIDAR scanners are used for vehicle automation. A vehicle is localized by comparing imagery from a monocular camera against several candidate views, seeking to maximize normalized mutual information (NMI) (as outlined in FIG. 1). [0074] In collecting each dataset, we made two passes through the same environment (on separate days) and aligned the two together using our offline SLAM procedure outlined in §III. This allowed us to build a prior map ground-mesh on the first pass through the environment. Then, the subsequent pass would be well localized with respect to the ground-mesh, providing sufficiently accurate ground-truth in the experiment. [0008] The present teachings leverage a graphics processing unit (GPU) so that one can generate several synthetic, pin-hole camera images, which we can then directly compare against streaming vehicle imagery. [0054] Given a query camera pose parameterized as [R|t], where R and t are the camera's rotation and translation, respectively, our goal is to provide a synthetic view of our world from that vantage point. We Use OpenGL, which is commonly used for visualization utilities, in a robotics context to simulate a pin-hole camera model. [0089] By maximizing normalized mutual information, we are able to register a video stream to our prior map. Our system is aided by a GPU implementation, leveraging OpenGL to generate synthetic views of the environment; this implementation is able to provide corrective positional updates at ˜10 Hz). Therefore, it would have been obvious to a person of ordinary skill in the art before the time of filing to modify the method of Wrenninge with the teachings of Eustice to apply a SAM algorithm to the scan data and track 3D location data corresponding to the virtual representation subset with respect to at least one vantage point of the virtual camera because "The graphics processing unit accesses a database of prior map information and generates a synthetic image that is then compared to the real-time visual camera data to determine corrected position data…A corrective system for applying navigation of the vehicle based on the determined camera position can be used in some embodiments" [Abstract] and "This significantly simpler approach avoids over-engineering the problem by formulating a slightly more computationally expensive solution that is still real-time tractable on a mobile-grade GPU and capable of high accuracy localization" [0008]. Regarding claim 13, Wrenninge teaches the device of claim 11. Wrenninge does not explicitly teach wherein, to generate the virtual representation of the scene in the virtual environment based on the scan data, the at least one processing device is further configured to: apply a smoothing and mapping (SAM) algorithm to the scan data; and track three-dimensional (3D) location data corresponding to at least one of the object or the virtual representation subset with respect to at least one vantage point of the virtual camera based on applying the SAM algorithm to the scan data. Eustice, in the same field of endeavor of scanning data analysis, teaches wherein, to generate the virtual representation of the scene in the virtual environment based on the scan data, the at least one processing device is further configured to: apply a smoothing and mapping (SAM) algorithm to the scan data ([0046] We use the state-of-the-art in nonlinear least-squares, pose-graph SLAM and measurements from our survey vehicle's 3D LIDAR scanners to produce a map of the 3D structure in a self-consistent frame. We construct a pose-graph to solve the full SLAM problem, as shown in FIG. 2, where nodes in the graph are poses (X) and edges are either odometry constraints (U), laser scan-matching constraints (Z), or GPS prior constraints (G). These constraints are modeled as Gaussian random variables; resulting in a nonlinear least-squares optimization problem that we solve with the incremental smoothing and mapping (iSAM) algorithm); and track three-dimensional (3D) location data corresponding to at least one of the object or the virtual representation subset with respect to at least one vantage point of the virtual camera based on applying the SAM algorithm to the scan data ([0009] According to the present teachings, 3D prior maps (augmented with surface reflectivities) constructed by a survey vehicle equipped with 3D LIDAR scanners are used for vehicle automation. A vehicle is localized by comparing imagery from a monocular camera against several candidate views, seeking to maximize normalized mutual information (NMI) (as outlined in FIG. 1). [0074] In collecting each dataset, we made two passes through the same environment (on separate days) and aligned the two together using our offline SLAM procedure outlined in §III. This allowed us to build a prior map ground-mesh on the first pass through the environment. Then, the subsequent pass would be well localized with respect to the ground-mesh, providing sufficiently accurate ground-truth in the experiment. [0008] The present teachings leverage a graphics processing unit (GPU) so that one can generate several synthetic, pin-hole camera images, which we can then directly compare against streaming vehicle imagery. [0054] Given a query camera pose parameterized as [R|t], where R and t are the camera's rotation and translation, respectively, our goal is to provide a synthetic view of our world from that vantage point. We Use OpenGL, which is commonly used for visualization utilities, in a robotics context to simulate a pin-hole camera model. [0089] By maximizing normalized mutual information, we are able to register a video stream to our prior map. Our system is aided by a GPU implementation, leveraging OpenGL to generate synthetic views of the environment; this implementation is able to provide corrective positional updates at ˜10 Hz). Therefore, it would have been obvious to a person of ordinary skill in the art before the time of filing to modify the device of Wrenninge with the teachings of Eustice to apply a SAM algorithm to the scan data and track 3D location data corresponding to the virtual representation subset with respect to at least one vantage point of the virtual camera because "The graphics processing unit accesses a database of prior map information and generates a synthetic image that is then compared to the real-time visual camera data to determine corrected position data…A corrective system for applying navigation of the vehicle based on the determined camera position can be used in some embodiments" [Abstract] and "This significantly simpler approach avoids over-engineering the problem by formulating a slightly more computationally expensive solution that is still real-time tractable on a mobile-grade GPU and capable of high accuracy localization" [0008]. Regarding claim 20, Wrenninge teaches the medium of claim 18. Wrenninge does not explicitly teach wherein, to generate the virtual representation of the scene in the virtual environment based on the scan data, the at least one computing device is further configured to: apply a smoothing and mapping (SAM) algorithm to the scan data; and track three-dimensional (3D) location data corresponding to at least one of the object or the virtual representation subset with respect to at least one vantage point of the virtual camera based on applying the SAM algorithm to the scan data. Eustice, in the same field of endeavor of scanning data analysis, teaches wherein, to generate the virtual representation of the scene in the virtual environment based on the scan data, the at least one computing device is further configured to: apply a smoothing and mapping (SAM) algorithm to the scan data ([0046] We use the state-of-the-art in nonlinear least-squares, pose-graph SLAM and measurements from our survey vehicle's 3D LIDAR scanners to produce a map of the 3D structure in a self-consistent frame. We construct a pose-graph to solve the full SLAM problem, as shown in FIG. 2, where nodes in the graph are poses (X) and edges are either odometry constraints (U), laser scan-matching constraints (Z), or GPS prior constraints (G). These constraints are modeled as Gaussian random variables; resulting in a nonlinear least-squares optimization problem that we solve with the incremental smoothing and mapping (iSAM) algorithm); and track three-dimensional (3D) location data corresponding to at least one of the object or the virtual representation subset with respect to at least one vantage point of the virtual camera based on applying the SAM algorithm to the scan data ([0009] According to the present teachings, 3D prior maps (augmented with surface reflectivities) constructed by a survey vehicle equipped with 3D LIDAR scanners are used for vehicle automation. A vehicle is localized by comparing imagery from a monocular camera against several candidate views, seeking to maximize normalized mutual information (NMI) (as outlined in FIG. 1). [0074] In collecting each dataset, we made two passes through the same environment (on separate days) and aligned the two together using our offline SLAM procedure outlined in §III. This allowed us to build a prior map ground-mesh on the first pass through the environment. Then, the subsequent pass would be well localized with respect to the ground-mesh, providing sufficiently accurate ground-truth in the experiment. [0008] The present teachings leverage a graphics processing unit (GPU) so that one can generate several synthetic, pin-hole camera images, which we can then directly compare against streaming vehicle imagery. [0054] Given a query camera pose parameterized as [R|t], where R and t are the camera's rotation and translation, respectively, our goal is to provide a synthetic view of our world from that vantage point. We Use OpenGL, which is commonly used for visualization utilities, in a robotics context to simulate a pin-hole camera model. [0089] By maximizing normalized mutual information, we are able to register a video stream to our prior map. Our system is aided by a GPU implementation, leveraging OpenGL to generate synthetic views of the environment; this implementation is able to provide corrective positional updates at ˜10 Hz). Therefore, it would have been obvious to a person of ordinary skill in the art before the time of filing to modify the medium of Wrenninge with the teachings of Eustice to apply a SAM algorithm to the scan data and track 3D location data corresponding to the virtual representation subset with respect to at least one vantage point of the virtual camera because "The graphics processing unit accesses a database of prior map information and generates a synthetic image that is then compared to the real-time visual camera data to determine corrected position data…A corrective system for applying navigation of the vehicle based on the determined camera position can be used in some embodiments" [Abstract] and "This significantly simpler approach avoids over-engineering the problem by formulating a slightly more computationally expensive solution that is still real-time tractable on a mobile-grade GPU and capable of high accuracy localization" [0008]. Claims 8 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Wrenninge in view of Wang (US20230135234A1). Regarding claim 8, Wrenninge teaches the method of claim 1. Wrenninge further teaches wherein applying the label data to the virtual representation subset to create the labeled virtual representation subset corresponding to the object comprises: generating, by the computing device, a bounding box around the virtual representation subset ([0070] In alternative variations, labelling can be performed on a basis other than a pixel-by-pixel basis; for example, labelling can include automatically generating a bounding box around objects depicted in the image, a bounding polygon of any other suitable shape, a centroid point, a silhouette or outline, a floating label, and any other suitable annotation, wherein the annotation includes label metadata such as the object class and/or other object metadata); and associating, by the computing device, metadata indicative of the object with at least one of the virtual representation subset, the vertex annotation, or the bounding box ([0123] In some implementations, instance metadata is generated for each instance of a non-static class (e.g., pedestrians, riders, cars, truck, buses, trains, motorcycles, and bicycles). In some implementations, the instance metadata may specific bounding boxes, class information, a fractional occlusion, and a fractional truncation). Wrenninge does not explicitly teach applying, by the computing device, a vertex annotation along vertices of the virtual representation subset. Wang, in the same field of endeavor of image annotation, teaches applying, by the computing device, a vertex annotation along vertices of the virtual representation subset ([0132] In some embodiments, the LiDAR data may be annotated to identify points on a 3D surface of interest (e.g., a 3D road surface). Generally, annotations may be synthetically produced (e.g., generated from computer models or renderings)…and/or a combination thereof (e.g., a human identifies vertices of polylines, a machine generates polygons using polygon rasterizer). Therefore, it would have been obvious to a person of ordinary skill in the art before the time of filing to modify the method of Wrenninge with the teachings of Wang to apply a vertex annotation "to identify 3D points on the 3D surface of interest, and the identified 3D points may be projected to generate a dense representation of the 3D surface structure" [Abstract]. Regarding claim 16, Wrenninge teaches the device of claim 11. Wrenninge further teaches wherein, to apply the label data to the virtual representation subset to create the labeled virtual representation subset corresponding to the object, the at least one processing device is further configured to: generate a bounding box around the virtual representation subset ([0070] In alternative variations, labelling can be performed on a basis other than a pixel-by-pixel basis; for example, labelling can include automatically generating a bounding box around objects depicted in the image, a bounding polygon of any other suitable shape, a centroid point, a silhouette or outline, a floating label, and any other suitable annotation, wherein the annotation includes label metadata such as the object class and/or other object metadata); and associate metadata indicative of the object with at least one of the virtual representation subset, the vertex annotation, or the bounding box ([0123] In some implementations, instance metadata is generated for each instance of a non-static class (e.g., pedestrians, riders, cars, truck, buses, trains, motorcycles, and bicycles). In some implementations, the instance metadata may specific bounding boxes, class information, a fractional occlusion, and a fractional truncation). Wrenninge does not explicitly teach apply a vertex annotation along vertices of the virtual representation subset. Wang, in the same field of endeavor of image annotation, teaches apply a vertex annotation along vertices of the virtual representation subset ([0132] In some embodiments, the LiDAR data may be annotated to identify points on a 3D surface of interest (e.g., a 3D road surface). Generally, annotations may be synthetically produced (e.g., generated from computer models or renderings)…and/or a combination thereof (e.g., a human identifies vertices of polylines, a machine generates polygons using polygon rasterizer). Therefore, it would have been obvious to a person of ordinary skill in the art before the time of filing to modify the device of Wrenninge with the teachings of Wang to apply a vertex annotation "to identify 3D points on the 3D surface of interest, and the identified 3D points may be projected to generate a dense representation of the 3D surface structure" [Abstract]. Claims 9-10 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Wrenninge in view of Versace (US20170024877A1). Regarding claim 9, Wrenninge teaches the method of claim 1. Wrenninge does not explicitly teach wherein applying the labeled virtual representation subset to the object depicted in the video data based on the correlation of the scanner, the camera, and the virtual camera comprises: applying, by the computing device, the labeled virtual representation subset to the object depicted in one or more video frames of the video data based on a correlation of pose data respectively corresponding to the scanner, the camera, and the virtual camera. Versace, in the same field of endeavor of image sensor analysis, teaches wherein applying the labeled virtual representation subset to the object depicted in the video data based on the correlation of the scanner, the camera, and the virtual camera comprises: applying, by the computing device, the labeled virtual representation subset to the object depicted in one or more video frames of the video data based on a correlation of pose data respectively corresponding to the scanner, the camera, and the virtual camera ([0030] The technology described herein provide a unified mechanism for identifying, learning, localizing, and tracking objects in an arbitrary sensory system, including data streams derived from static/pan-tilt cameras (e.g., red-green-blue (RGB) cameras, or other cameras), wireless sensors (e.g., Bluetooth), multi-array microphones, depth sensors, infrared (IR) sensors (e.g., IR laser projectors), monochrome or color CMOS sensors, and mobile robots with similar or other sensors packs (e.g., LIDAR, IR, RADAR), and virtual sensors in virtual environments (e.g., video games or simulated reality), or other networks of sensors. [0049] OpenEye may be used in both artificial environments (e.g., synthetically generated environments via video-game engine) and natural environments. OpenEye learns incrementally about its visual input, and identifies and categorizes object identities and object positions. OpenEye can operate with or without supervision—it does not require a manual labeling of object(s) of interest to learn object identity. OpenEye can accept user input to verbally label objects. [0092] FIGS. 2A-2C depicts aspects of the What and Where systems shown in FIG. 1 for an OpenSense architecture that processes visual data (aka an OpenEye system). FIG. 2A shows the Environment Module (120) and the Where System (130), which collectively constitute the Where Pathway (140). The environment module 120 includes an RGB image sensor 100, which may acquire still and/or video images. [0096] The object layer 290 uses the representations from the view layer 280 to learn pose-invariant object representations by associating different view prototypes from the view layer 280 according to their temporal continuity provided by the reset signal from the Where system 130. This yields an identity confidence measure, which can be fed into a name layer 300 that groups different objects under the same user label. [0047] Although FIG. 1 illustrates OpenSense system with three sensory inputs, the OpenSense system can be generalized to arbitrary numbers and types of sensory inputs (e.g., static/pan-tilt cameras, wireless sensors, multi-array microphone, depth sensors, IR laser projectors, monochrome CMOS sensors, and mobile robots with similar or other sensors packs—e.g., LIDAR, IR, RADAR, and virtual sensors in virtual environments). Therefore, it would have been obvious to a person of ordinary skill in the art before the time of filing to modify the method of Wrenninge with the teachings of Versace to apply the labeled virtual representation subset to the object depicted in the video data based on the correlation of pose data of the scanner, the camera, and the virtual camera because "For a mobile robot to operate autonomously, it should be able to learn about, locate, and possibly avoid objects as it moves within its environment. For example, a ground mobile/air/underwater robot may acquire images of its environment, process them to identify and locate objects, then plot a path around the objects identified in the images. Additionally, such learned objects may be located in a map (e.g., a world-centric, or allocentric human-readable map) for further retrieval in the future, or to provide additional information of what is preset in the environment to the user. In some cases, a mobile robot may include multiple cameras, e.g., to acquire sterescopic image data that can be used to estimate the range to certain items within its field of view. A mobile robot may also use other sensors, such as RADAR or LIDAR, to acquire additional data about its environment… To date, however, sensory processing of visual, auditory, and other sensor information (e.g., LIDAR, RADAR) is conventionally based on “stovepiped,” or isolated processing, with little interactions between modules. For this reason, continuous fusion and learning of pertinent information has been an issue" [0003-0004]. Regarding claim 10, Wrenninge teaches the method of claim 1. Wrenninge does not explicitly teach mapping, by the computing device, first time series pose data of the scanner to second time series pose data of the camera and third time series pose data of the virtual camera to correlate the scanner, the camera, and the virtual camera. Versace, in the same field of endeavor of image sensor analysis, teaches mapping, by the computing device, first time series pose data of the scanner to second time series pose data of the camera and third time series pose data of the virtual camera to correlate the scanner, the camera, and the virtual camera ([0030] The technology described herein provide a unified mechanism for identifying, learning, localizing, and tracking objects in an arbitrary sensory system, including data streams derived from static/pan-tilt cameras (e.g., red-green-blue (RGB) cameras, or other cameras), wireless sensors (e.g., Bluetooth), multi-array microphones, depth sensors, infrared (IR) sensors (e.g., IR laser projectors), monochrome or color CMOS sensors, and mobile robots with similar or other sensors packs (e.g., LIDAR, IR, RADAR), and virtual sensors in virtual environments (e.g., video games or simulated reality), or other networks of sensors. [0041] The joint Where system tells one or more of the other sensory modules in the OpenSense framework (auditory, RADAR, etc): “all focus at x=12, y=32, z=31.” The auditory system responds to this command by suppressing anything in the auditory data stream that is not in x=12, y=32, z=31, e.g., by using Interaural Time Differences (ITD) to pick up signals from one location, and suppress signals from other locations. Similarly, the RADAR system may focus only on data acquired from sources at or near x=12, y=32, z=31, e.g., by processing returns from one or more appropriate azimuths, elevations, and/or range bins. [0096] The object layer 290 uses the representations from the view layer 280 to learn pose-invariant object representations by associating different view prototypes from the view layer 280 according to their temporal continuity provided by the reset signal from the Where system 130). Therefore, it would have been obvious to a person of ordinary skill in the art before the time of filing to modify the method of Wrenninge with the teachings of Versace to map time series pose data of the scanner, camera, and virtual camera because "To date, however, sensory processing of visual, auditory, and other sensor information (e.g., LIDAR, RADAR) is conventionally based on “stovepiped,” or isolated processing, with little interactions between modules. For this reason, continuous fusion and learning of pertinent information has been an issue. Additionally, learning has been treated mostly as an off-line method, which happens in a separate time frame with respect to performance of tasks by the robot." [0004]. Regarding claim 17, Wrenninge teaches the device of claim 11. Wrenninge does not explicitly teach wherein, to apply the labeled virtual representation subset to the object depicted in the video data based on the correlation of the scanner, the camera, and the virtual camera, the at least one processing device is further configured to: apply the labeled virtual representation subset to the object depicted in one or more video frames of the video data based on a correlation of pose data respectively corresponding to the scanner, the camera, and the virtual camera. Versace, in the same field of endeavor of image sensor analysis, teaches wherein, to apply the labeled virtual representation subset to the object depicted in the video data based on the correlation of the scanner, the camera, and the virtual camera, the at least one processing device is further configured to: apply the labeled virtual representation subset to the object depicted in one or more video frames of the video data based on a correlation of pose data respectively corresponding to the scanner, the camera, and the virtual camera ([0030] The technology described herein provide a unified mechanism for identifying, learning, localizing, and tracking objects in an arbitrary sensory system, including data streams derived from static/pan-tilt cameras (e.g., red-green-blue (RGB) cameras, or other cameras), wireless sensors (e.g., Bluetooth), multi-array microphones, depth sensors, infrared (IR) sensors (e.g., IR laser projectors), monochrome or color CMOS sensors, and mobile robots with similar or other sensors packs (e.g., LIDAR, IR, RADAR), and virtual sensors in virtual environments (e.g., video games or simulated reality), or other networks of sensors. [0049] OpenEye may be used in both artificial environments (e.g., synthetically generated environments via video-game engine) and natural environments. OpenEye learns incrementally about its visual input, and identifies and categorizes object identities and object positions. OpenEye can operate with or without supervision—it does not require a manual labeling of object(s) of interest to learn object identity. OpenEye can accept user input to verbally label objects. [0092] FIGS. 2A-2C depicts aspects of the What and Where systems shown in FIG. 1 for an OpenSense architecture that processes visual data (aka an OpenEye system). FIG. 2A shows the Environment Module (120) and the Where System (130), which collectively constitute the Where Pathway (140). The environment module 120 includes an RGB image sensor 100, which may acquire still and/or video images. [0096] The object layer 290 uses the representations from the view layer 280 to learn pose-invariant object representations by associating different view prototypes from the view layer 280 according to their temporal continuity provided by the reset signal from the Where system 130. This yields an identity confidence measure, which can be fed into a name layer 300 that groups different objects under the same user label. [0047] Although FIG. 1 illustrates OpenSense system with three sensory inputs, the OpenSense system can be generalized to arbitrary numbers and types of sensory inputs (e.g., static/pan-tilt cameras, wireless sensors, multi-array microphone, depth sensors, IR laser projectors, monochrome CMOS sensors, and mobile robots with similar or other sensors packs—e.g., LIDAR, IR, RADAR, and virtual sensors in virtual environments). Therefore, it would have been obvious to a person of ordinary skill in the art before the time of filing to modify the device of Wrenninge with the teachings of Versace to apply the labeled virtual representation subset to the object depicted in the video data based on the correlation of pose data of the scanner, the camera, and the virtual camera because "For a mobile robot to operate autonomously, it should be able to learn about, locate, and possibly avoid objects as it moves within its environment. For example, a ground mobile/air/underwater robot may acquire images of its environment, process them to identify and locate objects, then plot a path around the objects identified in the images. Additionally, such learned objects may be located in a map (e.g., a world-centric, or allocentric human-readable map) for further retrieval in the future, or to provide additional information of what is preset in the environment to the user. In some cases, a mobile robot may include multiple cameras, e.g., to acquire sterescopic image data that can be used to estimate the range to certain items within its field of view. A mobile robot may also use other sensors, such as RADAR or LIDAR, to acquire additional data about its environment… To date, however, sensory processing of visual, auditory, and other sensor information (e.g., LIDAR, RADAR) is conventionally based on “stovepiped,” or isolated processing, with little interactions between modules. For this reason, continuous fusion and learning of pertinent information has been an issue" [0003-0004]. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Gupta (US11417069B1) teaches object labeling and localization using cameras, scanners, and virtual cameras. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jacqueline R Zak whose telephone number is (571)272-4077. The examiner can normally be reached M-F 9-5. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Emily Terrell can be reached at (571) 270-3717. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JACQUELINE R ZAK/Examiner, Art Unit 2666 /EMILY C TERRELL/Supervisory Patent Examiner, Art Unit 2666
Read full office action

Prosecution Timeline

Oct 25, 2024
Application Filed
Aug 18, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737878
BONE FRACTURE RISK PREDICTION USING LOW-RESOLUTION CLINICAL COMPUTED TOMOGRAPHY (CT) SCANS
3y 11m to grant Granted Sep 15, 2026
Patent 12711760
METHODS AND SYSTEMS FOR USE IN PROCESSING IMAGES RELATED TO CROP PHENOLOGY
3y 10m to grant Granted Aug 18, 2026
Patent 12705755
PRECISE BOOK PAGE BOUNDARY DETECTION USING DEEP MACHINE LEARNING MODEL AND IMAGE PROCESSING ALGORITHMS
2y 11m to grant Granted Aug 11, 2026
Patent 12652373
IMAGE PROCESSING METHOD AND APPARATUS, DEVICE, AND MEDIUM
3y 2m to grant Granted Jun 09, 2026
Patent 12644773
TEMPERATURE CONTROL SYSTEM, TEMPERATURE CONTROL METHOD AND TEMPERATURE CONTROL PROGRAM FOR FACILITY EQUIPMENT
3y 6m to grant Granted Jun 02, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
64%
Grant Probability
77%
With Interview (+12.7%)
3y 2m (~1y 3m remaining)
Median Time to Grant
Low
PTA Risk
Based on 36 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month