DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Joint Inventors
This application currently names joint inventors. In considering patentability of the claims, the Examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the Examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Information Disclosure Statement
The information disclosure statements (IDS) submitted on 04/09/2025 and 07/08/2025, were filed before the mailing of a First Office Action on the Merits. The submissions are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statements are being considered by the examiner.
Priority/Benefit
Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55. The instant application is a 371 national stage of PCT/EP2023/075514, which has an effective filing date of 09/15/2023, as well as also claiming benefit to provisional 63/407,129, which has an effective filing date of 09/15/2022. The examiner has checked and verified that the subject matter of the instant application is supported by the earlier filed provisional, and as such, the earlier filed date of 09/15/2022 is granted.
Status of Claims
This action is in response to Applicant’s filing on 03/14/2025. Claims 1-20 are pending and examined below.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-4, 6-9, 12, 16-18, and 20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Toshev et al. US 20200114506 A1, herein referred to as Toshev.
Regarding claim 1,
Toshev discloses the following:
one or more computers (Paragraph 0025)
system may include one or more computers
obtaining a plurality of images of a scene in a real-world environment with which a robot will interact and, for each image, corresponding camera data comprising a viewpoint of a camera that captured the image (Paragraphs 0016-0017, 0020, 0034, 0065)
multiple images from different cameras may be utilized as ‘query’ images for a given scene
each of the query images may be obtained from vision components of a robot
training a scene synthesis machine learning model using the plurality of images and the corresponding camera data, wherein the scene synthesis machine learning model is configured to receive a scene input that comprises a camera viewpoint and to generate as output a synthetic image of the scene from the camera viewpoint (Fig. 2, Paragraph 0036)
a machine learning model may be trained on query images (input) to generate synthetic images (output) of a scene
these synthetic images have a given viewpoint based on associated vision components of the robot
these synthetic images may be of a simulated environment
generating, using at least synthetic images generated by the scene synthesis machine learning model, training data for training a policy neural network for use in controlling the robot in the real-world environment to perform one or more tasks, wherein the policy neural network is configured to receive a policy input comprising an observation characterizing a current state of the environment (Paragraphs 0037, 0041)
training data may be generated based on the synthetic images
these synthetic images (and associated viewpoints) are used for a policy neural network for controlling the robot
this policy neural network may be trained based on the synthetic image output from the synthetic model
these synthetic image inputs may be considered as observations of the environment
generate as output a policy output defining an action to be performed by the robot in response to the observation, wherein the observation comprises an image of the environment captured by a robot camera of the robot (Paragraph 0041)
the output of the policy network is a robot control action
this control action is in response to the image input into the policy network at a given time step
wherein generating the training data comprises: generating, from synthetic images generated by the scene synthesis machine learning model, observations of scenes in a simulation of the environment being interacted with by a model of the robot (Paragraphs 0018, 0036, 0041)
training data for the policy network may include synthetic images for a given timestamp
these synthetic images may be generated from a simulated environment (see above rationale)
this simulated environment is one in which the robot interacts
Regarding claim 2,
Toshev discloses all the limitations of claim 1. Toshev further discloses the following:
training the policy neural network on the training data (Paragraphs 0037, 0041)
the policy network may be trained based on input synthetic images (observations) which are generally considered as training data
Regarding claim 3,
Toshev discloses all the limitations of claim 2. Toshev further discloses the following:
after the training, controlling the robot in the real-world environment using the policy neural network (Paragraph 0041)
after training, the policy network outputs a robot action
Regarding claim 4,
Toshev discloses all the limitations of claim 1. Toshev further discloses the following:
obtaining a video of the scene in the real-world environment (Paragraphs 0036-0037)
videos of actual robotic actions may be utilized for training and input
selecting, as the plurality of images, a plurality of the video frames from the video (Paragraphs 0036-0037)
images may be used for generating synthetic images
these images may be from a given video (sequence of images)
Regarding claim 6,
Toshev discloses all the limitations of claim 1. Toshev further discloses the following:
controlling the model of the robot in the simulation of the environment using the policy neural network at each of a plurality of time steps, comprising, at each time step: obtaining, from a simulator, an input camera viewpoint based on a location of the robot camera at the time step within a state of the simulation of the real-world environment at the time step (Fig. 2, Paragraphs 0039-0041, 0043)
a robot may be controlled in a simulated environment to perform various actions through the policy network
the simulated environment may be generated through a simulator; the simulated environment may reflect the real environment
each action may be associated with a given point of view based on the camera used (see claim 1 rationale)
each action may be paired with a given observation and embedded into a given timestep
generating, using the scene synthesis model, a synthetic image of the scene from the input camera viewpoint (Paragraphs 0037, 0041)
synthetic images may be generated based on training data which can include a given camera and its point of view
generating an input image for the time step from at least the synthetic image of the scene (Fig. 2, Paragraphs 0039-0041, 0043)
synthetic images may be input into the policy network for each time step
these images may be of a synthetic scene
processing an observation comprising the input image using the policy neural network to generate a policy output (Fig. 2, Paragraphs 0039-0041, 0043)
the policy network may utilize the input images to generate a policy output, which can be a robotic action
selecting an action using the policy output (Fig. 2, Paragraphs 0039-0041, 0043)
an action may be selected from the policy network outputs
providing, to the simulator, the selected action for use in controlling the model of the robot to update the state of the simulation (Paragraphs 0051-0053)
simulated actions of a robot may be utilized for the simulator and its training
this training effectively updates the simulation as it is based on previous and learned results
generating a respective training example for each of the time steps that comprises the observation for the time step and the selected action for the time step (Paragraph 0054-0060)
episodes may be simulated, wherein each episode contains multiple time steps
each time step may be associated with a given scene and robotic action
each time step and associated data may be used for further training which means that the training data is effectively an example
Regarding claim 7,
Toshev discloses all the limitations of claim 6. Toshev further discloses the following:
obtaining, from the simulator, a respective rendering of one or more dynamic objects in the environment at the time step (Paragraphs 0054-0060)
the simulator may render multiple dynamic objects in a given scene for each time step
generating the input image for the time step by combining the synthetic image of the scene and the respective renderings (Paragraphs 0054-0060)
images of the various object in the environment, including the synthetic images, may be combined to create the simulated scene image as a whole
this simulated scene image may be used for input for training the recurrent network model
Regarding claim 8,
Toshev discloses all the limitations of claim 6. Toshev further discloses the following:
wherein the scene synthesis model is configured to receive camera viewpoints in a first reference frame and wherein the simulator operates in a world reference frame (Paragraphs 0041, 0063)
vision components of the robot may be used to generate fixed viewpoints
these viewpoints are based on the robot which includes a robot frame
wherein obtaining, from a simulator, an input camera viewpoint based on a location of the robot camera at the time step within the simulation of the real-world environment comprises: receiving, from the simulator, an initial camera viewpoint in the world reference frame (Paragraphs 0036, 0053-0055, 0063)
simulated viewpoints may be based on the simulated environment which can be considered a ‘world’ frame as it is different than that of the robot frame
generating the input camera viewpoint by mapping the initial camera viewpoint from the world reference frame to the first reference frame (Paragraphs 0036, 0041, 0053-0055, 0063)
simulated viewpoints may be utilized by the model to generate robotic actions
this is performed by utilizing the initial viewpoint of the robot camera to generate a simulated viewpoint in the simulated environment (world frame)
Regarding claim 9,
Toshev discloses all the limitations of claim 6. Toshev further discloses the following:
at each time step, receiving, from the simulator, a respective reward for each of the one or more tasks, wherein the training example includes the respective rewards (Paragraphs 0054-0060)
each time step may be associated with a given simulated robotic action
each simulated robotic action may be associated with a reward
each time step may be considered a training example (see claim 6 rationale)
Regarding claim 12,
Toshev discloses all the limitations of claim 1. Toshev further discloses the following:
wherein the observation further comprises data from a gyroscope of the robot, an accelerometer of the robot, or both (Paragraph 0101)
the robot may include accelerometers, gyroscopes, etc.
Regarding claim 16, a portion of the claim limitations are similar to those in claim 1 and are rejected using the same rationale as seen above in claim 1. Additionally, Toshev discloses one or more computers (Paragraph 0025; system may include one or more computers), and one or more storage devices (Paragraph 0025; system may include one or more readable storage mediums).
Regarding claim 17, the claim limitations are similar to those in claim 16 and are rejected using the same rationale as seen above in claim 16.
Regarding claim 18, the claim limitations are similar to those in claim 4 and are rejected using the same rationale as seen above in claim 4.
Regarding claim 20, the claim limitations are similar to those in claim 6 and are rejected using the same rationale as seen above in claim 6.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 5, 10, 14, and 19 are rejected under 35 U.S.C. 103 as being obvious over Toshev and in view of Sundaralingam et al., US 20240066710 A1, herein referred to as Sundaralingam (effective filing date of 08/29/2022 from provisional 63/373,846; examiner has checked and verified subject matter is supported).
Regarding claim 5, Toshev discloses all the limitations of claim 4. Toshev further discloses determining camera data for each of the plurality of images (Paragraphs 0006, 0034-0036, 0063; data concerning the vision components for capturing images may be determined; such data can include viewpoint, type of sensor, location, etc.), but fails to disclose determining camera data for each of the plurality of images using Structure-from-Motion (SfM). However, Sundaralingam, in an analogous field of endeavor, teaches determining camera data for each of the plurality of images using Structure-from-Motion (SfM) (Paragraph 0039; SfM may be used to determine relative camera poses). Therefore, from the teaching of Sundaralingam, it would have been obvious to one of ordinary skill in the art before the effective filing date to have modified, with a reasonable expectation for success, the robotic system of Toshev to include determining camera data for each of the plurality of images using Structure-from-Motion (SfM), as taught/suggested by Sundaralingam. The motivation to do so would be to utilize a well-known method to determine pose data of a camera. This can lead to more accurate robotic control and can increase the quality of the machine learning inputs and outputs.
Regarding claim 10, Toshev discloses all the limitations of claim 1. Toshev further discloses generating, using the trained scene synthesis model, synthetic data of a scene (Paragraphs 0036-0037, 0041; a machine learning model may output synthetic images of a scene, amongst other data), but fails to disclose generating, using the trained scene synthesis model, a mesh of the scene, and providing the mesh to the simulator for use in modeling collisions when updating the state of the simulation. However, Sundaralingam teaches generating, using the trained scene synthesis model, a mesh of the scene, and providing the mesh to the simulator for use in modeling collisions when updating the state of the simulation (Paragraphs 0042, 0046-0049; meshes may be generated for objects and the robot to determine potential collisions in the environment). Therefore, from the teaching of Sundaralingam, it would have been obvious to one of ordinary skill in the art before the effective filing date to have modified, with a reasonable expectation for success, the robotic system of Toshev to include generating, using the trained scene synthesis model, a mesh of the scene, and providing the mesh to the simulator for use in modeling collisions when updating the state of the simulation, as taught/suggested by Sundaralingam. The motivation to do so would be to generate accurate representations of each object in the scene to accurately determine if collisions are possible.
Regarding claim 14, Toshev discloses all the limitations of claim 1. Toshev further discloses a scene synthesis model (at least Paragraph 0036; a network may be trained to generate synthetic scenes/images), but fails to disclose wherein the scene synthesis model is a Neural Radiance Field (NeRF) model. However, Sundaralingam teaches wherein the scene synthesis model is a Neural Radiance Field (NeRF) model (at least Paragraph 0044; a neural radiance field (NeRF) may be utilized to output data related to an image). Therefore, from the teaching of Sundaralingam, it would have been obvious to one of ordinary skill in the art before the effective filing date to have modified, with a reasonable expectation for success, the robotic system of Toshev to include wherein the scene synthesis model is a Neural Radiance Field (NeRF) model, as taught/suggested by Sundaralingam. The motivation to do so would be to use a well-known method for generating a synthetic scene/image. This can allow for better representations of the environment for real and simulated robotic control.
Regarding claim 19, the claim limitations are similar to those in claim 5 and are rejected using the same rationale as seen above in claim 5.
Claim 13 is rejected under 35 U.S.C. 103 as being obvious over Toshev, and in view of Iqbal et al., US 20200061811 A1, herein referred to as Iqbal.
Regarding claim 13, Toshev discloses all the limitations of claim 2. Toshev further discloses wherein training the policy neural network comprises: training the policy neural network through reinforcement learning (Paragraph 0041; reinforcement learning can be used to train the policy network), but fails to disclose wherein training the policy neural network comprises: training the policy neural network through reinforcement learning with domain randomization. However, Iqbal, in an analogous field of endeavor, teaches wherein training the policy neural network comprises: training the policy neural network through reinforcement learning with domain randomization (Paragraph 0071; domain randomization may be used for a network). Therefore, from the teaching of Iqbal, it would have been obvious to one of ordinary skill in the art before the effective filing date to have modified, with a reasonable expectation for success, the robotic invention of Toshev to include wherein training the policy neural network comprises: training the policy neural network through reinforcement learning with domain randomization, as taught/suggested by Iqbal. The motivation to do so would be to increase the robustness of the policy network which can lead to better network outputs.
Allowable Subject Matter
Claims 11 and 15 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
The following is a statement of reasons for the indication of allowable subject matter:
Regarding claim 11, the examiner has performed a thorough search and has not found a piece of prior art, either alone or in combination with other prior art, that discloses, teaches, suggests, or renders obvious the claim limitations. The closest prior art combination, Toshev and Sundaralingam, teaches generating initial meshes in a first frame (see claim 10 rationale), but fails to disclose generating the mesh by mapping vertices in the initial mesh from the first reference frame to the world reference frame of the simulator, and providing the mesh to the simulator for use in modeling collisions when updating the state of the simulation. These features are novel in that they allow specific mapping of mesh vertices for use in the simulator. This can allow for the simulator to accurately represent objects for use in collision prediction. Examiner further notes that another piece of prior art, Gao et al. (‘K-VIL: Keypoints-based Visual Imitation Learning’), discloses meshes having similar correspondences being mapped to similar descriptors (Page 3). However, Gao only discloses a mapping of the meshes to a value (correspondence), and does not disclose a mapping of the mesh vertices to a global frame, and then providing that global mesh to update a state of a simulator for simulating collisions. Although Gao does disclose collision simulation between objects (as does Sundaralingam in claim 10), neither reference teaches a global mesh being generated based on mapped vertices, nor do they teach the global mesh being used to update a simulation state. As stated above, these features are novel in that they ultimately allow for better environment representation as well as better collision prediction.
Regarding claim 15, the examiner has performed a thorough search and has not found a piece of prior art, either alone or in combination with other prior art, that discloses, teaches, suggests, or renders obvious the claim limitations. The closest prior art combination, Toshev and Gao, teaches camera parameters that specify intrinsics of the camera that captures a plurality of images (Gao, Page 3), and generating observations by providing scene inputs (Toshev, see at least claim 1 rationale). However, neither Toshev nor Gao teach the camera that captured a plurality of images being different from the robot camera, wherein the scene input further comprises input camera parameters that specify intrinsics of an input camera that the synthetic image generated by the scene synthesis machine learning should match, and generating each of the observations by providing scene inputs that include input camera parameters that specify intrinsics of the robot camera instead of intrinsics of the camera that captured the plurality of images. These features are novel in that it allows the robotic system to utilize external cameras to generate the synthetic images/scene, ensuring that the camera intrinsics match between the camera and the inputs, and further utilizing intrinsics of the robot camera for generating the observations. This can ensure that the inputs to the model are correct and can lead to better outputs, specifically synthetic images/scenes. Examiner notes that Gao does teach camera intrinsic and extrinsic data being ‘mapped’ to an object, but does not teach the intrinsic data being on a different camera, nor that it matches with the synthetic scene/image after generation of the synthetic scene/image.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CHRISTOPHER ALLEN BUKSA whose telephone number is (571)272-5346. The examiner can normally be reached M-F 7:30 AM-4:30 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Thomas Worden can be reached at (571) 272-4876. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/CHRISTOPHER A BUKSA/Examiner, Art Unit 3658