Prosecution Insights
Last updated: October 01, 2026
Application No. 19/266,785

GENERATING GRASP POSES FOR CONTROLLING ROBOTS USING DIFFUSION MODELS

Non-Final OA §102§103§112
Filed
Jul 11, 2025
Priority
Oct 30, 2024 — provisional 63/713,898
Examiner
KHAYER, SOHANA T
Art Unit
3657
Tech Center
3600 — Transportation & Electronic Commerce
Assignee
NVIDIA Corporation
OA Round
1 (Non-Final)
82%
Grant Probability
Favorable
1-2
OA Rounds
1y 5m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 82% — above average
82%
Career Allowance Rate
263 granted / 321 resolved
+29.9% vs TC avg
Strong +19% interview lift
Without
With
+18.7%
Interview Lift
resolved cases with interview
Typical timeline
2y 8m
Avg Prosecution
31 currently pending
Career history
350
Total Applications
across all art units

Statute-Specific Performance

§101
4.2%
-35.8% vs TC avg
§103
50.4%
+10.4% vs TC avg
§102
12.5%
-27.5% vs TC avg
§112
27.6%
-12.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 321 resolved cases

Office Action

§102 §103 §112
DETAILED ACTION Remarks This non-final office action is in response to the application filled on 07/11/2025. Claims 1-20 are pending and examined below. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Acknowledgment is made of applicant’s claim for domestic benefit under 35 U.S.C. 119 (e). The provisional application No. 63/713,898, was filed on 10/30/2024. Information Disclosure Statement As of date of this action, IDS filled has been annotated and considered. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION. —The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim(s) 9, 10, 15 and 16 is/are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, or for pre-AIA the applicant regards as the invention. Regarding claim 9 (and similarly claim 15), which recites “a first label included in the one or more labels” is not clear. It is not clear whether label is referring first label/ number one/initial grasp or first layer of machine learning or something else. Dependent claim(s) 10 and 16 is/are also rejected because they do not resolve their parent deficiencies. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 1, 6, 7, 11, 13 and 20 is/are rejected under 35 U.S.C. 102(a)(2) as being anticipated by US 2025/0042024 (“Dijkman”). Regarding claim 1 (and similarly claim 11 and 20), Dijkman discloses a computer-implemented method for training a robot grasp diffusion model (see at least [0027], where “The affordance models 120 may be pre-trained on one or more large-scale datasets and/or the affordable actions (e.g., and associated action parameters, probabilities of success, etc.) may be learned online in an end-to-end interactive fashion (e.g., via reinforcement learning).”; see also [0026], where “the affordance models 120 include a set of convolutional neural networks (CNNs), such as a set of fully convolutional neural networks (FCNNs). However, in various aspects, any type of machine learning model capable of being trained to identify features in image or other sensor data may be implemented, such as, for example, a transformer neural network model, a recurrent neural network (RNN), an autoencoder, a diffusion model, etc.”; see also [0024]), the method comprising: performing, based on grasp data that includes one or more first robot grasp poses, one or more operations to train an untrained diffusion model to generate a trained diffusion model (see at least [0024], where “an affordable action generated by an affordance model 120 may include parameters indicating that a grasping action can be performed at (x, y, z) coordinates corresponding to the location of a handle of a mug, an orientation (e.g., of a robotic grasper) at which the handle is to be grasped by the device, and/or a force (e.g., applied by a robotic grasper to the handle) with which the handle is to be grasped by the device.”; orientation is interpreted as grasp pose. See also [0027], where diffusion model is trained using dataset and actions e.g., action parameters, probabilities of success. So, untrained diffusion model is trained using grasping action parameters. See also fig 5, block 520, where first set of actions is generated by first machine learning model. First machine learning model is interpreted as diffusion model.); generating, using the trained diffusion model, one or more second robot grasp poses (see at least [0091], where “a second set of affordable actions may be generated based on processing the second data via the first set of machine learning models.”); simulating the one or more second robot grasp poses to generate one or more labels indicating if the one or more second robot grasp poses are successful robot grasp poses (see at least [0058], where “the control system can score the possible configurations (e.g., each combination of a location and a set of action parameters) and select the highest-valued configuration (e.g., the location and set of action parameters having the highest score) to test.”; see also [0019], where “perform actions associated with a task with a higher success rate”; see also [0046]); and performing, based on the one or more second robot grasp poses and the one or more labels, one or more operations to train an untrained machine learning model to generate a trained machine learning model (see at least fig 5, block 530, where first selected action is generated by second machine learning model; second machine learning model is interpreted as machine learning model. see also [0027] and [0114-0115]), wherein the trained diffusion model and the trained machine learning model are used to process sensor data to generate a robot grasp plan for causing a robot to perform at least part of a task (see at least [0076], where “The data may include, for example, image data, sensor data”; see also fig 6). Regarding claim 6, Dijkman further discloses method wherein performing the one or more operations to train the untrained machine learning model is based on the one or more first robot grasp poses (see at least [0026] and fig 5). Regarding claim 7 (and similarly claim 13), Dijkman further discloses a method wherein generating the one or more second robot grasp poses comprises performing at least one of one or more first denoising steps using the trained diffusion model to generate a translation component of the one or more second robot grasp poses or one or more second denoising steps using the trained diffusion model to generate a translation component of the one or more second robot grasp poses (see at least [0091], where “a second set of affordable actions may be generated based on processing the second data via the first set of machine learning models”; second set of actions is interpreted as second robot grasp poses). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 2 is/are rejected under 35 U.S.C. 103 as being unpatentable over US 2025/0042024 (“Dijkman”), as applied to claim 1 above, and further in view of US 2025/0353169 (“Zhou”). Regarding claim 2, Dijkman does not disclose claim 2. However, Zhou discloses a method wherein the one or more first robot grasp poses include at least one of one or more grasp poses for an antipodal gripper or one or more grasp poses for a suction-based gripper (see at least [0039], where “robots may be equipped with an end effector 106 that takes the form of a claw with two opposing “fingers” or “digits.” Such a claw is one type of “gripper” known as an “impactive” gripper. Other types of grippers may include but are not limited to “ingressive” (e.g., physically penetrating an object using pins, needles, etc.), “astrictive” (e.g., using suction or vacuum to pick up an object), or “contigutive” (e.g., using surface tension, freezing or adhesive to pick up object).”; see also [0001]). Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to have modified Dijkman to incorporate the teachings of Zhou by including the above feature for providing slip -free grasping by including gripper type for pose determination. Claim(s) 3, 8, 9, 12, 14 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over US 2025/0042024 (“Dijkman”), as applied to claim 1 and 11 above, and further in view of US 2025/0249595 (“Li”). Regarding claim 3 (and similarly claim 12), Dijkman does not disclose claim 3. However, Li discloses a method wherein performing the one or more operations to train the untrained diffusion model to generate the trained diffusion model comprises: generating, based on object geometry data included in the grasp data and using an encoder, an object geometry embedding (see at least [0072], where “The machine may be of various sizes in order to accommodate the handling of different types or sizes of materials or objects. For example, the manipulator may be sized to handle large objects or small objects (e.g., the manipulator may have an open size of 5 mm to 70 mm and fingers having a width of 5 mm to 10 mm). Thus, the machine, being controlled by an ML model adapted for the parameters or dimensions of the machine, may enable a grasp fidelity, force, and/or friction during the performance of different tasks (e.g., manipulation of non-rigid materials).”); performing, based on the object geometry embedding and a third robot grasp pose, one or more forward diffusion steps using the untrained diffusion model to generate a predicted noise (see at least [0058], where “predicted noise”; noise is predicted that means forward diffusion steps is used. See also [0102], where “the depth data identifies the positions of the one or more robotic arms with respect to points on a three-dimensional representation of the object.”; three-dimensional object representation is interpreted as object geometry embedding), wherein the third robot grasp pose is generated by adding noise to a first robot grasp pose included in the one or more robot grasp poses (see at least [0058], where “The outputs 360 may include velocity fields when using flow matching models (e.g., also referred to as vector fields), raw/normalized action sequences when using autoregressive models, predicted noise when using diffusion models”); calculating, based on the predicted noise and the noise, a loss (see at least [0050], where “The generative model may compute internal parameters (e.g., weights), which map the inputs to the outputs while minimizing differences (e.g., loss) between an output generated by the model and a real data output (e.g., training the model).”); and updating, based on the loss, one or more parameters of the untrained diffusion model (see at least [0072], where “the machine control system 130 can train, create, generate, update, and/or modify the generative ML model to output control instructions for various types of machines, including machines that include or perform tasks using a manipulator at the end of a robotic arm.”; see also [0051]). Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to have modified Dijkman to incorporate the teachings of Li by including the above feature for providing more accurate grasp pose by considering object geometry and noise. Regarding claim 8 (and similarly claim 14), Li further discloses a method wherein the one or more labels include at least one of: a binary label indicating a success or a failure associated with a second robot grasp pose included in the one or more second robot grasp poses (see at least [0063], where “the feedback loop or module may operate based on a state of the environment, where the machine runs a policy after repetition of a successful task or unsuccessful task.”); or a continuous-valued score reflecting at least one of a grasp stability or one or more contact force margins associated with a second robot grasp pose included in the one or more second robot grasp poses. Regarding claim 9 (and similarly claim 15), as best understood in view of indefiniteness rejection explained above, Li further discloses a method wherein performing the one or more operations to train the untrained machine learning model to generate the trained machine learning model comprises: generating, based on object geometry data and using an encoder, an object geometry embedding (see citation on claim 3); generating, based on the object geometry embedding, a third robot grasp pose (see citation on claim 3); generating, based on the third robot gasp pose and using the untrained machine learning model, a predicted grasp pose score (see at least [0072], where successful task and unsuccessful task is interpreted grasp score); calculating, based on a first label included in the one or more labels and the predicted grasp pose score, a loss (see at least citation on claim 3); and updating, based on the loss, one or more parameters of the untrained machine learning model (see at citation on claim 3). Regarding claim 19, Li further discloses a system wherein the one or more labels include at least one of one or more positive robot grasp labels or one or more negative robot grasp labels (see at least [0063], where successful task is interpreted as positive robot grasp and unsuccessful task is interpreted as negative grasp). Claim(s) 4 is/are rejected under 35 U.S.C. 103 as being unpatentable over US 2025/0042024 (“Dijkman”), as applied to claim 1 above, and in view of US 2025/0249595 (“Li”), as applied to claim 3 above, and further in view of US 2025/0336043 (“Vasconcelos”). Regarding claim 4, Dijkman in view of Li does not disclose claim 4. However, Vasconcelos discloses a method wherein the loss comprises a denoising loss that measures an L2 norm of a difference between the predicted noise and the noise (see at least [0064], where “The denoising objective on which the system 100 trains the initial denoising neural network using can be any of a variety of appropriate objectives such as a mean squared error objective (to minimize the difference between predicted estimate of a noise component of the noisy initial image and the true noise component), or score matching objective (to estimate the score function (defined as the gradient of the log density) of the perturbed data distribution at different noise levels).”; see also [0124], where “represents the squared norm of the difference between the true noise ϵ and the target denoising neural network estimate of a noise component of the noisy target image”; see also [0166], [0170] and [0198]). Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to have modified Dijkman in view of Li to incorporate the teachings of Vasconcelos by including the above feature for eliminating large and glaring artifacts in a model. Claim(s) 5 is/are rejected under 35 U.S.C. 103 as being unpatentable over US 2025/0042024 (“Dijkman”), as applied to claim 1 above, and in view of US 2025/0249595 (“Li”), as applied to claim 3 above, and further in view of US 2019/0147234 (“Kicanaoglu”). Regarding claim 5, Dijkman in view of Li does not disclose claim 5. However, Vasconcelos discloses a method wherein calculating the loss comprises at least one of calculating a first loss for a rotation component of the first robot grasp pose or calculating a second loss for a translation component of the first robot grasp pose (see at least [0144], where “calculate a loss between the rotated second pose unit 1276 and the real forty-degree pose.”; see also [0145] and [0167]). Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to have modified Dijkman in view of Li to incorporate the teachings of Kicanaoglu by including the above feature for reducing errors during grasps by calculating loss between poses. Claim(s) 10 and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over US 2025/0042024 (“Dijkman”), as applied to claim 1 and 11 above, and in view of US 2025/0249595 (“Li”), as applied to claim 9 and 15 above, and further in view of US 2023/0368414 (“Afrooze”). Regarding claim 10 (and similarly claim 16), Dijkman in view of Li does not disclose claim 10. However, Afrooze discloses a method wherein the loss comprises a binary cross-entropy loss measuring a divergence between the predicted grasp score and the first label (see at least [0138], where “The neural network may, for example be trained by constructing labels for binary cross entropy loss where “true” labels are coordinate frames within the near threshold, “false” labels are coordinate frames beyond the far threshold, and coordinate frames in between the near and far thresholds are ignored.”; see also [0212]). Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to have modified Dijkman in view of Li to incorporate the teachings of Afrooze by including the above feature for increasing grasp success. Claim(s) 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over US 2025/0042024 (“Dijkman”), as applied to claim 11 above, and in view of US 2025/0249595 (“Li”), as applied to claim 15 above, and further in view of US 2024/0198530 (“Ugalde Diaz”). Regarding claim 17, Dijkman in view of Li does not disclose claim 17. However, Ugalde Diaz discloses a system wherein the loss comprises a first loss penalizing a confident incorrect grasp pose score generated by the untrained machine learning model more than a correct grasp pose score generated by the untrained machine learning model (see at least [0034], [0041] and [0051]). Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to have modified Dijkman in view of Li to incorporate the teachings of Ugalde Diaz by including the above feature for selecting accurate grasp poses by eliminating unsuccessful grasp poses. Claim(s) 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over US 2025/0042024 (“Dijkman”), as applied to claim 11 above, and in view of US 2025/0249595 (“Li”), as applied to claim 15 above, and further in view of US 2023/0017505 (“Menon”). Regarding claim 18, Dijkman in view of Li does not disclose claim 18. However, Menon discloses a system wherein calculating the loss comprises: calculating one or more first losses over one or more batches of grasp pose scores (see at least [0018], [0037] and [0052]. Dijkman discloses a system that determine grasp poses, see citation above. Menon discloses a system that generate scores for respective tasks and loss function between categories. So, it would be obvious to calculate losses utilizing Menon for generated grasp pose disclosed by Dijkman); and calculating, based on the one or more first losses, an average loss (see at least [0035], where “the loss function can be the sum or the average of the cross-entropy losses for the training examples in a batch sampled from the training data.”). Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to have modified Dijkman in view of Li to incorporate the teachings of Menon by including the above feature for applying the model for unseen environments. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to SOHANA TANJU KHAYER whose telephone number is (408)918-7597. The examiner can normally be reached Monday - Thursday, 7 am-5.30 pm, PT. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Abby Lin can be reached at 5712703976. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SOHANA TANJU KHAYER/ Primary Examiner, Art Unit 3657
Read full office action

Prosecution Timeline

Jul 11, 2025
Application Filed
Jul 23, 2026
Non-Final Rejection mailed — §102, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12746077
SYSTEM WITH REMOVABLE HUBS FOR MANUAL AND ROBOTIC PROCEDURE
2y 9m to grant Granted Sep 29, 2026
Patent 12748430
AUTONOMOUS ROBOT SYSTEM, AND METHOD FOR CONTROLLING AUTONOMOUS ROBOT
2y 9m to grant Granted Sep 29, 2026
Patent 12722832
METHODS AND SYSTEMS FOR USE IN PROCESSING SEEDS
3y 1m to grant Granted Sep 01, 2026
Patent 12725112
INTEGRATED ROOFING ACCESSORIES FOR UNMANNED VEHICLE NAVIGATION AND METHODS AND SYSTEMS INCLUDING THE SAME
1y 9m to grant Granted Sep 01, 2026
Patent 12715133
Setting Device, Setting Method, And Non-Transitory Computer-Readable Storage Medium Storing Setting Program
2y 5m to grant Granted Aug 25, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
82%
Grant Probability
99%
With Interview (+18.7%)
2y 8m (~1y 5m remaining)
Median Time to Grant
Low
PTA Risk
Based on 321 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month