DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Joint Inventors
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on July 24th, 2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-4, 7-10, 11-14 and 17-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Analysis of the claims in view of MPEP § 2106.04 are provided below.
Regarding Claim 1 : A method comprising:
receiving training data comprising a plurality of RGB-D images of an object at a plurality of time steps, and a plurality of robot actions associated with the object at the plurality of time steps; and
optimizing, using the training data, a dynamics function to predict a future state of the object based on a current state of the object and a robot action,
wherein a state of the object is estimated as a plurality of particles comprising 3D Gaussians using particle filtering.
Step 1: Statutory Category – Yes
The claim recites a process, which falls within one of the four statutory categories. MPEP § 2106.03.
Step 2A Prong One Evaluation: Judicial Exception – Yes
The Office submits that the foregoing underlined limitation(s) constitute judicial exceptions in terms of “mental processes” because under broadest reasonable interpretation, the claim covers performance using mental processes.
The claim recites limitations of receiving and optimizing data comprising images of an object and associated actions with the respective image, with an estimation of an object state. These limitations as drafted, are simple processes that under their broadest reasonable interpretation, covers performance of the limitations in the mind without further reciting significant structure or application to classify as an inventive concept. Nothing in the claim precludes the elements of the limitations from being performed in the mind. For example, a person could receive images of an object and have a preconceived mental process regarding what actions can and should be associated with the object by simply observing and analyzing said images. A person can also reasonably optimize a function in their mind for predicting future states of the object based on a combination of its current state and any one of the well-known and understood potential robot actions. Furthermore, a person can reasonably base the state of the object on an estimation of its state using particle filtering, which represented in the mind can simply be excluding the particles in the received image that do not directly represent the target object. Thus, this step recites a mental process.
Each of the limitations identified above falls within at least one of collecting, observing, and evaluating information, all of which fall within the bucket of “Mental Processes” of abstract ideas.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong Two Evaluation: Practical Application – No
Claim 1 is evaluated whether as a whole it integrates the recited judicial exception into a practical application. As noted in MPEP § 2106.04, it must be determined whether any additional elements in the claim beyond the abstract idea integrate the exception into a practical application in a manner that imposes a meaningful limit on the judicial exception. The courts have indicated that additional elements merely using a computer or generic component to implement an abstract idea, adding insignificant extra-solution activity, or generally linking use of a judicial exception to a particular technological environment or field of use do not integrate a judicial exception into a “practical application”.
In the present case, the additional limitations beyond the above-noted abstract ideas are as follows (where the bolded portions are the “additional limitations” while the underlined portions continue to represent the “abstract idea”).
The additional elements recited in Claim 1 do not integrate the judicial exception(s) into a practical application, and are either (a) merely using generic components to implement an abstract idea or (b) lacking any supportive structure to implement an abstract idea. The involvement of these elements is nothing more than generally linking the abstract idea into a particular field of use. The claim is directed to an abstract idea.
Step 2B Evaluation: Inventive Concept – No
Claim 1 is evaluated as to whether the claims as a whole amount to significantly more than the recited exception (i.e., whether any additional element, or combination of additional elements, adds an inventive concept to the claim).
As discussed with respect to Step 2A Prong Two, the additional elements are merely generally linking the abstract idea into a particular field of use. The same analysis applies here in 2B.
For these reasons, there is no inventive concept in the claim, and thus it is ineligible.
Regarding Claims 2-4 and 7-10, these claims generally only further limit the abstract idea by introducing additional steps that can be performed mentally. While some of these claims include additional elements not previously discussed, they are discussed in a substantially similar manner as the additional elements in Claim 1, such as optimizing different layers of some mental function, predicting or updating future object states, and receiving secondary data for further estimation and prediction. Therefore, similar analysis can be used that would arrive at the same conclusion that these claims are still ineligible.
Regarding Claims 11-14 and 17-20, other than falling under different statutory categories, all claimed limitations are substantially similar to that of Claims 1-4 and 7-10, and therefore, similar analysis can apply and these claims are also ineligible.
Examiner notes that physical operation or execution of the claimed plurality of robot actions to physically affect the robot is suggested as part of intended use language in Claim 1. In the event the claim is amended to positively recite the robot’s physical operation or execution of claimed robot actions as implemented using the information received and processed, it will likely overcome the 101 rejection(s) noted above.
Claim Objections
Claims 1, 3-4, 8, 10, 13-14 and 18-19 are objected to because of the following informalities:
Claim 1 Line 7: “3D” should be revised to “three-dimensional (3D)”.
Claim 3 Line 1: “the robot actions” should be revised to “the plurality of robot actions” to avoid lacking antecedent basis.
Claim 4 Line 1: “the particles” should be revised to “the plurality of particles” to avoid lacking antecedent basis.
Claim 8 Line 2: “the object at first time step” should be revised to “the object at a first time step”.
Claim 10 Line 4: “the Gaussians” should be revised to “the 3D Gaussians”.
Claim 13 Line 1: “the robot actions” should be revised to “the plurality of robot actions” to avoid lacking antecedent basis.
Claim 14 Line 1: “the particles” should be revised to “the plurality of particles” to avoid lacking antecedent basis.
Claim 18 Line 3: “the object at first time step” should be revised to “the object at a first time step”.
Claim 19 Line 5: “the Gaussians” should be revised to “the 3D Gaussians”.
Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-2, 6-7, 11-12, 16-17 and 20 (along with claims 3-5, 8-10, 13-15 and 18-19 due to dependency) are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
The terms “optimizing” in claims 1-2 and 6-7 and “optimize” in claims 11-12, 16-17 and 20, are relative terms which render the claims indefinite. The terms “optimizing” and “optimize” are not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1, 3-9, 11, 13-18 and 20 are rejected under 35 U.S.C. 103 as being obvious over Matsumara et al. (US Patent Pub. No. 2024/0061431 A1), herein “Matsumara”, published February 22nd, 2024, in view of Zhang et al. (“AdaptiGraph: Material-Adaptive Graph-Based Neural Dynamics for Robotic Manipulation”), herein “Zhang”, published July 10th, 2024.
Regarding Claims 1, 11 and 20, Matsumara discloses a method, computing device comprising one or more processors, and non-transitory computer readable storage medium storing a program that when executed by a processor, causes the processor to:
receiving training data comprising a plurality of RGB-D images of an object at a plurality of time steps, and a plurality of robot actions associated with the object at the plurality of time steps (See 0040, “[…] observation device 113 is an image pickup device that acquires an image and video data, a device that observes a surrounding environment […]” See also 0049, “[…] a storage for storing input and output information serving as learning data and a function for performing learning […] learning of the model is performed by using a method generally applied to the learning of the model with the history of the observation and the action as input and output […]” See also 0052, “[…] acquisition information acquired by a depth camera is a two-dimensional image […]” See also 0059-0060, “[…] pieces of information corresponding to three times or more may be used […] position of the object is computed at a plurality of times, the other detection/tracking unit 200 has an object detection and tracking function.”); and
optimizing, using the training data, a dynamics function to predict a future state of the object based on a current state of the object and a robot action (See 0010, “[…] predicts future states of the plurality of detected objects from the first state information by using a first model for predicting a future state of an object […]” See also 0049, “[…] where the model is a neural network, learning can be performed by using an error back propagation method or the like […] input and output information serving as learning data and a function for performing learning.” See also 0054, “[…] prediction model 500 receives, as input, self-state information 501 indicating the current state of the robot, action candidate information 502 indicating an action assumed by the robot […]”).
But does not explicitly disclose wherein a state of the object is estimated as a plurality of particles comprising 3D Gaussians using particle filtering.
Zhang, in a similar field of endeavor, teaches a state of the object is estimated as a plurality of particles comprising 3D Gaussians using particle filtering (See Fig. 4 shown below and Section 1, “[…] employed Graph Neural Networks (GNN) to model environments as 3D particles […] encoding the material type and physical property variables into particles in the graph, the model learns material-specific dynamic functions that predict different physical behaviors for objects […]” Examiner notes Fig. 4 shows the process for rendering images of varying objects, representing their physical parameters using dynamic prediction, and modeling the resulting scene from the image sample, thus being the same as 3D Gaussians).
PNG
media_image1.png
516
704
media_image1.png
Greyscale
In view of Zhang’s teachings, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to include, with the process for learning from observations and action histories of a robot as disclosed by Matsumara, a particle dynamics model conditioned on physical parameters of an object, with a reasonable expectation of success, since both references are directed towards predicting future object and environmental states from observes states and robot actions; and AdaptiGraph provides particle-based dynamics implementation using the same type of prediction models in the robot system of Matsumara.
Regarding Claims 3 and 13, Matsumara further discloses the method of claim 1 and computing device of claim 11, wherein the robot actions are represented by a second set of Gaussians (See 0049 and 0054 as referenced above. See also 0076, “[…] the action determination unit 240 computes entropy (average information amount) of the other-prediction information 802 as the other uncertainty […]” Examiner notes the action determination unit clearly calculates a probability density based on the relationship of the prediction information and parameters of uncertainty).
Regarding Claims 4 and 14, Matsumara further discloses the method of claim 1 and computing device of claim 11, wherein opacities of the particles comprise importance weights describing contributions of each Gaussian (See 0078, “[…] represent the weights of the other-self-prediction information Bi and the other uncertainty Ci for each other, and are given by other weight information 281. The weight of the other may be determined in advance in accordance with the type of the other, and the detected other may be labeled with reference to at least one of the type and the feature of the object of the classification definition information 280 to determine the value of the weight of each of the others.”).
Regarding Claims 5 and 15, Matsumara further discloses the method of claim 1 and computing device of claim 11, wherein the dynamics function comprises a neural network comprising:
an object encoder to encode the RGB-D images of the object (See 0103, “[…] synthesizing images from different viewpoints from a plurality of acquired images […] a technique of outputting an observation image from another viewpoint in a manner that a plurality of images, viewpoint information of the images, and another viewpoint information from which an image is to be estimated are input by using a neural network learned by images of a plurality of viewpoints and viewpoint information of the images […]”);
an action encoder to encode the robot actions (See 0068, “[…] to predict the future state by using the prediction model 500, it is necessary to input the action candidate information 502 of the own person or robot.”);
a grid interaction network to determine a grid solution based on the grid features (See 0054, “[…] the neural network can be configured by, for example, a graph neural network or the like capable of handling graph representation in which a plurality of objects are set as vertices (nodes) and a relationship between the objects is set as an edge.”); and
an object decoder to generate output dynamics based on the updated particle features (See 0049 and 0054 as referenced above. See also 0054, “[…] prediction model 500 can be configured not by a model that predicts one prediction value but by a Bayesian model that predicts as a probability distribution […] where a future position is predicted as a probability distribution, prediction information is output as a probability that a target robot or person exists at each surrounding position in the future.”).
But does not explicitly disclose a particle-to-grid module to convert particle features to grid features; and
a grid-to-particle module to convert the grid solution to updated particle features.
Zhang, in a similar field of endeavor, teaches a particle-to-grid module to convert particle features to grid features (See Section 1, “[…] provide such graph-based models to adapt to objects and tasks involving diverse materials and varying physical properties, such as manipulating ropes with different stiffness and granular media with different granularity […] encode this variation using […] the physical property variable, and integrate the variable into a Graph-Based Neural Dynamics (GBND) framework […] By encoding the material type and physical property variables into particles in the graph, the model learns material-specific dynamic functions that predict different physical behaviors for objects […]”); and
a grid-to-particle module to convert the grid solution to updated particle features (See Section 3 Reference B, “[…] for any global 3D translation added to the particle locations, the predictions should also be translated identically […] by passing the position difference of receiver and sender particles to the edge encoder […] perform backpropagation through time to optimize model parameters.” See also Section 3 Reference C, “[…] an inverse optimization pipeline […] the robot updates its estimate of the object by minimizing the dynamics prediction error from previous interactions.”).
In view of Zhang’s teachings, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to include, with the process for learning from observations and action histories of a robot as disclosed by Matsumara, a particle dynamics model included in the neural network, with a reasonable expectation of success, since both references are directed towards predicting future object and environmental states from observes states and robot actions; and AdaptiGraph provides particle-based dynamics implementation using the same type of prediction models in the robot system of Matsumara.
Regarding Claims 6 and 16, Matsumara further discloses the method of claim 5 and computing device of claim 15, wherein optimizing the dynamics function comprises learning parameters associated with the object encoder, the action encoder, the grid interaction network, and the object decoder (See 0049 as referenced above. See also 0047, “[…] the model information 282 includes configuration information of a neural network, learned parameter information, and the like.” See also 0055, “The prediction model 500 performs learning by using the action of the robot and the history information of the observation of the states of the robot and the others at that time […] the model may be updated by performing learning online at an appropriate time from the observation information obtained […]”).
Regarding Claims 7 and 17, Matsumara further discloses the method of claim 1 and computing device of claim 11, further comprising optimizing the dynamics function to predict the future state of the object by:
predicting the future state of the object at a next time step (See 0059-0060 as referenced above. See also 0010, “[…] predicts future states of the plurality of detected objects from the first state information by using a first model for predicting a future state of an object […]”); and
updating the future state of the object at the next time step based on a likelihood function (See 0055 and 0076 as referenced above. See also 0076, “[…] other uncertainty 903 represents the uncertainty of the prediction of the others […]” Examiner notes an uncertainty factor in determining prediction information is a likelihood function).
Regarding Claims 8 and 18, Matsumara further discloses the method of claim 1 and computing device of claim 11, further comprising:
receiving a second plurality of RGB-D images of the object at first time step (See 0040, 0049, 0052 and 0059-0060 as referenced above. Examiner notes a plurality of images are received at a plurality of time steps, thus being a mere duplication of the essential working parts or steps and choosing an order of the already available information is a simply, conventional design choice);
receiving a second robot action associated with the object at the first time step (See 0040, 0049, 0052 and 0059-0060 as referenced above. Examiner notes a plurality of images are received at a plurality of time steps, thus being a mere duplication of the essential working parts or steps);
estimating a state of the object at the first time step based on the second plurality of RGB- D images as a plurality of 3D Gaussians using particle filtering (See 0010, 0049 and 0054 as referenced above); and
predicting a second state of the object at a second time step based on the state of the object at the first time step, the second robot action, and the dynamics function (See 0010, 0049 and 0054 as referenced above).
Regarding Claim 9, Matsumara further discloses the method of claim 8, further comprising adjusting weights of the 3D Gaussians based on a resampling function (See 0078 as referenced above and, “[…] where a person is prioritized over a robot, the weight information may be determined to increase the weight of the person depending on whether the type of the other person is a person or a robot.” See also 0088-0090, “[…] regarding the weight value, the weight value may be determined in accordance with the correspondence between the determined rank and the weight by ranking a value obtained by normalizing the prediction error […] other-weight computation unit 1102 may determine the weight value in consideration of both the prediction error and the position information of the other obtained from the other detection/tracking unit […]”).
Claims 2 and 12 are rejected under 35 U.S.C. 103 as being obvious over Matsumara et al. (US Patent Pub. No. 2024/0061431 A1) in view of Zhang et al. (“AdaptiGraph: Material-Adaptive Graph-Based Neural Dynamics for Robotic Manipulation”) as applied to claims 1 and 11 above, and further in view of Florin et al. (US Patent Pub. No. 2007/0098221 A1), herein “Florin”.
Regarding Claims 2 and 12, Matsumara in view of Zhang does not explicitly teach the method of claim 1 and computing device of claim 11, further comprising optimizing parameters of the dynamics function against a rendering loss and a physical constraint loss.
Florin, in a similar field of endeavor, teaches optimizing parameters of the dynamics function against a rendering loss and a physical constraint loss (See 0032, “[…] parameters of such deformation […] recovered through […] additional regularization constraints […]” See also 0037, “[…] unknown parameters of such over-constrained system can be determined through a robust least square minimization […]” See also 0040, “[…] consider the Euclidean distance between the prediction and the actual observations in the image coordinate system to be the error metric of the auto-regressive model […]” Examiner notes the image observation error metric provides the same parameter as a rendering loss).
In view of Florin’s teachings, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to include, with the process for learning from observations and action histories of a robot and optimizing the learned dynamics model as disclosed by Matsumara in view of Zhang, the optimization to include observation-based errors and physical constraints, with a reasonable expectation of success, since minimizing predicted and observed errors improves reliability in overall dynamics prediction as taught by Florin. Furthermore, the system taught already optimizes its learned dynamics model, and incorporating physical constraint penalties would further constrain the model to generate physically permissible predictions.
Claims 10 and 19 are rejected under 35 U.S.C. 103 as being obvious over Matsumara et al. (US Patent Pub. No. 2024/0061431 A1) in view of Zhang et al. (“AdaptiGraph: Material-Adaptive Graph-Based Neural Dynamics for Robotic Manipulation”) as applied to claims 9 and 18 above, and further in view of Panagiotis et al. (“Reducing the Memory Footprint of 3D Gaussian Splatting”), herein “Panagiotis”, published June 24th, 2024.
Regarding Claims 10 and 19, Matsumara in view of Zhang does not explicitly teach the method of claim 9 and computing device of claim 18, wherein the resampling function performs the steps of:
merging one or more of the 3D Gaussians having an opacity below a first predetermined threshold; and
splitting one or more of the Gaussians having a ratio of a maximum to minimum eigenvalue greater than a second predetermined threshold.
Panagiotis, in a similar field of endeavor, teaches the resampling function performs the steps of:
merging one or more of the 3D Gaussians having an opacity below a first predetermined threshold (See Section 4.1, “During optimization, the original 3DGS approach regularly culls primitives that fall below a specified opacity threshold, as they contribute little to the final image […] estimate this spatial redundancy and combine this information with low-opacity filtering […] make this decision for each Gaussian primitive g in space […] a 3D region around g of that extent should be occupied by a low number of primitives; densely-packed clusters of Gaussians in that region […]” See also Section 6.2, “[…] full method that combines Opacity (1st row) and Redundancy (2nd row) […]”); and
splitting one or more of the Gaussians having a ratio of a maximum to minimum eigenvalue greater than a second predetermined threshold (See Section 4.1 as referenced above and, “[…] corresponding to the Gaussian’s scale/rotation […]” See also Section 2, “[…] cull redundant Gaussians […]” Examiner notes the scale of the Gaussian is its eigenvalue).
In view of Panagiotis’ teachings, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to include, with the process for learning from observations and action histories of a robot and optimizing the learned dynamics model as disclosed by Matsumara in view of Zhang, the merging and splitting using opacity-based Gaussian selection, with a reasonable expectation of success, since it is recognized that Gaussian opacity provides a measure of a primitive’s significance, and conventional particle filtering requires preferential retention and removal of particles according to their weights. Using the known opacity measure as the particle weight would provide a predictable mechanism for controlling the Gaussian particles.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Bryant Tang whose telephone number is (571)270-0145. The examiner can normally be reached M-F 8-5 CST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Thomas Worden can be reached at (571)272-4876. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/BRYANT TANG/Examiner, Art Unit 3658
/THOMAS E WORDEN/Supervisory Patent Examiner, Art Unit 3658