DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
Receipt is acknowledged of claim amendments with associated arguments/remarks, received June 29, 2026. Claims 1-20 are pending in which claims 1, 4, 6-11, 14-15, 18-20 were amended.
Response to Arguments
Applicant’s arguments, see pg , filed June 29, with respect to the rejections of claim 1-13, 15-20 under 35 U has been fully considered and is persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration of the claim amendments that changed the scope and interpretation of at least the independent claim limitations, a new grounds of rejection for the independent amended claims is made under Yoshitake et al (TransPoser: Transformer as an Optimizer for Joint Object Shape and Pose Estimation) in view of Hampali et al (US 2022/0301304).
All arguments were addressed.
Information Disclosure Statement
The information disclosure statements (IDS) submitted on 04/01/2026 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statements are considered by examiner.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-3, 7-13, 16, 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Yoshitake et al (TransPoser: Transformer as an Optimizer for Joint Object Shape and Pose Estimation, cited in Non-Final 03/27/2026) in view of Hampali et al (US 2022/0301304).
Regarding Claim 1, Yoshitake et al teach a computer-implemented (DeepDDF and TransPoser executed on a GeForce RTX 2080 Ti GPU; 6.2 Real-world Data ¶ 3) method for determining object poses (use of Deep Directional Distance Function (DeepDDF) neural network and TransPoser transformer for a joint object pose and shape reconstruction, executed on a computer; Fig 1-3 and 3. Joint Shape and Pose Estimation-5. TransPoser), the method comprising:
receiving a first image of an object (a first RGB-D observed depth image is captured and input to the DeepDDF; Fig 1, 2 and 3. Joint Shape and Pose Estimation ¶ 3, 4. Deep DDF ¶ 2);
sampling an initial pose of the object (the Deep DDF will use 3D directional sampling rays for a given viewpoint and viewpoint; Fig 1, 2 and 4. Deep DDF ¶ 4); and
performing one or more operations to update the initial pose (additional operations are performed (decoding latent code into 3D voxel array, feature vector encoding, using TransPoser for additional shape and pose updates; Fig 1-3 and 3. Joint Shape and Pose Estimation ¶ 5-6, 4. Deep DDR ¶ 5, 5. TransPoser ¶ 1) to determine
pose of an object based at least on a cropping of the first image conditioned on the initial pose (from the observed image of the object, a predicted output depth image is generated by the DeepDDF; Fig 1, 2 and 3. Joint Shape and Pose Estimation ¶ 3, 4. Deep DDF ¶ 4),
a first rendered image of the object in the initial pose (from the observed input image and the predicted output depth image an image is rendered that represents the difference; Fig 1, 2 and 3. Joint Shape and Pose Estimation ¶ 3, 4. Deep DDF ¶ 4), and
one or more transformer encoders (the TransPoser contains multiple transformer encoders and is used to optimize the object pose; Fig 1, 3 and 3. Joint Shape and Pose Estimation ¶ 5, 5. TransPoser ¶ 1).
Yoshitake et al does not teach the pose of an object based at least on a cropping of the first image conditioned on the initial pose.
Hampali et al is analogous art pertinent to the technological problem addressed in the current application and teaches pose of an object based at least on a cropping of the first image conditioned on the initial pose (hand pose is estimated using a joint vector representation based on cropped image patches representing the object (hand) used for pose evaluation (using a transformer architecture machine learning system 257; ¶ [0096]-[0098]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the current application to combine the teachings of Yoshitake et al with Hampali et al including pose of an object based at least on a cropping of the first image conditioned on the initial pose. By using cropped image regions, such as a patch focused on the region of interest for pose analysis, the only region tested is focused on, thereby improving computational speed and reducing data analysis while improving the analysis accuracy by the machine learning model, as recognized by Hampali et al (¶ [0100]).
Regarding Claim 2, Yoshitake et al in view of Hampali et al teach the computer-implemented method of claim 1 (as described above), further comprising:
sampling one or more additional initial poses of the object (Yoshitake et al, additional images of the object are taken (see Fig 1a new image 2, new image 3) from the given pose of the object from different perspectives; Fig 1, 2 are input to the DeepDDF; Fig 1, 2 and 3. Joint Shape and Pose Estimation ¶ 3, 4. Deep DDF ¶ 2);
performing one or more operations to update the one or more additional initial poses to generate one or more additional poses of the object (Yoshitake et al, additional operations are performed (decoding latent code into 3D voxel array, feature vector encoding, using TransPoser for additional shape and pose updates of the new images from the additional perspectives; Fig 1-3 and 3. Joint Shape and Pose Estimation ¶ 5-6, 4. Deep DDR ¶ 5, 5. TransPoser ¶ 1); and
selecting a second pose of the object from the first pose (interpreted as a second pose in addition to the first pose (see k=1, k=2 of Yoshitake Fig 1a) and the one or more additional poses (Yoshitake et al, the multiple observation of the multiple Kt views of the estimate xt are computed with the self-attention in the encoder transforms for each view to determine the additional views; Fig 1, 2 and 3. Joint Shape and Pose Estimation ¶ 3-5, 4. Deep DDF ¶ 4, 5. TransPoser ¶ 6).
Regarding Claim 3, Yoshitake et al in view of Hampali et al teach the computer-implemented method of claim 2 (as described above), wherein selecting the second pose comprises:
generating a ranking (described by applicant as a vector of floating point numbers to estimate poses; specification ¶ [0070]) for each of the first pose and the one or more additional poses using a hierarchy of self-attention layers (Yoshitake et al, a feature vector encoding the 3D shape corresponding to z is stored from the DeepDDF (4. Deep DDF ¶ 5) and the Transformer Encoder uses the vector (denoted x in the TransPoser) to estimate and refine the current estimate xt at step t where the encoder computes self-attention between tokens of multiple views for the current estimate; Fig 3, 4 and 5. TransPoser ¶ 1-2, 6); and
selecting one of the first pose or the one or more additional poses that is associated with a highest ranking as the second pose (Yoshitake et al, an optimization is performed with weighting between the views of the object to consolidate information across the views with the self-attention; Fig 3, 4 and 5. TransPoser ¶ 6-7).
Regarding Claim 7, Yoshitake et al in view of Hampali et al teach the computer-implemented method of claim 1 (as described above), further comprising:
performing one or more operations to generate a model of the object (Yoshitake et al, additional operations are performed (decoding latent code into 3D voxel array, feature vector encoding, using TransPoser for additional shape and pose updates to output a category-level object shape and pose estimation; Fig 1-3 and 3. Joint Shape and Pose Estimation ¶ 1-6, 4. Deep DDR ¶ 5, 5. TransPoser ¶ 1); and
rendering the first rendered image based at least on the model of the object (Yoshitake et al, the output from the DeepDDF combined with TransPoser renders a category-level object shape representation and pose estimation; Fig 1-3 and 3. Joint Shape and Pose Estimation ¶ 1-6, 5. TransPoser ¶ 1).
Regarding Claim 8, Yoshitake et al in view of Hampali et al teach the computer-implemented method of claim 1 (as described above), wherein the first rendered image is rendered based on at least a computer-aided design (CAD) model of the object (Yoshitake et al, DeepDDF and TransPoser are trained using CAD image models and used as the ground truth data 6.1 DeepDDF ¶ 1, 6.2 TransPoser ¶ 3).
Regarding Claim 9, Yoshitake et al in view of Hampali et al teach the computer-implemented method of claim 1 (as described above), further comprising:
receiving a second image of the object, wherein the first image and the second image are consecutive frames of a video (Yoshitake et al, the consecutive frames (Camera Ray Sampler that samples volume for a given viewpoint and viewing direction Fig 2, 4. Deep DDF ¶ 1) of a given pose may be from RGB-D videos of real-world indoor scenes; 6.2 TransPoser Real World Data ¶ 1); and
updating the first pose to determine a second pose of the object within the second image based at least on the second image (Yoshitake et al, the camera ray sampler data is used to determine the multi-view 3D representation based on the 2D image-space data based on the sampled feature vectors Gz; Fig 1, 2 and 4. Deep DDF ¶ 4-5), a second rendered image of the object in the first pose (Yoshitake et al, new additional images are rendered from the additional perspectives; Fig 1-3 and 3. Joint Shape and Pose Estimation ¶ 5-6, 4. Deep DDR ¶ 5, 5. TransPoser ¶ 1), and the one or more transformer encoders (Yoshitake et al, the TransPoser contains multiple transformer encoders and is used to optimize the object pose from the given image of the object pose perspective; Fig 1, 3 and 3. Joint Shape and Pose Estimation ¶ 5, 5. TransPoser ¶ 1).
Regarding Claim 10, Yoshitake et al in view of Hampali et al teach the computer-implemented method of claim 1 (as described above), further comprising at least one of rendering an image, controlling a vehicle, or controlling a robot based at least on the first pose of the object (Yoshitake et al, the use of DeepDDF and TransPoser is used for joint object shape and pose estimation with applications for robotics, VR/AR (image rendering) and autonomous driving; Fig 1 and 1. Introduction ¶ 6).
Regarding Claim 11, Yoshitake et al teach one or more non-transitory computer-readable media storing program instructions (computer program code for use of Deep Directional Distance Function (DeepDDF) neural network and TransPoser transformer for a joint object pose and shape reconstruction, understood as stored on memory and executed with processor of GeForce RTX 2080 Ti GPU (see NPL for “geforce rtx 2080 ti”); Fig 1-3 and 3. Joint Shape and Pose Estimation-5. TransPoser) that, when executed by at least one processor (DeepDDF and TransPoser executed on a GeForce RTX 2080 Ti GPU; 6.2 Real-world Data ¶ 3), cause the at least one processor to perform the steps of: identical steps as claim 1 (as described above).
Regarding Claim 12, Yoshitake et al in view of Hampali et al teach the one or more non-transitory computer-readable media of claim 11 (as described above), with further steps identical to claim 2 (as described above).
Regarding Claim 13, Yoshitake et al in view of Hampali et al teach the one or more non-transitory computer-readable media of claim 12 (as described above), with further steps identical to claim 3 (as described above).
Regarding Claim 16, Yoshitake et al in view of Hampali et al teach the one or more non-transitory computer-readable media of claim 11 (as described above), wherein the one or more transformer encoders includes a first transformer encoder used to generate a position update to the initial pose and a second transformer encoder used to generate a rotation update to the initial pose (Yoshitake et al, the TransPoser uses multi-view sequential observations with the first transformer et1 analyzing the first positional view and pose update and the second transformer et2 analyzing the second positional view, rotated from the first view, k=1, k=2 Fig 1) and pose update; Fig 1, 3 and 5. TransPoser ¶ 2-6).
Regarding Claim 18, Yoshitake et al teach the one or more non-transitory computer-readable media of claim 11 (as described above), with further steps identical to claim 7 (as described above).
Regarding Claim 19, Yoshitake et al teach the one or more non-transitory computer-readable media of claim 11 (as described above), with further steps identical to claim 9 (as described above).
Regarding Claim 20, Yoshitake et al teach a system, comprising: one or more memories storing instructions (DeepDDF and TransPoser executed on a GeForce RTX 2080 Ti GPU with includes a GDDR6 memory (see NPL for “geforce rtx 2080 ti”); 6.2 Real-world Data ¶ 3); and one or more processors that are coupled to the one or more memories (DeepDDF and TransPoser executed on a GeForce RTX 2080 Ti GPU; 6.2 Real-world Data ¶ 3) and, when executing the instructions (computer program code for use of Deep Directional Distance Function (DeepDDF) neural network and TransPoser transformer for a joint object pose and shape reconstruction, executed on a computer; Fig 1-3 and 3. Joint Shape and Pose Estimation-5. TransPoser), are configured to: perform the steps identical to the steps of claim 1 (as described above).
Claims 4, 6, 15, 17 are rejected under 35 U.S.C. 103 as being unpatentable over Yoshitake et al (TransPoser: Transformer as an Optimizer for Joint Object Shape and Pose Estimation, cited in Non-Final 03/27/2026) in view of Hampali et al (US 2022/0301304) and Nguyen et al (PIZZA: A Powerful Image-only Zero-Shot Zero-CAD Approach to 6 DoF Tracking).
Regarding Claim 4, Yoshitake et al in view of Hampali et al teach the computer-implemented method of claim 3 (as described above), including cropping of the first image conditioned on the initial pose (described aby Hampali et al above).
Yoshitake et al in view of Hampali et al does not teach wherein generating the ranking comprises processing, via the hierarchy of self-attention layers, one or more second images that are cropped from the first image based at least on the first pose and the one or more additional poses, and one or more third images of the object that are rendered based on at least the first pose and the one or more additional poses.
Nguyen et al is analogous art pertinent to the technological problem addressed in the current application and teaches generating the ranking comprises processing, via the hierarchy of self-attention layers , one or more second images that are cropped from the first image based at least on the first pose and the one or more additional poses (objects are identified in the image with bounding boxes and cropped as second images from the first image and are based on different poses (Icrop1, Icrop2, Icrop3), are embedded and input to the multi-head self-attention transformer as an object query; Fig 5 and 3.4 Proposed Architecture), and one or more third images of the object that are rendered based on at least the first pose and the one or more additional poses (the interframe correlation is determined in the self-attention transformer encoder T as an object query to predict relative pose based on rotation and translation estimation and a sequence of 6D poses are predicted by chaining relative poses predicted from consecutive frames; Fig 5 and 3.4 Proposed Architecture ¶ 1).
It would have been obvious to one of ordinary skill in the art to combine the teachings of Yoshitake et al in view of Hampali et al with Nguyen et al including generating the ranking comprises processing, via the hierarchy of self-attention layers, one or more second images that are cropped from the first image based on the first pose and the one or more additional poses, and one or more third images of the object that are rendered based on the first pose and the one or more additional poses. By cropping the object and then using a transformer encoder, pose estimation is performed using consecutive frames which may then allow for recursive analysis and reduce frame drifting tracking of an object over iterations, thereby improving the object pose estimation and tracking, as recognized by Nguyen et al (1. Introduction ¶ 4-6).
Regarding Claim 6, Yoshitake et al in view of Hampali et al teach the computer-implemented method of claim 3 (as described above).
Yoshitake et al in view of Hampali et al does not teach wherein performing the one or more operations to update the initial pose comprises: rendering the object based at least on the initial pose to generate the first rendered image; cropping the first image based at least on the initial pose to generate a cropped image; and processing the first rendered image and the cropped image using a trained machine learning model that comprises the one or more transformer encoders.
Nguyen is analogous art pertinent to the technological problem addressed in the current application and teaches performing the one or more operations to update the initial pose comprises:
rendering the object based at least on the initial pose to generate the first rendered image (objects are identified in the image with the object detector backbone to generate the first pose in the first frame; Fig 1-3 and 3.1 Unseen Object Tracking and Segmentation ¶ 1-4);
cropping the first image based at least on the initial pose to generate a cropped image (the objects are identified in the image with bounding boxes and cropped to generate a cropped image of the object; Fig 3 and 3.1 Unseen Object Tracking and Segmentation ¶ 1-4); and
processing the first rendered image and the cropped image using a trained machine learning model that comprises the one or more transformer encoders (the cropped object with the backgrounding masking is estimated to determine object translation with a transformer encoder; Fig 5 and 3.4 Proposed Architecture).
It would have been obvious to one of ordinary skill in the art to combine the teachings of Yoshitake et al in view of Hampali et al with Nguyen et al including wherein performing the one or more operations to update the initial pose comprises: rendering the object based at least on the initial pose to generate the first rendered image; cropping the first image based at leaston the initial pose to generate a cropped image; and processing the first rendered image and the cropped image using a trained machine learning model that comprises the one or more transformer encoders. By cropping the object and then using a transformer encoder, pose estimation is performed using consecutive frames which may then allow for recursive analysis and reduce frame drifting tracking of an object over iterations, thereby improving the object pose estimation and tracking, as recognized by Nguyen et al (1. Introduction ¶ 4-6).
Regarding Claim 15, Yoshitake et al in view of Hampali et al teach the one or more non-transitory computer-readable media of claim 11 (as described above), with further steps identical to claim 6 (as described above).
Regarding Claim 17, Yoshitake et al in view of Hampali et al and Nguyen et al teach the one or more non-transitory computer-readable media of claim 15 (as described above), wherein the instructions, when executed by the at least one processor, further cause the at least one processor to perform the step of performing one or more operations to train a machine learning model that includes the one or more transformer encoders (Yoshitake et al, the TransPoser includes training the model with the multiple transformer encoders based on a number of iterations to optimize the shape and pose; Fig 3-5 and 6.2 TransPoser Ablation Studies ¶ 1-3).
Claims 5 is rejected under 35 U.S.C. 103 as being unpatentable over Yoshitake et al (TransPoser: Transformer as an Optimizer for Joint Object Shape and Pose Estimation) in view of Hampali et al (US 2022/0301304) and Watson et al (US 2024/0244322, cited in Non-Final 03/27/2026).
Regarding Claim 5, Yoshitake et al in view of Hampali et al teach the computer-implemented method of claim 2 (as described above), wherein the initial pose and the one or more additional initial poses include one or more rotations that are sampled on an icosphere (Yoshitake et al, sampling is performed with viewpoints on a fixed-radius sphere, referred to as a canonical sphere; Fig 1 and 4. Deep DDF ¶ 7).
Yoshitake et al in view of Hampali et al does not explicitly teach sampling is on an icosphere.
Watson is analogous art pertinent to the technological problem addressed in the current application and teaches sampling is performed on an icosphere (projections are generated on the object with modeling tools including icospheres; Fig 3, 4 and ¶ [0046]).
It would have been obvious to one of ordinary skill in the art to substitute the canonical sphere teachings of Yoshitake et al in view of Hampali et al with the icosphere teaching of Watson. Properties of icospheres provide an efficient means of querying an image and reconstructing the object image to maintain spatial alignment with maintaining different resolutions, thereby improving and optimizing computations, while accounting for camera rotation and optical flow estimation during motion, as recognized by Watson et al ¶ [0046]-[0048]).
Allowable Subject Matter
Claim 14 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Regarding Claim 14, the prior art was not identified to teach, suggest or provide motivations to combine the following limitations with the limitations in which the claim rely in a non-obvious manner:
The one or more non-transitory computer-readable media of claim 13, wherein the instructions, when executed by the at least one processor, further cause the at least one processor to perform the step of performing one or more operations to train a machine learning model that comprises the hierarchy of self-attention layers based on a pose-conditioned triplet loss and a plurality of pairs of pose samples, each pair of pose samples in the plurality of pairs of pose samples including a positive pose sample from a viewpoint that is less than a threshold from a corresponding ground truth pose.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Wen et al (US 2024/0169563, application 18/509627, cited in Non-Final 03/27/2026) from the same applicant and co-inventors teach a method and system of 3D reconstruction and tracking of objected used for mixed reality, which is distinct from the current application focused on determining object poses and using an encoder to render an image of the object based on the pose.
Tremblay et al (US 2024/0123620, application 18/219,031, cited in Non-Final 03/27/2026) teach a system and method for grasp pose prediction, with claims focused on using a first neural network and a second neural network to generate grasp proposal code using image data and neural radiance fields analysis for determining the grasping poses for objects.
Jantos et al (PoET: Pose Estimation Transformer for Single-View, Multi-Object 6D Pose Estimation, cited in Non-Final 03/27/2026) teach a system and method where objects are identified in image data with bounding boxes (effectively cropped as second images from the first image) and positional encoding are determined, which is then passed to the multi-head attention transformer as an object query, with a 6D pose estimation performed.
Zach (US 2016/0275686, cited in Non-Final 03/27/2026) teach a method and system for object pose recognition, including capturing a plurality of images of the object pose based on a given perspective and ranking the images to select an image from the candidate location but does not teach the pose ranking is performed with a self-attention transformer encoder.
Description of a Nvidia “geforce rtx 2080 ti” that discloses computer hardware components, including memory and processor, cited in Non-Final 03/27/2026.
Applicants’ amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KATHLEEN M BROUGHTON whose telephone number is (571)270-7380. The examiner can normally be reached Monday-Friday 8:00-5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, John Villecco can be reached at (571) 272-7319. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/KATHLEEN M BROUGHTON/Primary Examiner, Art Unit 2661