Prosecution Insights
Last updated: October 01, 2026
Application No. 18/844,310

Multi-Stage Object Pose Estimation

Final Rejection §103§112
Filed
Sep 05, 2024
Priority
Mar 11, 2022 — EU 22161679.0 +1 more
Examiner
BEZUAYEHU, SOLOMON G
Art Unit
Tech Center
Assignee
Siemens Aktiengesellschaft
OA Round
2 (Final)
76%
Grant Probability
Favorable
3-4
OA Rounds
1y 2m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 76% — above average
76%
Career Allowance Rate
480 granted / 634 resolved
+15.7% vs TC avg
Strong +30% interview lift
Without
With
+29.9%
Interview Lift
resolved cases with interview
Typical timeline
3y 2m
Avg Prosecution
37 currently pending
Career history
667
Total Applications
across all art units

Statute-Specific Performance

§101
17.2%
-22.8% vs TC avg
§103
52.7%
+12.7% vs TC avg
§102
12.7%
-27.3% vs TC avg
§112
10.0%
-30.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 634 resolved cases

Office Action

§103 §112
DETAILED ACTION Response to Arguments Applicant's arguments filed with respect to claims 1-6, and 8-12 have been fully considered but are moot in view of the new ground(s) of rejection. The rejections are necessitated due to claim amendments. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 8 and 9 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. In this case, they depend on cancelled claim 7. For examination purposes, the examiner assumed that the claims depend on claim 1. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1 and 2 are rejected under 35 U.S.C. 103 as being unpatentable over Wiedemann et al. (Pub. No. 2009/0096790 hereinafter “Wied”) in view of Venkataraman et al. (Pub. No. US 2022/0414928, hereinafter “Ven”) further in view of Yang et al. (Pub. No. US 2019/0197196). Regarding claim 1, Wied teaches method for estimating a multi-dimensional pose (3d poses) of an object based on an image of the object [Para. 12 “This invention provides a system and method for recognizing a 3D object in a single camera image and for determining the 3D pose of the object with respect to the camera coordinate system.”], the method comprising: in a preparational stage, providing the image depicting the object and a plurality of templates (2d projections) TEMPL(i) including generating the templates TEMPL(i) from a 3D model of the object in a rendering procedure (rendering), wherein different templates TEMPL(i), TEMPL(j) with i≠j of the plurality are generated by rendering from different known virtual viewpoints vVIEW(i), VVIEW (j) on the model [Para. 9 “View-based recognition techniques are based on the comparison of the 2D search image with 2D projections of the object seen from various viewpoints”; Para. 11 “In another form of the view-based recognition, the 2D projections are created by rendering a 3D model of the 3D object from different viewpoints”. Regarding claim limitation “i≠j”, Wied addresses it by generating multiple rendered views/projections of the same 3D object from different viewpoints (para. 11), and different views of the object are generated …… by placing virtual cameras around the 3D object (para. 64); because the viewpoints are plural and different, at least two distinct templates exist. Therefore, selecting two different templates in that plurality corresponds to different indices (i≠j)]; matching templates (2d projections) wherein at least one template TEMPL(J) from the plurality of templates TEMPL(i) is identified which matches best (most similar) with the image [Para. 9 “View-based recognition techniques are based on the comparison of the 2D search image with 2D projections of the object seen from various viewpoints. The desired 3D pose of the object is the pose that was used to create the 2D projection that is the most similar to the 2D search image”; determining correspondence by comparing a representation of the identified template TEMPL(J) with a representation of the image to determine 2D-3D-correspondences between pixels in the image and 3D model of the object [Para. 4 “If the 3D coordinates of the features are known, the 3D pose of the object can be computed directly from a sufficiently large set (e.g., four points) of those 2D-3D correspondences.”; para. 91 “The visible projected model edges are sampled to discrete points using a suitable sampling distance, e.g., 1 pixel.”; “Only the correspondences with an angle difference below a threshold are accepted as valid correspondences”; and Para. “the hidden-line computation requires a significant amount of computation time, which in some cases is too slow for real-time computations, especially when using a complex 3D model that consist of many edges”]; and estimating a multi-dimensional pose (3d pose) based on the 2D-3D-correspondences [Para. 4 “If the 3D coordinates of the features are known, the 3D pose of the object can be computed directly from a sufficiently large set (e.g., four points) of those 2D-3D correspondences”]. However, Wied doesn’t explicitly teach determining correspondence by computing 2D-2D correspondences between representation of the identified template TEMPL(J) and the representation of the image by a trained network. Ven teaches determining correspondence (dense correspondences between the rendered image 731 and the observed image 732) by computing 2D-2D correspondences between representation of the identified template TEMPL(J) (rendered image 731 of the object in the scene on the initial estimated pose) and the representation of the image (observed image 732) by a trained network (retraining of existing neural network backbones) [Para. 159, 160, and 162]. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Wied’s local edge-correspondence stage by using Ven’s retained neural network backbone to receive the rendered template image and observed image and compute dense 2D-2D correspondences. This modification improves Wiedemann by increasing correspondence density and learned tolerance to image variation, thereby supplying more reliable matches for pose refinement. Wied in view of Ven doesn’t explicitly teach processing the 2D-2D correspondences using the known virtual viewpoint vVIEW(J) to provide 2D-3D-correspondences between pixels in the image belonging to the object and voxels of the 3D model of the object. Yang teaches processing (inversely convert) the 2D-2D correspondences (correspondences between 2D locations of edges of the target object and 2D locations of 2D model points of the target object) using the known virtual viewpoint vVIEW(J) (virtual viewpoint of the present embodiment; on the basis of the view) to provide 2D-3D-correspondences (obtain the 3D model point P_i) between pixels in the image belonging to the object (2D locations of edges of the target object) and voxels of the 3D model of the object (2D model points P_i) [Para. 102, 144 and 167]. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Wied’s matched view correspondence refinement stage, modified by Ven, by using Yang’s location correspondence determination to inversely convert the matched 2D model points, on the basis of the associated view, into 3D model points corresponding to the image points. This modification improves Wied by furnishing explicitly image to model 2D-3D correspondences for direct geometric pose estimation, thereby making the refinement input more geometrically constrained. Regarding claim 2, Wied in view of Declerck teaches all claim limitation as stated above. Furthermore, Declerck teaches wherein object detection includes generating a segmentation mask from the image identifying those pixels of the image which belong to the object [Para. 60, fig. 1, 3 and related description]. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Wied’s pose-refinement (correspondences-determination) step that samples projected model edges at 1 pixel and accepts valid correspondences by incorporating Declerck’s teaching of determining a voxel corresponding to a given pixel and voxels for volumetric 3d model representations, thereby meeting the claimed voxels based correspondence requirement while retaining Wied’s pose estimation based on the resulting correspondences. Claims 3-5, and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Wiedemann et al. (Pub. No. 2009/0096790 hereinafter “Wied”) in view of Venkataraman et al. (Pub. No. US 2022/0414928, hereinafter “Ven”) further in view of Yang et al. (Pub. No. US 2019/0197196) further in view of Degol et al. (Pub. No. US 2022/0284233). Regarding claim 3, Wied in view of Ven further in view of Yang doesn’t explicitly teach the claim limitations. However, Degol teaches wherein generation of the segmentation mask includes performing a semantic segmentation by dense matching of features of the image to an object descriptor tensor ok representing the plurality of templates TEMPL(i) [Para. 18, 19, 23, fig. 2 and related description]. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Wied view of Ven further in view of Yang’s template comparison step to use Degol’s dense feature map matching via a 4D correlation map between image features, thereby enabling more robust pixel level segmentation support for the template pipeline, thereby improving object pixel identification prior to pose estimation. Regarding claim 4, Wied in view of Ven further in view of Yang doesn’t explicitly teach the claim limitations. However, Degol teaches computing an object descriptor tensor ok=F FE k (MOD) for the model of the object utilizing a feature extractor F FE k from the templates TEMPL(i), wherein the object descriptor tensor ok represents all templates TEMPL(i), computing feature maps fk = FFE(IMA) utilizing a feature extractor FFE for the image; and computing the binary segmentation mask based on a correlation tensor ck which results from a comparison of image features expressed by the feature maps fk of the image IMA and the object descriptor tensor ok [Para. 23, fig. 1, 3 and related description]. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Wied view of Ven further in view of Yang’s template matching workflow by incorporating Degol’s feature extractor generated feature maps and correlation map computation to compare image features against template derived feature representations for dense matching. This modification improves robustness of downstream correspondence/pose computation. Regarding claim 5, Wied in view of Ven further in view of Yang doesn’t explicitly teach the claim limitations. However, Degol teaches calculating per-pixel correlations between the feature maps fk of the image IMA and features from the object descriptor tensor ok, wherein each pixel in the feature maps fk of the image IMA is matched to the object descriptor ok, which results in the correlation tensor ck, wherein a particular correlation tensor value for a particular pixel (h, w) of the image IMA is defined as ch,w,x,y,z k = corr (fh,w k, ox,y,z k), wherein corr represents a correlation function, preferably according to a Pearson correlation [Para. 25, fig. 2 and related description]. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Wied view of Ven further in view of Yang’s template comparison step to use Degol’s dense feature map matching via a 4D correlation map between image features, thereby enabling more robust pixel level segmentation support for the template pipeline. Regarding claim 11, Wied in view of Ven further in view of Yang doesn’t explicitly teach the claim limitations. However, Degol teaches computing for each template a template feature map at least for a template foreground section of the respective template which contains the pixels belonging to the model; computing a feature map at least for an image foreground section of the image which contains the pixels belonging to the object; and computing for each template a similarity for the respective template feature map and the image feature map; wherein the template for which the highest similarity is determined is chosen to be the identified template [Para. 20-24]. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Wied view of Ven further in view of Yang’s template comparison step to use Degol’s dense feature map matching via a 4D correlation map between image features, thereby enabling more robust pixel level segmentation support for the template pipeline. Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Wiedemann et al. (Pub. No. 2009/0096790 hereinafter “Wied”) in view of Venkataraman et al. (Pub. No. US 2022/0414928, hereinafter “Ven”) further in view of Yang et al. (Pub. No. US 2019/0197196) further in view of Xu (Pub. No. US 2023/0252667). Regarding claim 6, Wied in view of Ven further in view of Yang doesn’t explicitly teach the claim limitations. However, Xu teaches applying a PnP+RANSAC procedure to estimate the multi-dimensional pose from the 2D-3D-correspondences 2D3D [Para. 16, fig. 3-5, and related description]. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Wied view of Ven further in view of Yang’s pose computation from 2D-3D correspondences by incorporating Xu’s teaching of running a PnP algorithm in a RANSAC loop on the correspondence produced in Wied’s pipeline. This medication improves Wied by reducing sensitivity to outlier correspondences, thereby improving pose accuracy. Claims 8-10 are rejected under 35 U.S.C. 103 as being unpatentable over Wiedemann et al. (Pub. No. 2009/0096790 hereinafter “Wied”) in view of Venkataraman et al. (Pub. No. US 2022/0414928, hereinafter “Ven”) further in view of Yang et al. (Pub. No. US 2019/0197196) and further in view Degol et al. (Pub. No. US 2022/0284233). Regarding claim 8, Wied in view of Ven further in view of doesn’t explicitly teach the claim limitations. However, Degol teaches correlating the representation of the image and the representation of the identified template [Para. 23, fig. 2, 4 and related description]; and computing the 2D-2D-correspondences are computed based on the correlation result [Para. 23, fig. 2, 4 and related description]. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Wied view of Ven further in view of Yang further in view of Xu’s template comparison step to use Degol’s dense feature map matching via a 4D correlation map between image features, thereby enabling more robust pixel level segmentation support for the template pipeline. Regarding claim 9, Wied in view of Ven further in view of Yang doesn’t explicitly teach the claim limitations. However, Degol teaches the representation of the image comprises an image feature map of at least a section of the image which includes pixels belonging to the object [para. 20, 23, fig. 1, 2 and related description]; and the representation of the identified template is a template feature map of the identified template [para. 23, 25, fig. 1, 2 and related description]. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Wied view of Ven further in view of Yang’s template comparison step to use Degol’s dense feature map matching via a 4D correlation map between image features, thereby enabling more robust pixel level segmentation support for the template pipeline. Regarding claim 10, Wied in view of Ven further in view of Yang doesn’t explicitly teach the claim limitations. However, Degol teaches computing a 2D-2D-correlation tensor ck by matching each pixel of one of the feature maps with all pixels of the respective other feature map [Para. 25]; and processing the 2D-2D-correlation tensor ck to determine the 2D-2D-correspondences [Para. 23]. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Wied view of Ven further in view of Yang’s template comparison step to use Degol’s dense feature map matching via a 4D correlation map between image features, thereby enabling more robust pixel level segmentation support for the template pipeline. Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Wiedemann et al. (Pub. No. 2009/0096790 hereinafter “Wied”) in view of Venkataraman et al. (Pub. No. US 2022/0414928, hereinafter “Ven”) further in view of Yang et al. (Pub. No. US 2019/0197196) further in view of Degol et al. (Pub. No. US 2022/0284233) and further in view of Sedky et al. (Pub. No. US 2012/0008858). Regarding claim 12, Wied in view of Ven further in view of Yang further in view of Degol doesn’t explicitly teach the claim limitations. However, Sedky teaches cropping the image foreground section from the image by utilizing the segmentation mask generated [Para. 7]. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Wied view of Ven further in view of Yang further in view of Degol’s search image processing by incorporating Sedky’s teaching of cropping out the foreground object using the segmentation output to form an image foreground section prior to matching/pose estimation. This modification improves Wied in view of Declerck by reducing background clutter and computation on irrelevant pixels, thereby improving efficiently. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SOLOMON G BEZUAYEHU whose telephone number is (571)270-7452. The examiner can normally be reached on Monday-Friday 10 AM-8 PM. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Oneal Mistry can be reached on 313-446-4912. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 888-786-0101 (IN USA OR CANADA) or 571-272-4000. /SOLOMON G BEZUAYEHU/ Primary Examiner, Art Unit 2666
Read full office action

Prosecution Timeline

Sep 05, 2024
Application Filed
May 19, 2026
Non-Final Rejection mailed — §103, §112
Aug 05, 2026
Response Filed
Sep 25, 2026
Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749214
THREE-DIMENSIONAL POSE ESTIMATION USING TWO-DIMENSIONAL IMAGES
3y 2m to grant Granted Sep 29, 2026
Patent 12731374
SAFE CONTROL/MONITORING OF A COMPUTER-CONTROLLED SYSTEM
3y 2m to grant Granted Sep 08, 2026
Patent 12725271
IMAGE QUALITY ENHANCING
4y 2m to grant Granted Sep 01, 2026
Patent 12721537
SYSTEM AND METHOD FOR MEASURING BLOOD FLOW VELOCITY ON A MICROFLUIDIC CHIP
3y 3m to grant Granted Sep 01, 2026
Patent 12718592
INFORMATION COLLECTION SYSTEM, SERVER, AND INFORMATION COLLECTION METHOD
3y 3m to grant Granted Aug 25, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
76%
Grant Probability
99%
With Interview (+29.9%)
3y 2m (~1y 2m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 634 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month