Prosecution Insights
Last updated: October 01, 2026
Application No. 18/931,491

DEPTH ESTIMATION USING ODOMETRY AND HAND TRACKING

Non-Final OA §103
Filed
Oct 30, 2024
Priority
Sep 16, 2024 — GR 20240100633
Examiner
YANG, JIANXUN
Art Unit
2662
Tech Center
2600 — Communications
Assignee
Snap Inc.
OA Round
1 (Non-Final)
74%
Grant Probability
Favorable
1-2
OA Rounds
8m
Est. Remaining
93%
With Interview

Examiner Intelligence

Grants 74% — above average
74%
Career Allowance Rate
491 granted / 663 resolved
+12.1% vs TC avg
Strong +19% interview lift
Without
With
+19.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
46 currently pending
Career history
700
Total Applications
across all art units

Statute-Specific Performance

§101
4.6%
-35.4% vs TC avg
§103
66.2%
+26.2% vs TC avg
§102
5.9%
-34.1% vs TC avg
§112
17.4%
-22.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 663 resolved cases

Office Action

§103
DETAILED ACTION The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims 1-20 are pending. Claim Rejections - 35 USC § 103 The following is a quotation of pre-AIA 35 U.S.C. 103(a) which forms the basis for all obviousness rejections set forth in this Office action: (a) A patent may not be obtained though the invention is not identically disclosed or described as set forth in section 102 of this title, if the differences between the subject matter sought to be patented and the prior art are such that the subject matter as a whole would have been obvious at the time the invention was made to a person having ordinary skill in the art to which said subject matter pertains. Patentability shall not be negatived by the manner in which the invention was made. Claim(s) 1-4, 8-14 and 18-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Tran et al (US20210065391A1) in view of Iqbal et al (US20210117661A1). Regarding claims 1, 19 and 20, Tran teaches a system comprising: at least one processor; and at least one memory component storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising: accessing a two-dimensional (2D) camera image captured by a camera on an augmented reality (AR) head-mounted device; (Tran, "capturing a sequence of RGB images from an unlabeled monocular video stream obtained by a monocular camera", [0004]; "augmented reality applications 915", [0085]; accessing 2D camera images (RGB images) from a monocular camera for augmented reality applications, which implicitly encompasses standard AR hardware such as head-mounted devices) generating a first set of tracked three-dimensional (3D) points using an odometry system on the AR head-mounted device on the 2D camera image; (Tran, "employ an SLAM system, e.g., the RGB-D version of ORB-SLAM, to process the pseudo RGB-D data, yielding camera poses as well as 3D map points", [0044]; "sparse 3D feature point estimates from geometric SLAM", [0026]; using a SLAM (odometry) system on the camera images to generate tracked 3D map points/feature estimates) generating a second set of tracked 3D points based on one or more images captured by the camera; Tran does not expressly disclose but Iqbal teaches: (Iqbal, "Estimating a 3D pose of an object from an image captured by a monocular camera.", [0022]; "3D coordinates of object keypoints are estimated relative to the camera position.", [0024]; Tran teaches tracking general 3D map points (see above, [0044, 0026]), while Iqbal teaches using a neural network to estimate a second set of tracked 3D points (3D keypoints) for specific objects like a hand from the camera images) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to incorporate the teachings of Iqbal into the system or method of Tran in order to track both static scene structures (via SLAM) and dynamic objects (via neural networks) to enable rich AR interactions. The combination of Tran and Iqbal also teaches other enhanced capabilities. The combination of Tran and Iqbal further teaches: creating a sparse depth image by projecting the first and second set of tracked 3D points onto the 2D camera image; and (Tran, "reliable sparse depth estimation", [0041]; Iqbal, "the relationship between the 3D location Pk and corresponding 2D projection pk can be written as follows under a perspective projection", [0030]; Tran teaches sparse depth estimation from SLAM, and Iqbal teaches projecting 3D keypoints onto a 2D image projection. Combining the projected 2D coordinates of both the SLAM tracked points (Tran) and the dynamic object keypoints (Iqbal) onto the 2D camera image would create a combined sparse depth image representing all tracked elements) generating a metric depth estimation by inputting the 2D camera image and the sparse depth image into a first machine learning model. (Tran, "use depth maps from the CNN-based depth network and run pRGBD-SLAM and the exemplary embodiments inject the outputs of pRGBD-SLAM... to fine-tune the depth network parameters to improve the depth prediction.", [0027]; inputting the RGB images and the sparse depth data (SLAM outputs) into a CNN-based depth network (first machine learning model) to refine and generate improved depth predictions) Regarding claim 2, the combination of Tran and Iqbal teaches its/their respective base claim(s). The combination further teaches the system of claim 1, wherein the camera includes a monocular camera, wherein the 2D camera image includes intensity information. (Tran, "unlabeled monocular video stream obtained by a monocular camera", [0004]; "minimize the pixel intensity discrepancies", [0036]; capturing image sequences from a monocular camera and utilizing pixel intensity information for depth estimation) Regarding claim 3, the combination of Tran and Iqbal teaches its/their respective base claim(s). The combination further teaches the system of claim 1, wherein the 2D camera image includes a color image of a current view of a user of the AR head-mounted device. (Tran, "sequence of RGB images", [0004]; capturing RGB images, which are color images of the camera's current view) Regarding claim 4, the combination of Tran and Iqbal teaches its/their respective base claim(s). The combination further teaches the system of claim 1, wherein generating the first set of tracked 3D points using the odometry system includes tracking spatial movement of the 3D coordinates as a user of the AR head-mounted device moves. (Tran, "camera motion that induces multiple-view geometric constraints ... estimate the ego-motion of the agent", [0019]; the SLAM/odometry system tracks the spatial ego-motion of the agent (user) to simultaneously recover 3D scene structure and movement) Regarding claim 8, the combination of Tran and Iqbal teaches its/their respective base claim(s). The combination further teaches the system of claim 1, wherein generating the second set of tracked 3D points is by inputting the one or more images captured by the camera into a second machine learning model, the second machine learning model is trained for near field 3D point detection. (Tran, "generates reliable depth estimates for nearby points", [0041]; utilizing the machine learning depth network specifically because it generates reliable depth estimates for nearby (near field) points) Regarding claim 9, the combination of Tran and Iqbal teaches its/their respective base claim(s). The combination further teaches the system of claim 1, wherein generating the second set of tracked 3D points is by inputting the one or more images captured by the camera into a second machine learning model, the second machine learning model is trained for detecting 3D points for objects in motion, wherein the odometry system is optimized for static objects. (Tran, "persistent 3D points that are visible across many frames", [0041]; Iqbal, "a 3D pose of an object, such as a hand or body (human, animal, robot, etc.) ", [0003]; Tran teaches SLAM optimized for persistent (static) 3D points. Iqbal teaches training the neural network model to detect 3D points of dynamic, moving objects like humans or animals. Incorporating Iqbal's dynamic object tracking model with Tran's static SLAM would be obvious for a robust AR system) Regarding claim 10, the combination of Tran and Iqbal teaches its/their respective base claim(s). The combination further teaches the system of claim 1, wherein generating the second set of tracked 3D points is by inputting the one or more images captured by the camera into a second machine learning model, the second machine learning model is trained to detect one or more hands of a user of the AR head-mounted device. (Iqbal, "A deep neural network-based system is described for estimating a 3D pose of an object from an image captured by a monocular camera... the object may be any object represented by a structural skeletal model, including a human hand", [0022]; "A hand pose is represented by a set of points in 3D space, called keypoints ... A neural network architecture learns to generate a depth value for each keypoint in the captured image", [0005]; using a deep neural network (machine learning model) that takes an input image captured by a camera and is trained to detect and estimate the 3D pose/keypoints (second set of tracked 3D points) of a user's hand) Regarding claim 11, the combination of Tran and Iqbal teaches its/their respective base claim(s). The combination further teaches the system of claim 10, wherein generating the second set of tracked 3D points is by inputting the one or more images captured by the camera into a second machine learning model, the second machine learning model outputs the second set of tracked 3D points that include at least joint positions of a detected hand of the user. (Iqbal, "fixed set of points in 3D space, usually joints, called landmarks or keypoints.", [0003]; the tracked 3D points for the hand model correspond directly to the joint positions of the detected hand) Regarding claim 12, the combination of Tran and Iqbal teaches its/their respective base claim(s). The combination further teaches the system of claim 1, wherein the one or more images includes the 2D camera image. (Tran, "depth maps 107 are then fed together with the RGB images 103", [0031]; feed the same initial 2D RGB camera image into the machine learning module along with the depth data) Regarding claim 13, the combination of Tran and Iqbal teaches its/their respective base claim(s). The combination further teaches the system of claim 1, wherein the one or more images are of a different resolution than the 2D camera image. (Tran, "multiscale strategy", [0063]; employing a multiscale strategy, which involves processing images at different resolutions to improve depth and feature estimation) Regarding claim 14, the combination of Tran and Iqbal teaches its/their respective base claim(s). The combination further teaches the system of claim 1, wherein the one or more images are of a different field of view than the 2D camera image. (Tran, "horizontal focal length and b is the virtual stereo baseline.", [0045]; transforms the image data into a virtual stereo setting with a virtual baseline, effectively utilizing a different field of view for depth processing) Regarding claim 18, the combination of Tran and Iqbal teaches its/their respective base claim(s). The combination further teaches the system of claim 1, wherein the operations further comprise applying a global correction factor to the metric depth estimation by determining a difference between points on the sparse depth image and the metric depth estimation. (Tran, "compute the absolute scale only once by using additional cues such as known object sizes.", [0037]; "inherent scale ambiguity of depth recovery", [0021]; applying an absolute scale correction factor to the depth estimations to overcome scale ambiguity and align the relative depth maps with real-world scales) Claim(s) 5-7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Tran et al (US20210065391A1) in view of Iqbal et al (US20210117661A1) and further in view of Pan et al (US20210375054A1). Regarding claim 5, the combination of Tran and Iqbal teaches its/their respective base claim(s). The combination further teaches the system of claim 1, wherein generating the first set of tracked 3D points using the odometry system includes applying one or more computer vision algorithms to estimate the AR head-mounted device's motion and applying an inertial measurement unit that includes one or more accelerometers or gyroscopes that measure acceleration and rotation respectively to determine changes in position of the AR head-mounted device. (Tran, " Structure from Motion (SfM), which aims to estimate the ego-motion of an agent (e.g., vehicle, robot, etc.) and three-dimensional (3D) scene structure of an environment by using the input of one or multiple cameras.", [0003]; "The user input devices 642 can be any of ... a motion sensing device", [0070]; Pan, "update the position of a user's device relative to the 3D model using a combination of the device's camera stream, and accelerometer and gyro information in real-time.", [0015]; "receiving an IMU pose determined from data generated by an inertial measurement unit including motion sensors", [0018]; "The motion sensing components 232 include acceleration sensor components (e.g., accelerometers 246), rotation sensor components (e.g., gyroscopes 250)", [0038]; Tran teaches using computer vision algorithms (SfM/SLAM) to estimate the device's ego-motion and further receives input from a "motion sensing device". Pan teaches that for AR tracking, it is highly advantageous to use a hybrid odometry approach that combines visual camera tracking with an Inertial Measurement Unit (IMU) comprising accelerometers and gyroscopes to robustly determine changes in position/pose in real-time) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to implement the motion sensing device of Tran as an IMU having accelerometers and gyroscopes, as taught by Pan, to provide stable and accurate tracking of the AR device even during rapid movements where visual-only tracking might fail. The combination of Tran, Iqbal and Pan also teaches other enhanced capabilities. Regarding claim 6, the combination of Tran and Iqbal teaches its/their respective base claim(s). The combination of Tran, Iqbal and Pan further teaches the system of claim 1, wherein generating the first set of tracked 3D points using the odometry system includes tracking corners of objects in view in the 2D camera image. (Pan, "The generation of a 3D model is referred to as “mapping” and typically involves locating recognizable features in the real world and recording them in the 3D model. While the features recorded in the 3D model are typically referred to as “landmarks,” they may be little more than points or edges corresponding to corners or edges of structures or items in the real world.", [0010]; Tran teaches generating a 3D model via SLAM mapping ([0044]). Pan teaches that SLAM mapping features ("landmarks") correspond to corners of structures in the real world. Incorporating Pan's tracking corners into Tran's SLAM system would provide reliable and easily identifiable static reference points for robust odometry) Regarding claim 7, the combination of Tran and Iqbal teaches its/their respective base claim(s). The combination of Tran, Iqbal and Pan further teaches the system of claim 1, wherein generating the first set of tracked 3D points using the odometry system includes tracking edges of objects in view in the 2D camera image. (Pan, "The generation of a 3D model is referred to as “mapping” and typically involves locating recognizable features in the real world and recording them in the 3D model. While the features recorded in the 3D model are typically referred to as “landmarks,” they may be little more than points or edges corresponding to corners or edges of structures or items in the real world.", [0010]; SLAM mapping features track edges of structures in the real world. Incorporating Pan's tracking edges into Tran's SLAM system would provide reliable linear reference features for the odometry tracking) Claim(s) 15-17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Tran et al (US20210065391A1) in view of Iqbal et al (US20210117661A1) and further in view of Das et al (US20210042996A1). Regarding claim 15, the combination of Tran and Iqbal teaches its/their respective base claim(s). The combination does not expressly disclose but Das teaches the system of claim 1, wherein the operations further comprise: identifying a boundary based on the second set of tracked 3D points; and removing tracked 3D points within the boundary in the first set of tracked 3D points to generate a modified first set of tracked 3D points, wherein creating the sparse depth image by projecting the first and second set of tracked 3D points onto the 2D camera image includes projecting the modified first set of tracked 3D points onto the 2D camera image. (Das, "using bounding box based cropping technique on objects in the input image sequences, key point detection of the objects", [0012]; "gives the bounding boxes on the objects in the images. Further the images are cropped by the bounding boxes and passed to a stacked hourglass network™, which gives key-points on object", [0040]; identifying an object boundary (bounding box) and cropping the region to isolate object keypoints (second set), effectively segregating/removing the background SLAM points (first set) inside that boundary to process the object independently) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine Das's bounding box cropping and point segregation with Tran and Iqbal to remove background SLAM points from the dynamic object's boundary to prevent map corruption, resulting in a modified first set of points. The combination of Tran, Iqbal and Das also teaches other enhanced capabilities. Regarding claim 16, the combination of Tran and Iqbal teaches its/their respective base claim(s). The combination of Tran, Iqbal and Das further teaches the system of claim 1, wherein the operations further comprise: identifying a boundary based on the second set of tracked 3D points; removing tracked 3D points within the boundary in the first set of tracked 3D points to generate a modified first set of tracked 3D points; and adding the second set of tracked 3D points to the modified first set of tracked 3D points to generate a third set of tracked 3D points, wherein creating the sparse depth image by projecting the first and second set of tracked 3D points onto the 2D camera image includes projecting the third set of tracked 3D points onto the 2D camera image. (Das, "perform a joint optimization by minimizing a resultant cost function to generate an optimized 3D map of the area of interest, wherein joint optimization comprises adding constraints to the bundle adjustment by integrating the plurality of objects in the SLAM", [0013]; "optimizing for all the 3D map points, the key frame poses, and the objects (Bundle adjustment) as a result, the optimized 3D map of edge points, key frame poses, object shape parameters and object poses are received as output.", [0066]; taking the segregated SLAM points (modified first set) and integrating/adding the object keypoints (second set) to generate a jointly optimized 3D map (third set). It would be obvious to integrate the object points back into the main SLAM map to produce a complete and semantically meaningful 3D representation) Regarding claim 17, the combination of Tran, Iqbal and Das teaches its/their respective base claim(s). The combination further teaches the system of claim 16, wherein the operations further comprise: removing depth data from the metric depth estimation that corresponds to the boundary to generate an updated metric depth estimation; and generating a 3D virtual representation of the scene shown in the 2D camera image by applying the updated metric depth estimation. (Das, "creating a variety of applications in augmented reality ... objects in the map is embedded along with the 3D structure obtained from the monocular SLAM framework, which gives a better and more meaningful visualization", [0068]; updating the map by embedding the isolated object models into the 3D structure yields a better 3D virtual representation/visualization of the scene for AR. It would be obvious to apply this to Tran's metric depth estimation by updating the depth data inside the object boundary with the object-specific keypoints, thereby generating an updated 3D virtual representation) Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to JIANXUN YANG whose telephone number is (571)272-9874. The examiner can normally be reached on MON-FRI: 8AM-5PM Pacific Time. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amandeep Saini can be reached on (571)272-3382. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272- 1000. /JIANXUN YANG/ Primary Examiner, Art Unit 2662 8/8/2026
Read full office action

Prosecution Timeline

Oct 30, 2024
Application Filed
Aug 12, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743866
FUNDAMENTAL MATRIX GENERATION APPARATUS, CONTROL METHOD, AND COMPUTER-READABLE MEDIUM
3y 0m to grant Granted Sep 22, 2026
Patent 12743791
SYSTEMS AND METHODS FOR MULTI-BRANCH VIDEO OBJECT DETECTION FRAMEWORK
3y 3m to grant Granted Sep 22, 2026
Patent 12743884
VERSATILE ACTION MODELS (VAMOS) FOR VIDEO UNDERSTANDING
2y 6m to grant Granted Sep 22, 2026
Patent 12738036
IMAGE PROCESSING APPARATUS, IMAGE PROCESSING METHOD, AND IMAGE PROCESSING PROGRAM
2y 5m to grant Granted Sep 15, 2026
Patent 12731362
IMAGE PROCESSING DEVICE, IMAGE PROCESSING METHOD, AND PROGRAM
3y 2m to grant Granted Sep 08, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
74%
Grant Probability
93%
With Interview (+19.3%)
2y 7m (~8m remaining)
Median Time to Grant
Low
PTA Risk
Based on 663 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month