Prosecution Insights
Last updated: August 16, 2026
Application No. 19/052,050

CAMERA-SPACE HAND MESH PREDICTION WITH DIFFERENTIAL GLOBAL POSITIONING

Non-Final OA §103
Filed
Feb 12, 2025
Priority
Feb 12, 2024 — GR 20240100094
Examiner
LIU, GORDON G
Art Unit
Tech Center
Assignee
Niantic, Inc.
OA Round
1 (Non-Final)
83%
Grant Probability
Favorable
1-2
OA Rounds
8m
Est. Remaining
98%
With Interview

Examiner Intelligence

Grants 83% — above average
83%
Career Allowance Rate
574 granted / 692 resolved
+22.9% vs TC avg
Moderate +15% lift
Without
With
+15.0%
Interview Lift
resolved cases with interview
Fast prosecutor
2y 2m
Avg Prosecution
36 currently pending
Career history
717
Total Applications
across all art units

Statute-Specific Performance

§101
7.2%
-32.8% vs TC avg
§103
77.3%
+37.3% vs TC avg
§102
3.5%
-36.5% vs TC avg
§112
2.6%
-37.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 692 resolved cases

Office Action

§103
DETAILED ACTION The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims 1-20 are pending under this Office action. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 9, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Ge, etc. (US 20200402305 A1) in view of Spurr, etc. (US 20210233273 A1). Regarding claim 1, Ge teaches that a computer-implemented method (See Ge: Fig. 1, and [0019], “FIG. 1 is a block diagram showing an example messaging system 100 for exchanging data (e.g., messages and associated content) over a network 106. The messaging system 100 includes multiple client devices 102, each of which hosts a number of applications, including a messaging client application 104 and a AR/VR application 105. Each messaging client application 104 is communicatively coupled to other instances of the messaging client application 104, the AR/VR application 105, and a messaging server system 108 via a network 106 (e.g., the Internet)”) comprising: accessing an image of a hand captured by a camera (See Ge: Fig. 1, and [0022], “In order for AR/VR application 105 to generate the 3D hand model directly from a captured RGB image, the AR/VR application 105 obtains one or more trained machine learning techniques from the hand shape and pose estimation system 124 and/or messaging server system 108. Hand shape and pose estimation system 124 trains the machine learning techniques to generate the 3D hand model in two training phases. In the first training phase, the hand shape and pose estimation system 124 obtains a first plurality of input images that include synthetic representations of a hand. These synthetic representations include a depiction of an animated hand and also provide the ground truth information about the shape and pose of the animated hand (e.g., a 3D mesh and 3D pose ground truth information). A first machine learning technique (e.g., a two-stacked hourglass network) is initially trained based on a first feature (e.g., a heat-map loss) of the first plurality of images. A second machine learning technique (e.g., a 3D pose regressor network) is initially trained, separately from the first machine learning technique, based on a second feature (e.g., a 3D pose loss) of the first plurality of images. The first and second machine learning techniques are then trained together with a graph CNN based on the first plurality of images with the combined mesh, pose and heat-map losses. In the second training phase, a second plurality of input images are obtained that include real-world depictions of a hand and reference 3D depth maps captured using a depth camera or sensor. A pseudo-ground truth mesh of the real-world depictions of the hand is generated using the graph CNN trained in the first training phase. The first and second machine learning techniques are then trained together with the graph CNN based on the pseudo-ground truth mesh, the real-world depictions of the hand, and the reference 3D depth maps. In some implementations, in the second training phase, the first machine learning technique (e.g., the stacked hourglass network) is first trained with the first feature of the second plurality of images (e.g., the heat-map loss) and then all the machine learning techniques are trained or fine-tuned based on a heat-map loss, a depth loss, and a mesh loss”); predicting a set of 2D keypoints and a set of root-relative 3D vertices based on the image (See Ge: Figs. 1-5, and [0023], “For example, the input RGB image depicting a computer-generated hand pose is passed through a two-stacked hourglass network for 2D hand pose estimation. The estimated 2D heat maps, combined with the image feature maps, are encoded as latent feature vectors by a residual network. The latent feature vector is then input to a graph CNN to infer the 3D coordinates of the mesh vertices. Finally, the 3D hand pose is linearly regressed from the 3D hand mesh. Specifically, the machine learning techniques are initially trained on the synthetic dataset in a fully supervised manner with heat-map loss, 3D mesh loss, and 3D pose loss. The machine learning techniques are then optimized using real-world datasets that include a depth map without 3D mesh or 3D pose ground truth information. Particularly, the networks are fine-tuned in a weakly-supervised manner by rendering the full 3D hand mesh to a depth map and minimizing the depth map loss against the reference depth map. These processes are described in more detail below in connection with FIGS. 6 and 7”; [0041], “In some implementations, the first plurality of images (also referred to as synthetic images), stored in synthetic and real hand training images 209, provides the labels of both 3D hand joint locations and full 3D hand meshes. A 3D hand model is generated, rigged with joints, and then photorealistic textures are applied on the 3D hand model as well as natural lighting using high-dynamic range (HDR) images. The variations of the hand are modeled by creating blend shapes with different shapes and ratios, and then random weights are applied to the blend shapes. Hand poses from 500 common hand gestures and 1000 unique camera viewpoints are created and captured in the first plurality of images. To simulate real-world diversity, 30 lightings and five skin colors are used. The hand is rendered using global illumination. In some implementations, the first plurality of images includes 375,000 hand RGB images with large variations. In some embodiments, only a portion (e.g., 315,000) of the first plurality of images are used in the first training phase to train the machine learning techniques. During training or before, each rendered hand in the first plurality of images is cropped from the image and blended with a randomly selected background image (e.g., a city image, a living room image, or any other suitable image obtained randomly or pseudo-randomly from a background image server(s)). To do this, the system obtains an image that contains a rendered 3D hand mesh, the 3D hand mesh is cropped and extracted from the image, a background image is randomly selected, and the cropped 3D hand mesh is combined with the selected background image and stored as a new image to be used in the first training phase. Particularly, the first plurality of images used to train the machine learning techniques in the first training phase is modified to include a variety of simulated or computer-generated hand models overlaid on top of or blended with a background image to provide a more realistic representation”; [0068], “SVMs are supervised learning models with associated learning algorithms that are configured to recognize patterns. Given a set of training examples, with each marked as belonging to one of two categories, an SVM training algorithm builds a model that assigns new examples into one category or the other, making it a non-probabilistic binary linear classifier. An SVM model is a representation of the examples as points in space, mapped so that the examples of the separate categories are divided by a clear gap that is as wide as possible. New examples are then mapped into that same space and predicted to belong to a category based on which side of the gap they fall on”; and [0069], “An embodiment of the graph CNN module 418 is shown in FIG. 5. Particularly, the graph CNN module 418 generates 3D coordinates of vertices in the hand mesh and estimates the 3D hand pose from the mesh. In this way, the graph CNN module 418 models, based on features extracted by other machine learning technique modules of FIG. 5, a post of a hand depicted in a monocular image by adjusting skeletal joint positions of a 3D hand mesh and also models a shape of the hand in the monocular image by adjusting blend shape values of the 3D hand mesh representing surface features of the hand depicted in the monocular image. The resulting 3D hand mesh is then generated for display. An illustrative 3D mesh 803 with its different viewpoints 804 is shown in FIG. 8. A 3D mesh can be represented by an undirected graph”. Note that features (skeleton joint positions/vertices extraction from the hand image and modeling the hand mesh with joints/keypoints, which is inherently root-relative via skeletal structure, using CNN is mapped to the predicting the 2D keypoints and a set of root-relative 3D vertices); obtaining, using the set of root-relative 3D vertices, a set of root-relative 3D keypoints that correspond to the set of 2D keypoints (See Ge: Fig. 5, and [0069], “An embodiment of the graph CNN module 418 is shown in FIG. 5. Particularly, the graph CNN module 418 generates 3D coordinates of vertices in the hand mesh and estimates the 3D hand pose from the mesh. In this way, the graph CNN module 418 models, based on features extracted by other machine learning technique modules of FIG. 5, a post of a hand depicted in a monocular image by adjusting skeletal joint positions of a 3D hand mesh and also models a shape of the hand in the monocular image by adjusting blend shape values of the 3D hand mesh representing surface features of the hand depicted in the monocular image. The resulting 3D hand mesh is then generated for display. An illustrative 3D mesh 803 with its different viewpoints 804 is shown in FIG. 8. A 3D mesh can be represented by an undirected graph”. Note that adjusting the skeletal joint positions from vertices and deriving keypoints and joints from mesh vertices are mapped to obtain a set of root-relative 3D keypoints); generating a global camera space mesh prediction of the hand in 3D space based on the set of 2D keypoints and the set of root-relative 3D keypoints (See Ge: Figs. 5-8, and [0081], “At operation 605, the hand shape and pose estimation system 124 generates, for display, the 3D hand mesh adjusted to model the pose and shape of the hand depicted in the monocular image. For example, a 3D mesh 803 (FIG. 8) is adjusted to model the pose (e.g., using 3D pose graph 805) and shape of a hand depicted in given hand training image data 801. This 3D mesh 803 and 3D pose graph 805 are returned to client device 102 for presentation in the AR/VR application 105”); and outputting a virtual element in a virtual space (See Ge: Fig. 1, and [0015], “Typically, VR and AR systems display a 3D hand representation of a real-world user's hand by focusing on sparse 3D hand joint locations and ignore dense 3D hand shapes. Specifically, these systems capture an image using a red, green, blue (RGB) camera and determine the joint positions of the hand from the image. Given the diversity and complexity of real-world hand shapes, such typical systems simply obtain a generic 3D hand model and use the determined joint positions to fit the obtained generic hand model to resemble the joint positions (e.g., the finger positions) of the real-world hand. Such generic representations of the hand fail to consider the shape of the hand and other surface features of the hand, so the 3D hand model that is presented is not very accurate. This makes user interactions with content in the VR and AR systems more difficult and less realistic, which detracts from the overall user experience”; and Figs. 9-10, and [0091], “FIG. 9 is a block diagram illustrating an example software architecture 906, which may be used in conjunction with various hardware architectures herein described. FIG. 9 is a non-limiting example of a software architecture and it will be appreciated that many other architectures may be implemented to facilitate the functionality described herein. The software architecture 906 may execute on hardware such as machine 1000 of FIG. 10 that includes, among other things, processors 1004, memory 1014, and input/output (I/O) components 1018. A representative hardware layer 952 is illustrated and can represent, for example, the machine 1000 of FIG. 10. The representative hardware layer 952 includes a processing unit 954 having associated executable instructions 904. Executable instructions 904 represent the executable instructions of the software architecture 906, including implementation of the methods, components, and so forth described herein. The hardware layer 952 also includes memory and/or storage modules memory/storage 956, which also have executable instructions 904. The hardware layer 952 may also comprise other hardware 958”; and [0097], “The applications 916 may use built-in operating system functions (e.g., kernel 922, services 924, and/or drivers 926), libraries 920, and frameworks/middleware 918 to create UIs to interact with users of the system. Alternatively, or additionally, in some systems, interactions with a user may occur through a presentation layer, such as presentation layer 914. In these systems, the application/component “logic” can be separated from the aspects of the application/component that interact with a user”. Note that the animated model of a hand for user interaction with the 3D hand mesh in the virtual space VR/AR application is mapped to the outputting a virtual element in a virtual space) based on the global camera space mesh prediction of the hand. However, Ge fails to explicitly disclose that based on the global camera space mesh prediction of the hand. However, Spurr teaches that based on the global camera space mesh prediction of the hand (See Spurr: Fig. 1, and [0068], “Vision-based reconstruction of the 3-D pose of human hands is a difficult problem that has applications in many domains. In at least one embodiment, a marker-less approach relies on multiple cameras, or a single depth camera. Estimating the full 3-D pose and dense surface of human hands from 2-D imagery alone is often challenging due to the dexterity of the human hand, self-occlusions, varying lighting conditions and interactions with objects. Moreover, a given 2-D point in the image plane may correspond to multiple 3-D points in world space, all of which project onto that same 2-D point. This sometimes makes 3-D hand pose estimation from monocular imagery an ill-posed inverse problem in which depth and the resulting scale ambiguity pose a significant difficulty”. Note that multiple cameras in the 3D environments is mapped to the global camera space). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention was effectively filed to modify Ge to have based on the global camera space mesh prediction of the hand as taught by Spurr in order to improve the robustness and stability of the system (See Spurr: Fig. 1, and [0066], “At least one embodiment does not require any depth data, nor does the embodiment make use of MANO or synthetic data. Instead, the embodiment relies solely on novel formulation of loss functions. At least one embodiment enhances accuracy and robustness of 3-D hand pose estimation using additional weakly-labeled and unlabeled data. In general, this provides improvements over methods that only make use of additional 2-D data without enhancing the 3-D estimation. At least one embodiment allows enhanced solutions for human machine interaction. Products that require human input in the way of pointing, as an example, will benefit”). Ge teaches a method and system that may generate 3D hand mesh by adjusting the skeletal joint positions of a 3D hand mesh using the trained CNN model to process the 2D hand images; while Spurr teaches a system and method that may determine the pose of a human hand from 2D images captured by multiple cameras in order to estimate the full 3D pose and dense surface of human hands more accurately. Therefore, it is obvious to one of ordinary skill in the art to modify Ge by Spurr to reconstruct the 3D hand mesh more accurately in camera space (using multiple cameras). The motivation to modify Ge by Spurr is “Use of known technique to improve similar devices (methods, or products) in the same way”. Regarding claim 9, Ge and Spurr teach all the features with respect to claim 1 as outlined above. Further, Ge and Spurr teach that a non-transitory computer-readable medium storing instructions that, when executed by a computing system, cause the computing system to perform operations (See Ge: Fig. 1, and [0019], “FIG. 1 is a block diagram showing an example messaging system 100 for exchanging data (e.g., messages and associated content) over a network 106. The messaging system 100 includes multiple client devices 102, each of which hosts a number of applications, including a messaging client application 104 and a AR/VR application 105. Each messaging client application 104 is communicatively coupled to other instances of the messaging client application 104, the AR/VR application 105, and a messaging server system 108 via a network 106 (e.g., the Internet)”) comprising: accessing an image of a hand captured by a camera (See Ge: Fig. 1, and [0022], “In order for AR/VR application 105 to generate the 3D hand model directly from a captured RGB image, the AR/VR application 105 obtains one or more trained machine learning techniques from the hand shape and pose estimation system 124 and/or messaging server system 108. Hand shape and pose estimation system 124 trains the machine learning techniques to generate the 3D hand model in two training phases. In the first training phase, the hand shape and pose estimation system 124 obtains a first plurality of input images that include synthetic representations of a hand. These synthetic representations include a depiction of an animated hand and also provide the ground truth information about the shape and pose of the animated hand (e.g., a 3D mesh and 3D pose ground truth information). A first machine learning technique (e.g., a two-stacked hourglass network) is initially trained based on a first feature (e.g., a heat-map loss) of the first plurality of images. A second machine learning technique (e.g., a 3D pose regressor network) is initially trained, separately from the first machine learning technique, based on a second feature (e.g., a 3D pose loss) of the first plurality of images. The first and second machine learning techniques are then trained together with a graph CNN based on the first plurality of images with the combined mesh, pose and heat-map losses. In the second training phase, a second plurality of input images are obtained that include real-world depictions of a hand and reference 3D depth maps captured using a depth camera or sensor. A pseudo-ground truth mesh of the real-world depictions of the hand is generated using the graph CNN trained in the first training phase. The first and second machine learning techniques are then trained together with the graph CNN based on the pseudo-ground truth mesh, the real-world depictions of the hand, and the reference 3D depth maps. In some implementations, in the second training phase, the first machine learning technique (e.g., the stacked hourglass network) is first trained with the first feature of the second plurality of images (e.g., the heat-map loss) and then all the machine learning techniques are trained or fine-tuned based on a heat-map loss, a depth loss, and a mesh loss”); predicting a set of 2D keypoints and a set of root-relative 3D vertices based on the image (See Ge: Figs. 1-5, and [0023], “For example, the input RGB image depicting a computer-generated hand pose is passed through a two-stacked hourglass network for 2D hand pose estimation. The estimated 2D heat maps, combined with the image feature maps, are encoded as latent feature vectors by a residual network. The latent feature vector is then input to a graph CNN to infer the 3D coordinates of the mesh vertices. Finally, the 3D hand pose is linearly regressed from the 3D hand mesh. Specifically, the machine learning techniques are initially trained on the synthetic dataset in a fully supervised manner with heat-map loss, 3D mesh loss, and 3D pose loss. The machine learning techniques are then optimized using real-world datasets that include a depth map without 3D mesh or 3D pose ground truth information. Particularly, the networks are fine-tuned in a weakly-supervised manner by rendering the full 3D hand mesh to a depth map and minimizing the depth map loss against the reference depth map. These processes are described in more detail below in connection with FIGS. 6 and 7”; [0041], “In some implementations, the first plurality of images (also referred to as synthetic images), stored in synthetic and real hand training images 209, provides the labels of both 3D hand joint locations and full 3D hand meshes. A 3D hand model is generated, rigged with joints, and then photorealistic textures are applied on the 3D hand model as well as natural lighting using high-dynamic range (HDR) images. The variations of the hand are modeled by creating blend shapes with different shapes and ratios, and then random weights are applied to the blend shapes. Hand poses from 500 common hand gestures and 1000 unique camera viewpoints are created and captured in the first plurality of images. To simulate real-world diversity, 30 lightings and five skin colors are used. The hand is rendered using global illumination. In some implementations, the first plurality of images includes 375,000 hand RGB images with large variations. In some embodiments, only a portion (e.g., 315,000) of the first plurality of images are used in the first training phase to train the machine learning techniques. During training or before, each rendered hand in the first plurality of images is cropped from the image and blended with a randomly selected background image (e.g., a city image, a living room image, or any other suitable image obtained randomly or pseudo-randomly from a background image server(s)). To do this, the system obtains an image that contains a rendered 3D hand mesh, the 3D hand mesh is cropped and extracted from the image, a background image is randomly selected, and the cropped 3D hand mesh is combined with the selected background image and stored as a new image to be used in the first training phase. Particularly, the first plurality of images used to train the machine learning techniques in the first training phase is modified to include a variety of simulated or computer-generated hand models overlaid on top of or blended with a background image to provide a more realistic representation”; [0068], “SVMs are supervised learning models with associated learning algorithms that are configured to recognize patterns. Given a set of training examples, with each marked as belonging to one of two categories, an SVM training algorithm builds a model that assigns new examples into one category or the other, making it a non-probabilistic binary linear classifier. An SVM model is a representation of the examples as points in space, mapped so that the examples of the separate categories are divided by a clear gap that is as wide as possible. New examples are then mapped into that same space and predicted to belong to a category based on which side of the gap they fall on”; and [0069], “An embodiment of the graph CNN module 418 is shown in FIG. 5. Particularly, the graph CNN module 418 generates 3D coordinates of vertices in the hand mesh and estimates the 3D hand pose from the mesh. In this way, the graph CNN module 418 models, based on features extracted by other machine learning technique modules of FIG. 5, a post of a hand depicted in a monocular image by adjusting skeletal joint positions of a 3D hand mesh and also models a shape of the hand in the monocular image by adjusting blend shape values of the 3D hand mesh representing surface features of the hand depicted in the monocular image. The resulting 3D hand mesh is then generated for display. An illustrative 3D mesh 803 with its different viewpoints 804 is shown in FIG. 8. A 3D mesh can be represented by an undirected graph”. Note that features (skeleton joint positions/vertices extraction from the hand image and modeling the hand mesh with joints/keypoints, which is inherently root-relative via skeletal structure, using CNN is mapped to the predicting the 2D keypoints and a set of root-relative 3D vertices); obtaining, using the set of root-relative 3D vertices, a set of root-relative 3D keypoints that correspond to the set of 2D keypoints (See Ge: Fig. 5, and [0069], “An embodiment of the graph CNN module 418 is shown in FIG. 5. Particularly, the graph CNN module 418 generates 3D coordinates of vertices in the hand mesh and estimates the 3D hand pose from the mesh. In this way, the graph CNN module 418 models, based on features extracted by other machine learning technique modules of FIG. 5, a post of a hand depicted in a monocular image by adjusting skeletal joint positions of a 3D hand mesh and also models a shape of the hand in the monocular image by adjusting blend shape values of the 3D hand mesh representing surface features of the hand depicted in the monocular image. The resulting 3D hand mesh is then generated for display. An illustrative 3D mesh 803 with its different viewpoints 804 is shown in FIG. 8. A 3D mesh can be represented by an undirected graph”. Note that adjusting the skeletal joint positions from vertices and deriving keypoints and joints from mesh vertices are mapped to obtain a set of root-relative 3D keypoints); generating a global camera space mesh prediction of the hand in 3D space based on the set of 2D keypoints and the set of root-relative 3D keypoints (See Ge: Figs. 5-8, and [0081], “At operation 605, the hand shape and pose estimation system 124 generates, for display, the 3D hand mesh adjusted to model the pose and shape of the hand depicted in the monocular image. For example, a 3D mesh 803 (FIG. 8) is adjusted to model the pose (e.g., using 3D pose graph 805) and shape of a hand depicted in given hand training image data 801. This 3D mesh 803 and 3D pose graph 805 are returned to client device 102 for presentation in the AR/VR application 105”); and outputting a virtual element in a virtual space (See Ge: Fig. 1, and [0015], “Typically, VR and AR systems display a 3D hand representation of a real-world user's hand by focusing on sparse 3D hand joint locations and ignore dense 3D hand shapes. Specifically, these systems capture an image using a red, green, blue (RGB) camera and determine the joint positions of the hand from the image. Given the diversity and complexity of real-world hand shapes, such typical systems simply obtain a generic 3D hand model and use the determined joint positions to fit the obtained generic hand model to resemble the joint positions (e.g., the finger positions) of the real-world hand. Such generic representations of the hand fail to consider the shape of the hand and other surface features of the hand, so the 3D hand model that is presented is not very accurate. This makes user interactions with content in the VR and AR systems more difficult and less realistic, which detracts from the overall user experience”; and Figs. 9-10, and [0091], “FIG. 9 is a block diagram illustrating an example software architecture 906, which may be used in conjunction with various hardware architectures herein described. FIG. 9 is a non-limiting example of a software architecture and it will be appreciated that many other architectures may be implemented to facilitate the functionality described herein. The software architecture 906 may execute on hardware such as machine 1000 of FIG. 10 that includes, among other things, processors 1004, memory 1014, and input/output (I/O) components 1018. A representative hardware layer 952 is illustrated and can represent, for example, the machine 1000 of FIG. 10. The representative hardware layer 952 includes a processing unit 954 having associated executable instructions 904. Executable instructions 904 represent the executable instructions of the software architecture 906, including implementation of the methods, components, and so forth described herein. The hardware layer 952 also includes memory and/or storage modules memory/storage 956, which also have executable instructions 904. The hardware layer 952 may also comprise other hardware 958”; and [0097], “The applications 916 may use built-in operating system functions (e.g., kernel 922, services 924, and/or drivers 926), libraries 920, and frameworks/middleware 918 to create UIs to interact with users of the system. Alternatively, or additionally, in some systems, interactions with a user may occur through a presentation layer, such as presentation layer 914. In these systems, the application/component “logic” can be separated from the aspects of the application/component that interact with a user”. Note that the animated model of a hand for user interaction with the 3D hand mesh in the virtual space VR/AR application is mapped to the outputting a virtual element in a virtual space) based on the global camera space mesh prediction of the hand (See Spurr: Fig. 1, and [0068], “Vision-based reconstruction of the 3-D pose of human hands is a difficult problem that has applications in many domains. In at least one embodiment, a marker-less approach relies on multiple cameras, or a single depth camera. Estimating the full 3-D pose and dense surface of human hands from 2-D imagery alone is often challenging due to the dexterity of the human hand, self-occlusions, varying lighting conditions and interactions with objects. Moreover, a given 2-D point in the image plane may correspond to multiple 3-D points in world space, all of which project onto that same 2-D point. This sometimes makes 3-D hand pose estimation from monocular imagery an ill-posed inverse problem in which depth and the resulting scale ambiguity pose a significant difficulty”. Note that multiple cameras in the 3D environments is mapped to the global camera space). Regarding claim 17, Ge and Spurr teach all the features with respect to claim 1 as outlined above. Further, Ge and Spurr teach that a client device for predicting a hand mesh, the client device (See Ge: Fig. 1, and [0019], “FIG. 1 is a block diagram showing an example messaging system 100 for exchanging data (e.g., messages and associated content) over a network 106. The messaging system 100 includes multiple client devices 102, each of which hosts a number of applications, including a messaging client application 104 and a AR/VR application 105. Each messaging client application 104 is communicatively coupled to other instances of the messaging client application 104, the AR/VR application 105, and a messaging server system 108 via a network 106 (e.g., the Internet)”) comprising: a display (See Ge: Fig. 9, and [0093], “The operating system 902 may manage hardware resources and provide common services. The operating system 902 may include, for example, a kernel 922, services 924, and drivers 926. The kernel 922 may act as an abstraction layer between the hardware and the other software layers. For example, the kernel 922 may be responsible for memory management, processor management (e.g., scheduling), component management, networking, security settings, and so on. The services 924 may provide other common services for the other software layers. The drivers 926 are responsible for controlling or interfacing with the underlying hardware. For instance, the drivers 926 include display drivers, camera drivers, Bluetooth® drivers, flash memory drivers, serial communication drivers (e.g., Universal Serial Bus (USB) drivers), Wi-Fi® drivers, audio drivers, power management drivers, and so forth depending on the hardware configuration”; and [0094], “The libraries 920 provide a common infrastructure that is used by the applications 916 and/or other components and/or layers. The libraries 920 provide functionality that allows other software components to perform tasks in an easier fashion than to interface directly with the underlying operating system 902 functionality (e.g., kernel 922, services 924 and/or drivers 926). The libraries 920 may include system libraries 944 (e.g., C standard library) that may provide functions such as memory allocation functions, string manipulation functions, mathematical functions, and the like. In addition, the libraries 920 may include API libraries 946 such as media libraries (e.g., libraries to support presentation and manipulation of various media format such as MPREG4, H.264, MP3, AAC, AMR, JPG, PNG), graphics libraries (e.g., an OpenGL framework that may be used to render two-dimensional and three-dimensional in a graphic content on a display), database libraries (e.g., SQLite that may provide various relational database functions), web libraries (e.g., WebKit that may provide web browsing functionality), and the like. The libraries 920 may also include a wide variety of other libraries 948 to provide many other APIs to the applications 916 and other software components/modules”); a camera (See Ge: Fig. 7, and [0087], “At operation 705, the hand shape and pose estimation system 124 obtains a second plurality of input images that include real-world depictions of a hand and reference 3D depth maps captured using a depth camera. For example, machine learning techniques network 410 receives real hand training image data 403. An illustrative synthetic hand training image data 801 and its corresponding output is shown in a second row 820 of FIG. 8”); one or more processors (See Ge: Fig. 9, and [0091], “FIG. 9 is a block diagram illustrating an example software architecture 906, which may be used in conjunction with various hardware architectures herein described. FIG. 9 is a non-limiting example of a software architecture and it will be appreciated that many other architectures may be implemented to facilitate the functionality described herein. The software architecture 906 may execute on hardware such as machine 1000 of FIG. 10 that includes, among other things, processors 1004, memory 1014, and input/output (I/O) components 1018. A representative hardware layer 952 is illustrated and can represent, for example, the machine 1000 of FIG. 10. The representative hardware layer 952 includes a processing unit 954 having associated executable instructions 904. Executable instructions 904 represent the executable instructions of the software architecture 906, including implementation of the methods, components, and so forth described herein. The hardware layer 952 also includes memory and/or storage modules memory/storage 956, which also have executable instructions 904. The hardware layer 952 may also comprise other hardware 958”); and memory storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations (See Ge: Fig. 9, and [0091], “FIG. 9 is a block diagram illustrating an example software architecture 906, which may be used in conjunction with various hardware architectures herein described. FIG. 9 is a non-limiting example of a software architecture and it will be appreciated that many other architectures may be implemented to facilitate the functionality described herein. The software architecture 906 may execute on hardware such as machine 1000 of FIG. 10 that includes, among other things, processors 1004, memory 1014, and input/output (I/O) components 1018. A representative hardware layer 952 is illustrated and can represent, for example, the machine 1000 of FIG. 10. The representative hardware layer 952 includes a processing unit 954 having associated executable instructions 904. Executable instructions 904 represent the executable instructions of the software architecture 906, including implementation of the methods, components, and so forth described herein. The hardware layer 952 also includes memory and/or storage modules memory/storage 956, which also have executable instructions 904. The hardware layer 952 may also comprise other hardware 958”) comprising: accessing an image of a hand captured by the camera (See Ge: Fig. 1, and [0022], “In order for AR/VR application 105 to generate the 3D hand model directly from a captured RGB image, the AR/VR application 105 obtains one or more trained machine learning techniques from the hand shape and pose estimation system 124 and/or messaging server system 108. Hand shape and pose estimation system 124 trains the machine learning techniques to generate the 3D hand model in two training phases. In the first training phase, the hand shape and pose estimation system 124 obtains a first plurality of input images that include synthetic representations of a hand. These synthetic representations include a depiction of an animated hand and also provide the ground truth information about the shape and pose of the animated hand (e.g., a 3D mesh and 3D pose ground truth information). A first machine learning technique (e.g., a two-stacked hourglass network) is initially trained based on a first feature (e.g., a heat-map loss) of the first plurality of images. A second machine learning technique (e.g., a 3D pose regressor network) is initially trained, separately from the first machine learning technique, based on a second feature (e.g., a 3D pose loss) of the first plurality of images. The first and second machine learning techniques are then trained together with a graph CNN based on the first plurality of images with the combined mesh, pose and heat-map losses. In the second training phase, a second plurality of input images are obtained that include real-world depictions of a hand and reference 3D depth maps captured using a depth camera or sensor. A pseudo-ground truth mesh of the real-world depictions of the hand is generated using the graph CNN trained in the first training phase. The first and second machine learning techniques are then trained together with the graph CNN based on the pseudo-ground truth mesh, the real-world depictions of the hand, and the reference 3D depth maps. In some implementations, in the second training phase, the first machine learning technique (e.g., the stacked hourglass network) is first trained with the first feature of the second plurality of images (e.g., the heat-map loss) and then all the machine learning techniques are trained or fine-tuned based on a heat-map loss, a depth loss, and a mesh loss”); predicting a set of 2D keypoints and a set of root-relative 3D vertices based on the image (See Ge: Figs. 1-5, and [0023], “For example, the input RGB image depicting a computer-generated hand pose is passed through a two-stacked hourglass network for 2D hand pose estimation. The estimated 2D heat maps, combined with the image feature maps, are encoded as latent feature vectors by a residual network. The latent feature vector is then input to a graph CNN to infer the 3D coordinates of the mesh vertices. Finally, the 3D hand pose is linearly regressed from the 3D hand mesh. Specifically, the machine learning techniques are initially trained on the synthetic dataset in a fully supervised manner with heat-map loss, 3D mesh loss, and 3D pose loss. The machine learning techniques are then optimized using real-world datasets that include a depth map without 3D mesh or 3D pose ground truth information. Particularly, the networks are fine-tuned in a weakly-supervised manner by rendering the full 3D hand mesh to a depth map and minimizing the depth map loss against the reference depth map. These processes are described in more detail below in connection with FIGS. 6 and 7”; [0041], “In some implementations, the first plurality of images (also referred to as synthetic images), stored in synthetic and real hand training images 209, provides the labels of both 3D hand joint locations and full 3D hand meshes. A 3D hand model is generated, rigged with joints, and then photorealistic textures are applied on the 3D hand model as well as natural lighting using high-dynamic range (HDR) images. The variations of the hand are modeled by creating blend shapes with different shapes and ratios, and then random weights are applied to the blend shapes. Hand poses from 500 common hand gestures and 1000 unique camera viewpoints are created and captured in the first plurality of images. To simulate real-world diversity, 30 lightings and five skin colors are used. The hand is rendered using global illumination. In some implementations, the first plurality of images includes 375,000 hand RGB images with large variations. In some embodiments, only a portion (e.g., 315,000) of the first plurality of images are used in the first training phase to train the machine learning techniques. During training or before, each rendered hand in the first plurality of images is cropped from the image and blended with a randomly selected background image (e.g., a city image, a living room image, or any other suitable image obtained randomly or pseudo-randomly from a background image server(s)). To do this, the system obtains an image that contains a rendered 3D hand mesh, the 3D hand mesh is cropped and extracted from the image, a background image is randomly selected, and the cropped 3D hand mesh is combined with the selected background image and stored as a new image to be used in the first training phase. Particularly, the first plurality of images used to train the machine learning techniques in the first training phase is modified to include a variety of simulated or computer-generated hand models overlaid on top of or blended with a background image to provide a more realistic representation”; [0068], “SVMs are supervised learning models with associated learning algorithms that are configured to recognize patterns. Given a set of training examples, with each marked as belonging to one of two categories, an SVM training algorithm builds a model that assigns new examples into one category or the other, making it a non-probabilistic binary linear classifier. An SVM model is a representation of the examples as points in space, mapped so that the examples of the separate categories are divided by a clear gap that is as wide as possible. New examples are then mapped into that same space and predicted to belong to a category based on which side of the gap they fall on”; and [0069], “An embodiment of the graph CNN module 418 is shown in FIG. 5. Particularly, the graph CNN module 418 generates 3D coordinates of vertices in the hand mesh and estimates the 3D hand pose from the mesh. In this way, the graph CNN module 418 models, based on features extracted by other machine learning technique modules of FIG. 5, a post of a hand depicted in a monocular image by adjusting skeletal joint positions of a 3D hand mesh and also models a shape of the hand in the monocular image by adjusting blend shape values of the 3D hand mesh representing surface features of the hand depicted in the monocular image. The resulting 3D hand mesh is then generated for display. An illustrative 3D mesh 803 with its different viewpoints 804 is shown in FIG. 8. A 3D mesh can be represented by an undirected graph”. Note that features (skeleton joint positions/vertices extraction from the hand image and modeling the hand mesh with joints/keypoints, which is inherently root-relative via skeletal structure, using CNN is mapped to the predicting the 2D keypoints and a set of root-relative 3D vertices); obtaining, using the set of root-relative 3D vertices, a set of root-relative 3D keypoints that correspond to the set of 2D keypoints (See Ge: Fig. 5, and [0069], “An embodiment of the graph CNN module 418 is shown in FIG. 5. Particularly, the graph CNN module 418 generates 3D coordinates of vertices in the hand mesh and estimates the 3D hand pose from the mesh. In this way, the graph CNN module 418 models, based on features extracted by other machine learning technique modules of FIG. 5, a post of a hand depicted in a monocular image by adjusting skeletal joint positions of a 3D hand mesh and also models a shape of the hand in the monocular image by adjusting blend shape values of the 3D hand mesh representing surface features of the hand depicted in the monocular image. The resulting 3D hand mesh is then generated for display. An illustrative 3D mesh 803 with its different viewpoints 804 is shown in FIG. 8. A 3D mesh can be represented by an undirected graph”. Note that adjusting the skeletal joint positions from vertices and deriving keypoints and joints from mesh vertices are mapped to obtain a set of root-relative 3D keypoints); generating a global camera space mesh prediction of the hand in 3D space based on the set of 2D keypoints and the set of root-relative 3D keypoints (See Ge: Figs. 5-8, and [0081], “At operation 605, the hand shape and pose estimation system 124 generates, for display, the 3D hand mesh adjusted to model the pose and shape of the hand depicted in the monocular image. For example, a 3D mesh 803 (FIG. 8) is adjusted to model the pose (e.g., using 3D pose graph 805) and shape of a hand depicted in given hand training image data 801. This 3D mesh 803 and 3D pose graph 805 are returned to client device 102 for presentation in the AR/VR application 105”); and displaying, on the display (See Ge: Fig. 1, and [0036], “The database 120 also stores annotation data, in the example form of filters, in an annotation table 212. Database 120 also stores annotated content received in the annotation table 212. Filters for which data is stored within the annotation table 212 are associated with and applied to videos (for which data is stored in a video table 210) and/or images (for which data is stored in an image table 208). Filters, in one example, are overlays that are displayed as overlaid on an image or video during presentation to a recipient user. Filters may be of various types, including user-selected filters from a gallery of filters presented to a sending user by the messaging client application 104 when the sending user is composing a message. Other types of filters include geolocation filters (also known as geo-filters), which may be presented to a sending user based on geographic location. For example, geolocation filters specific to a neighborhood or special location may be presented within a UI by the messaging client application 104, based on geolocation information determined by a Global Positioning System (GPS) unit of the client device 102. Another type of filter is a data filter, which may be selectively presented to a sending user by the messaging client application 104, based on other inputs or information gathered by the client device 102 during the message creation process. Examples of data filters include current temperature at a specific location, a current speed at which a sending user is traveling, battery life for a client device 102, or the current time”), a virtual element in a virtual space (See Ge: Fig. 1, and [0015], “Typically, VR and AR systems display a 3D hand representation of a real-world user's hand by focusing on sparse 3D hand joint locations and ignore dense 3D hand shapes. Specifically, these systems capture an image using a red, green, blue (RGB) camera and determine the joint positions of the hand from the image. Given the diversity and complexity of real-world hand shapes, such typical systems simply obtain a generic 3D hand model and use the determined joint positions to fit the obtained generic hand model to resemble the joint positions (e.g., the finger positions) of the real-world hand. Such generic representations of the hand fail to consider the shape of the hand and other surface features of the hand, so the 3D hand model that is presented is not very accurate. This makes user interactions with content in the VR and AR systems more difficult and less realistic, which detracts from the overall user experience”; and Figs. 9-10, and [0091], “FIG. 9 is a block diagram illustrating an example software architecture 906, which may be used in conjunction with various hardware architectures herein described. FIG. 9 is a non-limiting example of a software architecture and it will be appreciated that many other architectures may be implemented to facilitate the functionality described herein. The software architecture 906 may execute on hardware such as machine 1000 of FIG. 10 that includes, among other things, processors 1004, memory 1014, and input/output (I/O) components 1018. A representative hardware layer 952 is illustrated and can represent, for example, the machine 1000 of FIG. 10. The representative hardware layer 952 includes a processing unit 954 having associated executable instructions 904. Executable instructions 904 represent the executable instructions of the software architecture 906, including implementation of the methods, components, and so forth described herein. The hardware layer 952 also includes memory and/or storage modules memory/storage 956, which also have executable instructions 904. The hardware layer 952 may also comprise other hardware 958”; and [0097], “The applications 916 may use built-in operating system functions (e.g., kernel 922, services 924, and/or drivers 926), libraries 920, and frameworks/middleware 918 to create UIs to interact with users of the system. Alternatively, or additionally, in some systems, interactions with a user may occur through a presentation layer, such as presentation layer 914. In these systems, the application/component “logic” can be separated from the aspects of the application/component that interact with a user”. Note that the animated model of a hand for user interaction with the 3D hand mesh in the virtual space VR/AR application is mapped to the outputting a virtual element in a virtual space) based on the global camera space mesh prediction of the hand (See Spurr: Fig. 1, and [0068], “Vision-based reconstruction of the 3-D pose of human hands is a difficult problem that has applications in many domains. In at least one embodiment, a marker-less approach relies on multiple cameras, or a single depth camera. Estimating the full 3-D pose and dense surface of human hands from 2-D imagery alone is often challenging due to the dexterity of the human hand, self-occlusions, varying lighting conditions and interactions with objects. Moreover, a given 2-D point in the image plane may correspond to multiple 3-D points in world space, all of which project onto that same 2-D point. This sometimes makes 3-D hand pose estimation from monocular imagery an ill-posed inverse problem in which depth and the resulting scale ambiguity pose a significant difficulty”. Note that multiple cameras in the 3D environments is mapped to the global camera space). Claims 2-3 and 10-11 are rejected under 35 U.S.C. 103 as being unpatentable over Ge, etc. (US 20200402305 A1) in view of Spurr, etc. (US 20210233273 A1), further in view of Ali, etc. (US 20220277489 A1). Regarding claim 2, Ge and Spurr teach all the features with respect to claim 1 as outlined above. However, Ge, modified by Spurr, fails to explicitly disclose that the computer-implemented method of claim 1, wherein the image is a single RGB image, and wherein the method further comprises: rectifying the image by establishing a canonical camera space and mapping predictions back to an original camera space, wherein the set of 2D keypoints and the set of root-relative 3D vertices are predicted using the rectified image. However, Ali teaches that that the computer-implemented method of claim 1, wherein the image is a single RGB image, and wherein the method further comprises: rectifying the image by establishing a canonical camera space and mapping predictions back to an original camera space, wherein the set of 2D keypoints and the set of root-relative 3D vertices are predicted using the rectified image (See Ali: Fig. 6, and [0122], “At block 604, the process 600 can include determining, based on the image and the metadata, first 3D mesh parameters of a first 3D mesh of the target. The first 3D mesh parameters and the first 3D model can correspond to a first reference frame associated with the image and/or the image capture device. In some examples, the first reference frame can be a coordinate reference frame of the image capture device. In some cases, the first 3D mesh parameters can be determined using a neural network system (e.g., network 210, network 410, neural network 510)”; [0124], “At block 606, the process 600 can include determining, based on the first 3D mesh parameters, second 3D mesh parameters (e.g., mesh parameters 302, real-world frame output 560) for a second 3D mesh of the target. The second 3D mesh parameters and the second 3D mesh can correspond to a second reference frame. In some examples, the second reference frame can include a 3D coordinate system of a real-world scene in which the target is located. In some cases, a neural network system (e.g., network 210, network 410, network 510) can infer a rigid transformation to determine a different reference frame (e.g., the second reference frame). In some examples, a neural network system can infer a rigid transformation between the first reference frame and the second reference frame (e.g., between a camera frame and a real-world frame)”; and [0125], “In some examples, determining the second 3D mesh parameters can include transforming one or more of the first 3D mesh parameters from the first reference frame to the second reference frame. For example, determining the second 3D mesh parameters can include transforming rotation, translation, location, and/or pose parameters from the first reference frame to the second reference frame. As another example, determining the second 3D mesh parameters can include transforming the first 3D mesh from the first reference frame to the second reference frame. In some cases, determining the second 3D mesh parameters can include determining a rotation and translation of the first 3D mesh from the first reference frame to the second reference frame”. Note that the first reference frame associated with the captured image and the camera intrinsic is mapped to the rectifying the image in the canonical camera space, the second reference frame is mapped to the global camera space). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention was effectively filed to modify Ge to have the computer-implemented method of claim 1, wherein the image is a single RGB image, and wherein the method further comprises: rectifying the image by establishing a canonical camera space and mapping predictions back to an original camera space, wherein the set of 2D keypoints and the set of root-relative 3D vertices are predicted using the rectified image as taught by Ali in order to improve accuracy and consistency of 3D modeling of the object in an accurate and consistent manner, and hence reducing computational complexity of a system (See Ali: Fig. 1, and [0047], “As further described herein, the use of non-parametric mesh models can help increase the accuracy and results of meshes generated by the image processing system 100, and the use of parametric mesh models at inference time can increase the modeling efficiency, increase flexibility and scalability, reduce the size of representation of 3D mesh models, reduce latency, reduce power/resource use/requirements at the device (e.g., the image processing system 100), etc. In some examples, the image processing system 100 can use non-parameterized mesh models to learn a better fitting capacity and/or performance, and can learn output parameters to drive the modeling of the mesh. The image processing system 100 can efficiently and accurately use parameterized mesh models at inference time, and can regress model parameters using one or more neural networks”). Ge teaches a method and system that may generate 3D hand mesh by adjusting the skeletal joint positions of a 3D hand mesh using the trained CNN model to process the 2D hand images; while Ali teaches a system and method that may generate the 3D hand mesh model by transforming the 3D hand mesh from the canonical camera space to the global camera space (real world environment). Therefore, it is obvious to one of ordinary skill in the art to modify Ge by Ali to reconstruct the 3D hand mesh more accurately in camera space by transforming the 3D hand mesh constructed in the canonical camera space. The motivation to modify Ge by Ali is “Use of known technique to improve similar devices (methods, or products) in the same way”. Regarding claim 3, Ge, Spurr, and Ali teach all the features with respect to claim 2 as outlined above. Further, Ali teaches that the computer-implemented method of claim 2, wherein rectifying the image comprises: resizing the image with a ratio that converts camera parameters to an original set of camera parameters (See Ali: Fig. 4, and [0091], “The metadata 406 can include the same or different type of information as the metadata 204 in FIG. 2. In some examples, the metadata 406 can include a radial distortion associated with the image capture device (and/or a lens of the image capture device) that captured the image data associated with the cropped image 402, a focal length associated with the image capture device, an optical center associated with the image capture device (and/or a lens of the image capture device), a crop size of the hand 404, a size and/or location of a bounding box containing the hand 404, a scaling ratio of the cropped image 402 (e.g., relative to the hand 404 and/or the uncropped image) and/or the hand 404 (e.g., relative to the cropped image 402 and/or the uncropped image), a distance of a point and/or region of the hand 404 to an optical point and/or region of a lens associated with the image capture device (e.g., a distance of a center of the hand 404 to an optical center of the lens), and/or any other metadata and/or device (e.g., image capture device) calibration information”. Note that the scaling ratio is mapped to the resizing ratio). Regarding claim 10, Ge and Spurr teach all the features with respect to claim 9 as outlined above. Further, Ali teaches that the non-transitory computer-readable medium of claim 9, wherein the image is a single RGB image, and wherein the operations further comprise: rectifying the image by establishing a canonical camera space and mapping predictions back to an original camera space, wherein the set of 2D keypoints and the set of root-relative 3D vertices are predicted using the rectified image (See Ali: Fig. 6, and [0122], “At block 604, the process 600 can include determining, based on the image and the metadata, first 3D mesh parameters of a first 3D mesh of the target. The first 3D mesh parameters and the first 3D model can correspond to a first reference frame associated with the image and/or the image capture device. In some examples, the first reference frame can be a coordinate reference frame of the image capture device. In some cases, the first 3D mesh parameters can be determined using a neural network system (e.g., network 210, network 410, neural network 510)”; [0124], “At block 606, the process 600 can include determining, based on the first 3D mesh parameters, second 3D mesh parameters (e.g., mesh parameters 302, real-world frame output 560) for a second 3D mesh of the target. The second 3D mesh parameters and the second 3D mesh can correspond to a second reference frame. In some examples, the second reference frame can include a 3D coordinate system of a real-world scene in which the target is located. In some cases, a neural network system (e.g., network 210, network 410, network 510) can infer a rigid transformation to determine a different reference frame (e.g., the second reference frame). In some examples, a neural network system can infer a rigid transformation between the first reference frame and the second reference frame (e.g., between a camera frame and a real-world frame)”; and [0125], “In some examples, determining the second 3D mesh parameters can include transforming one or more of the first 3D mesh parameters from the first reference frame to the second reference frame. For example, determining the second 3D mesh parameters can include transforming rotation, translation, location, and/or pose parameters from the first reference frame to the second reference frame. As another example, determining the second 3D mesh parameters can include transforming the first 3D mesh from the first reference frame to the second reference frame. In some cases, determining the second 3D mesh parameters can include determining a rotation and translation of the first 3D mesh from the first reference frame to the second reference frame”. Note that the first reference frame associated with the captured image and the camera intrinsic is mapped to the rectifying the image in the canonical camera space, the second reference frame is mapped to the global camera space). Regarding claim 11, Ge, Spurr, and Ali teach all the features with respect to claim 10 as outlined above. Further, Ali teaches that the non-transitory computer-readable medium of claim 10, wherein rectifying the image comprises: resizing the image with a ratio that converts camera parameters to an original set of camera parameters (See Ali: Fig. 4, and [0091], “The metadata 406 can include the same or different type of information as the metadata 204 in FIG. 2. In some examples, the metadata 406 can include a radial distortion associated with the image capture device (and/or a lens of the image capture device) that captured the image data associated with the cropped image 402, a focal length associated with the image capture device, an optical center associated with the image capture device (and/or a lens of the image capture device), a crop size of the hand 404, a size and/or location of a bounding box containing the hand 404, a scaling ratio of the cropped image 402 (e.g., relative to the hand 404 and/or the uncropped image) and/or the hand 404 (e.g., relative to the cropped image 402 and/or the uncropped image), a distance of a point and/or region of the hand 404 to an optical point and/or region of a lens associated with the image capture device (e.g., a distance of a center of the hand 404 to an optical center of the lens), and/or any other metadata and/or device (e.g., image capture device) calibration information”. Note that the scaling ratio is mapped to the resizing ratio). Claims 7-8, 15-16, and 18-19 are rejected under 35 U.S.C. 103 as being unpatentable over Ge, etc. (US 20200402305 A1) in view of Spurr, etc. (US 20210233273 A1), further in view of Guleryuz (US 20190180473 A1). Regarding claim 7, Ge and Spurr teach all the features with respect to claim 1 as outlined above. However, Ge, modified by Spurr, fails to explicitly disclose that the computer-implemented method of claim 1, further comprising: identifying from the set of 2D keypoints, a subset of 2D keypoints that belong to a wrist of the hand of a user; determining a correspondence between keypoints in a 3D mesh template of a human wrist and keypoints in the subset of 2D keypoints that belong to the wrist of the hand of the user; and predicting a 3D mesh of the wrist of the hand of the user based on the correspondence. However, Guleryuz teaches that the computer-implemented method of claim 1, further comprising: identifying from the set of 2D keypoints, a subset of 2D keypoints that belong to a wrist of the hand of a user (See Guleryuz: Figs. 1-3, and [0025]. “The thumb triangle 305 is defined by vertices at the wrist location 310, the palm knuckle 311 of the index finger, and a palm knuckle 330 of the thumb. The plane that includes the thumb triangle 305 is defined by unit vectors 315, 335, which are represented by the parameters u.sub.I, u.sub.T, respectively. As discussed herein, the distance 320 from the wrist location 310 to the palm knuckle 311 of the index finger is represented by the parameter I. A distance 340 from the wrist location 310 to the palm knuckle 330 of the thumb is represented by the parameter T. Thus, the location of the palm knuckle 330 relative to the wrist location 310 is given by a vector having a direction u.sub.T and a magnitude of T The thumb triangle 305 differs from the palm triangle 300 in that the thumb triangle 305 is compressible and can have zero area. Values of the parameters that define the thumb triangle 305 are learned using 2D images of the hand while the hand is held in the set of training poses”. Note that the wrist is the key identified landmarks and it is a subset of the keypoints in the hand mesh reconstruction); determining a correspondence between keypoints in a 3D mesh template of a human wrist and keypoints in the subset of 2D keypoints that belong to the wrist of the hand of the user (See Guleryuz: Figs. 1-3, and [0024]. “The palm triangle 300 is defined by vertices at the wrist location 310 and palm knuckles 311, 312, 313, 314 (collectively referred to herein as “the palm knuckles 311-314”) of the hand. The plane that includes the palm triangle 300 is defined by unit vectors 315, 316, which are represented by the parameters u.sub.I, u.sub.L, respectively. A distance 320 from the wrist location 310 to the palm knuckle 311 of the index finger is represented by the parameter I and a distance 325 from the wrist location 310 to the palm knuckle 314 of the little finger (or pinky finger) is represented by the parameter L. Thus, the location of the palm knuckle 311 relative to the wrist location 310 is given by a vector having a direction u.sub.I and a magnitude of I. The location of the palm knuckle 314 relative to the wrist location 310 is given by a vector having a direction tu, and a magnitude of L. The location of the palm knuckle 312 of the middle finger is defined as”. Note that the palm template matching starts from the writ 310, and this is mapped to “determining a correspondence between keypoints in a 3D mesh template of a human wrist and keypoints in the subset of 2D keypoints”); and predicting a 3D mesh of the wrist of the hand of the user based on the correspondence (See Guleryuz: Fig. 5, and [0029]. “FIG. 5 illustrates a skeleton model 500 of a finger in a finger pose plane according to some embodiments. The skeleton model 500 represent some embodiments of the skeleton model 415 shown in FIG. 4. The skeleton model 500 also represents portions of some embodiments of the skeleton model 110 shown in FIG. 1 and the skeleton model 210 shown in FIG. 2. The skeleton model 500 includes a palm knuckle 501, a first joint knuckle 502, a second joint knuckle 503, and a fingertip 504. The skeleton model 500 is characterized by a length of a phalanx 510 between the palm knuckle 501 and the first joint knuckle 502, a length of a phalanx 515 between the first joint knuckle 502 and the second joint knuckle 503, and a link of a phalanx 520 between the second joint knuckle 503 and the fingertip 504”. Note that the skeletal 3D hand model reconstruction is mapped to predicting a 3D mesh of the wrist of the hand of the user). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention was effectively filed to modify Ge to have the computer-implemented method of claim 1, further comprising: identifying from the set of 2D keypoints, a subset of 2D keypoints that belong to a wrist of the hand of a user; determining a correspondence between keypoints in a 3D mesh template of a human wrist and keypoints in the subset of 2D keypoints that belong to the wrist of the hand of the user; and predicting a 3D mesh of the wrist of the hand of the user based on the correspondence as taught by Guleryuz in order to enable determining closed curve based on lengths of phalanxes of finger and anatomical constraints on relative angles between phalanxes to detect lengths of phalanxes of fingers and thumb from set of training images of hand (See Guleryuz: Fig. 6, and [0032], “FIG. 6 is a representation of an LUT 600 that is used to look up 2D coordinates of a finger in a finger pose plane based on a relative location of a palm knuckle and a tip of the finger according to some embodiments. The vertical axis of the LUT 600 represents a displacement of the tip of the finger relative to the palm knuckle in the vertical direction. The horizontal axis of the LUT represents a displacement of the tip of the finger relative to the palm knuckle in the horizontal direction. The closed curve 605 represents an outer boundary of the possible locations of the tip of the finger relative to the palm knuckle. The closed curve 605 is therefore determined based on lengths of the phalanxes of the finger and anatomical constraints on relative angles between the phalanxes due to limits on the range of motion of the corresponding joints. Locations within the closed curve 605 represent possible relative positions of the tip of the finger and the palm knuckle”). Ge teaches a method and system that may generate 3D hand mesh by adjusting the skeletal joint positions of a 3D hand mesh using the trained CNN model to process the 2D hand images; while Guleryuz teaches a system and method that explicitly teaches wrist subset identification, 3D wrist template correspondence and wrist mesh prediction for reducing drift and improving correspondence under oclussion. Therefore, it is obvious to one of ordinary skill in the art to modify Ge by Guleryuz to perform wrist subset feature extraction, and 3D wrist mesh reconstruction to improve the accuracy of the3D hand mesh reconstruction. The motivation to modify Ge by Guleryuz is “Use of known technique to improve similar devices (methods, or products) in the same way”. Regarding claim 8, Ge, Spurr, and Guleryuz teach all the features with respect to claim 7 as outlined above. Further, Ge and Guleryuz teach that the computer-implemented method of claim 7, wherein outputting the virtual element in the virtual space based the global camera space mesh prediction of the hand (See Guleryuz: Fig. 1, and [0002], “The location of human body parts, and especially human hands, in three-dimensional (3D) space is a useful driver of numerous applications. Virtual reality (VR), augmented reality (AR), or mixed reality (MR) applications use a representation of a user's hands to facilitate interaction with virtual objects, to select items from a virtual memory, to place objects in the user's virtual hand, to provide a user interface by drawing a menu on one hand and selecting elements of the menu with another hand, and the like. Gesture interactions add an additional modality of interaction with automated household assistants such as Google Home or Nest. Security or monitoring systems use 3D representations of human hands or other body parts to detect and signal anomalous situations. In general, a 3D representation of the location of human hands or other body parts provides an additional modality of interaction or detection that is used instead of or in addition to existing modalities such as voice communication, touchscreens, keyboards, computer mice, and the like. However, computer systems do not always implement 3D imaging devices. For example, devices such as smart phones, tablet computers, and head mounted devices (HMD) typically implement lightweight imaging devices such as two-dimensional (2D) cameras”; and Fig. 9, and [0046], “At block 930, the processor generates a 3D skeleton model that represents the 3D pose of the hand. To generate the 3D skeleton model, the processor rotates the 2D coordinates of the fingers and thumb based on the orientations of the palm triangle and the thumb triangle, respectively. The 3D skeleton model is determined by combining the palm triangle, the orientation of the palm triangle, the thumb triangle, the orientation of the thumb triangle, and the rotated 2D finger coordinates of the fingers and thumb”) comprises: placing the virtual element (See Ge: Fig. 1, and [0025], “The messaging server system 108 supports various services and operations that are provided to the messaging client application 104. Such operations include transmitting data to, receiving data from, and processing data generated by the messaging client application 104. This data may include message content, client device information, geolocation information, media annotation and overlays, virtual objects, message content persistence conditions, social network information, and live event information, as examples. Data exchanges within the messaging system 100 are invoked and controlled through functions available via user interfaces (UIs) of the messaging client application 104”) on the wrist of the hand of the user based on the global camera space mesh prediction of the hand and the 3D mesh of the wrist (See Guleryuz: Figs. 10-11, and [0052], “At block 1110, the processor identifies a first set of 3D keypoints that are compliant with the 3D skeleton model of the hand. For example, the first set of 3D keypoints represents keypoints corresponding to tips of the fingers and thumb, joints of the fingers and thumb, palm knuckles of the fingers and thumb, and a wrist location defined by the 3D skeleton model of the hand. In some embodiments, the first set of 3D keypoints includes the skeleton-compliant keypoint 1020 shown in FIG. 10”. Note that the final 3D mesh output is used to real time virtual object placement and interactions such as the virtual object attached to or controlled by the wrist region of the 3D hand mesh model). Regarding claim 15, Ge and Spurr teach all the features with respect to claim 9 as outlined above. Further, Guleryuz teaches that the non-transitory computer-readable medium of claim 9, wherein the operations further comprise: identifying from the set of 2D keypoints, a subset of 2D keypoints that belong to a wrist of the hand of a user (See Guleryuz: Figs. 1-3, and [0025]. “The thumb triangle 305 is defined by vertices at the wrist location 310, the palm knuckle 311 of the index finger, and a palm knuckle 330 of the thumb. The plane that includes the thumb triangle 305 is defined by unit vectors 315, 335, which are represented by the parameters u.sub.I, u.sub.T, respectively. As discussed herein, the distance 320 from the wrist location 310 to the palm knuckle 311 of the index finger is represented by the parameter I. A distance 340 from the wrist location 310 to the palm knuckle 330 of the thumb is represented by the parameter T. Thus, the location of the palm knuckle 330 relative to the wrist location 310 is given by a vector having a direction u.sub.T and a magnitude of T The thumb triangle 305 differs from the palm triangle 300 in that the thumb triangle 305 is compressible and can have zero area. Values of the parameters that define the thumb triangle 305 are learned using 2D images of the hand while the hand is held in the set of training poses”. Note that the wrist is the key identified landmarks and it is a subset of the keypoints in the hand mesh reconstruction); determining a correspondence between keypoints in a 3D mesh template of a human wrist to keypoints in the subset of 2D keypoints that belong to the wrist of the hand of the user (See Guleryuz: Figs. 1-3, and [0024]. “The palm triangle 300 is defined by vertices at the wrist location 310 and palm knuckles 311, 312, 313, 314 (collectively referred to herein as “the palm knuckles 311-314”) of the hand. The plane that includes the palm triangle 300 is defined by unit vectors 315, 316, which are represented by the parameters u.sub.I, u.sub.L, respectively. A distance 320 from the wrist location 310 to the palm knuckle 311 of the index finger is represented by the parameter I and a distance 325 from the wrist location 310 to the palm knuckle 314 of the little finger (or pinky finger) is represented by the parameter L. Thus, the location of the palm knuckle 311 relative to the wrist location 310 is given by a vector having a direction u.sub.I and a magnitude of I. The location of the palm knuckle 314 relative to the wrist location 310 is given by a vector having a direction tu, and a magnitude of L. The location of the palm knuckle 312 of the middle finger is defined as”. Note that the palm template matching starts from the writ 310, and this is mapped to “determining a correspondence between keypoints in a 3D mesh template of a human wrist and keypoints in the subset of 2D keypoints”); and predicting a 3D mesh of the wrist of the hand of the user based on the correspondence (See Guleryuz: Fig. 5, and [0029]. “FIG. 5 illustrates a skeleton model 500 of a finger in a finger pose plane according to some embodiments. The skeleton model 500 represent some embodiments of the skeleton model 415 shown in FIG. 4. The skeleton model 500 also represents portions of some embodiments of the skeleton model 110 shown in FIG. 1 and the skeleton model 210 shown in FIG. 2. The skeleton model 500 includes a palm knuckle 501, a first joint knuckle 502, a second joint knuckle 503, and a fingertip 504. The skeleton model 500 is characterized by a length of a phalanx 510 between the palm knuckle 501 and the first joint knuckle 502, a length of a phalanx 515 between the first joint knuckle 502 and the second joint knuckle 503, and a link of a phalanx 520 between the second joint knuckle 503 and the fingertip 504”. Note that the skeletal 3D hand model reconstruction is mapped to predicting a 3D mesh of the wrist of the hand of the user). Regarding claim 16, Ge, Spurr, and Guleryuz teach all the features with respect to claim 15 as outlined above. Further, Ge and Guleryuz teach that the non-transitory computer-readable medium of claim 15, wherein outputting the virtual element in the virtual space based the global camera space mesh prediction of the hand (See Guleryuz: Fig. 1, and [0002], “The location of human body parts, and especially human hands, in three-dimensional (3D) space is a useful driver of numerous applications. Virtual reality (VR), augmented reality (AR), or mixed reality (MR) applications use a representation of a user's hands to facilitate interaction with virtual objects, to select items from a virtual memory, to place objects in the user's virtual hand, to provide a user interface by drawing a menu on one hand and selecting elements of the menu with another hand, and the like. Gesture interactions add an additional modality of interaction with automated household assistants such as Google Home or Nest. Security or monitoring systems use 3D representations of human hands or other body parts to detect and signal anomalous situations. In general, a 3D representation of the location of human hands or other body parts provides an additional modality of interaction or detection that is used instead of or in addition to existing modalities such as voice communication, touchscreens, keyboards, computer mice, and the like. However, computer systems do not always implement 3D imaging devices. For example, devices such as smart phones, tablet computers, and head mounted devices (HMD) typically implement lightweight imaging devices such as two-dimensional (2D) cameras”; and Fig. 9, and [0046], “At block 930, the processor generates a 3D skeleton model that represents the 3D pose of the hand. To generate the 3D skeleton model, the processor rotates the 2D coordinates of the fingers and thumb based on the orientations of the palm triangle and the thumb triangle, respectively. The 3D skeleton model is determined by combining the palm triangle, the orientation of the palm triangle, the thumb triangle, the orientation of the thumb triangle, and the rotated 2D finger coordinates of the fingers and thumb”) comprises: placing the virtual element (See Ge: Fig. 1, and [0025], “The messaging server system 108 supports various services and operations that are provided to the messaging client application 104. Such operations include transmitting data to, receiving data from, and processing data generated by the messaging client application 104. This data may include message content, client device information, geolocation information, media annotation and overlays, virtual objects, message content persistence conditions, social network information, and live event information, as examples. Data exchanges within the messaging system 100 are invoked and controlled through functions available via user interfaces (UIs) of the messaging client application 104”) on the wrist of the hand of the user based on the global camera space mesh prediction of the hand and the 3D mesh of the wrist (See Guleryuz: Figs. 10-11, and [0052], “At block 1110, the processor identifies a first set of 3D keypoints that are compliant with the 3D skeleton model of the hand. For example, the first set of 3D keypoints represents keypoints corresponding to tips of the fingers and thumb, joints of the fingers and thumb, palm knuckles of the fingers and thumb, and a wrist location defined by the 3D skeleton model of the hand. In some embodiments, the first set of 3D keypoints includes the skeleton-compliant keypoint 1020 shown in FIG. 10”. Note that the final 3D mesh output is used to real time virtual object placement and interactions such as the virtual object attached to or controlled by the wrist region of the 3D hand mesh model). Regarding claim 18, Ge and Spurr teach all the features with respect to claim 17 as outlined above. Further, Guleryuz teaches that the client device of claim 17, wherein the operations further comprise: identifying from the set of 2D keypoints, a subset of 2D keypoints that belong to a wrist of the hand of a user (See Guleryuz: Figs. 1-3, and [0025]. “The thumb triangle 305 is defined by vertices at the wrist location 310, the palm knuckle 311 of the index finger, and a palm knuckle 330 of the thumb. The plane that includes the thumb triangle 305 is defined by unit vectors 315, 335, which are represented by the parameters u.sub.I, u.sub.T, respectively. As discussed herein, the distance 320 from the wrist location 310 to the palm knuckle 311 of the index finger is represented by the parameter I. A distance 340 from the wrist location 310 to the palm knuckle 330 of the thumb is represented by the parameter T. Thus, the location of the palm knuckle 330 relative to the wrist location 310 is given by a vector having a direction u.sub.T and a magnitude of T The thumb triangle 305 differs from the palm triangle 300 in that the thumb triangle 305 is compressible and can have zero area. Values of the parameters that define the thumb triangle 305 are learned using 2D images of the hand while the hand is held in the set of training poses”. Note that the wrist is the key identified landmarks and it is a subset of the keypoints in the hand mesh reconstruction); determining a correspondence between keypoints in a 3D mesh template of a human wrist to keypoints in the subset of 2D keypoints that belong to the wrist of the hand of the user (See Guleryuz: Figs. 1-3, and [0024]. “The palm triangle 300 is defined by vertices at the wrist location 310 and palm knuckles 311, 312, 313, 314 (collectively referred to herein as “the palm knuckles 311-314”) of the hand. The plane that includes the palm triangle 300 is defined by unit vectors 315, 316, which are represented by the parameters u.sub.I, u.sub.L, respectively. A distance 320 from the wrist location 310 to the palm knuckle 311 of the index finger is represented by the parameter I and a distance 325 from the wrist location 310 to the palm knuckle 314 of the little finger (or pinky finger) is represented by the parameter L. Thus, the location of the palm knuckle 311 relative to the wrist location 310 is given by a vector having a direction u.sub.I and a magnitude of I. The location of the palm knuckle 314 relative to the wrist location 310 is given by a vector having a direction tu, and a magnitude of L. The location of the palm knuckle 312 of the middle finger is defined as”. Note that the palm template matching starts from the writ 310, and this is mapped to “determining a correspondence between keypoints in a 3D mesh template of a human wrist and keypoints in the subset of 2D keypoints”); and predicting a 3D mesh of the wrist of the hand of the user based on the correspondence (See Guleryuz: Fig. 5, and [0029]. “FIG. 5 illustrates a skeleton model 500 of a finger in a finger pose plane according to some embodiments. The skeleton model 500 represent some embodiments of the skeleton model 415 shown in FIG. 4. The skeleton model 500 also represents portions of some embodiments of the skeleton model 110 shown in FIG. 1 and the skeleton model 210 shown in FIG. 2. The skeleton model 500 includes a palm knuckle 501, a first joint knuckle 502, a second joint knuckle 503, and a fingertip 504. The skeleton model 500 is characterized by a length of a phalanx 510 between the palm knuckle 501 and the first joint knuckle 502, a length of a phalanx 515 between the first joint knuckle 502 and the second joint knuckle 503, and a link of a phalanx 520 between the second joint knuckle 503 and the fingertip 504”. Note that the skeletal 3D hand model reconstruction is mapped to predicting a 3D mesh of the wrist of the hand of the user). Regarding claim 19, Ge, Spurr, and Guleryuz teach all the features with respect to claim 18 as outlined above. Further, Ge and Guleryuz teach that the client device of claim 18, wherein outputting the virtual element in the virtual space based the global camera space mesh prediction of the hand (See Guleryuz: Fig. 1, and [0002], “The location of human body parts, and especially human hands, in three-dimensional (3D) space is a useful driver of numerous applications. Virtual reality (VR), augmented reality (AR), or mixed reality (MR) applications use a representation of a user's hands to facilitate interaction with virtual objects, to select items from a virtual memory, to place objects in the user's virtual hand, to provide a user interface by drawing a menu on one hand and selecting elements of the menu with another hand, and the like. Gesture interactions add an additional modality of interaction with automated household assistants such as Google Home or Nest. Security or monitoring systems use 3D representations of human hands or other body parts to detect and signal anomalous situations. In general, a 3D representation of the location of human hands or other body parts provides an additional modality of interaction or detection that is used instead of or in addition to existing modalities such as voice communication, touchscreens, keyboards, computer mice, and the like. However, computer systems do not always implement 3D imaging devices. For example, devices such as smart phones, tablet computers, and head mounted devices (HMD) typically implement lightweight imaging devices such as two-dimensional (2D) cameras”; and Fig. 9, and [0046], “At block 930, the processor generates a 3D skeleton model that represents the 3D pose of the hand. To generate the 3D skeleton model, the processor rotates the 2D coordinates of the fingers and thumb based on the orientations of the palm triangle and the thumb triangle, respectively. The 3D skeleton model is determined by combining the palm triangle, the orientation of the palm triangle, the thumb triangle, the orientation of the thumb triangle, and the rotated 2D finger coordinates of the fingers and thumb”) comprises: placing the virtual element (See Ge: Fig. 1, and [0025], “The messaging server system 108 supports various services and operations that are provided to the messaging client application 104. Such operations include transmitting data to, receiving data from, and processing data generated by the messaging client application 104. This data may include message content, client device information, geolocation information, media annotation and overlays, virtual objects, message content persistence conditions, social network information, and live event information, as examples. Data exchanges within the messaging system 100 are invoked and controlled through functions available via user interfaces (UIs) of the messaging client application 104”) on the wrist of the hand of the user based on the global camera space mesh prediction of the hand and based on the 3D mesh of the wrist (See Guleryuz: Figs. 10-11, and [0052], “At block 1110, the processor identifies a first set of 3D keypoints that are compliant with the 3D skeleton model of the hand. For example, the first set of 3D keypoints represents keypoints corresponding to tips of the fingers and thumb, joints of the fingers and thumb, palm knuckles of the fingers and thumb, and a wrist location defined by the 3D skeleton model of the hand. In some embodiments, the first set of 3D keypoints includes the skeleton-compliant keypoint 1020 shown in FIG. 10”. Note that the final 3D mesh output is used to real time virtual object placement and interactions such as the virtual object attached to or controlled by the wrist region of the 3D hand mesh model). Allowable Subject Matter Claims 4-6, 12-14, and 20 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The best arts searched, Ge, etc. (US 20200402305 A1), Spurr, etc. (US 20210233273 A1), and Ali, etc. (US 20220277489 A1), do not teach the cited limitations of “the computer-implemented method of claim 2, further comprising: predicting a set of weights based on the rectified image, wherein each weight of the set of weights represents a confidence in a prediction of a corresponding keypoint of the set of 2D keypoints, the set of weights accounting for keypoint landmark correspondence inaccuracy due to occlusion of the hand in the image.” Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to GORDON G LIU whose telephone number is (571)270-0382. The examiner can normally be reached Monday - Friday 8:00-5:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Devona E Faulk can be reached at 571-272-7515. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /GORDON G LIU/ Primary Examiner, Art Unit 2618
Read full office action

Prosecution Timeline

Feb 12, 2025
Application Filed
Jul 28, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705808
AI-BASED VISUAL CONTENT COLLAGE GENERATION
2y 5m to grant Granted Aug 11, 2026
Patent 12705751
MAP SCENE RENDERING METHOD AND APPARATUS, SERVER, TERMINAL, COMPUTER-READABLE STORAGE MEDIUM, AND COMPUTER PROGRAM PRODUCT
2y 4m to grant Granted Aug 11, 2026
Patent 12705804
AUGMENTED REALITY BASED AIMING OF LIGHT FIXTURES
2y 4m to grant Granted Aug 11, 2026
Patent 12700152
AI-BASED AVATAR CREATION USING STYLE TRANSFER AND SUBJECT IMAGES
2y 5m to grant Granted Aug 04, 2026
Patent 12694614
IMAGING DEVICE, AND CONTROL METHOD OF IMAGING DEVICE
2y 1m to grant Granted Jul 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
83%
Grant Probability
98%
With Interview (+15.0%)
2y 2m (~8m remaining)
Median Time to Grant
Low
PTA Risk
Based on 692 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month