Prosecution Insights
Last updated: August 14, 2026
Application No. 19/034,224

THREE-DIMENSIONAL RECONSTRUCTION METHOD AND APPARATUS, PRODUCT INFORMATION PROCESSING METHOD AND APPARATUS, AND DEVICE AND STORAGE MEDIUM

Non-Final OA §103
Filed
Jan 22, 2025
Priority
Oct 14, 2022 — CN 202211257959.4 +1 more
Examiner
LIU, GORDON G
Art Unit
Tech Center
Assignee
Taobao (China) Software Co. Ltd.
OA Round
1 (Non-Final)
83%
Grant Probability
Favorable
1-2
OA Rounds
7m
Est. Remaining
98%
With Interview

Examiner Intelligence

Grants 83% — above average
83%
Career Allowance Rate
574 granted / 692 resolved
+22.9% vs TC avg
Moderate +15% lift
Without
With
+15.0%
Interview Lift
resolved cases with interview
Fast prosecutor
2y 2m
Avg Prosecution
35 currently pending
Career history
717
Total Applications
across all art units

Statute-Specific Performance

§101
7.2%
-32.8% vs TC avg
§103
77.3%
+37.3% vs TC avg
§102
3.5%
-36.5% vs TC avg
§112
2.6%
-37.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 692 resolved cases

Office Action

§103
DETAILED ACTION The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims 1-20 are pending under this Office action. Claim Objections Claim 13 is objected to because of the following informalities: “SMPL” may be “SMPL (Skinned Multi-Person Linear Model)”. Appropriate correction is required. Claim 17 is objected to because of the following informalities: “The method according to claim 14” may be “The method according to claim 16”. Appropriate correction is required. Claim Interpretation— 35 USC 112 (f) The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. Claim Limitation Interpreted under 35 USC 112(f) Use of the word "means" (or "step for") in a claim with functional language creates a rebuttable presumption that the claim element is to be treated in accordance with 35 U.S.C. 112(f) (pre-AIA 35 U.S.C. 112, sixth paragraph). The presumption that 35 U.S.C. 112(f) (pre-AIA 35 U.S.C. 112, sixth paragraph) is invoked is rebutted when the function is recited with sufficient structure, material, or acts within the claim itself to entirely perform the recited function. Absence of the word "means" (or "step for") in a claim creates a rebuttable presumption that the claim element is not to be treated in accordance with 35 U.S.C. 112(f) (pre-AIA 35 U.S.C. 112, sixth paragraph). The presumption that 35 U.S.C. 112(f) (pre-AIA 35 U.S.C. 112, sixth paragraph) is not invoked is rebutted when the claim element recites function but fails to recite sufficiently definite structure, material or acts to perform that function. Claim elements in this application that use the word "means" (or "step for") are presumed to invoke 35 U.S.C. 112(f) except as otherwise indicated in an Office action. Similarly, claim elements that do not use the word "means" (or "step for") are presumed not to invoke 35 U.S.C. 112(f) except as otherwise indicated in an Office action. Claim limitations " an image acquisition unit, configured to ", " a feature extraction unit, configured to ", “a vector concatenation unit, configured to”, “a parameter regression unit, configured to”, and " a masking processing unit, configured to" have been interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because it uses/they use a generic placeholder "an image acquisition unit /a feature extraction unit/a vector concatenation unit/ a parameter regression unit / a masking processing unit " coupled with functional language "configured to acquire / input / concatenate/ input / apply " without reciting sufficient structure to achieve the function. Furthermore, the generic placeholder is not preceded by a structural modifier. Since the claim limitations invokes 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, claim 20 has been interpreted to cover the corresponding structure described in the specification that achieves the claimed function, and equivalents thereof. A review of the specification shows that the following appears to be the corresponding structure described in the specification for the 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph limitation: " an image acquisition unit, configured to acquire " is interpreted as to " image acquisition unit 81 is configured to acquire a plurality of frames of images of the target object, as well as the three-dimensional model description information corresponding to the target object " (See Specification: Fig. 8, and [0115], " image acquisition unit 81 is configured to acquire a plurality of frames of images of the target object, as well as the three-dimensional model description information corresponding to the target object"). " a feature extraction unit, configured to input " is interpreted as to " feature extraction unit 82 is configured to input a plurality of frames of images into a feature extraction network to extract features, thereby obtaining the feature vectors of the plurality of frames of image" (See Specification: Fig. 8, and [0116], " feature extraction unit 82 is configured to input a plurality of frames of images into a feature extraction network to extract features, thereby obtaining the feature vectors of the plurality of frames of image"). " a vector concatenation unit, configured to concatenate " is interpreted as to " vector concatenation unit 83 is configured to concatenate the feature vectors of the plurality of frames of images to obtain the target concatenated feature vector" (See Specification: Fig. 8, and [0117], " vector concatenation unit 83 is configured to concatenate the feature vectors of the plurality of frames of images to obtain the target concatenated feature vector"). " a parameter regression unit, configured to input" is interpreted as to " parameter regression unit 84 is configured to input the target concatenated feature vector into a parameter regression network to predict a plurality of control parameters for model control based on the number of parameters. The plurality of control parameters include pose control parameters and shape control parameters" (See Specification: Fig. 8, and [0118], parameter regression unit 84 is configured to input the target concatenated feature vector into a parameter regression network to predict a plurality of control parameters for model control based on the number of parameters. The plurality of control parameters include pose control parameters and shape control parameters"). " a masking processing unit, configured to apply " is interpreted as to " masking processing unit 85 is configured to apply a masking operation to the initial three-dimensional model of the target object based on the pose control parameters and shape control parameters, thereby obtaining the target three-dimensional model of the target object" (See Specification: Fig. 8, and [0119], " masking processing unit 85 is configured to apply a masking operation to the initial three-dimensional model of the target object based on the pose control parameters and shape control parameters, thereby obtaining the target three-dimensional model of the target object"). If applicant wishes to provide further explanation or dispute the examiner's interpretation of the corresponding structure, applicant must identify the corresponding structure with reference to the specification by page and line number, and to the drawing, if any, by reference characters in response to this Office action. If applicant does not intend to have the claim limitations treated under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may amend the claim(s) so that it/they will clearly not invoke 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, or present a sufficient showing that the claim recites/recite sufficient structure, material, or acts for performing the claimed function to preclude application of 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. For more information, see MPEP § 2173 et seq. and Supplementary Examination Guidelines for Determining Compliance With 35 U.S.C. 112 and for Treatment of Related Issues in Patent Applications, 76 FR 7162, 7167 (Feb. 9, 2011). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-4, 10, 12-15, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Guler, etc. (US 20210241522 A1) in view of Faras, etc. (US 20210118160 A1), further in view of Gonzalez, etc. (US 20180211438 A1). Regarding claim 1, Guler teaches that a method for three-dimensional reconstruction (See Guler: Fig. 7, and [0101], “FIG. 7 shows an example of a computational unit/apparatus. The apparatus 700 comprises a processing apparatus 702 and memory 704 according to an exemplary embodiment. Computer-readable code 706 may be stored on the memory 704 and may, when executed by the processing apparatus 702, cause the apparatus 700 to perform methods as described herein”), comprising: obtaining a plurality of frames of images containing a target object, and three-dimensional model description information of the target object (See Guler: Figs. 1-3, and [0034], “FIG. 1 shows a schematic overview of a method 100 of part-based three-dimensional reconstruction 102 from a two-dimensional image 104. A keypoint detector indicates 2D keypoint (e.g. joint) positions, driving feature pooling. These in turn are used to determine 3D model parameters of the parametric model (e.g. 3D SMPL joint angles). The method may be implemented on one or more computing devices. The method extracts localized features around key points (in the example shown, human joints) following a part-based modelling paradigm. The resulting model may be referred to as a fused three dimensional representation”; [0035], “The method uses a tight prior model of the object shape (such as a deformable parametric model for human body reconstruction) in tandem with a multitude of pose estimation methods to derive accurate monocular 3D reconstruction. In embodiments relating to 3D human reconstruction, the articulated nature of the human body is taken into account, which can substantially improve performance over a monolithic baseline. A refinement procedure may also be used that allows the shape prediction results of a single-shot system to be adapted so as to meet geometric constraints imposed by complementary, fully-convolutional networks/neural network branches”; and [0049], “At operation 2.1, a two-dimensional image is received. As described above, the 2D image may comprise a set of pixel values corresponding to a two-dimensional array. In some embodiments, a single image is received. Alternatively, a plurality of images may be received. If multiple images are taken (for example, of a person moving), the three dimensional representation may account for these changes”. Note that the human body is mapped to the target object); inputting the plurality of frames of images into a feature extraction network to extract features, to obtain feature vectors of the plurality of frames of images (See Guler: Figs. 1-3, and [0037], “A two dimensional image 104 is received, and input into a neural network 106. The two-dimensional image, I, comprises a set of pixel values corresponding to a two-dimensional array. For example, in a colour image. I∈custom-character.sup.H×W×3, where H is the height of the image in pixels, W is the height of the image in pixels and the image has three colour channels (e.g. RGB or CIELAB). The two dimensional image may, in some embodiments, be in black-and-white/greyscale”; [0044], “Each feature extracted around a 2D keypoint/joint can deliver a separate parametric deformable model parameter estimate. But intuitively a 2D keypoint/joint should have a stronger influence on parametric deformable model parameters that are more relevant to it—for instance a left wrist joint should be affecting the left arm parameters, but not those of kinetically independent parts such as the right arm, head, or feet. Furthermore, the fact that some keypoints/joints can be missing from an image means that we cannot simply concatenate the features in a larger feature vector, but need to use a form that handles gracefully the potential absence of parts”; and [0036], “The 3D reconstruction method relies on a prior model of the target shape. The prior model may be a parametric model of the object being reconstructed. The type of prior model used may depend on the object being reconstructed. For example, in embodiments where the shape is a human body, the human body may be parametrised using a parametric deformable model. In the following, the method will be described with reference to a parametric deformable model, though it will be appreciated that other parametric shape models may alternatively be used. In some embodiments, SMPL is used. SMPL uses two parameter vectors to capture pose and shape: θ comprises 23 3D-rotation matrices corresponding to each joint in a kinematic tree for the human pose, and β captures shape variability across subjects in terms of a 10-dimensional shape vector. SMPL uses these parameters to obtain a triangulated mesh of the human body through linear skinning and blend shapes as a differentiable function of θ, β. In some embodiments, the parametric deformable model may be modified to allow it to be used more efficiently with convolutional neural networks, as described below in relation to FIG. 3”), and concatenating the feature vectors of the plurality of frames of images to generate a target concatenated feature vector (See Guler: Figs. 1-3, and [0047], “An arg-soft-max operation may be used over angle clusters, and information from multiple 2D joints is fused: w.sup.i.sub.k,j indicates a score that 2D joint j assigns to cluster k for the i-th model parameter, θ.sup.i. The clusters can be defined as described below in relation to the “mixture-of-experts prior”. The neighbourhood of i is constructed by inspecting which model parameters directly influence human 2D joints, based on kinematic tree dependencies. Joints are found in the image by taking the maximum of a 2D joint detection module. If the maximum is below a threshold (for example, 0:4 in some embodiments) it is considered consider that a joint is not observed in the image. In that case, every summand corresponding to a missing joint is excluded from the summations in the numerator and denominator, so that the above equation still delivers a distribution over angles. If all elements of N(i) are missing, θ.sup.i is set to a resting pose”; [0056], “At operation 2.6, the respective three-dimensional representations are combined to generate a fused three-dimensional representation of the object”; and [0058], The method may further comprise applying the fused three-dimensional representation of the object to a part-based three-dimensional shape representation model. An example of such a model is a kinematic tree”. Note that combining and fusion is mapped to concatenating); inputting the target concatenated feature vector into a parameter regression network, and predicting, based on the three-dimensional model description information, a plurality of control parameters for model control (See Guler: Figs. 1-5, and [0036], “The 3D reconstruction method relies on a prior model of the target shape. The prior model may be a parametric model of the object being reconstructed. The type of prior model used may depend on the object being reconstructed. For example, in embodiments where the shape is a human body, the human body may be parametrised using a parametric deformable model. In the following, the method will be described with reference to a parametric deformable model, though it will be appreciated that other parametric shape models may alternatively be used. In some embodiments, SMPL is used. SMPL uses two parameter vectors to capture pose and shape: θ comprises 23 3D-rotation matrices corresponding to each joint in a kinematic tree for the human pose, and β captures shape variability across subjects in terms of a 10-dimensional shape vector. SMPL uses these parameters to obtain a triangulated mesh of the human body through linear skinning and blend shapes as a differentiable function of θ, β. In some embodiments, the parametric deformable model may be modified to allow it to be used more efficiently with convolutional neural networks, as described below in relation to FIG. 3”; and [0093], “At operation 5.5, parameters of the fused three-dimensional representation are adjusted based on the error value. The adjustment may be performed using learning based optimisation methods”), wherein the plurality of control parameters include pose control parameters and shape control parameters; and applying a masking operation to an initial three-dimensional model of the target object based on the pose control parameters and the shape control parameters to generate a target three-dimensional model of the target object, wherein the initial three-dimensional model is obtained based on the three-dimensional model description information (See Guler: Figs. 4-5, and [0069], “Additional sets of data, such as DensePose and 3D joint estimation, can be used to increase the accuracy of 3D reconstruction. This may be done both by introducing additional losses that can regularize the training in a standard multi-task learning setup and by providing bottom-up, CNN-based cues to drive the top-down, model-based 3D reconstruction. For this, rather than relying on a CNN-based system to deliver accurate results in a single shot at test time, predictions from the CNN are used as an initialization to an iterative fitting procedure that uses the pose estimation results delivered by a 2D decoder to drive the model parameter updates. This can result in an effective 3D reconstruction. For this, the parametric model parameters are updated so as to align the parametric model-based and CNN-based pose estimates, as captured by a combination of a Dense Pose-based loss and the distances between the parametric model-based and CNN-based estimates of the 3D joints. This allows the parametric model parameters to be updated on-the-fly, so as to better match the CNN-based localization results”; [0073], “The method 400 proceeds in a similar way to the method described above with respect to FIGS. 1-3. A two dimensional image 404 is received, and input into a neural network 406. The neural network 406 processes the input image 404 to generate an initial 3D model 408 (i.e. a fused three-dimensional model/3D surface estimate) of objects in the image (e.g. a human body) using part-based 3D shape reconstruction 410 (for example, as described above). The neural network further outputs 2D keypoint locations (e.g. 2D joints), 3D keypoint locations (e.g. 3D joints) and a DensePose estimate. Each of these may be output from a different head 406b-e of the neural network 406. The neural network 406 may comprise one or more convolutional and/or deconvolutional layers/networks. The initial layers of the neural network 406 may be shared and form a backbone network 406a. In the example shown, the backbone network 406a is a ResNet-50 network, though other networks may alternatively be used”; and [0074], “The initial 3D model 408 may be refined using a 3D shape refinement process 412. The 3D shape refinement process 412 takes as input a 3D model, 2D keypoint locations, 3D keypoint locations and the DensePose estimate for the corresponding 2D image, and uses them to refine the 3D model to generate a refined 3D model of the object 402”. Note that the pose and shape parameters are inputted to the CNN model, and the output of the CNN model is also fed back to the CNN model iteratively to improve the accuracy of the CNN reconstruction results; and the refinement process is fundamentally the masking operation but a new art will be used for clarity). However, Guler fails to explicitly disclose that wherein the plurality of control parameters include pose control parameters and shape control parameters; and applying a masking operation to an initial three-dimensional model of the target object. However, Faras teaches that wherein the plurality of control parameters include pose control parameters and shape control parameters (See Faras: Figs. 1-4, and [0039], “Localization 204 of FIG. 2 is used to determine a 3D map and/or poses, which may be important factors of creating a 3D representation. Some embodiments described herein arise from the recognition that, in image processing operations to create a 3D representation of a subject from images captured by an image capture device, the 3D representation may be degraded if the corresponding pose of the image capture device and/or a related 3D map cannot be accurately determined. Embodiments described herein are thus directed to using improved techniques to combine recursive and non-recursive approaches to creating and/or updating a 3D map and/or determining accurate estimated poses of the image capture device. Recursive techniques relate to or involve the repeated application of a rule, definition, or procedure to successive results. Any of the operations described herein as being recursive may be performed in a causal manner on the poses and operations described as being non-recursive may be performed in an acausal manner on the poses, respectively”; and [0041], “FIGS. 4A and 4B illustrate localization of 3D poses, corresponding to block 204 of FIG. 2. Referring now to FIG. 4A, several 2D images of an object such as the face of a person have been collected during a portion of a scan. The poses 1 to 19 are estimated at various camera viewpoints of the 2D images. A 3D map 410 of the object including various 3D points 420 is constructed. Referring now to FIG. 4B, the scan is continued by capturing additional 2D images for which poses 20 to 96 are estimated. The 3D map 410 includes additional 3D points 420 that have been triangulated from the additional 2D images”. Note that the pose and shape of the human body is recursively updated, and this is explicitly mapped to the regression network for pose and shape parameters). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention was effectively filed to modify Guler to have wherein the plurality of control parameters include pose control parameters and shape control parameters as taught by Faras in order to minimize prediction error (See Faras: Fig. 9, and [0049], “FIG. 9 illustrates creation of a camera model. Referring to FIG. 9, a feature point or landmark in a 2D image 130 may have a measurement û 942 that corresponds to a pixel coordinate. The 2D point u 940 is the predicted image coordinate of the 3D point. A statistical goal is to reduce or minimize the prediction error, which is the distance between û 942 and u 940. (R, z) 930 represents the pose of the camera where R is the angular orientation and z is the camera position. z may be a vector with three components. A projection of x 910 depends on the coordinates and the pose. x920 is a 3D point in the 3D map. In other words, a collection of x 3D points forms the 3D map”). Guler teaches a method and system that may reconstruct three-dimensional models of objects from two-dimensional images by extracting features from the 2D input images, via neural networks, estimating the 3D representation of the objects from the extracted features, concatenating the estimated parameters, and generated the 3D model of the objects based on the fused 3D representation of the object; while Faras teaches a system and method that may reconstruct the 3D human body model with pose and shape parameters recursively estimated and updated to improve the accuracy of the 3D model reconstruction. Therefore, it is obvious to one of ordinary skill in the art to modify Guler by Faras to perform recursive (regression) updating the pose and shape parameters in 3D object model reconstruction. The motivation to modify Guler by Faras is “Use of known technique to improve similar devices (methods, or products) in the same way”. However, Guler, modified by Faras, fails to explicitly disclose that applying a masking operation to an initial three-dimensional model of the target object. However, Gonzalez teaches that applying a masking operation to an initial three-dimensional model of the target object (See Gonzalez: Figs. 1-3, and [0026], “The segment extractor 114 can receive the validated and scaled ROIs 112 and generate one or more binary segments 116, textured segments 118, and segment characterizations 120 based on the validated and scaled ROIs 112. For example, the segment extractor 114 can extract a binary segment representing the shape of an object from a raw image 108. For example, the binary segment may be a two-dimensional shape corresponding to the object's 2D projection on an image plane. The binary segment may have attributes such as a center of mass, bounding box, total pixels, principal axes, and a binary mask. In some examples, the binary mask may be applied to a source color image to generate a texture image as described below. In some examples, the segment extractor 114 may also include adaptive filters and graph-based segmentation that can be used to prevent and solve problems caused by image compression artifacts and unpropitious lighting and material combinations. The segment extractor 114 can thus perform fast and fully automatic removal of compression artifacts. In some examples, the segment extractor 114 can also perform speckle removal by adaptive rejection of pixels using their 8-connectivity degree in a recursive fashion. In some examples, the speckle removal can be further refined by smoothing expansion and contraction. The processing performed by the segment extractor 114 may generally be referred to herein as a signal phase. Because the signal phase may be the foundation for all subsequent phases, the segment extractor 114 may thus include high quality smoothing and acutance-preserving contouring. The operation of the segment extractor 114 is described in detail with respect to the example segment extractor of FIG. 2 below. The final output of the segment extractor may be a binary segment 116, a textured segment 118, and a binary segment with segment characterization 120”; and [0046], “In FIG. 3, an ROI 112 is received. For example, the ROI may be a region corresponding to a light bulb identified in a 2D color image. The ROI 112 may then be desaturated by a desaturator to produce a grayscale image, or desaturated ROI 304. An intensity enhancer may then improve image intensity of the desaturated ROI 304 to generate an intensity enhanced ROI 306. An adaptive binarization may then be applied to the intensity enhanced ROI 306 to generate a binary image 308. For example, the adaptive binarization may remove gradients and result in an image of black and white pixels. A segmentation can then be applied to the binary image 308 by a segmenter to extract a binary segment 310. Hole removal may then be performed by a hole remover on the extract binary segments 310 to generate a binary mask image 312. Speckle removal may then be performed by a speckle remover on the binary mask image 312 to generate a speckle-less binary mask image 314. For example, the speckle remover may remove spurious contour artifacts that may have resulted from image compression. An expander may then apply diffusion on the speckle-less binary mask image 314 to generate a diffused binary mask image 316. A contraction may then apply a non-adaptive binarization on the diffused binary mask image 316 to generate a binary image 318. The binary image 318 may be used to generate a binary segment 116 and a textured segment 118. For example, the binary segment 116 may be generated as described above with respect to FIG. 1. In some examples, the texture image 118 can be generated by applying the binary mask 216 to the source image to select pixels from the source image masked by the binary mask 216”). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention was effectively filed to modify Guler to have applying a masking operation to an initial three-dimensional model of the target object as taught by Gonzalez in order to enable the efficient generation of 3D meshes of various resolutions (See Gonzalez: Fig. 1, and [0019], “The techniques described herein thus may enable unlimited-resolution 3D models to be generated automatically from single 2D images of free-form revolution objects. The size of the 3D models may be relatively small, and thus a large amount of 3D models may be efficiently stored. In some examples, the size of the 3D models may be several orders of magnitude smaller than 3D models generated by other methods. For example, each 3D model may be defined by sets of equations that can be used to generate 3D meshes with varying resolution and tessellation. In some examples, the 3D model may be used efficiently to create mesh models for a variety of applications, including 3D printing, virtual reality, augmented reality, robotics, computer automated design (CAD) and product, etc. The techniques described herein thus enable efficient generation of 3D meshes of various resolutions. For example, a mesh of any resolution and amount of tessellation can be generated based on the 3D model, a resolution parameter, and a tessellation parameter”). Guler teaches a method and system that may reconstruct three-dimensional models of objects from two-dimensional images by extracting features from the 2D input images, via neural networks, estimating the 3D representation of the objects from the extracted features, concatenating the estimated parameters, and generated the 3D model of the objects based on the fused 3D representation of the object; while Gonzalez teaches a system and method that may generate the 3D model based on the contour and skeleton cue obtained from the input image though binary image masking, segmentation, and texture segmentation. Therefore, it is obvious to one of ordinary skill in the art to modify Guler by Gonzalez to extract features from the input images using masking operation. The motivation to modify Guler by Gonzalez is “Use of known technique to improve similar devices (methods, or products) in the same way”. Regarding claim 2, Guler, Faras, and Gonzalez teach all the features with respect to claim 1 as outlined above. Further, Guler and Faras teach that the method according to claim 1, wherein inputting the plurality of frames of images into the feature extraction network to extract features, to obtain feature vectors of the plurality of frames of images (See Guler: Figs. 1-3, and [0036], “The 3D reconstruction method relies on a prior model of the target shape. The prior model may be a parametric model of the object being reconstructed. The type of prior model used may depend on the object being reconstructed. For example, in embodiments where the shape is a human body, the human body may be parametrised using a parametric deformable model. In the following, the method will be described with reference to a parametric deformable model, though it will be appreciated that other parametric shape models may alternatively be used. In some embodiments, SMPL is used. SMPL uses two parameter vectors to capture pose and shape: θ comprises 23 3D-rotation matrices corresponding to each joint in a kinematic tree for the human pose, and β captures shape variability across subjects in terms of a 10-dimensional shape vector. SMPL uses these parameters to obtain a triangulated mesh of the human body through linear skinning and blend shapes as a differentiable function of θ, β. In some embodiments, the parametric deformable model may be modified to allow it to be used more efficiently with convolutional neural networks, as described below in relation to FIG. 3”), comprises: for each frame of the plurality of frames of images, inputting the frame into a feature extraction module of the feature extraction network to perform feature extraction and obtain an image feature map for the frame (See Guler: Figs. 1-3, and [0049], “At operation 2.1, a two-dimensional image is received. As described above, the 2D image may comprise a set of pixel values corresponding to a two-dimensional array. In some embodiments, a single image is received. Alternatively, a plurality of images may be received. If multiple images are taken (for example, of a person moving), the three dimensional representation may account for these changes”; and [0045], “A linear layer 112 applied on top of the pooled extracted features is used to estimate parameters 114 of the parametric deformable model. The linear layer may, for example, output weights w.sup.i.sub.k,j 116 for use in determining the parametric deformable model parameters 11”); inputting camera pose data recorded at the time the frame was captured into a camera parameter fusion module of the feature extraction network to perform feature extraction and obtain a camera pose feature map for the frame (See Faras: Figs. 1-2, and [0026], “Applications such as 3D imaging, mapping, and navigation may use Simultaneous Localization and Mapping (SLAM). SLAM relates to constructing or updating a map of an unknown environment while simultaneously keeping track of an object's location within it and/or estimating the pose of the camera with respect to the object or scene. This computational problem is complex since the object may be moving and the environment may be changing. 2D images of real objects and/or 3D object may be captured with the objective of creating a 3D image that is used in real-world applications such as augmented reality, 3D printing, and/or 3D visualization with different perspectives of the real objects. The 3D objects may be characterized by features that are specific locations on the physical object in the 2D images that are of importance for the 3D representation such as corners, edges, center points, or object-specific features on a physical object such as a face that may include nose, ears, eyes, mouth, etc. There are several algorithms used for solving this computational problem associated with 3D imaging, using approximations in tractable time for certain environments. Popular approximate solution methods include the particle filter and Extended Kalman Filter (EKF). The particle filter, also known as a Sequential Monte Carlo (SMC) linearizes probabilistic estimates of data points. The Extended Kalman Filter is used in non-linear state estimation in applications including navigation systems such as Global Positioning Systems (GPS), self-driving cars, unmanned aerial vehicles, autonomous underwater vehicles, planetary rovers, newly emerging domestic robots, medical devices inside the human body, and/or image processing systems. Image processing systems may perform 3D pose estimation using SLAM techniques by performing a transformation of an object in a 2D image to produce a 3D object. However, existing techniques such as SMC and EKF may be insufficient in accurately estimating and positioning various points in a 3D object based on information discerned from 2D objects and may be computationally inefficient in real time”); concatenating the image feature map and the camera pose feature map of each frame using a feature concatenation module of the feature extraction network to obtain a concatenated feature map for each frame (See Guler: Figs. 1-3, and [0044], “Each feature extracted around a 2D keypoint/joint can deliver a separate parametric deformable model parameter estimate. But intuitively a 2D keypoint/joint should have a stronger influence on parametric deformable model parameters that are more relevant to it—for instance a left wrist joint should be affecting the left arm parameters, but not those of kinetically independent parts such as the right arm, head, or feet. Furthermore, the fact that some keypoints/joints can be missing from an image means that we cannot simply concatenate the features in a larger feature vector, but need to use a form that handles gracefully the potential absence of parts”); and performing dimensionality reduction on the concatenated feature map of each frame using a dimensionality reduction module of the feature extraction network to obtain a feature vector for each frame (See Faras: Figs. 1-3, and [0040], “More particularly, a robust and accurate method that can deliver real-time pose estimates and/or a 3D map for 3D reconstruction and provide enough information for camera calibration is described in various embodiments. The inventive concepts described herein combine a non-recursive initialization phase with a recursive sequential updating (tracking phase) system. Initialization of the 3D map or structure may be based on the scene or the scene structure that is discerned from a series of 2D images or frames. Sequential tracking or sequential updating may also be referred to as recursive pose and positioning. During the initialization phase, a non-recursive initialization of the 3D map and the poses is used to localize the camera for 2D frames. An initial map of the scene, which is represented by a set of 3D coordinates corresponding to salient image points that are tracked between sequential frames, is constructed and the camera poses (orientation and position of the camera along its trajectory) are computed. Criteria, such as, for example, the number of tracked points or the pose change, are used to decide if the current frame should become a key-frame. Key frames are selected as representative sets of frames to be used in the localization. If a given frame is selected as a key frame, a local/global bundle adjustment (BA) may be used to refine the key-frames positions and/or to refine or triangulate new 3D points. During this processing a global feature database may be created and populated with globally optimized landmarks. Each landmark may be associated with some stored information such as the related 3D coordinates, a list of frames/key-frames where it was visible, and/or a reference patch. After the initialization phase, a set of anchor landmarks may be available when the sequential updating and/or tracking phase is entered. A fully recursive system, also based on feature tracking, may be used to localize the camera. In particular, the initial set of global features may reduce and/or remove the known drift problem of localization with recursive systems”). Regarding claim 3, Guler, Faras, and Gonzalez teach all the features with respect to claim 2 as outlined above. Further, Guler and Gonzalez teach that the method according to claim 2, wherein for each frame of the plurality of frames of images, inputting the frame into the feature extraction module of the feature extraction network to perform feature extraction and obtain an image feature map for the frame, comprises: for each frame of the plurality of frames of images, inputting the frame into a skip connection layer of the feature extraction module to perform multi-resolution feature map extractions and skip connections of feature maps with the same resolution, to obtain a second intermediate feature map for the frame (See Gonzalez: Fig. 1, and [0025], “The receiver, validator, and scaler 110 can perform preprocessing on the received image 104, 108. For example, the receiver, validator, and scaler 110 can perform detection of image compression, content validation, and spatial scaling. In some examples, the receiver, validator, and scaler 110 can detect a region of interest (ROI) 112 in a received image and a compression ratio to automatically adjust parameter boundaries for subsequent adaptive algorithms. The region of interest 112 may be a rectangular region of an image representing an object to be modeled. In some examples, the object to be modeled can be determined based on pixel color. For example, the region of interest may be a rectangular region including pixels that have a different color than a background color. In some examples, at least one of the pixels of the object to be modeled may touch each edge of the edge of the region of interest. In some examples, the receiver, validator, and scaler 110 can validate the region of interest. For example, the receiver, validator, and scaler 110 can compare a detected region of interest with the corresponding image including the region of interest. In some examples, if the region of interest is the same as the corresponding image, then the image may be rejected. In some examples, if the region of interest is not the same as the corresponding image, then the region of interest may be further processed as described below. The receiver, validator, and scaler 110 can then add a spatial-convolutional margin to the ROI 112. For example, the spatial-convolutional margin can be a margin of pixels that can be used during filtering of the region of interest. For example, the spatial-convolutional margin may be a margin of 5-7 pixels that can be used by the filter to prevent damage to content in the region of interest. The receiver, validator, and scaler 110 may then scale the ROI 112 to increase the resolution of the region of interest. For example, the receiver, validator, and scaler 110 can scale content of the ROI 112 by a factor of 2× using bi-cubic interpolation to increase saliency representative and discrete-contour precision. The scaling may be used to improve the quality of the filtering process”); and inputting the second intermediate feature map of the frame into a downsampling layer of the feature extraction module to perform M downsampling operations, where M is a positive integer greater than or equal to 1, to obtain the image feature map for the frame (See Guler: Figs. 8A-B, and [0097], “In some embodiments, the neural networks used in the methods described above comprise an ImageNet pre-trained ResNet-50 network as a system backbone. Each dense prediction task may use a deconvolutional head, following the architectural choices of “Simple baselines for human pose estimation and tracking” (B. Xiao et al., arXiv:1804.06208, 2018), with threes 4×4 deconvolutional layers applied with batch-norm and ReLU. These are followed by a linear layer to obtain outputs of the desired dimensionality. Each deconvolution layer has a stride of two, leading to an output resolution of 64×64 given 256×256 images as input”. Note that the input image is progressively downsampling by a factor of 2 till the desired resolution, the factor 2 is mapped to M >=1). Regarding claim 4, Guler, Faras, and Gonzalez teach all the features with respect to claim as outlined above. Further, Guler and Gonzalez teach that the method according to claim 3, wherein the skip connection layer adopts an encoder-decoder structure, and inputting the frame into the skip connection layer of the feature extraction module to perform multi-resolution feature map extractions and skip connections of feature maps with the same resolution to obtain a second intermediate feature map for the frame, comprises: inputting the frame into the encoder of the skip connection layer to encode the frame and obtain an initial feature map for the frame, and sequentially performing N downsampling operations on the initial feature map to obtain a first intermediate feature map, where N is a positive integer (See Guler: Figs. 1-2, and [0060], “Operations 2.3 to 2.6 may be performed by a neural network, such as an encoder neural network. The neural network may have multiple heads, each outputting data relating to a different aspect of the method”; and [0097], “In some embodiments, the neural networks used in the methods described above comprise an ImageNet pre-trained ResNet-50 network as a system backbone. Each dense prediction task may use a deconvolutional head, following the architectural choices of “Simple baselines for human pose estimation and tracking” (B. Xiao et al., arXiv:1804.06208, 2018), with threes 4×4 deconvolutional layers applied with batch-norm and ReLU. These are followed by a linear layer to obtain outputs of the desired dimensionality. Each deconvolution layer has a stride of two, leading to an output resolution of 64×64 given 256×256 images as input”); and inputting the first intermediate feature map into the decoder of the skip connection layer, sequentially performing N upsampling operations on the first intermediate feature map (See Gonzalez: Figs. 1 and 4, and [0025], “The receiver, validator, and scaler 110 can perform preprocessing on the received image 104, 108. For example, the receiver, validator, and scaler 110 can perform detection of image compression, content validation, and spatial scaling. In some examples, the receiver, validator, and scaler 110 can detect a region of interest (ROI) 112 in a received image and a compression ratio to automatically adjust parameter boundaries for subsequent adaptive algorithms. The region of interest 112 may be a rectangular region of an image representing an object to be modeled. In some examples, the object to be modeled can be determined based on pixel color. For example, the region of interest may be a rectangular region including pixels that have a different color than a background color. In some examples, at least one of the pixels of the object to be modeled may touch each edge of the edge of the region of interest. In some examples, the receiver, validator, and scaler 110 can validate the region of interest. For example, the receiver, validator, and scaler 110 can compare a detected region of interest with the corresponding image including the region of interest. In some examples, if the region of interest is the same as the corresponding image, then the image may be rejected. In some examples, if the region of interest is not the same as the corresponding image, then the region of interest may be further processed as described below. The receiver, validator, and scaler 110 can then add a spatial-convolutional margin to the ROI 112. For example, the spatial-convolutional margin can be a margin of pixels that can be used during filtering of the region of interest. For example, the spatial-convolutional margin may be a margin of 5-7 pixels that can be used by the filter to prevent damage to content in the region of interest. The receiver, validator, and scaler 110 may then scale the ROI 112 to increase the resolution of the region of interest. For example, the receiver, validator, and scaler 110 can scale content of the ROI 112 by a factor of 2× using bi-cubic interpolation to increase saliency representative and discrete-contour precision. The scaling may be used to improve the quality of the filtering process”; and [0058], “The shaper sampler and blender 420 can sample the continuous contour. For example, shaper sampler and blender 420 can sample the continuous contour for explicit mesh tessellation. In some examples, the sampling can be driven by a desired target resolution. Evaluation of the contour points within overlapping regions can be performed by using a kernel weighted combination of radial polynomials”), and, during each upsampling operation, performing a skip connection with the feature map of the same resolution obtained through the downsampling operations in the encoder, to obtain the second intermediate feature map for the frame (See Guler: Figs. 1-2, and [0069], “Additional sets of data, such as DensePose and 3D joint estimation, can be used to increase the accuracy of 3D reconstruction. This may be done both by introducing additional losses that can regularize the training in a standard multi-task learning setup and by providing bottom-up, CNN-based cues to drive the top-down, model-based 3D reconstruction. For this, rather than relying on a CNN-based system to deliver accurate results in a single shot at test time, predictions from the CNN are used as an initialization to an iterative fitting procedure that uses the pose estimation results delivered by a 2D decoder to drive the model parameter updates. This can result in an effective 3D reconstruction. For this, the parametric model parameters are updated so as to align the parametric model-based and CNN-based pose estimates, as captured by a combination of a Dense Pose-based loss and the distances between the parametric model-based and CNN-based estimates of the 3D joints. This allows the parametric model parameters to be updated on-the-fly, so as to better match the CNN-based localization results”). Regarding claim 10. , Guler, Faras, and Gonzalez teach all the features with respect to claim 1 as outlined above. Further, Guler and Faras teach that the method according to claim 1, wherein inputting the target concatenated feature vector into the parameter regression network to predict a plurality of control parameters for model control based on a number of parameters (See Guler: Figs. 1-5, and [0036], “The 3D reconstruction method relies on a prior model of the target shape. The prior model may be a parametric model of the object being reconstructed. The type of prior model used may depend on the object being reconstructed. For example, in embodiments where the shape is a human body, the human body may be parametrised using a parametric deformable model. In the following, the method will be described with reference to a parametric deformable model, though it will be appreciated that other parametric shape models may alternatively be used. In some embodiments, SMPL is used. SMPL uses two parameter vectors to capture pose and shape: θ comprises 23 3D-rotation matrices corresponding to each joint in a kinematic tree for the human pose, and β captures shape variability across subjects in terms of a 10-dimensional shape vector. SMPL uses these parameters to obtain a triangulated mesh of the human body through linear skinning and blend shapes as a differentiable function of θ, β. In some embodiments, the parametric deformable model may be modified to allow it to be used more efficiently with convolutional neural networks, as described below in relation to FIG. 3”; and [0093], “At operation 5.5, parameters of the fused three-dimensional representation are adjusted based on the error value. The adjustment may be performed using learning based optimisation methods”), comprises: inputting the target concatenated feature vector into the parameter regression network (See Faras: Figs. 1-4, and [0039], “Localization 204 of FIG. 2 is used to determine a 3D map and/or poses, which may be important factors of creating a 3D representation. Some embodiments described herein arise from the recognition that, in image processing operations to create a 3D representation of a subject from images captured by an image capture device, the 3D representation may be degraded if the corresponding pose of the image capture device and/or a related 3D map cannot be accurately determined. Embodiments described herein are thus directed to using improved techniques to combine recursive and non-recursive approaches to creating and/or updating a 3D map and/or determining accurate estimated poses of the image capture device. Recursive techniques relate to or involve the repeated application of a rule, definition, or procedure to successive results. Any of the operations described herein as being recursive may be performed in a causal manner on the poses and operations described as being non-recursive may be performed in an acausal manner on the poses, respectively”; and [0041], “FIGS. 4A and 4B illustrate localization of 3D poses, corresponding to block 204 of FIG. 2. Referring now to FIG. 4A, several 2D images of an object such as the face of a person have been collected during a portion of a scan. The poses 1 to 19 are estimated at various camera viewpoints of the 2D images. A 3D map 410 of the object including various 3D points 420 is constructed. Referring now to FIG. 4B, the scan is continued by capturing additional 2D images for which poses 20 to 96 are estimated. The 3D map 410 includes additional 3D points 420 that have been triangulated from the additional 2D images”. Note that the pose and shape of the human body is recursively updated, and this is explicitly mapped to the regression network for pose and shape parameters); and performing at least one multilayer perceptron (MLP) operation on the target concatenated feature vector based on the number of parameters to obtain a plurality of control parameters for model control (See Guler: Fig.1, and [0039], “Neural networks comprise a plurality of layers of nodes, each node associated with one or more parameters. The parameters of each node of a neural network may comprise one or more weights and/or biases. The nodes take as input one or more outputs of nodes in the previous layer, or, in an initial layer, input data. The one or more outputs of nodes in the previous layer are used by each node to generate an activation value using an activation function and the parameters of the neural network. One or more of the layers of the neural network may be convolutional layers”). Regarding claim 12, Guler, Faras, and Gonzalez teach all the features with respect to claim 1 as outlined above. Further, Guler teaches that the method according to claim 1, wherein after obtaining the target three-dimensional model (See Guler: Fig. 1, and [0012], “According to a further aspect, this specification describes a method for updating a fused three-dimensional representation of an object comprised in a two-dimensional image, the method comprising: projecting the fused three-dimensional representation on to a two-dimensional image plane resulting in a projected representation; comparing respective positions of the projected representation with the object in the two-dimensional image; determining an error value based on the comparison; and adjusting parameters of the fused three-dimensional representation based on the error value, wherein the comparing, measuring and adjusting are iterated until the measured error value is below a predetermined threshold value or a threshold number of iterations is surpassed”), the method further comprises: for each frame of the plurality of frames of images, adapting the target three-dimensional model to the target object in the frame based on the camera pose data recorded at the time the frame was captured, and selecting a product compatible with the target object based on an adaptation result (See Guler: Fig. 1, and [0035], “The method uses a tight prior model of the object shape (such as a deformable parametric model for human body reconstruction) in tandem with a multitude of pose estimation methods to derive accurate monocular 3D reconstruction. In embodiments relating to 3D human reconstruction, the articulated nature of the human body is taken into account, which can substantially improve performance over a monolithic baseline. A refinement procedure may also be used that allows the shape prediction results of a single-shot system to be adapted so as to meet geometric constraints imposed by complementary, fully-convolutional networks/neural network branches”; [0030], “Exploiting the fact that parametric deformable models of the human body model provides a low-dimensional, differentiable representation of the human body, these works have trained systems to regress model parameters by minimizing the re-projection error between parametric deformable model of human body-based 3D keypoints and 2D joint annotations, human segmentation masks and 3D volume projections or even refining to the level of body parts”; and [0093], “At operation 5.5, parameters of the fused three-dimensional representation are adjusted based on the error value. The adjustment may be performed using learning based optimisation methods”); and/or inputting any frame of the plurality of frames of images into a depth estimation network to estimate size information of the target object, and annotating the target three-dimensional model based on the estimated size information of the target object; and/or for each frame of the plurality of frames of images, adapting the target three-dimensional model to the target object in the frame based on the camera pose data recorded at the time the frame was captured, and measuring shape parameters of the target object based on the adaptation result. Regarding claim 13, Guler, Faras, and Gonzalez teach all the features with respect to claim 1 as outlined above. Further, Guler teaches that the method according to claim 1, wherein the target object is a foot object, hand object, head object, elbow object, or leg object on a human body, and the three-dimensional model description information corresponding to the target object is determined based on an SMPL model (See Guler: Fig. 1, and [0044], “Each feature extracted around a 2D keypoint/joint can deliver a separate parametric deformable model parameter estimate. But intuitively a 2D keypoint/joint should have a stronger influence on parametric deformable model parameters that are more relevant to it—for instance a left wrist joint should be affecting the left arm parameters, but not those of kinetically independent parts such as the right arm, head, or feet. Furthermore, the fact that some keypoints/joints can be missing from an image means that we cannot simply concatenate the features in a larger feature vector, but need to use a form that handles gracefully the potential absence of parts”; and [0036], “The 3D reconstruction method relies on a prior model of the target shape. The prior model may be a parametric model of the object being reconstructed. The type of prior model used may depend on the object being reconstructed. For example, in embodiments where the shape is a human body, the human body may be parametrised using a parametric deformable model. In the following, the method will be described with reference to a parametric deformable model, though it will be appreciated that other parametric shape models may alternatively be used. In some embodiments, SMPL is used. SMPL uses two parameter vectors to capture pose and shape: θ comprises 23 3D-rotation matrices corresponding to each joint in a kinematic tree for the human pose, and β captures shape variability across subjects in terms of a 10-dimensional shape vector. SMPL uses these parameters to obtain a triangulated mesh of the human body through linear skinning and blend shapes as a differentiable function of θ, β. In some embodiments, the parametric deformable model may be modified to allow it to be used more efficiently with convolutional neural networks, as described below in relation to FIG. 3”). Regarding claim 14 , Guler, Faras, and Gonzalez teach all the features with respect to claim 1 as outlined above. Further, Guler teaches that the non-transitory computer-readable storage medium configured with instructions executable by one or more processors to cause the one or more processors to perform the method according to claim 1 (See Guker: Fig. 7, and [0104], “The various example embodiments described herein are described in the general context of method steps or processes, which may be implemented in one aspect by a computer program product, embodied in a computer-readable medium, including computer-executable instructions, such as program code, executed by computers in networked environments. A computer-readable medium may include removable and non-removable storage devices including, but not limited to, Read Only Memory (ROM), Random Access Memory (RAM), compact discs (CDs), digital versatile discs (DVD), etc. Generally, program modules may include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. Computer-executable instructions, associated data structures, and program modules represent examples of program code for executing steps of the methods disclosed herein. The particular sequence of such executable instructions or associated data structures represents examples of corresponding acts for implementing the functions described in such steps or processes”). Regarding claim 15, Guler, Faras, and Gonzalez teach all the features with respect to claim 1 as outlined above. Further, Guler teaches that an electronic device (See Guler: Fig. 7, and [0101], “FIG. 7 shows an example of a computational unit/apparatus. The apparatus 700 comprises a processing apparatus 702 and memory 704 according to an exemplary embodiment. Computer-readable code 706 may be stored on the memory 704 and may, when executed by the processing apparatus 702, cause the apparatus 700 to perform methods as described herein”) comprising: one or more processors (See Guler: Fig. 7, and [0102], “The processing apparatus 702 may be of any suitable composition and may include one or more processors of any suitable type or suitable combination of types. Indeed, the term “processing apparatus” should be understood to encompass computers having differing architectures such as single/multi-processor architectures and sequencers/parallel architectures. For example, the processing apparatus may be a programmable processor that interprets computer program instructions and processes data. The processing apparatus may include plural programmable processors. Alternatively, the processing apparatus may be, for example, programmable hardware with embedded firmware. The processing apparatus may alternatively or additionally include Graphics Processing Units (GPUs), or one or more specialized circuits such as field programmable gate arrays FPGA, Application Specific Integrated Circuits (ASICs), signal processing devices etc. In some instances, processing apparatus may be referred to as computing apparatus or processing means”); and one or more computer-readable memories coupled to the one or more processors and having instructions stored thereon that are executable by the one or more processors to perform the method according to claim 1 (See Guler: Fig. 7, and [0103], “The processing apparatus 702 is coupled to the memory 704 and is operable to read/write data to/from the memory 704. The memory 704 may comprise a single memory unit or a plurality of memory units, upon which the computer readable instructions (or code) is stored. For example, the memory may comprise both volatile memory and non-volatile memory. In such examples, the computer readable instructions/program code may be stored in the non-volatile memory and may be executed by the processing apparatus using the volatile memory for temporary storage of data or data and instructions. Examples of volatile memory include RAM. DRAM, and SDRAM etc. Examples of non-volatile memory include ROM, PROM, EEPROM, flash memory, optical storage, magnetic storage, etc.”). Regarding claim 20, Guler, Faras, Gonzalez, and Ayush teach all the features with respect to claim 1 as outlined above. Further, Guler, Faras, and Gonzalez teach that the three-dimensional reconstruction device (See Guler: Fig. 7, and [0101], “FIG. 7 shows an example of a computational unit/apparatus. The apparatus 700 comprises a processing apparatus 702 and memory 704 according to an exemplary embodiment. Computer-readable code 706 may be stored on the memory 704 and may, when executed by the processing apparatus 702, cause the apparatus 700 to perform methods as described herein”), comprising: an image acquisition unit (See Faras: Fig. 1, and [0034], “FIG. 1 illustrates a user taking pictures with a camera at various locations around the object. Although the foregoing examples discuss the images acquired from a camera, the images that are processed may be previously residing in memory or the images be sent to the processing unit for processing according to various embodiments described herein. Furthermore, a face of a person is discussed herein as an example object, but the techniques described herein may apply to any object for which a 2D image can be acquired. Referring now to FIG. 1, a user 110 has a camera 100 for which that they initiate a photographic session of an object 135, such as a person's face, at location 120a. Relative movement between the camera 100 and the object 135 takes place”), configured to acquire a plurality of frames of images of a target object and three-dimensional model description information corresponding to the target object (See Guler: Figs. 1-3, and [0034], “FIG. 1 shows a schematic overview of a method 100 of part-based three-dimensional reconstruction 102 from a two-dimensional image 104. A keypoint detector indicates 2D keypoint (e.g. joint) positions, driving feature pooling. These in turn are used to determine 3D model parameters of the parametric model (e.g. 3D SMPL joint angles). The method may be implemented on one or more computing devices. The method extracts localized features around key points (in the example shown, human joints) following a part-based modelling paradigm. The resulting model may be referred to as a fused three dimensional representation”; [0035], “The method uses a tight prior model of the object shape (such as a deformable parametric model for human body reconstruction) in tandem with a multitude of pose estimation methods to derive accurate monocular 3D reconstruction. In embodiments relating to 3D human reconstruction, the articulated nature of the human body is taken into account, which can substantially improve performance over a monolithic baseline. A refinement procedure may also be used that allows the shape prediction results of a single-shot system to be adapted so as to meet geometric constraints imposed by complementary, fully-convolutional networks/neural network branches”; and [0049], “At operation 2.1, a two-dimensional image is received. As described above, the 2D image may comprise a set of pixel values corresponding to a two-dimensional array. In some embodiments, a single image is received. Alternatively, a plurality of images may be received. If multiple images are taken (for example, of a person moving), the three dimensional representation may account for these changes”. Note that the human body is mapped to the target object); a feature extraction unit, configured to input the plurality of frames of images into a feature extraction network to extract features and obtain feature vectors of the plurality of frames of images (See Guler: Figs. 1-3, and [0037], “A two dimensional image 104 is received, and input into a neural network 106. The two-dimensional image, I, comprises a set of pixel values corresponding to a two-dimensional array. For example, in a colour image. I∈custom-character.sup.H×W×3, where H is the height of the image in pixels, W is the height of the image in pixels and the image has three colour channels (e.g. RGB or CIELAB). The two dimensional image may, in some embodiments, be in black-and-white/greyscale”; [0044], “Each feature extracted around a 2D keypoint/joint can deliver a separate parametric deformable model parameter estimate. But intuitively a 2D keypoint/joint should have a stronger influence on parametric deformable model parameters that are more relevant to it—for instance a left wrist joint should be affecting the left arm parameters, but not those of kinetically independent parts such as the right arm, head, or feet. Furthermore, the fact that some keypoints/joints can be missing from an image means that we cannot simply concatenate the features in a larger feature vector, but need to use a form that handles gracefully the potential absence of parts”; and [0036], “The 3D reconstruction method relies on a prior model of the target shape. The prior model may be a parametric model of the object being reconstructed. The type of prior model used may depend on the object being reconstructed. For example, in embodiments where the shape is a human body, the human body may be parametrised using a parametric deformable model. In the following, the method will be described with reference to a parametric deformable model, though it will be appreciated that other parametric shape models may alternatively be used. In some embodiments, SMPL is used. SMPL uses two parameter vectors to capture pose and shape: θ comprises 23 3D-rotation matrices corresponding to each joint in a kinematic tree for the human pose, and β captures shape variability across subjects in terms of a 10-dimensional shape vector. SMPL uses these parameters to obtain a triangulated mesh of the human body through linear skinning and blend shapes as a differentiable function of θ, β. In some embodiments, the parametric deformable model may be modified to allow it to be used more efficiently with convolutional neural networks, as described below in relation to FIG. 3”); a vector concatenation unit, configured to concatenate the feature vectors of the plurality of frames of images to obtain a target concatenated feature vector (See Guler: Figs. 1-3, and [0047], “An arg-soft-max operation may be used over angle clusters, and information from multiple 2D joints is fused: w.sup.i.sub.k,j indicates a score that 2D joint j assigns to cluster k for the i-th model parameter, θ.sup.i. The clusters can be defined as described below in relation to the “mixture-of-experts prior”. The neighbourhood of i is constructed by inspecting which model parameters directly influence human 2D joints, based on kinematic tree dependencies. Joints are found in the image by taking the maximum of a 2D joint detection module. If the maximum is below a threshold (for example, 0:4 in some embodiments) it is considered consider that a joint is not observed in the image. In that case, every summand corresponding to a missing joint is excluded from the summations in the numerator and denominator, so that the above equation still delivers a distribution over angles. If all elements of N(i) are missing, θ.sup.i is set to a resting pose”; [0056], “At operation 2.6, the respective three-dimensional representations are combined to generate a fused three-dimensional representation of the object”; and [0058], The method may further comprise applying the fused three-dimensional representation of the object to a part-based three-dimensional shape representation model. An example of such a model is a kinematic tree”. Note that combining and fusion is mapped to concatenating); a parameter regression unit, configured to input the target concatenated feature vector into a parameter regression network, and predict, based on a number of parameters, a plurality of control parameters for model control (See Guler: Figs. 1-5, and [0036], “The 3D reconstruction method relies on a prior model of the target shape. The prior model may be a parametric model of the object being reconstructed. The type of prior model used may depend on the object being reconstructed. For example, in embodiments where the shape is a human body, the human body may be parametrised using a parametric deformable model. In the following, the method will be described with reference to a parametric deformable model, though it will be appreciated that other parametric shape models may alternatively be used. In some embodiments, SMPL is used. SMPL uses two parameter vectors to capture pose and shape: θ comprises 23 3D-rotation matrices corresponding to each joint in a kinematic tree for the human pose, and β captures shape variability across subjects in terms of a 10-dimensional shape vector. SMPL uses these parameters to obtain a triangulated mesh of the human body through linear skinning and blend shapes as a differentiable function of θ, β. In some embodiments, the parametric deformable model may be modified to allow it to be used more efficiently with convolutional neural networks, as described below in relation to FIG. 3”; and [0093], “At operation 5.5, parameters of the fused three-dimensional representation are adjusted based on the error value. The adjustment may be performed using learning based optimisation methods”), wherein the plurality of control parameters include pose control parameters and shape control parameters (See Faras: Figs. 1-4, and [0039], “Localization 204 of FIG. 2 is used to determine a 3D map and/or poses, which may be important factors of creating a 3D representation. Some embodiments described herein arise from the recognition that, in image processing operations to create a 3D representation of a subject from images captured by an image capture device, the 3D representation may be degraded if the corresponding pose of the image capture device and/or a related 3D map cannot be accurately determined. Embodiments described herein are thus directed to using improved techniques to combine recursive and non-recursive approaches to creating and/or updating a 3D map and/or determining accurate estimated poses of the image capture device. Recursive techniques relate to or involve the repeated application of a rule, definition, or procedure to successive results. Any of the operations described herein as being recursive may be performed in a causal manner on the poses and operations described as being non-recursive may be performed in an acausal manner on the poses, respectively”; and [0041], “FIGS. 4A and 4B illustrate localization of 3D poses, corresponding to block 204 of FIG. 2. Referring now to FIG. 4A, several 2D images of an object such as the face of a person have been collected during a portion of a scan. The poses 1 to 19 are estimated at various camera viewpoints of the 2D images. A 3D map 410 of the object including various 3D points 420 is constructed. Referring now to FIG. 4B, the scan is continued by capturing additional 2D images for which poses 20 to 96 are estimated. The 3D map 410 includes additional 3D points 420 that have been triangulated from the additional 2D images”. Note that the pose and shape of the human body is recursively updated, and this is explicitly mapped to the regression network for pose and shape parameters); and a masking processing unit, configured to apply a masking operation on an initial three-dimensional model of the target object (See Gonzalez: Figs. 1-3, and [0026], “The segment extractor 114 can receive the validated and scaled ROIs 112 and generate one or more binary segments 116, textured segments 118, and segment characterizations 120 based on the validated and scaled ROIs 112. For example, the segment extractor 114 can extract a binary segment representing the shape of an object from a raw image 108. For example, the binary segment may be a two-dimensional shape corresponding to the object's 2D projection on an image plane. The binary segment may have attributes such as a center of mass, bounding box, total pixels, principal axes, and a binary mask. In some examples, the binary mask may be applied to a source color image to generate a texture image as described below. In some examples, the segment extractor 114 may also include adaptive filters and graph-based segmentation that can be used to prevent and solve problems caused by image compression artifacts and unpropitious lighting and material combinations. The segment extractor 114 can thus perform fast and fully automatic removal of compression artifacts. In some examples, the segment extractor 114 can also perform speckle removal by adaptive rejection of pixels using their 8-connectivity degree in a recursive fashion. In some examples, the speckle removal can be further refined by smoothing expansion and contraction. The processing performed by the segment extractor 114 may generally be referred to herein as a signal phase. Because the signal phase may be the foundation for all subsequent phases, the segment extractor 114 may thus include high quality smoothing and acutance-preserving contouring. The operation of the segment extractor 114 is described in detail with respect to the example segment extractor of FIG. 2 below. The final output of the segment extractor may be a binary segment 116, a textured segment 118, and a binary segment with segment characterization 120”; and [0046], “In FIG. 3, an ROI 112 is received. For example, the ROI may be a region corresponding to a light bulb identified in a 2D color image. The ROI 112 may then be desaturated by a desaturator to produce a grayscale image, or desaturated ROI 304. An intensity enhancer may then improve image intensity of the desaturated ROI 304 to generate an intensity enhanced ROI 306. An adaptive binarization may then be applied to the intensity enhanced ROI 306 to generate a binary image 308. For example, the adaptive binarization may remove gradients and result in an image of black and white pixels. A segmentation can then be applied to the binary image 308 by a segmenter to extract a binary segment 310. Hole removal may then be performed by a hole remover on the extract binary segments 310 to generate a binary mask image 312. Speckle removal may then be performed by a speckle remover on the binary mask image 312 to generate a speckle-less binary mask image 314. For example, the speckle remover may remove spurious contour artifacts that may have resulted from image compression. An expander may then apply diffusion on the speckle-less binary mask image 314 to generate a diffused binary mask image 316. A contraction may then apply a non-adaptive binarization on the diffused binary mask image 316 to generate a binary image 318. The binary image 318 may be used to generate a binary segment 116 and a textured segment 118. For example, the binary segment 116 may be generated as described above with respect to FIG. 1. In some examples, the texture image 118 can be generated by applying the binary mask 216 to the source image to select pixels from the source image masked by the binary mask 216”) based on the pose control parameters and shape control parameters to generate a target three-dimensional model of the target object, wherein the initial three-dimensional model is obtained based on the three-dimensional model description information (See Guler: Figs. 4-5, and [0069], “Additional sets of data, such as DensePose and 3D joint estimation, can be used to increase the accuracy of 3D reconstruction. This may be done both by introducing additional losses that can regularize the training in a standard multi-task learning setup and by providing bottom-up, CNN-based cues to drive the top-down, model-based 3D reconstruction. For this, rather than relying on a CNN-based system to deliver accurate results in a single shot at test time, predictions from the CNN are used as an initialization to an iterative fitting procedure that uses the pose estimation results delivered by a 2D decoder to drive the model parameter updates. This can result in an effective 3D reconstruction. For this, the parametric model parameters are updated so as to align the parametric model-based and CNN-based pose estimates, as captured by a combination of a Dense Pose-based loss and the distances between the parametric model-based and CNN-based estimates of the 3D joints. This allows the parametric model parameters to be updated on-the-fly, so as to better match the CNN-based localization results”; [0073], “The method 400 proceeds in a similar way to the method described above with respect to FIGS. 1-3. A two dimensional image 404 is received, and input into a neural network 406. The neural network 406 processes the input image 404 to generate an initial 3D model 408 (i.e. a fused three-dimensional model/3D surface estimate) of objects in the image (e.g. a human body) using part-based 3D shape reconstruction 410 (for example, as described above). The neural network further outputs 2D keypoint locations (e.g. 2D joints), 3D keypoint locations (e.g. 3D joints) and a DensePose estimate. Each of these may be output from a different head 406b-e of the neural network 406. The neural network 406 may comprise one or more convolutional and/or deconvolutional layers/networks. The initial layers of the neural network 406 may be shared and form a backbone network 406a. In the example shown, the backbone network 406a is a ResNet-50 network, though other networks may alternatively be used”; and [0074], “The initial 3D model 408 may be refined using a 3D shape refinement process 412. The 3D shape refinement process 412 takes as input a 3D model, 2D keypoint locations, 3D keypoint locations and the DensePose estimate for the corresponding 2D image, and uses them to refine the 3D model to generate a refined 3D model of the object 402”. Note that the pose and shape parameters are inputted to the CNN model, and the output of the CNN model is also fed back to the CNN model iteratively to improve the accuracy of the CNN reconstruction results; and the refinement process is fundamentally the masking operation but a new art will be used for clarity). Claims 16-19 are rejected under 35 U.S.C. 103 as being unpatentable over Guler, etc. (US 20210241522 A1) in view of Faras, etc. (US 20210118160 A1), further in view of Gonzalez, etc. (US 20180211438 A1), and Ayush, etc. (US 20190378204 A1). Regarding claim 16, Guler, Faras, and Gonzalez teach all the features with respect to claim 1 as outlined above. Further, Guler, Faras, and Gonzalez teach that the method for processing product information (See Guler: Fig. 7, and [0101], “FIG. 7 shows an example of a computational unit/apparatus. The apparatus 700 comprises a processing apparatus 702 and memory 704 according to an exemplary embodiment. Computer-readable code 706 may be stored on the memory 704 and may, when executed by the processing apparatus 702, cause the apparatus 700 to perform methods as described herein”), comprising: obtaining a plurality of frames of images containing a fitting subject, and three-dimensional model description information corresponding to the fitting subject (See Guler: Figs. 1-3, and [0034], “FIG. 1 shows a schematic overview of a method 100 of part-based three-dimensional reconstruction 102 from a two-dimensional image 104. A keypoint detector indicates 2D keypoint (e.g. joint) positions, driving feature pooling. These in turn are used to determine 3D model parameters of the parametric model (e.g. 3D SMPL joint angles). The method may be implemented on one or more computing devices. The method extracts localized features around key points (in the example shown, human joints) following a part-based modelling paradigm. The resulting model may be referred to as a fused three dimensional representation”; [0035], “The method uses a tight prior model of the object shape (such as a deformable parametric model for human body reconstruction) in tandem with a multitude of pose estimation methods to derive accurate monocular 3D reconstruction. In embodiments relating to 3D human reconstruction, the articulated nature of the human body is taken into account, which can substantially improve performance over a monolithic baseline. A refinement procedure may also be used that allows the shape prediction results of a single-shot system to be adapted so as to meet geometric constraints imposed by complementary, fully-convolutional networks/neural network branches”; and [0049], “At operation 2.1, a two-dimensional image is received. As described above, the 2D image may comprise a set of pixel values corresponding to a two-dimensional array. In some embodiments, a single image is received. Alternatively, a plurality of images may be received. If multiple images are taken (for example, of a person moving), the three dimensional representation may account for these changes”. Note that the human body is mapped to the target object); inputting the plurality of frames of images into a feature extraction network to extract features, to obtain feature vectors of the plurality of frames of images (See Guler: Figs. 1-3, and [0037], “A two dimensional image 104 is received, and input into a neural network 106. The two-dimensional image, I, comprises a set of pixel values corresponding to a two-dimensional array. For example, in a colour image. I∈custom-character.sup.H×W×3, where H is the height of the image in pixels, W is the height of the image in pixels and the image has three colour channels (e.g. RGB or CIELAB). The two dimensional image may, in some embodiments, be in black-and-white/greyscale”; [0044], “Each feature extracted around a 2D keypoint/joint can deliver a separate parametric deformable model parameter estimate. But intuitively a 2D keypoint/joint should have a stronger influence on parametric deformable model parameters that are more relevant to it—for instance a left wrist joint should be affecting the left arm parameters, but not those of kinetically independent parts such as the right arm, head, or feet. Furthermore, the fact that some keypoints/joints can be missing from an image means that we cannot simply concatenate the features in a larger feature vector, but need to use a form that handles gracefully the potential absence of parts”; and [0036], “The 3D reconstruction method relies on a prior model of the target shape. The prior model may be a parametric model of the object being reconstructed. The type of prior model used may depend on the object being reconstructed. For example, in embodiments where the shape is a human body, the human body may be parametrised using a parametric deformable model. In the following, the method will be described with reference to a parametric deformable model, though it will be appreciated that other parametric shape models may alternatively be used. In some embodiments, SMPL is used. SMPL uses two parameter vectors to capture pose and shape: θ comprises 23 3D-rotation matrices corresponding to each joint in a kinematic tree for the human pose, and β captures shape variability across subjects in terms of a 10-dimensional shape vector. SMPL uses these parameters to obtain a triangulated mesh of the human body through linear skinning and blend shapes as a differentiable function of θ, β. In some embodiments, the parametric deformable model may be modified to allow it to be used more efficiently with convolutional neural networks, as described below in relation to FIG. 3”), and concatenating the feature vectors of the plurality of frames of images to generate a target concatenated feature vector concatenating the feature vectors of the plurality of frames of images to generate a target concatenated feature vector (See Guler: Figs. 1-3, and [0047], “An arg-soft-max operation may be used over angle clusters, and information from multiple 2D joints is fused: w.sup.i.sub.k,j indicates a score that 2D joint j assigns to cluster k for the i-th model parameter, θ.sup.i. The clusters can be defined as described below in relation to the “mixture-of-experts prior”. The neighbourhood of i is constructed by inspecting which model parameters directly influence human 2D joints, based on kinematic tree dependencies. Joints are found in the image by taking the maximum of a 2D joint detection module. If the maximum is below a threshold (for example, 0:4 in some embodiments) it is considered consider that a joint is not observed in the image. In that case, every summand corresponding to a missing joint is excluded from the summations in the numerator and denominator, so that the above equation still delivers a distribution over angles. If all elements of N(i) are missing, θ.sup.i is set to a resting pose”; [0056], “At operation 2.6, the respective three-dimensional representations are combined to generate a fused three-dimensional representation of the object”; and [0058], The method may further comprise applying the fused three-dimensional representation of the object to a part-based three-dimensional shape representation model. An example of such a model is a kinematic tree”. Note that combining and fusion is mapped to concatenating); inputting the target concatenated feature vector into a parameter regression network, to predict, based on the three-dimensional model description information, a plurality of control parameters for model control (See Guler: Figs. 1-5, and [0036], “The 3D reconstruction method relies on a prior model of the target shape. The prior model may be a parametric model of the object being reconstructed. The type of prior model used may depend on the object being reconstructed. For example, in embodiments where the shape is a human body, the human body may be parametrised using a parametric deformable model. In the following, the method will be described with reference to a parametric deformable model, though it will be appreciated that other parametric shape models may alternatively be used. In some embodiments, SMPL is used. SMPL uses two parameter vectors to capture pose and shape: θ comprises 23 3D-rotation matrices corresponding to each joint in a kinematic tree for the human pose, and β captures shape variability across subjects in terms of a 10-dimensional shape vector. SMPL uses these parameters to obtain a triangulated mesh of the human body through linear skinning and blend shapes as a differentiable function of θ, β. In some embodiments, the parametric deformable model may be modified to allow it to be used more efficiently with convolutional neural networks, as described below in relation to FIG. 3”; and [0093], “At operation 5.5, parameters of the fused three-dimensional representation are adjusted based on the error value. The adjustment may be performed using learning based optimisation methods”), wherein the plurality of control parameters include pose control parameters and shape control parameters (See Faras: Figs. 1-4, and [0039], “Localization 204 of FIG. 2 is used to determine a 3D map and/or poses, which may be important factors of creating a 3D representation. Some embodiments described herein arise from the recognition that, in image processing operations to create a 3D representation of a subject from images captured by an image capture device, the 3D representation may be degraded if the corresponding pose of the image capture device and/or a related 3D map cannot be accurately determined. Embodiments described herein are thus directed to using improved techniques to combine recursive and non-recursive approaches to creating and/or updating a 3D map and/or determining accurate estimated poses of the image capture device. Recursive techniques relate to or involve the repeated application of a rule, definition, or procedure to successive results. Any of the operations described herein as being recursive may be performed in a causal manner on the poses and operations described as being non-recursive may be performed in an acausal manner on the poses, respectively”; and [0041], “FIGS. 4A and 4B illustrate localization of 3D poses, corresponding to block 204 of FIG. 2. Referring now to FIG. 4A, several 2D images of an object such as the face of a person have been collected during a portion of a scan. The poses 1 to 19 are estimated at various camera viewpoints of the 2D images. A 3D map 410 of the object including various 3D points 420 is constructed. Referring now to FIG. 4B, the scan is continued by capturing additional 2D images for which poses 20 to 96 are estimated. The 3D map 410 includes additional 3D points 420 that have been triangulated from the additional 2D images”. Note that the pose and shape of the human body is recursively updated, and this is explicitly mapped to the regression network for pose and shape parameters); applying a masking operation to an initial three-dimensional model of the fitting subject (See Gonzalez: Figs. 1-3, and [0026], “The segment extractor 114 can receive the validated and scaled ROIs 112 and generate one or more binary segments 116, textured segments 118, and segment characterizations 120 based on the validated and scaled ROIs 112. For example, the segment extractor 114 can extract a binary segment representing the shape of an object from a raw image 108. For example, the binary segment may be a two-dimensional shape corresponding to the object's 2D projection on an image plane. The binary segment may have attributes such as a center of mass, bounding box, total pixels, principal axes, and a binary mask. In some examples, the binary mask may be applied to a source color image to generate a texture image as described below. In some examples, the segment extractor 114 may also include adaptive filters and graph-based segmentation that can be used to prevent and solve problems caused by image compression artifacts and unpropitious lighting and material combinations. The segment extractor 114 can thus perform fast and fully automatic removal of compression artifacts. In some examples, the segment extractor 114 can also perform speckle removal by adaptive rejection of pixels using their 8-connectivity degree in a recursive fashion. In some examples, the speckle removal can be further refined by smoothing expansion and contraction. The processing performed by the segment extractor 114 may generally be referred to herein as a signal phase. Because the signal phase may be the foundation for all subsequent phases, the segment extractor 114 may thus include high quality smoothing and acutance-preserving contouring. The operation of the segment extractor 114 is described in detail with respect to the example segment extractor of FIG. 2 below. The final output of the segment extractor may be a binary segment 116, a textured segment 118, and a binary segment with segment characterization 120”; and [0046], “In FIG. 3, an ROI 112 is received. For example, the ROI may be a region corresponding to a light bulb identified in a 2D color image. The ROI 112 may then be desaturated by a desaturator to produce a grayscale image, or desaturated ROI 304. An intensity enhancer may then improve image intensity of the desaturated ROI 304 to generate an intensity enhanced ROI 306. An adaptive binarization may then be applied to the intensity enhanced ROI 306 to generate a binary image 308. For example, the adaptive binarization may remove gradients and result in an image of black and white pixels. A segmentation can then be applied to the binary image 308 by a segmenter to extract a binary segment 310. Hole removal may then be performed by a hole remover on the extract binary segments 310 to generate a binary mask image 312. Speckle removal may then be performed by a speckle remover on the binary mask image 312 to generate a speckle-less binary mask image 314. For example, the speckle remover may remove spurious contour artifacts that may have resulted from image compression. An expander may then apply diffusion on the speckle-less binary mask image 314 to generate a diffused binary mask image 316. A contraction may then apply a non-adaptive binarization on the diffused binary mask image 316 to generate a binary image 318. The binary image 318 may be used to generate a binary segment 116 and a textured segment 118. For example, the binary segment 116 may be generated as described above with respect to FIG. 1. In some examples, the texture image 118 can be generated by applying the binary mask 216 to the source image to select pixels from the source image masked by the binary mask 216”) based on the pose control parameters and shape control parameters to generate a target three-dimensional model of the fitting subject, wherein the initial three-dimensional model is generated based on the three-dimensional model description information (See Guler: Figs. 4-5, and [0069], “Additional sets of data, such as DensePose and 3D joint estimation, can be used to increase the accuracy of 3D reconstruction. This may be done both by introducing additional losses that can regularize the training in a standard multi-task learning setup and by providing bottom-up, CNN-based cues to drive the top-down, model-based 3D reconstruction. For this, rather than relying on a CNN-based system to deliver accurate results in a single shot at test time, predictions from the CNN are used as an initialization to an iterative fitting procedure that uses the pose estimation results delivered by a 2D decoder to drive the model parameter updates. This can result in an effective 3D reconstruction. For this, the parametric model parameters are updated so as to align the parametric model-based and CNN-based pose estimates, as captured by a combination of a Dense Pose-based loss and the distances between the parametric model-based and CNN-based estimates of the 3D joints. This allows the parametric model parameters to be updated on-the-fly, so as to better match the CNN-based localization results”; [0073], “The method 400 proceeds in a similar way to the method described above with respect to FIGS. 1-3. A two dimensional image 404 is received, and input into a neural network 406. The neural network 406 processes the input image 404 to generate an initial 3D model 408 (i.e. a fused three-dimensional model/3D surface estimate) of objects in the image (e.g. a human body) using part-based 3D shape reconstruction 410 (for example, as described above). The neural network further outputs 2D keypoint locations (e.g. 2D joints), 3D keypoint locations (e.g. 3D joints) and a DensePose estimate. Each of these may be output from a different head 406b-e of the neural network 406. The neural network 406 may comprise one or more convolutional and/or deconvolutional layers/networks. The initial layers of the neural network 406 may be shared and form a backbone network 406a. In the example shown, the backbone network 406a is a ResNet-50 network, though other networks may alternatively be used”; and [0074], “The initial 3D model 408 may be refined using a 3D shape refinement process 412. The 3D shape refinement process 412 takes as input a 3D model, 2D keypoint locations, 3D keypoint locations and the DensePose estimate for the corresponding 2D image, and uses them to refine the 3D model to generate a refined 3D model of the object 402”. Note that the pose and shape parameters are inputted to the CNN model, and the output of the CNN model is also fed back to the CNN model iteratively to improve the accuracy of the CNN reconstruction results; and the refinement process is fundamentally the masking operation but a new art will be used for clarity); and providing target product information compatible with the fitting subject based on the target three-dimensional model. However, Guler, modified by Faras and Gonzalez, fails to explicitly disclose that providing target product information compatible with the fitting subject based on the target three-dimensional model. However, Ayush teaches that providing target product information compatible with the fitting subject based on the target three-dimensional model (See Ayush: Figs. 1-6, and [0024], “In addition, the AR product recommendation system can analyze the viewpoint generated from the camera feed to identify one or more real-world objects. For instance, the AR product recommendation system can utilize a region-based convolutional neural network (“R-CNN”) to generate proposed regions of the viewpoint with corresponding probabilities of containing objects. Indeed, the AR product recommendation system can utilize an R-CNN to generate bounding boxes around regions of the viewpoint. The AR product recommendation system can further generate a confidence score (e.g., a probability) as well as an object label for each bounding box that indicates a likelihood of the bounding box containing a real-world object that corresponds to the given object label”; [0047], “As illustrated in FIG. 1, the environment includes the server(s) 104. The server(s) 104 may generate, store, receive, and transmit electronic data, such as AR content, digital video, digital images, metadata, etc. For example, the server(s) 104 may receive data from the user client device 108 in the form of a camera feed. In addition, the server(s) 104 can transmit data to the user client device 108 to provide an AR representation of a recommended product within a user's view of the camera feed. For example, the server(s) 104 can communicate with the user client device 108 to transmit and/or receive data via the network 116. In some embodiments, the server(s) 104 comprises a content server. The server(s) 104 can also comprise an application server, a communication server, a web-hosting server, a social networking server, or a digital content campaign serve”; and [0080], “As shown, the AR product recommendation system 102 can access a 3D collection 502 of three-dimensional models within the model database 112. In addition, the AR product recommendation system 102 can select a three-dimensional model from among the plurality of three-dimensional models within the 3D collection 502. To select a matching three-dimensional model, the AR product recommendation system 102 can analyze three-dimensional models within the model database 112 that match the object label associated with the identified real-world object 302. For example, the real-world object within the bounding box 402 has an object label of “chair.” Thus, the AR product recommendation system 102 analyzes three-dimensional models that are of the same class or that have the same label—the AR product recommendation system 102 analyzes chairs within the model database 112”). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention was effectively filed to modify Guler to have providing target product information compatible with the fitting subject based on the target three-dimensional model as taught by Ayush in order to enable generating effective product recommendations in a real-world environment by an AR product recommendation system (See Ayush: Fig. 1, and [0021], “This disclosure describes one or more embodiments of an augmented reality (“AR”) product recommendation system that accurately and flexibly generates AR product recommendations that match the style of surrounding real-world objects. In particular, the disclosed AR product recommendation system detects objects shown within an AR scene (received from a user's client device) and, based on the detected objects, selects products with matching styles to recommend to the user. For example, to generate an AR product recommendation, the AR product recommendation system identifies a real-world object within a camera view of a user client device to replace with an AR product recommendation. The AR product recommendation system further determines an object class associated with the identified real-world object. Based on the identified real-world object, the AR product recommendation system determines a three-dimensional model from a model database that matches the identified real-world object. The AR product recommendation system then utilizes a style similarity algorithm to generate a recommended product that is similar in style to the object being replaced from the camera view”). Guler teaches a method and system that may reconstruct three-dimensional models of objects from two-dimensional images by extracting features from the 2D input images, via neural networks, estimating the 3D representation of the objects from the extracted features, concatenating the estimated parameters, and generated the 3D model of the objects based on the fused 3D representation of the object; while Ayush teaches a system and method that may provide target product information for the target subject to try-on based on the subject 3D model reconstructed from the captured images. Therefore, it is obvious to one of ordinary skill in the art to modify Guler by Ayush to use the 3D body model to find the fitting try-on target products for the users or the target subject. The motivation to modify Guler by Ayush is “Use of known technique to improve similar devices (methods, or products) in the same way”. Regarding claim 17, Guler, Faras, Gonzalez, and Ayush teach all the features with respect to claim 16 as outlined above. Further, Ayush teaches that the method according to claim 14, wherein providing target product information compatible with the fitting subject based on the target three-dimensional model (See Ayush: Fig. 1, and [0028], “Upon generating recommended products, the AR product recommendation system can further generate AR representations of the recommended products to provide for display to the user. In particular, based on replacing the identified real-world object with a stylistically similar recommended product, the AR product recommendation system can utilize a color compatibility algorithm to generate a color compatibility score as a basis for generating an AR representation of a recommended product that matches a color theme present in the viewpoint. Thus, based on the color compatibility score as well as the determined similarity score, the AR product recommendation system can embed the color-compatible AR representation of the recommended product within the real-world environment of the camera feed. For example, the AR product recommendation system can remove an identified real-world object from the depicted real-world environment and replace the real-world object with a stylistically similar AR representation of a recommended product (e.g., stylistically similar to the real-world object) that is also color compatible with the real-world environment. Thus, in the view of the user, the AR representation matches not only the style of the real-world environment but also the colors of the real-world environment”) comprises: selecting, from a plurality of candidate product information items, the product information whose product three-dimensional model has the highest compatibility with the target three-dimensional model, based on the target three-dimensional model and the product three-dimensional models corresponding to the plurality of candidate product information items, and providing a selected target product information to the fitting subject (See Ayush: Fig. 1, and [0030], “Additionally, the AR product recommendation system can further determine overall scores for the recommended products. To elaborate, the AR product recommendation system can weight a style similarity score and can further weight a color compatibility score. To determine the weights, the AR product recommendation system can utilize a rank support vector machine (“SVM”) algorithm that employs a pair-wise ranking method. Based on the scores and their respective weights, the AR product recommendation system can determine an overall score for each recommended product. In addition, the AR product recommendation system can select a number of products that correspond to top overall scores to provide AR product recommendations to a user client device (e.g., via emails, push notifications, AR environments, etc.).”); or customizing a product three-dimensional model compatible with the target three-dimensional model for the fitting subject based on model parameters corresponding to the target three-dimensional model and a selected product type, and providing the product information corresponding to the customized product three-dimensional model as the target product information to the fitting subject (See Ayush: Figs. 1-2, and [0039], “As mentioned, the AR product recommendation system analyzes the viewpoint to identify a real-world object. As used herein, the term “real-world object” (or sometimes simply “object”) refers to an object that is depicted within the camera feed. In particular, a real-world object can refer to a physical object that exists in the physical world. For example, a real-world object may include, but is not limited to, accessories, animals, clothing, cosmetics, footwear, fixtures, furnishings, furniture, hair, people, physical human features, vehicles, or any other physical object that exists outside of a computer. In some embodiments, a digital image depicts real objects within an AR scene. The AR product recommendation system can identify and analyze a real-world object to identify a style of the real-world object to which the AR product recommendation system can match recommended products”; and [0058], “Although FIG. 2 and subsequent figures illustrate a room with furniture where the AR product recommendation system 102 generates an AR chair to recommend to a user, in some embodiments the AR product recommendation system 102 generates recommended products apart from furniture. Indeed, the AR product recommendation system 102 can analyze a camera feed that depicts any real-world environment such as an outdoor scene, a person wearing a particular style of clothing, or some other scene. Accordingly, the AR product recommendation system 102 can generate recommended products (and AR representations of those products) based on the real-world environment of the camera feed—e.g., to recommend products such as clothing items that are similar to the style of clothing worn by a group of people, accessories that match an outfit worn by an individual, landscaping items that match outdoor scenery of a house, etc.”). Regarding claim 18 , Guler, Faras, Gonzalez, and Ayush teach all the features with respect to claim 16 as outlined above. Further, Guler and Faras teach that the method according to claim 16, wherein inputting the plurality of frames of images into the feature extraction network to extract features, to obtain feature vectors of the plurality of frames of images (See Guler: Figs. 1-3, and [0036], “The 3D reconstruction method relies on a prior model of the target shape. The prior model may be a parametric model of the object being reconstructed. The type of prior model used may depend on the object being reconstructed. For example, in embodiments where the shape is a human body, the human body may be parametrised using a parametric deformable model. In the following, the method will be described with reference to a parametric deformable model, though it will be appreciated that other parametric shape models may alternatively be used. In some embodiments, SMPL is used. SMPL uses two parameter vectors to capture pose and shape: θ comprises 23 3D-rotation matrices corresponding to each joint in a kinematic tree for the human pose, and β captures shape variability across subjects in terms of a 10-dimensional shape vector. SMPL uses these parameters to obtain a triangulated mesh of the human body through linear skinning and blend shapes as a differentiable function of θ, β. In some embodiments, the parametric deformable model may be modified to allow it to be used more efficiently with convolutional neural networks, as described below in relation to FIG. 3”), comprises: for each frame of the plurality of frames of images, inputting the frame into a feature extraction module of the feature extraction network to perform feature extraction and obtain an image feature map for the frame (See Guler: Figs. 1-3, and [0049], “At operation 2.1, a two-dimensional image is received. As described above, the 2D image may comprise a set of pixel values corresponding to a two-dimensional array. In some embodiments, a single image is received. Alternatively, a plurality of images may be received. If multiple images are taken (for example, of a person moving), the three dimensional representation may account for these changes”; and [0045], “A linear layer 112 applied on top of the pooled extracted features is used to estimate parameters 114 of the parametric deformable model. The linear layer may, for example, output weights w.sup.i.sub.k,j 116 for use in determining the parametric deformable model parameters 11”); inputting camera pose data recorded at the time the frame was captured into a camera parameter fusion module of the feature extraction network to perform feature extraction and obtain a camera pose feature map for the frame (See Faras: Figs. 1-2, and [0026], “Applications such as 3D imaging, mapping, and navigation may use Simultaneous Localization and Mapping (SLAM). SLAM relates to constructing or updating a map of an unknown environment while simultaneously keeping track of an object's location within it and/or estimating the pose of the camera with respect to the object or scene. This computational problem is complex since the object may be moving and the environment may be changing. 2D images of real objects and/or 3D object may be captured with the objective of creating a 3D image that is used in real-world applications such as augmented reality, 3D printing, and/or 3D visualization with different perspectives of the real objects. The 3D objects may be characterized by features that are specific locations on the physical object in the 2D images that are of importance for the 3D representation such as corners, edges, center points, or object-specific features on a physical object such as a face that may include nose, ears, eyes, mouth, etc. There are several algorithms used for solving this computational problem associated with 3D imaging, using approximations in tractable time for certain environments. Popular approximate solution methods include the particle filter and Extended Kalman Filter (EKF). The particle filter, also known as a Sequential Monte Carlo (SMC) linearizes probabilistic estimates of data points. The Extended Kalman Filter is used in non-linear state estimation in applications including navigation systems such as Global Positioning Systems (GPS), self-driving cars, unmanned aerial vehicles, autonomous underwater vehicles, planetary rovers, newly emerging domestic robots, medical devices inside the human body, and/or image processing systems. Image processing systems may perform 3D pose estimation using SLAM techniques by performing a transformation of an object in a 2D image to produce a 3D object. However, existing techniques such as SMC and EKF may be insufficient in accurately estimating and positioning various points in a 3D object based on information discerned from 2D objects and may be computationally inefficient in real time”); concatenating the image feature map and the camera pose feature map of each frame using a feature concatenation module of the feature extraction network to obtain a concatenated feature map for each frame (See Guler: Figs. 1-3, and [0044], “Each feature extracted around a 2D keypoint/joint can deliver a separate parametric deformable model parameter estimate. But intuitively a 2D keypoint/joint should have a stronger influence on parametric deformable model parameters that are more relevant to it—for instance a left wrist joint should be affecting the left arm parameters, but not those of kinetically independent parts such as the right arm, head, or feet. Furthermore, the fact that some keypoints/joints can be missing from an image means that we cannot simply concatenate the features in a larger feature vector, but need to use a form that handles gracefully the potential absence of parts”); and performing dimensionality reduction on the concatenated feature map of each frame using a dimensionality reduction module of the feature extraction network to obtain a feature vector for each frame (See Faras: Figs. 1-3, and [0040], “More particularly, a robust and accurate method that can deliver real-time pose estimates and/or a 3D map for 3D reconstruction and provide enough information for camera calibration is described in various embodiments. The inventive concepts described herein combine a non-recursive initialization phase with a recursive sequential updating (tracking phase) system. Initialization of the 3D map or structure may be based on the scene or the scene structure that is discerned from a series of 2D images or frames. Sequential tracking or sequential updating may also be referred to as recursive pose and positioning. During the initialization phase, a non-recursive initialization of the 3D map and the poses is used to localize the camera for 2D frames. An initial map of the scene, which is represented by a set of 3D coordinates corresponding to salient image points that are tracked between sequential frames, is constructed and the camera poses (orientation and position of the camera along its trajectory) are computed. Criteria, such as, for example, the number of tracked points or the pose change, are used to decide if the current frame should become a key-frame. Key frames are selected as representative sets of frames to be used in the localization. If a given frame is selected as a key frame, a local/global bundle adjustment (BA) may be used to refine the key-frames positions and/or to refine or triangulate new 3D points. During this processing a global feature database may be created and populated with globally optimized landmarks. Each landmark may be associated with some stored information such as the related 3D coordinates, a list of frames/key-frames where it was visible, and/or a reference patch. After the initialization phase, a set of anchor landmarks may be available when the sequential updating and/or tracking phase is entered. A fully recursive system, also based on feature tracking, may be used to localize the camera. In particular, the initial set of global features may reduce and/or remove the known drift problem of localization with recursive systems”). Regarding claim 19, Guler, Faras, Gonzalez, and Ayush teach all the features with respect to claim 18 as outlined above. Further, Guler and Gonzalez teach that the method according to claim 18, wherein for each frame of the plurality of frames of images, inputting the frame into the feature extraction module of the feature extraction network to perform feature extraction and obtain an image feature map for the frame, comprises: for each frame of the plurality of frames of images, inputting the frame into a skip connection layer of the feature extraction module to perform multi-resolution feature map extractions and skip connections of feature maps with the same resolution, to obtain a second intermediate feature map for the frame (See Gonzalez: Fig. 1, and [0025], “The receiver, validator, and scaler 110 can perform preprocessing on the received image 104, 108. For example, the receiver, validator, and scaler 110 can perform detection of image compression, content validation, and spatial scaling. In some examples, the receiver, validator, and scaler 110 can detect a region of interest (ROI) 112 in a received image and a compression ratio to automatically adjust parameter boundaries for subsequent adaptive algorithms. The region of interest 112 may be a rectangular region of an image representing an object to be modeled. In some examples, the object to be modeled can be determined based on pixel color. For example, the region of interest may be a rectangular region including pixels that have a different color than a background color. In some examples, at least one of the pixels of the object to be modeled may touch each edge of the edge of the region of interest. In some examples, the receiver, validator, and scaler 110 can validate the region of interest. For example, the receiver, validator, and scaler 110 can compare a detected region of interest with the corresponding image including the region of interest. In some examples, if the region of interest is the same as the corresponding image, then the image may be rejected. In some examples, if the region of interest is not the same as the corresponding image, then the region of interest may be further processed as described below. The receiver, validator, and scaler 110 can then add a spatial-convolutional margin to the ROI 112. For example, the spatial-convolutional margin can be a margin of pixels that can be used during filtering of the region of interest. For example, the spatial-convolutional margin may be a margin of 5-7 pixels that can be used by the filter to prevent damage to content in the region of interest. The receiver, validator, and scaler 110 may then scale the ROI 112 to increase the resolution of the region of interest. For example, the receiver, validator, and scaler 110 can scale content of the ROI 112 by a factor of 2× using bi-cubic interpolation to increase saliency representative and discrete-contour precision. The scaling may be used to improve the quality of the filtering process”); and inputting the second intermediate feature map of the frame into a downsampling layer of the feature extraction module to perform M downsampling operations, where M is a positive integer greater than or equal to 1, to obtain the image feature map for the frame (See Guler: Figs. 8A-B, and [0097], “In some embodiments, the neural networks used in the methods described above comprise an ImageNet pre-trained ResNet-50 network as a system backbone. Each dense prediction task may use a deconvolutional head, following the architectural choices of “Simple baselines for human pose estimation and tracking” (B. Xiao et al., arXiv:1804.06208, 2018), with threes 4×4 deconvolutional layers applied with batch-norm and ReLU. These are followed by a linear layer to obtain outputs of the desired dimensionality. Each deconvolution layer has a stride of two, leading to an output resolution of 64×64 given 256×256 images as input”. Note that the input image is progressively downsampling by a factor of 2 till the desired resolution, the factor 2 is mapped to M >=1). Allowable Subject Matter Claim 5 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The most relevant arts searched, Guler, etc. (US 20210241522 A1), Faras, etc. (US 20210118160 A1), and Gonzalez, etc. (US 20180211438 A1), do not teach that claimed limitation of “the method according to claim 4, wherein the encoder comprises an encoding submodule and N downsampling submodules connected in sequence, and inputting the frame into the encoder of the skip connection layer to encode the frame to obtain an initial feature map for the frame, and sequentially performing N downsampling operations on the initial feature map to obtain a first intermediate feature map, comprises: inputting the frame into the encoding submodule to perform encoding and obtain the initial feature map for the frame; performing N downsampling operations on the initial feature map using the N downsampling submodules to obtain the first intermediate feature map; wherein, in each downsampling submodule: performing convolution operations on the input using K1 convolution units, each corresponding to a target convolution parameter, to generate an intermediate feature map to be activated, where Kl is a positive integer greater than or equal to 2; and activating the intermediate feature map to be activated using an activation function to generate an output of each convolution unit.” Claim 6 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The most relevant arts searched, Guler, etc. (US 20210241522 A1), Faras, etc. (US 20210118160 A1), and Gonzalez, etc. (US 20180211438 A1), do not teach that claimed limitation of “the method according to claim 3, wherein the downsampling layer comprises M downsampling submodules connected in sequence, and inputting the second intermediate feature map of the frame into the downsampling layer of the feature extraction module to perform M downsampling operations to obtain the image feature map for the frame, comprises: performing M downsampling operations on the second intermediate feature map using the M downsampling submodules to obtain the image feature map for the frame; wherein, in each downsampling submodule, performing convolution operations on the input using K2 convolution units connected in sequence, each corresponding to a target convolution parameter, to generate an intermediate feature map to be activated, where K2 is a positive integer greater than or equal to 2; and activating the intermediate feature map to be activated using an activation function to generate the output of each convolution unit.” Claims 7-9 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The most relevant arts searched, Guler, etc. (US 20210241522 A1), Faras, etc. (US 20210118160 A1), and Gonzalez, etc. (US 20180211438 A1), do not teach that claimed limitation of “the method according to claim 2, wherein inputting camera pose data recorded at the time the frame was captured into the camera parameter fusion module of the feature extraction network to perform feature extraction and obtain a camera pose feature map for the frame, comprises: inputting the camera pose data recorded at the time the frame was captured into the camera parameter fusion module of the feature extraction network, wherein the camera pose data includes at least two types of pose angles; performing trigonometric processing based on at least two pose angles of the at least two types of pose angles and relationships between the at least two pose angles to obtain a plurality of pose representation parameters; and processing the plurality of pose representation parameters using a multilayer perceptron (MLP) network in the camera parameter fusion module to obtain the camera pose feature map for the frame.” Claim 11 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The most relevant arts searched, Guler, etc. (US 20210241522 A1), Faras, etc. (US 20210118160 A1), and Gonzalez, etc. (US 20180211438 A1), do not teach that claimed limitation of “the method according to claim 1, wherein the plurality of frames of images include a current frame image and at least one historical frame image; inputting the plurality of frames of images into the feature extraction network to extract features to obtain feature vectors of the plurality of frames of images, comprises: each time, inputting a current frame image into the feature extraction network to extract features and obtain a feature vector for the current frame image; and concatenating the feature vectors of the plurality of frames of images to obtain a target concatenated feature vector, comprises: using a predetermined sliding window to retrieve a feature vector of at least one historical frame image from a specified storage space; and concatenating the feature vector of the current frame image and the feature vector of at least one historical frame image to obtain the target concatenated feature vector.” Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to GORDON G LIU whose telephone number is (571)270-0382. The examiner can normally be reached Monday - Friday 8:00-5:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Devona E Faulk can be reached at 571-272-7515. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /GORDON G LIU/Primary Examiner, Art Unit 2618
Read full office action

Prosecution Timeline

Jan 22, 2025
Application Filed
Jul 23, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705808
AI-BASED VISUAL CONTENT COLLAGE GENERATION
2y 5m to grant Granted Aug 11, 2026
Patent 12705751
MAP SCENE RENDERING METHOD AND APPARATUS, SERVER, TERMINAL, COMPUTER-READABLE STORAGE MEDIUM, AND COMPUTER PROGRAM PRODUCT
2y 4m to grant Granted Aug 11, 2026
Patent 12705804
AUGMENTED REALITY BASED AIMING OF LIGHT FIXTURES
2y 4m to grant Granted Aug 11, 2026
Patent 12700152
AI-BASED AVATAR CREATION USING STYLE TRANSFER AND SUBJECT IMAGES
2y 5m to grant Granted Aug 04, 2026
Patent 12694614
IMAGING DEVICE, AND CONTROL METHOD OF IMAGING DEVICE
2y 1m to grant Granted Jul 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
83%
Grant Probability
98%
With Interview (+15.0%)
2y 2m (~7m remaining)
Median Time to Grant
Low
PTA Risk
Based on 692 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month