Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claim8 and 17 objected to because of the following informalities: Claims 8 and 17 list "the first set of morphable model coefficients based on (i) a set of features associated with the first image and (iii) the one or more position encodings", but no (ii) is referenced. (iii) should be changed to (ii) in both claims, or another item labeled (ii) should be introduced.. Appropriate correction is required.
Claim 18 objected to because of the following informalities: Claim 18 states “to perform the step of synthesizing”, there is no "step of synthesizing" introduced prior to this limitation. Limitation should be corrected to "to perform a step of synthesizing" (emphasis added). Appropriate correction is required.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1, 10-11, 14, and 17-18 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Ferman et al (US 12387419 B2, hereinafter Ferman).
Regarding claim 1, Ferman teaches A computer-implemented method for performing landmark detection, the method comprising (Col 10 Line 39-43 “The methods described herein may be performed using any suitable computing apparatus. For example, computing system 600 includes a power supply 602, one or more processors 604, memory 606, and input and output devices 608”):
generating, via execution of a first machine learning model, a first set of morphable model coefficients associated with a first object depicted in a first image (Col 4 Line 30-34 “A 3D landmark regressor may be a machine-learned model arranged to process an image (or a sequence of images) depicting an object to determine locations in a three-dimensional coordinate system of a set of landmarks associated with the object.”, Col 6 Line 57-62 “The 3D landmark regressor 118 may be configured to localize a predetermined set of 3D landmarks, for example by determining offsets to a 3D landmark template. The 3D landmark regressor 118 may be configured to process individual images or sequences of images depicting an object.”);
determining one or more three-dimensional (3D) landmarks on the first object based on the first set of morphable model coefficients (Col 7 Line 39-43 “The candidate 3D landmarks may be predicted as offsets to a template, which may be defined as the landmark-wise mean of the 3D pseudo-labels 116 obtained during 3D landmark optimization”);
projecting the one or more 3D landmarks onto the first image to generate one or more two-dimensional (2D) landmarks (Col 9 Line 18-22 “The output data 404 is used for projecting and masking 406, in which an unoccluded subset of the set of candidate 3D landmarks is identified and at least the unoccluded subset is projected to 2D using the candidate camera pose”, Col 9 Line 43-45 “The result of the projecting and masking 406 is a set of masked candidate 2D landmarks 408.”);
and training the first machine learning model based on one or more losses associated with the one or more 2D landmarks to generate a first trained machine learning model (Col 9 Line 61 – Col 10 Line 2 “Accordingly, the loss function 124 may include a term which evaluates a deviation, error, or difference between the set of masked candidate 2D landmarks 408 (corresponding to an unoccluded subset of the candidate 3D landmarks), and a corresponding subset of the 2D pseudo-labels 410. By augmenting the loss function 124 with such a term, the 3D landmark regressor 118 may be jointly trained on rendered views of synthetic 3D objects and images of real-world objects”).
Regarding claim 10, Ferman teaches the computer-implemented method of claim 1, and further teaches wherein the first object comprises at least one of a face, a body, or a body part (Col 11 Line 55-58 “some or all of the disclosed techniques may equally be applied to objects other than human faces, for example to animal faces, entire human or animal bodies, vehicles, etc.”)
Regarding claim 11, the non-transitory computer readable media (Ferman claim 1 “one or more non-transitory computer-readable media”) claim 11 is similar in scope to the computer-implemented method claim 1, and is rejected under similar rationale.
Regarding claim 14, Ferman teaches the one or more non-transitory computer-readable media of claim 11, and further teaches wherein determining the one or more 3D landmarks comprises: determining, via execution of a second machine learning model, a pose of the first object (Col 7 Line 36-39 “The landmark tokens and the pose tokens are respectively routed to landmark MLP heads 308 and pose MLP heads 310 to predict 3D landmarks 312 and a 3D pose 314”, where MLP is a multi-layer perceptron, which is a machine learning model.); and generating the one or more 3D landmarks based on an evaluation of a morphable model using the first set of morphable model coefficients and the pose of the first object (Col 7 Line 36-48 “The landmark tokens and the pose tokens are respectively routed to landmark MLP heads 308 and pose MLP heads 310 to predict 3D landmarks 312 and a 3D pose 314. The candidate 3D landmarks may be predicted as offsets to a template, which may be defined as the landmark-wise mean of the 3D pseudo-labels 116 obtained during 3D landmark optimization. The pose may be predicted using any suitable representation, for example via a 6D rotation representation and 3D translation vector from which the camera extrinsic matrix can be computed. The pose 310 may optionally be used to project the 3D landmarks 312 into the screen space of the image 302 to generate projected landmarks 316”).
Regarding claim 17, the non-transitory computer readable media claim 17 is similar in scope to the computer-implemented method claim 8, and is rejected under similar rationale.
Regarding claim 18, Ferman teaches the one or more non-transitory computer-readable media of claim 11, and further teaches wherein the instructions further cause the one or more processors to perform the step of synthesizing the first image based on at least one of a reconstructed facial geometry, a facial texture associated with the reconstructed facial geometry, an artist-created asset, an environment map, or a set of camera parameters (Col 4 Line 59-62 “The system 100 stores data defining a generative model 102 arranged to render a set of 2D views 104 of a synthetic 3D object in dependence on conditioning data 106 and a set of camera poses 108”, Col 5 Line 22-26 “the conditioning data 106 may determine the facial identity, the facial expression, lighting levels, other visual effects, and/or other characteristics of the 3D object that are independent of the camera pose.”, where camera poses correspond to camera parameters.)
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 2-3 are rejected under 35 U.S.C. 103 as being unpatentable over Ferman as applied to claim 1 above, and further in view of Revaud et al (US 20240404104 A1, hereinafter Revaud).
Regarding claim 2, Ferman teaches the computer-implemented method of claim 1, and further teaches further comprising: determining, via execution of a second machine learning model, a first set of parameters used to determine at least one of the one or more 3D landmarks or the one or more 2D landmarks (Col 7 Line 35-39 “a multi-layer perceptron (MLP), with (optionally) layer-normalization applied prior to each layer. The landmark tokens and the pose tokens are respectively routed to landmark MLP heads 308 and pose MLP heads 310 to predict 3D landmarks 312 and a 3D pose 314.”, Col 8 Line 46-48 “The pose 310 may optionally be used to project the 3D landmarks 312 into the screen space of the image 302 to generate projected landmarks 316.”), but fails to explicitly teach training the second machine learning model based on the one or more losses.
In related field of endeavor, Revaud teaches training a multi-layer perceptron based on one or more losses ([0155] “A vision transformer 1306 is used to encode all query images 110 and database images or reference images 120. In more details, each image is divided into non-overlapping patches, and a linear projection encodes them into patch features. A series of transformer blocks is then applied on these features: each block consists of multi-head self-attention and a multilayer perceptron (MLP) decoder”, [0161] “The model 1301 is trained using two complementary regression losses. The first regression loss is a 3D regression loss. The second regression loss is a 2D reprojection loss”, where multilayer perceptrons are components of the model 1301)
It would have been obvious to one of ordinary skill in the art prior to the time of filing to have modified Ferman to include training a multi-layer perceptron based on one or more losses as taught by Revaud. Doing so would provide a model tuned for PnP and regularized to prevent under-confidence ([0165] “By default, the 2D reprojection loss is used for the confidence, since it makes more sense for PnP (Perspective-n-Point). The second term of the loss acts as a regularizer, so as to prevent the model from getting under-confident everywhere.”)
Regarding claim 3, Ferman as modified by Revaud teaches the computer-implemented method of claim 2. Ferman further teaches wherein the first set of parameters comprises at least one of a head pose or a camera parameter (Col 7 Line 37-39 “landmark tokens and the pose tokens are respectively routed to landmark MLP heads 308 and pose MLP heads 310 to predict 3D landmarks 312 and a 3D pose 314”, Col 8 Line 5-8 “the 3D landmark regressor 118 may determine a candidate camera pose (not shown) along with the candidate 3D landmarks 120”).
Claims 4-5, 12-13 are rejected under 35 U.S.C. 103 as being unpatentable over Ferman as applied to claim 1 and 11 above, and further in view of Sommerlade et al (US 11657557 B2, hereinafter Sommerlade) and Revaud.
Regarding claim 4, Ferman teaches the computer-implemented method of claim 1, and further teaches determining, via execution of a second machine learning model, one or more additional 2D landmarks on a second object depicted in a second image (Claim 1 “processing, using a two-dimensional landmark regressor, the plurality of two-dimensional views to generate respective sets of two-dimensional landmarks”, Claim 5 “processing, using the two-dimensional landmark regressor, a first image depicting a second object to determine a first set of two-dimensional landmarks”, where multiple images and multiple objects may be processed); but fails to explicitly teach generating the first image from a region within the second image that includes a subset of the one or more additional 2D landmarks corresponding to the first object; and training the second machine learning model based on the one or more losses.
In related field of endeavor, Sommerlade teaches a second object depicted in a second image and generating the first image from a region within the second image that includes a subset of the one or more additional 2D landmarks corresponding to the first object (Col 9 Line 5-8 “the processing comprises step S15A in which an image is segmented into a set of semantically meaningful regions, a non-exhaustive list of which are: face, body, hair, garments, background”, Col 11 Line 35-36 “The image is processed in step S42 to extract a face region of interest”, where face region corresponds to first region);
It would have been obvious to one of ordinary skill in the art prior to the time of filing to have modified Ferman to include a second object depicted in a second image; generating the first image from a region within the second image that includes a subset of the one or more additional 2D landmarks corresponding to the first object as taught by Sommerlade. Doing so would provide an extracted region of interest for further processing (Col 11 Line 33-35 “Thus, in step S41 an image is provided. The image is processed in step S42 to extract a face region of interest”)
Ferman as modified by Sommerlade fails to explicitly teach training the second machine learning model based on the one or more losses. In related field of endeavor, Revaud teaches training the second machine learning model based on the one or more losses ([0161] “The model 1301 is trained using two complementary regression losses. The first regression loss is a 3D regression loss. The second regression loss is a 2D reprojection loss”)
It would have been obvious to one of ordinary skill in the art prior to the time of filing to have further modified Ferman and Sommerlade to include training the second machine learning model based on the one or more losses as taught by Revaud. Doing so would provide a model tuned for PnP and regularized to prevent under-confidence ([0165] “By default, the 2D reprojection loss is used for the confidence, since it makes more sense for PnP (Perspective-n-Point). The second term of the loss acts as a regularizer, so as to prevent the model from getting under-confident everywhere.”)
Regarding claim 5, Ferman as modified by Sommerlade and Revaud teach the computer-implemented method of claim 4. Ferman further teaches the first object comprises a body part included in the body (Col 11 Line 55-58 “For example, some or all of the disclosed techniques may equally be applied to objects other than human faces, for example to animal faces, entire human or animal bodies”, where human face is a body part included in an entire human body), but fails to explicitly teach the second object comprises a body.
Sommerlade further teaches the second object comprises a body (Col 9 Line 5-8 “the processing comprises step S15A in which an image is segmented into a set of semantically meaningful regions, a non-exhaustive list of which are: face, body, hair, garments, background”)
It would have been obvious to one of ordinary skill in the art prior to the time of filing to have further modified Ferman in view of Sommerlade and Revaud to include the second object comprises a body as taught by Sommerlade. Doing so would allow for modeling of both face and body geometry (Col 9 Line 33-37 “modelling of face and/or body geometry may be performed (step S15B). In some embodiments, a parametric model may be used to infer information about the underlying 3D geometry of face and body of the user”)
Regarding claim 12, Ferman teaches the one or more non-transitory computer-readable media of claim 11, and further teaches wherein the instructions further cause the one or more processors to perform the steps of: determining, via execution of a second machine learning model, one or more additional 2D landmarks on a second object depicted in a second image (Claim 1 “processing, using a two-dimensional landmark regressor, the plurality of two-dimensional views to generate respective sets of two-dimensional landmarks”, Claim 5 “processing, using the two-dimensional landmark regressor, a first image depicting a second object to determine a first set of two-dimensional landmarks”, where multiple images and multiple objects may be processed); but fails to explicitly teach generating the first image based on a bounding box within the second image that includes a subset of the one or more additional 2D landmarks corresponding to the first object; and training the second machine learning model based on the one or more losses. In related field of endeavor, Sommerlade teaches generating the first image 15A in which an image is segmented into a set of semantically meaningful regions, a non-exhaustive list of which are: face, body, hair, garments, background. This can be achieved by standard object segmentation methods”);
It would have been obvious to one of ordinary skill in the art prior to the time of filing to have modified Ferman to include generating the first image
Ferman as modified by Sommerlade fails to explicitly teach segmentation based on a bounding box and training the second machine learning model based on the one or more losses. In related field of endeavor, Revaud teaches segmentation based on a bounding box ([0149] “Gen6D uses detection and retrieval to initialize the pose of a query image and then refines it by regressing the pose residual. However, it requires an accurately detected 2D bounding box for pose initialization”, [0151] “The method only uses the 3D bounding box of objects in the reference images to estimate 2D-3D correspondence of query image”, [0176] “the mesh is used to compute a 3D bounding box aligned with the 3 axes. The scale of the shape proxy is set accordingly to the 3D bounding box. The generated 3D proxy shape is then transformed using the object pose”) and training the second machine learning model based on the one or more losses ([0161] “The model 1301 is trained using two complementary regression losses. The first regression loss is a 3D regression loss. The second regression loss is a 2D reprojection loss”)
It would have been obvious to one of ordinary skill in the art prior to the time of filing to have further modified Ferman and Sommerlade to include a bounding box and training the second machine learning model based on the one or more losses as taught by Revaud. Doing so would provide a segmentation method and a model tuned for PnP and regularized to prevent under-confidence ([0165] “By default, the 2D reprojection loss is used for the confidence, since it makes more sense for PnP (Perspective-n-Point). The second term of the loss acts as a regularizer, so as to prevent the model from getting under-confident everywhere.”)
Regarding claim 13, Ferman as modified by Sommerlade and Revaud teach the one or more non-transitory computer-readable media of claim 12, and Ferman further teaches the first object comprises a face on the body (Col 11 Line 55-58 “For example, some or all of the disclosed techniques may equally be applied to objects other than human faces, for example to animal faces, entire human or animal bodies”, where human face is a body part included in an entire human body), but fails to explicitly teach wherein the second object comprises a body. Sommerlade further teaches the second object comprises a body (Col 9 Line 5-8 “the processing comprises step S15A in which an image is segmented into a set of semantically meaningful regions, a non-exhaustive list of which are: face, body, hair, garments, background”)
It would have been obvious to one of ordinary skill in the art prior to the time of filing to have further modified Ferman in view of Sommerlade and Revaud to include the second object comprises a body as taught by Sommerlade. Doing so would allow for modeling of both face and body geometry (Col 9 Line 33-37 “modelling of face and/or body geometry may be performed (step S15B). In some embodiments, a parametric model may be used to infer information about the underlying 3D geometry of face and body of the user”)
Claims 6-7, 15-16, 20 are rejected under 35 U.S.C. 103 as being unpatentable over Ferman as applied to claim 1 and 11 above, and further in view of Chen et al (US 10796480 B2, hereinafter Chen).
Regarding claim 6, Ferman teaches the computer-implemented method of claim 1, and further teaches further comprising: generating, via execution of the first trained machine learning model, a second set of morphable model coefficients associated with a second object depicted in a second image (Col 4 Line 30-34 “A 3D landmark regressor may be a machine-learned model arranged to process an image (or a sequence of images) depicting an object to determine locations in a three-dimensional coordinate system of a set of landmarks associated with the object.”, Col 6 Line 57-62 “The 3D landmark regressor 118 may be configured to localize a predetermined set of 3D landmarks, for example by determining offsets to a 3D landmark template. The 3D landmark regressor 118 may be configured to process individual images or sequences of images depicting an object.”), but fails to explicitly teach generating a 3D shape associated with the second object based on the second set of morphable model coefficients.
In related field of endeavor, Chen teaches generating a 3D shape associated with the second object based on the second set of morphable model coefficients (Col 14 Line 2-8 “In the first stage, we find an approximate head geometry as an initialisation using a generative shape prior that models shape variation of an object category (i.e. the face) in the low dimensional subspace with a dimension reduction method. The full head geometry of the user can be reconstructed from this low dimensional shape prior with a small number of parameters.”, Col 14 Line 36-37 “We select the appropriate 3D shape prior from the library based on matching a user's attributes”).
It would have been obvious to one of ordinary skill in the art prior to the time of filing to have modified Ferman to include generating a 3D shape associated with the second object based on the second set of morphable model coefficients as taught by Chen. Doing so would recover missing depth information of a user’s face more accurately (Col 14 Line 38-41 “This method can recover the missing depth information of the user's face more accurately. It is useful for building a product that will work across different ethnic regions”)
Regarding claim 7, Ferman as modified by Chen teaches the computer-implemented method of claim 6. Chen further teaches wherein generating the 3D shape comprises at least one of: generating an animation associated with the second object based on the 3D shape (Col 42 Line 60-62 “Rendering and animation module, renders the animation (e.g. animated GIF) from the 3D face model(s) given the specified background image and head pose sequences.”); applying the second set of morphable model coefficients to a morphable model of a third object (Col 42 Line 34-36 “face transfers: i.e. transfer the face appearance from one to the other, or merges the face appearance of two or more users by e.g. averaging, blending, and morphing.”, where transferring the face appearance results in generating a 3D shape associated with a third object); or editing the 3D shape based on the second set of morphable model coefficients (Col 14 Line 6-8 “The full head geometry of the user can be reconstructed from this low dimensional shape prior with a small number of parameters.”).
It would have been obvious to one of ordinary skill in the art prior to the time of filing to have modified Ferman to include generating an animation associated with the second object based on the 3D shape, applying the second set of morphable model coefficients to a morphable model of a third object, or editing the 3D shape based on the second set of morphable model coefficients as taught by Chen. Doing so would provide personalization applications to allow users to create and share personalized 3D models (Col 42 Line 10-16 “In this section, we describe other examples of personalisation applications which derived from the personalised 3D face/head reconstruction techniques as described in Section 2. By integrating them with commercial social network websites and/or messenger applications on the mobile platforms, it allows users to create, visualize, and share their personalised 3D models conveniently.”), and recover missing depth information of a user’s face more accurately (Col 14 Line 38-41 “This method can recover the missing depth information of the user's face more accurately. It is useful for building a product that will work across different ethnic regions”)
Regarding claim 15, Ferman teaches the one or more non-transitory computer-readable media of claim 11, and further teaches wherein the instructions further cause the one or more processors to perform the steps of: generating, via execution of the first trained machine learning model, a second set of morphable model coefficients associated with a second object depicted in a second image (Col 4 Line 30-34 “A 3D landmark regressor may be a machine-learned model arranged to process an image (or a sequence of images) depicting an object to determine locations in a three-dimensional coordinate system of a set of landmarks associated with the object.”, Col 6 Line 57-62 “The 3D landmark regressor 118 may be configured to localize a predetermined set of 3D landmarks, for example by determining offsets to a 3D landmark template. The 3D landmark regressor 118 may be configured to process individual images or sequences of images depicting an object.”), but fails to explicitly teach generating a 3D shape associated with a third object based on the second set of morphable model coefficients. In related field of endeavor, Chen teaches generating a 3D shape associated with a third object based on the second set of morphable model coefficients (Col 14 Line 2-8 “In the first stage, we find an approximate head geometry as an initialisation using a generative shape prior that models shape variation of an object category (i.e. the face) in the low dimensional subspace with a dimension reduction method. The full head geometry of the user can be reconstructed from this low dimensional shape prior with a small number of parameters.”, Col 42 Line 34-36 “face transfers: i.e. transfer the face appearance from one to the other, or merges the face appearance of two or more users by e.g. averaging, blending, and morphing.”, where transferring the face appearance results in generating a 3D shape associated with a third object)
It would have been obvious to one of ordinary skill in the art prior to the time of filing to have modified Ferman to include generating a 3D shape associated with a third object based on the second set of morphable model coefficients as taught by Chen. Doing so would provide personalization applications to allow users to create and share personalized 3D models (Col 42 Line 10-16 “In this section, we describe other examples of personalisation applications which derived from the personalised 3D face/head reconstruction techniques as described in Section 2. By integrating them with commercial social network websites and/or messenger applications on the mobile platforms, it allows users to create, visualize, and share their personalised 3D models conveniently.”)
Regarding claim 16, Ferman as modified by Chen teaches the one or more non-transitory computer-readable media of claim 15. Ferman further teaches an identity coefficient or an expression coefficient (Col 5 Line 18-24 “The conditioning data 106 may include one or more of text data, image data, random noise, and/or any other data which can be provided to the generative model 102 to affect the properties of a sample (3D object) drawn from the generative model 102. In the case of human faces, the conditioning data 106 may determine the facial identity, the facial expression”), but fails to explicitly teach wherein the second set of morphable model coefficients comprise at least one of an identity coefficient or an expression coefficient. In related field of endeavor, Chen teaches wherein the second set of morphable model coefficients comprise at least one of an identity coefficient or an expression coefficient (Col 9 Line 15-20 “we use a machine learning attribute classifier to analyze the user's photo and predict his/her attributes (e.g. ethnics and gender) from the image, and then select the appropriate 3D shape prior from the library to recover the missing depth information of the user's face more accurately”).
It would have been obvious to one of ordinary skill in the art prior to the time of filing to have further modified Ferman and Chen to include wherein the second set of morphable model coefficients comprise at least one of an identity coefficient or an expression coefficient as taught by Chen. Doing so would help to select an appropriate 3D shape prior (Col 9 Line 18-20 “then select the appropriate 3D shape prior from the library to recover the missing depth information of the user's face more accurately”).
Regarding claim 20, Ferman teaches A system, comprising: one or more memories that store instructions, and one or more processors that are coupled to the one or more memories and, when executing the instructions, are configured to perform the steps of (Claim 1 “A system comprising: one or more processors; and one or more non-transitory computer-readable media storing instructions which, when executed by the one or more processors, cause the one or more processors to perform operations comprising:”):
determining a first machine learning model, wherein the first machine learning model is trained based on one or more losses associated with one or more landmarks corresponding to one or more sets of morphable model coefficients generated by the first machine learning model (Claim 1 “updating the three-dimensional landmark regressor based at least in part on a loss function comprising a term that evaluates a deviation between the candidate set of three-dimensional landmarks and the fitted set of three-dimensional landmarks”, Col 6 Line 57-60 “The 3D landmark regressor 118 may be configured to localize a predetermined set of 3D landmarks, for example by determining offsets to a 3D landmark template”) from a set of training images (Claim 1 “rendering a plurality of two-dimensional views of a three-dimensional object generated by a generative model”, where the plurality of two-dimensional views corresponds to a set of training images);
generating, via execution of the first machine learning model, a second set of morphable model coefficients associated with an object depicted in an image (Claim 5 “The system of claim 1, wherein the object is a first object, the candidate set of three-dimensional landmarks is a first candidate set of three-dimensional landmarks, and the operations comprise: processing, using the two-dimensional landmark regressor, a first image depicting a second object to determine a first set of two-dimensional landmarks; processing, using the three-dimensional landmark regressor, at least the first image to determine a second candidate set of three-dimensional landmarks and a candidate camera pose relative to the second object;”); but fails to explicitly teach generating a 3D shape associated with the object based on the second set of morphable model coefficients.
In related field of endeavor, Chen teaches generating a 3D shape associated with the object based on the second set of morphable model coefficients (Col 14 Line 2-8 “In the first stage, we find an approximate head geometry as an initialisation using a generative shape prior that models shape variation of an object category (i.e. the face) in the low dimensional subspace with a dimension reduction method. The full head geometry of the user can be reconstructed from this low dimensional shape prior with a small number of parameters.”, Col 14 Line 36-37 “We select the appropriate 3D shape prior from the library based on matching a user's attributes”).
It would have been obvious to one of ordinary skill in the art prior to the time of filing to have modified Ferman to include generating a 3D shape associated with the object based on the second set of morphable model coefficients as taught by Chen. Doing so would recover missing depth information of a user’s face more accurately (Col 14 Line 38-41 “This method can recover the missing depth information of the user's face more accurately. It is useful for building a product that will work across different ethnic regions”)
Claims 8-9 are rejected under 35 U.S.C. 103 as being unpatentable over Ferman as applied to claim 1 above, and further in view of Bradley et al (US 12198225 B2, hereinafter Bradley).
Regarding claim 8, Ferman teaches the computer-implemented method of claim 1, and further teaches wherein generating the first set of morphable model coefficients comprises: converting one or more points on a canonical shape positions can be sampled or defined, and each shape in the domain is represented as a set of offsets from a corresponding set of positions on the canonical shape”); and generating, via execution of the first machine learning model, the first set of morphable model coefficients based on (i) a set of features associated with the first image and (iii) the one or more position encodings (Col 4 Line 34-37 “The 3D landmark regressor may be capable of localizing landmarks for one or more classes of object, and the number and definition of landmarks may be predetermined for a given class of object” where class of object corresponds to features associated with the first image, Col 6 Line 57-60 “The 3D landmark regressor 118 may be configured to localize a predetermined set of 3D landmarks, for example by determining offsets to a 3D landmark template”).
Ferman fails to explicitly teach converting one or more points on a canonical shape into one or more position encodings, but in related field of endeavor Bradley teaches converting one or more points on a canonical shape into one or more position encodings (Col 6 Line 16-19 “Position MLP 306 in encoder 204 converts a set of canonical shape positions 232 in canonical shape 220 into a corresponding set of position tokens”)
It would have been obvious to one of ordinary skill in the art prior to the time of filing to have modified Ferman to include converting one or more points on a canonical shape into one or more position encodings as taught by Bradley. Doing so would allow distribution of representations of shape positions in a way that optimizes for shape modeling (Col 6 Line 39-44 “This conversion of canonical shape positions 232 and offsets 228 into position tokens 336 and offset tokens 338, respectively, allows transformer 200 to distribute representations of canonical shape positions 232 and offsets 228 in a way that optimizes for the shape modeling task.”)
Regarding claim 9, Ferman as modified by Bradley teaches the computer-implemented method of claim 8, and Ferman further teaches wherein the first set of morphable model coefficients comprises at least one coefficient for each point included in the one or more points (Col 6 Line 57-60 “The 3D landmark regressor 118 may be configured to localize a predetermined set of 3D landmarks, for example by determining offsets to a 3D landmark template”, where each landmark has an offset).
Claim 19 is rejected under 35 U.S.C. 103 as being unpatentable over Ferman as applied to claim 11 above, and further in view of Marks et al (US 11127164 B2, hereinafter Marks).
Regarding claim 19, Ferman teaches the one or more non-transitory computer-readable media of claim 11, and further teaches wherein the losses comprise a Laplacian log likelihood (Col 8 Line 29-32 “the loss function term may correspond to a Laplacian Log Likelihood (LLL) objective parametrized by a predicted Cholesky factorization of landmark covariances”) but fails to explicitly teach wherein the one or more losses comprise a Gaussian negative likelihood loss. In related field of endeavor, Marks teaches wherein the one or more losses comprise a Gaussian negative likelihood loss (Col 11 Line 15-18 “the method of training the neural network is an Uncertainty with Gaussian Log-Likelihood Loss (UGLLI) method, which is a neural network-based method”).
It would have been obvious to one of ordinary skill in the art prior to the time of filing to have modified Ferman to include wherein the one or more losses comprise a Gaussian negative likelihood loss as taught by Marks. Doing so would provide a model which yields more accurate results (Col 11 Line 19-20 “The UGLLI method yields more accurate results”)
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Lee et al (US 11830132 B2) teaches extracting 2D feature points from a 2D image of a face and deriving a set of 3D feature points from the 2D set and a standard face model, generating a 3D face model from the 3D set, and projecting the 3D face model onto the two-dimensional image.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JOHN PATRICK GOCO whose telephone number is (571)272-5872. The examiner can normally be reached M-Th, 7:00 am - 5:00 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kee Tung can be reached at (571)272-7794. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/J.P.G./ Examiner, Art Unit 2611
/KEE M TUNG/ Supervisory Patent Examiner, Art Unit 2611