DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) filed on January 20, 2026 has been considered by the examiner.
Response to Amendment
Applicants Amendments filed on February 20, 2026, has been entered and made of record.
Currently pending Claim(s): 1-20
Independent Claim(s): 1, 10, 16
Amended Claim(s): 1, 2, 6, 10, 11, 15, and 19
This office action is responsive to the Applicant’s Arguments/Remarks Made in an Amendment
received on February 20, 2026.
Specification Objections
In view of Applicant’s amendments to the Specification, the objections are withdrawn.
Claim Objections
In view of Applicant’s amendments to Claim 6, the objection is withdrawn
Response to Arguments
In view of amendments filed on February 20, 2026, the Applicant has amended independent Claim 1 to recite the additional imitation of “wherein the set of points corresponds to a user-defined layout of a set of landmarks that is selectable at runtime”. Originally, (in the claim set dated November 8, 2023) Claim 1 recited, “A computer-implemented method, comprising: receiving an input image including one or more facial representations and a set of points associated with a 3D canonical shape, wherein the set of points are selectable at runtime; extracting a set of features from the input image that represent at least one facial representation included in the one or more facial representations; and determining a set of landmarks on the at least one facial representation based on the set of features and the set of points, wherein each landmark in the set of landmarks is associated with at least one point in the set of points”, and was rejected over Cheul (KR Pub No. 2017/0006219) in view of Varanasi (US Pub No 2017/0278302). Claims 8 and 17 were rejected over Cheul in view of Varanasi and further in view of Tuzel (US Pub No 2017/0083751).
As discussed in the following paragraphs, the combination of Cheul and Varanasi does render obvious the new limitation of a set of user-defined layout of set of landmarks. However, the Applicant’s amendment necessitated the new grounds of rejection presented in this Office Action. Upon conducting a new search, the Examiner argues that the newly amended Claims 1, 10, and 16 are unpatentable over Liu et al. (US Pub No 2022/0292773) .
In view of Applicant Arguments/Remarks filed on February 20, 2026, with respect to the claims the Applicant explained (on page 10 of Applicant Arguments/Remarks) that the vertex adjustment signal taught by Cheul is not equivalent to a ‘user defined layout’ of landmarks because Cheul, “does not teach or suggest receiving any user input for specifying the landmark points”. The Examiner agrees. Although Cheul teaches that landmark points may be adjusted by a user, the initial landmark positions are determined through an algorithm which determines the best template landmark positions with respect to a facial image (see paragraph [0021] of Cheul).
The Applicant further argued that (on pages 10 and 11) that neither Varanasi nor Tuzel teaches the newly added limitation. The Examiner agrees. Varanasi teaches that landmark’s may be automatically determined based on the location of facial features in an image (see paragraph [0047]), but fails to explicitly teach that the user may specify a landmark location. Tuzel similarly teaches that landmark positions may be automatically determined based on the location of features on a 3D model (see paragraph [0010]), but fails to explicitly teach that the user may directly specify a landmark location.
Thus, the Applicant’s Amendments necessitated the new grounds of rejection presented in this Office Action, and the independent Claims 1, 10 and 16 are rejected under 35 USC 103 as being unpatentable over Liu and Varanasi. Therefore, the rejections to the dependent claims are maintained.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claim(s) 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. (US Pub No 2022/0292773), hereinafter Liu in view of Varanasi (US Pub No 2017/0278302), hereinafter Varanasi.
As to Claim 1, Liu teaches a computer-implemented (see Fig.34, image processing apparatus 3400) method (see paragraph [0002], “The present disclosure relates generally to image technologies, and in particular, to image processing and head/facial model formation methods and systems”) comprising:
receiving an input image including one or more facial representations (see abstract, “a method of generating a three-dimensional (3D) head deformation model that includes: receiving a two-dimensional (2D) facial image”)
and a set of points associated with a 3D canonical shape, wherein the set of points corresponds to a user-defined layout of a set of landmarks that is selectable at runtime (see paragraph [0021], “a set of user-provided keypoint annotations located on a plurality of vertices of a mesh of a 3D head template model”, and see Fig 23, Image 2304, where a set of points are manually marked on a 3D template)
and extracting a set of features from the input image that represent at least one facial representation included in the one or more facial representations (see paragraph [0013], “Apart from the facial keypoint detection, in some embodiments, multi-task learning and transfer learning solutions are implemented for facial feature classification tasks, so that more information can be extracted from an input face image, which is complementary to the keypoints information. The detected facial keypoints with the predicted facial features together are valuable to computers or mobile games for creating the face avatar of the players”).
Liu fails to teach determining a set of landmarks on the facial representation based on the set of features and the set of points, wherein each landmark in the set of landmarks is associated with at least one point in the set of points. However, Varanasi teaches that 3D points can be projected onto 2D images to determine 2D landmarks (see paragraph [0052], “Thus as an output from Saragih’s Face Tracker a triangulated 3D point cloud is obtained, the 2x4 Projection Matrix and the corresponding images with projected landmark points for every frame of the video”, and see Fig. 10C with landmark points shown on 2D images). Varanasi and Liu are combinable because both are from analogous fields of facial feature extraction. Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Liu and Varanasi. The motivation for doing so would be to allow the user to track landmarks across a range of different facial expressions to build a model that is capable of displaying several facial expressions. Varanasi teaches in paragraph [0002], “Mechanisms for registering different images and videos to a common 3D geometric model, can lead to several interesting applications. For example, semantically rich video editing applications can be developed, such as changing the facial expression of the person in a given image or even making the person appear younger. However, for in order to realize any such applications, firstly, a 3D face registration algorithm is required that robustly estimates a registered 3D mesh in correspondence to an input image”. Thus, one of ordinary skill in the art would have been motivated to track landmarks across a range of facial expressions. Additionally, Varanasi further teaches in paragraph [0054], “However these face databases are expensive and building them is a time-consuming effort. So instead, a simple blendshape model is used showing facial expressions of a single person”. By combining multiple expressions into one model, time and database storage is saved. Thus, it would have been obvious to combine the teachings of Liu with the teachings of Varanasi to obtain the invention as claimed in Claim 1.
As to Claim 2, Liu fails to teach further comprising encoding the set of points based on a latent representation to generate a set of position queries, wherein the set of landmarks are generated using the set of position queries. However, Varanasi teaches that the points can be encoded to generate position queries (see paragraph [0063], “Facial feature points of the 3D face model are grouped into face regions 810, and the corresponding landmark points of the face tracker are grouped into corresponding regions 820 as shown in FIG. 8 . For each region, a local affine warp T; is computed that maps a region from the face model to the corresponding region of the output of the face tracker This local affine warp is composed of a global rigid transformation and scaling (that affects all the vertices) and a residual local affine transform”). The facial landmark’s locations are determined based on the regions, which are derived from the output of the face tracker. A 3D point cloud is the output of the face tracker, (see paragraph [0079]), and this 3D point cloud can be considered a latent representation. These points are then encoded through the local affine warps which are then later used to create a 3D mesh, which can be projected over images to visualize landmarks (see paragraph [0087], “The following step involves projecting the meshes onto the image frames in order to build up a correspondence between the pixels of the Kth frame and the vertices in the Kth 3D blendshape model”). Thus, it would have been obvious to one of ordinary skill before the effective filing date of the claimed invention to combine the teachings of Liu with the teachings of Varanasi. The motivation for doing so would be to allow the user to build meshes of facial representations which can blend several different facial expressions. Varanasi teaches in paragraph [0043], “The method includes applying a 3D mesh blend shape model of a human face that is parameterized to blend between different facial expressions.” This prevents the user from having to build multiple face databases with multiple expressions. Varanasi further teaches in paragraph [0054], “However these face databases are expensive and building them is a time-consuming effort. So instead, a simple blendshape model is used showing facial expressions of a single person.” By combining multiple expressions into one model, time and database storage is saved. Thus, it would have been obvious to one of ordinary skill to combine the teachings of Liu with the teachings of Varanasi in order to obtain the invention as claimed in Claim 2.
As to Claim 3, Liu in view of Varanasi teaches the computer-implemented method of Claim 1, wherein the 3D canonical shape comprises a fixed 3D object model of a face, (see Liu, Fig. 23, Image 2304, where a template face is shown).
As to Claim 4, Liu in view of Varanasi teaches the computer-implemented method of Claim 1, wherein the set of points are positioned on or around the 3D canonical shape based on a desired layout of the set of landmarks (see Liu, see paragraph [0021], “a set of user-provided keypoint annotations located on a plurality of vertices of a mesh of a 3D head template model”, and see Fig 23, Image 2304, where a set of points are manually marked on a 3D template).
As to Claim 5, Liu in view of Varanasi teaches the computer-implemented method of Claim 1, wherein the input image comprises a two-dimensional image captured by an image capture device (see Liu, paragraph [0021], “According to a third aspect of the present application, a method of generating a three-dimensional (3D) head deformation model, includes: receiving a two-dimensional (2D) facial image”, and see paragraph [0231], “The system and method do not require the face to be directly facing the camera”, where the camera is the ‘image capture device’).
As to Claim 6, Liu fails to teach generating a facial segmentation mask associated with the at least one facial representation based on the set of landmarks, wherein the facial segmentation mask divides the at least one face into semantically meaningful regions. However, Varanasi teaches that regions facial regions may be found from an input image, (see paragraph [0020], “computing , a set of localized affine transformations connecting a set of facial regions of the said 3D facial model to the sets of feature points defining the sparse facial landmarks”, and see Fig. 9C where different facial regions are shown on a model). Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Liu and Varanasi. The motivation for doing so would be allow for different regions of the face to be tracked across different facial expressions. Varanasi teaches in paragraph [0092], “Using the technique of face registration by localized affine warps according to the embodiment of the invention, a dense registration of the different regions in the face model to a given input face image is obtained, as illustrated in FIG . 10 . FIG . 10 the registered 3D face model to different face input images”. By tracking the regions of the face, a 3D model can be obtained which can show a range of different facial expressions without requiring a large database (see paragraph [0104] “Embodiments of the invention provide that produces a dense 3D mesh output , but which is computationally fast and has little overhead. Moreover, embodiments of the invention do not require a 3D face database . Instead, it may use a 3D face model showing expression changes from one single person as a reference person, which is far easier to obtain”). Thus, it would have been obvious to combine the teachings of Liu with the teachings of Varanasi to obtain the invention as claimed in Claim 6.
As to Claim 7, Liu fails to teach receiving a second input image including the at least one facial representation wherein the second input image is captured at a different point in time from the input image. However, Varanasi teaches that several images can be received (see paragraph [0042], “In a general embodiment the invention involves inputting a monocular face video comprising a sequence of captured images of a face and tracking facial landmarks”).
Liu fails to teach extracting a second set of features from the second input image that represent the at least one facial representation. However, Varanasi teaches that features can be extracted from a series of images, (see paragraph [0047], “In step S102 2D facial landmark features are tracked through the sequence of images in acquisition step 1”, and see Fig. 2 where multiple frames are capture over time).
Liu fails to teach determining a second set of landmarks on the at least one facial representation based on the second set of features and the set of points wherein each landmark in the second set of landmarks is associated with at least one point in the set of points. Varanasi teaches that different landmarks are tracked, (see paragraph [0047], “Thus as an output from Saragih’s Face Tracker a triangulated 3D point cloud is obtained , the 2x4 Projection Matrix and the corresponding images with projected landmark points for every frame of the video”).
Liu fails to teach comparing a first landmark in the set of landmarks and a second landmark in the second set of landmarks to perform facial tracking operations. However, Varanasi teaches that multiple landmarks can be tracked across several images, (see paragraph [0016], “tracking a set of facial landmarks in a sequence of facial images of a target person”, where sequence of images implies that multiple landmarks across images can be compared). Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Liu with Varanasi. The motivation for doing so would be to allow the user to build meshes of facial representations which can blend several different facial expressions. Liu teaches in paragraph [0043], “The method includes applying a 3D mesh blend shape model of a human face that is parameterized to blend between different facial expressions.” This prevents the user from having to build multiple face databases with multiple expressions. To model expressions, landmarks must be captured across multiple images that show the expressions. Varanasi further teaches in paragraph [0054], “However these face databases are expensive and building them is a time-consuming effort. So instead, a simple blendshape model is used showing facial expressions of a single person.” By combining multiple facial expressions into a single model, time and database storage is saved. Thus, it would have been obvious to combine the teachings of Liu with the teachings of Varanasi to obtain the invention as claimed in Claim 7.
As to Claim 9, Liu fails to teach that the set of landmarks are determined using one or more trained machine learning models. However, Varanasi teaches in paragraph [0048], “In one embodiment of the invention, the 2D landmark features are tracked using a sparse spatial feature tracking algorithm, for example Saragih's face tracker (“Face alignment through subspace constrained mean-shifts”… Alternatively, other techniques used in the computer vision such as dense optical flow, particle filters may be applied for facial landmark tracking. The Saragih tracking algorithm uses a sparse set of 66 points on the face including the eyes, nose, mouth, face boundary and the eye brows. The algorithm is based upon a Point Distribution model (PDM)”, where a Point Distribution Model is a machine learning model”). Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Liu and Varanasi. The motivation for doing so would be to automatically obtain 2D landmarks across multiple frames in video data, (see paragraph [0052], “Thus as an output from Saragih’s Face Tracker a triangulated 3D point cloud is obtained, the 2x4 Projection Matrix and the corresponding images with the projected landmark points for every frame of the video”). Thus, it would have been obvious to combine the teachings of Liu and Varanasi in order to obtain the invention as claimed in Claim 9.
As to Claim 10, Claim 10 claims one or more non-transitory computer readable media (see Liu paragraph [0024], “a non-transitory computer readable storage medium stores a plurality of programs for execution”,),
that when executed by one or more processors (see Liu paragraph [0024], “The programs, when executed by the one or more processing units, cause the electronic apparatus to perform the one or more methods as described above”), cause the processors to perform the same computerized method as claimed in Claim 1. Therefore, the rejection and rationale are analogous to that made in Claim 1.
As to Claim 11, Claim 11 claims the same limitation as Claim 2, and is dependent on a similarly rejected independent claim. Therefore, the rejection and rationale are analogous to that made in Claim 2.
As to Claim 12, Claim 12 claims the same limitation as Claim 3, and is dependent on a similarly rejected independent claim. Therefore, the rejection and rationale are analogous to that made in Claim 3.
As to Claim 13, Claim 13 claims the same limitation as Claim 4, and is dependent on a similarly rejected independent claim. Therefore, the rejection and rationale are analogous to that made in Claim 4.
As to Claim 14, Claim 14 claims the same limitation as Claim 5, and is dependent on a similarly rejected independent claim. Therefore, the rejection and rationale are analogous to that made in Claim 5.
As to Claim 15, Claim 13 claims the same limitation as Claim 6, and is dependent on a similarly rejected independent claim. Therefore, the rejection and rationale are analogous to that made in Claim 6.
As to Claim 16, Claim 16 claims the same limitation as Claim 7, and is dependent on a similarly rejected independent claim. Therefore, the rejection and rationale are analogous to that made in Claim 7.
As to Claim 18, Claim 18 claims the same limitation as Claim 9, and is dependent on a similarly rejected independent claim. Therefore, the rejection and rationale are analogous to that made in Claim 9.
As to Claim 19, Claim 19 claims a computer system, comprising: one or more memories and one or more processors (see Varanasi paragraph [0023], “According to a fifth aspect of the present application, an electronic apparatus includes one or more processing units, memory and a plurality of programs stored in the memory. The programs, when executed by the one or more processing units, cause the electronic apparatus to perform the one or more methods as described above”)
for executing the same computerized method as claimed in Claim 1. Therefore, the rejection and rationale are analogous to that made in Claim 1.
As to Claim 20, Claim 20 claims the same limitation as Claim 2, and is dependent on a similarly rejected independent claim. Therefore, the rejection and rationale are analogous to that made in Claim 2.
Claim(s) 8 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. (US Pub No 2022/0292773), hereinafter Liu, in view of Varanasi et al. (US 2017/0278302), hereinafter Varanasi and further in view of Tuzel et al. (US Pub No 2017/0083751), hereinafter Tuzel.
As to Claim 8, Liu in view of Varanasi fails to receiving an annotated image including one or more landmarks; and determining, via query optimization, the set of points based on the one or more landmarks. However, Tuzel teaches receiving annotated images including one or more landmarks and later obtaining global parameters (points associated with a 3D model), (see paragraph [0010], “The method acquires an image of a face, e.g., using a camera or obtaining a previously captured image. The input to the method includes an initial estimate of a set of locations of the landmarks, referred to as initial landmark locations. The set of initial landmark locations is globally aligned to a set of landmark locations of a face with a prototype shape to obtain global alignment parameters”, and [0030], “The initial landmark locations can be marked manually or automatically, e.g., using a. facial part detection algorithm”) Tuzel further teaches in paragraph [0032], that “the prototype face shape could, for example, be obtained or adapted from an existing 2D or 3D face model”. Tuzel is combinable with Liu and Varanasi because all three are from the analogous art of facial feature detection and landmark estimation. Thus, it would be obvious to one of ordinary skill of the art before the effective filing date of the claimed invention to combine the teachings of Tuzel with the teachings of Liu and Varanasi. The motivation for doing so would be improved alignment of landmarks on facial images. Tuzel teaches in paragraph [0055], “Our model significantly improves upon the alignment accuracy and robustness of the prior art.” Thus, it would have been obvious to combine the teachings of Tuzel with the teachings of Liu and Varanasi to obtain the invention as claimed in Claim 8.
As to Claim 17, Claim 17 claims the same limitation as Claim 8, and is dependent on a similarly rejected independent claim. Therefore, the rejection and rationale are analogous to that made in Claim
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Lin et al. (Kevin Lin, Lijuan Wang, and Zicheng Liu, “End-to-end human pose and mesh reconstruction with transformers”, In CVPR, pages 1954–1963, 2021) teaches a transformer that uses an image and 3D queries on a canonical form to create mesh reconstructions.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SOUMYA THOMAS whose telephone number is (571)272-8639. The examiner can normally be reached M-F 8:30-5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Mehmood can be reached at (571) 272-2976. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/S.T./Examiner, Art Unit 2664
/JENNIFER MEHMOOD/Supervisory Patent Examiner, Art Unit 2664