DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Applicant’s submission filed on 07/13/2026 has been entered. Claims 1-2, 7, 10-11, 16, and 19 were amended. Claims 8 and 17 were canceled. Claims 1-7, 9-16, and 18-20 are pending in the application.
Claim Objections
Claims 9 and 18 are objected to because of the following informalities: Claims 9 and 18 depend on claims 8 and 17, respectively, which have been canceled. (For examination, claims 9 and 18 are considered to depend on claims 7 and 16, respectively.) Appropriate correction is required.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-2, 10-11, and 19-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al. (US 2014/0043329) in view of Peng et al. ("Sparse-to-dense multi-encoder shape completion of unstructured point cloud," IEEE Access 8 (2020): 30969-30978), Ambrus et al. (US 2023/0177850), Martinez (US 2019/0114824), and Yebes Torres et al. (US 2024/0096125).
Regarding claim 1, Wang teaches/suggests: A computer-implemented method for completing three dimensional face reconstruction comprising:
receiving image data associated with multiple two dimensional non-frontal face images (Wang [0029] “The camera obtains at least one 2D image 102” [0126] “approximately 30 photos around the face of the user may be taken”);
analyzing the image data and extracting two dimensional facial features (Wang [0031] “Personalized avatar generation component 112 detects face regions in the 2D images 102 and reconstructs a face mesh … landmark feature points between the 2D face model and 3D face model may be detected and registered”);
constructing sparse three dimensional facial feature point clouds based on the two dimensional facial features (Wang [0031] “sparse point clouds of the user's face will be recovered accordingly” [0029] “Personalized facial components 106 comprise a 3D morphable model”); and
generate a three dimensional facial feature point cloud of complete facial features (Wang [0031] “a dense point cloud for the 2D face model may be estimated based on multi-view images with a bundle adjustment approach” [0029] “Personalized facial components 106 comprise a 3D morphable model”),
wherein the three dimensional facial feature point cloud is utilized to control a computing device to complete a downstream task (Wang [0029] “The personalized facial components 106 may be used in other application programs, processing systems, and/or processing devices as desired”),
Wang does not teach/suggest:
inputting the sparse three dimensional facial feature point clouds into an encoder-decoder architecture to generate a three dimensional facial feature point cloud of complete facial features,
Peng, however, teaches/suggests an encoder-decoder architecture (Peng §III-C ¶2 “we adopt PointNet++ [12] as the encoder to encode and decode the sparse point cloud ... we use a three-layer MLP to decode the expanded feature and reduce the feature dimension to 3 to regress the 3D coordinates of the final point cloud”). Before the effective filing date of the claimed invention, the substitution of one known element (the PointNet++ of Peng) for another (the bundle adjustment of Wang) would have been obvious to one of ordinary skill in the art because such substitutions would have yielded predictable results, namely, to generate the dense point cloud. As such, Wang as modified by Peng teaches/suggests:
inputting the sparse three dimensional facial feature point clouds into an encoder-decoder architecture to generate a three dimensional facial feature point cloud of complete facial features (Wang [0031] “sparse point clouds of the user's face will be recovered accordingly” Peng §III-C ¶2 “we adopt PointNet++ [12] as the encoder to encode and decode the sparse point cloud ... we use a three-layer MLP to decode the expanded feature and reduce the feature dimension to 3 to regress the 3D coordinates of the final point cloud”),
Peng further teaches/suggests in parallel (Peng §III-C ¶2 “The hierarchical feature learning structure of PointNet++ has been proven to be able to learn the local and global features of point cloud simultaneously”). Wang as modified by Peng does not teach/suggest including an encoder employing two or more types of neural networks in parallel. Ambrus and Martinez, teach/suggest two or more types of neural networks (Ambrus [0080] “Another implementation of the aggregation layer 860 may be a dynamic graph (DG) convolutional neural network (CNN) that maintains a permutation invariance of point sets; however, the aggregation layer 860 may be designed to capture a local geometric structure by encoding features in edges between points” Martinez [0009] “The innovation enables mapping with a feed-forward neural network that defines two criteria, one that learns to detect important shape landmark points on an image”). Before the effective filing date of the claimed invention, the substitution of one known element (the dynamic graph CNN of Ambrus and the feed-forward NN of Martinez) for another (the PointNet++ of Peng) would have been obvious to one of ordinary skill in the art because such substitutions would have yielded predictable results, namely, to learn the local and global features.
As such, Wang as modified by Peng, Ambrus, and Martinez teaches/suggests an encoder employing two or more types of neural networks in parallel (Peng §III-C ¶2 “The hierarchical feature learning structure of PointNet++ has been proven to be able to learn the local and global features of point cloud simultaneously” Ambrus [0080] “Another implementation of the aggregation layer 860 may be a dynamic graph (DG) convolutional neural network (CNN) that maintains a permutation invariance of point sets; however, the aggregation layer 860 may be designed to capture a local geometric structure by encoding features in edges between points” Martinez [0009] “The innovation enables mapping with a feed-forward neural network that defines two criteria, one that learns to detect important shape landmark points on an image”).
Wang, Peng, Ambrus, and Martinez are silent regarding:
wherein the encoder of the encoder-decoder architecture encodes node features along with information of neighbor nodes to use both local and global information to learn an overall geometry of the face of an individual who is being captured within the images.
Yebes Torres, however, teaches/suggests encodes node features along with information of neighbor nodes to use both local and global information (Yebes Torres [0116] “The GAN-based model 414 is used to compute hidden representations of each node 410 in the graph 406 by attending over its neighbors nodes 410 (e.g., a local aspect) and the global node, which causes the GAN-based model 414 to learn contextualized information in the document from both local and global aspects”). Before the effective filing date of the claimed invention, it would have been obvious for one of ordinary skill in the art to modify the encoding of Wang as modified by Peng, Ambrus, and Martinez to include both local and global aspects as taught/suggested by Yebes Torres to learn contextualized information. As such, Wang as modified by Peng, Ambrus, Martinez, and Yebes Torres teaches/suggests:
wherein the encoder of the encoder-decoder architecture encodes node features along with information of neighbor nodes to use both local and global information to learn an overall geometry of the face of an individual who is being captured within the images (Wang [0031] “sparse point clouds of the user's face will be recovered accordingly” Peng §III-C ¶2 “we adopt PointNet++ [12] as the encoder to encode and decode the sparse point cloud ... we use a three-layer MLP to decode the expanded feature and reduce the feature dimension to 3 to regress the 3D coordinates of the final point cloud” Yebes Torres [0116] “The GAN-based model 414 is used to compute hidden representations of each node 410 in the graph 406 by attending over its neighbors nodes 410 (e.g., a local aspect) and the global node, which causes the GAN-based model 414 to learn contextualized information in the document from both local and global aspects”).
Regarding claim 2, Wang as modified by Peng, Ambrus, Martinez, and Yebes Torres teaches/suggests: The computer-implemented method of claim 1, wherein the multiple two dimensional non-frontal face images include occlusions that are caused by at least one of: the individual who is being captured within the images and an object that is located in between at least one camera and the individual who is being captured within the images (Wang [0057] “FIG. 10 shows faces wearing sunglasses and faces being occluded by a hand or hair”).
Claims 10 and 11 recite limitation(s) similar in scope to those of claims 1 and 2, respectively, and are rejected for the same reason(s). Wang as modified by Peng, Ambrus, Martinez, and Yebes Torres further teaches/suggests a memory storing instructions and a processor (Wang Fig. 18: memory 1812 and processor(s) 1802).
Claim 19 recites limitation(s) similar in scope to those of claim 1, and is rejected for the same reason(s). Wang as modified by Peng, Ambrus, Martinez, and Yebes Torres further teaches/suggests a non-transitory computer readable storage medium storing instructions (Wang Fig. 18: memory 1812).
Regarding claim 20, Wang as modified by Peng, Ambrus, Martinez, and Yebes Torres teaches/suggests: The non-transitory computer readable storage medium of claim 19, wherein an output vector of the encoder is fed into a decoder of the encoder-decoder architecture to decode sparse three dimensional facial feature point clouds and to generate the three dimensional facial feature point cloud of complete facial features (Wang [0031] “sparse point clouds of the user's face will be recovered accordingly” Peng §III-A ¶1 “we encode the input point cloud to get a high-dimensional feature vector, and then decode the feature vector to output a sparse point cloud with complete shape ... and get the final dense complete point cloud through encoding, feature expansion and decoding”). The same rationale to combine as set forth in the rejection of claim 1 is incorporated herein.
Claim(s) 3-5 and 12-14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al. (US 2014/0043329) in view of Peng et al. ("Sparse-to-dense multi-encoder shape completion of unstructured point cloud," IEEE Access 8 (2020): 30969-30978), Ambrus et al. (US 2023/0177850), Martinez (US 2019/0114824), and Yebes Torres et al. (US 2024/0096125) as applied to claims 2 and 11 above, and further in view of Hu et al. (US 2021/0209839).
Regarding claim 3, Wang, Peng, Ambrus, Martinez, and Yebes Torres are silent regarding: The computer-implemented method of claim 2, wherein analyzing the image data includes extracting a fixed number of facial landmarks, wherein the fixed number of facial landmarks include the occlusions and a shape completion matrix is used to estimate true locations of occluded facial feature points. Hu, however, teaches/suggests extracting a fixed number of facial landmarks (Hu [0048] “the circuitry 202 may acquire a plurality of pre-defined landmark points on the aligned 3D mean-shape model 324”), wherein the fixed number of facial landmarks include the occlusions and a shape completion matrix is used to estimate true locations of occluded facial feature points (Hu [0058] “As the second aligned 3D mean-shape model 402B may be associated with the non-frontal view of the face of the user 110 in which the face of the user 110 may be titled towards the left-side, a portion of the left-side of the face may be occluded … vertices 410A, 410B, 410C, 410D, 410E, and 410F may be the left-most vertices on the parallel lines in the second aligned 3D mean-shape model 402B. The circuitry 202 may determine landmark points 412A, 412B, 412C, 412D, 412E, and 412F in the second 2D projection 406B as landmark points on the contour of the face, which may correspond to the vertices 410A, 410B, 410C, 410D, 410E, and 410F”). Before the effective filing date of the claimed invention, it would have been obvious for one of ordinary skill in the art to modify the feature points of Wang as modified by Peng, Ambrus, Martinez, and Yebes Torres to be extracted as taught/suggested by Hu for the shape completion.
Regarding claim 4, Wang as modified by Peng, Ambrus, Martinez, Yebes Torres, and Hu teaches/suggests: The computer-implemented method of claim 3, wherein facial features that correspond to the facial landmarks in the multiple two dimensional non-frontal face images are matched (Wang [0097] “a set of corresponding points in the multiple images may be tracked as t.sub.k={x1.sub.k, x2.sub.k, x3.sub.k, . . . } which depict the same 3D point in the first image, second image, and third image”).
Regarding claim 5, Wang as modified by Peng, Ambrus, Martinez, Yebes Torres, and Hu teaches/suggests: The computer-implemented method of claim 4, wherein constructing the sparse three dimensional facial feature point clouds includes using the matching correspondences of the facial features to the facial landmarks in the multiple two dimensional non-frontal face images to create a three dimensional reconstruction of sparse feature points (Wang [0097] “a set of corresponding points in the multiple images may be tracked as t.sub.k={x1.sub.k, x2.sub.k, x3.sub.k, . . . } which depict the same 3D point in the first image, second image, and third image”).
Claims 12-14 recite limitation(s) similar in scope to those of claims 3-5, respectively, and are rejected for the same reason(s).
Claim(s) 6-7, 9, 15-16, and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al. (US 2014/0043329) in view of Peng et al. ("Sparse-to-dense multi-encoder shape completion of unstructured point cloud," IEEE Access 8 (2020): 30969-30978), Ambrus et al. (US 2023/0177850), Martinez (US 2019/0114824), Yebes Torres et al. (US 2024/0096125), and Hu et al. (US 2021/0209839) as applied to claims 5 and 14 above, and further in view of Zhang et al. (US 2023/0206603).
Regarding claim 6, Wang, Peng, Ambrus, Martinez, Yebes Torres, and Hu are silent regarding: The computer-implemented method of claim 5, wherein inputting the sparse three dimensional facial feature point clouds into the encoder-decoder architecture include inputting the sparse three dimensional facial feature point clouds with a variable number of points to generate a complete dense point cloud of a missing part of the face of the individual who is being captured within the images. Zhang, however, teaches/suggests point clouds with a variable number of points (Zhang [0065] “Three missing point clouds of different scales generated by sampling the farthest point are input into the multi-resolution encoder module to extract feature”). Before the effective filing date of the claimed invention, it would have been obvious for one of ordinary skill in the art to modify the sparse point clouds of Wang as modified by Peng, Ambrus, Martinez, Yebes Torres, and Hu to have different scales as taught/suggested by Zhang for the shape completion.
As such, Wang as modified by Peng, Ambrus, Martinez, Yebes Torres, Hu, and Zhang teaches/suggests inputting the sparse three dimensional facial feature point clouds with a variable number of points to generate a complete dense point cloud of a missing part of the face of the individual who is being captured within the images (Wang [0031] “sparse point clouds of the user's face will be recovered accordingly” Peng §III-C ¶2 “we adopt PointNet++ [12] as the encoder to encode and decode the sparse point cloud ... we use a three-layer MLP to decode the expanded feature and reduce the feature dimension to 3 to regress the 3D coordinates of the final point cloud” Zhang [0065] “Three missing point clouds of different scales generated by sampling the farthest point are input into the multi-resolution encoder module to extract feature”).
Regarding claim 7, Wang, Peng, Ambrus, Martinez, Yebes Torres, Hu, and Zhang teaches/suggests: The computer-implemented method of claim 6, wherein the encoder of the encoder-decoder architecture employs a graph convolutional neural network to understand a specific geometry of the sparse three dimensional facial feature point clouds and uses a fully connected class of a feedforward artificial neural network to learn the overall geometry of the face of the individual who is being captured within the images (Wang [0031] “sparse point clouds of the user's face will be recovered accordingly” Peng §III-C ¶2 “The hierarchical feature learning structure of PointNet++ has been proven to be able to learn the local and global features of point cloud simultaneously” Ambrus [0080] “Another implementation of the aggregation layer 860 may be a dynamic graph (DG) convolutional neural network (CNN) that maintains a permutation invariance of point sets; however, the aggregation layer 860 may be designed to capture a local geometric structure by encoding features in edges between points” Martinez [0009] “The innovation enables mapping with a feed-forward neural network that defines two criteria, one that learns to detect important shape landmark points on an image” [0044] “the deep neural network used four convolutional layers, two max pooling layers and two fully connected layers”). The same rationale to combine as set forth in the rejection of claim 1 is incorporated herein.
Regarding claim 9, Wang, Peng, Ambrus, Martinez, Yebes Torres, Hu, and Zhang teaches/suggests: The computer-implemented method of claim [7], wherein an output vector of the encoder is fed into a decoder of the encoder-decoder architecture to decode sparse three dimensional facial feature point clouds and to generate the three dimensional facial feature point cloud of complete facial features (Wang [0031] “sparse point clouds of the user's face will be recovered accordingly” Peng §III-A ¶1 “we encode the input point cloud to get a high-dimensional feature vector, and then decode the feature vector to output a sparse point cloud with complete shape ... and get the final dense complete point cloud through encoding, feature expansion and decoding”). The same rationale to combine as set forth in the rejection of claim 1 is incorporated herein.
Claims 15-16 and 18 recite limitation(s) similar in scope to those of claims 6-7 and 9, respectively, and are rejected for the same reason(s).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
US 2022/0351863 – GNN including local and global features
US 2023/0351112 – local relationship between nodes and global characteristic
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ANH-TUAN V NGUYEN whose telephone number is 571-270-7513. The examiner can normally be reached on M-F 9AM-5PM ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, JASON CHAN can be reached on 571-272-3022. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ANH-TUAN V NGUYEN/
Primary Examiner, Art Unit 2619