Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-4, 6-15, and 17-20 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Smith et al. (Patent No.: US 11,854,146).
Regarding claim 1, Smith discloses an image processing method, comprising: obtaining at least two images of a target object to be processed (see Fig. 2A, “2D Body Image” 202-1 to 202-N) , wherein the at least two images to be processed include images captured from at least two different shooting angles (e.g. the 2D images including front, side and back); inputting the at least two images to be processed into a target three-dimensional (3D) reconstruction model to reconstruct a target 3D body shape corresponding to the target object (see Fig. 2A, “3D modeling 210” and “3D Refinement 208”), wherein the target 3D reconstruction model is trained based on M preset dimensional key points, and M is a positive integer (see Fig. 1B and col. 5, lines 16-24: “the 2D body image or segmented silhouette of the 2D body image may be processed to determine 2D landmarks 151 of the body. For example, the image 104 may be provided to a trained machine learning model, such as a CNN that is trained to determine 2D landmarks of a body represented in an image. Based on the provided input, the CNN may generate an output indicating the location (e.g., x, y coordinates, or pixels) corresponding to the 2D landmarks for which the CNN was trained”); and outputting the target 3D body shape (Fig. 2A, 3D body shape “214”).
Regarding claim 2, Smith discloses the method according to claim 1, wherein inputting the at least two images to be processed into the target 3D reconstruction model to reconstruct the target 3D body shape corresponding to the target object comprises: determining a global feature (se Fig. 1A, “SHAPE Parameters”), a shooting parameter (Fig. 1A, “Camera parameters”), and a posture parameter (Fig. 1A, “3D POSE”) corresponding to the target object based on the at least two images to be processed and through the target 3D reconstruction model; and reconstructing the target 3D body shape based on the global feature, shooting parameter, and posture parameter through the target 3D reconstruction model (Fig. 1A, “CNN 106A” and “3D model refinement 108”).
Regarding claim 3, Smith discloses the method according to claim 2, wherein determining the global feature corresponding to the target object through the target 3D reconstruction model comprises: identifying, in the at least two images to be processed, an image region where the target object is located; segmenting the image region from the at least two images to be processed to obtain at least two target image patches corresponding to the target object (Fig. 1A, “semantic segmentation 104”, and also see col. 3, lines 59-66: “the segmented silhouette may be segmented into one or more body segments, such as hair segment 104-1, head segment 104-2, neck segment 104-3, upper clothing segment 104-4, upper left arm 104-5, lower left arm 104-6, left hand 104-7, torso 104-8, upper right arm 104-9, lower right arm 104-10, right hand 104-11, lower clothing 104-12, upper left leg 104-13, upper right leg 104-16, etc.” ); determining a set of semantic features corresponding to each of the target image patches (Fig. 1A, “semantic segmentation 104”); and fusing at least two sets of the semantic features to obtain the global feature (see Fig. 1A, “CNN 106A, SHAPE Parameter”).
Regarding claim 4, Smith discloses the method according to claim 3, wherein determining a set of semantic features corresponding to each of the target image patches comprises: inputting each of the target image patches into a deep encoder based on a convolutional neural network (CNN) to extract a set of semantic features corresponding to the target image patch, wherein the target 3D reconstruction model comprises the deep encoder based on the convolutional neural network (see Fig. 1A, “semantic segmentation” and also see col. 19, lines 50-59, “the neural network is trained to identify the target output, to within an acceptable level of error. In unsupervised learning of an identity function, such as that which is typically performed by a sparse autoencoder, target output of the training set is the input, and the neural network is trained to recognize the input as such. Sparse autoencoders employ backpropagation in order to train the autoencoders to recognize an approximation of an identity function for an input, or to otherwise approximate the input”).
Regarding claims 6, Smith discloses the method according to claim 2, wherein determining the shooting parameter and posture parameter corresponding to the target object through the target 3D reconstruction model comprises: performing prediction processing on the at least two images to be processed using an encoder and a regressor to obtain the shooting parameter and posture parameter, wherein the target 3D reconstruction model comprises the encoder and the regressor (see Fig. 1A, “3D POSE”, Fig. 2A, “”207 Predict Body Parameters” and Abstract “Described are systems and methods directed to generation of a dimensionally accurate three-dimensional (“3D”) model of a body, such as a human body, based on two-dimensional (“2D”) images of at least a portion of that body. A user may use a 2D camera, such as a digital camera typically included in many of today's portable devices (e.g., cell phones, tablets, laptops, etc.) and obtain a series of 2D body images of at least a portion of their body from different views with respect to the camera. The 2D body images may then be used to generate a plurality of predicted body parameters corresponding to the body represented in the 2D body images. Those predicted body parameters may then be further processed to generate a dimensionally accurate 3D model or avatar of the body of the user”).
Regarding claim 7, Smither discloses the method according to claim 2, wherein reconstructing the target 3D body shape based on the global feature, shooting parameter, and posture parameter through the target 3D reconstruction model comprises: fitting the global feature, shooting parameter, and posture parameter to obtain a predicted parameter; and inputting the predicted parameter into a human body reconstruction model to obtain the target 3D body shape, wherein the target 3D reconstruction model comprises the human body reconstruction model (see Fig. 1A and Abstract).
Regarding claim 8, Smith discloses the method according to claim 1, further comprising, before inputting the at least two images to be processed into the target 3D reconstruction model: inputting a sample image (e.g. initial partial body images) into a preset 3D reconstruction model to obtain a sample 3D body shape; projecting the sample 3D body shape onto a 2D plane to obtain sample projection dimensional key points corresponding to the sample 3D body shape (see Fig. 1B, key points 151-1…..151-18); comparing the sample projection dimensional key points with preset dimensional key points to determine a loss value; and adjusting the parameters of the preset 3D reconstruction model based on the loss value and performing iterative training to obtain the target 3D reconstruction model (see Fig. 1A, “3D refinement”).
Regarding claim 9, Smith discloses the method according to claim 1, wherein outputting the target 3D body shape comprises: performing preset processing on the target 3D body shape to obtain a 3D body shape to be displayed, wherein the preset processing includes at least one of clothing addition or facial processing; and displaying the 3D body shape to be displayed (see Fig. 1A, “112 texture augmentation”).
Regarding claim 10, Smith discloses a non-transitory computer-readable storage medium configured with instructions executable by one or more processors to cause the one or more processors to perform the method of claim 1 (see col. 31, lines 4-22).
Regarding claim 11, Smith discloses an electronic device comprising: one or more processors; and one or more computer-readable memories coupled to the one or more processors and having instructions stored thereon that are executable by the one or more processors to perform the method of claim 1 (see col. 31, lines 4-22).
Regarding claim 12, Smith discloses an image processing method, comprising: obtaining, in response to a touch operation on a first control at least two images of a target object to be processed (col. 1, line 65 to col. 2, line 3; and col. 22, lines 19-24), wherein the at least two images include images captured from at least two different shooting angles (e.g. the 2D images including front, side and back); inputting, in response to a touch operation on a second control, the at least two images to be processed into a target 3D reconstruction model to reconstruct a target 3D body shape corresponding to the target object (see Figs. 1A and 2A) , wherein the target 3D reconstruction model is trained based on M preset dimensional key points , and M is a positive integer (see Fig. 1B and col. 5, lines 16-24: “the 2D body image or segmented silhouette of the 2D body image may be processed to determine 2D landmarks 151 of the body. For example, the image 104 may be provided to a trained machine learning model, such as a CNN that is trained to determine 2D landmarks of a body represented in an image. Based on the provided input, the CNN may generate an output indicating the location (e.g., x, y coordinates, or pixels) corresponding to the 2D landmarks for which the CNN was trained”); and outputting and displaying the target 3D body shape (Fig. 2A, 3D body shape “214”).
Regarding claim 13, Smith discloses the method according to claim 12, wherein , in response to a touch operation on a second control, the at least two images to be processed into a target 3D reconstruction model to reconstruct a target 3D body shape corresponding to the target object comprises: determining a global feature (se Fig. 1A, “SHAPE Parameters”), a shooting parameter (Fig. 1A, “Camera parameters”), and a posture parameter (Fig. 1A, “3D POSE”) corresponding to the target object based on the at least two images to be processed and through the target 3D reconstruction model; and reconstructing the target 3D body shape based on the global feature, shooting parameter, and posture parameter through the target 3D reconstruction model (Fig. 1A, “CNN 106A” and “3D model refinement 108”).
Regarding claim 14, Smith discloses the method according to claim 13, wherein determining the global feature corresponding to the target object through the target 3D reconstruction model comprises: identifying, in the at least two images to be processed, an image region where the target object is located; segmenting the image region from the at least two images to be processed to obtain at least two target image patches corresponding to the target object; determining a set of semantic features corresponding to each of the target image patches; and fusing at least two sets of the semantic features to obtain the global feature (see Fig. 1A, “Semantic Segmentation 104).
Regarding claim 15, Smith discloses the method according to claim 14, wherein determining a set of semantic features corresponding to each of the target image patches comprises: inputting each of the target image patches into a deep encoder based on a convolutional neural network (CNN) to extract a set of semantic features corresponding to the target image patch, wherein the target 3D reconstruction model comprises the deep encoder based on the convolutional neural network (see Fig. 1A, “semantic segmentation” and also see col. 19, lines 50-59, “the neural network is trained to identify the target output, to within an acceptable level of error. In unsupervised learning of an identity function, such as that which is typically performed by a sparse autoencoder, target output of the training set is the input, and the neural network is trained to recognize the input as such. Sparse autoencoders employ backpropagation in order to train the autoencoders to recognize an approximation of an identity function for an input, or to otherwise approximate the input”).
Regarding claim 17, Smith discloses the method according to claim 13, wherein determining the shooting parameter and posture parameter corresponding to the target object through the target 3D reconstruction model comprises: performing prediction processing on the at least two images to be processed using an encoder and a regressor to obtain the shooting parameter and posture parameter, wherein the target 3D reconstruction model comprises the encoder and the regressor (see Fig. 1A, “3D POSE”, Fig. 2A, “”207 Predict Body Parameters” and Abstract “Described are systems and methods directed to generation of a dimensionally accurate three-dimensional (“3D”) model of a body, such as a human body, based on two-dimensional (“2D”) images of at least a portion of that body. A user may use a 2D camera, such as a digital camera typically included in many of today's portable devices (e.g., cell phones, tablets, laptops, etc.) and obtain a series of 2D body images of at least a portion of their body from different views with respect to the camera. The 2D body images may then be used to generate a plurality of predicted body parameters corresponding to the body represented in the 2D body images. Those predicted body parameters may then be further processed to generate a dimensionally accurate 3D model or avatar of the body of the user”).
Regarding claim 18, Smith discloses the method according to claim 13, wherein reconstructing the target 3D body shape based on the global feature, shooting parameter, and posture parameter through the target 3D reconstruction model comprises: fitting the global feature, shooting parameter, and posture parameter to obtain a predicted parameter; and inputting the predicted parameter into a human body reconstruction model to obtain the target 3D body shape, wherein the target 3D reconstruction model comprises the human body reconstruction model (see Fig. 1A and Abstract).
Regarding claim 19, Smith discloses the method according to claim 12, further comprising, before inputting the at least two images to be processed into a target 3D reconstruction model: inputting a sample image (e.g. initial partial body images) into a preset 3D reconstruction model to obtain a sample 3D body shape; projecting the sample 3D body shape onto a 2D plane to obtain sample projection dimensional key points corresponding to the sample 3D body shape (see Fig. 1B, key points 151-1…..151-18); comparing the sample projection dimensional key points with preset dimensional key points to determine a loss value; and adjusting the parameters of the preset 3D reconstruction model based on the loss value and performing iterative training to obtain the target 3D reconstruction model (see Fig. 1A, “3D refinement”).
Regarding claim 20, Smith discloses an image processing device, comprising: an acquisition module (e.g. 2D Body images 202-1-202-N in Fig. 1A), configured to obtain at least two images of a target object to be Client processed, wherein the at least two images include images captured from at least two different shooting angles (e.g. front, side and back of human body); a reconstruction module, configured to input the at least two images to be processed into a target 3D reconstruction model to reconstruct a target 3D body shape corresponding to the target object (Fig. 2A, “3D modeling 210”, “3D Model Refinement 208), wherein the target 3D reconstruction model is trained based on M preset dimensional key points, and M is a positive integer; and an output module (see Fig. 1B, key points 151-1…151-18), configured to output the target 3D body shape (Fig. 2A, 3D body 214)..
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 5 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Smith et al. (Patent No.: US 11,854,146) in view of Kamiyama et al. (Pub. No. US 2020/0319015).
Regarding claim 5, Smith fails to disclose the method according to claim 3, wherein fusing the at least two sets of the semantic features to obtain the global feature comprises: inputting the at least two sets of semantic features as at least two sequences into a feature fusion module to obtain the global feature, wherein the target 3D reconstruction model comprises the feature fusion module. However, Kamiyama is cited to teach converting key point annotations in the 2D images to 3D body part volume estimates by using machine learning (see para# 0140 of Kamiyama) which is similar to Smith. Kamiyama further teaches the neural networks including data fusion module (see para # 0106).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Smith with the features of fusion module as taught by Kamiyama so as to enhance the machine learning for 2D body images.
Regarding claim 16, it has the similar scope as claim 5. Thus, claim 16 is rejected for the same reason as claim 5 above.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Chen et al. (Patent No.: US 10,796,480) is cited teach a method of an image file of a personalized 3D head model of a user, the method comprising the steps of: (i) acquiring at least one 2D image of the user's face; (ii) performing automated face 2D landmark recognition based on the at least one 2D image of the user's face; (iii) providing a 3D face geometry reconstruction.
Agrawal et al. (Patent No. : US 11,861,860) is cited to teach systems and methods to determine one or more body dimensions of a body based on a processing of one or more two-dimensional images that include a representation of the body. Body dimensions include any length, circumference, etc., of any part of a body, such as shoulder circumference, chest circumference, waist circumference, hip circumference, inseam length, bicep circumference, leg circumference, etc.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to XIAO M WU whose telephone number is (571)272-7761. The examiner can normally be reached Monday to Friday 7:30am to 4pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexander Beck can be reached at 571-272-3750. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/XIAO M WU/Supervisory Patent Examiner, Art Unit 2613