Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claim 13 objected to because of the following informalities:
Claim 13 Line 1 recites, “The method of claim 11, where in the global model is an encoder-decoder network . . .” Even though, it can be inferred that the “global model” referenced in Claim 13 is referring to the “global machine learning model” mentioned within the base claim of Claim 11, to keep the limitations clear and consistent, the term should be changed to “The method of claim 11, where in the global machine learning model is an encoder-decoder network . . . “ Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim 20 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter.
Claim 20 recites a computer readable medium. The broadest reasonable interpretation of a claim drawn to a computer readable medium (also called computer readable storage medium, machine readable medium and other such variations) typically covers forms of non-transitory tangible media and transitory propagating signals per se in view of the ordinary and customary meaning of computer readable media, particularly when the specification is silent. See MPEP 2111.01. When the broadest reasonable interpretation of a claim covers a signal per se, the claim must be rejected under 35 U.S.C. 101 as covering non-statutory subject matter. The USPTO recognizes that applicants may have claims directed to computer readable storage media that cover signals per se, which the USPTO must reject under 35 U.S.C. 101 as covering both non-statutory subject matter and statutory subject matter. A claim drawn to such a computer readable medium that covers both transitory and non-transitory embodiments may be amended to narrow the claim to cover only statutory embodiments to avoid a rejection under 35 U.S.C. $ 101 by adding the limitation "non-transitory" to the claim. Such an amendment would typically not raise the issue of new matter, even when the specification is silent because the broadest reasonable interpretation relies on the ordinary and customary meaning that includes signals per se.
Applicant's specification in Page 9 Lines 32 – Page 10 Line 4 recites, “Suitably, the computer program is stored on a non-transitory carrier medium in machine or device readable form, such as in the form of a computer readable medium. In examples, the medium can be one or more of a solid-state memory, magnetic memory such as disk or tape, optically or magneto-optically readable memory such as compact disk or digital versatile disk etc., and the processing device utilises the program or a part thereof to configure it for operation. The computer program may be supplied from a remote source embodied in a communications medium such as an electronic signal, radio frequency carrier wave or optical carrier wave. Such carrier media are also envisaged.” Based on the specifications, it is not very clear that CRM excludes communications mediums as one can interpret that communication mediums is part of the CRM. Thus, having “non-transitory” added to the CRM is needed.
As an additional note, a non-transitory computer readable storage medium having executable programming instructions stored thereon is considered statutory as non-transitory computer readable media excludes transitory data signals.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 11-12, 17, 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (WO 2021184933 A1) (Hereinafter referred to as Chen) in view of Liu et al. (“Interactive 3D Modeling with a Generative Adversarial Network”) (Hereinafter referred to Liu).
Regarding Claim 11, Chen discloses A computer-implemented method for 3D reconstruction of a physical object, the method comprising: (See Abstract, “A three-dimensional human body model reconstruction method”)
obtaining an initial global reconstruction of the physical object in a 3D space, inferred by a global machine learning model; (See Page 18 Paragraph 6, “S404: Input the mask or the target image into the reconstruction module, and perform three-dimensional reconstruction on the target image or the mask to obtain a total human body three-dimensional model used to describe the characteristics of the entire human body.”
Also see Page 18 Paragraph 8, “The pre-trained third neural network model is used to obtain a three-dimensional model of the total human body according to the input mask or target image. The total human body three-dimensional model refers to a complete human body three-dimensional model, or refers to a human body three dimensional model including at least two human body features.” Note that the total human body three-dimensional model would correspond to “an initial global reconstruction of the physical object in 3D space” and that the pre-trained third neural network model corresponds to the “global machine learning model”.)
providing, to a user, an initial visualisation of the physical object based on the reconstruction; (See Page 13 Paragraph 5, “Finally, the 1/0 interface 112 returns the processing result, such as the mask, semantic segmentation mask, total human body three-dimensional model or sub-human three-dimensional model obtained as described above, to the client device 140 for subsequent processing or directly presented to the user.”)
receiving, from the user, at least one indication of at least one point of interest (See Page 3 Paragraph 4, “Human body features with different semantics can distinguish human body features well. Users can determine which semantic human body features have high requirements for high resolution according to their needs, and which semantic human body features do not require high resolution.”
Also see Page 20 Paragraph 4, “In a possible implementation, users have the highest resolution requirements for the head human body 3D model, followed by the resolution requirements of the hand human body 3D model. The resolution requirement of the three-dimensional human body model of the head and other parts of the human body is the lowest.”
In this case, users can indicate which body part has a higher resolution. This can be considered as a user indicating a point of interest.)
resampling at least one first subsection of the physical object based on the at least one point of interest to obtain local data, wherein the local data is associated with the subsection based on spatial information that associates the local data with a point in 3D space; (See Page 2-3 Paragraph 9, “Specifically, after acquiring a target image that includes human body characteristics that is captured immediately or selected from stored pictures. . . perform three-dimensional reconstruction on the target image or mask to obtain a total human body three-dimensional model, which is used to describe the characteristics of the entire human body; perform semantic segmentation on the target image to obtain a human body with different semantics Feature semantic segmentation image, . . . perform three-dimensional reconstruction on the semantic segmentation image or semantic segmentation mask to obtain a sub-human three-dimensional model. . . where the resolution of the total human three-dimensional model is different from that of the sub-human three-dimensional model Therefore, the different feature parts of the fused human body three-dimensional model have different resolutions.” Here, Chen teaches that 3D reconstruction is performed on a target image and that one can perform semantic segmentation on the target image to get a semantic segmentation image (local data) which features a human body part.
Also see Page 20 Paragraph 4, “In a possible implementation, users have the highest resolution requirements for the head human body 3D model, followed by the resolution requirements of the hand human body 3D model. The resolution requirement of the three-dimensional human body model of the head and other parts of the human body is the lowest.”
In summary, Chen teaches to perform 3D reconstruction on a target image to obtain a total body 3D model, performing semantic segmentation on the target image to obtain semantic segmentation images that contain specific body parts, and having the user be able to indicate which body parts should have what resolution. These teaching closely resemble the limitation of resampling the physical object according to the points of interest, as Chen teaches a user indicating body parts which are specifically segmented and reconstructed based on higher or lower resolution requirements.)
inputting the resampled local data and spatial information into a local feature machine learning model to obtain at least one local 3D reconstruction of the physical object, wherein the local feature machine learning model has been trained to output a physical object reconstruction from local data of resampled subsections, (See Page 19 Paragraph 8, “S405: Input the semantic segmentation mask or the semantic segmentation image into the reconstruction module, and perform three-dimensional reconstruction on the semantic segmentation image or the semantic segmentation mask to obtain a sub-human body three-dimensional model used to describe human body features with different semantics.” Also see Page 19 Paragraph 9, “In a possible implementation, the reconstruction module may be a pre-trained third neural network model.”
Note that the sub-human body three-dimensional model corresponds to “at least one local 3D reconstruction of the physical object”. Also note that although Chen doesn’t explicitly teach to use a different machine learning model trained specifically for subsections, but instead uses the same model for reconstructing the whole human body, since Chen does teach that the pre-trained third neural network is capable of performing the functions of both the global and local feature machine learning models, as the pre-trained third neural network can reconstruct both the whole human body or just a body part, then the above claim limitations are still taught, as the pre-trained third neural network can be considered as both a global and local feature machine learning model. Note that the training of the pre-trained third neural network is implied.)
and wherein the 3D coordinate system of the local 3D reconstruction aligns with the global 3D reconstruction; and merging the global 3D reconstruction with the local 3D reconstruction. (See Page 21 Paragraph 8, “S407: Suture the total human body 3D model and the sub-human body 3D model to obtain a multi-resolution fused human body 3D model.” Note that the merging of the models would implicitly need the alignment of 3D coordinate systems.)
However, Chen fails to explicitly disclose receiving, from the user, at least one indication of at least one point of interest in the visualisation;
Liu additionally teaches obtaining an initial global reconstruction an object in a 3D space, inferred by a global machine learning model, providing, to a user, an initial visualisation of the object based on the reconstruction (See Page 1 Fig. 1, “Figure 1. Interactive 3D modeling with a GAN. The user iteratively makes edits to a voxel grid with a simple painting interface and then hits a SNAP command to refine the current shape. The SNAP command projects the current shape into a latent vector shape manifold learned with a GAN, and then generates a new shape with the generator network. SNAP aims to increase the realism of the user’s input, while maintaining similarity.”
Also see Page 8 Fig. 13 row 1 which is an input voxel based 3D model of a user created object. Then see Row 2 which is the model after it was reconstructed and remodeled by the GAN. Thus, the 3D models shown in row 2 can be considered as the initial visualization provided to the user.)
receiving, from the user, at least one indication of at least one point of interest in the visualisation; (See Page 1 Fig. 1, “Figure 1. Interactive 3D modeling with a GAN. The user iteratively makes edits to a voxel grid with a simple painting interface and then hits a SNAP command to refine the current shape. The SNAP command projects the current shape into a latent vector shape manifold learned with a GAN, and then generates a new shape with the generator network. SNAP aims to increase the realism of the user’s input, while maintaining similarity.”
Also see Page 8 Fig. 13, “The user paints an initial shape (top) and then alternates between snapping it (solid arrows) and adding/removing voxels (dotted arrows). After each snap, the resulting object conforms roughly to the specifications of the user.” Further see row 3 showing that the user can edit the 3D model of the initial visualization in row 2 by adding or removing voxels. This would correspond to the idea of a user indicating points of interest in the visualization. Note that after making the edits, Liu teaches to using a GAN (a machine learning model) to further reconstruct the object.
Most importantly, Liu teaches the idea of a user interactive remodeling process which allows users to edit localized portions of the 3D model. Even though Chen does teach users being able to reconstruct body parts to have different resolutions, Chen did not have the exact mechanism of an editing process of the user indicating which portions to edit based on a presented visualization. Liu supplements Chen with this idea of incorporating the user being able to editing and resample a visualized model.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Chen with Liu to include allowing a user to indicate points of interest in a provided visualization.
The motivation to combine Chen with Liu would have been obvious as both arts are within the field of 3D modeling with machine learning models (See Liu Page 1 Fig. 1). The benefit would have been to meet a demand for interactive tools which allows for users to create 3D models (See Liu Page 1 Left Column Introduction Section, Paragraph 1). Note that the idea of presenting a visualization of a 3D model for further editing is already a well-known and common technique (See Liu Page 8 Fig. 13) and thus applying it to Chen would have been obvious for someone of ordinary skill in the art.
Regarding Claim 12, Chen in view of Liu disclose The method of claim 11, wherein the steps of the method are performed iteratively using the merged reconstruction as the initial global reconstruction in the next iteration until receiving, from the user, an indication to stop. (See Chen Page 20 Paragraph 4, “In a possible implementation, users have the highest resolution requirements for the head human body 3D model, followed by the resolution requirements of the hand human body 3D model. The resolution requirement of the three-dimensional human body model of the head and other parts of the human body is the lowest.” In this possible implementation, the head, hand, and body have different resolution. This would imply that the steps of segmenting the body part for reconstruction can be performed interactively until the user’s requirements are satisfied.
See Liu Page 1 Fig. 1, “Figure 1. Interactive 3D modeling with a GAN. The user iteratively makes edits to a voxel grid with a simple painting interface and then hits a SNAP command to refine the current shape. The motivation to combine would have been similar to that of Claim 11 rejection motivation.)
Regarding Claim 17, Chen in view of Liu disclose The method of claim 11 wherein resampling the physical object based on the at least one point of interest comprises resampling a subspace in the global 3D space centred on the at least one point of interest. (See Chen Page 3 Paragraph 4 teaching that users can indicate what human body features to have high requirements for resolution. Also see Chen Page 19 Paragraph 8 teaching semantic segmentation and performing 3D reconstruction to obtain a sub-human body 3D model. Lastly see Liu Fig. 1 and Fig. 13 teaching that a user can interactively make edits by adding or removing voxels from the mode.
Thus, with the combination of Chen and Liu, the user can interactively indicate a part of the human body to have a higher resolution, and thus have that body part reconstructed with a higher resolution by segmenting that part from the target image (resampling a subspace in the global 3D space centred on the point of interest). The motivation to combine would have been similar to that of Claim 11 rejection motivation.)
Regarding Claim 19, Chen in view of Liu disclose A computer system comprising a processor and a memory storing instructions executable by the processor to cause the processor to perform the method of claim 11. (See Chen Page 6 Paragraph 8, “The device for reconstructing a three-dimensional human body model includes a memory and a processor; the memory is coupled to the processor; the memory is used to store computer program code, the The computer program code includes computer instructions; when the processor executes the computer instructions . . .” Also note that the above limitations are similar to those of Claim 11 and is therefore rejected under a similar rationale as that of Claim 11.)
Regarding Claim 20, Chen in view of Liu disclose A computer readable medium comprising computer program code to, when loaded on and executed by a computer, causes the computer to carry out the method of claim 11. (See Chen Page 6 Paragraph 10, “In a fifth aspect of the present application, a computer-readable storage medium is provided. The computer-readable storage medium includes computer instructions. When the computer instructions are executed on a human body three-dimensional model reconstruction device, . . .” Also note that the above limitations are similar to those of Claim 11 and is therefore rejected under a similar rationale as that of Claim 11.)
Allowable Subject Matter
Claims 13-16 and 18 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Regarding Claim 13, the cited prior art does not disclose or render obvious the combination of elements cited in the claims as a whole. Specifically, the cited prior art fails to disclose or render obvious the limitations: wherein the global model is an encoder-decoder network comprising a global encoder trained to infer a global latent code from the physical object data, and wherein the local feature machine learning model comprises a local feature encoder-decoder network, the local feature encoder-decoder network comprising: a local feature encoder trained to infer a local feature latent code from the resampled local data and spatial information; and a local feature decoder trained to infer a representation of the physical object in the 3D space from a combination of the local feature latent code and the global latent code. Thus, Claim 13 contains allowable subject matter.
Regarding Claims 14-16, Claims 14-16 are dependent upon the base claim of Claim 13 and thus would also contain allowable subject matter.
Regarding Claim 18, the cited prior art does not disclose or render obvious the combination of elements cited in the claims as a whole. Specifically, the cited prior art fails to disclose or render obvious the limitations: wherein the local and global 3D reconstructions are each a scalar field representing an occupation probability of a point in space and wherein merging the global 3D reconstruction with the local 3D reconstruction comprises: combining the scalar field values of the global 3D reconstruction with the scalar field values of the local 3D reconstruction; and, optionally, extracting a probability iso-surface from the combined scalar field to represent the shape of the physical object for visualisation. Thus, Claim 18 contains allowable subject matter.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Reference Chen et al. (US 20210286977 A1) is made of record as an art that teaches generating a three-dimensional face model based on a face image using a face model generation model (See Fig. 4), performing feature extraction on the face image extracting global and local features (See [0047]), wherein the feature extraction can be performed by using an encoder and results in global and local feature vectors (See [0059]), and a decoder used to calculated a 3D face model parameter (See [0059]), the 3D face model parameter is then used in the face model generation model to output a 3D face (See Fig. 4).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to THANG G HUYNH whose telephone number is (571)272-5432. The examiner can normally be reached Mon-Thu 7:30am-4:30pm EST | Fri 7:30am-11:30am EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kee Tung can be reached at (571)272-7794. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/T.G.H./Examiner, Art Unit 2611
/KEE M TUNG/Supervisory Patent Examiner, Art Unit 2611