DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 01/18/2024 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement has been considered by the examiner.
Specification
The lengthy specification has not been checked to the extent necessary to determine the presence of all possible minor errors. Applicant’s cooperation is requested in correcting any errors of which applicant may become aware in the specification.
Claim Rejections - 35 USC § 102
1 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
2 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
3 Claim(s) 1-3, 7, 9, 12-14, 18, and 20 is/are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Chen et al. (US 20230274492 A1).
4 Regarding claim 1, Chen teaches a computer-implemented method to perform image synthesis, the method comprising:
obtaining, by a computing system comprising one or more computing devices, data descriptive of a pose from which to render a synthetic image of an object ([Abstract] reciting “A trained network can generate a basis shared by all shape textures, and can predict input-specific coefficients to construct the output texture for each shape as a linear combination of the basis images, then deform the texture to match the pose of the input.”; [0041] reciting “The network can be trained in stages, such as three stages, due at least in part to a trade-off between the quality of the texture alignment and the level of distortion. In some cases, aligning textures may involve requires heavy distortion in the texture image, such as when aligning a sedan with a van. In many instances, a reduced amount of distortion may be desirable as it can result in a reduced aliasing effect when rendering the textures, and can simplify post-processing.”);
generating, by the computing system using an image generation model, a three-dimensional location for each of a plurality of pixels of the synthetic image of the object ([0024] reciting “In at least one embodiment, content synthesis can be used to generate new instances of content, such as new two-dimensional (2D) or three-dimensional (3D) views or representations of new objects, or objects that are based upon combinations of features or aspects of one or more other objects…”; [0030] reciting “The encoder 214 can extract features from the input image 212 and generate one or more encodings or embeddings, as may correspond to one or more vectors or points in a latent space. In this example, the encoder 214 can predict the coefficients 218 to use to weight the basis images in order to produce the output image, or determine the output colors 220 for pixels of the output image.”);
mapping, by the computing system using a machine-learned correspondence network, the three-dimensional location of each pixel to a two-dimensional coordinate in a two-dimensional canonical coordinate space ([Abstract] reciting “Approaches presented herein can utilize a network that learns to embed three-dimensional (3D) coordinates on a surface of one or more 3D shapes into an aligned two-dimensional (2D) texture space, where corresponding parts of different 3D shapes can be mapped to the same location in a texture image.;”);
retrieving, by the computing system, a texture value from a set of texture data for each pixel of the synthetic image based on the two-dimensional coordinate for such pixel in the two-dimensional canonical coordinate space ([0044] reciting “These coordinate mappings and feature encoding(s) can be used to generate 508 a 2D texture image that corresponds to the first object and includes data values for at least some of the extracted features. For a front texture image, the texture image can include data values for features visible from a single point of view (e.g., the entire front of the object), and a back texture image can include data values visible from an opposite point of view, such that the pair of front and back texture images include data values representing a 360 degree view of the first object.”); and
rendering, by the computing system, the synthetic image of the object using the retrieved texture values for the plurality of pixels ([Abstract] reciting “Alignment can be performed using a texture alignment module that generates a set of basis images for synthesizing textures. A trained network can generate a basis shared by all shape textures, and can predict input-specific coefficients to construct the output texture for each shape as a linear combination of the basis images, then deform the texture to match the pose of the input.”).
5 Regarding claim 2, Chen teaches the computer-implemented method of claim 1 (see claim 1 rejection above), wherein the set of texture data is editable and has been edited by a user ([0024] reciting “Further, such synthesis enables specific types of objects to be modified to have specific visual appearances or aspects without the need for an artist to manually perform any part of the generation process. Such a synthesis process can also provide a user with the ability to quickly modify the appearance of an object, or generate new objects, by changing the reference objects used in the process.”; [0160] reciting “In at least one embodiment, user 1510 may interact with a GUI via computing device 1508 to edit or fine-tune (auto)annotations. In at least one embodiment, a polygon editing feature may be used to move vertices of a polygon to more accurate or fine-tuned locations.”).
6 Regarding claim 3, Chen teaches the computer-implemented method of claim 1 (see claim 1 rejection above), further comprising generating, by the computing system, the set of texture data from one or more input images of the object ([Abstract] reciting “Alignment can be performed using a texture alignment module that generates a set of basis images for synthesizing textures. A trained network can generate a basis shared by all shape textures, and can predict input-specific coefficients to construct the output texture for each shape as a linear combination of the basis images, then deform the texture to match the pose of the input.”), wherein generating the set of texture data comprises:
obtaining, by the computing system, the image generation model ([0009] reciting “FIG. 6 illustrates components of a distributed system that can be utilized to train or perform inferencing using a generative model, according to at least one embodiment”);
training, by the computing system using the one or more input images, the image generation model to generate synthetic images of the object ([Abstract] reciting “A trained network can generate a basis shared by all shape textures, and can predict input-specific coefficients to construct the output texture for each shape as a linear combination of the basis images, then deform the texture to match the pose of the input.”; [0023] reciting “…a trained machine learning model can take, as input, a three-dimensional 3D digital or virtual object and can generate one or more texture images (e.g., front and back texture images) that represent the texture (e.g., visual features) extracted from the object.”);
generating, by the computing system using the image generation model, one or more views of the object from one or more poses, wherein a set of three-dimensional points is associated with each of the one or more views ([0024] reciting “In at least one embodiment, content synthesis can be used to generate new instances of content, such as new two-dimensional (2D) or three-dimensional (3D) views or representations of new objects, or objects that are based upon combinations of features or aspects of one or more other objects, or the same object in a different pose or state…These generated or synthesized views or representations can be used for many different applications or use cases, such as for inclusion in video games, animation, virtual/augmented/enhanced reality applications, or environment simulation, as may be useful for training robotic, autonomous, or security systems.”);
training, by the computing system, the correspondence network to map from three-dimensional space to the two-dimensional canonical coordinate space based on the one or more views of the object (see previous statements for claim 3 for similar rejections); and
using, by the computing system, the trained correspondence network to extract the set of texture data from the one or more views or the one or more input images, wherein the set of texture data is expressed in the two-dimensional canonical coordinate space ([0023] reciting “a trained machine learning model can take, as input, a three-dimensional 3D digital or virtual object and can generate one or more texture images (e.g., front and back texture images) that represent the texture (e.g., visual features) extracted from the object. A texture alignment module can be used to generate a set of basis images for synthesizing textures.”).
7 Regarding claim 7, Chen teaches the computer-implemented method of claim 3, wherein (see claims 1 and 3 rejections above):
generating, by the computing system using the image generation model, the one or more views of the object from one or more poses comprises generating, by the computing system using the image generation model, multiple views of the object from multiple poses ([0024] reciting “In at least one embodiment, content synthesis can be used to generate new instances of content, such as new two-dimensional (2D) or three-dimensional (3D) views or representations of new objects, or objects that are based upon combinations of features or aspects of one or more other objects, or the same object in a different pose or state (e.g., a person having a shape of an adult rather than a child, or having changed shape through muscle gain or weight loss, among other such options). These generated or synthesized views or representations can be used for many different applications or use cases”);
training, by the computing system, the correspondence network to map from three-dimensional space to the two-dimensional canonical coordinate space based on the one or more views of the object comprises training, by the computing system, the correspondence network to map from three-dimensional space to the two-dimensional canonical coordinate space based on the multiple views of the object ([0024] reciting “These generated or synthesized views or representations can be used for many different applications or use cases, such as for inclusion in video games, animation, virtual/augmented/enhanced reality applications, or environment simulation, as may be useful for training robotic, autonomous, or security systems.”; [0027] reciting “For example, a neural network can be trained and/or used to embed multi-dimensional (e.g., three-dimensional (3D)) coordinates on a surface of one or more multi-dimensional shapes into an aligned two dimensional (2D) space, such as a UV space where U and V denote axes of a texture (to differentiate from X, Y, and Z coordinates in model space). Prior approaches to texture representation for 3D shapes, in operations such as texture transfer and synthesis, applied spherical texture maps, which may lead to heavy distortion, or used continuous texture fields that yield smooth outputs lacking details.”); and
using, by the computing system, the trained correspondence network to extract the set of texture data from the one or more views or the one or more input images comprises using, by the computing system, the trained correspondence network to extract the set of texture data from the multiple views ([0030] reciting “The encoder 214 can extract features from the input image 212 and generate one or more encodings or embeddings, as may correspond to one or more vectors or points in a latent space.”).
8 Regarding claim 9, Chen teaches the computer-implemented method of claim 3 (see claims 1 and 3 rejections above), wherein using, by the computing system, the trained correspondence network to extract the set of texture data from the one or more views or the one or more input images comprises using, by the computing system, the trained correspondence network to extract the set of texture data from both the one or more views and the one or more input images ([Abstract] reciting “A trained network can generate a basis shared by all shape textures, and can predict input-specific coefficients to construct the output texture for each shape as a linear combination of the basis images, then deform the texture to match the pose of the input.”; [0023] reciting “In at least one embodiment, a trained machine learning model can take, as input, a three-dimensional 3D digital or virtual object and can generate one or more texture images (e.g., front and back texture images) that represent the texture (e.g., visual features) extracted from the object.”; [0024] reciting “In at least one embodiment, content synthesis can be used to generate new instances of content, such as new two-dimensional (2D) or three-dimensional (3D) views or representations of new objects, or objects that are based upon combinations of features or aspects of one or more other objects, or the same object in a different pose or state (e.g., a person having a shape of an adult rather than a child, or having changed shape through muscle gain or weight loss, among other such options). These generated or synthesized views or representations can be used for many different applications or use cases, such as for inclusion in video games, animation, virtual/augmented/enhanced reality applications, or environment simulation, as may be useful for training robotic, autonomous, or security systems.”).
9 Claim 12 has similar limitations as of claim 1, therefore it is rejected under the same rationale as claim 1.
10 Claim 13 has similar limitations as of claim 2, therefore it is rejected under the same rationale as claim 2.
11 Claim 14 has similar limitations as of claim 3, therefore it is rejected under the same rationale as claim 3.
12 Claim 18 has similar limitations as of claim 7, therefore it is rejected under the same rationale as claim 7.
13 Claim 20 has similar limitations as of claim 9, therefore it is rejected under the same rationale as claim 9.
Claim Rejections - 35 USC § 103
14 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
15 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
16 Claim(s) 4-6 and 15-17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (US 20230274492 A1) in view of Xu et al. (US 20240265621 A1).
17 Regarding claim 4, Chen teaches the computer-implemented method of claim 1 (see claim 1 rejection above), but does not explicitly teach wherein the image generation model comprises a neural radiance field (NERF) model.
18 Xu teaches wherein the image generation model comprises a neural radiance field (NERF) model ([0003] reciting “In one example embodiment disclosed and recited herein, a method for generating controllable three-dimensional (3D) synthesized images includes conditioning 3D representations of a head geometry in a canonical space based on feature vectors and camera viewing parameters”; [0045] reciting “Volume renderer 150 may be programmed, designed, or otherwise configured to receive the feature map from decoder 145 and then map the decoded neural radiance field from the canonical space to the observation space, i.e., guided by the volumetric correspondence, mapping to output the geometry and image in the observation space.”).
19 It would have been obvious to one with ordinary skill before the effective filing date of the claimed invention, to have modified the method (taught by Chen) to incorporate the teachings of Xu to provide a method that incorporates a type of neural radiance that can function with a type of image generation model/method, like the model taught by Chen. Doing so would provide various synthesis with full control on camera poses, facial expressions, head shape, articulated neck and jaw poses as stated by Xu ([Abstract] recited).
20 Regarding claim 5, Chen teaches the computer-implemented method of claim 1 (see claim 1 rejection above), but does not explicitly teach wherein the image generation model comprises a tri-plane representation.
21 Xu teaches wherein the image generation model comprises a tri-plane representation ([0003] reciting “In one example embodiment disclosed and recited herein, a method for generating controllable three-dimensional (3D) synthesized images includes conditioning 3D representations of a head geometry in a canonical space based on feature vectors and camera viewing parameters”; [0042] reciting “Generator 135 may refer to an image generator, as part of EG3D framework 115, which be programmed, designed, or otherwise configured to generate tri-plane neural volume representations of a human head…”).
22 It would have been obvious to one with ordinary skill before the effective filing date of the claimed invention, to have modified the method (taught by Chen) to incorporate the teachings of Xu to provide a method that incorporates a type of tri-plane representation that can function with a type of image generation model/method, like the model taught by Chen. Doing so would provide various synthesis with full control on camera poses, facial expressions, head shape, articulated neck and jaw poses as stated by Xu ([Abstract] recited).
23 Regarding claim 6, Chen teaches the computer-implemented method of claim 1 (see claim 1 rejection above), but does not explicitly teach wherein the image generation model is trained using generative latent optimization.
24 Xu teaches wherein the image generation model is trained using generative latent optimization ([0003] reciting “In one example embodiment disclosed and recited herein, a method for generating controllable three-dimensional (3D) synthesized images includes conditioning 3D representations of a head geometry in a canonical space based on feature vectors and camera viewing parameters”; [0057] reciting “However, such design results in entanglement of expression control with the generative latent code z, inducing potential identity or appearance changes when varying expressions.”; [0068] reciting “The embodiments described and recited herein also support 3D-aware face reenactment of a single-view portrait to a video sequence. To that end, optimization may be performed in a latent Z+ space to find the corresponding latent embedding z, with FLAME parameter and camera pose estimated from the input portrait.”).
25 It would have been obvious to one with ordinary skill before the effective filing date of the claimed invention, to have modified the method (taught by Chen) to incorporate the teachings of Xu to provide a method that incorporates a type of generative latent representation that can function with a type of image generation model/method, like the model taught by Chen. Doing so would provide various synthesis with full control on camera poses, facial expressions, head shape, articulated neck and jaw poses as stated by Xu ([Abstract] recited).
26 Claim 15 has similar limitations as of claim 4, therefore it is rejected under the same rationale as claim 4.
27 Claim 16 has similar limitations as of claim 5, therefore it is rejected under the same rationale as claim 5.
28 Claim 17 has similar limitations as of claim 6, therefore it is rejected under the same rationale as claim 6.
29 Claim(s) 8 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (US 20230274492 A1) in view of Du et al. (US 20220157017 A1).
30 Regarding claim 8, Chen teaches the computer-implemented method of claim 7 (see claims 1, 3, and 7 rejections above), wherein the multiple views comprise a frontal view ([0023] reciting “In at least one embodiment, a trained machine learning model can take, as input, a three-dimensional 3D digital or virtual object and can generate one or more texture images (e.g., front and back texture images) that represent the texture (e.g., visual features) extracted from the object.”),
31 Chen does not explicitly teach wherein the multiple views comprise a frontal view, a left view, a right view, a top view, and a bottom view.
32 Du teaches wherein the multiple views comprise a frontal view, a left view, a right view, a top view, and a bottom view ([0019] reciting “Here, the two-dimensional images corresponding to the at least two viewing angles may include a front view, a top view, a left side view, a right side view, and a bottom view, etc., as well as a plan view, an elevation view, an oblique view, a perspective view, a cross-sectional view, an exploded view, a partial view, an enlarged view and so on.”).
33 It would have been obvious to one with ordinary skill before the effective filing date of the claimed invention, to have modified the method (taught by Chen) to incorporate the teachings of Du to provide a method that incorporates many other types of views to go along with the frontal views provided by Chen for the various image generation models. Doing so would allow other specific views like plan view, an elevation view, an oblique view, a perspective view, a cross-sectional view, an exploded view, a partial view, an enlarged view and so on as stated by Du ([0019] recited).
34 Claim 19 has similar limitations as of claim 8, therefore it is rejected under the same rationale as claim 8.
35 Claim(s) 10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (US 20230274492 A1) in view of Burns et al. (US 20160292907 A1).
36 Regarding claim 10, Chen teaches the computer-implemented method of claim 1 (see claim 1 rejection above), wherein retrieving, by the computing system, the texture value from the set of texture data for each pixel of the synthetic image based on the two-dimensional coordinate for such pixel in the two-dimensional canonical coordinate space comprises performing a in the two-dimensional canonical coordinate space ([0044] reciting “These coordinate mappings and feature encoding(s) can be used to generate 508 a 2D texture image that corresponds to the first object and includes data values for at least some of the extracted features. For a front texture image, the texture image can include data values for features visible from a single point of view (e.g., the entire front of the object), and a back texture image can include data values visible from an opposite point of view, such that the pair of front and back texture images include data values representing a 360 degree view of the first object.”).
37 Chen does not explicitly teach wherein retrieving, by the computing system, the texture value from the set of texture data for each pixel of the synthetic image based on the two-dimensional coordinate for such pixel in the two-dimensional canonical coordinate space comprises performing a nearest neighbor interpolation over multiple texture values retrieved from a neighborhood in the two-dimensional canonical coordinate space.
38 Burns teaches wherein retrieving, by the computing system, the texture value from the set of texture data for each pixel of the synthetic image based on the two-dimensional coordinate for such pixel in the two-dimensional canonical coordinate space comprises performing a nearest neighbor interpolation over multiple texture values retrieved from a neighborhood in the two-dimensional canonical coordinate space ([Abstract] reciting “In some embodiments, the graphics unit also includes texture processing circuitry configured to perform different types of interpolation for pixels in the group of pixels.”; [0049] reciting “As an illustration of the two-dimensional case, FIG. 5 shows an exemplary set of texels used to render the FIG. 3 element (a) relative to a plurality of pixels in a screen space. The dashed lines in FIG. 5 outline an exemplary area that falls within the interpolation width (e.g., based on parameter P) according to some embodiments of an adaptive interpolation technique. In some embodiments, adaptive interpolation utilizes a single piece-wise interpolation function that that degenerates to nearest-neighbor interpolation in some places (e.g., for end groups of pixels) and non-nearest-neighbor interpolation in other places (e.g., for an intermediate group of pixels such as those falling between the dashed lines in FIG. 5). Thus, in the illustrated embodiment, TPU 165 may implement an adaptive interpolation function that results in using a nearest-neighbor approach for pixel T and using a non-nearest-neighbor approach (e.g., bilinear) for pixel Q.”).
39 It would have been obvious to one with ordinary skill before the effective filing date of the claimed invention, to have modified the method (taught by Chen) to incorporate the teachings of Burns to provide a type of nearest neighbor interpolation for the texture values for the pixels, and 2d spaces that are provided by the teachings of Chen. Doing so would provide desired visual effect without visual artifacts as stated by Burns ([0006] recited).
40 Claim(s) 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (US 20230274492 A1) in view of Zhang et al. (US 20220394293 A1).
41 Regarding claim 11, Chen teaches the computer-implemented method of claim 1 (see claim 1 rejection above), but does not explicitly teach wherein the set of texture data is structured as a K-d tree and wherein retrieving, by the computing system, the texture value from the set of texture data for each pixel of the synthetic image based on the two-dimensional coordinate for such pixel in the two-dimensional canonical coordinate space comprises querying the K-d tree.
42 Zhang teaches wherein the set of texture data is structured as a K-d tree and wherein retrieving, by the computing system, the texture value from the set of texture data for each pixel of the synthetic image based on the two-dimensional coordinate for such pixel in the two-dimensional canonical coordinate space comprises querying the K-d tree ([0037] reciting “For example, the geometry image generation module 306 and the texture image generation module 308 may exploit the 3D to 2D mapping computed during the packing process of the patch packing module 304 to store the geometry and texture of the point cloud as images (a.k.a. layers).”; [0062] reciting “In the current design of TMC2, the recoloring process may be rather complex because a K-dimensional (KD)-tree data structure is utilized in the nearest neighbor search and the recoloring operation is applied to every point in the reconstructed point cloud. In embodiments, the entire recoloring process may be bypassed by generating texture maps from the original point cloud directly.”).
43 It would have been obvious to one with ordinary skill before the effective filing date of the claimed invention, to have modified the method (taught by Chen) to incorporate the teachings of Zhang to provide a method that incorporates k-dimensional trees to obtain various texture data for the synthetic images based on the 2d maps provided by the teachings of Chen. Doing so would allow various processes like recoloring to apply to larger geometry distortions as stated by Zhang ([0062] recited).
Conclusion
44 Any inquiry concerning this communication or earlier communications from the examiner should be directed to JOHNNY TRAN LE whose telephone number is (571)272-5680. The examiner can normally be reached Mon-Thu: 7:30am-5pm; First Fridays Off; Second Fridays: 7:30am-4pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kent Chang can be reached at (571) 272-7667. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JOHNNY T LE/ Examiner, Art Unit 2614
/KENT W CHANG/ Supervisory Patent Examiner, Art Unit 2614