DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Election/Restrictions
Applicant's election with traverse of Species I regarding an image editing method, represented through Claims 1-5 and 9-11 in the reply filed on February 12th, 2026 is acknowledged. The traversal is on the ground(s) that the Office Action mailed December 12th 2025 fails to establish a serious search and/or examination burden for Claims 1-12.
Species I, as depicted by aspects of Figures 1, 2A and 2B, described in Paragraphs 0018-0019, and defined respectively by Claims 1-5 and 9-11, describes an image editing method by deriving and editing latent representations of image data in a feature space for image reconstruction.
Species II, as depicted in Figure 3B, described in Paragraphs 0035-0041, and defined respectively by Claims 6-8 and 12, describes a UV map generation method in which images from multiple viewpoints are used to construct a texture representative of multiple viewpoints.
The species were previously considered by the examiner to be independent or distinct because the claims for the different species recite mutually exclusive characteristics of such species; Species I is an image editing method, whereas Species II is a UV map generation method. In addition, these species are not obvious variants of each other based on the current record.
The Applicant argues that the specification describes a single integrated “map and edit” approach in which latent representations are mapped between a first and second representation space to enable explicit, parameterized edits including view-angle based edits, and using the resulting multi-view images as a precursor to UV map generation; The Applicant subsequently argues that in view of the integrated framework, there would be substantial overlap because of the same core technology and components are implemented throughout, such as latent representations in multiple spaces, encoder and decoder mappings between those spaces, view-angled conditioned generation of multi-view images, and downstream construction of a UV map from those multi-view images.
The Examiner agrees with the Applicant’s argument of a single integrated “map and edit” approach, which would allow prior art for Species II to be potentially relevant to Species I. Therefore, Species I, consisting of Claims 1-5 and 9-11, and Species II, consisting of Claims 6-8 and 12, will be examined together. Species III, consisting of Claims 13-20, are deemed to be withdrawn.
The updated requirement is deemed proper and is therefore made FINAL.
Claim Objections
Claims 3, 7, 8, 9 and 11 are objected to because of the following informalities:
Regarding Claim 3, “wherein the generating the output image is further based on the second input image” should be “wherein the generating of the output image is further based on the second input image”. Appropriate correction is required.
Regarding Claim 7, “wherein the generating the final UV map” should be “wherein the generating of the final UV map”. Appropriate correction is required.
Regarding Claim 8, “wherein the generating the final UV map” should be “wherein the generating of the final UV map”. Appropriate correction is required.
Regarding Claim 9, “wherein the generating the plurality of latent representations associated with the plurality of view angles” should be “wherein the generating of the plurality of latent representations associated with the plurality of view angles”. Appropriate correction is required.
Regarding Claim 11, “wherein the generating the plurality of images” should be “wherein the generating of the plurality of images”. Appropriate correction is required.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-4 are rejected under 35 U.S.C. 103 as being unpatentable over Karras.A (A Style-Based Generator Architecture for Generative Adversarial Networks, 2019), in view of Huang (Arbitrary Style Transfer in Real-time with Adaptive Instance Normalization, 2017), and Park (US 20210358177 A1).
Regarding Claim 1, Karras.A teaches a method for image editing, the method comprising:
Receiving, via a data interface, an input image (Figure 3: “Two sets of images were generated from their respective latent codes (sources A and B)”. Refer to Figure 3 images, which show the respective images of the latent codes. A data interface, in its broadest reasonable interpretation, is any entity capable of data transfer, and is hence inherent with regards to the receiving of the input images related to sources A and B);
Generating, via a first encoder, a first latent representation in a first representation space based on the input image (Figure 1 illustrates that an initial latent representation z is mapped to an intermediary latent space, where latent representation z is in the latent space Z (image space). An encoder is inherent for retrieving the latent representation in the image space from an image);
Generating, via a second encoder, a second latent representation in a second representation space based on the first latent representation (Figure 1: “we first map the input to an intermediate latent space W”; Figure 1 clearly illustrates deriving intermediate latent representation w, in intermediate latent space W, from latent representation z, which is latent space Z).
Karras.A does not explicitly teach generating a third latent representation in the second representation space based on the second latent representation, generating a fourth latent representation in the first representation space based on the second latent representation via a first decoder, or generating a fifth latent representation in the first representation space based on the second latent representation via a second decoder, although it is heavily implicit.
However, Huang teaches generating a third latent representation in the second representation space based on the second latent representation (Figure 2: “An overview of our style transfer algorithm. We use the first few layers of a fixed VGG-19 network to encode the content and style images. An AdaIN layer is used to perform style transfer in the feature space. A decoder is learned to invert the AdaIN output to the image spaces”; Section 6.1, Architecture: “After encoding the content and style images in feature space” Notes: Similar to Karas, the latent representations in the first representation space correlate with the image space, and the latent representations in the second representation space are an intermediary feature space in which style transfer is performed. Huang clearly states that AdaIN (which is utilized implicitly in the same manner in Keras) is used for style transfer in the feature space, where an image is stylized according to a style image, as displayed in Figure 2. Note that the images (which have corresponding latent representations in the image space) are encoded into the feature space, where encoding into another space inherently produces a latent representation based on the previous latent representation);
Generating, via a first decoder, a fourth latent representation in the first representation space based on the second latent representation (Figure 2: “A decoder is learned to invert the AdaIN output to the image spaces”. Notes: While Huang doesn’t explicitly state that a fourth latent representation based on the second latent representation in the first representation space is generated, it is obvious that an image converted to the feature space (second latent representation) can be fed into AdaIN with its own image as a style image resulting in no style change, and decoded back into the image space (first representation space), resulting in the fourth latent representation);
Generating, via a second decoder, a fifth latent representation in the first representation space based on the third latent representation (Figure 2: “A decoder is learned to invert the AdaIN output to the image spaces”. Notes: fifth latent representation is the output of AdaIN decoded into the image space (first representation space), where the third latent representation is the AdaIN output before being decoded into the image space);
Karras.A and Huang are considered analogous in the art with respect to style transferring between images using AdaIN. The generation of a third latent representation and fourth and fifth latent representation as specified are heavily implicit in Karras.A; Huang definitively demonstrates that AdaIN is utilized in the art for generating a third latent representation in the second representation space (feature space) and fourth and fifth latent representation in the first representation space (image space). AdaIN is well known in the art for style transfer, as apparent in Karras.A and Huang.
Therefore, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the image editing method via AdaIN of Karras.A with the style transfer mechanisms involving third, fourth, and fifth latent representations via AdaIN of Huang; Doing so would yield the predictable result of producing latent representations representative of style transfer images and their associated features in both the image space and feature space.
Karras.A as modified does not teach computing a difference between the fourth latent representation and the fifth representation, generating a sixth latent representation based on the fourth latent representation and the difference, or generating, via a third decoder, an output image based on the sixth latent representation.
However, Park teaches computing a difference between the fourth latent representation and the fifth latent representation (Paragraph [0101]: “the deep image manipulation system can generate an attribute direction by determining a difference between the latent codes of the two groups of digital images”. Notes: the latent codes of the two groups of digital images are in the image space, and hence are analogous with the fourth and fifth latent representations of Karras.A);
Generating a sixth latent representation based on the first latent representation and the difference (Paragraph [0101]: “Moving a spatial code or a global code in the attribute direction increases the attribute, while moving a spatial code or global code in a direction opposite to the attribute direction reduces the presence of the attribute in a resulting image”. Notes: spatial code and global code are latent codes (representations) of an image); and
Generating, via a third decoder, an output image based on the sixth latent representation (Paragraph [0101]: “Moving a spatial code or a global code in the attribute direction increases the attribute, while moving a spatial code or global code in a direction opposite to the attribute direction reduces the presence of the attribute in a resulting image”. Notes: the third decoder is inherent to displaying an image given its corresponding latent representation in the image space).
Karras.A as modified and Park are considered analogous in the art with respect to modifying an image with respect to another image. A motivation for modifying an image (via its latent code) through a difference derived from the image and a second image is to demonstrate levels of change, as evident in Park. With regards to Karras.A as modified, an image may be heavily or less heavily influenced by the style of the style image utilizing the method of Park; one would be motivated to do so to obtain a stylized image that appeals to the viewer’s aesthetic or preference for style injection.
Therefore, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the image editing method of Karras.A as modified with the image editing method involving a difference between images of interest corresponding with a difference between the fourth and fifth latent representation, generation of a sixth latent representation, and image generation via a third decoder of Park; Doing so would yield the predictable result of producing stylized images of varying levels of style injection via the image editing method of Karras.A as modified.
Regarding Claim 2, the method of Claim 1 is rejected over Karras.A as modified.
Karras.A as modified teaches the first decoder being the same as the second decoder (Huang, Figure 2: “A decoder is learned to invert the AdaIN output to the image spaces”. Notes: While Huang doesn’t explicitly state that a fourth latent representation based on the second latent representation in the first representation space is generated, it is obvious that an image converted to the feature space (second latent representation) can be fed into AdaIN with its own image as a style image resulting in no style change, and decoded back into the image space (first representation space), resulting in the fourth latent representation. Hence, the first and second decoder are the same).
Regarding Claim 3, the method of Claim 1 is rejected over Karras.A as modified.
Karras.A as modified teaches receiving, via the data interface, a second input image (Huang, Figure 2 illustrates inputting a first and second input image); and
Inputting the second input image to the third decoder (Park, Paragraph [0101]: “the deep image manipulation system can generate an attribute direction by determining a difference between the latent codes of the two groups of digital images. Moving a spatial code or a global code in the attribute direction increases the attribute, while moving a spatial code or global code in a direction opposite to the attribute direction reduces the presence of the attribute in a resulting image”. Notes: the second input image isn’t directly inputted into the third decoder itself, but is used for calculating a difference between the first and second input image, which is utilized to generate the output image; thus, in its broadest reasonable interpretation, the second input serves as an input for the third decoder),
Wherein the generating the output image is further based on the second input image (Park, Paragraph [0101]: “the deep image manipulation system can generate an attribute direction by determining a difference between the latent codes of the two groups of digital images. Moving a spatial code or a global code in the attribute direction increases the attribute, while moving a spatial code or global code in a direction opposite to the attribute direction reduces the presence of the attribute in a resulting image”).
Regarding Claim 4, the method of Claim 1 is rejected over Karras.A as modified.
Karras.A as modified teaches generating the third latent representation by modifying at least one parameter associated with the second latent representation (Karras.A, Figure 3: “Copying the styles corresponding to coarse spatial resolutions (42– 82) brings high-level aspects such as pose, general hair style, face shape, and eyeglasses from source B”. Notes: The broadest reasonable interpretation of a parameter is an input that has control over an aspect of an output; in this case, the “style” is a particular subset (parameter) of the second latent representation that has control over a particular aspect, where when the value of the subset is changed, a corresponding change to the particular aspect is observable when the third latent representation is visualized).
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Karras.A (A Style-Based Generator Architecture for Generative Adversarial Networks, 2019), in view of Huang (Arbitrary Style Transfer in Real-time with Adaptive Instance Normalization, 2017), and Park (US 20210358177 A1), in further view of Karras.B (Supplemental Material: A style-Based Generator Architecture for Generative Adversarial Networks).
Regarding Claim 5, the method of Claim 4 is rejected over Karras.A as modified.
Karras.A as modified teaches generating a third latent representation.
Karras.A as modified does not explicitly teach modifying at least one parameter associated with the second latent representation.
However, Karras.B teaches a view angle parameter (Section 3, Other Datasets: “The accompanying video provides results for style mixing and stochastic variation tests. As can be seen therein, in case of BEDROOM the coarse styles basically control the viewpoint of the camera, middle styles select the particular furniture, and fine styles deal with colors and smaller details of materials. In CARS the effects are roughly similar. Stochastic variation affects primarily the fabrics in BEDROOM, backgrounds and headlamps in CARS, and fur, background, and interestingly, the positioning of paws in CATS”).
Karras.A as modified and Karras.B are considered analogous in the art, as Karras.B is the supplemental material of Karras.A. A motivation in the art for combining the view angle parameter of Karras.B with the general image editing method of Karras.A is to produce a plurality of view angles of an image.
Therefore, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the generation of the third latent representation by modifying parameters associated with the second latent representation in the feature space of Karras.A as modified with the explicit inclusion of view angle as a parameter of Karras.B; Doing so would yield the predictable result of a comprehensive image editing method including adjusting the view angle.
Claim 6 and 9-12 are rejected under 35 U.S.C. 103 as being unpatentable over Karras.A (A Style-Based Generator Architecture for Generative Adversarial Networks, 2019), in view of Huang (Arbitrary Style Transfer in Real-time with Adaptive Instance Normalization, 2017), Park (US 20210358177 A1), and Karras.B (Supplemental Material: A style-Based Generator Architecture for Generative Adversarial Networks, 2019), in view of Na (Facial UV map completion for pose-invariant face recognition: a novel adversarial approach based on coupled attention residual UNets, 2020).
Regarding Claim 6, Karras.A as modified teaches a method comprising:
Receiving, via a data interface, an input image (Karras.A, Figure 3: “Two sets of images were generated from their respective latent codes (sources A and B)”. Refer to Karras.A, Figure 3 images, which show the respective images of the latent codes. A data interface, in its broadest reasonable interpretation, is any entity capable of data transfer, and is hence inherent with regards to the receiving of the input images related to sources A and B);
Generating, via an encoder, a first latent representation based on the input image (Karras.A, Figure 1 illustrates that an initial latent representation z is mapped to an intermediary latent space, where latent representation z is in the latent space Z (image space));
Generating, based on the first latent representation, a plurality of latent representations associated with a plurality of view angles (Karras.B, Section 3, Other Datasets: “The accompanying video provides results for style mixing and stochastic variation tests. As can be seen therein, in case of BEDROOM the coarse styles basically control the viewpoint of the camera” Notes: a plurality of latent representations can clearly be derived by adjusting the coarse styles resulting in different viewpoints, as the coarse styles are used to adjust the latent representation in the feature space);
Generating, via a decoder, a plurality of images in the plurality of view angles based on the plurality of latent representations (Karras.B, Section 3, Other Datasets: “The accompanying video provides results for style mixing and stochastic variation tests. As can be seen therein, in case of BEDROOM the coarse styles basically control the viewpoint of the camera”; Huang, Figure 2: “A decoder is learned to invert the AdaIN output to the image spaces”. Notes: as previously analyzed, adjusting style in the feature space using AdaIN is well established. The latent code in the feature space is adjusted with respect to the view angle, and decoded to image space. Lastly, a decoder is inherent to displaying an image from a corresponding latent representation in the image space).
Karras.A as modified does not teach generating a final UV map based on the plurality of images.
However, Na teaches generating a final UV map based on a plurality of images corresponding to different view angles (Figure 5: “The creation of ground-truth complete UV maps. Three facial images with yaw angles of O, -30, and +30 degrees, are fed to the 3DDFA model to create three incomplete UV maps which are then merged by Poisson blending to generate the ground-truth complete UV map”; Refer to Figure 5 for a visualization).
Karras.A as modified and Na are considered analogous in the art with respect to generating different view angle images from an input image. A common motivation in the art is to generate detailed and accurate UV maps for subsequent use with a 3D mesh to generate a 3D model. Utilizing multiple view angle images improves the quality of the subsequent final UV map; this is evident in Na. Hence, one would be motivated to utilize the method of Karras.A as modified to generate a plurality of images depicting different view angles, and use the derived images to generate a UV map as detailed in Na.
Therefore, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the image editing method resulting in multiple images from different view angles of Karras.A as modified with the use of multiple images from different view angles for generating a final UV map of Na.
Regarding Claim 9, the method of Claim 6 is rejected over Karras.A as modified.
Karras.A as modified teaches a method wherein the first latent representation is in a first representation space (Karras.A, Figure 1 illustrates that an initial latent representation z is mapped to an intermediary latent space, where latent representation z is in the latent space Z (image space));
Wherein the generating of the plurality of latent representations associated with the plurality of view angles (Karras.B, Section 3, Other Datasets: “The accompanying video provides results for style mixing and stochastic variation tests. As can be seen therein, in case of BEDROOM the coarse styles basically control the viewpoint of the camera” Notes: a plurality of latent representations can clearly be derived by adjusting the coarse styles resulting in different viewpoints, as the coarse styles are used to adjust the latent representation in the feature space) includes:
Generating, via a second encoder, a second latent representation in a second representation space based on the first latent representation (Karras.A, Figure 1: “we first map the input to an intermediate latent space W”; Karras.A, Figure 1 clearly illustrates deriving intermediate latent representation w, in intermediate latent space W (feature space), from latent representation z, which is latent space Z);
Generating a first plurality of edited latent representations in the second representation space based on the second latent representation (Huang, Figure 2: “An overview of our style transfer algorithm. We use the first few layers of a fixed VGG-19 network to encode the content and style images. An AdaIN layer is used to perform style transfer in the feature space”; Karras.B, Section 3, Other Datasets: “The accompanying video provides results for style mixing and stochastic variation tests. As can be seen therein, in case of BEDROOM the coarse styles basically control the viewpoint of the camera” Notes: a plurality of latent representations can clearly be derived by adjusting the coarse styles resulting in different viewpoints, as the coarse styles are used to adjust the latent representation in the feature space. As noted previously, AdaIN is utilized by both Karras.A, Karras.B (being supplementary material to Karras.A), and Huang, and is commonly used for style transfer. AdaIN serves the same purpose in all three references; Huang is utilized to more clearly demonstrate how AdaIN interacts with latent representations and their associated image and feature space);
Generating, via a second decoder, a second plurality of edited latent representations in the first representation space based on the first plurality of edited latent representations (Huang, Figure 2: “A decoder is learned to invert the AdaIN output to the image spaces”; Karras.B, Section 3, Other Datasets: “The accompanying video provides results for style mixing and stochastic variation tests. As can be seen therein, in case of BEDROOM the coarse styles basically control the viewpoint of the camera” Notes: a plurality of latent representations can clearly be derived by adjusting the coarse styles resulting in different viewpoints, as the coarse styles are used to adjust the latent representation in the feature space);
Generating, via a third decoder, a third latent representation in the first representation space based on the second latent representation (Huang, Figure 2: “A decoder is learned to invert the AdaIN output to the image spaces”. Notes: While Huang doesn’t explicitly state that a third latent representation based on the second latent representation in the first representation space is generated, it is obvious that an image converted to the feature space (second latent representation) can be fed into AdaIN with its own image as a style image resulting in no style change, and decoded back into the image space (first representation space), resulting in the third latent representation);
Computing a plurality of vector directions based on a comparison of the third latent representation and the second plurality of edited latent representations (Park, Paragraph [0101]: “the deep image manipulation system can generate an attribute direction by determining a difference between the latent codes of the two groups of digital images”; Karras.B, Section 3, Other Datasets: “The accompanying video provides results for style mixing and stochastic variation tests. As can be seen therein, in case of BEDROOM the coarse styles basically control the viewpoint of the camera” Notes: a plurality of latent representations can clearly be derived by adjusting the coarse styles resulting in different viewpoints, as the coarse styles are used to adjust the latent representation in the feature space. The latent codes of the two groups of digital images are in the image space, and hence are analogous with the second plurality of edited latent representations of Karras.B. Furthermore, different attribute directions may be determined depending on the number of groups of digital images observed); and
Generating the plurality of latent representations based on the first latent representations and the plurality of vector directions (Park, Paragraph [0101]: “Moving a spatial code or a global code in the attribute direction increases the attribute, while moving a spatial code or global code in a direction opposite to the attribute direction reduces the presence of the attribute in a resulting image”. Notes: spatial code and global code are latent codes (representations) of an image)).
Regarding Claim 10, the method of Claim 9 is rejected over Karras.A as modified.
Karras.A as modified teaches a method wherein the second decoder is the same as the third decoder (Huang, Figure 2: “An overview of our style transfer algorithm. We use the first few layers of a fixed VGG-19 network to encode the content and style images. An AdaIN layer is used to perform style transfer in the feature space. A decoder is learned to invert the AdaIN output to the image spaces”. Notes: As noted previously, the second and third decoders take as input AdaIN output shaped latent representation (edited and non-edited latent representations in the second representations pace), and decodes it into the image space; hence, the second and third decoder are the same).
Regarding Claim 11, the method of Claim 6 is rejected over Karras.A as modified.
Karras.A as modified teaches a method further comprising:
Receiving, via the data interface, a second input image (Huang, Figure 2 illustrates inputting a first and second input image); and
Inputting the second input image to the decoder, wherein the generating of the plurality of images is further based on the second input image (Park, Paragraph [0101]: “the deep image manipulation system can generate an attribute direction by determining a difference between the latent codes of the two groups of digital images. Moving a spatial code or a global code in the attribute direction increases the attribute, while moving a spatial code or global code in a direction opposite to the attribute direction reduces the presence of the attribute in a resulting image”; Karras.B, Section 3, Other Datasets: “The accompanying video provides results for style mixing and stochastic variation tests. As can be seen therein, in case of BEDROOM the coarse styles basically control the viewpoint of the camera”. Notes: the second input image isn’t directly inputted into the decoder itself, but is used for calculating a difference between the first and second input image, which is utilized to generate the output image; thus, in its broadest reasonable interpretation, the second input serves as an input for the decoder).
Regarding Claim 12, the method of Claim 6 is rejected over Karras.A as modified.
Karras.A as modified does not explicitly teach generating a 3D model by applying the final UV map to a 3D mesh.
However, it is well-known in the art that a 3D model can be generated by applying a UV map to a 3D mesh (The intended use of UV mapping is to map a texture to a 3D mesh, which describes 3D coordinate points of a 3D model).
A common motivation in the art is to apply contextually relevant or otherwise appropriate textures to a 3D mesh; textures derived from 2D images are mapped to a 3D mesh using a UV map.
Therefore, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the generation of a final UV map of Karras.A as modified with the well-known use of UV maps for generating 3D models via applying a UV map to a 3D mesh; Doing so would yield the predictable result of obtaining a 3D model with the texture mapped from an associated 2D image.
Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Karras.A (A Style-Based Generator Architecture for Generative Adversarial Networks, 2019), in view of Huang (Arbitrary Style Transfer in Real-time with Adaptive Instance Normalization, 2017), Park (US 20210358177 A1), Karras.B (Supplemental Material: A style-Based Generator Architecture for Generative Adversarial Networks, 2019), and Na (Facial UV map completion for pose-invariant face recognition: a novel adversarial approach based on coupled attention residual UNets, 2020), in further view of Pavllo (Convolutional Generation of Textured 3D Meshes, 2020) and Chen (Towards High-Fidelity Face Self-Occlusion Recovery via Multi-View Residual-Based GAN Inversion, 2022).
Regarding Claim 7, the method of Claim 6 is rejected over Karras.A as modified.
Karras.A as modified teaches a method for generating a final UV map from a plurality of UV maps based on a plurality of images (Na, Figure 5: “The creation of ground-truth complete UV maps. Three facial images with yaw angles of O, -30, and +30 degrees, are fed to the 3DDFA model to create three incomplete UV maps which are then merged by Poisson blending to generate the ground-truth complete UV map”; Refer to Na, Figure 5 for a visualization”).
Karras.A as modified does not teach generating a plurality of 3D meshes based on the plurality of images, generating a plurality of UV maps based on the plurality of images and the plurality of 3D meshes, computing a respective visibility score for each of the plurality of 3D meshes, or generating a plurality of visibility masks based on the plurality of 3D meshes and the respective visibility scores.
However, Pavllo teaches generating a plurality of 3D meshes based on a plurality of images (Figure 1 demonstrates generating a 3D mesh (predicted mesh) from an image. Notes: Given multiple images, Pavllo’s method of generating a 3D mesh from an image can clearly be repeated for different images to generate multiple 3D meshes);
Generating a plurality of UV maps based on the plurality of images and the plurality of 3D meshes (Figure 1 clearly demonstrates obtaining a UV map from an image. Section 2, Related Work, Sub-Section 3D mesh generation: “Our work is based on GANs and can explicitly generate high-resolution texture maps which are then mapped to the mesh via UV mapping. Notes: textures are mapped to 3D meshes via a UV map, which inherently ties a texture/UV map with its corresponding 3D mesh. Notes: as previously noted, Pavllo’s method can clearly be repeated on multiple images to produce a plurality of UV maps based on a plurality of images and associated 3D meshes”);
Computing a respective visibility score for each of the plurality of 3D meshes (Section 3.1, Pose-independent dataset, Sub-Section Pose-independent representation: “The result is the projection of the natural image onto the UV map. However, as can be seen in the figure, this process erroneously projects occluded vertices (the back of the car in the example), which should ideally be masked out as visual information associated with them is not available in the 2D image. We therefore mask the projection using a binary visibility mask, which describes what parts of the mesh are visible in UV space. The mask is obtained by rendering the mesh using a dummy texture (e.g. all white) and computing its gradient with respect to the texture (we provide implementation details in the Appendix A.2). Only texels (pixels of the texture) that contribute to the final image (i.e. visible ones) will have non-zero gradients, therefore we obtain the visibility mask by thresholding these gradients”. Notes: As noted previously, a plurality of 3D meshes can be obtained via Pavllo’s method by repeating on multiple images. Visibility score, in its broadest reasonable interpretation, is a metric for judging the visibility. Pavllo’s visibility mask is derived by checking whether texels have non-zero gradients with respect to a dummy texture; hence, the set of numbers depicting the visibility of each texel is considered a visibility score); and
Generating a plurality of visibility masks based on the plurality of 3D meshes and the respective visibility scores (Section 3.1, Pose-independent dataset, Sub-Section Pose-independent representation “We therefore mask the projection using a binary visibility mask, which describes what parts of the mesh are visible in UV space. The mask is obtained by rendering the mesh using a dummy texture (e.g. all white) and computing its gradient with respect to the texture (we provide implementation details in the Appendix A.2). Only texels (pixels of the texture) that contribute to the final image (i.e. visible ones) will have non-zero gradients, therefore we obtain the visibility mask by thresholding these gradients”. Notes: as noted previously, a plurality of visibility masks can be derived from the plurality of 3D meshes and respective visibility scores when Pavllo’s method is repeated on multiple images).
Karras.A as modified and Pavllo are considered analogous in the art with respect to generating UV maps based on images. One would be motivated to derive 3D meshes and visibility scores and masks corresponding to a plurality of UV maps derived from a plurality of images for the purpose of generating a 3D model, as the most common method of model generation is to map a texture to a 3D mesh via UV mapping for a detailed 3D model with a texture depicted in a 2D image. Hence, One would be motivated to utilize Pavllo’s detailed method for generating a UV map, 3D mesh, visibility score and visibility mask, and apply its use to multiple contextually related images; as demonstrated by Karras.A as modified, combining multiple UV maps that are contextually related (such as different view angles of a person) already exists in the art.
Karras.A as modified does not explicitly teach generating the final UV map based on the plurality of UV maps and the plurality of visibility masks.
However, Chen teaches generating the final UV map based on the plurality of UV maps and the plurality of visibility masks (Figure 2 demonstrates generating a final texture (UV map) based on UV samples of images from different view angles and their associated visibility masks).
Karras.A as modified and Chen are considered analogous in the art with respect to combining UV maps to form a final UV map. One would be motivated to utilize the method of forming a final UV map based on a plurality of UV maps and associated visibility masks, as doing so would result in a cohesive and accurately generated final UV map with definitive information regarding visibility of certain regions of the UV maps corresponding with the 3D meshes.
Therefore, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the generation of a plurality of UV maps and visibility masks of Karras.A as modified with the method of forming a final UV map using visibility masks of Chen; Doing so would yield the predictable result of more accurately generating a final UV map.
Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Karras.A (A Style-Based Generator Architecture for Generative Adversarial Networks, 2019), in view of Huang (Arbitrary Style Transfer in Real-time with Adaptive Instance Normalization, 2017), Park (US 20210358177 A1), Karras.B (Supplemental Material: A style-Based Generator Architecture for Generative Adversarial Networks, 2019), and Na (Facial UV map completion for pose-invariant face recognition: a novel adversarial approach based on coupled attention residual UNets, 2020), Pavllo (Convolutional Generation of Textured 3D Meshes, 2020) and Chen (Towards High-Fidelity Face Self-Occlusion Recovery via Multi-View Residual-Based GAN Inversion, 2022), in further view of Stack Overflow (Element-wise multiplication with Keras, 2018).
Regarding Claim 8, the method of Claim 7 is rejected over Karras.A as modified.
Karras.A as modified teaches generating the final UV map based on the application of each of visibility masks of the plurality of visibility masks with a respective UV map of the plurality of UV maps (Figure 2 demonstrates generating a final texture (UV map) based on UV samples of images from different view angles and their associated visibility masks).
Karras.A as modified does not explicitly teach multiplying each visibility mask of the plurality of visibility masks with a corresponding UV map of the plurality of UV maps.
However, Stack Overflow teaches multiplying a visibility mask with a corresponding UV map (Mark F.: “I have a RGB image of shape (256,256,3) and I have a weight mask of shape (256,256). How do I perform the element-wise multiplication between them with Keras? (all channels share the same mask)”. Notes: multiplying a mask to data is well-known in the art and largely inherent to the application of a mask to data (image). A visibility mask and UV map are analogous in structure to the weight mask and RGB image).
Karras.A as modified and Stack Overflow are considered analogous in the art with respect to obtaining masked data given an image in the form of a latent representation. One would be motivated to use the method of applying a mask to data of Stack Overflow when applying a visibility mask to a UV map, as applying a mask via multiplication is well-known in the art, to create a more accurate UV map.
Therefore, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the use of a visibility mask and corresponding UV map for generating a final UV map of Karras.A as modified with the well-known method of applying a mask to data via multiplication of Stack Overflow; Doing so would yield the predictable result of a final UV map generated through conventional mask application.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to RAYMOND CHUN LAM LI whose telephone number is (571)272-5124. The examiner can normally be reached M-F 8:30-5.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kent Chang can be reached at 571-272-7667. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/RAYMOND CHUN LAM LI/Examiner, Art Unit 2614
/KENT W CHANG/Supervisory Patent Examiner, Art Unit 2614