Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
Election/Restrictions
Applicant’s election without traverse of Invention I, claims 3-9 in the reply filed on 5/5/2026 is acknowledged.
As a result of Applicant’s election, Invention I, claims 1-9, 13-20, are examined in the present office action, and claims 10-12 have been withdrawn for further consideration as being directed to non-elected Invention. Applicant should note that the non-elected claims will be rejoined if the linking claims 1-2 are later found as an allowable claims.
Claim Objections
Claims 14-15 are objected to because of the following informalities: Claims 14-15 recites the limitation “via the data interface a target..” in line 2. It should be comma between interface and a (via the data interface, a target…) Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 8 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 8 depends from claim 7 and claim 7 depends from claim 5 recites the limitation "an encoder" in line 2. It is unclear if “an encoder” is referring to “an encoder” in claim 5 or something else.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
1. Claim 1 is rejected under 35 U.S.C. 103 as being unpatentable over Gu, Jiatao, et al., IDS, "Stylenerf: A style-based 3d-aware generator for high-resolution image synthesis." arXiv preprint arXiv:2110.08985 (2021) (“Gu”) further in view of Kuang et al., U.S Patent Application Publication No.20240193850 (“Kuang”)
Regarding independent claim 1, Gu teaches a method of image generation ( abstract, “We propose StyleNeRF, a 3D-aware generative model for photo-realistic high resolution image synthesis with high multi-view consistency, which can be trained on unstructured 2D images. Existing approaches either cannot synthesize high resolution images with fine details or yield noticeable 3D-inconsistent artifacts. In addition, many of them lack control over style attributes and explicit 3D camera poses. StyleNeRF integrates the neural radiance field (NeRF) into a style-based generator to tackle the aforementioned challenges, i.e., improving rendering ef ficiency and 3D consistency for high-resolution image generation. We perform volume rendering only to produce a low-resolution feature map and progressively apply upsampling in 2D to address the first issue.”), the method comprising:
receiving, via data interface, a plurality of control parameters and a view direction (section 3 METHOD 3.1 IMAGE SYNTHESIS AS NEURAL IMPLICIT FIELD RENDERING Style-based Generative Neural Radiance Field We start by modeling a 3D scene as neural radiance field (NeRF, Mildenhall et al., 2020). It is typically parameterized as multilayer perceptrons (MLPs), which takes the position x ∈ R3 and viewing direction d ∈ S2 as input);
generating a plurality of predicted densities based on a plurality of positions and the plurality of control parameters (3 METHOD 3.1 IMAGE SYNTHESIS AS NEURAL IMPLICIT FIELD RENDERING Style-based Generative Neural Radiance Field We start by modeling a 3D scene as neural radiance field (NeRF, Mildenhall et al., 2020). It is typically parameterized as multilayer perceptrons (MLPs), which takes the position x ∈ R3 and viewing direction d ∈ S2 as input, and predicts the density σ(x) ∈ R+ and view-dependent color c(x,d) ∈ R3. To model high-frequency details, follwing NeRF (Mildenhall et al., 2020), we map each dimension of x and d with Fourier features…”; see section Synthesis Network , “Considering that our training images generally have unbounded background, we choose NeRF++ (Zhang et al., 2020), a variant of NeRF, as the StyleNeRF backbone. NeRF++ consists of a foreground NeRF in a unit sphere and a background NeRF represented with inverted sphere parameterization. As shown in Figure 3, two MLPs are used to predict the density where the background network has fewer parameters than the foreground one. Then a shared MLP is employed for color prediction. Each style-conditioned block consists of an affine transformation layer and a 1 ×1 convolution layer (Conv). The Conv weights are modulated with the affine-transformed styles, and then demodulated for computation. leaky ReLU is used as non-linear activation. The number of blocks depends on the input and target image resolutions.”);and
generating an image based on the plurality of predicted densities and the view direction (3.1 IMAGE SYNTHESIS AS NEURAL IMPLICIT FIELD RENDERING Style-based Generative Neural Radiance Field We start by modeling a 3D scene as neural ra diance field (NeRF, Mildenhall et al., 2020). It is typically parameterized as multilayer perceptrons (MLPs), which takes the position x ∈ R3 and viewing direction d ∈ S2 as input, and predicts the density σ(x) ∈ R+ and view-dependent color c(x,d) ∈ R3. To model high-frequency details, follwing NeRF (Mildenhall et al., 2020), we map each dimension of x and d with Fourier features : ζL(x) = sin(20x),cos(20x),...,sin(2L−1x),cos(2L−1x) We formalize StyleNeRF representations by conditioning NeRF with style vectors w as follows: φn w(x) = gn w ◦gn−1 w ◦...◦g1 w ◦ζ(x), where w = f(z),z ∈ Z (1) (2) Similar as StyleGAN2 (Karras et al., 2020b), f is a mapping network that maps noise vectors from the spherical Gaussian space Z to the style space W; gi w(.) is the ith layer MLP whose weight matrix is modulated by the input style vector w. φn w(x) is the n-th layer feature of that point. We then use the extracted features to predict the density and color, respectively: σw(x) = hσ ◦φnσ w (x), cw(x,d) = hc ◦[φncw(x),ζ (d)], (3) where hσ and hc can be a linear projection or 2-layer MLPs. Different from the original NeRF, we assume nc > nσ for Equation (3) as the visual appearance generally needs more capacity to model than the geometry. The first min(nσ,nc) layers are shared in the network. Volume Rendering Image synthesis is modeled as volume rendering from a given camera pose p ∈ P. For simplicity, we assume a camera is located on the unit sphere pointing to the origin with a fixed field of view (FOV). We sample the camera’s pitch & yaw from a uniform or Gaussian distribution depending on the dataset. To render an image I ∈ RH×W×3, we shoot a camera ray r(t) = o+td(ois the camera origin) for each pixel, and then calculate the color using the volume rendering equation: ∞ INeRF w (r)= pw(t)cw(r(t),d)dt, where pw(t) = exp − 0 t σw(r(s))ds · σw(r(t)) (4) 0 In practice, the above equation is discretized by accumulating sampled points along the ray. Following NeRF (Mildenhall et al., 2020), stratified and hierarchical sampling are applied for more accurate discrete approximation to the continuous implicit function. Challenges Compared to 2D generative models (e.g., StyleGANs (Karras et al., 2019; 2020b)), the images generated by NeRF-based models have 3D consistency, which is guaranteed by modeling the image synthesis as a physics process, and the neural 3D scene representation is invariant across different viewpoints. However, the drawbacks are apparent: these models cost much more computation to render an image at the exact resolution. For example, 2D GANs are 100 ∼ 1000 times more efficient to generate a 10242 image than NeRF-based models. Furthermore, NeRF consumes much more memory to cache the intermediate results for gradient back-propagation during training, making it difficult to train on high-resolution images. Both of these restrict the scope of applying NeRF-based models in high-quality image synthesis, especially at the training stage when calculating the objective function over the whole image is crucial.”) Gu is understood to be silent on the remaining limitations of claim 1.
In the same field of endeavor, Kuang teaches receiving, via a data interface ([0055] The user interface manager 502 is a component of the neural decomposition rendering system 500 configured to allow users to provide input image data. In some embodiments, the user interface manager 502 provides a user interface through which the user provides the input images 518 representing a scene, as discussed above. For example, the user interface enables the user to upload the images and/or download the images from a local or remote storage location (e.g., by providing an address (e.g., a URL or other endpoint) associated with an image source). In some embodiments, the user interface can enable a user to link an image capture device, such as a camera or other hardware to capture image data and provided it to the neural decomposition rendering system 500. In some embodiments, the user interface manager 502 also enables the user to provide a specific viewpoint for a view to be synthesized. Additionally, the user interface manager 502 allows users to request neural decomposition rendering system 500 to produce an appearance decomposition for a synthesized scene relating to the input images 518. In some embodiments, the user interface manager 502 enables the user to directly edit the resulting appearance decomposition. Alternatively, as discussed above, the appearance decomposition can be edited in a graphics design system separate from the neural decomposition rendering system 500.”), a plurality of control parameters and a view direction (see at least [0022] Newer techniques have utilized neural networks, rather than meshes, to perform object capture. For example, a neural network can be used to encode the geometry and appearance of an object and can be used to render the object in an environment. One such technique is the neural radiance field (“NeRF”) that synthesizes novel views of complex scenes by optimizing an underlying continuous volumetric scene function using a sparse set of input views. The NeRF algorithm represents a scene using two fully connected (non-convolutional) deep networks, whose input are continuous coordinates consisting of a spatial location and viewing direction. The output of one of the networks is the volume density and the other is a view-dependent emitted radiance at that spatial location. The synthesized views are generated using classic volume rendering techniques to project the output colors and densities into an image.”; [0026] More specifically, embodiments improve upon prior techniques by introducing neural basis decomposition into a fully connected neural network to change the appearance of an object in a scene while leaving other objects unchanged. For example, embodiments replace the color prediction of NeRF with neural basis functions that allow for colors and/or shading of a synthesized scene to be editable. In some embodiments, the neural basis function has two parameters: a global parameter representing the fundamental properties of the basis (e.g., the colors of a color palette) and an influence parameter representing the influence of a basis at a certain 3D point.”);
generating a plurality densities based on a plurality of positions and the plurality of control parameters (see at least [0040] As in prior techniques, such as NeRF, for geometry calculations, a scene density network F.sub.σ 125 is used to regress the volume density σ at any 3D point x=(x, y, z). In contrast to prior techniques, for appearance, a plurality of models is used. For example, a color palette model f.sub.c(x, d) 131 is used to decompose the colors of a scene, and a shading model f.sub.s(x, d)135 is used to decompose the shading effects of a scene using the 2D view-dependent radiance given a viewpoint d=(0, 0).”); and
generating an image based on the plurality of densities and the view direction (see at least [0037] The resulting volume density values from the density manager 120 and the appearance decomposition from the appearance manager 130 are provided to the volume rendering module 140. The volume rendering module 140 implements conventional volume rendering techniques to generate the output 150. This includes generating a synthesized scene 152 at the appropriate viewpoint. The volume rendering module 140 generates an editable output image because the appearance is decomposed into appearance decomposition 152, thereby allowing conventional editing techniques to change the appearance of an object on the synthesized scene 152.”)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify a style-based 3D-aware generator for high-resolution image synthesis of Gu with including user interface as seen in Kuang because this modification would achieve the expected benefits of performing tasks quickly and easily of use.
Thus, the combination of Gu and Kuang teaches a method of image generation, the method comprising: receiving, via a data interface, a plurality of control parameters and a view direction; generating a plurality of predicted densities based on a plurality of positions and the plurality of control parameters; and generating an image based on the plurality of predicted densities and the view direction.
2. Claims 2-4 and 9 are rejected under 35 U.S.C. 103 as being unpatentable over Gu, Jiatao, et al., IDS, "Stylenerf: A style-based 3d-aware generator for high-resolution image synthesis." arXiv preprint arXiv:2110.08985 (2021) (“Gu”) in view of Kuang et al., U.S Patent Application Publication No.20240193850 (“Kuang”) further in view of Zhou, P , Xie, L , Ni, B , Tian, Q , IDS, “Cips-3d A 3d-aware generator of gans based on conditionally-independent pixel synthesis” arXiv (2021) (“Zhou”)
Regarding claim 2, Gu and Kuang teach the method of claim 1, wherein the generating the plurality of predicted densities includes:
generating a vector representation of each position of the plurality of positions (see at least 3.1 IMAGE SYNTHESIS AS NEURAL IMPLICIT FIELD RENDERING of Gu, “Style-based Generative Neural Radiance Field We start by modeling a 3D scene as neural radiance field (NeRF, Mildenhall et al., 2020). It is typically parameterized as multilayer perceptrons (MLPs), which takes the position x ∈ R3 and viewing direction d ∈ S2 as input, and predicts the density σ(x) ∈ R+ and view-dependent color c(x,d) ∈ R3. To model high-frequency details, follwing NeRF (Mildenhall et al., 2020), we map each dimension of x and d with Fourier features : ζL(x) = sin(20x),cos(20x),...,sin(2L−1x),cos(2L−1x) We formalize StyleNeRF representations by conditioning NeRF with style vectors w as follows: φn w(x) = gn w ◦gn−1 w ◦...◦g1 w ◦ζ(x), where w = f(z),z ∈ Z (1) (2) Similar as StyleGAN2 (Karras et al., 2020b), f is a mapping network that maps noise vectors from the spherical Gaussian space Z to the style space W; gi w(.) is the ith layer MLP whose weight matrix is modulated by the input style vector w. φn w(x) is the n-th layer feature of that point. We then use the extracted features to predict the density and color, respectively: σw(x) = hσ ◦φnσ w (x), cw(x,d) = hc ◦[φncw(x),ζ (d)], (3) where hσ and hc can be a linear projection or 2-layer MLPs. Different from the original NeRF, we assume nc > nσ for Equation (3) as the visual appearance generally needs more capacity to model than the geometry. The first min(nσ,nc) layers are shared in the network”; [0060] of Kuang “In some embodiments, the sub-networks, the scene density network 510, the color palette model 512, and the shading model 514 are designed as MLP networks. Embodiments use unit vectors to represent a 3D point u=(x, y, z) and a viewing angle d=(θ, Ø). Embodiments use positional encoding to deduce high-frequency geometry and appearance details. In particular, embodiments apply positional encoding for the scene density network 510, the color palette model 512, and the shading model 514 on their respective input components, including u and d.”); and
updating the vector representation via a series of modulation blocks to provide an updated vector representation, wherein each modulation block of the series of modulation blocks uses a different respective subset of the plurality of control parameters (see at least 3.1 IMAGE SYNTHESIS AS NEURAL IMPLICIT FIELD RENDERING of Gu “Style-based Generative Neural Radiance Field We start by modeling a 3D scene as neural radiance field (NeRF, Mildenhall et al., 2020). It is typically parameterized as multilayer perceptrons (MLPs), which takes the position x ∈ R3 and viewing direction d ∈ S2 as input, and predicts the density σ(x) ∈ R+ and view-dependent color c(x,d) ∈ R3. To model high-frequency details, follwing NeRF (Mildenhall et al., 2020), we map each dimension of x and d with Fourier features : ζL(x) = sin(20x),cos(20x),...,sin(2L−1x),cos(2L−1x) We formalize StyleNeRF representations by conditioning NeRF with style vectors w as follows: φn w(x) = gn w ◦gn−1 w ◦...◦g1 w ◦ζ(x), where w = f(z),z ∈ Z (1) (2) Similar as StyleGAN2 (Karras et al., 2020b), f is a mapping network that maps noise vectors from the spherical Gaussian space Z to the style space W; gi w(.) is the ith layer MLP whose weight matrix is modulated by the input style vector w. φn w(x) is the n-th layer feature of that point. We then use the extracted features to predict the density and color, respectively: σw(x) = hσ ◦φnσ w (x), cw(x,d) = hc ◦[φncw(x),ζ (d)], (3) where hσ and hc can be a linear projection or 2-layer MLPs. Different from the original NeRF, we assume nc > nσ for Equation (3) as the visual appearance generally needs more capacity to model than the geometry. The first min(nσ,nc) layers are shared in the network.”; 3.3 PRESERVING 3D CONSISTENCY Upsampler design Up-sampling in 2D space causes multi-view inconsistency in general; however, the specific design choice of the upsampler determines how much such inconsistency is introduced. As our model is directly derived from NeRF, MLP (1 × 1 Conv) is the basic building block. With MLPs, however, pixel-wise learnable upsamplers such as pixelshuffle (Shi et al., 2016) or LIIF (Chen et al., 2021b) produce “chessboard” or “texture sticking” artifacts due to its tendency proposed to use non-learnable upsamplers that interpolate the feature map with pre-defined low pass filters (e.g. bilinear interpolation). While these upsamplers can produce smoother outputs, we observed non-removable “bubble” artifacts in both the feature maps and output images. We conjec ture it is due to the lack of local variations when combining MLPs with fixed low-pass filters. We achieve the balance between consistency and image quality by combining these two approaches..”; Synthesis Network Considering that our training images generally have unbounded background, we choose NeRF++ (Zhang et al., 2020), a variant of NeRF, as the StyleNeRF backbone. NeRF++ consists of a foreground NeRF in a unit sphere and a background NeRF represented with inverted sphere parameterization. As shown in Figure 3, two MLPs are used to predict the density where the background network has fewer parameters than the foreground one. Then a shared MLP is employed for color prediction. Each style-conditioned block consists of an affine transformation layer and a 1 ×1 convolution layer (Conv). The Conv weights are modulated with the affine-transformed styles, and then demodulated for computation. leaky ReLU is used as non-linear activation. The number of blocks depends on the input and target image resolutions.; [0060] of Kuang “In some embodiments, the sub-networks, the scene density network 510, the color palette model 512, and the shading model 514 are designed as MLP networks. Embodiments use unit vectors to represent a 3D point u=(x, y, z) and a viewing angle d=(θ, Ø). Embodiments use positional encoding to deduce high-frequency geometry and appearance details. In particular, embodiments apply positional encoding for the scene density network 510, the color palette model 512, and the shading model 514 on their respective input components, including u and d.) In addition, the same motivation is used as the rejection for claim 1. Both Gu and Kuang are understood to be silent on the remaining limitations of claim 2.
In the same field of endeavor, Zhou teaches generating a vector representation of each position of the plurality of positions; and updating the vector representation via a series of modulation blocks to provide an updated vector representation, wherein each modulation block of the series of modulation blocks uses a different respective subset of the plurality of control parameters (whole paper, see at least section 3. Method, Modulated SIREN Block, “ The vanilla NeRF is restricted to a specific scene with fixed geometry. Because the generated images are various for GANs, the shape of each image is different. To render different shapes with one NeRF, we condition the NeRF network on a noise vector zs so that different shapes can be obtained by sampling zs. In particular, like StyleGAN [32], we use a mapping network ms : Zs → Ws tomapzs to ws, and use ws to modulate the feature maps of the NeRF network (see Fig. 2a). We adopt the strategy of pi-GAN [11]: modulating features with FiLM [17,47], followed by a SIREN activation function [54]. The modulated SIREN block is given by where γ=Affine(ws) and β=Affine(ws) represent frequency and phase, respectively. The block contains a FC layer with W and b as the weight matrix and the bias. As shown in Fig. 2b, the NeRF network only contains three SIREN blocks to minimize run time memory complexity. Our experiments show that the viewing direction will cause inconsistencies of face identities under multiple views (seeFig.11a). Thus we do not take d as input, which is different from the original NeRF[39]. Furthermore, instead of predicting the color c, we let the NeRF network predicta more general featurev[42]. As are result, the proposed NeRF function is given by g:Rdim(x)×Rdim(zs)→R+×Rdim(v), (x,zs)→(σ,v), (3) where v is a feature vector corresponding to point x. zs is a shape code that is shared by all pixels of a generated image..)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify a style-based 3D-aware generator for high-resolution image synthesis of Gu, Kuang with including modulated SIREN block as seen in Zhou because this modification would outputs the frequencies and phase shifts ( section Modulated SIREN Block of Zhou).
Thus, the combination of Gu, Kuang and Zhou teaches wherein the generating the plurality of predicted densities includes: generating a vector representation of each position of the plurality of positions; and updating the vector representation via a series of modulation blocks to provide an updated vector representation, wherein each modulation block of the series of modulation blocks uses a different respective subset of the plurality of control parameters.
Regarding claim 3, Gu, Kuang and Zhou teach the method of claim 2, wherein the updating the vector representation includes: generating, by each modulation block of the series of modulation blocks, a plurality of frequency values and a plurality of offset values based on an affine transformation of the respective subset of the plurality of control parameters; and updating the vector representation based on the plurality of frequency values and the plurality of offset values ((see at least 3.1 IMAGE SYNTHESIS AS NEURAL IMPLICIT FIELD RENDERING of Gu, “Style-based Generative Neural Radiance Field We start by modeling a 3D scene as neural radiance field (NeRF, Mildenhall et al., 2020). It is typically parameterized as multilayer perceptrons (MLPs), which takes the position x ∈ R3 and viewing direction d ∈ S2 as input, and predicts the density σ(x) ∈ R+ and view-dependent color c(x,d) ∈ R3. To model high-frequency details, follwing NeRF (Mildenhall et al., 2020), we map each dimension of x and d with Fourier features : ζL(x) = sin(20x),cos(20x),...,sin(2L−1x),cos(2L−1x) We formalize StyleNeRF representations by conditioning NeRF with style vectors w as follows: φn w(x) = gn w ◦gn−1 w ◦...◦g1 w ◦ζ(x), where w = f(z),z ∈ Z (1) (2) Similar as StyleGAN2 (Karras et al., 2020b), f is a mapping network that maps noise vectors from the spherical Gaussian space Z to the style space W; gi w(.) is the ith layer MLP whose weight matrix is modulated by the input style vector w. φn w(x) is the n-th layer feature of that point. We then use the extracted features to predict the density and color, respectively: σw(x) = hσ ◦φnσ w (x), cw(x,d) = hc ◦[φncw(x),ζ (d)], (3) where hσ and hc can be a linear projection or 2-layer MLPs. Different from the original NeRF, we assume nc > nσ for Equation (3) as the visual appearance generally needs more capacity to model than the geometry. The first min(nσ,nc) layers are shared in the network”; Synthesis Network Considering that our training images generally have unbounded background, we choose NeRF++ (Zhang et al., 2020), a variant of NeRF, as the StyleNeRF backbone. NeRF++ consists of a foreground NeRF in a unit sphere and a background NeRF represented with inverted sphere parameterization. As shown in Figure 3, two MLPs are used to predict the density where the background network has fewer parameters than the foreground one. Then a shared MLP is employed for color prediction. Each style-conditioned block consists of an affine transformation layer and a 1 ×1 convolution layer (Conv). The Conv weights are modulated with the affine-transformed styles, and then demodulated for computation. leaky ReLU is used as non-linear activation. The number of blocks depends on the input and target image resolutions”; [0042] of Kuang “ In some embodiments, the shading model 135 decomposes a scene's appearance into different shading terms (e.g., diffuse and specular terms). Each neural basis can represent a shading component for a specific frequency domain. The higher the frequency is, the glossier the shading effect is. Specifically, the shading model 135 controls the frequency domain of a neural basis by adding the basis property θ to the Sinusoidal Positional Embedding of a viewing direction r. The embedding is changed such that θ.sub.i becomes a scalar parameter between 0 and 1. ϕ.sub.i is set to the blending weight neural network 136 and the phase function neural network 137 with the updated position embedding. Accordingly, the color palette model 131 produces a color decomposition, and the shading model 135 produces a shading decomposition that can support realistic rendering while also allowing for appearance editing of the color and/or shading appearance of a synthesized object in a rendered scene.; section 3. Method, Modulated SIREN Block of Zhou, “ The vanilla NeRF is restricted to a specific scene with fixed geometry. Because the generated images are various for GANs, the shape of each image is different. To render different shapes with one NeRF, we condition the NeRF network on a noise vector zs so that different shapes can be obtained by sampling zs. In particular, like StyleGAN [32], we use a mapping network ms : Zs → Ws tomapzs to ws, and use ws to modulate the feature maps of the NeRF network (see Fig. 2a). We adopt the strategy of pi-GAN [11]: modulating features with FiLM [17,47], followed by a SIREN activation function [54]. The modulated SIREN block is given by where γ=Affine(ws) and β=Affine(ws) represent frequency and phase, respectively. The block contains a FC layer with W and b as the weight matrix and the bias. As shown in Fig. 2b, the NeRF network only contains three SIREN blocks to minimize run time memory complexity. Our experiments show that the viewing direction will cause inconsistencies of face identities under multiple views (seeFig.11a). Thus we do not take d as input, which is different from the original NeRF[39]. Furthermore, instead of predicting the color c, we let the NeRF network predicta more general featurev[42]. As are result, the proposed NeRF function is given by g:Rdim(x)×Rdim(zs)→R+×Rdim(v), (x,zs)→(σ,v), (3) where v is a feature vector corresponding to point x. zs is a shape code that is shared by all pixels of a generated image..”) In addition, the same motivation is used as the rejection for claim 2.
Regarding claim 4, Gu, Kuang and Zhou teach the method of claim 3, wherein the updating the vector representation further includes: adding, by each modulation block of the series of modulation blocks, an input vector representation as input to each respective modulation block to an output vector representation as modulated by each respective modulation block (see at least 3.1 IMAGE SYNTHESIS AS NEURAL IMPLICIT FIELD RENDERING of GU “Style-based Generative Neural Radiance Field We start by modeling a 3D scene as neural radiance field (NeRF, Mildenhall et al., 2020). It is typically parameterized as multilayer perceptrons (MLPs), which takes the position x ∈ R3 and viewing direction d ∈ S2 as input, and predicts the density σ(x) ∈ R+ and view-dependent color c(x,d) ∈ R3. To model high-frequency details, follwing NeRF (Mildenhall et al., 2020), we map each dimension of x and d with Fourier features : ζL(x) = sin(20x),cos(20x),...,sin(2L−1x),cos(2L−1x) We formalize StyleNeRF representations by conditioning NeRF with style vectors w as follows: φn w(x) = gn w ◦gn−1 w ◦...◦g1 w ◦ζ(x), where w = f(z),z ∈ Z (1) (2) Similar as StyleGAN2 (Karras et al., 2020b), f is a mapping network that maps noise vectors from the spherical Gaussian space Z to the style space W; gi w(.) is the ith layer MLP whose weight matrix is modulated by the input style vector w. φn w(x) is the n-th layer feature of that point. We then use the extracted features to predict the density and color, respectively: σw(x) = hσ ◦φnσ w (x), cw(x,d) = hc ◦[φncw(x),ζ (d)], (3) where hσ and hc can be a linear projection or 2-layer MLPs. Different from the original NeRF, we assume nc > nσ for Equation (3) as the visual appearance generally needs more capacity to model than the geometry. The first min(nσ,nc) layers are shared in the network”; Synthesis Network Considering that our training images generally have unbounded background, we choose NeRF++ (Zhang et al., 2020), a variant of NeRF, as the StyleNeRF backbone. NeRF++ consists of a foreground NeRF in a unit sphere and a background NeRF represented with inverted sphere parameterization. As shown in Figure 3, two MLPs are used to predict the density where the background network has fewer parameters than the foreground one. Then a shared MLP is employed for color prediction. Each style-conditioned block consists of an affine transformation layer and a 1 ×1 convolution layer (Conv). The Conv weights are modulated with the affine-transformed styles, and then demodulated for computation. leaky ReLU is used as non-linear activation. The number of blocks depends on the input and target image resolutions; [0060] of Kuang “In some embodiments, the sub-networks, the scene density network 510, the color palette model 512, and the shading model 514 are designed as MLP networks. Embodiments use unit vectors to represent a 3D point u=(x, y, z) and a viewing angle d=(θ, Ø). Embodiments use positional encoding to deduce high-frequency geometry and appearance details. In particular, embodiments apply positional encoding for the scene density network 510, the color palette model 512, and the shading model 514 on their respective input components, including u and d.; see section 3. Method of Zhou, Efficient Implementation for ModFC Like the NeRF network, the INR network also adopts a style-based architecture. As shown in Fig. 2d, a mapping network ma : Za →Wa turns za into wa, where the stochasticity of appearance comes from the code za. Then, wa is mapped to style vectors using affine layers (i.e., FC layer). The style vectors are injected into the INR network using Modulated Fully Connected (ModFC) layers..) In addition, the same motivation us used as the rejection for claim 2.
Regarding claim 9, Gu, Kuang and Zhou teach the method of claim 2, wherein the generating the plurality of predicted densities includes generating each density of the plurality of predicted densities via a neural network based transformation based on the updated vector representation (see at least 3.1 IMAGE SYNTHESIS AS NEURAL IMPLICIT FIELD RENDERING of Gu, “Style-based Generative Neural Radiance Field We start by modeling a 3D scene as neural radiance field (NeRF, Mildenhall et al., 2020). It is typically parameterized as multilayer perceptrons (MLPs), which takes the position x ∈ R3 and viewing direction d ∈ S2 as input, and predicts the density σ(x) ∈ R+ and view-dependent color c(x,d) ∈ R3. To model high-frequency details, follwing NeRF (Mildenhall et al., 2020), we map each dimension of x and d with Fourier features : ζL(x) = sin(20x),cos(20x),...,sin(2L−1x),cos(2L−1x) We formalize StyleNeRF representations by conditioning NeRF with style vectors w as follows: φn w(x) = gn w ◦gn−1 w ◦...◦g1 w ◦ζ(x), where w = f(z),z ∈ Z (1) (2) Similar as StyleGAN2 (Karras et al., 2020b), f is a mapping network that maps noise vectors from the spherical Gaussian space Z to the style space W; gi w(.) is the ith layer MLP whose weight matrix is modulated by the input style vector w. φn w(x) is the n-th layer feature of that point. We then use the extracted features to predict the density and color, respectively: σw(x) = hσ ◦φnσ w (x), cw(x,d) = hc ◦[φncw(x),ζ (d)], (3) where hσ and hc can be a linear projection or 2-layer MLPs. Different from the original NeRF, we assume nc > nσ for Equation (3) as the visual appearance generally needs more capacity to model than the geometry. The first min(nσ,nc) layers are shared in the network”; [0040] of Kuang “As in prior techniques, such as NeRF, for geometry calculations, a scene density network F.sub.σ 125 is used to regress the volume density σ at any 3D point x=(x, y, z). In contrast to prior techniques, for appearance, a plurality of models is used. For example, a color palette model f.sub.c(x, d) 131 is used to decompose the colors of a scene, and a shading model f.sub.s(x, d)135 is used to decompose the shading effects of a scene using the 2D view-dependent radiance given a viewpoint d=(0, 0). 044] As shown in FIG. 3, embodiments include a decomposed neural representation comprising multiple neural networks (e.g., MLPs, CNNs, or combinations of these and/or other neural networks) for neural decomposition rendering. As in prior techniques, such as NeRF, for geometry calculations, a scene density network F.sub.σ125 is used to regress the volume density a at any 3D point x=(x, y, z). In contrast to prior techniques, for appearance, a plurality of models are used. For example, a color palette model f.sub.c(x, d) 131 is used to decompose the colors of a scene, and a shading model f.sub.s(x, d)135 is used to decompose the shading effects of a scene using the 2D view-dependent radiance given a viewpoint d=(θ,ϕ).; see section 3. Method, 3.1. NeRF Network for 3D Shape, Volume Rendering of Zhou “As described in Eq.(3),the neural radiance field represents a scene as the volume density σ and feature vector v at any point in space. Let o be the camera origin. For each pixel, we cast aray r(t)=o+td from origin o towards the pixel. We sample points along the camera rayr(t) and transform the 3D coordinates of the points into volume densities and feature vectors using Eq.(3).Using the classical volume rendering[28], the over all feature vector Vr corresponding to the rayr(t) is given by….”) In addition, the same motivation is used as the rejection for claim 2.
3. Claims 5-8 are rejected under 35 U.S.C. 103 as being unpatentable over Gu, Jiatao, et al., IDS, "Stylenerf: A style-based 3d-aware generator for high-resolution image synthesis." arXiv preprint arXiv:2110.08985 (2021) (“Gu”) in view of Kuang et al., U.S Patent Application Publication No.20240193850 (“Kuang”) further in view of Zhou, P , Xie, L , Ni, B , Tian, Q , IDS, “Cips-3d A 3d-aware generator of gans based on conditionally-independent pixel synthesis” arXiv (2021) (“Zhou”) further in view Sargent, Kyle, et al. "Vq3d: Learning a 3d-aware generative model on imageNet." Proceedings of the IEEE/CVF International Conference on Computer Vision. Feb 14, 2023.(“Sargent”) further in view of WEBER et al., U.S Patent Application Publication No.20240078726 (“WEBER”)
Regarding claim 5, Gu, Kuang and Zhou teach the method of claim 3, wherein the generated image is a first image (see at least 3.2 APPROXIMATION FOR HIGH-RESOLUTION IMAGE GENERATION of Gu, Figure 2: Internal representations and the outputs from StyleNeRF trained with different upsampling operators. With LIIF, patterns stick to the same pixel coordinates when the viewpoint changes; With bilinear interpolation, bubble-shape artifacts can be seen on the feature maps and images. Our proposed upsampler faithfully preserves 3D consistency while getting rid of bubble-shape artifacts”; see section 3. Method of Zhou),further comprising:
generating a second image based on the plurality of predicted densities and a canonical view direction (see at least section NeRF path regularization of Gu “We propose a new regularization term to enforce 3D consistency, which regularizes the model output to match the original path (Equation (4)). In this way, the final outputs can be closer to the NeRF results, which have multi-view consistency. This is implemented by sub-sampling pixels on the output and comparing them against those generated by NeRF: LNeRF-path = 1 |S| (i,j)∈S IApprox w (Rin)[i,j]−INeRF w (Rout[i,j]) 2 , (8) where S is the set of randomly sampled pixels; Rin and Rout are the corresponding rays of the pixels in the low-resolution image generated via NeRF and high-resolution output of StyleNeRF”; see section 3. Method of Zhou);
generating, via an encoder, a latent representation of the first image (whole paper, see at least Figure 3 of Gu);
generating, via a neural network based transformation model, an updated latent representation of the first image; and updating parameters of the neural network based transformation model based on a comparison of the updated latent representation of the first image and the latent representation of the second image (whole paper, see at least section NeRF path regularization of Gu “We propose a new regularization term to enforce 3D consistency, which regularizes the model output to match the original path (Equation (4)). In this way, the final outputs can be closer to the NeRF results, which have multi-view consistency. This is implemented by sub-sampling pixels on the output and comparing them against those generated by NeRF: LNeRF-path = 1 |S| (i,j)∈S IApprox w (Rin)[i,j]−INeRF w (Rout[i,j]) 2 , (8) where S is the set of randomly sampled pixels; Rin and Rout are the corresponding rays of the pixels in the low-resolution image generated via NeRF and high-resolution output of StyleNeRF”; section 3. Method 3, Efficient Implementation for ModFC of Zhou “ Like the NeRF network, the INR network also adopts a style-based architecture. As shown in Fig. 2d, a mapping network ma : Za →Wa turns za into wa, where the stochasticity of ap pearance comes from the code za. Then, wa is mapped to style vectors using affine layers (i.e., FC layer). The style vectors are injected into the INR network using Modulated Fully Connected (ModFC) layers. CIPS [4] regards ModFC as a special case of 1 × 1 convolutional layer and implements ModFC with the off the-shelf modulated convolutional layer [33]. The modulated convolution is implemented using grouped convolution, which is not efficient for ModFC. In fact, we can directly utilize the batch matrix multiplication (bmm) to implement ModFC more efficiently. As shown in Fig. 4, ModFC consists of Mod, Demod, and a batch matrix multiplication operation. Mathematically, let W ∈ Rdin×dout be the weights of a fully connected layer, S ∈ Rb×din be a batch of style vectors, and X ∈ Rb×n×din be the input with n being the length of the sequence. We first resize W and S to shapes of 1×din ×dout and b×din ×1, respectively. The Mod operation is given by W= W ⊗ S, where ⊗ stands for tensor-broadcasting multiplication and W ∈Rb×din×dout. The Demod operation is given by W =W⊗ din (W·,din,·)2+ −1 2,where is as mall constant and W ∈ Rb×din×dout. Finally, we use W to linearly map the input X ∈ Rb×n×din (i.e., Y = X×W , and Y ∈ Rb×n×dout), which is achieved through the batch matrix multiplication function1. Experiments substantiate that this implementation is more efficient than the implementation using grouped convolution (see Fig. 9) In addition, the same motivation is used as the rejection for claim 2. Gu, Kuang and Zhou are understood to be silent on the remaining limitations of claim 4.
In the same field of endeavor, Sargent teaches generating a second image based on a canonical view direction (see at least 3.2. Training, 1. Good reconstruction from a canonical view. On Im ageNet, ground truth camera extrinsics are unknown and probably not even well-defined due to the presence of deformable and ambiguous object categories and scenes without salient objects. Therefore, we simply fix a single ‘canonical pose’ for reconstruction, and our cri terion is that our conditional NeRF-based autoencoder should successfully reconstruct the dataset from this view. 2. Reasonable novel views. We expect that images de coded at novel views within a specified range of the canonical view will have similar quality to images de coded at the canonical view.”);
generating, via an encoder, a latent representation of the first image (see at least 3 Model, 3.1. Overview of VQ3D, “Stage 2 is a generative autoregressive transformer which predicts sequences of latent tokens. A diagram of the inputs, outputs, and architecture is shown in the bottom of Figure 2. The architectural and training details are generally the same as [43]. We train it on the sequences of latent codes produced by our Stage 1 encoder. After training, the autoregressive transformer can be used to generate totally new 3D images by first sampling a sequence of latent tokens and then applying”); generating, via the encoder, a latent representation of the second image ((see at least 3 Model, 3.1. Overview of VQ3D, “Stage 2 is a generative autoregressive transformer which predicts sequences of latent tokens. A diagram of the inputs, outputs, and architecture is shown in the bottom of Figure 2. The architectural and training details are generally the same as [43]. We train it on the sequences of latent codes produced by our Stage 1 encoder. After training, the autoregressive transformer can be used to generate totally new 3D images by first sampling a sequence of latent tokens and then applying”)
generating, via a neural network based transformation model, an updated latent representation of the first image (see at least 3. Model.. 3.3. Architecture “A full architecture diagram is shown in Figure 2. Similar to [43], we leverage the powerful vision transformer [10] architecture in both the encoder and decoder. Different from [43], which is trained on 2D images, we utilize a novel decoder with 3D inductive bias to facilitate the learning of 3D representations. We now give an overview of the individual components of our architecture.”; Autoregressive transformer. We train transformer [39] to autoregressively predict the next image token. We follow the hyperparameters in the base model of VIM [43]. For Ima geNet, we train a conditional model, and for other datasets we train unconditional generative models);
Therefore, in combination of Gu, Kuang, Zhou, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify a style-based 3D-aware generator for high-resolution image synthesis of with applying canonical view as seen in Sargent because this modification would reconstruct the dataset from the canonical view (3.2. Training of Sargent) Gu, Kuang, Zhou and Sargent are understood to be silent on the remaining limitations of claim 5.
In the same field of endeavor, WEBER teaches generating, via an encoder, a latent representation of the first image (see at least [0027] Encoder 204 converts each of images 222(1)-222(X) into corresponding latent representations 224(1)-224(X) (each of which is referred to individually as latent representation 224). Decoder 206 converts latent representations 224 into output data 228 that includes images 226(1)-226(X) (each of which is referred to as image 226) with a different facial identity 218 from that of images 222. [0028] In one or more embodiments, machine learning model 200 is an autoencoder that includes encoder 204 and decoder 206. Encoder 204 includes convolutional layers and/or other types of neural network layers that convert a 2D image 222 that depicts facial identity 250 into a corresponding latent representation 224. This latent representation 224 includes a fixed-length vector representation of the visual attributes in image 222 in a lower-dimensional space”);
generating, via the encoder, a latent representation of the second image (see at least [0027] Encoder 204 converts each of images 222(1)-222(X) into corresponding latent representations 224(1)-224(X) (each of which is referred to individually as latent representation 224). Decoder 206 converts latent representations 224 into output data 228 that includes images 226(1)-226(X) (each of which is referred to as image 226) with a different facial identity 218 from that of images 222. [0028] In one or more embodiments, machine learning model 200 is an autoencoder that includes encoder 204 and decoder 206. Encoder 204 includes convolutional layers and/or other types of neural network layers that convert a 2D image 222 that depicts facial identity 250 into a corresponding latent representation 224. This latent representation 224 includes a fixed-length vector representation of the visual attributes in image 222 in a lower-dimensional space”)
generating, via a neural network based transformation model, an updated latent representation of the first image (see at least [0039] For example, training engine 122 could perform a forward pass that uses encoder 204 to convert a batch of multiscopic training images 230 into corresponding training latent representations 212 and uses one or more instances of decoder 206 to convert training latent representations 212 into decoder output 210. Training engine 122 could compute a mean squared error (MSE), L1 loss, and/or another type of reconstruction loss 232 that measures the differences between multiscopic training images 230 and decoder output 210 that corresponds to reconstructions of multiscopic training images 230. Training engine 122 could then perform a backward pass that backpropagates reconstruction loss 232 across layers of machine learning model 200 and uses stochastic gradient descent to update parameters (e.g., neural network weights) of machine learning model 200 based on the negative gradients of the backpropagated reconstruction loss 232. Training engine 122 could repeat the forward and backward passes with different batches of training data 214 and/or over a number of training epochs and/or iterations until reconstruction loss 232 falls below a threshold and/or another condition is met. [0040] By training machine learning model 200 in a way that minimizes reconstruction loss 232, training engine 122 generates a trained encoder 204 that learns lower-dimensional latent representations 224 of multiscopic training images 230. Training engine 122 also generates one or more instances of a trained decoder 206 that learn different facial identities 216 associated with multiscopic training images 230 and can reconstruct the original multiscopic training images 230 from the corresponding latent representations 224 outputted by the trained encoder 204.”);
and updating parameters of the neural network based transformation model based on a comparison of the updated latent representation of the first image and the latent representation of the second image ([0039] For example, training engine 122 could perform a forward pass that uses encoder 204 to convert a batch of multiscopic training images 230 into corresponding training latent representations 212 and uses one or more instances of decoder 206 to convert training latent representations 212 into decoder output 210. Training engine 122 could compute a mean squared error (MSE), L1 loss, and/or another type of reconstruction loss 232 that measures the differences between multiscopic training images 230 and decoder output 210 that corresponds to reconstructions of multiscopic training images 230. Training engine 122 could then perform a backward pass that backpropagates reconstruction loss 232 across layers of machine learning model 200 and uses stochastic gradient descent to update parameters (e.g., neural network weights) of machine learning model 200 based on the negative gradients of the backpropagated reconstruction loss 232. Training engine 122 could repeat the forward and backward passes with different batches of training data 214 and/or over a number of training epochs and/or iterations until reconstruction loss 232 falls below a threshold and/or another condition is met. [0040] By training machine learning model 200 in a way that minimizes reconstruction loss 232, training engine 122 generates a trained encoder 204 that learns lower-dimensional latent representations 224 of multiscopic training images 230. Training engine 122 also generates one or more instances of a trained decoder 206 that learn different facial identities 216 associated with multiscopic training images 230 and can reconstruct the original multiscopic training images 230 from the corresponding latent representations 224 outputted by the trained encoder 204. ; [0047] “Training engine 122 can also repeat training of machine learning model 200 using reconstruction loss 232 and/or geometry loss 234. For example, training engine 122 could alternate between one training stage that updates parameters of machine learning model 200 based on reconstruction loss 232 and another training stage that updates parameters of machine learning model 200 based on geometry loss 234 until one or both losses fall below corresponding thresholds. In another example, training engine 122 could perform one or more stages that train machine learning model 200 using a combination (e.g., sum) of reconstruction loss 232 and geometry loss 234. Thus, by training machine learning model 200 using both reconstruction loss 232 and geometry loss 234, training engine 122 improves the performance of machine learning model 200 both in reconstructing images of faces and in generating images with changed facial identities 216 in a geometrically consistent way.” [0057] In step 304, training engine 122 executes one or more decoder neural networks that convert the set of latent representations into a set of output images. Continuing with the above example, training engine 122 could input some or all of the latent representations into each decoder neural network. Training engine 122 could use one or more convolutional layers and/or other types of neural network layers in the decoder neural network to convert the latent representations into the output images. [0058] As mentioned above, the input images can depict multiple facial identities (e.g., combinations of personal identities, ages, lighting conditions, etc.). Each of these facial identities can be learned by a different decoder neural network. As a result, training engine 122 can perform step 304 by inputting different subsets of latent representations associated with different facial identities into different decoder neural networks and using each decoder neural network to convert the corresponding subset of latent representations into output images. [0059] Alternatively, the same decoder neural network can be used to convert all latent representations generated by the encoder neural network into output images. In this scenario, one or more sets of dense layers are used to generate AdaIN coefficients, neural network weights, and/or other parameters that control the operation of the decoder neural network in converting the latent representations into the output images. These parameters can be adapted to different facial identities associated with the original input images.”)
Therefore, in combination of Gu, Kuang, Zhou, Sargent, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify a style-based 3D-aware generator for high-resolution image synthesis of Gu with applying encoder-decoder machine learning model as seen in WEBER because this modification would generate images with different facial identities ([0030]) of WEBER).
Thus, the combination of Gu, Kuang, Zhou, Sargent and WEBER teaches wherein the generated image is a first image, further comprising: generating a second image based on the plurality of predicted densities and a canonical view direction; generating, via an encoder, a latent representation of the first image; generating, via the encoder, a latent representation of the second image; generating, via a neural network based transformation model, an updated latent representation of the first image; and updating parameters of the neural network based transformation model based on a comparison of the updated latent representation of the first image and the latent representation of the second image.
Regarding claim 6, Gu, Kuang, Zhou, Sargent and WEBER teach he method of claim 5, further comprising: generating, via a decoder, a third image based on the updated latent representation of the first image; generating, via the decoder, a fourth image based on the latent representation of the second image; and updating parameters of the neural network based transformation model based on a comparison of the third image and the fourth image (see at least section 3. Method of Gu, see Section 3.Method of Zhou, Section 3. Model of Sargent; [0027],[0039] of WEBER “ For example, training engine 122 could perform a forward pass that uses encoder 204 to convert a batch of multiscopic training images 230 into corresponding training latent representations 212 and uses one or more instances of decoder 206 to convert training latent representations 212 into decoder output 210. Training engine 122 could compute a mean squared error (MSE), L1 loss, and/or another type of reconstruction loss 232 that measures the differences between multiscopic training images 230 and decoder output 210 that corresponds to reconstructions of multiscopic training images 230. Training engine 122 could then perform a backward pass that backpropagates reconstruction loss 232 across layers of machine learning model 200 and uses stochastic gradient descent to update parameters (e.g., neural network weights) of machine learning model 200 based on the negative gradients of the backpropagated reconstruction loss 232. Training engine 122 could repeat the forward and backward passes with different batches of training data 214 and/or over a number of training epochs and/or iterations until reconstruction loss 232 falls below a threshold and/or another condition is met. [0040] By training machine learning model 200 in a way that minimizes reconstruction loss 232, training engine 122 generates a trained encoder 204 that learns lower-dimensional latent representations 224 of multiscopic training images 230. Training engine 122 also generates one or more instances of a trained decoder 206 that learn different facial identities 216 associated with multiscopic training images 230 and can reconstruct the original multiscopic training images 230 from the corresponding latent representations 224 outputted by the trained encoder 204. ; [0047] “Training engine 122 can also repeat training of machine learning model 200 using reconstruction loss 232 and/or geometry loss 234. For example, training engine 122 could alternate between one training stage that updates parameters of machine learning model 200 based on reconstruction loss 232 and another training stage that updates parameters of machine learning model 200 based on geometry loss 234 until one or both losses fall below corresponding thresholds. In another example, training engine 122 could perform one or more stages that train machine learning model 200 using a combination (e.g., sum) of reconstruction loss 232 and geometry loss 234. Thus, by training machine learning model 200 using both reconstruction loss 232 and geometry loss 234, training engine 122 improves the performance of machine learning model 200 both in reconstructing images of faces and in generating images with changed facial identities 216 in a geometrically consistent way.” [0057] In step 304, training engine 122 executes one or more decoder neural networks that convert the set of latent representations into a set of output images. Continuing with the above example, training engine 122 could input some or all of the latent representations into each decoder neural network. Training engine 122 could use one or more convolutional layers and/or other types of neural network layers in the decoder neural network to convert the latent representations into the output images. [0058] As mentioned above, the input images can depict multiple facial identities (e.g., combinations of personal identities, ages, lighting conditions, etc.). Each of these facial identities can be learned by a different decoder neural network. As a result, training engine 122 can perform step 304 by inputting different subsets of latent representations associated with different facial identities into different decoder neural networks and using each decoder neural network to convert the corresponding subset of latent representations into output images. [0059] Alternatively, the same decoder neural network can be used to convert all latent representations generated by the encoder neural network into output images. In this scenario, one or more sets of dense layers are used to generate AdaIN coefficients, neural network weights, and/or other parameters that control the operation of the decoder neural network in converting the latent representations into the output images. These parameters can be adapted to different facial identities associated with the original input images.”) In addition, the same motivation is used as the rejection for claim 5.
Regarding claim 7, Gu, Kuang, Zhou, Sargent and WEBER teach the method of claim 5, wherein the view direction is a first view direction (see at least section 3. Method of Gu; at least [0022] of Kuang “ Newer techniques have utilized neural networks, rather than meshes, to perform object capture. For example, a neural network can be used to encode the geometry and appearance of an object and can be used to render the object in an environment. One such technique is the neural radiance field (“NeRF”) that synthesizes novel views of complex scenes by optimizing an underlying continuous volumetric scene function using a sparse set of input views. The NeRF algorithm represents a scene using two fully connected (non-convolutional) deep networks, whose input are continuous coordinates consisting of a spatial location and viewing direction. The output of one of the networks is the volume density and the other is a view-dependent emitted radiance at that spatial location. The synthesized views are generated using classic volume rendering techniques to project the output colors and densities into an image.”, further comprising: receiving, via the data interface( see at least [0055] of Kuang), a second view direction (see at least section 3. Method of Gu. [0022] of Kuang; section 3. Method of Zhou);
generating a third image based on the plurality of predicted densities and the second view direction (see at least section 3. Method of Gu;. [0034]of Kuang “he radiance field module 110 implements ray marching techniques to generate rays that pass through each pixel and sample the ray at a plurality of 3D points. The result is a plurality of neural radiance fields, each having a three-dimensional coordinate and a viewing direction. The neural radiance fields are provided to a plurality of neural networks to produce an editable synthesized scene. Unlike prior techniques using neural radiance fields, which use two separate neural networks (density MLP and color MLP to represent scene geometry (density MLP) and appearance (color MLP) separately, embodiments described herein augment the color MLP with multiple neural networks (neural basis MLPs) for the scene appearance (in addition to the color MLP). These neural networks can represent a decomposed basis of the 3D scene appearance. As discussed, instead of training neural networks to determine a static radiance and color of a scene, embodiments train neural networks to decompose the appearance so that the scene's colors and/or shading is editable.”; section 3. Method of Zhou);
generating a target pose latent representation of the first image based on the updated latent representation of the first image, the second view direction, and a learnable parameter matrix (whole paper, section 3. Method, Efficient Implementation for ModFC of Zhou “ Like the NeRF network, the INR network also adopts a style-based architecture. As shown in Fig. 2d, a mapping network ma : Za →Wa turns za into wa, where the stochasticity of ap pearance comes from the code za. Then, wa is mapped to style vectors using affine layers (i.e., FC layer). The style vectors are injected into the INR network using Modulated Fully Connected (ModFC) layers. CIPS [4] regards ModFC as a special case of 1 × 1 convolutional layer and implements ModFC with the off the-shelf modulated convolutional layer [33]. The modulated convolution is implemented using grouped convolution, which is not efficient for ModFC. In fact, we can directly utilize the batch matrix multiplication (bmm) to implement ModFC more efficiently. As shown in Fig. 4, ModFC consists of Mod, Demod, and a batch matrix multiplication operation. Mathematically, let W ∈ Rdin×dout be the weights of a fully connected layer, S ∈ Rb×din be a batch of style vectors, and X ∈ Rb×n×din be the input with n being the length of the sequence. We first resize W and S to shapes of 1×din ×dout and b×din ×1, respectively. The Mod operation is given by W= W ⊗ S, where ⊗ stands for tensor-broadcasting multiplication and W ∈Rb×din×dout. The Demod operation is given by W =W⊗ din (W·,din,·)2+ −1 2,where is as mall constant and W ∈ Rb×din×dout. Finally, we use W to linearly map the input X ∈ Rb×n×din (i.e., Y = X×W , and Y ∈ Rb×n×dout), which is achieved through the batch matrix multiplication function1. Experiments substantiate that this implementation is more efficient than the implementation using grouped convolution (see Fig. 9); ee at least [0027],[0030] of WEBER “ For example, facial identity 218 could be represented by a set of densely connected neural network layers (also referred to herein as “dense layers”) that operate in conjunction with decoder 206. Input into each dense layer could include an output of a corresponding layer of decoder 206. In response to the input, the dense layer generates parameters that control the operation of one or more subsequent layers within decoder 206 and/or the values of weights within the subsequent layer(s). These parameters could include adaptive instance normalization (AdaIN) coefficients that represent facial identity 218. These parameters could also, or instead, include weights associated with neurons within the subsequent layer(s) of decoder 206. The outputted AdaIN coefficients and/or weights would be used to modify convolutional operations performed by the subsequent layer(s) in decoder 206, thereby allowing decoder 206 to generate images 226 with different facial identities. [0048] Execution engine 124 uses the trained machine learning model 200 to convert multiscopic data 220 that includes images 222 depicting a certain facial identity 250 and a certain performance (e.g., behavior, facial expression, pose, etc.) from multiple viewpoints into output data 228 that includes images 226 depicting a different facial identity 218 and the same performance from the same viewpoints. More specifically, execution engine 124 receives a set of multiscopic images 222 of a face at a certain time with a certain facial identity 250. Each image 222 in the set of multiscopic images depicts the face from a different viewpoint. Execution engine 124 uses the trained encoder 204 to convert each image 222 into a corresponding latent representation 224. Execution engine 124 also uses the trained decoder 206 to convert each latent representation 224 into a corresponding image 226 that includes the same performance and viewpoint as the original image represented by the latent representation but depicts a face with a different facial identity 218 than that of the original image. As mentioned above, this facial identity 218 can be represented by a certain instance of decoder 206, an identity vector, and/or a set of dense layers that generate AdaIN coefficients or neural network weights that control the operation of decoder 206.”; [0057] In step 304, training engine 122 executes one or more decoder neural networks that convert the set of latent representations into a set of output images. Continuing with the above example, training engine 122 could input some or all of the latent representations into each decoder neural network. Training engine 122 could use one or more convolutional layers and/or other types of neural network layers in the decoder neural network to convert the latent representations into the output images. [0058] As mentioned above, the input images can depict multiple facial identities (e.g., combinations of personal identities, ages, lighting conditions, etc.). Each of these facial identities can be learned by a different decoder neural network. As a result, training engine 122 can perform step 304 by inputting different subsets of latent representations associated with different facial identities into different decoder neural networks and using each decoder neural network to convert the corresponding subset of latent representations into output images. [0059] Alternatively, the same decoder neural network can be used to convert all latent representations generated by the encoder neural network into output images. In this scenario, one or more sets of dense layers are used to generate AdaIN coefficients, neural network weights, and/or other parameters that control the operation of the decoder neural network in converting the latent representations into the output images. These parameters can be adapted to different facial identities associated with the original input images.”;
generating, via a decoder, a fourth image based on the target pose latent representation of the first image; and updating parameters of the learnable parameter matrix based on a comparison of the third image and the fourth image (see at least section 3. Method of Gu,”; whole paper, section 3. Method, Efficient Implementation for ModFC of Zhou; Section 3. Model of Sargent; [0030] of WEBER “ For example, facial identity 218 could be represented by a set of densely connected neural network layers (also referred to herein as “dense layers”) that operate in conjunction with decoder 206. Input into each dense layer could include an output of a corresponding layer of decoder 206. In response to the input, the dense layer generates parameters that control the operation of one or more subsequent layers within decoder 206 and/or the values of weights within the subsequent layer(s). These parameters could include adaptive instance normalization (AdaIN) coefficients that represent facial identity 218. These parameters could also, or instead, include weights associated with neurons within the subsequent layer(s) of decoder 206. The outputted AdaIN coefficients and/or weights would be used to modify convolutional operations performed by the subsequent layer(s) in decoder 206, thereby allowing decoder 206 to generate images 226 with different facial identities. [0048] Execution engine 124 uses the trained machine learning model 200 to convert multiscopic data 220 that includes images 222 depicting a certain facial identity 250 and a certain performance (e.g., behavior, facial expression, pose, etc.) from multiple viewpoints into output data 228 that includes images 226 depicting a different facial identity 218 and the same performance from the same viewpoints. More specifically, execution engine 124 receives a set of multiscopic images 222 of a face at a certain time with a certain facial identity 250. Each image 222 in the set of multiscopic images depicts the face from a different viewpoint. Execution engine 124 uses the trained encoder 204 to convert each image 222 into a corresponding latent representation 224. Execution engine 124 also uses the trained decoder 206 to convert each latent representation 224 into a corresponding image 226 that includes the same performance and viewpoint as the original image represented by the latent representation but depicts a face with a different facial identity 218 than that of the original image. As mentioned above, this facial identity 218 can be represented by a certain instance of decoder 206, an identity vector, and/or a set of dense layers that generate AdaIN coefficients or neural network weights that control the operation of decoder 206.”; [0057] In step 304, training engine 122 executes one or more decoder neural networks that convert the set of latent representations into a set of output images. Continuing with the above example, training engine 122 could input some or all of the latent representations into each decoder neural network. Training engine 122 could use one or more convolutional layers and/or other types of neural network layers in the decoder neural network to convert the latent representations into the output images. [0058] As mentioned above, the input images can depict multiple facial identities (e.g., combinations of personal identities, ages, lighting conditions, etc.). Each of these facial identities can be learned by a different decoder neural network. As a result, training engine 122 can perform step 304 by inputting different subsets of latent representations associated with different facial identities into different decoder neural networks and using each decoder neural network to convert the corresponding subset of latent representations into output images. [0059] Alternatively, the same decoder neural network can be used to convert all latent representations generated by the encoder neural network into output images. In this scenario, one or more sets of dense layers are used to generate AdaIN coefficients, neural network weights, and/or other parameters that control the operation of the decoder neural network in converting the latent representations into the output images. These parameters can be adapted to different facial identities associated with the original input images.”) In addition, the same motivation is used as the rejection for claim 5.
Regarding claim 8, Gu, Kuang, Zhou, Sargent and WEBER teach the method of claim 7, further comprising: generating, via an encoder, a latent representation of the third image; and updating parameters of the learnable parameter matrix based on a comparison of the target pose latent representation of the first image and the latent representation of the third image (see at least section 3. Method of Gu,”; whole paper, section 3. Method, Efficient Implementation for ModFC of Zhou; Section 3. Model of Sargent; [0027][0030] of WEBER “ For example, facial identity 218 could be represented by a set of densely connected neural network layers (also referred to herein as “dense layers”) that operate in conjunction with decoder 206. Input into each dense layer could include an output of a corresponding layer of decoder 206. In response to the input, the dense layer generates parameters that control the operation of one or more subsequent layers within decoder 206 and/or the values of weights within the subsequent layer(s). These parameters could include adaptive instance normalization (AdaIN) coefficients that represent facial identity 218. These parameters could also, or instead, include weights associated with neurons within the subsequent layer(s) of decoder 206. The outputted AdaIN coefficients and/or weights would be used to modify convolutional operations performed by the subsequent layer(s) in decoder 206, thereby allowing decoder 206 to generate images 226 with different facial identities. [0048] Execution engine 124 uses the trained machine learning model 200 to convert multiscopic data 220 that includes images 222 depicting a certain facial identity 250 and a certain performance (e.g., behavior, facial expression, pose, etc.) from multiple viewpoints into output data 228 that includes images 226 depicting a different facial identity 218 and the same performance from the same viewpoints. More specifically, execution engine 124 receives a set of multiscopic images 222 of a face at a certain time with a certain facial identity 250. Each image 222 in the set of multiscopic images depicts the face from a different viewpoint. Execution engine 124 uses the trained encoder 204 to convert each image 222 into a corresponding latent representation 224. Execution engine 124 also uses the trained decoder 206 to convert each latent representation 224 into a corresponding image 226 that includes the same performance and viewpoint as the original image represented by the latent representation but depicts a face with a different facial identity 218 than that of the original image. As mentioned above, this facial identity 218 can be represented by a certain instance of decoder 206, an identity vector, and/or a set of dense layers that generate AdaIN coefficients or neural network weights that control the operation of decoder 206.”; [0057] In step 304, training engine 122 executes one or more decoder neural networks that convert the set of latent representations into a set of output images. Continuing with the above example, training engine 122 could input some or all of the latent representations into each decoder neural network. Training engine 122 could use one or more convolutional layers and/or other types of neural network layers in the decoder neural network to convert the latent representations into the output images. [0058] As mentioned above, the input images can depict multiple facial identities (e.g., combinations of personal identities, ages, lighting conditions, etc.). Each of these facial identities can be learned by a different decoder neural network. As a result, training engine 122 can perform step 304 by inputting different subsets of latent representations associated with different facial identities into different decoder neural networks and using each decoder neural network to convert the corresponding subset of latent representations into output images. [0059] Alternatively, the same decoder neural network can be used to convert all latent representations generated by the encoder neural network into output images. In this scenario, one or more sets of dense layers are used to generate AdaIN coefficients, neural network weights, and/or other parameters that control the operation of the decoder neural network in converting the latent representations into the output images. These parameters can be adapted to different facial identities associated with the original input images.”) In addition, the same motivation is used as the rejection for claim 5.
4. Claims 13, 16-17,19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Gu, Jiatao, et al., IDS, "Stylenerf: A style-based 3d-aware generator for high-resolution image synthesis." arXiv preprint arXiv:2110.08985 (2021) (“Gu”) in view of WEBER et al., U.S Patent Application Publication No.20240078726 (“WEBER”) further in view of LEE et al., U.S Patent Application Publication No.20200234066 (”LEE”)
Regarding independent claim 13, Gu teaches method of image generation ( abstract, “We propose StyleNeRF, a 3D-aware generative model for photo-realistic high resolution image synthesis with high multi-view consistency, which can be trained on unstructured 2D images. Existing approaches either cannot synthesize high resolution images with fine details or yield noticeable 3D-inconsistent artifacts. In addition, many of them lack control over style attributes and explicit 3D camera poses. StyleNeRF integrates the neural radiance field (NeRF) into a style-based generator to tackle the aforementioned challenges, i.e., improving rendering ef ficiency and 3D consistency for high-resolution image generation. We perform volume rendering only to produce a low-resolution feature map and progressively apply upsampling in 2D to address the first issue.”), the method comprising:
receiving, via a data interface, an input image and a view direction (see at least section 3 METHOD 3.1 IMAGE SYNTHESIS AS NEURAL IMPLICIT FIELD RENDERING Style-based Generative Neural Radiance Field We start by modeling a 3D scene as neural radiance field (NeRF, Mildenhall et al., 2020). It is typically parameterized as multilayer perceptrons (MLPs), which takes the position x ∈ R3 and viewing direction d ∈ S2 as input” which it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to recognize taking the position x ∈ R3 and viewing direction d ∈ S2 as input of Gu with using data interface because this modification would achieve the expected benefits of eliminating the need for manual data entry, manual reconciliation, and redundant processing );
generating, via an encoder, a latent representation of the input image (see at least 3 METHOD 3.1 IMAGE SYNTHESIS AS NEURAL IMPLICIT FIELD RENDERING , “Style-based Generative Neural Radiance Field We start by modeling a 3D scene as neural radiance field (NeRF, Mildenhall et al., 2020). It is typically parameterized as multilayer perceptrons (MLPs), which takes the position x ∈ R3 and viewing direction d ∈ S2 as input, and predicts the density σ(x) ∈ R+ and view-dependent color c(x,d) ∈ R3. To model high-frequency details, follwing NeRF (Mildenhall et al., 2020), we map each dimension of x and d with Fourier features…”; see section Synthesis Network , “Considering that our training images generally have unbounded background, we choose NeRF++ (Zhang et al., 2020), a variant of NeRF, as the StyleNeRF backbone. NeRF++ consists of a foreground NeRF in a unit sphere and a background NeRF represented with inverted sphere parameterization. As shown in Figure 3, two MLPs are used to predict the density where the background network has fewer parameters than the foreground one. Then a shared MLP is employed for color prediction. Each style-conditioned block consists of an affine transformation layer and a 1 ×1 convolution layer (Conv). The Conv weights are modulated with the affine-transformed styles, and then demodulated for computation. leaky ReLU is used as non-linear activation. The number of blocks depends on the input and target image resolutions.”, Figure 3);
generating, via a neural network based transformation model, an updated latent representation of the input image (whole paper, see at least 3.1 IMAGE SYNTHESIS AS NEURAL IMPLICIT FIELD RENDERING Style-based Generative Neural Radiance Field We start by modeling a 3D scene as neural radiance field (NeRF, Mildenhall et al., 2020). It is typically parameterized as multilayer perceptrons (MLPs), which takes the position x ∈ R3 and viewing direction d ∈ S2 as input, and predicts the density σ(x) ∈ R+ and view-dependent color c(x,d) ∈ R3. To model high-frequency details, follwing NeRF (Mildenhall et al., 2020), we map each dimension of x and d with Fourier features : ζL(x) = sin(20x),cos(20x),...,sin(2L−1x),cos(2L−1x) We formalize StyleNeRF representations by conditioning NeRF with style vectors w as follows: φn w(x) = gn w ◦gn−1 w ◦...◦g1 w ◦ζ(x), where w = f(z),z ∈ Z (1) (2) Similar as StyleGAN2 (Karras et al., 2020b), f is a mapping network that maps noise vectors from the spherical Gaussian space Z to the style space W; gi w(.) is the ith layer MLP whose weight matrix is modulated by the input style vector w. φn w(x) is the n-th layer feature of that point. We then use the extracted features to predict the density and color, respectively: σw(x) = hσ ◦φnσ w (x), cw(x,d) = hc ◦[φncw(x),ζ (d)], (3) where hσ and hc can be a linear projection or 2-layer MLPs. Different from the original NeRF, we assume nc > nσ for Equation (3) as the visual appearance generally needs more capacity to model than the geometry. The first min(nσ,nc) layers are shared in the network. Volume Rendering Image synthesis is modeled as volume rendering from a given camera pose p ∈ P. For simplicity, we assume a camera is located on the unit sphere pointing to the origin with a fixed field of view (FOV). We sample the camera’s pitch & yaw from a uniform or Gaussian distribution depending on the dataset. To render an image I ∈ RH×W×3, we shoot a camera ray r(t) = o+td(ois the camera origin) for each pixel, and then calculate the color using the volume rendering equation: ∞ INeRF w (r)= pw(t)cw(r(t),d)dt, where pw(t) = exp − 0 t σw(r(s))ds · σw(r(t)) (4) 0 In practice, the above equation is discretized by accumulating sampled points along the ray. Following NeRF (Mildenhall et al., 2020), stratified and hierarchical sampling are applied for more accurate discrete approximation to the continuous implicit function. Challenges Compared to 2D generative models (e.g., StyleGANs (Karras et al., 2019; 2020b)), the images generated by NeRF-based models have 3D consistency, which is guaranteed by modeling the image synthesis as a physics process, and the neural 3D scene representation is invariant across different viewpoints. However, the drawbacks are apparent: these models cost much more computation to render an image at the exact resolution. For example, 2D GANs are 100 ∼ 1000 times more efficient to generate a 10242 image than NeRF-based models. Furthermore, NeRF consumes much more memory to cache the intermediate results for gradient back-propagation during training, making it difficult to train on high-resolution images. Both of these restrict the scope of applying NeRF-based models in high-quality image synthesis, especially at the training stage when calculating the objective function over the whole image is crucial.”) Gu is understood to be silent on the remaining limitations of claim 13.
In the same field of endeavor, WEBER teaches generating, via an encoder, a latent representation of the input image (see at least [0027] Encoder 204 converts each of images 222(1)-222(X) into corresponding latent representations 224(1)-224(X) (each of which is referred to individually as latent representation 224). Decoder 206 converts latent representations 224 into output data 228 that includes images 226(1)-226(X) (each of which is referred to as image 226) with a different facial identity 218 from that of images 222. [0028] In one or more embodiments, machine learning model 200 is an autoencoder that includes encoder 204 and decoder 206. Encoder 204 includes convolutional layers and/or other types of neural network layers that convert a 2D image 222 that depicts facial identity 250 into a corresponding latent representation 224. This latent representation 224 includes a fixed-length vector representation of the visual attributes in image 222 in a lower-dimensional space”);
generating, via a neural network based transformation model, an updated latent representation of the input image (see at least [0039] For example, training engine 122 could perform a forward pass that uses encoder 204 to convert a batch of multiscopic training images 230 into corresponding training latent representations 212 and uses one or more instances of decoder 206 to convert training latent representations 212 into decoder output 210. Training engine 122 could compute a mean squared error (MSE), L1 loss, and/or another type of reconstruction loss 232 that measures the differences between multiscopic training images 230 and decoder output 210 that corresponds to reconstructions of multiscopic training images 230. Training engine 122 could then perform a backward pass that backpropagates reconstruction loss 232 across layers of machine learning model 200 and uses stochastic gradient descent to update parameters (e.g., neural network weights) of machine learning model 200 based on the negative gradients of the backpropagated reconstruction loss 232. Training engine 122 could repeat the forward and backward passes with different batches of training data 214 and/or over a number of training epochs and/or iterations until reconstruction loss 232 falls below a threshold and/or another condition is met. [0040] By training machine learning model 200 in a way that minimizes reconstruction loss 232, training engine 122 generates a trained encoder 204 that learns lower-dimensional latent representations 224 of multiscopic training images 230. Training engine 122 also generates one or more instances of a trained decoder 206 that learn different facial identities 216 associated with multiscopic training images 230 and can reconstruct the original multiscopic training images 230 from the corresponding latent representations 224 outputted by the trained encoder 204. “;
generating a target pose latent representation of the input image based on the updated latent representation of the input image, the view direction and learnable parameter (see at least [0048] Execution engine 124 uses the trained machine learning model 200 to convert multiscopic data 220 that includes images 222 depicting a certain facial identity 250 and a certain performance (e.g., behavior, facial expression, pose, etc.) from multiple viewpoints into output data 228 that includes images 226 depicting a different facial identity 218 and the same performance from the same viewpoints. More specifically, execution engine 124 receives a set of multiscopic images 222 of a face at a certain time with a certain facial identity 250. Each image 222 in the set of multiscopic images depicts the face from a different viewpoint. Execution engine 124 uses the trained encoder 204 to convert each image 222 into a corresponding latent representation 224. Execution engine 124 also uses the trained decoder 206 to convert each latent representation 224 into a corresponding image 226 that includes the same performance and viewpoint as the original image represented by the latent representation but depicts a face with a different facial identity 218 than that of the original image. As mentioned above, this facial identity 218 can be represented by a certain instance of decoder 206, an identity vector, and/or a set of dense layers that generate AdaIN coefficients or neural network weights that control the operation of decoder 206.”; [0057] In step 304, training engine 122 executes one or more decoder neural networks that convert the set of latent representations into a set of output images. Continuing with the above example, training engine 122 could input some or all of the latent representations into each decoder neural network. Training engine 122 could use one or more convolutional layers and/or other types of neural network layers in the decoder neural network to convert the latent representations into the output images. [0058] As mentioned above, the input images can depict multiple facial identities (e.g., combinations of personal identities, ages, lighting conditions, etc.). Each of these facial identities can be learned by a different decoder neural network. As a result, training engine 122 can perform step 304 by inputting different subsets of latent representations associated with different facial identities into different decoder neural networks and using each decoder neural network to convert the corresponding subset of latent representations into output images. [0059] Alternatively, the same decoder neural network can be used to convert all latent representations generated by the encoder neural network into output images. In this scenario, one or more sets of dense layers are used to generate AdaIN coefficients, neural network weights, and/or other parameters that control the operation of the decoder neural network in converting the latent representations into the output images. These parameters can be adapted to different facial identities associated with the original input images.”), and
generating, via a decoder, an output image based on the target pose latent representation of the input image ([0042] After machine learning model 200 has been trained to reconstruct multiscopic training images 230 and/or non-multiscopic training images, training engine 122 further trains machine learning model 200 based on a geometry loss 234 between training geometries 236 associated with different sets of multiscopic training images 230 and output geometries 208 associated with decoder output 210 generated by machine learning model 200 from these sets of multiscopic training images 230. More specifically, training engine 122 inputs a given set of multiscopic training images 230 into encoder 204 and uses encoder 204 to convert the inputted multiscopic training images 230 into corresponding training latent representations 212. Training engine 122 also uses one or more instances of decoder 206 associated with one or more facial identities 216 to convert each of training latent representations 212 into decoder output 210 that includes an output image. This output image would depict the performance from the original image that was converted into the training latent representation and a selected facial identity that differs from the original image. [0050] After machine learning model 200 generates output data 228 that includes a set of images 226 with a different facial identity 218 from the original facial identity 250 depicted within a corresponding set of input images 222, execution engine 124 can use geometry estimation model 202 to assess a geometric consistency 244 between images 222 and images 226. In some embodiments, geometric consistency 244 includes a measure of the distance or difference between a first geometry 240 associated with two or more images 222 and a second geometry 242 associated with two or more corresponding images 226. For example, geometric consistency 244 could include an aggregation (e.g., average, weighted average, sum, weighted sum, etc.) of distances between 3D points in geometry 240 and corresponding 3D points in geometry 242.”)”
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify a style-based 3D-aware generator for high-resolution image synthesis of Gu with applying encoder-decoder machine learning model as seen in WEBER because this modification would generate images with different facial identities ([0030]) of WEBER. Both Gu and WEBER are understood to be silent on the remaining limitations of claim 13.
In the same field of endeavor, LEE teaches receiving, via a data interface ([0070] The functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in hardware, an example hardware configuration may comprise a processing system in a device. The processing system may be implemented with a bus architecture. The bus may include any number of interconnecting buses and bridges depending on the specific application of the processing system and the overall design constraints. The bus may link together various circuits including a processor, machine-readable media, and a bus interface. The bus interface may connect a network adapter, among other things, to the processing system via the bus. The network adapter may implement signal processing functions. For certain aspects, a user interface (e.g., keypad, display, mouse, joystick, etc.) may also be connected to the bus. The bus may also link various other circuits such as timing sources, peripherals, voltage regulators, power management circuits, and the like, which are well known in the art, and therefore, will not be described any further.”), an input image and a view direction ([0044] Representatively, an input image sequence 402 of the CNN-LSTM framework 400 is a chunk of a video sequence, typically sampled by window-sliding along the temporal direction”);
generating, via an encoder, a latent representation of the input image ([0046] head, Each image is forwarded to certain layers of a CNN 410 (e.g., CNN.sub.head) to obtain deep features in the l.sup.th layers, denoted by Z.sup.l={z.sub.t.sup.l}.sub.t=1.sup.T. The input of the spatial attention model 420, z.sub.t.sup.l, is forwarded to two 2-D convolutional layers, and the output is concatenated with last hidden variable h.sub.t−1 from an LSTM network 430. Then, the concatenated tensors are forwarded to the fully connected layers (FC), the hyperbolic tangent (tanh) layer, and the softmax layer, to obtain the attention weights α.sub.t. The attention weights element-wise product with z.sub.t.sup.l is computed to obtain selective deep features, which are then forwarded to the remaining portions of layers of the CNN 410 (e.g., CNN.sub.tail) for latent features Z.sup.f={z.sub.t.sup.f}.sub.i=1.sup.T. The latent features are then forwarded to the LSTM network 430 for encoding temporal dependencies.[0047] The input and the output of the LSTM network 430 are recurrent over time steps. At time step t, the input is the latent feature z.sub.t.sup.f from the last fully connected layer of the CNN 410, the hidden variable h.sub.t−1, and the internal states c.sub.t−1, while the output is the hidden unit h.sub.t and a memory cell c.sub.t for the next time step. Both h.sub.t and c.sub.t are updated and then passed to the LSTM network 430 at each time step. [0048] The LSTM network 430 outputs a set of hidden variables H={h.sub.t}.sub.i=1.sup.T and a set of memory cells C={c.sub.t}.sub.i=1.sup.T, which are used in the temporal attention model 440. The temporal attention model 440 calculates the attention by dot-production between decoder context and encoder representations. Instead of multiple decoder layers, however, a single layer is used for the output of the LSTM network 430. The inputs h.sub.t and c.sub.t, are fused to be a set of state summaries D={d.sub.t}.sub.i=1.sup.T. Then, H and D perform a matrix production, where the results are followed by a softmax layer for attention selection. The attention weights are then applied to H. The adjusted hidden variables H′ are followed by fully connected layers and the tanh layer to obtain class probability distribution P={P.sub.t}.sub.i=1.sup.T.”;
generating a target pose latent representation of the input image based on the updated latent representation of the input image and a learnable parameter matrix ([0046] head, Each image is forwarded to certain layers of a CNN 410 (e.g., CNN.sub.head) to obtain deep features in the l.sup.th layers, denoted by Z.sup.l={z.sub.t.sup.l}.sub.t=1.sup.T. The input of the spatial attention model 420, z.sub.t.sup.l, is forwarded to two 2-D convolutional layers, and the output is concatenated with last hidden variable h.sub.t−1 from an LSTM network 430. Then, the concatenated tensors are forwarded to the fully connected layers (FC), the hyperbolic tangent (tanh) layer, and the softmax layer, to obtain the attention weights α.sub.t. The attention weights element-wise product with z.sub.t.sup.l is computed to obtain selective deep features, which are then forwarded to the remaining portions of layers of the CNN 410 (e.g., CNN.sub.tail) for latent features Z.sup.f={z.sub.t.sup.f}.sub.i=1.sup.T. The latent features are then forwarded to the LSTM network 430 for encoding temporal dependencies. [0047] The input and the output of the LSTM network 430 are recurrent over time steps. At time step t, the input is the latent feature z.sub.t.sup.f from the last fully connected layer of the CNN 410, the hidden variable h.sub.t−1, and the internal states c.sub.t−1, while the output is the hidden unit h.sub.t and a memory cell c.sub.t for the next time step. Both h.sub.t and c.sub.t are updated and then passed to the LSTM network 430 at each time step. ; [0051] Based on the outputs provided by the LSTM network 430, the temporal attention model 440, and the spatial attention model 420 are integrated into the CNN-LSTM framework 400. For example, a temporal attention β.sub.t,u corresponding to the t.sup.th state summary and the u.sup.th hidden variable is computed as a dot-product between d.sub.t and h.sub.u, and then followed by a soft-selection:
[00002]βt,u=exp(dt.Math.hu).Math.k=1T.Math.exp(dt.Math.hk).(4)
The state summary d.sub.t at time step t is defined by:
d.sub.t=W.sub.d.sup.hh.sub.t+W.sub.d.sup.c tanh(c.sub.t)+b.sub.d.sup.hc, (5) Where W.sub.d.sup.h, W.sub.d.sup.c are the learnable parameter matrices, and b.sub.d is a bias vector. This implies how the u.sup.th input contributes to the t.sup.th output. Therefore, the output hidden variable is adjusted according to β.sub.t,u:..”; 0052] The adjusted hidden variable h′.sub.t is then forwarded to the fully connected layers and tanh layer to obtain final prediction p.sub.t for the t.sup.th time step:
p.sub.t=W.sub.p tanh (W.sub.p.sup.hh′.sub.i+W.sub.p.sup.cc.sub.j+b.sub.p.sup.hc)+b.sub.p, (7)
Where W.sub.p, W.sub.p.sup.h, W.sub.p.sup.c are the learnable parameter matrices, and b.sub.p.sup.hc, b.sub.a are the bias vectors.”); and generating, via a decoder, an output image based on the target pose latent representation of the input image ([0048] The LSTM network 430 outputs a set of hidden variables H={h.sub.t}.sub.i=1.sup.T and a set of memory cells C={c.sub.t}.sub.i=1.sup.T, which are used in the temporal attention model 440. The temporal attention model 440 calculates the attention by dot-production between decoder context and encoder representations. Instead of multiple decoder layers, however, a single layer is used for the output of the LSTM network 430. The inputs h.sub.t and c.sub.t, are fused to be a set of state summaries D={d.sub.t}.sub.i=1.sup.T. Then, H and D perform a matrix production, where the results are followed by a softmax layer for attention selection. The attention weights are then applied to H. The adjusted hidden variables H′ are followed by fully connected layers and the tanh layer to obtain class probability distribution P={P.sub.t}.sub.i=1.sup.T”)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify a style-based 3D-aware generator for high-resolution image synthesis of Gu and WEBER with using temporal attention model as seen in LEE because this modification would calculate the attention by dot-production between decoder context and encoder representations ([0048] of LEE).
Thus, the combination of Gu, WEBER and LEE teaches a method of image generation, the method comprising: receiving, via a data interface, an input image and a view direction; generating, via an encoder, a latent representation of the input image; generating, via a neural network based transformation model, an updated latent representation of the input image; generating a target pose latent representation of the input image based on the updated latent representation of the input image, the view direction, and a learnable parameter matrix; and generating, via a decoder, an output image based on the target pose latent representation of the input image.
Regarding claim 16, Gu, WEBER and LEE teach the method of claim 13, wherein generating the target pose latent representation of the input image includes summing the updated latent representation with a product of the view direction and the learnable parameter matrix ((3 METHOD 3.1 IMAGE SYNTHESIS AS NEURAL IMPLICIT FIELD RENDERING of Gu , “Style-based Generative Neural Radiance Field We start by modeling a 3D scene as neural radiance field (NeRF, Mildenhall et al., 2020). It is typically parameterized as multilayer perceptrons (MLPs), which takes the position x ∈ R3 and viewing direction d ∈ S2 as input, and predicts the density σ(x) ∈ R+ and view-dependent color c(x,d) ∈ R3. To model high-frequency details, follwing NeRF (Mildenhall et al., 2020), we map each dimension of x and d with Fourier features…”; see section Synthesis Network , “Considering that our training images generally have unbounded background, we choose NeRF++ (Zhang et al., 2020), a variant of NeRF, as the StyleNeRF backbone. NeRF++ consists of a foreground NeRF in a unit sphere and a background NeRF represented with inverted sphere parameterization. As shown in Figure 3, two MLPs are used to predict the density where the background network has fewer parameters than the foreground one. Then a shared MLP is employed for color prediction. Each style-conditioned block consists of an affine transformation layer and a 1 ×1 convolution layer (Conv). The Conv weights are modulated with the affine-transformed styles, and then demodulated for computation. leaky ReLU is used as non-linear activation. The number of blocks depends on the input and target image resolutions.”, Figure 3; [0031] of WEBER “ Continuing with the above example, when facial identity 218 corresponds to an interpolation between two or more facial identities represented by two or more sets of dense layers, a weighted combination of AdaIN coefficients (or neural network weights) outputted by the two or more sets of dense layers could be used to control the convolutional operations performed by the layers of decoder 206. Thus, a given facial identity 218 that corresponds to 80% of a first facial identity associated with a first set of dense layers and 20% of a second facial identity associated with a second set of dense layers could be represented by a weighted sum of AdaIN coefficients (or neural network weights) outputted by the first and second sets of dense layers. Within the weighted sum, the AdaIN coefficients (or neural network weights) outputted by the first set of dense layers would be combined with a weight of 0.8, and the AdaIN coefficients (or neural network weights) outputted by the second set of dense layers would be combined with a weight of 0.2.”; [0046] of LEE “head, Each image is forwarded to certain layers of a CNN 410 (e.g., CNN.sub.head) to obtain deep features in the l.sup.th layers, denoted by Z.sup.l={z.sub.t.sup.l}.sub.t=1.sup.T. The input of the spatial attention model 420, z.sub.t.sup.l, is forwarded to two 2-D convolutional layers, and the output is concatenated with last hidden variable h.sub.t−1 from an LSTM network 430. Then, the concatenated tensors are forwarded to the fully connected layers (FC), the hyperbolic tangent (tanh) layer, and the softmax layer, to obtain the attention weights α.sub.t. The attention weights element-wise product with z.sub.t.sup.l is computed to obtain selective deep features, which are then forwarded to the remaining portions of layers of the CNN 410 (e.g., CNN.sub.tail) for latent features Z.sup.f={z.sub.t.sup.f}.sub.i=1.sup.T. The latent features are then forwarded to the LSTM network 430 for encoding temporal dependencies. [0047] The input and the output of the LSTM network 430 are recurrent over time steps. At time step t, the input is the latent feature z.sub.t.sup.f from the last fully connected layer of the CNN 410, the hidden variable h.sub.t−1, and the internal states c.sub.t−1, while the output is the hidden unit h.sub.t and a memory cell c.sub.t for the next time step. Both h.sub.t and c.sub.t are updated and then passed to the LSTM network 430 at each time step. ; [0051] Based on the outputs provided by the LSTM network 430, the temporal attention model 440, and the spatial attention model 420 are integrated into the CNN-LSTM framework 400. For example, a temporal attention β.sub.t,u corresponding to the t.sup.th state summary and the u.sup.th hidden variable is computed as a dot-product between d.sub.t and h.sub.u, and then followed by a soft-selection:
[00002]βt,u=exp(dt.Math.hu).Math.k=1T.Math.exp(dt.Math.hk).(4)
The state summary d.sub.t at time step t is defined by:
d.sub.t=W.sub.d.sup.hh.sub.t+W.sub.d.sup.c tanh(c.sub.t)+b.sub.d.sup.hc, (5) Where W.sub.d.sup.h, W.sub.d.sup.c are the learnable parameter matrices, and b.sub.d is a bias vector. This implies how the u.sup.th input contributes to the t.sup.th output. Therefore, the output hidden variable is adjusted according to β.sub.t,u:..”; 0052] The adjusted hidden variable h′.sub.t is then forwarded to the fully connected layers and tanh layer to obtain final prediction p.sub.t for the t.sup.th time step:
p.sub.t=W.sub.p tanh (W.sub.p.sup.hh′.sub.i+W.sub.p.sup.cc.sub.j+b.sub.p.sup.hc)+b.sub.p, (7)
Where W.sub.p, W.sub.p.sup.h, W.sub.p.sup.c are the learnable parameter matrices, and b.sub.p.sup.hc, b.sub.a are the bias vectors.”) In addition, the same motivation is used as the rejection for claim 13.
Regarding independent claim 17, Gu teaches method of image generation ( abstract, “We propose StyleNeRF, a 3D-aware generative model for photo-realistic high resolution image synthesis with high multi-view consistency, which can be trained on unstructured 2D images. Existing approaches either cannot synthesize high resolution images with fine details or yield noticeable 3D-inconsistent artifacts. In addition, many of them lack control over style attributes and explicit 3D camera poses. StyleNeRF integrates the neural radiance field (NeRF) into a style-based generator to tackle the aforementioned challenges, i.e., improving rendering ef ficiency and 3D consistency for high-resolution image generation. We perform volume rendering only to produce a low-resolution feature map and progressively apply upsampling in 2D to address the first issue.”), the method comprising:
receiving, via a data interface, a plurality of control parameters and a view direction (section 3 METHOD 3.1 IMAGE SYNTHESIS AS NEURAL IMPLICIT FIELD RENDERING Style-based Generative Neural Radiance Field We start by modeling a 3D scene as neural radiance field (NeRF, Mildenhall et al., 2020). It is typically parameterized as multilayer perceptrons (MLPs), which takes the position x ∈ R3 and viewing direction d ∈ S2 as input” which it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to recognize taking the position x ∈ R3 and viewing direction d ∈ S2 as input of Gu with using data interface because this modification would achieve the expected benefits of eliminating the need for manual data entry, manual reconciliation, and redundant processing );
generating, via a mapping network, a latent representation of an image based on the plurality of control parameters (3 METHOD 3.1 IMAGE SYNTHESIS AS NEURAL IMPLICIT FIELD RENDERING , “Style-based Generative Neural Radiance Field We start by modeling a 3D scene as neural radiance field (NeRF, Mildenhall et al., 2020). It is typically parameterized as multilayer perceptrons (MLPs), which takes the position x ∈ R3 and viewing direction d ∈ S2 as input, and predicts the density σ(x) ∈ R+ and view-dependent color c(x,d) ∈ R3. To model high-frequency details, follwing NeRF (Mildenhall et al., 2020), we map each dimension of x and d with Fourier features…”; see section Synthesis Network , “Considering that our training images generally have unbounded background, we choose NeRF++ (Zhang et al., 2020), a variant of NeRF, as the StyleNeRF backbone. NeRF++ consists of a foreground NeRF in a unit sphere and a background NeRF represented with inverted sphere parameterization. As shown in Figure 3, two MLPs are used to predict the density where the background network has fewer parameters than the foreground one. Then a shared MLP is employed for color prediction. Each style-conditioned block consists of an affine transformation layer and a 1 ×1 convolution layer (Conv). The Conv weights are modulated with the affine-transformed styles, and then demodulated for computation. leaky ReLU is used as non-linear activation. The number of blocks depends on the input and target image resolutions.”, Figure 3);and
generating, via a neural network based transformation model, an updated latent representation of the input image (whole paper, see at least 3.1 IMAGE SYNTHESIS AS NEURAL IMPLICIT FIELD RENDERING Style-based Generative Neural Radiance Field We start by modeling a 3D scene as neural radiance field (NeRF, Mildenhall et al., 2020). It is typically parameterized as multilayer perceptrons (MLPs), which takes the position x ∈ R3 and viewing direction d ∈ S2 as input, and predicts the density σ(x) ∈ R+ and view-dependent color c(x,d) ∈ R3. To model high-frequency details, follwing NeRF (Mildenhall et al., 2020), we map each dimension of x and d with Fourier features : ζL(x) = sin(20x),cos(20x),...,sin(2L−1x),cos(2L−1x) We formalize StyleNeRF representations by conditioning NeRF with style vectors w as follows: φn w(x) = gn w ◦gn−1 w ◦...◦g1 w ◦ζ(x), where w = f(z),z ∈ Z (1) (2) Similar as StyleGAN2 (Karras et al., 2020b), f is a mapping network that maps noise vectors from the spherical Gaussian space Z to the style space W; gi w(.) is the ith layer MLP whose weight matrix is modulated by the input style vector w. φn w(x) is the n-th layer feature of that point. We then use the extracted features to predict the density and color, respectively: σw(x) = hσ ◦φnσ w (x), cw(x,d) = hc ◦[φncw(x),ζ (d)], (3) where hσ and hc can be a linear projection or 2-layer MLPs. Different from the original NeRF, we assume nc > nσ for Equation (3) as the visual appearance generally needs more capacity to model than the geometry. The first min(nσ,nc) layers are shared in the network. Volume Rendering Image synthesis is modeled as volume rendering from a given camera pose p ∈ P. For simplicity, we assume a camera is located on the unit sphere pointing to the origin with a fixed field of view (FOV). We sample the camera’s pitch & yaw from a uniform or Gaussian distribution depending on the dataset. To render an image I ∈ RH×W×3, we shoot a camera ray r(t) = o+td(ois the camera origin) for each pixel, and then calculate the color using the volume rendering equation: ∞ INeRF w (r)= pw(t)cw(r(t),d)dt, where pw(t) = exp − 0 t σw(r(s))ds · σw(r(t)) (4) 0 In practice, the above equation is discretized by accumulating sampled points along the ray. Following NeRF (Mildenhall et al., 2020), stratified and hierarchical sampling are applied for more accurate discrete approximation to the continuous implicit function. Challenges Compared to 2D generative models (e.g., StyleGANs (Karras et al., 2019; 2020b)), the images generated by NeRF-based models have 3D consistency, which is guaranteed by modeling the image synthesis as a physics process, and the neural 3D scene representation is invariant across different viewpoints. However, the drawbacks are apparent: these models cost much more computation to render an image at the exact resolution. For example, 2D GANs are 100 ∼ 1000 times more efficient to generate a 10242 image than NeRF-based models. Furthermore, NeRF consumes much more memory to cache the intermediate results for gradient back-propagation during training, making it difficult to train on high-resolution images. Both of these restrict the scope of applying NeRF-based models in high-quality image synthesis, especially at the training stage when calculating the objective function over the whole image is crucial.”) Gu is understood to be silent on the remaining limitations of claim 13.
In the same field of endeavor, WEBER teaches generating, via a mapping network, a latent representation of an image based on the plurality of control parameters (see at least [0027] Encoder 204 converts each of images 222(1)-222(X) into corresponding latent representations 224(1)-224(X) (each of which is referred to individually as latent representation 224). Decoder 206 converts latent representations 224 into output data 228 that includes images 226(1)-226(X) (each of which is referred to as image 226) with a different facial identity 218 from that of images 222. [0028] In one or more embodiments, machine learning model 200 is an autoencoder that includes encoder 204 and decoder 206. Encoder 204 includes convolutional layers and/or other types of neural network layers that convert a 2D image 222 that depicts facial identity 250 into a corresponding latent representation 224. This latent representation 224 includes a fixed-length vector representation of the visual attributes in image 222 in a lower-dimensional space”);
generating, via a neural network based transformation model, an updated latent representation based on the latent representation of the image (see at least [0039] For example, training engine 122 could perform a forward pass that uses encoder 204 to convert a batch of multiscopic training images 230 into corresponding training latent representations 212 and uses one or more instances of decoder 206 to convert training latent representations 212 into decoder output 210. Training engine 122 could compute a mean squared error (MSE), L1 loss, and/or another type of reconstruction loss 232 that measures the differences between multiscopic training images 230 and decoder output 210 that corresponds to reconstructions of multiscopic training images 230. Training engine 122 could then perform a backward pass that backpropagates reconstruction loss 232 across layers of machine learning model 200 and uses stochastic gradient descent to update parameters (e.g., neural network weights) of machine learning model 200 based on the negative gradients of the backpropagated reconstruction loss 232. Training engine 122 could repeat the forward and backward passes with different batches of training data 214 and/or over a number of training epochs and/or iterations until reconstruction loss 232 falls below a threshold and/or another condition is met. [0040] By training machine learning model 200 in a way that minimizes reconstruction loss 232, training engine 122 generates a trained encoder 204 that learns lower-dimensional latent representations 224 of multiscopic training images 230. Training engine 122 also generates one or more instances of a trained decoder 206 that learn different facial identities 216 associated with multiscopic training images 230 and can reconstruct the original multiscopic training images 230 from the corresponding latent representations 224 outputted by the trained encoder 204. ;
generating a target pose latent representation based on the latent representation, the view direction, and a learnable parameter matrix (see at least [0030] For example, facial identity 218 could be represented by a set of densely connected neural network layers (also referred to herein as “dense layers”) that operate in conjunction with decoder 206. Input into each dense layer could include an output of a corresponding layer of decoder 206. In response to the input, the dense layer generates parameters that control the operation of one or more subsequent layers within decoder 206 and/or the values of weights within the subsequent layer(s). These parameters could include adaptive instance normalization (AdaIN) coefficients that represent facial identity 218. These parameters could also, or instead, include weights associated with neurons within the subsequent layer(s) of decoder 206. The outputted AdaIN coefficients and/or weights would be used to modify convolutional operations performed by the subsequent layer(s) in decoder 206, thereby allowing decoder 206 to generate images 226 with different facial identities. [0048] Execution engine 124 uses the trained machine learning model 200 to convert multiscopic data 220 that includes images 222 depicting a certain facial identity 250 and a certain performance (e.g., behavior, facial expression, pose, etc.) from multiple viewpoints into output data 228 that includes images 226 depicting a different facial identity 218 and the same performance from the same viewpoints. More specifically, execution engine 124 receives a set of multiscopic images 222 of a face at a certain time with a certain facial identity 250. Each image 222 in the set of multiscopic images depicts the face from a different viewpoint. Execution engine 124 uses the trained encoder 204 to convert each image 222 into a corresponding latent representation 224. Execution engine 124 also uses the trained decoder 206 to convert each latent representation 224 into a corresponding image 226 that includes the same performance and viewpoint as the original image represented by the latent representation but depicts a face with a different facial identity 218 than that of the original image. As mentioned above, this facial identity 218 can be represented by a certain instance of decoder 206, an identity vector, and/or a set of dense layers that generate AdaIN coefficients or neural network weights that control the operation of decoder 206.”; [0057] In step 304, training engine 122 executes one or more decoder neural networks that convert the set of latent representations into a set of output images. Continuing with the above example, training engine 122 could input some or all of the latent representations into each decoder neural network. Training engine 122 could use one or more convolutional layers and/or other types of neural network layers in the decoder neural network to convert the latent representations into the output images. [0058] As mentioned above, the input images can depict multiple facial identities (e.g., combinations of personal identities, ages, lighting conditions, etc.). Each of these facial identities can be learned by a different decoder neural network. As a result, training engine 122 can perform step 304 by inputting different subsets of latent representations associated with different facial identities into different decoder neural networks and using each decoder neural network to convert the corresponding subset of latent representations into output images. [0059] Alternatively, the same decoder neural network can be used to convert all latent representations generated by the encoder neural network into output images. In this scenario, one or more sets of dense layers are used to generate AdaIN coefficients, neural network weights, and/or other parameters that control the operation of the decoder neural network in converting the latent representations into the output images. These parameters can be adapted to different facial identities associated with the original input images.”), and generating, via a decoder, an output image based on the target pose latent representation (see at least [0042] After machine learning model 200 has been trained to reconstruct multiscopic training images 230 and/or non-multiscopic training images, training engine 122 further trains machine learning model 200 based on a geometry loss 234 between training geometries 236 associated with different sets of multiscopic training images 230 and output geometries 208 associated with decoder output 210 generated by machine learning model 200 from these sets of multiscopic training images 230. More specifically, training engine 122 inputs a given set of multiscopic training images 230 into encoder 204 and uses encoder 204 to convert the inputted multiscopic training images 230 into corresponding training latent representations 212. Training engine 122 also uses one or more instances of decoder 206 associated with one or more facial identities 216 to convert each of training latent representations 212 into decoder output 210 that includes an output image. This output image would depict the performance from the original image that was converted into the training latent representation and a selected facial identity that differs from the original image. [0050] After machine learning model 200 generates output data 228 that includes a set of images 226 with a different facial identity 218 from the original facial identity 250 depicted within a corresponding set of input images 222, execution engine 124 can use geometry estimation model 202 to assess a geometric consistency 244 between images 222 and images 226. In some embodiments, geometric consistency 244 includes a measure of the distance or difference between a first geometry 240 associated with two or more images 222 and a second geometry 242 associated with two or more corresponding images 226. For example, geometric consistency 244 could include an aggregation (e.g., average, weighted average, sum, weighted sum, etc.) of distances between 3D points in geometry 240 and corresponding 3D points in geometry 242.”; [0059] Alternatively, the same decoder neural network can be used to convert all latent representations generated by the encoder neural network into output images. In this scenario, one or more sets of dense layers are used to generate AdaIN coefficients, neural network weights, and/or other parameters that control the operation of the decoder neural network in converting the latent representations into the output images. These parameters can be adapted to different facial identities associated with the original input images.)” In addition, the same motivation is used as the rejection for claim 13. Gu and WEBER are understood to be silent on the remaining limitations of claim 17
In the same field of endeavor, LEE teaches receiving, via a data interface ([0070] The functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in hardware, an example hardware configuration may comprise a processing system in a device. The processing system may be implemented with a bus architecture. The bus may include any number of interconnecting buses and bridges depending on the specific application of the processing system and the overall design constraints. The bus may link together various circuits including a processor, machine-readable media, and a bus interface. The bus interface may connect a network adapter, among other things, to the processing system via the bus. The network adapter may implement signal processing functions. For certain aspects, a user interface (e.g., keypad, display, mouse, joystick, etc.) may also be connected to the bus. The bus may also link various other circuits such as timing sources, peripherals, voltage regulators, power management circuits, and the like, which are well known in the art, and therefore, will not be described any further.”), an input image and a view direction ([0044] Representatively, an input image sequence 402 of the CNN-LSTM framework 400 is a chunk of a video sequence, typically sampled by window-sliding along the temporal direction”);
generating, via a neural network based transformation model, an updated latent representation based on the latent representation of the image ([0046] head, Each image is forwarded to certain layers of a CNN 410 (e.g., CNN.sub.head) to obtain deep features in the l.sup.th layers, denoted by Z.sup.l={z.sub.t.sup.l}.sub.t=1.sup.T. The input of the spatial attention model 420, z.sub.t.sup.l, is forwarded to two 2-D convolutional layers, and the output is concatenated with last hidden variable h.sub.t−1 from an LSTM network 430. Then, the concatenated tensors are forwarded to the fully connected layers (FC), the hyperbolic tangent (tanh) layer, and the softmax layer, to obtain the attention weights α.sub.t. The attention weights element-wise product with z.sub.t.sup.l is computed to obtain selective deep features, which are then forwarded to the remaining portions of layers of the CNN 410 (e.g., CNN.sub.tail) for latent features Z.sup.f={z.sub.t.sup.f}.sub.i=1.sup.T. The latent features are then forwarded to the LSTM network 430 for encoding temporal dependencies.[0047] The input and the output of the LSTM network 430 are recurrent over time steps. At time step t, the input is the latent feature z.sub.t.sup.f from the last fully connected layer of the CNN 410, the hidden variable h.sub.t−1, and the internal states c.sub.t−1, while the output is the hidden unit h.sub.t and a memory cell c.sub.t for the next time step. Both h.sub.t and c.sub.t are updated and then passed to the LSTM network 430 at each time step. [0048] The LSTM network 430 outputs a set of hidden variables H={h.sub.t}.sub.i=1.sup.T and a set of memory cells C={c.sub.t}.sub.i=1.sup.T, which are used in the temporal attention model 440. The temporal attention model 440 calculates the attention by dot-production between decoder context and encoder representations. Instead of multiple decoder layers, however, a single layer is used for the output of the LSTM network 430. The inputs h.sub.t and c.sub.t, are fused to be a set of state summaries D={d.sub.t}.sub.i=1.sup.T. Then, H and D perform a matrix production, where the results are followed by a softmax layer for attention selection. The attention weights are then applied to H. The adjusted hidden variables H′ are followed by fully connected layers and the tanh layer to obtain class probability distribution P={P.sub.t}.sub.i=1.sup.T.”;
generating a target pose latent representation based on the latent representation, the view direction, and a learnable parameter matrix ([0046] head, Each image is forwarded to certain layers of a CNN 410 (e.g., CNN.sub.head) to obtain deep features in the l.sup.th layers, denoted by Z.sup.l={z.sub.t.sup.l}.sub.t=1.sup.T. The input of the spatial attention model 420, z.sub.t.sup.l, is forwarded to two 2-D convolutional layers, and the output is concatenated with last hidden variable h.sub.t−1 from an LSTM network 430. Then, the concatenated tensors are forwarded to the fully connected layers (FC), the hyperbolic tangent (tanh) layer, and the softmax layer, to obtain the attention weights α.sub.t. The attention weights element-wise product with z.sub.t.sup.l is computed to obtain selective deep features, which are then forwarded to the remaining portions of layers of the CNN 410 (e.g., CNN.sub.tail) for latent features Z.sup.f={z.sub.t.sup.f}.sub.i=1.sup.T. The latent features are then forwarded to the LSTM network 430 for encoding temporal dependencies. [0047] The input and the output of the LSTM network 430 are recurrent over time steps. At time step t, the input is the latent feature z.sub.t.sup.f from the last fully connected layer of the CNN 410, the hidden variable h.sub.t−1, and the internal states c.sub.t−1, while the output is the hidden unit h.sub.t and a memory cell c.sub.t for the next time step. Both h.sub.t and c.sub.t are updated and then passed to the LSTM network 430 at each time step. ; [0051] Based on the outputs provided by the LSTM network 430, the temporal attention model 440, and the spatial attention model 420 are integrated into the CNN-LSTM framework 400. For example, a temporal attention β.sub.t,u corresponding to the t.sup.th state summary and the u.sup.th hidden variable is computed as a dot-product between d.sub.t and h.sub.u, and then followed by a soft-selection:
[00002]βt,u=exp(dt.Math.hu).Math.k=1T.Math.exp(dt.Math.hk).(4)
The state summary d.sub.t at time step t is defined by:
d.sub.t=W.sub.d.sup.hh.sub.t+W.sub.d.sup.c tanh(c.sub.t)+b.sub.d.sup.hc, (5) Where W.sub.d.sup.h, W.sub.d.sup.c are the learnable parameter matrices, and b.sub.d is a bias vector. This implies how the u.sup.th input contributes to the t.sup.th output. Therefore, the output hidden variable is adjusted according to β.sub.t,u:..”; 0052] The adjusted hidden variable h′.sub.t is then forwarded to the fully connected layers and tanh layer to obtain final prediction p.sub.t for the t.sup.th time step:
p.sub.t=W.sub.p tanh (W.sub.p.sup.hh′.sub.i+W.sub.p.sup.cc.sub.j+b.sub.p.sup.hc)+b.sub.p, (7)
Where W.sub.p, W.sub.p.sup.h, W.sub.p.sup.c are the learnable parameter matrices, and b.sub.p.sup.hc, b.sub.a are the bias vectors.”); and generating, via a decoder, an output image based on the target pose latent representation ([0048] The LSTM network 430 outputs a set of hidden variables H={h.sub.t}.sub.i=1.sup.T and a set of memory cells C={c.sub.t}.sub.i=1.sup.T, which are used in the temporal attention model 440. The temporal attention model 440 calculates the attention by dot-production between decoder context and encoder representations. Instead of multiple decoder layers, however, a single layer is used for the output of the LSTM network 430. The inputs h.sub.t and c.sub.t, are fused to be a set of state summaries D={d.sub.t}.sub.i=1.sup.T. Then, H and D perform a matrix production, where the results are followed by a softmax layer for attention selection. The attention weights are then applied to H. The adjusted hidden variables H′ are followed by fully connected layers and the tanh layer to obtain class probability distribution P={P.sub.t}.sub.i=1.sup.T”) In addition, the same motivation is used as the rejection for claim 13.
Thus, the combination of Gu, WEBER and LEE teaches a method of image generation, the method comprising: receiving, via a data interface, a plurality of control parameters and a view direction; generating, via a mapping network, a latent representation of an image based on the plurality of control parameters; generating, via a neural network based transformation model, an updated latent representation based on the latent representation of the image; generating a target pose latent representation based on the latent representation, the view direction, and a learnable parameter matrix; and generating, via a decoder, an output image based on the target pose latent representation.
Regarding claim 19, Gu, WEBER and LEE teach the method of claim 17, wherein the neural network based transformation model is a fully-connected multi-layer perceptron (whole paper, see at least 3.1 IMAGE SYNTHESIS AS NEURAL IMPLICIT FIELD RENDERING of Gu “Style-based Generative Neural Radiance Field We start by modeling a 3D scene as neural radiance field (NeRF, Mildenhall et al., 2020). It is typically parameterized as multilayer perceptrons (MLPs), which takes the position x ∈ R3 and viewing direction d ∈ S2 as input, and predicts the density σ(x) ∈ R+ and view-dependent color c(x,d) ∈ R3. To model high-frequency details, follwing NeRF (Mildenhall et al., 2020), we map each dimension of x and d with Fourier features : ζL(x) = sin(20x),cos(20x),...,sin(2L−1x),cos(2L−1x) We formalize StyleNeRF representations by conditioning NeRF with style vectors w as follows: φn w(x) = gn w ◦gn−1 w ◦...◦g1 w ◦ζ(x), where w = f(z),z ∈ Z (1) (2) Similar as StyleGAN2 (Karras et al., 2020b), f is a mapping network that maps noise vectors from the spherical Gaussian space Z to the style space W; gi w(.) is the ith layer MLP whose weight matrix is modulated by the input style vector w. φn w(x) is the n-th layer feature of that point. We then use the extracted features to predict the density and color, respectively: σw(x) = hσ ◦φnσ w (x), cw(x,d) = hc ◦[φncw(x),ζ (d)], (3) where hσ and hc can be a linear projection or 2-layer MLPs. Different from the original NeRF, we assume nc > nσ for Equation (3) as the visual appearance generally needs more capacity to model than the geometry. The first min(nσ,nc) layers are shared in the network. Volume Rendering Image synthesis is modeled as volume rendering from a given camera pose p ∈ P. For simplicity, we assume a camera is located on the unit sphere pointing to the origin with a fixed field of view (FOV). We sample the camera’s pitch & yaw from a uniform or Gaussian distribution depending on the dataset. To render an image I ∈ RH×W×3, we shoot a camera ray r(t) = o+td(ois the camera origin) for each pixel, and then calculate the color using the volume rendering equation: ∞ INeRF w (r)= pw(t)cw(r(t),d)dt, where pw(t) = exp − 0 t σw(r(s))ds · σw(r(t)) (4) 0 In practice, the above equation is discretized by accumulating sampled points along the ray. Following NeRF (Mildenhall et al., 2020), stratified and hierarchical sampling are applied for more accurate discrete approximation to the continuous implicit function. Challenges Compared to 2D generative models (e.g., StyleGANs (Karras et al., 2019; 2020b)), the images generated by NeRF-based models have 3D consistency, which is guaranteed by modeling the image synthesis as a physics process, and the neural 3D scene representation is invariant across different viewpoints. However, the drawbacks are apparent: these models cost much more computation to render an image at the exact resolution. For example, 2D GANs are 100 ∼ 1000 times more efficient to generate a 10242 image than NeRF-based models. Furthermore, NeRF consumes much more memory to cache the intermediate results for gradient back-propagation during training, making it difficult to train on high-resolution images. Both of these restrict the scope of applying NeRF-based models in high-quality image synthesis, especially at the training stage when calculating the objective function over the whole image is crucial.”) In addition, the same motivation is used as the rejection for claim 13
Regarding claim 20, Gu, WEBER and LEE teach the method of claim 17, wherein generating the target pose latent representation includes summing the updated latent representation with a product of the view direction and the learnable parameter matrix (whole paper, see at least 3 METHOD 3.1 IMAGE SYNTHESIS AS NEURAL IMPLICIT FIELD RENDERING of Gu , “Style-based Generative Neural Radiance Field We start by modeling a 3D scene as neural radiance field (NeRF, Mildenhall et al., 2020). It is typically parameterized as multilayer perceptrons (MLPs), which takes the position x ∈ R3 and viewing direction d ∈ S2 as input, and predicts the density σ(x) ∈ R+ and view-dependent color c(x,d) ∈ R3. To model high-frequency details, follwing NeRF (Mildenhall et al., 2020), we map each dimension of x and d with Fourier features…”; see section Synthesis Network , “Considering that our training images generally have unbounded background, we choose NeRF++ (Zhang et al., 2020), a variant of NeRF, as the StyleNeRF backbone. NeRF++ consists of a foreground NeRF in a unit sphere and a background NeRF represented with inverted sphere parameterization. As shown in Figure 3, two MLPs are used to predict the density where the background network has fewer parameters than the foreground one. Then a shared MLP is employed for color prediction. Each style-conditioned block consists of an affine transformation layer and a 1 ×1 convolution layer (Conv). The Conv weights are modulated with the affine-transformed styles, and then demodulated for computation. leaky ReLU is used as non-linear activation. The number of blocks depends on the input and target image resolutions.”, Figure 3; [0031] of WEBER “ Continuing with the above example, when facial identity 218 corresponds to an interpolation between two or more facial identities represented by two or more sets of dense layers, a weighted combination of AdaIN coefficients (or neural network weights) outputted by the two or more sets of dense layers could be used to control the convolutional operations performed by the layers of decoder 206. Thus, a given facial identity 218 that corresponds to 80% of a first facial identity associated with a first set of dense layers and 20% of a second facial identity associated with a second set of dense layers could be represented by a weighted sum of AdaIN coefficients (or neural network weights) outputted by the first and second sets of dense layers. Within the weighted sum, the AdaIN coefficients (or neural network weights) outputted by the first set of dense layers would be combined with a weight of 0.8, and the AdaIN coefficients (or neural network weights) outputted by the second set of dense layers would be combined with a weight of 0.2.”; [0046] of LEE “head, Each image is forwarded to certain layers of a CNN 410 (e.g., CNN.sub.head) to obtain deep features in the l.sup.th layers, denoted by Z.sup.l={z.sub.t.sup.l}.sub.t=1.sup.T. The input of the spatial attention model 420, z.sub.t.sup.l, is forwarded to two 2-D convolutional layers, and the output is concatenated with last hidden variable h.sub.t−1 from an LSTM network 430. Then, the concatenated tensors are forwarded to the fully connected layers (FC), the hyperbolic tangent (tanh) layer, and the softmax layer, to obtain the attention weights α.sub.t. The attention weights element-wise product with z.sub.t.sup.l is computed to obtain selective deep features, which are then forwarded to the remaining portions of layers of the CNN 410 (e.g., CNN.sub.tail) for latent features Z.sup.f={z.sub.t.sup.f}.sub.i=1.sup.T. The latent features are then forwarded to the LSTM network 430 for encoding temporal dependencies. [0047] The input and the output of the LSTM network 430 are recurrent over time steps. At time step t, the input is the latent feature z.sub.t.sup.f from the last fully connected layer of the CNN 410, the hidden variable h.sub.t−1, and the internal states c.sub.t−1, while the output is the hidden unit h.sub.t and a memory cell c.sub.t for the next time step. Both h.sub.t and c.sub.t are updated and then passed to the LSTM network 430 at each time step. ; [0051] Based on the outputs provided by the LSTM network 430, the temporal attention model 440, and the spatial attention model 420 are integrated into the CNN-LSTM framework 400. For example, a temporal attention β.sub.t,u corresponding to the t.sup.th state summary and the u.sup.th hidden variable is computed as a dot-product between d.sub.t and h.sub.u, and then followed by a soft-selection:
[00002]βt,u=exp(dt.Math.hu).Math.k=1T.Math.exp(dt.Math.hk).(4)
The state summary d.sub.t at time step t is defined by:
d.sub.t=W.sub.d.sup.hh.sub.t+W.sub.d.sup.c tanh(c.sub.t)+b.sub.d.sup.hc, (5) Where W.sub.d.sup.h, W.sub.d.sup.c are the learnable parameter matrices, and b.sub.d is a bias vector. This implies how the u.sup.th input contributes to the t.sup.th output. Therefore, the output hidden variable is adjusted according to β.sub.t,u:..”; 0052] The adjusted hidden variable h′.sub.t is then forwarded to the fully connected layers and tanh layer to obtain final prediction p.sub.t for the t.sup.th time step:
p.sub.t=W.sub.p tanh (W.sub.p.sup.hh′.sub.i+W.sub.p.sup.cc.sub.j+b.sub.p.sup.hc)+b.sub.p, (7)
Where W.sub.p, W.sub.p.sup.h, W.sub.p.sup.c are the learnable parameter matrices, and b.sub.p.sup.hc, b.sub.a are the bias vectors.”) In addition, the same motivation is used as the rejection for claim 13.
5. Claims 14-15 are rejected under 35 U.S.C. 103 as being unpatentable over Gu, Jiatao, et al., IDS, "Stylenerf: A style-based 3d-aware generator for high-resolution image synthesis." arXiv preprint arXiv:2110.08985 (2021) (“Gu”) in view of WEBER et al., U.S Patent Application Publication No.20240078726 (“WEBER”) further in view of LEE et al., U.S Patent Application Publication No.20200234066 (”LEE”) further in view of Chakrabarty et al. U.S Patent Application Publication No.20220101577 (“Chakrabarty”)
Regarding claim 14, Gu, WEBER and LEE teach the method of claim 13, further comprising: receiving, via the data interface a target image; and updating parameters of at least one of the neural network based transformation model or the learnable parameter matrix based on a comparison of the target image and the output image (whole paper, see at least Discriminator & Objectives of Gu “We use the same discriminator as StyleGAN2. Following previous works (Chan et al., 2021; Niemeyer & Geiger, 2021b), StyleNeRF adopts a non-saturating GAN ob jective with R1 regularization (Mescheder et al., 2018). A new NeRF path regularization is employed to enforce 3D consistency. The final loss function is defined as follows (D is the discriminator and Gis the generator including the mapping and synthesis networks): L(D,G) = Ez∼Z,p∼P [f(D(G(z,p))]+EI∼pdata f(−D(I)+λ ∇D(I) 2)+β·LNeRF-path (9) where f(u) = −log(1+exp(−u)), and pdata is the data distribution. We set β = 0.2 and λ = 0.5. Progressive Training We train StyleNeRF progressively from low to high resolution, which makes the training more stable and efficient. We observed in the experiments that were directly training for the highest resolution easily makes the model fail to capture the object geometry. We suspect it is because both Equations (5) and (6) are just approximations to the original NeRF. Therefore, inspired by Karras et al. (2017), we propose a new three-stage progressive training strategy: For the first T1 images, we train StyleNeRF without approximation at low-resolution; then, during T1 ∼ T2 images, both the generator and discriminator linearly increase the output resolutions until reaching the target resolution; At last, we fix the architecture and continue training the model at the highest resolution until T3 images. Please refer to the Appendix A.4 for more details.”; [0047] of WEBER “Training engine 122 can also repeat training of machine learning model 200 using reconstruction loss 232 and/or geometry loss 234. For example, training engine 122 could alternate between one training stage that updates parameters of machine learning model 200 based on reconstruction loss 232 and another training stage that updates parameters of machine learning model 200 based on geometry loss 234 until one or both losses fall below corresponding thresholds. In another example, training engine 122 could perform one or more stages that train machine learning model 200 using a combination (e.g., sum) of reconstruction loss 232 and geometry loss 234. Thus, by training machine learning model 200 using both reconstruction loss 232 and geometry loss 234, training engine 122 improves the performance of machine learning model 200 both in reconstructing images of faces and in generating images with changed facial identities 216 in a geometrically consistent way.”) In addition, the same motivation is used as the rejection for claim 13. Gu, WEBER and LEE are understood to be silent on the remaining limitations of claim 14.
In the same field of endeavor, Chakrabarty teaches receiving, via the data interface ([0042] As shown in FIG. 1, the server device(s) 102 include a digital graphics system 104 which further includes the digital hairstyle transfer system 106. Indeed, in some embodiments, the digital hairstyle transfer system 106 generates a transferred hairstyle image by transferring a hairstyle depicted within a target image onto a person depicted within a source image in accordance with one or more embodiments. Furthermore, in one or more embodiments, the digital hairstyle transfer system 106 displays the transferred hairstyle image on a graphical user interface of the client device 110 (e.g., via a digital image hairstyle transfer application 112). a target image ([0049] In particular, as shown in FIG. 2, the digital hairstyle transfer system 106 identifies (or receives) a source image 202 depicting a first person with a first hairstyle and also identifies (or receives) a target image 204 depicting a second person with a second hairstyle. Subsequently (as shown in FIG. 2), the digital hairstyle transfer system 106 generates a transferred hairstyle image 206 from the source image 202 and the target image 204 (in accordance with one or more embodiments herein). Indeed, as illustrated in FIG. 2, the digital hairstyle transfer system 106 generates a generated transferred hairstyle image 206 that depicts the first person from the source image 202 with the second hairstyle from the target image 204.; and
updating parameters of at least one of the neural network based transformation model or the learnable parameter matrix based on a comparison of the target image and the output image ([0089] More specifically, the digital hairstyle transfer system 106 utilizes the face-generative neural network to project the source image into a first iteration of a source latent vector using an encoder network of the face-generative neural network. The digital hairstyle transfer system 106 then utilizes the decoder network of the face-generative neural network to decode the source latent vector into a first iteration of a synthesized source image. Furthermore, as mentioned above, the digital hairstyle transfer system 106 utilizes a loss function (e.g., an optimizer) to optimize the source latent vector projected from the source image. In particular, during an iteration, of the digital hairstyle transfer system 106 utilizes a loss function to determine a loss between the first iteration of a synthesized source image and the input image. The digital hairstyle transfer system 106 backpropagates the loss to the first iteration of the source latent vector to update the first iteration of the source latent vector to generate a second iteration of the source latent vector. The digital hairstyle transfer system 106 then iteratively repeats this process (e.g., decodes the second iteration of the source latent vector into a second synthesized source image, compares the second synthesized source image to the source image 402 to generate a loss, backpropagates the loss to update the second iteration of the source latent vector). [0153] In one or more embodiments, the act 1306 includes synthesizing a hairstyle-transfer latent vector utilizing a face-generative neural network by iteratively adjusting the hairstyle-transfer latent vector utilizing at least one loss between an iteration of the transferred hairstyle image and a face mask from a source image or a second hairstyle mask from a target image. In some embodiments, the act 1306 includes determining a source loss based on a comparison between an iteration of a transferred hairstyle image and a face mask from a source image. For example, a face mask from a source image corresponds to regions depicting facial attributes within the source image. Furthermore, in some embodiments, the act 1306 includes determining a target loss based on a comparison between an iteration of a transferred hairstyle image and a second hairstyle mask from a target image. For instance, a second hairstyle mask from a target image corresponds to regions depicting the second hairstyle within the target image. In one or more embodiments, the act 1306 includes utilizing a discriminator to determine a discriminator loss from an iteration of a transferred hairstyle image. In some embodiments, the act 1306 includes iteratively adjusting a hairstyle-transfer latent vector utilizing a combination of a source loss, a target loss, and/or a discriminator loss.”)
Therefore, in combination of Gu, WEBER and LEE, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify a style-based 3D-aware generator for high-resolution image synthesis of Gu with iteratively repeating compares the second synthesized source image to the source image to generate a loss, backpropagates the loss to update the second iteration of the source latent vector as seen in Chakrabarty because this modification would generate a second iteration of the source latent vector ([0089] of Chakrabarty).
Thus, the combination of Gu, WEBER, LEE and Chakrabarty teaches further comprising: receiving, via the data interface a target image; and updating parameters of at least one of the neural network based transformation model or the learnable parameter matrix based on a comparison of the target image and the output image.
Regarding claim 15, Gu, WEBER, LEE teach the method of claim 13, further comprising: receiving, via the data interface a target image; generating, via the encoder, a latent representation of the target image and updating parameters of at least one of the neural network based transformation model or the learnable parameter matrix based on a comparison of the target pose latent representation of the input image and the latent representation of the target image (whole paper, see at least Discriminator & Objectives of Gu “We use the same discriminator as StyleGAN2. Following previous works (Chan et al., 2021; Niemeyer & Geiger, 2021b), StyleNeRF adopts a non-saturating GAN ob jective with R1 regularization (Mescheder et al., 2018). A new NeRF path regularization is employed to enforce 3D consistency. The final loss function is defined as follows (D is the discriminator and Gis the generator including the mapping and synthesis networks): L(D,G) = Ez∼Z,p∼P [f(D(G(z,p))]+EI∼pdata f(−D(I)+λ ∇D(I) 2)+β·LNeRF-path (9) where f(u) = −log(1+exp(−u)), and pdata is the data distribution. We set β = 0.2 and λ = 0.5. Progressive Training We train StyleNeRF progressively from low to high resolution, which makes the training more stable and efficient. We observed in the experiments that were directly training for the highest resolution easily makes the model fail to capture the object geometry. We suspect it is because both Equations (5) and (6) are just approximations to the original NeRF. Therefore, inspired by Karras et al. (2017), we propose a new three-stage progressive training strategy: For the first T1 images, we train StyleNeRF without approximation at low-resolution; then, during T1 ∼ T2 images, both the generator and discriminator linearly increase the output resolutions until reaching the target resolution; At last, we fix the architecture and continue training the model at the highest resolution until T3 images. Please refer to the Appendix A.4 for more details.”; [0047] of WEBER “Training engine 122 can also repeat training of machine learning model 200 using reconstruction loss 232 and/or geometry loss 234. For example, training engine 122 could alternate between one training stage that updates parameters of machine learning model 200 based on reconstruction loss 232 and another training stage that updates parameters of machine learning model 200 based on geometry loss 234 until one or both losses fall below corresponding thresholds. In another example, training engine 122 could perform one or more stages that train machine learning model 200 using a combination (e.g., sum) of reconstruction loss 232 and geometry loss 234. Thus, by training machine learning model 200 using both reconstruction loss 232 and geometry loss 234, training engine 122 improves the performance of machine learning model 200 both in reconstructing images of faces and in generating images with changed facial identities 216 in a geometrically consistent way.”) In addition, the same motivation is used as the rejection for claim 13. Gu, WEBER and LEE are understood to be silent on the remaining limitations of claim 15.
In the same field of endeavor, Chakrabarty teaches receiving, via the data interface ([0042] As shown in FIG. 1, the server device(s) 102 include a digital graphics system 104 which further includes the digital hairstyle transfer system 106. Indeed, in some embodiments, the digital hairstyle transfer system 106 generates a transferred hairstyle image by transferring a hairstyle depicted within a target image onto a person depicted within a source image in accordance with one or more embodiments. Furthermore, in one or more embodiments, the digital hairstyle transfer system 106 displays the transferred hairstyle image on a graphical user interface of the client device 110 (e.g., via a digital image hairstyle transfer application 112). a target image ([0049] In particular, as shown in FIG. 2, the digital hairstyle transfer system 106 identifies (or receives) a source image 202 depicting a first person with a first hairstyle and also identifies (or receives) a target image 204 depicting a second person with a second hairstyle. Subsequently (as shown in FIG. 2), the digital hairstyle transfer system 106 generates a transferred hairstyle image 206 from the source image 202 and the target image 204 (in accordance with one or more embodiments herein). Indeed, as illustrated in FIG. 2, the digital hairstyle transfer system 106 generates a generated transferred hairstyle image 206 that depicts the first person from the source image 202 with the second hairstyle from the target image 204.; generating, via the encoder, a latent representation of the target image; and updating parameters of at least one of the neural network based transformation model or the learnable parameter matrix based on a comparison of the target pose latent representation of the input image and the latent representation of the target image. ([0089] More specifically, the digital hairstyle transfer system 106 utilizes the face-generative neural network to project the source image into a first iteration of a source latent vector using an encoder network of the face-generative neural network. The digital hairstyle transfer system 106 then utilizes the decoder network of the face-generative neural network to decode the source latent vector into a first iteration of a synthesized source image. Furthermore, as mentioned above, the digital hairstyle transfer system 106 utilizes a loss function (e.g., an optimizer) to optimize the source latent vector projected from the source image. In particular, during an iteration, of the digital hairstyle transfer system 106 utilizes a loss function to determine a loss between the first iteration of a synthesized source image and the input image. The digital hairstyle transfer system 106 backpropagates the loss to the first iteration of the source latent vector to update the first iteration of the source latent vector to generate a second iteration of the source latent vector. The digital hairstyle transfer system 106 then iteratively repeats this process (e.g., decodes the second iteration of the source latent vector into a second synthesized source image, compares the second synthesized source image to the source image 402 to generate a loss, backpropagates the loss to update the second iteration of the source latent vector). [0153] In one or more embodiments, the act 1306 includes synthesizing a hairstyle-transfer latent vector utilizing a face-generative neural network by iteratively adjusting the hairstyle-transfer latent vector utilizing at least one loss between an iteration of the transferred hairstyle image and a face mask from a source image or a second hairstyle mask from a target image. In some embodiments, the act 1306 includes determining a source loss based on a comparison between an iteration of a transferred hairstyle image and a face mask from a source image. For example, a face mask from a source image corresponds to regions depicting facial attributes within the source image. Furthermore, in some embodiments, the act 1306 includes determining a target loss based on a comparison between an iteration of a transferred hairstyle image and a second hairstyle mask from a target image. For instance, a second hairstyle mask from a target image corresponds to regions depicting the second hairstyle within the target image. In one or more embodiments, the act 1306 includes utilizing a discriminator to determine a discriminator loss from an iteration of a transferred hairstyle image. In some embodiments, the act 1306 includes iteratively adjusting a hairstyle-transfer latent vector utilizing a combination of a source loss, a target loss, and/or a discriminator loss.”) In addition, the same motivation is used as the rejection for claim 14.
Thus, the combination of Gu, WEBER, LEE and Chakrabarty teaches further comprising: receiving, via the data interface a target image; generating, via the encoder, a latent representation of the target image; and updating parameters of at least one of the neural network based transformation model or the learnable parameter matrix based on a comparison of the target pose latent representation of the input image and the latent representation of the target image.
6. Claim 18 is rejected under 35 U.S.C. 103 as being unpatentable over Gu, Jiatao, et al., IDS, "Stylenerf: A style-based 3d-aware generator for high-resolution image synthesis." arXiv preprint arXiv:2110.08985 (2021) (“Gu”) in view of WEBER et al., U.S Patent Application Publication No.20240078726 (“WEBER”) further in view of LEE et al., U.S Patent Application Publication No.20200234066 (”LEE”) further in view of Jongsma et al, U.S Patent Application Publication No.20230054256 (“Jongsma”)
Regarding claim 18, Gu, WEBER and LEE teach the method of claim 17, wherein the view direction (whole paper, see at least 3.1 IMAGE SYNTHESIS AS NEURAL IMPLICIT FIELD RENDERING of Gu “Style-based Generative Neural Radiance Field We start by modeling a 3D scene as neural radiance field (NeRF, Mildenhall et al., 2020). It is typically parameterized as multilayer perceptrons (MLPs), which takes the position x ∈ R3 and viewing direction d ∈ S2 as input, and predicts the density σ(x) ∈ R+ and view-dependent color c(x,d) ∈ R3. To model high-frequency details, follwing NeRF (Mildenhall et al., 2020), we map each dimension of x and d with Fourier features : ζL(x) = sin(20x),cos(20x),...,sin(2L−1x),cos(2L−1x) We formalize StyleNeRF representations by conditioning NeRF with style vectors w as follows: φn w(x) = gn w ◦gn−1 w ◦...◦g1 w ◦ζ(x), where w = f(z),z ∈ Z (1) (2) Similar as StyleGAN2 (Karras et al., 2020b), f is a mapping network that maps noise vectors from the spherical Gaussian space Z to the style space W; gi w(.) is the ith layer MLP whose weight matrix is modulated by the input style vector w. φn w(x) is the n-th layer feature of that point. We then use the extracted features to predict the density and color, respectively: σw(x) = hσ ◦φnσ w (x), cw(x,d) = hc ◦[φncw(x),ζ (d)], (3) where hσ and hc can be a linear projection or 2-layer MLPs. Different from the original NeRF, we assume nc > nσ for Equation (3) as the visual appearance generally needs more capacity to model than the geometry. The first min(nσ,nc) layers are shared in the network. Volume Rendering Image synthesis is modeled as volume rendering from a given camera pose p ∈ P. For simplicity, we assume a camera is located on the unit sphere pointing to the origin with a fixed field of view (FOV). We sample the camera’s pitch & yaw from a uniform or Gaussian distribution depending on the dataset. To render an image I ∈ RH×W×3, we shoot a camera ray r(t) = o+td(ois the camera origin) for each pixel, and then calculate the color using the volume rendering equation: ∞ INeRF w (r)= pw(t)cw(r(t),d)dt, where pw(t) = exp − 0 t σw(r(s))ds · σw(r(t)) (4) 0 In practice, the above equation is discretized by accumulating sampled points along the ray. Following NeRF (Mildenhall et al., 2020), stratified and hierarchical sampling are applied for more accurate discrete approximation to the continuous implicit function. Challenges Compared to 2D generative models (e.g., StyleGANs (Karras et al., 2019; 2020b)), the images generated by NeRF-based models have 3D consistency, which is guaranteed by modeling the image synthesis as a physics process, and the neural 3D scene representation is invariant across different viewpoints. However, the drawbacks are apparent: these models cost much more computation to render an image at the exact resolution. For example, 2D GANs are 100 ∼ 1000 times more efficient to generate a 10242 image than NeRF-based models. Furthermore, NeRF consumes much more memory to cache the intermediate results for gradient back-propagation during training, making it difficult to train on high-resolution images. Both of these restrict the scope of applying NeRF-based models in high-quality image synthesis, especially at the training stage when calculating the objective function over the whole image is crucial.”) Gu, WEBER and LEE are understood to be silent on the remaining limitations of claim 18..
In the same field of endeavor, Jongsma teaches wherein the view direction is represented as a pitch value and a yaw value ([0076] In one example, the registration algorithm 84 uses a set of synthetic view projection images 52 generated from the DGM 50. Each of the synthetic 52 view projection images is generated as if it is acquired from a predetermined viewpoint (e.g. position coordinates of the camera image centre, relative to a DGM-fixed or other earth-fixed reference frame) and viewing direction (e.g. pitch, yaw, and roll angles for of the optical axis through this image centre, relative to this fixed reference frame).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify a style-based 3D-aware generator for high-resolution image synthesis of Gu, WEBER and LEE with view direction is represented as a pitch value and a yaw value as seen in Jongsma because this modification would achieve the expected benefits of providing precise 3D spatial orientation.
Thus, the combination of Gu, WEBER, LEE and Jongsma teaches wherein the view direction is represented as a pitch value and a yaw value.
Contact
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SARAH LE whose telephone number is (571)270-7842. The examiner can normally be reached Monday: 8AM-4:30PM EST, Tuesday: 8 AM-3:30PM EST, Wednesday: 8AM-2:30PM EST, Thursday and Friday off.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kent Chang can be reached at (571) 272-7667. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SARAH LE/Primary Examiner, Art Unit 2614