CTNF 18/903,151 CTNF 88707 Notice of Pre-AIA or AIA Status 07-03-aia AIA 15-10-aia The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA. DETAILED ACTION Claim Rejections - 35 USC § 112 07-30-02 AIA The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. 07-34-01 Claims 13-14 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. 07-34-05 AIA Claim 13 recites the limitation " the content input" in line 6 and “the style embedding ” in line 11 . There is insufficient antecedent basis for this limitation in the claim. 07-34-05 AIA Claim 14 recites the limitation " the text prompt " in line 4 . There is insufficient antecedent basis for this limitation in the claim. Claim Rejections - 35 USC § 103 07-06 AIA 15-10-15 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. 07-20-aia AIA The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 07-23-aia AIA The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. 07-21-aia AIA 1. Claims 1-2, 6-18 are reject ed under 35 U.S.C. 103 as being unpatentable over Karpma n et al., U.S Patent No.11,995,803 (“Karpman”) in view of Liu et al, U.S Patent Application Publication No.20210358164 (“Liu”) Regard ing independent claim 1, Karpman teaches method for image generation (abstract, ”In some embodiments, a method receives a text prompt and executes a text encoder on the text prompt to generate an embedding representation. A set of base images is generated based on the embedding representation and parameters of a base image generation model. A high resolution model is executed to upsample one or more base images in the set of base images based on parameters of the high resolution model to generate a set of final images.”) comprising: obtaining a text prompt and a style input, wherein the text prompt describes image content and the style input describes an image style (see at least col.20, lines 19-22 as shown in Fig. 4A “Referring back to FIG. 4A, the interactive text field 402 prompts and enables the user to input text that describes an image they wish to generate via the text-to-image diffusion model 112”; col.20, lines “; col.20, lines 56-67 “The generation interface 400 also includes a style menu 404 that enables the user to browse and select among pre-set image styles for the image generation request . In the example of FIG. 4A, the style menu 404 displays a set (e.g., array) of style option tiles, each style option tile including a text description of the image style (e.g., anime, Van Gogh, oil painting, line drawing, digital art, etc.) and a sample image in the corresponding style. In response to detecting a horizontal swipe (or scroll) input on a display or trackpad over the style menu area, the software application layer 124 can transition the display of the style menu to replace currently displayed style option tiles with other style option tiles”); generating, using a text encoder, a text embedding based on the text prompt, wherein the text embedding represents the image content (see at least col.2, lines 51-63 “ As shown in FIG. 1, the system includes a text-to-image diffusion model 112. Text-to-image diffusion model 112 may be a probabilistic generative model used to generate image data. In some embodiments, text-to-image diffusion model 112 may include multiple sub-models that improve the image generation. For example, text-to-image diffusion model 112 may define a (set of) pre-trained text encoders 118 (e.g., one or more pre-trained language models), base image diffusion models 120, and high-resolution diffusion models 116. Text encoders 118 interpret a text query and generate an embedding of the text query . Base image diffusion models 120 generate a base image (e.g., an initial, low-resolution image) from the embedding”); generating, using a style encoder , a style embedding based on the style input, wherein the style embedding represents the image style (see at least col.4,lines 46-57 “During pre-training and/or training stages, the system can execute the multi-modal encoder-decoder 126 on text prompts included in the training corpus to generate embeddings used by the text-to-image diffusion model 112 to generate images (e.g., in addition to embeddings generated by the pre-trained text encoder 118, instead of embeddings generated by the pre-trained text encoder 118 or via a dropout method). In some implementations, a multimodal encoder-decoder 126 includes and/or interfaces with a vision transformer configured to divide an input image into segments and generate a corresponding sequence of embeddings for the input image”;); and generating, using an image generation model, a synthetic image based on the text embedding and the style embedding, wherein the text embedding is provided to the image generation model at a first step and the style embedding is provided to the image generation model at a second step after the first step (col.5, lines 15-62 “ In some implementations, the base image diffusion model 120 defines a deep learning network (e.g., a convolutional neural network, a residual neural network, etc.) configured (e.g., through the training described) to generate images from random (e.g., Gaussian) noise based on text prompts and/or descriptions . The base image diffusion model 120 can include a U-net architecture (e.g., Efficient U-Net) defined from residual and multi-head attention blocks that enable the base image diffusion model 120 to progressively denoise (e.g., infill, generate, augment) image data according to cross-attention inputs based on the text prompt. T he base image diffusion model 120 can therefore: receive one or more text embeddings from the set of pre-trained text encoders 118 ; receive and/or initialize a (randomly sampled) noise distribution at a preset resolution (e.g., 64 pixels by 64 pixels); and transform the noise distribution into a base image at the preset resolution based on the one or more text embeddings and parameters, weights, and/or paths corresponding to an iterative denoising process learned by the base image diffusion model 120 during training. The system can then pass the base image to the set of high-resolution diffusion models 116 for upsampling and output. In some implementations, e ach high-resolution diffusion model 116 in the set of high-resolution diffusion models 116 defines a deep learning network configured to receive a low-resolution base image (e.g., 64 pixels by 64 pixels, 256 pixels by 256 pixels) and generate a higher-resolution version (e.g., copy) of the base image (e.g., 256 pixels by 256 pixels, 1024 pixels by 1024 pixels) . Generally, a high-resolution version shares a similar architecture with the base image diffusion (e.g., U-net, efficient U-net). However, self-attention layers in the base diffusion model architecture can be omitted to improve memory efficiency and inference time. During training, the set of high-resolution diffusion models 116 can be conditioned on text information (e.g., text descriptions of training images), noise augmentations, and/or visual information (e.g., embeddings of low-resolution images generated by the multimodal encoder-decoder) . Thus, during operation, each high-resolution diffusion model 116 can implement an iterative denoising process similar to the base image diffusion model 120 in order to progressively upsample generated base images to higher resolution, infill, infer, and/or generate additional visual detail and/or texture, and remove visual artifacts generated by the base image diffusion model 120. As discussed above, high-resolution diffusion models 116 operate in the pixel space to upsample the base images output by base image diffusion models 120 .” Where visual information (e.g., embeddings of low-resolution images generated by the multimodal encoder-decoder is considered as style embedding ). Karpman is understood to be silent on the remaining limitations of claim 1. In the same field of endeavor, Liu teaches obtaining a style input, wherein the style input describes an image style ([0091]as show in Fig.4A “ In at least one embodiment, a content input 404 is an image or image data of any type containing one or more objects to which a style from a style input 406 is to be applied by a generator 402. In at least one embodiment, a style input 406 is one or more images or image data of any type containing a style to be applied to a content input 404 by a generator 402. When multiple style inputs 406 are used, in an embodiment, a generator 402 extracts a style from each style input 406 and uses an average style to be applied to a content input 404.”); generating, using a encoder, a embedding, wherein the embedding represents the image content ([0092] In at least one embodiment, a generator 402 comprises a content encoder E.sub.c 408, a style encoder E.sub.s 410, and an image decoder F 412. In at least one embodiment, a content encoder E.sub.c 408 is data values and a set of software instructions that, when executed, implement one or more neural networks or models that take a content input x.sub.c 404 and output a content embedding z.sub.c. In at least one embodiment, a content encoder E.sub.c 408 utilizes vanilla convolutional layers. In at least one embodiment, a content embedding z.sub.c is a vector or set of values containing continuous numbers that represent information about a content input x.sub.c 404, such as features or objects to which a style is to be applied ” ) .; generating, using a style encoder, a style embedding based on the style input, wherein the style embedding represents the image style ([0093] In at least one embodiment, a style encoder E.sub.s 410 is data values and a set of software instructions that, when executed, implement one or more neural networks or models that take one or more style inputs x.sub.s 406 and output a style embedding z.sub.s. In at least one embodiment, a style embedding z.sub.s is a vector or set of numbers containing continuous values that represent style information about a style input x.sub.s 406. In at least one embodiment, a style embedding z.sub.s contains other information about a style input x.sub.s 406 such as pose of objects contained in said style input x.sub.s 406. In at least one embodiment, a style encoder E.sub.s 410 that takes two or more style images as inputs 406 outputs an individual style embedding z.sub.s.sup.i for each input i, and a generator 402 then averages each style embedding to determine z.sub.s that will be applied to a content input 404.”); and generating, using an image generation model, a synthetic image based on the embedding and the style embedding ([0094] In at least one embodiment, an image decoder F 412 generates an output x 416 using a content embedding z.sub.c and information from a style embedding z.sub.s, as described above. In at least one embodiment, an image decoder F 412 is data values and a set of software instructions that, when executed, implement one or more neural networks or models that construct an output 416 based on a content embedding z.sub.c and adaptive instance normalization (AdaIN) parameters 414. In at least one embodiment, an image decoder F 412 utilizes vanilla convolutional layers. In at least one embodiment, AdaIN parameters 414 are numerical data values generated based on a style embedding z.sub.s output from a style encoder E.sub.s 410. In at least one embodiment, mean and scale parameters of AdaIN parameters 414 are computed or generated based on a style embedding z.sub.s output from a style encoder E.sub.s 414 .” ) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of claimed invention to modify the method of generating images base on text prompt of Karpman with obtaining content input and style input as seen in Liu because this modification would generate a styled output image based on a style of a first image and content of a second image ([0001] of Liu) Thus, the combination of Karpman and Liu teaches a method for image generation, comprising: obtaining a text prompt and a style input, wherein the text prompt describes image content and the style input describes an image style; generating, using a text encoder, a text embedding based on the text prompt, wherein the text embedding represents the image content; generating, using a style encoder, a style embedding based on the style input, wherein the style embedding represents the image style; and generating, using an image generation model, a synthetic image based on the text embedding and the style embedding, wherein the text embedding is provided to the image generation model at a first step and the style embedding is provided to the image generation model at a second step after the first step . Regarding claim 2, Karpman and Liu teach t he method of claim 1, wherein generating the synthetic image comprises: performing, using the image generation model, a reverse diffusion process including a plurality of diffusion time steps ( see at least col. 13, lines 51-67-col.14, lines 1-3 “The system 100 can then execute the text-to-image diffusion model 112 on images within the modified training corpus to train the base image diffusion model 120. Generally, during pre-training, the system can iteratively add or inject noise (e.g., Gaussian blur) to training images and execute the base image diffusion model 120 on (or otherwise expose the image model to) each step in this iterative noising process with a training objective to reverse this noising process in a way that maximizes the likelihood of recovering the initial training image(s ). Thus, the base image diffusion model 120 can sample and parameterize each step of this iterative noising process in order to infer weights, parameters, and/or paths for reversing each noising iteration back towards the initial distribution (e.g., the unmodified training image). During the iterative noising and/or denoising processes, the system 100 can also inject timestep information into embedding representations of training images (e.g., via a timestep encoding vector) in order to condition the base image diffusion on representations of the (re)generated image at each timestep . ”) , wherein the text embedding is provided to the image generation model during a first portion of the plurality of diffusion time steps including the first step (see at least col.5, lines 15- 36 “In some implementations, the base image diffusion model 120 defines a deep learning network (e.g., a convolutional neural network, a residual neural network, etc.) configured (e.g., through the training described) to generate images from random (e.g., Gaussian) noise based on text prompts and/or descriptions. The base image diffusion model 120 can include a U-net architecture (e.g., Efficient U-Net) defined from residual and multi-head attention blocks that enable the base image diffusion model 120 to progressively denoise (e.g., infill, generate, augment) image data according to cross-attention inputs based on the text prompt. The base image diffusion model 120 can therefore: receive one or more text embeddings from the set of pre-trained text encoders 118; receive and/or initialize a (randomly sampled) noise distribution at a preset resolution (e.g., 64 pixels by 64 pixels); and transform the noise distribution into a base image at the preset resolution based on the one or more text embeddings and parameters, weights, and/or paths corresponding to an iterative denoising process learned by the base image diffusion model 120 during training. The system can then pass the base image to the set of high-resolution diffusion models 116 for upsampling and output.”) and both the text embedding and the style embedding are provided to the image generation model during a second portion of the plurality of diffusion time steps including the second step and following the first portion (see at least col.5, lines 37-62 “In some implementations, each high-resolution diffusion model 116 in the set of high-resolution diffusion models 116 defines a deep learning network configured to receive a low-resolution base image (e.g., 64 pixels by 64 pixels, 256 pixels by 256 pixels) and generate a higher-resolution version (e.g., copy) of the base image (e.g., 256 pixels by 256 pixels, 1024 pixels by 1024 pixels). Generally, a high-resolution version shares a similar architecture with the base image diffusion (e.g., U-net, efficient U-net). However, self-attention layers in the base diffusion model architecture can be omitted to improve memory efficiency and inference time. During training, the set of high-resolution diffusion models 116 can be conditioned on text information (e.g., text descriptions of training images), noise augmentations, and/or visual information (e.g., embeddings of low-resolution images generated by the multimodal encoder-decoder). T hus, during operation, each high-resolution diffusion model 116 can implement an iterative denoising process similar to the base image diffusion model 120 in order to progressively upsample generated base images to higher resolution, infill, infer, and/or generate additional visual detail and/or texture, and remove visual artifacts generated by the base image diffusion model 120 . As discussed above, high-resolution diffusion models 116 operate in the pixel space to upsample the base images output by base image diffusion models 120 . ”) Regarding claim 6, Karpman and Liu teach the method of claim 1, wherein generating the style embedding comprises: encoding the style input using a multimodal image encoder to obtain the style embedding, wherein the style input comprises an image ( see at least col.4,lines 46-57 of Karpman “During pre-training and/or training stages, the system can execute the multi-modal encoder-decoder 126 on text prompts included in the training corpus to generate embeddings used by the text-to-image diffusion model 112 to generate images (e.g., in addition to embeddings generated by the pre-trained text encoder 118, instead of embeddings generated by the pre-trained text encoder 118 or via a dropout method). In some implementations, a multimodal encoder-decoder 126 includes and/or interfaces with a vision transformer configured to divide an input image into segments and generate a corresponding sequence of embeddings for the input image”; Liu : [0091] “In at least one embodiment, a content input 404 is an image or image data of any type containing one or more objects to which a style from a style input 406 is to be applied by a generator 402. In at least one embodiment, a style input 406 is one or more images or image data of any type containing a style to be applied to a content input 404 by a generator 402. When multiple style inputs 406 are used, in an embodiment, a generator 402 extracts a style from each style input 406 and uses an average style to be applied to a content input 404.” [0093] In at least one embodiment, a style encoder E.sub.s 410 is data values and a set of software instructions that, when executed, implement one or more neural networks or models that take one or more style inputs x.sub.s 406 and output a style embedding z.sub.s. In at least one embodiment, a style embedding z.sub.s is a vector or set of numbers containing continuous values that represent style information about a style input x.sub.s 406. In at least one embodiment, a style embedding z.sub.s contains other information about a style input x.sub.s 406 such as pose of objects contained in said style input x.sub.s 406. In at least one embodiment, a style encoder E.sub.s 410 that takes two or more style images as inputs 406 outputs an individual style embedding z.sub.s.sup.i for each input i, and a generator 402 then averages each style embedding to determine z.sub.s that will be applied to a content input 404.”) In addition, the same motivation is used as the rejection for claim 1. Regarding claim 7, Karpman and Liu teach the method of claim 1, wherein: the style embedding is in a multimodal embedding space (see at least col.11, lines 61-67-col.12, lines 1-7 of Karpman “At Block M130, the method generates a final training corpus by executing a multimodal encoder-decoder 126 on the initial training corpus to process text captions, such as to (i) generate text captions describing each image in the set of training images and (ii) identify and remove misaligned text captions associated with images in initial training corpus. As described above, the text-to-image diffusion model 112 can include and/or interface with a multimodal encoder-decoder 126 that is configured to both generate a visual feature embedding of input images in a text-image embedding space (e.g., a high dimensional abstract vector space) and decode visual feature embeddings into a natural language (e.g., text) description of corresponding images”; [0062] of Liu “In at least one embodiment, training 110 translates one or more style images 104, 106, 108 into latent space 114 representing said one or more style images 104, 106, 108. In at least one embodiment, latent space 114 representing one or more style images 104, 106, 108 is a set of numerical values in which similar data points are closer together in space, such as data points representing style information from said one or more style images 104, 106, 108. In at least one embodiment, latent space 114 representing one or more style images 104, 106, 108 is a vector, matrix, or any other data type capable of representing any of said one or more style images 104, 106, 108. In at least one embodiment, latent space 114 representing one or more style images 104, 106, 108 contains numerical representations or data corresponding to features specific to each of said one or more style images 104, 106, 108 that is to be applied to features of a content image 102. In at least one embodiment, training 110 translates one or more style images 104, 106, 108 to latent space 114 using one or more neural networks or models, as described below in conjunction with FIGS. 4 and 5. [0093] In at least one embodiment, a style encoder E.sub.s 410 is data values and a set of software instructions that, when executed, implement one or more neural networks or models that take one or more style inputs x.sub.s 406 and output a style embedding z.sub.s. In at least one embodiment, a style embedding z.sub.s is a vector or set of numbers containing continuous values that represent style information about a style input x.sub.s 406. In at least one embodiment, a style embedding z.sub.s contains other information about a style input x.sub.s 406 such as pose of objects contained in said style input x.sub.s 406. In at least one embodiment, a style encoder E.sub.s 410 that takes two or more style images as inputs 406 outputs an individual style embedding z.sub.s.sup.i for each input i, and a generator 402 then averages each style embedding to determine z.sub.s that will be applied to a content input 404.”) In addition, the same motivation is used as the rejection for claim 1. Regarding claim 8, Karpman and Liu teach the method of claim 1, wherein: the style embedding is an image embedding comprising semantic information of the style input (see at least col.4, lines 17-26 of Karpman “Generally, each pre-trained text encoder 118 defines a model, such as a (large) language model, that is configured to generate embeddings (e.g., a set of vectors) that encode semantic concepts represented by text from a high dimensional vector space into a latent space . The set of pre-trained text encoders 118 can include text-only language models (e.g., T5-XXL, BERT) and multi-modal language models (e.g., CLIP). The different models may add different understanding of the text prompt, which improves the embeddings.”; (see at least col.4,lines 46-57 “During pre-training and/or training stages, the system can execute the multi-modal encoder-decoder 126 on text prompts included in the training corpus to generate embeddings used by the text-to-image diffusion model 112 to generate images (e.g., in addition to embeddings generated by the pre-trained text encoder 118, instead of embeddings generated by the pre-trained text encoder 118 or via a dropout method). In some implementations, a multimodal encoder-decoder 126 includes and/or interfaces with a vision transformer configured to divide an input image into segments and generate a corresponding sequence of embeddings for the input image” [0093] of Liu “In at least one embodiment, a style encoder E.sub.s 410 is data values and a set of software instructions that, when executed, implement one or more neural networks or model s that take one or more style inputs x.sub.s 406 and output a style embedding z.sub.s. In at least one embodiment, a style embedding z.sub.s is a vector or set of numbers containing continuous values that represent style information about a style input x.sub.s 406. In at least one embodiment, a style embedding z.sub.s contains other information about a style input x.sub.s 406 such as pose of objects contained in said style input x.sub.s 406. In at least one embodiment, a style encoder E.sub.s 410 that takes two or more style images as inputs 406 outputs an individual style embedding z.sub.s.sup.i for each input i, and a generator 402 then averages each style embedding to determine z.sub.s that will be applied to a content input 404.”) In addition, the same motivation is used as the rejection for claim 8. Regarding claim 9, Karpman and Liu teach t he method of claim 1, wherein obtaining the text prompt and the style input comprises: extracting the text prompt and the style input from a user input (see at least col.20, lines 19-22 as shown in Fig. 4A of Karpman “Referring back to FIG. 4A, the interactive text field 402 prompts and enables the user to input text that describes an image they wish to generate via the text-to-image diffusion model 112”; col.20, lines “; col.20, lines 56-67 “The generation interface 400 also includes a style menu 404 that enables the user to browse and select among pre-set image styles for the image generation request . In the example of FIG. 4A, the style menu 404 displays a set (e.g., array) of style option tiles, each style option tile including a text description of the image style (e.g., anime, Van Gogh, oil painting, line drawing, digital art, etc.) and a sample image in the corresponding style. In response to detecting a horizontal swipe (or scroll) input on a display or trackpad over the style menu area, the software application layer 124 can transition the display of the style menu to replace currently displayed style option tiles with other style option tiles”; [0097] of Liu “In at least one embodiment, a content input x.sub.c 420 is an image or image data of any type containing one or more objects to which a style from a style input x.sub.s 422 is to be applied by a generator G 418. In at least one embodiment, a style input x.sub.s 422 is one or more images or image data of any type containing a style to be applied to a content input x.sub.c 420 by a generator G 418. When multiple style inputs 422 are used, in an embodiment, a generator G 418 extracts a style from each style input x.sub.s 422 and uses an average style to be applied to a content input x.sub.c 420.) In addition, the same motivation is used as the rejection for claim 1. Regarding claim 10, Karpman and Liu teach t he method of claim 1, wherein obtaining the style input comprises: wherein obtaining the style input comprises: providing a plurality of predetermined image styles; and receiving a user input selecting at least one of the plurality of predetermined image styles (see at least col.20, lines 56-67-col.21,lines 1-27 “The generation interface 400 also includes a style menu 404 that enables the user to browse and select among pre-set image styles for the image generation request. In the example of FIG. 4A, the style menu 404 displays a set (e.g., array) of style option tiles, each style option tile including a text description of the image style (e.g., anime, Van Gogh, oil painting, line drawing, digital art, etc.) and a sample image in the corresponding style. In response to detecting a horizontal swipe (or scroll) input on a display or trackpad over the style menu area, the software application layer 124 can transition the display of the style menu to replace currently displayed style option tiles with other style option tiles. Additionally or alternatively, in response to detecting a user input (e.g., a touch input, a click) on a “see all” affordance (e.g., button, link) 410 located adjacent to the style menu, the software application layer can expand the style menu in an overlay tab. FIG. 4C depicts an example of the overlay tab 412 for the style menu according to some embodiments. Overlay tab 412 can display a larger number of style option tiles 1 to 12 (and/or display style option tiles at a larger size or resolution) to enable the user to browse the set of supported pre-set image styles quickly and efficiently (e.g., via vertical swipe or scroll inputs). In response to detecting a user input (e.g., a touch input, a click) on a style option tile, the software application layer 124 can visually emphasize display of the selected style option tile—such as by shading, highlighting, or displaying a box around the selected style option tile—in order to provide visual feedback to the user. Thus, the software application layer 124 and inspiration affordance can assist the user in deciding on a type of image they wish to generate by providing visual examples of (possible) image styles and enabling users to select among pre-set styles without requiring the user to develop and/or phrase image generation prompts that include stylistic constraints, thereby reducing the cognitive burden on the user in operating the software application layer and increasing the likelihood that the image subsequently generated by the text-to-image diffusion model 112 will align with user expectations”; [0060] of Liu “ In at least one embodiment, training 110 is a process that trains one or more neural networks or models to translate a content image 102 based on one or more styles 120, 122, 124. In at least one embodiment, training 110 is accomplished by execution of a set of software instructions that, when executed, implement a training framework, as described below in conjunction with FIG. 2. In at least one embodiment, training 110 takes, as input, a content image 102. In at least one embodiment, a content image 102 is any type of image file. In at least one embodiment, a content image 102 contains one or more objects to which a style 120, 122, 124 is to be applied. In at least one embodiment, training 110 takes, as input, one or more style images 104, 106, 108. In at least one embodiment, one or more style images 104, 106, 108 are image data of any file type. In at least one embodiment, one or more style images 104, 106, 108 comprise one or more objects that contain a style 120, 122, 124 to be learned by one or more neural networks or models.”) In addition, the same motivation is used as the rejection for claim 1. Regarding claim 11, Karpman and Liu teach the method of claim 10, wherein obtaining the style input comprises: displaying a plurality of preview images corresponding to the plurality of predetermined image styles, respectively (see at least Karpman: col.20, lines 56-67-col.21,lines 1-27 “The generation interface 400 also includes a style menu 404 that enables the user to browse and select among pre-set image styles for the image generation reques t. In the example of FIG. 4A, the style menu 404 displays a set (e.g., array) of style option tiles, each style option tile including a text description of the image style (e.g., anime, Van Gogh, oil painting, line drawing, digital art, etc.) and a sample image in the corresponding style . In response to detecting a horizontal swipe (or scroll) input on a display or trackpad over the style menu area, the software application layer 124 can transition the display of the style menu to replace currently displayed style option tiles with other style option tiles. Additionally or alternatively, in response to detecting a user input (e.g., a touch input, a click) on a “see all” affordance (e.g., button, link) 410 located adjacent to the style menu, the software application layer can expand the style menu in an overlay tab. FIG. 4C depicts an example of the overlay tab 412 for the style menu according to some embodiments. Overlay tab 412 can display a larger number of style option tiles 1 to 12 (and/or display style option tiles at a larger size or resolution) to enable the user to browse the set of supported pre-set image styles quickly and efficiently (e.g., via vertical swipe or scroll inputs). In response to detecting a user input (e.g., a touch input, a click) on a style option tile, the software application layer 124 can visually emphasize display of the selected style option tile—such as by shading, highlighting, or displaying a box around the selected style option tile—in order to provide visual feedback to the user. Thus, the software application layer 124 and inspiration affordance can assist the user in deciding on a type of image they wish to generate by providing visual examples of (possible) image styles and enabling users to select among pre-set styles without requiring the user to develop and/or phrase image generation prompts that include stylistic constraints, thereby reducing the cognitive burden on the user in operating the software application layer and increasing the likelihood that the image subsequently generated by the text-to-image diffusion model 112 will align with user expectations”; col.23, lines 24-38 “Concurrently (e.g., after the base diffusion model has generated the base image and while the executing the high-resolution diffusion model(s), after the communication interface 122 has output the generated image(s)), the software application layer 124 can transition display of the waiting screen to the results interface. FIG. 4D depicts an example of the results interface 414 according to some embodiments. Results interface 414 displays (a preview of) the image(s) 1 to 6 generated by the text-to-image diffusio n model 112 at 416 in response to the image generation request and a description field 418 that displays the text prompt submitted to the model. If the image generation request included a style selected at the generation interface, the description tile can also display a style tag 420 identifying the style of “Mystical” for the generated image(s) 1 to 6.”) Regarding claim 12, Karpman and Liu teach the method of claim 1, wherein obtaining the style input comprises: generating the style input based on a style image (see at least col.20, lines 56-67-col.21,lines 1-27 “The generation interface 400 also includes a style menu 404 that enables the user to browse and select among pre-set image styles for the image generation request. In the example of FIG. 4A, the style menu 404 displays a set (e.g., array) of style option tiles, each style option tile including a text description of the image style (e.g., anime, Van Gogh, oil painting, line drawing, digital art, etc.) and a sample image in the corresponding style. In response to detecting a horizontal swipe (or scroll) input on a display or trackpad over the style menu area, the software application layer 124 can transition the display of the style menu to replace currently displayed style option tiles with other style option tiles. Additionally or alternatively, in response to detecting a user input (e.g., a touch input, a click) on a “see all” affordance (e.g., button, link) 410 located adjacent to the style menu, the software application layer can expand the style menu in an overlay tab. FIG. 4C depicts an example of the overlay tab 412 for the style menu according to some embodiments. Overlay tab 412 can display a larger number of style option tiles 1 to 12 (and/or display style option tiles at a larger size or resolution) to enable the user to browse the set of supported pre-set image styles quickly and efficiently (e.g., via vertical swipe or scroll inputs). In response to detecting a user input (e.g., a touch input, a click) on a style option tile, the software application layer 124 can visually emphasize display of the selected style option tile—such as by shading, highlighting, or displaying a box around the selected style option tile—in order to provide visual feedback to the user. Thus, the software application layer 124 and inspiration affordance can assist the user in deciding on a type of image they wish to generate by providing visual examples of (possible) image styles and enabling users to select among pre-set styles without requiring the user to develop and/or phrase image generation prompts that include stylistic constraints, thereby reducing the cognitive burden on the user in operating the software application layer and increasing the likelihood that the image subsequently generated by the text-to-image diffusion model 112 will align with user expectations”; [0060] of Liu “ In at least one embodiment, training 110 is a process that trains one or more neural networks or models to translate a content image 102 based on one or more styles 120, 122, 124. In at least one embodiment, training 110 is accomplished by execution of a set of software instructions that, when executed, implement a training framework, as described below in conjunction with FIG. 2. In at least one embodiment, training 110 takes, as input, a content image 102. In at least one embodiment, a content image 102 is any type of image file. In at least one embodiment, a content image 102 contains one or more objects to which a style 120, 122, 124 is to be applied. I n at least one embodiment, training 110 takes, as input, one or more style images 104, 106, 108. In at least one embodiment, one or more style images 104, 106, 108 are image data of any file type . In at least one embodiment, one or more style images 104, 106, 108 comprise one or more objects that contain a style 120, 122, 124 to be learned by one or more neural networks or models.”) In addition, the same motivation is used as the rejection for claim 1. Regarding independent claim 13, Karpman teaches a non-transitory computer readable medium storing code for media processing, the code comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations (col.24, lines 8-33 as shown in Fig. 5 “The processor 501 may perform operations such as those described herein. Instructions for performing such operations may be embodied in the memory 503, on one or more non-transitory computer readable media, or on some other storage device. Various specially configured devices can also be used in place of or in addition to the processor 501. Memory 503 may be random access memory (RAM) or other dynamic storage devices. Storage device 505 may include a non-transitory computer-readable storage medium holding information, instructions, or some combination thereof, for example instructions that when executed by the processor 501, cause processor 501 to be configured or operable to perform one or more operations of a method as described herein .” ) comprising: obtaining an input and a style input, wherein the input depicts image content and the style input describes an image style (col.20, lines 19-22 as shown in Fig. 4A “Referring back to FIG. 4A, the interactive text field 402 prompts and enables the user to input text that describes an image they wish to generate via the text-to-image diffusion model 112”; col.20, lines “; col.20, lines 56-67 “The generation interface 400 also includes a style menu 404 that enables the user to browse and select among pre-set image styles for the image generation request. In the example of FIG. 4A, the style menu 404 displays a set (e.g., array) of style option tiles, each style option tile including a text description of the image style (e.g., anime, Van Gogh, oil painting, line drawing, digital art, etc.) and a sample image in the corresponding style. In response to detecting a horizontal swipe (or scroll) input on a display or trackpad over the style menu area, the software application layer 124 can transition the display of the style menu to replace currently displayed style option tiles with other style option tiles”) generating, using an image generation model, a first intermediate output based on the content input during a first stage of a diffusion process (see at least col. 4, lines 45-67-col.5, lines 1-5 “During pre-training and/or training stages, the system can execute the multi-modal encoder-decoder 126 on text prompts included in the training corpus to generate embeddings used by the text-to-image diffusion model 112 to generate images (e.g., in addition to embeddings generated by the pre-trained text encoder 118 , instead of embeddings generated by the pre-trained text encoder 118 or via a dropout method). In some implementations, a multimodal encoder-decoder 126 includes and/or interfaces with a vision transformer configured to divide an input image into segments and generate a corresponding sequence of embeddings for the input image. The multimodal encoder-decoder 126 can operate in different ways, such as: (1) as a unimodal encoder that generates image embeddings given an image input and generates text embeddings given a text input (e.g., generates separate image and text embeddings); (2) as an image-aware text encoder that modifies embeddings generated by the pre-trained text encoder and/or unimodal encoder module to include visual information based on visual features of an input image (e.g., via cross-attention to the input image); and (3) as an image-aware text decoder that generates a text (e.g., natural language) description of an input image based on image-text embedding representations. Generally, the image-aware text encoder and the image aware text decoder can share a common architecture and/or parameters (e.g., with the exception of the cross-attention and self-attention layers) in order to improve training efficiency”; col.5, lines 15-36 “In some implementations, the base image diffusion model 120 defines a deep learning network (e.g., a convolutional neural network, a residual neural network, etc.) configured (e.g., through the training described) to generate images from random (e.g., Gaussian) noise based on text prompts and/or descriptions. The base image diffusion model 120 can include a U-net architecture (e.g., Efficient U-Net) defined from residual and multi-head attention blocks that enable the base image diffusion model 120 to progressively denoise (e.g., infill, generate, augment) image data according to cross-attention inputs based on the text prompt. The base image diffusion model 120 can therefore: r eceive one or more text embeddings from the set of pre-trained text encoders 118 ; receive and/or initialize a (randomly sampled) noise distribution at a preset resolution (e.g., 64 pixels by 64 pixels); and transform the noise distribution into a base image at the preset resolution based on the one or more text embeddings and parameters, weights, and/or paths corresponding to an iterative denoising process learned by the base image diffusion model 120 during training. The system can then pass the base image to the set of high-resolution diffusion models 116 for upsampling and output .”) generating, using the image generation model, a second intermediate output based on the first intermediate output and the style input during a second stage of the diffusion process ((see at least col.5, lines 37-62 “In some implementations, each high-resolution diffusion model 116 in the set of high-resolution diffusion models 116 defines a deep learning network configured to receive a low-resolution base image (e.g., 64 pixels by 64 pixels, 256 pixels by 256 pixels) and generate a higher-resolution version (e.g., copy) of the base image (e.g., 256 pixels by 256 pixels, 1024 pixels by 1024 pixels). Generally, a high-resolution version shares a similar architecture with the base image diffusion (e.g., U-net, efficient U-net). However, self-attention layers in the base diffusion model architecture can be omitted to improve memory efficiency and inference time. During training, the set of high-resolution diffusion models 116 can be conditioned on text information (e.g., text descriptions of training images), noise augmentations, and/or visual information (e.g., embeddings of low-resolution images generated by the multimodal encoder-decoder). T hus, during operation, each high-resolution diffusion model 116 can implement an iterative denoising process similar to the base image diffusion model 120 in order to progressively upsample generated base images to higher resolution, infill, infer, and/or generate additional visual detail and/or texture, and remove visual artifacts generated by the base image diffusion model 120 . As discussed above, high-resolution diffusion models 116 operate in the pixel space to upsample the base images output by base image diffusion models 120 .” Where visual information (e.g., embeddings of low-resolution images generated by the multimodal encoder-decoder is considered as style embedding ) Karpman is understood to be silent on the remaining limitations of claim 13. In the same field of endeavor, Liu teaches obtaining an image input and a style input, wherein the image input depicts image content and the style input describes an image style ([0091]as show in Fig.4A “ In at least one embodiment, a content input 404 is an image or image data of any type containing one or more objects to which a style from a style input 406 is to be applied by a generator 40 2. In at least one embodiment, a style input 406 is one or more images or image data of any type containing a style to be applied to a content input 404 by a generator 402. When multiple style inputs 406 are used, in an embodiment, a generator 402 extracts a style from each style input 406 and uses an average style to be applied to a content input 404.”); generating, using the image generation model, a synthetic image wherein the style embedding is provided step of generating the synthetic image (0081] In at least one embodiment, one input data set 304 is equivalent or similar to a baseline data set 302 . In at least one embodiment, an input data set 304 is used to train a generator 308 . In at least one embodiment, a generator 308 in a GAN provides as output a probabilistic distribution 310 . In at least one embodiment, a generator 308 in a GAN operating on image content and styles outputs a generated image instead of or in addition to probabilistic values 310 . In at least one embodiment, output from a generator 308 is provided as input to a discriminator 312 for training purposes. In at least one embodiment, a discriminator 312 provides loss information 316 to a generator 308 in a GAN in order to update weights through backpropagation in a generator 308 .”; [0094] In at least one embodiment, an image decoder F 412 generates an output x 416 using a content embedding z.sub.c and information from a style embedding z.sub.s, as described above. In at least one embodiment, an image decoder F 412 is data values and a set of software instructions that, when executed, implement one or more neural networks or models that construct an output 416 based on a content embedding z.sub.c and adaptive instance normalization (AdaIN) parameters 414. In at least one embodiment, an image decoder F 412 utilizes vanilla convolutional layers. In at least one embodiment, AdaIN parameters 414 are numerical data values generated based on a style embedding z.sub.s output from a style encoder E.sub.s 410. In at least one embodiment, mean and scale parameters of AdaIN parameters 414 are computed or generated based on a style embedding z.sub.s output from a style encoder E.sub.s 414 . ) In addition, the same motivation is used as the rejection for claim 1. Thus, the combination of Karpman and Liu teaches a non-transitory computer readable medium storing code for media processing, the code comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising: obtaining an image input and a style input, wherein the image input depicts image content and the style input describes an image style; generating, using an image generation model, a first intermediate output based on the content input during a first stage of a diffusion process; generating, using the image generation model, a second intermediate output based on the first intermediate output and the style input during a second stage of the diffusion process; and generating, using the image generation model, a synthetic image based on the second intermediate output, wherein the style embedding is provided at a second step of generating the synthetic image after a first step. Regarding claim 14, Karpman and Liu teach t he non-transitory computer readable medium of claim 13, the code further comprising instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising: generating, using a text encoder, a text embedding based on the text prompt, wherein the text embedding represents the image content and wherein the first intermediate output is based on the text embedding (see at least col.20, lines 19-22 as shown in Fig. 4A of Karpman “Referring back to FIG. 4A, the interactive text field 402 prompts and enables the user to input text that describes an image they wish to generate via the text-to-image diffusion model 112”; col.20, lines “; col.20, lines 56-67 “The generation interface 400 also includes a style menu 404 that enables the user to browse and select among pre-set image styles for the image generation request . In the example of FIG. 4A, the style menu 404 displays a set (e.g., array) of style option tiles, each style option tile including a text description of the image style (e.g., anime, Van Gogh, oil painting, line drawing, digital art, etc.) and a sample image in the corresponding style. In response to detecting a horizontal swipe (or scroll) input on a display or trackpad over the style menu area, the software application layer 124 can transition the display of the style menu to replace currently displayed style option tiles with other style option tiles”; col.5, lines 15-36 “In some implementations, the base image diffusion model 120 defines a deep learning network (e.g., a convolutional neural network, a residual neural network, etc.) configured (e.g., through the training described) to generate images from random (e.g., Gaussian) noise based on text prompts and/or descriptions. The base image diffusion model 120 can include a U-net architecture (e.g., Efficient U-Net) defined from residual and multi-head attention blocks that enable the base image diffusion model 120 to progressively denoise (e.g., infill, generate, augment) image data according to cross-attention inputs based on the text prompt. The base image diffusion model 120 can therefore: r eceive one or more text embeddings from the set of pre-trained text encoders 118 ; receive and/or initialize a (randomly sampled) noise distribution at a preset resolution (e.g., 64 pixels by 64 pixels); and transform the noise distribution into a base image at the preset resolution based on the one or more text embeddings and parameters, weights, and/or paths corresponding to an iterative denoising process learned by the base image diffusion model 120 during training. The system can then pass the base image to the set of high-resolution diffusion models 116 for upsampling and output .”) ; and generating, using a style encoder, a style embedding based on the style input, wherein the style embedding represents the image style (see at least col.4,lines 46-57 of Karpman “During pre-training and/or training stages, the system can execute the multi-modal encoder-decoder 126 on text prompts included in the training corpus to generate embeddings used by the text-to-image diffusion model 112 to generate images (e.g., in addition to embeddings generated by the pre-trained text encoder 118, instead of embeddings generated by the pre-trained text encoder 118 or via a dropout method). In some implementations, a multimodal encoder-decoder 126 includes and/or interfaces with a vision transformer configured to divide an input image into segments and generate a corresponding sequence of embeddings for the input image” ;[0093] of Liu “ In at least one embodiment, a style encoder E.sub.s 410 is data values and a set of software instructions that, when executed, implement one or more neural networks or models that take one or more style inputs x.sub.s 406 and output a style embedding z.sub.s. In at least one embodiment, a style embedding z.sub.s is a vector or set of numbers containing continuous values that represent style information about a style input x.sub.s 406. In at least one embodiment, a style embedding z.sub.s contains other information about a style input x.sub.s 406 such as pose of objects contained in said style input x.sub.s 406. In at least one embodiment, a style encoder E.sub.s 410 that takes two or more style images as inputs 406 outputs an individual style embedding z.sub.s.sup.i for each input i, and a generator 402 then averages each style embedding to determine z.sub.s that will be applied to a content input 404.”) and the second intermediate output is based on the style embedding ( (see at least col.5, lines 37-62 of Karpman “In some implementations, each high-resolution diffusion model 116 in the set of high-resolution diffusion models 116 defines a deep learning network configured to receive a low-resolution base image (e.g., 64 pixels by 64 pixels, 256 pixels by 256 pixels) and generate a higher-resolution version (e.g., copy) of the base image (e.g., 256 pixels by 256 pixels, 1024 pixels by 1024 pixels). Generally, a high-resolution version shares a similar architecture with the base image diffusion (e.g., U-net, efficient U-net). However, self-attention layers in the base diffusion model architecture can be omitted to improve memory efficiency and inference time. During training, the set of high-resolution diffusion models 116 can be conditioned on text information (e.g., text descriptions of training images), noise augmentations, and/or visual information (e.g., embeddings of low-resolution images generated by the multimodal encoder-decoder). T hus, during operation, each high-resolution diffusion model 116 can implement an iterative denoising process similar to the base image diffusion model 120 in order to progressively upsample generated base images to higher resolution, infill, infer, and/or generate additional visual detail and/or texture, and remove visual artifacts generated by the base image diffusion model 120 . As discussed above, high-resolution diffusion models 116 operate in the pixel space to upsample the base images output by base image diffusion models 120 . ” Where visual information (e.g., embeddings of low-resolution images generated by the multimodal encoder-decoder is considered as style embedding ) In addition, the same motivation is used as the rejection for claim 1. Regarding independent claim 15, Karpman teaches a system for image generation, comprising: a memory component; and a processing device coupled to the memory component, the processing device configured to perform operations ( see at least col.24, lines 8-47 of Karpman “FIG. 5 illustrates one example of a computing device according to some embodiments. According to various embodiments, a system 500 suitable for implementing embodiments described herein includes a processor 501, a memory module 503, a storage device 505, an interface 511, and a bus 515 (e.g., a PCI bus or other interconnection fabric.) comprising: Remaining limitations of claim 15 is similar in scope to claim 1 and therefore rejected under the same rationale. Regarding claim 16, Karpman and Liu teach t he system of claim 15, the system further comprising: an image encoder trained to generate an image embedding based on an image (see at least Liu [0060] In at least one embodiment, training 110 is a process that trains one or more neural networks or models to translate a content image 102 based on one or more styles 120, 122, 124. In at least one embodiment, training 110 is accomplished by execution of a set of software instructions that, when executed, implement a training framework, as described below in conjunction with FIG. 2. In at least one embodiment, training 110 takes, as input, a content image 102. In at least one embodiment, a content image 102 is any type of image file. In at least one embodiment, a content image 102 contains one or more objects to which a style 120, 122, 124 is to be applied. In at least one embodiment, training 110 takes, as input, one or more style images 104, 106, 108. In at least one embodiment, one or more style images 104, 106, 108 are image data of any file type. In at least one embodiment, one or more style images 104, 106, 108 comprise one or more objects that contain a style 120, 122, 124 to be learned by one or more neural networks or models. [0092] In at least one embodiment, a generator 402 comprises a content encoder E.sub.c 408 , a style encoder E.sub.s 410, and an image decoder F 412. In at least one embodiment, a content encoder E.sub.c 408 is data values and a set of software instructions that, when executed, implement one or more neural networks or models that take a content input x.sub.c 404 and output a content embedding z.sub.c. In at least one embodiment, a content encoder E.sub.c 408 utilizes vanilla convolutional layers. In at least one embodiment, a content embedding z.sub.c is a vector or set of values containing continuous numbers that represent information about a content input x.sub.c 404, such as features or objects to which a style is to be applied .” Where a content encoder is considered as image encoder) In addition, the same motivation is used as the rejection for claim 1. Regarding claim 17, Karpman and Liu teach t he system of claim 15, the system further comprising: a user interface configured to display the synthetic image to a user (see at least col.23, lines 24-38 of Karpman “Concurrently (e.g., after the base diffusion model has generated the base image and while the executing the high-resolution diffusion model(s), after the communication interface 122 has output the generated image(s)), the software application layer 124 can transition display of the waiting screen to the results interface . FIG. 4D depicts an example of the results interface 414 according to some embodiments. Results interface 414 displays (a preview of) the image(s) 1 to 6 generated by the text-to-image diffusio n model 112 at 416 in response to the image generation request and a description field 418 that displays the text prompt submitted to the model. If the image generation request included a style selected at the generation interface, the description tile can also display a style tag 420 identifying the style of “Mystical” for the generated image(s) 1 to 6.”) Regarding claim 18, Karpman and Liu teach the system of claim 15, wherein: the image generation model comprises a diffusion model (see at least col.2, lines 34-42 as shown in Fig.1 (item 112) of Karpman “As described below, the web intelligence engine 108 may include a set of web crawler software modules configured to automatically review content, including locating, fetching, ingesting and/or downloading image content and associated alternate text hosted on web pages or other content. The set of storage devices no may store information for content, such as a repository, database, and/or index of content (e.g., images, text, and/or text-image pairs), that is used to train the text-to-image diffusion model 112 .”) 07-21-aia AIA 2. Claim s 3-5 are rejected under 35 U.S.C. 103 as being unpatentable over Karpman et al., U.S Patent No.11,995,803 (“Karpman”) in view of Liu et al, U.S Patent Application Publication No.20210358164 (“Liu”) further in view of Pan et al, U.S Patent Application Publication No.20250014233 (“Pan”) further in view of RAMESH et al, U.S Patent Application Publication No.2024/0331237 (“REMASH”) Regarding claim 3, Karpman and Liu teach the method of claim 1, wherein generating the style embedding comprises: encoding the style input using a multimodal text encoder to obtain a style text embedding; and converting the text embedding to the style embedding using an embedding conversion model (see at least col.4, lines 46-67-col.5, lines 1-6 of Karpman “During pre-training and/or training stages, the system can execute the multi-modal encoder-decoder 126 on text prompts included in the training corpus to generate embeddings used by the text-to-image diffusion model 112 to generate images ( e.g., in addition to embeddings generated by the pre-trained text encoder 118, instead of embeddings generated by the pre-trained text encoder 118 or via a dropout method). In some implementations, a multimodal encoder-decoder 126 includes and/or interfaces with a vision transformer configured to divide an input image into segments and generate a corresponding sequence of embeddings for the input image. T he multimodal encoder-decoder 126 can operate in different ways , such as: (1) as a unimodal encoder that generates image embeddings given an image input and generates text embeddings given a text input (e.g., generates separate image and text embeddings); ( 2) as an image-aware text encoder that modifies embeddings generated by the pre-trained text encoder and/or unimodal encoder module to include visual information based on visual features of an input image (e.g., via cross-attention to the input image) ; and (3) as an image-aware text decoder that generates a text (e.g., natural language) description of an input image based on image-text embedding representations. Generally, the image-aware text encoder and the image aware text decoder can share a common architecture and/or parameters (e.g., with the exception of the cross-attention and self-attention layers) in order to improve training efficiency . ” [0092] of Liu In at least one embodiment, a generator 402 comprises a content encoder E.sub.c 408, a style encoder E.sub.s 410, and an image decoder F 412. In at least one embodiment, a content encoder E.sub.c 408 is data values and a set of software instructions that, when executed, implement one or more neural networks or models that take a content input x.sub.c 404 and output a content embedding z.sub.c. In at least one embodiment, a content encoder E.sub.c 408 utilizes vanilla convolutional layers. In at least one embodiment, a content embedding z.sub.c is a vector or set of values containing continuous numbers that represent information about a content input x.sub.c 404, such as features or objects to which a style is to be applied ” ) In addition, the same motivation is used as the rejection for claim 1. Both Karpman and Liu is understood to be silent on the remaining limitations of claim 3. In the same field of endeavor, Pan teaches wherein generating the style embedding comprises: encoding the style input using a text encoder to obtain a style text embedding, wherein the style input comprises text ([0079] At block 520 (Generate clean objects), the processor may generate one or more clean objects by performing a forward generation process of a diffusion model with a control signal. The control signal is configured to control a visual effect of the clean objects. In one example embodiments, the generating of the clean objects can include generating a final image after a completion of iteratively de-noising a starting noise by performing the forward generation process of the diffusion model. For example, the de-noising module 250 can perform the forward generation process 350 with the starting noise 320 to generate the final image 340, as shown in FIGS. 2 and 3. In one example embodiment, the diffusion model can be a pre-trained DPM (e.g., a text-to-image diffusion-based generative model) with a first input as a content conditioner and a second input as a visual effect conditioner to generate the clean objects. The visual effect conditioner may be represented as a text embedding as denoted as “#” in FIG. 5B. The text embedding can be combined (e.g., concatenated) with embeddings of other text prompts to generate the clean objects . For example, the diffusion model may receive a first text prompt (e.g., “a cute totoro in a yard”) as a content conditioner and a second text prompt as a visual effect conditioner (e.g., “bokeh”) to the pre-trained text-to-image diffusion-based generative model, and the corresponding embeddings can be concatenated to generate a clean image. The reference object and the clean objects may be obtained using the pre-trained text-to-image diffusion-based generative model with the same first input (e.g., “a cute totoro in a yard”). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of claimed invention to modify the method of generating images base on text prompt of Karpman and Liu with style input comprise text as seen in Pan because this modification would provide a visual effect conditioner via text prompt ([0079] of Pan) In the same field of endeavor, RAMESH teaches encoding the input using a multimodal text encoder to obtain a text embedding, wherein the input comprises text ([0038] FIG. 2 illustrates a flow chart of an exemplary method for generating an image corresponding to a text input, consistent with embodiments of the present disclosure. The process shown in FIG. 2 or any of its constituent steps may be implemented using operating environment 300, system 400, or any component thereof. In some embodiments, the process shown in FIG. 2 may represent the workflow of inference for generating an image corresponding to a text input. The steps illustrated in FIG. 2 are exemplary and steps may be added, merged, divided, duplicated, repeated (e.g., as part of a machine learning process), modified, performed sequentially, performed in parallel, and/or deleted in some embodiments. FIG. 2 includes steps 202-216. At step 202, a method may involve accessing a text description. At step 204, a method may involve inputting the text description into a text encoder. At step 206, a method may involve receiving, from the text encoder, a text embedding. At step 208, a method may involve inputting at least one of the text description or the text embedding into a first sub-model.”) and converting the text embedding to the style embedding using an embedding conversion model ([0036] FIG. 1B illustrates a functional diagram for generating an image corresponding to a text input, consistent with embodiments of the present disclosure. A s discussed herein, training an image generation model may involve training an image encoder 134 and a text encoder 132. Joint representation 137 may be a latent space representation of the text embedding 136 and the image embedding 135 . In some embodiments, upon completion of training, text encoder 132 and image encoder 134 may be fixed and may be used as part of machine learning models. In some embodiments, an image generation request 128 comprises a request or query (e.g., an electronic transmission of information) to input/output devices 318. For example, a user may initiate image generation request 128 by interacting with input/output devices 318. Generating an image corresponding to a text description may involve accessing a text description 130 and inputting the text description 130 to the text encoder 132. Text encoder 132 may generate a text embedding 136, which may be an input to first sub-model 138. [0041] For example, as referenced in FIG. 1B, a first sub-model 138 may be configured to generate a corresponding image embedding 140. In some embodiments, prior to generating corresponding image embedding 140, first sub-model 138 may include a prior model, such as an autoregressive prior. The prior model may be configured to encode text description 130 or text embedding 136 via a transformer. In some embodiments, first sub-model 138 may comprise an autoregressive prior. For example, the autoregressive prior may encode text description 130 or text embedding 136 via a transformer. I n some embodiments, first sub-model 138, prior to generating corresponding image embedding 140, first sub-model 138 may include a diffusion prior model, which may include a transformer. The output of first sub-model may be the corresponding image embedding 140 . For example, the diffusion prior or the autoregressive prior may generate the corresponding image embedding 140, as discussed herein. In some embodiments, corresponding image embedding 140 may be different from image embedding 135 resultant from training of image encoder 134. Corresponding image embedding 140 may be accessible to other models or to be used for image generation .”; [0044] As referenced in FIG. 2, at step 210, a method may involve generating, from at least one of the text description or the text embedding, a corresponding image embedding”) Therefore, in combination of Karpman and Liu. it would have been obvious to one of ordinary skill in the art before the effective filling date of claimed invention to modify the method of generating images base on text prompt of Karpman and style input comprise text of Pan with transforming text embedding to image embedding as seen in RAMESH because this modification would generate, based on at least one of the text description or the text embedding, a corresponding image embedding (abstract of KAMESH) Thus, the combination of Karpman, Liu, Pan and KAMESH teaches wherein generating the style embedding comprises: encoding the style input using a multimodal text encoder to obtain a style text embedding, wherein the style input comprises text; and converting the style text embedding to the style embedding using an embedding conversion model. Regarding claim 4, Karpman, Liu, Pan and KAMESH teach the method of claim 3, wherein: the style text embedding is in a multimodal embedding space (see at least col.11, lines 61-67-col.12, lines 1-7 of Karpman “At Block M130, the method generates a final training corpus by executing a multimodal encoder-decoder 126 on the initial training corpus to process text captions, such as to (i) generate text captions describing each image in the set of training images and (ii) identify and remove misaligned text captions associated with images in initial training corpus. As described above, the text-to-image diffusion model 112 can include and/or interface with a multimodal encoder-decoder 126 that is configured to both generate a visual feature embedding of input images in a text-image embedding space (e.g., a high dimensional abstract vector space) and decode visual feature embeddings into a natural language (e.g., text) description of corresponding images”; col. 13, lines 21-50 “At Block M140, during a pre-training stage, the method executes the text-to-image diffusion model 112 on the final training corpus to infer an initial set of image generation parameters. Generally, the text-to-image diffusion model 112 can be initialized with (a subset of) parameters of the multimodal encoder decoder 126 to condition the base image diffusion model 120 for vision-language tasks (e.g., to transfer vision-understanding learned by the multimodal encoder-decoder 126 during training and/or through filtering and captioning the initial training corpus). During pre-training, the system 100 can execute the set of pre-trained text encoders 118 on captions within the modified training corpus to generate one or more text embeddings (e.g., vector representations) of captions associated with each training image. More specifically, for each text caption in the modified training corpus, the system can: execute a first pre-trained text encoder in the set of pre-trained text encoders 118 (e.g., selected from one of: the unimodal encoder of the multimodal encoder-decoder, a pre-trained large language model such as T5, CLIP, BERT, etc.) on a text prompt to generate a first corresponding text embedding in a first embedding spac e; execute a second pre-trained text encoder in the set of pre-trained text encoders 118 on the text caption to generate a second corresponding text embedding in a second embedding space different from the first embedding space. The system can then store and/or queue the set of text embeddings associated with each image in the modified training corpus in order to condition the base image diffusion model 120 on these different text embeddings (e.g., representations) during pre-training.;” ([0079] of Pan; RAMESH [0036] FIG. 1B illustrates a functional diagram for generating an image corresponding to a text input, consistent with embodiments of the present disclosure. As discussed herein, training an image generation model may involve training an image encoder 134 and a text encoder 132. Joint representation 137 may be a latent space representation of the text embedding 136 and the image embedding 135 . In some embodiments, upon completion of training, text encoder 132 and image encoder 134 may be fixed and may be used as part of machine learning models. In some embodiments, an image generation request 128 comprises a request or query (e.g., an electronic transmission of information) to input/output devices 318. For example, a user may initiate image generation request 128 by interacting with input/output devices 318. Generating an image corresponding to a text description may involve accessing a text description 130 and inputting the text description 130 to the text encoder 132. Text encoder 132 may generate a text embedding 136, which may be an input to first sub-model 138. [0052] It is appreciated that embedding images and text to the same latent space contributes to enabling language-guided image manipulations. Mapping embeddings of images and text to the same latent space, together with language-guided modeling forms a non-conventional and non-generic arrangement, which contributes to the capability to manipulate an existing image based on textual information. For example, an image may be modified to reflect a new text description. Disclosed embodiments may involve accessing the text embedding corresponding to the output image. In some embodiments, the text embedding corresponding to the output image may include the text embedding received from the text encoder in the first sub-model. For example, the text embedding may be based on the input or caption which the output image is based on. In some embodiments, the text embedding may correspond to a baseline, such as a generic caption, or an empty caption. Some embodiments may involve, based on a second text description, accessing a second text embedding. A second text description may include the new text caption or description which the modified image should reflect. Accessing the second text embedding may include inputting the second description into the text encoder, and receiving the resultant text embedding. Disclosed embodiments may involve generating a vector representation of the text embedding corresponding to the output image and the second text embedding. In some embodiments, the vector representation, which may be referred to as a difference vector or text diff, may be the normalization of the difference between the second text embedding and the text embedding corresponding to the output image. Disclosed embodiments may involve performing an interpolation between the image embedding of the output image and the vector representation of the text embedding corresponding to the output image and the second text embedding. Performing an interpolation may involve rotating between the image embedding and the difference vector with spherical interpolation, such as spherical linear interpolation. Based on the interpolation, disclosed embodiments may involve generating a modified instance of the output image. Based on the interpolation may involve the interpolation producing intermediate representations, such as interpolates or trajectories of embeddings based on an angle. Generating the modified instance of the output image may inversion to reconstruct the image, or applying the interpolates to a decoder or diffusion decoder model, as discussed herein.”) In addition, the same motivation is used as the rejection for claim 3. Regarding claim 5, Karpman, Liu, Pan and KAMESH teach the method of claim 3, wherein: the style text embedding is based on the text prompt ([0079] At block 520 (Generate clean objects), the processor may generate one or more clean objects by performing a forward generation process of a diffusion model with a control signal. The control signal is configured to control a visual effect of the clean objects. In one example embodiments, the generating of the clean objects can include generating a final image after a completion of iteratively de-noising a starting noise by performing the forward generation process of the diffusion model. For example, the de-noising module 250 can perform the forward generation process 350 with the starting noise 320 to generate the final image 340, as shown in FIGS. 2 and 3. In one example embodiment, the diffusion model can be a pre-trained DPM (e.g., a text-to-image diffusion-based generative model) with a first input as a content conditioner and a second input as a visual effect conditioner to generate the clean objects. The visual effect conditioner may be represented as a text embedding as denoted as “#” in FIG. 5B. The text embedding can be combined (e.g., concatenated) with embeddings of other text prompts to generate the clean objects . For example, the diffusion model may receive a first text prompt (e.g., “a cute totoro in a yard”) as a content conditioner and a second text prompt as a visual effect conditioner (e.g., “bokeh”) to the pre-trained text-to-image diffusion-based generative model, and the corresponding embeddings can be concatenated to generate a clean image. The reference object and the clean objects may be obtained using the pre-trained text-to-image diffusion-based generative model with the same first input (e.g., “a cute totoro in a yard”).;[0053] of KAMESH “ For example, a model may generate an output image based on a caption of “a photo of an antique car,” through embodiments of the present disclosure. The model may generate the text embedding for the caption, which corresponds to the output image. It may be desired to modify the output image of the antique car, such that the modified image reflects the second text description of “a modern car”. The second text description may be inputted into the text encoder, resulting in a second text embedding. The difference vector of the text embedding and the second text embedding may be generated, and the interpolation can then be performed. The generated modified instance of the output image may then be an image of a car with modern features. As such, embodiments of the present disclosure may allow visualization of changes occurring as the image is modified, such as generating visualizations of before and after modified instances of an image. It is appreciated that this capability of generating modified instances of output images improves image generation machine learning model output by enabling more user guidance and allowing a user to provide language to direct the model to generate a desired image output.”) In addition, the same motivation is used as the rejection for claim 3 . 07-21-aia AIA 3. Claim 19 is rejected under 35 U.S.C. 103 as being unpatentable over Karpman et al., U.S Patent No.11,995,803 (“Karpman”) in view of Liu et al, U.S Patent Application Publication No.20210358164 (“Liu”) further in view of RAMESH et al, U.S Patent Application Publication No.2024/0331237 (“REMASH”) Regarding claim 19, Karpman and Liu teach the system of claim 15, wherein: the style encoder further comprises an embedding conversion model configured to convert the style text embedding to the style embedding, wherein the embedding conversion model comprises an autoregressive model or a diffusion model (see at least col.4, lines 46-67-col.5, lines 1-6 of Karpman “During pre-training and/or training stages, the system can execute the multi-modal encoder-decoder 126 on text prompts included in the training corpus to generate embeddings used by the text-to-image diffusion model 112 to generate images ( e.g., in addition to embeddings generated by the pre-trained text encoder 118, instead of embeddings generated by the pre-trained text encoder 118 or via a dropout method). In some implementations, a multimodal encoder-decoder 126 includes and/or interfaces with a vision transformer configured to divide an input image into segments and generate a corresponding sequence of embeddings for the input image. T he multimodal encoder-decoder 126 can operate in different ways , such as: (1) as a unimodal encoder that generates image embeddings given an image input and generates text embeddings given a text input (e.g., generates separate image and text embeddings); ( 2) as an image-aware text encoder that modifies embeddings generated by the pre-trained text encoder and/or unimodal encoder module to include visual information based on visual features of an input image (e.g., via cross-attention to the input image) ; and (3) as an image-aware text decoder that generates a text (e.g., natural language) description of an input image based on image-text embedding representations. Generally, the image-aware text encoder and the image aware text decoder can share a common architecture and/or parameters (e.g., with the exception of the cross-attention and self-attention layers) in order to improve training efficiency . ” [0092] of Liu In at least one embodiment, a generator 402 comprises a content encoder E.sub.c 408, a style encoder E.sub.s 410, and an image decoder F 412. In at least one embodiment, a content encoder E.sub.c 408 is data values and a set of software instructions that, when executed, implement one or more neural networks or models that take a content input x.sub.c 404 and output a content embedding z.sub.c. In at least one embodiment, a content encoder E.sub.c 408 utilizes vanilla convolutional layers. In at least one embodiment, a content embedding z.sub.c is a vector or set of values containing continuous numbers that represent information about a content input x.sub.c 404, such as features or objects to which a style is to be applied ” ) In addition, the same motivation is used as the rejection for claim 1. Both Karpman and Liu is understood to be silent on the remaining limitations of claim 19. In the same field of endeavor, KAMESH teaches the style encoder further comprises an embedding conversion model configured to convert the style text embedding to the style embedding (see at least [0036] FIG. 1B illustrates a functional diagram for generating an image corresponding to a text input, consistent with embodiments of the present disclosure. A s discussed herein, training an image generation model may involve training an image encoder 134 and a text encoder 132. Joint representation 137 may be a latent space representation of the text embedding 136 and the image embedding 135 . In some embodiments, upon completion of training, text encoder 132 and image encoder 134 may be fixed and may be used as part of machine learning models. In some embodiments, an image generation request 128 comprises a request or query (e.g., an electronic transmission of information) to input/output devices 318. For example, a user may initiate image generation request 128 by interacting with input/output devices 318. Generating an image corresponding to a text description may involve accessing a text description 130 and inputting the text description 130 to the text encoder 132. Text encoder 132 may generate a text embedding 136, which may be an input to first sub-model 138. [0041] For example, as referenced in FIG. 1B, a first sub-model 138 may be configured to generate a corresponding image embedding 140. In some embodiments, prior to generating corresponding image embedding 140, first sub-model 138 may include a prior model, such as an autoregressive prior. The prior model may be configured to encode text description 130 or text embedding 136 via a transformer. In some embodiments, first sub-model 138 may comprise an autoregressive prior. For example, the autoregressive prior may encode text description 130 or text embedding 136 via a transformer. I n some embodiments, first sub-model 138, prior to generating corresponding image embedding 140, first sub-model 138 may include a diffusion prior model, which may include a transformer. The output of first sub-model may be the corresponding image embedding 140 . For example, the diffusion prior or the autoregressive prior may generate the corresponding image embedding 140, as discussed herein. In some embodiments, corresponding image embedding 140 may be different from image embedding 135 resultant from training of image encoder 134. Corresponding image embedding 140 may be accessible to other models or to be used for image generation .”; [0044] As referenced in FIG. 2, at step 210, a method may involve generating, from at least one of the text description or the text embedding, a corresponding image embedding”) , wherein the embedding conversion model comprises an autoregressive model or a diffusion model ( RAMESH [0036] FIG. 1B illustrates a functional diagram for generating an image corresponding to a text input, consistent with embodiments of the present disclosure. As discussed herein, training an image generation model may involve training an image encoder 134 and a text encoder 132. Joint representation 137 may be a latent space representation of the text embedding 136 and the image embedding 135 . In some embodiments, upon completion of training, text encoder 132 and image encoder 134 may be fixed and may be used as part of machine learning models. In some embodiments, an image generation request 128 comprises a request or query (e.g., an electronic transmission of information) to input/output devices 318. For example, a user may initiate image generation request 128 by interacting with input/output devices 318. Generating an image corresponding to a text description may involve accessing a text description 130 and inputting the text description 130 to the text encoder 132. Text encoder 132 may generate a text embedding 136, which may be an input to first sub-model 138. [0052] It is appreciated that embedding images and text to the same latent space contributes to enabling language-guided image manipulations. Mapping embeddings of images and text to the same latent space, together with language-guided modeling forms a non-conventional and non-generic arrangement, which contributes to the capability to manipulate an existing image based on textual information. For example, an image may be modified to reflect a new text description. Disclosed embodiments may involve accessing the text embedding corresponding to the output image. In some embodiments, the text embedding corresponding to the output image may include the text embedding received from the text encoder in the first sub-model. For example, the text embedding may be based on the input or caption which the output image is based on. In some embodiments, the text embedding may correspond to a baseline, such as a generic caption, or an empty caption. Some embodiments may involve, based on a second text description, accessing a second text embedding. A second text description may include the new text caption or description which the modified image should reflect. Accessing the second text embedding may include inputting the second description into the text encoder, and receiving the resultant text embedding. Disclosed embodiments may involve generating a vector representation of the text embedding corresponding to the output image and the second text embedding. In some embodiments, the vector representation, which may be referred to as a difference vector or text diff, may be the normalization of the difference between the second text embedding and the text embedding corresponding to the output image. Disclosed embodiments may involve performing an interpolation between the image embedding of the output image and the vector representation of the text embedding corresponding to the output image and the second text embedding. Performing an interpolation may involve rotating between the image embedding and the difference vector with spherical interpolation, such as spherical linear interpolation. Based on the interpolation, disclosed embodiments may involve generating a modified instance of the output image. Based on the interpolation may involve the interpolation producing intermediate representations, such as interpolates or trajectories of embeddings based on an angle. Generating the modified instance of the output image may inversion to reconstruct the image, or applying the interpolates to a decoder or diffusion decoder model, as discussed herein.”) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of claimed invention to modify the method of generating images base on text prompt of Karpman and Liu with transforming text embedding to image embedding as seen in RAMESH because this modification would generate, based on at least one of the text description or the text embedding, a corresponding image embedding (abstract of KAMESH) Thus, the combination of Karpman, Liu an KAMESH teaches wherein: the style encoder further comprises an embedding conversion model configured to convert the style text embedding to the style embedding, wherein the embedding conversion model comprises an autoregressive model or a diffusion model . 07-21-aia AIA 4. Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Karpman et al., U.S Patent No.11,995,803 (“Karpman”) in view of Liu et al, U.S Patent Application Publication No.20210358164 (“Liu”) further in view of Pan et al, U.S Patent Application Publication No.20250014233 (“Pan”) Regarding claim 20, Karpman, Liu teach the system of claim 15, wherein: the style encoder further comprises a multimodal text encoder configured to obtain a style text embedding, wherein the style input comprises text or an image (see at least col.4, lines 46-67-col.5, lines 1-6 of Karpman “During pre-training and/or training stages, the system can execute the multi-modal encoder-decoder 126 on text prompts included in the training corpus to generate embeddings used by the text-to-image diffusion model 112 to generate images ( e.g., in addition to embeddings generated by the pre-trained text encoder 118, instead of embeddings generated by the pre-trained text encoder 118 or via a dropout method). In some implementations, a multimodal encoder-decoder 126 includes and/or interfaces with a vision transformer configured to divide an input image into segments and generate a corresponding sequence of embeddings for the input image. T he multimodal encoder-decoder 126 can operate in different ways , such as: (1) as a unimodal encoder that generates image embeddings given an image input and generates text embeddings given a text input (e.g., generates separate image and text embeddings); ( 2) as an image-aware text encoder that modifies embeddings generated by the pre-trained text encoder and/or unimodal encoder module to include visual information based on visual features of an input image (e.g., via cross-attention to the input image) ; and (3) as an image-aware text decoder that generates a text (e.g., natural language) description of an input image based on image-text embedding representations. Generally, the image-aware text encoder and the image aware text decoder can share a common architecture and/or parameters (e.g., with the exception of the cross-attention and self-attention layers) in order to improve training efficiency . ”; Liu : [0091] “In at least one embodiment, a content input 404 is an image or image data of any type containing one or more objects to which a style from a style input 406 is to be applied by a generator 402. In at least one embodiment, a style input 406 is one or more images or image data of any type containing a style to be applied to a content input 404 by a generator 402. When multiple style inputs 406 are used, in an embodiment, a generator 402 extracts a style from each style input 406 and uses an average style to be applied to a content input 404.” [0093] “In at least one embodiment, a style encoder E.sub.s 410 is data values and a set of software instructions that, when executed, implement one or more neural networks or models that take one or more style inputs x.sub.s 406 and output a style embedding z.sub.s. In at least one embodiment, a style embedding z.sub.s is a vector or set of numbers containing continuous values that represent style information about a style input x.sub.s 406. In at least one embodiment, a style embedding z.sub.s contains other information about a style input x.sub.s 406 such as pose of objects contained in said style input x.sub.s 406. In at least one embodiment, a style encoder E.sub.s 410 that takes two or more style images as inputs 406 outputs an individual style embedding z.sub.s.sup.i for each input i, and a generator 402 then averages each style embedding to determine z.sub.s that will be applied to a content input 404.”) In addition, the same motivation is used as the rejection for claim 1. Both Karpman, Liu are understood to be silent on the remaining limitations of claim 20. In the same field of endeavor, Pan teaches t he style encoder further comprises a text encoder configured to obtain a style text embedding, wherein the style input comprises text or an image (see at least [0074] In the embodiment depicted in FIG. 4B, the images 402, 412, 422, and 432 are generated using a diffusion model before optimizing the network weights thereof, which may be a pre-trained DPM (e.g., a text-to-image diffusion-based generative model) with an input as a content conditioner. For example, a user may input a text prompt (e.g., “apple”) as a content conditioner to the pre-trained text-to-image diffusion-based generative model to generate a content of an image consistent with the text prompt . The generated images 402, 412, 422, and 432 may not have the target or reference style of the reference images 403 and 405. That is, the images 402, 412, 422, and 432 generated before optimizing the network weights of the diffusion model may not have the visual appearance or unique visual characteristics of the target or reference object. [0079] At block 520 (Generate clean objects), the processor may generate one or more clean objects by performing a forward generation process of a diffusion model with a control signal. The control signal is configured to control a visual effect of the clean objects. In one example embodiments, the generating of the clean objects can include generating a final image after a completion of iteratively de-noising a starting noise by performing the forward generation process of the diffusion model. For example, the de-noising module 250 can perform the forward generation process 350 with the starting noise 320 to generate the final image 340, as shown in FIGS. 2 and 3. In one example embodiment, the diffusion model can be a pre-trained DPM (e.g., a text-to-image diffusion-based generative model) with a first input as a content conditioner and a second input as a visual effect conditioner to generate the clean objects. The visual effect conditioner may be represented as a text embedding as denoted as “#” in FIG. 5B. The text embedding can be combined (e.g., concatenated) with embeddings of other text prompts to generate the clean objects . For example, the diffusion model may receive a first text prompt (e.g., “a cute totoro in a yard”) as a content conditioner and a second text prompt as a visual effect conditioner (e.g., “bokeh”) to the pre-trained text-to-image diffusion-based generative model, and the corresponding embeddings can be concatenated to generate a clean image. The reference object and the clean objects may be obtained using the pre-trained text-to-image diffusion-based generative model with the same first input (e.g., “a cute totoro in a yard”). Processing may proceed from block 520 to block 530”) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of claimed invention to modify the method of generating images base on text prompt of Karpman and Liu with style input comprise text as seen in Pan because this modification would provide a visual effect conditioner via text prompt ([0079] of Pan) Thus, the combination of Karpman, Liu and Pan teaches the style encoder further comprises a multimodal text encoder configured to obtain a style text embedding, wherein the style input comprises text or an image . Contact Any inquiry concerning this communication or earlier communications from the examiner should be directed to SARAH LE whose telephone number is (571)270-7842. The examiner can normally be reached Monday: 8AM-4:30PM EST, Tuesday: 8 AM-3:30PM EST, Wednesday: 8AM-2:30PM EST, Thursday and Friday off. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kent Chang can be reached at (571) 272-7667. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SARAH LE/Primary Examiner, Art Unit 2614 Application/Control Number: 18/903,151 Page 2 Art Unit: 2614 Application/Control Number: 18/903,151 Page 3 Art Unit: 2614 Application/Control Number: 18/903,151 Page 4 Art Unit: 2614 Application/Control Number: 18/903,151 Page 5 Art Unit: 2614 Application/Control Number: 18/903,151 Page 6 Art Unit: 2614 Application/Control Number: 18/903,151 Page 7 Art Unit: 2614 Application/Control Number: 18/903,151 Page 8 Art Unit: 2614 Application/Control Number: 18/903,151 Page 9 Art Unit: 2614 Application/Control Number: 18/903,151 Page 10 Art Unit: 2614 Application/Control Number: 18/903,151 Page 11 Art Unit: 2614 Application/Control Number: 18/903,151 Page 12 Art Unit: 2614 Application/Control Number: 18/903,151 Page 13 Art Unit: 2614 Application/Control Number: 18/903,151 Page 14 Art Unit: 2614 Application/Control Number: 18/903,151 Page 15 Art Unit: 2614 Application/Control Number: 18/903,151 Page 16 Art Unit: 2614 Application/Control Number: 18/903,151 Page 17 Art Unit: 2614 Application/Control Number: 18/903,151 Page 18 Art Unit: 2614 Application/Control Number: 18/903,151 Page 19 Art Unit: 2614 Application/Control Number: 18/903,151 Page 20 Art Unit: 2614 Application/Control Number: 18/903,151 Page 21 Art Unit: 2614 Application/Control Number: 18/903,151 Page 22 Art Unit: 2614 Application/Control Number: 18/903,151 Page 23 Art Unit: 2614 Application/Control Number: 18/903,151 Page 24 Art Unit: 2614 Application/Control Number: 18/903,151 Page 25 Art Unit: 2614 Application/Control Number: 18/903,151 Page 26 Art Unit: 2614 Application/Control Number: 18/903,151 Page 27 Art Unit: 2614 Application/Control Number: 18/903,151 Page 28 Art Unit: 2614 Application/Control Number: 18/903,151 Page 29 Art Unit: 2614 Application/Control Number: 18/903,151 Page 30 Art Unit: 2614 Application/Control Number: 18/903,151 Page 31 Art Unit: 2614 Application/Control Number: 18/903,151 Page 32 Art Unit: 2614 Application/Control Number: 18/903,151 Page 33 Art Unit: 2614 Application/Control Number: 18/903,151 Page 34 Art Unit: 2614 Application/Control Number: 18/903,151 Page 35 Art Unit: 2614 Application/Control Number: 18/903,151 Page 36 Art Unit: 2614 Application/Control Number: 18/903,151 Page 37 Art Unit: 2614 Application/Control Number: 18/903,151 Page 38 Art Unit: 2614 Application/Control Number: 18/903,151 Page 39 Art Unit: 2614 Application/Control Number: 18/903,151 Page 40 Art Unit: 2614 Application/Control Number: 18/903,151 Page 41 Art Unit: 2614 Application/Control Number: 18/903,151 Page 42 Art Unit: 2614 Application/Control Number: 18/903,151 Page 43 Art Unit: 2614 Application/Control Number: 18/903,151 Page 44 Art Unit: 2614 Application/Control Number: 18/903,151 Page 45 Art Unit: 2614 Application/Control Number: 18/903,151 Page 46 Art Unit: 2614 Application/Control Number: 18/903,151 Page 47 Art Unit: 2614 Application/Control Number: 18/903,151 Page 48 Art Unit: 2614 Application/Control Number: 18/903,151 Page 49 Art Unit: 2614 Application/Control Number: 18/903,151 Page 50 Art Unit: 2614 Application/Control Number: 18/903,151 Page 51 Art Unit: 2614 Application/Control Number: 18/903,151 Page 52 Art Unit: 2614 Application/Control Number: 18/903,151 Page 53 Art Unit: 2614 Application/Control Number: 18/903,151 Page 54 Art Unit: 2614 Application/Control Number: 18/903,151 Page 55 Art Unit: 2614