CTNF 18/753,108 CTNF 97656 DETAILED ACTION Notice of Pre-AIA or AIA Status 07-03-aia AIA 15-10-aia The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA. Information Disclosure Statement The information disclosure statement (IDS) submitted are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Objections 07-29-01 AIA Claim 4 is objected to because of the following informalities: Add “and” at the end of first second limitation . Appropriate correction is required. Claim Rejections - 35 USC § 112 07-30-02 AIA The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. 07-34-01 Claims 2 and 3 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 2 recites the limitation “the first term”. There is insufficient antecedent basis for this limitation in the claim as there is no prior definition for “first term” in claim 1. However, claim 1 defines “a term” in fifth limitation and it is unclear and confusing to one of the ordinary skill in the art if the applicant is referring to “term” as “first term”. Claim 3 recites the limitation “the first term” in the preamble. There is insufficient antecedent basis for this limitation in the claim as there is no prior definition for “first term” in claim 1. However, claim 1 defines “term” in fifth limitation and it is unclear and confusing to one of the ordinary skill in the art if the applicant is referring to “term” as “first term”. Appropriate corrections are needed. Claim Rejections - 35 USC § 102 07-06 AIA 15-10-15 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. 07-07-aia AIA 07-07 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – 07-08-aia AIA (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. 07-12-aia AIA (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. 07-14-aia AIA (g)(1) during the course of an interference conducted under section 135 or section 291, another inventor involved therein establishes, to the extent permitted in section 104, that before such person’s invention thereof the invention was made by such other inventor and not abandoned, suppressed, or concealed, or (2) before such person’s invention thereof, the invention was made in this country by another inventor who had not abandoned, suppressed, or concealed it. In determining priority of invention under this subsection, there shall be considered not only the respective dates of conception and reduction to practice of the invention, but also the reasonable diligence of one who was first to conceive and last to reduce to practice, from a time prior to conception by the other. A rejection on this statutory basis (35 U.S.C. 102(g) as in force on March 15, 2013) is appropriate in an application or patent that is examined under the first to file provisions of the AIA if it also contains or contained at any time (1) a claim to an invention having an effective filing date as defined in 35 U.S.C. 100(i) that is before March 16, 2013 or (2) a specific reference under 35 U.S.C. 120, 121, or 365(c) to any patent or application that contains or contained at any time such a claim. 07-103 AIA The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action. 07-15 AIA Claim s 1 – 3, 7, 8, and 13 are rejected under 35 U.S.C. 102( a)(1 ) as being anticipated by Chefer, H., Alaluf, Y., Vinker, Y., Wolf, L., & Cohen-Or, D. (2023). Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models. ACM transactions on Graphics (TOG) , 42 (4), 1-10; hereafter referred to as Chefer) . Regarding Claim 1 , Chefer teaches: A computer-implemented method, comprising: generating an image (Chefer, Fig. 3, image x), wherein the image is generated by a neural network (Chefer, Fig. 3, page 148:3, col. 1, section 3, UNet network) and wherein the generating of the images includes the following steps: providing a randomly drawn image or a representation of the randomly drawn image as input of a sequence of layers of the neural network, wherein the sequence of layers includes at least one cross-attention layer (Chefer, Fig. 3, page 148:3, col. 1, section 3, “an encoder E is trained to map a given image 𝑥 ∈ X into a spatial latent code 𝑧 =E( 𝑥 ).. produce a denoised version of an input latent 𝑧𝑡 at each time step 𝑡 . During the denoising process, the diffusion model can be conditioned on an additional input vector….this additional input is typically a text encoding produced by a pre-trained CLIP text encoder… Diffusion is performed using the cross-attention mechanism. The denoising UNet network consists of self-attention layers followed by cross-attention layers”); providing a first input to the cross-attention layer (Chefer, Fig. 3, cross attention), wherein the first input is either a representation of the input of the sequence of layers determined by layers of the sequence of layers preceding the cross-attention layer or the first input is the input to the sequence of layers (Chefer, page 148:3, section 3, “an encoder E is trained to map a given image 𝑥 ∈ X into a spatial latent code 𝑧 =E( 𝑥 ). A decoder D is then tasked with reconstructing the input image such that D(E( 𝑥 ))≈ 𝑥 ”), and providing a text embedding characterizing a description of the image to be generated as a second input to the cross-attention layer (Fig. 3, “Given a prompt(e.g. “A lion with a crown”),we extract the subject tokens (lion, crown), and their corresponding attention maps( 𝐴 2 𝑡 , 𝐴 5 𝑡 )”); determining, by the cross-attention layer, an attention map based on the first input and the second input (Chefer, page 148:4, section 4, “Extracting the Cross-Attention Maps. Given the input text prompt P, we consider the set of all subject tokens (e.g., nouns) S = { 𝑠 1, ..., 𝑠𝑘 } present in P…to extract a spatial attention map for each token 𝑠 ∈ S, indicating the influence of the token 𝑠 on each image patch… perform a forward pass through the pre-trained UNet network using 𝑧𝑡 and P (Step 1 in Algorithm 1)… the resulting cross attention map obtained. The resulting aggregated map 𝐴𝑡 contains 𝑁 spatial attention maps, one for each of the tokens of P”); optimizing the input provided to the sequence of layers based on a loss function to determine an optimized input, wherein the loss function includes a term characterizing a negative total variation of the attention map (Chefer, page 148:4, section 4, “our optimization encourages the existence of at least one patch of 𝐴𝑠 𝑡 with a high activation value. Therefore, we define the loss quantifying this desired behavior”; Fig. 3, optimization enhances the maximal activation for the most neglected token at timestep 𝑡 and updates the latent code 𝑧 t”; the negative total variation of the attention map is computed as per equation 2”); determining an output of the sequence of layers based on the optimized input (Chefer, page 148:4, Algorithm 1, output step 16, Zt-1”) ; and determining the image based on the determined output of the sequence of layers ( Chefer, reconstructed image D(E( 𝑥 )) ≈ 𝑥 , Fig. 3 final cross attention maps). Regarding Claim 2 , Chefer teaches the method of claim 1, wherein the text embedding includes embeddings for a plurality of subjects included in the description, a respective attention map is determined for each subject of the plurality of subjects by the cross-attention layer, and a negative total variation is determined for each respective attention map corresponding to a subject from the plurality of subjects, and wherein the first term characterizes a minimal negative total variation among the determined negative total variations (Chefer, page 148:3, section 3, “An attention map 𝐴𝑡 ∈ R 𝑃 × 𝑃 × 𝑁 is calculated over linear projections of the intermediate features( 𝑄 )and text embed ding( 𝐾 ),as illustrated in the second row of Figure3. A 𝑡 defines a distribution over the text tokens for each spatial patch ( 𝑖 , 𝑗 ). Specifically, 𝐴 𝑡 [ 𝑖 , 𝑗 , 𝑛 ] denotes the probability assigned to token n for the ( 𝑖 , 𝑗 )-th spatial patch of the intermediate feature map”; Chefer, page 148:4, col. 2, section 4, “each subject token in 𝑆 , our optimization encourages the existence of at least one patch of 𝐴𝑠 𝑡 with a high activation value. Therefore, we define the loss quantifying this desired behavior as L =max 𝑠∈𝑆 L 𝑠 where L 𝑠 = 1−max ( 𝐴𝑠 𝑡 )”). Regarding Claim 3 , Chefer teaches the method of claim 1, wherein the first term is characterized by the formula: PNG media_image1.png 50 512 media_image1.png Greyscale wherein At is the attention map and expression At[i,j,s] characterizes the attention map at position i,j for the embedding s of a set of subjects S included in the description (Chefer, page 148:3, section 3, “An attention map 𝐴𝑡 ∈ R 𝑃 × 𝑃 × 𝑁 is calculated over linear projections of the intermediate features( 𝑄 )and text embed ding( 𝐾 ),as illustrated in the second row of Figure3. A 𝑡 defines a distribution over the text tokens for each spatial patch ( 𝑖 , 𝑗 ). Specifically, 𝐴 𝑡 [ 𝑖 , 𝑗 , 𝑛 ] denotes the probability assigned to token n for the ( 𝑖 , 𝑗 )-th spatial patch of the intermediate feature map”; page 148:4, col. 2, section 4, equation 2”). Regarding Claim 7 , Chefer teaches the method of claim 1, where the optimized input is determined by minimizing the loss function using a gradient descent method (Chefer, page 184:4, col. 2, “computed our loss L, we shift the current latent 𝑧𝑡 by 𝑧 ′ 𝑡 ← 𝑧𝑡 − 𝛼𝑡 · ∇𝑧𝑡 L, (3) where 𝛼𝑡 is a scalar defining the step size of the gradient update”). Regarding Claim 8 , Chefer teaches the method of claim 1, wherein the neural network is a latent diffusion model, including a stable diffusion model or a normalizing flow (Chefer, page 148:3, col.1 “Latent Diffusion Models. We apply our method over the state-of the-art Stable Diffusion model (SD). Instead of operating in the image space, SD operates in the latent space of an autoencoder”). Regarding Claim 13 , Chefer teaches: A non-transitory machine-readable storage medium on which is stored a computer program, the computer program, when executed by a processor (Chefer, Computing methodologies → Computer graphics; Image processing”), causing the processor to perform the following steps: generating an image (Chefer, Fig. 3, image x), wherein the image is generated by a neural network (Chefer, Fig. 3, page 148:3, col. 1, section 3, UNet network) and wherein the generating of the images includes the following steps: providing a randomly drawn image or a representation of the randomly drawn image as input of a sequence of layers of the neural network, wherein the sequence of layers includes at least one cross-attention layer (Chefer, Fig. 3, page 148:3, col. 1, section 3, “an encoder E is trained to map a given image 𝑥 ∈ X into a spatial latent code 𝑧 =E( 𝑥 ).. produce a denoised version of an input latent 𝑧𝑡 at each time step 𝑡 . During the denoising process, the diffusion model can be conditioned on an additional input vector….this additional input is typically a text encoding produced by a pre-trained CLIP text encoder… Diffusion is performed using the cross-attention mechanism. The denoising UNet network consists of self-attention layers followed by cross-attention layers”); providing a first input to the cross-attention layer (Chefer, Fig. 3, cross attention), wherein the first input is either a representation of the input of the sequence of layers determined by layers of the sequence of layers preceding the cross-attention layer or the first input is the input to the sequence of layers (Chefer, page 148:3, section 3, “an encoder E is trained to map a given image 𝑥 ∈ X into a spatial latent code 𝑧 =E( 𝑥 ). A decoder D is then tasked with reconstructing the input image such that D(E( 𝑥 ))≈ 𝑥 ”), and providing a text embedding characterizing a description of the image to be generated as a second input to the cross-attention layer (Fig. 3, “Given a prompt(e.g. “A lion with a crown”),we extract the subject tokens (lion, crown), and their corresponding attention maps( 𝐴 2 𝑡 , 𝐴 5 𝑡 )”); determining, by the cross-attention layer, an attention map based on the first input and the second input (Chefer, page 148:4, section 4, “Extracting the Cross-Attention Maps. Given the input text prompt P, we consider the set of all subject tokens (e.g., nouns) S = { 𝑠 1, ..., 𝑠𝑘 } present in P…to extract a spatial attention map for each token 𝑠 ∈ S, indicating the influence of the token 𝑠 on each image patch… perform a forward pass through the pre-trained UNet network using 𝑧𝑡 and P (Step 1 in Algorithm 1)… the resulting cross attention map obtained. The resulting aggregated map 𝐴𝑡 contains 𝑁 spatial attention maps, one for each of the tokens of P”); optimizing the input provided to the sequence of layers based on a loss function to determine an optimized input, wherein the loss function includes a term characterizing a negative total variation of the attention map (Chefer, page 148:4, section 4, “our optimization encourages the existence of at least one patch of 𝐴𝑠 𝑡 with a high activation value. Therefore, we define the loss quantifying this desired behavior”; Fig. 3, optimization enhances the maximal activation for the most neglected token at timestep 𝑡 and updates the latent code 𝑧 t”; the negative total variation of the attention map is computed as per equation 2”); determining an output of the sequence of layers based on the optimized input (Chefer, page 148:4, Algorithm 1, output step 16, Zt-1”) ; and determining the image based on the determined output of the sequence of layers ( Chefer, reconstructed image D(E( 𝑥 )) ≈ 𝑥 , Fig. 3 final cross attention maps) . Claim Rejections - 35 USC § 103 07-06 AIA 15-10-15 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. 07-20-aia AIA The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 07-103 AIA The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action. 07-23-aia AIA The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. 07-21-aia AIA Claim s 4 and 5 are rejected under 35 U.S.C. 103 as being unpatentable over Chefer et al. (Chefer, H., Alaluf, Y., Vinker, Y., Wolf, L., & Cohen-Or, D. (2023). Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models. ACM transactions on Graphics (TOG) , 42 (4), 1-10; hereafter referred to as Chefer) in view of Rassin et al. (Rassin, R., Hirsch, E., Glickman, D., Ravfogel, S., Goldberg, Y., & Chechik, G. (2023). Linguistic binding in diffusion models: Enhancing attribute correspondence through attention map alignment. Advances in Neural Information Processing Systems , 36 , 3536-3559; hereafter referred to as Rassin) . Regarding Claim 4 , Chefer teaches the method of claim 1, wherein the text embedding includes an embedding for a subject included in the description and an embedding for an attribute included in the description and describing the subject (Chefer, page 148:6, section 5, results, “to test correct attribute binding, the prompts should contain a variety of attributes matched to the subject tokens. Specifically, we consider three types of text prompts: (i) “a [animalA] and a [animalB]”, (ii) “a [animal] and a [color][object]”, and (iii) “a [colorA][objectA] and a [colorB ][objectB]”), and wherein the method further comprises the following steps: determining, by the cross-attention layer, a first attention map for the attribute and a second attention map for the subject (Chefer, Fig. 3, page 148:3, cross-attention maps for all the words in the prompt”; Chefer, page 148:6, col. 1, “both the cat and the frog are accurately localized in the attention maps”); However, Chefer fails to explicitly teach: optimizing the input provided to the sequence of layers based on the loss function wherein the loss function includes a second term, wherein the second term characterizes a difference between the first attention map and the second attention map. In the same field of endeavor, Rassin teaches: optimizing the input provided to the sequence of layers based on the loss function wherein the loss function includes a second term, wherein the second term characterizes a difference between the first attention map and the second attention map (Rassin, Fig. 2 and 3, page 3, loss function, “Our first loss aims to minimize that distance (maximize the overlap) over all pairs of modifiers and their corresponding entity nouns, equation 3”; “included in the generated image, and our loss depends on pairwise relations of linguistically-related words and aims to align the diffusion process to the linguistic-structure of the prompt”). Chefer and Rassin are considered analogous art as they are reasonably pertinent to the same field of endeavor of image processing. Therefore, it would have been obvious to one of the ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Chefer with the invention of Rassin to optimize the input provided to the sequence of layers based on the loss function that includes a second term; doing so can enable generation of higher-quality images (Rassin, Abstract); thus, one of the ordinary skill in the art would have been motivated to combine the references. Regarding Claim 5 , Chefer in view of Rassin teaches the method of claim 4, wherein the difference is a Jensen-Shannon divergence (Rassin, page 4, “a measure of distance between attention maps we use a symmetric Kullback-Leibler divergence dist(Ai,Aj) = 1 2 DKL(Ai||Aj) + 1 2 DKL(Aj||Ai), where Ai, Aj are attention maps normalized to a sum of 1, i and j are generic indices”; “ the Jensen–Shannon divergence is based on the Kullback–Leibler divergence,” (Jensen–Shannon divergence – Wikipedia https://en.wikipedia.org/wiki/Jensen%E2%80%93Shannon_divergence) . 07-21-aia AIA Claim s 9 – 11 are rejected under 35 U.S.C. 103 as being unpatentable over Chefer et al. (Chefer, H., Alaluf, Y., Vinker, Y., Wolf, L., & Cohen-Or, D. (2023). Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models. ACM transactions on Graphics (TOG) , 42 (4), 1-10; hereafter referred to as Chefer) in view of Mariani et al. (Mariani, G., Scheidegger, F., Istrate, R., Bekas, C., & Malossi, C. (2018). Bagan: Data augmentation with balancing gan. arXiv preprint arXiv:1803.09655 ; hereafter referred to as Mariani) . Regarding Claim 9, Chefer teaches the method of claim 1, but fails to explicitly teach: training or testing an image classification system and/or image regression system using the generated image. In the same field of endeavor, Mariani teaches: training or testing an image classification system and/or image regression system using the generated image (Mariani, Section 5.2, “assess the accuracy of a deep-learning classifier trained on an augmented dataset… train the considered generative models, 4) augment the imbalanced data set to restore its balance by means of the generative models”). Chefer and Mariani are considered analogous art as they are reasonably pertinent to the same field of endeavor of image processing. Therefore, it would have been obvious to one of the ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Chefer with the invention of Mariani to train an image classification system and/or image regression system using the generated image; doing so can enable generation of images with higher classification accuracy (Mariani, section 5.2); thus, one of the ordinary skill in the art would have been motivated to combine the references. Regarding Claim 10, Chefer in view of Mariani teaches the method of claim 9, further comprising: determining a control signal of an actuator and/or a display based on an output of the trained image classification system and/or image regression system (Mariani, section 5.2, “Accuracy results averaged over the different classes are shown in Figure 9. The proposed BAGAN methodology returns the best accuracy”; Output (accuracy results) of the trained classifier is shown in Fig. 9). Regarding Claim 11 , Chefer teaches: A training system, configured to: generate an image (Chefer, Fig. 3, image x), wherein the image is generated by a neural network (Chefer, Fig. 3, page 148:3, col. 1, section 3, UNet network) and wherein the generating of the images includes the following steps: provide a randomly drawn image or a representation of the randomly drawn image as input of a sequence of layers of the neural network, wherein the sequence of layers includes at least one cross- attention layer (Chefer, Fig. 3, page 148:3, col. 1, section 3, “an encoder E is trained to map a given image 𝑥 ∈ X into a spatial latent code 𝑧 =E( 𝑥 ).. produce a denoised version of an input latent 𝑧𝑡 at each time step 𝑡 . During the denoising process, the diffusion model can be conditioned on an additional input vector….this additional input is typically a text encoding produced by a pre-trained CLIP text encoder… Diffusion is performed using the cross-attention mechanism. The denoising UNet network consists of self-attention layers followed by cross-attention layers”); provide a first input to the cross-attention layer (Chefer, Fig. 3, cross attention), wherein the first input is either a representation of the input of the sequence of layers determined by layers of the sequence of layers preceding the cross-attention layer or the first input is the input to the sequence of layers (Chefer, page 148:3, section 3, “an encoder E is trained to map a given image 𝑥 ∈ X into a spatial latent code 𝑧 =E( 𝑥 ). A decoder D is then tasked with reconstructing the input image such that D(E( 𝑥 ))≈ 𝑥 ”), and providing a text embedding characterizing a description of the image to be generated as a second input to the cross-attention layer (Fig. 3, “Given a prompt(e.g. “A lion with a crown”),we extract the subject tokens (lion, crown), and their corresponding attention maps( 𝐴 2 𝑡 , 𝐴 5 𝑡 )”); determine, by the cross-attention layer, an attention map based on the first input and the second input (Chefer, page 148:4, section 4, “Extracting the Cross-Attention Maps. Given the input text prompt P, we consider the set of all subject tokens (e.g., nouns) S = { 𝑠 1, ..., 𝑠𝑘 } present in P…to extract a spatial attention map for each token 𝑠 ∈ S, indicating the influence of the token 𝑠 on each image patch… perform a forward pass through the pre-trained UNet network using 𝑧𝑡 and P (Step 1 in Algorithm 1)… the resulting cross attention map obtained. The resulting aggregated map 𝐴𝑡 contains 𝑁 spatial attention maps, one for each of the tokens of P”); optimize the input provided to the sequence of layers based on a loss function to determine an optimized input, wherein the loss function includes a term characterizing a negative total variation of the attention map (Chefer, page 148:4, section 4, “our optimization encourages the existence of at least one patch of 𝐴𝑠 𝑡 with a high activation value. Therefore, we define the loss quantifying this desired behavior”; Fig. 3, optimization enhances the maximal activation for the most neglected token at timestep 𝑡 and updates the latent code 𝑧 t”; the negative total variation of the attention map is computed as per equation 2”); determine an output of the sequence of layers based on the optimized input (Chefer, page 148:4, Algorithm 1, output step 16, Zt-1”) ; and determine the image based on the determined output of the sequence of layers ( Chefer, reconstructed image D(E( 𝑥 )) ≈ 𝑥 , Fig. 3 final cross attention maps). However, Chefer fails to explicitly teach: train an image classification system and/or image regression system using the generated image. In the same field of endeavor, Mariani teaches: train an image classification system and/or image regression system using the generated image (Mariani, Section 5.2, “assess the accuracy of a deep-learning classifier trained on an augmented dataset… train the considered generative models, 4) augment the imbalanced data set to restore its balance by means of the generative models”). Chefer and Mariani are considered analogous art as they are reasonably pertinent to the same field of endeavor of image processing. Therefore, it would have been obvious to one of the ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Chefer with the invention of Mariani to train an image classification system and/or image regression system using the generated image; doing so can enable generation of images with higher classification accuracy (Mariani, section 5.2); thus, one of the ordinary skill in the art would have been motivated to combine the references . 07-21-aia AIA Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Mariani et al. (Mariani, G., Scheidegger, F., Istrate, R., Bekas, C., & Malossi, C. (2018). Bagan: Data augmentation with balancing gan. arXiv preprint arXiv:1803.09655 ; hereafter referred to as Mariani) in view of Chefer et al. (Chefer, H., Alaluf, Y., Vinker, Y., Wolf, L., & Cohen-Or, D. (2023). Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models. ACM transactions on Graphics (TOG) , 42 (4), 1-10; hereafter referred to as Chefer) . Regarding Claim 12 , Mariani teaches: A control system configured to determine a control signal of an actuator and/or a display based on an output of the trained image classification system and/or image regression system, the image classification system and/or the image regression system being trained by a training system (Mariani, section 5.2, “Accuracy results averaged over the different classes are shown in Figure 9. The proposed BAGAN methodology returns the best accuracy”; Output (accuracy results) of the trained classifier is shown in Fig. 9) , configured to: train an image classification system and/or image regression system using the generated image (Mariani, Section 5.2, “assess the accuracy of a deep-learning classifier trained on an augmented dataset… train the considered generative models, 4) augment the imbalanced data set to restore its balance by means of the generative models”). However, Mariani fails to explicitly teach: generate an image, wherein the image is generated by a neural network and wherein the training system is configured to: provide a randomly drawn image or a representation of the randomly drawn image as input of a sequence of layers of the neural network, wherein the sequence of layers includes at least one cross-attention layer, provide a first input to the cross-attention layer, wherein the first input is either a representation of the input of the sequence of layers determined by layers of the sequence of layers preceding the cross-attention layer or the first input is the input to the sequence of layers, and provide a text embedding characterizing a description of the image to be generated as a second input to the cross-attention layer, determine, by the cross-attention layer, an attention map based on the first input and the second input, optimize the input provided to the sequence of layers based on a loss function to determine an optimized input, wherein the loss function includes a term characterizing a negative total variation of the attention map, determine an output of the sequence of layers based on the optimized input, and determine the image based on the determined output of the sequence of layers. In the same field of endeavor, Chefer teaches: generate an image (Chefer, Fig. 3, image x), wherein the image is generated by a neural network (Chefer, Fig. 3, page 148:3, col. 1, section 3, UNet network) and wherein the generating of the images includes the following steps: provide a randomly drawn image or a representation of the randomly drawn image as input of a sequence of layers of the neural network, wherein the sequence of layers includes at least one cross-attention layer (Chefer, Fig. 3, page 148:3, col. 1, section 3, “an encoder E is trained to map a given image 𝑥 ∈ X into a spatial latent code 𝑧 =E( 𝑥 ).. produce a denoised version of an input latent 𝑧𝑡 at each time step 𝑡 . During the denoising process, the diffusion model can be conditioned on an additional input vector….this additional input is typically a text encoding produced by a pre-trained CLIP text encoder… Diffusion is performed using the cross-attention mechanism. The denoising UNet network consists of self-attention layers followed by cross-attention layers”); provide a first input to the cross-attention layer (Chefer, Fig. 3, cross attention), wherein the first input is either a representation of the input of the sequence of layers determined by layers of the sequence of layers preceding the cross-attention layer or the first input is the input to the sequence of layers (Chefer, page 148:3, section 3, “an encoder E is trained to map a given image 𝑥 ∈ X into a spatial latent code 𝑧 =E( 𝑥 ). A decoder D is then tasked with reconstructing the input image such that D(E( 𝑥 ))≈ 𝑥 ”), and providing a text embedding characterizing a description of the image to be generated as a second input to the cross-attention layer (Fig. 3, “Given a prompt(e.g. “A lion with a crown”),we extract the subject tokens (lion, crown), and their corresponding attention maps( 𝐴 2 𝑡 , 𝐴 5 𝑡 )”); determine, by the cross-attention layer, an attention map based on the first input and the second input (Chefer, page 148:4, section 4, “Extracting the Cross-Attention Maps. Given the input text prompt P, we consider the set of all subject tokens (e.g., nouns) S = { 𝑠 1, ..., 𝑠𝑘 } present in P…to extract a spatial attention map for each token 𝑠 ∈ S, indicating the influence of the token 𝑠 on each image patch… perform a forward pass through the pre-trained UNet network using 𝑧𝑡 and P (Step 1 in Algorithm 1)… the resulting cross attention map obtained. The resulting aggregated map 𝐴𝑡 contains 𝑁 spatial attention maps, one for each of the tokens of P”); optimize the input provided to the sequence of layers based on a loss function to determine an optimized input, wherein the loss function includes a term characterizing a negative total variation of the attention map (Chefer, page 148:4, section 4, “our optimization encourages the existence of at least one patch of 𝐴𝑠 𝑡 with a high activation value. Therefore, we define the loss quantifying this desired behavior”; Fig. 3, optimization enhances the maximal activation for the most neglected token at timestep 𝑡 and updates the latent code 𝑧 t”; the negative total variation of the attention map is computed as per equation 2”); determine an output of the sequence of layers based on the optimized input (Chefer, page 148:4, Algorithm 1, output step 16, Zt-1”) ; and determine the image based on the determined output of the sequence of layers ( Chefer, reconstructed image D(E( 𝑥 )) ≈ 𝑥 , Fig. 3 final cross attention maps). Mariani and Chefer are considered analogous art as they are reasonably pertinent to the same field of endeavor of image processing. Therefore, it would have been obvious to one of the ordinary skill in the are before the effective filing date of the claimed invention to modify the invention of Mariani with the invention of Chefer to generate an image, wherein the image is generated by a neural network by providing a first input to the cross-attention layer and provide a text embedding characterizing a description of the image to be generated as a second input to the cross-attention layer; determine, by the cross-attention layer, an attention map based on the first input and the second input; optimize the input provided to the sequence of layers based on a loss function; determine an output of the sequence of layers based on the optimized input, and determine the image based on the determined output of the sequence of layers; doing so can strengthen the text conditioning along the image generation process. (Chefer, section 7); thus, one of the ordinary skill in the art would have been motivated to combine the references. Objected Claims Claim 6 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims and none of the cited prior arts teach the claim limitations recited in claim 6 . Conclusion 07-96 AIA The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. US 20220277218 A1: A transformer based vision-linguistic (VL) model and training technique uses a number of different image patches covering the same portion of an image, along with a text description of the image to train the model. The model and pre-training techniques may be used in domain specific training of the model. The model can be used for fine-grained image-text tasks in the fashion domain. Contact Information Any inquiry concerning this communication or earlier communications from the examiner should be directed to VAISALI RAO KOPPOLU whose telephone number is (571)270-0273. The examiner can normally be reached Monday - Friday 8:30 - 5. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Mehmood can be reached at (571) 272-2976 . The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. VAISALI RAO. KOPPOLU Examiner Art Unit 2664 /VAISALI RAO KOPPOLU/Examiner of Art Unit 2664 Application/Control Number: 18/753,108 Page 2 Art Unit: 2664 Application/Control Number: 18/753,108 Page 3 Art Unit: 2664 Application/Control Number: 18/753,108 Page 4 Art Unit: 2664 Application/Control Number: 18/753,108 Page 5 Art Unit: 2664 Application/Control Number: 18/753,108 Page 6 Art Unit: 2664 Application/Control Number: 18/753,108 Page 7 Art Unit: 2664 Application/Control Number: 18/753,108 Page 8 Art Unit: 2664 Application/Control Number: 18/753,108 Page 9 Art Unit: 2664 Application/Control Number: 18/753,108 Page 10 Art Unit: 2664 Application/Control Number: 18/753,108 Page 11 Art Unit: 2664 Application/Control Number: 18/753,108 Page 12 Art Unit: 2664 Application/Control Number: 18/753,108 Page 13 Art Unit: 2664 Application/Control Number: 18/753,108 Page 14 Art Unit: 2664 Application/Control Number: 18/753,108 Page 16 Art Unit: 2664 Application/Control Number: 18/753,108 Page 17 Art Unit: 2664 Application/Control Number: 18/753,108 Page 18 Art Unit: 2664 Application/Control Number: 18/753,108 Page 19 Art Unit: 2664 Application/Control Number: 18/753,108 Page 20 Art Unit: 2664