Prosecution Insights
Last updated: October 01, 2026
Application No. 18/908,075

CONTROLLING COMPOSITION AND STRUCTURE IN GENERATED IMAGES

Non-Final OA §103
Filed
Oct 07, 2024
Priority
Oct 06, 2023 — provisional 63/588,610
Examiner
LI, JAI WEI TOMMY
Art Unit
2613
Tech Center
2600 — Communications
Assignee
Adobe Inc.
OA Round
2 (Non-Final)
Grant Probability
Favorable
2-3
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-62.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
32 currently pending
Career history
33
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment The objection to the specifications have been withdrawn in view of the applicants amendments filed 07/01/2026. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1 and 6-10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lianghua Huang et al. "Composer: Creative and Controllable Image Synthesis with Composable Conditions", 22 Feb 2023, arXiv, 2302.09778v2 in view of Huberman-Spiegelglas et al. “An Edit Friendly DDPM Noise Space: Inversion and Manipulations”, 12 Apr 2023, arXiv, 2304.06140v1. Regarding claim 1, Huang discloses a method comprising (sec 3.3, “There are two methods to colorize an image x according to palette p using Composer: one entails conditioning the sampling process on both the grayscale version of x and p, while the other involves applying a reconfiguration (Section 2.1) on x in terms of color palette”; also, sec 2.3, “We use diffusion models to recompose images from a set of representations.”): content input indicates an image element (sec 3.3, “Pose transfer: The CLIP embedding of an image captures its style and semantics, enabling Composer to modify the pose of an object without compromising its identity.”; also, sec 3.3, “Style transfer: Composer roughly disentangles the content and style representations, which allows us to transfer the style of image x1 to another image x2 by simply conditioning on the style representations of x1 and the content representations of x2.”; also, provided image comprises a series of elements such as style and semantics) and the composition input indicates a target composition of the image element (sec 3.3, “We use the object’s segmentation map to represent its pose and the image embedding to capture its semantics, then leverage the reconfiguration approach described in Section 2.1 to modify the pose of the object (Figure 5d).”; also, sec 2.2, “Depthmap: We use a pretrained monocular depth estimation model (Ranftl et al., 2022) to extract the depthmap of an image, which roughly captures the image’s layout.”; also, depth estimation model extracts the depthmap from the given content input in order to get the images layout); encoding the composition input to obtain a composition embedding representing the target composition (sec 2.2, “Localized conditioning: For localized representations in cluding sketches, segmentation masks, depthmaps, intensity images, and masked images, we project them into uniform dimensional embeddings with the same spatial size as the noisy latent xt using stacked convolutional layers.”); content input indicating the image element (sec 3.3, “Pose transfer: The CLIP embedding of an image captures its style and semantics, enabling Composer to modify the pose of an object without compromising its identity. We use the object’s segmentation map to represent its pose and the image embedding to capture its semantics, then leverage the reconfiguration approach described in Section 2.1 to modify the pose of the object (Figure 5d)”; also, sec 2.3, “In addition, we project image embeddings and color palettes into eight extra tokens and concatenate them with CLIP word embeddings, which are then used as the context for cross-attention in GLIDE, similar to unCLIP (Ramesh et al., 2022).”) and the composition embedding (sec 2.2, “We then compute the sum of these embeddings and concatenate the result to xt before feeding it into the UNet. Since the embeddings are additive, it is easy to accommodate for missing conditions or to incorporate new localized conditions”; also, sec 2.3, “In addition, we project image embeddings and color palettes into eight extra tokens and concatenate them with CLIP word embeddings, which are then used as the context for cross-attention in GLIDE, similar to unclip”), wherein the synthetic image depicts the image element with the target composition (sec 3.3, “Pose transfer: The CLIP embedding of an image captures its style and semantics, enabling Composer to modify the pose of an object without compromising its identity. We use the object’s segmentation map to represent its pose and the image embedding to capture its semantics, then leverage the reconfiguration approach described in Section 2.1 to modify the pose of the object (Figure 5d)”, also, sec 3.2, “Using Composer, we can create new images that are similar to a given image but vary in certain aspects by conditioning on a specific subset of its representations.”). Huang does not disclose a obtaining a content input, a composition input, and a noise map, and generating, using an image generation model, a synthetic image by denoising the noise map. However, in a similar field of endeavor, Huberman-spiegelglas discloses a obtaining a content input, a composition input, and a noise map map (sec 5, “Suppose we are given a real image x0, a text prompt describing it psrc, and a target text prompt ptar. To modify the image according to these prompts, we extract the edit-friendly noise maps {xT,zT,...,z1}, while injecting psrc to the denoiser. We then fix those noise maps and generate an image while in jecting ptar to the denoiser.”; also, sec 1, “For example, one property we want from an inversion in the context of text-conditional models, is that fixing the noise maps and changing the text-prompt would lead to an artifact-free image, where the semantics correspond to the new text but the structure remains similar to that of the input image”), and generating, using an image generation model, a synthetic image by denoising the noise map (sec 4, “We then fix those noise maps and generate an image while injecting ptar to the denoiser.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Huang's invention of obtaining a content signal that carries the semantics, style, and identity of a depicted element and a structural representation that captures the layout or pose of that element, projecting the structural representation into a composition embedding with stacked convolutional layers, and concatenating that embedding to the latent that the generative network processes, with the features of Huberman-Spiegelglas's invention of being given a real image and a text prompt, extracting noise maps, and then fixing those noise maps and generating an image by injecting a prompt to the denoiser. The combination would have been obvious because Huang extracts its representations from images its own pipeline already holds and describes the thing being denoised only as a latent, leaving the operator with neither a stated way to supply the signals that drive a particular generation nor a stated starting point for that generation, and Huberman-Spiegelglas supplies both in the same field and for the same class of model, being given a real image and a text prompt describing it, and holding a set of noise maps fixed and generating the image from them through the denoiser under a target prompt. Huberman-Spiegelglas further states that fixing the noise maps and changing the text prompt yields an image whose semantics correspond to the new text while the structure remains similar to that of the input image, which is the same division of labor Huang draws between its content signal and its structural map. Huang already conditions its network on embeddings it projects from its representations, so the noise maps Huberman-Spiegelglas holds fixed and denoises enter Huang's pipeline at the same place Huang already feeds its noisy latent, and a person of ordinary skill would have expected the predictable result that the conditioned denoising proceeds from those noise maps to the finished image. Regarding claim 6, Huang as modified by Huberman-Spiegelglas discloses the method of claim 1, wherein Huang further discloses: the composition input comprises a depth map, an edge map, pose information, layout information, or any combination thereof (sec 2.2, “Depthmap: We use a pretrained monocular depth estimation model (Ranftl et al., 2022) to extract the depthmap of an image, which roughly captures the image’s layout”; also, sec 3.3, “We use the object’s segmentation map to represent its pose and the image embedding to capture its semantics, then leverage the reconfiguration approach described in Section 2.1 to modify the pose of the object (Figure 5d).”). Regarding claim 7, Huang as modified by Huberman-Spiegelglas discloses the method of claim 1, wherein Huang further discloses: the image element comprises a style attribute, an identity of an object, a lighting attribute, a texture attribute, a scene attribute, or a combination thereof (sec 2.2, “Semantics and style: We use the image embedding extracted by the pretrained CLIP ViT-L/14@336px (Radford et al., 2021) model to represent the semantics and style of an image, similar to unCLIP (Ramesh et al., 2022).”; also, sec 3.3, “Composer roughly disentangles the content and style representations, which allows us to transfer the style of image x1 to another image x2 by simply conditioning on the style representations of x1 and the content representations of x2”; also, sec 3.3, “The CLIP embedding of an image captures its style and semantics, enabling Composer to modify the pose of an object without compromising its identity”). Regarding claim 8, Huang as modified by Huberman-Spiegelglas discloses the method of claim 1, wherein Huang further discloses: the image generation model is trained to generate images depicting the image element (sec 2.3, “We use diffusion models to recompose images from a set of representations.”; also, sec 3.1, “For the base model, we pretrain it with 1M steps on the full dataset using only image embeddings as the condition, and then finetune the model on a subset of 60M examples (excluding LAION images with aesthetic scores below 7.0) from the original dataset for 200K steps with all conditions enabled.”; also, sec 2.3, “It is essential to devise a joint training strategy that enables the model to learn to decode images from a variety of combinations of conditions.”). Regarding claim 9, Huang as modified by Huberman-Spiegelglas discloses the method of claim 1, wherein Huang further discloses obtaining the composition input comprises: extracting the composition input from the composition image (sec 3.3, “We use the object’s segmentation map to represent its pose and the image embedding to capture its semantics, then leverage the reconfiguration approach described in Section 2.1 to modify the pose of the object (Figure 5d).”; also, sec 2.2, “Sketch: We apply an edge detection model (Su et al., 2021) followed by a sketch simplification algorithm (Simo-Serra et al., 2017) to extract the sketch of an image. Sketches capture local details of images and have less semantics.”; also, sec 2.2, “Depthmap: We use a pretrained monocular depth estimation model (Ranftl et al., 2022) to extract the depthmap of an image, which roughly captures the image’s layout.”). Huang does not fully disclose obtaining a composition image. However, in a similar field of endeavor Huberman-Spiegelglas discloses obtaining a composition image (sec 5, “Suppose we are given a real image x0, a text prompt describing it psrc, and a target text prompt ptar. To modify the image according to these prompts, we extract the edit-friendly noise maps {xT,zT,...,z1}, while injecting psrc to the denoiser. We then fix those noise maps and generate an image while in jecting ptar to the denoiser.”; also, sec 1, “For example, one property we want from an inversion in the context of text-conditional models, is that fixing the noise maps and changing the text-prompt would lead to an artifact-free image, where the semantics correspond to the new text but the structure remains similar to that of the input image”). Regarding claim 10, Huang as modified by Huberman-Spiegelglas discloses the method of claim 9, further comprising: generating the content input based on the composition image (sec 2.2, “Semantics and style: We use the image embedding extracted by the pretrained CLIP ViT-L/14@336px (Radford et al., 2021) model to represent the semantics and style of an image, similar to unCLIP (Ramesh et al., 2022)”; also, sec 2.2, “We decompose an image into decoupled representations which capture various aspects of it. We describe eight representations we use in this work, where all of them are extracted on-the-fly during training.”; also, sec 3.2, “Variations: Using Composer, we can create new images that are similar to a given image but vary in certain aspects by conditioning on a specific subset of its representations. By carefully selecting combinations of different representations”). Claim(s) 2 and 4 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lianghua Huang et al. "Composer: Creative and Controllable Image Synthesis with Composable Conditions", 22 Feb 2023, arXiv, 2302.09778v2 as modified by Huberman-Spiegelglas et al. “An Edit Friendly DDPM Noise Space: Inversion and Manipulations”, 12 Apr 2023, arXiv, 2304.06140v1, further in view of Ruiz et al. “DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation”, 15 Mar 2023, arXiv, 2208.12242v2. Regarding claim 2, Huang as modified by Huberman-Spiegelglas discloses the method of claim 1, . Huang as modified by Huberman-Spiegelglas does not disclose wherein: the content input comprises text describing the image element. However, in a similar field of endeavor, Ruiz discloses wherein: the content input comprises text describing the image element (sec 3.2, “In order to bypass the overhead of writing detailed image descriptions for a given image set we opt for a simpler approach and label all input images of the subject “a [identifier] [class noun]”, where [identifier] is a unique identifier linked to the subject and [class noun] is a coarse class descriptor of the subject (e.g. cat, dog, watch, etc.).”; also, sec 4.4, “We can generate novel images for a specific subject in different contexts (Figure 7) with descriptive prompts (“a [V] [class noun] [context description]”).”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified Huang in view of Huberman-Spiegelglas, in which a content signal carrying the semantics, style, and identity of a depicted element conditions the denoising of the noise maps alongside an encoded structural map, with the features of Ruiz's invention of a descriptive text prompt of the form "a [V] [class noun] [context description]" in which the identifier and the class noun name the particular subject to be depicted and the context description states the scene the subject is to be placed into. The combination would have been obvious because Huang's content signal is an embedding projected from a whole reference image, which delivers whatever the reference image happened to contain and gives the operator no way to state in text which particular element is to appear in the generated image, while Ruiz addresses exactly that shortfall by writing the element itself into the prompt, labeling the images of the subject with an identifier that is linked to that subject and a class noun that describes it, and then generating that same subject in a new context from a prompt built out of those terms. Huang already accepts captions as one of its eight representations and already concatenates the tokens it projects with CLIP word embeddings used as the cross-attention context, so a prompt naming the element enters Huang's pipeline through an interface Huang has already built, and a person of ordinary skill would have expected the predictable result that the generated image depicts the element the text names. Regarding claim 4, Huang as modified by Huberman-Spiegelglas discloses the method of claim 1, content input comprises a nonce token representing the image element. However, in a similar field of endeavor, Ruiz discloses wherein: the content input comprises a nonce token representing the image element (sec 3.2, “Our approach is to find rare tokens in the vocabulary, and then invert these tokens into text space, in order to minimize the probability of the identifier having a strong prior.”; also, sec 3.2, “This motivates the need for an identifier that has a weak prior in both the language model and the diffusion model.”; also, sec 3.2, “In order to bypass the overhead of writing detailed image descriptions for a given image set we opt for a simpler approach and label all input images of the subject “a [identifier] [class noun]”, where [identifier] is a unique identifier linked to the subject and [class noun] is a coarse class descriptor of the subject (e.g. cat, dog, watch, etc.).”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified Huang in view of Huberman-Spiegelglas, in which a content signal carrying the semantics, style, and identity of a depicted element conditions the denoising of the noise maps alongside an encoded structural map, with the features of Ruiz's invention of a rare token drawn from the tokenizer vocabulary as the identifier for that element. The combination would have been obvious because Huang's content signal is a whole image embedding, which supplies whatever element the reference image happened to contain and gives the operator no way to name one particular element in the text the model also accepts, and Ruiz addresses that shortfall with a rare token found in the vocabulary and inverted into text space so that the identifier does not carry a strong prior, stating that an identifier having a weak prior in both the language model and the diffusion model is what the task requires. Huang already encodes its captions with a pretrained CLIP model and already projects content signals into extra tokens concatenated with the CLIP word embeddings, so a rare vocabulary token enters Huang's pipeline at an interface Huang has already built, and a person of ordinary skill would have expected the predictable result that the element binds to a token no other meaning competes for. Claim(s) 12-17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lianghua Huang et al. "Composer: Creative and Controllable Image Synthesis with Composable Conditions", 22 Feb 2023, arXiv, 2302.09778v2 in view of Huberman-Spiegelglas et al. “An Edit Friendly DDPM Noise Space: Inversion and Manipulations”, 12 Apr 2023, arXiv, 2304.06140v1 and Yu et al. (U.S. Pub. No 20240386623). Regarding claim 12, Huang discloses extracting a composition input from the composition image, wherein the composition input represents the target composition (sec 2.2, “Depthmap: We use a pretrained monocular depth estimation model (Ranftl et al., 2022) to extract the depthmap of an image, which roughly captures the image’s layout.”; also, sec 3.3, “We use the object’s segmentation map to represent its pose and the image embedding to capture its semantics, then leverage the reconfiguration approach described in Section 2.1 to modify the pose of the object (Figure 5d)”); generating a content input based on the composition image, wherein the content input indicates an image element (sec 3.2, “Variations: Using Composer, we can create new images that are similar to a given image but vary in certain aspects by conditioning on a specific subset of its representations.”; also, sec 3.2, “Specifically, given an image x, we can obtain its latent xT by applying DDIM inversion conditioned on a set of its representations ci; we then apply DDIM sampling starting from xT conditioned on a modified set of representations cj to obtain a variant of the image ˆx.”; also, sec 2.1, “Semantics and style: We use the image embedding extracted by the pretrained CLIP ViT-L/14@336px (Radford et al., 2021) model to represent the semantics and style of an image, similar to unclip”; also, sec 3.3, “Pose transfer: The CLIP embedding of an image captures its style and semantics, enabling Composer to modify the pose of an object without compromising its identity.”); and composition input and the content input indicating the image element (sec 2.3, “Localized conditioning: For localized representations in cluding sketches, segmentation masks, depthmaps, intensity images, and masked images, we project them into uniform dimensional embeddings with the same spatial size as the noisy latent xt using stacked convolutional layers. We then compute the sum of these embeddings and concatenate the result to xt before feeding it into the UNet.”; also, sec 2.3, “In addition, we project image embeddings and color palettes into eight extra tokens and concatenate them with CLIP word embeddings, which are then used as the context for cross-attention in GLIDE, similar to unclip”), wherein the synthetic image depicts the image element with the target composition (sec 3.3, “Pose transfer: The CLIP embedding of an image captures its style and semantics, enabling Composer to modify the pose of an object without compromising its identity. We use the object’s segmentation map to represent its pose and the image embedding to capture its semantics, then leverage the reconfiguration approach described in Section 2.1 to modify the pose of the object (Figure 5d).”). Huang does not disclose a non-transitory computer readable medium storing code for image processing, the code comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising: obtaining a composition image with a target composition, and generating, using an image generation model, a synthetic image by denoising a noise map. However, in a similar field of endeavor, Huberman-Spiegelglas discloses obtaining a composition image with a target composition, and generating, using an image generation model, a synthetic image by denoising a noise map (sec 5, “Suppose we are given a real image x0, a text prompt describing it psrc, and a target text prompt ptar. To modify the image according to these prompts, we extract the edit-friendly noise maps {xT,zT,...,z1}, while injecting psrc to the denoiser. We then fix those noise maps and generate an image while in jecting ptar to the denoiser.”; also, sec 1, “For example, one property we want from an inversion in the context of text-conditional models, is that fixing the noise maps and changing the text-prompt would lead to an artifact-free image, where the semantics correspond to the new text but the structure remains similar to that of the input image”; sec 5, “We then fix those noise maps and generate an image while in jecting ptar to the denoiser.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Huang's invention of extracting a structural representation and a content embedding from one image and conditioning a generative network on both, with the features of Huberman-Spiegelglas's invention of being given a real image and a text prompt, extracting noise maps, and then fixing those noise maps and generating an image by injecting a prompt to the denoiser. The combination would have been obvious because Huang extracts its representations from images its own pipeline already holds and never states either how the source image is supplied or the starting point from which its network generates, while Huberman-Spiegelglas states both for the same class of model, being given a real image whose structure the generated image is to follow, and holding a set of noise maps fixed and generating the image from them through the denoiser, and a person of ordinary skill would have expected the predictable result that the conditioned denoising proceeds from those noise maps to the finished image. Yu discloses a non-transitory computer readable medium storing code for image processing, the code comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations (para 50, “In some examples, memory 720 may include non-transitory, tangible, machine readable media that includes executable code that when run by one or more processors (e.g., processor 710) may cause the one or more processors to perform the methods described in further detail herein.”; also, para 50, “For example, as shown, memory 720 includes instructions for image generation module 730 that may be used to implement and/or emulate the systems and models, and/or to implement any of the methods described further herein.”; also, para 82, “One or more of the processes of method 900 may be implemented, at least in part, in the form of executable code stored on non-transitory, tangible, machine-readable media that when run by one or more processors may cause the one or more processors to perform one or more of the processes”; also, para 48, “Some common forms of machine-readable media may include floppy disk, flexible disk, hard disk, magnetic tape, any other magnetic medium, CD-ROM, any other optical medium, punch cards, paper tape, any other physical medium with patterns of holes, RAM, PROM, EPROM, FLASH-EPROM, any other memory chip or cartridge, and/or any other medium from which a processor or computer is adapted to read.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified Huang in view of Huberman-Spiegelglas, in which a structural representation and a content embedding extracted from one image condition the generation of an image from noise maps held fixed, with the features of Yu's invention of executable code stored on a non-transitory, tangible, machine-readable medium that, when run by one or more processors, causes those processors to perform the image generation processes. The combination would have been obvious because Huang describes a trained multibillion parameter generative network but never states where the code that runs it resides, and Yu addresses the same problem in the same field, generating images with a denoising diffusion model conditioned on a text prompt and on an input conditioning image such as a sketch or a depth map, and states expressly that its image generation module is held as instructions in memory and stored as executable code on a machine-readable medium from which a processor or computer is adapted to read. A person of ordinary skill implementing Huang's pipeline as a distributable product would have stored it exactly as Yu describes, with the predictable result that the pipeline runs wherever the medium is loaded. Regarding claim 13, Huang as modified by Huberman-Spiegelglas and Yu discloses the non-transitory computer readable medium of claim 12, the code further comprising instructions, when executed by the at least one processor, cause the at least one processor to perform operations where Huang further comprising: encoding the composition input to obtain a composition embedding, wherein the synthetic image is generated based on the composition embedding (sec 2.3, “Localized conditioning: For localized representations in cluding sketches, segmentation masks, depthmaps, intensity images, and masked images, we project them into uniform dimensional embeddings with the same spatial size as the noisy latent xt using stacked convolutional layers. We then compute the sum of these embeddings and concatenate the result to xt before feeding it into the UNet. Since the embeddings are additive, it is easy to accommodate for missing conditions or to incorporate new localized conditions.”). Regarding claim 14, Huang as modified by Huberman-Spiegelglas and Yu discloses the non-transitory computer readable medium of claim 12, wherein Huang further discloses: the composition input comprises a depth map, an edge map, pose information, layout information, or any combination thereof (sec 2.2, “Depthmap: We use a pretrained monocular depth estimation model (Ranftl et al., 2022) to extract the depthmap of an image, which roughly captures the image’s layout.”; also, sec 3.3, “We use the object’s segmentation map to represent its pose and the image embedding to capture its semantics, then leverage the reconfiguration approach described in Section 2.1 to modify the pose of the object (Figure 5d)”). Regarding claim 15, Huang as modified by Huberman-Spiegelglas and Yu discloses the non-transitory computer readable medium of claim 12, wherein Huang further discloses: the image element comprises a style attribute, an identity of an object, a lighting attribute, a texture attribute, a scene attribute, or a combination thereof (sec 2.2, “Semantics and style: We use the image embedding extracted by the pretrained CLIP ViT-L/14@336px (Radford et al., 2021) model to represent the semantics and style of an image, similar to unCLIP (Ramesh et al., 2022).”; also, sec 3.3, “Style transfer: Composer roughly disentangles the content and style representations, which allows us to transfer the style of image x1 to another image x2 by simply conditioning on the style representations of x1 and the content representations of x2.”). Regarding claim 16, Huang as modified by Huberman-Spiegelglas and Yu discloses the non-transitory computer readable medium of claim 12, wherein Huang further discloses: the image generation model is trained to generate images depicting the image element (sec 2.3, “We use diffusion models to recompose images from a set of representations.”; also, sec 3.1, “For the base model, we pretrain it with 1M steps on the full dataset using only image embeddings as the condition, and then finetune the model on a subset of 60M examples (excluding LAION images with aesthetic scores below 7.0) from the original dataset for 200K steps with all conditions enabled.”; also, sec 2.3, “Joint training strategy: It is essential to devise a joint training strategy that enables the model to learn to decode images from a variety of combinations of conditions.”). Regarding claim 17, Huang discloses content input indicates an image element and the composition input indicates a target composition of the image element (sec 3.3, “Pose transfer: The CLIP embedding of an image captures its”; also, sec 3.3, “We use the object’s segmentation map to represent its pose and the image embedding to capture its semantics, then leverage the reconfiguration approach described in Section 2.1 to modify the pose of the object (Figure 5d).”); encoding the composition input to obtain a composition embedding representing the target composition (sec 2.3, “Localized conditioning: For localized representations in cluding sketches, segmentation masks, depthmaps, intensity images, and masked images, we project them into uniform dimensional embeddings with the same spatial size as the noisy latent xt using stacked convolutional layers.”; also, sec 3.2, “We then compute the sum of these embeddings and concatenate the result to xt before feeding it into the UNet. Since the embeddings are additive, it is easy to accommodate for missing conditions or to incorporate new localized conditions.”; also, sec 3.2, “We use the object’s segmentation map to represent its pose and the image embedding to capture its semantics, then leverage the reconfiguration approach described in Section 2.1 to modify the pose of the object (Figure 5d)”; also, sec 2.2, “Depthmap: We use a pretrained monocular depth estimation model (Ranftl et al., 2022) to extract the depthmap of an image, which roughly captures the image’s layout.”); content input indicating the image element (sec 3.3, “Pose transfer: The CLIP embedding of an image captures its style and semantics, enabling Composer to modify the pose of an object without compromising its identity. We use the object’s segmentation map to represent its pose and the image embedding to capture its semantics, then leverage the reconfiguration approach described in Section 2.1 to modify the pose of the object (Figure 5d)”; also, sec 2.3, “In addition, we project image embeddings and color palettes into eight extra tokens and concatenate them with CLIP word embeddings, which are then used as the context for cross-attention in GLIDE, similar to unCLIP (Ramesh et al., 2022).”) and the composition embedding (sec 2.3, “We then compute the sum of these embeddings and concatenate the result to xt before feeding it into the UNet. Since the embeddings are additive, it is easy to accommodate for missing conditions or to incorporate new localized conditions”; also, sec 2.3, “In addition, we project image embeddings and color palettes into eight extra tokens and concatenate them with CLIP word embeddings, which are then used as the context for cross-attention in GLIDE, similar to unCLIP (Ramesh et al., 2022).”), wherein the synthetic image depicts the image element with the target composition (sec 3.3, “Pose transfer: The CLIP embedding of an image captures its style and semantics, enabling Composer to modify the pose of an object without compromising its identity. We use the object’s segmentation map to represent its pose and the image embedding to capture its semantics, then leverage the reconfiguration approach described in Section 2.1 to modify the pose of the object (Figure 5d)”). Huang does not disclose an apparatus comprising: at least one processor; at least one memory including instructions executable by the at least one processor to perform operations comprising: obtaining a content input, a composition input, and a noise map and generating, using an image generation model, a synthetic image by denoising the noise map. However, in a similar field of endeavor, Huberman-Spiegelglas discloses obtaining a content input, a composition input, and a noise map (sec 5, “Suppose we are given a real image x0, a text prompt describing it psrc, and a target text prompt ptar. To modify the image according to these prompts, we extract the edit-friendly noise maps {xT,zT,...,z1}, while injecting psrc to the denoiser. We then fix those noise maps and generate an image while in jecting ptar to the denoiser.”; also, sec 1, “For example, one property we want from an inversion in the context of text-conditional models, is that fixing the noise maps and changing the text-prompt would lead to an artifact-free image, where the semantics correspond to the new text but the structure remains similar to that of the input image”) and generating, using an image generation model, a synthetic image by denoising the noise map (Sec 5, “We then fix those noise maps and generate an image while in jecting ptar to the denoiser.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Huang's invention of obtaining a content signal and a structural representation, projecting the structural representation into a composition embedding, and conditioning the generative network on both, with the features of Huberman-Spiegelglas's invention of being given a real image and a text prompt, extracting noise maps, and then fixing those noise maps and generating an image by injecting a prompt to the denoiser. The combination would have been obvious because Huang extracts its representations from images its own pipeline already holds and describes the thing being denoised only as a latent, so it states neither how the signals that drive a particular generation are supplied nor the starting point of that generation, while Huberman-Spiegelglas supplies both in the same field and for the same class of model, being given a real image and a text prompt describing it, and holding a set of noise maps fixed and generating the image from them through the denoiser under a target prompt. Huberman-Spiegelglas further states that fixing the noise maps and changing the text prompt yields an image whose semantics correspond to the new text while the structure remains similar to that of the input image, which is the same division of labor Huang draws between its content signal and its structural map. Huang already conditions its network on embeddings it projects from its representations, so the noise maps Huberman-Spiegelglas holds fixed and denoises enter Huang's pipeline at the same place Huang already feeds its noisy latent, and a person of ordinary skill would have expected the predictable result that the conditioned denoising proceeds from those noise maps to the finished image. Yu discloses an apparatus comprising: at least one processor; at least one memory including instructions executable by the at least one processor to perform operations (para 47, “As shown in FIG. 7A, computing device 700 includes a processor 710; also, para 50, “In some examples, memory 720 may include non-transitory, tangible, machine readable media that includes executable code that when run by one or more processors (e.g., processor 710) may cause the one or more processors to perform the methods described in further detail herein.”; also, para 50, “For example, as shown, memory 720 includes instructions for image generation module 730 that may be used to implement and/or emulate the systems and models, and/or to implement any of the methods described further herein.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified Huang in view of Huberman-Spiegelglas, in which a content signal and a composition embedding projected from a structural representation condition the generation of an image from noise maps held fixed, with the features of Yu's invention of a computing device having a processor coupled to a memory that holds the instructions for an image generation module which the processor executes to carry out the generation operations. The combination would have been obvious because Huang recites the operations and the network but never the machine that runs them, and Yu addresses the same problem in the same field, generating images with a denoising diffusion model conditioned on a text prompt and on an input conditioning image such as a sketch or a depth map, and states that its image generation module is held in memory as instructions executed by the coupled processor. A person of ordinary skill building Huang's pipeline into a working machine would have organized it exactly that way, with the predictable result that the processor performs the recited operations. Claim(s) 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lianghua Huang et al. "Composer: Creative and Controllable Image Synthesis with Composable Conditions", 22 Feb 2023, arXiv, 2302.09778v2 as modified by Huberman-Spiegelglas et al. “An Edit Friendly DDPM Noise Space: Inversion and Manipulations”, 12 Apr 2023, arXiv, 2304.06140v1 and Yu et al. (U.S. Pub. No. 20240386623), further in view of Zhang et al., “Adding Conditional Control to Text-to-Image Diffusion Models”, 26 Nov 2023, arXiv, 2302.05543v3. Regarding claim 20, Huang as modified by Huberman-Spiegelglas and Yu discloses the apparatus of claim 17, does not disclose wherein: the image generation model comprises a ControlNet. However, in a similar field of endeavor, Zhang discloses wherein: the image generation model comprises a ControlNet (sec 3.2, “In particular, we use ControlNet to create a trainable copy of the 12 encoding blocks and 1 middle block of Stable Diffusion.”; also, sec 3.2, “The outputs are added to the 12 skip-connections and 1 middle block of the U-net.”; also, sec 4.5, “Since ControlNets do not change the network topology of pretrained SD models, it can be directly applied to various models in the stable diffusion community, such as Comic Diffusion[61] and Pro togen 3.4[16], in Figure12”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified Huang in view of Huberman-Spiegelglas, further in view of Yu, in which a processor executes instructions held in memory to project a structural representation into a composition embedding and to generate an image from noise maps held fixed and conditioned on that embedding and on a content signal, with the features of Zhang's invention of a control branch built as a trainable copy of the generative network's encoding blocks whose outputs are added into that network's skip connections. The combination would have been obvious because Huang trains its conditioning modules jointly with a multi billion parameter base model, while Zhang identifies the cost that imposes and offers the alternative of copying the encoding blocks of an already pretrained generative model and leaving the original parameters locked, and states that because this arrangement does not change the network topology it can be applied directly to models already in circulation. A person of ordinary skill wanting Huang's conditioning behavior without retraining a base model would have adopted Zhang's copied and locked arrangement, with the predictable result that the generative network carries the control branch among its own blocks. Allowable Subject Matter Claim 3, 11, 18, and 19 objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter: Claim 3 requires that the composition embedding, which claim 1 defines as the product of encoding the composition input, be based on the content input. The closest art on the composition side encodes only the structural map, projecting sketches, segmentation masks, depthmaps, intensity images, and masked images into uniform dimensional embeddings with stacked convolutional layers, and it routes the content signal to the network by a separate path, adding it to the timestep embedding and concatenating projected tokens with the word embeddings used for cross attention. Its composition embedding therefore carries nothing derived from the content input. The closest art on the grounded generation side fuses the text feature of a named entity with that entity's bounding box coordinates through a multi-layer perceptron to form a grounding token, but the bounding box it encodes is a grounding spatial configuration supplied alongside the entity rather than a composition input extracted from or representing a composition image, and the grounding token it produces conditions the network through gated self-attention rather than serving as the encoding of a composition input. Reaching claim 3 from that art would require treating the bounding box as the composition input and the grounding token as the composition embedding, neither of which the art states. The art therefore reaches an embedding that combines text with spatial coordinates, but does not reach an embedding produced by encoding the composition input and made to depend on the content input. Claim 11 requires obtaining an adherence factor input that indicates a level of adherence to the composition input, and generating the synthetic image based on that input. The closest art discloses an edge guidance scale that the authors selected for inference and that may be modified based on the user requirement to balance between edge fidelity and realism. That quantity is a guidance parameter applied to a guidance signal, not an input obtained by the method, and the reference expresses it as a tradeoff between edge alignment and realism rather than as a level of adherence to a composition input. The next closest art discloses a weight for blending among plural composed adapters, leaves the single condition path unweighted, and states that the manual adjustment such weights require is a limitation to be eliminated, which teaches away from the claimed arrangement. The art therefore reaches a scalar that shifts a generated image along a fidelity axis, but does not reach an adherence factor obtained as an input and indicating a level of adherence to the composition input. Claim 18 requires a text encoder configured to generate a text embedding from the content input, that is, a component of the apparatus whose output is a text embedding and whose input is the content input that indicates the image element. The closest art on the composition side represents its captions with sentence and word embeddings extracted by a pretrained multimodal model, but it names no text encoder as a component of anything and draws those captions from an image-text training corpus rather than from a content input supplied for a particular generation. The closest art on the control branch side names a text encoder and states that text prompts are encoded using it, but never identifies the output of that encoder as an embedding, the word appearing nowhere in that reference outside its discussion of the work of others. The closest art on the personalization side names a text encoder and a text prompt together, but states that the encoder produces a conditioning vector rather than an embedding, and does so in a passage the authors expressly introduce as background on text-to-image diffusion models rather than as a description of their own system. A search of United States patent documents returned no document filed early enough to have published before the effective filing date that contains a text encoder, a text embedding, and a diffusion model together, and no United States claim set reciting both a text encoder and a text embedding. The art therefore reaches a text encoder that encodes a prompt, and separately reaches text embeddings as a quantity, but does not reach a text encoder recited as a component that generates a text embedding from the content input. Claim 19 requires that the image generation model be trained to generate images having a plurality of different image elements based on a plurality of different nonce tokens, respectively, that is, one trained model in which several distinct rare tokens each recall a different image element. The closest art that reaches a rare token at all binds one such token to one element per finetuned model. Its collection of thirty subjects is an evaluation corpus rather than a single model's vocabulary, its ablations report finetuning across fifteen subjects and across five subjects as separate runs, and its description of the effect of the number of training images states that separate models were trained for each subject. The closest art on the composition side conditions generation on continuous embeddings projected from whole images and from captions and discloses no identifier token of any kind, so it cannot supply a plurality of them. The art therefore reaches a single token to element binding but does not reach one model holding several such bindings at once, which is what the respective plurality of claim 19 requires. Response to Arguments Applicant’s arguments filed 07/01/2026 have been fully considered and are persuasive. Applicant argues at Remarks pages 12 to 14 that Zhang does not teach or disclose denoising the noise map based on the content input indicating the image element and the composition embedding, as recited in amended claims 1, 12, and 17. This argument is persuasive as to Zhang. Zhang has no detailed description sentence reciting a noise map that is obtained at generation time and denoised, so it cannot supply the denoising of a noise map, and Zhang has no personalization of any kind, so it cannot supply a content input that indicates a particular image element. The rejection of claims 1, 12, and 17 over Zhang as previously stated has therefore been withdrawn, Zhang being retained only for the ControlNet of claim 20. Upon further consideration, claims 1, 6, 7, 8, 9, and 10 are newly rejected under 35 U.S.C. 103 over Huang in view of Huberman-Spiegelglas, and claims 12, 13, 14, 15, 16, 17, and 18 are newly rejected under 35 U.S.C. 103 over Huang in view of Huberman-Spiegelglas, further in view of Yu, both as set forth above. Huang discloses a content input that is an image embedding representing the semantics and style of an image, a composition input that is a depthmap capturing the image's layout or a segmentation map representing an object's pose, the projection of that structural representation into an embedding using stacked convolutional layers, the summing of those embeddings and their concatenation to the noisy latent before the network, and an output image depicting the same object with a modified pose without compromising its identity, and Huberman-Spiegelglas supplies the noise maps that are held fixed and denoised to generate the image. Applicant argues at Remarks page 15 that Ruiz only teaches that a text prompt and a noise input are fed to a diffusion model and is silent about generating a synthetic image by denoising the noise map based on the content input and the composition embedding. This argument has been considered but is moot because it does not apply to the new combination of references being used in the current rejection. Ruiz is now relied upon only for the descriptive text prompt naming the subject of claim 2 and the nonce token of claim 4, and the noise map is now supplied by Huberman-Spiegelglas, which extracts noise maps, fixes them, and generates an image from them through the denoiser. Applicant argues at Remarks pages 15 to 17 that the conditioning vector cited from Zhang for claim 11 is the composition input itself and not a separate input indicating a level of adherence to the composition input. This argument is persuasive. The rejection of claim 11 as previously stated has been withdrawn. Upon further consideration a new ground of rejection of claim 11 under 35 U.S.C. 103 over Huang in view of Huberman-Spiegelglas, further in view of Voynov is made as set forth above, Voynov reciting an edge guidance scale that the user may modify to balance between edge fidelity and realism and on whose value the generated image depends. Applicant argues at Remarks page 17 that claims 2-4, 6-11, 13-16, and 18-20 are allowable by virtue of depending from a patentable independent claim. This argument is not persuasive as to claims 2, 3, 4, 6-11, 13-16, 18, and 20, because the independent claims are not allowable for the reasons given above. New grounds of rejection are made above for claim 2 over Huang in view of Huberman-Spiegelglas, further in view of Ruiz, for claim 3 over Huang in view of Huberman-Spiegelglas, further in view of Li, for claim 4 over Huang in view of Huberman-Spiegelglas, further in view of Ruiz, and for claim 20 over Huang in view of Huberman-Spiegelglas, further in view of Yu and Zhang. Claim 19 is treated as allowable subject matter above. Conclusion The following prior art made of record and not relied upon is considered pertinent to applicant’s disclosure: Andrey Voynov et al. " Sketch-Guided Text-to-Image Diffusion Models", 24 Nov 2022, arXiv, 2211.13752v1, Yuheng Li et al, “GLIGEN: Open-Set Grounded Text-to-Image Generation”, 17, Apr 2023, arXiv, 2301.07093v2, Chong Molu et al, T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models”, 20, Mar 2023, arXiv, 2302.08453v2, Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jai Li whose telephone number is (571)272-1170. The examiner can normally be reached Mon-Thu between 06:00-16:00 EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Xiao Wu can be reached at (571)272-7761. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JAI W LI/Junior Examiner, Art Unit 2613 /XIAO M WU/Supervisory Patent Examiner, Art Unit 2613
Read full office action

Prosecution Timeline

Oct 07, 2024
Application Filed
Apr 01, 2026
Non-Final Rejection mailed — §103
Jun 22, 2026
Interview Requested
Jun 29, 2026
Applicant Interview (Telephonic)
Jun 29, 2026
Examiner Interview Summary
Jul 01, 2026
Response Filed
Aug 07, 2026
Non-Final Rejection mailed — §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

2-3
Expected OA Rounds
Grant Probability
Moderate
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month