DETAILED ACTION
DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1 is/are rejected under 35 U.S.C. 102(a)(1) as being unpatentable by Wu et al. (“DiffuMask: Synthesizing Images with Pixel-level Annotations for Semantic Segmentation Using Diffusion Models”).
Regarding claim 1, Wu teaches:
A method for generating a synthetic image and a semantic segmentation mask thereof, comprising:
generating, by a language model, a text prompt, wherein the text prompt comprises a caption and class labels; (page 1211, left bottome: “
PNG
media_image1.png
290
588
media_image1.png
Greyscale
”FIG. 4) and
generating, by a diffusion model, the synthetic image and the semantic segmentation mask thereof based on the text prompt.(page 1208, right, under 3. Methodology: “In this paper, we explore simultaneously generating images and the semantic mask described in the text prompt with the existing pre-trained diffusion model.” FIG. 4)
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 2-5, 12-14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wu in view of Rombach et al. (“High-Resolution Image Synthesis with Latent Diffusion Models”) and further in view of Borse et al. (US 2025/0131606 A1).
Regarding claim 2, Wu teaches:
The method of claim 1, wherein the generating, by the diffusion model, the synthetic image and the semantic segmentation mask thereof based on the text prompt comprises:
encoding, by a text encoder, the text prompt into a text embedding; (FIG. 4, the Text Encoder encodes the sampled prompt) and
diffusion model of Stable Diffusion model (page 1208, right, bottom. )
However, Wu does not, but Rombach teaches the details of the Stable Diffusion model
outputting, by the diffusion model, a final latent state reflecting a content encoded in the text embedding from an initial latent state after a predetermined number of denoising steps, wherein the diffusion model comprises a predetermined number of …and cross-attention layers, at each denoising step, the … and cross-attention layers transform a latent state of a current step to a latent state of a next step.(page 10687, FIG. 3: “
PNG
media_image2.png
310
588
media_image2.png
Greyscale
”
And corresponding paragraphs.)
Wu teaches using a Stable Diffusion model is used to generate the images and semantic mask. Rombach teaches the details of the Stable Diffusion model.
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to have combined the teachings of Wu with the specific teachings of Rombach to accurately and efficiently generate the semantic mask.
However, Wu in view of Rombach does not explicitly, but Borse teaches:
self-attention ([0063], “UNet architectures in current state-of-the-art (SOTA) models, for example, Stable Diffusion and variants, adopt a combination of attention layers and convolutional layers at each stage of the UNet. Current UNet architectures adopt global self-attention and cross-attention operations at all spatial resolutions”)
Wu teaches using Stable Diffusion model. Rombach teaches the Stable diffusion model includes a U-net. Borse teaches the U-net includes self-attention and cross-attention layers.
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to have combined the teachings of Wu in view of Rombach with the specific teachings of Borse to build a detailed stable diffusion model to accurately and efficiently generate the semantic mask.
Regarding claim 3, Wu in view of Rombach and Borse teaches:
The method of claim 2, wherein the initial latent state followsN(0,1), which is a standard Gaussian distribution.(Wu, page 1208, right bottom: “
PNG
media_image3.png
144
578
media_image3.png
Greyscale
”)
Regarding claim 4, Wu in view of Rombach and Borse teaches:
The method of claim 3, wherein the transforming, at each denoising step, the latent state of the current step to the latent state of the next step comprises: generating, by each self-attention layer, a self-attention map for capturing a pairwise similarity between positions within a latent state of the current layer in order to enhance a local feature with a global context in a latent state of next layer; and generating, by each cross-attention layer, a cross-attention map for modeling the relationship between each position of the latent state of the current layer and each token of the text embedding so that the latent state of the next layer expresses more of the content encoded in the text embedding.(Borse ([0063], “UNet architectures in current state-of-the-art (SOTA) models, for example, Stable Diffusion and variants, adopt a combination of attention layers and convolutional layers at each stage of the UNet. Current UNet architectures adopt global self-attention and cross-attention operations at all spatial resolutions” the combination of claim 2 is incorporated here.)
Regarding claim 5, Wu in view of Rombach and Borse teaches:
The method of claim 4, wherein the generating, by each cross-attention layer, the cross-attention map, comprises: generating a class-specific text prompt based on each class label; (Wu section 3.4 teaches generate a prompt for each class and FIG. 4) and generating the cross-attention map based on the class-specific text prompts.( Wu Section 3.1: “Text-guided generative models (e.g., Imagen [53], Stable Diffusion [49]) use a text prompt P to guide the content related image I generation from a random gaussian image noise z, where visual and textual embedding are fused using the spatial cross-attention. Specifically, Stable Diffusion [49] consists of a text encoder, a variational autoencoder (VAE), and a U-shaped network [50]. The interaction between the text and vision occurs in the U-Net for the latent vectors at each time step, where cross-attention layers are used to fuse the embeddings of the visual and textual features and produce spatial attention maps for each textual token.”)
Claims 12-14 recites similar limitations of claim 2, 4-5 respectively, thus are rejected accordingly.
Claim(s) 11, 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wu.
Regarding claim 11, Wu teaches:
A system for generating a synthetic image and a semantic segmentation mask thereof comprising: one or more processors; and a computer-readable medium having instructions stored there on, which, when executed by the one or more processors, cause the system to perform operations comprising (Wu teaches using a system to do tests on the method. It would be obvious to the system to have computer and memory to run the method instruction to perform the method.) The rest of claim 11 recites similar limitations of claim 1, thus are rejected accordingly.
Claims 20 recites similar limitations of claim 11, thus are rejected accordingly.
Claim(s) 6, 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wu in view of Rombach and Borse and further in view of Campagnolo et al. (US 2024/0354991 A1).
Regarding claim 6, Wu in view of Rombach and Borse teaches:
The method of claim 5, further comprising averaging the cross-attention maps over layers and steps to generate an average cross-attention map; (Wu FIG. 2 and page 1207, right: “visualization of cross attention map between text token and vision. 8 × 8, 16 × 16, 32 × 32, and 64 × 64, as four different resolutions, are extracted from different layers of the U-Net of Stable Diffusion [49]. 8×8 feature map is the lowest resolution, including obvious class-discriminative location. 32 × 32 and 64 × 64 feature maps include high resolution and highlight fine-grained details. The average map shows the possibility for us to use for semantic segmentation, where it is class-discriminative and fine-grained.”) averaging the self-attention maps over layers and steps to generate an average self-attention map;(Borse teaches Unet has both cross-attention layers and self-attention layers. Combining with the teachings of Wu to average the self-attention output maps to generate fine-grained results.) and
However, Wu in view of Rombach and Borse does not, but Campagnolo teaches
enhancing the average cross-attention map based on the average self-attention map.(FIG. 3, [0033], “Regarding outputs, the self-attention network 308 outputs are combined with the cross-attention network 306 for forming mean and standard deviation-related pairs about the conditioned latent representations.”)
Wu in view of Rombach and Borse teaches Unet has both cross-attention layers and self-attention layers. Campagnolo teaches the output of cross-attention layers output are enhanced based on the self-attention map.
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to have combined the teachings of Wu in view of Rombach and Borse with the specific teachings of Campagnolo to build a fine-grained results.
Claims 15 recites similar limitations of claim 6, thus are rejected accordingly.
Allowable Subject Matter
Claims 7-10, 16-19 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
The following is a statement of reasons for the indication of allowable subject matter: none of the references along or in combination teaches the limitations of “wherein the enhancing the average cross-attention map based on the average self-attention map comprises: powering the average self-attention map to a predetermined exponent to generate a powered self-attention map; and multiplying the powered self-attention map to the average cross-attention map to generate an enhanced cross-attention map.” Recited in claim 7 and similarly recited in claim 16.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to YANNA WU whose telephone number is (571)270-0725. The examiner can normally be reached Monday-Thursday 8:00-5:30 ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alicia Harrington can be reached at 5712722330. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/YANNA WU/Primary Examiner, Art Unit 2615