Prosecution Insights
Last updated: October 01, 2026
Application No. 18/956,138

JOINT INTRINSIC LAYERS FROM LATENT DIFFUSION MODELS

Non-Final OA §103
Filed
Nov 22, 2024
Examiner
WU, YANNA
Art Unit
2615
Tech Center
2600 — Communications
Assignee
Adobe Inc.
OA Round
1 (Non-Final)
81%
Grant Probability
Favorable
1-2
OA Rounds
3m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 81% — above average
81%
Career Allowance Rate
369 granted / 456 resolved
+18.9% vs TC avg
Strong +34% interview lift
Without
With
+34.4%
Interview Lift
resolved cases with interview
Fast prosecutor
2y 2m
Avg Prosecution
23 currently pending
Career history
474
Total Applications
across all art units

Statute-Specific Performance

§101
9.8%
-30.2% vs TC avg
§103
69.7%
+29.7% vs TC avg
§102
6.8%
-33.2% vs TC avg
§112
8.1%
-31.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 456 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Comments The response to the restriction requirement has been acknowledged. Applicant elected species I (Claims 1-6, and 15-20) without traverse. Applicant is reminded to cancel the non-elected claims. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-4, 6, 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Aggarwal et al. (US 2022/0343561 A1). Regarding claim 1, Aggarwal teaches: A method for image processing, comprising: obtaining an input image depicting a scene and an input prompt indicating an intrinsic modality of the input image, (FIG. 2, provide an image 200. Provide a source color input text 205. [0039], “The user may say a color to segment the portions and then use a color text (i.e., basic, complex or specific colors) to replace the segmented regions.”) encoding, using a conditional image encoder, the input image to obtain a condition embedding representing the intrinsic modality of the input image; ([0037], “At operation 210, the system segments the colors in the image. The color segmentation is performed by extracting color embeddings for the unique pixels in an image using the color pixel encoder.”) and generating, using an image generation model, a synthetic output based on the input prompt and the condition embedding, wherein the synthetic output comprises a visual representation of the intrinsic modality of the input image.(FIG. 2, [0039], “At operation 220, the system replaces the source color with the target color to create an adjusted image. Different lighting and shadows in the images are preserved when the hue part of a pixel's hue, saturation, and lightness (HSL) value is replaced. Some embodiments of the present disclosure are used for style editing for real-world images where distinct colors are present. The user may say a color to segment the portions and then use a color text (i.e., basic, complex or specific colors) to replace the segmented regions.”) On the different embodiment/example, Aggarwal teaches: wherein the intrinsic modality determines how the scene interacts with light; ([0040], “In some embodiments, when replacing a color, the hue dimension may be replaced, while retaining variations in shades and lightness of a color in the masked portion of the image. For example, a user may be provided with controls to adjust portions of the image based on color dominance and control the saturation (shade) and lightness of the replacing colors.”) It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to have combined the different embodiments/examples of Aggarwal to allow the method to be used to replace a variety of information of an image. Regarding claim 2, Aggarwal teaches: The method of claim 1, further comprising: encoding the input prompt to obtain a text embedding, wherein the synthetic output is generated based on the text embedding. ([0036], “At operation 205, the user provides a speech or text input with a source color. The speech input is provided to a multi-lingual text encoder to convert text into a color embedding.”) Regarding claim 3, Aggarwal teaches: The method of claim 1, wherein encoding the input image comprises: encoding the input image to obtain a plurality of condition embeddings corresponding to a plurality of intrinsic modalities; and selecting the condition embedding from the plurality of condition embeddings based on the input prompt. ([0037], “At operation 210, the system segments the colors in the image. The color segmentation is performed by extracting color embeddings for the unique pixels in an image using the color pixel encoder.” [0044], “Segmented image 305 is an intermediate image produced by a color replacement system of the present disclosure. In the example scenario of FIG. 3, the segmented image 305 is segmented into two regions; light and dark regions. The light regions have been determined to not be a target color. The dark regions have been determined to be a target color. Therefore, the dark region will be replaced with a source color. In some examples, an image segmentation mask may be presented to a user to make it more clear which portions of the image will be replaced with another color.”) Regarding claim 4, Aggarwal teaches: The method of claim 1, wherein generating the synthetic output comprises: generating control guidance based on the condition embedding; and providing the control guidance to the image generation model.([0038], “At operation 215, the user provides a speech or text input with a target color. The speech input is provided to a multi-lingual text encoder to convert text into a color embedding.” [0039], “The user may say a color to segment the portions and then use a color text (i.e., basic, complex or specific colors) to replace the segmented regions. Some embodiments of the present disclosure are used to do palette mapping (i.e., map multiple painting colors to a different set of colors and transfer the original image according to color texts provided by a user). A user may adjust the saturation and lightness of the replaced color regions as the hue part of a color is replaced.”) Regarding claim 6, Aggarwal teaches: The method of claim 1, wherein: the intrinsic modality is selected from a set comprising at least one of a shading modality, an albedo modality, and a normal modality. ([0040], “In some embodiments, when replacing a color, the hue dimension may be replaced, while retaining variations in shades and lightness of a color in the masked portion of the image. For example, a user may be provided with controls to adjust portions of the image based on color dominance and control the saturation (shade) and lightness of the replacing colors.”) Regarding claim 15, Aggarwal teaches: A system comprising: a memory component; and a processing device coupled to the memory component, the processing device configured to perform operations comprising: (fig. 4) The rest of claim 15 recites similar limitations of claim 1, thus is rejected accordingly. Claim(s) 5 is/are rejected under 35 U.S.C. 103 as being unpatentable over Aggarwal in view of Smock et al. (US 2024/0331235 A1). Regarding claim 5, Aggarwal teaches: The method of claim 1, wherein generating the synthetic output comprises: However, Aggarwal does not, but Smock teaches: obtaining a noise map; and denoising the noise map based on the condition embedding to generate the synthetic output. ([0022], “Stable Diffusion consists of three parts: a variational autoencoder (VAE), U-Net, and an optional text encoder. The VAE encoder compresses the image from pixel space to a smaller dimensional latent space, capturing a more fundamental semantic meaning of the image. Gaussian noise is iteratively applied to the compressed latent representation during forward diffusion. The U-Net block, composed of a ResNet backbone, denoises the output from forward diffusion backwards to obtain a latent representation. Finally, the VAE decoder generates the final image by converting the representation back into pixel space. The denoising step can be flexibly conditioned on a string of text, an image, or another modality. The encoded conditioning data is exposed to denoising U-Nets via a cross-attention mechanism.”) Aggarwal and Smock both teach a method of generating an output image based on the input image and text prompt. Smock teaches more specifically using a machine learning model in the process, where noise and denoise process are used in the machine learning model. It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to have combined the teachings of Aggarwal with the specific teachings of Smock to accurately generate the output image based on input prompt. Claim(s) 16-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Aggarwal in view of Zhang et al. (“Adding Conditional Control to Text-to-Image Diffusion Models” from IDS). Regarding claim 16, Aggarwal teaches: The system of claim 15, wherein: However, Aggarwal does not, but Zhang teaches: the conditional image encoder comprises a plurality of output blocks corresponding to a plurality of intrinsic modalities, respectively. (FIG. 3 shows the image encoder has encoder block A, B, C.”) Aggarwal and Zhang both teach a method of generating an output image based on the input image and text prompt and encoding input images. Zhang teaches a specific encoding model using a diffusion model, where encoding image into several blocks. It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to have combined the teachings of Aggarwal with the specific teachings of Zhang to use a specific diffusion model to accurately generate an output image based on input prompt. Regarding claim 17, Aggarwal in view of Zhang teaches: The system of claim 16, wherein: the conditional image encoder comprises a preliminary encoder trained to provide a preliminary condition embedding to each of the plurality of output blocks. (Zhang, FIG. 3, page 5, top left: “ PNG media_image1.png 210 330 media_image1.png Greyscale ” The combination rationale of claim 16 is incorporated here.) Regarding claim 18, Aggarwal in view of Zhang teaches: The system of claim 16, wherein: each of the plurality of output blocks comprises a transformer network. (Zhang, FIG. 3 shows the image encoder has encoder block A, B, C.”, each encoder corresponds to a transformer network. The combination rationale of claim 16 is incorporated here.) Regarding claim 19, Aggarwal teaches: The system of claim 15, wherein: However, Aggarwal does not, but Zhang teaches: the image generation model comprises a diffusion network. (FIG. 3, stable diffusion network.) Aggarwal and Zhang both teach a method of generating an output image based on the input image and text prompt. Zhang teaches more specifically using a diffusion model in the process. It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to have combined the teachings of Aggarwal with the specific teachings of Zhang to accurately generate the output image based on input prompt. Regarding claim 20, Aggarwal teaches: The system of claim 15, However, Aggarwal does not, but Zhang teaches: wherein: the image generation model includes a control network. (FIG. 3, b controlNet) Aggarwal and Zhang both teach a method of generating an output image based on the input image and text prompt. Zhang teaches more specifically using a control network in the process. It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to have combined the teachings of Aggarwal with the specific teachings of Zhang to accurately generate the output image based on input prompt. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to YANNA WU whose telephone number is (571)270-0725. The examiner can normally be reached Monday-Thursday 8:00-5:30 ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alicia Harrington can be reached at 5712722330. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /YANNA WU/Primary Examiner, Art Unit 2615
Read full office action

Prosecution Timeline

Nov 22, 2024
Application Filed
Jul 02, 2026
Non-Final Rejection mailed — §103
Sep 24, 2026
Applicant Interview (Telephonic)
Sep 24, 2026
Examiner Interview Summary

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743561
Generating Technical Drawings From Building Information Models
2y 0m to grant Granted Sep 22, 2026
Patent 12724582
SYSTEMS, APPARATUSES, AND METHODS FOR REAL-TIME COLLABORATION WITH A GRAPHICAL RENDERING PROGRAM
3y 0m to grant Granted Sep 01, 2026
Patent 12725352
INFORMATION PROCESSING DEVICE AND INFORMATION PROCESSING METHOD
2y 1m to grant Granted Sep 01, 2026
Patent 12718475
SITE MODEL UPDATING METHOD AND SYSTEM
3y 2m to grant Granted Aug 25, 2026
Patent 12711690
SPECULATIVE EXECUTION OF HIT AND INTERSECTION SHADERS ON PROGRAMMABLE RAY TRACING ARCHITECTURES
1y 10m to grant Granted Aug 18, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
81%
Grant Probability
99%
With Interview (+34.4%)
2y 2m (~3m remaining)
Median Time to Grant
Low
PTA Risk
Based on 456 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month