Prosecution Insights
Last updated: October 01, 2026
Application No. 18/812,113

NEURAL ARCHITECTURE SEARCH FOR IMAGE GENERATION MODELS

Non-Final OA §103
Filed
Aug 22, 2024
Examiner
SUN, HAI TAO
Art Unit
2616
Tech Center
2600 — Communications
Assignee
Adobe Inc.
OA Round
1 (Non-Final)
74%
Grant Probability
Favorable
1-2
OA Rounds
5m
Est. Remaining
98%
With Interview

Examiner Intelligence

Grants 74% — above average
74%
Career Allowance Rate
363 granted / 493 resolved
+11.6% vs TC avg
Strong +25% interview lift
Without
With
+24.7%
Interview Lift
resolved cases with interview
Typical timeline
2y 6m
Avg Prosecution
42 currently pending
Career history
529
Total Applications
across all art units

Statute-Specific Performance

§101
7.1%
-32.9% vs TC avg
§103
68.6%
+28.6% vs TC avg
§102
1.4%
-38.6% vs TC avg
§112
16.2%
-23.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 493 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Election/Restrictions The applicant elected Group I encompassing claims 1-5. Additionally, Applicant has canceled claims 6-20 corresponding to the non-elected claims, and add new claims 21-32. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-5 and 21-32 are rejected under 35 U.S.C. 103 as being unpatentable over Cui (US 20240346799 A1) in view of Park (US 20230360169 A1), and further in view of Franco (US 20250045930 A1). Regarding to claim 1 (Original), Cui discloses a method (Fig. 1; [0031]: the electronic device 110 inputs the target image 101 into the image segmentation model 112; the image segmentation model 112 outputs an image segmentation result 102 corresponding to the target image 101; [0032]: the image segmentation result 102 may be a segmentation map, a text sequence, an attention map, etc. associated with the target image 101; [0033]: the image segmentation model 112 is any neural network that performs image segmentation; [0092]: the electronic device 110 extracts an image feature representation of a target image using a trained image encoder; [0093]: the electronic device 110 generates, using a trained text encoder, a text feature representation corresponding to a name of the class, and determines a candidate segmentation map for the target image) comprising: obtaining a target quality level and an input prompt describing an image element ([0025]: receive the user's active request; Fig. 1; [0051]: receive the target image 301, and obtain the image feature representation 302; the target image 301 contains quality level, such as, resolution; Fig. 3; [0056]: obtain and receive texts of the name of class, such as bicycle, animal, shoes, and so on; PNG media_image1.png 216 474 media_image1.png Greyscale ; Fig. 4; [0058]: fill the name of class 402 into at least one prompt template respectively; for each class, the filling unit 420 fills the name of class 402 into the at least one prompt template through prompt engineering, to obtain the at least one text sequence 303 containing the name of class; PNG media_image2.png 414 342 media_image2.png Greyscale ; [0059]: the filling unit 420 fills the bicycle into the at least one prompt template to obtain a plurality of text sequences; [0074]: the image resolution of the upsampled attention map 404 is the same as that of the target image 301; the input target image contains quality information, such as resolution; [0101]: the electronic device 110 selects, from the plurality of classes, at least one class related to the target image); selecting an attention map size based on the target quality level (Fig. 4; [0072]: the plurality of image features 401 are L×L image features; obtain the attention map 404 with a size of L×L; [0073]: a size corresponding to the target image 301; the candidate segmentation map generation unit 450 may upsample the attention map 404 to a selected size corresponding to the target image 301 to obtain an upsampled attention map 404); generating, using an image generation model, an attention map having the attention map size selected based on the target quality level ([0067]: the determination unit 330 determines an attention map based on both of the feature representations; Fig. 4; [0072]: the plurality of image features 401 are L×L image features, the attention map determination unit 440 performs calculation on the text feature representation 304 and the L×L image features to obtain the attention map 404; [0073]: the candidate segmentation map generation unit 450 may upsample the attention map 404 to a size corresponding to the target image 301 to obtain an upsampled attention map 404; [0098]: upsampling the attention map to a size corresponding to the target image, to obtain an upsampled attention map; [0112]: upsample the attention map to a size corresponding to the target image to obtain an upsampled attention map); and Cui fails to explicitly disclose: generating, using the image generation model, a synthetic image based on the input prompt and the attention map, wherein the synthetic image depicts the image element with the target quality level. In same field of endeavor, Park teaches: generating, using the image generation model, an image based on the input prompt and the attention map, wherein the image depicts the image element with the target quality level ([0014]: generate a second image having a second quality that is higher quality than a first quality of the first image, the generating the second image being based on the attention map, the second feature information, and the fourth feature information; [0066]: generate the second image 102 by performing denoising for removing noise while maintaining fine edges and textures of the first image 101; the second image 102 has a higher quality than the first image 101; the second image 102 may have one or more characteristics or features that are improved when compared to the first image 101; Fig. 15; [0232]: the image processing apparatus 100 generates a second image, based on the attention map obtained in operation S1530, the second feature information obtained in operation S1540, and the fourth feature information obtained in operation S1550; [0233]: the image processing apparatus 100 generates the second image by performing upsampling, an AdaIN operation, a convolution operation, a channel-split SFT operation, etc. using the attention map, the second feature information, and the fourth feature information). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Cui to include generating, using the image generation model, an image based on the input prompt and the attention map, wherein the image depicts the image element with the target quality level as taught by Park. The motivation for doing so would have been to generate mapping between input data and output data by a neural network; to generate the second image 102 by performing denoising for removing noise while maintaining fine edges and textures of the first image 101; to generate a second image, based on the attention map obtained in operation S1530 as taught by Park in paragraphs [0004], [0066], and [0232]. Cui in view of Park fails to explicitly disclose image is synthetic image. In same field of endeavor, Franco teaches: image is synthetic image and generate a synthetic image (Fig. 2A; [0034]: produce segmentation masks and synthetic image as illustrated in Fig. 2A; PNG media_image3.png 228 632 media_image3.png Greyscale ; Fig. 10A; Fig. 10B; [0081]: generate masks for synthetic images of diverse styles; generate synthetic images as illustrated in Fig. 10A and Fig. 10B; PNG media_image4.png 180 662 media_image4.png Greyscale PNG media_image5.png 202 642 media_image5.png Greyscale ). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Cui in view of Park to include image is synthetic image and generate a synthetic image as taught by Franco. The motivation for doing so would have been to produce a cross-attention map for each token; to generating a label assignment for a segment of a segmentation mask; to associate a segment from a segment mask for the image with a token based on the token correspondence maps; to achieve this level of performance in a pure zero-shot manner without any language dependency or auxiliary images as taught by Franco in paragraphs [0061], [0063], [0070], and [0073]. Regarding to claim 2 (Original), Cui in view of Park and Franco discloses the method of claim 1, further comprising: obtaining performance information (Cui; [0029]: determine the performance, i.e. performance information, of the model; [0045]: extract an image feature representation of a target image; [0032]: the target image 101 may be any image in any format, any size, and any color; [0044]: the image segmentation performance of the trained image segmentation model is poor, i.e. performance information; the performance of the model, segmentation performance, and the size of image are performance information; [0052]: extract a plurality of image features 401 for a plurality of image blocks of an input target image 301), wherein the attention map size is selected based on the performance information (Cui; Fig. 4; [0052]: extract a plurality of image features 401 for a plurality of image blocks of an input target image 301; the size of the target image 301 is L×L, that is, the target image 301 contains L×L image blocks; [0072]: the plurality of image features 401 are L×L image features; obtain the attention map 404 with a size of L×L; [0073]: a size corresponding to the target image 301). Cui in view of Park and Franco further discloses obtaining performance information (Franco; [0043]: the iterative attention merger 154 reduces the set of proposed objects by merging similar aggregated attention maps; [0072]: measure segmentation performance of the example implementation; [0078]: implementations uses N=3 for a better system latency and performance tradeoff; [0079]: a range τ∈[0.9, 1.1] yields reasonable performance). Same motivation of claim 1 is applied here. Regarding to claim 3 (Original), Cui in view of Park and Franco discloses the method of claim 1, further comprising: selecting a subnet of a base image generation model with the selected attention map size based on the target quality level (Cui; [0048]: neural networks, such as a Unet architecture, a SegNet architecture, a Fully Convolutional Network (FCN), a Convolutional Neural Network (CNN), a Recurrent Neural Network (RNN), etc.), wherein the image generation model comprises the subnet of the base image generation model (Cui; [0048]: a specific type of model structure may be selected according to actual application requirements). Cui in view of Park and Franco further discloses selecting a subnet of a base image generation model with the selected attention map size based on the target quality level (Franco; [0031]: the U-Net architecture consists of a stack of modular blocks. 16 of them have two major components: a ResNet layer and a Transformer layer 172; the Transformer layer 172 uses two types of attention mechanisms, self-attention, and cross-attention; PNG media_image6.png 342 288 media_image6.png Greyscale ; [0058]: the description generator 160 provides the prompt as input to the diffusion model 170 with the input image 105; [0063]: there are a total 16 cross-attention layers in the diffusion model 170, resulting in 16 cross-attention tensors). Same motivation of claim 1 is applied here. Regarding to claim 4 (Original), Cui in view of Park and Franco discloses the method of claim 1, wherein selecting the attention map size comprises: selecting a number of tokens for a key object (Cui; [0051]: extract token-wise image features of the target image 301, where each token corresponds to an image block, and the image feature representation 302 are obtained by aggregating the image features token-by-token; [0061]: extract token-by-token text features of the text sequence, each token corresponding to one word; the text feature representation 304 are obtained by aggregating the text features token-by-token; [0062]: a plurality of text units in the text sequence 303 containing the “name of class” can be tokenized and converted into an embedding vector representation), wherein the attention map comprises a product of the key object and a query object (Cui; [0072]: the attention map determination unit 440 performs calculation on the text feature representation 304 and the L×L image features to obtain the attention map 404; perform an inner product on the text feature representation 304 and the L×L image features respectively; perform the inner product operation on the text feature representation 304 and the L×L image features, to obtain the attention map 404 with a size of L×L; [0075]: perform the inner product operation on the image feature representation 302 and the text feature representation 304 corresponding to the class). Cui in view of Park and Franco further discloses selecting a number of tokens for a key object (Franco; [0033]: where Q represents the number of tokenized words in the prompt; the cross-attention map 220 is associated with the token “person” and the cross-attention map 225 is associated with the token “vehicle”). Same motivation of claim 1 is applied here. Regarding to claim 5 (Original), Cui in view of Park and Franco discloses the method of claim 4, wherein selecting the attention map size (same as rejected in claim 4) comprises: computing a product of the attention map and the value object (Cui; [0051]: extract token-wise image features of the target image 301, where each token corresponds to an image block, and the image feature representation 302 may be obtained by aggregating the image features token-by-token; [0061]: extract token-by-token text features of the text sequence, each token corresponding to one word; the text feature representation 304 are obtained by aggregating the text features token-by-token; [0062]: a plurality of text units in the text sequence 303 containing the “name of class” can be tokenized and converted into an embedding vector representation; [0072]: the attention map determination unit 440 performs calculation on the text feature representation 304 and the L×L image features to obtain the attention map 404; perform an inner product on the text feature representation 304 and the L×L image features respectively; perform the inner product operation on the text feature representation 304 and the L×L image features, to obtain the attention map 404 with a size of L×L; [0075]: perform the inner product operation on the image feature representation 302 and the text feature representation 304 corresponding to the class). Cui in view of Park and Franco further discloses: selecting a number of tokens for a value object corresponding to the number of tokens for the key object (Franco; [0019]: identify a segment as associated with a particular token; [0033]: where Q represents the number of tokenized words in the prompt; the cross-attention map 220 is associated with the token “person” and the cross-attention map 225 is associated with the token “vehicle”; [0058]: the words can also be referred to as tokens; a token represents a beginning of a sentence, a token represents an end of a sentence, or a token represents punctuation; [0059]: generate a set of tokens from the prompt.). Same motivation of claim 1 is applied here. Regarding to claim 21 (New), Cui discloses a non-transitory computer readable medium storing code for image processing, the code comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations (Fig. 1; [0031]: after obtaining a target image 101, the electronic device 110 inputs the target image 101 into the image segmentation model 112 and the image segmentation model 112 then outputs an image segmentation result 102 corresponding to the target image 101; [0032]: the image segmentation result 102 may be a segmentation map, a text sequence, an attention map, etc. associated with the target image 101; [0033]: the image segmentation model 112 is any neural network that performs image segmentation; [0092]: the electronic device 110 extracts an image feature representation of a target image using a trained image encoder; [0093]: the electronic device 110 generates, using a trained text encoder, a text feature representation corresponding to a name of the class, and determines a candidate segmentation map for the target image; Fig. 8; [0119-0121]: the components of the electronic device 800 includes one or more processors or processing units 810, a memory 820, a storage device 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860; the processing unit 810 may be an actual or virtual processor and can execute various processes according to the programs stored in the memory 820; the electronic device 800 typically includes a variety of computer storage medium; the electronic device 800 includes additional removable/non-removable, transitory/non-transitory, volatile/non-volatile storage medium.) comprising: The rest claim limitations are similar to claim limitations recited in claim 1. Therefore, same rational used to reject claim 1 is also used to reject claim 21. Regarding to claim 22 (New), Cui in view of Park and Franco discloses the non-transitory computer readable medium of claim 21, the code further comprising instructions executable by the at least one processor to perform operations (same as rejected in claim 21) comprising: The rest claim limitations are similar to claim limitations recited in claim 4. Therefore, same rational used to reject claim 4 is also used to reject claim 22. Regarding to claim 23 (New), Cui in view of Park and Franco discloses the non-transitory computer readable medium of claim 21, the code further comprising instructions executable by the at least one processor to perform operations (same as rejected in claim 21) comprising: The rest claim limitations are similar to claim limitations recited in claim 3. Therefore, same rational used to reject claim 3 is also used to reject claim 23. Regarding to claim 24 (New), Cui in view of Park and Franco discloses the non-transitory computer readable medium of claim 21, wherein selecting the attention map size (same as rejected in claim 21) comprises: The rest claim limitations are similar to claim limitations recited in claim 4. Therefore, same rational used to reject claim 4 is also used to reject claim 24. Regarding to claim 25 (New), Cui in view of Park and Franco discloses the non-transitory computer readable medium of claim 24, wherein selecting the attention map size (same as rejected in claim 21) comprises: The rest claim limitations are similar to claim limitations recited in claim 5. Therefore, same rational used to reject claim 5 is also used to reject claim 25. Regarding to claim 26 (New), Cui discloses a system (Fig. 1; [0031]: after obtaining a target image 101, the electronic device 110 inputs the target image 101 into the image segmentation model 112 and the image segmentation model 112 then outputs an image segmentation result 102 corresponding to the target image 101; [0032]: the image segmentation result 102 may be a segmentation map, a text sequence, an attention map, etc. associated with the target image 101; [0033]: the image segmentation model 112 is any neural network that performs image segmentation; [0092]: the electronic device 110 extracts an image feature representation of a target image using a trained image encoder; [0093]: the electronic device 110 generates, using a trained text encoder, a text feature representation corresponding to a name of the class, and determines a candidate segmentation map for the target image; Fig. 8; [0119-0121]: the components of the electronic device 800 includes one or more processors or processing units 810, a memory 820, a storage device 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860; the processing unit 810 may be an actual or virtual processor and can execute various processes according to the programs stored in the memory 820; the electronic device 800 typically includes a variety of computer storage medium; the electronic device 800 includes additional removable/non-removable, transitory/non-transitory, volatile/non-volatile storage medium.) comprising: a memory component (Fig. 8; [0119-0121]: the components of the electronic device 800 includes one or more processors or processing units 810, a memory 820, a storage device 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860; the processing unit 810 may be an actual or virtual processor and can execute various processes according to the programs stored in the memory 820; the electronic device 800 typically includes a variety of computer storage medium; the electronic device 800 includes additional removable/non-removable, transitory/non-transitory, volatile/non-volatile storage medium.); and a processing device coupled to the memory component, the processing device configured to perform operations comprising (Fig. 8; [0119-0121]: the components of the electronic device 800 includes one or more processors or processing units 810, a memory 820, a storage device 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860; the processing unit 810 may be an actual or virtual processor and can execute various processes according to the programs stored in the memory 820; the electronic device 800 typically includes a variety of computer storage medium; the electronic device 800 includes additional removable/non-removable, transitory/non-transitory, volatile/non-volatile storage medium. ): the rest claim limitations are similar to claim limitations recited in claim 1. Therefore, same rational used to reject claim 1 is also used to reject claim 26. Regarding to claim 27 (New), Cui in view of Park and Franco discloses the system of claim 26, wherein the processing device configured to perform operations further (same as rejected in claim 26) comprising: The rest claim limitations are similar to claim limitations recited in claim 2. Therefore, same rational used to reject claim 2 is also used to reject claim 27. Regarding to claim 28 (New), Cui in view of Park and Franco discloses the system of claim 26, wherein the processing device configured to perform operations further (same as rejected in claim 26) comprising: The rest claim limitations are similar to claim limitations recited in claim 3. Therefore, same rational used to reject claim 3 is also used to reject claim 28. Regarding to claim 29 (New), Cui in view of Park and Franco discloses the system of claim 26, wherein selecting the attention map size (same as rejected in claim 26) comprises: The rest claim limitations are similar to claim limitations recited in claim 4. Therefore, same rational used to reject claim 4 is also used to reject claim 29. Regarding to claim 30 (New), Cui in view of Park and Franco discloses the system of claim 29, wherein selecting the attention map size (same as rejected in claim 26) comprises: The rest claim limitations are similar to claim limitations recited in claim 5. Therefore, same rational used to reject claim 5 is also used to reject claim 30. Regarding to claim 31 (New), Cui in view of Park and Franco discloses the system of claim 26, wherein: the image generation model comprises a U-Net (Franco; [0030]: a latent space diffusion U-Net; [0031]: the U-Net architecture consists of a stack of modular blocks). Same motivation of claim 1 is applied here. Regarding to claim 32 (New), Cui in view of Park and Franco discloses the system of claim 26, wherein: the image generation model comprises a diffusion model (Franco; [0030]: a stable diffusion model consists of an encoder-decoder module and a latent space diffusion U-Net). Same motivation of claim 1 is applied here. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Hai Tao Sun whose telephone number is (571)272-5630. The examiner can normally be reached 9:00AM-6:00PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Hajnik can be reached at 5712727642. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /HAI TAO SUN/Primary Examiner, Art Unit 2616
Read full office action

Prosecution Timeline

Aug 22, 2024
Application Filed
Aug 12, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12738052
SYSTEMS AND METHODS FOR IDENTIFYING TREES AND ESTIMATING TREE HEIGHTS AND OTHER TREE PARAMETERS
2y 1m to grant Granted Sep 15, 2026
Patent 12731299
GRAPHICS RENDERING SYSTEM AND METHOD
2y 8m to grant Granted Sep 08, 2026
Patent 12731333
GAZE-BASED AUDIO SWITCHING AND 3D SIGHT LINE TRIANGULATION MAP
2y 6m to grant Granted Sep 08, 2026
Patent 12718493
PREDICTION OF CONTACT POINTS BETWEEN 3D MODELS
3y 6m to grant Granted Aug 25, 2026
Patent 12718477
DYNAMIC (4D) SCENE RECONSTRUCTION USING MULTIPLE NEURAL RADIANCE FIELDS
2y 5m to grant Granted Aug 25, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
74%
Grant Probability
98%
With Interview (+24.7%)
2y 6m (~5m remaining)
Median Time to Grant
Low
PTA Risk
Based on 493 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month