Prosecution Insights
Last updated: August 17, 2026
Application No. 18/955,806

CONTROLLABLE IMAGE SYNTHESIS FOR TRANSFORMER-BASED IMAGE GENERATION MODELS

Non-Final OA §102§103
Filed
Nov 21, 2024
Examiner
WELCH, DAVID T
Art Unit
2613
Tech Center
2600 — Communications
Assignee
Adobe Inc.
OA Round
1 (Non-Final)
81%
Grant Probability
Favorable
1-2
OA Rounds
1y 3m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 81% — above average
81%
Career Allowance Rate
256 granted / 315 resolved
+19.3% vs TC avg
Strong +27% interview lift
Without
With
+26.8%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
33 currently pending
Career history
345
Total Applications
across all art units

Statute-Specific Performance

§101
11.7%
-28.3% vs TC avg
§103
49.0%
+9.0% vs TC avg
§102
21.2%
-18.8% vs TC avg
§112
12.0%
-28.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 315 resolved cases

Office Action

§102 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Objections Claim 19 is objected to because of the following informalities: this claim uses the acronym “VQGAN” which should be preceded by its generic terminology. The Examiner suggests amending the acronym to read --Vector Quantized Generative Adversarial Network (VQGAN)--. Appropriate correction is required. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 1, 7, and 8 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Gafni et al. (US Patent Application Pub. No. 2024/0221235), referred herein as Gafni. Regarding claim 1, Gafni teaches a method comprising: obtaining a condition map comprising a spatial representation of a target image structure, and encoding, using a condition encoder of an image generation model, the condition map to obtain a condition sequence of tokens representing the target image structure (figs 2 and 3; paragraph 31, lines 9-11; paragraph 33, lines 1-8; paragraph 34, lines 1-17; paragraph 43, lines 7-14; paragraph 59, lines 1-9; a condition map is obtained that represents a target image structure, which is then encoded to obtain a condition sequence of tokens representing the target image structure); generating, using a transformer of the image generation model, an output sequence of tokens based on the condition sequence of tokens and a preliminary sequence of tokens from a discrete codebook (figs 2 and 3; paragraph 31, lines 11-16; paragraph 33, lines 1-8; paragraph 41; paragraph 43, lines 11-25; paragraph 59, lines 9-13; a transformer uses the condition sequence of tokens and a prior sequence of tokens from a codebook [a fixed set of possible tokens, per applicant’s specification] to generate an output sequence of tokens); and generating, using a decoder the image generation model, a synthetic image based on the output sequence of tokens, wherein the synthetic image depicts a scene with the target image structure (fig 2; paragraph 31, the last 6 lines; paragraph 37; paragraph 59, lines 13-16; a decoder generates a synthetic image depicting a scene with the target image structure based on the output sequence of tokens). Regarding claim 7, Gafni teaches the method of claim 1, wherein generating the output sequence of tokens comprises: combining the condition sequence of tokens and the preliminary sequence of tokens to obtain a combined sequence of tokens, wherein the output sequence of tokens is based on the combined sequence of tokens (paragraph 31, lines 11-16; paragraph 33, lines 1-8; paragraph 41; paragraph 43, lines 11-25; paragraph 59, lines 9-13). Regarding claim 8, Gafni teaches the method of claim 1, wherein: the condition map comprises an edge map, a spatial color map, or a depth map (paragraph 31, lines 9-11; paragraph 34, lines 1-17). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 2-4 and 10-12, 14, and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Gafni, in view of Sada (U.S. Patent Application Publication No. 2026/0051143), referred herein as Sada. Regarding claim 2, Gafni teaches the method of claim 1, further comprising: performing a process on the condition sequence of tokens to obtain a subsequent condition sequence of tokens, wherein the output sequence of tokens is generated based on the subsequent condition sequence of tokens (paragraph 31, lines 11-16; paragraph 33, lines 1-8; paragraph 41; paragraph 43, lines 11-25; paragraph 59, lines 9-13). Gafni does not explicitly teach performing a linear attention process. However, in a similar field of endeavor, Sada teaches a method comprising receiving input image data representing a target image, encoding the input to obtain a sequence of tokens, and generating an output image based on the sequence of tokens (paragraph 33; paragraphs 35 and 36; paragraph 44), and further comprising performing a linear attention process on the sequence of tokens (figs 8 and 9; paragraphs 75 and 88). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the linear attention process of Sada with the token processing of Gafni because this helps reduce computational complexity, thereby improving processing speed and efficiency (see, for example, Sada, paragraph 25, lines 1-3; paragraph 28; paragraphs 86 and 114). Regarding claim 3, Gafni teaches the method of claim 1, wherein generating the output sequence of tokens comprises: performing a process on the preliminary sequence of tokens (paragraph 31, lines 11-16; paragraph 33, lines 1-8; paragraph 41; paragraph 43, lines 11-25; paragraph 59, lines 9-13). Gafni does not explicitly teach performing a linear attention process. However, in a similar field of endeavor, Sada teaches a method comprising receiving input image data representing a target image, encoding the input to obtain a sequence of tokens, and generating an output image based on the sequence of tokens (paragraph 33; paragraphs 35 and 36; paragraph 44), and further comprising performing a linear attention process on the sequence of tokens (figs 8 and 9; paragraphs 75 and 88). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the linear attention process of Sada with the token processing of Gafni because this helps reduce computational complexity, thereby improving processing speed and efficiency (see, for example, Sada, paragraph 25, lines 1-3; paragraph 28; paragraphs 86 and 114). Regarding claim 4, Gafni in view of Sada teaches the method of claim 3, wherein: the linear attention process comprises a bidirectional generation process (Gafni, fig 2, parallel token processing; paragraph 31, lines 9-13; paragraph 44, lines 1-6; Sada, paragraphs 75 and 88; the motivation to combine is similar to that discussed above in the rejection of claim 3). Regarding claim 10, Gafni teaches a non-transitory computer readable medium storing code for image processing, the code comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising (fig 10; paragraph 63, lines 1-9): encoding, using a first process, a condition map to obtain a condition sequence of tokens representing a target image structure (figs 2 and 3; paragraph 31, lines 9-11; paragraph 33, lines 1-8; paragraph 34, lines 1-17; paragraph 43, lines 7-14; paragraph 59, lines 1-9; a condition map is obtained that represents a target image structure, which is then encoded to obtain a condition sequence of tokens representing the target image structure); generating, using a second process, an output sequence of tokens based on the condition sequence of tokens and a preliminary sequence of tokens (figs 2 and 3; paragraph 31, lines 11-16; paragraph 33, lines 1-8; paragraph 41; paragraph 43, lines 11-25; paragraph 59, lines 9-13; a transformer uses the condition sequence of tokens and a prior sequence of tokens to generate an output sequence of tokens); and generating, using an image generation model, a synthetic image based on the output sequence of tokens, wherein the synthetic image depicts a scene with the target image structure (fig 2; paragraph 31, the last 6 lines; paragraph 37; paragraph 59, lines 13-16; a decoder generates a synthetic image depicting a scene with the target image structure based on the output sequence of tokens). Gafni does not explicitly teach performing linear attention processes. However, in a similar field of endeavor, Sada teaches a system for performing a method comprising receiving input image data representing a target image, encoding the input to obtain a sequence of tokens, and generating an output image based on the sequence of tokens (paragraph 33; paragraphs 35 and 36; paragraph 44), and further comprising performing linear attention processes on the sequence of tokens (figs 8 and 9; paragraphs 75 and 88). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the linear attention processes of Sada with the token processing of Gafni because this helps reduce computational complexity, thereby improving processing speed and efficiency (see, for example, Sada, paragraph 25, lines 1-3; paragraph 28; paragraphs 86 and 114). Regarding claim 11, Gafni in view of Sada teaches the non-transitory computer readable medium of claim 10, wherein: the first linear attention process comprises an autoregressive generation process (Gafni, paragraph 31, lines 9-16; paragraph 41; Sada, paragraphs 75 and 88; the motivation to combine is similar to that discussed above in the rejection of claim 3). Regarding claim 12, Gafni in view of Sada teaches the non-transitory computer readable medium of claim 10, wherein: the first linear attention process comprises a bidirectional generation process (Gafni, fig 2, parallel token processing; paragraph 31, lines 9-13; paragraph 44, lines 1-6; Sada, paragraphs 75 and 88; the motivation to combine is similar to that discussed above in the rejection of claim 3). Regarding claim 14, Gafni in view of Sada teaches the non-transitory computer readable medium of claim 10, wherein generating the output sequence of tokens comprises: combining the condition sequence of tokens and the preliminary sequence of tokens to obtain a combined sequence of tokens, wherein the output sequence of tokens is based on the combined sequence of tokens (Gafni, paragraph 31, lines 11-16; paragraph 33, lines 1-8; paragraph 41; paragraph 43, lines 11-25; paragraph 59, lines 9-13). Regarding claim 15, Gafni in view of Sada teaches the non-transitory computer readable medium of claim 10, wherein: the condition map comprises an edge map, a spatial color map, or a depth map (Gafni, paragraph 31, lines 9-11; paragraph 34, lines 1-17). Claims 5, 6, 9, and 16-20 are rejected under 35 U.S.C. 103 as being unpatentable over Gafni, in view of Fan et al. (U.S. Patent Application Publication No. 2022/0277218), referred herein as Fan. Regarding claim 5, Gafni teaches the method of claim 1, wherein: a token of the condition sequence of tokens comprises an index corresponding to an image patch and a token of the output sequence of tokens comprises a token from the discrete codebook with the index indicating the image patch (paragraph 17, lines 9-14; paragraph 31, lines 9-16). One skilled in the art could infer that correspondence to an image patch comprises correspondence to an image patch location; however, Gafni does not explicitly discuss that a token corresponds to an image patch location. However, in a similar field of endeavor, Fan teaches a method comprising obtaining a condition map representing a target image, encoding the condition map to generate a condition sequence of tokens, generating an output sequence of tokens based on the condition sequence of tokens and a preliminary sequence of tokens using a transformer, and generating a synthetic image based on the output sequence of tokens (paragraph 62, lines 6-16 and the last 7 lines; paragraph 63, the last 8 lines; paragraph 67, lines 1-11; paragraph 69, lines 1-8; paragraph 70; paragraph 71, lines 1-6), wherein tokens comprise an index corresponding to a patch location (paragraph 67, lines 1-11; paragraph 69, the last 9 lines; paragraph 77). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the token/image patch location correspondence of Fan with the token/image patch correspondence of Gafni because this helps to greatly improve the model’s ability to associate tokens with patches, increasing both the efficiency and accuracy with which it does so, and ultimately improving the model’s output results (see, for example, Fan, paragraphs 107 and 108). Regarding claim 6, Gafni teaches the method of claim 1, but does not teach the method, wherein: each of the preliminary sequence of tokens comprises a mask token. However, in a similar field of endeavor, Fan teaches a method comprising obtaining a condition map representing a target image, encoding the condition map to generate a condition sequence of tokens, generating an output sequence of tokens based on the condition sequence of tokens and a preliminary sequence of tokens using a transformer, and generating a synthetic image based on the output sequence of tokens (paragraph 62, lines 6-16 and the last 7 lines; paragraph 63, the last 8 lines; paragraph 67, lines 1-11; paragraph 69, lines 1-8; paragraph 70; paragraph 71, lines 1-6), wherein the preliminary sequence of tokens comprises a mask token (paragraph 63; paragraph 69, lines 1-23). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the mask tokens of Fan with the tokens of Gafni because this helps to greatly improve the model’s ability to process the tokens and their effect on the target image, increasing both the efficiency and accuracy with which it does so, and ultimately improving the model’s output results (see, for example, Fan, paragraphs 107 and 108). Regarding claim 9, Gafni teaches the method of claim 1, wherein: the image generation model is trained using a training set including an image and a training condition map comprising a spatial representation of an image structure of the image (paragraph 31, lines 9-16; paragraph 33, lines 1-8; paragraph 34, lines 1-17; paragraph 43, lines 7-14; paragraph 59, lines 1-9). Gafni does not explicitly teach utilizing a masked image. However, in a similar field of endeavor, Fan teaches a method comprising obtaining a condition map representing a target image, encoding the condition map to generate a condition sequence of tokens, generating an output sequence of tokens based on the condition sequence of tokens and a preliminary sequence of tokens using a transformer, and generating a synthetic image based on the output sequence of tokens (paragraph 62, lines 6-16 and the last 7 lines; paragraph 63, the last 8 lines; paragraph 67, lines 1-11; paragraph 69, lines 1-8; paragraph 70; paragraph 71, lines 1-6), wherein the training comprises using a training set including a masked image (paragraph 63; paragraph 69, lines 1-23). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the masked training image of Fan with the training image of Gafni because this helps to greatly improve the model’s ability to process the tokens and their effect on the target image, increasing both the efficiency and accuracy with which it does so, and ultimately improving the model’s output results (see, for example, Fan, paragraphs 107 and 108). Regarding claim 16, Gafni teaches a system comprising: a memory component; and a processing device coupled to the memory component, the processing device configured to perform operations comprising (fig 10; paragraph 63, lines 1-9): obtaining a condition map comprising a spatial representation of a target image structure, and encoding, using a condition encoder of an image generation model, the condition map to obtain a condition sequence of tokens representing the target image structure (figs 2 and 3; paragraph 31, lines 9-11; paragraph 33, lines 1-8; paragraph 34, lines 1-17; paragraph 43, lines 7-14; paragraph 59, lines 1-9; a condition map is obtained that represents a target image structure, which is then encoded to obtain a condition sequence of tokens representing the target image structure), wherein a token of the condition sequence of tokens comprises an index corresponding to an image patch (paragraph 17, lines 9-14; paragraph 31, lines 9-16); generating, using a transformer of the image generation model, an output sequence of tokens based on the condition sequence of tokens and a preliminary sequence of tokens from a discrete codebook (figs 2 and 3; paragraph 31, lines 11-16; paragraph 33, lines 1-8; paragraph 41; paragraph 43, lines 11-25; paragraph 59, lines 9-13; a transformer uses the condition sequence of tokens and a prior sequence of tokens from a codebook [a fixed set of possible tokens, per applicant’s specification] to generate an output sequence of tokens), wherein a token of the output sequence of tokens comprises a token from the discrete codebook with the index indicating the image patch (paragraph 17, lines 9-14; paragraph 31, lines 9-16); and generating, using a decoder the image generation model, a synthetic image based on the output sequence of tokens, wherein the synthetic image depicts a scene with the target image structure (fig 2; paragraph 31, the last 6 lines; paragraph 37; paragraph 59, lines 13-16; a decoder generates a synthetic image depicting a scene with the target image structure based on the output sequence of tokens). One skilled in the art could infer that correspondence to an image patch comprises correspondence to an image patch location; however, Gafni does not explicitly discuss that a token corresponds to an image patch location. However, in a similar field of endeavor, Fan teaches a system for performing a method comprising obtaining a condition map representing a target image, encoding the condition map to generate a condition sequence of tokens, generating an output sequence of tokens based on the condition sequence of tokens and a preliminary sequence of tokens using a transformer, and generating a synthetic image based on the output sequence of tokens (paragraph 62, lines 6-16 and the last 7 lines; paragraph 63, the last 8 lines; paragraph 67, lines 1-11; paragraph 69, lines 1-8; paragraph 70; paragraph 71, lines 1-6), wherein tokens comprise an index corresponding to a patch location (paragraph 67, lines 1-11; paragraph 69, the last 9 lines; paragraph 77). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the token/image patch location correspondence of Fan with the token/image patch correspondence of Gafni because this helps to greatly improve the model’s ability to associate tokens with patches, increasing both the efficiency and accuracy with which it does so, and ultimately improving the model’s output results (see, for example, Fan, paragraphs 107 and 108). Regarding claim 17, Gafni in view of Fan teaches the system of claim 16, wherein: the condition encoder comprises a plurality of linear attention blocks (Fan, paragraph 67, lines 1-6; paragraph 74; paragraph 75, lines 1-5; paragraph 77; the motivation to combine is similar to that discussed above in the rejection of claim 16). Regarding claim 18, Gafni in view of Fan teaches the system of claim 16, wherein: the condition encoder has a same architecture as the transformer of the image generation model (Gafni, paragraph 41; paragraph 44, lines 1-16; GPT utilizes BPE, thus the encoder and transformer are of the same architecture; Fan, paragraph 62, the last 7 lines; paragraph 64; BERT utilizes Wordpiece, thus the encoder and transformer are of the same architecture). Regarding claim 19, Gafni in view of Fan teaches the system of claim 16, wherein: the decoder comprises a VQGAN architecture (Gafni, paragraph 38, lines 1-22; paragraph 40, lines 1-8; paragraph 50, lines 6-9). Regarding claim 20, Gafni in view of Fan teaches the system of claim 16, further comprising: an encoder configured to generate a sequence of tokens representing an input image (Gafni, paragraph 31, lines 9-11; paragraph 59, lines 1-9). Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over Gafni, in view of Sada, and further in view of Fan. Regarding claim 13, Gafni in view of Sada teaches the non-transitory computer readable medium of claim 10, but does not explicitly teach the method, wherein: each of the preliminary sequence of tokens comprises a mask token. However, in a similar field of endeavor, Fan teaches a system for performing a method comprising obtaining a condition map representing a target image, encoding the condition map to generate a condition sequence of tokens, generating an output sequence of tokens based on the condition sequence of tokens and a preliminary sequence of tokens using a transformer, and generating a synthetic image based on the output sequence of tokens (paragraph 62, lines 6-16 and the last 7 lines; paragraph 63, the last 8 lines; paragraph 67, lines 1-11; paragraph 69, lines 1-8; paragraph 70; paragraph 71, lines 1-6), wherein the preliminary sequence of tokens comprises a mask token (paragraph 63; paragraph 69, lines 1-23). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the mask tokens of Fan with the tokens of Gafni in view of Sada because this helps to greatly improve the model’s ability to process the tokens and their effect on the target image, increasing both the efficiency and accuracy with which it does so, and ultimately improving the model’s output results (see, for example, Fan, paragraphs 107 and 108). Conclusion The following prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Mitchell (U.S. Patent Application Publication No. 2025/0022185); Fine-tuning images generated by artificial intelligence based on aesthetic and accuracy metrics and systems and methods for the same. Karpman (U.S. Patent No. 12,499,519); Training and deployment of image generation models. Brégier (U.S. Patent Application Publication No. 2026/0134624); Conditional human mesh recovery in multi-person scenes. Lawrence et al. (Attending to Future Tokens for Bidirectional Sequence Generation); 9th International Joint Conference on Natural Language Processing; 2019. Chang et al. (Muse: Text-To-Image Generation via Masked Generative Transformers); arXiv; January 2023. Any inquiry concerning this communication or earlier communications from the examiner should be directed to DAVID T WELCH whose telephone number is (571)270-5364. The examiner can normally be reached on Monday-Thursday, 8:30-5:30 EST, and alternate Fridays, 9:00-2:30 EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Xiao Wu can be reached on 571-272-7761. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. DAVID T. WELCH Primary Examiner Art Unit 2613 /DAVID T WELCH/Primary Examiner, Art Unit 2613
Read full office action

Prosecution Timeline

Nov 21, 2024
Application Filed
Jul 17, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12682570
PRECOMPUTED CELL GENERATION AND DISPLAY METHODS AND SYSTEMS
1y 9m to grant Granted Jul 14, 2026
Patent 12670633
IMAGE OPTIMIZATIONS FOR MODERN WEB
2y 9m to grant Granted Jun 30, 2026
Patent 12664703
SYNCHRONIZING IMAGE SIGNAL PROCESSOR AND IMAGE SENSOR CONFIGURATIONS
2y 4m to grant Granted Jun 23, 2026
Patent 12663859
COMMUNICATION METHOD, WEARABLE DEVICE, AND STORAGE MEDIUM
2y 3m to grant Granted Jun 23, 2026
Patent 12659450
METHOD, AN APPARATUS AND A COMPUTER PROGRAM PRODUCT FOR REAL-TIME STREAMING OF VOLUMETRIC VIDEO
2y 3m to grant Granted Jun 16, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
81%
Grant Probability
99%
With Interview (+26.8%)
3y 0m (~1y 3m remaining)
Median Time to Grant
Low
PTA Risk
Based on 315 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month