Prosecution Insights
Last updated: August 17, 2026
Application No. 18/677,874

GENERATIVE ARTIFICAL INTELLIGENCE VISUAL EFFECTS

Non-Final OA §103
Filed
May 30, 2024
Priority
Apr 22, 2024 — IN 202411031947
Examiner
HE, WEIMING
Art Unit
2611
Tech Center
2600 — Communications
Assignee
Adobe Inc.
OA Round
3 (Non-Final)
46%
Grant Probability
Moderate
3-4
OA Rounds
1y 2m
Est. Remaining
59%
With Interview

Examiner Intelligence

Grants 46% of resolved cases
46%
Career Allowance Rate
193 granted / 417 resolved
-15.7% vs TC avg
Moderate +13% lift
Without
With
+12.9%
Interview Lift
resolved cases with interview
Typical timeline
3y 4m
Avg Prosecution
28 currently pending
Career history
456
Total Applications
across all art units

Statute-Specific Performance

§101
8.2%
-31.8% vs TC avg
§103
61.5%
+21.5% vs TC avg
§102
10.9%
-29.1% vs TC avg
§112
15.2%
-24.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 417 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 6/10/2026 has been entered. Response to Amendment The amendment filed on 6/10/2026 has been entered and made of record. Claims 1, 13 and 16-17 are amended. Claims 2 and 20 are cancelled. Claims 1, 3-19 and 21 are pending. Response to Arguments Applicant’s arguments with respect to claims 1, 13 and 16 have been considered but they are not persuasive. Applicant asserts that both Chen and Smock merely describe a text and image encoder, and are silent to use of "a generative machine-learning model" nor the basis for generating the claimed "contribution image embedding" using "the text specifying the visual effect, the text specifying the shape, and the mask" as recited in Claim 1 as amended (p. 9 of Remarks). The above argument is directed to two claimed features: a.) generative machine-learning (ML) model; b) text-to-image ML model. Here, it is well-known that an encoder of a generative ML model may process the input text/image data to generate text/image embedding, while the text prompt may specify the visual effect, shape and mask. For example, applicant discloses “To do so in one or more examples, the visual effect generation system first employs a generative model that is conditioned on both text and image embeddings, e.g., is a text-to-image machine-learning model such as "Dall-E." A contribution image embedding is generated by the generative model based on the prompt, e.g., identifying the shape and the visual effect” in [0026]. Examine notices that Chen discloses “To address these issues, the proposed AI-based consistent font visual effect generation pipeline introduces a shape-adaptive generative model that adapts the existing rectangular canvases into irregularly shaped canvas (i.e., non-rectangular canvas), and performs visual content creation on the irregularly shaped canvas. The shape-adaptive generative model includes a generation model, a refinement model, and a visual effect transfer scheme… In one implementation, the dataset of shape-adaptive mask-image-text triplets is generated by another text-to-image model (e.g., DALL•E3) to train the generation model to generate content inside irregular (i.e., non-rectangle) canvas” in [0015]; one example of text prompt in Table 1; style prompt, font mask and conditional diffusion model in [0039]; text prompt embedding in [0059]. Please refer to Smock’s Fig 4 for text/image embedding as below: PNG media_image1.png 437 696 media_image1.png Greyscale Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 4, and 6-9 are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (US 2025/0322570 A1) in view of Smock et al. (US 2024/0331235 A1). As to Claim 1, Chen teaches A method comprising: receiving, by a processing device, a prompt including text specifying a visual effect and text specifying a shape (Chen discloses “The visual effects are expressed via a combination of objects, scales, styles, colors, patterns, lighting, and the like” in [0001]; a style prompt in [0039]; one example of text prompt in Table 1 in [0045]); forming, by the processing device, a mask defining a portion of digital content based on an object selected from the digital content (Chen discloses “and then crops the font image 152 with a font mask 153 of the reference character (e.g., R), to generate a styled reference character image 155” in [0039]; “a font mask 160 of the target character (e.g., D) in [0040]); generating, by the processing device, a contribution image embedding using generative artificial intelligence implemented using a generative machine-learning models that is conditioned on text and image embeddings, the generating of the contribution image embedding by the generative machine-learning model based on the text specifying the visual effect, the text specifying the shape, and the mask (Chen discloses “The consistent font visual effect generation pipeline leverages the advanced capabilities of text-to-image model(s) 126a and vision generative model(s) 126b, to generate character images based on a style prompt at runtime” in [0038]; “the pipeline applies a text-to-image model 126a (e.g., a conditional diffusion model) to iteratively generate salient content of a croissant” in [0039]; “The final output is a combination of the refined image 169 (i.e., the RGB part) and the refined mask 170 (i.e., the A part)), together as a RGBA image that the user requested. Accordingly, a second stage of the visual effect transfer scheme is accomplished by generating the refined target character image with more defined and natural croissant edges in a white background” in [0044]. Here, it is well-known that an image/text encoder is configured to receive image/text to generate image/text embedding. For example, Smock discloses “The image encoder 406 uses machine vision techniques to recognize features in molecular images and generate the image embedding 408. Examples implementations of an image encoder 406 that may be used are provided in Stable Diffusion and InstructPix2Pix” in [0057], see also Fig 4 below, see also above Response to Arguments. PNG media_image2.png 437 710 media_image2.png Greyscale ); receiving, by the processing device, the contribution image embedding by a diffusion model over one or more diffusion iterations, the diffusion model is separate from the generative machine-learning model; and generating, by the processing device, the visual effect over one or more diffusion iterations by a diffusion model based at least in part on the contribution image embedding (Chen discloses “the proposed AI-based consistent font visual effect generation pipeline introduces a shape-adaptive generative model…and performs visual content creation on the irregularly shaped canvas… The generation model is a text-to-image model (e.g., a conditional diffusion model) trained with a dataset of shape-adaptive mask-image-text triplets to generate content within a font-shaped canvas as a styled character image” in [0015]; “to train a conditional text-to-image model including any backbone models including various versions of stable diffusion models into the generation model of the shape-adaptive generative model” in [0023]; “By applying different linear projections, the pipeline transforms Φ' into the query embedding space Q, and the text prompt embedding (or pixel embedding) into the key embedding space K and value embedding space V for cross-attention (or self-attention)” in [0059]; “Secondly, the pipeline applies a text-to-image model 126a (e.g., a conditional diffusion model) to iteratively generate salient content of a croissant and concentrate the salient content within the font mask 153 of the reference character in a depth-to-image (i.e., a conditional text-to-image) process 151 (e.g., for 20 steps/times) using a conditional text2img model (e.g., the conditional diffusion model) to generate a first image 152 of the reference character (e.g., R).” in [0039]; see also image embedding in Fig 1B. Please note that Chen discloses a text-to-image model (a shape-adaptive generative model) can be performed by a diffusion-based text-to-image generation models in [0014]); presenting, by the processing device, the digital content as having the visual effect applied to the portion of the digital content for display in a user interface (Chen, Fig 1B as shown below: PNG media_image3.png 597 702 media_image3.png Greyscale ). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the invention of Chen with the teaching of Smock so as to explain the generation of image embedding by image encoder. As to Claim 4, Chen in view of Smock teaches The method as described in claim 1, wherein the generating the visual effect over the one or more diffusion iterations includes adjusting an amount of noise (Chen discloses “The option of setting the noise strengths with different values within SGM and SRM achieves better results.” in [0061].) As to Claim 6, Chen in view of Smock teaches The method as described in claim 1, further comprising expanding the prompt to include at least one additional item of text using a machine-learning model and wherein the generating of the visual effect is further based on the at least one additional item of text (Chen discloses conditional text to image model in [0046]; a first prompt and a second prompt in Fig 4.) As to Claim 7, Chen in view of Smock teaches The method as described in claim 1, further comprising receiving a user input selecting the object from the digital content via the user interface and wherein the generating is performed responsive to the receiving of the user input (Chen discloses “In this example, the chat pane 225 shows a prompt enter box 225c with instructions of "(1) Enter a style text, (2) Select a style image, or (3) Upload a style image" and several style images for the user to select. After the user selects a style image of an easter egg, the consistent font visual effect generation application can generate characters with the easter egg style for the user as the embodiments discussed above.” in [0070].) As to Claim 8, Chen in view of Smock teaches The method as described in claim 1, wherein the forming of the mask includes forming a binary mask by recoloring the portion of the digital content using a first color and remaining portions of the digital content using a second color (Chen discloses in [0058]. PNG media_image4.png 91 638 media_image4.png Greyscale ) As to Claim 9, Chen in view of Smock teaches The method as described in claim 1, wherein the digital content includes a table and the portion is defined between cells of the table (Chen discloses “The term "style" refers to the distinctive visual characteristics of an image. These characteristics can include color palette, texture (e.g., brushstrokes), composition, layout, structure, scale, typography, level of details and abstraction, whitespace, overall mood, atmosphere, and the like.” in [0030]. Here, color palette can be a table, while each color is the cell of the table.) Claims 3 and 12-15 are rejected under 35 U.S.C. 103 as being unpatentable over Chen in view of Smock and Choi et al. (Custom-Edit: Text-Guided Image Editing with Customized Diffusion Models). As to Claim 3, Chen in view of Smock teaches The method as described in claim 1, wherein the generating the visual effect over the one or more diffusion iterations includes a first said diffusion iteration in which the contribution image embedding is applied and a second said diffusion iteration in which the contribution image embedding is removed (Chen discloses “The generation model is a text-to-image model (e.g., a conditional diffusion model) trained with a dataset of shape-adaptive mask-image-text triplets to generate content within a font-shaped canvas as a styled character image.” in [0015]. Choi further discloses “The attention map editing operation Edit includes two sub-operations, namely prompt refinement and word swap. word swap refers to replacing cross-attention maps of words in the original prompt with other words, while prompt refinement refers to adding cross-attention maps of new words to the prompt while preserving attention maps of the common words.” at p. 6. Here, Choi’s swapping processing is applied to add or remove the image embedding for respective diffusion iterations based on the word swap.) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the invention of Chen and Smock with the teaching of Choi so as to perform both operations of prompt refinement and word swap for the limited number of reference images (Choi, p. 2). As to Claim 12, Chen in view of Smock teaches The method as described in claim 1. The combination of Choi further teaches comprising receiving an edit input that alters the portion of the digital content and reapplying the visual effect to the altered portion (Chen discloses conditional text to image model in [0046]; a first prompt and a second prompt in Fig 4. Choi discloses text-guided image editing as shown in Fig 3.) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the invention of Chen and Smock with the teaching of Choi so as to perform a text-guided editing to allow fine-grained editing with textural prompt. Claim 13 recites similar limitations as claims 1 & 3 but in a device form. Therefore, the same rationale used for claims 1 & 3 is applied. Claim 14 is rejected based upon similar rationale as Claim 1. Claim 15 is rejected based upon similar rationale as Claim 9. Claims 5, 16-19 and 21 are rejected under 35 U.S.C. 103 as being unpatentable over Chen in view of Smock and Watson et al. (US 2021/0390681 A1). As to Claim 5, Chen in view of Smock teaches The method as described in claim 4. The combination of Watson further teaches wherein the adjusting is based on a user input received via a control in the user interface (Chen discloses “The option of setting the noise strengths with different values within SGM and SRM achieves better results.” in [0061]. Watson further discloses “In addition, the parameters indicating the augmentation limits (e.g., for generating training images) may be defined under a user's control. FIG. 9 is an example of an image augmentation graphical user interface 900 that can be used to define the parameters indicating the augmentation limits” in [0079]; “The noise augmentation functionality enables a user to specify a parameter indicating an amount of noise for generating augmented training images. More particularly, the graphical user interface 900 includes a slider 934 with a handle 936 that can be selected and moved to set a value 938 of the parameter indicating the maximum amount of digital noise (e.g., 10)…” in [0082], see also Fig 9 below: PNG media_image5.png 773 559 media_image5.png Greyscale It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the invention of Chen and Smock with the teaching of Watson so as to provide a graphical user interface to enable user to specify a parameter indicating an amount of noise for generating augmented training images (Watson, [0082]). Claim 16 recites similar limitations as claims 1 & 5 further with user input to select image (Chen, [0070, 0114]), but in a computer readable medium form. Therefore, the same rationale used for claims 1 & 5 is applied. Claim 17 is rejected based upon similar rationale as Claim 1. Claim 18 is rejected based upon similar rationale as Claim 3. Claim 19 is rejected based upon similar rationale as Claim 4. As to Claim 21, Chen in view of Smock and Watson teaches The one or more computer-readable storage media as described in claim 19, wherein the diffusion model is separate from the generative machine-learning model (Chen discloses a text-to-image model (a shape-adaptive generative model) can be performed by a diffusion-based text-to-image generation models in [0014].) Claims 10-11 are rejected under 35 U.S.C. 103 as being unpatentable over Chen in view of Smock and Fisher et al. (US 2021/0073323 A1). As to Claim 10, Chen in view of Smock teaches The method as described in claim 1, wherein the object is a vector object (Chen discloses a AI-based font effect generation system as shown in Fig 1B. It is well-known that a character is a vector object. For example, Fisher discloses “In other words, a captured font can be generated by applying the optimized parameters to other characters to produce vectorized versions of all characters” in [0024].) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the invention of Chen and Smock with the teaching of Fisher so as to generate a vectorized versions of the characters (Fisher, [0024]). As to Claim 11, Chen in view of Smock and Fisher teaches The method as described in claim 10, further comprising generating the vector object from a raster object (Fisher discloses “to automatically generate a font from an image by creating a rasterized mask and then running standard vectorization algorithms on the rasterizations” in [0019].) Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Chen in view of Smock, Watson and Couairon et al. (DIFFEDIT: DIFFUSION-BASED SEMANTIC IMAGE EDITING WITH MASK GUIDANCE, cited in IDS). As to Claim 20, Chen in view of Smock and Watson teaches The one or more computer-readable storage media as described in claim 19. The combination of Couairon further teaches wherein the noise is Gaussian noise (Chen discloses adding noise within diffusion model as shown in Fig 1E. Couairon further discloses “Diffusion models are an especially interesting class of model for image editing because of their iterative denoising process starting from random Gaussian noise. This process can be guided through a variety of techniques, like CLIP guidance” at p.2.) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the invention of Chen, Smock and Watson with the teaching of Couairon because Gaussian noise is easy to work with mathematically and allow the model to learn how to reverse the process of adding noise to create structured data. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to WEIMING HE whose telephone number is (571)270-1221. The examiner can normally be reached on Monday-Friday, 8:30am-5:00pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tammy Goddard can be reached on 571-272-7773. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see https://ppair-my.uspto.gov/pair/PrivatePair. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /WEIMING HE/ Primary Examiner, Art Unit 2611
Read full office action

Prosecution Timeline

Show 2 earlier events
Mar 17, 2026
Response Filed
Apr 14, 2026
Examiner Interview (Telephonic)
Apr 17, 2026
Final Rejection mailed — §103
Jun 08, 2026
Examiner Interview Summary
Jun 08, 2026
Applicant Interview (Telephonic)
Jun 10, 2026
Request for Continued Examination
Jun 14, 2026
Response after Non-Final Action
Jul 30, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12703209
VEHICLE MONITORING SYSTEM
1y 6m to grant Granted Aug 11, 2026
Patent 12639877
REFINEMENT OF FACIAL KEYPOINT METADATA GENERATION FOR VIDEO CONFERENCING OR OTHER APPLICATIONS
3y 6m to grant Granted May 26, 2026
Patent 12632615
DATA SERIALIZATION EXTRUSION FOR CONVERTING TWO-DIMENSIONAL IMAGES TO THREE-DIMENSIONAL GEOMETRY
5y 11m to grant Granted May 19, 2026
Patent 12633000
TEXT-TO-IMAGE SYNTHESIS UTILIZING DIFFUSION MODELS WITH TEST-TIME ATTENTION SEGREGATION AND RETENTION OPTIMIZATION
2y 11m to grant Granted May 19, 2026
Patent 12608891
INFORMATION PROCESSING DEVICE, HEAD-MOUNTED DISPLAY DEVICE, CONTROL METHOD OF INFORMATION PROCESSING DEVICE, AND NON-TRANSITORY COMPUTER READABLE MEDIUM WITH WHITE-BALANCE CORRECTION VALUE CORRESPONDING TO COLOR TEMPERATURE OF ENVIRONMENT LIGHT-SOURCE
2y 9m to grant Granted Apr 21, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
46%
Grant Probability
59%
With Interview (+12.9%)
3y 4m (~1y 2m remaining)
Median Time to Grant
High
PTA Risk
Based on 417 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month