Prosecution Insights
Last updated: October 01, 2026
Application No. 19/094,499

TYPEAHEAD IMAGE GENERATION

Non-Final OA §103
Filed
Mar 28, 2025
Priority
Apr 17, 2024 — provisional 63/635,550
Examiner
PARK, HYORIM NMN
Art Unit
Tech Center
Assignee
Meta Platforms Inc.
OA Round
1 (Non-Final)
75%
Grant Probability
Favorable
1-2
OA Rounds
5m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 75% — above average
75%
Career Allowance Rate
3 granted / 4 resolved
+15.0% vs TC avg
Strong +38% interview lift
Without
With
+37.5%
Interview Lift
resolved cases with interview
Fast prosecutor
2y 0m
Avg Prosecution
23 currently pending
Career history
23
Total Applications
across all art units

Statute-Specific Performance

§101
3.9%
-36.1% vs TC avg
§103
63.6%
+23.6% vs TC avg
§102
20.9%
-19.1% vs TC avg
§112
10.9%
-29.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 4 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. Information Disclosure Statement The information disclosure statement (IDS) submitted on 08/12/2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Objections Claims 19 and 20 objected to because of the following informalities: Claim 19 line 1 cites “A non-transitory computer readable medium comprising stored instructions that when executed effectuates:” which should be rewritten such as A non-transitory computer readable medium comprising stored instructions when executed by a processor, comprises:” Claim 20 line 2 cites “wherein the stored instructions when executed further effectuates:” which should be rewritten such as “wherein the stored instructions when executed by a processor further comprises:” Appropriate correction is required. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-3, 8-14, and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Li, Yanyu, et al. "Snapfusion: Text-to-image diffusion model on mobile devices within two seconds." Advances in Neural Information Processing Systems 36 (2023): 20662-20678.; IDS REF (hereinafter Li) in view of Wang, Zijie J., et al. "Diffusiondb: A large-scale prompt gallery dataset for text-to-image generative models." Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers). 2023. (hereinafter Wang). Regarding Claim 1, Li discloses A method comprising: (3 Architecture Optimizations "we propose an architecture-evolving method that preserves the performance of the pre-trained UNet model while gradually improving its efficacy.") receiving, via a user interface during a prompting session, a text prompt describing an image; generating, via a trained diffusion model, the image representative of the text prompt; (Figure 3; Figure 9; Abstract "we present a generic approach that, for the first time, unlocks running text-to-image diffusion models on mobile devices in less than 2 seconds."; 1 Introduction "We propose a novel evolving training framework to obtain an efficient UNet that performs better than the original Stable Diffusion v1.51 while being significantly faster. We also introduce a data distillation pipeline to compress and accelerate the image decoder."; 5.1 Text-to-Image Generation "Our model can generate images from text prompts with high fidelity. More examples are shown in Fig. 9." PNG media_image1.png 242 553 media_image1.png Greyscale ) Li does not disclose determining, via the trained diffusion model, a reconciled risk score based on a determined risk score of the text prompt and a determined risk score of the generated image; and causing, via the trained diffusion model in response to the determined reconciled risk score, to (i) approve the generated image in an instance in which the determined reconciled risk score meets or exceeds a predetermined threshold, or (ii) deny the generated image in an instance in which the determined reconciled risk score fails to meet the predetermined threshold. Wang teaches determining, via the trained diffusion model, a reconciled risk score based on a determined risk score of the text prompt and a determined risk score of the generated image; and (Fig. 2; 2.3 Identifying NSFW Content “The Stable Diffusion Discord server prohibits generating NSFW images (StabilityAI, 2022a). Also, Stable Diffusion has a built-in NSFW filter that automatically blurs generated images if it detects NSFW content. However, we find DIFFUSIONDB still includes NSFW images that were not detected by the built-in filter or removed by server moderators. To help researchers filter these images, we apply state-of-the-art NSFW classifiers to compute NSFW scores for each prompt and image.”; PNG media_image2.png 234 639 media_image2.png Greyscale ) causing, via the trained diffusion model in response to the determined reconciled risk score, to (i) approve the generated image in an instance in which the determined reconciled risk score meets or exceeds a predetermined threshold, or (ii) deny the generated image in an instance in which the determined reconciled risk score fails to meet the predetermined threshold. (2.3 Identifying NSFW Content “The Stable Diffusion Discord server prohibits generating NSFW images (StabilityAI, 2022a). Also, Stable Diffusion has a built-in NSFW filter that automatically blurs generated images if it detects NSFW content. However, we find DIFFUSIONDB still includes NSFW images that were not detected by the built-in filter or removed by server moderators. To help researchers filter these images, we apply state-of-the-art NSFW classifiers to compute NSFW scores for each prompt and image. Researchers can determine a suitable threshold to filter out potentially unsafe data for their tasks.”) As both Li and Wang are from the same field of endeavor, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Li to include determining, via the trained diffusion model, a reconciled risk score based on a determined risk score of the text prompt and a determined risk score of the generated image; and causing, via the trained diffusion model in response to the determined reconciled risk score, to (i) approve the generated image in an instance in which the determined reconciled risk score meets or exceeds a predetermined threshold, or (ii) deny the generated image in an instance in which the determined reconciled risk score fails to meet the predetermined threshold, in the context of using diffusion model to create images, according to the teaching of Wang, in order to filter out potentially unsafe or harmful content (Fig. 2 of Wang). Regarding claim 2, Li does not disclose The method of claim 1, wherein the text prompt comprises a seed indicating a constant attribute associated with the image for a duration of the prompting session. Wang teaches The method of claim 1, wherein the text prompt comprises a seed indicating a constant attribute associated with the image for a duration of the prompting session. (Fig. 2; “DIFFUSIONDB contains 14 million Stable Diffusion images, 1.8 million unique text prompts, and all model hyperparameters: seed, step, CFG scale, sampler, and image size.”; 2.1 Collecting User Generated Images “The bot then replies with the generated images and used random seeds.”) As both Li and Wang are from the same field of endeavor, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Li to include wherein the text prompt comprises a seed indicating a constant attribute associated with the image for a duration of the prompting session, in the context of using diffusion model to create images, according to the teaching of Wang, in order to link images to prompts and hyperparameters including seeds (2 Constructing DIFFUSIONDB of Wang) to help the model with consistency. Regarding claim 3, Li does not disclose The method of claim 1, further comprising: transmitting, via the user interface, an indication to enhance the received text prompt describing the image. Wang teaches The method of claim 1, further comprising: transmitting, via the user interface, an indication to enhance the received text prompt describing the image. (4 Enabling New Research Directions “Prompt Autocomplete. With DIFFUSIONDB, researchers can develop an autocomplete system to help users construct prompts. For example, one can use the prompt corpus to train an n-gram model to predict likely words following a prompt part. Alternatively, researchers can use semantic autocomplete (Hyvönen and Mäkelä, 2006) by categorizing prompt keywords into ontological categories such as subject, style, quality, repetition, and magic terms (Oppenlaender, 2022). This allows the system to suggest related keywords from unspecified categories, for example suggesting style keyword “depth of field” and a magic keyword “award-winning” to improve the quality of generated images. Additionally, researchers can also use DIFFUSIONDB to study prompt auto-replace by distilling effective prompt patterns and creating a “translation” model that replaces weaker prompt keywords with more effective ones.”) As both Li and Wang are from the same field of endeavor, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Li to include transmitting, via the user interface, an indication to enhance the received text prompt describing the image, in the context of using diffusion model to create images, according to the teaching of Wang, in order to enable users to construct prompts more efficiently (4 Enabling New Research Directions of Wang). Regarding claim 8, the combination of Li and Wang discloses The method of claim 1, further comprising: causing the approved generated image to be displayed on the user interface. (Li, Figure 1; Figure 9; PNG media_image3.png 246 641 media_image3.png Greyscale PNG media_image4.png 806 791 media_image4.png Greyscale ) Regarding claim 9, Li does not disclose The method of claim 1, wherein the denial is a discard or a hold of the generated image. Wang teaches The method of claim 1, wherein the denial is a discard or a hold of the generated image. (2.3 Identifying NSFW Content “The Stable Diffusion Discord server prohibits generating NSFW images (StabilityAI, 2022a). Also, Stable Diffusion has a built-in NSFW filter that automatically blurs generated images if it detects NSFW content. However, we find DIFFUSIONDB still includes NSFW images that were not detected by the built-in filter or removed by server moderators. To help researchers filter these images, we apply state-of-the-art NSFW classifiers to compute NSFW scores for each prompt and image. Researchers can determine a suitable threshold to filter out potentially unsafe data for their tasks.” As both Li and Wang are from the same field of endeavor, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Li to include wherein the denial is a discard or a hold of the generated image, in the context of using diffusion model to create images, according to the teaching of Wang, in order to filter out potentially unsafe or harmful content (Fig. 2 of Wang). Regarding claim 10, the combination of Li and Wang discloses The method of claim 1, wherein the trained diffusion model is located on a server operably coupled to the user interface. (Li, Abstract “We achieve so by introducing efficient network architecture and improving step distillation.”; 1 Introduction “Through the improved Step distillation and network architecture development for the diffusion model, our introduced model, SnapFusion, generates a 512 × 512 image from the text on mobile devices in less than 2 seconds, while with image quality similar to Stable Diffusion v1.5 [4] (see example images from our approach in Fig. 1).”; 2.2 Benchmark and Analysis “Here we comprehensively study the parameter and computation intensity of the SD-v1.5. The in-depth analysis helps us understand the bottleneck to deploying text-to-image diffusion models on mobile devices from the scope of network architecture and algorithm paradigms. Meanwhile, the micro-level breakdown of the networks serves as the basis of the architecture redesign and search.”) Regarding claim 11, the combination of Li and Wang discloses The method of claim 8, wherein the trained diffusion model comprises a trained student diffusion model distilled with any one or more of backward distillation, shifted reconstruction loss, or noise correction. (Li, 1 Introduction “Finally, we explore the training strategies for step distillation, especially the best teacher-student paradigm for training the on-device model.”; Figure 3; 4 Step Distillation “We follow the research direction of step distillation [33], where the inference steps are reduced by distilling the teacher, e.g., at 32 steps, to a student that runs at fewer steps, e.g., 16 steps.”; C Detailed Derivations of Step Distillation; 5.2 Ablation Analysis “Namely, cross-attention is responsible for semantic coherency (Fig. 5(c)-(e)), while ResNet blocks capture local information and are critical to the reconstruction of details (Fig. 5(f)-(h)), especially in the output upsampling stage.”) Regarding claims 12 and 19, claim 12 is the system claims of method claim 1 except a non-transitory memory comprising instructions stored thereon; and at least one processor, operably coupled to the non-transitory memory, configured to execute the instructions comprising (5 Experiment of Li) and claim 19 is the non-transitory computer readable medium claim of method claim 1. SnapFusion (1 Introduction of Li)is a text-to image diffusion model running on a mobile phone, so the form of non-transitory readable medium is implicit. The claims are similar in scope to claim 1 and they are rejected under similar rationale as claim 1. Regarding claims 13-14 and 20, claims 13-14 are system claims of method claims 2 and 3 and claim 20 is the non-transitory computer readable medium claim (SnapFusion (1 Introduction of Li) is a text-to image diffusion model running on a mobile phone, so the form of non-transitory readable medium is implicit.) of claim 8. The claims are similar in scope to claims 2-3 and 8 respectively and they are rejected under similar rationale as claims 2-3 and 8 respectively. Claims 4 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Li, Yanyu, et al. "Snapfusion: Text-to-image diffusion model on mobile devices within two seconds." Advances in Neural Information Processing Systems 36 (2023): 20662-20678.;IDS REF (hereinafter Li) in view of Wang, Zijie J., et al. "Diffusiondb: A large-scale prompt gallery dataset for text-to-image generative models." Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers). 2023. (hereinafter Wang) and further in view of Liu, Vivian. "Beyond text-to-image: Multimodal prompts to explore generative AI." Extended abstracts of the 2023 CHI conference on human factors in computing systems. 2023. (hereinafter Liu). Regarding claim 4, the combination of Li and Wang does not disclose The method of claim 1, wherein: the text prompt comprises a first text prompt and a second text prompt; and the generated image of the second text prompt is different from the generated image of the first text prompt. Liu teaches The method of claim 1, wherein: the text prompt comprises a first text prompt and a second text prompt; and the generated image of the second text prompt is different from the generated image of the first text prompt. (Figure 3; See “a sports car built like a lego building block set” and generate image. See “a single sports car built like a lego building block. set view from top” and generate image. Examiner’s note: the generated image of the second text prompt (“a single sports car built like a lego building block. set view from top” is different from the generate image of the first text prompt (“a sports car built like a lego building block set”)) PNG media_image5.png 808 328 media_image5.png Greyscale ) As Li, Wang and Liu are from the same field of endeavor, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination of Li and Wang to include the text prompt comprises a first text prompt and a second text prompt; and the generated image of the second text prompt is different from the generated image of the first text prompt, in the context of text-to-image generation, according to the teaching of Liu, in order to provide great utility in using text prompts and that AI augmented workflows can increase momentum on creative tasks for end users (Abstract of Liu). Regarding claim 15, claim 15 is the system claim of method claim 4 and is accordingly rejected under same rationale. Claims 5, 7, 16 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Li, Yanyu, et al. "Snapfusion: Text-to-image diffusion model on mobile devices within two seconds." Advances in Neural Information Processing Systems 36 (2023): 20662-20678.; IDS REF (hereinafter Li) in view of Wang, Zijie J., et al. "Diffusiondb: A large-scale prompt gallery dataset for text-to-image generative models." Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers). 2023. (hereinafter Wang) and further in view of Inognic “Controlling Typeahead Search Trigger in the Advanced Lookup - Microsoft Dynamics 365 CRM Tips and Tricks," 22 July 2021. https://www.inogic.com/blog/2021/07/controlling-typeahead-search-trigger-in-the-advanced-lookup/; IDS REF (hereinafter Inognic). Regarding claim 5, the combination of Li and Wang disclose The method of claim 1, further comprising: prior to transmitting the text prompt to the trained diffusion model, (Li, Figure 3) The combination of Li and Wang does not disclose determining the text prompt comprises a threshold number of characters. Inognic teaches determining the text prompt comprises a threshold number of characters. ("Using the “Minimum number of characters to trigger typeahead search” field, you can trigger a typeahead search in the Search box present in the Advanced Lookup window based on the number of characters you enter.") As Li, Wang and Inognic are from the same field of endeavor, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination of Li and Wang to include determining the text prompt comprises a threshold number of characters, in the context of entering text, according to the teaching of Inognic, in order to set minimum number of characters while entering text prompts to trigger image generation. Regarding claim 7, the combination of Li and Wang does not disclose The method of claim 5, wherein the determining the text prompt comprises monitoring a predetermined amount of time elapsed after receiving the text prompt. Inognic teaches The method of claim 5, wherein the determining the text prompt comprises monitoring a predetermined amount of time elapsed after receiving the text prompt. (In addition to the same, the “Delay between character inputs that will trigger a search” field is used to trigger a typeahead search based on delay occurred while entering the input characters in the Search box present in the Advanced Lookup window.”) As Li, Wang and Inognic are from the same field of endeavor, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination of Li and Wang to include wherein the determining the text prompt comprises monitoring a predetermined amount of time elapsed after receiving the text prompt, in the context of entering text, according to the teaching of Inognic, in order to trigger search after the determined amount of time and to prevent delay after receiving the text prompt. Regarding claims 16 and 18, claims 16 and 18 are system claims of method claims 5 and 7. The claims are similar in scope to claims 5 and 7 and they are rejected under similar rationale as claims 5 and 7 respectively. Claims 6 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Li, Yanyu, et al. "Snapfusion: Text-to-image diffusion model on mobile devices within two seconds." Advances in Neural Information Processing Systems 36 (2023): 20662-20678.; IDS REF (hereinafter Li) in view of Wang, Zijie J., et al. "Diffusiondb: A large-scale prompt gallery dataset for text-to-image generative models." Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers). 2023. (hereinafter Wang) and further in view of Inognic “Controlling Typeahead Search Trigger in the Advanced Lookup - Microsoft Dynamics 365 CRM Tips and Tricks," 22 July 2021. https://www.inogic.com/blog/2021/07/controlling-typeahead-search-trigger-in-the-advanced-lookup/; IDS REF (hereinafter Inognic) and Bhatia et al. (US 20120278350 A1) (hereinafter Bhatia). Regarding claim 6, the combination of Li, Wang, and Igocnic does not disclose The method of claim 5, wherein the determining the text prompt comprises evaluating the text prompt for one or more characters indicating insufficient data to generate the image. Bhatia discloses The method of claim 5, wherein the determining the text prompt comprises evaluating the text prompt for one or more characters indicating insufficient data to generate the image. ([0028]-[0029] “Consider the phrase “president of usa.” Each of the possible bi-grams from this phrase (“president of”, “of usa”) starts or ends with a stop word and is, thus, an incomplete phrase and not desirable as query completions. One possible solution can be to remove all the stop-words from corpus before extracting N-grams. However, removing stop-words may also lead to loss of semantics and make the resulting suggestions harder to understand. For example, compare “president usa” with “president of usa” and “president in usa.” If one removes the stop-words, the second and third phrase will both reduce to the first phrase even though they mean different things. In order to avoid such difficulties, in accordance with at least one embodiment of the invention, instead of skipping over and discarding the adjacent words, whenever a stop-word is encountered there is a jump over to the next word, and the stop-word is retained. Accordingly, the resulting phrases do not start or end with stop words. Thus, in the above example, there will only be one bi-gram (“president of usa”). Note that now the order of an N-gram is not the number of words in the N-gram, but the number of non stop-words.” As Li, Wang, Inognic, and Bhatia are from the same field of endeavor, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination of Li, Wang, and Inognic to include wherein the determining the text prompt comprises evaluating the text prompt for one or more characters indicating insufficient data, in the context of entering text, according to the teaching of Bhatia, in order to understand the complete meaning of text prompt entered by users. Thus, it would have been further obvious to one of ordinary skill in the art at the effective filing date of the claimed invention to indicate insufficient data to generate the image by the teaching of Bhatia, to increase the probability that the image produced by LI in view of Wang and Inognic is in accordance the user’s prompt. Regarding claim 17, claim 17 is the system claim of method claim 6 and is accordingly rejected under same rationale. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Hyorim Park whose telephone number is (571)272-3859. The examiner can normally be reached Monday - Friday. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alicia Harrington can be reached at (571) 272-2330. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Hyorim Park/Examiner, Art Unit 2615 /ALICIA M HARRINGTON/Supervisory Patent Examiner, Art Unit 2615
Read full office action

Prosecution Timeline

Mar 28, 2025
Application Filed
Sep 03, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12675952
IMAGE PROCESSING APPARATUS, IMAGE PROCESSING METHOD, AND STORAGE MEDIUM
2y 1m to grant Granted Jul 07, 2026
Study what changed to get past this examiner. Based on 1 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
75%
Grant Probability
99%
With Interview (+37.5%)
2y 0m (~5m remaining)
Median Time to Grant
Low
PTA Risk
Based on 4 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month