Prosecution Insights
Last updated: October 02, 2026
Application No. 18/736,340

Token Pruning for Image Generation

Non-Final OA §103
Filed
Jun 06, 2024
Examiner
STATZ, BENJAMIN TOM
Art Unit
2611
Tech Center
2600 — Communications
Assignee
Adobe Inc.
OA Round
3 (Non-Final)
29%
Grant Probability
At Risk
3-4
OA Rounds
6m
Est. Remaining
59%
With Interview

Examiner Intelligence

Grants only 29% of cases
29%
Career Allowance Rate
2 granted / 7 resolved
-33.4% vs TC avg
Strong +30% interview lift
Without
With
+30.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
17 currently pending
Career history
39
Total Applications
across all art units

Statute-Specific Performance

§101
3.5%
-36.5% vs TC avg
§103
67.7%
+27.7% vs TC avg
§102
7.7%
-32.3% vs TC avg
§112
11.8%
-28.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 7 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant’s arguments, see pg. 10, filed 06/15/2026, with respect to the objection to claim 15 have been fully considered and are persuasive. The objection to claim 15 has been withdrawn. Applicant’s arguments, see pg. 10-14, filed 06/15/2026, with respect to the rejection of claims 1 and 15 under 35 U.S.C. 103 have been fully considered and are persuasive. The rejection of claims 1 and 15 under 35 U.S.C. 103 has been withdrawn. Applicant's arguments, see pg. 14-21, filed 06/15/2026, with respect to the rejection of claim 12 under 35 U.S.C. 103 have been fully considered but they are not persuasive. Firstly, applicant argues that the combination of Bolya in view of Thorsley and Yang does not teach the limitation “pruning the plurality of tokens according to a denoising-steps-aware pruning (DSAP) schedule based on the denoising timestep” as “one coherent claim feature”. Applicant states, correctly, that Yang does not teach pruning tokens, and that Thorsley does not teach a DSAP schedule. However, the method by which a set of tokens’ importance score is determined is not necessarily intrinsically connected to the method that selects which tokens are analyzed to determine their importance score. Additionally, applicant cites paragraph [0045] of the specification, which describes “a token pruning scheme applied within each denoising step of the diffusion process, and an adaptive pruning schedule across different denoising steps, such as a Denoising-Steps-Aware Pruning (DSAP) schedule”, and paragraphs [0042] and [0029] which describe the adaptive pruning in further detail, as evidence of distinction compared to the cited prior art. However, the concept of adaptive pruning as claimed is taught by Yang as seen in fig. 3, which explains how machine learning is used to determine a level of pruning for each step of the denoising process. Secondly, applicant argues that there is no motivation to combine the teachings of Thorsley, which teaches token pruning based on an attention map, and Yang, which teaches optimizing pruning using a DSAP schedule, stating that there is no “reasonable expectation of success”. Applicant notes again that the cited references do not mention DSAP token pruning as a whole, and argues that combining the inventions of Thorsley and Yang would require “substantial redesign” of Yang. In response to applicant's argument that the invention of Yang would require "substantial redesign" in order to be combined with the teachings of Thorsley, the test for obviousness is not whether the features of a secondary reference may be bodily incorporated into the structure of the primary reference; nor is it that the claimed invention must be expressly suggested in any one or all of the references. Rather, the test is what the combined teachings of the references would have suggested to those of ordinary skill in the art. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981). In this case, Yang teaches a method of optimizing a pruning procedure. Since the goal of token pruning itself, in this case, is to improve the efficiency and resource consumption of a diffusion model, one of ordinary skill in the art would have reasonably examined methods of streamlining the pruning system itself to further improve efficiency and reduce resource consumption. Therefore, despite not explicitly pertaining to tokens, Yang may be considered analogous art to the invention. One of ordinary skill in the art may not have attempted to incorporate the entire invention of Yang into the invention of Thorsley as-is, but may have been motivated to adapt some of its methodology. Therefore, the rejection of claim 12 under 35 U.S.C. 103 is maintained. Applicant's arguments, see pg. 14-21, filed 06/15/2026, with respect to the newly added claim 21 have been fully considered but they are not persuasive. The claim limitations are not explicitly stated by Yang, but they are implied based on the stated characteristics of the invention (see “Claim Rejections - 35 USC § 103” section for more detail). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 12-14 and 21 is/are rejected under 35 U.S.C. 103 as being unpatentable over Bolya et al. ("Token Merging for Fast Stable Diffusion". arXiv preprint (30 Mar 2023). https://doi.org/10.48550/arXiv.2303.17604; hereinafter "Bolya") in view of Thorsley et al. (US 20220374766 A1; hereinafter "Thorsley") and Yang et al. ("Denoising Diffusion Step-aware Models". arXiv preprint (24 May 2024). https://arxiv.org/abs/2310.03337v5; hereinafter "Yang"). Regarding claim 12, Bolya teaches: obtain an input prompt (fig. 2 “Prompt”, pg. 2 “Each transformer block has the standard self attention [22] and multi-layer perception (mlp) modules, with the addition of a cross attention module to condition on the prompt.”); generate a plurality of tokens for an attention layer of the generative machine learning model based on an intermediate noise map (pg. 2 “Stable Diffusion uses a U-Net [16] with transformer-based blocks. Thus, it first encodes the current noised image as a set of tokens, then passes it through a series of transformer blocks.”; transformer blocks contain attention layers as shown in fig. 2); denoise, using the generative machine learning model, the intermediate noise map based on the pruned set of tokens to obtain a denoised map (pg. 2 “Diffusion models [4,20,21] generate images by repeatedly denoising some initial noise over some number of diffusion steps.”); and generate, using the generative machine learning model, a synthetic output based on the denoised map (examples of image generation included in all figures except for fig. 2). Bolya does not explicitly teach: A non-transitory computer readable medium storing code for a generative machine learning model, the code comprising instructions executable by at least one processor to: generate, using the attention layer, an attention map based on the plurality of tokens; prune the plurality of tokens based on the attention map to obtain a pruned set of tokens; Thorsley teaches: A non-transitory computer readable medium storing code for a generative machine learning model, the code comprising instructions executable by at least one processor ([0073] “Embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., one or more modules of computer-program instructions, encoded on computer-storage medium for execution by, or to control the operation of data-processing apparatus… The computer-storage medium can also be, or be included in, one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices).”) to: generate, using the attention layer, an attention map based on the plurality of tokens ([0024] “Transformer deep-learning models utilize a self-attention mechanism, differentially weighing the significance of each part of the input data. A self-attention mechanism provides context for any position in the input sequence, allowing for parallelization during training of a deep neural network. Self-attention mechanisms receive input sequences that are converted into tokens. Each token is provided a probability and scored based on a relevant metric (also known as the attention-probability and as the attention-score or importance-score)”; [0031] to [0037] describes using a self-attention mechanism to calculate an attention matrix (or map) which is later used to relate the attention of each token to each other token, specifically referencing a “scaled dot-product attention layer 131”); prune the plurality of tokens based on the attention map to obtain a pruned set of tokens ([0043] “Each head has the attention probabilities determined using the process described with respect to FIG. 1C with the scaled dot-product attention layers 131… The token importance score s.sup.(l) 340 may be calculated from the mean 330 of the attention probability over all the heads.”, where element 131 references the process of generating an attention matrix/map); [0044] “Alternatively, using a threshold token pruning operation 360, a token may be kept at 362 if the importance score of the token exceeds an absolute threshold value, and may be pruned at 364 otherwise.”). Bolya and Thorsley are both analogous to the claimed invention because they are in the same field of increasing the efficiency of a transformer neural network. Furthermore, Bolya also suggests (but does not explicitly teach) the use of token pruning to speed up neural network processing (pg. 1 col. 2 section 1 “Introduction”), as taught by Thorsley. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the invention of Bolya with the teachings of Thorsley to use the token pruning method of Thorsley, rather than the token merging method of Bolya, to reduce computations associated with tokens. The motivation would have been to try an alternate method of increasing the efficiency of the neural network. The combination of Bolya and Thorsley does not explicitly teach: prune the plurality of tokens based on the attention map to obtain a pruned set of tokens by identifying a denoising timestep and pruning the plurality of tokens according to a denoising-steps-aware pruning (DSAP) schedule based on the denoising timestep. Yang teaches: prune the plurality of tokens based on the attention map to obtain a pruned set of tokens by identifying a denoising timestep (pg. 3 “Conventional diffusion models use a heavy network for all denoising steps, ignoring the differences between the steps. However, we hypothesize that some steps in the generation process may be easier to process, and that even a lightweight network can handle these steps.”; section 4 “Step-aware Network for Diffusion Models” discusses determining which steps to prune) and pruning the plurality of tokens according to a denoising-steps-aware pruning (DSAP) schedule based on the denoising timestep (pg. 2 “We show the feasibility of this approach and present the Denoising Diffusion Step-aware Models (DDSM). In DDSM, the neural network is variable and slimmable at different steps to avoid redundant computation at unimportant steps. We determine the capacity of the neural network at each step via evolutionary search and prune the network to various scales accordingly.”). Yang is analogous to the claimed invention because it is in the same field of pruning a diffusion model to improve efficiency. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the invention of Bolya in view of Thorsley with the teachings of Yang to apply Yang’s steps-aware pruning to the token pruning system of Bolya in view of Thorsley. The motivation would have been “to avoid redundant computation at unimportant steps” (Yang pg. 2). Regarding claim 13, the combination of Bolya in view of Thorsley and Yang teaches: The non-transitory computer readable medium of claim 12, wherein pruning the plurality of tokens comprises: computing an importance score for each of the plurality of tokens based on the attention map (Thorsley [0043] “Each head has the attention probabilities determined using the process described with respect to FIG. 1C with the scaled dot-product attention layers 131… The token importance score s.sup.(l) 340 may be calculated from the mean 330 of the attention probability over all the heads.”, where element 131 references the process of generating an attention matrix/map); and identifying a threshold importance score (Thorsley [0044] “Alternatively, using a threshold token pruning operation 360, a token may be kept at 362 if the importance score of the token exceeds an absolute threshold value, and may be pruned at 364 otherwise.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the invention of Bolya in view of Thorsley and Yang with the additional teachings of Thorsley to compute importance scores and a threshold importance score for token pruning. The motivation would have been to include a method of establishing metrics or criteria by which tokens are selected for pruning. Regarding claim 14, the combination of Bolya in view of Thorsley and Yang teaches: The non-transitory computer readable medium of claim 12, wherein generating the synthetic output comprises: identifying a plurality of pruned tokens (Bolya pg. 2 section 3.1 “Defining Unmerging”: “And if we have information about what tokens we merged, we have enough information to then unmerge those same tokens.”; pruned tokens could be identified in the same manner); generating a plurality of replacement tokens corresponding to the plurality of pruned tokens; and adding the plurality of replacement tokens to the pruned set of tokens to obtain an augmented set of tokens (Bolya pg. 2 section 3.1 “Defining Unmerging” describes duplicating existing merged tokens, which were identified based on similarity, and copying them into previously filled token slots; a similar procedure could be applied to pruned tokens). Regarding claim 21, the combination of Bolya in view of Thorsley and Yang teaches: The non-transitory computer readable medium of claim 12, wherein the DSAP schedule comprises: pruning the plurality of tokens more aggressively at a later denoising timestep than the denoising timestep (Yang fig. 3 shows how the optimal pruning strategy is determined for each step; a later step may have a more heavily pruned network than an earlier step). Allowable Subject Matter Claims 1-11, 15-16, and 18-20 allowed. The following is a statement of reasons for the indication of allowable subject matter: Regarding claim 1, the prior art of record, alone or in combination, fails to teach, disclose, or render obvious the combination of limitations set forth in claim 1. The closest prior art includes: Bolya et al. ("Token Merging for Fast Stable Diffusion". arXiv preprint (30 Mar 2023). https://doi.org/10.48550/arXiv.2303.17604), Thorsley et al. (US 20220374766 A1), Wang et al. ("Zero-TPrune: Zero-Shot Token Pruning through Leveraging of the Attention Graph in Pre-Trained Transformers". arXiv preprint (27 May 2023). https://arxiv.org/abs/2305.17328v1), and Guo et al. ("ELIP: Efficient Language-Image Pre-training with Fewer Vision Tokens". arXiv preprint (17 Nov 2023). https://arxiv.org/abs/2309.16738v2). The documents cited above, in combination, partially teach the limitations of claim 1 (see citations from previous rejections), but fail to teach or render obvious the following subject matter in claim 1: “wherein the plurality of tokens are pruned by performing a weighted page rank process using a function that computes an importance score of a query token based on the attention map and an importance score of a key token having a modality different from the query token”. In particular, Wang et al. teaches computing importance scores for token pruning via a weighted page rank process, but it only teaches a self-attention mechanism and not a cross-attention mechanism, meaning that all tokens are of the same modality. On the other hand, Guo et al. teaches token pruning with a multi-modal attention map, but it provides no indication of compatibility with the weighted page rank process of Wang. Therefore, even in combination, the cited references do not render claim 1 obvious. Regarding claims 2-11, they are dependent on claim 1 which has been identified as allowable. Regarding claim 15, it is interpreted to be a corresponding apparatus to the method of claim 1, which has been identified as allowable. Therefore, claim 15 is allowable. Regarding claims 16 and 18-20, they are dependent on claim 15 which has been identified as allowable. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to BENJAMIN STATZ whose telephone number is (571)272-6654. The examiner can normally be reached Mon-Fri 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tammy Goddard can be reached at (571)272-7773. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /BENJAMIN TOM STATZ/ Examiner, Art Unit 2611 /TAMMY GODDARD/ Supervisory Patent Examiner, Art Unit 2611
Read full office action

Prosecution Timeline

Show 4 earlier events
Mar 24, 2026
Response Filed
Apr 14, 2026
Final Rejection mailed — §103
Jun 05, 2026
Interview Requested
Jun 11, 2026
Examiner Interview Summary
Jun 11, 2026
Applicant Interview (Telephonic)
Jun 15, 2026
Request for Continued Examination
Jun 16, 2026
Response after Non-Final Action
Sep 23, 2026
Non-Final Rejection mailed — §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
29%
Grant Probability
59%
With Interview (+30.0%)
2y 10m (~6m remaining)
Median Time to Grant
High
PTA Risk
Based on 7 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month