Prosecution Insights
Last updated: October 02, 2026
Application No. 18/222,725

VISION TRANSFORMER FOR IMAGE GENERATION

Final Rejection §101§103
Filed
Jul 17, 2023
Priority
Dec 02, 2022 — provisional 63/429,933
Examiner
SORRIN, AARON JOSEPH
Art Unit
2672
Tech Center
2600 — Communications
Assignee
NVIDIA Corporation
OA Round
2 (Final)
76%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 76% — above average
76%
Career Allowance Rate
57 granted / 75 resolved
+14.0% vs TC avg
Strong +42% interview lift
Without
With
+42.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
32 currently pending
Career history
102
Total Applications
across all art units

Statute-Specific Performance

§101
20.0%
-20.0% vs TC avg
§103
37.1%
-2.9% vs TC avg
§102
14.1%
-25.9% vs TC avg
§112
28.0%
-12.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 75 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) submitted on 06/08/2026 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Response to Arguments Objections to the Specification and Claims are withdrawn. Rejections under 35 USC 112(b) and 35 USC 112(d) are withdrawn. Applicant's arguments regarding Prior Art rejections have been fully considered but they are not persuasive. On Pages 7-9 of Remarks argues that Dhariwal in view of Bao do not disclose “calculating one or more attention values based on the one or more time embeddings. In particular, Applicant argues: “In regard to claim 1, Dhariwal in view of Bao fails to teach calculating one or more attention values based on the one or more time embeddings, as in Applicant's claim. The Office Action acknowledges that Dhariwal fails to describe calculating one or more attention values based on the one or more time embeddings and relies on Bao, citing FIG. 1 and section 2, paragraph 1 where Bao mentions treating time embeddings as tokens. However, Bao whether considered singly or in combination with Dhariwal fails to mention anything about calculating attention values based on time embeddings. Bao does not describe any manner in which time embeddings are used to calculate attention values. In fact, Bao is completely silent regarding how attention values are calculated. The Office Action also states that FIG. 1 of Bao "shows that embeddings (i.e., time embeddings) are used to calculate the multi-head attention" (parenthesis in original, Office Action, page 12, lines 1-5). However, just because Bao treats time embeddings as tokens does not teach that attention values are calculated based on the time embedding tokens - especially since Bao is silent regarding how any attention values are calculated. For example, the portion of FIG. 1 relied on does not show the calculation of attention values based on time embeddings. Instead, FIG. 1 merely shows an "embeddings" block feeding into a transformer block that also has a "Multi-Head Attention" block. Bao does not mention that the embeddings block in FIG. 1 includes time embeddings (as opposed to label embeddings or noisy patches which Bao also treats as embeddings). Bao also fails to describe how the embeddings are used within FIG. 1. Merely including an embeddings block within a block diagram does not teach the specific feature of calculating one or more attention values based on the one or more time embeddings, as in Applicant's claim. Moreover, the combination of Dhariwal and Bao also fails to teach this feature of Applicant's claim. Since Dhariwal and Bao are both silent regarding calculating one or more attention values based on the one or more time embeddings, no combination of the two would include such a feature.” Bao is very clear that time embeddings are used for the calculation of attention. Figure 1, as referenced in the Non-Final Rejection, shows that embeddings are input, normalized, and fed into a multi-head attention block. A multi-head attention block, by definition, calculates attention values. These embeddings are explicitly described as including time embeddings in both the caption and also Section 2 Paragraph 1, “We first attempt to train a diffusion model using a vanilla ViT [3] on CIFAR10. For simplicity, we treat everything including the time embedding, label embedding and patches of the noisy image as tokens. With carefully tuned hyperparameters, a 13-layer ViT of size 41M achieves a FID 5.97, which is significantly better than 20.20 of the prior ViT-based diffusion models [18]. We conjecture that this is mainly because our model is larger.” Attention values are therefore calculated based on time embeddings. In view of the above, Applicant’s arguments regarding other independent and dependent claims are moot. Applicant's arguments regarding 35 USC 101 rejections have been fully considered but they are not persuasive. On Page 9, Applicant argues, “The Office Action rejected claims 1-20 under 35 U.S.C. § 101 as allegedly directed to a judicial exception. Applicant traverses this rejection for at least the following reason. Applicant respectfully disagrees with this rejection. However, Applicant has amended the claims to further clarify eligibility. Applicant notes that determining time embeddings associated with a resolution level of an input image (that has noise), and calculating attention values based on the time embeddings, as well as generating an output image based on the calculated attention values including reducing the noise from the input image by denoising the input image is not an abstract idea. To the contrary, the combination of features recited in Applicant's claims is a specific solution and practical application of reducing noise in images. The claims recite a specific "ordered combination" of features that is "significantly more" than an abstract idea. Given the claims reflect a specific ordered combination of features for a practical application and improvement in computer technology, as explained above, the pending claims are patent eligible. Moreover, give the improvement in computer technology, the recent precedential PTAB Decision Ex Parte Desjardins et al., Appeal 2024-000567 (ARP Sept. 26, 2025) (precedential) is directly on point.” The Applicant’s arguments and Specification are silent regarding the current challenges/limitations that exists in any particular field, and how the limitations of the presented claims are solving/overcoming them. Merely reducing noise does not amount to a specific solution or practical application. The ordered combination of steps amounts to reducing noise based on attention values and embeddings, which is a highly generic order of operations in the field of image processing using routinely used attention models. It is also unclear from Applicant’s arguments what the “improvement in computer technology” is in the claims. Specification The abstract of the disclosure is objected to because the amended abstract was not provided on a separate sheet. A corrected abstract of the disclosure is required and must be presented on a separate sheet, apart from any other text. See MPEP § 608.01(b). Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101. Claim 1 is rejected under 35 U.S.C. 101 because the claimed invention is directed to the abstract idea of generating a denoised image using mathematical processes, without significantly more. The claim recites: “A computer-implemented method comprising: determining one or more time embeddings associated with a resolution level of an input image having noise; calculating one or more attention values based on the one or more time embeddings; and generating an output image having reduced noise” The limitations, as drafted, are processes that, under their broadest reasonable interpretation, amount to mathematical calculations. Time embeddings can be determined mathematically, calculating attention values using the time embeddings amounts to performing mathematical operations to vectors, and generating the output image amounts to a weighted sum. This judicial exception is not integrated into a practical application. In particular, the claim describes a computer-implemented method. The computer is recited at a high level of generality such that it amounts to any generic computer. Accordingly, the additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements are recited at a high-level of generality. It is therefore a judicial exception that is not integrated into a practical application, and does not include additional elements that are sufficient to amount to significantly more than the judicial exception. This claim is not patent eligible. Claim 2 is rejected under 35 U.S.C. 101 because the claimed invention is directed to the abstract idea of calculating attention values at a plurality of stages, where each stage has a pixel number based on a portion of the input image, which amounts to further mathematical processes. The claim is not patent eligible. Claim 3 is rejected under 35 U.S.C. 101 because the claimed invention is directed to the abstract idea of calculating the attention values based on a machine learning model. The machine learning model is recited at a high level of generality such that it amounts to a generic machine learning model, which fails to meaningfully limit the performance of the abstract idea. The claim is not patent eligible. Claim 4 is rejected under 35 U.S.C. 101 because the claimed invention is directed to the abstract idea of describing the time embeddings as time steps of a generic machine learning model. The machine learning model is recited at a high level of generality such that it amounts to a generic machine learning model, which fails to meaningfully limit the performance of the abstract idea. The claim is not patent eligible. Claim 5 is rejected under 35 U.S.C. 101 because the claimed invention is directed to the abstract idea of outputting an image which amounts insignificant extra-solution activity. The claim is not patent eligible. Claim 6 is rejected under 35 U.S.C. 101 because the claimed invention is directed to the abstract idea of performing calculation using a GPU, which is recited at a high level of generality such that it amounts to no more than a generic GPU. The claim is not patent eligible. Claim 7 is rejected under 35 U.S.C. 101 because the claimed invention is directed to the abstract idea of calculating attention values using time embeddings (mathematical process) and self-attention blocks of a diffusion model. The self-attention blocks of a diffusion model is recited at a high-level of generality such that it amounts to blocks of a generic diffusion model. The claim is not patent eligible. Claims 8-14 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a non-transitory computer readable storage medium storing instructions analogous to the method of claims 1-7. The non-transitory computer readable storage medium is recited at a level of generality such that it fails to meaningfully limit the performance of the abstract idea. The claims are not patent eligible. Claims 15-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a system with a processor, with elements analogous to the method of claims 1-7. The processor is recited at a level of generality such that it fails to meaningfully limit the performance of the abstract idea. The claims are not patent eligible. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1, 3-8, and 10-15 and 17-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Dhariwal (Diffusion Models Beat GANs on Image Synthesis) in view of Bao (All are Worth Words: a ViT Backbone for Score-based Diffusion Models). Regarding claim 1, Dhariwal teaches “A computer-implemented method comprising: determining one or more time embeddings associated with a resolution level of an input image having noise; calculating one or more attention values based on the one or more time embeddings;” (Dhariwal, Algorithms 1 and 2 each show input of a noisy image xT and output of a denoised image x0. Also, alternatively, note that the algorithms disclose a looping process wherein each xt for a plurality of timesteps represent images that could be considered input images.; Tables 3 and 5 show FID (Fréchet Inception Distance) which indicates quality of de-noising an initial noisy image in the context of the paper; Section 3.1 further describes, “We also experiment with a layer [43] that we refer to as adaptive group normalization (AdaGN), which incorporates the timestep and class embedding into each residual block after a group normalization operation [69], similar to adaptive instance norm [27] and FiLM [48]. We define this layer as AdaGN(h,y) = ys GroupNorm(h)+yb, where his the intermediate activations of the residual block following the first convolution, and y = [ys,yb] is obtained from a linear projection of the timestep and class embedding. We had already seen AdaGN improve our earliest diffusion models, and so had it included by default in all our runs. In Table 3, we explicitly ablate this choice, and find that the adaptive group normalization layer indeed improved FID. Both models use 128 base channels and 2 residual blocks per resolution, multi-resolution attention with 64 channels per head, and BigGAN up/downsampling, and were trained for 700K iterations. In the rest of the paper, we use this final improved model architecture as our default: variable width with 2 residual blocks per resolution, multiple heads with 64 channels per head, attention at 32, 16 and 8 resolutions, BigGAN residual blocks for up and downsampling, and adaptive group normalization for injecting timestep and class embeddings into residual blocks.” Note that timestep embeddings are injected into each residual block, and there are two residual blocks per resolution, therefore all timesteps are associated with a resolution.) While Dhariwal discloses calculating attention values at different resolutions, (Dhariwal, Section 3.1 Paragraphs 2-3, “We had already seen AdaGN improve our earliest diffusion models, and so had it included by default in all our runs. In Table 3, we explicitly ablate this choice, and find that the adaptive group normalization layer indeed improved FID. Both models use 128 base channels and 2 residual blocks per resolution, multi-resolution attention with 64 channels per head, and BigGAN up/downsampling, and were trained for 700K iterations.In the rest of the paper, we use this final improved model architecture as our default: variable width with 2 residual blocks per resolution, multiple heads with 64 channels per head, attention at 32, 16 and 8 resolutions, BigGAN residual blocks for up and downsampling, and adaptive group normalization for injecting timestep and class embeddings into residual blocks.”) it does not disclose “calculating one or more attention values based on the one or more time embeddings;” Bao teaches “calculating one or more attention values based on the one or more time embeddings;” (Bao, Figure 1 and Section 2 Paragraph 1, “We first attempt to train a diffusion model using a vanilla ViT [3] on CIFAR10. For simplicity, we treat everything including the time embedding, label embedding and patches of the noisy image as tokens. With carefully tuned hyperparameters, a 13-layer ViT of size 41M achieves a FID 5.97, which is significantly better than 20.20 of the prior ViT-based diffusion models [18]. We conjecture that this is mainly because our model is larger.” Figure 1, right, shows that embeddings (i.e. time embeddings) are used to calculate the multi-head attention.) It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to incorporate the use of time embeddings for multi-head attention calculation, taught by Bao, into the attention calculation of Dhariwal. The motivation for doing so would have been to improve the overall denoising by calculating time-contextualized attention values. Diffusion models for denoising depend on time step, for example each time step is associated with a different level of noise. Computing attention weights based on time step thus enables improved denoising by factoring in the respective stage. Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Dhariwal with the above teaching of Bao to fully disclose, “calculating one or more attention values based on the one or more time embeddings;” Dhariwal in view of Bao further disclose “and generating an output image having reduced noise compared to the input image by denoising the input image based on the calculated one or more attention values.” (Dhariwal, Algorithms 1 and 2 each show input of a noisy image xT and output of a denoised image x0. Also, alternatively, note that the algorithms disclose a looping process wherein each iteration outputs an image that could also be mapped to output images; Tables 3 and 5 show FID (Fréchet Inception Distance) which indicates quality of an output denoised image; Figure 1 and its caption show output denoised images using the diffusion model; similarly, see Figure 6.) Regarding claim 3, Dhariwal in view of Bao teach “The computer-implemented method of claim 1,” “wherein calculating the one or more attention values comprises using one or more machine learning models.” (Dhariwal, Section 3.1 Paragraphs 2-3, “We had already seen AdaGN improve our earliest diffusion models, and so had it included by default in all our runs. In Table 3, we explicitly ablate this choice, and find that the adaptive group normalization layer indeed improved FID. Both models use 128 base channels and 2 residual blocks per resolution, multi-resolution attention with 64 channels per head, and BigGAN up/downsampling, and were trained for 700K iterations. In the rest of the paper, we use this final improved model architecture as our default: variable width with 2 residual blocks per resolution, multiple heads with 64 channels per head, attention at 32, 16 and 8 resolutions, BigGAN residual blocks for up and downsampling, and adaptive group normalization for injecting timestep and class embeddings into residual blocks.”) Regarding claim 4, Dhariwal in view of Bao teach “The computer-implemented method of claim 1,” “wherein the one or more time embeddings represent a time step of one or more layers of a machine learning model. (Dhariwal, Section 3.1 Paragraphs 1-3, “We also experiment with a layer [43] that we refer to as adaptive group normalization (AdaGN), which incorporates the timestep and class embedding into each residual block after a group normalization operation [69], similar to adaptive instance norm [27] and FiLM [48]. We define this layer as AdaGN(h,y) = ys GroupNorm(h)+yb, where his the intermediate activations of the residual block following the first convolution, and y = [ys,yb] is obtained from a linear projection of the timestep and class embedding. We had already seen AdaGN improve our earliest diffusion models, and so had it included by default in all our runs. In Table 3, we explicitly ablate this choice, and find that the adaptive group normalization layer indeed improved FID. Both models use 128 base channels and 2 residual blocks per resolution, multi-resolution attention with 64 channels per head, and BigGAN up/downsampling, and were trained for 700K iterations. In the rest of the paper, we use this final improved model architecture as our default: variable width with 2 residual blocks per resolution, multiple heads with 64 channels per head, attention at 32, 16 and 8 resolutions, BigGAN residual blocks for up and downsampling, and adaptive group normalization for injecting timestep and class embeddings into residual blocks.” Accordingly, the timestep embeddings represent a timestep of the residual block, which is a layer of a machine learning model.) Regarding claim 5, Dhariwal in view of Bao teach “The computer-implemented method of claim 1,” “further comprising, causing the output image to be presented.” (Dhariwal, Figures 1 and 6 (middle).) Regarding claim 6, Dhariwal in view of Bao teach “The computer-implemented method of claim 1,” “wherein the one or more one or more attention values are calculated using one or more graphics processing units (GPUs). (Dhariwal, Section A.1 discloses the use of the NVIDIA Tesla V100 GPU.) Regarding claim 7, Dhariwal in view of Bao teach “The computer-implemented method of claim 1,” “wherein the one or more attention values are generated using one or more self-attention blocks of a diffusion model and the one or more time embeddings.” (Dhariwal, Section 3.1 Paragraph 3, “In the rest of the paper, we use this final improved model architecture as our default: variable width with 2 residual blocks per resolution, multiple heads with 64 channels per head, attention at 32, 16 and 8 resolutions, BigGAN residual blocks for up and downsampling, and adaptive group normalization for injecting timestep and class embeddings into residual blocks.”; Bao, Figure 1 and Section 2 Paragraph 1, “We first attempt to train a diffusion model using a vanilla ViT [3] on CIFAR10. For simplicity, we treat everything including the time embedding, label embedding and patches of the noisy image as tokens. With carefully tuned hyperparameters, a 13-layer ViT of size 41M achieves a FID 5.97, which is significantly better than 20.20 of the prior ViT-based diffusion models [18]. We conjecture that this is mainly because our model is larger.” Figure 1, right, shows that embeddings (i.e. time embeddings) are used to calculate the multi-head attention. Note that this was incorporated with rationale and motivation in the rejection of claim 1. As Dhariwal and Bao are combined in the rejection of claim 1, Dhariwal teaches self-attention blocks (32, 16, and 8 resolutions) of a diffusion model where attention values are generated, and Bao teaches the use of time embeddings in this attention value generation.) Regarding claims 8 and 10-14, these claims recite a non-transitory computer readable storage medium storing thereon executable instructions corresponding to the steps recited in Claims 1 and 3-7. Therefore, the recited programming instructions of these claims are mapped to the analogous steps in the corresponding method claims. Additionally, the rationale and motivation to combine the Dhariwal and Bao references apply here. Finally, Dhariwal in view of Bao discloses non-transitory computer readable storage medium storing thereon executable instructions (Dhariwal, the abstract discloses that the programming instructions are released at github. Storage on github servers amount to the storage of the instructions on a non-transitory computer readable storage medium.) Regarding claims 15 and 17-20, these claims recite a system comprising a processor with elements corresponding to the steps recited in Claims 1, 3, 4, 6, and 7. Therefore, the recited elements of these claims are mapped to the analogous steps in the corresponding method claims. Additionally, the rationale and motivation to combine the Dhariwal and Bao references apply here. Finally, Dhariwal in view of Bao discloses a system comprising a processor (Dhariwal, Section A.1 discloses the use of the NVIDIA Tesla V100 GPU.) Claim(s) 2, 9, and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Dhariwal in view of Bao further in view of Zhang (VSA: Learning Varied-Size Window Attention in Vision Transformers). Regarding claim 2, Dhariwal in view of Bao teach “The computer-implemented method of claim 1,” While Dhariwal in view of Bao disclose that calculating the one or more attention values comprise determining attention scores of at least a first and second number of pixels, (Dhariwal, Section 3.1 Paragraphs 2-3, “We had already seen AdaGN improve our earliest diffusion models, and so had it included by default in all our runs. In Table 3, we explicitly ablate this choice, and find that the adaptive group normalization layer indeed improved FID. Both models use 128 base channels and 2 residual blocks per resolution, multi-resolution attention with 64 channels per head, and BigGAN up/downsampling, and were trained for 700K iterations.In the rest of the paper, we use this final improved model architecture as our default: variable width with 2 residual blocks per resolution, multiple heads with 64 channels per head, attention at 32, 16 and 8 resolutions, BigGAN residual blocks for up and downsampling, and adaptive group normalization for injecting timestep and class embeddings into residual blocks.” Note that attention is calculated at multiple stages, each stage having a different number of pixels (32, 16, and 8 resolution).), Dhariwal in view of Bao do not expressly disclose that the numbers of pixels are corresponding to respective portions of the input image Zhang teaches the numbers of pixels at each stage of a self-attention model being based on respective portions of the input image (Zhang, Page 3 Paragraph 1, Figure 1b, and Figure 3b, “To this end, we propose a novel Varied-Size Window Attention (VSA) mechanism to learn adaptive window configurations from data. Different from the previous window-based transformers where query, key, and value tokens are all sampled from the same window as shown in Figure 1(a), VSA employs a window regression module to predict the size and location of the target window based on the tokens within each default window. Then, the key and values tokens are sampled from the target window. By adopting VSA independently for each attention head, it enables the attention layers to model long-term dependencies, capture rich context from diverse windows, and promote information exchange among overlapped windows, as illustrated in Figure 1(b). VSA is an easy-to-implementation module that can replace the window attention in state of-the-art representative models with minor modifications and negligible extra computational cost while improving their performance by a large margin, e.g., 1.1% for Swin-T on ImageNet classification. In addition, the performance gain increases when using larger images for training and test, as shown in Figure 2. With the larger images as input, Swin-T with predefined window sizes cannot adapt to large objects well, and the improvement brought by enlarging image sizes is marginal, i.e., a gain of 0.3% from 224 × 224 to 480 × 480. In contrast, the performance gain of VSA over Swin-T increases significantly from 1.1% to 1.9%, owing to the varied-size window attention. Besides, as VSA can effectively promote information exchange across overlapped windows via token sampling, it does not need the shifted windows mechanism in Swin.” As shown in Figure 3b, window size (number of pixels) varies based on respective portions of the input image.) It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to incorporate the target window-based attention calculation of Zhang into the self-attention stages of Dhariwal in view of Bao. The motivation for doing so would have been to generate more accurate attention values that are adaptive to image regional context. Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Dhariwal in view of Bao with the above additional teaching of Zhang to fully disclose, “wherein calculating the one or more attention values comprise determining attention scores of a plurality of portions of the input image, a first portion of the plurality of portions having a first number of pixels and a second portion of the plurality of portions having a second number of pixels. Regarding claim 9, this claim recites a non-transitory computer readable storage medium storing thereon executable instructions corresponding to the steps recited in claim 2. Therefore, the recited programming instructions of this claim are mapped to the analogous steps in the corresponding method claim. Note that while there are minor differences in wording between claims 2 and 9, these differences do not change the meaning of the analogous limitations in any meaningful way that prevents the rejection of claim 2 from applying directly to the rejection of claim 9. Additionally, the rationale and motivation to combine the Dhariwal, Bao, and Zhang references apply here. Finally, Dhariwal in view of Bao further in view of Zhang discloses non-transitory computer readable storage medium storing thereon executable instructions (Dhariwal, the abstract discloses that the programming instructions are released at github. Storage on github servers amount to the storage of the instructions on a non-transitory computer readable storage medium.) Regarding claim 16, this claim recites a system comprising a processor with elements corresponding to the steps recited in Claim 2. Therefore, the recited elements of this claim are mapped to the analogous steps in the corresponding method claim. Note that while there are minor differences in wording between claims 2 and 16, these differences do not change the meaning of the analogous limitations in any meaningful way that prevents the rejection of claim 2 from applying directly to the rejection of claim 16. Additionally, the rationale and motivation to combine the Dhariwal and Bao references apply here. Finally, Dhariwal in view of Bao discloses a system comprising a processor (Dhariwal, Section A.1 discloses the use of the NVIDIA Tesla V100 GPU.) Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to AARON JOSEPH SORRIN whose telephone number is (703)756-1565. The examiner can normally be reached Monday - Friday 9am - 5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Sumati Lefkowitz can be reached at (571) 272-3638. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /AARON JOSEPH SORRIN/ Examiner, Art Unit 2672 /SUMATI LEFKOWITZ/Supervisory Patent Examiner, Art Unit 2672
Read full office action

Prosecution Timeline

Jul 17, 2023
Application Filed
Apr 02, 2026
Non-Final Rejection mailed — §101, §103
Aug 03, 2026
Response Filed
Sep 24, 2026
Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743886
COMPONENT MOUNTING SYSTEM, IMAGE PROCESSING DEVICE, IMAGE PROCESSING METHOD, AND IMAGE PROCESSING SYSTEM
2y 10m to grant Granted Sep 22, 2026
Patent 12731269
ROTATION STATE ESTIMATION APPARATUS, METHOD THEREOF, AND PROGRAM
3y 1m to grant Granted Sep 08, 2026
Patent 12720096
IMAGE PROCESSING DEVICE, IMAGE DISPLAY SYSTEM, IMAGE PROCESSING METHOD, AND RECORDING MEDIUM
3y 0m to grant Granted Aug 25, 2026
Patent 12718945
A RADIOMIC-BASED MACHINE LEARNING ALGORITHM TO RELIABLY DIFFERENTIATE BENIGN RENAL MASSES FROM RENAL CELL CARCINOMA
2y 10m to grant Granted Aug 25, 2026
Patent 12705851
Method And System For Detecting, Quantifying, And Attributing Gas Emissions Of Industrial Assets
3y 6m to grant Granted Aug 11, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
76%
Grant Probability
99%
With Interview (+42.0%)
3y 0m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 75 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month