Prosecution Insights
Last updated: October 01, 2026
Application No. 18/882,629

DUAL FORMULATION FOR A COMPUTER VISION RETENTION MODEL

Non-Final OA §103§112
Filed
Sep 11, 2024
Priority
Oct 03, 2023 — provisional 63/542,256
Examiner
TORRES, JOSE
Art Unit
Tech Center
Assignee
NVIDIA Corporation
OA Round
1 (Non-Final)
82%
Grant Probability
Favorable
1-2
OA Rounds
11m
Est. Remaining
94%
With Interview

Examiner Intelligence

Grants 82% — above average
82%
Career Allowance Rate
532 granted / 649 resolved
+22.0% vs TC avg
Moderate +12% lift
Without
With
+12.2%
Interview Lift
resolved cases with interview
Typical timeline
2y 12m
Avg Prosecution
17 currently pending
Career history
670
Total Applications
across all art units

Statute-Specific Performance

§101
8.8%
-31.2% vs TC avg
§103
46.1%
+6.1% vs TC avg
§102
20.1%
-19.9% vs TC avg
§112
20.0%
-20.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 649 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Drawings The drawings are objected to as failing to comply with 37 CFR 1.84(p)(4) because reference character “732” has been used to designate both Software and Job Scheduler in Figure 7, and on Paragraph [0081] lines 1-2, Paragraph [0081] line 4, Paragraph [0081] line 5, Paragraph [0081] line 10, Paragraph [0081] lines 16-17, and Paragraph [0082] line 1. Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance. Specification The disclosure is objected to because of the following informalities: Paragraph [0052] line 1: “The output of the retention encoder with L layers ( Z L n is user in a classification MLP” should read -- The output of the retention encoder with L layers ( Z L n ) is user in a classification MLP -- Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 15 and 17 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 15 recites the limitation “the 1D retention decay” in line 1. There is insufficient antecedent basis for this limitation in the claim. Claim 17 recites the limitation “the 2D retention decay” in line 1. There is insufficient antecedent basis for this limitation in the claim. Appropriate correction is required. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 2, 4, 6, 11, 13, 20, 21, 23, 25, 31, and 32 are rejected under 35 U.S.C. 103 as being unpatentable over Hussain et al. (“Vision Transformer and Deep Sequence Learning for Human Activity Recognition in Surveillance Videos”, Computational Intelligence and Neuroscience, Volume 2022, Issue 1, Article ID 3454167, 4 April 2022, pp. 1-10) in view of Sultana et al. (U.S. Pub. No. 2024/0203098). As to claims 1, 20 and 31, Hussain et al. teaches a method (i.e., “Human Activity Recognition”, Abstract, p. 1), comprising: at a device (i.e., “The proposed method is implemented using Python (3.6 version) in Spyder integrated development environment”, 3. Experimental Results and Discussion, p. 6): processing/process an input representation of an image, using a retention encoder of a computer vision model (See for example, “In the first step, surveillance cameras capture video streams which are then fed to pre-trained ViT-Base-16 for frame-level spatial features extraction”, 2. The Proposed Activity Recognition Framework, p. 3) operating in accordance with a first formulation that at least in part includes a recurrent formulation (See for example, “the term Δt represents the input overtime and sigmoid activation function is represented by ∂ . Their weights and bias terms are represented by w and b, respectively. The forget gate ft at time t keeps the information of the previous frame that is needed and discard it otherwise. The output gate Ot keeps the information of the upcoming step and R is the recurrent unit having tanh activation function. It is computed from the input of the current frame and state of the previous frame St-1”, 2.4. Learning Long-Range Temporal Dependencies via LSTM, p. 5), to generate an encoded representation of the image (i.e., “The spatial features are stacked together to create a resultant feature vector from 30 consecutive frames”, 2. The Proposed Activity Recognition Framework, 2.1. Features Extraction Using Vision Transformer, 2.2. Linear Embedding Layer, p. 3); and processing/process the encoded representation of the image, using a multilayer perceptron (MLP) of the computer vision model (i.e., “The resultant embedding patches of Z0 (in (1)) are fed to the transformer encoder module, which consists of L identical layers as shown in Figure 2. Furthermore, each module is divided into two components such as multihead self-attention (MSA) block and multilayer perceptron (MLP)”, 2.3. Vision Transformer Encoder, p. 3), to generate an output particular to a defined computer vision task (i.e., 2.5. Modeling Human Activity Recognition via ViT and Multilayer LSTM, pp. 5-6). However, Hussain et al. does not explicitly disclose the system, comprising: a non-transitory memory storage comprising instructions, and one or more processors in communication with the memory, wherein the one or more processors execute the instructions/the non-transitory computer-readable media storing computer instructions which to be executed by one or more processors of a device, and wherein the computer vision model has been trained with the retention encoder operating in accordance with a second formulation that is a parallel formulation. Sultana et al. teaches a method (i.e., “method for a machine learning engine”, Abstract), comprising: at a device (i.e., “computer system 1900”, Paragraph [0113])/a system, comprising: a non-transitory memory storage (i.e., “a non-transitory computer readable storage medium storing program instructions”, Paragraph [0017]; and “The computer system 1900 includes main memory 1902, typically random access memory RAM, which contains the software being executed by the processing cores 1950 and GPUs 1912, as well as a non-volatile storage device 1904 for storing data and the software programs”, Paragraph [0113]) comprising instructions, and one or more processors (i.e., “one or more central processing units (CPU) 1950”, Paragraph [0113]) in communication with the memory, wherein the one or more processors execute the instructions/a non-transitory computer-readable media (i.e., “a non-transitory computer readable storage medium storing program instructions”, Paragraph [0017]) storing computer instructions which to be executed by one or more processors of a device (), and a computer vision model (i.e., “ERM-SDVIT”, Paragraph [0049]) that has been trained with a retention encoder operating in accordance with a second formulation that is a parallel formulation (See for example, “The present vision transformer seeks to achieve a performance gain and in fact provides plug-and-play DG approach for ERM-ViT, referred to as self-distillation for ViTs. The present self-distillation approach explicitly trains the model in a manner that facilitates exploiting cross-domain transferable features (FIG. 2). The present self-distillation approach alleviates the overfitting to the source domains by easing the mapping problem via non-zero entropy supervision of multiple feature pathways that are trained by a comparison with a sequential transformer pathway that is performed in parallel”, Paragraph [0055]). Hussain et al. and Sultana et al. are analogous art because they are from the field of digital image processing using computer vision models. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify Hussain et al. by incorporating the system, comprising: a non-transitory memory storage comprising instructions, and one or more processors in communication with the memory, wherein the one or more processors execute the instructions/the non-transitory computer-readable media storing computer instructions which to be executed by one or more processors of a device, and the computer vision model has been trained with the retention encoder operating in accordance with a second formulation that is a parallel formulation, as taught by Sultana et al. The suggestion/motivation for doing so would have been to implement the computer vision model on a conventional computing system, to avoid the introduction of any new parameters, and to alleviate the problem of overfitting source domains when processing images from different source domains. Therefore, it would have been obvious to combine Sultana et al. with Hussain et al. to obtain the invention as specified in claims 1, 20 and 31. As to claims 2 and 21, Hussain et al. teaches wherein the retention encoder includes a multi-head retention component (i.e., “each module is divided into two components such as multihead self-attention (MSA) block and multilayer perceptron (MLP)”, 2.3. Vision Transformer Encoder, p. 3). As to claims 4 and 23, Hussain et al. teaches wherein the retention encoder includes at least one layer comprised of a multi-head retention component and a multilayer perceptron (MLP) component (i.e., “each module is divided into two components such as multihead self-attention (MSA) block and multilayer perceptron (MLP)”, 2.3. Vision Transformer Encoder, p. 3). As to claim 6, Hussain et al. teaches wherein the first formulation includes only the recurrent formulation (i.e., “2.4. Learning Long-Range Temporal Dependencies via LSTM, p. 5). As to claims 11, 25 and 32, Hussain et al. teaches wherein the recurrent formulation computes retention based on at least one previous state (i.e., “The output gate Ot keeps the information of the upcoming step and R is the recurrent unit having tanh activation function. It is computed from the input of the current frame and state of the previous frame St-1”, 2.4. Learning Long-Range Temporal Dependencies via LSTM, p. 5). As to claim 13, Hussain et al. teaches wherein the input representation of the image is a sequence of patch and position embeddings having a class token appended at an end of the sequence (See for example, “it divides the input image into a number of patches that are linearly projected with learnable positional embedding to learn the order of patches followed by transformer encoder with multilayer perceptron for final classification”, 2.1. Features Extraction Using Vision Transformer, p. 3; and “Then, these embedded representations are concatenated together with learnable classification token vclass”, 2.2. Linear Embedding Layer, p. 3). Claims 5 and 24 are rejected under 35 U.S.C. 103 as being unpatentable over Hussain et al. in view of Sultana et al. as applied to claims 4 and 23 above, and further in view of Al-hammuri et al. (“Vision transformer architecture and applications in digital health: a tutorial and survey”, Visual Computing for Industry, Biomedicine, and Art, 10 July 2023, 6(1):14, pp. 1-28). The teachings of Hussain et al. and Sultana et al. have been discussed above. As to claims 5 and 24, Hussain et al. and Sultana et al. do not explicitly disclose wherein the retention encoder includes a plurality of layers each comprised of the multi-head retention component and the MLP component. Al-hammuri et al. teaches a retention encoder that includes a plurality of layers each comprised of the multi-head retention component and the MLP component (See for example, “Figure 2 shows a typical encoder architecture [8] that consists of a stack of N identical layers, with each layer containing two sublayers. The first sublayer performs the multihead self-attention (MSA), while the second sub-layer normalizes the output of the first sublayer and feeds it into the multilayer perceptron (MLP), which is a type of feedforward network”, Encoder architecture, p. 3). Hussain et al., Sultana et al. and Al-hammuri et al. are analogous art because they are from the field of digital image processing using computer vision models. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to further modify Hussain et al. and Sultana et al. by incorporating the retention encoder includes a plurality of layers each comprised of the multi-head retention component and the MLP component, as taught by Al-hammuri et al. The suggestion/motivation for doing so would have been to provide a computer vision model with superior performance. Therefore, it would have been obvious to combine Al-hammuri et al. with Hussain et al. and Sultana et al. to obtain the invention as specified in claims 5 and 24. Claims 18, 19, 29, 30, 36, and 37 are rejected under 35 U.S.C. 103 as being unpatentable over Hussain et al. in view of Sultana et al. as applied to claims 1, 20 and 31 above, and further in view of Carion et al. (“End-to-End Object Detection with Transformers”, arXiv:2005.12872v3, 20 May 2020, pp. 1-26). The teachings of Hussain et al. and Sultana et al. have been discussed above. As to claims 18, 19, 29, 30, 36, and 37, Hussain et al. and Sultana et al. do not explicitly disclose wherein the defined computer vision task is object detection and instance segmentation/semantic segmentation. Carion et al. teaches a defined computer vision task that is object detection and instance segmentation/semantic segmentation (See for example, Fig. 6, p. 13; “we visualize decoder attentions in Fig. 6 coloring attention maps for each predicted object in different colors. We observe that decoder attention is fairly local, meaning that it mostly attends to object extremities such as heads or legs. We hypothesise that after the encoder has separated instances via global attention, the decoder only needs to attend to the extremities to extract the class and object boundaries”, Number of decoder layers, pp. 10-11; and “Fig.9: Qualitative results for panoptic segmentation generated by DETR-R101. DETR produces aligned mask predictions in a unified manner for things and stuff”, p. 15). Hussain et al., Sultana et al. and Carion et al. are analogous art because they are from the field of digital image processing using computer vision models. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to further modify Hussain et al. and Sultana et al. by incorporating the defined computer vision task is object detection and instance segmentation/semantic segmentation, as taught by Carion et al. The suggestion/motivation for doing so would have been to accurately visualize panoptic segmentations in a unified manner. Therefore, it would have been obvious to combine Carion et al. with Hussain et al. and Sultana et al. to obtain the invention as specified in claims 18, 19, 29, 30, 36, and 37. Allowable Subject Matter Claims 3, 7-10, 12, 14, 16, 22, 26-28, and 33-35 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Claims 15 and 17 would be allowable if rewritten to overcome the rejection(s) under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), 2nd paragraph, set forth in this Office action and to include all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter: the closest prior art made of record fails to disclose, teach, and/or suggest, inter alia, the method of claim 1, and further comprising, wherein the first formulation includes a combination of the parallel formulation and the recurrent formulation, or wherein the parallel formulation computes retention without regard to at least one previous state, or wherein the retention encoder is configured for one-dimensional or two-dimensional retention, and the method of claim 2, and further comprising, wherein the multi-head retention component uses a causal retention decay mask, as claimed. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to JOSE M TORRES whose telephone number is (571)270-1356. The examiner can normally be reached Monday thru Friday; 10:00 AM to 6:00 PM EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Mehmood can be reached at 571-272-2976. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JOSE M TORRES/Examiner, Art Unit 2664 09/16/2026
Read full office action

Prosecution Timeline

Sep 11, 2024
Application Filed
Sep 18, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12725249
SYSTEMS AND METHODS FOR IDENTIFYING IMAGES OF POLYPS
3y 6m to grant Granted Sep 01, 2026
Patent 12690922
PATIENT-SPECIFIC MEDICAL SYSTEMS, DEVICES, AND METHODS
2y 0m to grant Granted Jul 28, 2026
Patent 12675875
MACHINE LEARNING ON MULTIMODAL CHEMICAL AND WHOLE SLIDE IMAGING DATA FOR PREDICTING CANCER LABELS DIRECTLY FROM TISSUE IMAGES
3y 0m to grant Granted Jul 07, 2026
Patent 12670587
AUTOMATIC ANNOTATION OF CONDITION FEATURES IN MEDICAL IMAGES
3y 0m to grant Granted Jun 30, 2026
Patent 12657708
IMAGE PROCESSING APPARATUS, METHOD FOR OPERATING IMAGE PROCESSING APPARATUS, AND PROGRAM FOR OPERATING IMAGE PROCESSING APPARATUS
2y 9m to grant Granted Jun 16, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
82%
Grant Probability
94%
With Interview (+12.2%)
2y 12m (~11m remaining)
Median Time to Grant
Low
PTA Risk
Based on 649 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month