Prosecution Insights
Last updated: August 17, 2026
Application No. 18/776,010

SYSTEMS, METHODS, AND APPARATUSES FOR IMPLEMENTING STEPWISE INCREMENTAL PRE-TRAINING FOR INTEGRATING DISCRIMINATIVE, RESTORATIVE, AND ADVERSARIAL LEARNING INTO AN AI MODEL

Non-Final OA §103§112
Filed
Jul 17, 2024
Priority
Jul 17, 2023 — provisional 63/514,037
Examiner
SHARIFF, MICHAEL ADAM
Art Unit
2672
Tech Center
2600 — Communications
Assignee
Arizona Board of Regents on Behalf of Arizona State University
OA Round
1 (Non-Final)
82%
Grant Probability
Favorable
1-2
OA Rounds
8m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 82% — above average
82%
Career Allowance Rate
101 granted / 124 resolved
+19.5% vs TC avg
Strong +24% interview lift
Without
With
+23.6%
Interview Lift
resolved cases with interview
Typical timeline
2y 9m
Avg Prosecution
17 currently pending
Career history
139
Total Applications
across all art units

Statute-Specific Performance

§101
10.1%
-29.9% vs TC avg
§103
47.9%
+7.9% vs TC avg
§102
20.4%
-19.6% vs TC avg
§112
18.1%
-21.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 124 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/process/file/efs/guidance/eTD-info-I.jsp. Claims 1, 7, and 13 are rejected on the ground of nonstatutory double patenting as being unpatentable over claim 1 of U.S. Patent Application No.: 2026/0119883 in which the issue fee has not been paid. Although the claims are not identical, they are not patently distinct from each other because they recite substantially the same limitations in slight broader verbiage. Claim 1 is used as the representative of claims 7 and 13 reciting similar limitations but different statutory categories. Summary of double patenting claims (the similarities are bolded and differences are italicized) Instant application Patent Application No.: 2026/0119883 1. A system comprising: a memory to store instructions; a processor to execute the instructions stored in the memory to perform the following operations: receiving at the system a training dataset comprising a plurality of medical images for training a unified artificial intelligence (AI) model; executing stepwise incremental pre-training operations to train the unified AI model, comprising: pre-training a discriminative encoder via discriminative learning, yielding a pre-trained discriminative encoder; attaching the pre-trained discriminative encoder to a restorative decoder to form a skip-connected encoder-decoder; pre-training the skip-connected encoder-decoder via joint discriminative and restorative learning, yielding the pre-trained discriminative encoder and a pre-trained restorative decoder; and associating the pre-trained skip-connected encoder-decoder with an adversarial encoder; finalizing training of the AI model by performing discriminative, restorative, and adversarial learning on the training dataset using the unified AI model, yielding the pre-trained discriminative encoder, the pre-trained restorative decoder, and a pre-trained adversarial encoder; and applying, via the unified AI model, each of discriminative, restorative, and adversarial learning operations through the discriminative encoder, the restorative decoder, and the adversarial encoder, generated via training of the unified AI model, for classifying and annotating medical images. 1. A framework for pre-training a self-supervised machine learning model for analysis of digital images, comprising: a discriminative encoder trained via discriminative learning, yielding a pre-trained discriminative encoder; a restorative decoder attached to the pre-trained discriminative encoder, forming a skip-connected encoder-decoder, the skip-connected encoder-decoder jointly trained via discriminative learning and restorative learning, yielding a pre-trained encoder-decoder; and an adversarial encoder associated with the pre-trained encoder-decoder and jointly trained with the pre-trained encoder-decoder via discriminative, restorative, and adversarial learning. Claims 1, 7, and 13 are rejected on the ground of nonstatutory double patenting as being unpatentable over claim 1 of U.S. Patent No.: 12,572,813 in which the issue fee has been paid. Although the claims are not identical, they are not patently distinct from each other because they recite substantially the same limitations in slight broader verbiage. Claim 1 is used as the representative of claims 7 and 13 reciting similar limitations but different statutory categories. Summary of double patenting claims (the similarities are bolded and differences are italicized) Instant application Patent No.: 12,572,813 1. A system comprising: a memory to store instructions; a processor to execute the instructions stored in the memory to perform the following operations: receiving at the system a training dataset comprising a plurality of medical images for training a unified artificial intelligence (AI) model; executing stepwise incremental pre-training operations to train the unified AI model, comprising: pre-training a discriminative encoder via discriminative learning, yielding a pre-trained discriminative encoder; attaching the pre-trained discriminative encoder to a restorative decoder to form a skip-connected encoder-decoder; pre-training the skip-connected encoder-decoder via joint discriminative and restorative learning, yielding the pre-trained discriminative encoder and a pre-trained restorative decoder; and associating the pre-trained skip-connected encoder-decoder with an adversarial encoder; finalizing training of the AI model by performing discriminative, restorative, and adversarial learning on the training dataset using the unified AI model, yielding the pre-trained discriminative encoder, the pre-trained restorative decoder, and a pre-trained adversarial encoder; and applying, via the unified AI model, each of discriminative, restorative, and adversarial learning operations through the discriminative encoder, the restorative decoder, and the adversarial encoder, generated via training of the unified AI model, for classifying and annotating medical images. 1. A method comprising: receiving a medical image; integrating Self-Supervised machine Learning (SSL) instructions for performing a discriminative learning operation, a restorative learning operation, and an adversarial learning operation into a model for processing the received medical image; configuring the model with each of a discriminative encoder, a restorative decoder, and an adversarial encoder; configuring each of the discriminative encoder and the restorative decoder of the model to be skip connected, forming an encoder-decoder of the model; performing step-wise incremental training to incrementally train each of the discriminative encoder, the restorative decoder, and the adversarial encoder, by performing the following operations: pre-training the discriminative encoder via discriminative learning; attaching the pre-trained discriminative encoder to the restorative decoder to configure the encoder-decoder as a pre-trained encoder-decoder of the model; training the pre-trained encoder-decoder of the model using joint discriminative and restorative learning; and associating the pre-trained encoder-decoder of the model with the adversarial encoder; performing training of the pre-trained encoder-decoder associated with the adversarial encoder through discriminative, restorative, and adversarial learning to render a trained model for the processing of the received medical image; and processing the medical image through the model using the trained model. Claim Objections Claims 4, 10, and 16 are objected to because of the following informalities: 1) the claim term “the United framework for 3D medical imaging” should recite “a United framework for 3D medical imaging” for proper antecedent basis; and 2) the claim term “Swein UNETR” should recite “Swin UNETR” for proper spelling. Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. Claims 5, 11, and 17 are rejected under 35 U.S.C. 112(b) as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor regards as the invention. Claim 5, as the representative claim for analysis, recites “training a discriminative encoder via discriminative learning; attaching the pre-trained discriminative encoder to a restorative decoder to form an encoder-decoder; training the encoder-decoder via combined discriminative and restorative learning; and associating a pre-trained auto-encoder with the adversarial-encoder; training the associated adversarial-encoder via full discriminative, restorative, and adversarial training”; this claim is indefinite because independent claim 1, from which claim 5 depends from already recites “pre-training a discriminative encoder via discriminative learning, yielding a pre-trained discriminative encoder; attaching the pre-trained discriminative encoder to a restorative decoder to form a skip-connected encoder-decoder; pre-training the skip-connected encoder-decoder via joint discriminative and restorative learning, yielding the pre-trained discriminative encoder and a pre-trained restorative decoder; and associating the pre-trained skip-connected encoder-decoder with an adversarial encoder”; therefore it is unclear whether the “training”, “a discriminative encoder”, “a restorative decoder”, “the encoder-decoder”, are the same as the “pre-training”, the discriminative encoder”, “the restorative decoder”, and “the encoder-decoder” already introduced in claim 1. Further, it appears the only further delineation in claim 5 that further specifies aspects of independent claim 1, without repeating aspects of independent claim 1 are the claim limitations “associating a pre-trained auto-encoder with the adversarial-encoder; training the associated adversarial-encoder via full discriminative, restorative, and adversarial training” where the auto-encoder is used as a specific form of the encoder-decoder introduced in independent claim 1; therefore, for examination, Examiner will be interpreting the claim as reciting “the system of claim 1, wherein the skip-connected encoder-decoder is an autoencoder that is associated with the adversarial-encoder; and wherein executing stepwise incremental pre-training operations to train the unified AI model further comprises: training the associated adversarial-encoder via full discriminative, restorative, and adversarial training”. Proper corrections are requested. Claims 4, 10, and 16 are rejected under 35 U.S.C. 112(b) as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor regards as the invention. The claims recite “performing self-supervised learning (SSL) training via 3D adapted SSL training operations including Rotation, Jigsaw, Rubik’s Cube, Deep Clustering, TransVW, MoCo, BYOL, PCRL, and Swein UNETR; and wherein each of the 3D adapted SSL training operations are augmented with supplemental 3D compatible components within the United framework for 3D medical imaging” which introduces numerous 3D adapted SSL training operations including in the conjunctive; however upon review of the present specification, para. [0036], para. [0051]-[0060], and FIG. 2A-2I show that each of these nine SSL training operations are separately used and change the model significantly between each operation and are never used concurrently together; further since the specification confirms a disjunctive claim interpretation, this is how Examiner will interpret the claims. Therefore, for examination, the claims shall recite “performing self-supervised learning (SSL) training via 3D adapted SSL training operations including Rotation, Jigsaw, Rubik’s Cube, Deep Clustering, TransVW, MoCo, BYOL, PCRL, or Swein UNETR; and wherein each of the 3D adapted SSL training operations are augmented with supplemental 3D compatible components within the United framework for 3D medical imaging”. Proper corrections are requested. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-3, 6-9, 12-15, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over non-patent literature "Multi-task Collaborative Pre-training and Individual-adaptive-tokens Fine-tuning: A Unified Framework for Brain Representation Learning"; arXiv preprint arXiv:2306.11378 (2023) (Jiang et al.) (hereinafter Jiang), in view of non-patent literature "Representation recovering for self-supervised pre-training on medical images"; 2023 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV); IEEE, 2023 (Yan et al.) (hereinafter Yan). Regarding claim 1, Jiang teaches a system comprising: a memory to store instructions; a processor to execute the instructions stored in the memory to perform the following operations: (Jiang, abstract: “Here, we propose MCIAT, a unified framework that combines Multi-task Collaborative pre-training and Individual-Adaptive-Tokens fine-tuning. Specifically, we first synthesize restorative learning, age prediction auxiliary learning and adversarial learning as a joint proxy task for deep semantic representation learning. Then, a mutual-attention-based token selection method is proposed to highlight discriminative features. The proposed MCIAT achieves state-of-the-art diagnosis performance on the ADHD-200 dataset compared with several sMRI-based approaches and shows superior generalization on the MCIC and OASIS datasets. Moreover, we studied 12 behavioral tasks and found significant associations between cognitive functions and MCIAT established representations, which interpretability of our proposed framework”; image processing/computer vision for medical images as laid out in the abstract of Jiang assumes using a computer having a processor and a memory to conduct the processing of the images to create a united framework for deep learning specifically using restorative learning, discriminative learning, and adversarial learning) receiving at the system a training dataset comprising a plurality of medical images for training a unified artificial intelligence (AI) model (Jiang, page 8, Section 6) Linear Prob Classification on ADHD-200; Table V; FIG. 6: “The linear-prob classification results of ADHD-200 are shown in Table V. In particular, we freeze the pre-trained models as the feature extractors and feed them with all the data in the ADHD-200 training and testing sets. Then, the encoded features of the training set are used to train a linear SVM as a classifier, which is tested on the encoded features of the testing set in each transformer layer.”; PNG media_image1.png 464 750 media_image1.png Greyscale ”); executing stepwise incremental pre-training operations to train the unified AI model, comprising: (Jiang, page 2, left-side col., para. 4; page 6, Section B. IV. Results and Discussion, Part B. Ablation Study, para. 1; page 2, right-side col., para. 2: “We propose a ViT-based MCIAT, which leverages restorative learning, age prediction auxiliary learning, and adversarial learning to foster Multi-task Collaborative pre-training and combines Individual-Adaptive Tokens fine-tuning realized by a mutual attention mechanism. Specifically, we first reshape an input structural magnetic resonance imaging (sMRI) into several patches, and a random mask divides the patches into visible patches and masked patches. The pre-training stage has four proxy tasks: (1) Infer the mask target from the visible context, forcing the model to build long-range dependencies and understand contextual semantic information. (2) Restore the distorted visible patches back to the original pixels, encouraging the model to focus on the structural information hidden in image details. (3) Predict the age of the input sample from the visible patches, naturally revealing the biological information generated during brain development and aging. (4) Reinforce the restoration quality by adversarial learning, making it essential for the model to establish a latent space that matches the distribution of brain sMRI to fool the discriminator with images reconstructed from the latent space.”; “We perform a thorough ablation study to show how each component contributes to MCIAT. For pre-training, we start with a ViT backbone and incrementally add restorative learning branch A, age prediction auxiliary learning, restorative learning branch B, and adversarial learning.”; “(1) We propose MCIAT, a unified framework that combines multi-task collaborative pre-training and individual-adaptive tokens fine-tuning, which exhibits superior performance in downstream tasks”) pre-training a discriminative encoder via discriminative learning, yielding a pre-trained discriminative encoder (Jiang, page 4, Section 2) Individual-Adaptive Tokens Fine-Tuning; page 6, Section 6) Linear-prob Classification on ADHD-200; FIG. 2; FIG. 1: “For fine-tuning, a new classifier with initialized weights is added to the pre-trained encoder, which can extract subtle yet discriminative features but needs to further highlight and strengthen task-specific features in fine-tuning. We propose a mutual-attention-based token selection (MATS) method and insert it into the pre-trained ViT backbone, where we adaptively and individually select the informative tokens from each transformer layer (see Fig. 2 for more details).”; “Our proposed MCIAT achieves a 66.67% linear-prob accuracy, surpassing most of the compared methods, indicating that the pre-trained encoder has learned how to represent biological patterns of the brain organization in healthy people and thus can discriminate intra-class variation and inter-class variation”; see pre-trained encoders in FIG. 2 and FIG. 1 below: PNG media_image2.png 543 955 media_image2.png Greyscale ; PNG media_image3.png 594 993 media_image3.png Greyscale ); attaching the pre-trained discriminative encoder to a restorative decoder to form an encoder-decoder; pre-training encoder-decoder via joint discriminative and restorative learning, yielding the pre-trained discriminative encoder and a pre-trained restorative decoder (Jiang, page 3, right-side col; FIG. 1; page 4, left-side col., para. 1-2: “ PNG media_image4.png 472 551 media_image4.png Greyscale ”; PNG media_image5.png 1080 534 media_image5.png Greyscale ”; as shown above, at the encoder step there is a loss (see eq. (1) above) that trains the discriminative aspect, there is an encoder/decoder restorative loss (see eq. (5)) that trains the combined encoder/decoder during the pre-training; see FIG. 1 above showing the restorative decoders attached to the discriminative encoders); and associating the pre-trained encoder-decoder with an adversarial encoder (Jiang, page 4, left-side col., para. 3; FIG. 1; page 2, left-side col., para. 4; page 4, left-side col., para. 4: “ PNG media_image6.png 199 537 media_image6.png Greyscale ”; “A shared ViT encoder, three independent lightweight decoders, and a convolutional neural network (CNN) discriminator are used for pre-training.”; see the discriminator shown in FIG. 1 above; in self-supervised learning, the encoder maps inputs to a latent space, the decoder reconstructs data or features, and the CNN discriminator classifies real versus fake representations or outputs, forcing high-fidelity learning; a Convolutional Neural Network (CNN) discriminator used alongside a self-supervised encoder-decoder for discriminative, restorative, and adversarial learning acts as the adversarial component of an adversarial encoder-decoder setup; however, it is technically the adversary/discriminator network, distinct from the primary encoder that maps inputs to latent space; therefore, the CNN discriminator meets the broadest reasonable interpretation for the claim term “adversarial encoder”; PNG media_image7.png 429 538 media_image7.png Greyscale ); finalizing training of the AI model by performing discriminative, restorative, and adversarial learning on the training dataset using the unified AI model, yielding the pre-trained discriminative encoder, the pre-trained restorative decoder, and a pre-trained adversarial encoder (Jiang, page 8, Section 6) Linear Prob Classification on ADHD-200; Table V; FIG. 6; page 5, left-side col., para. 2: “The linear-prob classification results of ADHD-200 are shown in Table V. In particular, we freeze the pre-trained models as the feature extractors and feed them with all the data in the ADHD-200 training and testing sets. Then, the encoded features of the training set are used to train a linear SVM as a classifier, which is tested on the encoded features of the testing set in each transformer layer.”; see Table V above; “In this study, 636 cognitively healthy adults from Cam-CAN were used for collaborative pre-training. For ADHD-200, we used the same general training/testing sets as in previous studies: 768 training subjects (488 normal control subjects (NCs), 280 ADHD patients) and 171 test subjects (94 NCs, 77 ADHD patients) collected from eight independent sites. For the MCIC and OASIS datasets, we performed ten-fold cross-validation on 190 subjects (99 NCs, 91 SZ patients) and 280 subjects (190 NCs, 90 preclinical ADs) and reported the averaged accuracy (ACC), sensitivity (SEN), specificity (SPE), and area under the receiver operating characteristic (ROC) curve (AUC)”); and applying, via the unified AI model, each of discriminative, restorative, and adversarial learning operations through the discriminative encoder, the restorative decoder, and the adversarial encoder, generated via training of the unified AI model, for classifying and annotating medical images (Jiang, page 5, Section V. Results and Discussion; page 9, left-side col., para. 1; FIG. 8; page 9, right-side col., para. 2: “A. Classification Performance: “The classification results of ADHD-200 are shown in Table I”; “In Fig. 8, we choose 6 subjects (3 ADHDs, 3 NCs) and visualize the locations of the tokens that are the first to be selected for each subject in each layer. It suggests that there are both inter-class variation and intra-class variation, indicating that the token selection strategy is highly individualized.”; PNG media_image8.png 276 622 media_image8.png Greyscale ; “Second, we verify the effectiveness of the MCIAT on multiple classification tasks, while a wide range of downstream tasks (i.e., brain tumor segmentation, multiple MRI modality synthesis, etc.) may make the performance of the model more convincing.”). Jiang fails to teach attaching the pre-trained discriminative encoder to a restorative decoder to form a skip-connected encoder-decoder. Yan teaches attaching the pre-trained discriminative encoder to a restorative decoder to form a skip-connected encoder-decoder (Yan, page 2687, left-side col., para. 1; right-side col. Section 4.3 Fine-tuning Stage; FIG. 2: “The contrastive pre-training stage is designed for two main purposes. One is to provide learnable feature maps, which will be utilized as inputs in the later generative pre-training stage. The other is to generate different levels of feature maps, which follows the common design of the U-Net model family. These feature maps will be used through skip connections to provide alternative paths for the gradient in the later fine-tuning stage.”; PNG media_image9.png 324 523 media_image9.png Greyscale ”; PNG media_image10.png 622 1028 media_image10.png Greyscale ). It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the step of attaching the pre-trained discriminative encoder to a restorative decoder to form an encoder-decoder, as taught by Jiang, to include attaching the pre-trained discriminative encoder to a restorative decoder to form a skip-connected encoder-decoder, as taught by Yang. The suggestion/motivation for doing so would have been that “By adding skip connections in U-Net from convolutional encoder Ɛc to convolutional decoder Dc, DSC score of both random initialized method and our RepRec get improvements by 1.47% and 1.82% respectively; this shows skip connections from encoder to decoder are also essential in segmentation tasks.” (Yang, page 2691, Section 5.4.2 Choices of Decoder Dc). Jiang, in view of Yang, teaches attaching the pre-trained discriminative encoder to a restorative decoder to form a skip-connected encoder-decoder; pre-training the skip-connected encoder-decoder via joint discriminative and restorative learning, yielding the pre-trained discriminative encoder and a pre-trained restorative decoder; and associating the pre-trained skip-connected encoder-decoder with an adversarial encoder (Yan, page 2687, left-side col., para. 1; right-side col. Section 4.3 Fine-tuning Stage; Jiang, page 3, right-side col; FIG. 1; page 4, left-side col., para. 3; FIG. 1; page 2, left-side col., para. 4; page 4, left-side col., para. 4; the encoder-decoder model of Jiang is modified by Yang to include the U-net neural network architecture which is an encoder/decoder model having skip-connections; both Jiang and Yang do self-supervised learning for medical images, pre-training of the encoder-decoder model, and use ViT (Vision Transformer) encoders). Therefore, it would have been obvious to combine Jiang, with Yang, to obtain the invention as specified in claim 1. Regarding claim 2, Jiang, in view of Yang, teaches the system of claim 1, further comprising: outputting the trained AI model for use with medical image analysis (Jiang, page 5, Section V. Results and Discussion; page 9, left-side col., para. 1; FIG. 8; page 9, right-side col., para. 2; see rejection of claim 1 above discussing medical images throughout; “Second, we verify the effectiveness of the MCIAT on multiple classification tasks, while a wide range of downstream tasks (i.e., brain tumor segmentation, multiple MRI modality synthesis, etc.) may make the performance of the model more convincing.”). Regarding claim 3, Jiang, in view of Yang, teaches the system of claim 1, wherein executing the stepwise incremental pre-training operations to train the unified AI model comprises performing self-supervised learning (SSL) training at each of the pre-training operations (Jiang, page 8, Section 5) Ablation Study for Representation Analysis; Fig. 3; FIG. 4: “We conduct a representation analysis ablation study to explore how each component affects the representation learning preferences of MCIAT. The results are illustrated in Fig. 3 and Fig. 4. We make the following observations: … (3) Visible portion restoration (restorative learning branch B) improves both the correlations of learned representations with cognitive functions and morphological estimates, suggesting that such a task can serve as a basic self-supervised pre-training component in a representation learning framework.”; PNG media_image11.png 295 1065 media_image11.png Greyscale ; PNG media_image12.png 292 1053 media_image12.png Greyscale ). Regarding claim 6, Jiang, in view of Yang, teaches the system of claim 1, wherein executing stepwise incremental pre-training operations to train the unified AI model generates as an output a stable trained AI model (Jiang, page 2, right-side col., para. 3; page 8, Section D. Representation Analysis for the Fine-tuned Models: “(2) We studied 12 behavioral tasks to explore the associations between deep brain representation and behavioral functions. We used multiple representation analysis and visualization methods to verify the interpretability of our MCIAT and demonstrate that the proposed model represents the semantic information relevant to cognitive functions involving comprehension, reasoning, memory, emotion, and learning.”; “In Section IV-C, we have observed that: (1) Correlations between MCIAT-extracted representations and cognitive functions tend to strengthen with increasing depth of the transformer layers, whereas for morphological estimates, the correlation strength tends to oscillate downward or is relatively stable, whether or not the effects of age are controlled for.”). With regards to claims 7-9, and 12, they recite the functions of the apparatus of claims 1-3 and 6, respectively, as processes. Thus, the analyses in rejecting claims 1-3 and 6 are equally applicable to claims 7-9 and 12, respectively. Regarding claim 13, Jiang teaches a non-transitory computer readable storage media having instructions stored thereupon that, when executed by a system having at least a processor and a memory therein, cause the processor to perform the following operations: (Jiang, abstract: “Structural magnetic resonance imaging (sMRI) provides accurate estimates of the brain's structural organization and learning invariant brain representations from sMRI is an enduring issue in neuroscience … Here, we propose MCIAT, a unified framework that combines Multi-task Collaborative pre-training and Individual-Adaptive-Tokens fine-tuning. Specifically, we first synthesize restorative learning, age prediction auxiliary learning and adversarial learning as a joint proxy task for deep semantic representation learning. Then, a mutual-attention-based token selection method is proposed to highlight discriminative features.”; it is implicit from the abstract, that this MCIAT model is an image processing deep learning framework executed on a computer with a processor a memory storing instructions to execute the method of discriminating features in medical images). With regards to the remaining limitations of independent claim 13, dependent claims 14-15, and 18, they recite the functions of the apparatus of claims 1-3 and 6, respectively, as non-transitory computer readable media storing instructions. Thus, the analyses in rejecting claims 1-3 and 6 are equally applicable to the remaining limitations of independent claim 13 and dependent claims 14-15, and 18, respectively. Claims 4, 10, and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Jiang, in view Yan, and in view of non-patent literature "Self-supervised pre-training of swin transformers for 3d medical image analysis"; 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE, 2022 (Tang et al.) (hereinafter Tang). Regarding claim 4, Jiang, in view of Yang, teaches the system of claim 1. Jiang, in view of Yang fails to teach wherein executing the stepwise incremental pre-training operations to train the unified AI model comprises performing self-supervised learning (SSL) training via 3D adapted SSL training operations including Rotation, Jigsaw, Rubik’s Cube, Deep Clustering, TransVW, MoCo, BYOL, PCRL, and Swein UNETR; and wherein each of the 3D adapted SSL training operations are augmented with supplemental 3D compatible components within the United framework for 3D medical imaging. Tang teaches wherein executing the stepwise incremental pre-training operations to train the unified AI model comprises performing self-supervised learning (SSL) training via 3D adapted SSL training operations including Rotation, Jigsaw, Rubik’s Cube, Deep Clustering, TransVW, MoCo, BYOL, PCRL, and Swein UNETR; and wherein each of the 3D adapted SSL training operations are augmented with supplemental 3D compatible components within the United framework for 3D medical imaging (Tang, abstract; FIG. 1: “Vision Transformers (ViT)s have shown great performance in self-supervised learning of global and local representations that can be transferred to downstream applications. Inspired by these results, we introduce a novel self-supervised learning framework with tailored proxy tasks for medical image analysis. Specifically, we propose: (i) a new 3D transformer-based model, dubbed Swin UNEt TRansformers (Swin UNETR), with a hierarchical encoder for self-supervised pretraining; (ii) tailored proxy tasks for learning the underlying pattern of human anatomy. We demonstrate successful pre-training of the proposed model on 5,050 publicly available computed tomography (CT) images from various body organs. The effectiveness of our approach is validated by fine-tuning the pre-trained models on the Beyond the Cranial Vault (BTCV) Segmentation Challenge with 13 abdominal organs and segmentation tasks from the Medical Segmentation Decathlon (MSD) dataset.”; PNG media_image13.png 575 568 media_image13.png Greyscale ; “Swein UNETR” is misspelled in the claim; see claim objection of claim 4 above for correct spelling of “Swin UNETR” as recited in Tang; Examiner interpretation of this claim is in the disjunctive (Rotation, Jigsaw, Rubik’s Cube, Deep Clustering, TransVW, MoCo, BYOL, PCRL, or Swin UNETR” rather than conjunctive using “and” including all the SSL training operations; see 35 U.S.C. 112(b) for reasoning behind this claim interpretation). It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the step of executing the stepwise incremental pre-training operations to train the unified AI model, as taught by Jiang, in view of Yang, to include performing self-supervised learning (SSL) training via 3D adapted SSL training operations including Rotation, Jigsaw, Rubik’s Cube, Deep Clustering, TransVW, MoCo, BYOL, PCRL, or Swin UNETR, and wherein each of the 3D adapted SSL training operations are augmented with supplemental 3D compatible components within the United framework for 3D medical imaging, as taught by Tang. The suggestion/motivation for doing so would have been that “medical image analysis has not benefited from these advances in general computer vision due to: (1) large domain gap between natural images and medical imaging modalities, like computed tomography (CT) and magnetic resonance imaging (MRI); (2) lack of cross-plane contextual information when applied to volumetric (3D) images (such as CT or MRI); the latter is a limitation of 2D transformer models for various medical imaging tasks such as segmentation; prior studies have demonstrated the effectiveness of supervised pre-training in medical imaging for different applications, but creating expert-annotated 3D medical datasets at scale is a non-trivial and time-consuming effort; to tackle these limitations, we propose a novel self-supervised learning framework for 3D medical image analysis; first, we propose a new architecture dubbed Swin UNEt TRansformers (Swin UNETR) with a Swin Transformer encoder that directly utilizes 3D input patches” (Tang, page 20699, left-side col., para. 3-4). Therefore, it would have been obvious to combine Jiang and Yang, with Tang, to obtain the invention as specified in claim 4. With regards to claim 10, it recites the functions of the apparatus of claim 4 as a process. Thus, the analysis in rejecting claim 4 is equally applicable to claim 10. With regards to claim 16, it recites the functions of the apparatus of claim 4 as a non-transitory computer readable medium storing instructions. Thus, the analysis in rejecting claim 4 is equally applicable to claim 16. Claims 5, 11, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Jiang, in view Yan, and in view of non-patent literature "Masked autoencoders are scalable vision learners"; 2022 IEEE/CVF conference on computer vision and pattern recognition (CVPR); IEEE, 2022 (He et al.) (hereinafter He). Regarding claim 5, Jiang, in view of Yan, teaches the system of claim 1, wherein executing stepwise incremental pre-training operations to train the unified AI model comprises: training a discriminative encoder via discriminative learning (Jiang, page 4, Section 2) Individual-Adaptive Tokens Fine-Tuning; page 6, Section 6) Linear-prob Classification on ADHD-200; FIG. 2; FIG. 1; see rejection of claim 1 above); and attaching the pre-trained discriminative encoder to a restorative decoder to form an encoder-decoder; training the encoder-decoder via combined discriminative and restorative learning (Jiang, page 3, right-side col; FIG. 1; see rejection of claim 1 above); and training the associated adversarial-encoder via full discriminative, restorative, and adversarial training (Jiang, page 4, left-side col., para. 3; FIG. 1; page 2, left-side col., para. 4; page 4, left-side col., para. 4; see collaborative pre-training above for total combined loss used for training, and Ladv in FIG. 1 that is the loss used to train the CNN discriminator (adversarial encoder)). Jiang, in view of Yan, fails to teach a pre-trained auto-encoder. He teaches a pre-trained auto-encoder (He, abstract; FIG. 1: “This paper shows that masked autoencoders (MAE) are scalable self-supervised learners for computer vision. Our MAE approach is simple: we mask random patches of the input image and reconstruct the missing pixels. It is based on two core designs. First, we develop an asymmetric encoder-decoder architecture, with an encoder that operates only on the visible subset of patches (without mask tokens), along with a lightweight decoder that reconstructs the original image from the latent representation and mask tokens. Second, we find that masking a high proportion of the input image, e.g., 75%, yields a nontrivial and meaningful self-supervisory task. Coupling these two designs enables us to train large models efficiently and effectively: we accelerate training (by 3× or more) and improve accuracy. Our scalable approach allows for learning high-capacity models that generalize well: e.g., a vanilla ViT-Huge model achieves the best accuracy (87.8%) among methods that use only ImageNet-1K data. Transfer performance in downstream tasks outperforms supervised pretraining and shows promising scaling behavior.”; PNG media_image14.png 624 642 media_image14.png Greyscale ). It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the pre-trained encoder-decoder, as taught by Jiang, in view of Yang, to be a pre-trained auto-encoder, as taught by He. The suggestion/motivation for doing so would have been that for “a masked autoencoder (MAE) for visual representation learning … shifting the mask tokens to the small decoder in our asymmetric encoder-decoder results in a large reduction in computation; under this design, a very high masking ratio (e.g., 75%) can achieve a win-win scenario: it optimizes accuracy while allowing the encoder to process only a small portion (e.g., 25%) of patches; this can reduce overall pre-training time by 3x or more and likewise reduce memory consumption, enabling us to easily scale our MAE to large models.” (He, page 15980, left-side col., para. 3; page 15980, right-side col., para. 1). Jiang, in view of Yang, and in view He, teaches associating a pre-trained auto-encoder with the adversarial-encoder (Jiang, page 4, left-side col., para. 3; FIG. 1; page 2, left-side col., para. 4; page 4, left-side col., para. 4; He, abstract; FIG. 1; both Jiang, in view of Yang, and He, teach self-supervised learning for image processing using encoder-decoder models and specifically ViT (Vision Transformer) encoders; however, unlike Jiang, in view of Yang, He teaches utilizing an autoencoder rather than a separate encoder/decoder network; therefore, when Jiang, in view of Yang is modified by He, the auto-encoder may be used rather than a distinct encoder and decoder; an autoencoder is a specific, symmetric type of encoder-decoder network where the target output is identical to the input, whereas an encoder-decoder is a general framework mapping an input to a different output domain; therefore, Jiang, in view of Yang, and in view of He, associates the autoencoder from He to the adversarial encoder (CNN discriminator) from Jiang). Therefore, it would have been obvious to combine Jiang and Yang, with He, to obtain as specified in claim 5. With regards to claim 11, it recites the functions of the apparatus of claim 5 as a process. Thus, the analysis in rejecting claim 5 is equally applicable to claim 11. With regards to claim 17, it recites the functions of the apparatus of claim 5 as a non-transitory computer readable medium storing instructions. Thus, the analysis in rejecting claim 5 is equally applicable to claim 17. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: non-patent literature "Self-supervised domain adaptation for diabetic retinopathy grading using vessel image reconstruction"; German Conference on Artificial Intelligence (Künstliche Intelligenz); Cham: Springer International Publishing, 2021 (Nguyen et al.) is relevant for teaching self-supervised learning incorporating encoders, decoders, discriminators, pre-training, and medical images. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHAEL ADAM SHARIFF whose telephone number is 571-272-9741. The examiner can normally be reached M-F 8:30-5PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Sumati Lefkowitz can be reached on 571-272-3638. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MICHAEL ADAM SHARIFF/ Examiner, Art Unit 2672
Read full office action

Prosecution Timeline

Jul 17, 2024
Application Filed
Aug 06, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12700223
METHOD OF EXTRACTING UNSUIITABLE AND DEFECTIVE DATA FROM PLURALITY OF PIECES OF TRAINING DATA USED FOR LEARNING OF MACHINE LEARNING MODEL, INFORMATION PROCESSING DEVICE, AND NON-TRANSITORY COMPUTER-READABLE STORAGE MEDIUM STORING COMPUTER PROGRAM
3y 8m to grant Granted Aug 04, 2026
Patent 12688681
METHOD AND APPARATUS FOR CORRECTING ERRORS IN OUTPUTS OF MACHINE LEARNING MODELS
2y 9m to grant Granted Jul 21, 2026
Patent 12666928
APPARATUS AND METHODS FOR THREE DIMENSIONAL RETICLE DEFECT SMART REPAIR
4y 11m to grant Granted Jun 23, 2026
Patent 12657908
AUTOMATED IMAGING SYSTEM FOR OBJECT FOOTPRINT DETECTION AND A METHOD THEREOF
3y 9m to grant Granted Jun 16, 2026
Patent 12646208
METHOD AND APPARATUS OF DETERMINING DISPENSING POSITION, DISPENSING SYSTEM FOR BATTERIES, ELECTRONIC DEVICE AND MEDIUM
2y 4m to grant Granted Jun 02, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
82%
Grant Probability
99%
With Interview (+23.6%)
2y 9m (~8m remaining)
Median Time to Grant
Low
PTA Risk
Based on 124 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month