Prosecution Insights
Last updated: October 02, 2026
Application No. 18/346,374

MULTITASK MACHINE LEARNING USING DISJOINT DATASETS

Final Rejection §103
Filed
Jul 03, 2023
Examiner
ACOSTA, RILEY SULLIVAN
Art Unit
2143
Tech Center
2100 — Computer Architecture & Software
Assignee
Qualcomm Incorporated
OA Round
2 (Final)
100%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 100% — above average
100%
Career Allowance Rate
1 granted / 1 resolved
+45.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 1m
Avg Prosecution
21 currently pending
Career history
6
Total Applications
across all art units

Statute-Specific Performance

§101
25.0%
-15.0% vs TC avg
§103
54.4%
+14.4% vs TC avg
§102
8.7%
-31.3% vs TC avg
§112
9.8%
-30.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment The amendment filed on 08/05/2026 has been entered. Claims 1-2, 5-6, 10-11, 14-15, 19-20, 23-24, & 28-30 are amended. Claims 1-30 are pending in the application. Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Such claim limitations are: means for accessing, means for accessing, means for generating, means for generating, means for generating, means for updating, means for aggregating, means for generating, means for generating in claims 28-30. Because these claim limitations are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, they are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have these limitations interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitations to avoid them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitations recite sufficient structure to perform the claimed function so as to avoid them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 5-6, 8, 10, 14-15, 17, 19, 23-24, & 26 are rejected under 35 U.S.C. 103 as being unpatentable over Hong et al. ("BEYOND WITHOUT FORGETTING: MULTI-TASK LEARNING FOR CLASSIFICATION WITH DISJOINT DATASETS", arXiv) (Year: 2020), hereafter Hong, and in view of Ranjan et al. ("An All-In-One Convolutional Neural Network for Face Analysis", Center for Automation Research, UMIACS, arXiv) (Year: 2016), hereafter Ranjan. Hong was cited in the previous office action dated 05/18/2026. Regarding independent claim 1, Hong teaches a system comprising: a memory comprising computer-executable instructions; and one or more processors configured to execute the computer-executable instructions and cause the processing system to ([Sec. 3.1 & Alg. 1] discusses a convolutional neural network with shared parameters that is trained via an iterative optimization algorithm; thus, the network necessarily would comprise processors executing stored instructions and memory storing model parameters and the instructions); access a first dataset comprising one or more labeled exemplars for a first machine learning task ([Sec. 3.1] discusses accessing a first dataset comprised of labeled exemplars for task A); access a second dataset comprising one or more labeled exemplars for a second machine learning task ([Sec. 3.1] discusses accessing a second dataset comprised of labeled exemplars for task B); generate a first task-specific combined loss for the first machine learning task using both the first and second datasets ([Sec. 3.1-3.2 & Eq. 1] discusses generating a combined loss at a given training stage using both the first and second datasets); wherein to generate the first task-specific combined loss, the one or more processors are configured to execute the computer-executable instructions to cause the processing system to: generate a first supervised loss for the first machine learning task based on the one or more labeled exemplars from the first dataset ([Sec. 3.1-3.2 & Fig. 1] discusses generating a supervised loss using cross-entropy based on the first dataset); and generate a first self-supervised loss for the first machine learning task based on the one or more labeled exemplars from the second dataset ([Sec. 3.2, 4, Eq. 5-6, & Fig. 1] discusses generating a self-supervised loss using cross-entropy by using the second dataset to obtain soft label vectors); update one or more parameters of a multitask machine learning model based on the first task-specific combined loss ([Sec. 3.2, 4, & Alg. 1] discusses updating the parameters of the model based on the combined loss); generate a second task-specific combined loss for the second machine learning task using both the first and second datasets ([Sec. 3 & Alg. 1] discusses training until complete; thus, the same training process completed to generate the first task’s combined loss is repeated for the second machine learning task); and update the one or more parameters of the multitask machine learning model based on the second task-specific combined loss ([Sec. 3 & Alg. 1] discusses training until complete; thus, part of the training process includes updating parameters based on the loss at each iteration). Hong does not explicitly teach generate a first task-specific combined loss for the first machine learning task during a first training iteration; generate a second task-specific combined loss for the second machine learning task during a second training iteration. However, in a similar field of endeavor, Ranjan teaches a system for training disjoint datasets simultaneously during first, second, and any subsequent training iterations for each different machine learning task ([Abstract, Sec. 3, & Fig. 3] discusses simultaneously generating combined loss terms for a task using multiple disjoint datasets within a single training iteration). Because Hong teaches a system that accesses a first and second dataset, generates a first task combined loss using both datasets to generate a first supervised loss using the first dataset, and a self-supervised loss using the second dataset, updating the parameters based on the combined loss, and repeating this process for a second task combined loss; and Ranjan teaches generating both a first supervised loss using the first dataset, and a self-supervised loss using the second dataset during a single training iteration, and repeating this process for subsequent iterations, accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate generating task-specific combined losses for a machine learning task during the same training iteration using both the first and second datasets simultaneously as taught by Ranjan into Hong’s system, with a reasonable expectation of success, to teach a system that accesses a first dataset comprising one or more labeled exemplars for a first machine learning task; accesses a second dataset comprising one or more labeled exemplars for a second machine learning task; generates a first task-specific combined loss for the first machine learning task during a first training iteration using both the first and second datasets, generates a first supervised loss for the first machine learning task based on the one or more labeled exemplars from the first dataset; generates a first self-supervised loss for the first machine learning task based on the one or more labeled exemplars from the second dataset; updates one or more parameters of a multitask machine learning model based on the first task-specific combined loss; generates a second task-specific combined loss for the second machine learning task during a second training iteration using both the first and second datasets; and updates the one or more parameters of the multitask machine learning model based on the second task-specific combined loss. This combination would have been motivated by the desire to implement simultaneous training, a well-known alternative strategy to the alternating epoch system, as noted by Hong (Hong [Sec. 1]), and reduce over-fitting in the layers (Ranjan [Sec. 1]). Regarding dependent claim 5, the combination of Hong and Ranjan teaches the claimed invention as claimed in claim 1, including wherein to generate the second task-specific combined loss, the one or more processors are configured to execute the computer-executable instructions to cause the processing system to generate a second supervised loss for the second machine learning task based on the one or more labeled exemplars from the second dataset (Hong [Sec. 3.2 & Equation 2 & 5] discusses generating a second supervised loss which is based off of DB and thus, represents the one or more labeled exemplars from the second dataset). Regarding dependent claim 6, the combination of Hong and Ranjan teaches the claimed invention as claimed in claim 1, including wherein to generate the second task-specific combined loss, the one or more processors are configured to execute the computer-executable instructions to cause the processing system to generate a second self-supervised loss for the second machine learning task based on the one or more labeled exemplars from the first dataset (Hong [Sec. 1, Pg. 1] discusses that once the first self-supervised loss is generated based on the second dataset, the same method is applied to generate a second self-supervised loss based on the first dataset). Regarding dependent claim 8, the combination of Hong and Ranjan teaches the claimed invention as claimed in claim 1, including wherein the multitask machine learning model comprises an encoder component shared by both the first and second machine learning tasks, a first decoder component for the first machine learning task, and a second decoder component for the second machine learning task (Hong [Sec. 3.1 & 5.2] discusses the use of convolutional layers of VGG as shared layers and two FC layers as task-specific layers and thus, represents an encoder shared between the two “layers” or first and second machine learning tasks as well as a separate encoder for each layer or machine learning task). Regarding claims 10, 14-15, & 17, claims 10, 14-15, & 17 are method claims that are substantially the same as the system of claims 1, 5-6, & 8. Therefore, claims 10, 14-15, & 17 are rejected for the same reasons as claims 1, 5-6, & 8. Regarding claims 19, 23-24, & 26, claims 19, 23-24, & 26 are non-transitory computer-readable storage medium claims that are substantially the same as the system of claims 1, 5-6, & 8. Therefore, claims 19, 23-24, & 26 are rejected for the same reasons as claims 1, 5-6, & 8. Claims 2-4, 11-13, 20-22, 28-30 are rejected under 35 U.S.C. 103 as being unpatentable over Hong et al. ("BEYOND WITHOUT FORGETTING: MULTI-TASK LEARNING FOR CLASSIFICATION WITH DISJOINT DATASETS", arXiv) (Year: 2020), hereafter Hong, in view of Ranjan et al. ("An All-In-One Convolutional Neural Network for Face Analysis", Center for Automation Research, UMIACS, arXiv) (Year: 2016), hereafter Ranjan, as applied in claims 1, 10, and 19 above, and further in view of Tarvainen et al. ("Mean Teachers are Better Role Models: Weight-Averaged Consistency Targets Improve Semi-Supervised Deep Learning Results", arXiv:1703.01780v6 [cs.NE], 16 April 2018, pp.1-16), hereafter Tarvainen. Tarvainen was cited in IDS filed 10/06/2023. Regarding dependent claim 2, the combination of Hong and Ranjan teaches the claimed invention as claimed in claim 1, including wherein to generate the first task-specific combined loss, the one or more processors are configured to execute the computer-executable instructions to cause the processing system to aggregate the first supervised loss and the first self-supervised loss based at least in part on a first weight for the first self-supervised loss (Hong [Sec. 4] discusses assigning a weight to each training sample to be used in loss calculations for the supervised and self-supervised loss and thus, the first supervised loss and first self-supervised loss are aggregated based in part on the weight for the first self-supervised loss). The combination of Hong and Ranjan does not explicitly teach the first weight being determined based on a current epoch of training the multitask machine learning model. However, in the same field of endeavor, Tarvainen teaches aggregating the first supervised loss and first self-supervised loss based on a first weight and that weight is determined based on a current epoch of training ([Sec. 3] discusses assigning a mean squared error as the consistency cost for training and the weight value ramps up from 0 to its final value based on each epoch). Because the combination of Hong and Ranjan teaches aggregating losses based in part on an assigned weight, and Tarvainen teaches aggregating losses based in part on a first weight and the weight being determined based on a current epoch of training, accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate determining the weight based on the current epoch of training as taught by Tarvainen into the combination of Hong and Ranjan’s system, with a reasonable expectation of success, to teach wherein to generate the first task-specific combined loss, the one or more processors are configured to execute the computer-executable instructions to cause the processing system to aggregate the first supervised loss and the first self-supervised loss based at least in part on a first weight for the first self-supervised loss, the first weight being determined based on a current epoch of training the multitask machine learning model. This combination would have been motivated by the desire to have a more accurate model and aggregate information after every step (Tarvainen [Sec. 2]). Regarding dependent claim 3, the combination of Hong, Ranjan, and Tarvainen teaches the claimed invention as claimed in claim 3, including wherein the first weight is assigned a relatively lower value during relatively earlier epochs of training the multitask machine learning model, as compared to relatively later epochs of training the multitask machine learning model (Tarvainen [Sec. 3] discusses the weight (mean squared error) being ramped up from 0 during the first 80 epochs and thus, the weight is assigned a relatively lower value during earlier epochs as compared to relatively later epochs). Regarding dependent claim 4, the combination of Hong, Ranjan, and Tarvainen teaches the claimed invention as claimed in claim 4, including wherein the first supervised loss and the first self-supervised loss are aggregated based further on a second weight for the first machine learning task, the second weight having a constant value during training of the multitask machine learning model (Hong [Sec. 3.2-4 & 5.4] discusses assigning a weight and that weight remaining constant throughout training and thus, represents aggregating losses based further on a second weight, that weight having a constant value during training). Regarding claims 11-13, claims 11-13 are method claims that are substantially the same as the system of claims 2-4. Therefore, claims 11-13 are rejected for the same reasons as claims 2-4. Regarding claims 20-22, claims 20-22 are non-transitory computer-readable storage medium claims that are substantially the same as the system of claims 2-4. Therefore, claims 20-22 are rejected for the same reasons as claims 2-4. Regarding independent claim 28, Hong teaches a processing system comprising: means for accessing a first dataset comprising one or more labeled exemplars for a first machine learning task; means for accessing a second dataset comprising one or more labeled exemplars for a second machine learning task ([Sec. 3.1] discusses the system obtaining and accessing a first dataset comprised of labeled exemplars for task A, and the system obtaining and accessing a second dataset comprised of labeled exemplars for task B); means for generating a first task-specific combined loss for the first machine learning task using both based on the first and second datasets, comprising: ([Sec. 3.1-3.2 & Eq. 1] discusses generating a combined loss at a given training stage using both the first and second datasets); means for generating a first supervised loss for the first machine learning task based on the one or more labeled exemplars from the first dataset ([Sec. 3.1-3.2 & Fig. 1] discusses the system generating a supervised loss using cross-entropy based on the first dataset); generating a first self-supervised loss for the first machine learning task based on the one or more labeled exemplars from the second dataset ([Sec. 3.2, 4 & Fig. 1] discusses the system generating a self-supervised loss using cross-entropy by using the second dataset to obtain soft label vectors); means for updating one or more parameters of a multitask machine learning model based on the first task-specific combined loss ([Sec. 3.2, 4, & Alg. 1] discusses updating the parameters of the model based on the combined loss); means for generating a second task-specific combined loss for the second machine learning task using both the first and second datasets; and means for updating the one or more parameters of the multitask machine learning model based on the second task-specific combined loss ([Sec. 3 & Alg. 1] discusses training until complete; thus, the same training process completed to generate the first task’s combined loss is repeated for the second machine learning task). Hong does not explicitly teach means for generating a first self-supervised loss for the first machine learning task based on the one or more labeled exemplars from the second dataset; generate a first task-specific combined loss for the first machine learning task during a first training iteration; generate a second task-specific combined loss for the second machine learning task during a second training iteration. However, in a similar field of endeavor, Ranjan teaches a system for training disjoint datasets simultaneously during first, second, and any subsequent training iterations for each different machine learning task ([Abstract, Sec. 3, & Fig. 3] discusses simultaneously generating combined loss terms for a task using multiple disjoint datasets within a single training iteration). Because Hong teaches a system comprising accessing a first and second dataset, generating a first task combined loss using both datasets to generate a first supervised loss using the first dataset, and a self-supervised loss using the second dataset, updating the parameters based on the combined loss, and repeating this process for a second task combined loss; and Ranjan teaches generating both a first supervised loss using the first dataset, and a self-supervised loss using the second dataset during a single training iteration, and repeating this process for subsequent iterations, accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate generating task-specific combined losses for a machine learning task during the same training iteration using both the first and second datasets simultaneously as taught by Ranjan into Hong’s system, with a reasonable expectation of success, to teach a system with means for accessing a first dataset comprising one or more labeled exemplars for a first machine learning task; means for accessing a second dataset comprising one or more labeled exemplars for a second machine learning task; means for generating a first task-specific combined loss for the first machine learning task during a first training iteration using both the first and second datasets; means for generating a first supervised loss for the first machine learning task based on the one or more labeled exemplars from the first dataset; generating a first self-supervised loss for the first machine learning task based on the one or more labeled exemplars from the second dataset; means for updating one or more parameters of a multitask machine learning model based on the first task-specific combined loss; means for generating a second task-specific combined loss for the second machine learning task during a second training iteration using both the first and second datasets; and means for updating the one or more parameters of the multitask machine learning model based on the second task-specific combined loss. This combination would have been motivated by the desire to implement simultaneous training, a well-known alternative strategy to the alternating epoch system, as noted by Hong (Hong [Sec. 1]), and reduce over-fitting in the layers (Ranjan [Sec. 1]). The combination of Hong and Ranjan does not explicitly teach means for generating a first self-supervised loss for the first machine learning task based on the one or more labeled exemplars from the second dataset. However, in the same field of endeavor, Tarvainen teaches the necessary structure for generating a first self-supervised loss (Tarvainen [Sec. 2] discusses the structure of a teacher and student model to generate a self-supervised loss using EMA of weights, per the specification, and thus, teaches a means for generating a first self-supervised loss for the first machine learning task based on the one or more labeled exemplars from the second dataset). Because the combination of Hong and Ranjan teaches a system comprising accessing a first and second dataset, generating a first task combined loss using both datasets to generate both a first supervised loss using the first dataset and a self-supervised loss using the second dataset during a single training iteration, updating the parameters based on the combined loss, and repeating this process for a second task combined loss; and Tarvainen teaches exponential moving average for generating a first self-supervised loss for the first machine learning task, accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teacher and student model using EMA for weights as taught by Tarvainen into the combination of Hong and Ranjan’s computer-implemented system, with a reasonable expectation of success, to teach means for accessing a first dataset comprising one or more labeled exemplars for a first machine learning task; means for accessing a second dataset comprising one or more labeled exemplars for a second machine learning task; means for generating a first task-specific combined loss for the first machine learning task during a first training iteration using both the first and second datasets; means for generating a first supervised loss for the first machine learning task based on the one or more labeled exemplars from the first dataset; means for generating a first self-supervised loss for the first machine learning task based on the one or more labeled exemplars from the second dataset; means for updating one or more parameters of a multitask machine learning model based on the first task-specific combined loss; means for generating a second task-specific combined loss for the second machine learning task during a second training iteration using both the first and second datasets; and means for updating the one or more parameters of the multitask machine learning model based on the second task-specific combined loss. This combination would have been motivated by the desire to aggregate information after every step and update the weights with EMA to reduce noise (Tarvainen [Sec. 2-3]). Regarding dependent claim 29, the combination of Hong, Ranjan, and Tarvainen teaches the claimed invention as claimed in claim 29, including wherein the means for generating the first task-specific combined loss comprises means for aggregating the first supervised loss and the first self-supervised loss based at least in part on a first weight for the first self-supervised loss, the first weight being determined based on a current epoch of training the multitask machine learning model (Hong [Sec. 4] discusses the system assigning a weight to each training sample to be used in loss calculations for the supervised and self-supervised loss and thus, the first supervised loss and first self-supervised loss are aggregated based in part on the weight for the first self-supervised loss; Tarvainen [Sec. 2-3] discusses assigning a mean squared error as the consistency cost for training and the weight value ramps up from 0 to its final value based on each epoch and thus, the first weight is being determined based on a current epoch of training). Regarding dependent claim 30, the combination of Hong, Ranjan, and Tarvainen teaches the claimed invention as claimed in claim 30, including wherein the means for generating the second task-specific combined loss comprise: means for generating a second supervised loss for the second machine learning task based on the one or more labeled exemplars from the second dataset (Hong [Sec. 3.2] discusses the system generating a second supervised loss using cross-entropy for dataset B and the second machine learning task); and means for generating a second self-supervised loss for the second machine learning task based on the one or more labeled exemplars from the first dataset (Hong [Sec. 4 & Eq. 5] discusses the system generating a second self-supervised loss based on the first dataset; and Tarvainen [Sec. 2] discusses the structure of a teacher and student model to generate a self-supervised loss using EMA of weights, per the specification, and thus, the combination teaches a means for generating a second self-supervised loss for the second machine learning task based on the one or more labeled exemplars from the first dataset). Claims 7, 16, & 25 are rejected under 35 U.S.C. 103 as being unpatentable over Hong, in view of Ranjan, as applied in claims 1, 10, and 19 above, and further in view of Rabadán et al. ("Dense FixMatch: A Simple Semi-Supervised Learning Method for Pixel-wise Prediction Tasks," 18 October 2022, pp.1-12, [Pages 2, 22, and 34.]), hereafter Rabadán. Rabadán was cited in IDS filed 10/06/2023. Regarding dependent claim 7, the combination of Hong and Ranjan teaches the claimed invention as claimed in claim 1, including: generate a first output based on the first labeled exemplar augmented according to a first set of augmentations (Hong [Sec. 1] discusses uses dataset B with pseudo labels to augment task A, and generate a first output based on that set of augmentations); generate a second output based on the first labeled exemplar augmented according to a second set of augmentations (Hong [Sec. 1] discusses generating output based on augmentations to task A; correspondingly, a second output is generated based on using dataset A with pseudo labels to augment task B, and generate a second output based on the second set of augmentations); generate a pseudo-label (Hong [Sec. 4] discusses generating a pseudo-label to supervise and augment each task and compare the pseudo-label and the second output (Hong [Sec. 4, Eq. 4, 5] discusses comparing the pseudo-label to the second output in loss functions to take the minimum cross-entropy loss). The combination of Hong and Ranjan does not explicitly teach generate a pseudo-label based on modifying the first output using the first and second sets of augmentations. However, in the same field of endeavor, Rabadán teaches a system for a semi-supervised learning method including generating a pseudo label based on modifying the first output using the first and second sets of augmentations ([Sec. 3, Pg. 4-5, Eq. 5] discusses generating a pseudo-label by modifying the first output, using the first and second sets of augmentations). Because the combination of Hong and Ranjan teaches generating a first output based on a first set of augmentations, generating a second output based on a second set of augmentations, generating a pseudo-label, and comparing the pseudo-label and the second output; and Rabadán teaches generating a pseudo-label based on modifying the first output using the first and second sets of augmentation, accordingly, it would have been obvious to a person having ordinary skill in the art, before the effective filing date, to incorporate generating the pseudo label based on modifying the first output as taught by Rabadán into the combination of Hong and Ranjan’s computer-implemented system, with a reasonable expectation of success, to teach wherein to generate the first self-supervised loss, the one or more processors are configured to execute the computer-executable instructions to cause the processing system to, for a first labeled exemplar from the second dataset: generate a first output based on the first labeled exemplar augmented according to a first set of augmentations; generate a second output based on the first labeled exemplar augmented according to a second set of augmentations; generate a pseudo-label based on modifying the first output using the first and second sets of augmentations; and compare the pseudo-label and the second output. This combination would have been motivated by the desire to define a consistency objective between the two views for any dense or structured task, including semantic segmentation, object detection, and instance segmentation, while still being able to use different geometric transformations in both augmentation pipelines (Rabadán [Pg. 4]). Regarding claim 16, claim 16 is a method claim that is substantially the same as the system of claim 7. Therefore, claim 16 is rejected for the same reasons as claim 7. Regarding claim 25, claim 25 is a non-transitory computer-readable storage medium claim that is substantially the same as the system of claim 7. Therefore, claim 25 is rejected for the same reasons as claim 7. Claims 9, 18, & 27 are rejected under 35 U.S.C. 103 as being unpatentable over Hong, in view of Ranjan, as applied in claims 1, 10, and 19 above, and further in view of Ghiasi et al. ("Multi-Task Self-Training for Learning General Representations", arXiv) (Year: 2021), hereafter Ghiasi. Regarding dependent claim 9, the combination of Hong and Ranjan teaches the claimed invention as claimed in claim 1, including: first and second machine learning tasks (Hong [Sec. 1] discusses taking two datasets corresponding to two machine learning tasks). The combination of Hong and Ranjan does not explicitly teach wherein the first and second machine learning tasks are computer vision tasks and comprise at least one of: monocular depth estimation, semantic segmentation, object detection, surface normal estimation, or edge detection. However, in the same field of endeavor, Ghiasi teaches a multi-task training for learning general representations where the teacher models include four important tasks ([Sec. 1 & 3.1] discusses the four machine learning tasks being classification, object detection, semantic segmentation, and depth estimation). Because the combination of Hong and Ranjan teaches the use of two machine learning tasks, and Ghiasi teaches using machine learning tasks comprising object detection, semantic segmentation, and depth estimation, accordingly, it would have been obvious to a person having ordinary skill in the art, before the effective filing date, to incorporate machine learning tasks comprising either object detection, semantic segmentation, and depth estimation as taught by Ghiasi into the combination of Hong and Ranjan’s system, with a reasonable expectation of success, to teach wherein the first and second machine learning tasks are computer vision tasks and comprise at least one of: monocular depth estimation, semantic segmentation, object detection, surface normal estimation, or edge detection. This combination would have been motivated by the desire to learn from a set of teachers that provide rich training signals with their pseudo labels (Ghiasi [3.1]). Regarding claim 18, claim 18 is a method claim that is substantially the same as the system of claim 9. Therefore, claim 18 is rejected for the same reasons as claim 9. Regarding claim 27, claim 27 is a non-transitory computer-readable storage medium claim that is substantially the same as the system of claim 9. Therefore, claim 27 is rejected for the same reasons as claim 9. Response to Arguments Applicant’s amendments and remarks filed 08/05/2026, on pages 11-16, traversing the 35 U.S.C. 101 rejections set forth in the Office Action dated 05/18/2026, are persuasive and hereby withdrawn. Applicant’s arguments with respect to claims 1, 10, 19, & 28 have been considered but are moot because of the new ground of rejection (see rejection above). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Augenstein et al. ("Multi-task learning of pairwise sequence classification tasks over disparate label spaces", arXiv) (Year: 2018) (Cited in the previous office action) ([Abstract] We combine multi-task learning and semi supervised learning by inducing a joint embedding space between disparate label spaces and learning transfer functions between label embeddings, enabling us to jointly leverage unlabelled data and auxiliary, annotated datasets). Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to RILEY S ACOSTA whose telephone number is (571)272-8714. The examiner can normally be reached Monday-Thursday 6am-4pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer N Welch can be reached at (571)272-7212. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /RILEY S ACOSTA/Examiner, Art Unit 2143 /JENNIFER N WELCH/Supervisory Patent Examiner, Art Unit 2143
Read full office action

Prosecution Timeline

Jul 03, 2023
Application Filed
May 18, 2026
Non-Final Rejection mailed — §103
Aug 05, 2026
Response Filed
Sep 08, 2026
Final Rejection mailed — §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
100%
Grant Probability
99%
With Interview (+0.0%)
3y 1m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 1 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month