Prosecution Insights
Last updated: October 02, 2026
Application No. 18/596,414

DOMAIN GENERALIZATION FOR MACHINE LEARNING MODELS

Non-Final OA §101§103
Filed
Mar 05, 2024
Examiner
HAN, KYU HYUNG
Art Unit
Tech Center
Assignee
International Business Machines Corporation
OA Round
1 (Non-Final)
44%
Grant Probability
Moderate
1-2
OA Rounds
1y 7m
Est. Remaining
80%
With Interview

Examiner Intelligence

Grants 44% of resolved cases
44%
Career Allowance Rate
7 granted / 16 resolved
-16.2% vs TC avg
Strong +37% interview lift
Without
With
+36.7%
Interview Lift
resolved cases with interview
Typical timeline
4y 2m
Avg Prosecution
26 currently pending
Career history
44
Total Applications
across all art units

Statute-Specific Performance

§101
29.6%
-10.4% vs TC avg
§103
59.7%
+19.7% vs TC avg
§102
2.2%
-37.8% vs TC avg
§112
8.6%
-31.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 16 resolved cases

Office Action

§101 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections – 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1: Claims 1-7 are method claims. Claims 8-20 are machine/system/product claims. Therefore, claims 1-20 are directed to either a process, machine, manufacture or composition of matter. With respect to claim 1: Step 2A – Prong 1: A method of training a machine learning model, the method comprising: processing, …, a first plurality of images belonging to a first domain … (mental process – a person can manually process images belonging to a first domain with the assistance of a pen/paper.) processing, …, a second plurality of images belonging to a second domain … (mental process – a person can manually process images belonging to a second domain with the assistance of a pen/paper.) generating, …, a compound error metric from a plurality of error metrics derived from results generated from the processing of the first network and the processing of the second network; (mental process – a person can manually generate a compound error metric from a plurality of error metrics derived from results generated from the processing of the first network and the processing of the second network with the assistance of a pen/paper.) updating, …, weights of the first network based on the compound error metric; (mental process – a person can manually update weights of the first network based on the compound error metric with the assistance of a pen/paper.) and updating, …, weights of the second network using a moving average technique that is dependent on the weights of the first network as updated. (mental process – a person can manually update weights of the second network using a moving average technique that is dependent on the weights of the first network as updated with the assistance of a pen/paper.) Step 2A – Prong 2: This judicial exception is not integrated into a practical application. … using computer hardware … (mere instructions to apply the exception using a generic computer component – computer applies exception) … through a first network; (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: High level recitation of training the machine learning engine to process images.); With respect to claim 2: Step 2A – Prong 1: The method of claim 1, wherein the first plurality of images and the second plurality of images are ordered according to class as processed by the first network and the second network. (mental process – a person can recognize that the first plurality of images and the second plurality of images are ordered according to class as processed by the first network and the second network.) With respect to claim 3: Step 2A – Prong 1: The method of claim 1, wherein: the first network includes an online encoder, an online projector, an online predictor, and a classifier; (mental process – a person can recognize that the first network includes an online encoder, an online projector, an online predictor, and a classifier.) and the second network includes a target encoder and a target projector. (mental process – a person can recognize that the second network includes a target encoder and a target projector.) With respect to claim 4: Step 2A – Prong 1: The method of claim 3, wherein the classifier is configured to receive embeddings generated by the online encoder. (mental process – a person can recognize that the classifier is configured to receive embeddings generated by the online encoder.) With respect to claim 5: Step 2A – Prong 1: The method of claim 1, wherein the plurality of error metrics includes a selected error metric generated based on classification results from a classifier of the first network and ground truth labels of the first plurality of images. (mental process – a person can recognize that the plurality of error metrics includes a selected error metric generated based on classification results from a classifier of the first network and ground truth labels of the first plurality of images.) With respect to claim 6: Step 2A – Prong 1: The method of claim 1, wherein the plurality of error metrics include a supervised contrastive loss comprising: an intra-domain error generated by comparing images in same mini-batches having same labels only for the first plurality of images; (mental process – a person can recognize that the plurality of error metrics include a supervised contrastive loss comprising an intra-domain error generated by comparing images in same mini-batches having same labels only for the first plurality of images.) and an inter-domain error generated by comparing mini-batches of images of the first plurality of images with mini-batches of images of the second plurality of images, wherein the inter-domain error compares images with same labels. (mental process – a person can recognize that the inter-domain error is generated by comparing mini-batches of images of the first plurality of images with mini-batches of images of the second plurality of images, wherein the inter-domain error compares images with same labels.) With respect to claim 7: Step 2A – Prong 1: The method of claim 1, wherein the plurality of error metrics includes a selected error metric generated based on a prediction similarity between the first network and the second network. (mental process – a person can recognize that the plurality of error metrics includes a selected error metric generated based on a prediction similarity between the first network and the second network.) Claim 8 is substantially similar to claim 1, but has the following additional elements: With respect to claim 8: Step 2A – Prong 2: A system for training a machine learning model, comprising: one or more processors configured to execute operations including (mere instructions to apply the exception using a generic computer component – processor applies exception) Claims 9-14 are rejected on the same grounds under 35 U.S.C. 101 as claims 2-7 as they are substantially similar, respectively. Mutatis mutandis. Claim 15 is substantially similar to claim 1, but has the following additional elements: With respect to claim 15: Step 2A – Prong 2: A computer program product comprising one or more computer readable storage mediums having program instructions embodied therewith, wherein the program instructions are executable by one or more processors to cause the one or more processors to execute operations comprising (mere instructions to apply the exception using a generic computer component – processor applies exception) Claims 16-20 are rejected on the same grounds under 35 U.S.C. 101 as claims 2, 3, 5, 6, 7 as they are substantially similar, respectively. Mutatis mutandis. Claim Rejections – 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-2, 5-6, 8-9, 12-13, 15-16, 18-19 are rejected under 35 U.S.C. 103 as being unpatentable over Zhou et al. (US20220269946) hereinafter known as Zhou in view of Das et al. (“Weakly-Supervised Domain Adaptive Semantic Segmentation with Prototypical Contrastive Learning”) hereinafter known as Das. Regarding independent claim 1, Zhou teaches: A method of training a machine learning model, the method comprising: processing, using computer hardware, a first plurality of images belonging to a first domain through a first network; (Zhou [0023]: “In one embodiment, the query sample … may be input to the online network … while the set of positive instances … and the set of negative instance … may be input to the target network” Zhou teaches that the query images are processed through the first neural network, which is the online network.) processing, using the computer hardware, a second plurality of images belonging to a second domain through a second network; (Zhou [0023]: “In one embodiment, the query sample … may be input to the online network … while the set of positive instances … and the set of negative instance … may be input to the target network” Zhou teaches that the positive and negative key images are processed through the second neural network, which is the target network.) Zhou does not explicitly teach: generating, using the computer hardware, a compound error metric from a plurality of error metrics derived from results generated from the processing of the first network and the processing of the second network; updating, using the computer hardware, weights of the first network based on the compound error metric; and updating, using the computer hardware, weights of the second network using a moving average technique that is dependent on the weights of the first network as updated. However, Das teaches: generating, using the computer hardware, a compound error metric from a plurality of error metrics derived from results generated from the processing of the first network and the processing of the second network; (Das [Page 15437, Col. 2, Paragraph 3.3]: “align the features from the student network with prototypes from the teacher network using both Intra Domain and Inter Domain Alignment” Das teaches using both intra domain and inter domain alignment to calculate two losses corresponding to both.) updating, using the computer hardware, weights of the first network based on the compound error metric; (Das [Page 15437, Col. 2, Equation 13]: Das teaches combining both losses into a final training loss.) and updating, using the computer hardware, weights of the second network using a moving average technique that is dependent on the weights of the first network as updated. (Das [Page 15437, Col. 2, Equation 13]: “we employ a teacher network, which is an exponential moving average of the segmentation network (also called student network) during training. Specifically, the weights of the teacher network are updated using the student weights” Das teaches that the weights of the teacher network are updated using the student weights using an exponential moving average.) Zhou and Das are in the same field of endeavor as the present invention, as the references are directed to training image-processing neural networks through contrastive feature learning. It would have been obvious, before the effective filing date of the claimed invention, to a person of ordinary skill in the art, to combine the processing of images by the neural networks as taught in Zhou with updating the weights of the network via a moving average of a teacher/student pair of networks as taught in Das. Das provides this additional functionality. As such, it would have been obvious to one of ordinary skill in the art to modify the teachings of Zhou to include teachings of Das because the combination would allow for a pair of networks that update based on the other. This has the potential benefit of training the network to better predict domains that it hasn’t seen already. Regarding dependent claim 2, Zhou and Das teach: The method of claim 1, wherein the first plurality of images and the second plurality of images are ordered according to class as processed by the first network and the second network. (Zhou [0026]: “Due to the shared instance set B, all queries have unified label definition and their labels can be linearly combined to form the self-label” Zhou teaches that the images of a set can form a label that describes itself.) The reasons to combine are substantially similar to those of claim 1. Regarding dependent claim 5, Zhou and Das teach: The method of claim 1, wherein the plurality of error metrics includes a selected error metric generated based on classification results from a classifier of the first network and ground truth labels of the first plurality of images. (Das [Page 15437, Col. 1, Paragraph 1]: “We use the pixelwise logit score m(i,k), to obtain an image level prediction probability, pk t that is used for computing the image loss” Das teaches an image loss that’s computed from prediction probabilities. Das [Page 15438, Col. 2, Paragraph 3]: “combine the image loss from classification layer with prototype based image loss from Eq. (6) to get the final image loss … image labels are obtained from the available class labels from the ground truth labels” Das teaches that the image labels are obtained from a ground truth.) The reasons to combine are substantially similar to those of claim 1. Regarding dependent claim 6, Zhou and Das teach: The method of claim 1, wherein the plurality of error metrics include a supervised contrastive loss comprising: an intra-domain error generated by comparing images in same mini-batches having same labels only for the first plurality of images; (Das [Page 15438, Equation 10]: Das teaches contrastive loss. Das [Page 15437, Col. 2, last paragraph]: “Specifically, for a given pixel feature, we construct a positive pair of the pixel features with its corresponding class prototype while a negative pair with different class prototypes” Das constructs class prototypes from a training batch and separately defines intra-domain loss.) and an inter-domain error generated by comparing mini-batches of images of the first plurality of images with mini-batches of images of the second plurality of images, wherein the inter-domain error compares images with same labels. (Das [Page 15438, Col. 1, Paragraph 2]: “Given a feature fi from the source domain, we construct a positive pair with the prototype from the same class from the target domain and a negative pair with the prototype from a different class in the target domain” Das constructs an inter-domain term between a source feature and a target prototype from the same class.) The reasons to combine are substantially similar to those of claim 1. Claim 8 is substantially similar to claim 1, but has the following additional elements: Regarding independent claim 8, Zhou and Das teach: A system for training a machine learning model, comprising: one or more processors configured to execute operations including (Zhou [0050]: “processor 510 may be representative of one or more central processing units, multi-core processors, microprocessors” Zhou teaches a processor that can execute instructions.) The reasons to combine are substantially similar to those of claim 1. Claims 9, 12, 13 are rejected on the same grounds under 35 U.S.C. 103 as claims 2, 5, 6 as they are substantially similar, respectively. Mutatis mutandis. Claim 15 is substantially similar to claim 1, but has the following additional elements: Regarding independent claim 15, Zhou and Das teach: A system for training a machine learning model, comprising: one or more processors configured to execute operations including (Zhou [0050]: “processor 510 may be representative of one or more central processing units, multi-core processors, microprocessors” Zhou teaches a processor that can execute instructions.) The reasons to combine are substantially similar to those of claim 1. Claims 16, 18, 19 are rejected on the same grounds under 35 U.S.C. 103 as claims 2, 5, 6 as they are substantially similar, respectively. Mutatis mutandis. Claims 3, 4, 7, 10, 11, 14, 17, 20 are rejected under 35 U.S.C. 103 as being unpatentable over Zhou in view of Das in view of Grill et al. (“Bootstrap Your Own Latent: A New Approach to Self-Supervised Learning”) hereinafter known as Grill. Regarding dependent claim 3, Zhou and Das teach: … … and a classifier; (Zhou [0061]: “For classification, a linear classifier is trained upon ResNet50 100 epochs by SGD” Zhou teaches a linear classifier that’s trained on the ResNet50.) … Zhou and Das do not explicitly teach: The method of claim 1, wherein: the first network includes an online encoder, an online projector, an online predictor, … … and the second network includes a target encoder and a target projector. However, Grill teaches: The method of claim 1, wherein: the first network includes an online encoder, an online projector, an online predictor, … (Grill [Page 3, Section 3.1]: “The online network is defined by a set of weights θ and is comprised of three stages: an encoder fθ, a projector gθ and a predictor qθ” Grill teaches that the network comprises of an encoder, projector and a predictor.) … and the second network includes a target encoder and a target projector. (Grill [Page 4, Paragraph 2]: “target network outputs yξ = ∆ fξ(v) and the target projection zξ = ∆ gξ(y)” Grill teaches that the network includes a target encoded and a target projector.) Grill is in the same field as the present invention, since it is directed to training image representation models using paired online and target networks. It would have been obvious, before the effective filing date of the claimed invention, to a person of ordinary skill in the art, to combine the pair of networks that update based on the other as taught in Zhou as modified by Das with the online encoder, projector, and predictor with the estimated moving average as taught in Grill. Grill provides this additional functionality. As such, it would have been obvious to one of ordinary skill in the art to modify the teachings of Zhou as modified by Das to include teachings of Grill because the combination would allow for asymmetric online architecture where the online predictor matches representations generated by the EMA target network. This has the potential benefit of making the models more robust across different image representations and domains.s Regarding dependent claim 4, Zhou, Das, and Grill teach: The method of claim 3, wherein the classifier is configured to receive embeddings generated by the online encoder. (Grill [Page 3, Section 3.1]: “The online network is defined by a set of weights θ and is comprised of three stages: an encoder fθ, a projector gθ and a predictor qθ” Grill teaches that the network encodes into embeddings, which may subsequently be used by other systems, including a classifier.) The reasons to combine are substantially similar to those of claim 3. Regarding dependent claim 7, Zhou, Das, and Grill teach: The method of claim 1, wherein the plurality of error metrics includes a selected error metric generated based on a prediction similarity between the first network and the second network. (Grill [Page 4, Figure 2]: Grill teaches minimizing a similarity loss between the online prediction and the target projection.) The reasons to combine are substantially similar to those of claim 3. Claims 10, 11, 14 are rejected on the same grounds under 35 U.S.C. 103 as claims 3, 4, 7 as they are substantially similar, respectively. Mutatis mutandis. Claims 17, 20 are rejected on the same grounds under 35 U.S.C. 103 as claims 3, 7 as they are substantially similar, respectively. Mutatis mutandis. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to KYU HYUNG HAN whose telephone number is (703) 756-5529. The examiner can normally be reached on MF 9-5. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached on (571) 270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Kyu Hyung Han/ Examiner Art Unit 2123 /ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123
Read full office action

Prosecution Timeline

Mar 05, 2024
Application Filed
Aug 31, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12750204
MACHINE LEARNING NETWORK EXTENSION BASED ON HOMOMORPHIC ENCRYPTION PACKINGS
4y 3m to grant Granted Sep 29, 2026
Patent 12743739
SYSTEM AND METHOD FOR BALANCING CONTAINERIZED APPLICATION OFFLOADING AND BURST TRANSMISSION FOR THERMAL CONTROL
4y 8m to grant Granted Sep 22, 2026
Patent 12651157
METHODS AND SYSTEMS FOR GENERATING THE GRADIENTS OF A LOSS FUNCTION WITH RESPECT TO THE WEIGHTS OF A CONVOLUTION LAYER
4y 2m to grant Granted Jun 09, 2026
Patent 12585928
HARDWARE ARCHITECTURE FOR INTRODUCING ACTIVATION SPARSITY IN NEURAL NETWORK
4y 10m to grant Granted Mar 24, 2026
Patent 12387101
SYSTEMS AND METHODS FOR PRUNING BINARY NEURAL NETWORKS GUIDED BY WEIGHT FLIPPING FREQUENCY
4y 3m to grant Granted Aug 12, 2025
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
44%
Grant Probability
80%
With Interview (+36.7%)
4y 2m (~1y 7m remaining)
Median Time to Grant
Low
PTA Risk
Based on 16 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month