Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
Response to Amendment
This communication is in response to the amendment filed on 03/06/2026 for the application No. 18/421,711, Claims 1-20 are currently pending and have been examined. Claims 1-20 have been rejected as follow,
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1- 20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter.
Claims 1-20 are not compliant with 101, according with the last “2019 Revised Patent Subject Matter Eligibility Guidance” (2019 PEG), published in the MPEP 2103 through 2106.07(c). The claims has been amended and Examiner’s analysis is presented below in all the claims.
Claim 1: Step 1 of 2019 PGE, does the claim fall within a Statutory Category? Yes. The claim recites a method.
Step 2A - Prong 1: Is a Judicial Exception recited in the claim? Yes. The claim recites the limitations of
“b) computing, …. on the batch from dataset, a first loss using a first loss function, wherein the first loss function comprises a term that calculates a natural logarithm of a prediction probability to a power of beta, where beta is a tunable parameter and is a real number; and c) updating weights … based on the first loss”
The “computing, updating” limitations, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitations as certain methods of organizing human activity, advertising, marketing or sales activities or behaviors. The method for training a machine learning for predicting click-through-rate for online advertisement. Thus, the claim recites an abstract idea.
This method also comprises “ ML model, first loss function, calculates a natural logarithm of a prediction probability, second loss function “ which can be considered a mathematic formula or equation with fall within the mathematical concepts grouping of abstract ideas. (see MPEP 2106.04).
Step 2A - Prong 2: Integrated into a Practical Application? No. The claim recites additional limitations, such as,
“a) receiving a batch from a dataset “. These are limitations toward accessing or receiving data (gathering data).
The Examiner analyses other supplementary elements in the claim in view of the instant disclosure:
“training a machine learning (ML) model”, “the (ML) model”, “computer-implemented, ,to perform a data processing task, by one or more processors, wherein the ML model is a neural network model; training, by the one or more processors, the ML model;” These elements are recited in a very generic way, and the amended limitations provided more structure to the claim.
The Examiner gives the broadest reasonable interpretation to the above elements. They are insignificant extra-solution activity. See MPEP 2106.05(g).
The combination of these additional elements can also be considered no more than mere instructions “to apply” the exception, See MPEP 2106.05(f).
Accordingly, even in combination, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea.
The claim as a whole does not integrate the method of organizing human activity into a practical application. OR the claim as a whole does not integrate the method of claiming mathematical concepts into a practical application. Thus, the claim is ineligible because is directed to the recited judicial exception (abstract idea).
Step 2B : claim provides an inventive concept? No.
As discussed with respect to Step 2A Prong Two, the additional elements in the claim,
“training a machine learning (ML) model”, “the (ML) model”, “computer-implemented, ,to perform a data processing task, by one or more processors, wherein the ML model is a neural network model; training, by the one or more processors, the ML model;” amount to no more than mere instructions to apply the exception. i.e., mere instructions to apply an exception using generic hardware and software cannot integrate a judicial exception into a practical application at Step 2A or provide an inventive concept in Step 2B.
Under the 2019 PEG, a conclusion that an additional element is insignificant extra-solution activity in Step 2A should be re-evaluated in Step 2B.
Here, the limitations:
“training a machine learning (ML) model”, “the (ML) model”, “computer-implemented, ,to perform a data processing task, by one or more processors, wherein the ML model is a neural network model; training, by the one or more processors, the ML model;” were considered to be extra-solution activity in Step 2A, and thus it is re-evaluated in Step 2B to determine if it is more than what is well-understood, routine, conventional activity in the field.
Other limitations in the claim, such as:
“a) receiving a batch from a dataset “. These are limitations toward accessing or receiving data (gathering data). Accessing or transmitting data is very well understood, routine and conventional computer task activity; It represents insignificant extra solution activity. Mere data-gathering step[s] cannot make an otherwise nonstaturory claim statutory In re Grams,888 F.2d 835, 840 (Fed. Cir. 1989) (quoting In re Meyer, 688 F.2d 789, 794 (CCPA 1982)).
Further, the instant specification does not provide any indication that the elements
“training a machine learning (ML) model”, “the (ML) model”, “computer-implemented, ,to perform a data processing task, by one or more processors, wherein the ML model is a neural network model; training, by the one or more processors, the ML model;” were are anything other than generic software and generic computer elements, and the OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network); and v. Presenting offers and gathering statistics, OIP Techs., 788 F.3d at 1362-63, 115 USPQ2d at 1092-93; court decisions cited in MPEP 2106.05(d)(II) indicate that merely computer receives and sends information over a network and presenting or displaying information, is a well‐understood, routine, conventional function when it is claimed in a merely generic manner (as it is here).
Accordingly, a conclusion that the “training a machine learning (ML) model”, “the (ML) model”, “computer-implemented, ,to perform a data processing task, by one or more processors, wherein the ML model is a neural network model; training, by the one or more processors, the ML model;” limitations (pointed above) are well-understood, routine, conventional activity is supported under Berkheimer Option 2. The claim is ineligible.
Additionally, the Examiner notes that generic elements such as
“training a machine learning (ML) model”, “the (ML) model” as claimed here, are well-understood, routine, conventional elements and activity, see for example the references:
- “Instance Weighting in Neural Networks for Click-Through Rate Prediction”. This article discloses “The instances on which a learning algorithm is most undecided can be weighted more during training to guide the learning model towards spending more effort on the difficult instances. We introduce three instance weighting algorithms to weigh the loss obtained”. Abstract. All the elements in the instant claim are fully supported under Berkheimer Option 2.
Claim 15: Step 1 of 2019 PGE, does the claim fall within a Statutory Category? Yes. The claim recites a computer readable medium.
Step 2A - Prong 1: Is a Judicial Exception recited in the claim ? Yes. Because the same reasons pointed above.
Step 2A - Prong 2: Integrated into a Practical Application? No. Because the same reasons pointed above.
Step 2B : claim provides an inventive concept? No. Because the same reasons pointed above. The claim is ineligible.
Claim 20 Step 1 of 2019 PGE, does the claim fall within a Statutory Category? Yes. The claim recites an apparatus.
Step 2A - Prong 1: Is a Judicial Exception recited in the claim ? Yes. Because the same reasons pointed above.
Step 2A - Prong 2: Integrated into a Practical Application? No. Because the same reasons pointed above.
Step 2B : claim provides an inventive concept? No. Because the same reasons pointed above. The claim is ineligible.
Dependent claims 2-14, 16-19, the claims recite elements such as “computing, based on predictions by the ML model on a second dataset for validation, a second loss using a second loss function; determining, based on the second loss, convergence of the ML model; and output the ML model based on determining that the ML model is converged”, etc. These elements do not integrate the system of organizing human activity into a practical application. The claims are ineligible.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-7, 11-18 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over US PG. Pub. No. 20220114444 ((WEINZAEPFEL called herein WEIN) in view of https://www.kaggle.com/code/dansbecker/what-is-log-loss/notebook (Kaggle).
As to claims 1, 15 and 20, WEIN discloses A computer-implemented method for training a machine learning (ML) model to perform a data processing task, the method (abstract and paragraph 2), comprising:
a) receiving, by one or more processors, a batch from a dataset to train the ML model;
(see Fig. 4 steps 410-430) wherein the ML model is a neural network model (“training a neural network to perform a task”, paragraph 2.
“[0101] FIG. 4 is a flow diagram of a method 400 of training a neural network using the SuperLoss function described above. At 410, a batch of data samples to be processed by the neural network is obtained, where a batch comprises a number of randomly-selected data samples”, paragraph 101 and Fig. 4);
b) computing, by the one or more processors, based on predictions by the ML model on the batch from dataset, a first loss using a first loss function, wherein the first loss function comprises a term that calculates a natural logarithm of a prediction probability [to a power of beta]
(“[0051] In further features, automatically computing the weight value includes, by the second loss function, automatically computing the weight value based on a regularization term given by λ(log σ*).sup.2, where σ* is the confidence value, λ is the regularization hyperparameter [tunable parameter], and log represents the logarithm function”, paragraph 51);
where [beta] is a tunable parameter and is a real number;
(see “[0045] In further features, automatically computing the weight value includes, by the second loss function, setting the weight value one of (a) based on and (b) equal to, a minimum one of: custom-character−τ; and λ(custom-character−τ), where custom-character is the first loss, τ is the threshold value, and λ is the regularization hyperparameter that is between 0 and 1”, paragraph 45.
“[0102] At 420, a task loss is computed using a first loss function corresponding to the task to be performed by the neural network. The task loss corresponds to an error in the prediction (relative to a target prediction) for each data sample of the batch. …”, paragraph 102);
And c) training, by the one or more processors, the ML model by updating weights in the ML model based on the first loss.
(see “selectively updating a trainable parameter of the neural network based on the weight value”, paragraph 56 and Fig. 4 elements 440 and 450, paragraph 63.
See also “0102] At 420, a task loss is computed using a first loss function corresponding to the task to be performed by the neural network. The task loss corresponds to an error in the prediction (relative to a target prediction) for each data sample of the batch. The task loss is then input into a second loss function (the SuperLoss function). Based on the task loss, the SuperLoss function computes a second loss (the super loss) for each data sample of the batch at 430. …”, paragraph 102 and Fig. 4).
WEIN does not expressly disclose
… to a power of beta
But from WEIN teaching of a function “ log represents the logarithm function” , “λ is the regularization hyperparameter”, [tunable parameter], (see paragraph 51). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate a loss function log (prediction probability ) to a power of beta [tunable parameter], as it is an obvious variation of the standard log loss or logarithmic losses, used in classification and probability estimation in order to get a model for improving result prediction accuracy.( see WEIN paragraph 5), and the results would have been predictable.
Further, Kaggle teaches that Log Loss is the most important classification metric based on probabilities (Kaggle page 1) . It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate a loss function, log (prediction probability ), to a power of beta [tunable parameter], as it is an obvious variation of the standard log loss or logarithmic losses, used in classification and probability estimation in order to get a model for improving result prediction accuracy.( see WEIN paragraph 5 ), and the results would have been predictable.
As to claim 15, it comprises the same limitations than claim 1 above, therefore it is rejected in the same manner. Further, the claim comprises a non-transitory computer readable medium with instructions stored thereon, wherein the instructions, when executed by one or more processors, causing the one or more processors to perform operations (WEIN paragraphs 36 and 147).
As to claim 20, it comprises the same limitations than claim 1 above, therefore it is rejected in the same manner. Further, the claim comprises an apparatus, comprising:
one or more memories storing instructions; and one or more processors, wherein the one or more processors are configured to execute the instructions to cause the apparatus to train a machine learning (ML) model by performing operations (Fig. 11, paragraph 71 and associated disclosure).
As to claim 2, WEIN discloses further comprising:
computing, based on predictions by the ML model on a second dataset for validation, a second loss using a second loss function (see computing superloss, paragraph 102 and element 4490 of Fig. 4);
determining, based on the second loss, convergence of the ML model;
(see element 450 in Fig. 4
“[0025] In further features, automatically computing a weight of the data sample based on the task loss computed for the data sample may include increasing the weight of the data sample if the task loss is below a threshold value and decreasing the weight of the data sample if the task loss is above the threshold value”, paragraph 25.
“[0050] In further features, automatically computing the weight value includes, by the second loss function, automatically computing the weight value based on a loss amplifying term given by σ*(custom-character−τ), where σ* is the confidence value, custom-character is the first loss, and τ is the threshold value”, paragraph 50. See also paragraphs 44-45.
“[0069] FIG. 10 is a plot showing model convergence during training on the noisy Landmarks-full dataset “, Fig. 10 and associated disclosure); and
output the ML model based on determining that the ML model is converged.
(Fig. 10 and associated disclosure);
As to claim 3, WEIN discloses further comprising:
wherein the second loss function is the same as the first loss function, or
wherein the second loss function is different from the first loss function.
(“At 420, a task loss is computed using a first loss function corresponding to the task to be performed by the neural network. The task loss corresponds to an error in the prediction (relative to a target prediction) for each data sample of the batch. The task loss is then input into a second loss function (the SuperLoss function). Based on the task loss, the SuperLoss function computes a second loss (the super loss) for each data sample of the batch at 430. …”, paragraph 102 and Fig. 4).
As to claim 4, WEIN discloses
wherein the first loss function causes a reduced loss for a first respective prediction of the predictions with a corresponding probability above a predetermined threshold, wherein the first loss function causes an increased loss for a second respective prediction of the predictions with a corresponding probability below the predetermined threshold, and wherein the predetermined threshold is obtained based on a comparison between the first loss function and a cross entropy (CE) loss function without the power term f3.
(paragraphs 23-25).
As to claim 5, WEIN discloses
wherein the second dataset is different from the batch from the dataset.
(see element 420 is different of element 440 ).
As to claims 6 and 16, WEIN discloses
wherein the first loss function is a power cross entropy (PCE) loss function, expressed by:
PNG
media_image1.png
390
1708
media_image1.png
Greyscale
(“classification, the Cross-Entropy loss (CE) may be straightforwardly input to the SuperLoss: SL*.sub.CE=SL*(custom-character.sub.CE(ƒ(x), y)). The threshold value τ may be fixed and set to τ=log C, where C is the number of classes, representing the cross-entropy of a uniform prediction and hence a natural boundary between correct and incorrect prediction”, paragraph 108).
As to claims 7 and 17, WEIN discloses
wherein the first loss function is a power focus loss (PPL) loss function, expressed by:
PNG
media_image2.png
264
1433
media_image2.png
Greyscale
(“[0023] In a feature, a computer-implemented method for training a neural network to perform a data processing task is provided. The method includes, for each data sample of a set of labeled data samples: computing a task loss for the data sample using a first loss function for the data processing task…”, paragraph 23.
“[0038] In a feature, a computer-implemented method for training a neural network to perform a data processing task includes: for each data sample of a set of labeled data samples: by a first loss function for the data processing task, computing a first loss for that data sample…”, paragraph 38.
Further WEIN’s system in paragraph 110 “[0110] Regarding object detection, the SuperLoss function may be applied on the box classification component of two object detection frameworks, such as the faster recursive convolutional neural network (Faster R-CNN) framework. The Faster R-CNN framework is described in …., and Lin, T. Y., et al., Focal Loss for dense object detection, ICCV, 2017, which are incorporated herein in their entireties. …”, paragraph 110.
“Figure 1. We propose a novel loss we term the Focal Loss that
adds a factor (1 pt) to the standard cross entropy criterion.
Setting > 0 reduces the relative loss for well-classified examples
(pt > :5), putting more focus on hard, misclassified examples. As
our experiments will demonstrate, the proposed focal loss enables
training highly accurate dense object detectors in the presence of
vast numbers of easy background examples”, “Focal Loss for dense object detection incorporated by reference”, Col 1, page 1 and numeral “3. Focal loss”, Col 1 and Col 2 in page 3).
As to claim 11, WEIN discloses
wherein the dataset comprises a plurality of samples labeled with a plurality of classes, and wherein first samples of the plurality of samples associated with a first class of the plurality of classes are significantly fewer than second samples of the plurality of samples associated with a second class of the plurality of classes.
(WEIN system discloses multiple data samples labeled with a plurality of classes, “…for each data sample of a set of labeled data samples: by a first loss function for the data processing task, computing a first loss for that data sample…”, abstract and paragraph 23.
“C is the number of classes”, paragraphs 12. See also paragraphs “sample …has label”, paragraphs 10-11.
“[0075] The neural network is configured to process each of the data samples to generate a prediction (e.g., a label) for the respective data sample. The loss function corresponding to the task that the neural network is being trained to perform (also referred to herein as a first loss function) indicates the error between the prediction output by the neural network based on a data sample and a target value for the data sample. For example, … the neural network generates a prediction for the label of the data sample. The predicted label for each data sample is then compared to the ground truth label of the data sample. The difference (the error) between the ground truth label and the predicted label is a task loss output by the neural network. …”, paragraph 75.
Further, WEIN and “Focal Loss for dense object detection incorporated by reference” discloses imbalance in the data [Examiner interprets as a first class of the plurality of classes are significantly fewer than second samples of the plurality of samples associated with a second class ] see page 3, Col. 2, where the weighting value of the samples are between zero and 1).
As to claims 12 and 18, WEIN discloses
determining a first value for the tunable parameter f3 in the first loss function for training the ML model; and
adjusting the tunable parameter f3 in the first loss function from the first value to a second value during training of the ML model.
(“[0114] The neural network model trained with the original task loss is referred to as the baseline. The protocol involved first training the baseline and tuning its hyperparameters (e.g., learning rate, weight decay, etc….”, paragraph 114).
As to claim 13, WEIN discloses
wherein the first loss function further comprises one or more other tunable parameters different from the tunable parameter f3, the method further compnsmg:
determining third values for the one or more other tunable parameters in the first loss function for training the ML model; and
adjusting the one or more other tunable parameters in the first loss function from the third values to fourth values during training of the ML model.
(“[0114] The neural network model trained with the original task loss is referred to as the baseline. The protocol involved first training the baseline and tuning its hyperparameters (e.g., learning rate, weight decay, etc.)[Examiner interprets as determining third values for the one or more other tunable parameters in the first loss function for training the ML model] …. For a fair comparison between the baseline and the SuperLoss function, the model is trained with the SuperLoss function with the same hyperparameters…”, paragraph 114)
As to claim 14, WEIN discloses
dynamically adjusting at least one of the tunable parameter f3 and the one or more other tunable parameters in the first loss functions at different stages of the training
(“…The protocol involved first training the baseline and tuning its hyperparameter…”, paragraph 114).
Claims 8-10 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over US PG. Pub. No. 20220114444 ((WEINZAEPFEL called herein WEIN) in view of https://www.kaggle.com/code/dansbecker/what-is-log-loss/notebook (Kaggle)
in view of US PG. Pub. No. 20130346182 (Cheng).
As to claims 8 and 19, WEIN discloses
wherein the ML model is trained …
(“…data … may be one of image samples, video samples, text content samples and audio samples.”, paragraphs 23-29.
“…a neural network trained using the method above to perform a data processing task is provided. The data processing task may be an image processing task. The image processing task may be one of classification, regression, object detection and image retrieval”, paragraph 35).
Wein does not teach but Cheing discloses
trained for predicting click- through-rate (CTR) for online advertising, and wherein the CTR indicates a percentage of impressions of advertisements that are actually clicked on by users.
(Cheng teaches “… a click prediction model for improving click prediction accuracy….”, abstract.
“…ads can choose to pay per impression ("CPM"). However, advertisers may also prefer to pay if the ad attracted the user's attention…”, paragraphs 12, 24, 28-29,
“…A maximum entropy algorithm may be used for this supervised learning task because of its simplicity and strength in combining diverse features and large scale learning. The maximum-entropy model, also known as logistic regression, may have the following form…”, paragraph 43.
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Cheng’s teaching with the teaching of WEIN. One would have been motivated to use functionality to support
analysis of historical click data for advertisements in order to provide a click prediction model (see at least claim 13 of Cheng).
As to claim 9, WEIN discloses
wherein the dataset comprises samples with binary values, and wherein a value of "O" indicates a non-click prediction, and a value of" l" indicate a click prediction.
(see WEIN paragraphs 14 and 16, WEIN uses values between o and 1 for confidence and reliability variables in WEIN’s system.
See also “3.1. Balanced Cross Entropy
A common method for addressing class imbalance is to introduce a weighting factor … In practice may be set by inverse class frequency or treated as a hyperparameter to set by cross validation. For notational convenience, we define t analogously to how we defined pt. We write the -balanced CE loss as:… This loss is a simple extension to CE that we consider as an experimental baseline for our proposed focal loss”, “Focal Loss for dense object detection incorporated by reference”, page 3, Col 2).
As to claim 10, WEIN discloses
wherein the number of samples with a value "l" in the dataset is smaller than the number of samples with a value "O" by a threshold multiple.
(WEIN and “Focal Loss for dense object detection incorporated by reference” discloses imbalance in the data see page 3, Col. 2, where the weighting value of the samples are between zero and 1).
PNG
media_image3.png
331
652
media_image3.png
Greyscale
Response to Arguments
Applicant’s arguments of 03/06/2026 have been very carefully considered but are not persuasive.
Applicant argues (page 7-11)
Rejections under §101
The Office action rejected claims 1-20 under 35 U.S.C. §101 as being directed to an
abstract idea without significantly more. Solely for the sake of brevity, Applicant has not
repeated the entire remarks set forth in the response to previous Office action filed on
September 23, 2025, but still respectfully submits that claims 1-20 are directed to patent eligible
subject matter pursuant to the 2019 guidance and MPEP § 2106….
Step 2A - Prong One
The Office action alleges that claim 1 recites an abstract idea. Specifically, the Office
states that the features of claim 1 are abstract ideas directed to "certain methods of organizing
human activity" and "mathematical concepts." See Office action, p. 3.
Applicant respectfully submits that the features of claim 1 are not directed to certain
methods of organizing human activity. MPEP § 2106.04(a)(2) states that the grouping of certain
methods of organizing human activity is limited to "activity that falls within the enumerated
sub-groupings of... managing personal behavior and relationships or interactions between
people, and is not to be expanded beyond these enumerated sub-groupings except in rare
circumstances." MPEP § 2106.04(a)(2) further states that this sub-grouping includes social
activities, teachings, and following rules or instructions….
In response the Examiner asserts that the rejection of the claims under 101 is maintained because the claims are directed to an abstract idea, which is a judicial exception.
The Examiner acknowledges that the applicant amended the claims, this gave some structure to the claims but with a very generic computer components. The Facially sufficient analysis conducted, determined in step 2A, the claims are directed to a method for training a machine learning for predicting click-through-rate for online advertisement. Thus, the claim recites an abstract idea. since, step 2A is
determined to be directed to a judicial exception, then (Step 2B) the claim is analyzed to determine that the claim recites any element or combination of elements such that the claim as a whole amounts to more than the judicial exception itself . Here the claim elements additional to the abstract idea are “training a machine learning (ML) model”, “the (ML) model”, “computer-implemented, ,to perform a data processing task, by one or more processors, wherein the ML model is a neural network model; training, by the one or more processors, the ML model;” These elements are recited in a very generic way. The claim(s) do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the additional elements are generic computer components claimed to perform their basic functions of the invention.
All the elements were analyzed individually and as a whole. No element has been ignored. A prima facie of unpatentability has been established. The prima facie case of unpatentability is a bar towards 101 eligibility.
Furthermore, contrary to the opinions of the Office, adjustment of weights does not
comprise a mathematical relationship, formula, or calculation. In Example 38 of the USPTO
provided Subject Matter Eligibility Examples: Abstract Ideas ("Example 38"), the Office
indicates that claims that describe using a normally distributed pseudo random number
generator to determine a randomized working value of each element is not directed to a
mathematical concept, because the claims of Example 38 does not recite a mathematical
relationship, formula, or calculation. Similar to Example 38, the features of the independent
claims merely recite using loss computed by a first loss function to update weights in the ML
model. But the independent claim 1 does not recite a mathematical relationship, formula, or
calculation. Consequently, Applicant submits that the Office erroneously determines that the
claimed subject matter merely recites mathematical concepts.
Therefore, the claims are patent eligible under Prong One of Step 2A.
In response the Examiner asserts that this case is not rejected under 101 only because the invention ability to run on a generic processor or general purpose computer, but also because the detail facially sufficient analysis provided above, where the Examiner looked both the instant claims and the specification to elaborate Examiner's facially sufficient analysis. The additional elements in the instant claims do not provide significantly more to the abstract idea identified above, as the additional elements do not: Improve another technology or technical field; Improve the functioning of a computer itself; Add a specific limitation other than what is well-understood, routine, and conventional in the field; Add meaningful limitations that amount to more than generally linking the use of the exception to a particular technological environment; Improve computer related technology by allowing computer performance of a function not previously performable by a computer. (see MPEP 2106.05). The prima facie of unpatentability established in this case is not in error because the carefully and detail facially analysis conducted to establish the prima facie of unpatentability.
Applicant argues that this case is similar to Example 38, the Examiner does not agree that the instant claims have the same fact pattern than Example 38 -Simulating an Analog Audio Mixer -. Instead, the Examiner finds that Example 39, -Method for Training a Neural Network for Facial Detection – comprises training a neural network model. In Example 39, the claim 1 does not recite any of the judicial exceptions enumerated in the 2019 PEG. For instance, the claim does not recite any mathematical relationships, formulas, or calculations. While some of the limitations may be based on mathematical concepts, the mathematical concepts are not recited in the claims. Further, the claim does not recite a mental process because the steps are not practically performed in the human mind. Finally, the claim does not recite any method of organizing human activity such as a fundamental economic concept or managing interactions between people. Thus, the claim is eligible because it does not recite a judicial exception. Contrary, the instant claims predict probability of interaction, which is a human activity, advertising, marketing or it can be considered also mathematical concepts grouping of abstract ideas. (see MPEP 2106.04).
Next, the Examiner directs applicant’s attention to Example 47, the training requires mathematical calculations, like in the instant claims , for that reason it is not eligible for 101 protection.
Step 2A - Prong Two
Moreover, regardless of whether the independent claims are viewed as being directed
to an exception under Prong One of Step 2A of the Subject Matter Eligibility Test, it is
respectfully submitted that independent claims provide a practical application under Prong Two
of Step 2A of the Subject Matter Eligibility Test. As set forth in the 2019 Guidance, an element
or combination of elements that reflect a technical improvement (e.g., to a computer or other
technology) provide that a claim is directed to a practical application.
The Office action states that the claims do not contain any additional elements that
would integrate the judicial exception into practical application. See Office action, pp. 3-4.
Assuming arguendo that the claims are not automatically patent eligible after the Prong
One analysis described above, Applicant submits that additional elements recited in the claims
are being ignored by the Office, and that the claims recite features that integrate any alleged
abstract idea into a practical application, making the claims patent eligible. As noted above,
amended claim 1 recites "training, by the one or more processors, the ML model by updating
weights in the ML model based on the first loss."
MPEP § 2106.0S(a) explicitly states that "[a]n indication that the claimed invention
provides an improvement can include a discussion in the specification that identifies a technical
problem and explains the details of an unconventional technical solution expressed in the claim,
or identifies technical improvements realized by the claim over the prior art." Here, the
published specification 1 describes a number of technical problems that are overcome by the
features of the claims…..
In response the Examiner asserts that at Step 2A Prong Two or Step 2B, the facially sufficient analysis presented above concludes that the exception, recited in the instant claims, is not integrated into a practical application or that the additional elements do not amount to significantly more than the exception.
The Examiner asserts that the additional limitations to the abstract idea recited herein, “training a machine learning (ML) model”, “the (ML) model”, “computer-implemented, ,to perform a data processing task, by one or more processors, wherein the ML model is a neural network model; training, by the one or more processors, the ML model;” are well-understood, routine, conventional activities in Step 2B. again , in the facially sufficient analysis above, the elements were considered insignificant extra-solution activity or well- understood, routine, or conventional activity evidence was provided according to the MPEP 2106.05(h), 2106.05(f) and 2106.05(g). A prima facie of unpatentability has been established.
Lasty, In response the Examiner asserts per MPEP 2106 an invention must have to comply with the Subject Matter Eligibility test under Alice framework, see also (2019 PEG), but also with the MPEP 2111 that provides that claims must be given their broadest reasonable interpretation and it is generally considered improper to read limitations contained in the specification into the claims. See In re Prater, 415 F.2d 1393, 162 USPQ 541 (CCPA 1969) and In re Winkhaus, 527 F.2d 637, 188 USPQ 129 (CCPA 1975), which discuss the premise that one cannot rely on the specification to impart limitations to the claim that are not recited in the claim. It seems that applicant wants that the Examiner reads limitations from the specification into the claims.
Applicant also refers to the recent Ex Parte Desjardins Memo for further guidance that
claims that provide improvements to machine learning like the present claims provide a
practical application. The Ex Parte Desjardins Memo provides that the Appeals Review Panel
("ARP") evaluated the claims as a whole and discerned that a certain claim limitation of the
application in issue reflected the improvement disclosed in the specification, and as such, the
claims as a whole integrated what would otherwise be a judicial exception into a practical
application at Step 2A - Prong Two. The ARP, in determining this holding, also determined
that the specification identified improvements as to how the machine learning model itself
operated, including training machine learning model to learn new tasks while protecting
knowledge about previous tasks to overcome the problem of "catastrophic forgetting"
encountered in continual learning systems. Similarly, as discussed above and in the present
application, the claims are directed to improving existing machine learning models using a
power loss function, which decreases the loss for confident predictions and increases the loss
for error-prone predictions, to dynamically adjust parameters of the machine learning model
and increase click-through rate prediction accuracy for more difficult, less common instances.
Thus, especially when reviewed under the recent guidance provided in the Ex Parte
Desjardins Memo, the claims integrate any alleged abstract idea into a practical application by
providing improvements to machine learning models, in particular the training of the machine
learning models (e.g., a technological field), which is evidenced by the technological problems
and solutions described in the specification, and which are reflected in the current claims. It is
respectfully submitted that the Step 2A - Prong Two analysis included in the Office action is
not in line with the subsequently issued guidance provided in the Ex Parte Desjardins Memo
by relying on an oversimplification of the claims that evaluates the claims a high level of
generality without considering the claim as a whole and dismisses the potentially meaningful
technical limitations, which have been described in the published specification, without
adequate explanation. Accordingly, it is respectfully submitted that the § 101 rejections should
be withdrawn, and it is further respectfully submitted that the claims integrate the alleged
abstract idea into a practical application based on the updated guidance provided by the Ex
Parte Desjardins Memo for at least the reasons provided above.
In response the Examiner asserts that based in the facially sufficient analysis presented above to establish the prima facie of unpatentability in this case, the instant claim 1 does not appear to show any improvement technique to the machine learning model. In other words, always a machine learning model needs to update weights and parameters to reflect new patters and increment learning. That is considered to be extra-solution activity in Step 2A, and also that is well understood, routine and conventional activity supported under Berkheimer Option 2, when this activity is re-evaluated in Step 2B of the 2019 PEG.
The Examiner notes that Not all the inventions that claim training a machine learning model and updating parameters or weights are eligible for patent protection.
In Desjadins claim 1, the ARP evaluated claim 1 as a whole and found that the claim in question recited not merely abstract algorithms but represented a "technical improvement in training the machine learning model itself". Specifically, the ARP decision evaluated the claim as a whole in discerning at least the limitation “adjust the first values of the plurality of parameters to optimize performance of the machine learning model on the second machine learning task while protecting performance of the machine learning model on the first machine learning task” reflected the improvement disclosed in the specification.
Thus, comparing the instant claim 1 to Desjandins claim 1, there is not any improvement to the way of training the machine learning model itself in the instant claim. Again, always a machine learning model requires the generic steps of training and updating to incorporate fresh input of new data to combat performance degradation. For that reason, always the parameters (weights) are adjusted to reflect the new patterns.
In summary, merely training and updating a neural network model to combat performance degradation, like is recited in the instant claims, cannot make an otherwise nonstaturory claim statutory.
Applicant argues (page 11)
Rejections under § 103
The Office action rejected claims 1-7, 11-18, and 20 under 35 U.S.C. §103 as being
unpatentable over Weinzaepfel (U.S. Patent Pub. No. 2022/0114444) in view of Kaggle
(https://www.kaggle.com/code/dansbecker/what-is-log-loss/notebook). In addition, the Office
action rejected claims 8-10 and 19 under 35 U.S.C. §103 as being unpatentable over
Weinzaepfel and Kaggle in view of Cheng (U.S. Patent Pub. No. 2013/0346182).
Claim 1 recites:….
In claim 1, the first loss function comprises a term that calculates a natural logarithm,
e.g., "log( )", of a prediction probability, e.g., "p", to a power of /J, where fJ is a tunable
parameter and is a real number. Written as an equation, the first loss function in claim 1
includes a term that calculates:
log(p8)
None of the cited references discloses or suggests a loss function that includes such a term.
Applicant’s arguments filed have been fully considered. A prima facie of obviousness has been established. The legal concept of prima facie is herein supported by evidence (see MPEP 2142).
The combination WEIN, Kaggle and Cheng, comprises a prima facie of obviousness . The prima facie discloses all the limitations on the claims. The Examiner respectfully notes that Applicant has not provided persuasive rebuttal evidence to overcome the prima facie case. The combination set for the rejection produce results that are predictable. When a prima facie case is established, the burden shifts to applicant to come forward with rebuttal evidence or argument to overcome the prima facie case. The Examiner respectfully notes that Applicant has not provided rebuttal evidence to overcome the prima facie case.
For instance, the loss function, as described in Weinzaepfel (i.e., A(log 0*)2) includes a
regularization hyperparameter that A is multiplied to the output of the log function. By contrast,
in claim 1, the log function is applied to the prediction probability, e.g., "p", taken to a power
of beta. Even if it is assumed that the hyperparameter A of Weinzaepfel teaches the power of beta,
which is not conceded, Weinzaepfel makes no mention of the hyperparameter being related to
the confidence value 0*. Because there is no mention of the confidence value of Weinzaepfel
being raised to any power, let alone a tunable power, one of skill in the art would not look to
the hyperparameter A of Weinzaepfel to show or suggest the power of beta as recited in
independent claim 1.
In response the Examiner asserts first of all, the Examiner is given the broadest reasonable interpretation to the claims consistent with the interpretation that those skilled in the art (See MPEP 2111). Secondly, it is improper to import claim limitations from the specification, therefore the Examiner is given the broadest reasonable interpretation to the claims.
Thirdly, regarding to the argument “…Weinzaepfel makes no mention of the hyperparameter being related to the confidence value 0*. Because there is no mention of the confidence value of Weinzaepfel being raised to any power, let alone a tunable power …”, it seems that the applicant is arguing limitations that are not in the instant claims.
And Finally, WEIN teaches in paragraph 51 a log of a confidence value or probability, raised to a parameter, this teaching is similar in mathematical behavior than, an algorithm of a prediction probability to a power parameter. Moreover, applicant is claiming the term “ tunable parameter” , WEIN discloses in his system the term “regularization hyperparameter”, both terms are often used interchangeably in machine learning to describe settings that are configured by a human before training. Therefore, there is not novelty in the claimed limitation.
Additionally, the Office suggests that Kaggle teaches this subject matter at page 1. See
Office action, p. 11. Applicant respectfully disagrees because the cited portion of Kaggle is
completely silent regarding raising the parameter of the log likelihood function to any power,
let alone a power of beta as recited in independent claim 1. At best, Kaggle merely describes
utilizing a logarithmic function to estimate losses, but is silent regarding whether the input of
the loss function is raised to any power, let alone a tunable parameter power of /J, as recited in
independent claim 1.
Cheng fails to cure the deficiencies of Weinzaepfel and Kaggle with respect to the claim
terms discussed above.
As the foregoing illustrates, the cited references fail to disclose or suggest each and
every feature of claim 1. Therefore, the cited references cannot render obvious claim 1.
The Examiner assets that one cannot show nonobviousness by attacking references individually where the rejections are based on combinations of references. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981); In re Merck & Co., 800 F.2d 1091, 231 USPQ 375 (Fed. Cir. 1986). See MPEP 2145.
The prima facie of obviousness established in this case is combination WEIN, Kaggle and Cheng. The prima facie discloses all the limitations on the claims.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
“Instance Weighting in Neural Networks for Click-Through Rate Prediction”. IEEE. 2023. This article discloses “The instances on which a learning algorithm is most undecided can be weighted more during training to guide the learning model towards spending more effort on the difficult instances. We introduce three instance weighting algorithms to weigh the loss obtained. All of these instance weighting methods improve loss, AUC, and F1 results. We demonstrate the improvements on four different classifiers and on two different datasets. The improvements in loss reach 2.8% and in AUC reach 0.73% for Masknet on the Avazu dataset.”
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MARIA VICTORIA VANDERHORST whose telephone number is (571)270-3604. The examiner can normally be reached on business hours from Monday through Friday from 8:30 AM to 4:30 PM.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ashraf Waseem can be reached on 571-270-3948. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MARIA V VANDERHORST/ Primary Examiner, Art Unit 3621 5/1/2026