DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Claims 1-20 are subject to review.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on September 2, 2022 is being considered by the examiner.
Response to Arguments
Applicant's arguments filed 9/11/2025 have been fully considered.
Regarding applicants’ remarks directed to the rejection of claims under 35 USC 112(b), after further review of the amended limitations, the rejection made in the previous rejection has been withdrawn.
Rejection of claims under 35 USC 101: Abstract idea, see pages 7-8 of filed remarks.
Regarding applicants’ remarks directed to the rejection of claims under 35 USC 101: Abstract idea, the examiner maintains the rejection made in the previous rejection.
Applicant argues that claims are directed to patent eligible subject matter because the claims when examined as a whole are directed to an improvement in technology or technical field. The applicant also quotes the USPOTO Memorandum to conclude that the analysis was insufficient.
Examiner disagrees with the applicant’s allegations for the noted reasons below.
First, the examiner notes that the MPEP serves as the official guidance for establishing the rules for examination and this is also reiterated in the cited Memorandum, issued August 4th , 2025:
“ … This memorandum is not intended to announce any new USPTO practice or procedure and is meant to be consistent with existing USPTO guidance. Examiners should consult the specific MPEP sections referenced below for more thorough information on each topic…Analysis of claim as a whole: The analysis in Step 2A Prong Two considers the claim as a whole. The way in which the additional elements use or interact with the exception may integrate the judicial exception into a practical application. Accordingly, the additional limitations should not be evaluated in a vacuum, completely separate from the recited judicial exception. Instead, the analysis should take into consideration all the claim limitations and how these limitations interact… The examiner is reminded to consult the specification to determine whether the disclosed invention improves technology or a technical field, and evaluate the claim to ensure it reflects the disclosed improvement. The specification does not need to explicitly set forth the improvement, but it must describe the invention such that the improvement would be apparent to one of ordinary skill in the art….”
Second, the MPEP 2106.04(d)(1) discloses the evaluation of claimed improvements in the functioning of a computer or improvement to a technical field in step 2A prong two. The MPEP section discloses “if the specification explicitly sets forth an improvement but in a conclusory manner (i.e., a bare assertion of an improvement without the detail necessary to be apparent to a person of ordinary skill in the art), the examiner should not determine the claim improves technology. Second, if the specification sets forth an improvement in technology, the claim must be evaluated to ensure that the claim itself reflects the disclosed improvement. That is, the claim includes the components or steps of the invention that provide the improvement described in the specification…”
The applicant remarks are deficient, they do not provide any recitation in the specification that supports the allegations that the claims as a whole are directed to an improvement and the claim limitations recite a Mathematical concept: mathematical relationships, see MPEP § 2106.04(a)(2), subsection I. A mathematical relationship is a relationship between variables or numbers. A mathematical relationship may be expressed in words or using mathematical symbols. The claims in the instant case include limitations directed to a mathematical relationship, per the MPEP guidance.
The additional element that must be considered, in step 2A prong two and 2B , as the obtaining step was examined by the examiner per the noted pages in applicant’s response. The examiner has documented per the MPEP that the noted addition element is directed to activity the courts have deemed well-known, routine and conventional, See MPEP 2106.05(d)(II) : Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TLI Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016) (using a telephone for image transmission); OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network); but see DDR Holdings, LLC v. Hotels.com, L.P., 773 F.3d 1245, 1258, 113 USPQ2d 1097, 1106 (Fed. Cir. 2014) ("Unlike the claims in Ultramercial, the claims at issue here specify how interactions with the Internet are manipulated to yield a desired result‐‐a result that overrides the routine and conventional sequence of events ordinarily triggered by the click of a hyperlink." (emphasis added));
These types of claim limitations do not amount to what the courts have deemed significantly more, per MPEP guidelines, and thus the claims when considered as a whole and in combination do not amount to eligible subject matter.
The rejection made in the pervious office analysis documented by the previous examiner was prima facia enough to support the rejection of claims under 35 USC 101 Abstract idea.
The rejection made in the pervious office action has been maintained.
Rejection of claims under 35 USC 103, See pages 8-10 of filed remarks.
Regarding applicants’ remarks directed to the rejection of claims under 35 USC 103, the examiner maintains the rejection made in the previous rejection.
The applicant argues that the examiner has failed to make a prima facia case of obviousness because the documented motivation were interpreted as mere descriptions of implementing choices disclosed by the cited prior art.
Examiner notes that the requirements for making a prima facia case of obviousness are documented in MPEP 2143. And per applicant’s own remarks, pages 9-11, the motivation statement indicates the rationale for a more desirable performance, such as “In other words, the gradient solver may use the implicit function gradient directly to iteratively improve the outer problem solution, which may be similar to gradient descent method."
The examiner has followed guidelines per MPEP 2143 (I)(G) where the courts have provided that such documentation is sufficient for making a prima facie case of obvious per the citation below as the examiner has noted that one of ordinary skill in the art would recognized the cited prior arts as analogous and has summarized the benefits.
See the relevant support of examiner’s rejecter per MPEP 2143 (I)(G)
“The courts have made clear that the teaching, suggestion, or motivation test is flexible and an explicit suggestion to combine the prior art is not necessary. The motivation to combine may be implicit and may be found in the knowledge of one of ordinary skill in the art, or, in some cases, from the nature of the problem to be solved. Id. at 1366, 80 USPQ2d at 1649. "[A]n implicit motivation to combine exists not only when a suggestion may be gleaned from the prior art as a whole, but when the ‘improvement’ is technology-independent and the combination of references results in a product or process that is more desirable, for example because it is stronger, cheaper, cleaner, faster, lighter, smaller, more durable, or more efficient. Because the desire to enhance commercial opportunities by improving a product or process is universal—and even common-sensical—we have held that there exists in these situations a motivation to combine prior art references even absent any hint of suggestion in the references themselves. In such situations, the proper question is whether the ordinary artisan possesses knowledge and skills rendering him capable of combining the prior art references." Id. at 1368, 80 USPQ2d at 1651.”
Once a prima facie case of obviousness is established, the burden shifts to the applicant to come forward with arguments and/or evidence to rebut the prima facie case. See, e.g., In re Dillon, 919 F.2d 688, 692, 16 USPQ2d 1897, 1901 (Fed. Cir. 1990) (en banc).
The applicant’s argument, in the instant case, amount to mere allegations of patentability and the rejections made in the pervious office action has been maintained.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 of the subject matter eligibility test (see MPEP 2106.03).
Claims 1-14 are directed to a “method” which describes one of the four statutory categories of patentable subject matter, i.e., a process. 35 U.S.C. 100(b).
Claims 15-20 are directed to a “computing system” which describes one of the four statutory categories of patentable subject matter, i.e., a machine.
Regarding Claim 1:
Step 2A of the subject matter eligibility test (see MPEP 2106.04).
Prong One:
Claim 1 recites (“sets forth” or “describes”) the abstract idea, substantially as follows:
“initializing, by the computing system, an initial plurality of machine-learned model parameters and an initial at least one threshold such that the initial plurality of machine-learned model parameters and the initial at least one threshold satisfy a constraint function;” – This limitation is recited at a high level of generality and is directed to the abstract idea of a mental process i.e., initializing parameters (concepts performed in the human mind, including evaluation and judgment [MPEP 2106.04(a)(2) III. C.]), and may be performed with the aid of pen and paper, or using a computer as a tool.
“determining, by the computing system, a gradient of an objective function with respect to the plurality of machine-learned model parameters at a current optimization step based at least in part on an implicit function of the at least one threshold as a function of the plurality of machine- learned model parameters; and” – is directed to the abstract idea of mathematical concepts (see MPEP 2106.04(a)(2)) as it is describing calculating a gradient of the objective function, is considered to be mathematical calculation.
“Updating, by the computing system, the plurality of machine-learned model parameters and the at least one threshold based at least in part on the gradient.” – This limitation is recited at a high level of generality and is directed to the abstract idea of a mental process i.e., updating a parameter and threshold based on, in part, a gradient (concepts performed in the human mind, including observation, evaluation and judgment [MPEP 2106.04(a)(2) III. C.]), and may be performed with the aid of pen and paper, or using a computer as a tool.
Prong Two: Claim 1 does not include additional elements that integrate the mental process into a practical application.
“Obtaining, by a computing system comprising one or more computing devices, data indicative of a plurality of machine-learned model parameters and at least one threshold comprising a machine-learned model;” – is merely a recitation of an insignificant extra-solution data gathering (see MPEP § 2106.05(g)), which does not integrate a judicial exception into practical application.
Therefore, the additional elements, alone or in combination, do not integrate the abstract idea into a practical application. See MPEP 2106.04(d).
Step 2B of the subject matter eligibility test (see MPEP 2106.05).
When considered individually or in combination, the additional limitations and elements of claim 1 does not amount to significantly more than the judicial exception for the reasons as discussed above as to why the additional limitations do not integrate the abstract idea into a practical application.
“Obtaining, by a computing system comprising one or more computing devices, data indicative of a plurality of machine-learned model parameters and at least one threshold comprising a machine-learned model;” – is merely a recitation of an insignificant extra-solution data gathering (see MPEP § 2106.05(g)), which does not integrate a judicial exception into practical application. Further, the insignificant extra-solution data gathering is also WURC, see MPEP § 2106.05(d)(II) “The courts have recognized the following computer functions as well‐understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity. i. Receiving or transmitting data over a network”. Therefore, the limitation does not integrate the abstract idea in to a practical application, nor does significantly more.
Claim 2 is clarifying the mathematical operation to determine the gradient and therefore is directed to the abstract idea of mathematical concepts (see MPEP 2106.04(a)(2)) as it is describing how to calculate the gradient, including calculating a derivative and a gradient of the derivative, and is therefore considered to be mathematical calculation.
Claim 3 is clarifying the mathematical operation to determine the gradient and therefore is directed to the abstract idea of mathematical concepts (see MPEP 2106.04(a)(2)) as it is clarifying the derivative of the implicit function used to calculate the gradient, is considered to be mathematical calculation.
Claim 4 is clarifying the mathematical operation to determine the gradient and therefore is directed to the abstract idea of mathematical concepts (see MPEP 2106.04(a)(2)) as it is clarifying the derivative of the objective function used to calculate the gradient, is considered to be mathematical calculation.
Claim 5 is clarifying the mathematical operation to determine the gradient and therefore is directed to the abstract idea of mathematical concepts (see MPEP 2106.04(a)(2)) as it is clarifying the derivative of the objective function, sum of the ratios of partial derivatives and constraint function used to calculate the gradient, is considered to be mathematical calculation.
Claim 6 is clarifying the initializing is from a parameter distribution and therefore is directed to the abstract idea of mathematical concepts (see MPEP 2106.04(a)(2)) as it is clarifying the initialization is from a parameter distribution is considered to be mathematical relationship.
Claim 7 is clarifying how to perform the updating of parameters and is therefore directed to the abstract idea of a mental process i.e., updating a parameter based on, in part, the gradient (concepts performed in the human mind, including observation and evaluation [MPEP 2106.04(a)(2) III. C.]), and may be performed with the aid of pen and paper, or using a computer as a tool.
Claim 8 is clarifying how to perform the updating of the at least one threshold and is therefore directed to the abstract idea of a mental process i.e., updating a at least one threshold based on, in part, the gradient (concepts performed in the human mind, including observation and evaluation [MPEP 2106.04(a)(2) III. C.]), and may be performed with the aid of pen and paper, or using a computer as a tool.
Claim 9 is clarifying how to perform the updating of the at least one threshold and is therefore directed to the abstract idea of a mental process i.e., updating a at least one threshold based on, in part, the gradient (concepts performed in the human mind, including observation and evaluation [MPEP 2106.04(a)(2) III. C.]), and may be performed with the aid of pen and paper, or using a computer as a tool.
Claim 10 is clarifying the mathematical operation to determine the gradient and therefore is directed to the abstract idea of mathematical concepts (see MPEP 2106.04(a)(2)) as it is clarifying gradient includes a regularization cost and is considered to be mathematical calculation or stating a mathematical formula in prose.
Claim 11 is clarifying that the objective function or constraint function comprises a smooth differentiable surrogate and is therefore directed to the abstract idea of mathematical concepts (see MPEP 2106.04(a)(2)) as it is considered to be stating a mathematical formula in prose, additionally as the objective function or constraint function is used to calculate the gradient the claim is further clarifying the mathematical calculation of claim 1.
Claim 12 is clarifying that smooth differentiable surrogate is a sigmoid function or a softplus function and is therefore directed to the abstract idea of mathematical concepts (see MPEP 2106.04(a)(2)) as it is considered to be stating a mathematical formula in prose, additionally as the objective function or constraint function, which the smooth differentiable surrogate is a part of, is used to calculate the gradient the claim is further clarifying the mathematical calculation of claim 1.
Claim 13 is stating that the objective function and constraint function are selected for at least precision at fixed recall metric, FNR at fixed FPR metric, the precision at K metric, the AUC-PR metric, the AUC-ROC metric or fairness criterion metric, this limitation is recited at a high level of generality and is directed to the abstract idea of a mental process i.e., selecting an objective function and constraint function based on certain metrics (concepts performed in the human mind, including observation and evaluation [MPEP 2106.04(a)(2) III. C.]).
Claim 14 is reciting is merely a recitation of an insignificant extra-solution data transmitting (see MPEP § 2106.05(g)), which does not integrate a judicial exception into practical application. Further, the insignificant extra-solution of data transmitting is also WURC, see MPEP § 2106.05(d)(II) “The courts have recognized the following computer functions as well‐understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity. i. Receiving or transmitting data over a network”. Therefore, the limitation does not integrate the abstract idea in to a practical application, nor does significantly more.
Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A Prong 2. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B.
Regarding Claim 15:
Step 2A of the subject matter eligibility test (see MPEP 2106.04).
Prong One:
Claim 15 recites (“sets forth” or “describes”) the abstract idea, substantially as follows:
“initializing, by the computing system, an initial plurality of machine-learned model parameters and an initial at least one threshold such that the initial plurality of machine-learned model parameters and the initial at least one threshold satisfy a constraint function;” – This limitation is recited at a high level of generality and is directed to the abstract idea of a mental process i.e., initializing parameters (concepts performed in the human mind, including evaluation and judgment [MPEP 2106.04(a)(2) III. C.]), and may be performed with the aid of pen and paper, or using a computer as a tool.
“determining, by the computing system, a gradient of an objective function with respect to the plurality of machine-learned model parameters at a current optimization step based at least in part on an implicit function of the at least one threshold as a function of the plurality of machine- learned model parameters; and” – is directed to the abstract idea of mathematical concepts (see MPEP 2106.04(a)(2)) as it is describing calculating a gradient of the objective function, is considered to be mathematical calculation.
“updating, by the computing system, the plurality of machine-learned model parameters and the at least one threshold based at least in part on the gradient.” – This limitation is recited at a high level of generality and is directed to the abstract idea of a mental process i.e., updating a parameter and threshold based on, in part, a gradient (concepts performed in the human mind, including observation and evaluation [MPEP 2106.04(a)(2) III. C.]), and may be performed with the aid of pen and paper, or using a computer as a tool.
Prong Two: Claim 15 does not include additional elements that integrate the mental process into a practical application.
“one or more processors; and” – does not particularly nor specifically identify a machine because a processor is a generic computer component within a generic computer environment and does not integrate the abstract idea in to a practical application.
“one or more computer-readable memory devices storing instructions that, when implemented, cause the one or more processors to perform operations, the operations comprising:” – does not particularly nor specifically identify a machine because a computer-readable memory is a generic computer component within a generic computer environment and does not integrate the abstract idea in to a practical application.
“obtaining, by a computing system comprising one or more computing devices, data indicative of a plurality of machine-learned model parameters and at least one threshold comprising a machine-learned model;” – is merely a recitation of an insignificant extra-solution data gathering (see MPEP § 2106.05(g)), which does not integrate a judicial exception into practical application.
Therefore, the additional elements, alone or in combination, do not integrate the abstract idea into a practical application. See MPEP 2106.04(d).
Step 2B of the subject matter eligibility test (see MPEP 2106.05).
When considered individually or in combination, the additional limitations and elements of claim 15 does not amount to significantly more than the judicial exception for the reasons as discussed above as to why the additional limitations do not integrate the abstract idea into a practical application.
“one or more processors; and” – Invokes a computer merely as a tool for performing an existing process (see MPEP 2106.05(f)(2)) and therefore fails to amount to significantly more than the judicial exception.
“one or more computer-readable memory devices storing instructions that, when implemented, cause the one or more processors to perform operations, the operations comprising:” – Invokes a computer merely as a tool for performing an existing process (see MPEP 2106.05(f)(2)) and therefore fails to amount to significantly more than the judicial exception.
“obtaining, by a computing system comprising one or more computing devices, data indicative of a plurality of machine-learned model parameters and at least one threshold comprising a machine-learned model;” – is merely a recitation of an insignificant extra-solution data gathering (see MPEP § 2106.05(g)), which does not integrate a judicial exception into practical application. Further, the insignificant extra-solution data gathering is also WURC, see MPEP § 2106.05(d)(II) “The courts have recognized the following computer functions as well‐understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity. i. Receiving or transmitting data over a network”. Therefore, the limitation does not integrate the abstract idea in to a practical application, nor does significantly more.
Claim 16 is clarifying that the objective function or constraint function comprises a smooth differentiable surrogate and is therefore directed to the abstract idea of mathematical concepts (see MPEP 2106.04(a)(2)) as it is considered to be stating a mathematical formula in prose, additionally as the objective function or constraint function is used to calculate the gradient the claim is further clarifying the mathematical calculation of claim 15.
Claim 17 is clarifying the mathematical operation to determine the gradient and therefore is directed to the abstract idea of mathematical concepts (see MPEP 2106.04(a)(2)) as it is describing how to calculate the gradient, including calculating a derivative and a gradient of the derivative, and is therefore considered to be mathematical calculation.
Claim 18 is clarifying the mathematical operation to determine the gradient and therefore is directed to the abstract idea of mathematical concepts (see MPEP 2106.04(a)(2)) as it is clarifying the derivative of the implicit function used to calculate the gradient, is considered to be mathematical calculation.
Claim 19 is clarifying the mathematical operation to determine the gradient and therefore is directed to the abstract idea of mathematical concepts (see MPEP 2106.04(a)(2)) as it is clarifying the derivative of the objective function used to calculate the gradient, is considered to be mathematical calculation.
Claim 20 is clarifying the mathematical operation to determine the gradient and therefore is directed to the abstract idea of mathematical concepts (see MPEP 2106.04(a)(2)) as it is clarifying the derivative of the objective function, sum of the ratios of partial derivatives and constraint function used to calculate the gradient, is considered to be mathematical calculation.
Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A Prong 2. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B.
Claim Rejections – 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1, 2, 3, 4, 7, 11-13, 15-19 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Eban, et al. (2017) “Scalable Learning of Non-Decomposable Objectives,” arXiv:1608.04802v2 (hereinafter referred to as “Eban”) in view of Alesiani, Francesco (US 20220277859 A1) (hereinafter referred to as “Alesiani”).
Regarding claim 1, Eban recites “A computer-implemented method for optimizing machine-learned models by non-decomposable objectives with improved performance, the method comprising:” (Eban at pg. 2, cl. 2: Finally, and most importantly, our bounds give rise to an optimization approach for non-decomposable learning metrics that is highly scalable and that is applicable to truly large datasets)
“obtaining, by a computing system comprising one or more computing devices, data indicative of a plurality of machine-learned model parameters and at least one threshold comprising a machine-learned model;” (Eban at pg. 3, cl. 1; 3 Building Block Bounds: Definition 3.1. A classification rule fb is characterized by a score function f : X → R, and a threshold b ∈ R, indicating that classification is done according to f(x) ≥ b. Note that we intentionally separate the parameters of the models embedded in f (which could be a linear model or a deep neural-net), and the decision threshold b. The former provides a score which defines a ranking over examples, while the latter defines a decision boundary on the score that separates examples that are predicted to be relevant (positive) from those that are not.) [f represents model parameters and b represents threshold.]
“initializing, by the computing system, an initial plurality of machine-learned model parameters and an initial at least one threshold such that the initial plurality of machine-learned model parameters and the initial at least one threshold satisfy a constraint function;” (Eban at pg. 7, cl. 1-2; 7.3 JFT: To learn a model using our approach, we starting training from the pre-trained parameters (training from scratch is a multi-month process), and to optimize the AUCPR objective for several days. See also Eban at pg. 5, cl. 1; 5 Maximizing AUCPR:
PNG
media_image1.png
87
318
media_image1.png
Greyscale
.) [Training a model from pre-trained parameters is functionally equivalent to initializing an initial plurality of machine-learned model parameters. AUCPR includes at which serves as a discrete precision threshold, i.e., an initial at least one threshold and s.t. [subject to] the constraint that precision P(fbt) must be greater than or equal to threshold at, i.e., where the initial at least one threshold satisfy a constraint function.]
“updating, by the computing system, the plurality of machine-learned model parameters and the at least one threshold based at least in part on the gradient.” (Eban at pg. 5, cl. 2:
PNG
media_image2.png
131
324
media_image2.png
Greyscale
See also Eban at pg. 4, cl. 2:
PNG
media_image3.png
166
326
media_image3.png
Greyscale
.)1 [By solving this optimization problem via stochastic gradient descent, we are updating f, i.e. the model parameters, and bi, i.e., at least one threshold, based on in-part the SGD, i.e., the gradient.]
However, while Eban does recite calculating a gradient of the Lagrangian function L using stochastic gradient descent, Eban does not explicitly recite “determining, by the computing system, a gradient of an objective function with respect to the plurality of machine-learned model parameters at a current optimization step based at least in part on an implicit function of the at least one threshold as a function of the plurality of machine-learned model parameters; and”
On the other hand, Alesiani recites “determining, by the computing system, a gradient of an objective function with respect to the plurality of machine-learned model parameters at a current optimization step based at least in part on an implicit function of the at least one threshold as a function of the plurality of machine-learned model parameters; and” (Alesiani at 0203: Regarding continuous variables, to evaluate the gradient of the variables z versus the loss function L, the gradients of the two output variables x, y through the two optimization problems may be propagated. The implicit function theorem may be used to approximate locally the function z[Wingdings font/0xE0](x, y). Thus the following main results may be determined. See also Alesiani at 0204:
PNG
media_image4.png
177
478
media_image4.png
Greyscale
) [The relationship between the input parameters (z) and the bilevel solution variables (x, y) are solved for the gradient backpropagation of dzL that relies on implicit formula F and G. Therefore, determining the gradient dZL relies on parameters (z), y is a threshold because y represents the inner problem solution, which is an intermediate variable whose value influences the outer solution for x, i.e., is a threshold. The implicit function refers to the relationship between z and the optimal solution variables (x, y) of the bilevel problem optimally condition equations F(x, y, z) = 0 and G(x, y, z) = 0.]
Eban and Alesiani are analogous arts in machine learning frameworks to solve optimization problems. A person skilled in the art, before the effective filing date of the present application, would modify Eban with Alesiani to recite determining, by the computing system, a gradient of an objective function with respect to the plurality of machine-learned model parameters at a current optimization step based at least in part on an implicit function of the at least one threshold as a function of the plurality of machine-learned model parameters with the motivation being “(Alesiani at 0093) The gradient solver (e.g., the sub-processor 204) may use the implicit theorem of differentiability and the KKT optimality conditions to solve the continuous version and the random smoothness and linear function for the discrete case. For instance, the gradient solver computes or estimates the total gradient of the solution (e.g., the variables x and y) with respect to the output from the neural network (e.g., the variable z). This may be performed by defining a system of equations defined by the implicit function for the continuous case. Whereas for the discrete variable case, this may be performed by estimating the total gradient by propagating the gradient estimation from the inner level to the outer level. The estimation of the gradient may be based on the perturbation or change of variable principles. In other words, the gradient solver may use the implicit function gradient directly to iteratively improve the outer problem solution, which may be similar to gradient descent method.”
Regarding claim 2, Eban in view of Alesiani recites “The computer-implemented method of claim 1, wherein determining the gradient comprises” and Alesiani further recites “determining, by the computing system, a derivative of the implicit function with respect to the plurality of model parameters based at least in part on a derivative of the constraint function with respect to the plurality of model parameters; and” (Alesiani at 204:
PNG
media_image5.png
176
463
media_image5.png
Greyscale
) [∇xF and ∇xG represent derivatives of the implicit function. F(x,y,z)=0 and G(x,y,z)=0 are constraint functions as they constrain the variables x and y in terms of z additionally Eban recites a constraint function, see Eban at pg. 5, cl. 1; 5 Maximizing AUCPR:
PNG
media_image1.png
87
318
media_image1.png
Greyscale
. Calculating a gradient for an optimization problem, would involve calculating a derivative based, in part, of the constraint function.]
“determining, by the computing system, the gradient based at least in part on the derivative of the implicit function.” (Alesiani at 204:
PNG
media_image5.png
176
463
media_image5.png
Greyscale
) [dzL, the gradient, is determined based at least in part ∇xF and ∇xG, i.e., the derivative of the implicit function.] The motivation rationale applied in claim 1 is similarly applicable to claim 2.
Regarding claim 3, Eban in view of Alesiani recite “The computer-implemented method of claim 2,” and Alesiani further recites “wherein the derivative of the implicit function comprises the derivative of the constraint function divided by a partial derivative of the constraint function with respect to the at least one threshold.” (Alesiani at 204:
PNG
media_image5.png
176
463
media_image5.png
Greyscale
See also Alesiani at 226: ) [F(x,y,z)=0 and G(x,y,z)=0 are constraint functions as they constrain the variables x and y in terms of z.
PNG
media_image6.png
74
73
media_image6.png
Greyscale
is dividing the derivative of the constraint function with a partial derivative because multiply by the inverse of a matrix is analogous to division.] The motivation rationale applied in claim 1 is similarly applicable to claim 2.
Eban further recites “with respect to the at least one threshold.” (Eban at pg. 5, cl. 1: …where π is the positive class prior and R@Pa(f) denotes the recall we achieve when using f as a score function with b=b(α) is a threshold which achieves precision a. …)
Regarding claim 4, Eban in view of Alesiani recite “The computer-implemented method of claim 2,” and Alesiani further recites “wherein the gradient comprises a derivative of the objective function with respect to the plurality of model parameters and the multiplication of the derivative of the implicit function with a partial derivative of the objective function with respect to the at least one threshold.” (See Alesiani at 0236:
PNG
media_image7.png
301
578
media_image7.png
Greyscale
) [dxf represents the total derivative of the objection function (f), ∇xf is the derivative of the objective function ∇yf is a partial derivative of the objective function with respect to the at least one threshold and ∇yy is the derivative of the implicit function, y is a threshold because it was determined by x and satisfies a condition specifically the optimality condition.] The motivation rationale applied to claim 1 is applicable to claim 4.
Regarding claim 7, Eban in view of Alesiani recite “The computer-implemented method of claim 1,” and Eban further recites “wherein updating the plurality of machine-learned model parameters comprises adjusting values of the plurality of machine-learned model parameters at a previous optimization step based at least in part on the gradient.” (Eban at pg. 4, cl. 2:
PNG
media_image8.png
108
343
media_image8.png
Greyscale
) [f and λ represent model parameters and ∇L represents the gradient of the objective function and (t+1) represents the previous optimization step, i.e., the values of machine-learned model parameters are adjusted at a previous optimization step based at least in part on the gradient.]
Regarding claim 11, Eban in view of Alesiani recite “The computer-implemented method of claim 1,” and Eban further recites “wherein at least one of the objective function or the constraint function comprises a smooth differentiable surrogate of the at least one of the objective function or the constraint function.” (Eban at pg. 3, cl. 2: Now it is natural to bound these quantities by using a surrogate for the zero-one loss function such as the hinge loss:… for simplicity but other losses could be used in all the results presented below (with the exception of the linear-fractional transformation of the Fβ score in section 6). In the case of convex surrogates for the zero-one loss such as the log-loss or the smooth-hingeloss [21] we get convex (and smooth) optimization problems.)2
Regarding claim 12, Eban in view of Alesiani recite “The computer-implemented method of claim 11,” and Eban further recites “wherein the smooth differentiable surrogate comprises at least one of a sigmoid function or a softplus function.” (Eban at pg. 3, cl. 2: We note that in what follows we use the hinge-loss for simplicity but other losses could be used in all the results presented below (with the exception of the linear-fractional transformation of the Fβ score in section 6). In the case of convex surrogates for the zero-one loss such as the log-loss or the smooth-hingeloss [21] we get convex (and smooth) optimization problems.) [Log-loss utilizes a sigmoid function.]
Regarding claim 13, Eban in view of Alesiani recite “The computer-implemented method of claim 1,” and Eban further recites “wherein the objective function and the constraint function are selected for at least one of the precision at fixed recall metric, the FNR at fixed FPR metric, the precision at k metric, the AUC-PR metric, the AUC-ROC metric, or the fairness criterion metric.” (Eban at pg. 3, cl. 2: In this section we show how the building block bounds of (2) can provide a concave lower bound on the objective of maximum recall with at least α precision. A similar derivation could also be used to provide a bound on maximum precision given a minimum desired recall. Aside from the stand-alone usefulness of the P@R and R@P metrics, the developments here will underlie the construction for optimizing the maximum AUCPR objective that we present in the next section. See also Eban at pg. 5, cl. 1: We are now ready to use our derivation of the R@P optimization objective in the previous section order to construct a concave lower-bound surrogate for AUCPR. A similar derivation could be used for AUCROC optimization.) [The objective function and the constraint function are selected for at least precision at fixed recall or AUC-PR or AUC-ROC.]
Regarding claim 15, claim 15 is the system embodiment of claim 1 with substantially similar limitations as claim 1. Therefore, claim 15 is rejected for the same rationale as claim 1. Additionally, claim 15 recites “A computing system for optimizing machine-learned models by non-decomposable objectives with improved performance, the computing system comprising:” (Eban at pg. 1, cl. 2: Machine learning models underlie most modern automated retrieval systems. The quality of such systems is evaluated using ranking-based measures such as area under the ROC curve (AUCROC) or, as is more appropriate in the common scenario of few relevant items, measures such as area under the precision recall curve (AUCPR, also known as average precision), mean average precision (MAP), precision at a fixed recall rate (P@R), etc. In fraud detection, for example, we would like to constrain the fraction of customers that are falsely identified as fraudsters, while maximizing the recall of true ones. See also Eban at pg. 6, cl. 2, Eban at pg. 7, cl. 1) “one or more processors; and” (Eban at pg. 6, cl. 2: All the models were trained on a single tesla k40 GPU for about eight hours. See also Eban at pg. 7, cl. 1: The ImageNet experiments were trained with 50 tesla K40 GPU replicas for three days performing about 5M mini-batch updates.) “one or more computer-readable memory devices storing instructions that, when implemented, cause the one or more processors to perform operations, the operations comprising:” (Eban at pg. 6, cl. 2: All the models were trained on a single tesla k40 GPU for about eight hours. See also Eban at pg. 7, cl. 1: The ImageNet experiments were trained with 50 tesla K40 GPU replicas for three days performing about 5M mini-batch updates.)
Claim 16 is the system embodiment of claim 11 with substantially similar limitations. Therefore, claim 16 is rejected by the same rationale as claim 11.
Claim 17 is the system embodiment of claim 2 with substantially similar limitations. Therefore, claim 17 is rejected by the same rationale as claim 2.
Claim 18 is the system embodiment of claim 3 with substantially similar limitations. Therefore, claim 18 is rejected by the same rationale as claim 3.
Claim 19 is the system embodiment of claim 4 with substantially similar limitations. Therefore, claim 19 is rejected by the same rationale as claim 4.
Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Eban and Alesiani in further view Ding, et al. (US 20200134468 A1) (hereinafter referred to as “Ding”).
Regarding claim 6, Eban in view of Alesiani recite “The computer-implemented method of claim 1,” however neither Eban nor Alesiani recite “wherein initializing the initial plurality of machine-learned model parameters and the initial at least one threshold comprises sampling, by the computing system, the initial plurality of machine-learned model parameters and the initial at least one threshold from a parameter distribution comprising the machine- learned model parameters.”
On the other hand, Ding recites “wherein initializing the initial plurality of machine-learned model parameters and the initial at least one threshold comprises sampling, by the computing system, the initial plurality of machine-learned model parameters and the initial at least one threshold from a parameter distribution comprising the machine- learned model parameters.” (Ding at Fig. 4:
PNG
media_image9.png
510
546
media_image9.png
Greyscale
See also Ding at 0055: Epsilon dictionary (epsdct): epsdct saves the suitable perturbation length of a given training example that was used to perturb it the last time when it is encountered. When an example is met for the first time, this value is initialized as mineps (a hyperparameter that is minimum perturbation length).) [Randomly initialize network N, i.e., initializing the initial plurality of machine-learned model parameters, perturbation lengths is a threshold value because it is the minimum perturbation that triggers a classification change, every element of e of Emin and Emax (and also Einc) are examples of a degenerate distribution, i.e., the initial plurality of machine-learned model parameters and the initial at least one threshold from a parameter distribution comprising the machine- learned model parameters.]
Eban, Alesiani and Ding are analogous arts in machine learning frameworks to solve optimization problems. A person skilled in the art, before the effective filing date of the present application, would modify Eban and Alesiani with Ding to recite wherein initializing the initial plurality of machine-learned model parameters and the initial at least one threshold comprises sampling, by the computing system, the initial plurality of machine-learned model parameters and the initial at least one threshold from a parameter distribution comprising the machine- learned model parameters with the motivation being “(0009) Because of its “direct” margin maximization nature, MMA training is an improvement over alternate approaches of adversarial training which have the inherent problem that a perturbation length ∈ has to be set and fixed throughout the training process, where ∈ is often set arbitrarily” and “(0011) MMA training resolves problems associated with said fixed perturbation magnitude ∈ in the sense that: 1) the approach dynamically determines (e.g., maximizes) the margin, the “current robustness” of the data, instead of robustness with regards to a predefined magnitude; 2) the margin is determined for each data point, therefore each sample's robustness could be maximized individually.”
Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Eban and Alesiani in further view of Kotriwala, et al. (US 20230214724 A1) (hereinafter referred to as “Kotriwala”).
Regarding claim 8, Eban in view of Alesiani recite “The computer-implemented method of claim 1,” while Eban recites updating the threshold (see pg. 5, cl. 2:
PNG
media_image10.png
104
332
media_image10.png
Greyscale
) neither Eban nor Alesiani explicitly recite “comprises adjusting values of the at least one threshold at a previous optimization step based at least in part on an inner product of the derivative of the implicit function and a parameter delta of the plurality of machine-learned model parameters.”
On the other hand, Kotriwala recites “comprises adjusting values of the at least one threshold at a previous optimization step based at least in part on an inner product of the derivative of the implicit function and a parameter delta of the plurality of machine-learned model parameters.” (Kotriwala at 0056: Besides the increased flexibility that remembrance factors deliver, they can also be leveraged by the MuL unit 122 to reduce the computational ML training effort by using mathematical sensitivity analysis of the optimal ML weights (where “weight” refers here to the degrees of freedom in the ML model 102, not the feature weights of an explanation 110) with respect to the continuous remembrance factors, based for example on the Implicit Function Theorem applied to the optimality conditions of the ML training loss minimization problem. The MuL unit 122 may thus change the optimal weights in correspondence to changes in the remembrance factors, in order to avoid complete retraining in favor of incremental training. The application of the Implicit Function Theorem yields that the inverse of the so-called Hessian matrix of the training objective function yields the correct scaling of increments in the remembrance factors to the optimal ML weights increments. Hence, approximately optimal ML weight increments can be computed by iterative linear algebra methods, such as the conjugate gradients method, which only require Hessian matrix-vector products that modern ML frameworks can compute efficiently for mini-batch approximation of the Hessian.) [The change of the remembrance factors, which are data sample weightings, is functionally equivalent to a parameter delta as both reference a chance in the model parameters, the calculation of Hessian matrix-vector products, involve the calculation of inner products as a Hessian matrix-vector product refers to each element of a vector from the product is calculated as the inner product or dot product of a row with a vector.]
Eban, Alesiani and Kotriwala are analogous arts in machine learning frameworks to solve optimization problems. A person skilled in the art, before the effective filing date of the present application, would modify Eban and Alesiani with Kotriwala to recite comprises adjusting values of the at least one threshold at a previous optimization step based at least in part on an inner product of the derivative of the implicit function and a parameter delta of the plurality of machine-learned model parameters with the motivation being “(0056) Hence, approximately optimal ML weight increments can be computed by iterative linear algebra methods, such as the conjugate gradients method, which only require Hessian matrix-vector products that modern ML frameworks can compute efficiently for mini-batch approximation of the Hessian.”
Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Eban and Alesiani in further view of Narasimhan, et al. (26 May 2015) “Optimizing Non-Decomposable Performance Measures: A Tale of Two Classes,” arXiv:1505.06812 (hereinafter referred to as “Narasimhan”).
Regarding claim 9, Eban in view of Alesiani recite “The computer-implemented method of claim 1,” however neither Eban nor Alesiani explicitly recite “wherein updating the at least one threshold comprises, at regular optimization steps, setting the at least one threshold such that the at least one threshold satisfies the constraint function for the plurality of machine-learned model parameters at the regular optimization steps.”
On the other hand, Narasimhan recites “wherein updating the at least one threshold comprises, at regular optimization steps, setting the at least one threshold such that the at least one threshold satisfies the constraint function for the plurality of machine-learned model parameters at the regular optimization steps.” (Narasimhan at pg. 6, cl. 2, Algorithm 3:
PNG
media_image11.png
428
416
media_image11.png
Greyscale
) [At each epoch, i.e., optimization step,
PNG
media_image12.png
34
88
media_image12.png
Greyscale
, i.e., ve is functionally equivalent to a threshold because v represents a challenge level or a target value for the performance measure (i.e., a threshold) and
PNG
media_image12.png
34
88
media_image12.png
Greyscale
is the constraint function because it represents an estimate of the F1-measure.]
Eban, Alesiani and Narasimhan are analogous arts in machine learning frameworks to solve optimization problems. A person skilled in the art, before the effective filing date of the present application, would modify Eban and Alesiani with Narasimhan to recite wherein updating the at least one threshold comprises, at regular optimization steps, setting the at least one threshold such that the at least one threshold satisfies the constraint function for the plurality of machine-learned model parameters at the regular optimization steps with the motivation being “(pg. 6, cl. 2) …optimize these performance measures in an online stochastic manner. To this end, we observe that the AMP algorithm can be executed in an online fashion by using stochastic updates to train the intermediate models. The resulting algorithm STAMP, is presented in Algorithm 3” and “(pg. 8, cl. 2) Figures 3 and 4 report the performance of the STAMP method applied to pseudo-linear functions. Similar to the concave measures, STAMP was found to provide competitive accuracies as compared to the baseline methods but require at least 3 − 4× less computational time.”
Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Eban and Alesiani in further view of Bai, et al. (US 20220398480 A1).
Regarding claim 10, Eban in view of Alesiani recite “The computer-implemented method of claim 1,” and however neither Eban nor Alesiani explicitly recite “wherein the gradient comprises a regularization cost that penalizes, with respect to the plurality of model parameters, the derivative of the constraint function with respect to the at least one threshold.”
On the other hand, Bai recites “wherein the gradient comprises a regularization cost that penalizes, with respect to the plurality of model parameters, the derivative of the constraint function with respect to the at least one threshold.” (Bai at 0003: In one or more illustrative examples, a method for regularized training of a Deep Equilibrium Model (DEQ) is provided. A regularization term is computed using a predefined quantity of random samples and the Jacobian matrix of the DEQ, the regularization term penalizing the spectral radius of the Jacobian matrix. The regularization term is included in an original loss function of the DEQ to form a regularized loss function. A gradient of the regularized loss function is computed with respect to model parameters of the DEQ. The gradient is used to update the model parameters. See also Bai at 0046:
PNG
media_image13.png
89
531
media_image13.png
Greyscale
)
[
PNG
media_image14.png
86
141
media_image14.png
Greyscale
is the regularization term, z* is the threshold and g is the constraint function since we are trying to find a z* where g is satisfied. The plurality of model parameters is θ in the function fθ.]
Eban, Alesiani and Bai are analogous arts in machine learning frameworks to solve optimization problems. A person skilled in the art, before the effective filing date of the present application, would modify Eban and Alesiani with Bai to recite wherein the gradient comprises a regularization cost that penalizes, with respect to the plurality of model parameters, the derivative of the constraint function with respect to the at least one threshold with the motivation being “(0026) There are two immediate benefits of the resulting more stable dynamics. First, solving a DEQ requires far fewer iterations than before, which makes regularized DEQs significantly faster than their unregularized counterparts. Second, this class of model becomes much less brittle to architectural variants that would otherwise break the DEQ.” And “(0047) …without regularization, a DEQ model that stops after a fixed number T of solver iterations exhibits increasingly poor convergence, accompanied by a growing ∥Jƒθ∥F at these fixed points that empirically signals the growing instability. Therefore, by constraining the Jacobian's Frobenius norm, DEQs can be optimized for stabler and simpler dynamics whose fixed points are easier to solve for.”
Claim 14 is rejected under 35 U.S.C. 103 as being unpatentable over Eban and Alesiani in further view of Shpurov, et al. (US 20200244435 A1) (hereinafter referred to as “Shpurov”).
Regarding claim 14, Eban in view of Alesiani recite “The computer-implemented method of claim 1,” however neither Eban nor Alesiani recite “wherein: obtaining the data indicative of the plurality of machine-learned model parameters and the at least one threshold comprising the machine-learned model comprises receiving, by the computing system, the data from a second computing system; and wherein the computer-implemented method further comprises, subsequent to updating the plurality of machine-learned model parameters and the at least one threshold based at least in part on the gradient, providing, by the computing system, data indicative of the plurality of machine-learned model parameters and the at least one threshold to the second computing system.”
On the other hand, Shpurov recites “wherein: obtaining the data indicative of the plurality of machine-learned model parameters and the at least one threshold comprising the machine-learned model comprises receiving, by the computing system, the data from a second computing system; and wherein the computer-implemented method further comprises, subsequent to updating the plurality of machine-learned model parameters and the at least one threshold based at least in part on the gradient, providing, by the computing system, data indicative of the plurality of machine-learned model parameters and the at least one threshold to the second computing system.” (Shpurov at 0035: Training engine 112 may also … package the determined model coefficients or parameters and the determined threshold into corresponding portions of modelling data 118. See also Shpurov at 0070: Referring to FIG. 3A, third computing system 302 may receive discrete elements of modelling data from first computing system 102… As described herein, modelling data 122 may include scaled (or unscaled) modeled coefficients or parameters and a threshold value specifying a first predictive fraud model privately trained by first computing system 102, e.g., based on locally maintained elements of confidential transaction data using any of the exemplary processes described herein. See also Shpurov at 0037.) [The model parameters and threshold are determined after training and is referred to as “modelling data,” i.e., subsequent to updating the model. The modelling data is then provided to the computing system that trained, i.e., the first computing to system, to another computing system such as the third computing system, i.e., the second computing system.]
Eban, Alesiani and Shpurov are analogous arts in efficient methods of machine learning frameworks. A person skilled in the art, before the effective filing date of the present application, would modify Eban and Alesiani with Shpurov to recite wherein: obtaining the data indicative of the plurality of machine-learned model parameters and the at least one threshold comprising the machine-learned model comprises receiving, by the computing system, the data from a second computing system; and wherein the computer-implemented method further comprises, subsequent to updating the plurality of machine-learned model parameters and the at least one threshold based at least in part on the gradient, providing, by the computing system, data indicative of the plurality of machine-learned model parameters and the at least one threshold to the second computing system with the motivation being “(0051) Further, in some instances, first computing system 102 may perform the homomorphic computations that apply the corresponding predictive fraud model to encrypted transaction data 220 may be implemented in conjunction within one or more additional computing systems operations within environment 100, e.g., in parallel across a distributed computing system, such as, but not limited to a cloud-based network. See also 0002.”
Allowable Subject Matter
Claims 5 and 20 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all the limitations of the base claim and an intervening claims provided the 35 USC 101.
Regarding claim 5, Eban in view of Alesiani recite “The computer-implemented method of claim 2,” and Eban further recites “wherein the at least one threshold comprises a plurality of thresholds,” (Eban at pg. 5, cl. 1: …where π is the positive class prior and R@Pa(f) denotes the recall we achieve when using f as a score function with b=b(α) is a threshold which achieves precision a. … To apply our bounds to the objective of maximizing AUCPR(f), we first approximate the integral in Equation 10 by a discrete sum over a set of precision anchor values A = { π = α0 < α1 < α2 < … < αk}…See also Eban at pg. 5, cl. 2:
PNG
media_image15.png
65
90
media_image15.png
Greyscale
) [The threshold b is calculated for each anchor point/saddle point at each α, i.e., a plurality of thresholds.]
However, neither Eban nor Alesiani recite “and wherein the gradient comprises a derivative of the objective function with respect to the plurality of model parameters and the sum of the ratios of the partial derivatives of the objective function and the constraint function with respect to each threshold of the plurality of thresholds multiplied by the constraint function of the threshold.” The examiner has found the distinct features of applicant’s claimed invention over the prior art is the explicit claiming of the aforementioned limitations specified in claim 5. When viewed individually or in as a combination with prior art of record, the limitations specified in claim 5 are distinct. Claim 20 is the system embodiment of claim 5 with substantially similar limitations. Therefore, claim 20 is objected to by the same rationale as claim 5.
Examiner’s Note
Examiner cites particular columns, paragraphs, figures and line numbers in the references as applied to the claims below for the convenience of the applicant. Although the specified citations are representative of the teachings in the art and are applied to the specific limitations within the individual claim, other passages and figures may apply as well. It is respectfully requested that, in preparing responses, the applicant fully consider the references in their entirety as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior art or disclosed by the examiner. The entire reference is considered to provide disclosure relating to the claimed invention.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Eban et al. (US 20190266513): teaches the systems and methods of the present disclosure can transform a constrained optimization problem into an unconstrained optimization problem that is solved more efficiently and generally than the constrained optimization problem. And optimizing the unconstrained objective function in which the decision threshold of the machine-learned classification model is expressed as the estimator of the quantile function can include, for each of a plurality of iterations: determining a gradient of the unconstrained objective function in which the decision threshold of the machine-learned classification model is expressed as the estimator of the quantile function.
Tavker et al. (NPL: Consistent Plug-in Classifiers for Complex Objectives and Constraints): teaches in Secs 2 & 3.3-3.4 : considering performance metrics ¯ and constraint functions φk’s that are general functions of the confusion matrix of classifier h. This includes several common examples, including those that are non-decomposable and cannot be expressed as a simple expectation of errors on individual examples… we adopt the Frank-Wolfe based approach of Gidel et al. (2018) [15] that enables optimization of a convex objective over the intersection of two convex sets with access to only linear minimization oracles for the individual sets… We then define the augmented Lagrangian…
Narasimhan (NPL: Learning with complex loss functions and constraints): teaches in Sec. 3.1: We start with the case where ψ is convex over CD. Introducing Lagrange multipliers λ = [λ1,...,λK] ∈ RK+ for the constraints, we formulate the Lagrangian … and the optimization problem… We can now apply a gradient ascent procedure to maximize F over λ… where w 2 Rd denotes the Lagrange multipliers for the equality constraints and λ>0 is a constant. Gidel at al. (2018) [15] propose a simple gradient ascent step for w, a linear minimization step for u over C and a linear minimization step for v over F… The classifier g⇤ defined above is a deterministic classifier that thresholds the conditional probability ⌘ based on the example-dependent loss matrix L(x)…
Ramaswamy et al. (NPL: Consistent Classification Algorithms for Multi-class Non-Decomposable Performance Metrics): teaches in Sec. 3: a process where the authors provide a generic framework for studying a multi-class non-decomposable performance metric, where they view the problem of finding the optimal classifier for a non-decomposable metric as an optimization problem over the space of all confusion matrices that are attainable under the given distribution.
Blondel, et al. (31 May 2021) “Efficient and Modular Implicit Differentiation,” arXiv:2105.15183v1 – recites calculating gradient in optimization problems by relying on implicit function theorem. By using the implicit function theorem, gradient expressions are derived that include partial derivatives for the optimization problem, which may be viewed as an objective function or constraint function.
Katanforoosh & Kunin, "Initializing neural networks", deeplearning.ai, 2019 – broadly discusses initializing machine learning models including initializing models from distributions such as normal distributions, zero distribution or uniform distribution.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to OLUWATOSIN ALABI whose telephone number is (571)272-0516. The examiner can normally be reached Monday-Friday, 8:00am-5:00pm EST..
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michael Huntley can be reached at (303) 297-4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/OLUWATOSIN ALABI/ Primary Examiner, Art Unit 2129
1 Notably, “As before” is referring to Eban at pg. 4, cl. 2, which is describing how to optimize the Recall@Precision (R@P) with stochastic gradient descent (SGD) updates, based on the gradient of the Lagrangian function L.
2 For the purposes of compact prosecution, the secondary reference, Alesiani also recites “wherein at least one of the objective function or the constraint function comprises a smooth differentiable surrogate of the at least one of the objective function or the constraint function.” (Alesiani at 0093: The gradient solver (e.g., the sub-processor 204) may use the implicit theorem of differentiability and the KKT optimality conditions to solve the continuous version and the random smoothness and linear function for the discrete case. For instance, the gradient solver computes or estimates the total gradient of the solution (e.g., the variables x and y) with respect to the output from the neural network (e.g., the variable z). This may be performed by defining a system of equations defined by the implicit function for the continuous case. Whereas for the discrete variable case, this may be performed by estimating the total gradient by propagating the gradient estimation from the inner level to the outer level. The estimation of the gradient may be based on the perturbation or change of variable principles. In other words, the gradient solver may use the implicit function gradient directly to iteratively improve the outer problem solution, which may be similar to gradient descent method. Otherwise, the gradient solver may use the stochastic gradient descent ascent method for minimum maximum (min-max) problems. See also 0212, 0105.)