DETAILED ACTION
This action is responsive to the application filed on 04/16/2026. Claims 1-2, 4-5, 9-20, and 22-25 are pending and have been examined. This action is Final.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Applicant’s claim for the benefit of a prior-filed application under 35 U.S.C. 119(e) or under 35 U.S.C.
120, 121, 365(c), or 386(c) is acknowledged.
Response to Arguments
Argument 1: The applicant argues (see pages 10-12) that the claims as amended are not directed to an abstract idea because independent claims 1, 11, and 16 now recite determining uncertainty values based on first inputs corresponding to incoming calls to a call center, routing an incoming call to a first device associated with an operator of the call center when the machine learning system classifies the corresponding input into the first portion of the outputs, and routing the incoming call to a second device instead of any device associated with any operator of the call center when the input is classified into the second portion of the outputs. The applicant contends that these features are directed to a practical application relating to control of particular devices, namely routing calls to different devices based on classification results from a machine learning system, rather than to a mental process, and that they result in improvement to overall call center performance as described at paragraph 56 of the specification, which states that a call center can “save time, cost, difficulty, for example where the second portion is routed to a human expert, the second portion can have a higher proportion of ML outputs that actually needed the human expert’s attention.” The applicant further asserts that claim 21 has been canceled without prejudice or disclaimer, obviating the rejection as to that claim, and that claims 2, 4-5, 9-10, 12-15, 17-20, and 22-25 recite patentable subject matter for at least the reason that they depend from patentable base claims.
Examiner Response to Argument 1: The applicant’s argument is persuasive. In view of the amendments to independent claims 1, 11, and 16, and the cancellation of claim 21, the rejection of claims 1-5 and 9-24 under 35 U.S.C. 101 is hereby withdrawn.
Argument 2: The applicant argues (see pages 12-19) that the cited references, alone or in combination, fail to teach or suggest iteratively performing an update operation on the adjustable penalty value based on a hyperparameter wherein the update operation is selected from a group consisting of incrementing the adjustable penalty value by a value of the hyperparameter and decrementing the adjustable penalty value by the value of the hyperparameter, as now recited by each independent claim. Specifically, Kung is said to disclose only early exiting once a threshold confidence level is reached and to be silent as to any such penalty update. Chen is said to update loss weights to move gradient norms toward a target for each task without performing the claimed hyperparameter-based increment or decrement. Kang is said to relate only to partitioning deep neural network computation between devices and not to cure that deficiency; Geifman is said to use λ merely as a constraint weighting parameter that was set to a fixed value of 32 and α to 0.5 for all experiments rather than iteratively updated. Wegkamp is said to use a rejection cost d that “should be known a priori” and therefore to teach away from adjusting that parameter. Ziyin is said to disclose only a selection function compared against a predetermined threshold, and Cortes is said to permit abstention only “at the price of a fixed cost” and likewise to teach away from the claimed iterative update. The applicant further asserts that Lin, Han, and Carpenter do not supply the features missing from the base combinations, that the rejection of claim 21 is obviated by its cancellation, and that the dependent claims are patentable for at least the reason that they depend from patentable base claims.
Examiner Response to Argument 2: The examiner has considered the argument set forth above, however it’s not persuasive. The rejection of the independent claims no longer relies on Kung or Chen for the recited penalty update. New reference Thulasidasan is relied upon for a loss function combining a conventional cross-entropy term with an adjustable abstention penalty and for adapting that penalty during training, teaching that the network is given “an option to abstain on a confusing training sample thereby mitigating the misclassification loss but incurring an abstention penalty,” that “the approach described in this paper employs abstention during training as well as inference,” and that the penalty is derived from a running batch average of the conventional loss, where “β̃ is a smoothed moving average of the α threshold (initialized to 0), and updated at every mini-batch iteration.” Barnes is relied upon for the claimed convergence of an uncertainty-derived indicator to a target value, teaching that “The abstention loss is designed to abstain on a user-defined fraction of the samples via a PID controller” and that the penalty “can be adaptively modified throughout training so that the network abstains on a specified percent of the training samples,” which is the opposite of a penalty held constant. Veeraraghavan is relied upon for the claimed direction of the update, teaching that “if an insult rate is higher than an insult rate equilibrium, S240 may function to increase a fraud or threat score threshold (e.g., from 90 ⇒ 95)” and, where the measured rate falls below that equilibrium, that the threshold is correspondingly updated so that more activity is evaluated adversely. The assertions that Geifman, Wegkamp, and Cortes teach away are likewise not persuasive, because those references are relied upon only for initializing the penalty to a dominating value, for the floor beneath which the penalty may not fall, and for the per-sample abstention cost, respectively, and a reference that employs a fixed value for one purpose does not criticize, discredit, or otherwise discourage the separate iterative adjustment taught by Barnes and Veeraraghavan (see MPEP 2145(X)(D)). Carpenter is relied upon for the newly recited call center limitations, teaching that the router will “compare the confidence values to a pre-determined threshold” and route the call to that destination where exactly one candidate results, and otherwise “punt them to a human operator,” and Riahi is relied upon for the phone tree of claim 25. Accordingly, the cited combinations teach or render obvious the claimed subject matter as recited, and the dependent claims are not separately patentable for the reasons set forth in the rejections below.
Claim Rejections - 35 U.S.C. 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1, 11, and 16 are rejected under 35 U.S.C. 103 as being unpatentable over the NPL reference “Combating Label Noise in Deep Learning Using Abstention” by Thulasidasan et al. (referred herein as Thulasidasan) in view of the NPL reference “Vector-Based Natural Language Call Routing” by Carpenter et al. (referred herein as Carpenter) in view of the NPL reference “Controlled Abstention Neural Networks for Identifying Skillful Predictions for Classification Problems” by Barnes et al. (referred herein as Barnes) in view of US 11,068,910 B1 by Veeraraghavan et al. (referred herein as Veeraraghavan) in view of the NPL reference “BranchyNet: Fast Inference via Early Exiting from Deep Neural Networks” by Kung et al. (referred herein as Kung) further in view of the NPL reference “Neurosurgeon: Collaborative Intelligence Between the Cloud and Mobile Edge” by Kang et al. (referred herein as Kang).
Regarding claim 1, Thulasidasan teaches:
A device, comprising: at least one processor; and at least one memory that stores executable instructions that, when executed by the at least one processor, facilitate performance of operations, comprising: ([Thulasidasan, page 5, sec. 3.1] we use a deep convolutional network employing the VGG-16 (Simonyan &Zisserman, 2014) architecture, implemented in the Py Torch (Paszke et al., 2017) framework…We perform abstention-free training…We train the network for 200 epochs using SGD accelerated with Nesterov momentum and employ a weight decay of .0005, initial learning rate of 0.1 and learning rate annealing using an annealing factor of 0.5 at epoch 60, 120 and 160.” AND [Thulasidasan, page 2, sec 1] “Section 5 has further discussions on abstention behavior in the context of memorization.” AND [Thulasidasan, page 7, sec 4] “To identify the samples for elimination, we train the DAC, observing the performance of the non-abstaining part of the DAC on a validation set (which we assume to be clean).”, wherein the examiner interprets the network implemented in a software framework and executed over 200 training epochs to be the same as executable instructions stored in at least one memory and executed by at least one processor because they are both program code held in a machine and carried out by that machine’s computing hardware.)
determining, based on first inputs, provided to a machine learning system…uncertainty values corresponding to outputs of the machine learning system; ([Thulasidasan, page 2, sec 1-2] “abstention training allows the DNN to learn features that are indicative of unreliable training signals which are thus likely to lead to uncertain pre dictions…This gives the DNN an option to abstain on a confusing training sample thereby mitigating the misclassification loss but incurring an abstention penalty.”, wherein the examiner interprets the abstention output the network produces for a given “unreliable” training sample to be the same as an uncertainty value corresponding to an output of the machine learning system because they are both a value generated by the model from an input that expresses the degree to which the model’s own prediction for that input is unreliable.)
updating a machine learning model employed by the machine learning system based on the uncertainty values and a loss function, the loss function being a function of at least a conventional loss function and an adjustable penalty value, ([Thulasidasan, page 3, sec 2-2.1] “ℒ(x_j)=(1-p_{k+1})(-∑_{i=1}^{k}t_i log(p_i/(1-p_{k+1})))+α log(1/(1-p_{k+1}))…We perform abstention-free training for L initial epochs (a warm-up period) to accelerate learning, triggering abstention from epoch L +1 onwards. At the start of abstention, α is initialized to a much smaller value than the threshold ˜β to encourage abstention on all but the easiest of examples learnt so far. As the learning progresses on the true classes, abstention is reduced. We linearly ramp up α over the remaining epochs (updating once per epoch) to a final value of α final In the experiments in the subsequent sections, we illustrate how the DAC, when trained with this loss function, learns representations for abstention remarkably well.” AND [Thulasidasan, page 2, sec. 2] “If α is very large, there is a high penalty for abstention driving p_{k+1} to zero and recovering the standard unmodified cross-entropy loss; in such case, the model learns to never abstain. With α very small, the classifier may abstain on everything.”, wherein the examiner interprets the disclosed training objective, which sums a cross-entropy term over the non-abstaining classes with a term weighting α by the abstention output and “We linearly ramp up α over the remaining epochs (updating once per epoch) to a final value of αfinal In the experiments in the subsequent sections, we illustrate how the DAC, when trained with this loss function, learns representations for abstention remarkably well.”, to be the same as a loss function being a function of at least a conventional loss function and an adjustable penalty value because they are both a single training objective built from an ordinary classification loss combined with a separately settable cost assessed against the model for declining to predict.)
wherein updating the machine learning model comprises updating the adjustable penalty value based on a mean value of the conventional loss function as applied to a batch of second inputs processed before the first inputs; ([Thulasidasan, page 2, sec. 2] “Section 2 describes the loss function formulation and an algorithm for automatically tuning abstention behavior.” AND [Thulasidasan, page 8, sec. 6] “Furthermore, the loss function formulation is simple to implement and can work with any existing DNN architecture; this makes the DAC a useful addition to real-world deep learning pipelines.” AND [Thulasidasan, page 3, sec. 2.1] “β is a smoothed moving average of the α threshold (initialized to 0), and updated at every mini-batch iteration.”, wherein the examiner interprets the loss function formulation algorithm tuning (updating) abstention behavior to be the same as updating the adjustable penalty value based on a mean value of the conventional loss function as applied to a batch of second inputs processed before the first inputs because they are both directed to loss functions and updating/tuning inputted behavior based on a mean value (smoothed moving average).)
segregating, during training of the machine learning system, the outputs of the machine learning system into a first portion of the outputs and a second portion of the outputs based on the uncertainty values, ([Thulasidasan, page 2, sec. 1-2] “In contrast to the above, the approach described in this paper employs abstention during training as well as inference… Absence of the abstaining output (i.e., pk+1 = 0) recovers exactly the usual cross-entropy; otherwise, the abstention mass has been normalized out of the k class probabilities.” AND [Thulasidasan, page 7, sec. 4] “we train the DAC, observing the performance of the non-abstaining part of the DAC on a validation set”, wherein the examiner interprets the division of the network’s outputs into an abstaining set and a non-abstaining set while training is ongoing to be the same as segregating, during training, the outputs into a first portion and a second portion based on the uncertainty values because they are both the sorting of the model’s outputs into two groups during training according to the model’s own measure of how unreliable each output is.)
Thulasidasan does not teach corresponding to incoming calls to a call center…wherein the updating of the adjustable penalty value comprises iteratively performing an update operation on the adjustable penalty value based on a hyperparameter to cause an indicator value to converge to a target value across iterations of evaluating the loss function, the indicator value being based on a mean of the uncertainty values comprising a respective uncertainty value in each iteration of the evaluating the loss function…routing, in response to the machine learning system classifying an input of the first inputs into one of the first portion of the outputs, an incoming call associated with the input to a first device associated with an operator of the call center; and routing, in response to the machine learning system classifying the input into one of the second portion of the outputs, the incoming call to a second device instead of any device associated with any operator of the call center…wherein the update operation is selected from a group consisting of incrementing the adjustable penalty value by a value of the hyperparameter and decrementing the adjustable penalty value by the value of the hyperparameter…the segregating of the outputs comprising truncating processing, using the machine learning system, of respective ones of the first inputs corresponding to the second portion of the outputs … in response to determining that the respective ones of the first inputs correspond to respective ones of the uncertainty values that are greater than a sufficiency threshold…and redirecting the processing to an entity not associated with the machine learning system.
Carpenter teaches:
corresponding to incoming calls to a call center ([Carpenter, page 1, Abstract] “Evaluation of the call router performance over a financial services call center using both accurate transcriptions of calls and fairly noisy speech recognizer output demonstrated robustness in the face of speech recognition errors.”, wherein the examiner interprets the caller requests that the disclosed router receives and classifies at a financial services call center to be the same as first inputs corresponding to incoming calls to a call center because they are both data drawn from telephone calls arriving at a call center and supplied to a classifier for a routing decision.)
routing, in response to the machine learning system classifying an input of the first inputs into one of the first portion of the outputs, an incoming call associated with the input to a first device associated with an operator of the call center ([Carpenter, page 13, sec. 4.1.3] “Once we have obtained a confidence value for each destination, the final step in the routing process is to compare the confidence values to a pre-determined threshold and return those destinations whose confidence values are greater than the threshold as candidate destinations.” AND [Carpenter, page 4, sec. 4] “Given a caller request, the routing module selects a set of candidate destinations to which it believes the call can reasonably be routed. If there is exactly one such destination, the call is routed to that destination and the caller notified;”, wherein the examiner interprets sending the call to the single destination whose confidence value exceeds the threshold to be the same as routing the incoming call to a first device associated with an operator of the call center because they are both delivery of the call to the station of the call center representative that the classifier has identified as the correct handler for that call.)
routing, in response to the machine learning system classifying the input into one of the second portion of the outputs, the incoming call to a second device instead of any device associated with any operator of the call center ([Carpenter, page 4, sec. 4] “In addition to notifying the caller of a selected destination or querying the caller for further information, an automatic call router should also be able to identify when it is unable to handle a call and route the call to a human operator for further processing.” AND [Carpenter, page 4] “For calls that do not satisfy either criterion, the call router should simply punt them to a human operator.” AND [Carpenter, page 1, Abstract] “Furthermore, our system showed a substantial improvement in performance over existing systems by correctly routing 93.8% of the calls after punting 10.2% of all calls to a human operator on transcription”, wherein the examiner interprets diverting the calls the router cannot confidently classify to the fallback human handler rather than to any of the routing destinations to be the same as routing the incoming call to a second device instead of any device associated with any operator of the call center because they are both delivery of the low-confidence call to a station other than that of any of the call center representatives to which confidently classified calls are sent.)
Thulasidasan and Carpenter do not teach wherein the updating of the adjustable penalty value comprises iteratively performing an update operation on the adjustable penalty value based on a hyperparameter to cause an indicator value to converge to a target value across iterations of evaluating the loss function, the indicator value being based on a mean of the uncertainty values comprising a respective uncertainty value in each iteration of the evaluating the loss function… wherein the update operation is selected from a group consisting of incrementing the adjustable penalty value by a value of the hyperparameter and decrementing the adjustable penalty value by the value of the hyperparameter…the segregating of the outputs comprising truncating processing, using the machine learning system, of respective ones of the first inputs corresponding to the second portion of the outputs…in response to determining that the respective ones of the first inputs correspond to respective ones of the uncertainty values that are greater than a sufficiency threshold… and redirecting the processing to an entity not associated with the machine learning system.
Barnes teaches:
wherein the updating of the adjustable penalty value comprises iteratively performing an update operation on the adjustable penalty value based on a hyperparameter to cause an indicator value to converge to a target value across iterations of evaluating the loss function, the indicator value being based on a mean of the uncertainty values comprising a respective uncertainty value in each iteration of the evaluating the loss function ([Barnes, page 2, Abstract] “The abstention loss is designed to abstain on a user-defined fraction of the samples via a PID controller.” AND [Barnes, page 7, sec. 3.2.2] “The parameter α in Eq. [3] determines how much the network is penalized for abstaining. α can be adaptively modified throughout training so that the network abstains on a specified percent of the training samples…Because of this, we evaluate the PID terms on 6 consecutive batches (32×6=192 samples); this strategy leads to more stable behavior of the abstention fraction, but it does not impede training.”, wherein the examiner interprets adjusting α across training under controller gains so that the abstention fraction measured over the accumulated batches settles at the level the user specified to be the same as iteratively performing an update operation on the adjustable penalty value based on a hyperparameter to cause an indicator value, based on a mean of the uncertainty values, to converge to a target value because they are both repeated adjustment of the penalty term by a settable control constant until an average taken over the model’s per-sample uncertainty outputs reaches a preselected value.)
Thulasidasan, Carpenter, and Barnes do not teach wherein the update operation is selected from a group consisting of incrementing the adjustable penalty value by a value of the hyperparameter and decrementing the adjustable penalty value by the value of the hyperparameter…the segregating of the outputs comprising truncating processing, using the machine learning system, of respective ones of the first inputs corresponding to the second portion of the outputs…in response to determining that the respective ones of the first inputs correspond to respective ones of the uncertainty values that are greater than a sufficiency threshold…and redirecting the processing to an entity not associated with the machine learning system.
Veeraraghavan teaches wherein the update operation is selected from a group consisting of incrementing the adjustable penalty value by a value of the hyperparameter and decrementing the adjustable penalty value by the value of the hyperparameter ([Veeraraghavan, page 15, col. 17, lines 25-28] “For example, if an insult rate is higher than an insult rate equilibrium, S240 may function to increase a fraud or threat score threshold (e.g., from 90 ⇒ 95) for blocking transactions to reduce the number of transactions.” AND [Veeraraghavan, page 15, col. 17, lines 33-37] “In another embodiment , if a computed insult rate is lower than an insult rate equilibrium for a given subscriber , S240 may function to update one or more adverse decision 35 thresholds of an automated workflow such that more activity evaluated within the automated workflow are evaluated adversely.”, wherein the examiner interprets raising the machine-learning parameter by a set amount when the measured rate sits above the equilibrium and lowering it by that amount when the measured rate sits below to be the same as an update operation selected from a group consisting of incrementing the adjustable penalty value by a value of the hyperparameter and decrementing the adjustable penalty value by the value of the hyperparameter because they are both a two-way adjustment of the governing parameter by a fixed quantity according to which side of the target the measured rate falls on.)
Thulasidasan, Carpenter, Barnes, and Veeraraghavan do not teach the segregating of the outputs comprising truncating processing, using the machine learning system, of respective ones of the first inputs corresponding to the second portion of the outputs…in response to determining that the respective ones of the first inputs correspond to respective ones of the uncertainty values that are greater than a sufficiency threshold.
Kung teaches the segregating of the outputs comprising truncating processing, using the machine learning system, of respective ones of the first inputs corresponding to the second portion of the outputs…in response to determining that the respective ones of the first inputs correspond to respective ones of the uncertainty values that are greater than a sufficiency threshold ([Kung, page 2, sec 1] “If the entropy of a test sample is below a learned threshold value, meaning that the classifier is confident in the prediction, the sample exits the network with the prediction result at this exit point, and is not processed by the higher network layers. If the entropy value is above the threshold, then the classifier at this exit point is deemed not confident, and the sample continues to the next exit point in the network.”, wherein the examiner interprets halting a sample’s progress through the remaining network layers once its entropy has been compared against the learned threshold to be the same as truncating processing of respective ones of the first inputs in response to determining that they correspond to uncertainty values greater than a sufficiency threshold because they are both the stopping of further machine-learning computation on a selected input once that input’s uncertainty measure has been tested against a threshold.)
Thulasidasan, Carpenter, Barnes, Veeraraghavan, and Kung do not teach and redirecting the processing to an entity not associated with the machine learning system.
Kang teaches and redirecting the processing to an entity not associated with the machine learning system. ([Kang, page 623, sec. 5.3] “NSmobile sends the output of that layer from the mobile device to NSserver residing on the server side. NSserver then executes the remaining DNN layers.”, wherein the examiner interprets handing the partially completed computation to the server-side component, which carries out the remaining layers, to be the same as redirecting the processing to an entity not associated with the machine learning system because they are both transfer of work that the original model has stopped performing to a separate entity that continues it.)
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, and the instant application are analogous art because they are all directed to machine learning systems that use a measure of prediction uncertainty to decide whether to continue processing an input or divert it elsewhere.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the abstention-penalty training approach disclosed by Thulasidasan to include the “call router” disclosed by Carpenter. One would be motivated to do so to effectively apply the abstention decision to live customer traffic so that only the calls the classifier can handle reliably are sent to a representative, as suggested by Carpenter ([Carpenter, page 1, Abstract] “Furthermore, our system showed a substantial improvement in performance over existing systems by correctly routing 93.8% of the calls after punting 10.2% of all calls to a human operator on transcription”.)
It would have also been obvious to a person of ordinary skill in the art before the effective filing date of the invention to include the “PID controller” disclosed by Barnes. One would be motivated to do so to efficiently hold the proportion of diverted inputs at the level the operator has selected instead of tuning the penalty by hand, as suggested by Barnes ([Barnes, page 2, Abstract] “The abstention loss is designed to abstain on a user-defined fraction of the samples via a PID controller.”)
It would have also been obvious to a person of ordinary skill in the art before the effective filing date of the invention to include the two-way threshold adjustment disclosed by Veeraraghavan. One would be motivated to do so to reliably drive the measured rate back toward its target by a predictable amount in whichever direction the measurement deviates, as suggested by Veeraraghavan (([Veeraraghavan, page 15, col. 17, lines 25-28] “For example, if an insult rate is higher than an insult rate equilibrium, S240 may function to increase a fraud or threat score threshold (e.g., from 90 ⇒ 95) for blocking transactions to reduce the number of transactions.”)
It would have also been obvious to a person of ordinary skill in the art before the effective filing date of the invention to include the early exiting technique disclosed by Kung. One would be motivated to do so to efficiently avoid spending further layer computation on an input the model has already declined to predict, as suggested by Kung ([Kung, page 2, sec 1] “If the entropy of a test sample is below a learned threshold value, meaning that the classifier is confident in the prediction, the sample exits the network with the prediction result at this exit point, and is not processed by the higher network layers. If the entropy value is above the threshold, then the classifier at this exit point is deemed not confident, and the sample continues to the next exit point in the network.”)
It would have also been obvious to a person of ordinary skill in the art before the effective filing date of the invention to include the technique for shifting processing to a remote device disclosed by Kang. One would be motivated to do so to effectively distribute computational workloads across multiple processing entities and reduce local resource consumption, as suggested by Kang ([Kang, page 615] “We evaluate Neurosurgeon on a state-of-the-art mobile development platform and show that it improves end-to-end latency by 3.1× on average and up to 40.7×, reduces mobile energy consumption by 59.5% on average and up to 94.7%, and improves data center throughput by 1.5× on average and up to 6.7×.”
Claims 11 and 16 are analogous to claim 1, aside from claim type and minute differences, and thus the same rejection applies as above.
Claims 2, 5, and 17are rejected under 35 U.S.C. 103 as being unpatentable over Thulasidasan in view of Carpenter in view of Barnes in view of Veeraraghavan in view of Kung in view of Kang further in view of NPL reference “SelectiveNet: A Deep Neural Network with an Integrated Reject Option”, by Geifman et. al. (referred herein as Geifman).
Regarding claim 2, Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, and Kang teaches The device of claim 1 (see rejection of claim 1).
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, and Kang do not teach wherein the loss function facilitates determining a loss vector based on results of the conventional loss function and the uncertainty value.
Geifman teaches wherein the loss function facilitates determining a loss vector based on results of the conventional loss function and the uncertainty value. ([Geifman, page 3, sec 4.2] “This results in the following unconstrained objective, which is averaged over the samples in Sm, Eq, 2] where c is the target coverage, λ is a hyperparameter controlling the relative importance of the constraint, and Ψ is a quadratic penalty function.”, wherein the examiner interprets the objective that combines r̂ℓ (a conventional prediction loss) with λΨ(·) (a tunable penalty driven by the selection head g) to be the same as a loss function being a function of at least a conventional loss function and an adjustable penalty value because they are both directed to updating the model using a base task loss plus an adjustable penalty that is set from uncertainty-related values.)
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, Geifman, and the instant application are analogous art because they are all directed to employing loss functions that combine a conventional prediction loss with an uncertainty-driven term.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the device employing a machine-learning model of claim 1 disclosed by Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, and Kang to include the confidence term, g(x), in the loss function as disclosed by Geifman [Eq. (2)]. One would be motivated to do so to efficiently enforce a target acceptance/coverage level while training the model, as suggested by Geifman ([Geifman, page 3] “a deep neural architecture allowing end-to-end optimization of selective models…to train it for any desired target coverage...L(f,g) = r̂(f,g|S_m) + λΨ(c - φ̂(g|S_m))” and “a neural network with an integrated reject option, designed to be optimal for the required coverage slice.”).
Regarding claim 5, Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, and Geifman teaches The device of claim 2, (see rejection of claim 2).
Geifman further teaches wherein the adjustable penalty value, in a first iteration of evaluating the loss function, is initially set to at least a defined value, resulting in the adjustable penalty value term initially dominating the loss function. ([Geifman, page 3, section 4.2] “To enforce the coverage constraint, we utilize a variant of the well-known Interior Point Method (IPM). This results in the following unconstrained objective, which is averaged over the samples in Sm
PNG
media_image1.png
57
287
media_image1.png
Greyscale
where c is the target coverage, λ is a hyperparameter controlling the relative importance of the constraint, and Ψ is a quadratic penalty function…… λ was set to 32 (this value was found to be large enough to preserve the constraint through the training process).”, wherein the examiner interprets “λ was set to 32 (this value was found to be large enough to preserve the constraint through the training process)” to be the same as “the adjustable… value [loss, penalty, any value]… initially set to at least a defined value, resulting in the adjustable penalty value term initially dominating the loss function” as both are used to initialize to a defined value where the penalty component (i.e. ψ) dominates the loss function over the empirical risk term (i.e. r) during the initial phases of the training process and then later updates adjust the balance.)
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, Geifman, and the instant application are analogous art because they are all directed to tuning or adjusting loss components during the training of a machine learning system based on model uncertainty.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the device of claim 2 disclosed by Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, and Geifman to include the value adjustment approach disclosed by Geifman. One would be motivated to do so to effectively improve the model’s ability to account for prediction uncertainty during training, as suggested by Geifman ([Geifman, page 3, section 4.2] “To enforce the coverage constraint, we utilize a variant of the well-known Interior Point Method (IPM).
Regarding claim 17, Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, and Kang teaches The non-transitory machine-readable storage medium of claim 16, (see rejection of claim 16).
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, and Kang do not teach wherein the updating of the penalty value improves an optimization of the loss function according to a defined performance criterion.
Geifman teaches wherein the updating of the penalty value improves an optimization of the loss function according to a defined performance criterion. ([Geifman, Abstract] “SelectiveNet is trained to optimize both classification (or regression) and rejection simultaneously, end-to-end. The result is a deep neural network that is optimized over the covered domain.” wherein the examiner interprets adjusting the rejection mechanism (e.g. a penalty or cost function), to “optimize both classification (or regression) and rejection simultaneously” to generate a “neural network that is optimized over the covered domain” is the same as “updating the penalty value improves an optimization of the loss function according to a defined performance criterion” because both are optimizing a loss function according to a performance criterion.)
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, Geifman, and the instant application are analogous art because they are all directed to adaptive adjustment of a penalty term in response to uncertainty in order to optimize a learning objective.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the CRM claim 16 disclosed by Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, and Kang to include the “optimize both classification (or regression) and rejection simultaneously, end-to-end” disclosed by Geifman. One would be motivated to do so to effectively optimize/update that can be applied to a loss function, as suggested by Geifman ([Geifman, Abstract] “SelectiveNet is trained to optimize both classification (or regression) and rejection simultaneously, end-to-end. The result is a deep neural network that is optimized over the covered domain.”)
Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Thulasidasan in view of Carpenter in view of Barnes in view of Veeraraghavan, in view of Kung in view of Kang in view of in view of Geifman further in view of US 10452992 B2 by Lee et al. (referred herein as Lee).
Regarding claim 10, Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, and Kang teaches The device of claim 1 (see rejection of claim 1).
Barnes further teaches wherein the target value is selectable based on user input ([Barnes, Abstract] “The abstention loss is designed to abstain on a user-defined fraction of the samples via a PID controller.”, wherein the examiner interprets the fraction of samples the user defines for the controller to hold to be the same as a target value selectable based on user input because they are both a target quantity whose magnitude is chosen by the operator rather than fixed by the system.)
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, and Kang do not teach received via a user interface, or and wherein the hyperparameter is less than 1.
Lee teaches received via a user interface ([Lee, page 78, col. 5-6, lines 65-67, 1] “FIGS. 64a and 64b impact of a change to a prediction interpretation threshold value, indicated by a client via a particular control of an interactive graphical interface, on a set of model quality metrics” AND [Lee, page 5 col. 5, lines 55-60] “FIG. 62 illustrates an example system environment in which a machine learning service implements an interactive graphical interface enabling clients to explore tradeoffs between various prediction quality metric goals, and to modify settings that can be used for interpreting model execution results”, wherein the examiner interprets the client’s manipulation of a control of the interactive graphical interface to set the threshold value the model applies to be the same as a value received via a user interface because they are both the entry of an operator-chosen quantity into the system through a displayed control rather than by fixed configuration.)
Geifman teaches and wherein the hyperparameter is less than 1 ([Geifman, page 1, Introduction] “In all our experiment we used α = 0.5 without any hyperparameter optimization.”, wherein the examiner interprets the weighting hyperparameter fixed at 0.5 to be the same as a hyperparameter that is less than 1 because they are both a control constant whose magnitude is a fraction below unity.)
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, Geifman, Lee, and the instant application are analogous art because they are all directed to selecting the control constants that govern how a machine learning model trades prediction against abstention.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the abstention-penalty training approach disclosed by Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, and Kang to include the fraction of samples the user defines for the controller to hold disclosed by Barnes. One would be motivated to do to efficiently apply abstention penalty training as a way of applying user input target values, as suggested by Barnes ([Barnes, Abstract] “The abstention loss is designed to abstain on a user-defined fraction of the samples via a PID controller.”)
It would have also been obvious to a person of ordinary skill in the art before the effective filing date of the invention to include the “α = 0.5” weighting disclosed by Geifman. One would be motivated to do so to efficiently obtain a step magnitude small enough that the controlled quantity approaches its target without overshooting, as suggested by Geifman ([Geifman, page 1, Introduction] “without any hyperparameter optimization.”).
It would have also been obvious to a person of ordinary skill in the art before the effective filing date of the invention to include the “interactive graphical interface” disclosed by Lee. One would be motivated to do so to effectively permit the operator to set and revise the target value during use rather than reconfiguring the system, as suggested by Lee [Lee, page 5 col. 5, lines 55-60] “FIG. 62 illustrates an example system environment in which a machine learning service implements an interactive graphical interface enabling clients to explore tradeoffs between various prediction quality metric goals, and to modify settings that can be used for interpreting model execution results”.)
Claims 15 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Thulasidasan in view of Carpenter in view of Barnes in view of Veeraraghavan in view of Kung in view of Kang further in view of NPL reference “Deep gamblers: Learning to abstain with portfolio theory.”, by Ziyin et. al. (referred herein as Ziyin).
Regarding claim 15, Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, and Kang teaches The method of claim 11 (see rejection of claim 11).
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, and Kang do not teach further comprising tuning, by the system, the penalty value based on a mean of previous uncertainty values, wherein the tuning comprises incrementing the penalty value by a first amount in response to the mean of the previous uncertainty values being less than a target value, and wherein the tuning comprises decrementing the penalty value by a second amount in response to the mean of the previous uncertainty values being greater than the target value.
Ziyin teaches further comprising tuning, by the system, the penalty value based on a mean of previous uncertainty values, wherein the tuning comprises incrementing the penalty value by a first amount that is less than 1 in response to the mean of the previous uncertainty values being less than a target value, and wherein the tuning comprises decrementing the penalty value by a second amount that is less than 1 in response to the mean of the previous uncertainty values being greater than the target value. ([Ziyin, page 2] “A prediction model augmented with a rejection option is a pair of functions (f, g) such that gₕ : X → R is a selection function … i.e., the model abstains from making a prediction when the selection function g(x) falls below a predetermined threshold h. We call g(x) the uncertainty score of x … The covered dataset is defined to be {x : gₕ(x) ≥ h}, and the coverage is the ratio … One may trade-off coverage for lower risk … We gradually decrease the threshold h … This shows how we might calibrate threshold h to control coverage.”, wherein the examiner interprets that the gradual adjustment of the threshold h based on aggregated uncertainty scores corresponds to tuning the penalty value based on a mean of previous uncertainty values, because both describe feedback-driven calibration of a control parameter to maintain target coverage or confidence. The examiner further interprets the phrase “gradually decrease the threshold h” as corresponding to incrementing or decrementing by small fractions (less than 1) where fine-grained, progressive threshold updates to achieve smooth adjustment toward a target coverage are achieved.)
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, Ziyin, and the instant application are analogous art, because they are all directed to adjusting, on the basis of accumulated uncertainty measures, the control parameter that governs how readily a machine learning model declines to predict.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of claim 11 disclosed by Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, and Kang to include the threshold calibration disclosed by Ziyin. One would be motivated to do so to effectively hold the proportion of inputs on which the model declines to predict at the level the system is aiming for, as suggested by Ziyin ([Ziyin, page 2] “This shows how we might calibrate threshold h to control coverage.”).
Claim 20 is analogous to claim 15, aside from claim type and minute differences, and thus the same rejection applies as above.
Claim 4 and 9 are rejected under 35 U.S.C. 103 as being unpatentable over Thulasidasan in view of Carpenter in view of Barnes in view of Veeraraghavan in view of Kung in view of in view of Kang further in view of NPL reference “Classification with a Reject Option using a Hinge Loss.”, by Wegkemp et. al. (referred herein as Wegkemp). Claim 18 is rejected under 35 U.S.C. 103 as being unpatentable over Thulasidasan in view of Carpenter in view of Barnes in view of Veeraraghavan in view of Kung in view of in view of Kang in view of Geifman further in view of Wegkemp.
Regarding claim 4, Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, and Kang teaches The device of claim 2 (see rejection of claim 2).
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, and Kang do not teach wherein the loss function is Lnew(xj)=(1-puncert)*Lold(xj)+puncert*Cpenalty, wherein puncert is the uncertainty value, wherein Lold(xj) are the results of the conventional loss function, and wherein Cpenalty is the adjustable penalty value.
Wegkamp teaches wherein the loss function is Lnew(xj)=(1-puncert)*Lold(xj)+puncert*Cpenalty, wherein puncert is the uncertainty value, wherein Lold(xj) are the results of the conventional loss function, and wherein Cpenalty is the adjustable penalty value. ([Wegkamp, page 2] Introduction] “We propose to incorporate the reject option into our classification scheme by using a threshold value 0 ≤ δ < 1 as follows. Given a discriminant function f : X → R, we report sgn(f(x))) ∈ {-1,1} if | f(x)| > δ, but we withhold decision if | f(x)| ≤ δ and report r. In this note, we assume that the cost of making a wrong decision is 1 and the cost of using the reject option is d > 0. The appropriate risk function is then
PNG
media_image2.png
175
648
media_image2.png
Greyscale
.” wherein the examiner interprets P{|Y f(X)| ≤ δ} to be the same as puncert (probability / uncertainty of the ambiguous region) because both have to do with uncertainty. Further clarification is seen from the specification of the instant application “[0023] As an example, Uncertainty-type Loss Function (ULF) component 260 can employ the formula Lnew(xj)=(1-puncert)*Lold(xj)+puncert*Cpenalty where L(x,) is conventional loss function data, e.g., CLF data from CLF component 250, where puncert information embodied in UPS 240, where Cpenalty is a penalty value discussed in further detail elsewhere herein, and where L, (x,) is uncertainty based loss data derived from CLF data.”. The examiner further interprets P{Y f(X) < -δ} to be the same as Lold(xj) because they are both related to the ordinary classification error of a neural network / ML model. Furthermore, the constant d in the above formula is interpreted to be the same as Cpenalty because they both have to do with cost of abstaining. The examiner finally interprets that together with previous interpretations, Ld,δ (f) is the same as Lnew(xi) because they are both associated with the total loss.”).
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, Wegkamp, and the instant application are analogous art because they are all directed to loss functions that integrate uncertainty and cost-sensitive abstention mechanisms in machine learning systems.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the device of claim 2 disclosed by Kung, Chen, Kang, and Geifman to include the “cost of using the reject option is d > 0” disclosed by Wegkamp. One would be motivated to do so to effectively account for the cost tradeoff associated with abstaining from low-confidence predictions, as suggested by Wegkamp ([Wegkamp, page 2] “In this note, we assume that the cost of making a wrong decision is 1 and the cost of using the reject option is d > 0.”).
Claim 18 is analogous to claim 4, aside from claim type and minute differences, and thus the same rejection applies as above.
Regarding claim 9, Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, and Kang teaches The device of claim 1 (see rejection of claim 1).
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, and Kang do not teach wherein the updating of the adjustable penalty value comprises restricting the adjustable penalty value from traversing a floor penalty value.
Wegkamp teaches wherein the updating of the adjustable penalty value comprises restricting the adjustable penalty value from traversing a floor penalty value. ([Wegkamp, page 2, Introduction] “we assume that the cost of making a wrong decision is 1 and the cost of using the reject option is d > 0, i.e., a minimum penalty that cannot be crossed] ...[Eq. 1] Ld,δ(f)=Pr{Yf(X)<-δ}+dPr{∣Yf(X)∣≤δ}”, wherein the examiner interprets this fixed positive reject cost as a floor on the penalty term (a lower bound that the penalty cannot go below), which restricts the penalty from traversing that floor during training/optimization of the objective, thereby meeting the claim language that the updating “comprises restricting … from traversing a floor penalty value.”)
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, Wegkamp, and the instant application are analogous art because each addresses loss-function penalties that shape model training or behavior, including enforcing a floor on the penalty.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the device of claim 1 disclosed by Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, and Kang to include Wegkamp’s fixed positive reject cost d>0. One would be motivated to do so to effectively prevent collapse of the penalty to zero and preserve a meaningful abstain trade-off, i.e., to efficiently enforce a floor on the penalty during updates, as suggested by Wegkamp ([Wegkamp, page 2, Introduction] “we assume that the cost of making a wrong decision is 1 and the cost of using the reject option is d > 0.”)
Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Thulasidasan in view of Carpenter in view of Barnes in view of Veeraraghavan in view of Kung in view of Kang in view of Geifman further in view of Wegkamp.
Regarding claim 12, Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, and Kang teaches The method of claim 11 (see rejection of claim 1, it is analogous to claim 11 and 16).
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, and Kang do not teach wherein the updating the penalty value and the adapting the machine learning model increase an efficacy of the loss function according to a defined performance criterion, wherein the loss function is Lnew(xj)=(1-puncert)*Lold(xj)+puncert*Cpenalty, wherein puncert is the uncertainty value, wherein Lold(xj) are second results of the conventional loss function, and wherein Cpenalty is the penalty value.
Geifman teaches wherein the updating the penalty value and the adapting the machine learning model increase an efficacy of the loss function according to a defined performance criterion, ([Geifman, page 1, Introduction] “A selective loss function that optimizes a specified coverage slice using a variant of the interior point optimization method”, wherein the examiner interprets the “selective loss function that optimizes a specified coverage slice” to be the same as the “adapting of the machine learning model increase an efficacy of the loss function according to a defined performance criterion”, as they both optimize the selective loss, matching the notion that adapting the model increases an “efficacy” of the loss function.
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, Geifman, and the instant application are analogous art, because they are all directed to updating a clue and machine learning adapting for optimal performance.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of claim 11 disclosed by Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, and Kang to include the “selective loss function that optimizes a specified coverage slice using a variant of the interior point optimization method” disclosed by Geifman. One would be motivated to do so to efficiently use a loss function to optimize loss in hopes to improve performance/efficacy as suggested by Geifman ([Geifman, page 1, Introduction] “A selective loss function that optimizes a specified coverage slice using a variant of the interior point optimization method”.)
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, and Geifman do not teach wherein the loss function is Lnew(xj)=(1-puncert)*Lold(xj)+puncert*Cpenalty, wherein puncert is the uncertainty value, wherein Lold(xj) are second results of the conventional loss function, and wherein Cpenalty is the penalty value.
Wegkamp teaches wherein the loss function is Lnew(xj)=(1-puncert)*Lold(xj)+puncert*Cpenalty, wherein puncert is the uncertainty value, wherein Lold(xj) are second results of the conventional loss function, and wherein Cpenalty is the penalty value. ([Wegkamp, page 2, Introduction] “We propose to incorporate the reject option into our classification scheme by using a threshold value 0 ≤ δ < 1 as follows. Given a discriminant function f : X → R, we report sgn(f(x))) ∈ {-1,1} if | f(x)| > δ, but we withhold decision if | f(x)| ≤ δ and report r. In this note, we assume that the cost of making a wrong decision is 1 and the cost of using the reject option is d > 0. The appropriate risk function is then
PNG
media_image3.png
175
648
media_image3.png
Greyscale
.”, wherein the examiner interprets P{|Y f(X)| ≤ δ} to be the same as puncert (probability / uncertainty of the ambiguous region) because both have to do with uncertainty. Further clarification is seen from the specification of the instant application “[0023] As an example, Uncertainty-type Loss Function (ULF) component 260 can employ the formula Lnew(xj)=(1-puncert)*Lold(xj)+puncert*Cpenalty where L(x,) is conventional loss function data, e.g., CLF data from CLF component 250, where puncert information embodied in UPS 240, where Cpenalty is a penalty value discussed in further detail elsewhere herein, and where L, (x,) is uncertainty based loss data derived from CLF data.”. The examiner further interprets P{Y f(X) < -δ} to be the same as Lold(xj) because they are both related to the ordinary classification error of a neural network / ML model. Furthermore, the constant d in the above formula is interpreted to be the same as Cpenalty because they both have to do with cost of abstaining. The examiner finally interprets that together with previous interpretations, Ld,δ (f) is the same as Lnew(xi) because they are both associated with the total loss.”).
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, Geifman, Wegkamp, and the instant application are analogous art because they are all directed to loss functions that integrate uncertainty and cost-sensitive abstention mechanisms in machine learning systems.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of claim 11 disclosed by Ziyin, Cortes, Kung, and Kang, and the coverage constraint optimization disclosed by Geifman and to include the “cost of using the reject option is d > 0” disclosed by Wegkamp. One would be motivated to do so to effectively account for the cost tradeoff associated with abstaining from low-confidence predictions, as suggested by Wegkamp (Wegkamp, page 2, Introduction] “In this note, we assume that the cost of making a wrong decision is 1 and the cost of using the reject option is d > 0.”)
Claim 13 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Thulasidasan in view of Carpenter in view of Barnes in view of Veeraraghavan in view of Kung in view of Kang further in view of NPL reference “Boosting with abstention.” by Cortes et. al. (referred herein as Cortes).
Regarding claim 13, Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, and Kang teaches The method of claim 11, (see rejection of claim 11).
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, and Kang do not teach wherein the updating of the penalty value comprises evaluating the formula Cpenalty = mean(Lold(xbatch_j)), and wherein Lold(xbatch_j) are the first results of the conventional loss function.
Cortes teaches wherein the updating of the penalty value comprises evaluating the formula Cpenalty = mean(Lold(xbatch_j)), and wherein Lold(xbatch_j) are the first results of the conventional loss function. ([Cortes, page 2, section 2.1] “The abstention cost c(x) is assumed known to the learner. In the following, we assume that c is a constant function…
PNG
media_image4.png
55
550
media_image4.png
Greyscale
.”, wherein the examiner interprets defining a constant abstention cost c(x) that is applied within the per-sample conventional loss function to be the same as evaluating the formula for Cpenalty as a mean of the first results of the conventional loss function, where the “abstention cost” is the same as “penalty value” as they both are a “cost” in governing the contribution of terms.)
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, Cortes, and the instant application are analogous art because they are all directed to modifying or adjusting a penalty value in a loss function to manage model uncertainty or abstention behavior during learning.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of claim 11 disclosed by Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, and Kang to include the definition of a constant abstention cost used in the per-sample loss function disclosed by Cortes. One would be motivated to do so to effectively make a determination of when to abstain, as suggested by Ziyin (Ziyin, page 2] section 2) “the model abstains from making a prediction when the selection function g(x) falls below a predetermined threshold h.”)
Regarding Claim 19, Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, and Kang teaches The non-transitory machine-readable storage medium of claim 16, (see rejection of claim 16).
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, and Kang do not teach wherein the updating the penalty value comprises evaluating the formula Cpenalty=max(mean(Lold(xbatch_j)), Cmin), wherein Cpenalty is the penalty value, wherein Lold(xbatch_j) are the previous results of the conventional loss function, and wherein Cmin is a minimum allowable penalty value.
Cortes teaches wherein the updating the penalty value comprises evaluating the formula Cpenalty=max(mean(Lold(xbatch_j)), Cmin), wherein Cpenalty is the penalty value, wherein Lold(xbatch_j) are the previous results of the conventional loss function, and wherein Cmin is a minimum allowable penalty value. [Cortes, page 2, section 2.1] “The abstention cost c(x) is assumed known to the learner. In the following, we assume that c is a constant function…
PNG
media_image4.png
55
550
media_image4.png
Greyscale
and [Cortes, page 8, section 5] “For each cost c, the hyperparameter configuration was chosen to be the set of parameters that attained the smallest average rejection loss on the validation set. For that set of parameters we report the results on the test set.”, wherein the examiner interprets the constant abstention cost c selected to minimize average rejection loss and applied in the per-example abstention loss function to be the same as Cpenalty, the penalty value that is updated based on a mean of previous conventional loss results and constrained by a minimum allowable value, as both terms are directed to a cost that modulates the loss function during optimization.)
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, Cortes, and the instant application are analogous art because they are all directed to machine learning systems that incorporate selective prediction with adjustable or learnable abstention penalties within a loss function.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the machine-readable storage medium of claim 16 disclosed by Kung, Chen, and Kang to include the abstention-aware loss function using a fixed cost disclosed by Cortes. One would be motivated to do so to effectively use the cost function to calculate loss as suggested by Cortes ([Cortes, page 2, section 2.1] “The abstention cost c(x) is assumed known to the learner. In the following, we assume that c is a constant function…[Cortes, page 8, section 5] “For each cost c, the hyperparameter configuration was chosen to be the set of parameters that attained the smallest average rejection loss on the validation set.”)
Claim 14 is rejected under 35 U.S.C. 103 as being unpatentable over Thulasidasan in view of Carpenter in view of Barnes in view of Veeraraghavan in view of Kung in view of Kang in view of Cortes further in view of Wegkamp.
Regarding claim 14, Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, and Cortes teaches The method of claim 13 (see rejection of claim 13).
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, and Cortes do not teach wherein the updating of the penalty value comprises preventing the penalty value from dropping below a minimum value according to the formula C_penalty = max(C_penalty, C_min), wherein C_min is the minimum value.
Wegkamp teaches wherein the updating of the penalty value comprises preventing the penalty value from dropping below a minimum value according to the formula C_penalty = max(C_penalty, C_min), wherein C_min is the minimum value. ([Wegkamp, page 2, Introduction] “In this note, we assume that the cost of making a wrong decision is 1 and the cost of using the reject option is d > 0.”, wherein the examiner interprets holding the rejection cost strictly above zero to be the same as preventing the penalty value from dropping below a minimum value because they are both the enforcement of a lowest permitted magnitude for the cost charged for declining to predict)
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, Cortes, Wegkamp, and the instant application are analogous art because they are all directed to bounding the cost a machine learning model is charged for declining to predict.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the abstention-penalty training approach disclosed by Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, and Cortes to include the “cost of using the reject option is d > 0” constraint disclosed by Wegkamp. One would be motivated to do so to effectively keep the incentive to predict intact across the whole of training, as suggested by Wegkamp ([Wegkamp, page 2, Introduction] “the cost of using the reject option is d > 0.”).
Claims 22 and 23 are rejected under 35 U.S.C. 103 as being unpatentable over Thulasidasan in view of Carpenter in view of Barnes in view of Veeraraghavan in view of Kung in view of Kang in view of Geifman and further in view of the NPL reference “Pareto Multi-Task Learning” by Lin et al. (referred herein as Lin).
Regarding claim 22, Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, and Geifman teaches The device of claim 2 (see rejection of claim 2).
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, and Geifman do not teach repeating the updating of the machine learning model until the loss vector and the adjustable penalty value are stabilized according to a defined stability criterion.
Lin teaches repeating the updating of the machine learning model until the loss vector and the adjustable penalty value are stabilized according to a defined stability criterion. ([Lin, page 5] “The gradient-based update rule is θrt+1 = θrt + ηr drt and will be stopped once a feasible solution is found or a predefined number of iterations is met.” AND [Lin, page 7] “where we adaptively assign the weights λi by solving the following problem in each iteration:”, wherein the examiner interprets continuing the parameter and weight updates until the stated stopping condition is met to be the same as repeating the updating until the loss vector and the adjustable penalty value are stabilized according to a defined stability criterion because they are both the continuation of iterative training updates until a preset condition signals that further updating is unnecessary)
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, Geifman, Lin, and the instant application are analogous art because they are all directed to iteratively updating a machine learning model and the weights applied to its loss terms until a stopping condition is met.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the device of claim 2 disclosed by Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, and Geifman to include the stopping condition disclosed by Lin. One would be motivated to do so to efficiently end training once further updates no longer change the solution, as suggested by Lin ([Lin, page 5] “will be stopped once a feasible solution is found or a predefined number of iterations is met.”).
Regarding claim 23, Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, Geifman, and Lin teaches The device of claim 22 (see rejection of claim 22).
Lin further teaches in response to the adjustable penalty value being determined to be stabilized according to the defined stability criterion, modifying the adjustable penalty value by an incremental value based on a difference between loss values associated with the loss vector and a defined preferred loss value. ([Lin, page 7] “where we adaptively assign the weights λi by solving the following problem in each iteration:”, wherein the examiner interprets recomputing, at each iteration, the weight assigned to each loss term from how the present loss gradients compare with the targeted balance to be the same as modifying the adjustable penalty value by an incremental value based on a difference between loss values and a defined preferred loss value because they are both adjustment of the weight applied to the objective by an amount derived from the gap between the current losses and the loss the system is aiming for)
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, Geifman, Lin, and the instant application are analogous art because they are all directed to adjusting the weight applied to a loss term from the gap between the current losses and a targeted loss.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the abstention-penalty training approach disclosed by Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, Geifman, and Lin to include the adaptive weight assignment disclosed by Lin. One would be motivated to do so to effectively bring the achieved losses to the balance the system is aiming for once the model has settled, as suggested by Lin ([Lin, page 7] “we adaptively assign the weights λi by solving the following problem in each iteration.”).
Claim 24 is rejected under 35 U.S.C. 103 as being unpatentable over Thulasidasan in view of Carpenter in view of Barnes in view of Veeraraghavan in view of Kung in view of Kang in in view of Geifman in view of Lin in view of NPL reference “GradNorm: Gradient Normalization for Adaptive Loss Balancing in Deep Multitask Networks” by Chen et. al. (referred herein as Chen) further in view of the NPL reference “Adaptive Transfer Learning on Graph Neural Networks” by Han et al. (referred herein as Han).
Regarding claim 24, Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, Geifman, and Lin teaches The device of claim 22 (see rejection of claim 22).
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, Geifman, and Lin do not teach wherein the updating of the machine learning model comprises: setting the adjustable penalty value to a mean value of the conventional loss function as applied to a first batch of the second inputs processed before the first inputs, updating the loss vector based on the adjustable penalty value; repeating the setting of the adjustable penalty value and the updating of the loss vector for other batches of the second inputs, other than the first batch, until the adjustable penalty value is stabilized according to a defined stability criterion.
Chen teaches:
wherein the updating of the machine learning model comprises: setting the adjustable penalty value to a mean value of the conventional loss function as applied to a first batch of the second inputs processed before the first inputs, ([Chen, page 8] “We introduced GradNorm, an efficient algorithm for tuning loss weights in a multi-task learning setting based on balancing the training rates of different tasks.”, wherein the examiner interprets “tuning loss weights” based on task-training signals aggregated at the batch level to be the same as setting the adjustable penalty value to a mean value of the conventional loss function as applied to a first batch of the second inputs processed before the first inputs because they are both directed to determining a loss-side weight from batch-derived loss information obtained from inputs processed earlier in training.)
updating the loss vector based on the adjustable penalty value; ([Chen, page 3] “GradNorm is then implemented as an L1 loss function Lgrad between the actual and target gradient norms…The computed gradients ∇wiLgrad are then applied via standard update rules to update each wi”, wherein the examiner interprets “implemented as an L1 loss function Lgrad” and “applied via standard update rules to update each wi” to be the same as updating the loss vector based on the adjustable penalty value because they are both directed to modifying the loss-side quantities used in training as a function of the current adjustable penalty-related weight.)
repeating the setting of the adjustable penalty value and the updating of the loss vector for other batches of the second inputs, other than the first batch, ([Chen, page 3] “we update our loss weights wi(t) to move gradient norms towards this target for each task”, wherein the examiner interprets “update our loss weights wi(t)” to be the same as repeating the setting of the “adjustable penalty value and the updating of the loss vector for other batches of the second inputs, other than the first batch” because they are both directed to performing the loss-weight adjustment iteratively over successive training iterations/batches rather than only once.)
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, Geifman, Lin, and Chen do not teach until the adjustable penalty value is stabilized according to a defined stability criterion.
Han teaches until the adjustable penalty value is stabilized according to a defined stability criterion. ([Han, page 6, sec 3.2] “GNN is trained by joint losses of both the target and auxiliary tasks, given the weights of different tasks are determined by the weighting model, we define the multi-task objective of GNN as follows:…We train both networks in an iterative manner until convergence.”, wherein the examiner interprets “We train both networks in an iterative manner until convergence” to be the same as until the adjustable penalty value is stabilized according to a defined stability criterion because they are both directed to repeatedly updating a training-related adjustable weighting/penalty term under a defined iterative process until a termination condition reflecting stabilization or convergence is satisfied.)
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, Geifman, Lin, Chen, Han, and the instant application are analogous art because they are all directed to adaptive machine learning training using batch-based loss information to iteratively update weighting.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the device of claim 22 disclosed by Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, Geifman, and Lin to include the “tuning loss weights” technique disclosed by Chen. One would be motivated to do so to efficiently set and re-set the adjustable penalty value from the loss information gathered over each successive batch rather than fixing that value in advance, as suggested by Chen ([Chen, page 3] “we update our loss weights wi(t) to move gradient norms towards this target for each task.”).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to include the adaptive tuning technique disclosed by Han. One would be motivated to do so to effectively ensure that the iteratively updated adjustable penalty value reaches a stable condition before termination of training updates, as suggested by Han ([Han, page 6, sec. 3.2] “We train both networks in an iterative manner until convergence.”).
Claim 25 is rejected under 35 U.S.C. 103 as being unpatentable over Thulasidasan in view of Carpenter in view of Barnes in view of Veeraraghavan, Kung, and Kang, and further in view of US 9386152B2 by Riahi et al. (referred herein as Riahi).
Regarding claim 25, Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, and Kang teaches The device of claim 1 (see rejection of claim 1).
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, and Kang do not teach wherein the first inputs are provided via respective originators of the incoming calls to the call center via navigation of a phone tree.
Riahi teaches wherein the first inputs are provided via respective originators of the incoming calls to the call center via navigation of a phone tree. ([Riahi, Abstract] “an interactive voice response (IVR) node configured to engage in an incoming interaction from a customer to the contact center by presenting set scripts to the customer and receiving corresponding responses”, wherein the examiner interprets the responses a customer supplies to the scripted prompts the interactive voice response node presents during an incoming interaction to be the same as first inputs provided via respective originators of the incoming calls through navigation of a phone tree because they are both the data a caller enters by working through a sequence of automated prompts on an incoming call)
Thulasidasan, Carpenter, Barnes, Veeraraghavan, Kung, Kang, Riahi, and the instant application are analogous art because they are all directed to obtaining, from an incoming call, the input on which an automated system bases a routing decision.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the abstention-penalty training approach disclosed by Thulasidasan to include the “interactive voice response (IVR) node” disclosed by Riahi. One would be motivated to do so to effectively collect a structured input from the caller before any routing decision is made, as suggested by Riahi ([Riahi, Abstract] “an interactive voice response (IVR) node configured to engage in an incoming interaction from a customer to the contact center by presenting set scripts to the customer and receiving corresponding responses”).
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DEVAN KAPOOR whose telephone number is (703)756-1434. The examiner can normally be reached Monday - Friday: 9:00AM - 5:00 PM EST (times may vary).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Yi can be reached at (571) 270-7519. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DEVAN KAPOOR/Examiner, Art Unit 2126
/DAVID YI/Supervisory Patent Examiner, Art Unit 2126