DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim 1-20 is rejected under 35 USC § 101 because claimed invention is directed to the abstract idea without significantly more.
As to claim 1:
Step 1 Analysis: Is the claim to a process, machine, manufacture or composition of matter? See MPEP § 2106.03.
Claim 1 is a method claim, therefore it falls under one of four categories of statutory subject matter.
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation if the difference is greater than a particular threshold: identifying, as an anomalous set, said each data item, the corresponding data item in the second target data, and a corresponding data item in the second input data; ” is an abstract idea of a mental process. See MPEP § 2106.04(a)(2)(III).
The limitation adding the anomalous set to a set of anomaly data; generating second training data based on the first training data and the set of anomaly data; ” is an abstract idea of mathematical concept. See MPEP § 2106.04(a)(2)(I)(A).
The limitation training a second machine-learned model based on the second training data; based on one or more anomalous sets in the set of anomaly data and the second machine- learned model, is an abstract idea of mathematical concept. See MPEP § 2106.04(a)(2)(I)(A).
The limitation computing a feature attribution value for each of one or more features of the second machine-learned model; is an abstract idea of mathematical concept. See MPEP § 2106.04(a)(2)(I)(A).
Step 2A Prong Two Analysis: Does the claim recite additional elements that integrate the judicial exception into a practical application? See MPEP § 2106.04(d).
The limitation “ training a first machine-learned model based on first training data that comprises first input data and first target data ; wherein the method is performed by one or more computing devices. is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §§ 2106.04(d), 2106.05(f)(2).
The limitation generating, using the first machine-learned model, based on the first input data, first output data; generating, using the first machine-learned model, based on second input data, second output data; for each data item in the second output data: ” is an additional element that amounts to adding insignificant extra-solution activity of mere input/output to the judicial exception. The claim recites learning the model from parameters and outputting a model from learned parameters. See MPEP §§ 2106.04(d), 2106.05(g).
The limitation generating a difference between said each data item and a corresponding data item in second target data that corresponds to the second output data is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §§ 2106.04(d), 2106.05(f)(2).
Step 2B Analysis: Does the claim recite additional elements that amount to significantly more than the judicial exception? See MPEP § 2106.05.
The limitation “ training a first machine-learned model based on first training data that comprises first input data and first target data ; wherein the method is performed by one or more computing devices. is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §§ 2106.04(d), 2106.05(f)(2).
Therefore, in examining elements as recited by the limitations individually and as an ordered combination, as a whole the independent claim limitations do not recite what have the courts have identified as “significantly more”.
As to claim 2
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation “wherein generating the second training data comprises: for each anomalous set in the set of anomaly data:” ” is the abstract idea of a mental process. See MPEP § 2106.04(a)(2)(III).
The limitation “computing a difference between a data item from the second output data and a corresponding data item in the second target data; including the difference in a training instance; adding the training instance to the second training data.” is an abstract idea of mathematical concept. See MPEP § 2106.04(a)(2)(I)(A).
As to claim 3
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation “wherein generating the second training data comprises: for each data item in the first target data” is the abstract idea of a mental process. See MPEP § 2106.04(a)(2)(III).
The limitation “ computing a difference between (1) said each data item in the first target data and (2) a corresponding data item in the first output data; including the difference in a training instance; adding the training instance to the second training data.” is an abstract idea of mathematical concept. See MPEP § 2106.04(a)(2)(I)(A).
As to claim 4
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation “wherein the second machine-learned model is a regression model. is an abstract idea of mathematical concept. See MPEP § 2106.04(a)(2)(I)(A).
As to claim 5
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation “wherein computing the feature attribution for each of the one or more features of the second machine-learned model comprises computing a feature attribution for each feature of the second machine-learned model.” is an abstract idea of mathematical concept. See MPEP § 2106.04(a)(2)(I)(A).
As to claim 6
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation” wherein the first feature attribution value is a Shapley value.“ is an abstract idea of mathematical concept. See MPEP § 2106.04(a)(2)(I)(A).
As to claim 7
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation “wherein computing the feature attribution value is performed for a plurality of anomalous sets in the set of anomaly data, .“ is an abstract idea of mathematical concept. See MPEP § 2106.04(a)(2)(I)(A).
The limitation “wherein a plurality of feature attribution values are generated for a particular feature of the second machine-learned model” is the abstract idea of a mental process. See MPEP § 2106.04(a)(2)(III).
The limitation “performing an aggregation operation on the plurality of feature attribution values to generate a global value for the particular feature.“ is an abstract idea of mathematical concept. See MPEP § 2106.04(a)(2)(I)(A).
As to claim 8
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation “wherein the aggregation operation is a mean operation or a median operation. “ is an abstract idea of mathematical concept. See MPEP § 2106.04(a)(2)(I)(A).
As to Claim 9
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation “wherein computing the feature attribution value is performed for a plurality of anomalous sets in the set of anomaly data” is an abstract idea of mathematical concept. See MPEP § 2106.04(a)(2)(I)(A).
The limitation” for each feature of the plurality of features, performing an aggregation operation on the plurality of feature attribution values that correspond to said each feature to generate a global value for said each feature.” is an abstract idea of mathematical concept. See MPEP § 2106.04(a)(2)(I)(A).
Step 2A Prong Two Analysis: Does the claim recite additional elements that integrate the judicial exception into a practical application? See MPEP § 2106.04(d).
The limitation “wherein a plurality of feature attribution values is generated for each feature of a plurality of features of the second machine-learned model” is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §§ 2106.04(d), 2106.05(f)(2).
Step 2B Analysis: Does the claim recite additional elements that amount to significantly more than the judicial exception? See MPEP § 2106.05.
The limitation “wherein a plurality of feature attribution values is generated for each feature of a plurality of features of the second machine-learned model” is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or merely uses a computer in its ordinary capacity as a tool to perform an existing process , which the courts have recognized as well‐understood, routine, and conventional when they are claimed in a generic manner. See MPEP § 2106.05(d)(II).
As to Claim 10
Step 2A Prong Two Analysis: Does the claim recite additional elements that integrate the judicial exception into a practical application? See MPEP § 2106.04(d).
The limitation “ ranking the plurality of features based on the global value for each feature in the plurality of features.” ” is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §§ 2106.04(d), 2106.05(f)(2).
Step 2B Analysis: Does the claim recite additional elements that amount to significantly more than the judicial exception? See MPEP § 2106.05
The limitation “ ranking the plurality of features based on the global value for each feature in the plurality of features.” ” is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or merely uses a computer in its ordinary capacity as a tool to perform an existing process which the courts have recognized as well‐understood, routine, and conventional when they are claimed in a generic manner. See MPEP § 2106.05(d)(II).
As to Claim 11
Step 2A Prong Two Analysis: Does the claim recite additional elements that integrate the judicial exception into a practical application? See MPEP § 2106.04(d).
The limitation “causing, to be displayed, on a screen of a computing device, a name for each feature of the plurality of features and the global value for each feature in the plurality of features. ” is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. (i.e. mere data output). See MPEP §§ 2106.04(d), 2106.05(g).
Step 2B Analysis: Does the claim recite additional elements that amount to significantly more than the judicial exception? See MPEP § 2106.05
The limitation “causing, to be displayed, on a screen of a computing device, a name for each feature of the plurality of features and the global value for each feature in the plurality of features. ” is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. (i.e. mere data output). which the courts have recognized as well‐understood, routine, and conventional when they are claimed in a generic manner. See MPEP § 2106.05(d)(II). See MPEP §§ 2106.04(d), 2106.05(g).
As to Claim 12
Step 1 Analysis: Is the claim to a process, machine, manufacture or composition of matter? See MPEP § 2106.03.
Claim 12 is a method claim, therefore it falls under one of four categories of statutory subject matter.
Step 2A Prong Two Analysis: Does the claim recite additional elements that integrate the judicial exception into a practical application? See MPEP § 2106.04(d).
The limitation “One or more non-transitory storage media storing instructions which, when executed by one or more computing devices” is additional element is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §§ 2106.04(d), 2106.05(f)(2)
Step 2B Analysis: Does the claim recite additional elements that amount to significantly more than the judicial exception? See MPEP § 2106.05.
The limitation “One or more non-transitory storage media storing instructions which, when executed by one or more computing devices” is the additional claim elements of one or more computing devices; a memory coupled to at least one of the computing devices; a stored instructions stored in the memory and executed by at least one of the computing devices are not sufficient to amount to significantly more than the judicial exception since these additional claim elements are recited at a high level of generality (i.e. using a generic processor and generic memory.
And for all other claim elements of claim 12 they are rejected using the PEG analysis of claim 1 since they are analogous claims.
As to Claim 13
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation “wherein generating the second training data comprises: for each anomalous set in the set of anomaly data:” ” is the abstract idea of a mental process. See MPEP § 2106.04(a)(2)(III).
The limitation “computing a difference between a data item from the second output data and a corresponding data item in the second target data; including the difference in a training instance; adding the training instance to the second training data.” is an abstract idea of mathematical concept. See MPEP § 2106.04(a)(2)(I)(A).
As to claim 14
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation “wherein generating the second training data comprises: for each data item in the first target data” is the abstract idea of a mental process. See MPEP § 2106.04(a)(2)(III).
The limitation “ computing a difference between (1) said each data item in the first target data and (2) a corresponding data item in the first output data; including the difference in a training instance; adding the training instance to the second training data.” is an abstract idea of mathematical concept. See MPEP § 2106.04(a)(2)(I)(A).
As to claim 15
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation “wherein computing the feature attribution for each of the one or more features of the second machine-learned model comprises computing a feature attribution for each feature of the second machine-learned model.” is an abstract idea of mathematical concept. See MPEP § 2106.04(a)(2)(I)(A).
As to claim 16
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation” wherein the first feature attribution value is a Shapley value.“ is an abstract idea of mathematical concept. See MPEP § 2106.04(a)(2)(I)(A).
As to claim 17
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation “wherein computing the feature attribution value is performed for a plurality of anomalous sets in the set of anomaly data, .“ is an abstract idea of mathematical concept. See MPEP § 2106.04(a)(2)(I)(A).
The limitation “wherein a plurality of feature attribution values are generated for a particular feature of the second machine-learned model” is the abstract idea of a mental process. See MPEP § 2106.04(a)(2)(III).
The limitation “performing an aggregation operation on the plurality of feature attribution values to generate a global value for the particular feature.“ is an abstract idea of mathematical concept. See MPEP § 2106.04(a)(2)(I)(A).
As to claim 18
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation “wherein the aggregation operation is a mean operation or a median operation. “ is an abstract idea of mathematical concept. See MPEP § 2106.04(a)(2)(I)(A).
As to Claim 19
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation “wherein computing the feature attribution value is performed for a plurality of anomalous sets in the set of anomaly data” is an abstract idea of mathematical concept. See MPEP § 2106.04(a)(2)(I)(A).
The limitation” for each feature of the plurality of features, performing an aggregation operation on the plurality of feature attribution values that correspond to said each feature to generate a global value for said each feature.” is an abstract idea of mathematical concept. See MPEP § 2106.04(a)(2)(I)(A).
Step 2A Prong Two Analysis: Does the claim recite additional elements that integrate the judicial exception into a practical application? See MPEP § 2106.04(d).
The limitation “wherein a plurality of feature attribution values is generated for each feature of a plurality of features of the second machine-learned model” is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §§ 2106.04(d), 2106.05(f)(2).
Step 2B Analysis: Does the claim recite additional elements that amount to significantly more than the judicial exception? See MPEP § 2106.05.
The limitation “wherein a plurality of feature attribution values is generated for each feature of a plurality of features of the second machine-learned model” is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or merely uses a computer in its ordinary capacity as a tool to perform an existing process , which the courts have recognized as well‐understood, routine, and conventional when they are claimed in a generic manner. See MPEP § 2106.05(d)(II).
As to Claim 20
Step 2A Prong Two Analysis: Does the claim recite additional elements that integrate the judicial exception into a practical application? See MPEP § 2106.04(d).
The limitation “causing, to be displayed, on a screen of a computing device, a name for each feature of the plurality of features and the global value for each feature in the plurality of features. ” is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. (i.e. mere data output). See MPEP §§ 2106.04(d), 2106.05(g).
Step 2B Analysis: Does the claim recite additional elements that amount to significantly more than the judicial exception? See MPEP § 2106.05
The limitation “causing, to be displayed, on a screen of a computing device, a name for each feature of the plurality of features and the global value for each feature in the plurality of features. ” is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. (i.e. mere data output). which the courts have recognized as well‐understood, routine, and conventional when they are claimed in a generic manner. See MPEP § 2106.05(d)(II). See MPEP §§ 2106.04(d), 2106.05(g).
Claim Rejections – 35 USC § 103
The following is a quotation of 35 U.S.C. 103, which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Claims 1-3 , 12-14 are rejected under 35 U.S.C. 103 as being unpatentable over Moyne et. al.
Pre-Grant Publication NO. US 2023/0126028 A1 (“Moyne”) in view of Faigon et. al. Patent No. US 10,270,788 B2 (“Faigon ”) -and in view of Cretu et. al. , Pre-Grant Publication No. US 20080262985 A1 (“Cretu”)
As to Claim 1
Moyne teaches “training a first machine-learned model based on first training data that comprises first input data and first target data” (Moyne, “paragraph [0057], In some embodiments, predictive system 110 further includes server machine 170 and server machine 180. Server machine 170 includes a data set generator 172 that is capable of generating data sets (e.g., a set of data inputs and a set of target outputs) to train, [on first training data that comprises first input data and first target data] validate, and/or test one or more models, such as machine learning models. Model 190 may include one or more machine learning models[a first machine-learned model based], or may be other types of models such as statistical models. Models incorporating machine learning may be trained using input data and in some cases target output data.”).
2. generating, using the first machine-learned model, based on the first input data, first output data; “ (Moyne, “paragraph [0057], In some embodiments, predictive system 110 further includes server machine 170 and server machine 180. Server machine 170 includes a data set generator 172 that is capable of generating data sets (e.g., a set of data inputs and a set of target outputs) to train, validate, and/or test one or more models, such as machine learning models. Model 190 may include one or more machine learning models, [a first machine-learned model based] or may be other types of models such as statistical models .Models incorporating machine learning may be trained using input data [, based on the first input data,] and in some cases target output data.[ first output data;]”).
3. generating, using the first machine-learned model, based on second input data, second output data; “(Moyne, “Paragraph,[0061] Model 190 [using the first machine-learned model] may refer to the model artifact that is created [generating ]by the training engine182 using a training set that includes data inputs [based on second input data], and corresponding target outputs).[second output data (correct answers for respective training inputs).”).
4. if the difference is greater than a particular threshold: identifying, as an anomalous set, (Moyne “Paragraph [0101] As shown in FIG. 3B by the regions of different patterns, feature extraction may be used to determine the boundaries of particular shapes. Current trace data for analysis may be compared to historical trace data associated with previously produced substrates, trace data from other sensors, trace data from other recipe operations, or the like. In some embodiments, a threshold value can be established, and data values higher or lower than that threshold value. [if the difference is greater than a particular threshold may be identified as anomalous
[identifying, as an anomalous set]”).
5. said each data item, the corresponding data item in the second target data, and a corresponding data item in the second input data;(Moyne “Paragraph [0061] Model 190 may refer to the model artifact that is created by the training engine182 using a training set that includes data inputs[said each data item, corresponding data item in the second input data;] and corresponding target outputs [corresponding data item in the second target data], (correct answers for respective training inputs). Patterns in the data sets can be found that map the data input to the target output (the correct answer), and the machine learning model is provided mappings that captures these patterns.”).
6. wherein the method is performed by one or more computing devices.(Moyne “ Paragraph [0065], In some embodiments, the functions [wherein the method is performed] of client device 120, [one or more computing devices] predictive server 112, server machine 170, and server machine 180 may be provided by a fewer number of machines.” ).
Moyne does not explicitly teach :
for each data item in the second output data: generating a difference between said each data item and a corresponding data item in second target data that corresponds to the second output data;
generating second training data based on the first training data and the set of anomaly data; training a second machine-learned model based on the second training data;
adding the anomalous set to a set of anomaly data
based on one or more anomalous sets in the set of anomaly data and the second machine- learned model, computing a feature attribution value for each of one or more features of the second machine-learned model;
Faigon teaches “for each data item in the second output data: generating a difference between said each data item and a corresponding data item in second target data that corresponds to the second output data;” ( Faigon,“paragraph [62] ,At the second stage, anomaly detection engine 142 evaluates the values [for each data item in the second output data: in the “actual value” [said each data item corresponds to the second output data ] column against [generating a difference between] corresponding values in the “predicted value” column. [corresponding data item in second target data]
PNG
media_image1.png
926
700
media_image1.png
Greyscale
Examiner notes : Under BRI predicted value in the fig-6 is interpreted as second target data and actual value is interpreted as second output data.
generating second training data based on the first training data and the set of anomaly data; training a second machine-learned model based on the second training data; (Faigon, “paragraph [31] Further, the models 214 [training a second machine-learned model ] are updated with the detected anomalies by [and the set of anomaly data] the OMLwatcher /modeler 212 so that if the anomalous behavior becomes sufficiently frequent over time, then it can be detected as normal behavior rather than anomalous behavior. In one implementation, the anomaly events are transferred to the models 214 as “*.4train” files 206. [based on the second training data;]”).
Examiner notes: Under BRI model 214 is interpreted as second machine learned model which is trained using first training data set and generated anomaly data i.e. 4train” files 206, corresponding to the Claim.
based on one or more anomalous sets in the set of anomaly data and the second machine- learned model, computing a feature attribution value for each of one or more features of the second machine-learned model; (Faigon , “Paragraph [64], At the fourth stage, once a real suspected anomaly is identified (i.e., large relative-error, predicted value smaller than the observed/actual value and seasoned space ID),[ based on one or more anomalous sets in the set of anomaly data the second machine- learned model,] anomaly detection engine 142 evaluates the individual weights or so-called “likelihood coefficients” of the features of the event that caused the anomaly. This is done to identify [computing] so-called “smoking-gun features” or lowest likelihood coefficient feature-value pairs [value for each of one or more features of the] that have very low weights [a feature attribution value]compared to other features in the same event.”).
Faigon and Moyne are related to the same field of endeavor (i.e. Machine Learning). In view of the teachings of Faigon it would have been obvious for a person of ordinary skill in the art to apply the teachings of Faigon to Moyne before the effective filing date of the claimed invention in order to Improve efficiency and reduce compute cost by using anomaly- related training data to better guide model evaluation.( faigon, “[abs] It further includes determining an anomaly score for a production event based on calculated likelihood coefficients of categorized feature-value pairs and a prevalencist probability value of the production event comprising the coded features-value pairs.). Furthermore, a person of ordinary skill in the art would be motivated to apply the teachings of Faigon to Moyne to detect the anomalies faster in critical sectors (i.e. Healthcare, finance, security) and also learn which specific feature is causing large deviation.( faigon, “paragraph [14]Further, it relates to detecting anomalies in near real-time streams of security-related events of one or more tenants by transforming the events in categorized features and requiring a loss function analyzer to correlate, essentially through an origin, the categorized features with a target feature artificially labeled as a constant.” )
Cretu teaches adding the anomalous set to a set of anomaly data; (Cretu “Paragraph [0026 ]For example, FIG. 3 illustrates a method for sanitizing a local normal training data set 300 based on remote abnormal detection models 320 to generate a local normal sanitized anomaly detection model 360 and a local abnormal anomaly detection model 370. Data set 300 can be tested, at 330, against remote models 320. If a remote model of the models 320 indicates a hit on a data item (in this case, if a model 320 indicates a data item is anomalous), it can be added to anomalous data set 340.[adding the anomalous set to a set of anomaly data;]”).
Cretu and Moyne are related to the same field of endeavor (Anomaly Detection in ML). In view of the teachings of it Cretu would have been obvious for a person of ordinary skill in the art to apply the teachings of Cretu to Moyne before the effective filing date of the claimed invention in order to train the new model using the anomaly data to improve detection accuracy, reduce false negatives, and handle rare events more effectively than models trained solely on normal data. ( Cretu [0032] In some embodiments, if a remote model indicates that a local training data item is abnormal and/or a local normal model contains abnormal content, further testing can be performed.[ 0025], Various digital processing devices can share various abnormal, normal, and/or sanitized models and compare models to update at least one a local abnormal, normal, and/or sanitized model. ).”
As to claim 12
Moyne teaches” One or more non-transitory storage media storing instructions which, when executed by one or more computing devices, cause” (Moyne “ Paragraph [0005] In another aspect of the disclosure, a non-transitory machine-readable storage medium [non-transitory storage media] stores instructions which, [storing instructions which] when executed, . [when executed] cause a processing device[ by one or more computing devices, cause ] to perform operations.”).
And for all the other limitation of Claim 12, it is rejected under same basis as Claim 1. As the Claim are Analogous.
As to Claim 2 and Analogous Claim 13.
Moyne, as modified by Faigon and Cretu, teaches the method of claim 1.
Faigon further teaches “wherein generating the second training data comprises: for each anomalous set in the set of anomaly data.(Faigon,“paragraph [62] ,At the second stage, anomaly detection engine 142 evaluates the Values in the “actual value column against corresponding values in the “predicted value” column. [for each anomalous set in the set of anomaly data:]
PNG
media_image2.png
199
613
media_image2.png
Greyscale
[8]FIG. 7 shows a sample of anomaly output when the anomaly threshold is set to50× relative-error spike.
Examiner notes: Under BRI, generating the second data is interpreted as a sample of anomaly output in fig 6 corresponding to the claim.
2. computing a difference between (1) a data item from the second output data and (2) a corresponding data item in the second target data; including the difference in a training instance (Faigon,“paragraph [62] ,At the second stage, anomaly detection engine 142 evaluates the values [data item in the second output data]: in the “actual value column against [ Computing difference between; including the difference in a training instance ] corresponding values in the “predicted value” column. [corresponding data item in second target data].
3. adding the training instance to the second training data.( Faigon “paragraph [64] Online machine learner 122 is used to learning the normal patterns (training) and for anomaly detection (testing new data against know patterns). In one implementation, every chunk of incoming data is used twice: (1) to look for anomalies in it and (2) to update the so-called known or normal behavior models incrementally.[adding the training instance to the second training data.]”).
Moyne, Faigon, and Cretu are combinable for the same rationale as set forth above with respect to Claim 1
As to Claim 3 and Analogous Claim 14.
Moyne, as modified by Faigon and Cretu, teaches the method of claim 1.
Moyne further teaches “wherein generating the second training data comprises (Moyne , “paragraph [0058] Data set generator 172 may receive the output of a trained machine learning model (e.g., a model trained to perform a first operation of trace data processing), collect that data into training, validation, and testing data sets, and use the data sets to train a second machine learning model [wherein generating the second training data comprises] (e.g., a model to be trained to perform a second operation of trace data processing).
Examiner notes: Under BRI the data generated by 172 from one of the machine learned model to trained a second machine learned model is interpreted as generating the second training data corresponding to the claim.
for each data item in the first target data: computing a difference between (1) said each data item in the first target data and (2) a corresponding data item in the first output data; including the difference in a training instance; (Moyne, “Paragraph[0059] For example, a first trained machine learning model 190 that was trained using a first set of elements of the training set may be validated using the first set of elements of the validation set. The validation engine 184 may be capable of validating a trained machine learning model 190 using a corresponding set of elements of the validation set [for each data item in the first target data] from data set generator 172. For example, a first trained machine learning model 190 that was trained using a first set of elements of the training set may be validated using the first set of elements of the validation set. The validation engine 184 may determine an accuracy [computing a difference between; including the difference in a training instance] of each of the trained machine learning models 190 a [corresponding data item in the first output data]; based on the corresponding sets of elements of the validation set.
Examiner notes: Under BRI validation test is interpreted as consisting of first target data and model 190 which is trained on first set of elements is interpreted as first output and validation engine 184 is interpreted as computing the difference between item in first target data corresponding to the first input data including the difference in training instance.
Faigon teaches adding the training instance to the second training data.( Faigon “paragraph [64] Online machine learner 122 is used to learning the normal patterns (training) and for anomaly detection (testing new data against know patterns). In one implementation, every chunk of incoming data is used twice: (1) to look for anomalies in it and (2) to update the so-called known or normal behavior models incrementally.[adding the training instance to the second training data.]”).
Examiner notes: Under BRI update the behavior incrementally is interpreted as adding the training instance to the training data i.e corresponding to the claim.
Moyne, Faigon, and Cretu are combinable for the same rationale as set forth above with respect to Claim 1
Claims 4-11, 15-20 are rejected rejected under 35 U.S.C. 103 as being unpatentable over (“Moyne”) in view (“Faigon,”) and in view of (“Cretu”) and in further view of Das et.al, Pre-Grant Publication No. US 20220172004 A1(“Das”).
As to Claim 4
Moyne, as modified by Faigon and Cretu, teaches the method of claim 1.
Moyne, as modified by Faigon and Cretu, does not teach explicitly teach :
wherein the second machine-learned model is a regression model.
Das teaches “wherein the second machine-learned model is a regression model” (Das, “paragraph, [0036] “Machine learning training stage 130 may take the prepared training data (e.g., which may have passed, been mitigated, or at least understood using pre-training bias measurement 120), and perform various machine learning techniques to train a machine learning model. Various different types of machine learning models[wherein the second machine-learned model is a] may be implemented (e.g., neural networks, support vector machines, linear regression,[ regression model] decision trees, naïve Bayes, nearest neighbor, q-learning, temporal difference, deep adversarial networks, among others).”).
Das and Moyne are related to the same field of endeavor (Anomaly detection). In view of the teachings of Das it would have been obvious for a person of ordinary skill in the art to apply the teachings of Das to Moyne before the effective filing date of the claimed invention in order to use regression model to detect the anomalies early by monitoring Bias metrics and feature attribution. ( Das ,[abs] “Bias metrics and feature attribution may be monitored for a machine learning model. A request to enable monitoring for bias metrics or feature attribution may be received.” If a divergence from reference data is detected, then a notification indicating the divergence may be sent. Furthermore, one in ordinary skill in the art would be motivated in order to save cost while detecting fraud transaction in variety of domains (such as financial services, healthcare) faster and efficiently. (Das, “paragraph [0001] Machine-learned models and data-driven systems have been increasingly used to help make decisions in application domains such as financial services, healthcare, education, and human resources. These applications have provided benefits such as improved accuracy, increased productivity, and cost savings.").
As to Claim 5 and Analogous Claim 15
Moyne, as modified by Faigon and Cretu, teaches the method of claim 1.
Das teaches “wherein computing the feature attribution for each of the one or more features of the second machine-learned model comprises computing a feature attribution for each feature of the second machine-learned model.” ( Das, “paragraph [0051]. In various embodiments, computing the Shapley values [wherein computing the feature attribution] involves considering all possible coalitions of features [computing a feature attribution for each feature]. This means that given d features, there are 2d such possible feature coalitions [for each of the one or more features] each corresponding to a potential model [of the second machine-learned model] that needs to be trained and evaluated. ”).
Moyne, Faigon, , Cretu and Das are combinable for the same rationale as set forth above with respect to Claim 4.
As to Claim 6 and Analogous Claim 16
Moyne, as modified by Faigon and Cretu, teaches the method of claim 1.
Moyne, as modified by Faigon and Cretu does not teach:
wherein the first feature attribution value is a Shapley value
Das teaches “wherein the first feature attribution value is a Shapley value. ( Das, “paragraph [0051] In various embodiments, feature attribution may be determined using Shapley values.”).
Das and Moyne are related to the same field of endeavor (Anomaly detection). In view of the teachings of Das it would have been obvious for a person of ordinary skill in the art to apply the teachings of Das to Moyne before the effective filing date of the claimed invention in order to improve fairness, consistency, and provide theoretically grounded explanations of how machine learning models make predictions. (Das,Paragraph,”[0048]Taking the example of a college admission scenario, con-sider a model with features {SAT Score, GPA, Class Rank}, where it is desirable to explain the model prediction for a candidate. The range of model prediction is 0-1. The pre-diction for a candidate is 0.95. Then in the game, the total payoff would be 0.95 and the players would be the three individual features. If for the candidate, the Shapley values are {0.65, 0.7, -0.4}, then it may be determined that the GPA affects the prediction the most, followed by the SAT score. It may also be determined that while GPA and SAT score affect the prediction positively, the class rank affects it negatively (note that a lower rank is better).”).
As to Claim 7 and Analogous Claim 17
Moyne, as modified by Faigon and Cretu, teaches the method of claim 1.
Faigon Further teaches “wherein computing the feature attribution value is performed for a plurality of anomalous sets in the set of anomaly data” (Faigon , “Paragraph [64], At the fourth stage, once a real suspected anomaly is identified (i.e., large relative-error, predicted value smaller than the observed/actual value and seasoned space ID), [is performed for a plurality of anomalous sets in the set of anomaly data] anomaly detection engine 142 evaluates the individual weights or so-called “likelihood coefficients” of the features of the event that caused the anomaly. This is done to identify [wherein computing] so-called “smoking-gun features” or lowest likelihood coefficient feature-value pairs that have very low weights [the feature attribution value]compared to other features in the same event.”).
Das teaches wherein a plurality of feature attribution values are generated for a particular feature of the second machine-learned model,(Das “paragraph [0133], In some embodiments, local feature attribution values may be generated [wherein a plurality of feature attribution values are generated] in order to provide an explanation for a specific inference [for a particular feature of the] performed by the trained machine learning model. [second machine-learned model]”).
performing an aggregation operation on the plurality of feature attribution values to generate a global value for the particular feature. (Das, “paragraph [0052] Global explanations of learning may be provided, in some embodiments, according to measurements 160. For example, global explanation of an ML models by aggregating [aggregation operation] the Shapely values over multiple instances. .[the plurality of feature attribution values] [132]For example, SHAP values may be generated to provide a global feature attribution [to generate a global value for the particular feature.] for the trained machine learning model, which may be calculated using distributed techniques discussed below with regard to FIG. 14.”).
Moyne, Faigon, , Cretu and Das are combinable for the same rationale as set forth above with respect to Claim 4.
As to Claim 8 and Analogous Claim 18
Moyne, as modified by Faigon , Cretu and Das and teaches the method of claim 7.
Das further teaches “wherein the aggregation operation is a mean operation or a median operation.” (Das “paragraph [0052], Different ways of aggregation [aggregation operation] may be implemented, in various embodiments, such as mean of absolute SHAP values for all instances, “median”: median of SHAP values for all instances, and mean of squared SHAP values for all instances. [is a mean operation or a median operation]”).
Moyne, Faigon, , Cretu and Das are combinable for the same rationale as set forth above with respect to Claim 6.
As to Claim 9 and Analogous Claim 19
Moyne, as modified by Faigon and Cretu, teaches the method of claim 1.
Faigon teaches “wherein computing the feature attribution value is performed for a plurality of anomalous sets in the set of anomaly data,”(Faigon , “Paragraph [64], At the fourth stage, once a real suspected anomaly is identified (i.e., large relative-error, predicted value smaller than the observed/actual value and seasoned space ID), [is performed for a plurality of anomalous sets in the set of anomaly data] anomaly detection engine 142 evaluates the individual weights or so-called “likelihood coefficients” of the features of the event that caused the anomaly. This is done to identify [wherein computing] so-called “smoking-gun features” or lowest likelihood coefficient feature-value pairs that have very low weights [the feature attribution value]compared to other features in the same event.”).
Das teaches wherein a plurality of feature attribution values is generated for each feature of a plurality of features of the second machine-learned model, (Das, “paragraph [0051], In various embodiments, computing the Shapley values involves [wherein a plurality of feature attribution values is generated] considering all possible coalitions of features. This means that given d features, there are 2d [for each feature of a plurality of features ] such possible feature coalitions, each corresponding to a potential model [second machine-learned model ] that needs to be trained and evaluated. “).
2. the method further comprising: for each feature of the plurality of features, performing an aggregation operation on the plurality of feature attribution values (Das, “Paragraph [0052]Global explanations of machine learning models may be provided, in some embodiments, according to feature attribution measurements 160. For example, global explanation of an ML model by aggregating the Shapley values : [for each feature of the plurality of features] over multiple instances. Different ways of aggregation [the method further comprising: performing an aggregation operation ]may be implemented, in various embodiments, such as mean of absolute SHAP values for all instances, "median": median of SHAP values for all instances, and mean of squared SHAP values for all instances.).”
that correspond to said each feature to generate a global value for said each feature.(Das, paragraph [132] For example, SHAP values [that correspond to said each feature]may be generated to provide a global feature attribution [to generate a global value for said each feature] for the trained machine learning model, which may be calculated using distributed techniques discussed below with regard to FIG. 14. In some embodiments, the feature attribution may be calculated using a specified aggregation techniques (e.g., mean, mean squared, or median value).”).
that correspond to said each feature to generate a global value for said each feature.(Das, paragraph [132] For example, SHAP values [that correspond to said each feature]may be generated to provide a global feature attribution [to generate a global value for said each feature] for the trained machine learning model, which may be calculated using distributed techniques discussed below with regard to FIG. 14. In some embodiments, the feature attribution may be calculated using a specified aggregation techniques (e.g., mean, mean squared, or median value).”).
Moyne, Faigon, , Cretu and Das are combinable for the same rationale as set forth above with respect to Claim 6.
As to Claim 10
Moyne, as modified by Faigon ,Cretu and Das teaches the method of claim 9.
Das teaches” further comprising: ranking the plurality of features based on the global value for each feature in the plurality of features”. (Das “paragraph ,0060,0061,0062,0063, “[0060] F=[f1, ... , fm] may be the list of features sorted respect to their attribution scores in the training data where m is the total number of features. [0061] a(f) may be a function that returns the feature attribution score on the training data given a feature f. [0062] F'= [f', ... , f' m] may be the list of features sorted [ranking the plurality of features ]with with respect to their attribution scores in the live data where m is the total number of features. [global value for each feature in the plurality of features]”).
Moyne, Faigon, , Cretu and Das are combinable for the same rationale as set forth above with respect to Claim 6.
As to Claim 11 and Analogous Claim 20
Moyne, as modified by Faigon ,Cretu and Das teaches the method of claim 9.
Das Further teaches “the causing, to be displayed, on a screen of a computing device, a name for each feature of the plurality of features and the global value for each feature in the plurality of features” (Das “paragraph, 0121, 0122, 0145 [0121], User interface elements to configure the view, such as monitoring job view properties 1032 may be implemented which may allow for subsets of bias metrics for feature data to be displayed [the causing, to be displayed] (e.g., age range of 20 to 50, gender=female, etc.).[ a name for each feature of the plurality of features] . [0122] Fairness and explainability monitoring 920 may obtain various reference feature attributions, from training reports or past measurements computed by feature attribution measurement 940 (which may perform global feature attribution measurement [ and the global value for each feature in the plurality of features]. according to the techniques discussed above with regard to FIGS. 1 and 3, such as by using SHAP values and generating comparisons using NDCG).”). [0145] Display(s) 2080 may include
standard computer monitor(s) and/or other display systems, [on a screen of a computing device] technologies or devices.
Moyne, Faigon, , Cretu and Das are combinable for the same rationale as set forth above with respect to Claim 6.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NIROJ KOIRALA whose telephone number is (571)270-0748. The examiner can normally be reached Monday -Friday 8am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, MICHAEL HUNTLEY can be reached on (303) 297-4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/N.K./Examiner, Art Unit 2129
/MICHAEL J HUNTLEY/Supervisory Patent Examiner, Art Unit 2129