DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This office action is in responsive to communication(s): original application filed on 01/04/2024, said application claims a priority filing date of 07/26/2021. Claims 1-7 are pending. Claim 1 is independent.
Drawings
The drawings are objected to under 37 CFR 1.83(a) because they fail to show details of S1 to S7 in FIG.1 as described in the specification. Any structural detail that is essential for a proper understanding of the disclosed invention should be shown in the drawing. MPEP § 608.02(d). Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as “amended.” If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance. In addition to Replacement Sheets containing the corrected drawing figure(s), applicant is required to submit a marked-up copy of each Replacement Sheet including annotations indicating the changes made to the previous version. The marked-up copy must be clearly labeled as “Annotated Sheets” and must be presented in the amendment or remarks section that explains the change(s) to the drawings. See 37 CFR 1.121(d)(1). Failure to timely submit the proposed drawing and marked-up copy will result in the abandonment of the application.
The drawings are objected to as failing to comply with 37 CFR 1.84(p)(4) because reference character “S4” has been used to designate both "The user 30 can decide to maintain the output label 22 or to process the output label 22" in ¶ [0044] and "the at least one additional machine learning model 40 is trained on the input data Z received from the user" in ¶ [0045]. Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
The drawings are objected to because (1) it is unclear that "the at least one additional machine learning model 40 is trained on the input data Z received from the user" in ¶ [0045] is for S4, S5, or both; (2) it is unclear what processes are involved in S7 of FIG. 1 because S7 only appear once in ¶ [0037] that "Figure 1 illustrates a flowchart of the method according to the embodiment of the invention with the method steps S1 to S7". Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as “amended.” If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Specification
The abstract of the disclosure is objected to because it appears that some words are missing at the end of "d." and before "e." (e.g., "… d. receiving at least one processed output label, at least one additional data item or at least one additional output label from the ??; e. training at least one additional machine learning model for anomaly detection …"). A corrected abstract of the disclosure is required and must be presented on a separate sheet, apart from any other text. See MPEP § 608.01(b).
The disclosure is objected to because of the following informalities:
in ¶ [0038], "Fig. 2 shows a schematic representation of the technical system according to an embodiment." appears to be "Fig. 2 shows a schematic representation of the technical system 1 according to an embodiment."
Appropriate correction is required.
Claim Objections
Claims 1 and 4-7 are objected to because of the following informalities:
in Claim 1, lines 3-5, "… providing a trained unsupervised machine learning model for anomaly detection; wherein the unsupervised machine learning model is trained on the basis of unlabeled training data …" appears to be "… providing a trained unsupervised machine learning model for anomaly detection; wherein the unsupervised machine learning model is trained on unlabeled training data …";
in Claim 1, lines 16-17, "… or maintaining the at least one determined output label unprocessed depending on the verification by the user …" appears to be "… or maintaining the at least one determined output label unprocessed depending on verification by the user …";
in Claim 4, lines 2-3, "… wherein the connection function is a logical function or logical operator, an AND and OR logical operator" appears to be "… wherein the connection function is a logical function or a logical operator with an AND logical operator and an OR logical operator";
in Claim 6, lines 2-5, "… a computer readable hardware storage device having computer readable program code stored therein, said program code executable by a processor of a computer system to implement a method for performing the steps according to claim 1 …" appears to be "… a computer readable hardware storage device having computer readable program code stored therein, said computer readable program code executable by a processor of a computer system to implement the method according to claim 1 …" (see also 112 rejection to Claim 6);
in Claim 7, lines 1-2, "A technical system, configured for performing the steps according to claim 1" appears to be "A technical system, configured for performing the method according to claim 1" (see also 101 rejection to Claim 7).
Appropriate correction is required.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: "technical system" in Claim 7.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure (i.e., FIG. 2 shows that "technical system 1" includes "combined machine learning model 10" which further includes "trained unsupervised machine learning model 20" and "additional machine learning model 40") described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-7 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 1 recites the limitation "… providing a trained unsupervised machine learning model for anomaly detection … training at least one additional machine learning model for anomaly detection … generating the combined machine learning model for anomaly detection …" in line 3-23, which rendering the claim indefinite because it is unclear whether these three instances of "anomaly detection" are the same or different. Clarification is required.
Claims 2-7 are rejected for fully incorporating the deficiency of their respective base claims.
Claim 2 recites the limitation "… wherein the processing comprises at least one processing step, selected from the group comprising …" in lines 2-3. There is insufficient antecedent basis for the limitations "the processing" and "the group" in the claim. Clarification is required.
Claim 6 recites the limitation "A computer program product, comprising a computer readable hardware storage device having computer readable program code stored therein, said program code executable by a processor of a computer system to implement a method for performing the steps according to claim 1 when the computer program product is running on a computer" in lines 1-6, which rendering the claim indefinite because it is unclear whether "a computer system" and "a computer" recited in the claim are same, different, or related (see also Claim Objections to Claim 6). Clarification is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim 7 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claim(s) does/do not fall within at least one of the four categories of patent eligible subject matter because "technical system" is claimed without reciting its structure and according to FIG. 2 with REFERENCE SIGNS in the specification, "technical system 1" includes "combined machine learning model 10" which further includes "trained unsupervised machine learning model 20" and "additional machine learning model 40", where the "combined machine learning model 10", "trained unsupervised machine learning model 20", and "additional machine learning model 40" can be software per se. and are not one of the four categories of patent eligible subject matter.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-3 and 5-7 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Givental et al. (US 2021/0279644 A1, filed on 03/06/2020), hereinafter Givental'644.
Independent Claim 1
Givental'644 discloses a computer-implemented method for generating a combined machine learning model (Givental'644, ¶ [0019]: leverage the strengths of both supervised and unsupervised machine learning in a hybrid approach to training a machine learning model to detect anomalies in computer system log data structures; combine unsupervised and supervised machine learning in a specific manner to provide a computer machine learning model that is able to detect anomalies in log data structures; ¶ [0024]: provide a mechanism that combines unsupervised and supervised machine learning to provide accurate anomaly detection in computer system log data structures; ¶ [0071] with 308 in FIG. 3: the weighted outputs of the individual unsupervised ML models are then combined using a combinatorial function, e.g., a weighted function which applies weights to each of the outputs of each of the unsupervised ML models in the ensemble, to generate a single ensemble output for the ensemble (step 308)), comprising:
a. providing a trained unsupervised machine learning model for anomaly detection; wherein the unsupervised machine learning model is trained on the basis of unlabeled training data (Givental'644, ¶ [0005] and ABSTRACT: implement an ensemble of unsupervised machine learning (ML) models; ¶¶ [0017]-[0020] and [0031]: the process of training a machine learning (ML) model involves providing an ML algorithm (that is, the learning algorithm) with training data to learn from; unsupervised machine learning algorithms do not require manual effort on labeling the historical computer system log data; an ensemble of a plurality of machine learning models trained using unsupervised machine learning algorithms, e.g., isolation forest, local outlier factor, one-class support vector machine (SVM), and/or the like, are generated with dynamic weighting; machine learning algorithms build a computer executed model based on sample data, known as "training data", in order to make predictions or decisions without being explicitly programmed to perform the task; ¶ [0024]: reduce the manual effort spent on labeling data for initial training of machine learning models by allowing the ensemble of unsupervised machine learning models to operate on unlabeled data; ¶¶ [0044]-[0048] with FIG. 1: the log data structures 104 obtained from the monitoring agents monitoring computing resources in the monitored computing environment 102 are first processed by the data cleaning and feature engineering engine 110 so as to convert log data of different formats into a predetermined common format that is able to be used as input to the unsupervised machine learning model ensemble 120; the data cleaning and feature engineering engine 110 provides a set of encoded features extracted from the log data structures after conversion to a common format to a training data database 160 for inclusion in a partially labeled training dataset 162 as unlabeled training data; an ensemble 120 of a plurality of machine learning models 122-128 are trained using unsupervised machine learning algorithms; the resulting machine learning models, e.g., isolation forest ML model 122, one-class SMV ML model 124, local outlier factor ML model 126, and other ML models 128, or a subset of the machine learning models 122-128, such as a "top X models" where X is any desired integer value; ¶ [0070] with 302 in FIG. 3: creating an initial ensemble of unsupervised ML models from a plurality of available unsupervised ML models; this initial selection of unsupervised ML models may be performed by randomly, based on user selections, or any other suitable methodology);
b. determining at least one output label for each data item of a plurality of data items of unlabeled application data by applying the trained unsupervised machine learning model on the unlabeled application data with the plurality of data items; wherein the at least one determined output label is an anomaly or a normal state (Givental'644, ¶ [0005] and ABSTRACT: processing, by the ensemble of unsupervised ML models, a portion of input data to generate an ensemble output; ¶¶ [0019]-[0021]: the log data structure is first parsed and processed into structured data of a predetermined format; this parsing and converting to a predetermined format may be done with regard to log data structures having different native formats and the pre-determined format may provide a common or universal format which is usable by downstream processes, such as an ensemble of machine learning models; the ensemble of unsupervised machine learning models is used on input data, such as the log data structure information stored in the pre-determined format data structure, to generate an initial output indicating a likelihood that an event represented in the log data is actually anomalous; based on the initial output generated by the ensemble specifying a likelihood, or probability, that an event in the log data is anomalous; ¶¶ [0041]-[0049] with FIG. 1: detecting anomalies in computing system or computing resource security event log data; a monitored computing system environment 102 comprises a plurality of computing system resources, such as computing devices, software executing on computing devices, data storage systems, data network components, e.g., switches, routers, etc., and any other known or later developed computing system resources from which and/or about which log data is generated, such as security event logs that may store access log data; the computer resource log data 104 is transmitted, such as via one or more data networks (not shown), to the hybrid machine learning (ML) anomaly detector 100; depending on the particular computing resources being monitored, the particular monitoring agents being employed, and the particular events being logged, the log data structures 104 may have differently formatted logged data; a common, "normalized" log format is utilized in order to forego the variety of data formats across vendor-specific security devices; a normalized log allows any downstream system to know how to read attributes such as IPs, event names, ports, etc.; the data cleaning and feature engineering engine 110 parses the log data and converts the log data into a commonly structured data frame; feature engineering is performed on the converted log data to drop useless features, e.g., implementing a "dropna" algorithm or the like, and extract new features to supplement the existing feature set present in the converted log data so as to represent the individual event logs; missing data imputation may also be implemented, such as a "fillna" algorithm or the like, to impute missing data in the data frame; additional data cleaning and feature engineering operations may be performed by the data cleaning and feature engineering engine 110 so as to format and extract features from the log data 104 for input to the unsupervised machine learning models 122-128 of the unsupervised machine learning model ensemble 120; the data cleaning and feature engineering engine 110 combines all required function in order to take a vendor-specific payload in a vendor-specified format and parse that down to a set of normalized fields, adding new fields (such as Geolocation), and splitting all attributes of the log into key/value pairs, so that they can be further extracted for use in the machine learning platform; the input provided to the ensemble 120 comprises a set of encoded features extracted from the log data structures after conversion to a common format; the data cleaning and feature engineering engine 110 provides this input to the ensemble 120 of unsupervised machine learning models 122-128 as input; ML models 122-126 may be selected, each operate on input data from the data cleaning and feature engineering engine 110 to generate an output result; initially, a subset of ML models may be randomly selected, manually selected by a security analyst, or the like, as an initial subset of ML models; ¶ [0056] with FIG. 1: the anomaly score 140 is used to label the log input data as to the particular classification for the log input data, e.g., "anomaly" or "normal"; an anomaly score 140 equal to or above an anomaly threshold value may indicate that the log input data represents an anomaly; an anomaly score 140 below the anomaly threshold may be indicative of log input data that is not anomalous, i.e. is "normal"; ¶¶ [0065]-[0067] with 202-212 in FIG. 2: the operation starts by having the data cleaning and feature engineering engine query the unstructured logs from various data sources of the monitored computing environment, or other computing system collecting such logs from the monitored computing environment; the unstructured logs are parsed and converted into structured log data frames; feature engineering of the data is performed so as to drop unnecessary features and extract new features to supplement existing features to represent the individual logs; the whole purpose of feature engineering is to establish new variables/features that can explain the observed behavior (target variable) in a better way; missing data imputation is performed, such as using a "fillna" algorithm or the like, to impute missing data in the data frame, e.g., geolocation, customer industry, threat category, etc.; feature encoding is performed, such as by using a label encoder algorithm to convert categorical variables into numerical values, and a one-hot encoding of those features whose values are lists and not single categorical values; the encoded log data is input to an ensemble of unsupervised machine learning models which process the encoded log data to generate output anomaly scores that are then weighted and combined to generate an anomaly score for the encoded log data; ¶¶ [0070]-[0071] with 304-308 in FIG. 3: input data is then processed via each of the unsupervised ML models in the ensemble and individual weights are applied to each of the outputs of the unsupervised ML models in the ensemble; the weights may also be set to an initial weight value, which may initially be all the same weight for each of the unsupervised ML models, or may be different based on an estimate as to which unsupervised ML models are more likely than others to provide correct outputs, e.g., based on user input indicating a preference of certain unsupervised ML models over others; the weighted outputs of the individual unsupervised ML models are then combined using a combinatorial function, e.g., a weighted function which applies weights to each of the outputs of each of the unsupervised ML models in the ensemble, to generate a single ensemble output for the ensemble);
c. transmitting the at least one determined output label to a user for verifying the at least one determined output label via a user interface (Givental'644, ¶ [0005] and ABSTRACT: outputting the ensemble output to an authorized user computing device to obtain user feedback from the authorized user via the user computing device; ¶¶ [0021]-[0022]: the events in the log data that have a high likelihood or probability (as measured by comparing the likelihood output of the ensemble to one or more predetermined threshold values) are provided to a human security analyst via a user interface output on a computing device associated with the security analyst; the newly detected log data having generated labels with high likelihood of representing anomalies are again sent to a human system analyst for review and responsive action; ¶¶ [0052], [0055]-[0057], and [0063]-[0064] with FIG.1: instances of outputs of the unsupervised machine learning models indicating potential security threats/attacks, unauthorized accesses, or the like, may be sent to a human security analyst, such as via a graphical user interface and workstation or other computing device associated with the human security analyst; such "normal" or "non-anomalous" outputs by the unsupervised machine learning models may be sent to the human security analyst only in instances where the confidence or probability associated with the classification is below a predetermined threshold, such that not all indications of "normal" or "non-anomalous" need be reviewed; the same is true for the "anomaly" or "anomalous" classifications, i.e. a threshold for confidence or probability value may be established and if the confidence/probability for the classification of "anomaly" or "anomalous" is equal to or above the threshold, then the human security analyst may be enlisted to provide feedback information; the anomaly score 140 generated by combining the weighted outputs for the various unsupervised machine learning models 122-128 in the ensemble 120 is used to identify input log data that corresponds with anomaly scores 140 indicative of a need for human security analyst review, e.g., logs associated with anomaly scores 140 equal to or above a threshold value; alternatively, a predetermined percentage of log data and corresponding anomaly scores 140 may be selected and provided to the security analyst via the workstation 155 for review; an additional threshold value may be used to select a subset of the log data inputs indicative of an anomaly that are to be reviewed by the security analyst and also provided as labeled data for inclusion in the partially labeled data 162; e.g., a selection threshold value may be set such that a relatively small subset of the log data having high anomaly scores 140, i.e. anomaly scores 140 equal to or above the anomaly threshold value, is selected for obtaining feedback information and for inclusion in the partially labeled data 162, e.g., a small percentage, e.g., approximately 2%, of the log data; the labeled log data 150, or the selected subset portion thereof, are used to provide the output to the workstation 155 for review by the security analyst 155; for those events indicating anomalies, corresponding logged anomalies are set to the STEM computing system 190 for further evaluation and responsive action; the threshold may be set such that only a small percentage of the logs and corresponding labels are sent to the security analyst workstation 155 for review; e.g., approximately only 2% of the log data and its corresponding labels are sent to the security analyst workstation 155 for review by the security analyst; ¶ [0068] with 214 in FIG. 2: a subset of the log data having anomaly scores equal to or greater than an anomaly threshold value, is selected for user feedback and inclusion in the partially labeled dataset; ¶¶ [0071]-[0072] with 310-312 in FIG. 3: a determination is then made as to whether the ensemble output indicates a need for user feedback (step 310); this determination may be, e.g., based on a comparison of the ensemble output to one or more criteria, e.g., threshold values or the like; if the one or more criteria are satisfied by the ensemble output, i.e. there is a determined need for user feedback, then the operation may continue with steps 312-324; if the one or more criteria are not satisfied, i.e. user feedback is determined to not be needed, the operation may skip steps 312-324 and continue at step 326; in response to determining that there is a need for user feedback, the ensemble output is output to a user workstation, such as a computing device associated with an authorized user, e.g., security analyst or the like (step 312));
d. receiving least one processed output label, at least one additional data item or at least one additional output label from the user via the user interface or maintaining the at least one determined output label unprocessed depending on the verification by the user (Givental'644, ¶ [0005] and ABSTRACT: the user feedback indicates a correctness of the ensemble output; ¶¶ [0021]-[0022]: so that the security analyst may review and label the input log data as either anomalous or not anomalous, e.g., "true threat" or "false positive; the newly detected log data having generated labels with high likelihood of representing anomalies are again sent to a human system analyst for review and responsive action; ¶ [0025]: reduce the manual efforts on identifying false positives generated by the ensemble of unsupervised machine learning models in that the analyst is only reviewing results the first time so as to provide labels to the unsupervised machine learning model, with subsequent results not needing to be reviewed by the human analyst; the potential for additional reduction in false positives comes from the fact that once the system is in production, the analysts' reviews both teach the system as well as earn the trust of the analysts for how the system is labelling "false positives"; as that trust is built, the analysts no longer have to examine every log and high-confidence anomalies can be automatically actioned as noise or true threats; ¶¶ [0043], [0052]-[0053], [0057], [0059], and [0063]-[0064] with FIG. 1: the security analyst computing device 155 may be integrated with the STEM computer system 190, or may be used to access the STEM computer system 190 to provide security analyst feedback; the human security analyst may verify whether the input is actually a true threat or false positive; the human security analyst may then provide an input via the graphical user interface to indicate their assessment of the input and thus, whether the corresponding unsupervised machine learning model(s) were able to generate the correct result; this feedback information may be used to set or adjust weights associated with the outputs of the various unsupervised machine learning models 122-128 of the ensemble 120; the human security analyst may also be provided with outputs indicating that the input to the unsupervised machine learning model is a "normal" or "non-anomalous" input so that the human security analyst may provide such feedback information as well, but this time with regard to whether the input is or is not a threat and thus indicate whether the output of the unsupervised machine learning model(s) represent a true non-threat or a false negative result; feedback information from the security analyst workstation 155 as to the correctness/incorrectness of the anomaly score 140 (or in other embodiments, the individual outputs of the unsupervised machine learning models 122-128 prior to combining them to generate the anomaly score 140); the security analyst 155 who can then either indicate agreement or non-agreement with the labels; if the security analyst reviews a high anomaly score event or log entry in the selected subset of labeled log data 150, and disagrees with the label generated as a result of the initial anomaly score 140, then the security analyst may indicate this disagreement and/or a correct label, and a corresponding default anomaly score will be associated with the event or log entry, e.g., if the security analyst indicates that the high anomaly score is a false positive, then the anomaly score may be changed to a zero anomaly score value which is associated with the correct label of "non-anomalous" in the selected subset of labeled log data 150 that is stored in the partially labeled data 162; significantly reduce the amount of manual effort spent on labeling data for training machine learning models; the only manual input required is the security analyst review of a selected subset of log data and their labels whose anomaly score 140 meets or exceeds a particular threshold value; reduce the manual effort for identifying false positives and provides a high detection accuracy that improves security coverage; ¶ [0068] with 216 in FIG. 2: user feedback information specifying correct labels for the subset of the log data is obtained and used to update the labels of this subset before inclusion of the subset as labeled data in the partially labeled dataset; ¶ [0072] with 314 in FIG. 3: user feedback is then received from the user workstation (step 314); the user feedback indicates a correct output that the ensemble should have generated; e.g., the user feedback may indicate agreement with the ensemble output or disagreement with the ensemble output and an indication of what the correct output should have been according to the user's evaluation);
e. training at least one additional machine learning model for anomaly detection in accordance with the at least one processed output label, the at least one additional data item or the at least one additional output label; f. generating the combined machine learning model for anomaly detection using a connection function based on the trained unsupervised machine learning model and the at least one trained additional machine learning model (Givental'644, ¶ [0005] and ABSTRACT: the method further comprises modifying at least one feature of the ensemble of unsupervised ML models based on the obtained user feedback to thereby generate a modified ensemble of unsupervised ML models; ¶ [0019]: leverage the strengths of both supervised and unsupervised machine learning in a hybrid approach to training a machine learning model to detect anomalies in computer system log data structures; combine unsupervised and supervised machine learning in a specific manner to provide a computer machine learning model that is able to detect anomalies in log data structures; ¶¶ [0021]-[0026]: the human security analyst response is stored in a training data database along with the unlabeled input log data for use as training data for training an anomaly detection machine learning model, i.e. performing supervised machine learning of an anomaly detection machine learning model; thus, the training data in the training data database comprises a hybrid of unlabeled log data and labeled log data, i.e. this hybrid is a partially labeled data set; the partially labeled data set is input to a semi-supervised machine learning model so that the semi-supervised machine learning model labels the unlabeled portion of the partially labeled data set based on similarities of the unlabeled data with labeled data in the partially labeled data set; in addition to the partially labeled data set being sent to the semi-supervised machine learning model, the feedback from the human security analysts with regard to the results of the ensemble unsupervised machine learning model labeling of the training data, i.e. the responses provided by the human security analysts in response to the user interface outputs generated for high likelihood anomalous log data, are provided to a dynamic weight generator to determine which unsupervised machine learning model(s) in the ensemble provided the most accurate outputs, e.g., most accurate classifications of logged events as "true threat"; the dynamic weight generator executes a weight generator supervised machine learning model, e.g., SVM, neural networks, etc., to assign weights to the top N number of machine learning models in the ensemble, where N is any desirable number of machine learning models for the particular implementation, e.g., top 3 machine learning models, where the "top" number refers to the models with a predetermined specified level of performance, e.g., the highest accuracy relative to all the models in the ensemble; provide a mechanism that combines unsupervised and supervised machine learning to provide accurate anomaly detection in computer system log data structures; results from the ensemble of unsupervised machine learning models are used to train a semi-supervised machine learning model to generate a classification output, and the feedback from the human analysts being used to generate dynamic weights for the machine learning models in the ensemble so as to make their classification outputs used by the semi-supervised machine learning model more accurate; ¶ [0047] with FIG. 1: the partially labeled training dataset 162 is used to train the semi-supervised machine learning model 170; ¶¶ [0049]-[0050] with FIG. 1: the evaluation of the ML models for selection in the subset of ML models may be performed using any suitable performance evaluation for the particular implementation; e.g., metrics may be maintained for each of the ML models with regard to their accuracy and/or precision in predicting classifications or labels for the inputs, as well as other performance metrics, and these metrics may be used as a basis for selecting the X number of ML models, e.g., X ML models having a relatively highest accuracy amongst the plurality of ML models 122-128; the accumulated performance metrics may be maintained by the hybrid machine learning (ML) anomaly detector 100, such as in the metrics storage and ML model selection engine 129, which also comprises the computer implemented logic for dynamically and automatically selecting a subset of ML models for inclusion in the ensemble 120; with regard to dynamically and automatically selecting a subset of ML models to include in the ensemble of unsupervised ML models 120, the metrics storage and ML model selection engine 129 may determine when the performance of an unsupervised ML model, e.g., unsupervised ML model 126, in the ensemble 120 is providing poor performance and may automatically replace this unsupervised ML model 126 with another unsupervised ML model, e.g., unsupervised ML model 128, in the available plurality of unsupervised ML models 122-128; e.g., threshold performance metric values may be established indicating an acceptable level of performance by the unsupervised ML models with regard to one or more performance metrics, e.g., accuracy, precision, etc.; if an unsupervised ML model in the ensemble 120 has an accumulated performance metric that equals or falls below the predetermined threshold value, then the unsupervised ML model may be removed from the ensemble 120 and/or replaced by another available unsupervised ML model; the modification to the ensemble 120 may also be communicated to the dynamic weights generator 130 so that corresponding weight values for the removed/replaced unsupervised ML model may be updated to reflect the removal/replacement; the weight value associated with the removed unsupervised ML model may be removed from the set of weights applied to the outputs of the unsupervised ML models in the ensemble 120, and if the unsupervised ML model is replaced, a corresponding weight for the replacement unsupervised ML model may be instigated in the set of weights, or a default weight value may be set for the replacement unsupervised ML model; ¶¶ [0053]-[0062] with FIG. 1: the outputs of the various unsupervised machine learning models 122-128 may be output to a dynamic weights generator 130 which applies weights 132-136 to the outputs prior to combining the results of the unsupervised machine learning models 122-128 to generate the anomaly score 140; the weights 132-136 are dynamically determined based on feedback information obtained from security analyst review of the outputs of the individual unsupervised machine learning models 122-128, or the final anomaly score 140 generated by a combination of the weighted outputs for the unsupervised machine learning models 122-128; the weights W1-W3 132-136 may be different from each other due to dynamic modifications of the individual weights 132-136 based on the dynamic weight generator 130 processing feedback information; in combining the weights to generate an initial anomaly score 140, any suitable function for the particular implementation may be used to combine the weighted anomaly scores from the individual unsupervised ML models 122-126 into a single anomaly score for the particular log entry/event that was evaluated by the ensemble 120; e.g., the function may be a sum of the individual unsupervised ML models 122-126 scores weighted by the corresponding weights 132-136; in other implementations, an average of the weighted anomaly scores may be utilized; in still other implementations, other functions involving each of the weighted anomaly scores generated by the individual unsupervised ML models 122-126 included in the ensemble 120 may be utilized; this feedback information is fed back into the dynamic weights generator 130 which determines how to modify the weights W1-W3 132-136 to increase the correctness of the anomaly scores 140 and corresponding labels generated by the ensemble 120; i.e., some models may operate better on different types of patterns of input data, e.g., different patterns in security log data. As a result, weights may need to be dynamically adjusted based on the patterns of log data input to the ensemble 120; the dynamic weights generator 130 dynamically adjusts the weights 132-136 by receiving the feedback information as to correctness, determining which unsupervised machine learning models 122-128 generated the correct output indicated in the feedback information and which generate the incorrect output; the weights of the unsupervised machine learning models that generate the correct output may be increased whereas the weights of the unsupervised machine learning models may be decreased; the amount of increase/decrease may be determined based on a desired function for the particular implementation; the security analyst's agreement or non-agreement with the anomaly score 140 output may be used to update performance metrics associated with the unsupervised ML models 122-128 that are part of the ensemble of unsupervised ML models 120 as stored in the metrics storage and ML model selection engine 129; those unsupervised ML models in the ensemble 120 that generated a correct output as indicated by the user feedback may have their performance metric(s) increased to represent that these unsupervised ML models are generating correct results; those unsupervised ML models in the ensemble 120 that generated an incorrect output as indicated by the user feedback may have their performance metric(s) decreased to represent that these unsupervised ML models are generating incorrect results; the amount of the increase/decrease may be a function of the amount of certainty the corresponding unsupervised ML model had in the output it generated, e.g., more certainty in an incorrect output may result in a larger decrease in the performance metric(s) and more certainty in a correct output may result in a larger increase in the performance metric(s); the labeled log data 150 is also stored as part of a partially labeled dataset 162 in a training dataset data storage 160; the labeled log data 150 that is stored in the partially labeled data 162 will include the correct label for each event or log entry in the labeled log data 150, as indicated either by the ensemble 120 generated initial anomaly score which is approved by the security analyst, or by the security analyst in the user feedback provided by the security analyst, as well as the anomaly score corresponding to the correct label, e.g., the initial anomaly score or an alternative anomaly score generated as a result of the user feedback; the partially labeled dataset 162 is input to a semi-supervised machine learning model 170 so that the semi-supervised machine learning model 170 labels the unlabeled portion of the partially labeled dataset 162 based on similarities of the unlabeled data with labeled data in the partially labeled data; this similarity measurement analysis may be performed for each pairing of unlabeled entry and labeled entry in the partially labeled dataset 162; based on the similarities of the unlabeled entry to the labeled entries, label propagation and label spreading operations are performed such that a labeled entry having a highest similarity measure is selected for the unlabeled entry and the corresponding label of the selected labeled entry is attributed to the unlabeled entry, e.g., log entry in the log data; this label propagation and label spreading may be performed with regard to each unlabeled event or log entry in the partially labeled data 162 until all of the events or log entries have a corresponding label; the semi-supervised machine learning model 170 also updates the anomaly scores for the partially labeled data 162 at least by generating new anomaly scores for the unlabeled events/log entries in the partially labeled data 162 based on the similarity analysis, to generate final anomaly scores 180 for each of the events/log entries in the log data 110; ¶¶ [0067]-[0068] with 220-222 in FIG. 2: the weights applied to the outputs of the various models are dynamically determined based on feedback information; any suitable function for combining the weighted outputs from the models may be used without departing from the spirit and scope of the present invention, e.g., simply sum, averaging of the weighted anomaly scores from the models, selection of a highest weighted anomaly score from the models, etc.; he feedback information is fed back into a dynamic weight generator which updates the weights applied to the outputs of the various unsupervised machine learning models of the ensemble based on whether or not they output a correct result; the partially labeled dataset, combining the subset of correctly labeled anomalous log data, and the encoded but unlabeled data from the data cleaning and feature engineering engine, is input to a semi-supervised machine learning model which labels unlabeled data based on similarities with the labeled data in the partially labeled dataset; ¶¶ [0069]-[0071] and [0073]-[0075] with 316-322 FIG. 3: with regard to the selection of unsupervised ML models to include in the ensemble and the adjustment of weights applied to the outputs of the selected unsupervised ML models when combining the outputs to generate a single ensemble output; these weight values may then be dynamically and automatically updated based on user feedback; the weighted outputs of the individual unsupervised ML models are then combined using a combinatorial function, e.g., a weighted function which applies weights to each of the outputs of each of the unsupervised ML models in the ensemble, to generate a single ensemble output for the ensemble; the user feedback is provided to a dynamic weight generator which updates the individual weights based on the user feedback; based on whether or not the corresponding unsupervised ML model generated a correct output individually, the unsupervised ML model's weight may be increased/decreased to give preference to unsupervised ML models that generate correct outputs and to not give preference to unsupervised ML models that generate incorrect outputs; in addition, the user feedback is provided to a metrics storage and ML model selection engine to thereby update performance metrics for the individual unsupervised ML models that are part of the ensemble; based on the updated performance metrics, a determination is made as to whether a change in the ensemble is to be performed; this determination may be made based on the performance metrics of unsupervised ML models in the ensemble satisfying one or more criteria for modifying the ensemble, e.g., performance metric(s) of an unsupervised ML model fall to or below a threshold performance value; in response to a determination that a change in the ensemble is to be performed, membership of the unsupervised ML models in the ensemble is automatically modified; this modification of the ensemble may include removing unsupervised ML models and/or replacement of the removed unsupervised ML models with other unsupervised ML models that are available for inclusion in the ensemble; the operation of steps 304-322 may then be repeated for each subsequent portion of input data until all portions of the input data have been processed); and
g. providing the combined machine learning model as output (Givental'644, ¶ [0005] and ABSTRACT: processing subsequent portions of input data using the modified ensemble of unsupervised ML models; ¶ [0023]: these dynamically determined weights are then used thereafter for predicting new data using the ensemble unsupervised machine learning operation described previously; ¶ [0027]: the illustrative embodiments may be implemented in a variety of different computing system environments for performing classifications of event patterns in input data indicative of a particular classification of interest; e.g., in a medical field, the combined unsupervised/supervised machine learning model mechanisms of the illustrative embodiments may be used as a basis for classifying medical image data that is input to the mechanisms of the illustrative embodiments, with regard to whether or not anomalous regions are present in the medical image data, such as cancerous tumors, blockages in blood vessels, or any other medical anomaly that may be detected in medical images; ¶ [0051] with FIG. 1: the output results of the unsupervised ML models 122-128 or a subset of unsupervised ML models 122-126, which are included in the in the ensemble 120 indicate, for the corresponding unsupervised ML model, the predicted classification of the input data generated by the evaluation of the combination of input features by that particular unsupervised ML model, e.g., each unsupervised ML model may predict whether the particular combination or pattern of features in the input data represent an anomaly or non-anomalous data; the output may be a binary output for configurations where there is only an indication of one class or a second class, e.g., "anomalous" or "non-anomalous", or may be a vector output having multiple vector slots in excess of two, each associated with a different classification; in some cases, the output may comprise one or more values representing a confidence or probability that the corresponding classification is correct, e.g., an output value of 0.85 indicates an 85% confidence or probability that the corresponding classification, e.g., anomalous, is a correct classification for the input data).
Claim 2
Givental'644 discloses all the elements as stated in Claim 1 and further disclose wherein the processing comprises at least one processing step, selected from the group comprising: adapting the determined at least one output label (Givental'644, ¶ [0005] and ABSTRACT: outputting the ensemble output to an authorized user computing device to obtain user feedback from the authorized user via the user computing device; the user feedback indicates a correctness of the ensemble output; ¶¶ [0021]-[0022]: the events in the log data that have a high likelihood or probability (as measured by comparing the likelihood output of the ensemble to one or more predetermined threshold values) are provided to a human security analyst via a user interface output on a computing device associated with the security analyst, so that they may review and label the input log data as either anomalous or not anomalous, e.g., "true threat" or "false positive; the newly detected log data having generated labels with high likelihood of representing anomalies are again sent to a human system analyst for review and responsive action; the newly detected log data having generated labels with high likelihood of representing anomalies are again sent to a human system analyst for review and responsive action; ¶¶ [0043], [0052]-[0053], [0055]-[0059], and [0063]-[0064] with FIG. 1: the security analyst computing device 155 may be integrated with the STEM computer system 190, or may be used to access the STEM computer system 190 to provide security analyst feedback; instances of outputs of the unsupervised machine learning models indicating potential security threats/attacks, unauthorized accesses, or the like, may be sent to a human security analyst, such as via a graphical user interface and workstation or other computing device associated with the human security analyst, so that the human security analyst may verify whether the input is actually a true threat or false positive; the human security analyst may then provide an input via the graphical user interface to indicate their assessment of the input and thus, whether the corresponding unsupervised machine learning model(s) were able to generate the correct result; this feedback information may be used to set or adjust weights associated with the outputs of the various unsupervised machine learning models 122-128 of the ensemble 120; the human security analyst may also be provided with outputs indicating that the input to the unsupervised machine learning model is a "normal" or "non-anomalous" input so that the human security analyst may provide such feedback information as well, but this time with regard to whether the input is or is not a threat and thus indicate whether the output of the unsupervised machine learning model(s) represent a true non-threat or a false negative result; such "normal" or "non-anomalous" outputs by the unsupervised machine learning models may be sent to the human security analyst only in instances where the confidence or probability associated with the classification is below a predetermined threshold, such that not all indications of "normal" or "non-anomalous" need be reviewed; the same is true for the "anomaly" or "anomalous" classifications, i.e. a threshold for confidence or probability value may be established and if the confidence/probability for the classification of "anomaly" or "anomalous" is equal to or above the threshold, then the human security analyst may be enlisted to provide feedback information; the weights 132-136 are dynamically determined based on feedback information obtained from security analyst review of the outputs of the individual unsupervised machine learning models 122-128, or the final anomaly score 140 generated by a combination of the weighted outputs for the unsupervised machine learning models 122-128; thus, a first weight W1 132 is applied to the outputs from the first unsupervised machine learning model (the isolation forest model in the depicted example) 122, a second weight W2 134 is applied to the outputs from the second unsupervised machine learning model (the one-class SVM model in the depicted example) 124, and a third weight W3 136 is applied to the outputs from the third unsupervised machine learning model (the local outlier factor model in the depicted example), while the other machine learning models 128 are not included in this particular grouping; the weights W1-W3 132-136 may be different from each other due to dynamic modifications of the individual weights 132-136 based on the dynamic weight generator 130 processing feedback information from the security analyst workstation 155 as to the correctness/incorrectness of the anomaly score 140 (or in other embodiments, the individual outputs of the unsupervised machine learning models 122-128 prior to combining them to generate the anomaly score 140); the anomaly score 140 generated by combining the weighted outputs for the various unsupervised machine learning models 122-128 in the ensemble 120 is used to identify input log data that corresponds with anomaly scores 140 indicative of a need for human security analyst review, e.g., logs associated with anomaly scores 140 equal to or above a threshold value; alternatively, a predetermined percentage of log data and corresponding anomaly scores 140 may be selected and provided to the security analyst via the workstation 155 for review; an additional threshold value may be used to select a subset of the log data inputs indicative of an anomaly that are to be reviewed by the security analyst and also provided as labeled data for inclusion in the partially labeled data 162; e.g., a selection threshold value may be set such that a relatively small subset of the log data having high anomaly scores 140, i.e. anomaly scores 140 equal to or above the anomaly threshold value, is selected for obtaining feedback information and for inclusion in the partially labeled data 162, e.g., a small percentage, e.g., approximately 2%, of the log data; the labeled log data 150, or the selected subset portion thereof, are used to provide the output to the workstation 155 for review by the security analyst 155 who can then either indicate agreement or non-agreement with the labels; the security analyst's agreement or non-agreement with the anomaly score 140 output may be used to update performance metrics associated with the unsupervised ML models 122-128 that are part of the ensemble of unsupervised ML models 120 as stored in the metrics storage and ML model selection engine 129; if the security analyst reviews a high anomaly score event or log entry in the selected subset of labeled log data 150, and disagrees with the label generated as a result of the initial anomaly score 140, then the security analyst may indicate this disagreement and/or a correct label, and a corresponding default anomaly score will be associated with the event or log entry, e.g., if the security analyst indicates that the high anomaly score is a false positive, then the anomaly score may be changed to a zero anomaly score value which is associated with the correct label of "non-anomalous" in the selected subset of labeled log data 150 that is stored in the partially labeled data 162; for those events indicating anomalies, corresponding logged anomalies are set to the STEM computing system 190 for further evaluation and responsive action; the only manual input required is the security analyst review of a selected subset of log data and their labels whose anomaly score 140 meets or exceeds a particular threshold value; e.g., the threshold may be set such that only a small percentage of the logs and corresponding labels are sent to the security analyst workstation 155 for review; e.g., approximately only 2% of the log data and its corresponding labels are sent to the security analyst workstation 155 for review by the security analyst; ¶ [0068] with 214-216 in FIG. 2: a subset of the log data having anomaly scores equal to or greater than an anomaly threshold value, is selected for user feedback and inclusion in the partially labeled dataset; user feedback information specifying correct labels for the subset of the log data is obtained and used to update the labels of this subset before inclusion of the subset as labeled data in the partially labeled dataset; ¶¶ [0071]-[0072] with 310-314 in FIG. 3: a determination is then made as to whether the ensemble output indicates a need for user feedback (step 310); this determination may be, e.g., based on a comparison of the ensemble output to one or more criteria, e.g., threshold values or the like; if the one or more criteria are satisfied by the ensemble output, i.e. there is a determined need for user feedback, then the operation may continue with steps 312-324; if the one or more criteria are not satisfied, i.e. user feedback is determined to not be needed, the operation may skip steps 312-324 and continue at step 326; in response to determining that there is a need for user feedback, the ensemble output is output to a user workstation, such as a computing device associated with an authorized user, e.g., security analyst or the like (step 312); user feedback is then received from the user workstation (step 314); the user feedback indicates a correct output that the ensemble should have generated; e.g., the user feedback may indicate agreement with the ensemble output or disagreement with the ensemble output and an indication of what the correct output should have been according to the user's evaluation).
Claim 3
Givental'644 discloses all the elements as stated in Claim 1 and further disclose wherein the at least one additional machine learning model is an unsupervised or supervised machine learning model (Givental'644, ¶¶ [0019], [0020], and [0023]-[0024]: leverage the strengths of both supervised and unsupervised machine learning in a hybrid approach to training a machine learning model to detect anomalies in computer system log data structures; combine unsupervised and supervised machine learning in a specific manner to provide a computer machine learning model that is able to detect anomalies in log data structures; an ensemble of a plurality of machine learning models trained using unsupervised machine learning algorithms, e.g., isolation forest, local outlier factor, one-class support vector machine (SVM), and/or the like, are generated with dynamic weighting; the responses provided by the human security analysts in response to the user interface outputs generated for high likelihood anomalous log data, are provided to a dynamic weight generator to determine which unsupervised machine learning model(s) in the ensemble provided the most accurate outputs, e.g., most accurate classifications of logged events as "true threat"; the dynamic weight generator executes a weight generator supervised machine learning model, e.g., SVM, neural networks, etc., to assign weights to the top N number of machine learning models in the ensemble, where N is any desirable number of machine learning models for the particular implementation, e.g., top 3 machine learning models, where the "top" number refers to the models with a predetermined specified level of performance, e.g., the highest accuracy relative to all the models in the ensemble; these dynamically determined weights are then used thereafter for predicting new data using the ensemble unsupervised machine learning operation described previously; provide a mechanism that combines unsupervised and supervised machine learning to provide accurate anomaly detection in computer system log data structures; ¶ [0071] with 308 in FIG. 3: the weighted outputs of the individual unsupervised ML models are then combined using a combinatorial function, e.g., a weighted function which applies weights to each of the outputs of each of the unsupervised ML models in the ensemble, to generate a single ensemble output for the ensemble (step 308)).
Claim 5
Givental'644 discloses all the elements as stated in Claim 1 and further disclose wherein the connection function is a weighted function (Givental'644, ¶¶ [0020] and [0023]: an ensemble of a plurality of machine learning models trained using unsupervised machine learning algorithms, e.g., isolation forest, local outlier factor, one-class support vector machine (SVM), and/or the like, are generated with dynamic weighting; the responses provided by the human security analysts in response to the user interface outputs generated for high likelihood anomalous log data, are provided to a dynamic weight generator to determine which unsupervised machine learning model(s) in the ensemble provided the most accurate outputs, e.g., most accurate classifications of logged events as "true threat"; the dynamic weight generator executes a weight generator supervised machine learning model, e.g., SVM, neural networks, etc., to assign weights to the top N number of machine learning models in the ensemble, where N is any desirable number of machine learning models for the particular implementation, e.g., top 3 machine learning models, where the "top" number refers to the models with a predetermined specified level of performance, e.g., the highest accuracy relative to all the models in the ensemble; these dynamically determined weights are then used thereafter for predicting new data using the ensemble unsupervised machine learning operation described previously; ¶¶ [0049]-[0050] with FIG. 1: the evaluation of the ML models for selection in the subset of ML models may be performed using any suitable performance evaluation for the particular implementation; e.g., metrics may be maintained for each of the ML models with regard to their accuracy and/or precision in predicting classifications or labels for the inputs, as well as other performance metrics, and these metrics may be used as a basis for selecting the X number of ML models, e.g., X ML models having a relatively highest accuracy amongst the plurality of ML models 122-128; the accumulated performance metrics may be maintained by the hybrid machine learning (ML) anomaly detector 100, such as in the metrics storage and ML model selection engine 129, which also comprises the computer implemented logic for dynamically and automatically selecting a subset of ML models for inclusion in the ensemble 120; with regard to dynamically and automatically selecting a subset of ML models to include in the ensemble of unsupervised ML models 120, the metrics storage and ML model selection engine 129 may determine when the performance of an unsupervised ML model, e.g., unsupervised ML model 126, in the ensemble 120 is providing poor performance and may automatically replace this unsupervised ML model 126 with another unsupervised ML model, e.g., unsupervised ML model 128, in the available plurality of unsupervised ML models 122-128; e.g., threshold performance metric values may be established indicating an acceptable level of performance by the unsupervised ML models with regard to one or more performance metrics, e.g., accuracy, precision, etc.; if an unsupervised ML model in the ensemble 120 has an accumulated performance metric that equals or falls below the predetermined threshold value, then the unsupervised ML model may be removed from the ensemble 120 and/or replaced by another available unsupervised ML model; the modification to the ensemble 120 may also be communicated to the dynamic weights generator 130 so that corresponding weight values for the removed/replaced unsupervised ML model may be updated to reflect the removal/replacement; the weight value associated with the removed unsupervised ML model may be removed from the set of weights applied to the outputs of the unsupervised ML models in the ensemble 120, and if the unsupervised ML model is replaced, a corresponding weight for the replacement unsupervised ML model may be instigated in the set of weights, or a default weight value may be set for the replacement unsupervised ML model; ¶¶ [0052]-[0054] and [0057]-[0058] with FIG. 1: the human security analyst may then provide an input via the graphical user interface to indicate their assessment of the input and thus, whether the corresponding unsupervised machine learning model(s) were able to generate the correct result; this feedback information may be used to set or adjust weights associated with the outputs of the various unsupervised machine learning models 122-128 of the ensemble 120; the outputs of the various unsupervised machine learning models 122-128 may be output to a dynamic weights generator 130 which applies weights 132-136 to the outputs prior to combining the results of the unsupervised machine learning models 122-128 to generate the anomaly score 140; the weights 132-136 are dynamically determined based on feedback information obtained from security analyst review of the outputs of the individual unsupervised machine learning models 122-128, or the final anomaly score 140 generated by a combination of the weighted outputs for the unsupervised machine learning models 122-128; the weights W1-W3 132-136 may be different from each other due to dynamic modifications of the individual weights 132-136 based on the dynamic weight generator 130 processing feedback information; in combining the weights to generate an initial anomaly score 140, any suitable function for the particular implementation may be used to combine the weighted anomaly scores from the individual unsupervised ML models 122-126 into a single anomaly score for the particular log entry/event that was evaluated by the ensemble 120; e.g., the function may be a sum of the individual unsupervised ML models 122-126 scores weighted by the corresponding weights 132-136; in other implementations, an average of the weighted anomaly scores may be utilized; in still other implementations, other functions involving each of the weighted anomaly scores generated by the individual unsupervised ML models 122-126 included in the ensemble 120 may be utilized; this feedback information is fed back into the dynamic weights generator 130 which determines how to modify the weights W1-W3 132-136 to increase the correctness of the anomaly scores 140 and corresponding labels generated by the ensemble 120; i.e., some models may operate better on different types of patterns of input data, e.g., different patterns in security log data; as a result, weights may need to be dynamically adjusted based on the patterns of log data input to the ensemble 120; the dynamic weights generator 130 dynamically adjusts the weights 132-136 by receiving the feedback information as to correctness, determining which unsupervised machine learning models 122-128 generated the correct output indicated in the feedback information and which generate the incorrect output; the weights of the unsupervised machine learning models that generate the correct output may be increased whereas the weights of the unsupervised machine learning models may be decreased; the amount of increase/decrease may be determined based on a desired function for the particular implementation; the security analyst's agreement or non-agreement with the anomaly score 140 output may be used to update performance metrics associated with the unsupervised ML models 122-128 that are part of the ensemble of unsupervised ML models 120 as stored in the metrics storage and ML model selection engine 129; those unsupervised ML models in the ensemble 120 that generated a correct output as indicated by the user feedback may have their performance metric(s) increased to represent that these unsupervised ML models are generating correct results; those unsupervised ML models in the ensemble 120 that generated an incorrect output as indicated by the user feedback may have their performance metric(s) decreased to represent that these unsupervised ML models are generating incorrect results; the amount of the increase/decrease may be a function of the amount of certainty the corresponding unsupervised ML model had in the output it generated, e.g., more certainty in an incorrect output may result in a larger decrease in the performance metric(s) and more certainty in a correct output may result in a larger increase in the performance metric(s); ¶¶ [0067]-[0068] with 220 in FIG. 2: the weights applied to the outputs of the various models are dynamically determined based on feedback information; any suitable function for combining the weighted outputs from the models may be used without departing from the spirit and scope of the present invention, e.g., simply sum, averaging of the weighted anomaly scores from the models, selection of a highest weighted anomaly score from the models, etc.; the feedback information is fed back into a dynamic weight generator which updates the weights applied to the outputs of the various unsupervised machine learning models of the ensemble based on whether or not they output a correct result; ¶¶ [0069]-[0071] and [0073]-[0075] with 316-322 FIG. 3: with regard to the selection of unsupervised ML models to include in the ensemble and the adjustment of weights applied to the outputs of the selected unsupervised ML models when combining the outputs to generate a single ensemble output; these weight values may then be dynamically and automatically updated based on user feedback; the weighted outputs of the individual unsupervised ML models are then combined using a combinatorial function, e.g., a weighted function which applies weights to each of the outputs of each of the unsupervised ML models in the ensemble, to generate a single ensemble output for the ensemble; the user feedback is provided to a dynamic weight generator which updates the individual weights based on the user feedback; based on whether or not the corresponding unsupervised ML model generated a correct output individually, the unsupervised ML model's weight may be increased/decreased to give preference to unsupervised ML models that generate correct outputs and to not give preference to unsupervised ML models that generate incorrect outputs; in addition, the user feedback is provided to a metrics storage and ML model selection engine to thereby update performance metrics for the individual unsupervised ML models that are part of the ensemble; based on the updated performance metrics, a determination is made as to whether a change in the ensemble is to be performed; this determination may be made based on the performance metrics of unsupervised ML models in the ensemble satisfying one or more criteria for modifying the ensemble, e.g., performance metric(s) of an unsupervised ML model fall to or below a threshold performance value; in response to a determination that a change in the ensemble is to be performed, membership of the unsupervised ML models in the ensemble is automatically modified; this modification of the ensemble may include removing unsupervised ML models and/or replacement of the removed unsupervised ML models with other unsupervised ML models that are available for inclusion in the ensemble; the operation of steps 304-322 may then be repeated for each subsequent portion of input data until all portions of the input data have been processed).
Claim 6
Givental'644 discloses all the elements as stated in Claim 1 and further disclose a computer program product, comprising a computer readable hardware storage device (Givental'644, ¶¶ [0089]-[0091] with 608/624/626 in FIG. 6) having computer readable program code stored therein (Givental'644, ¶ [0094] with FIG. 6: instructions for the operating system, the object oriented programming system, and applications or programs are located on storage devices), said program code executable by a processor (Givental'644, ¶¶ [0089] and [0093] with 606/610 in FIG. 6: processing unit 606; graphics processor 610; data processing system 600 may be a symmetric multiprocessor (SMP) system including a plurality of processors in processing unit 606) of a computer system (Givental'644, ¶¶ [0088] with 600 in FIG. 6 and ¶ [0084] with 504/530 in FIG. 5: data processing system 600 is an example of a computer, such as server 504 in FIG. 5; one or more of the computing devices, e.g., server 504, may be specifically configured to implement a hybrid machine learning anomaly detector 530) to implement a method for performing the steps according to claim 1 when the computer program product is running on a computer (Givental'644, ¶¶ [0033]-[0034]: the computer program product may include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present invention; the computer readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing; ¶¶ [0092]-[0094] with FIG. 6: an operating system runs on processing unit 606; instructions for the operating system, the object oriented programming system, and applications or programs are located on storage devices, such as HDD 626, and may be loaded into main memory 608 for execution by processing unit 606).
Claim 7
Givental'644 discloses all the elements as stated in Claim 1 and further discloses a technical system (Givental'644, ¶ [0043] with FIG. 1: the hybrid ML anomaly detector 100 comprises a data cleaning and feature engineering engine 110, an unsupervised machine learning model ensemble 120, a dynamic weights generator 130, and a semi-supervised ML model 170; ¶¶ [0084] and [0086] with 530 in FIG. 4: one or more of the computing devices, e.g., server 504, may be specifically configured to implement a hybrid machine learning anomaly detector 530; ¶ [0088] with FIG. 6: data processing system 600 is an example of a computer, such as server 504 in FIG. 5).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over Givental'644 in view of Muselli et al. (US 2022/0036137 A1, filed on 03/12/2021), hereinafter Muselli.
Claim 4
Givental'644 discloses all the elements as stated in Claim 1 and further disclose wherein the connection function is a logical function or logical operator, (Givental'644, ¶¶ [0020] and [0023]: an ensemble of a plurality of machine learning models trained using unsupervised machine learning algorithms, e.g., isolation forest, local outlier factor, one-class support vector machine (SVM), and/or the like, are generated with dynamic weighting; the responses provided by the human security analysts in response to the user interface outputs generated for high likelihood anomalous log data, are provided to a dynamic weight generator to determine which unsupervised machine learning model(s) in the ensemble provided the most accurate outputs, e.g., most accurate classifications of logged events as "true threat"; the dynamic weight generator executes a weight generator supervised machine learning model, e.g., SVM, neural networks, etc., to assign weights to the top N number of machine learning models in the ensemble, where N is any desirable number of machine learning models for the particular implementation, e.g., top 3 machine learning models, where the "top" number refers to the models with a predetermined specified level of performance, e.g., the highest accuracy relative to all the models in the ensemble; these dynamically determined weights are then used thereafter for predicting new data using the ensemble unsupervised machine learning operation described previously; ¶¶ [0049]-[0050] with FIG. 1: the evaluation of the ML models for selection in the subset of ML models may be performed using any suitable performance evaluation for the particular implementation; e.g., metrics may be maintained for each of the ML models with regard to their accuracy and/or precision in predicting classifications or labels for the inputs, as well as other performance metrics, and these metrics may be used as a basis for selecting the X number of ML models, e.g., X ML models having a relatively highest accuracy amongst the plurality of ML models 122-128; the accumulated performance metrics may be maintained by the hybrid machine learning (ML) anomaly detector 100, such as in the metrics storage and ML model selection engine 129, which also comprises the computer implemented logic for dynamically and automatically selecting a subset of ML models for inclusion in the ensemble 120; with regard to dynamically and automatically selecting a subset of ML models to include in the ensemble of unsupervised ML models 120, the metrics storage and ML model selection engine 129 may determine when the performance of an unsupervised ML model, e.g., unsupervised ML model 126, in the ensemble 120 is providing poor performance and may automatically replace this unsupervised ML model 126 with another unsupervised ML model, e.g., unsupervised ML model 128, in the available plurality of unsupervised ML models 122-128; e.g., threshold performance metric values may be established indicating an acceptable level of performance by the unsupervised ML models with regard to one or more performance metrics, e.g., accuracy, precision, etc.; if an unsupervised ML model in the ensemble 120 has an accumulated performance metric that equals or falls below the predetermined threshold value, then the unsupervised ML model may be removed from the ensemble 120 and/or replaced by another available unsupervised ML model; the modification to the ensemble 120 may also be communicated to the dynamic weights generator 130 so that corresponding weight values for the removed/replaced unsupervised ML model may be updated to reflect the removal/replacement; the weight value associated with the removed unsupervised ML model may be removed from the set of weights applied to the outputs of the unsupervised ML models in the ensemble 120, and if the unsupervised ML model is replaced, a corresponding weight for the replacement unsupervised ML model may be instigated in the set of weights, or a default weight value may be set for the replacement unsupervised ML model; ¶¶ [0052]-[0054] and [0057]-[0058] with FIG. 1: the human security analyst may then provide an input via the graphical user interface to indicate their assessment of the input and thus, whether the corresponding unsupervised machine learning model(s) were able to generate the correct result; this feedback information may be used to set or adjust weights associated with the outputs of the various unsupervised machine learning models 122-128 of the ensemble 120; the outputs of the various unsupervised machine learning models 122-128 may be output to a dynamic weights generator 130 which applies weights 132-136 to the outputs prior to combining the results of the unsupervised machine learning models 122-128 to generate the anomaly score 140; the weights 132-136 are dynamically determined based on feedback information obtained from security analyst review of the outputs of the individual unsupervised machine learning models 122-128, or the final anomaly score 140 generated by a combination of the weighted outputs for the unsupervised machine learning models 122-128; the weights W1-W3 132-136 may be different from each other due to dynamic modifications of the individual weights 132-136 based on the dynamic weight generator 130 processing feedback information; in combining the weights to generate an initial anomaly score 140, any suitable function for the particular implementation may be used to combine the weighted anomaly scores from the individual unsupervised ML models 122-126 into a single anomaly score for the particular log entry/event that was evaluated by the ensemble 120; e.g., the function may be a sum of the individual unsupervised ML models 122-126 scores weighted by the corresponding weights 132-136; in other implementations, an average of the weighted anomaly scores may be utilized; in still other implementations, other functions involving each of the weighted anomaly scores generated by the individual unsupervised ML models 122-126 included in the ensemble 120 may be utilized; this feedback information is fed back into the dynamic weights generator 130 which determines how to modify the weights W1-W3 132-136 to increase the correctness of the anomaly scores 140 and corresponding labels generated by the ensemble 120; i.e., some models may operate better on different types of patterns of input data, e.g., different patterns in security log data; as a result, weights may need to be dynamically adjusted based on the patterns of log data input to the ensemble 120; the dynamic weights generator 130 dynamically adjusts the weights 132-136 by receiving the feedback information as to correctness, determining which unsupervised machine learning models 122-128 generated the correct output indicated in the feedback information and which generate the incorrect output; the weights of the unsupervised machine learning models that generate the correct output may be increased whereas the weights of the unsupervised machine learning models may be decreased; the amount of increase/decrease may be determined based on a desired function for the particular implementation; the security analyst's agreement or non-agreement with the anomaly score 140 output may be used to update performance metrics associated with the unsupervised ML models 122-128 that are part of the ensemble of unsupervised ML models 120 as stored in the metrics storage and ML model selection engine 129; those unsupervised ML models in the ensemble 120 that generated a correct output as indicated by the user feedback may have their performance metric(s) increased to represent that these unsupervised ML models are generating correct results; those unsupervised ML models in the ensemble 120 that generated an incorrect output as indicated by the user feedback may have their performance metric(s) decreased to represent that these unsupervised ML models are generating incorrect results; the amount of the increase/decrease may be a function of the amount of certainty the corresponding unsupervised ML model had in the output it generated, e.g., more certainty in an incorrect output may result in a larger decrease in the performance metric(s) and more certainty in a correct output may result in a larger increase in the performance metric(s); ¶¶ [0067]-[0068] with 220 in FIG. 2: the weights applied to the outputs of the various models are dynamically determined based on feedback information; any suitable function for combining the weighted outputs from the models may be used without departing from the spirit and scope of the present invention, e.g., simply sum, averaging of the weighted anomaly scores from the models, selection of a highest weighted anomaly score from the models, etc.; the feedback information is fed back into a dynamic weight generator which updates the weights applied to the outputs of the various unsupervised machine learning models of the ensemble based on whether or not they output a correct result; ¶¶ [0069]-[0071] and [0073]-[0075] with 316-322 FIG. 3: with regard to the selection of unsupervised ML models to include in the ensemble and the adjustment of weights applied to the outputs of the selected unsupervised ML models when combining the outputs to generate a single ensemble output; these weight values may then be dynamically and automatically updated based on user feedback; the weighted outputs of the individual unsupervised ML models are then combined using a combinatorial function, e.g., a weighted function which applies weights to each of the outputs of each of the unsupervised ML models in the ensemble, to generate a single ensemble output for the ensemble; the user feedback is provided to a dynamic weight generator which updates the individual weights based on the user feedback; based on whether or not the corresponding unsupervised ML model generated a correct output individually, the unsupervised ML model's weight may be increased/decreased to give preference to unsupervised ML models that generate correct outputs and to not give preference to unsupervised ML models that generate incorrect outputs; in addition, the user feedback is provided to a metrics storage and ML model selection engine to thereby update performance metrics for the individual unsupervised ML models that are part of the ensemble; based on the updated performance metrics, a determination is made as to whether a change in the ensemble is to be performed; this determination may be made based on the performance metrics of unsupervised ML models in the ensemble satisfying one or more criteria for modifying the ensemble, e.g., performance metric(s) of an unsupervised ML model fall to or below a threshold performance value; in response to a determination that a change in the ensemble is to be performed, membership of the unsupervised ML models in the ensemble is automatically modified; this modification of the ensemble may include removing unsupervised ML models and/or replacement of the removed unsupervised ML models with other unsupervised ML models that are available for inclusion in the ensemble; the operation of steps 304-322 may then be repeated for each subsequent portion of input data until all portions of the input data have been processed).
Givental'644 fails to explicitly disclose an AND and OR logical operator.
Muselli teaches a system and a method for detecting anomalies (Muselli, ABSTRACT), wherein an AND and OR logical operator (Muselli, ¶¶ [0064]-[0069]: the classification method is a rule generation method generating a classification model in the form of one or more conditional rules of the type; the logical operators linking different conditions may be AND and/or OR; <conditions> is the logical product (AND) of mk conditions ckl, with l=1, … , mk, on the components xj, whereas <consequence> gives a class assignment y=
y
~
for the output; in general, according to the type of the variable xj a condition ckl in the premise of the rule has one of the following forms: a threshold condition xj > λ, xj ≥ μ, or λ ≤ xj ≤ μ, where λ and μ are two values in the domain Bj of xj, if xj is an ordered variable; a membership condition xj [Symbol font/0xCE]A, where λ is a non-empty subset of the domain Bj, if xj is a nominal variable; one of the advantages offered by a rule generation method is that the logical rules may be understood and interpreted by a human expert; in one embodiment, the rule generation method is based on decision trees techniques; in an embodiment, the rule generation method is based on a shadow clustering method; ¶ [0104]: a positive Boolean function f(z) can be simply written in its Disjunctive Normal Form (DNF) as a logical sum (OR) of logical products (AND) among some of the b components of string z).
Givental'644 and Muselli are analogous art because they are from the same field of endeavor, system and a method for detecting anomalies. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to apply the teaching of Muselli to Givental'644. Motivation for doing so would allow a human expert more easily understand and interpret/explain the results (Muselli, ¶¶ [0005] and [0069]).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Das et al. ("Active Anomaly Detection via Ensembles", arXiv:1809.06477, 09/17/2018) discloses in ABSTRACT and Section 1 of Page 1 that (1) consider the problem of anomaly detection, where the goal is to detect unusual but interesting data (referred to as anomalies) among the regular data (referred to as nominals); (2) study label-efficient active learning algorithms to improve unsupervised anomaly detector ensembles and address the above three shortcomings of prior work in a principled manner; (3) present an important insight into how anomaly detector ensembles are naturally suited for active learning, and why the greedy querying strategy of seeking labels for instances with the highest anomaly scores is efficient; (4) present a novel approach to describe the discovered anomalies and show that it can also be employed to improve the diversity of the instances presented to the analyst while achieving an anomaly discovery rate comparable to the greedy strategy; (5) present a novel algorithm to robustly detect drift in data streams and to adapt the detector in a principled manner; (6) present extensive empirical evidence in support of our insights and algorithms in both batch and streaming settings; and (7) also demonstrate the improvement in diversity through an objective measure relevant to real-world settings. Das further discloses in Section 2 of Pages 1-2 that (1) unsupervised anomaly detection algorithms are trained without labeled data, and have assumptions baked into the model about what defines an anomaly or a nominal; (2) active learning corresponds to the setup where the learning algorithm can selectively query a human analyst for labels of input instances to improve its prediction accuracy, wherein the overall goal is to minimize the number of queries to reach the target performance; and (3) the human analyst provides feedback to the algorithm on true labels, and if the algorithm makes wrong predictions, it updates its model parameters to be consistent with the analyst’s feedback. Das also discloses in Section 3 of Pages 2-3 that (1) instances labeled +1 represent the anomaly class and the label -1 represents the nominal class; (2) assume a linear model with weights w [Symbol font/0xCE]
R
m that will be used to combine the scores of m anomaly detectors as follows: Score(x) = w • z, where z [Symbol font/0xCE]
R
m correspond to the scores from anomaly detectors for instance x; (3) the goal of
A
is to learn optimal weights for maximizing the number of true anomalies shown to the analyst; (4) the common theme is that when the ensemble members are ideal, then the scores of true anomalies tend to lie in the farthest possible location in the positive direction of the uniform weight vector wunif by design; (5) when detecting anomalies with an ensemble of detectors: (a) it is compelling to always apply active learning; (b) the greedy strategy of labeling top ranked instances is efficient, and is therefore a good yardstick for evaluating the performance of other querying strategies as well; and (c) learning a decision boundary with active learning that generalizes to unseen data helps in limited-memory or streaming data settings; (6) tree-based ensemble detectors have several properties which make them ideal for active learning: (a) they can be employed to construct large ensembles inexpensively; (b) treating the nodes of tree as ensemble members allows us to both focus our feedback on fine-grained subspaces as well as increase the capacity of the model; and (c) since some of the tree-based models such as Isolation Forest (IFOR), HS Trees (HST), and RS Forest (RSF) are state-of-the-art unsupervised detectors, it is a significant gain if their performance can be improved with minimal feedback; (7) IFOR comprises of an ensemble of isolation trees, wherein each leaf node is assigned a score proportional to the length of the path from the root to itself; (8) by construction, leaves which correspond to subspaces that contain more anomalous instances have, on an average, smaller path lengths; (9) assign the leaf score as the negative path length; (10) as a result, anomalous instances have higher scores than the nominals; (11) we will represent the leaf-level scores by d; and (12) after constructing the trees, we extract the leaf nodes as the ensemble members, wherein each ensemble member assigns its leaf score as the anomaly score to an instance if the instance belongs to the corresponding subspace, else 0. Das further teaches in Section 4 of Pages 3-5 that (1) first present a novel formalism called compact description that describes groups of instances compactly using a tree-based model; (2) discuss a novel querying strategy that employs these descriptions to diversify the instances selected for labeling; (3) when the selected instance(s) are labeled by an analyst, the model updates the weights of ensembles to be consistent with all the instances labeled so far; (4) the algorithms to update the weights, in both batch and streaming data settings, are discussed next; (5) in the batch setting, the entire data is available at the outset; whereas, in the streaming setting, the data comes as a continuous stream; (6) the tree-based model assigns a weight and an anomaly score to each leaf; (7) denote the vector of leaf-level anomaly scores by d, and the overall anomaly scores of the subspaces (corresponding to the leaf-nodes) by a = {a1, …, am} = w ○ d, where denotes element-wise product operation; (8) the score ai provides a good measure of the relevance of the i-th subspace; (9) this relevance for each subspace is determined automatically through the label feedback; (10) our goal is to select a small subset of the most relevant and “compact” (by volume) subspaces which together contain all the instances in a group that we want to describe; (11) we treat this problem as a specific instance of the set covering problem; (12) the greedy strategy of selecting top-scored instances (referred as Select-Top) for labeling is efficient; (13) this strategy might lack diversity in the types of instances presented to the analyst; (14) the proposed strategy (Select-Diverse), which is described next, is intended to increase the diversity by employing tree-based ensembles to select groups of instances from subspaces that have minimum overlap; (15) Batch Active Learning (BAL): (a) the optimal hyperplane can now be learned efficiently by greedily asking the analyst to label the most anomalous instance in each feedback iteration; and (b) simplify the AAD formulation with a more scalable unconstrained optimization objective, and refer to this version as BAL; and (c) crucially, the ensemble weights are updated with an intent to maintain the hyperplane in the region of uncertainty through the entire budget B; (16) Streaming Active Learning (SAL): (a) in the streaming case, assume that the data is input to the algorithm continuously in windows of size K and is potentially unlimited; (b) initially, we train all the members of the ensemble with the first window of data; (c) when a new window of data arrives, the underlying tree model is updated as follows: (i) in case the model is an HST or RSF, only the node counts are updated while keeping the tree structures and weights unchanged; (ii) whereas, if the model is an IFOR, a subset of the current set of trees is replaced as shown in Update-Model (Algorithm 1); (iii) the updated model is then employed to determine which unlabeled instances to retain in memory, and which to “forget”; (iv) this step, referred to as Merge-and-Retain, applies the simple strategy of retaining only the most anomalous instances among those in the memory and in the current window, and discards the rest; (v) next, the weights are fine-tuned with analyst feedback through an active learning loop similar to the batch setting with a small budget Q; (vi) finally, the next window of data is read, and the process is repeated until the stream is empty or the total budget B is exhausted; and (vii) in the rest of this section, we will assume that the underlying tree model is IFOR; (17) when we replace a tree in Update-Model, its leaf nodes and corresponding weights get discarded; (18) on the other hand, adding a new tree implies adding all its leaf nodes with weights initialized to a default value v; (19) we first set
v
=
1
m
'
where m' is the total number of leaves in the new model, and then re-normalize the updated w to unit length; (20) SAL approach can be employed in two different situations: (a) limited memory with no concept drift, and (b) streaming data with concept drift; (21) the type of situation determines how fast the model needs to be updated in Update-Model; (22) if there is no concept drift, we need not update the model at all; (23) if there is a large change in the distribution of data from one window to the next, then a large fraction of members need to be replaced; (24) in general, it is hard to determine the true rate of drift in the data, and one approach is to replace, in the Update-Model step, a reasonable number (e.g. 20%) of older ensemble members with new members trained on new data; and (25) Drift Detection Algorithm: (a) Algorithm 1 presents a principled methodology that employs KL-divergence (denoted by DKL) to determine which trees should be replaced.
Zhang et al. (US 2023/0376026 A1, filed on 10/30/2020) discloses in ¶¶ [0033]-[0082] with FIGS.1-3 that (1) solve unsupervised learning tasks with supervised learning techniques by involving generic techniques to automate the model evaluation, feature selection, and explainable AI, which are usually available in supervised learning models, to solve unsupervised learning tasks; (2) failure detection by automating the manual process to detect failures accurately, efficiently, and effectively with anomaly detection models, and leveraging the introduced generic framework and solution architecture to apply supervised learning techniques (feature selection, model selection and explainable AI) to optimize and explain the anomaly detection models; (3) failure prediction via deriving signals/features within optimal feature windows and predicting rare failures within the optimal failure windows given the required response time by using both derived features and historical failures; (4) failure prevention via identifying the root cause of the predicted failures, automating the failure remediation recommendation by incorporating the domain knowledge, and suppressing alerts with an optimized, data-driven approach; (5) Sensor Data 100 are time series data collected from multiple sensors and will be the input in this solution, wherein the time series data is unlabeled, meaning that no manual process is required to label or tag the sensor data to indicate whether each data point corresponds to a failure or not; (6) Failure Detection 110 involves the following components configured to detect failures based on the input sensor data: (a) Feature Engineering 111 is used to derive features/signals which will be used to build failure detection and failure prediction models, which involves three sub-components: sensor selection, feature extraction, and feature selection; (b) Failure Detection 112 is configured to utilize an anomaly detection technique to detect rare failures in the industrial systems, wherein the detected rare failures are used as a target to build a failure prediction model, and the detected historical rare failures are also used to form features to build a failure prediction model; (7) Failure Prediction 120 involves the following components configured to predict failures with the features and detected failures: (a) Feature Transformer 121 transforms the features from the feature engineering module and detected failures into a format that can be consumed by the Long Short Term Memory (LSTM) Auto Encoder and LSTM Failure Prediction module; (b) Auto Encoder 122 is used to encode the derived features from the Feature Engineering component 111 and the detected rare failures to remove the redundant information in the time series data, wherein the encoded features keep the signals in the time series data and will be used to build failure prediction models; (c) Failure Prediction module 123 involves a deep Recurrent Neural Network (RNN) model with an LSTM network architecture, which is used to build the failure prediction model with the encoded features (as features), original features (as target), and detected failures (as target); (d) Predicted Failures 124 is one output of the failure prediction module 123, which is represented as a score to indicate the likelihood to be a failure; (e) Predicted Features 125 is another output of the failure prediction module 123, which is a set of features that has the same format as the output of the Feature Engineering module 111; (f) Detected Failures 126 is the output by applying the failure detection model to Predicted Features 125 and generating detected failure scores; and (g) Ensemble Failures 127 ensembles the output of the Predicted Failures 124 and Detected Failures 126 to form a single failure score, wherein different ensemble techniques can be used; e.g., the average value of Predicted Failures 125 and Detected Failures 126 can be used as a single failure score; (8) Failure Prevention 130 involves the following components configured to identify root causes, automate the remediation recommendations, and suppress the alerts: (a) Root Cause Analysis 131 is performed to automatically determine the root cause of the predicted failures; (b) Remediation Recommendation 132 is configured to automatically generate remediation actions against the predicted failures by incorporation of the domain knowledge; e.g., an alert is generated to notify the operators so that they can remediate or avoid the failures based on the root causes of the failures; (c) Alert Suppression 133 is configured to suppress alerts to avoid flooding the alert queue of the operator, which is done through an automated data-driven optimization technique; and (d) Alerts 134 are the final output of the solution, which include predicted failure scores, root causes, and remediation recommendations; (9) the solution architecture for applying model selection techniques of supervised learning to select the best unsupervised learning model(s), how the ensemble model works, and lastly the rationale behind this solution architecture are described with respect to FIG. 2: (a) the first step is to derive features from the given dataset which is done through the Feature Engineering module 111; (b) next, several unsupervised learning model algorithms are manually chosen and several parameter sets for each model algorithm are manually chosen as well as shown at 300, wherein each combination of model algorithm and parameter set will be used to build a model against the features derived from the feature engineering step as shown in FIG. 2; (c) however, due to the nature of unsupervised learning tasks, there are no ground truth facts that can be used to measure how the model performs, and hence implement a generic solution to evaluate how the model performs by stacking supervised learning models 301 on top of unsupervised learning models, wherein for each unsupervised learning model, the unsupervised learning model is applied to the features or data points to get the unsupervised results and such unsupervised results can involve which cluster each data point belongs to for clustering problems, or whether the data point indicates an anomaly for an anomaly detection problem, and so on; (d) such results and features will be the input for a supervised ensemble model, where features from the unsupervised teaming model will be used as features for supervised learning models, and results from the unsupervised learning model will be used as the target for supervised learning models; (e) the supervised ensembled models can be evaluated by comparing the target (results from the unsupervised learning model) and the predicted results from supervised ensemble models; (f) based on such evaluation results, which supervised ensemble model can produce the best evaluation results can thereby be identified; and (g) then, identify which unsupervised learning model corresponds to the best evaluation results at, and take that as the best unsupervised learning model with the best model parameter set, and output the model at 302; (10) implementation of a solution architecture for ensembling supervised learning models to train, select, and ensemble supervised learning models, wherein each "Ensemble Model xx" in FIG. 2 is represented by FIG. 3: (a) several supervised learning model algorithms are manually chosen and several parameter sets for each model algorithm are manually chosen as well; (b) select models with hyperparameter optimization, wherein several hyperparameter optimization techniques can be used, which include grid search, random search, Bayesian optimization, evolutional optimization, and reinforcement learning; (c) form the ensemble models 402, wherein the models from all the model algorithms are ensembled to form the final ensemble model 402 and ensemble is a process to combine or aggregate multiple individually trained models into one single model to make prediction for the unseen data which help reduce the generalization error of the prediction, assuming the base models are diverse and independent.; (d) different ensemble techniques can be used as follows: (i) classification models: the majority voting technique can be used to ensemble classification models; e.g., apply each model to the current feature set and get the predicted classes and the class that appears most frequently will be used for the final prediction of the instance; (ii) regression models: there are several techniques for ensembling regression models: <a>average for regression models: for each instance, apply each model to the current feature set and get the predicted value, and then, use the average of the predicted values from different models as the final prediction value; <b> trimmed average for regression models: for each instance, apply each model to the current feature set and get the predicted value, remove both the highest and the lowest prediction value(s) from the models and calculate the average of the remaining predicted values, and use the trimmed average value for the final prediction value; and <c> weighted average for regression models: for each instance, apply each model to the current feature set and get the predicted value, assign a weight to the predicted value based on the evaluation accuracy of the model, where the higher the accuracy of the model, the more weight that will be assigned to the predicted value from the model, and then, calculate the average of the weighted predicted values and use the weighted average value for the final prediction value, wherein the weights for different models need to be normalized so that the sum of the weights is equal to 1; (11) the evaluation of unsupervised learning model fu can be translated into the evaluation of the relationship between features and the results discovered by fu; (12) for this task, stack a set of supervised learning models by using the Features from Feature Engineering 400 (FIG. 3) as features F, and Results from Unsupervised Leaming Models 401 as target T to train the supervised learning models; (13) for the set of supervised learning models, several supervised learning model algorithms that are distinct in nature are chosen manually first, and then several parameter sets are chosen for each supervised learning model algorithm; (14) at the model algorithm level, hyperparameter optimization techniques can determine the best parameter set for each model algorithm; (15) let fs be the best model for each supervised learning model algorithm, wherein each fs can be considered an independent evaluator and yields an evaluation score for fu: if fs discovers the similar relationship as fu does from F and T, then the evaluation score will be high; otherwise, the score will be low; (16) for each supervised learning model fs, the model evaluation score of fs can be used as the evaluation score for unsupervised learning model fu: for each fs, the target T is computed by fu, while the predicted value is computed by fs; (17) the evaluation score for fs, which is computed as closeness between the target and predicted value, is essential to measure the similarity of relationships between F and T that are discovered by unsupervised learning model fu and supervised learning model fs; (18) at this point, several supervised learning models fs are obtained for each unsupervised model fu, and each fs gives an evaluation score for fu, and the scores will be aggregated or ensembled to determine whether the unsupervised learning model fu is a good model or not; (19) since the underlying model algorithms of fs are diverse and distinct in nature from each other, they may give different scores to fu; (20) there are two cases: (a) if most of fs yields a high score to fu, then the relationship between F and T is well-captured by fu, and fu is considered to be good model; and (b) if most of fs yields a low score to fu, the relationship between F and T is not well-captured by fu, and fu is considered to be a bad model; i.e., if and only if fu reveals the relationship of F and T to be good, most fs are able to capture the relationship in a similar way as fu does, and they can yield a good score to fu, and vice versa, if fu reveals the relationship of F and T to be bad, most fs will capture the relationships under F and T badly in different ways, and are not able to capture the relationship in a similar way as fu does, and most fs will yield a bad score to fu; (21) once the evaluation score for each unsupervised model fu is obtained, the final unsupervised learning model can be selected by utilizing the global best model, in which the example implementations select the model with the best score across the model algorithms and the parameter sets and use that as the final model, or alternatively, it can be selected by utilizing the local best model, in which the example implementations first select the model with the best score for each model algorithm, and then ensemble the models, each from a model algorithm; (22) Root Cause Analysis (RCA), on the other hand, is usually done at instance level, i.e., each prediction can have some root causes; (23) there are two broad families of models for RCA: Deterministic models and Probabilistic models, wherein (a) Deterministic models only handle certainty in the known facts or the inferences expressed in the supervised learning model; and (b) Probabilistic models are able to handle this uncertainty in the supervised learning model; (24) once root causes are identified, it can help derive recommendations to remediate or avoid the potential problems and risks; (25) an unsupervised model such as the "Isolation Forest" model can be utilized to perform anomaly detection on the features data, which are derived from the feature engineering module on the data; (25) the output of the anomaly detection will be anomaly scores for the instances in the features data; (26) a supervised model, such as the "Decision Tree" model can be used to perform regression tasks, where the features for the "Decision Tree" model is the same as the features for the "Isolation Forest", and target for the "Decision Tree" model is the anomaly scores which are output from the "Isolation Forest" model; (27) to explain the decision tree, feature importance can be calculated at the model level, and root cause can be identified at instance level; and (28) to calculate feature importance at model level, one implementation is to calculate the decrease in node impurity weighted by the probability of reaching that node, wherein the node impurity can be measure as a gini index, and the node probability can be calculated by the number of samples that reach the node, divided by the total number of samples so that the higher the feature importance value, the more important the feature.
Tang et al. ("Deep Anomaly Detection with Ensemble-Based Active Learning", 2020 IEEE International Conference on Big Data (Big Data), Dec. 10-13, 2020, pp. 1663-1670) discloses in ABSTRACT and Section I of Pages 1663-1664 that (1) existing unsupervised anomaly detection models, although they do not depend on the existence of labels to start training, can suffer from high false positive rates, and it is difficult to incorporate human experts’ feedback; (2) introduce a domain-agonistic, end-to-end, ensemble-based deep active learning framework for anomaly detection; (3) by leveraging unsupervised and semi-supervised anomaly detection models, the framework does not rely on labeled data to start training; (4) by using active learning, the framework incorporates human feedback in the workflow to reduce false positive rates, which is essential in many real-world applications; (5) use a neural network to ensemble various anomaly detectors for further performance improvement and leverage a reference score generator to enable an intuitive explanation of anomaly scores; (6) anything that deviates from the norm is by definition an anomaly, and thus anomalies do not have to be similar to the ones already seen; (7) many anomaly detection applications are domain specific, and it is then difficult for a single model to detect anomalies in all scenarios and across different domains; (8) the definition of anomaly could evolve over time, and a current notion of a normal or abnormal behavior might not be sufficiently representative in the future; (9) propose an end-to-end, ensemble-based, deep active learning model that learns to combine predictions from unsupervised or semi-supervised models via active learning; (10) proposed solution can operate with no or few labeled data, which does not rely on a particular definition or assumption of anomaly, and provide the flexibility to integrate as many unsupervised or semi-supervised anomaly detectors as possible to boost model performance; (11) in addition, by leveraging deep neural networks at different stages of our solution, reduce the level of effort on feature engineering, which has traditionally been driven mainly by domain knowledge; (12) last, our proposed solution utilizes active learning to use feedback to improve model performance iteratively and lower false positive rates; and (13) contributions of this paper are summarized as follows: (a) propose a domain-agonistic, end-to-end, ensemble-based deep active learning framework for anomaly detection, which could handle situations of few to no labeled training data, as labels are scarce in practice, this setting enables us to start training without relying on labels and also helps to control the false positive rate by prioritizing the most likely anomalous data points for reviews and solicitating feedback to improve models; and (b) propose a novel way that utilizes a Gaussian prior and Z-score-based deviation loss to ensemble anomaly detectors, which allows better flexibility to incorporate additional anomaly detectors in the workflow and therefore more robust performance when faced with new or unknown datasets and problems. Tang further discloses in Section II of Pages 1664-1665 that (1) depending on label availability, anomaly detection can be broadly categorized into three types: (a) supervised anomaly detection, (b) unsupervised anomaly detection, and (c) semi-supervised anomaly detection; (2) unsupervised anomaly detection identifies anomalies based solely on the intrinsic characteristics of the data; (3) traditional methods include classification-based models, proximity-based models, statistical models, information-theoretic models, and spectral models; (4) deep unsupervised anomaly detection models, such as Autoencoder (AE), first learn data representation by mapping input data to a latent space, then define an anomaly score in the latent space, in the case of AE a reconstructive error, to measure anomaly; (5) compared to traditional models, deep anomaly detection models can capture more complex relationships in the data, and they scale better in a high-dimensional data-rich environment; (6) one can also use a hybrid approach to learn feature representation in a latent space and apply traditional algorithms, like one-class SVM, on the learned features to identify anomalies; (7) many deep unsupervised anomaly detection models can be adapted to deep semi-supervised anomaly detection models; (8) deep semi-supervised anomaly detection models have used sparsely labeled data to either fine tune the decision threshold after feature representation learnt, or to integratedly learn both data representation and anomaly scores; (9) active learning is the process where an algorithm iteratively requests labels from human annotators to enrich training in order to achieve better performance; (10) active learning is potentially a good framework for anomaly detection for several reasons: (a) first, active learning addresses the problem of concept drift where data pattern is constantly evolving; (b) second, the cost of data labeling in anomaly detection can be prohibitive, and active learning addresses the issue of high labeling cost by requesting only a small number of labels to improve model performance; and (c) third, active learning incorporates human experts’ knowledge and experience in the workflow; (11) an ensemble of anomaly detectors is combined with active learning, where relevance of each anomaly detector is learned based on human feedback; and (12) adopt an ensemble-based active learning workflow and leverage a reference score generator to force that the anomaly scores of anomalies are significantly different from those of the normal objects, wherein the reference score generator not only offers a straightforward explanation on the scale of the anomaly score, but also uses information provided by class labels. Tang also discloses in Section III with FIGS. 1-2 of Pages 1665-1667 that (1) the objective of this paper is to train an anomaly detection model by leveraging few to no labeled data with active learning; (2) the ultimate goal is to learn a scoring function that assigns anomaly scores to data points in a way that a higher score indicates a higher chance of being an anomaly, and in an extreme case, where there are no labeled data to begin with; (3) Fig. 1 shows an overall workflow of our proposed approach, which consists of two modules: (a) the first module detects anomalies from raw client data using multiple unsupervised or semi-supervised models, namely anomaly detectors, and depending on the characteristics of client data, different models (e.g., clustering-based models, tree-based models, density-based models, time series-based models, graph-based models, kernel-based models) could be leveraged as long as they belong to either the unsupervised or the semi-supervised learning family, as we assume most client projects begin with few to no labeled historical data; and (b) the second module (the active learning module) takes the outputs from the anomaly detectors and the raw client data to learn a model to classify data points into anomalous or normal, and based on the same reason (few to no labeled data), the active learning model used in this module should belong to the semi-supervised learning family; (4) in addition, it is better to have more models implemented within this module because different models detect anomalies differently, to some extent it increases the diversity of results and yields better performance of an ensemble learner; (5) in addition, the model must work well with highly imbalanced data, as anomalies are scarce; (6) once the active learning model is trained, it will score all data points in the raw client data, wherein each data point will receive an anomaly score where a higher score indicates a higher chance to be an anomaly; (7) there are different query strategies in active learning to sample results for human reviews and solicit feedback, based upon the purpose and requirement; (8) Fig. 2 shows the structure of the proposed model, Ensemble-Based Deep Active Learning Model (EBDALM), which is used in the active learning module, which. consists of three major parts: (a) first use an ensemble learning network to learn an optimal strategy to combine anomaly scores generated by M anomaly detectors; (b) use a reference score generator to produce a reference score to compare the combined anomaly score with, wherein the reference score is generated by first sampling R random numbers from a Gaussian distribution with a pre-defined mean and variance, and then calculate the mean of these R numbers as the reference score; (c) given the combined anomaly score and the reference score, we define a loss function, wherein the purpose of this loss function is to penalize normal data if they deviate from the mean of the prior Gaussian distribution, and also to penalize anomalies if they do NOT deviate far enough from the mean of the prior Gaussian distribution; (9) subsequently, our workflow receives label feedback from SMEs and determines whether the ensemble made an error (i.e., anomalies are classified as normal objects and vice versa); (10) if there is an error, the weights will be updated through backpropagation to suppress all erroneous anomaly detection models with similar inputs in the future; (11) Algorithm 1 illustrates the steps to train our Ensemble-Based Deep Active Learning Model (EBDALM); and (12) Algorithm 2 shows the training algorithm of the entire framework as shown in Fig. 1.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HWEI-MIN LU whose telephone number is (313)446-4913. The examiner can normally be reached Mon - Fri: 9:00 AM - 6:00 PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Mariela D. Reyes can be reached at (571) 270-1006. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HWEI-MIN LU/Primary Examiner, Art Unit 2142