DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of Claims
This action is in reply to the amendment filed on 06/24/2026.
Claims 1-3, 8-10, and 15-17 have been amended and are hereby entered.
Claim 21 has been added.
Claim 20 has been canceled.
Claims 1-19 and 21 are currently pending and have been examined.
This action is made FINAL.
Response to Arguments
Applicant’s arguments, see page 8 of Remarks, filed 06/24/2026, with respect to the drawing objections have been fully considered and are persuasive. The drawing objections have been withdrawn.
Applicant’s arguments, see page 8 of Remarks, filed 06/24/2026, with respect to the specification objections have been fully considered and are persuasive. The specification objections have been withdrawn.
Applicant’s arguments, see page 8 of Remarks, filed 06/24/2026, with respect to the claim objections have been fully considered and are persuasive. The claim objections have been withdrawn.
Applicant’s arguments, see pages 8-9 of Remarks, filed 06/24/2026, with respect to the 35 U.S.C. 101 rejections of claims 1-20 have been fully considered but are not persuasive. The rejection of claim 20 has been withdrawn in light of the cancellation of claim 20, but the 35 U.S.C. 101 rejections of claims 1-19 have been maintained and new claim 21 stands rejected under 35 U.S.C. 101 as well.
First, Applicant argues that the amended claims do not recite a Mental Process across pages 8-9 of Remarks. Applicant argues that the classification of the population set of users and software applications, training of the machine learning classification engines, and execution of the machine learning classification engines in parallel all take the claims out of the realm of Mental Processes. Examiner respectfully disagrees. First, the broadest reasonable interpretation of “clustering analysis”, as no detail is recited into the claims about how the clustering is performed, covers clustering users and software programs in a manner that falls within the realm of mental processes. Indeed, Applicant’s specification [0057] recites “In the preprocessing phase, each of user data, application data and other data may be classified. Users may be classified as employees, and Human Resources (HR) records may be used as a population. Users/employees may be classified multiple times, with the resulting set of class labels stored for each in classification data store 1504. For example, each employee may be classified with respect to demographics, qualifications, job, job history, access level (e.g. permissions), and the like” and [0058] “Similarly, internal and external applications that are available to employees/users may also be classified multiple times with respect to criteria such as access, data, security, budget, integration pattern, and the like”. One of ordinary skill in the art would be capable of observing data/descriptions of users and software programs and classifying the users in software programs according to job title, demographics, or type of software application using their mind and possibly a pen and paper (see MPEP 2106.04(a)(2) III. B.) to generate multiple classifications. Regarding the training and executing of the machine learning classification models, Examiner points to MPEP 2106.04(a)(2) III.C. Here, the MPEP states “examiners should review the specification to determine if the claimed invention is described as a concept that is performed in the human mind and applicant is merely claiming that concept performed 1) on a generic computer, or 2) in a computer environment, or 3) is merely using a computer as a tool to perform the concept.” In the context of the claimed invention, the process being performed by the training and execution of the machine learning classification engines is assessing whether an event is anomalous and a confidence level in the determination. The “events” in question, under their broadest reasonable interpretation, include user actions, accessing data, security-related events, etc. (see specification [0048] and [0076]). Accordingly, a human could review a user’s access/actions in an application and determine whether, based on the group, job description, etc. of the user, the user’s actions were anomalous and how confident the person is in their assessment of whether the action is suspicious. Multiple people could make their judgment on whether the action is anomalous, resulting in a final anomaly score and a corresponding confidence in the determination.
The claimed invention, instead of using humans to assess whether the actions are anomalous, are using machine learning classification engines. No detail regarding the actual operation of machine learning classification engines is provided in the claim; claim 1 only recites that the event being audited goes in and anomaly and confidence scores come out. Paragraph [0081] of the specification alludes to adjusting parameters of the models, but [0082] recites “That being said, embodiments of the solution described herein rely on multiple different anomaly detection models, each trained and/or designed to focus on a specific goal”. As the models are merely described as “each trained and/or designed to focus on a specific goal” without further recitation of the design of the models beyond the various potential types of models in [0083], the machine learning classification engines are being used at a high level as tools to perform the mental process of determining whether an event is anomalous or not. Therefore, per MPEP 2106.04(a)(2) III.C., the claims recite a Mental Process.
Next, Applicant argues that the claims do not recite the Certain Method of Organizing Human Activity of the fundamental economic practice of mitigating risks. Applicant argues that the claims are directed to a specific technical system for anomaly detection. Examiner respectfully disagrees. First, Examiner refers back to the discussion above in which the machine learning classification engines are merely described as being designed to achieve a particular goal, with no further technical detail as to the construction or operation of the engines. Accordingly, the machine learning models do not recite a specific technical system that would preclude the claim from reciting the judicial exception. Next, Examiner points to specification [0003]-[0008], in which Applicant states that the present invention is directed to ensuring compliance with regulatory and compliance requirements. Ensuring regulatory compliance is mitigating the risk that the organization fails to comply with regulations and is subject to corresponding penalties/sanctions/etc. Accordingly, the claimed invention recites the fundamental economic concept of mitigating risks. As far as whether the claims are directed to the abstract ideas, analysis proceeds to Step 2A Prong Two.
Applicant next argues on page 9 that the claimed invention improves the functioning of a computer. Specifically, Applicant argues that the claimed invention improves upon the inaccuracy of conventional anomaly detection systems by classifying entity populations, filtering compliance events into subsets, and training multiple machine learning models to output confidence and anomaly scores. Applicant further argues that the machine learning models being trained on context-specific data sets, and the anomaly and confidence score outputs also improve accuracy by giving more weight to appropriate models and focusing model training on data more relevant to an event. Examiner respectfully disagrees that the claimed invention provides a technical improvement. MPEP 2106.05(a) II. recites “However, it is important to keep in mind that an improvement in the abstract idea itself (e.g. a recited fundamental economic concept) is not an improvement in technology”. In the context of the claimed invention, the classification of the population, filtering the compliance events into subsets (which Examiner notes is recited only in claim 21, not the independent claims), and using context-specific training data to train models are improvements to the judicial exception itself. Namely, the improvements of increased accuracy stem from a more targeted use of data than the generalized approach of conventional systems per specification [0043]. Clustering data and only using data more relevant to the particular event is not a technical improvement (i.e. improvement to computers, machine learning) but rather an improved usage of information in the abstract idea. The structures of the machine learning models themselves are not being improved in the claimed invention, but rather they are being fed more pertinent information, which leads to a more pertinent output. As discussed above, [0082]-[0083] of the specification recite that the structure of the machine learning models used depends on the goal of the model. Paragraph [0084] further states “In some embodiments, the particular anomaly detection algorithm used by ML detectors 1904 is not of particular materiality”. Accordingly, one of ordinary skill in the art would not have recognized the claimed invention as providing a technical improvement to machine learning itself. See MPEP 2106.05(f) “The recitation of claim limitations that attempt to cover any solution to an identified problem with no restriction on how the result is accomplished and no description of the mechanism for accomplishing the result, does not integrate a judicial exception into a practical application or provide significantly more because this type of recitation is equivalent to the words "apply it””. In the present case the machine learning models of the claimed invention are not particular models and are attempting to cover any machine learning model that can output an anomaly and confidence score. The training and execution of the machine learning models in parallel would be recognized as using the machine learning models as tools to arrive at the anomaly and confidence scores. See MPEP 2106.05 (f) “Use of a computer or other machinery in its ordinary capacity for economic or other tasks (e.g., to receive, store, or transmit data) or simply adding a general purpose computer or computer components after the fact to an abstract idea (e.g., a fundamental economic practice or mathematical equation) does not integrate a judicial exception into a practical application or provide significantly more”. In sum, the improvements in accuracy over conventional methods discussed by Applicant arise from the abstract idea, with the additional elements themselves being used as tools to perform the abstract idea and not providing a technical improvement.
Finally, Applicant argues that the claimed invention should be patent eligible for similar reasoning as that of SRI, namely that the instant claims are allegedly analogous to SRI’s detection of suspicious network activity by using network monitors to analyze network packets. Examiner respectfully disagrees. First, Examiner notes that the claimed invention is analyzing audit events that are “received” by the invention without an analogous structure to the system of network monitors in SRI. Additionally, the audit events themselves are not analogues to the network packets being monitored in SRI, as an audit event covers any action taken by a user and is not a particular part of network traffic as in SRI. As discussed above, the instant claims also do not recite a specific mechanism for identifying anomalous events because the algorithms of the machine learning classification engines processing events to determine anomaly and confidence scores are “not of particular materiality” per [0084] in the specification. Accordingly, Applicant’s argument that the claimed invention is patent eligible by analogous reasoning as that in SRI is not persuasive. The 35 U.S.C. 101 rejections have been maintained.
Applicant’s arguments, see pages 9-11 of Remarks, filed 06/24/2026, with respect to the 35 U.S.C. 103 rejections of claims 1-20 have been fully considered but are either moot or not persuasive. The rejection of claim 20 has been withdrawn in light of the cancellation of claim 20, but claims 1-19 still stand rejected under 35 U.S.C. 103, and new claim 21 stands rejected under 35 U.S.C. 103 as well.
On Pages 9-10 of Remarks Applicant argues that the Salunke reference fails to teach the amended pre-processing phase of claims 1 and new claim 21 and that there is no motivation to combine Torres Dho with Salunke. These arguments are moot, as the Williams, Jr. reference (U.S. Pre-Grant Publication No. 2015/0254555, hereafter known as Williams, Jr.) is used below to teach the amended pre-processing phase of amended claim 1 and new claim 21.
Next, Applicant argues that the Adamson reference does not teach the training of the plurality of the machine learning models. Examiner notes that Adamson was not and is not being used to teach the training process of the plurality of machine learning models. In the previous rejection of claim 1 the combination of Torres Dho and Salunke, and in the current rejection below the combination of Torres Dho and Williams, Jr., teaches the training of the machine learning classification engines. The Adamson reference was/is used to remedy the deficiency that the combination of Torres Dho and Salunke/Williams, Jr. do not explicitly teach that the events being analyzed are “compliance and audit” events. See “Adamson teaches the events being tracked by the system as compliance and audit events” in the rejection of claim 1 below and Page 21 of the 02/24/2026 Non-Final Rejection. Thus, Adamson is not being relied upon to teach the training process as Applicant is arguing on pages 10-11, and is instead merely being relied upon to teach that the type of events being analyzed for anomalies are “compliance and audit” events. Adamson explicitly teaches in Col. 5 line 63 thru Col. 6 line 3 that anomalies are being searched for as a part of compliance monitoring operations, which teaches the limitation of “compliance and audit” data that Adamson is being relied upon to teach. Applicant’s argument against the Adamson reference is not persuasive.
Finally, Applicant argues on Page 11 that Torres Dho does not teach the models outputting confidence scores as described in the amended limitation of claim 1. Examiner notes that the Williams, Jr. is also used to teach the newly added limitation of “wherein said confidence score represents a degree of confidence in the accuracy of said anomaly score” as is shown in the rejection of claim 1 below. Accordingly, Applicant’s argument regarding Torres Dho not teaching the amended confidence score is also moot.
Applicant’s arguments that the Muddu and Kumar references do not remedy the alleged deficiencies of Torres Dho, Salunke, and Adamson are also moot because the Muddu and Kumar references are not used to teach any of the argued limitations.
Accordingly, all of Applicant’s arguments are either moot or unpersuasive. Claims 1-19 and 21 stand rejected under 35 U.S.C. 103.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-19 and 21 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claims recite determining whether an event is anomalous.
As an initial matter, claims 1-7 and 21 fall into at least the process category of statutory subject matter. Claims 8-14 fall into at least the machine category of statutory subject matter. Finally, claims 15-19 fall into at least the manufacture category of statutory subject matter. Therefore, all claims fall into at least one of the statutory categories. Eligibility analysis proceeds to Step 2A.
In claim 1, the limitation of “A method of detecting anomalous behaviour in a network, the method comprising: in a pre-processing phase, classifying a population set comprising a plurality of users and a plurality of software applications to discover context-specific classes via a clustering analysis of each sub-population within said population set and storing said classification in a population data store, wherein each of said sub-populations is classified multiple times with respect to a plurality of classification context spaces”, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. That is, other than reciting “a network,” nothing in the claim element precludes the step from practically being performed in the mind. Similarly, the limitations of “receiving a current audit event; processing, by said plurality of machine learning classification engines executing in parallel, said current audit event to obtain an anomaly score and a confidence score from each respective classification engine, wherein said anomaly score represents a probability that said current audit event is anomalous, and wherein said confidence score represents a degree of confidence in the accuracy of said anomaly score; determine whether said current audit event is an anomalous event based on said respective anomaly scores and confidence scores”, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claims recite an abstract idea.
Additionally, claim 1 recites the concept of identifying risky anomalous actions in a network which is a certain method of organizing human activity including fundamental economic principles including mitigating risks. A method of detecting anomalous behaviour, the method comprising: in a pre-processing phase, classifying a population set comprising a plurality of users and a plurality of software applications to discover context-specific classes via a clustering analysis of each sub-population within said population set and storing said classification in a population data store, wherein each of said sub-populations is classified multiple times with respect to a plurality of classification context spaces; receiving a current audit event; processing said current audit event to obtain an anomaly score and a confidence score from each respective classification engine, wherein said anomaly score represents a probability that said current audit event is anomalous, and wherein said confidence score represents a degree of confidence in the accuracy of said anomaly score; determine whether said current audit event is an anomalous event based on said respective anomaly scores and confidence scores all, as a whole, fall under the category of fundamental economic principles including mitigating risks. The claim falls into the “Certain Methods of Organizing Human Activity” grouping of abstract ideas. Mere recitation of generic computer components does not remove the claim from this grouping. Accordingly, the claim recites an abstract idea.
This judicial exception is not integrated into a practical application. In particular, the claim recites the additional elements of a network, training a plurality of machine learning classification engines, executing a plurality of machine learning classification engines in parallel, and a software application. The recited additional elements are recited at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using generic computer components and generic machine learning. Accordingly, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The combination of these additional elements is also no more than mere instructions to apply the exception using generic computer components. Accordingly, even in combination, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements of a network, training a plurality of machine learning classification engines, executing a plurality of machine learning classification engines in parallel, and a software application amounts to no more than mere instructions to apply the exception using generic computer components. The combination of these additional elements is also no more than mere instructions to apply the exception using generic computer components. Mere instructions to apply an exception using generic computer components cannot provide an inventive concept. The claim is not patent eligible.
Claim 2 further limits the abstract idea of claim 1 while introducing the additional element of the plurality of machine learning classification models comprising one or more ensemble models. The claim does not integrate the abstract idea into a practical application because the element of the plurality of machine learning classification models comprising one or more ensemble models is recited at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using generic computer components and generic machine learning. Adding this new additional element into the additional element from claim 1 still amounts to no more than mere instructions to apply the exception using generic computer components and generic machine learning. The claim also does not amount to significantly more than the abstract idea because mere instructions to apply an exception using generic computer components and generic machine learning cannot provide an inventive concept. The claim is not patent eligible.
Claims 3 further limits the abstract idea of claim 1 without adding any new additional elements. Therefore, by the analysis of claim 1 above, this claim does not integrate the abstract idea into a practical application nor amount to significantly more than the abstract idea. The claim is not patent eligible.
Claim 4 further limits the abstract idea of claim 1 while introducing the additional element of the modification of machine learning classification models in response to assessment accuracy. The claim does not integrate the abstract idea into a practical application because the element of the modification of machine learning classification models in response to assessment accuracy is recited at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using generic computer components and generic machine learning. Adding this new additional element into the additional element from claim 1 still amounts to no more than mere instructions to apply the exception using generic computer components and generic machine learning. The claim also does not amount to significantly more than the abstract idea because mere instructions to apply an exception using generic computer components and generic machine learning cannot provide an inventive concept. The claim is not patent eligible.
Claims 5 further limits the abstract idea of claim 1 without adding any new additional elements. Therefore, by the analysis of claim 1 above, this claim does not integrate the abstract idea into a practical application nor amount to significantly more than the abstract idea. The claim is not patent eligible.
Claim 6 further limits the abstract idea of claim 1 while introducing the additional element of the machine learning classification models operating in parallel and independently from each other. The claim does not integrate the abstract idea into a practical application because the element of the machine learning classification models operating in parallel and independently from each other is recited at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using generic computer components and generic machine learning. Adding this new additional element into the additional element from claim 1 still amounts to no more than mere instructions to apply the exception using generic computer components and generic machine learning. The claim also does not amount to significantly more than the abstract idea because mere instructions to apply an exception using generic computer components and generic machine learning cannot provide an inventive concept. The claim is not patent eligible.
Claim 7 further limits the abstract idea of claim 4 while introducing the additional element of the parallel and separately executed feedback loops for each of the machine learning classification models. The claim does not integrate the abstract idea into a practical application because the element of the parallel and separately executed feedback loops for each of the machine learning classification models is recited at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using generic computer components and generic machine learning. Adding this new additional element into the additional element from claim 4 still amounts to no more than mere instructions to apply the exception using generic computer components and generic machine learning. The claim also does not amount to significantly more than the abstract idea because mere instructions to apply an exception using generic computer components and generic machine learning cannot provide an inventive concept. The claim is not patent eligible.
In claim 8, the limitation of “A system for detecting anomalous behaviour in a network, the system comprising: one or more processors; a non-transitory computer-readable storage medium having stored thereon processor-executable instructions that, when executed by said one or more processors, cause the one or more processors to perform a method comprising: in a pre-processing phase, classifying a population set comprising a plurality of users and a plurality of software applications to discover context-specific classes via a clustering analysis of each sub- population within said population set and storing said classification in a population data store, wherein each of said sub-populations is classified multiple times with respect to a plurality of classification context spaces”, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. That is, other than reciting “a system”, “a network”, “one or more processors”, “a non-transitory computer-readable storage medium having stored thereon processor-executable instructions that, when executed by said one or more processors, cause the one or more processors to perform a method,” nothing in the claim element precludes the step from practically being performed in the mind. Similarly, the limitations of “receiving a current audit event; processing, by said plurality of machine learning classification engines executing in parallel, said current audit event to obtain an anomaly score and a confidence score from each respective classification engine, wherein said anomaly score represents a probability that said current audit event is anomalous, and wherein said confidence score represents a degree of confidence in the accuracy of said anomaly score; determine whether said current audit event is an anomalous event based on said respective anomaly scores and confidence scores”, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claims recite an abstract idea.
Additionally, claim 8 recites the concept of identifying risky anomalous actions in a network which is a certain method of organizing human activity including fundamental economic principles including mitigating risks. Detecting anomalous behaviour, perform a method comprising: in a pre-processing phase, classifying a population set comprising a plurality of users and a plurality of software applications to discover context-specific classes via a clustering analysis of each sub-population within said population set and storing said classification in a population data store, wherein each of said sub-populations is classified multiple times with respect to a plurality of classification context spaces; receiving a current audit event; processing said current audit event to obtain an anomaly score and a confidence score from each respective classification engine, wherein said anomaly score represents a probability that said current audit event is anomalous, and wherein said confidence score represents a degree of confidence in the accuracy of said anomaly score; determine whether said current audit event is an anomalous event based on said respective anomaly scores and confidence scores all, as a whole, fall under the category of fundamental economic principles including mitigating risks. The claim falls into the “Certain Methods of Organizing Human Activity” grouping of abstract ideas. Mere recitation of generic computer components does not remove the claim from this grouping. Accordingly, the claim recites an abstract idea.
This judicial exception is not integrated into a practical application. In particular, the claim recites the additional elements of a system, a network, one or more processors, a non-transitory computer-readable storage medium having stored thereon processor-executable instructions that, when executed by said one or more processors, cause the one or more processors to perform a method, training a plurality of machine learning classification engines, executing a plurality of machine learning classification engines in parallel, and a software application. The recited additional elements are recited at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using generic computer components and generic machine learning. Accordingly, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The combination of these additional elements is also no more than mere instructions to apply the exception using generic computer components. Accordingly, even in combination, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements of a system, a network, one or more processors, a non-transitory computer-readable storage medium having stored thereon processor-executable instructions that, when executed by said one or more processors, cause the one or more processors to perform a method, training a plurality of machine learning classification engines, executing a plurality of machine learning classification engines in parallel, and a software application amounts to no more than mere instructions to apply the exception using generic computer components. The combination of these additional elements is also no more than mere instructions to apply the exception using generic computer components. Mere instructions to apply an exception using generic computer components cannot provide an inventive concept. The claim is not patent eligible.
Claim 9 further limits the abstract idea of claim 8 while introducing the additional element of the plurality of machine learning classification models comprising one or more ensemble models. The claim does not integrate the abstract idea into a practical application because the element of the plurality of machine learning classification models comprising one or more ensemble models is recited at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using generic computer components and generic machine learning. Adding this new additional element into the additional element from claim 8 still amounts to no more than mere instructions to apply the exception using generic computer components and generic machine learning. The claim also does not amount to significantly more than the abstract idea because mere instructions to apply an exception using generic computer components and generic machine learning cannot provide an inventive concept. The claim is not patent eligible.
Claims 10 further limits the abstract idea of claim 8 without adding any new additional elements. Therefore, by the analysis of claim 8 above, this claim does not integrate the abstract idea into a practical application nor amount to significantly more than the abstract idea. The claim is not patent eligible.
Claim 11 further limits the abstract idea of claim 8 while introducing the additional element of the modification of machine learning classification models in response to assessment accuracy. The claim does not integrate the abstract idea into a practical application because the element of the modification of machine learning classification models in response to assessment accuracy is recited at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using generic computer components and generic machine learning. Adding this new additional element into the additional element from claim 8 still amounts to no more than mere instructions to apply the exception using generic computer components and generic machine learning. The claim also does not amount to significantly more than the abstract idea because mere instructions to apply an exception using generic computer components and generic machine learning cannot provide an inventive concept. The claim is not patent eligible.
Claims 12 further limits the abstract idea of claim 8 without adding any new additional elements. Therefore, by the analysis of claim 8 above, this claim does not integrate the abstract idea into a practical application nor amount to significantly more than the abstract idea. The claim is not patent eligible.
Claim 13 further limits the abstract idea of claim 8 while introducing the additional element of the machine learning classification models operating in parallel and independently from each other. The claim does not integrate the abstract idea into a practical application because the element of the machine learning classification models operating in parallel and independently from each other is recited at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using generic computer components and generic machine learning. Adding this new additional element into the additional element from claim 8 still amounts to no more than mere instructions to apply the exception using generic computer components and generic machine learning. The claim also does not amount to significantly more than the abstract idea because mere instructions to apply an exception using generic computer components and generic machine learning cannot provide an inventive concept. The claim is not patent eligible.
Claim 14 further limits the abstract idea of claim 11 while introducing the additional element of the parallel and separately executed feedback loops for each of the machine learning classification models. The claim does not integrate the abstract idea into a practical application because the element of the parallel and separately executed feedback loops for each of the machine learning classification models is recited at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using generic computer components and generic machine learning. Adding this new additional element into the additional element from claim 11 still amounts to no more than mere instructions to apply the exception using generic computer components and generic machine learning. The claim also does not amount to significantly more than the abstract idea because mere instructions to apply an exception using generic computer components and generic machine learning cannot provide an inventive concept. The claim is not patent eligible.
In claim 15, the limitation of “A non-transitory computer-readable storage medium having stored thereon processor-executable instructions that, when executed by one or more processors, cause the one or more processors to perform a method comprising: in a pre-processing phase, classifying a population set comprising a plurality of users and a plurality of software applications to discover context-specific classes via a clustering analysis of each sub-population within said population set and storing said classification in a population data store, wherein each of said sub-populations is classified multiple times with respect to a plurality of classification context spaces”, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. That is, other than reciting “a non-transitory computer-readable storage medium having stored thereon processor-executable instructions that, when executed by one or more processors, cause the one or more processors to perform a method”, “one or more processors,” nothing in the claim element precludes the step from practically being performed in the mind. Similarly, the limitations of “receiving a current audit event; processing, by said plurality of machine learning classification engines executing in parallel, said current audit event to obtain an anomaly score and a confidence score from each respective classification engine, wherein said anomaly score represents a probability that said current audit event is anomalous, and wherein said confidence score represents a degree of confidence in the accuracy of said anomaly score; determine whether said current audit event is an anomalous event based on said respective anomaly scores and confidence scores”, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claims recite an abstract idea.
Additionally, claim 15 recites the concept of identifying risky anomalous actions in a network which is a certain method of organizing human activity including fundamental economic principles including mitigating risks. Perform a method comprising: in a pre-processing phase, classifying a population set comprising a plurality of users and a plurality of software applications to discover context-specific classes via a clustering analysis of each sub-population within said population set and storing said classification in a population data store, wherein each of said sub-populations is classified multiple times with respect to a plurality of classification context spaces; receiving a current audit event; processing said current audit event to obtain an anomaly score and a confidence score from each respective classification engine, wherein said anomaly score represents a probability that said current audit event is anomalous, and wherein said confidence score represents a degree of confidence in the accuracy of said anomaly score; determine whether said current audit event is an anomalous event based on said respective anomaly scores and confidence scores all, as a whole, fall under the category of fundamental economic principles including mitigating risks. The claim falls into the “Certain Methods of Organizing Human Activity” grouping of abstract ideas. Mere recitation of generic computer components does not remove the claim from this grouping. Accordingly, the claim recites an abstract idea.
This judicial exception is not integrated into a practical application. In particular, the claim recites the additional elements of one or more processors, a non-transitory computer-readable storage medium having stored thereon processor-executable instructions that, when executed by said one or more processors, cause the one or more processors to perform a method, training a plurality of machine learning classification engines, executing a plurality of machine learning classification engines in parallel, and a software application. The recited additional elements are recited at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using generic computer components and generic machine learning. Accordingly, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The combination of these additional elements is also no more than mere instructions to apply the exception using generic computer components. Accordingly, even in combination, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements of one or more processors, a non-transitory computer-readable storage medium having stored thereon processor-executable instructions that, when executed by said one or more processors, cause the one or more processors to perform a method, training a plurality of machine learning classification engines, executing a plurality of machine learning classification engines in parallel, and a software application amounts to no more than mere instructions to apply the exception using generic computer components. The combination of these additional elements is also no more than mere instructions to apply the exception using generic computer components. Mere instructions to apply an exception using generic computer components cannot provide an inventive concept. The claim is not patent eligible.
Claim 16 further limits the abstract idea of claim 15 while introducing the additional element of the plurality of machine learning classification models comprising one or more ensemble models. The claim does not integrate the abstract idea into a practical application because the element of the plurality of machine learning classification models comprising one or more ensemble models is recited at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using generic computer components and generic machine learning. Adding this new additional element into the additional element from claim 15 still amounts to no more than mere instructions to apply the exception using generic computer components and generic machine learning. The claim also does not amount to significantly more than the abstract idea because mere instructions to apply an exception using generic computer components and generic machine learning cannot provide an inventive concept. The claim is not patent eligible.
Claims 17 further limits the abstract idea of claim 15 without adding any new additional elements. Therefore, by the analysis of claim 15 above, this claim does not integrate the abstract idea into a practical application nor amount to significantly more than the abstract idea. The claim is not patent eligible.
Claim 18 further limits the abstract idea of claim 15 while introducing the additional element of the modification of machine learning classification models in response to assessment accuracy. The claim does not integrate the abstract idea into a practical application because the element of the modification of machine learning classification models in response to assessment accuracy is recited at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using generic computer components and generic machine learning. Adding this new additional element into the additional element from claim 15 still amounts to no more than mere instructions to apply the exception using generic computer components and generic machine learning. The claim also does not amount to significantly more than the abstract idea because mere instructions to apply an exception using generic computer components and generic machine learning cannot provide an inventive concept. The claim is not patent eligible.
Claim 19 further limits the abstract idea of claim 15 while introducing the additional element of the machine learning classification models operating in parallel and independently from each other. The claim does not integrate the abstract idea into a practical application because the element of the machine learning classification models operating in parallel and independently from each other is recited at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using generic computer components and generic machine learning. Adding this new additional element into the additional element from claim 15 still amounts to no more than mere instructions to apply the exception using generic computer components and generic machine learning. The claim also does not amount to significantly more than the abstract idea because mere instructions to apply an exception using generic computer components and generic machine learning cannot provide an inventive concept. The claim is not patent eligible.
Claim 21 further limits the abstract idea of claim 1 without adding any new additional elements. Therefore, by the analysis of claim 1 above, this claim does not integrate the abstract idea into a practical application nor amount to significantly more than the abstract idea. The claim is not patent eligible.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 3, 6, 8, 10, 13, 15, 17, 19, and 21 are rejected under 35 U.S.C. 103 as being unpatentable over Torres Dho et al. (U.S. Pre-Grant Publication No. 2023/0319083, hereafter known as Torres Dho) in view of Williams, Jr. et al. (U.S. Pre-Grant Publication No. 2015/0254555, hereafter known as Williams, Jr.) and Adamson et al. (U.S. Patent No. 12,341,797; hereafter known as Adamson).
Regarding claim 1, Torres Dho teaches:
A method of detecting anomalous behaviour in a network, the method comprising (see Fig. 2 and [0039]-[0051] for the overall method. See Fig. 1 and [0021]-[0022] for the host devices being monitored for anomalies being a part of a computer network)
receiving a current event (see steps 202 and 206 and [0040] "at operation 202, the process identifies current values for a set of metrics from a set of hosts. For example, the process may scan log files received from hosts 102a-n for entries with the most recent timestamp. As another example, the process may detect and record the most recent samples streamed from hosts 102a-n" and [0042] "At operation 206, the process generates a set of one or more current point-in-time value data frames. A data frame is a data structure, which may be implemented as a table or multidimensional array for storing the point-in-time values. With current point-in-time value data frames, a column in the data structure may correspond to different metrics, such as CPU utilization rates, memory throughput, active user sessions, etc.")
processing, by said plurality of machine learning classification engines executing in parallel, said current (see step 212 and [0045] "At operation 212, the process applies a set of outlier detection ML models to each data frame...Example ML algorithms/models include angle-based outlier detection (ABOD), clustering models (e.g., k-means clustering, k-mode clustering), k-nearest neighbors (KNN), principal component analysis (PCA), and support vector machines (SVM)...The ML model may classify the behavior of a host as an outlier or non-outlier based on the set of point-in-time values for the host relative to patterns in point-in-time values for other hosts" and [0046] "each of the applied ML models outputs a per-host value or score for each data frame that indicates a classification and/or probability that the host's behavior is an outlier...Probabilistic models may assign values between 0-1 based on the probability that the value is an outlier with 1 indicating a 100% probability, 0 indicating a 0% percent probability" for processing a current event state with a plurality of machine learning models and obtaining a probability (anomaly score) that the current event is anomalous. See Fig. 3 element 306 and [0054] "For each data frame, all of ML algorithms 306 are run to estimate a classification or probabilistic value. With AB OD, for instance, the process may identify outliers based on the variances of the angles and distances between data points within a data frame...With clustering, outliers may be detected based on the distance between the host (represented by the host's point-in-time values) at a point in time and the nearest cluster centroid. With KNN, outliers may be detected based on the distance between the host and the k nearest neighbors. With PCA, outliers may be detected based on a decomposition (e.g., an eigendecomposition) of the values into principal components and variance between the host's principal components from the principal components of other hosts. With SVM, outliers may be detected based on the position of a host relative to a hyperplane or boundary" for the different machine learning models running in parallel at step 306 and independently identifying anomalies using their own techniques)
determine whether said current (see steps 214 and 216 and [0047] "At operation 214, the process generates a set of anomaly scores for each host based on the output of the machine learning algorithms", [0048] for the weighting of contributions of the anomaly probabilities of each model based on the reliability of each particular model to obtain the anomaly score for a host, and [0050] "At operation 216, the process determines whether any hosts have an anomaly score satisfying a threshold value" for determining whether a current state is anomalous based on the overall score for the current events surpassing a threshold)
While Torres Dho teaches a plurality of machine learning models processing a current event in order to determine whether the event was anomalous, Torres Dho does not explicitly teach the pre-processing phase of classifying a population set to discover context-specific classes via clustering analysis, storing the classification in a population data store, and using the population data store and compliance and audit events to train the machine learning models. Torres Dho also does not explicitly teach the current event being processed as an audit event and the confidence score being representative of a degree of confidence in the anomaly score. Williams, Jr. teaches:
in a pre-processing phase, classifying a population set comprising a plurality of users and a plurality of software applications to discover context-specific classes via a clustering analysis of each sub-population within said population set and storing said classification in a population data store, wherein each of said sub-populations is classified multiple times with respect to a plurality of classification context spaces (see [0212]-[0214] for segmenting data into a per-user, per-user-group, and per-data-resource groupings. Also grouped by types and locations of data being accessed. Grouping users and data at multiple levels reads on the broadest reasonable interpretation of using a clustering analysis to classify sub-populations of users/data multiple times. See [0217] for additional context-specific information identified at Data Ingestion 504. For storing classifications in a data store, see [0220] “In at least one of the various embodiments, as new data is ingested, it may be added to the historical access records for users, groups, and content areas. The new data may be periodically added to the Training Corpus 508 containing sequences of access records” and [0222]. Also see [0084] “Data may be any medium, including but not limited to, hardcopy documents that have been scanned, photographs, digital files and media, sensor data, log files, survey data, database records, program code, or the like” by which Examiner is interpreting that the data being accessed is program code which reads on “a software application”)
training a plurality of machine learning classification engines based on said population data store and on received (see [0212]-[0215] for training anomaly detection models on data segmented into clusters and event patterns for the groups. Also see [0220] for retraining the models with new training data periodically)
obtain an anomaly score and a confidence score from each respective classification engine,…and wherein said confidence score represents a degree of confidence in the accuracy of said anomaly score… determine whether said current (see [0132] “In at least one of the various embodiments, one or more of the classifications/classes may be defined to be anomalies. In other cases, classes may be considered normal and/or expected” for classifying data as anomalous or normal, [0120]-[0122] including “In at least one of the various embodiments, each classification result generated by the fast learning model may be associated with a value (the confidence value or confidence score) that scores/ranks the how close the data matches the classifier”, and [0175]-[0176] for the generation of confidence scores along with anomaly classifications in each of multiple models to determine whether the event is anomalous)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the preprocessing of the users and accessed programs into clusters and training anomaly detection machine learning models based on the clusters of Williams, Jr. into the system of Torres Dho. As Williams, Jr. states in [0215] “per-user training may make detection of anomalous behavior more accurate and reduce false positives. Similarly, training using user-group access patterns, where the group may be defined by characteristics such as those who are in a similar job role, or users in the same or related organization unit(s), may enable detection of anomalies that are divergent relative to the group's normal behaviors. Training on data-resource access patterns may be designed to detect anomalies in access patterns to a group of related data, because related data is often accessed in similar, consistent ways.” Accordingly, one of ordinary skill in the art would have recognized that by incorporating the clustering of user and software access data and then training models based on the clustered data would have resulted in improved accuracy of the anomaly detection models of Torres Dho. Examiner notes here that the improved accuracy of the models provided by clustering is a motivation from Applicant’s specification as well, but the motivation to improve model accuracy is not impermissible hindsight because the Williams, Jr. reference explicitly calls attention to the improved accuracy provided by the clustering. Accordingly, one of ordinary skill in the art would have had the motivation to combine Torres Dho and Williams, Jr. without relying on Applicant’s specification.
It also would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the models outputting confidence scores reflecting a confidence in the model’s anomaly assessment of Williams, Jr. into the system of Torres Dho. As Williams Jr. discusses in [0121]-[0122], confidence scores reflect how well the data being processed matches the classifiers and the training data that was used to train the classifiers. By using a confidence score output from each model reflective of the model’s confidence in its classification of anomalous or normal data, paragraph [0176] recites that the higher confidence model can be used to generate the output to increase accuracy. In combination with Torres Dho, one of ordinary skill in the art would have recognized that outputting confidence values would allow the collection of models to be weighted according to confidence to provide a final determination as to whether the event was anomalous or not, with weight being responsive to how the event fits the models’ training.
The combination of Torres Dho and Williams, Jr. still does not explicitly teach the events being used to train the machine learning models and processed by the machine learning models as “compliance and audit” events. Adamson teaches the events being tracked by the system as compliance and audit events (see Col. 5 line 63 thru Col. 6 line 3 “Data processing resources 20 may be configured to perform various data processing operations with respect to data ingested by data ingestion resources 18, including data ingested and stored in data store 30. For example, data processing resources 20 may be configured to perform one or more data security monitoring and/or remediation operations, compliance monitoring operations, anomaly detection operations” and Col. 29 lines 20-27 “Two example kinds of anomalies that can be detected by data platform 12 include security anomalies (e.g., a user or process behaving in an unexpected manner) and devops/root cause anomalies (e.g., network congestion, application failure, etc.). Detected anomalies can be recorded and surfaced (e.g., to administrators, auditors, etc.), such as through alerts which are generated at 304 based on anomaly detection” for data being processed in the system being compliance and audit data. In combination with Torres Dho and Williams, Jr., the compliance and audit data are processed by the machine learning models and used to train the machine learning models).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include the data being used to train machine learning models and the data being processed by machine learning models to look for anomalies being compliance and audit data as taught by Adamson in the combination of Torres Dho and Williams, Jr., since the claimed invention is merely a combination of old elements, and in the combination each element merely would have performed the same function as it did separately, and one of ordinary skill in the art would have recognized that the results of the combination were predictable.
As discussed in Adamson, anomaly detection and compliance monitoring/auditing can be performed at the same time, with anomalies in a network being used as part of the compliance monitoring and auditing of network operations. Therefore, it would have been obvious to incorporate such compliance monitoring and auditing of Adamson into the combination of Torres Dho and Williams, Jr.
Regarding claim 3, the combination of Torres Dho, Williams, Jr., and Adamson teaches all of the limitations of claim 1 above. Torres Dho further teaches:
wherein said one or more machine learning classification engines is configured to detect anomalies in one or more of a user, an application in use, a location, a job role, and/or a demographic (see [0024] "Example metric readings may include...number of active user sessions" and [0023] "Example computing resources may include...application instances, and virtual machine instances" for metrics being analyzed for anomalies including users and applications in use)
Regarding claim 6, the combination of Torres Dho, Williams, Jr., and Adamson teaches all of the limitations of claim 1 above. Torres Dho further teaches:
wherein each of said plurality of machine learning classification engines executes in parallel and independently from other machine learning classification engines of said plurality of machine learning classification engines (see Fig. 3 element 306 and [0054] "For each data frame, all of ML algorithms 306 are run to estimate a classification or probabilistic value. With AB OD, for instance, the process may identify outliers based on the variances of the angles and distances between data points within a data frame...With clustering, outliers may be detected based on the distance between the host (represented by the host's point-in-time values) at a point in time and the nearest cluster centroid. With KNN, outliers may be detected based on the distance between the host and the k nearest neighbors. With PCA, outliers may be detected based on a decomposition (e.g., an eigendecomposition) of the values into principal components and variance between the host's principal components from the principal components of other hosts. With SVM, outliers may be detected based on the position of a host relative to a hyperplane or boundary" for the different machine learning models running in parallel at step 306 and independently identifying anomalies using their own techniques)
Regarding claim 8, Torres Dho teaches:
A system for detecting anomalous behaviour in a network, the system comprising: one or more processors (see Fig. 5 and [0101] “FIG. 5 illustrates a computer system upon which some embodiments may be implemented. Computer system 500 includes a bus 502 or other communication mechanism for communicating information, and a hardware processor 504 coupled with bus 502 for processing information. Hardware processor 504 may be, for example, a general-purpose microprocessor”)
a non-transitory computer-readable storage medium having stored thereon processor-executable instructions that, when executed by said one or more processors, cause the one or more processors to perform a method comprising (see [0102] “Computer system 500 also includes a main memory 506, such as a random-access memory (RAM) or other dynamic storage device, coupled to bus 502 for storing information and instructions to be executed by processor 504. Main memory 506 also may be used for storing temporary variables or other intermediate information during execution of instructions to be executed by processor 504. Such instructions, when stored in non-transitory storage media accessible to processor 504, render computer system 500 into a special-purpose machine that is customized to perform the operations specified in the instructions” and [0106])
Regarding the remaining limitations of claim 8, see the rejection of claim 1 above.
Regarding claim 10, the combination of Torres Dho, Williams, Jr., and Adamson teaches all of the limitations of claim 8 above. Regarding the limitations introduced in claim 10, see the rejection of claim 3 above.
Regarding claim 13, the combination of Torres Dho, Williams, Jr., and Adamson teaches all of the limitations of claim 8 above. Regarding the limitations introduced in claim 13, see the rejection of claim 6 above.
Regarding claim 15, Torres Dho teaches:
A non-transitory computer-readable storage medium having stored thereon processor-executable instructions that, when executed by one or more processors, cause the one or more processors to perform a method comprising (see [0115] “a non-transitory computer readable storage medium comprises instructions which, when executed by one or more hardware processors, causes performance of any of the operations described herein and/or recited in any of the claims” and [0106])
Regarding the remaining limitations of claim 15, see the rejection of claim 1 above.
Regarding claim 17, the combination of Torres Dho, Williams, Jr., and Adamson teaches all of the limitations of claim 15 above. Regarding the limitations introduced in claim 17, see the rejection of claim 3 above.
Regarding claim 19, the combination of Torres Dho, Williams, Jr., and Adamson teaches all of the limitations of claim 15 above. Regarding the limitations introduced in claim 19, see the rejection of claim 6 above.
Regarding claim 21, the combination of Torres Dho, Williams, Jr., and Adamson teaches all of the limitations of claim 1 above. As discussed above, Torres Dho does not explicitly teach a pre-processing phase of classifying populations into sub-populations and training the machine learning models based on the clustering. Accordingly, Torres Dho does not explicitly teach generating multi-label classes stored as key-value pairs in the population data store with a key being a unique identifier of an item and a value comprising one or more classification labels. Torres Dho also does not explicitly teach filtering the events into subsets when training the machine learning classification engines by joining the events to classification labels in the data store. Williams, Jr. further teaches:
wherein classifying each of said sub-populations multiple times generates multi-label classes stored as key-value pairs, wherein a key is a unique identifier of an item in the population, and wherein a value comprises one or more classification labels, and wherein said classifications are stored in said population data store (see [0216]-[0217] “Data Ingestion 504 is used to populate Training Corpus 508 and Testing Corpus 510 with data representations including, but not limited to, sequences of records containing:… 4) data about the user accessing the data or content, including, but not limited to, a hashed form of the user identity, an organizational group or department code, a code indicating role within the organization, a score computed based on the reporting structure indicating influence or level within the organization, and user tenure with the organization; 5) data about the groups of users accessing a group ID, and the level of data or content business sensitivity that the group is expected to access; 6) for application servers, a code or locality sensitive hash of the API or functionality used, the arguments used, as well as a description that the data returned from the server described by size and data type” for storing data pairing a user identity (the key) with classifications in which the user falls (value) along with a pairing of APIs with a description of the data accessed and storing the pairs in the Training Corpus)
and wherein said training said plurality of machine learning classification engines further comprises filtering said compliance and audit events into subsets by joining said compliance and audit events to said one or more classification labels in said population data store (see [0212]-[0214] “Training Corpus 508 and Testing Corpus 510 may be populated with data describing the historical access patterns of content on one or more servers, aggregated by users performing accesses and also by the types or location of storage areas of content being accessed. Data describing access patterns may be obtained from server access logs. In at least one of the various embodiments, Model(s) 518 may be constructed and trained in several configurations, including: …2) Anomaly detection, trained on data segmented into per-user, per-user-group, and per-data-resource groupings, with each grouping having a distinct trained model. By training a model on per-user data, anomaly detection is customized to a particular user's historical behavior patterns, as behavior varies widely from user to user” for training machine learning models by filtering associated event data to event data associated with an individual user or a labeled group of users)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the preprocessing of the users and accessed programs into clusters and training anomaly detection machine learning models based on the clusters of Williams, Jr. into the system of Torres Dho. As Williams, Jr. states in [0215] “per-user training may make detection of anomalous behavior more accurate and reduce false positives. Similarly, training using user-group access patterns, where the group may be defined by characteristics such as those who are in a similar job role, or users in the same or related organization unit(s), may enable detection of anomalies that are divergent relative to the group's normal behaviors. Training on data-resource access patterns may be designed to detect anomalies in access patterns to a group of related data, because related data is often accessed in similar, consistent ways.” Accordingly, one of ordinary skill in the art would have recognized that by incorporating the clustering of user and software access data and then training models based on the clustered data would have resulted in improved accuracy of the anomaly detection models of Torres Dho. Examiner notes here that the improved accuracy of the models provided by clustering is a motivation from Applicant’s specification as well, but the motivation to improve model accuracy is not impermissible hindsight because the Williams, Jr. reference explicitly calls attention to the improved accuracy provided by the clustering. Accordingly, one of ordinary skill in the art would have had the motivation to combine Torres Dho and Williams, Jr. without relying on Applicant’s specification.
Claims 2, 9, and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Torres Dho in view of Williams, Jr., Adamson, and Muddu (U.S. Patent No. 9,516,053; hereafter known as Muddu).
Regarding claim 2, the combination of Torres Dho, Williams, Jr., and Adamson teaches all of the limitations of claim 1 above. Torres Dho further teaches:
and determine whether said current audit event is anomalous based on the output of said one or more ensemble models (see steps 214 and 216 and [0047] "At operation 214, the process generates a set of anomaly scores for each host based on the output of the machine learning algorithms", [0048] for the weighting of contributions of the anomaly probabilities of each model based on the reliability of each particular model to obtain the anomaly score for a host, and [0050] "At operation 216, the process determines whether any hosts have an anomaly score satisfying a threshold value" for determining whether a current state is anomalous based on the overall score for the current events surpassing a threshold)
As discussed above and in Torres Dho [0047]-[0049], Torres Dho teaches calculating a weighted average of the individual anomaly scores based on confidence scores from the machine learning models to arrive at an overall anomaly score. The combination of the combination of Torres Dho, Williams, Jr., and Adamson thus does not explicitly teach an ensemble machine learning model that obtains an aggregate score based on the anomaly and confidence scores of the individual machine leaning classification models. Muddu teaches:
wherein said plurality of machine learning classification engines comprises one or more ensemble models, each of said one or more ensemble models configured to obtain an aggregate score based on respective anomaly and confidence scores from a subset of the plurality of classification engines (see Col. 105 line 66 thru Col. 106 line 14 "In some embodiments ensemble learning techniques can be applied to process the plurality of feature scores according to a plurality of models (including machine-learning models) to achieve better predictive performance in the anomaly scoring and reduce false positives. An example model suitable for ensemble learning is Random Forest. In such an embodiment, the process may involve, processing an entity profile according to a plurality of machine-learning models, assigning a plurality of intermediate anomaly scores, each of the plurality of intermediate anomaly scores based on processing of the entity profile according to one of the plurality of machine-learning models, processing the plurality of intermediate anomaly scores according to an ensemble-learning model, and assigning the anomaly score based on processing the plurality of intermediate anomaly scores")
Since each individual element and its function are shown in the prior art, albeit shown in separate references, the difference between the claimed subject matter and the prior art rests not on any individual element or function but in the very combination itself. That is in the substitution of the ensemble machine learning model combining intermediate anomaly scores to obtain a final anomaly score of Muddu for the weighted averaging of the intermediate anomaly scores of the combination of Torres Dho, Williams, Jr., and Adamson.
Thus, the simple substitution of one known element for another producing a predictable result renders the claim obvious.
Furthermore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the ensemble model of Muddu into the combination of Torres Dho, Williams, Jr., and Adamson because, as Muddu states in Col. 105 line 66 thru Col. 106 line 14 above, “ensemble learning techniques can be applied to process the plurality of feature scores according to a plurality of models (including machine-learning models) to achieve better predictive performance in the anomaly scoring and reduce false positives”. Therefore, the incorporation of the ensemble model would improve the combination by reducing the chance for false positives.
Regarding claim 9, the combination of Torres Dho, Williams, Jr., and Adamson teaches all of the limitations of claim 8 above. Regarding the limitations introduced in claim 9, see the rejection of claim 2 above.
Regarding claim 16, the combination of Torres Dho, Williams, Jr., and Adamson teaches all of the limitations of claim 15 above. Regarding the limitations introduced in claim 16, see the rejection of claim 2 above.
Claims 4, 7, 11, 14, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Torres Dho in view of Williams, Jr., Adamson, and Kumar et al. (U.S. Pre-Grant Publication No. 2024/0356945, hereafter known as Kumar).
Regarding claim 4, the combination of Torres Dho, Williams, Jr., and Adamson teaches all of the limitations of claim 1 above. Torres Dho further teaches feedback loops to refine downstream models that take in anomaly scores as inputs in [0060], but the combination of Torres Dho, Williams, Jr., and Adamson does not explicitly teach feedback loops for assessing the accuracy of the machine learning models in response to an assessment of their accuracy. Kumar teaches:
further comprising a feedback loop for assessing accuracy of said determination and modifying one or more of said machine learning classification engines in response to said assessment accuracy (see Fig. 5 and [0088] "Once received by the SOC, the security alert may be investigated by, for example, security analysts and/or other resources associated with the SOC. Based on this review, the security analysts and/or other resources associated with the SOC (generally, the SOC) may create feedback related to the detected anomaly and other detected anomalies", [0089] "in operation 504, one or more servers may receive the created feedback from the SOC. For example, with reference to FIG. 1, the data transform server 108, the ML server 116, the rule server 118, and/or the alert generation server 120 may receive such feedback. Then, in response to the feedback, upstream processes may be tweaked and improved for better model predictions in a recurrent manner", and [0090] "in operation 506, the models 110, 112, 114 of FIG. 1 may be optionally tuned, trained, etc. based on the received feedback. For instance, hyperparameters of the models 110, 112, 114, model families (e.g., in an ensemble), individual model weights in the ensemble, etc. may be modified based on the feedback to improve performance metrics")
It would have been obvious to one of ordinary skill in the art before thew effective filing date of the claimed invention to incorporate the feedback loops for the machine learning models of Kumar into the combination of Torres Dho, Williams, Jr., and Adamson. As Kumar states in [0089] “in response to the feedback, upstream processes may be tweaked and improved for better model predictions in a recurrent manner” and [0090] “hyperparameters of the models 110, 112, 114, model families (e.g., in an ensemble), individual model weights in the ensemble, etc. may be modified based on the feedback to improve performance metrics”. Therefore, one of ordinary skill in the art would have recognized that by incorporating the feedback loops of Kumar, the combined system would be able to improve the machine learning models to more accurately identify anomalies in a network.
Regarding claim 7, the combination of Torres Dho, Williams, Jr., Adamson, and Kumar teaches all of the limitations of claim 4 above. Torres Dho further teaches a plurality of the machine learning models operating independently and in parallel as discussed regarding claim 6 above. As discussed above regarding claim 4, the combination of Torres Dho, Williams, Jr., and Adamson does not explicitly teach feedback loops for assessing the accuracy of the machine learning models in response to an assessment of their accuracy. Accordingly, the combination of Torres Dho, Williams, Jr., and Adamson does not explicitly teach a plurality of these feedback loops for each machine learning model that operate in parallel and separately from the other feedback loops. Kumar further teaches:
wherein said feedback loop comprises a plurality of feedback loops for each respective machine learning classification engine of said plurality of machine learning classification engines, and wherein each of said plurality of feedback loops is executed in parallel and separately from others of said plurality of feedback loops (see [0071] "the models 110, 112, 114 (and/or other models herein) may be retrained. For example, the ML server 116 and/or the SOC 122 may detect whether performance of the ML models 110, 112, 114 falls below a defined threshold. Then, in response to the performance falling below the defined threshold, the models 110, 112, 114 may be retrained based on, for example, new normal patterns relating to remote and/or physical access" for a plurality of feedback loops for tuning the plurality of models running in parallel. See [0058] "any one of the unsupervised models 110, 112, 114 may be sufficiently trained when one or more of its performance metrics (e.g., accuracy, precision, recall, f1-score, etc.) on test data is about 10% or less, 5% or less, etc." for each of the models being retrained satisfactorily based on their own independent performance thresholds)
It would have been obvious to one of ordinary skill in the art before thew effective filing date of the claimed invention to incorporate the parallel and separately operating feedback loops for the machine learning models of Kumar into the combination of Torres Dho, Williams, Jr., and Adamson. As Kumar states in [0089] “in response to the feedback, upstream processes may be tweaked and improved for better model predictions in a recurrent manner” and [0090] “hyperparameters of the models 110, 112, 114, model families (e.g., in an ensemble), individual model weights in the ensemble, etc. may be modified based on the feedback to improve performance metrics”. Therefore, one of ordinary skill in the art would have recognized that by incorporating the feedback loops of Kumar, the combined system would be able to improve the machine learning models to more accurately identify anomalies in a network. Furthermore, by refining models when they fall below a particular threshold and refining the model until the individual model reaches a performance threshold, the combined system maintains system performance without wasting time and effort refining models that already meet a performance threshold.
Regarding claim 11, the combination of Torres Dho, Williams, Jr., and Adamson teaches all of the limitations of claim 8 above. Regarding the limitations introduced in claim 11, see the rejection of claim 4 above.
Regarding claim 14, the combination of Torres Dho, Williams, Jr., Adamson, and Kumar teaches all of the limitations of claim 11 above. Regarding the limitations introduced in claim 14, see the alternate rejection of claim 7 above.
Regarding claim 18, the combination of Torres Dho, Williams, Jr., and Adamson teaches all of the limitations of claim 15 above. Regarding the limitations introduced in claim 18, see the rejection of claim 4 above.
Claims 5 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Torres Dho in view of Williams, Jr., Adamson, and Broyda et al. (U.S. Pre-Grant Publication No. 2021/0004949, hereafter known as Broyda).
Regarding claim 5, the combination of Torres Dho, Williams, Jr., and Adamson teaches all of the limitations of claim 1 above. Torres Dho further teaches:
wherein each of said (see [0046] "Probabilistic models may assign values between 0-1 based on the probability that the value is an outlier with 1 indicating a 100% probability, 0 indicating a 0% percent probability, and values in between represent varying levels of probability increasing the closer the value is to 1" for anomaly scores between 0-1. See [0048] for weights of models being confidence scores. The weights being values within 0-1 is would have been obvious as part of routine optimization of Torres Dho (see MPEP 2144.05 II.). As all the models are weighted to contribute to an overall value, the absolute value of each of the weights does not matter to the functioning of the invention as much as the value of the weight relative to the sum of all weight values. In other words, a model with a weight of 30 out of an aggregate value of 100 has the same effect on the final result as a model with a weight of 0.3 out of an aggregate value of 1. Accordingly, the exact weight values, or the scale (out of 1) at which the weight values are issued, would be arrived at via routine optimization of the teachings of Torres Dho)
As discussed above regarding claim 1, Torres Dho does not explicitly teach the confidence score being representative of a degree of confidence in the anomaly score. Accordingly, Torres Dho does not explicitly teach the confidence score being a value between 0 and 1. Williams, Jr. teaches:
wherein each of said respective confidence scores is a value (see [0120]-[0122] including “In at least one of the various embodiments, each classification result generated by the fast learning model may be associated with a value (the confidence value or confidence score) that scores/ranks the how close the data matches the classifier”, and [0175]-[0176] for the generation of confidence scores along with anomaly classifications in each of multiple models)
It also would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the models outputting confidence scores reflecting a confidence in the model’s anomaly assessment of Williams, Jr. into the system of Torres Dho. As Williams Jr. discusses in [0121]-[0122], confidence scores reflect how well the data being processed matches the classifiers and the training data that was used to train the classifiers. By using a confidence score output from each model reflective of the model’s confidence in its classification of anomalous or normal data, paragraph [0176] recites that the higher confidence model can be used to generate the output to increase accuracy. In combination with Torres Dho, one of ordinary skill in the art would have recognized that outputting confidence values would allow the collection of models to be weighted according to confidence to provide a final determination as to whether the event was anomalous or not, with weight being responsive to how the event fits the models’ training.
The combination of Torres Dho, Williams, Jr., and Adamson still does not explicitly teach that the confidence scores have a value between 0 and 1. However, Broyda teaches confidence scores between 0 and 1 (see [0089] “The confidence score may be, for example, a value between zero and one”).
It would have been obvious to one of ordinary skill in the art at the time of the invention to include the confidence score being a value between 0 and 1 as taught by Broyda in the combination of Torres Dho, Williams, Jr., and Adamson, since the claimed invention is merely a combination of old elements, and in the combination each element merely would have performed the same function as it did separately, and one of ordinary skill in the art would have recognized that the results of the combination were predictable.
Regarding claim 12, the combination of Torres Dho, Williams, Jr., and Adamson teaches all of the limitations of claim 8 above. Regarding the limitations introduced in claim 12, see the alternate rejection of claim 5 above.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Kesin et al. (U.S. Patent No. 11,310,254) teaches detecting anomalous activity in a network by comparing a user’s behavior to behavior of other users in the user’s cohort
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHAEL C MORONEY whose telephone number is (571)272-4403. The examiner can normally be reached Mon-Fri 8:30-5:30.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Nathan Uber can be reached at (571) 270-3923. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/M.C.M./Examiner, Art Unit 3626
/EMMETT K. WALSH/Primary Examiner, Art Unit 3626