DETAILED ACTION
The action is in response to the original filing on March 6, 2023 and the Remarks and Amendments filed on April 13, 2026. Claims 1-22 are pending and have been considered below. Claims 1, 14, and 19 are independent claims. Claims 1, 14, and 19 are amended. Claims 5-6 are canceled. Claims 21 and 22 are new.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claim 22 objected to because of the following informalities: “the multiple anomaly detection model” should read “the multiple anomaly detection models.” Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
The terms the response time and the plurality of applications lack sufficient basis as there is no prior reference to these terms in claims 1 or 8. For examination purposes, these terms are interpreted to mean “a response time” and “a success rate,” respectively.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-22 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding claim 1:
Step 1 – Claim 1 is directed to a method: A method for machine learning-based application management…
Step 2A, Prong 1 – A judicial exception is recited in this claim as it recites mathematical concepts (see MPEP 2106.04(a)(2)(I)):
decomposing… log data associated with a plurality of applications into time-series data representing values of one or more key performance indicators (KPIs) over a time period associated with the log data… To “decompose” log data is to use mathematical calculations such as data parsing or aggregation to convert unstructured event-based data into structured time-series data. Hence “decomposing… log data associated with a plurality of applications into time-series data representing values of one or more key performance indicators (KPIs) over a time period associated with the log data” is a mathematical concept.
performing… clustering operations based on one or more temporal components derived from the time-series data to assign each of the plurality of applications to at least one of multiple training groups… To perform “clustering operations” to assign applications into groups is to aggregate data into sets, which is a mathematical calculation. Hence “performing… clustering operations based on one or more temporal components derived from the time-series data to assign each of the plurality of applications to at least one of multiple training groups” is a mathematical concept.
determining… a training sequence for the plurality of applications based on the multiple training groups… To “determine” a training sequence for the plurality of applications is to aggregate and append data into an ordered list, which is a mathematical calculation. Hence “determining… a training sequence for the plurality of applications based on the multiple training groups” is a mathematical concept.
to detect occurrence of an anomaly by a corresponding application based on received application data… To “detect” occurrence of an anomaly based on received application data, as understood in the present application’s specification, is “to check sparsity within the time- series data and use the central tendency for one of multiple different intervals for thresholding and prioritized detection of anomalies” (¶44). Using the “central tendency… for thresholding” is comparing data to a threshold value, which is a mathematical calculation. Hence, to “detect occurrence of an anomaly by a corresponding application based on received application data” is a mathematical concept.
Step 2A, Prong 2 – The following limitations are additional elements without significantly more than the abstract idea:
decomposing, by one or more processors… performing, by the one or more processors… determining, by the one or more processors… one or more processors used as mere tools to apply an exception are generic elements for performing or applying the abstract idea using a generic computing environment (see MPEP 2106.05(f)).
initiating, by the one or more processors, training of a plurality of anomaly detection models that correspond to the plurality of applications according to the training sequence… training of a plurality of models according to the training sequence is an attempt to use the plurality of models by merely applying the abstract idea (i.e., using the training sequence determined with math) without placing any limits on how the training is performed. Further, the limitation omits any details as to how “training of a plurality of anomaly detection models” solves a technical problem and instead recites only the idea of a solution or outcome (see MPEP 2106.05(f)). Thus, the limitation represents no more than mere instructions to implement the abstract idea which is equivalent to adding the words “apply it” to the recited judicial exception.
wherein the training of the plurality of anomaly detection models, according to the training sequence includes concurrently training one or more anomaly detection models of a first training group of the multiple training groups and one or more anomaly detection models of a second training group of the multiple training groups… describing “the training sequence” to include “concurrently training” is an attempt to limit the field of use without significantly more (see MPEP 2106.05(h)).
wherein training an anomaly detection model of the first training group comprises performing one or more same preprocessing operations, one or more same post-processing operations, or a combination thereof, than training an anomaly detection model of the second training group… performing “one or more same preprocessing operations…” for a model of a first training group when compared to a model of a second training group is an attempt to limit the field of use without significantly more (see MPEP 2106.05(h)).
wherein each anomaly detection model of the plurality of anomaly detection models comprises a machine learning (ML) model configured to detect occurrence of an anomaly… a machine learning model used as a mere tool to apply an exception is a generic element for performing or applying the abstract idea using a generic computing environment (see MPEP 2106.05(f)).
Step 2B: These elements are recited at such a high level of generality that they fail to integrate the abstract idea into a practical application, since they provide nothing more than mere instructions to implement an abstract idea on a generic computer (MPEP 2106.05(f)), only amount to data gathering or outputting without significantly more (MPEP 2106.05(g)), or attempt to limit the field of use without significantly more (MPEP 2106.05(h)). These limitations, taken either alone or in combination, fail to provide an inventive concept. Thus, the claim is not patent eligible.
Claims 2-13 and 21-22 recite limitations which further narrow the abstract idea of claim 1 by specifying more details of the mathematical concepts that occur:
Regarding claim 2, specifying wherein the one or more temporal components comprise trend components, seasonal components, cyclic components, or a combination thereof in this manner does not overcome the rejection of claim 1 as modifying the temporal components does not make “performing… clustering operations” to not be a mathematical concept.
Regarding claim 3, the claim further limits abstract idea of claim 1 to be based on a mental process: determining, by the one or more processors, training frequencies for the plurality of anomaly detection models based on the time-series data. For example, given a small enough plurality of anomaly detection models and simple enough time-series data, a human can reasonably perform determining training frequencies for the plurality of anomaly detection models (see MPEP 2106.04(a)(2)(III)). Additionally, the limitation generating, by the one or more processors, a training schedule for the plurality of anomaly detection models based on the training frequencies and the training sequence, the training schedule including the training sequence and one or more future training sequences amounts to necessary data outputting and is still insignificant extra-solution activity (see MPEP 2106.05(g)).
Regarding claim 6, specifying wherein the training of the plurality of anomaly detection models according to the training sequence includes training a first anomaly detection model of a first training group of the multiple training groups and a second anomaly detection model of the first training group in series in this manner does not overcome the rejection of claim 1 as modifying the training of the plurality of anomaly detection models does not make “to detect occurrence of an anomaly” to not be a mathematical concept.
Regarding claim 7, specifying wherein training the first anomaly detection model comprises performing one or more different preprocessing operations, one or more different post-processing operations, or a combination thereof, as training the second anomaly detection model in this manner does not overcome the rejection of claim 6 as modifying the training of an anomaly detection model does not make “to detect occurrence of an anomaly” to not be a mathematical concept.
Regarding claim 8, the claim further limits abstract idea of claim 1 to be based on a mental process: to identify one or more additional applications that are predicted to fail based on one or more detected anomalies output by the plurality of anomaly detection models. For example, given a small enough number of applications and detected anomalies, a human can reasonably perform identifying one or more additional applications that are predicted to fail (see MPEP 2106.04(a)(2)(III)). Additionally, generating, by the one or more processors, an application dependency graph based on the time-series data, the log data, or a combination thereof and initiating, by the one or more processors, training of a failure engine based on the application dependency graph to output indicators of applications that are predicted to fail amount to necessary data gathering and outputting and is still insignificant extra-solution activity (see MPEP 2106.05(g)), and wherein the failure engine executes a ML model configured to identify one or more additional applications… amounts to mere instructions to apply the judicial exception using a generic computing environment and is not indicative of significantly more (see MPEP 2106.05(f)).
Regarding claim 9, describing wherein the failure engine is further trained based on the application dependency graph to configure the failure engine to output failure scores corresponding to reasons for failure associated with the applications that are predicted to fail amounts to necessary data gathering and outputting and is still insignificant extra-solution activity (see MPEP 2106.05(g)).
Regarding claim 10, describing initiating, by the one or more processors, training of an application recovery model based on historical recovery action data, the log data, and the application dependency graph and wherein the application recovery model comprises an ML model configured to output recovery actions based on input indicators of applications that are predicted to fail amount to necessary data gathering and outputting and is still insignificant extra-solution activity (see MPEP 2106.05(g)).
Regarding claim 11, describing providing, by the one or more processors, current log data as input data to the plurality of anomaly detection models to generate one or more detected anomalies associated with one or more applications of the plurality of applications… providing, by the one or more processors, the one or more detected anomalies as input data to the failure engine to generate one or more indicators of applications that are predicted to fail and one or more failure scores corresponding to reasons for failure associated with the applications that are predicted to fail… providing, by the one or more processors, the one or more indicators of the applications that are predicted to fail as input data to the application recovery model to generate one or more recovery action recommendations… and displaying, by the one or more processors, a dashboard that indicates the applications that are predicted to fail, the one or more failure scores, the reasons for failure, the one or more recovery action recommendations, or a combination thereof amounts to necessary data gathering and outputting and is still insignificant extra-solution activity (see MPEP 2106.05(g)).
Regarding claim 12, describing initiating, by the one or more processors, automatic performance of an action indicated by the one or more recovery action recommendations amounts to mere automation of manual processes using a generic computer and is not sufficient to show an improvement in computer-functionality (see MPEP 2106.05(a)).
Regarding claim 13, specifying wherein the action comprises re-executing one or more of the applications that are predicted to fail, terminating one or more of the applications that are predicted to fail, or a combination thereof in this manner does not overcome the rejection of claim 12 as modifying the action does not make “to detect occurrence of an anomaly” to not be a mathematical concept.
Regarding claim 21, the claim further limits abstract ideas of claim 8 to be based on a mathematical concept and a mental process: wherein the plurality of edges corresponding to the KPIs are aggregated based on determining an average of number of hits, a maximum of the response time, and a minimum of the success rate by the plurality of applications… To “aggregate” edges corresponding to KPIs based on “determining an average… a maximum… and a minimum” recites mathematical calculations, which is a mathematical concept… wherein the plurality of nodes is classified to a normal value, near anomaly, or anomaly based on the aggregated KPIs for the corresponding plurality of applications… Given a small enough plurality of nodes and simple enough aggregated KPIs for the corresponding applications, a human can reasonably perform classifying the nodes within the human mind or with the aid of a pen and paper, which is a mental process. Furthermore, wherein the application dependency graph includes a plurality of nodes and a plurality of edges, wherein the plurality of nodes correspond to the plurality of applications, and the plurality of edges correspond to the KPIs, the plurality of nodes are linked by the plurality of edges represents dependency between the plurality of applications corresponding to the linked plurality of nodes is an attempt to limit the field of use without significantly more (see MPEP 2106.05(h)).
Regarding claim 22, the claim further limits abstract ideas of claim 8 to be based on a mathematical concept: wherein the application dependency graph is integrated with an anomaly inference pipeline formed from the multiple anomaly detection model, wherein integrating the application dependency graph with the anomaly detection models is based on summarizing output of the anomaly inference pipeline, wherein summarizing the output includes mean, median and central dispersion of the anomaly inference pipeline… To integrate an application dependency graph with an anomaly inference pipeline, wherein the integrating is based on summarizing an output of the anomaly inference pipeline, wherein the summarizing includes mean, median, and central dispersion is to perform statistical analyses on the output, which is a mathematical concept.
Claims 14-18 recite a system that parallels the method claims of 1 and 8-11, respectively. Therefore, the analysis discussed above with respect to claims 1 and 8-11 also applies to claims 14-18, respectively. Accordingly, claims 14-18 are rejected based on substantially the same rationale as set forth above with respect to claims 1 and 8-11, respectively.
Claims 19-20 recite a non-transitory computer-readable storage medium that parallels the method claims of 1 and 3, respectively. Therefore, the analysis discussed above with respect to claims 1 and 3 also applies to claims 19-20, respectively. Accordingly, claims 19-20 are rejected based on substantially the same rationale as set forth above with respect to claims 1 and 3, respectively.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 14, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Krishnan et al. (US 20180113773 A1, hereinafter Krishnan) in view of Priydarshi et al. (WO 2020171622 A1, hereinafter Priydarshi) and further in view of Sun et al. (“CoGNN: Efficient Scheduling for Concurrent GNN Training on GPUs,” 2022, hereinafter Sun).
Regarding claim 1:
Krishnan teaches a method for machine learning-based application management, the method comprising: decomposing, by one or more processors, log data associated with a plurality of applications (Fig. 1 – 100, ¶18 “FIG. 1 illustrates an example environment that employs an application failure prediction system (AFPS) 100 which uses a model to analyze logs of an application to predict the probability of application failure,” ¶41 “The AFPS 100 is enabled to proactively monitor and correct errors that occur during the course of application execution thereby ensuring the smooth running of the various applications,” wherein log data associated with a plurality of applications is implicit) into time-series data representing values of one or more key performance indicators (KPIs) over a time period associated with the log data (Fig. 1 – 102, 122, 162, Fig. 4 – 400-410, ¶36 “FIG. 4 is a flowchart 400 that details an example method of detecting potential application failures or malfunctions. The method of detecting potential application failures as detailed herein can be carried out by the processor 102… Real-time data 162 is received at block 402 during the course of execution of the application 122… If at block 404 it is determined that the real-time data 162 comprises unstructured data, it can be converted to structured data at block 406… The predictive data model 120 can then be applied to the real-time data 162 at block 408 and the anomalies are detected at 410… The anomalies may be detected, for example, by their characteristic temporal error patterns or other attributes,” wherein converting unstructured “real-time data,” or log data, into “structured data” encompasses the method comprising: decomposing, by one or more processors, log data… into time-series data and “characteristic temporal error patterns or other attributes” encompasses values of one or more key performance indicators (KPIs) over a time period associated with the log data when given its broadest reasonable interpretation).
Regarding the limitation performing, by the one or more processors, clustering operations based on one or more temporal components derived from the time-series data to assign each of the plurality of applications to at least one of multiple training groups, Krishnan teaches one or more temporal components derived from the time-series data (Fig. 4 – 410, ¶36 “characteristic temporal error patterns,” Fig. 5 – 504, ¶38 “FIG. 5 is a flowchart 500 that details one example of a method of estimating an anomaly score or the probability of application failure… Patterns of error codes which represent a temporal sequences of errors are therefore recognized at block 504”). However, Krishnan fails to teach performing, by the one or more processors, clustering operations based on one or more temporal components derived from the time-series data to assign each of the plurality of applications to at least one of multiple training groups.
Priydarshi, in the same field of endeavor, teaches performing, by the one or more processors (Fig. 2a – 100, 107, 213, ¶38 “the one or more modules 213 may be communicatively coupled to the processor 107 for performing one or more functions of the electronic device 100. The said modules 213 when configured with the functionality defined in the present disclosure will result in a novel hardware”), clustering operations based on application usage to assign each of the plurality of applications to at least one of multiple training groups (Fig. 2 – 109, 219, Fig. 5 – 505, ¶58 “the one or more applications are clustered into one or more groups by the clustering module 219 by using the learning model 109. The learning model 109 is trained dynamically based on the application usage pattern for clustering”).
Regarding the limitation determining, by the one or more processors, a training sequence for the plurality of applications based on the multiple training groups, Priydarshi teaches the plurality of applications based on the multiple training groups (Fig. 2c, ¶43 “the learning model 109 cluster the plurality of applications based on temporal and application usage pattern”). However, Priydarshi fails to teach the full limitation determining, by the one or more processors, a training sequence for the plurality of applications…
Sun, in the same field of endeavor, teaches determining, by the one or more processors (Page 10, Section F, ¶1 “CoGNN overhead can be divided into three parts, including PMC estimation, task grouping, and task scheduling. PMC estimation loads each model structure and traverses the computation graph through memory cost functions to update the PMC information. Task grouping re-orders the task queue and generates task groups according to the scheduling policy. Task scheduling iteratively allocates memory, dispatches workers, and synchronizes tasks within group… Figure 13 shows the breakdown of CoGNN processing, including PMC estimation and task scheduling,” Page 11, Figure 13 wherein determining, by the one or more processors is implicit), a training sequence for models based on task groups (Abstract: “Graph neural networks (GNNs) suffer from low GPU utilization due to frequent memory accesses. Existing concurrent training mechanisms cannot be directly adapted to GNNs… massive training tasks generated from scenarios such as hyper-parameter tuning require flexible scheduling strategies… we propose CoGNN that enables efficient management of GNN training tasks on GPUs… the CoGNN organizes the tasks in a queue and estimates the memory consumption of each task based on cost functions at operator basis. In addition, the CoGNN implements scheduling policies to generate task groups, which are iteratively submitted for execution,” Page 5, Col. 1, ¶2 “the CoGNN can support more general neural networks due to the versatility of its components”).
Regarding the limitation and initiating, by the one or more processors, training of a plurality of anomaly detection models that correspond to the plurality of applications according to the training sequence, Krishnan further teaches initiating, by the one or more processors, training of a plurality of anomaly detection models that correspond to the plurality of applications (Fig. 1 – 102, 120, 122, 124, 162, 164, Fig. 4 – 400-410, ¶20 “the features discussed herein are equally applicable when… executing the plurality of respective predictive data models corresponding to the plurality of applications,” ¶21 “The predictive data model 120 thus generated can be initially trained,” ¶36 “The predictive data model 120 can then be applied to the real-time data 162 at block 408 and the anomalies are detected at 410”). However, Krishnan fails to teach according to the training sequence.
Sun teaches training task groups according to the training sequence (Page 5, Col. 2, ¶2 “The task scheduler iteratively executes training tasks at the granularity of task groups”).
Regarding the limitation wherein the training of the plurality of anomaly detection models, according to the training sequence includes concurrently training one or more anomaly detection models of a first training group of the multiple training groups and one or more anomaly detection models of a second training group of the multiple training groups, Krishnan teaches the training of the plurality of anomaly detection models and one or more anomaly detection models (Fig. 1 – 102, 120, 122, 124, 162, 164, Fig. 4 – 400-410, ¶20-21, ¶36). However, Krishnan fails to teach wherein the training of the plurality of anomaly detection models according to the training sequence includes concurrently training one or more anomaly detection models of a first training group of the multiple training groups and one or more anomaly detection models of a second training group of the multiple training groups.
Priydarshi teaches a first training group of the multiple training groups and a second training group of the multiple training groups (Fig. 2a –219, Fig. 2c, ¶42 “The clustering module 219 may clusters the plurality of applications into one or more groups,” Figure 2c depicts tables with multiple Group ID’s, among which are Group ID 1 and 2, or a first training group of the multiple training groups and a second training group of the multiple training groups). However, Priydarshi fails to teach wherein the training of the plurality of anomaly detection models according to the training sequence includes concurrently training…
Sun teaches wherein training a plurality of models according to the training sequence includes concurrently training one model of a first training group and another model of the same training group (Page 4, Col. 1, Section B, ¶1 “concurrently training multiple GNN tasks on a single GPU is also a promising direction to improve GPU utilization,” Col. 2, Section A “we propose a concurrent GNN training framework CoGNN that organizes GNN training tasks into a queue and enables efficient scheduling and management under GPU co-location,” Page 5, Col. 1, ¶2 “The CoGNN combines spatial sharing and temporal sharing for more flexible task management,” Fig. 6 and Fig. 7 both depict executing training tasks from the same training task group simultaneously, Page 8, Col. 2, Section B, ¶3 “tasks within a group improve overall training throughput through spatial sharing”).
Regarding the limitation wherein training an anomaly detection model of the first training group comprises performing one or more same preprocessing operations, one or more same post-processing operations, or a combination thereof, than training an anomaly detection model of the second training group, Krishnan teaches training an anomaly detection model (Fig. 1 – 102, 120, 122, 124, 162, 164, Fig. 4 – 400-410, ¶20-21, ¶36). However, Krishnan fails to teach wherein training an anomaly detection model of the first training group comprises performing one or more same preprocessing operations, one or more same post-processing operations, or a combination thereof, than training an anomaly detection model of the second training group.
Priydarshi teaches the first training group and the second training group (Fig. 2a –219, Fig. 2c, ¶42 as explained above). However, Priydarshi fails to teach wherein training an anomaly detection model of the first training group comprises performing one or more same preprocessing operations, one or more same post-processing operations, or a combination thereof, than training an anomaly detection model of the second training group.
Sun teaches wherein training a model of the first training group comprises performing one or more same preprocessing operations, one or more same post-processing operations, or a combination thereof, than training another model of the same training group (Page 1, Col. 2, ¶2 “temporal sharing reduces pipelining latency by overlapping data preprocessing and computations,” Page 2, Col. 2, ¶2 “The CoGNN applies spatial sharing and temporal sharing to intra and inter task groups,” Page 7, Col. 2, ¶2 “The CoGNN consists of the scheduler process and worker process responsible for queue management and task execution, respectively… the scheduler process packs the tasks into a queue and loads the model structures… to generate task groups. At each group iteration, the scheduler process sends the hash indices of the tasks to the worker process. The two processes are then overlapped to reduce pipelining latency,” ¶3 “The scheduler process inserts parameter buffers and transfers the model parameters to GPU memory with synchronization through CUDA events. The worker process identifies tasks according to the hash indices and dispatches them to worker threads. Each worker thread loads the corresponding GNN model and attaches it to the CUDA stream… each worker thread adds hooks to wait for the scheduler to transfer the parameters needed for model execution. After the results return, the worker thread will send a “finish” signal to the scheduler process… When the number of received signals equals the group size, the scheduler process iterates to the execution of the next group,” Page 5, Col. 2, Fig. 7 depicts “overlapping” the “Insert comp. buffers” and “Dispatch workers” processes with the “Insert para. buffers” and “Transmit para.” processes for each group consisting of multiple training tasks, or performing one or more same preprocessing operations…).
Krishnan further teaches wherein each anomaly detection model of the plurality of anomaly detection models comprises a machine learning (ML) model configured to detect occurrence of an anomaly by a corresponding application based on received application data (Fig. 1 – 120, 124, 164, ¶21 “training data 124 which may comprise a subset of the application logs 164,” ¶23 “Anomalies may include a combination of error codes which the predictive data model 120 is trained to identify as leading to a high probability of application failure”).
Krishnan, Priydarshi, and Sun are analogous art to the claimed invention as all are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the application-based training groups of Priydarshi and the training scheduling framework of Sun with the methodology of Krishnan. The motivation to do so is to design a method that “optimizes user experience while using the application” (Priydarshi, ¶24) and to (Sun, Abstract: “achieve shorter completion and queuing time for training tasks”).
Regarding claim 14:
Krishnan teaches a system for machine learning-based application management (¶13 “An application failure prediction system (AFPS) disclosed herein is configured for accessing the real-time data from an application executing on a computing apparatus, predicting anomalies which may be indicative of potential application failures and implementing corrective actions to mitigate the occurrences of the anomalies”), the system comprising: a memory (Fig. 9 – 906, ¶50 “The computer-readable storage medium 906 may be any suitable medium which participates in providing instructions”).
Krishnan further teaches and one or more processors communicatively coupled to the memory, the one or more processors configured to (Fig. 1 – 100, Fig. 9 – 902, 906, ¶50 “The computer-readable storage medium 906… participates in providing instructions to the processor(s) 902 for execution… The instructions… stored on the computer readable medium 906 may include machine readable instructions… executed by the processor(s) 902 to perform the methods and functions for the AFPS 100 described herein”): decompose log data associated with a plurality of applications (Fig. 1 – 100, ¶18, ¶41, as explained above with respect to claim 1) into time-series data representing values of one or more key performance indicators (KPIs) over a time period associated with the log data (Fig. 1 – 102, 122, 162, Fig. 4 – 400-410, ¶36, as explained above with respect to claim 1).
Claim 14 recites a system that contains similar limitations to those of method claim of 1. Therefore, the analyses discussed above with respect to claim 1 apply to claim 14. Accordingly, claim 14 is rejected based on substantially the same rationale as set forth above with respect to claim 1.
Regarding claim 19:
Krishnan teaches a non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations for machine learning-based application management (Fig. 1 – 100, Fig. 9 – 902, 906, ¶50 “The computer-readable storage medium 906 may be any suitable medium which participates in providing instructions to the processor(s) 902 for execution… the computer readable medium 906 may be non-transitory… The instructions… stored on the computer readable medium 906 may include machine readable instructions… executed by the processor(s) 902 to perform the methods and functions for the (AFPS) 100 described herein”), the operations comprising: decomposing log data associated with a plurality of applications (Fig. 1 – 100, ¶18, ¶41, as explained above with respect to claim 1) into time-series data representing values of one or more key performance indicators (KPIs) over a time period associated with the log data (Fig. 1 – 102, 122, 162, Fig. 4 – 400-410, ¶36, as explained above with respect to claim 1).
Claim 19 recites a computer-readable storage medium that contains similar limitations to those of the method claim of 1 and the system claim of 14. Therefore, the analyses discussed above with respect to claims 1 and 14 apply to claim 19. Accordingly, claim 19 is rejected based on substantially the same rationale as set forth above with respect to claims 1 and 14.
Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Krishnan in view of Priydarshi and further in view of Sun, and further in view of Liu et al. (Machine Translation of CN 115220899 A, hereinafter Liu).
Regarding claim 6, Krishnan in view of Priydarshi and further in view of Sun teaches the method of claim 1 (and thus the rejection of claim 1 is incorporated).
Regarding the limitation wherein the training of the plurality of anomaly detection models according to the training sequence includes training a first anomaly detection model of a first training group of the multiple training groups and a second anomaly detection model of the first training group in series, Krishnan teaches the training of the plurality of anomaly detection models, training a first anomaly detection model, and training a second anomaly detection model (Fig. 1 – 120, ¶20 “the features discussed herein are equally applicable when… executing the plurality of respective predictive data models corresponding to the plurality of applications,” ¶21 “The predictive data model 120 thus generated can be initially trained,” wherein a “plurality of respective predictive data models” implies at least two models, or a first anomaly detection model and a second anomaly detection model). However, Krishnan fails to teach wherein the training of the plurality of anomaly detection models according to the training sequence includes training a first anomaly detection model of a first training group of the multiple training groups and a second anomaly detection model of the first training group in series.
Priydarshi teaches a first training group of the multiple training groups (Fig. 2a –219, Fig. 2c, ¶42 “The clustering module 219 may clusters the plurality of applications into one or more groups,” Figure 2c depicts a first training group of the multiple training groups). However, Priydarshi fails to teach wherein the training of the plurality of anomaly detection models according to the training sequence includes training a first anomaly detection model of a first training group of the multiple training groups and a second anomaly detection model of the first training group in series.
Liu, in the same field of endeavor, teaches wherein the training of a plurality of models according to the training sequence includes training a first model of a first training group and a second model of the first training group (¶42 “a target task group can be obtained, which includes multiple model training tasks to be processed. These model training tasks can be training tasks of models involving various deep learning methods”) in series (Fig. 3B, ¶57 “in one scheduling mode, after entering the (n-1)th training phase, task A is scheduled to resource 1, and the duration of task A using resource 1 is (t2-t1)… After tasks A, B, and C are completed, the nth training phase begins… Task B is scheduled to resource 3, and the duration of task B's use of resource 3 is (t3-t2)… The subsequent process follows the same pattern,” Figure 3B depicts each of the model training tasks taking up most of the training time during each respective stage of training, hence the training tasks are executed in series).
Krishnan, Priydarshi, and Liu are analogous art to the claimed invention as all are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the first training group of Priydarshi and the training in series of Liu with the plurality of anomaly detection models of Krishnan. The motivation to do so is to design a method that “optimizes user experience while using the application” (Priydarshi, ¶24) while “avoiding competition for model training resources between different model training tasks, improving the utilization rate of model training resources, and enhancing the efficiency of model training” (Liu, Machine Translation, ¶16).
Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Krishnan in view of Priydarshi and further in view of Sun, and further in view of Liu, and further in view of Dar et al. (US 20200134467 A1, hereinafter Dar).
Regarding claim 7, Krishnan in view of Priydarshi and further in view of Sun, and further in view of Liu teaches the method of claim 6 (and thus the rejection of claim 6 is incorporated).
Regarding the limitation wherein training the first anomaly detection model comprises performing one or more different preprocessing operations, one or more different post-processing operations, or a combination thereof, as training the second anomaly detection model, Krishnan teaches training the first anomaly detection model and training the second anomaly detection model (Fig. 1 – 120, ¶20-21 as explained above with respect to claim 6). However, Krishnan fails to teach wherein training the first anomaly detection model comprises performing one or more different preprocessing operations, one or more different post-processing operations, or a combination thereof, as training the second anomaly detection model.
Dar, in the same field of endeavor, teaches wherein training a first model comprises performing one or more different preprocessing operations, one or more different post-processing operations, or a combination thereof, as training another model (Fig. 2B, ¶52 “In the example of FIG. 2B, preprocessing batch A comprises… a non-sharable portion… The term “non-sharable portion” means a portion of the preprocessed data that is usable (e.g., sharable) as input to only a single NN computation task, for training of a single NN,” wherein the “non-shareable portion” of a “preprocessing batch” for “training of a single NN” is implied to comprise performing one or more different preprocessing operations when compared to training another model).
Krishnan and Dar are analogous art to the claimed invention as both are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the different preprocessing operations of Dar with the training of the anomaly detection models of Krishnan. The motivation to do so, as stated by Dar, is to “increase an availability of deep-learning-based products, such as artificial intelligence products which are based on learning methods, and improve hardware utilization” (Dar, ¶37).
Claims 8 and 9 are rejected under 35 U.S.C. 103 as being unpatentable over Krishnan in view of Priydarshi and further in view of Sun, and further in view of Mandal et al. (US 20230105304 A1, hereinafter Mandal).
Regarding claim 8, Krishnan in view of Priydarshi and further in view of Sun teaches the method of claim 1 (and thus the rejection of claim 1 is incorporated).
Regarding the limitation generating, by the one or more processors, an application dependency graph based on the time-series data, the log data, or a combination thereof, Krishnan teaches the time-series data, the log data, or a combination thereof (Fig. 1 –122, 162, Fig. 4 – 402-406, ¶36 “Real-time data 162 is received at block 402 during the course of execution of the application 122… If at block 404 it is determined that the real-time data 162 comprises unstructured data, it can be converted to structured data at block 406”). However, Krishnan fails to teach generating, by the one or more processors, an application dependency graph based on the time-series data…
Mandal, in the same field of endeavor, teaches generating, by the one or more processors (Fig. 7 – 710, ¶143 “CPU 710 may execute instructions stored in RAM 720 to provide several features of the present disclosure”), an application (Fig. 1 – 130, 160, ¶38 “Computing infrastructure 130 is a collection of nodes (160)… which are engineered to together host software applications,” ¶46 “software applications containing one or more components are deployed in nodes 160 of computing infrastructure 130… The components may include software/code modules of a software application,” wherein “component” encompasses application) dependency graph based on log data (Fig. 1 – 135, 150, Fig. 2 – 210, ¶59 “PRT 150 forms a causal dependency graph representing the usage dependencies among various components deployed in computing environment 135 during processing of prior user requests,” wherein “causal dependency graph” encompasses application dependency graph, ¶61 “PRT 150 receives real-time data such as… logs… during the processing of (current) user requests”).
Mandal further teaches and initiating, by the one or more processors, training of a failure engine based on the application dependency graph to output indicators of applications that are predicted to fail (Fig. 1 – 135, Fig. 2 – 220-260, Fig. 4 – 450A, Fig. 5A – 500, ¶95 “Incident predictor 450A takes as inputs… the causal dependency graph (500”)… and generates as outputs a probabilistic model, predicted future incidents/imminent performance issues… and severity of the predicted performance issues… the probabilistic model is a Markov network that corelates incidents to outliers occurring in the components deployed in computing environment 135,” ¶96 “the Markov network is trained to corelate the occurrences of the outliers to the eventual occurrences of the incidents, with the strength of correlation based on the strength of the relationships as indicated by causal dependency graph 500,” wherein the “incident predictor” encompasses a failure engine, ¶65 “performance issues may include… failure of the components,” wherein “performance issues” including “failure of the components” encompasses indicators of applications that are predicted to fail).
Regarding the limitation wherein the failure engine executes a ML model configured to identify one or more additional applications that are predicted to fail based on one or more detected anomalies output by the plurality of anomaly detection models, Krishnan teaches one or more detected anomalies output by the plurality of anomaly detection models (Fig. 1 – 120, ¶20 “the features discussed herein are equally applicable when… executing the plurality of respective predictive data models,” ¶23 “Anomalies in the real-time data 162 which can lead to application failures are identified by the predictive data model 120”). However, Krishnan fails to teach wherein the failure engine executes a ML model configured to identify one or more additional applications that are predicted to fail based on one or more detected anomalies…
Mandal teaches wherein the failure engine executes a ML model configured to identify one or more additional applications that are predicted to fail based on anomaly data (Fig. 1 – 135, Fig. 4 – 420, 450A, 460, Fig. 6E – 648, ¶65 as explained above, ¶95 “Incident predictor 450A takes as inputs… log events (stored in operational data 420) and generates as outputs a probabilistic model, predicted future incidents/imminent performance issues… and severity of the predicted performance issues… the probabilistic model is a Markov network that corelates incidents to outliers occurring in the components,” ¶109 “anomaly/outlier data (maintained as part of operational data 420),” ¶131 “Column 648 specifies a severity score indicating the severity of the incident, with a high value indicating a possible failure of multiple components of the software application or the application as a whole,” wherein an “incident predictor” that “generates” a “Markov network that correlates incidents to outliers occurring in the components” encompasses the failure engine executes a ML model configured to identify one or more additional applications that are predicted to fail).
Krishnan and Mandal are analogous art to the claimed invention as both are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the application dependency graph and failure engine of Mandal with the log data and anomaly detection models of Krishnan. The motivation to do so, as stated by Mandal, is to reduce “overall cost required to resolve failure by simply preventing failure from occurring. In addition, the overall productivity of the computing environment (135) increases by minimizing chances of interruption due to failure” (Mandal, Fig. 1 – 135, ¶54).
Regarding claim 9, Krishnan in view of Priydarshi and further in view of Sun and further in view of Mandal teaches the method of claim 8 (and thus the rejection of claim 8 is incorporated).
Regarding the limitation wherein the failure engine is further trained based on the application dependency graph to configure the failure engine to output failure scores corresponding to reasons for failure associated with the applications that are predicted to fail, Krishnan teaches failure scores corresponding to reasons for failure associated with the applications that are predicted to fail (Fig. 1 – 120, 162, Fig. 4 – 412, ¶36 “an anomaly is a potential application failure that the data model 120 is configured to detect… The anomaly score for an anomaly is calculated at block 412. For example, the anomaly score for a particular anomaly is calculated based on the occurrences of the various features corresponding to the anomaly in the predictive data model 120 within the real-time data 162… all the anomalies detected in the real-time data 162 can be simultaneously processed to obtain their anomaly scores,” wherein “anomaly scores” for an anomaly that are “calculated based on occurrences of the various features corresponding to the anomaly” encompasses failure scores corresponding to reasons for failure associated with the applications that are predicted to fail). However, Krishnan fails to teach wherein the failure engine is further trained based on the application dependency graph to configure the failure engine to output failure scores…
Mandal teaches wherein the failure engine is further trained based on the application dependency graph to configure the failure engine to output severity of predicted performance issues (Fig. 1 – 135, Fig. 4 – 450A, Fig. 5A – 500, Fig. 6E – 648, ¶95 “Incident predictor 450A takes as inputs… the causal dependency graph (500)… and generates as outputs a probabilistic model… and severity of the predicted performance issues… the probabilistic model is a Markov network,” ¶96 “the Markov network is trained to corelate the occurrences of the outliers to the eventual occurrences of the incidents, with the strength of correlation based on the strength of the relationships as indicated by causal dependency graph 500,” ¶131 “Column 648 specifies a severity score indicating the severity of the incident”).
Krishnan and Mandal are analogous art to the claimed invention as both are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the application dependency graph and failure engine of Mandal with the failure scores corresponding to reasons for failure and applications predicted to fail of Krishnan. The motivation to do so, as stated by Mandal, is to reduce “overall cost required to resolve failure by simply preventing failure from occurring. In addition, the overall productivity of the computing environment (135) increases by minimizing chances of interruption due to failure” (Mandal, Fig. 1 – 135, ¶54).
Claims 10-13 are rejected under 35 U.S.C. 103 as being unpatentable over Krishnan in view of Priydarshi and further in view of Sun, and further in view of Mandal, and further in view of Balasubramanian et al. (US 20200409715 A1, hereinafter Balasubramanian).
Regarding claim 10, Krishnan in view of Priydarshi and further in view of Sun and further in view of Mandal teaches the method of claim 9 (and thus the rejection of claim 9 is incorporated).
Regarding the limitation initiating, by the one or more processors, training of an application recovery model based on historical recovery action data, the log data, and the application dependency graph, Krishnan teaches historical recovery action data (Fig. 6 – 602-606, ¶39 “When an anomaly is detected, the action(s) to be implemented can be identified via recognizing similar features or patterns of error codes from the application logs 164 and retrieving the action or series of actions that were taken to address the anomaly. The model applicator 114 can be trained via, for example, un-supervised learning to identify similar anomalies and respective corrective actions that were earlier implemented. In an example, surveys may be collected from personnel who implement corrective actions”) and the log data (Fig. 1 – 122, 162, Fig. 4 – 402, ¶36 “Real-time data 162 is received at block 402 during the course of execution of the application 122”). However, Krishnan fails to teach training of an application recovery model based on historical recovery action data, the log data, and the application dependency graph.
Mandal teaches and the application dependency graph (Fig. 1 – 135, ¶59 “a causal dependency graph representing the usage dependencies among various components deployed in computing environment 135”). However, Mandal fails to teach initiating, by the one or more processors, training of an application recovery model based on….
Balasubramanian, in the same field of endeavor, teaches initiating, by the one or more processors (Fig. 1 – 101, 111, ¶46 “computing device 101 may include a processor 111… adapted to perform computations associated with machine learning”), training of an application recovery model based on past corrective actions (Fig. 12 – 1249, ¶142 “the monitoring device may utilize machine learning techniques to determine patterns of performance based on system state information associated with performance events. System state information for an event may be collected and used to train a machine learning model based on determining correlations between attributes of dependencies and the monitored application entering an unhealthy state. During later, similar events, the machine learning model may be used to generate a recommended action based on past corrective actions,” ¶153 “machine learning processes 1249 may also learn from the corrective actions associated with the event records and generate a recommendation that similar corrective action be taken when similar conditions arise at a later time”), event records or other system information (Fig. 12 – 1247-1249, ¶148 “The event records and other system information stored in smart database 1247 may be used by machine learning process 1249 to train a machine learning model and determine potential patterns of performance for the monitored application”), and application dependencies (¶112 “a machine learning model trained to identify correlations between the first dependency and the first application having an unhealthy operating status”).
Regarding the limitation wherein the application recovery model comprises an ML model configured to output recovery actions based on input indicators of applications that are predicted to fail, Krishnan teaches recovery actions based on input indicators of applications that are predicted to fail (Fig. 1 – 164, 170, Fig. 6 – 602-604, ¶20 “anomaly or a potential application failure,” ¶39 “The method begins at block 602 wherein the application logs 164 are accessed in order to identify solutions or corrective actions 170 to address the anomalies. When an anomaly is detected, the action(s) to be implemented can be identified via recognizing similar features or patterns of error codes from the application logs 164 and retrieving the action or series of actions that were taken to address the anomaly”). However, Krishnan fails to teach wherein the application recovery model comprises an ML model configured to output recovery actions…
Balasubramanian teaches wherein the application recovery model comprises an ML model configured to output recovery actions based on an application entering an unhealthy state (¶142 “System state information for an event may be collected and used to train a machine learning model based on determining correlations between attributes of dependencies and the monitored application entering an unhealthy state. During later, similar events, the machine learning model may be used to generate a recommended action based on past corrective actions”).
Krishnan, Mandal, and Balasubramanian are analogous art to the claimed invention as all are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the application dependency graph of Mandal and the recovery model of Balasubramanian with the log data, historical recovery action data, and recovery actions of Krishnan. The motivation to do so is to reduce “overall cost required to resolve failure by simply preventing failure from occurring. In addition, the overall productivity of the computing environment (135) increases by minimizing chances of interruption due to failure” (Mandal, Fig. 1 – 135, ¶54) while “allowing a more complete picture of the health and status of the target application and system” (Balasubramanian, ¶81).
Regarding claim 11, Krishnan in view of Priydarshi and further in view of Sun and further in view of Mandal and further in view of Balasubramanian teaches the method of claim 10 (and thus the rejection of claim 10 is incorporated).
Krishnan teaches providing, by the one or more processors, current log data as input data to the plurality of anomaly detection models to generate one or more detected anomalies associated with one or more applications of the plurality of applications (Fig. 1 –120, 122, 162, ¶20 “the features discussed herein are equally applicable when… executing the plurality of respective predictive data models corresponding to the plurality of applications,” ¶22 “real-time data 162 may be obtained… from the application 122 even as it is being generated,” ¶23 “Anomalies in the real-time data 162 which can lead to application failures are identified by the predictive data model 120”).
Regarding the limitation providing, by the one or more processors, the one or more detected anomalies as input data to the failure engine to generate one or more indicators of applications that are predicted to fail and one or more failure scores corresponding to reasons for failure associated with the applications that are predicted to fail, Krishnan teaches the one or more detected anomalies (Fig. 1 – 162, ¶23 “Anomalies in the real-time data 162 which can lead to application failures are identified”) and one or more indicators of applications that are predicted to fail (Fig. 1 – 164, ¶39 “an anomaly is detected… via recognizing similar features or patterns of error codes from the application logs 164”) and one or more failure scores corresponding to reasons for failure associated with the applications that are predicted to fail (Fig. 1 – 120, 162, Fig. 4 – 412, ¶36 “the anomaly score for a particular anomaly is calculated based on the occurrences of the various features corresponding to the anomaly in the predictive data model 120 within the real-time data 162… all the anomalies detected in the real-time data 162 can be simultaneously processed to obtain their anomaly scores”). However, Krishnan fails to teach providing, by the one or more processors, the one or more detected anomalies as input data to the failure engine to generate one or more indicators…
Mandal teaches providing, by the one or more processors, anomaly data as input data to the failure engine to generate predicted performance issues and severity of the predicted performance issues (Fig. 4 – 420, 450A, Fig. 6E – 648, ¶95 “Incident predictor 450A takes as inputs… log events (stored in operational data 420) and generates… predicted future incidents/imminent performance issues… and severity of the predicted performance issues,” ¶109 “anomaly/outlier data (maintained as part of operational data 420),” ¶131 “Column 648 specifies a severity score indicating the severity of the incident”).
Regarding the limitation providing, by the one or more processors, the one or more indicators of the applications that are predicted to fail as input data to the application recovery model to generate one or more recovery action recommendations, Krishnan teaches the one or more indicators of the applications that are predicted to fail (Fig. 1 – 164, ¶39) and one or more recovery action recommendations (Fig. 1 – 164, Fig. 6 – 608, 616, ¶40 “If it is determined at 608 that the action is not an automatic action, the procedure jumps to block 616 to transmit a message to the personnel. In an example, the message may include information regarding any solutions or corrective actions that were identified from the application logs 164,” wherein the “message” including “information regarding… corrective actions” encompasses recovery action recommendations). However, Krishnan fails to teach providing, by the one or more processors, the one or more indicators of the applications that are predicted to fail as input data to the application recovery model to generate one or more recovery action recommendations.
Balasubramanian teaches providing, by the one or more processors (Fig. 1 –111, ¶46 “a processor 111… adapted to perform… machine learning”), current operating status as input data to the application recovery model to generate recovery action recommendations (Fig. 12 – 1249, ¶155 “Once the model is trained on these patterns of performance, a current operating status of the monitored application and system may be used by the machine learning processes to generate predictions regarding the likelihood that the system will enter an unhealthy state. If the system is in an unhealthy state, or if conditions seem ripe for the system to enter an unhealthy state, the machine learning processes 1249 may generate a recommended action to restore the system to a healthy state”).
Krishnan further teaches and displaying, by the one or more processors, a dashboard that indicates the applications that are predicted to fail, the one or more failure scores, the reasons for failure, the one or more recovery action recommendations, or a combination thereof (Fig. 1 – 118, Fig. 8, ¶45 “FIG. 8 illustrates an example of the GUI 118… that allows a human user to monitor the real-time data 162… The predictors or features 802 for estimating the probability of application failure are shown on the right hand side of the GUI 118. The probability of each of the features indicating application failure can be indicated on the plot 804… A total anomaly score for the real-time data set that is currently being analyzed on the GUI 118 can be indicated via a torus 808. The color of the torus 808 indicates the status alert of the application 122 based on the information from the real-time data 162 currently being displayed”).
Krishnan, Mandal, and Balasubramanian are analogous art to the claimed invention as all are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the failure engine of Mandal and the recovery model of Balasubramanian with the log data, anomaly detection models, applications predicted to fail, and dashboard of Krishnan. The motivation to do so is to reduce “overall cost required to resolve failure by simply preventing failure from occurring. In addition, the overall productivity of the computing environment (135) increases by minimizing chances of interruption due to failure” (Mandal, Fig. 1 – 135, ¶54) while “allowing a more complete picture of the health and status of the target application and system” (Balasubramanian, ¶81).
Regarding claim 12, Krishnan in view of Priydarshi and further in view of Sun and further in view of Mandal and further in view of Balasubramanian teaches the method of claim 11 (and thus the rejection of claim 11 is incorporated).
Regarding the limitation initiating, by the one or more processors, automatic performance of an action indicated by the one or more recovery action recommendations, Krishnan teaches initiating, by the one or more processors, automatic performance of an action (Fig. 6 – 608-610, ¶41 “If the retrieved actions can be automatically executed… then such actions are automatically executed at block 610”). However, Krishnan fails to teach an action indicated by the one or more recovery action recommendations.
Balasubramanian teaches an action indicated by the one or more recovery action recommendations (¶155 “a recommended action to restore the system to a healthy state”).
Krishnan and Balasubramanian are analogous art to the claimed invention as both are from the same field of endeavor of machine learning. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the action indicated by the one or more recovery action recommendations of Balasubramanian with the automatic performance of Krishnan. The motivation to do so, as stated by Balasubramanian, is “allowing a more complete picture of the health and status of the target application and system” (Balasubramanian, ¶81).
Regarding claim 13, Krishnan in view of Priydarshi and further in view of Sun and further in view of Mandal and further in view of Balasubramanian teaches the method of claim 12 (and thus the rejection of claim 12 is incorporated).
Krishnan teaches wherein the action comprises re-executing one or more of the applications that are predicted to fail, terminating one or more of the applications that are predicted to fail, or a combination thereof (¶48 “different actions may be implemented. In an example, an action may be implemented on the application server, such as when the correction of the error requires a restart”).
Claim 21 is rejected under 35 U.S.C. 103 as being unpatentable over Krishnan in view of Priydarshi and further in view of Sun, and further in view of Mandal, and further in view of He et al. (“Graph based Incident Extraction and Diagnosis in Large-Scale Online Systems,” 2022, hereinafter He), and further in view of Iglesias et al. (US 20220166695 A1, hereinafter Iglesias).
Regarding claim 21, Krishnan in view of Priydarshi and further in view of Sun, and further in view of Mandal teaches the method of claim 8 (and thus the rejection of claim 8 is incorporated).
Regarding the limitation wherein the application dependency graph includes a plurality of nodes and a plurality of edges, wherein the plurality of nodes correspond to the plurality of applications, and the plurality of edges correspond to the KPIs, the plurality of nodes are linked by the plurality of edges represents dependency between the plurality of applications corresponding to the linked plurality of nodes, Krishnan teaches the plurality of applications (¶20 “the plurality of applications”) and the KPIs (¶36 “characteristic temporal error patterns or other attributes” as explained above with respect to claim 1). However, Krishnan fails to teach wherein the application dependency graph includes a plurality of nodes and a plurality of edges, wherein the plurality of nodes correspond to the plurality of applications, and the plurality of edges correspond to the KPIs, the plurality of nodes are linked by the plurality of edges represents dependency between the plurality of applications corresponding to the linked plurality of nodes.
Mandal teaches wherein the application dependency graph (Fig. 1 – 130, 160, ¶38, ¶59 all as explained above with respect to claim 6) includes a plurality of nodes and a plurality of edges, wherein the plurality of nodes correspond to components, and the plurality of edges correspond to invocations between nodes (¶60 “The causal dependency graph is typically in the form of a directed graph, with each node of the graph representing a corresponding component, and an edge of the graph connecting a first node and a second node in the directed graph indicating that a first component represented by the first node uses/consumes/invokes a second component represented by the second node”), the plurality of nodes are linked by the plurality of edges represents dependency between components corresponding to the linked plurality of nodes (Fig. 5A – 500, ¶92 “FIG. 5A illustrates a causal dependency graph (500) that captures the usage dependencies among components deployed in a computing environment during the processing of user requests in one embodiment. Causal dependency graph 500 is in the form of a directed graph, with each node of the graph representing a corresponding component… and an edge of the graph connecting a first node/component and a second node/component representing an invocation of the second component by the first component”).
Regarding the limitation wherein the plurality of edges corresponding to the KPIs are aggregated based on determining an average of number of hits, a maximum of the response time, and a minimum of the success rate by the plurality of applications, Krishnan teaches the plurality of applications (¶20). However, Krishnan fails to teach wherein the plurality of edges corresponding to the KPIs are aggregated based on determining an average of number of hits, a maximum of the response time, and a minimum of the success rate by the plurality of applications.
Mandal teaches the plurality of edges corresponding to the dependencies (Fig. 5A – 500, ¶60, ¶92 as explained above). However, Mandal fails to teach wherein the plurality of edges corresponding to the KPIs are aggregated based on determining an average of number of hits, a maximum of the response time, and a minimum of the success rate by the plurality of applications.
He, in the same field of endeavor, teaches wherein a graph’s edge features corresponding to KPIs are aggregated (Page 5, Col. 2, Section 4.3.3 ¶1 “Since each call-relationship can be represented as an edge in the issue impact topology, the constructed features are denoted as edge features… These features can describe the degree of abnormality in the KPI values,” Fig. 4, Caption: “An example of the edge features presenting the issue symptom information, which is constructed using the latest windows of KPIs,” Page 6, Col. 1, Section 4.3.4, ¶3 “this model includes an edge-conditioned graph layer… to aggregate both node features and edge features”) based on determining an average number of failures (Page 5, Fig. 4(a) depicts a graph with nodes “checkout” and “cart” and a directed edge from “checkout” to “cart,” Fig. 4(c) – “Compare Failure Count with its preceding values” depicts calculating the first value in an edge feature “E(checkout, cart)” using the average number of the most recent “Failure Counts”). However, He fails to teach determining an average number of hits, a maximum of the response time, and a minimum of the success rate by the plurality of applications.
Iglesias, in the same field of endeavor, teaches determining a number of hits (Fig. 2 – 207, ¶36 “KPIs 207 may include indicators of successful calls, such as… calls that were placed successfully (e.g., calls for which a called party acknowledged receipt of a call request), and/or other indicators of successful calls”), a maximum of the response time, and a minimum of the success rate by classification models (Fig. 1 – 101, 103, ¶18 “RAS 101 may identify failover conditions (e.g., based on KPIs associated with one or more VNFs)… classification models 103 may be used… to determine failover conditions… a “failover condition” may refer to a condition, set of conditions, criteria, or the like, that indicate that a VNF should be failed over… Such failover conditions may include… threshold values associated with one or more particular KPIs or metrics, such as a maximum latency threshold, a minimum throughput threshold, a maximum call failure rate threshold, a minimum call success rate threshold, and/or other suitable types of values, metrics, KPIs, etc.”).
Regarding the limitation and wherein the plurality of nodes is classified to a normal value, near anomaly, or anomaly based on the aggregated KPIs for the corresponding plurality of applications, Krishnan teaches and wherein a plurality of points is classified to a normal value, near anomaly, or anomaly based on an anomaly score for the corresponding plurality of applications (¶14 “ A graphical user interface (GUI) is configured to provide status alerts for the application in different colors based on the severity levels of the anomalies being detected… anomalies with anomaly scores less than the threshold may be determined to be of low severity… The GUI indicates anomalies with low severity in green color… anomalies with anomaly scores higher than the threshold but within a predetermined range of the threshold may be determined to be of moderate severity… The GUI may display anomalies with medium severity in amber color… Anomalies with very high anomaly scores may be determined to be highly severe thereby indicating an imminent application failure due to such high-severity anomalies. The GUI indicates such high-severity anomalies in red color,” Fig. 1 – 118, 122, Fig. 8 – 802, 804, 808, 822, 824, 826, ¶45 “FIG. 8 illustrates an example of the GUI 118… to determine the health of the application 122. The predictors or features 802 for estimating the probability of application failure… can be indicated on the plot 804 via points that are colored amber, green and red. The number of status alerts generated in each of the amber, red and green categories are shown on the strip 806… as indicated respectively by the shadings of the icons 822, 824 and 826. A total anomaly score for the real-time data set that is currently being analyzed on the GUI 118 can be indicated via a torus 808. The color of the torus 808 indicates the status alert of the application 122”). However, Krishnan fails to teach and wherein the plurality of nodes is classified to a normal value, near anomaly, or anomaly based on the aggregated KPIs for the corresponding plurality of applications.
Mandal teaches the plurality of nodes (Fig. 5A – 500, ¶60, ¶92). However, Mandal fails to teach and wherein the plurality of nodes is classified to a normal value, near anomaly, or anomaly based on the aggregated KPIs for the corresponding plurality of applications.
He teaches the aggregated KPIS for services (Page 5, Fig. 4 and Col. 2, Section 4.3.3, ¶1, Page 6, Col. 1, Section 4.3.4, ¶3 as explained above).
Krishnan, Mandal, He, and Iglesias are analogous art to the claimed invention as all are from the same field of endeavor of anomaly detection. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the dependency graph of Mandal, the aggregation of KPI edges of He, and the determined KPIs of Iglesias with the plurality of applications and the anomaly classification method of Krishnan. The motivation to do so is to reduce “overall cost required to resolve failure by simply preventing failure from occurring. In addition, the overall productivity of the computing environment (135) increases by minimizing chances of interruption due to failure” (Mandal, Fig. 1 – 135, ¶54), to ease “the burden of operators in the face of a flood of issues and related alert signals” (He, Abstract), and to better correlate KPIs to remediation models (Iglesias, ¶58 “AI/ML techniques or other suitable techniques… may enhance the accuracy of correlating KPIs… to particular classification models… and/or remediation models”).
Claim 22 is rejected under 35 U.S.C. 103 as being unpatentable over Krishnan in view of Priydarshi and further in view of Sun, and further in view of Mandal, and further in view of Chen et al. (“Graph-based Incident Aggregation for Large-Scale Online Service Systems,” 2021, hereinafter Chen).
Regarding claim 22, Krishnan in view of Priydarshi and further in view of Sun, and further in view of Mandal teaches the method of claim 8 (and thus the rejection of claim 8 is incorporated).
Regarding the limitation wherein the application dependency graph is integrated with an anomaly inference pipeline formed from the multiple anomaly detection model, wherein integrating the application dependency graph with the anomaly detection models is based on summarizing output of the anomaly inference pipeline, wherein summarizing the output includes mean, median and central dispersion of the anomaly inference pipeline, Krishnan teaches the multiple anomaly detection model and the anomaly detection models (¶20 “the plurality of respective data models”). However, Krishnan fails to teach wherein the application dependency graph is integrated with an anomaly inference pipeline formed from the multiple anomaly detection model, wherein integrating the application dependency graph with the anomaly detection models is based on summarizing output of the anomaly inference pipeline, wherein summarizing the output includes mean, median and central dispersion of the anomaly inference pipeline.
Mandal teaches wherein the application dependency graph (Fig. 1 – 130, 160, ¶38, ¶59 all as explained above with respect to claim 6) is integrated with an anomaly inference pipeline formed from anomaly detection components (Fig. 1 – 150, Fig. 3A – 300, Fig. 4 – 430, 440, 450A, Fig. 5A – 500, ¶82 “FIG. 4 is a block diagram illustrating an example implementation of proactive resolution tool (150). The block diagram is shown containing… outlier detector 430 (in turn shown containing metrics outlier detection engine 435A and logs outlier mining engine 435B), performance degradation predictors (PDP) 440 (in turn shown containing incident predictor 450A and response time predictor 450B),” ¶89 “PDP (performance degradation predictors) 440 is a set of engines/software modules that proactively predict potential future problems/imminent performance issues… PDP 440 is shown containing incident predictor 450A,” ¶91 “incident predictor 450A forms a causal dependency graph captures such causes and effects related to a software application (300) by representing usage dependencies among components,” ¶92 “FIG. 5A illustrates a causal dependency graph (500) that captures the usage dependencies among components,” wherein the “proactive resolution tool” depicted in Figure 4 encompasses an anomaly inference pipeline, given its broadest reasonable interpretation of an anomaly detection process with multiple sequential components), integrating the application dependency graph with anomaly detection components (Fig. 4 – 440, 450A, 470, Fig. 5A, ¶81 “PDP (performance degradation predictors) 440 is a set of engines/software modules… shown containing incident predictor 450A,” ¶91 “incident predictor 450A forms a causal dependency graph,” ¶109 “root cause analyzer 470 is implemented to take as inputs the imminent performance issues identified by performance degradation predictors 450… and generates as outputs a ranked list of probable root cause(s) for the predicted performance issues, corresponding confidence scores”), and the anomaly inference pipeline (Fig. 1 – 150, Fig. 4, ¶81-92 as explained above). However, Mandal fails to teach … wherein integrating the application dependency graph with the anomaly detection models is based on summarizing output of the anomaly inference pipeline, wherein summarizing the output includes mean, median and central dispersion of the anomaly inference pipeline.
Chen, in the same field of endeavor, teaches wherein integrating a failure impact graph is based on summarizing output of an incident aggregation framework (Page 1, Col. 2, ¶1-3 “Without automated incident aggregation, engineers may need to go through each incident to discover the existence of such a problem and collect all related incidents to understand it… we propose GRLIA… an incident aggregation framework to assist engineers in failure understanding and diagnosis,” wherein performing “incident aggregation” to simplify complex data into a format that human engineers can understand encompasses summarizing, Page 4, Fig. 3 and Col. 1, Section A, ¶1 “The overall framework of GRLIA… consists of four phases, i.e., service failure detection, failure-impact graph completion, graph representation learning, and online incident aggregation… we utilize the trends observed in KPI curves to auto-complete the failure-impact graphs. After obtaining the set of incidents associated with each failure, in the third phase, an embedding vector is learned for different types of incidents by leveraging existing graph representation learning models… Such representation encodes not only the temporal locality of incidents, but also their topological relationship… the learned incident representation will be employed for online incident aggregation by considering their cosine similarity and topological distance… GRLIA essentially learns the correlations among incidents, which are also applicable to the changed portion of the topology”), wherein summarizing the output includes mean, median, and central dispersion of the incident aggregation framework (Page 6, Col. 2, Section E, ¶1 “The EVT-based method also plays a role… by continuously monitoring the number of incidents per minute. If it alerts a failure, the online incident aggregation will be triggered,” Page 4, Fig. 3 and Col. 2, Section B, ¶2 “For time-series data, anomalies often manifest themselves as having a large magnitude of upward/downward changes. Extreme Value Theory (EVT)… is a popular statistical tool to identify data points with extreme deviations from the median of a probability distribution… to detect bursts in time series of the number of incidents per minute… The bursts are regarded as the occurrence of service failures… Fig. 3 (phase one) presents an example of service failure detection, where all abnormal spikes are successfully found by the decision boundary (the orange dashed line),” wherein central dispersion is implicit when measuring “extreme deviations” from the median, Page 5, Col. 2, ¶2-3 “the KPI trend similarity measures the underlying consistency of cloud components’ abnormal behaviors… KPIs should be utilized for similarity evaluation… EVT introduced in phase one is utilized again to detect anomalies for each KPI. Only the abnormal KPIs shared by two connected cloud components will be compared… when there exists more than one type of abnormal KPIs, we use the average similarity score,” Equation (3) depicts calculating a mean or the average or similarity scores between KPIs, Page 6, Col. 1, ¶2 “for each discovered community, the incidents inside it form the complete impact graph of the service failure”).
Krishnan, Mandal, and Chen are analogous art to the claimed invention as all are from the same field of endeavor of anomaly detection. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the integrated graph and pipeline of Mandal and the output summarization of Chen with the anomaly detection models Krishnan. The motivation to do so is to reduce “overall cost required to resolve failure by simply preventing failure from occurring. In addition, the overall productivity of the computing environment (135) increases by minimizing chances of interruption due to failure” (Mandal, Fig. 1 – 135, ¶54) and to better “assist engineers in failure understanding and diagnosis” (Chen, Page 11, Col. 2, Section VII, ¶1).
Response to Amendment
In review of Applicant’s amendments, filed April 13, 2026, the objections to the specification made in the previous office actions have been withdrawn.
Examiner notes Applicant’s remark that they submitted a replacement Figure 4. However, no such figure was found with the April 13, 2026 submission. Therefore, Figure 4 from March 6, 2023 is in effect, and the objection to such remains in place.
Response to Arguments
Applicant’s amendments and arguments, see page 15 filed on April 13, 2026, regarding the abstract idea rejections from the previous office action made under 35 U.S.C. 101 have been fully considered but are not persuasive.
On page 17 of the Remarks, Applicant asserts that “the anomaly detection model which is… trained concurrently with other machine learning models by performing preprocessing operations and post processing operations to detect an anomaly, is clearly not a mathematical concept.” Applicant cites as reasoning an excerpt from “reminders on subject matter eligibility of claims guideline dated August 4, 2025” and paragraphs 57 and 59 of the present application. Examiner agrees with the assertion. However, claim 1 still recites “decomposing… log data,” “clustering,” “determining… a training sequence,” and “to detect occurrence of an anomaly… based on received application data,” all of which can reasonably be interpreted as mathematical concepts. Thus, the claim is not subject matter eligible.
On page 18 of the Remarks, Applicant argues that “Even assuming the Examiner’s characterization is correct… Applicant asserts that claim 1, as a whole, integrates ‘a practical application’ at least because the additional elements of claim 1, apply or use the alleged judicial exception in some other meaningful way.” Applicant submits that the amended claim 1 is patent eligible because it directly improves the technology by improving the scheduling of training ML models that results in faster training. Examiner respectfully disagrees. Although applicant asserts that the amended features in claim 1 result in “faster and more resource-efficient training… reducing a multi-day or multi-week training time to a multi-hour training time” and that “The techniques provided herein provide the improved training efficiency and reduction in false positives of failure prediction without requiring changes to the code of the underlying applications and microservices,” as supported in paragraph 6 of the present application’s specification, the specification only sets forth the improvement in a conclusory manner, while the claim itself fails to reflect the disclosed improvements of “faster and more resource-efficient training” and “improved training efficiency and reduction in false positives” after training models concurrently and performing one or more same preprocessing operations, post-processing operations, or a combination thereof (see MPEP 2106.04(d)(1) “if the specification explicitly sets forth an improvement but in a conclusory manner (i.e., a bare assertion of an improvement without the detail necessary to be apparent to a person of ordinary skill in the art), the examiner should not determine the claim improves technology”). For example, it is unclear how a person of ordinary skill in the art would find it apparent that the amended claim 1 reduces “a multi-day or multi-week training time to a multi-hour training time” or predicts “application failures with an approximately 30-40% reduction in false positives compared to rule-based anomaly detection systems,” as explained in paragraph 6 of the specification.
In consideration of this conclusion, independent claims 1, 14, and 19 and their associated dependent claims 2-13 and 21-22, 15-18, and 20, respectively, are subject matter ineligible and, thus, the rejections under 35 U.S.C. 101 stand.
Applicant’s arguments, filed April 13, 2026 regarding the rejections from the previous office action made under 35 U.S.C. 103 have been fully considered but are moot as they do not apply to the reference Sun being used in the current rejections of 1, 14, and 19 and their associated dependent claims 2-13 and 21-22, 15-18, and 20, respectively, to teach the amended claim limitations directed to training models concurrently and performing one or more same preprocessing operations, post-processing operations, or a combination thereof.
On page 23 of the Remarks, Applicant argues that “Liu merely teaches that in one iteration of model training, stages need to be completed in sequence… but fails to teach about performing one or more same preprocessing operations, one or more same post-processing operations, or a combination thereof.” Applicant cites as reasoning paragraphs 32 and 58 of Liu, and asserts that “Liu merely discloses about training of deep learning models… At most, Liu discloses about scheduling mode of training and the utilization of model training resources is higher in one mode than the other.” Examiner agrees with the assertion. However, the new reference, Sun, implements “temporal sharing… by overlapping data preprocessing and computations” (Page 1, Col. 2, ¶2) into its concurrent training framework, which explicitly teaches the concurrent sharing of preprocessing operations of different training tasks on the same hardware resources, as opposed to the sharing of model training resources in sequence as described in Liu.
Additionally, the references He and Iglesias teach the limitations of the new claim 21 directed to “the application graph… wherein the plurality of edges corresponding to the KPIs are aggregated… and wherein the plurality of nodes is classified,” and the reference Chen teaches the limitation of the new claim 22 directed to “wherein the application dependency graph is integrated with an anomaly inference pipeline…”
With the addition of the reference Sun teaching the subject matter introduced in the amendments and the references He, Iglesias, and Chen teaching the subject matter in the new claims, the rejections under 35 U.S.C. 103 stand.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to WILLIAM M LEE whose telephone number is (571)272-4761. The examiner can normally be reached Mon-Fri. 8am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached at (571)272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/WILLIAM MICHAEL LEE/
Examiner, Art Unit 2145
/CESAR B PAULA/Supervisory Patent Examiner, Art Unit 2145