Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on July 2, 2026 has been entered.
Response to Arguments
Applicant's arguments filed July 2, 2026 have been fully considered but they are not persuasive. References Cited in the Office Action are: U.S. Pat. 12,050,715 to Ackerman et al. (hereinafter "Ackerman") U.S. Pat. Pub. 20220094713 to Lee et al. (hereinafter "Lee") U.S. Pat. Pub. 20230134546 to Gopalakrishnan et al. (hereinafter "Gopalakrishnan") U.S. Pat. 10,673,880 to Pratt et al. (hereinafter "Pratt") U.S. Pat. Pub. 20190034932 to Sadaghiani et al. (hereinafter "Sadaghiani").
In pages 1-5 of the remarks, Applicant states that the Office rejected claims 1-20 under 35 U.S.C. § 112(a) as "failing to comply with the written description requirement" because "[t]he algorithm or steps/procedures for these claimed functions is not explained at all or is not explained in sufficient detail." Office Action (“OA”), page 11. Applicant has amended the independent claims 1, 9, and 17 to remove the limitations of "'in response to a determination that the likelihood exceeds a predetermined threshold and a determination that a first of the host devices has been previously accessed by unauthorized malicious cybersecurity actors, instructing a first of the host devices to remain offline for a predetermined period of time, and instructing at least one of the host devices to implement additional authentication techniques"' and "'wherein the first host device has private user information stored thereon"' that were both described in the Specification of the Applicant in paragraph [0076], but that no other information is provided as to how the limitation is performed, OA, pages 11-12. Furthermore, Applicant states that the term “private user information”, specifically, the term “private” has been moved from the independent claims to a new dependent claim 21. Next, claim 1 has been amended to remove "wherein training of the second model during the second phase includes: generating a time-series matrix that is based on the embedding vectors generated by the first model formed by low-dimensional causal event streams with predetermined time-bins and labels […] and using the final CLS token embedding to pretrain the second model, wherein the pretraining comprises the second model classifying inputs as malicious or normal." Claim 4 has been amended to recite "...will occur within a second predetermined period of time." Claims 12 and 20 are similarly amended. These amendments are supported by the originally- filed Specification, inter alia, by as-filed claims 4, 12, and 20 as well as Specification para. [0009], [0028], and [0071]. Claims 6 and 14 now recite “wherein training of the second model during the second phase includes: determining labeled examples and causing the second model to classify whether each of the labeled examples represents abnormal or potentially malicious behavior, wherein a first of the labeled examples is based on an anomaly, and wherein a second of the labeled examples is based on detected malicious activities”. Lastly, claims 8 and 16 have been amended to remove the limitations of 'adding the determined embedding vectors that fall within a first of the time slots into a first sum embedding vector, wherein the first sum embedding vector is the only embedding vector of the first time slot', 'adding the determined embedding vectors that fall within a second of the time slots into a second sum embedding vector, wherein the second sum embedding vector is the only embedding vector of the second time slot', and 'stacking the sum embedding vectors into the two- dimensional matrix'. Applicant respectfully submits that these rejections under 35 U.S.C. § 112(a) are overcome at least by the amendments to the claims. Applicant therefore submits that, for at least the amendment made thereto, the claims are in condition for allowance and requests reconsideration and withdrawal of the rejections under 35 U.S.C. § 112. As a result of the amendments being made to the aforementioned claims recited above, Examiner withdraws the rejections made under 112(a) for claims 1-20.
Claims 1, 4, 7, 9, 12, 15, 17, and 20 stand rejected under 35 U.S.C. § 102(a) as anticipated by Ackerman. Claims 2, 3, 10, 11, 18, and 19 stand rejected under 35 U.S.C. § 103 as unpatentable over Ackerman in view of Lee. Claims 5, 6, 13, and 14 stand rejected under 35 U.S.C. § 103 as unpatentable over Ackerman in view of Gopalakrishnan. Claims 8 and 16 stand rejected under 35 U.S.C. § 103 as unpatentable over Ackerman in view of Pratt in further view of Sadaghiani in yet further view of Gopalakrishnan. Claim 20 has been canceled in the most recent amendments, with claim 21 being newly added. Without acquiescing to the rejection and solely to expedite prosecution, Applicant amends the independent claims to highlight identified novel subject matter. Specifically, independent claims 1, 9, and 17 have been amended to specify "wherein the hierarchical temporal event transformer model applies time-token embeddings to a hierarchical attention graph to generate a final classification (CLS) token embedding" to clarify advantages over the prior art. See Specification, para. [0069] and [0082] as well as FIG. 4C. Applicant states that the require generation of a CLS token embedding from time-token embeddings through a hierarchical attention-based process. Claims. Ackerman produces a classification result from decision tree processing. See Ackerman, col. 26, lines 53-64. The cited "class of the trees" output of Ackerman is not generated by applying time-token embeddings to a hierarchical attention graph to generate a final classification (CLS) token embedding as required by the claims. Id. The cited portions of Ackerman describe random forest classification outputs rather than the claimed hierarchical-attention-based processing pipeline. Ackerman, col. 26, lines 53- 64; Claims; Specification, para. [0082] and FIG. 4C. Applicant states that the claimed intermediate representation and processing architecture are fundamentally different. Next, Applicant states that the Office equates Ackerman's output of a class from a collection of decision trees with the claimed generation of a final classification (CLS) token embedding. Office Action, page 18. However, the claims do not merely require a classification result, but a specific processing architecture in which time-token embeddings are applied to a hierarchical attention graph to generate a final classification (CLS) token embedding. Applicant also states that the cited portions of Ackerman do not disclose the claimed intermediate representation, applying time-token embeddings to a hierarchical attention graph, or generating a final classification (CLS) token embedding through the claimed processing sequence. See Ackerman, e.g., col. 26, lines 53-64. Accordingly, the rejection relies on functional similarity rather than disclosure of the claimed structure and process. Furthermore, Applicant respectfully submits that the rejection does not establish that Ackerman discloses the claimed generation of a final classification (CLS) token embedding. The rejection cites col. 26, lines 59-64 as allegedly anticipating the CLS token of the pending claims, but that the cited passage instead describes a random forest classifier producing a classification result from a collection of decision trees. Ackerman, col. 26, lines 53-64. The cited passage does not disclose a CLS token embedding, time-token embeddings, or a hierarchical attention graph that generates the CLS token embedding as required by the claims. Ackerman, col. 26, lines 53-64. Applicant states that the cited disclosure does not teach the claimed intermediate representation or the claimed processing sequence used to generate classification, and therefore, does not establish anticipation.
Examiner reiterates that in claim 1’s limitation of “applying the time-token embeddings for the period of the recent logged event stream to a hierarchical attention graph to generate a final classification (CLS) token embedding”, the passage of Ackerman’s [Col. 26, lines 59-64] random forests are ensemble learning methods used for classification, and construct multitude of decisions trees at training time, outputting the class of the trees, this corresponds to generating a final classification (CLS) token, and in particular, the statement that ‘multi-dimensional description of events occurring over time for the entity generating of time-token embeddings for period of recent logged event streams when taking time stamps in event vectors 1110 into account, stated in [Col. 32, lines 29-31]’ occurs along with the statement that Ackerman’s passage of [Col. 32, lines 38-44] entity model 1120 can be a statistical model based on a history of events 1106 of event stream 1114 for the entity over time, such as via a window or rolling average of events 1106, corresponding to convolution of the time-series matrix. With the outputting of the class of the trees in particular, the class output corresponds to the final CLS token that is a part of the multi-dimensional description of events for the embeddings of the logged event streams. Furthermore, the ‘hierarchical attention graph’ is described as the random forests in Ackerman’s [Col. 26, lines 59-64], and the time-token embeddings are described in Ackerman’s [Col. 32, lines 29-31] in the multi-dimensional description of events occurring over time for the entity generating of time-token embeddings for period of recent logged event streams. As a result, Examiner maintains the rejections previously made to claims 1-19 under their respective prior art combinations. Claim 20 has been canceled.
Finally, in pages 8-10, Applicant states that claim 21 has been newly added, reciting "wherein the user information is private user information", with support in the Specification as described in paragraphs [0002]. [0052], and [0076]. Examiner states that as Ackerman’s section [Col. 34, lines 64-67] describes sensitive information as private user information stored on the host device, claim 21 is rejected as being anticipated by Ackerman.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1, 9, 17, and 21 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Ackerman et al (US 12050715 B2), hereinafter Ackerman.
Regarding claim 1, Ackerman discloses “a computer-implemented method, comprising: collecting historical event log data from host devices” ([Col. 30, lines 27-31] "FIG. 11 shows a system for event monitoring and response. In general, the system may include a number of compute instances 1102 that use local security agents 1108 to gather events 1106 from sensors 1104 into event vectors 1110, and then report these event vectors 1110 to a threat management facility 1112", wherein the event data being collected from sensors from a compute instance or compute instances, to which this process corresponds to the collection of event log data from host devices to then convert into vectors.);
“training a first model to convert textual log events of the historical event log data into event embedding vectors” ([Col. 30, lines 27-31] "FIG. 11 shows a system for event monitoring and response. In general, the system may include a number of compute instances 1102 that use local security agents 1108 to gather events 1106 from sensors 1104 into event vectors 1110", where events themselves are converted into event vectors, with [Col. 30, line 67-Col. 31, line 1] stating, 'events 1106 may be tokenized', and events can be assigned numbers or other identifiers to then be added to a vector, and in section [Col. 32, lines 21-24] and as shown in Fig. 11, a vector can be any size, and can encode any number of different events 1106 that may be tokenized, which allows the event vectors of Ackerman correspond to event embedded vectors of the applicant, such that the event vectors can go into a threat management facility, and be stored in an event stream for a second model of classifying behavior will utilize the vectors. Earlier in the prior art of Ackerman in [Col. 18, line 55-Col. 19, line 5], the security agent 306 is also described to apply machine learning models in order to detect the logs or other threats included in an endpoint 302, or other event types that may also occur while running programs, and logging those events.);
“training a second model to classify whether at least some of the event embedding vectors represent abnormal or potentially malicious behavior, wherein the second model is a hierarchical temporal event transformer model” ([Col. 30, lines 36-40] "The event stream 1114 may be analyzed with an analysis module 1118, which may in turn create entity models 1120 useful for detecting, e.g., unexpected variations in behavior of compute instances 1102… ", where the analysis module is capable of supplying the vectors and even preemptively analyzing some unusual behavior. [Col. 32, lines 38-44] "Each entity model 1120 may, for example, include a multi-dimensional description of events 1106 for an entity based on events 1106 occurring over time for that entity. This may be, e.g., a statistical model based on a history of events 1106 for the entity over time... entity models 1120 may, for example, be vector representations or the like of different events 1106 expected for or associated with an entity, and may also include information about the frequency, magnitude, or pattern of occurrence for each such event 1106", where the second model is capable of creating a statistical model that represents a time period of events that is then transformed into a vector for the statistical model, that takes into account events that have occurred various times in the events that the second model is being trained on. The categorization of the frequency, patterns of the events are important when detecting anomalies or dangerous behaviors of the events if they occur more frequently than other events, and therefore, corresponds to the hierarchical temporal event transformer model.),
“wherein the second model is trained in a first phase and a second phase and wherein the hierarchical temporal event transformer model applies time-token embeddings to a hierarchical attention graph to generate a final classification (CLS) token embedding” ([Col. 32, lines 33-37] Second model uses the vectors from the various entities to determine whether an events are considered safe and/or malicious, or otherwise suspicious behavior, based on how the events play out. Creating the entity models and determining normal behavior consists of the first phase. [Col. 33, lines 5-7] includes the training according to anomalies or dangerous behaviors for the second phase of training the second model. [Col. 26, lines 59-64] Random forests are ensemble learning methods used for classification, and construct multitude of decisions trees at training time, outputting the class of the trees, this corresponds to generating a final classification (CLS) token embedding of the Applicant. In the context of Ackerman, input of these random forests can include safe and unsafe samples of events, as stated in [Col. 26, lines 22-24]. Also, multi-dimensional description of events occurring over time for the entity generating of time-token embeddings for period of recent logged event streams when taking time stamps in event vectors 1110 into account, stated in [Col. 32, lines 29-31].),
“deploying the trained first model and the trained second model to predict a likelihood of a malicious cybersecurity event occurring within a first predetermined period of time that begins at the deployment of the trained models” ([Col. 34, lines 35-38] To deploy the first model and the second model in order to detect malicious or anomalous events, "the detection engine 1122 may compare new events 1106 generated by an entity, as recorded in the event stream 1114, to the entity model 1120 that characterizes a baseline of expected activity", which is able to obtain the event vectors from the data repository 1116, shown in Fig. 11. [Col. 34, lines 23-25] "Once an entity model 1120 has been created and a stable baseline established, the entity model 1120 may be deployed for use in monitoring prospective activity", in which the entity model being deployed corresponds to the trained second model being deployed to predict malicious events. With Fig. 3 as an enterprise network threat detection, various endpoint systems can be evaluated at one time, and the endpoint 302 is used to connect to the threat management facility 308, which is also in Fig. 11 as threat management facility 1112, which may also use a filter to "manage a flow of information from the data recorder 304 to a remote resource such as the threat detection tools 314 of the threat management facility 308", as stated in [Col. 19, lines 19-22]. When a connection has been established, a request for specific events, or events that can occur over a time frame, can monitor a data log, which is converted into an event vector and will be detected by a detection engine, can be done in a specified time frame from when it began.).
“wherein the likelihood is a classification output generated by the trained second model as a result of a two-dimensional matrix output by the trained first model being applied to the second model” ([Col. 34, lines 35-53] "The detection engine 1122 may compare new events 1106 generated by an entity, as recorded in the event stream 1114, to the entity model 1120 that characterizes a baseline of expected activity… comparison may use one or more vector distances such as a Euclidean distance, a Mahalanobis distance, a Minkowski distance, or any other suitable measurement of difference within the corresponding vector space. In another aspect, a k-nearest neighbor classifier may be used to calculate a distance between a point of interest and a training data set, or more generally to determine whether an event vector 1110 should be classified as within the baseline activity characterized by the entity model", where the comparison uses vector distances and operations to compare them, of which Minkowski, Euclidean, Mahalanobis are well known in the arts to perform operations in any number of dimensions, including two-dimensional matrices. [Col. 32, lines 38-41] Fig. 11, each entity model 1120 includes a multi-dimensional description of events 1106 based on event stream 1114, corresponding to a two-dimensional matrix output by the trained first model being applied to the trained second model. Furthermore, detection engine 1122 is applied to event stream 1114 to detect unusual or malicious activity based on entity models 1120, corresponding to a likelihood is a classification output generated by the trained second model of the Applicant.);
“and in response to a determination that the likelihood exceeds a predetermined threshold and a determination that a first of the host devices has been previously accessed by unauthorized malicious cybersecurity actors, instructing a first of the host devices to remain offline for a predetermined period of time and instructing at least one of the host devices to implement additional authentication techniques” ([Col. 9, lines 60-63] Risk scores can be determined for users by identity management facility 172, and [Col. 34, lines 12-16] also describes that different users may use the software differently, with behavior falling outside of the baseline corresponding to unauthorized malicious cybersecurity actors. [Col. 32, lines 38-41] Fig. 11, each entity model 1120 includes a multi-dimensional description of events 1106 based on event stream 1114, corresponding to a two-dimensional matrix output by the trained first model being applied to the trained second model, and in [Col. 35, lines 21-32] it is stated that when an event stream 1114 deviates from a baseline of expected activity that is described in the entity models 1120 for one or more entities, responses 1124 are initiated, including termination of network communication and quarantining the entity, corresponding to instructing a first of the host devices to remain offline for a predetermined period of time, in conjunction with filtering of a particular endpoints that should not exceed a predetermined time, such as an hour, stated in [Col. 37, lines 11-15]. [Col. 9, lines 60-63] The identity management facility 172 may determine a risk score for a user based on the events, with an identity provider also providing operations to confirm identity of the user based on the events received.),
“wherein the first host device has user information stored thereon” ([Col. 37, lines 11-15] Instructing a first of the host devices to remain offline for a predetermined period of time, in conjunction with filtering of a particular endpoints that should not exceed a predetermined time, such as an hour. [Col. 34, lines 64-67] Events that indicate malicious activity include transmitting sensitive information from an endpoint, where sensitive information corresponds to user information stored on the host device.).
Regarding claim 9, Ackerman discloses the computer-implemented method of claim 1 above. Ackerman also discloses "a computer program product, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions readable and/or executable by a computer to cause the computer to: collect historical event log data from host devices" ([Col. 56, lines 63-67] "Embodiments disclosed herein may include computer program products comprising computer-executable code or computer-usable code that, when executing on one or more computing devices, performs any and/or all of the steps thereof. The code may be stored in a non-transitory fashion in a computer memory, which may be a memory from which the program executes (such as random-access memory associated with a processor), or a storage device such as a disk drive, flash memory or any other optical, electromagnetic, magnetic, infrared, or other device or combination of devices", with the products being housed inside of non-transitory storage mediums such as disk drives and other forms of physical media. The computer is capable of reading the code of the program product and be executable on the computing devices. [Col. 30, lines 27-31] "FIG. 11 shows a system for event monitoring and response. In general, the system may include a number of compute instances 1102 that use local security agents 1108 to gather events 1106 from sensors 1104 into event vectors 1110, and then report these event vectors 1110 to a threat management facility 1112", wherein the event data being collected from sensors from a compute instance or compute instances, to which this process corresponds to the collection of event log data from host devices to then convert into vectors.);
Regarding claim 17, Ackerman discloses the computer-implemented method of claim 1 above. Ackerman also discloses “a system, comprising: a processor” ([Col. 56, lines 30-40] “The hardware may include a general-purpose computer and/or dedicated computing device. This includes realization in one or more microprocessors, microcontrollers... along with internal and/or external memory.”);
“and logic integrated with the processor, executable by the processor, or integrated with and executable by the processor, the logic being configured to: collect historical event log data from host devices” ([Col. 56, lines 30-40] "The hardware may include a general-purpose computer and/or dedicated computing device. This includes realization in one or more microprocessors, microcontrollers... along with internal and/or external memory. This may also, or instead, include one or more application specific integrated circuits, programmable gate arrays, programmable array logic components, or any other device or devices that may be configured to process electronic signals", with the logic expressed as specific integrated circuits, programmable gate arrays, or other devices configured to process electronic signals. [Col. 30, lines 27-31] "FIG. 11 shows a system for event monitoring and response. In general, the system may include a number of compute instances 1102 that use local security agents 1108 to gather events 1106 from sensors 1104 into event vectors 1110, and then report these event vectors 1110 to a threat management facility 1112", wherein the event data being collected from sensors from a compute instance or compute instances, to which this process corresponds to the collection of event log data from host devices to then convert into vectors.);
Regarding claim 21, Ackerman discloses the computer-implemented method of claim 1. Ackerman also discloses “wherein the user information is private user information” ([Col. 34, lines 64-67] Events that indicate malicious activity include transmitting sensitive information from an endpoint, where sensitive information corresponds to private user information stored on the host device.).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 2-4, 10-12, and 18-19 are rejected under 35 U.S.C. 103 as being unpatentable over Ackerman in view of Lee et al. (US 20220094713 A1), hereinafter Lee.
Regarding claim 2, Ackerman discloses the computer-implemented method of claim 1, as described above. Ackerman does not appear to teach ‘wherein the first model is trained to use natural language modeling to map tokens of the textual log events of the historical event log data to the event embedding vectors via a lookup table’. However, Lee teaches that, wherein the first model is trained to use natural language modeling to map tokens of the textual log events of the historical event log data to the event embedding vectors via a lookup table, wherein the trained models are deployed at the current time ([0051-0054] "FIG. 3 illustrates machine learning models. A first machine model 310 may be trained to perform a natural language task… A feature extractor may provide feature vectors representing natural language text to the embedding 312 of the first machine learning model 310", where the embedding of the first model will assist in mapping the inputs into vector space, and the files themselves have timestamp information so to keep track of the files and the event that will take the files and vectorize the information, which [0032] states information about files and time stamps. Furthermore, the vectors were stored in memory, and as paragraph [0048] states, "this feature vector 140 and/or the set of features may be stored in the memory 120. The feature extractor 112 also may determine contextual information for the file", to which the vector has information just like a file does, including a time stamp for the files it represents.).
Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of Ackerman, and Lee before them, to include Lee’s ‘wherein the first model is trained to use natural language modeling to map tokens of the textual log events of the historical event log data to the event embedding vectors via a lookup table’ in Ackerman’s ‘computer-implemented method, comprising: collecting historical event log data from host devices’. One would have been motivated to make such a combination to increase efficiency by having a "first model 310 may preprocess raw text and convert the text as a sequence of word tokens. The NLP may use a pre-defined vocabulary for tokenization as in traditional models, and/or sub-word tokenizers such as those employed in BERT and GPT". With that said, the embedding layer in the first model will take the input, that being raw text, and process that into vector embedding space, as taught by Lee [0054-0055]. With this in mind, the vectors that are used can be used in a machine learning model to then analyze the information to determine if the information present in the vector could either be anomalous or dangerous to the device or a network, as stated in Lee [0039].].
Regarding claim 3, Ackerman in view of Lee teaches the computer-implemented method of claim 2, as described above. Ackerman does not appear to teach, but Lee teaches that wherein the first model is a bidirectional encoder representations from transformers (BERT) model, wherein training the first model includes using masked-language modeling to learn to predict randomly masked words within a sentence that originates from the historical event log data, wherein the historical event log data is extracted from semi-structured event logs ([0053] "a BERT model may be trained to predict masked words in a sentence with large-scale datasets", where the BERT model is used on the first model to predict the masked word based on the context of a log or sentence, by taking into account words that appear behind of and ahead of the masked word that needs to be filled in. The BERT model is trained by sentences for predicting masked words in sentences from large-scale datasets, as stated in paragraphs [0052]-[0053]. [0033] Training data can include email messages that have been labeled as malicious or benign, and includes contextual information for the messages, such as time zone information, profile information, and even timestamps, which can correspond to historical event log data as training data. Furthermore, paragraph [0037] features a feature extractor 112 being configured to receive an analysis object, such as one or more of a file or a message, and paragraph [0014] further states that extracting words is used as training data for the models recited in paragraphs [0052]-[0053] of Lee. Finally, pre-training is used with a large unlabeled dataset and fine-tuning with a small, labelled dataset, corresponding to semi-structured event logs of the Applicant, stated in paragraph [0053] of Lee.).
Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of Ackerman, and Lee before them, to include Lee’s ‘wherein the first model is a bidirectional encoder representations from transformers (BERT) model, wherein training the first model includes using masked-language modeling to learn to predict randomly masked words within a sentence that originates from the historical event log data, wherein the historical event log data is extracted from semi-structured event logs’ in Ackerman’s system performing ‘computer-implemented method, comprising: collecting historical event log data from host devices’ and Lee’s ‘wherein the first model is trained to use natural language modeling to map tokens of the textual log events of the historical event log data to the event embedding vectors via a lookup table’. One would have been motivated to make such a combination to increase efficiency by utilizing a BERT model on the first machine learning model 310, by using a transformer, a "natural language processing model such as a Bidirectional Encoder Representations from Transformers (BERT) model, transformer layers can be replaced with simplified adapters without significant loss of predictive ability", which simplifies the process of training information, and therefore, detection of anomalous or malicious activity, as taught by Lee [0004].
Regarding claim 4, Ackerman in view of Lee teaches the computer-implemented method of claim 2 above. Ackerman also discloses wherein the first model is different than the second model, wherein training of the second model during the first phase includes: ([Col. 32, lines 33-37] "In general, an analysis module 1118 may analyze the event stream 1114 to identify patterns of events 1106 within the event stream 1114 useful for identifying unusual or suspicious behavior… may include creating entity models 1120 that characterize behavior of entities", where the second model uses the vectors from the various entities to determine whether an events are considered safe and/or malicious, or otherwise suspicious behavior, based on how the events play out. Creating the entity models and determining normal behavior consists of the first phase based on the applicant, as Ackerman [Col. 33, lines 5-7] states that, "once an entity model is created, the entity model may usefully be updated, which may occur at any suitable intervals according to, e.g., the length of time to obtain a stable baseline... or any other factors", which includes the training according to anomalies or dangerous behaviors for the second phase of training the second model.):
determining a subset of the event embedding vectors of the first model to use as training targets, and causing the second model to estimate whether events associated with the training targets will occur within a second predetermined period of time from the current time ([Col. 32, lines 45-57] "The entity models 1120 may, for example, be vector representations or the like of different events 1106 expected for or associated with an entity... entity model 1120 may be based on an entity type... which may have a related event schema that defines the types of events 1106 that are associated with that entity type. This may usefully provide a structural model for organizing events 1106 and characterizing an entity before any event vectors 1110 are collected, and/or for informing what events 1106 to monitor for or associate with a particular entity", where certain events are selected by the analysis module 1118 in Fig. 11 are highlighted for their frequency, type of activity, or other metrics that indicate an event's behavior and whether that would cause issues with threats appearing, or whether it is considered safe behavior within the entity that the second model, or entity model, has to work with and surveil the events while monitoring said events. [Col. 32, line 58-Col. 33, line 2] "As an event stream 1114 is collected, a statistical model or the like may be developed for each event 1106 represented within the entity model so that a baseline of expected activity can be created... The entity model may also or instead be created by observing activity by the entity... monitoring the entity for an hour, for a day, for a week, or over any other time interval suitable for creating a model with a sufficient likelihood of representing ordinary behavior to be useful as a baseline as contemplated herein", where the events that were specified to be closely monitored or otherwise taken into account in Ackerman [Col. 32, lines 45-57], and are monitored with a prediction to see if the events will occur within the time that the system is monitored for from the time that the monitoring begins.).
Regarding claim 10, Ackerman discloses the computer program product of claim 9 as described above. Ackerman in view of Lee also teaches the limitations of claim 2 above.
Regarding claim 11, Ackerman in view of Lee teaches the computer program product of claim 10 as described above. Ackerman in view of Lee also teaches the limitations of claim 3 above.
Regarding claim 12, Ackerman in view of Lee teaches the computer program product of claim 10 as described above. Ackerman also discloses the limitations of claim 4 above.
Regarding claim 18, Ackerman discloses the system of claim 17 as described above. Ackerman in view of Lee also teaches the limitations of claim 2 above.
Regarding claim 19, Ackerman in view of Lee teaches the system of claim 18 as described above. Ackerman in view of Lee also teaches the limitations of claim 3 above.
Claims 5-7, and 13-15 are rejected under 35 U.S.C. 103 as being unpatentable over Ackerman in view of Lee, and further in view of Gopalakrishnan et al. (US 20230134546 A1), hereinafter Gopalakrishnan.
Regarding claim 5, Ackerman in view of Lee teaches the method of claim 4 above. Ackerman in view of Lee does not appear to suggest, but Gopalakrishnan teaches also discloses “wherein labels identify key detection events from a base cybersecurity application as labels for events of interest in training” ([Col. 21, lines 12-17] Fig. 4, threat management system contains a coloring system 410 as a component to support threat detection, with the coloring system used to label or color software objects to improve tracking and detection of potentially harmful activity, such as labeling files, executables, processes, and so forth with suitable information. As a result, the coloring system is capable of identifying key detection events to label events that are potentially malicious and train the threat management system further.).
However, Gopalakrishnan teaches that, wherein training of the second model during the second phase includes: determining labeled examples, and causing the second model to classify whether the labeled examples represent abnormal or potentially malicious behavior ([0032] "the ML engine may use NLP to convert log text to numerical vectors and apply a trained classifier to the numerical vectors to generate a prediction", where the classifier is used to classify the data to generate a prediction of the model and the possibility of malware showing up. [0094] "FIG. 7 illustrates an example application of a model for analyzing a network threat associated with an event in accordance with some embodiments. Table 700 identifies a set of textual tokens and scores associated with an event log", where the vectors are classified as safe, anomalous, or dangerous based on the colors green, amber, or red as stated in both paragraph [0094] and [0067]. Fig. 7 shows an example of the second model of how the classification works.).
Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of Ackerman, Lee, and Gopalakrishnan before them, to include Gopalakrishnan’s ‘wherein training of the second model during the second phase includes: determining labeled examples, and causing the second model to classify whether the labeled examples represent abnormal or potentially malicious behavior’ in Ackerman’s ‘computer-implemented method, comprising: collecting historical event log data from host devices’. One would have been motivated to make such a combination to increase security, as when the vectors are determined by a weighting score in Fig. 7, via the vectors being determined by a possibility of whether a malicious process will happen, and by using, for instance, three trees as shown in Fig. 7, have the trees vote for whether an attack could happen, as taught in Gopalakrishnan [0094].
Regarding claim 6, Ackerman in view of Lee, and further in view of Gopalakrishnan teaches the method of claim 5 above. Ackerman also discloses “wherein training of the second model during the second phase includes: determining labeled examples and causing the second model to classify whether each of the labeled examples represents abnormal or potentially malicious behavior” ([Col. 34, lines 35-53] Entity model 1120 can determine whether an event vector 1110 should be classified as within the baseline activity characterized by entity model, a result of the output of the class of the trees, corresponding to classify input includes classifications of input as being malicious or normal.).
Ackerman in view of Lee does not appear to teach, but Gopalakrishnan teaches that, wherein a first of the labeled examples is based on an anomaly, wherein a second of the labeled examples is based on detected malicious activities ([0067] "For example, the trained classifier may assign a label of Green to activity not detected to be a network attack, Amber where a low risk of a network attack is predicted, and Red to activity with a high risk of network attack", wherein the label of amber could designate an anomaly, as while something is indicated as wrong by the system, the low risk would not make it enough of a risk for an attack to happen, which corresponds to an anomaly in the applicant.).
Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of Ackerman, Lee, and Gopalakrishnan before them, to include Gopalakrishnan’s ‘wherein a first of the labeled examples is based on an anomaly, wherein a second of the labeled examples is based on detected malicious activities’ in Ackerman’s ‘computer-implemented method, comprising: collecting historical event log data from host devices’. One would have been motivated to make such a combination to increase security, as when taking a look at Fig. 6, which details a process that leads into the results that are displayed in Fig. 7, an important aspect in labeling the state of a vector uses a color-coding system, with green denoting safe, amber as an anomaly, and red as a dangerous vector. With those in mind, they will be used to score and determine, with voting in Fig. 7, whether an attack could happen, as taught in Gopalakrishnan [0091].
Regarding claim 7, Ackerman in view of Lee, and further in view of Gopalakrishnan teaches the computer-implemented method of claim 6. Ackerman also discloses wherein the second model employs a neural-network architecture, wherein the first predetermined period of time ends an hour after the deployment of the trained models ([Col. 15, lines 54-56] "Classifiers may be used, such as neural network classifiers or other classifiers that may be trained by machine learning", in which the model used for training the vectors and determining a potential outcome or at least a training set will use neural network for its training. [Col. 34, lines 48-53] "a k-nearest neighbor classifier may be used to calculate a distance between a point of interest and a training data set, or more generally to determine whether an event vector 1110 should be classified as within the baseline activity characterized by the entity model", where the detection engine that has the vectors and the entity model will use in order to correlate behaviors exhibited by both to determine the state of the entity itself. [Col. 37, lines 11-15] Filtering of a particular endpoints that should not exceed a predetermined time after entity model is deployed, such as an hour, corresponding to a predetermined period of time ending an hour after deployment of trained models of the Applicant.).
Regarding claim 13, Ackerman in view of Lee teaches the computer program product of claim 12 as described above. Ackerman in view of Lee, and further in view of Gopalakrishnan also teaches the limitations of claim 5 above.
Regarding claim 14, Ackerman in view of Lee teaches the computer program product of claim 12 as described above. Ackerman in view of Lee, and further in view of Gopalakrishnan also teaches the limitations of claim 6 above.
Regarding claim 15, Ackerman in view of Lee, and further in view of Gopalakrishnan teaches the computer program product of claim 14 as described above. Ackerman in view of Lee, and further in view of Gopalakrishnan also teaches the limitations of claim 7 above.
Claims 8, and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Ackerman in view of Lee, further in view of Gopalakrishnan, and yet further in view of Pratt et al. (US 10673880 B1), hereinafter Pratt, and Sadaghiani et al. (US 20190034932 A1), hereinafter Sadaghiani.
Regarding claim 8, Ackerman in view of Lee, and further in view of Gopalakrishnan teaches the method of claim 7, as outlined above. Ackerman discloses that wherein, for the host devices, embedding vectors for a recent logged event stream ([Col. 34, lines 35-38] "The detection engine 1122 may compare new events 1106 generated by an entity, as recorded in the event stream 1114, to the entity model 1120 that characterizes a baseline of expected activity", where the event stream corresponds to the embedding vectors of a logged event stream, which in this part of the method, emphasizes newer events as to prioritize potential issues that may not have had attention to be trained on earlier.);
generating the two-dimensional matrix that is based on the determined embedding vectors for the recent logged event stream, wherein deployment of the trained second model includes: causing the two-dimensional matrix to be applied to the trained second model to generate a classification output that represents the likelihood ([Col. 30, lines 40-42] "A detection engine 1122 may be applied to the event stream 1114 in order to detect unusual or malicious activity, e.g. based on the entity models 1120 or any other techniques". [Col. 34, lines 35-53] "The detection engine 1122 may compare new events 1106 generated by an entity, as recorded in the event stream 1114, to the entity model 1120 that characterizes a baseline of expected activity… comparison may use one or more vector distances such as a Euclidean distance, a Mahalanobis distance, a Minkowski distance, or any other suitable measurement of difference within the corresponding vector space. In another aspect, a k-nearest neighbor classifier may be used to calculate a distance between a point of interest and a training data set, or more generally to determine whether an event vector 1110 should be classified as within the baseline activity characterized by the entity model", where the comparison uses vector distances and operations to compare them, of which Minkowski, Euclidean, Mahalanobis are well known in the arts to perform operations in any number of dimensions, including two-dimensional matrices. [Col. 32, lines 38-41] "Each entity model 1120 may, for example, include a multi-dimensional description of events 1106 for an entity based on events 1106 occurring over time for that entity", in which the multi-dimensional description of events could also consist of the various vectors that are included in the entity model for training, the comparison of the data in the detection engine, by the nature of how it would work with Euclidean distances or other vector distances, would also require the event vectors to be considered together as a multi-dimensional matrix, or in the case of how its drawn in Fig. 11, the event stream appears to have two dimensions. [Col. 35, lines 21-24] "where the event stream 1114 deviates from a baseline of expected activity that is described in the entity models 1120 for one or more entities, any number of responses may be initiated by the response facility 1124... ", where responses can vary, but would include further scans or alerts, as stated in Ackerman [Col. 35, lines 25-32].);
wherein the generating the two-dimensional matrix comprises: dividing, based on a predetermined time condition, a time period during which the recently logged event stream occurred into a plurality of time slots ([Col. 32, lines 9-11] Groups of events 1106 fall within a window of time that can be reported as an event vector 1110, with a window of time corresponding to a time slot within a predetermined time period of the Applicant.);
adding the determined embedding vectors that fall within a first of the time slots into a first sum embedding vector, wherein the first sum embedding vector is the only embedding vector of the first time slot ([Col. 32, lines 29-31] Event vectors 1110 are time stamped to record chronology, with a first event vector in the window of time being separate from a second event vector, stated in [Col. 32, lines 4-8]. The first event vector being time stamped associated with the window of time corresponds to the embedding vector in the first sum embedding vector as the only embedding vector of the first time slot of the Applicant.);
adding the determined embedding vectors that fall within a second of the time slots into a second sum embedding vector, wherein the second sum embedding vector is the only embedding vector of the second time slot ([Col. 32, lines 4-8] Second event vector 1110 can be created and reported along with other temporally adjacent events 1106, while separate from a first event vector. The second event vector being time stamped associated with the window of time corresponds to the embedding vector in the second sum embedding vector as the only embedding vector of the second time slot of the Applicant.);
and stacking the sum embedding vectors into the two-dimensional matrix, wherein the classification output is a numerical score of a predetermined range of potential numerical scores ([Col. 32, lines 38-41] Fig. 11, each entity model 1120 includes a multi-dimensional description of events 1106 based on event stream 1114, corresponding to a two-dimensional matrix output by the trained first model being applied to the trained second model).
Ackerman discloses the method of ‘wherein deployment of the trained first model includes: causing the trained first model to determine, for the host devices, embedding vectors for a recent logged event stream’. Ackerman does not explicitly teach the method of ‘wherein the classification output is a numerical score of a predetermined range of potential numerical scores’.
However, Pratt teaches that “wherein the classification output is a numerical score of a predetermined range of potential numerical scores, wherein bounds of the predetermined range are a score of one and a score of one hundred, wherein the score of one represents a relatively highest likelihood that the malicious cybersecurity event will occur within the first predetermined period of time” (Within either Fig. 28 and 29, which have two different outcomes for when a threat indicator goes off or not, we see that with Fig. 28 in particular, we have both an anomaly rule and an anomaly model at play, and when an anomaly is detected by a rule, it trains the model to lookout for anomalies similar to the anomaly 1 found earlier. When another anomaly is detected, it could raise a "threat indicator" alert as shown in the figure. Furthermore, as Fig. 30 relates to the figures and the elements found in Figs. 28-29 and in Fig. 18, the process can extend to other anomalies and even other models, but, according to [Col. 56, lines 40-46], when a threat indicator is found in an IT environment based on the events collected from other users or devices, "Process 3200 continues at step 3208 with generating a pattern matching score based on a result of the comparing. In some embodiments, the pattern matching score is a value in a set range. For example, the resulting pattern matching score may be a value between 0 and 10 with 0 being the least likely to be a threat and 10 being the most likely to be a threat".).
Ackerman in view of Pratt does not appear to teach, but Sadaghiani teaches the limitation of “wherein bounds of the predetermined range are a score of one and a score of one hundred, wherein the score of one represents a relatively highest likelihood” ([0046] Digital threat score range is between zero and 100, where a higher global digital threat score value indicates a higher likelihood that the digital event involves digital fraud and/or abuse. Alternatively, the higher threat score value indicates a lower likelihood of digital fraud and/or abuse, which makes zero the highest likelihood of abuse and/or fraud. This corresponds to one representing a highest likelihood, as the value of zero is closest to one.);
Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of Ackerman, Lee, Gopalakrishnan, Pratt, and Sadaghiani before them, to include Pratt’s ‘wherein the classification output is a numerical score of a predetermined range of potential numerical scores’ in Ackerman’s system performing ‘computer-implemented method, comprising: collecting historical event log data from host devices’, ‘wherein deployment of the trained first model includes: causing the trained first model to determine, for the host devices, embedding vectors for a recent logged event stream’, ‘generating a two-dimensional matrix that is based on the determined embedding vectors for the recent logged event stream, wherein deployment of the trained second model includes: causing the two-dimensional matrix to be applied to the trained second model to generate a classification output that represents the likelihood’. One would have been motivated to make such a combination to enhances security by showing a score to the user, and within "step 3208 with generating a pattern matching score based on a result of the comparing. In some embodiments, the pattern matching score is a value in a set range", with the conclusion of the process 3200 at "step 3210 with identifying a security threat if the pattern matching score satisfies a specified criterion", and in this example in the prior art, a score of 6 or greater indicates a threat being present in the entity, as stated in [Col. 56, lines 41-51]. This gives the user, the network, or the system an evaluation and lets the user know that an anomalous or malicious process is either running or stored, and will further be required to remediate the system if necessary.
Regarding claim 16, Ackerman in view of Lee, further in view of Gopalakrishnan teaches the computer program product of claim 15 as described above. Ackerman in view of Lee, further in view of Gopalakrishnan, and yet further in view of Pratt and Sadaghiani also teaches the limitations of claim 8 above.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Kumar et al. (US 20130298244 A1, "SYSTEMS AND METHODS FOR THREAT IDENTIFICATION AND REMEDIATION ")
Wu et al. (US 20220012499 A1, "SPATIAL-TEMPORAL GRAPH-TO-SEQUENCE LEARNING BASED GROUNDED VIDEO DESCRIPTIONS")
Yu et al. (US 20240290081 A1, "TRAINING AND USING A MODEL FOR CONTENT MODERATION OF MULTIMODAL MEDIA")
Ansel et al. (US 9483740 B1, "Automated Data Classification")
Wang et al. (NPL, "A complex history browsing text categorization method with improved BERT embedding layer", 2025)
Any inquiry concerning this communication or earlier communications from the examiner should be directed to TOMMY MARTINEZ whose telephone number is (703)756-5651. The examiner can normally be reached Monday thru Friday 8AM-5PM ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jorge L. Ortiz-Criado can be reached at (571) 272-7624 on Monday thru Friday, 7AM-7PM ET. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/T.M./ Examiner, Art Unit 2496 /ABU S SHOLEMAN/Primary Examiner, Art Unit 2496