DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s arguments, see pg. 9, filed 03/17/2026, with respect to claim 23 have been fully considered and are persuasive. The 35 USC 112 rejection of 01/12/2026 has been withdrawn.
Applicant's arguments filed 03/17/2026 have been fully considered but they are not persuasive.
Applicant’s arguments with respect to claims 1-20 have been considered. However, Strope has been remapped to teach the amended features.
Specifically:
wherein training repeatedly continues until a determination that an average difference between one or more inputs and one or more outputs of the anomaly detection model are less than a predetermined value; and (Strope, paragraph 0007, “Further, the response features of a training instance are applied as input to the response neural network model and a response vector generated over the response neural network model based on that input. A response score can then be determined based on comparison of the input vector and the response vector [one or more inputs and the one or more outputs]. For example, the response score can be based on the dot product of the input vector and the response vector. For instance, the dot product can result in a value from O to 1, with "1" indicating the highest likelihood a corresponding response is an appropriate response to a corresponding electronic communication and "O" indicating the lowest likelihood. Both the input neural network model and the response neural network model can then be updated based on comparison of: the response score ( and optionally additional response scores in batch techniques described herein); and a response score indicated by the training instance (e.g., a "1" or other "positive" response score for a positive training instance, a "O" or other "negative" response score for a negative training instance[less than the predetermined value, Examiner would like to point out that the predetermined value is the set being less than a predetermined value in paragraph 0068, the negative response is being interpreted as the less than value.]). For example, an error can be determined based on a difference between the response score and the indicated response score b[determination that the average difference], and the error backpropagated through both neural networks of the model.” And paragraph 0101, “The system may then identify a new batch of training instances, and restart method 500 for the new batch. Such training may continue until one or more criteria are satisfied [training repeatedly continues].”)
Applicant argues that Strope discusses multiple models which is not the case for the application, that there is only one model involved. With further search and consideration, examiner would like to point out that in the specification of the application, both paragraphs 0048 and 0082 both state that one or more models can be used in the invention. Paragraph 0101 of Strope teaches training continuing until a criteria is satisfied.
Jin has been mapped to the newly added claim 24, teaching the evaluation dataset.
Specifically:
wherein: the evaluation data set comprises a set of benign interactions and a set of fraudulent interactions; and the evaluation continues until a determination that the one or more outputs of the anomaly detection model during the evaluation accurately identify the set of benign interactions and the set of fraudulent interactions at least at a predetermined rate. (Jin, pg. 1012, Col. 1, paragraph 2, “Next we define two anomaly scores, the reconstruction anomaly score frec score(τ ) and the regression anomaly score freg score(τ ), as the metrics for evaluating the “degree of anomaly” of an observation Sτ . Note that there is more than one snippet that encompasses Sτ because we used a sliding window approach to generate the snippets.”)
Therefore, the 35 USC 103 rejection is maintained.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or non-obviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-2, 5-6, 11-12, 15-16 and 20-22 are rejected under 35 U.S.C. 103 as being unpatentable over Jadidi et al (“Flow-based Anomaly Detection Using Neural Network Optimized with GSA Algorithm” (2013), “Jadidi”), and in view of Strope et al (US Published Patent Application No. 20180240014, “Strope”) and in further view of Krishnamurthy et al (US Published Patent Application No. 20180068371, "Krishnamurthy") and in even further view of Shamiss et al (US Published Patent Application No. 20190370837, "Shamiss") and Siadati et al (Detecting Structurally Anomalous Logins Within Enterprise Networks, "Siadati").
In regard to claim 1 and analogous claims 11 and 20, Jadidi teaches receive interaction data representative of one or more interactions; (Jadidi, II. Dataset, pg. 77, “Sperotto’s data set [15] is divided into three categories: Malicious traffic, Side-effect traffic (this part is not by itself malicious, for example, ICMP, Auth/Ident and IRC traffic), Unknown traffic and uncorrelated alerts (in this part, the malicious or benign nature of the traffic cannot be identified).)
classify the one or more interactions as one of an anomalous interaction or a benign interaction using an anomaly detection model, (Jadidi, C. Proposed GSA-based flow anomaly detection system, pg. 79, “Two output nodes perform the classification of the flow-based traffic into malicious and benign subsets.”)
trained using a training data set comprising historical interaction data, (Jadidi, pg. 79, Col. 2, “GFADS is trained with preprocessed training data set and is evaluated in classifying the testing flows into malicious or benign subsets [trained using a training data set comprising historical interaction data].”
However, Jadidi does not explicitly teach identifies a similarity between the interaction data and known benign interactions, is trained using a training data set comprising historical interaction data including a set of event variables including at least one of location data, item information, and event specific data, wherein training includes a determination that one or more outputs of the anomaly detection model are less than a predetermined value, and wherein the anomaly detection model is evaluated using an evaluation data set, and wherein the evaluation data set and the training data set are different; and
generate an indication of authorization based on the classification of the interaction, wherein the indication of authorization: authorizes the interaction when the anomaly detection model classifies the interaction as benign, and denies the interaction when the anomaly detection model classifies the interaction as anomalous.
Strope teaches identifies a similarity between the interaction data and known benign interactions, (Strope, paragraph 0077, “In some of those and/or other implementations, the response indexing system 140 determines multiple clusters of response vectors, seeking to cluster similar [a similarity between] vectors together. The response indexing system 140 can build a tree or other structure to enable initial searching (e.g., by vector comparison scoring engine 124) for relevant response vectors by cluster [the interaction data and known benign interactions, the response is being interpreted as the interaction data and the known benign interaction data are being interpreted as the relevant response vectors in the search]. Such a tree or other structure can enable searching each cluster first to identify the most relevant cluster(s) as opposed to the more computationally inefficient searching of each response vector individually”)
wherein training repeatedly continues until a determination that an average difference between one or more inputs and one or more outputs of the anomaly detection model are less than a predetermined value; and (Strope, paragraph 0007, “Further, the response features of a training instance are applied as input to the response neural network model and a response vector generated over the response neural network model based on that input. A response score can then be determined based on comparison of the input vector and the response vector [one or more inputs and the one or more outputs]. For example, the response score can be based on the dot product of the input vector and the response vector. For instance, the dot product can result in a value from O to 1, with "1" indicating the highest likelihood a corresponding response is an appropriate response to a corresponding electronic communication and "O" indicating the lowest likelihood. Both the input neural network model and the response neural network model can then be updated based on comparison of: the response score ( and optionally additional response scores in batch techniques described herein); and a response score indicated by the training instance (e.g., a "1" or other "positive" response score for a positive training instance, a "O" or other "negative" response score for a negative training instance [less than the predetermined value, Examiner would like to point out that the predetermined value is the set being less than a predetermined value in paragraph 0068, the negative response is being interpreted as the less than value.]). For example, an error can be determined based on a difference between the response score and the indicated response score b[determination that the average difference], and the error backpropagated through both neural networks of the model.” And paragraph 0101, “The system may then identify a new batch of training instances, and restart method 500 for the new batch. Such training may continue until one or more criteria are satisfied [training repeatedly continues].”)
Jadidi and Strope are related to the same field of endeavor (i.e. neural network training). In view of the teachings of Strope, it would have been obvious for a person with ordinary skill in the art to apply the teachings of Strope to Jadidi before the effective filing date of the claimed invention in order to more efficiently identify similar data. (Strope, paragraph 0042, “It is noted that in some implementations, by comparing the input vector to response vectors associated with each of the clusters, a tree-based and/or other approach may be utilized to enable efficient identification of cluster(s) that are most relevant to the input vector, without necessitating comparison of the input vector to a response vector of each and every one of the clusters.”)
However, Jadidi and Strope do not explicitly teach including a set of event variables including at least one of location data, item information, and event specific data,
evaluated using an evaluation data set, and wherein the evaluation data set and the training data set are different; and
generate an indication of authorization based on the classification of the interaction, wherein the indication of authorization: authorizes the interaction when the anomaly detection model classifies the interaction as benign, and denies the interaction when the anomaly detection model classifies the interaction as anomalous.
Krishnamurthy teaches evaluated using an evaluation data set, and wherein the evaluation data set and the training data set are different; and (Krishnamurthy, paragraph 0037, “To evaluate different item recommendation models, a sample data set [evaluation data set] regarding movie views by users is used. The dataset comprises 10 million ratings and 100 thousand tag applications applied to 10 thousand movies by 72 thousand users. In order to evaluate performance, a recall@20 measure is used. For each user, the last movie that the user has seen is removed. The set ( of size K) of previously seen movies is then used [the training data set] to predict 20 movies that the model believes the user is most likely to watch. If the removed movie is present in the list of 20 recommended movies the prediction is counted as a success, else a failure. The metric is then simply the total percentage of successes over all users.”)
Jadidi, Strope and Krishnamurthy are related to the same field of endeavor (i.e. anomaly detection). In view of the teachings of Krishnamurthy it would have been obvious for a person with ordinary skill in the art to apply the teachings of Krishnamurthy to Jadidi and Strope before the effective filing date of the claimed invention in order to help find similarity between items (Krishnamurthy, paragraph 0016, “The item vector representations are then used by the at least one computing device to build an item similarity matrix that is able to define relationships between items.”)
However, Jadidi, Strope and Krishnamurthy do not explicitly teach including a set of event variables including at least one of location data, item information, and event specific data,
generate an indication of authorization based on the classification of the interaction, wherein the indication of authorization: authorizes the interaction when the anomaly detection model classifies the interaction as benign, and denies the interaction when the anomaly detection model classifies the interaction as anomalous.
Shamiss teaches including a set of event variables including at least one of location data, item information, and event specific data, (Shamiss, paragraph 0005, “A typical reverse logistics process for an article of consumer merchandise involves several steps. Generally executed in sequential order, the steps can comprise receiving an article in a return process, identifying the corresponding product based on the universal product code (UPC) of a returned article [including at least one of location data, item information, and event specific data], assigning a unique identification code (UID) to an identified article, grading the condition of a uniquely-identified article, evaluating the salability of a graded article, and selecting a disposition pathway for an evaluated article (see FIG. 1). A common disposition pathway is a listing in a secondary E-commerce marketplace.”)
Jadidi, Strope, Krishnamurthy and Shamiss are related to the same field of endeavor (i.e. anomaly detection). In view of the teachings of Shamiss it would have been obvious for a person with ordinary skill in the art to apply the teachings of Shamiss to Jadidi, Strope and Krishnamurthy before the effective filing date of the claimed invention in order to evaluate items autonomously. (Shamiss, paragraph 0010, “The present invention comprises an autonomous means of evaluating items of returned or liquidated consumer merchandise.”)
However, Jadidi, Strope, Krishnamurthy and Shamiss do not explicitly teach generate an indication of authorization based on the classification of the interaction, wherein the indication of authorization: authorizes the interaction when the anomaly detection model classifies the interaction as benign, and denies the interaction when the anomaly detection model classifies the interaction as anomalous.
However, Siadati, teaches generate an indication of authorization based on the classification of the interaction, wherein the indication of authorization: authorizes the interaction when the anomaly detection model classifies the interaction as benign, and denies the interaction when the anomaly detection model classifies the interaction as anomalous. (Siadati, pg. 1275, Col. 1, paragraph 3, “By computing the similarity of the new login with the class of normal logins, this component classifies new logins into one of two classes; benign or malicious. [based on the classification of the interaction]” and pg. 1282, Col. 2, paragraph 1, “Several mechanisms such as Access Control Lists in Linux and Active Directory in Windows systems [19] allow the network admins to enforce rules of access. Using the notion of groups of users and objects, network admins can allow [authorizes the interaction] or deny access of a group of users to a collection of resources [denies the interaction]. This mechanism is useful for stopping an employee from accessing data or resources he/she should have never access.”)
Jadidi, Strope, Krishnamurthy, Shamiss and Siadati are related to the same field of endeavor (i.e. anomaly detection). In view of the teachings of Siadati it would have been obvious for a person with ordinary skill in the art to apply the teachings of Siadati to Jadidi, Strope, Krishnamurthy and Shamiss before the effective filing date of the claimed invention in order to efficiently extract data needed to classify information. (Siadati, pg. 1274, Col. 1, Bullet 2, “We also provide an algorithm to automatically and efficiently extract login patterns from a large login data set.”)
In regard to claim 2 and analogous claim 12, Jadidi, Strope, Krishnamurthy, Shamiss and Siadati teach the system of claim 1.
Krishnamurthy further teaches wherein the interaction data includes at least one alphanumeric variable of the interaction, and wherein the anomaly detection model converts the at least one alphanumeric variable into a vector embedding. (Krishnamurthy, paragraph 0016, “A word representative of an item may be a single word for the item name, a concatenated item name, an item number for the item, or any other single string of letters without a space. For the example above, the sentences that are input into the word embedding model by the at least one computing device may look like "a b c" from the first user and "b c d" from the second user [at least one alphanumeric variable of the interaction]. So, this example includes two sentences corresponding to two user sessions [the interaction data]. Using the sentences, the at least one computing device employs the word embedding model to learn item vector representations of the items.” And paragraph 0030, “I order to build the item similarity matrix 118, the item recommendation module 104 inputs the session data 108 into a word embedding model [anomaly detection model] 212. The word embedding model 212 is maintained as data to produce item vector representations [configured to convert the at least one alphanumeric variable into a vector embedding.] for the items 116 that are in tum used by the item recommendation module 104 to create the item similarity matrix 118. This may be performed through comparisons of dot products of the item vector representations of the items, performing arithmetic on the item vector representations, and so on.”)
Jadidi and Krishnamurthy are combinable for the same rationale as set forth above with respect to claim 1.
In regard to claim 5 and analogous claim 15, Jadidi, Strope, Krishnamurthy, Shamiss and Siadati teach the system of claim 2.
Krishnamurthy further teaches wherein the at least one alphanumeric variable includes item variables, and wherein the anomaly detection model converts the item variables into an item embedding using an item vector conversion model. (Krishnamurthy, paragraph 0030, “I order to build the item similarity matrix 118, the item recommendation module 104 inputs the session data [includes item variables] 108 into a word embedding model 212. The word embedding model 212 is maintained as data to produce item vector representations for the items 116 that are in tum used by the item recommendation module 104 to create the item similarity matrix 118. This may be performed through comparisons of dot products of the item vector representations of the items, performing arithmetic on the item vector representations, and so on.” And paragraph 0034, “As discussed above, embedding of words is used by the computing device to produce item vector representations of the associated items. The item vectors are usable by an item recommendation system of a computing device to create an item similarity matrix through comparison of the item vectors, arithmetic that is based on the item vectors, and so on. [configured to convert the item variables into an item embedding using an item vector conversion model.]”)
Jadidi and Krishnamurthy are combinable for the same rationale as set forth above with respect to claim 1.
In regard to claim 6 and analogous claim 16, Jadidi, Strope, Krishnamurthy, Shamiss and Siadati teach the system of claim 5.
Krishnamurthy further teaches wherein the item vector conversion model includes at least one word2vec layer generates a vector representation of the item variables. (Krishnamurthy, paragraph 0035, “Word2Vec [at least one word2vec layer] consists of two distinct models (CBOW and skip-gram), each of which defines two training methods [the item vector conversion model] (with/without negative sampling) and other variations, such as hierarchical softmax. Both CBOW and skip-gram are shallow 2-layer neural network models. The CBOW model is used for item recommendations since it more intuitively captures the problem domain.” And paragraph 0036, “ In a typical CBOW embodiment the neural network is trained to predict the central word given the words that occur in a context window around it. The word representations are learned in such a way that a sequence of the words ( or items) in the embedment may have an effect on performance [to generate a vector representation of the item variables.]. However, for item recommendations it is beneficial to learn word embeddings in an order agnostic manner. In order to make the model less sensitive to these orderings, the Word2Vec model includes a number of random permutations of the items in a user session to the training corpus.”)
Jadidi and Krishnamurthy are combinable for the same rationale as set forth above with respect to claim 1.
In regard to claim 21, Jadidi, Strope, Krishnamurthy, Shamiss and Siadati teach the system of claim 1.
Shamiss further teaches wherein the historical interaction data includes a set of transaction data for a return interaction. (Shamiss, paragraph 0005, “A typical reverse logistics process for an article of consumer merchandise involves several steps. Generally executed in sequential order, the steps can comprise receiving an article in a return process, identifying the corresponding product based on the universal product code (UPC) of a returned article, assigning a unique identification code (UID) to an identified article, grading the condition of a uniquely-identified article, evaluating the salability of a graded article, and selecting a disposition pathway for an evaluated article [a set of transaction data for a return interaction.] (see FIG. 1). A common disposition pathway is a listing in a secondary E-commerce marketplace.”)
Jadidi and Shamiss are combinable for the same rationale as set forth above with respect to claim 1.
In regard to claim 22, Jadidi, Strope, Krishnamurthy, Shamiss and Siadati teach the system of claim 21.
Shamiss further teaches wherein the historical interaction data includes at least one of temporal data indicating a time of the return interaction, set location data indicating a place of interaction, item information, or return information. (Shamiss, paragraph 0005, “A typical reverse logistics process for an article of consumer merchandise involves several steps. Generally executed in sequential order, the steps can comprise receiving an article in a return process, identifying the corresponding product based on the universal product code (UPC) of a returned article, assigning a unique identification code (UID) to an identified article, grading the condition of a uniquely-identified article, evaluating the salability of a graded article, and selecting a disposition pathway for an evaluated article [return information] (see FIG. 1). A common disposition pathway is a listing in a secondary E-commerce marketplace.”)
Jadidi and Shamiss are combinable for the same rationale as set forth above with respect to claim 1.
In regard to claim 23, Jadidi, Strope, Krishnamurthy, Shamiss and Siadati teach the system of claim 1.
Strope further teaches wherein the training further continues until the anomaly detection model is generated in accordance with the determination that the average difference between the one or more inputs and the one or more outputs of the anomaly detection model are less than the predetermined value. (Strope, paragraph 0007, “Further, the response features of a training instance are applied as input to the response neural network model and a response vector generated over the response neural network model based on that input. A response score can then be determined based on comparison of the input vector and the response vector [one or more inputs and the one or more outputs]. For example, the response score can be based on the dot product of the input vector and the response vector. For instance, the dot product can result in a value from O to 1, with "1" indicating the highest likelihood a corresponding response is an appropriate response to a corresponding electronic communication and "O" indicating the lowest likelihood. Both the input neural network model and the response neural network model can then be updated based on comparison of: the response score ( and optionally additional response scores in batch techniques described herein); and a response score indicated by the training instance (e.g., a "1" or other "positive" response score for a positive training instance, a "O" or other "negative" response score for a negative training instance). For example, an error [less than the predetermined value.] can be determined based on a difference between the response score and the indicated response score b[determination that the average difference], and the error backpropagated through both neural networks of the model [training further continues].”)
Claims 3 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Jadidi, in view of Strope, in further view of Krishnamurthy, Shamiss and Siadati in further view of Emigh et al (US Patent No. 9264151, "Emigh").
In regard to claim 3 and analogous claim 13, Jadidi, Strope, Krishnamurthy, Shamiss and Siadati teach the system of claim 2.
However, Jadidi, Strope, Krishnamurthy, Shamiss and Siadati fail to teach wherein the at least one alphanumeric variable includes a store identifier, and wherein the anomaly detection model converts the store identifier into a store embedding using a store vector conversion model.
Emigh teaches wherein the at least one alphanumeric variable includes a store identifier, and wherein the anomaly detection model converts the store identifier into a store embedding using a store vector conversion model. (Emigh, Col. 8, lines 9-11, “In some embodiments, sonic signaling may encode a store identifier (ID) [a store identifier], for example a numeric identifier or an encoding of a string. [one alphanumeric variable]” And Col 38, lines 5-20, “In this example, frequencies 2-19 are used to populate a bit vector corresponding to the store identifier. For example, frequency 2 corresponds to the least significant bit of the store identifier while frequency 19 corresponds to the most significant bit of the store identifier. In such cases, if frequency 2 is on/off, the least significant bit of the store identifier is set to 1/0; if frequency 3 is on/off, the next bit of the store identifier is set to 1/0; and so on uutil the entire store identifier bit vector is populated [configured to convert the store identifier into a store embedding using a store vector conversion model]. Similar processes may be employed for frequencies 20-27 to populate the error detection and/or correction bit vector as well as for frequencies 28-32 to populate the department code bit vector. Although one example is described, any appropriate scheme to map frequencies to bit vectors may be employed in other embodiments.”)
Jadidi and Emigh are related to the same field of endeavor (i.e. detection models). In view of the teachings of Emigh it would have been obvious for a person with ordinary skill in the art to apply the teachings of Emigh to Jadidi before the effective filing date of the claimed invention in order to help make sure the identifiers are valid. (Emigh, Col. 38, lines 35-39, “In this case, the CRC for the store identifier bits is computed and compared with the decoded error detection/correction bits. If they match, the original store identifier bits are declared to be valid.”)
Claims 4 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Jadidi, in view of Strope, Krishnamurthy, Shamiss, Siadati and Emigh and in even further view of Cao et al ("Learning Neural Representations for Network Anomaly Detection", (2019), "Cao").
In regard to claim 4 and analogous claim 14, Jadidi, Strope, Krishnamurthy, Shamiss, Siadati and Emigh teach the system of claim 3.
However, Jadidi, Strope, Krishnamurthy, Shamiss, Siadati and Emigh fail to teach wherein the store vector conversion model includes at least one hidden layer generates a vector representation of the store identifier.
Cao teaches wherein the store vector conversion model includes at least one hidden layer generates a vector representation of the store identifier. (Cao, pg. 3075, Col. 2, “The middle hidden layer, sometimes called the bottleneck layer, like a nonlinear principal component analysis (PCA), compresses the redundancies while preserving and differentiating nonredundant information in the input [17].”
Jadidi and Cao are related to the same field of endeavor (i.e. anomaly detection). In view of the teachings of Cao it would have been obvious for a person with ordinary skill in the art to apply the teachings of Cao to Jadidi before the effective filing date of the claimed invention in order to improve the performance of the anomaly detection (Cao, pg. 3075, Col. 2, “Alternatively, the middle hidden layer of a trained AE can be used as a new feature representation (called a latent representation) for improving the performance of density-based anomaly detection [13] or anomaly detection based on self-organizing maps [23].”)
Claims 7 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Jadidi, in view of Strope, in further view of Krishnamurthy, Shamiss and Siadati and in even further view of Malhorta et al (US Published Patent Application No. 20160299938, “Malhorta”).
In regard to claim 7 and analogous claim 17, Jadidi, Strope, Krishnamurthy, Shamiss and Siadati teach the system of claim 1.
However, Jadidi, Strope, Krishnamurthy, Shamiss and Siadati fail to specifically teach wherein the anomaly detection model comprises a plurality of long short term memory (LSTM) cells.
Malhorta teaches wherein the anomaly detection model comprises a plurality of long short term memory (LSTM) cells. (Malhorta, paragraph 0009, “The anomaly is detected based on a prediction model by using a long short term memory (LSTM) neural network. [plurality of long short term memory (LSTM) cells.]” And paragraph 0060, “The anomaly detection system [anomaly detection model] 102 implements stacked LSTM networks that are able to learn higher level temporal patterns without prior knowledge of the pattern duration, and so the stacked LSTM networks may be a viable technique to model normal time-series behavior, which can then be used to detect anomalies.”)
Jadidi and Malhorta are related to the same field of endeavor (i.e. anomaly detection). In view of the teachings of Malhorta it would have been obvious for a person with ordinary skill in the art to apply the teachings of Malhorta to Jadidi before the effective filing date of the claimed invention in order to learn in an efficient manner and more accurately detect anomalous behavior (Malhorta, paragraph 0060, “In other words, dependencies among different dimensions of a multivariate timeseries data can be learnt by LSTM network which enables to learn normal behavior in an efficient manner and hence detect anomalous behavior more accurately.”)
Claims 8-9, 18 and 24 are rejected under 35 U.S.C. 103 as being unpatentable over Jadidi, in view of Strope, in further view of Krishnamurthy, Shamiss and Siadati and in even further view of Jin et al (“An Encoder-Decoder Based Approach for Anomaly Detection with Application in Additive Manufacturing” (2019), “Jin”).
In regard to claim 8, Jadidi, Strope, Krishnamurthy, Shamiss and Siadati teach the system of claim 1.
However, Jadidi, Strope, Krishnamurthy, Shamiss and Siadati fail to teach wherein the anomaly detection model comprises an encoder and decoder framework to generate an output tensor.
Jin teaches wherein the anomaly detection model comprises an encoder and decoder framework to generate an output tensor. (Jin, B. Encoder-Decoder architecture, “The encoder-decoder architecture has proven to be a useful approach for learning (deep) representations, and is widely used in various application domains of deep learning, including machine translation [3], and image denoising [26]. An encoder-decoder model generally consists of three parts: the encoder, the latent space representation, and the decoder. The purpose of the encoder network Enc is to transform the input data into a latent space representation z that is often a vector; the decoder network Dec then produces the output [configured to generate an output tensor] by decoding z. During training, the encoder and the decoder are trained together to minimize the empirical risk.” And C. Unsupervised anomaly detection with deep learning, “The encoder-decoder schemes for anomaly detection [anomaly detection model comprises an encoder and decoder framework] that appeared in literature in general fall into three categories, which differ in their prediction outputs: 1) autoencoder models [16], 2) prediction models, and 3) composite models [25] that performs both reconstruction and regression.”)
Jadidi and Jin are related to the same field of endeavor (i.e. anomaly detection). In view of the teachings of Jin it would have been obvious for a person with ordinary skill in the art to apply the teachings of Jin to Jadidi before the effective filing date of the claimed invention in order to minimize risk (JIn, B. Encoder-Decoder architecture, “During training, the encoder and the decoder are trained together to minimize the empirical risk.”)
In regard to claim 9, Jadidi, Strope, Krishnamurthy, Shamiss, Siadati and Jin teach the system of claim 8.
Strope further teaches wherein anomaly detection model classifies the one or more interactions based on a similarity between an input tensor generated from the interaction data and the output tensor. (Strope, paragraph 0076, “The response indexing system 140 generates the response index with response vectors 174 through processing of a large quantity (e.g., all) of the responses of responses database 172. The generated index 174 includes corresponding pre-determined response vectors and/or other values stored in association with each of the responses. For example, index 174 can have a stored association of "Response A" to a corresponding response vector, a stored association of "Response B" to a corresponding response vector, etc. The index 174 can have similar stored associations to each of a plurality of (thousands, hundreds of thousands, etc.) additional responses [classifies the one or more interactions based on a similarity].” and paragraph 0077, “In some of those and/or other implementations, the response indexing system 140 determines multiple clusters of response vectors, seeking to cluster similar vectors together [an input tensor generated from the interaction data and the output tensor]. The response indexing system 140 can build a tree or other structure to enable initial searching (e.g., by vector comparison scoring engine 124) for relevant response vectors by cluster. Such a tree or other structure can enable searching each cluster first to identify the most relevant cluster(s) as opposed to the more computationally inefficient searching of each response vector individually.”)
Jadidi and Jin are combinable for the same rationale as set forth above with respect to claim 8.
In regard to claim 18, Jadidi, Strope, Krishnamurthy, Shamiss, Siadati and Jin teach the system of claim 11.
Strope further teaches wherein anomaly detection model classifies the one or more interactions based on a similarity between an input tensor generated from the interaction data and the output tensor. (Strope, paragraph 0076, “The response indexing system 140 generates the response index with response vectors 174 through processing of a large quantity (e.g., all) of the responses of responses database 172. The generated index 174 includes corresponding pre-determined response vectors and/or other values stored in association with each of the responses. For example, index 174 can have a stored association of "Response A" to a corresponding response vector, a stored association of "Response B" to a corresponding response vector, etc. The index 174 can have similar stored associations to each of a plurality of (thousands, hundreds of thousands, etc.) additional responses [classifies the one or more interactions based on a similarity].” and paragraph 0077, “In some of those and/or other implementations, the response indexing system 140 determines multiple clusters of response vectors [an input tensor generated from the interaction data and the output tensor], seeking to cluster similar vectors together. The response indexing system 140 can build a tree or other structure to enable initial searching (e.g., by vector comparison scoring engine 124) for relevant response vectors by cluster. Such a tree or other structure can enable searching each cluster first to identify the most relevant cluster(s) as opposed to the more computationally inefficient searching of each response vector individually.”
Jin further teaches and wherein the output tensor is generated by the encoder-decoder framework. (Jin, B. Encoder-Decoder architecture, “The encoder-decoder architecture has proven to be a useful approach for learning (deep) representations, and is widely used in various application domains of deep learning, including machine translation [3], and image denoising [26]. An encoder-decoder model generally consists of three parts: the encoder, the latent space representation, and the decoder. The purpose of the encoder network Enc is to transform the input data into a latent space representation z that is often a vector; the decoder network Dec then produces the output [configured to generate an output tensor] by decoding z. During training, the encoder and the decoder are trained together to minimize the empirical risk.” And C. Unsupervised anomaly detection with deep learning, “The encoder-decoder schemes for anomaly detection [anomaly detection model comprises an encoder and decoder framework] that appeared in literature in general fall into three categories, which differ in their prediction outputs: 1) autoencoder models [16], 2) prediction models, and 3) composite models [25] that performs both reconstruction and regression.”)
Jadidi and Jin are combinable for the same rationale as set forth above with respect to claim 8.
In regard to claim 24, Jadidi, Strope, Krishnamurthy, Shamiss and Siadati teach the system of claim 1.
However, Jadidi, Strope, Krishnamurthy, Shamiss and Siadati fail to teach wherein: the evaluation data set comprises a set of benign interactions and a set of fraudulent interactions; and the evaluation continues until a determination that the one or more outputs of the anomaly detection model during the evaluation accurately identify the set of benign interactions and the set of fraudulent interactions at least at a predetermined rate.
Jin teaches wherein: the evaluation data set comprises a set of benign interactions and a set of fraudulent interactions; and the evaluation continues until a determination that the one or more outputs of the anomaly detection model during the evaluation accurately identify the set of benign interactions and the set of fraudulent interactions at least at a predetermined rate. (Jin, pg. 1012, Col. 1, paragraph 2, “Next we define two anomaly scores, the reconstruction anomaly score frec score(τ ) and the regression anomaly score freg score(τ ), as the metrics for evaluating the “degree of anomaly” of an observation Sτ . Note that there is more than one snippet that encompasses Sτ because we used a sliding window approach to generate the snippets. To get a single anomaly score taking into account the prediction errors from all relevant snippets, we define the anomaly score as the average prediction errors from all these snippets [at least at a predetermined rate].”)
Jadidi and Jin are combinable for the same rationale as set forth above with respect to claim 8.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SKYLAR K VANWORMER whose telephone number is (703)756-1571. The examiner can normally be reached M-F 6:00am to 3:00 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Usmaan Saeed can be reached on (571) 272-4046. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/S.K.V./Examiner, Art Unit 2146 /USMAAN SAEED/Supervisory Patent Examiner, Art Unit 2146