DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of Claims
Claims 1, 7, 11, 15, 19 – 20 were currently amended. Claims 1-20 are pending and examined herein.
Claims 3, 13 are rejected under 35 U.S.C. 112(d).
Claims 1-20 are rejected under 35 U.S.C. 103.
Response to Arguments
Applicant’s arguments, filed June 4th, 2026 regarding the 35 U.S.C. 103 rejection of claims 1 – 20 have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of Stojanovic et al. (U.S. Pub. 2019/0138538) and Mior et al. (NPL:” Column2Vec: Structural Understanding via Distributed Representations of Database Schemas”), as set forth below.
The new rejection does not rely on Wong’s aggregate transaction metrics as the claimed structural data. Instead, Stojanovic is relied for a profiling pipeline that generates metadata describing columns and data sources and source location information. Mior teaches generating embeddings from textual table and column names. Bremer teaches encoding textual information into vectors and separately discloses an autoencoder as a data representation model. The rejection explains why one of ordinary skill would have used Bremer’s autoencoder representation model to encode the textual structural metadata. Wong is still used for combining the feature vector portions and supplying the combined vector to a trained classifier. In the proposed combination, the classifier processes both the data value vectors and the structural metadata vectors, as required by the amended claims.
Claim Objections
Claims 11, 19, 20 are objected to because of the following informalities:
Claim 11 recites “memory storing instructions that … cause the computing device to: processing …” which is grammatically incorrect. Replacing “processing” to “process” like other limitations in claim 11 would be more accurate.
Claim 19 recites “wherein the first machine learning model comprises a non-neural network machine learning mode;” which appears to have a typo. “network machine learning mode” should be “network machine learning model”.
Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(d):
(d) REFERENCE IN DEPENDENT FORMS.—Subject to subsection (e), a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers.
The following is a quotation of pre-AIA 35 U.S.C. 112, fourth paragraph:
Subject to the following paragraph [i.e., the fifth paragraph of pre-AIA 35 U.S.C. 112], a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers.
Claims 3, 13 are rejected under 35 U.S.C. 112(d) or pre-AIA 35 U.S.C. 112, 4th paragraph, as being of improper dependent form for failing to further limit the subject matter of the claim upon which it depends, or for failing to include all the limitations of the claim upon which it depends. Independent claims 1 and 11 already require "an autoencoder of the embedding machine learning model." Claims 3 and 13 recite that the embedding machine learning model comprises "the autoencoder, a VAE, a Bert Model or a transformer model." Since "the autoencoder" is one of the recited alternatives and is already required by the respective independent claims, claims 3 and 13 do not further limit the subject matter of claims 1 and 11 . Applicant may cancel the claim(s), amend the claim(s) to place the claim(s) in proper dependent form, rewrite the claim(s) in independent form, or present a sufficient showing that the dependent claim(s) complies with the statutory requirements.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1 - 5, 8 - 14, 16 - 19 are rejected under 35 U.S.C. 103 as being unpatentable over Wong et al. (US 2023/0118240) in view of Bremer et al (US 2021/0374525), Reinders et al. (NPL: ”Neural Random Forest Imitation”), Stojanovic et al. (U.S. Pub. 2019/0138538), further in view of Mior et al. (NPL: ”Column2Vec: Structural Understanding via Distributed Representations of Database Schemas”).
Regarding Claim 1, Wong teaches
A computer-implemented method comprising: processing a plurality of data records using a first machine learning model; ([0044] of Wong states “The machine learning server 150 implements a machine learning system 160 for the processing of transaction data.” [0054] of Wong teaches that ML models can be generated based on a request, there can be a first request to generate a regular ML model to process transaction data)
, wherein the first machine learning model comprises a non-neural network machine learning model; ([0035] of Wong states “The term “machine learning model” is used herein to refer to at least a hardware-executed implementation of a machine learning model or function. Known models within the field of machine learning include logistic regression models, Naïve Bayes models, Random Forests, Support Vector Machines and artificial neural networks. Implementations of classifiers may be provided within one or more machine learning programming libraries including, but not limited to, scikit-learn, TensorFlow, and PyTorch.” Machine learning model Wong implements won’t be restricted to an neural network machine learning model.)
converting the first machine learning model to a neural network machine learning
model; ([0054] of Wong states “In one case, the model definition language comprises computer program code that is executable to implement one or more of training and inference of a defined machine learning model. The machine learning models may, for example, comprise, amongst others, artificial neural network architectures, ensemble models, regression models, decision trees such as random forests, graph models, and Bayesian networks. One example machine learning model based on an artificial neural network is described later with reference to FIGS. 6A to 6C.” Under BRI, this limitation can be interpreted as converting of a NN model, defined in a model definition language, into an executable NN model)
concatenating the first set of feature vectors with the second set of feature vectors to generate concatenated feature vectors; ([0111] of Wong states “The observable feature vector 916 and the context feature vector 926 are then combined to generate the overall feature vector 930. In one case, the observable feature vector 916 and the context feature vector 926 may be combined by concatenating the two feature vectors 916 and 926 to generate a longer vector.”)
generating, based on the converted first feature generator and the embedding feature generator, a combined machine learning model ([0054] of Wong teaches that machine learning system can be converted from machine learning model definitions to executable machine learning models. [0065] of Wong states “the machine learning systems may be implemented as a modular platform that allow for different machine learning models and configurations to be used to provide for the transaction processing.” [0110] of Wong states “The machine learning system 900 comprises an observable feature generator 910 and a context feature generator 930.” [0111] of Wong states “In FIG. 9 , the observable feature generator 910 receives transaction data 912, 914 and uses this to generate an observable feature vector 916. The context feature generator 930 receives ancillary data 922, 924 and uses this to generate a context feature vector 926. The observable feature vector 916 and the context feature vector 926 are then combined to generate the overall feature vector 930. In one case, the observable feature vector 916 and the context feature vector 926 may be combined by concatenating the two feature vectors 916 and 926 to generate a longer vector.”);
and training, using the concatenated feature vectors as input, the combined machine learning model, wherein the trained combined machine learning model enhances performance of the first machine learning model ([0110] of Wong states “FIG. 9 shows an example machine learning system 900 with an adaptation to allow for training based on unlabeled data.” It teaches combining two feature vectors to generate the overall feature vector 930 and use it as input to train the binary classifier. [0105] of Wong teaches “In other examples described herein an improved pipeline for training a machine learning system is presented. This improved pipeline allows the training of machine learning systems, such as those described herein, that are adapted to process transaction data… The improved pipeline for training a machine learning system involves adaptation to the feature vector generation process prior to applying the machine learning system and adaptation to the training procedure.“ that this improved pipeline allows the training of ML system that are adapted to process transaction data.)
by processing the concatenated feature vectors that incorporate both the data values via the neural network machine learning model and the structural data via the embedding machine learning model; ([0111] of Wong states “The observable feature vector 916 and the context feature vector 926 are then combined to generate the overall feature vector 930. In one case, the observable feature vector 916 and the context feature vector 926 may be combined by concatenating the two feature vectors 916 and 926 to generate a longer vector.” Wong teaches combining an observable feature vector generated from transaction data with a context feature vector and expressly states that the vectors may be concatenated. In the proposed combination, Bremer’s neural networks generate the first feature vectors from the record data values and Stojanovic, Moir and Bremer collectively provide the modified embedding model that generates the second feature vectors from the column names, table names, and textual source location metadata. Accordingly, Wong’s concatenated feature vector would incorporate both the data values through the neural network model and the structural data through the embedding model. Also, Wong’s trained classifier would process both portions of the concatenated vector as claimed herein.)
generating, using multiple-dimensional input data comprising the data values and the metadata comprising the structural data, and based on the trained combined machine learning model, labels corresponding to the plurality of data records; ([0112] of Wong states “For example, FIG. 9 shows the context feature generator 920 receiving ancillary data 922 and historical transaction data 924. The uniquely identifiable entity may be a particular end user (e.g., a card holder) or merchant and the ancillary data 922 may be data retrieved from records that are associated with the uniquely identifiable entity (e.g., so-called static data that is separate from the transaction data representing past transactions). The ancillary data 922 may comprise the ancillary data 146 or 242 as described previously. The historical transaction data 924 may comprise data associated with transactions that are outside of the temporal window, e.g. data derived from transactions that are outside the aforementioned predefined time range set with reference to a timestamp of the proposed transaction or the relative time range. The context feature generator 920 may be configured to compute aggregate metrics across the historical transaction data 924 (or retrieve pre-computed aggregate metrics) and to then include the aggregate metrics in the context feature vector. Aggregate metrics may comprise simple statistical metrics or more advanced neural network extracted features.” As Fig.12 shows method of training a machine learning system to detect anomalies within transaction data, it continues to use the system described in Fig. 9 which contains the combined machine learning model. [0122] of Wong states ”The method 1200 may be used to implement the pipeline 1000 shown in FIG. 10. At block 1202, the method 1200 comprises obtaining a training set of data samples. The data samples may comprise data samples such as 1012 in FIG. 10. Each data sample is derived, at least in part, from transaction data and is associated with one of a set of uniquely identifiable entities. For example, a data sample may have, or be retrieved based upon, one or more unique identifiers relating to a user or a merchant. In this method, at least a portion of the training set is unlabelled.“ [0123] of Wong states “At block 1204, the method 1200 comprises assigning a label indicating an absence of an anomaly to unlabelled data samples in the training set.“ [0126] of Wong states ”At block 1210, the method 1200 comprises assigning a label indicating a presence of an anomaly to the synthetic data samples” [0038] of Wong states “As discussed above, the term “tensor” is used, as per machine learning libraries, to refer to an array that may have multiple dimensions, e.g. a tensor may comprise a vector, a matrix or a higher dimensionality data structure. In preferred example, described tensors may comprise vectors with a predefined number of elements.” [0138] of Wong states “At block 1316, the received transaction data is selectively labelled based on the value output by the supervised machine learning system. This may be performed by the supervised machine learning system (e.g., as described with reference to FIG. 4 ) and/or by a separate computing device based on the output, such as the payment processor system 506 at blocks 528 or 552 in FIGS. 5A and 5B. The labelling may comprise sending a response to the original API request with the scalar value. It may also comprise applying one or more custom post-processing computations, such as application of a threshold to output a binary label of “anomaly” or “not an anomaly”.” With respect to the amended limitation, Wong teaches forming the classifier input by combining the observable feature vector and the context feature vector, including by concatenating the vectors. In the proposed combination, the observable portion would include Bremer’s feature vectors generated from the record data values, and the context portion would include the feature vectors generated from the column names, table names, and textual source location metadata using the modified embedding model discussed above. Therefore, the overall feature vector supplied to Won’s trained classifier would comprise multiple dimensional input data with both the data values and the structural metadata. Then it would generate the corresponding output labels as stated in [0110], [0138] of Wong.)
permitting, based on the labels and in real time, one or more transactions related to the plurality of data records to proceed. (Adding onto the system described from above limitation [0122-127], invention comprise the approving or declining of the transaction. [0044] and [0104] also mentions but [0139] seems more direct after the explanation of the supervised ML system from above limitation. [0139] In certain cases, block 1316 may comprise approving or declining the transaction based on the output of the supervised machine learning system. This may comprise generating control data to control whether at least one transaction within the transaction data is accepted or denied based on the value output by the supervised machine learning system. For example, in a simple case, a threshold may be applied to the output of the supervised machine learning system and values higher than the threshold (representing an “anomaly”) may be declined while those below the threshold may be approved (representing “normal” actions), which a suitable decision being made for values equal to the threshold. In certain cases, such as those illustrated in FIGS. 1A to 1C and FIGS. 5A to 5B, the transaction data may be received from a point-of-sale device in relation to a transaction to be approved. In these cases, block 1316 may comprise approving the transaction response to the value output by the supervised machine learning system being below a predefined threshold.)
Wong does not explicitly teach that
extracting, using a data profiler of an embedding machine learning model, metadata comprising structural data comprising column names, table names and locations of the plurality of data records;
generating, using data values in each column of the plurality of data records as input and via a converted first feature generator, a first set of feature vectors;
the converted first feature generator and the embedding feature generator can be machine learning models.
generating, using text representations of the metadata comprising the structural data as input and via the embedding machine learning model, a second set of feature vectors, wherein an autoencoder of the embedding machine learning model generates text embeddings corresponding to the structural data of the plurality of data records;
,wherein the neural network machine learning model mimics output from the first machine learning model;
However, Bremer teaches that
generating, using data values in each column of the plurality of data records as input and via a converted first feature generator, a first set of feature vectors; ([0041] of Bremer states “According to one embodiment, the trained data representation learning model is configured to output a feature vector of the set of feature vectors by generating for each attribute of the set of attributes an individual feature vector and combining the individual features vectors to obtain said feature vector. By processing the attributes at an individual level, this embodiment may enable to access features of the records in more details and may thus provide a reliable representation of the records.” [0052] of Bremer states “According to one embodiment, the trained data representation learning model comprises one trained neural network per attribute of the set of attributes, wherein the output of each feature vector of the set of feature vectors comprises: inputting the value of each attribute of the set of attributes into the associated trained neural network, receiving, in response to the inputting, an individual feature vector from each of the trained neural networks, and combining the individual feature vectors to obtain said feature vector.” [0111] of Wong states “In FIG. 9 , the observable feature generator 910 receives transaction data 912, 914 and uses this to generate an observable feature vector 916. The context feature generator 930 receives ancillary data 922, 924 and uses this to generate a context feature vector 926. The observable feature vector 916 and the context feature vector 926 are then combined to generate the overall feature vector 930. In one case, the observable feature vector 916 and the context feature vector 926 may be combined by concatenating the two feature vectors 916 and 926 to generate a longer vector. In other cases, combinatory logic and/or one or more neural network layers may be used that receive the observable feature vector 916 and the context feature vector 926 as input and map this input to the feature vector 930.” It would have been obvious to apply Bremer’s per attribute neural network processing to Wong’s transaction record fields because Bremer provides a known implementation for converting individual structured record values into feature vectors for downstream machine learning processing.)
the converted first feature generator and the embedding feature generator can be machine learning models (Fig. 6B, [0056] of Bremer states “A second subset of attribute level data representation learning models may be provided for the second subset of attributes, wherein each attribute level data representation learning model of the second subset is configured to generate a feature vector for a respective attribute of the second subset of attributes. A data representation learning model may be created such that it comprises the first trained data representation learning model and the second subset of attribute level data representation learning models. The created data representation learning model may be trained to generate the trained data representation learning model.” [0064] of Bremer states “In one example, the trained data representation learning model 120 may comprise multiple attribute level trained data representation learning models 121.1-121.N. Each of the attribute level trained data representation learning models 121.1-121.N may be associated with a respective attribute of the set of attributes a1 . . . aN. Each of the attribute level trained data representation learning models 121.1-121.N may be configured to receive a value of a respective attribute a1 . . . aN and to generate a corresponding individual feature vector.”, [0072] of Bremer states “The trained data representation learning model 120 may be configured to generate an individual feature vector for each received value. The individual feature vectors may be combined by the trained data representation learning model 120 in order to generate a feature vector that represents the record R1. This example may particularly be advantageous in case the set of attributes are of the same type. That is, a single trained data representation learning model (e.g., a single neural network) may validly generate feature vectors for different attributes of the same type.”).
Reinders teaches that
,wherein the neural network machine learning model mimics output from the first machine learning model; (Pg. 2 I. Introduction section of Reinders states “In this work, we present a transformation of random forests into neural networks which creates a very efficient neural network. We introduce a method for generating data from a random forest which creates any amount of input data and corresponding labels. With this data, a neural network is trained that learns to imitate the random forest.”)
Stojanovic teaches that
extracting, using a data profiler of an embedding machine learning model, metadata comprising structural data comprising column names, table names and locations of the plurality of data records; ([0054] of Stojanovic states “In some embodiments, metadata captured during enrichment associated with the ingested data source can be stored in the distributed storage system 105. System level metadata (e.g., that indicates the location of data sources, results, processing history, user sessions, execution history, and configurations, etc.) can be stored in the distributed storage system or in a separate repository accessible to the data enrichment service.” [0055] of Stojanovic states “The NL processors can automatically identify data source columns, determine the type of data in a particular column, name the column if no schema exists on input, and/or provide metadata describing the columns and/or data source. In some embodiments, the NL processors can identify and extract entities (e.g., people, places, things, etc.) from column text. NL processors can also identify and/or establish relationships within data sources and between data sources. As described further below, based on the extracted entities, the data can be repaired (e.g., to correct typographical or formatting errors) and/or enriched (e.g., to include additional related information to the extracted entities).” [0058] of Stojanovic states “For example, the prepare processing stage 108 can include ingest/prepare engines, a profiling engine and a recommendation engine.” [0062] of Stojanovic states “A profile engine can extract and/or generate metadata associated with the normalized data and a transform engine can transform (e.g., repair and/or enrich) the normalized data based on the metadata.” [0068] of Stojanovic states “A data enrichment request from the client 304 can identify a data source and/or particular data (tables, columns, files, or any other structured or unstructured data available through data sources 309 or client data store 307). Data enrichment service 302 may then access the identified data source to obtain the particular data specified in the data enrichment request. Data sources can be identified by address (e.g., URL), by storage provider name, or other identifier.” Stojanovic teaches a profiling pipeline that obtains column and source metadata and maintains table and source location identifiers. It would have been obvious to include the available table and source identifiers in the profile generated by Stojanovic’s profile engine so that the profile described both the structure and source of the profiled records.)
Mior teaches that
generating, using text representations of the metadata comprising the structural data as input and via the embedding machine learning model, a second set of feature vectors, wherein an autoencoder of the embedding machine learning model generates text embeddings corresponding to the structural data of the plurality of data records; (Pg. 2 of Mior states “To train the fastText model, we first construct documents based on the names of tables and columns of known database schemas. For example, one document might consist of the string "authors authorID firstName, lastName". Training on a collection of such documents allows us to generate embeddings, or word vectors for each of the table and column names in our training set. In addition, fastText also enables us to generate word vectors for terms outside of this vocabulary by looking at subword information.” [0068] of Stojanovic states “The received data may include structured data, unstructured data, or a combination thereof. Structure data may be based on data structures including, without limitation, an array, a record, a relational database table, a hash table, a linked list, or other types of data structures… Data sources can be identified by address (e.g., URL), by storage provider name, or other identifier. In some embodiments, access to a data source may be controlled by an access management service.” [0105] of Stojanovic states “In some embodiments, categorization module 318 can use an unsupervised machine learning tool, such as Word2Vec, to analyze an input data set. Word2Vec can receive a text input (e.g., a text corpus from a large data source) and generate a vector representation of each input word. The resulting model may then be used to identify how closely related are an arbitrary input set of words.” [0036] of Bremer states “Comparison functions like edit-distance and phonetic distance may work well on simple attributes like a first name attribute, but these comparison functions may not work on free-text fields like a 200-word product description. This may be solved using the feature vectors. For that, the content of the free-text description is encoded in a vector and close vectors refer to similar free-text descriptions.” [0051] of Bremer states “The comparison between feature vectors may be simplified using distances. In another example, the trained data representation learning model comprises an autoencoder.” It would have been obvious to use Bremer’s autoencoder as the representation learning model for generating the textual schema embeddings taught by Mior. It also would have been obvious to apply the same encoding process to Stojanovic’s textual source address because the table names, column names, and source address are textual identifiers describing the same dataset. Then, the model would generate the claimed second set of feature vectors representing the structural metadata.)
It would have been obvious to one with ordinary skill in the art before the effective filling date of the invention to combine the teachings of Wong, Reinders, Bremer, Stojanovic, Mior. Wong teaches combining different feature vector portions and providing the combined vector to a trained classifier. Bremer teaches generating feature vectors from individual record attributes using neural networks and separately discloses an autoencoder as a data representation model. Reinders teaches training a neural network to imitate a random forest, which is non-neural network machine learning model. Stojanovic teaches using a profile engine to generate metadata describing columns and data sources, including source location information. Mior teaches generating embeddings from textual table and column names. One with the ordinary skill in the art would have been motivated to incorporate the teachings of Wong with Bremer, Reinders, Stojanovic, Mior so that the classifier could process both the record values and the associated schema and source information. The combination would have been predictable use of each technique for known purpose and produced the expected result of providing both types of feature vectors to Wong’s classifier.
Regarding Claim 2, the rejection of Claim 1 is incorporated herein. Furthermore, the combination of Wong, Bremer, Reinder, Stojanovic, Mior teach
the first machine learning model comprises a decision tree model, a standard normal variate (SNV) model, a support vector machine (SVM) model or a random forest model ([0035] of Wong states “Known models within the field of machine learning include logistic regression models, Naïve Bayes models, Random Forests, Support Vector Machines and artificial neural networks”.)
Regarding Claim 3, the rejection of Claim 1 is incorporated herein. Furthermore, the combination of Wong, Bremer, Reinder, Stojanovic, Mior teach
the embedding machine learning model comprises the autoencoder, a variational autoencoder (VAE), a Bert Model or a transformer model ([0144] of Wong states “Approaches for performing unsupervised outlier detection include using tree-based isolation forests, generative adversarial networks or variational autoencoders.”)
Regarding Claim 4, the rejection of Claim 1 is incorporated herein. Furthermore, the combination of Wong, Bremer, Reinder, Stojanovic, Mior teach
the neural network machine learning model comprises a fully connected neural network (FCNN), a convolutional neural network (CNN), a recurrent neural network, or a feed forward neural network. ([0037] of Wong states “Neural network types include convolutional neural networks, recurrent neural networks, and feed-forward neural networks”.)
Regarding Claim 5, the rejection of Claim 1 is incorporated herein. Furthermore, the combination of Wong, Bremer, Reinder, Stojanovic, Mior teach
receiving the plurality of data records in a first data format; converting the plurality of data records from the first data format to a second data format; and generating, using the first machine learning model based on the plurality of data records in the second data format, a plurality of prediction labels ([0069] of Stojanovic states ”In some embodiments, data uploaded from the one or more data sources 309 can be modified into various different formats. The prepare engine 312 can convert the uploaded data into a common, normalized format, for processing by data enrichment service 302.” [0070] of Stojanovic states “ In one example, normalizing the data set to create a normalized data set includes modifying the data set having one format to an adjusted format as a normalized data set, the adjusted format being different from the format.” [0138] of Wong states “At block 1316, the received transaction data is selectively labelled based on the value output by the supervised machine learning system. This may be performed by the supervised machine learning system (e.g., as described with reference to FIG. 4 ) and/or by a separate computing device based on the output, such as the payment processor system 506 at blocks 528 or 552 in FIGS. 5A and 5B. The labelling may comprise sending a response to the original API request with the scalar value. It may also comprise applying one or more custom post-processing computations, such as application of a threshold to output a binary label of “anomaly” or “not an anomaly”.” It would have been obvious to apply Stojanovic’s format normalization process to Wong’s transaction records before sending the records to Wong’s machine learning model. Stojanovic expressly teaches converting received data into a common normalized format for subsequent processing. Therefore, the modified system would convert the records from their received format into a different processable format and then use Wong’s ML model to generate the prediction labels.)
Regarding Claim 8, the rejection of Claim 1 is incorporated herein. Furthermore, the combination of Wong, Bremer, Reinder, Stojanovic, Mior teaches
the metadata comprises a mean, a variance, a range or a length associated with a column in the plurality of data records. ([0093] of Stojanovic states “In addition to pattern identification, profile engine 326 can analyze data statistically. The profile engine 326 can characterize the content of large quantities of data, and can provide global statistics about the data and a per-column analysis of the data's content: e.g., its values, patterns, types, syntax, semantics, and its statistical properties. For example, numeric data can be analyzed statistically, including, e.g., N, mean, maximum, minimum, standard deviation, skewness, kurtosis, and/or a 20-bin histogram if N is greater than 100 and unique values is greater than K. Content may be classified for subsequent analysis.” Stojanovic teaches metadata comprising a mean associated with a column in the data records. )
Regarding Claim 9, the rejection of Claim 1 is incorporated herein. Furthermore, the combination of Wong, Bremer, Reinder, Stojanovic, Mior teaches
training, based on the plurality of data records and prediction labels generated by the first machine learning model, the neural network machine learning model to be a proxy to the first machine learning model. (Pg. 2 of Reinder states “We introduce a method for generating data from a random forest which creates any amount of input data and corresponding labels. With this data, a neural network is trained that learns to imitate the random forest.” Pg. 4 of Reinder states “These training examples are feed into the training process to teach the network predicting the same results as the random forest. To avoid overfitting, the data is generated on-the-fly so that each training example is unique. In this way, we learn an efficient representation of the decision boundaries and are able to transform random forest into neural networks.” Reinder’s random forest corresponds to the first, non-neural network machine learning model. The generated input samples correspond to the plurality of data records and the associated targets correspond to the class labels derived from the random forest. Reinders use those input target pairs to train the neural network to predict the same results as the random forest. Therefore, the trained neural network serves as a proxy for the first machine learning model.)
Regarding Claim 10, the rejection of Claim 1 is incorporated herein. Furthermore, the combination of Wong, Bremer, Reinder, Stojanovic, Mior teaches
receiving, as output from the first machine learning model and based on the plurality of data records, a plurality of prediction labels associated with the plurality of data records; ([0138] of Wong states “At block 1316, the received transaction data is selectively labelled based on the value output by the supervised machine learning system. This may be performed by the supervised machine learning system (e.g., as described with reference to FIG. 4).”)
providing, as input to a classification layer, the concatenated feature vectors; (FIG 9. Of Wong states 940 binary classifier is trained with 930 input vectors. [0111] of Wong states “In one case, the observable feature vector 916 and the context feature vector 926 may be combined by concatenating the two feature vectors 916 and 926” and these concatenated feature vectors are inputted as input vector 930 stated in [0111] of Wong “combinatory logic and/or one or more neural network layers may be used that receive the observable feature vector 916 and the context feature vector 926 as input and map this input to the feature vector 930.”)
receiving, as output from the classification layer and based on the concatenated feature vectors, a plurality of new prediction labels associated with the plurality of data records; (FIG 9. 950 the scalar output, stated in [0110] of Wong is indicative of a presence of an anomaly. If the binary classifier 940 is trained on data that has two assignable labels, resulting output will be in binary label.)
comparing the plurality of prediction labels with the plurality of new prediction labels; ([0117] of Wong states “In the case that the binary classifier 940 comprises a neural network architecture, during training, a prediction from the binary classifier 940 in the form of scalar output 950 may be compared within the assigned label 1076, e.g. in a loss function, and an error based on the difference between the scalar output 950 and one of the numeric values 0 or 1 in the label may be propagated back through the neural network architecture. In this case, a differential of the loss function with respect to the weights of the neural network architecture may be determined and used to update those weights”)
And training the combined machine learning model based on the comparison. ([0136] of Wong states “Training may comprise applying an available model fitting function in a machine learning computer program code library. In cases where the supervised machine learning system comprises a neural network architecture, training may comprise applying backpropagation with gradient descent, using a loss function based on a difference between a prediction output by the supervised machine learning system and the assigned labels. Training may be performed at a configuration stage prior to application of the method 1300.”)
Claims 11, 19 recite substantially similar subject matter as claim 1 respectively, and are rejected with the same rationale, mutatis mutandis.
Claims 12 – 14, 16 – 18 recite substantially similar subject matter as claims 2 – 4, 8 – 10 respectively, and are rejected with the same rationale, mutatis mutandis.
Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Wong et al. (US 2023/0118240 A1) in view of Bremer et al (US 2021/0374525), Tristan et al (US 2019/0095805 A1), Reinders et al. (NPL: ”Neural Random Forest Imitation”), Stojanovic et al. (U.S. Pub. 2019/0138538), Mior et al. (NPL: ”Column2Vec: Structural Understanding via Distributed Representations of Database Schemas”),, further in view of Goodsitt et al. (US 2020/0111019 A1).
Regarding Claim 6, the rejection of Claim 5 is incorporated herein. Furthermore, the combination of Wong, Bremer, Reinder, Stojanovic, Mior teach
the plurality of data records comprise transaction records, ([0002] of Wong states “The present invention relates to systems and methods for applying machine learning systems to transaction data”.)
The combination of Wong and Bremer do not appear to explicitly teach
and wherein the predicted labels comprise an indication whether the plurality of data records contain sensitive data.
However, Goodsitt—directed to analogous art—teaches
and wherein the predicted labels comprise an indication whether the plurality of data records contain sensitive data. ([0057] of Goodsitt states “The recurrent neural network can be configured to predict whether a character of a training sequence is part of a sensitive data portion”.)
It would have been obvious to one with ordinary skill in the art before the effective filling date of the invention to combine the teachings of Goodsitt with Wong, Reinders, Bremer, Stojanovic, Mior. Wong teaches combining different feature vector portions and providing the combined vector to a trained classifier. Bremer teaches generating feature vectors from individual record attributes using neural networks and separately discloses an autoencoder as a data representation model. Reinders teaches training a neural network to imitate a random forest, which is non-neural network machine learning model. Stojanovic teaches using a profile engine to generate metadata describing columns and data sources, including source location information. Mior teaches generating embeddings from textual table and column names. Goodsitt teaches predicting whether portions of data contain sensitive information. One with the ordinary skill in the art would have been motivated to incorporate the teachings of Goodsitt with Wong, Bremer, Reinders, Stojanovic, Mior so that a record could be identified when one or more portions of the record contain sensitive information. The combinations would have been predictable to allow records containing sensitive information to be identified for appropriate handling.
Claims 7 , 15, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Wong et al. (US 2023/0118240 A1) in view of Bremer et al (US 2021/0374525), Reinders et al. (NPL: ”Neural Random Forest Imitation”), Stojanovic et al. (U.S. Pub. 2019/0138538), Mior et al. (NPL: ”Column2Vec: Structural Understanding via Distributed Representations of Database Schemas”), further in view of Carrasco (US 2021/0012236 A1).
Regarding Claim 7, the rejection of Claim 1 is incorporated herein. Furthermore, the combination of Wong, Bremer, Reinder, Stojanovic, Mior teach
the structural data further comprises data sources, ([0068] of Stojanovic states “The received data may include structured data, unstructured data, or a combination thereof. Structure data may be based on data structures including, without limitation, an array, a record, a relational database table, a hash table, a linked list, or other types of data structures… A data enrichment request from the client 304 can identify a data source and/or particular data (tables, columns, files, or any other structured or unstructured data available through data sources 309 or client data store 307)… Data sources can be identified by address (e.g., URL), by storage provider name, or other identifier.”
The combination do not appear to explicitly teach
and correlation between features that are associated with the plurality of data records.
However, Carrasco—directed to analogous art—teaches
correlation between features that are associated with the plurality of data records. ([0064] of Carrasco states “On the other hand, feature meta-data can include standard statistical metrics (mean, average, maximum, minimum, and standard deviation) and the features' relationships with other features and models.” Therefore, it teaches relationships between features as feature metadata. It would have been obvious to represent the disclosed relationships between features using correlation.)
It would have been obvious to one with ordinary skill in the art before the effective filling date of the invention to combine the teachings of Carrasco with Wong, Reinders, Bremer, Stojanovic, Mior. Wong teaches combining different feature vector portions and providing the combined vector to a trained classifier. Bremer teaches generating feature vectors from individual record attributes using neural networks and separately discloses an autoencoder as a data representation model. Reinders teaches training a neural network to imitate a random forest, which is non-neural network machine learning model. Stojanovic teaches using a profile engine to generate metadata describing columns and data sources, including source location information. Mior teaches generating embeddings from textual table and column names. Carrasco teaches that feature metadata may include statistical metrics and relationships between features. One with the ordinary skill in the art would have been motivated to incorporate the teachings of Carrasco with Wong, Bremer, Reinders, Stojanovic, Mior so that the downstream model could also consider relationships among the record features. It would have been predictable combination to supply additional relationship or statistical metric information of data records for downstream processing.
Claim 15 recites substantially similar subject matter as claim 7 respectively, and is rejected with the same rationale, mutatis mutandis.
Claim 20 recite substantially similar subject matter as claim 7 and 8 combined respectively, and is rejected with the same rationale, mutatis mutandis.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to BYUNGKWON HAN whose telephone number is (571)272-5294. The examiner can normally be reached M-F: 9:00AM-6PM PST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Li B Zhen can be reached at (571)272-3768. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/BYUNGKWON HAN/ Examiner, Art Unit 2121
/Li B. Zhen/ Supervisory Patent Examiner, Art Unit 2121