Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant's arguments filed 5/16/2026 have been fully considered but they are not persuasive.
Regarding applicant arguments for 101, applicant argues in page 11-14 “Claims 1-20 stand rejected under 35 U.S.C. § 101 as allegedly being directed to non-statutory subject matter. Applicant respectfully traverses this rejection. Nevertheless, for the sole purpose of expediting allowance and without commenting on the propriety of the Examiner's rejections, Applicant herein amends the independent claims as shown above. Applicant respectfully submits that these amendments render the§ 101 rejection moot, as further explained below. …
Similarly, Applicant submits that here the Examiner has also described the claims at too high a level of abstraction. Specifically, the Examiner amounts the claimed features to ones that can be a mental process. This is an oversimplification of the claimed features at least because the claimed operations, particularly in light of the amendments, capture a technical benefit of providing a higher quality featurization approach that culminates in the selection and/or deployment of an associated machine learning model (see paragraphs [0006], [0007], [0022], and [0031] of Applicant's Specification).
Accordingly, because Applicant's claims reflect an improvement to a technology, they are
not directed to a judicial exception and should be found patent eligible. Applicant respectfully
requests that the Examiner withdraw the§ 101 rejection for at least these reasons.” Applicant argues that the claims were examined at to high a level of abstraction. Yet the claims continue to have abstract ideas as explain in the last office action. Also the claims do not provide specificity to how the evaluation of the machine learning models is performed as recited in the claim 1 and independent claims. Further the applicant argues amended limitations that have not been examined, therefore the argument is moot and not convincing.
Regarding applicant’s arguments for U.S.C. 103, applicant argues in page 14-17 “Claims 1, 4-6, 8, 11-13, 15, and 18-20 stand rejected under 35 U.S.C. § 103 as allegedly being obvious over a combination of Bavly and Gu. Applicant respectfully traverses the rejection and requests reconsideration in light of the amendments presented herein.
…
Consequently, Bavly has not been shown to teach or suggest "generating, by the large language model, a plurality of featurization approaches" where "a first featurization approach defines a first feature set, derived from the input dataset, that prioritizes a first aspect in the input dataset" and "a second featurization approach defines a second feature set, derived from the input
dataset, that is different than the first feature set and that prioritizes a second aspect in the input dataset," as recited in claim 1.
Applicant submits that Gu does not remedy the deficiencies in Bavly. Accordingly, Applicant respectfully requests that the Examiner withdraw the § 103 rejection of claim 1.” – The applicant argues that Bayley does not teach the amended sections and a large language model. Yet Bavly was never cited to teach a large language model and further applicant argues amended limitations. Therefore applicant argument is moot and not convincing.
Specification
The disclosure is objected to because of the following informalities:
The specification mentions in paragraph 0032 and 0033 recites data categories 206-218 as a feature. However FIG. 2A does not show any feature 218. FIG. 2A does show 206, 208, 210, 212, 214, and 216 data categories. The data categories should be 206-216.
The specification recites data storage 610 in paragraph 0070. However the image show it as 710. The data storage should be 710.
Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-18, 20, and 21 rejected under 35 U.S.C. 101 because the claimed invention is directed to abstract idea without significantly more. The claim(s) recite(s) significantly more. The subject matter eligibility test for products and process is describe below for claim 1 in view of dependent claims.
Regarding claim 1:
Step 1: Is the claim to a process machine manufacture or composition of matter?
Yes – Claim 1 recites a method, which a method falls under the statutory categories.
Step 2A Prong 1: Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes – The claim recites the following:
“evaluating a performance of each of the plurality of corresponding machine learning models implemented by the plurality candidate machine learning based on the evaluation metric;” - The limitations recites a mental process of evaluating the performance of each machine learning model based on the evaluation metric (see MPEP 2106.04(a)(2)III).
“selecting, for deployment to a distributed computing environment, a machine learning model from the plurality of corresponding machine learning models implemented by the plurality of candidate machine learning pipelines, the selected machine learning model having a higher performance in relation to the performances of other machine learning models in the plurality of corresponding machine learning models.”- The limitations recites a mental process of selecting based on higher performance in relation to other machine learning models (see MPEP 2106.04(a)(2)III).
Step 2 Prong 2: Does the claim recite additional elements that integrate the judicial exception into a particular application? No –
The claim includes the additional element(s):
“A method comprising: receiving an input dataset comprising a plurality of quantities and an evaluation metric at a large language model;”
The additional elements fall under Insignificant Extra-Solution Activity as mere data gathering by obtaining data and an evaluation metric at the large language model. See MPEP 2106.5(g).
“generating, by the large language model, a plurality of data transforms, each data transform of the plurality of data transforms formatting the input dataset for processing;”
The additional elements fall under “apply it” as using a generic computer to implement a large language model to generate a plurality of data transforms. See Mere Instructions to Apply an Exemption (see MPEP 2106.05(f)).
“generating, by the large language model, a plurality of featurization approaches, wherein: a first each featurization approach defining defines a first feature set for the input dataset comprising a constituent plurality of features, derived from the input dataset, that prioritizes a first aspect in the input dataset; and a second featurization approach defines a second feature set, derived from the input dataset, that is different than the first feature set and that prioritizes a second aspect in the input dataset;”
The additional elements fall under “apply it” as using a generic computer to generate two featurization approaches based on a input dataset. See Mere Instructions to Apply an Exemption (see MPEP 2106.05(f)).
“initializing a plurality of candidate machine learning pipelines, each candidate machine learning pipeline implementing a corresponding machine learning model utilizing a data transform of the plurality of data transforms and an associated featurization approach generated by the large language model;”
The additional elements fall under “apply it” as using a generic computer to initialize a plurality of candidate machine learning pipelines. See Mere Instructions to Apply an Exemption (see MPEP 2106.05(f)).
“configuring an automated machine learning training module with a plurality of corresponding machine learning models implemented by the plurality of candidate machine learning pipelines to process the input dataset;”
The additional elements fall under “apply it” as using a generic computer to configure an automated machine training module with a plurality of candidate machine models to process the input dataset. See Mere Instructions to Apply an Exemption (see MPEP 2106.05(f)).
Step 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception?
No - The claim does not include additional elements that are sufficient to amount to a significantly more than the judicial exemption. As an order whole, the claim is directed to using Large Lange Model to preprocess data for an automated machine learning method. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements of receiving, generating, initializing and configuring fall under using generic computer to apply an exemption and mere data gathering. The method does not improve on the function of a computer, transforms an article into another article, nor is it applied by a particular machine, making the claim not patent eligible.
Regarding claim 2:
Step 2A Prong 1:
“The method of claim 1, wherein a feature of the first feature set or the second feature set is a ratio of two quantities of the plurality of quantities.” – The limitation recites a mathematical relationship where a feature is a ratio between two quantities (see MPEP 2106.04(a)(2).
Step 2A Prong 2, Step 2B: The additional element(s):
No additional elements. The judicial exemptions do not integrate into a practical application nor provide an improvement. The process does not provide an inventive concept nor provides a practical application.
Regarding claim 3:
Step 2A Prong 1:
“The method of claim 1, wherein a feature of the first feature set or the second feature set is an aggregate quantity of a subset of the plurality of quantities.” – The limitation recites a mathematical calculation by where a feature is an aggregate quantity (see MPEP 2106.04(a)(2).
Step 2A Prong 2, Step 2B: The additional element(s):
No additional elements. The judicial exemptions do not integrate into a practical application nor provide an improvement. The process does not provide an inventive concept nor provides a practical application
Regarding claim 4:
Step 2A Prong 2, Step 2B: The additional element(s):
“The method of claim 1, wherein a feature of the first feature set or the second feature set is a subdivision extracted from a quantity of the plurality of quantities.”
The additional elements fall under Insignificant Extra-Solution Activity. See MPEP 2106.5(g). The judicial exemptions do not integrate into a practical application nor provide an improvement. The process does not provide an inventive concept nor provides a practical application.
Regarding claim 5:
Step 2A Prong 2, Step 2B: The additional element(s):
“The method of claim 1, wherein a feature of the cfirst feature set or the second feature set defines a characteristic of a quantity of the plurality of quantities.”
The additional elements fall under Insignificant Extra-Solution Activity. See MPEP 2106.5(g). The judicial exemptions do not integrate into a practical application nor provide an improvement. The process does not provide an inventive concept nor provides a practical application.
Regarding claim 6:
Step 2A Prong 2, Step 2B: The additional element(s):
“The method of claim 1, wherein the plurality of data transforms is generated based on a data type of the input dataset.”
The additional elements fall under Insignificant Extra-Solution Activity. See MPEP 2106.5(g). The judicial exemptions do not integrate into a practical application nor provide an improvement. The process does not provide an inventive concept nor provides a practical application.
Regarding claim 7:
Step 2A Prong 2, Step 2B: The additional element(s):
“The method of claim 1, wherein: the evaluation metric is selected based on a machine learning task associated with the input dataset; the machine learning task is a binary classification task identifying a malicious uniform resource locator; and the evaluation metric is an area under curve metric.”
The additional elements fall under “apply it” as using a generic computer to configure classify and use a determine evaluation metric See Mere Instructions to Apply an Exemption (see MPEP 2106.05(f)).
Claims 8-14 recite a system and are analogous to the method of claims 1-7. Therefore, the rejections of claim 1-7 above applies to claims 8-14.
Claims 15-18 and 20 recite a CRM and are analogous to the method of claims 1-7. Therefore, the rejections of claim 1-7 above applies to claims 15-20.
Regarding claim 21:
Step 2A Prong 2, Step 2B: The additional element(s):
“The method of claim 1, the evaluation metric is selected based on a machine learning task associated with the input dataset; the machine learning task is a binary classification task identifying a uniform resource locator as being malicious;
each of the plurality of corresponding machine learning models is configured to execute the binary classification task identifying the uniform resource locator as being malicious.”
The additional elements fall under “apply it” as using a generic computer to configure to use machine learning models to binary classify uniform resource locator. See Mere Instructions to Apply an Exemption (see MPEP 2106.05(f)).
“the first feature set and the second feature set include subdivisions of uniform resource locators;”
The additional elements fall under Insignificant Extra-Solution Activity. See MPEP 2106.5(g). The judicial exemptions do not integrate into a practical application nor provide an improvement. The process does not provide an inventive concept nor provides a practical application.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 4-6, 8, 11-13, 15 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Bavly et al. (US20210334693A1) (“Bavly”) in view of Polleri et al. (US11556862B2) (“Polleri”).
Regarding claim 1 and analogous claims 8 and 15, Bavly teaches A method comprising: receiving an input dataset comprising a plurality of quantities and an evaluation metric [at a large language model] (Bavly para 0022, The Auto-XAI module 212 may be configured to train the selection of models 210 with the feature datasets 206. The model training purpose for solving a developer's particular technical problem may be defined first to select particular models before training the selected models. For example, a model for predicting risk score may be selected and based on user transaction data and behaviors. The auto-XAI module 212 may be configured to extract a subset of models and parameters that may be offered as alternatives, one of which may be selected as the recommended XAI model based on a trade-off between model explainability and model performance.
Para 0031, At step 404, application server 120 may execute the Auto-XAI module 212 to train the selected models 210 with the respective input feature datasets 206 [receiving an input dataset comprising a plurality of quantities]. The auto-XAI module 212 may process and generate respective trained models 214 with a respective output for each model. Application server 120 may perform model evaluation 216 and model selection 218 of the pipeline platform 200 based on trained models' outputs
Para 0037 line 14, FIG. 5 shows example training results of four example models in accordance with some embodiments of the present disclosure. Each model trained may be optimized using auto-ML techniques. As illustrated in FIG. 5, outputs of the trained models may be used to evaluate the model performance. The model performance may be represented by an accuracy indicative of an accuracy value or performance score ( e.g., F 1 score) and explainability ( also referred to herein as explainability properties). The Fl score is a measure of accuracy of the trained model and may be defined as the weighted harmonic mean of the precision and recall of the trained model. The evaluation metrics may include accuracy, precision and recall, which may be interactively selected by an expert and or developer [an evaluation metric].);
initializing a plurality of candidate machine learning pipelines, each candidate machine learning pipeline implementing a corresponding machine learning model utilizing a data transform of the plurality of data transforms and an associated featurization approach generated [by the large language model] (Bavly para 0021, FIG. 2 is a conceptual diagram of an example machine learning pipeline platform of 200 to implement explainable machine learning in accordance with the disclosed principles. The platform 200 may include various software algorithms configured as computer programs ( e.g., software) executed on one or more computers, in which the systems, models, algorithms, processes, and embodiments can be implemented various functionalities as described below. The platform 200 may explore different modeling techniques ( e.g., machine learning algorithms or models) compatible with training feature dataset and evaluate the performances of the trained models [initializing a plurality of candidate machine learning pipelines, each candidate machine learning pipeline implementing a corresponding machine learning model].
para 0022 line 1-10, The platform 200 may receive and input original data 202 and may include, among other things, algorithms of various machine learning models 208 with the aim of providing one or more recommended explainable models 218 as described herein. For example, the platform 200 may further include an Auto-XAI module 212 (e.g., Auto-XAI module 124 in FIG. 1) to receive feature datasets 206 (after undergoing feature engineering 204, explained below in more detail) and a selection of models 210 output from the set of models 208.
Para 0025, At step 304, feature engineering 204 may be performed by the application server 120 to extract and construct a plurality of feature datasets 206, which may be used an input to the auto-XAI module 212. Appropriate features may be selected and extracted to be used as input feature datasets for training purposes. A search in the appropriate parameter space may be automatically conducted to perform feature selection, so that an expert or a developer may not be required to have an intimate understanding of each of the selected models. As part of step 306, Application server 120 may perform preprocessing operations by making slight additions and or modifications to the features to generate the feature datasets 206 [utilizing a data transform of the plurality of data transforms and an associated featurization approach generated]);
configuring an automated machine learning training module with a plurality of corresponding machine learning models implemented by the plurality of candidate machine learning pipelines to process the input dataset (Bavly Fig. 2,
PNG
media_image1.png
507
1071
media_image1.png
Greyscale
[configuring an automated machine learning training module with a plurality of corresponding machine learning models implemented]
Para 0031 line 1-8, At step 404, application server 120 may execute the Auto-XAI module 212 to train the selected models 210 with the respective input feature datasets 206. The auto-XAI module 212 may process and generate respective trained models 214 with a respective output for each model. Application server 120 may perform model evaluation 216 and model selection 218 of the pipeline platform 200 based on the trained models' outputs. The models 210 may be trained by varying their respective explainability properties [by the plurality of candidate machine learning pipelines to process the input dataset]);
evaluating a performance of each of the plurality of corresponding machine learning models implemented by the plurality candidate machine learning based on the evaluation metric (Bavly Para 0031, At step 404, application server 120 may execute the Auto-XAI module 212 to train the selected models 210 with the respective input feature datasets 206. The auto-XAI module 212 may process and generate respective trained models 214 with a respective output for each model. Application server 120 may perform model evaluation 216 and model selection 218 of the pipeline platform 200 based on the trained models' outputs [evaluating a performance of each of the plurality of corresponding machine learning models implemented]. The models 210 may be trained by varying their respective explainability properties.
para 0039 line 1-6, Returning again to FIGS. 2 and 4, at step 408, application server 120 may execute models or algorithms of the platform 200 to determine an explainable model 218 as a recommended model from the set of the trained models 214 based on at least one of the accuracy value and the explainability properties [by the plurality candidate machine learning based on the evaluation metric]);
selecting, for deployment to a distributed computing environment, a machine learning model from the plurality of corresponding machine learning models implemented by the plurality of candidate machine learning pipelines, the selected machine learning model having a higher performance in relation to the performances of other machine learning models in the plurality of corresponding machine learning models (Bavly para 0039, Returning again to FIGS. 2 and 4, at step 408, application server 120 may execute models or algorithms of the platform 200 to determine an explainable model 218 as a recommended model from the set of the trained models 214 based on at least one of the accuracy value and the explainability properties. The application server 120 may select and or determine the explainable model 218 from the set of trained models 214 based on a trade-off decision made between the accuracy and explainability properties of the trained models 214. The system may conduct model evaluation 216 by performing automated ranking and assessment of models and parameters so that the best list of possible options may be determined for the expert of developer based on the trade-off between performance and explainability. As a result, the system may only keep model options that are Pareto-optimal with respect to the explainability and multi-objective optimization [and selecting a machine learning model from the plurality of corresponding machine learning models implemented by the plurality of candidate machine learning pipelines,].
Para 0047, At step 612, the application server 120 may determine or select an explainable model 218 as the trained model with best explainability properties from the subset of the trained models.
Para 0048, In one embodiment, a typical case of a multi-objective process may be used to select acceptable models such that each model in the subset of trained models passes (i.e., exceeds) the predetermined accuracy threshold for one objective ( e.g., accuracy). The predetermined accuracy threshold may be set to have at least a percentage of accuracy or a predetermined performance score. For example, the model ranking may be conducted first based on accuracy values or performance scores when explainability is not important. Further, the best option from the remaining model options may be chosen based on another objective (e.g., explainability). The explainability of the subset of the trained models may be ranked or evaluated to determine models that exceed a predetermined explainability threshold. The most explainable or simplest model may be selected as the final explainable model 218 from the subset of the trained models.
Para 0049, The input-output relationship of each trained model may be used to show and or describe where each model fails or succeeds such that an expert and or developer may get a better understanding of the areas of failure. The model training results may be analyzed to show and determine the accuracy-explainability trade-off [the selected machine learning model having a higher performance in relation to the performances of other machine learning models in the plurality of corresponding machine learning models]
Para 0051, At step 614, the selected model with explainable AI may be deployed into a practical application, which may be used to provide real-time machine learning solutions in different technical and engineering areas [for deployment to a distributed computing environment].).
Bavly does not explicitly teach [receiving an input dataset comprising a plurality of quantities and an evaluation metric] at a large language model;
generating, by the large language model, a plurality of data transforms, each data transform of the plurality of data transforms formatting the input dataset for processing;
generating, by the large language model, a plurality of featurization approaches, wherein:
a first each featurization approach defining defines a first feature set for the input dataset comprising a constituent plurality of features, derived from the input dataset, that prioritizes a first aspect in the input dataset;
and a second featurization approach defines a second feature set, derived from the input dataset, that is different than the first feature set and that prioritizes a second aspect in the input dataset;
[and an associated featurization approach generated] by the large language model;
Polleri teaches [receiving an input dataset comprising a plurality of quantities and an evaluation metric] at a large language model (Polleri Col 9 line 4-11, At 206, the functionality includes receiving a third input of one or more performance requirements for the machine learning application. The third input can be entered as native language speech or text (e.g., through the use of a chatbot) or selected via an interface ( e.g., a graphical user interface). The performance requirements can include Quality of Service (QoS) metrics.
Col 21 line 2-6, The responses 412 provided by digital assistant 406 may also be in the form of natural language, which may involve natural language generation (NLG) processing performed by digital assistant 406.
Col 21 line 22-33, The NLU processing performed by a digital assistant, such as digital assistant 406, can include various NLP related processing such as sentence parsing (e.g., tokenizing, lemmatizing, identifying part-of-speech tags for the sentence, identifying named entities in the sentence, generating dependency trees to represent the sentence structure, splitting a sentence into clauses, analyzing individual clauses, resolving anaphoras, performing chunking, and the like). A digital assistant 406 may use an NLP engine and/or a machine learning model (e.g., an intent classifier) to map end user utterances to specific intents ( e.g., specific task/action or category of task/action that the chatbot can perform [and an evaluation metric at a large language model]) (Examiner Note: NLP functions as the LLM);
generating, by the large language model, a plurality of data transforms, each data transform of the plurality of data transforms formatting the input dataset for processing;
generating, by the large language model, a plurality of featurization approaches, wherein (Polleri Col 55 line 9-24, Some organizations store data from multiple clients, suppliers, and/or domains with customizable schemas. When developing a machine learning solution that works across these different data schemas, a reconciliation step typically is done, either manually or through a tedious extract, transform, and load (ETL) process. For a given a machine learning problem (e.g., "I would like to predict sales" or "who are the most productive employees?"), this service will crawl the entire data store across clients/suppliers/ domains and automatically detect equivalent entities ( e.g., adding a column for "location" or for "address" in a data structure). The service will also automatically select the features that are predictive for each individual use case (i.e., one client/supplier/domain), effectively making the machine learning solution client-agnostic for the application developer of the organization [generating, by the large language model,] [, each data transform of the plurality of data transforms formatting the input dataset for processing].
Col 55 line 25-30, Feature discovery is not limited to analyzing the feature name, but also the feature content. For example, this feature can detect dates, or a particular distribution of the data that fits a previous known feature with a very typical distribution. Combination of more than one of these factors can lead the system to match and discover even more features [a plurality of data transforms].
Fig. 14, 1410-1414
PNG
media_image2.png
374
503
media_image2.png
Greyscale
Col 56 line 56-56, At 1410, the functionality includes automatically detecting features from new data storage according to a weighted list. In various embodiments, the technique can identify metadata for the identification of features in the data. The technique can use the weighted list to determine which features to incorporate into the machine learning solution. Those features with a higher rankings, where results are closer to ground truth data, can better predict the desired machine learning solution. Therefore it would be advantageous for the machine learning application to incorporate these features [generating, by the large language model, a plurality of featurization approaches, wherein:]):
a first each featurization approach defining defines a first feature set for the input dataset comprising a constituent plurality of features, derived from the input dataset, that prioritizes a first aspect in the input dataset;
and a second featurization approach defines a second feature set, derived from the input dataset, that is different than the first feature set and that prioritizes a second aspect in the input dataset (Polleri Col 55 line 9-24, Some organizations store data from multiple clients, suppliers, and/or domains with customizable schemas. When developing a machine learning solution that works across these different data schemas, a reconciliation step typically is done, either manually or through a tedious extract, transform, and load (ETL) process. For a given a machine learning problem (e.g., "I would like to predict sales" or "who are the most productive employees?"), this service will crawl the entire data store across clients/suppliers/ domains and automatically detect equivalent entities ( e.g., adding a column for "location" or for "address" in a data structure) [aspect in the input dataset;]. The service will also automatically select the features that are predictive for each individual use case (i.e., one client/supplier/domain), effectively making the machine learning solution client-agnostic for the application developer of the organization.
Col 55 line 25-30, Feature discovery is not limited to analyzing the feature name, but also the feature content. For example, this feature can detect dates, or a particular distribution of the data that fits a previous known feature with a very typical distribution. Combination of more than one of these factors can lead the system to match and discover even more features [featurization approach].
Col 55 line 57-64, At 1402, the functionality includes receiving an instruction to design new machine learning application. In various embodiments, the instruction can be through a user interface. In various embodiments, the instruction can be received via a chatbot. The technique can employ natural language processing to determine the machine learning model, metrics that can be used to design the new machine application.
Col 56 line 57-67, At 1412, the functionality includes feeding features to a machine learning solution. The monitoring engine 156, shown in FIG. 1, can provide feedback to the model composition engine regarding the features to incorporate into the machine learning solution.
At 1414, the functionality includes updating weighted list from new data. When new data is added to the data storage, a matching service can automatically detect which features should be fed into the machine learning solution based at least in part on the weighted list previously computed. Based on the features found for the new data, the weighted list can be updated. This list can be regularly be updated based on the new data and used to improve feature selection of existing models [derived from the input dataset] (Examiner Note: The featurization approach is define by the users commands that allows the user to create a first or second featurization approach. The method is not limited to one featurization approach));
[and an associated featurization approach generated] by the large language model (Polleri Col 66 line 58-45, In the offline case, a user can use the adaptive pipeline composition service 1800 to define what are the library components 168, shown in FIG. 1, of a pipeline to solve a specified problem. Previous learnings/patterns of similar use cases are used to determine a pipeline 1836 for new specified);
Bavly and Polleri are considered to be analogous to the claim invention because they are in the same field of machine learning services. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to have modified Bavly in view of Polleri to disclose generating features based on the user data and using a chatbot. Doing so to review client’s data and automatically select the features that are predictive for each individual use case and effectively making a machine learning solution client-agnostic (Polleri Col 2 line 1-3
A chatbot can provide an intuitive interface to allow the data scientist to generate a machine learning application without considerable programming experience.
line 41 -50, A self-adjusting corporation-wide discovery and integration feature can review a client's data store, review the labels for the various data schema, and effectively map the client's data schema to classifications used by the machine learning model. The various techniques can automatically select the features that are predictive for each individual use case (i.e., one client), effectively making a machine learning solution client-agnostic for the application developer. A weighted list of common representations of each feature for a particular machine learning solution can be generated and stored).
Regarding claim 4 and analogous claims 11 and 18, Bavly in view of Polleri teach the method of claim 1.
Bavly teaches wherein a feature of thefirst feature set or the second feature set is a subdivision extracted from a quantity of the plurality of quantities (Bavly para 0025 line 1-10, At step 304, feature engineering 204 may be performed by the application server 120 to extract and construct a plurality of feature datasets 206 [from a quantity of the plurality of quantities], which may be used an input to the auto-XAI module 212. Appropriate features may be selected and extracted to be used as input feature datasets for training purposes [wherein a feature of the] first [feature set]. A search in the appropriate parameter space may be automatically conducted to perform feature selection, so that an expert or a developer may not be required to have an intimate understanding of each of the selected models).
Regarding claim 5 and analogous claims 12, Bavly in view of Polleri teach the method of claim 1.
Bavly teaches wherein a feature of the first feature set or the second feature set defines a characteristic of a quantity of the plurality of quantities (Bavly para 0025 line 10-15, As part of step 306, Application server 120 may perform preprocessing operations by making slight additions and or modifications to the features to generate the feature datasets 206. In one or more embodiments, a flag may be added to each feature of the dataset 206 to indicate whether the feature has a semantic representation or not [defines a characteristic of a quantity of the plurality of quantities]).
Regarding claim 6 and analogous claims 13 and 20, Bavly in view of Polleri teach the method of claim 1.
Bavly and Polleri are combined in the same rational as set forth above with respect to claim 1 and analogous claims 8 and 15.
Polleri further teaches wherein the plurality of data transforms is generated based on a data type of the input dataset (Polleri Col 58 line 5-15, At 1506, the functionality can include analyzing the data to extract one or more labels for the schema of the data set. The one or more labels can describe a type of data that is contained in that portion of the data set. For example, a data label such as "address" can include information regarding address entries [on a data type of the input dataset]. The one or more labels can be extracted along with a location for the corresponding data and stored in a memory. The labels can be part of the stored data. For example, the customer may have provided the labels for the data set. The technique can also include generating labels for the features discovered in the data set [wherein the plurality of data transforms is generated based].).
Claim(s) 2, 3, 9, 10, 16 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Bavly in view of Polleri and further in view of Wei Xu, Ling Huang, Armando Fox, David Patterson, and Michael I. Jordan. 2009. Detecting large-scale system problems by mining console logs. In Proceedings of the ACM SIGOPS 22nd symposium on Operating systems principles (SOSP '09). Association for Computing Machinery, New York, NY, USA, 117–132 (“Xu”).
Regarding claim 2 and analogous claims 9 and 16, Bavly in view of Polleri teach the method of claim 1.
Bavly and Polleri are combined in the same rational as set forth above with respect to claim 1 and analogous claims 8 and 15.
Bavly does not explicitly teach wherein a feature of the constituent plurality of features is a ratio of two quantities of the plurality of quantities.
However Xu teaches wherein a feature of the first feature set or the second feature set is an aggregate quantity of a subset of the plurality of quantities. (Xu Page 6, 4. Feature Creation, This section describes our technique for constructing features from parsed logs. We focus on two features, the state ratio vector and the message count vector [wherein a feature of the] [first feature set], based on state variables and identifiers (see Section 2.1), respectively. The state ratio vector is able to capture the aggregated behavior of the system over a time window. The message count vector helps detect problems related to individual operations. Both features describe message groups constructed to have strong correlations among their members. The features faithfully capture these correlations, which are often good indicators of runtime problems. Although these features are from the same log, and similar in structure, they are constructed independently, and have different semantics.
4.1 State variables and state ration vectors,
We construct state ratio vectors y to encode this correlation: Each state ratio vector represents a group of state variables in a time window, while each dimension of the vector corresponds to a distinct state variable value , and the value of the dimension is how many times this state value appears in the time window [features is a ratio of two quantities of the plurality of quantities]).
Bavly and Xu are considered to be analogous to the claim invention because they are in the same field of distributed machine learning. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to have modified Bavly in view of Xu to disclose generating features that is a ratio of two quantities. Doing so to generate sophisticated features without the use of human input (Xu Abstract line 9-23, We then analyze these features using machine learning to detect operational problems. We show that our method enables analyses that are impossible with previous methods because of its superior ability to create sophisticated features. We also show how to distill the results of our analysis to an operator-friendly one-page decision tree showing the critical messages associated with the detected problems. We validate our approach using the Darkstar online game server and the Hadoop File System, where we detect numerous real problems with high accuracy and few false positives. In the Hadoop case, we are able to analyze 24 million lines of console logs in 3 minutes. Our methodology works on textual console logs of any size and requires no changes to the service software, no human input, and no knowledge of the software’s internals.).
Regarding claim 3 and analogous claims 10 and 17, Bavly in view of Polleri teach the method of claim 1.
Bavly and Polleri are combined in the same rational as set forth above with respect to claim 1 and analogous claims 8 and 15.
Bavly and Xu are combine in the same rational as set forth above with respect to claim 2 and analogous claims 10 and 17.
Xu further teaches wherein a feature of the first feature set or the second feature set is an aggregate quantity of a subset of the plurality of quantities. (Xu Page 122, 4. Feature Creation, This section describes our technique for constructing features from parsed logs. We focus on two features, the state ratio vector and the message count vector [wherein a feature] [first feature set], based on state variables and identifiers (see Section 2.1), respectively. The state ratio vector is able to capture the aggregated behavior of the system over a time window. The message count vector helps detect problems related to individual operations. Both features describe message groups constructed to have strong correlations among their members. The features faithfully capture these correlations, which are often good indicators of runtime problems. Although these features are from the same log, and similar in structure, they are constructed independently, and have different semantics.
Page 122-123 4.2 Identifiers and message count vectors para 2-3, To form the message count vector, we first automatically discover identifiers, then group together messages with the same identifier values, and create a vector per group. Each vector dimension corresponds to a different message type, and the value of the dimension tells how many messages of that type appear in the message group. The structure of this feature is analogous to the bag of words model in information retrieval [6]. In our application, the “document” is the message group. The dimensions of the vector consist of the union of all useful message types across all groups (analogous to all possible “terms”), and the value of a dimension is the number of appearances of the corresponding message types in a group (corresponding to “term frequency”). Algorithm 1 summarizes our three-step process for feature construction. We now try to provide intuition behind the design choices in this algorithm [aggregate quantity of a subset of the plurality of quantities]).
Claim(s) 7, 14, and 21 are rejected under 35 U.S.C. 103 as being unpatentable over Bavly in view of Polleri and further in view of M. Darling, G. Heileman, G. Gressel, A. Ashok and P. Poornachandran, "A lexical approach for classifying malicious URLs," 2015 International Conference on High Performance Computing & Simulation (HPCS), Amsterdam, Netherlands, 2015, pp. 195-202, (“Darling”).
Regarding claim 7 and analogous claim 14, Bavly in view of Polleri teach the method of claim 1.
Bavly and Polleri are combined in the same rational as set forth above with respect to claim 1 and analogous claims 8 and 15.
Bavly does not explicitly teach wherein: the evaluation metric is selected based on a machine learning task associated with the input dataset; the machine learning task is a binary classification task identifying a malicious uniform resource locator; and the evaluation metric is an area under curve metric.
However Darling teaches wherein: the evaluation metric is selected based on a machine learning task associated with the input dataset; the machine learning task is a binary classification task identifying a malicious uniform resource locator; and the evaluation metric is an area under curve metric (Darling
Page 195 I Introduction para 7, In this paper we present an approach which uses an ngram model to develop a new classification system that adheres to the strict time-constraints required for a real-time system. The system increases accuracy on out-of-sample testing data while maintaining overall classification accuracy comparable to previous work [1]–[3], [6]. Our approach uses the J48 decision tree algorithm to perform classification of URLs using 16 features extracted from an n-gram model and 71 features from other lexical properties. J48 is an open-source implementation of the C4.5 algorithm [8].
page 198-199 E. Classification Algorithms, In this study we chose to explore several classification methods. As our baseline we built a linear classifier using regularized logistic regression. Logistic regression is a parametric model for binary classification where examples are classified by their distance from a decision boundary [the machine learning task is a binary classification]. We implemented the L1 regularized logistic regression model from the LibLinear package as was done by Ma et al. [1]. For the implementation of the classifier based on n-gram modeling we used the J48 algorithm (an implementation of the C4.5 decision tree). J48 is one of many classification algorithms available in the popular Weka machine learning suite [8]. We chose J48 due to its reputation of success in this domain [15] and after evaluating its performance in comparison to Bayesian Logistic Regression, Logistic Regression, Naive Bayes, and K-Nearest Neighbors classifiers. In order to evaluate performance of each algorithm, we compared their overall accuracy and receiver operating characteristic (ROC) curve [the evaluation metric is selected based on a machine learning task associated with the input dataset].
A classifier’s accuracy is not the only necessary metric to evaluate its performance. The ambiguity lies in the fact that classification accuracy will have equal misclassification costs: an accuracy of 99% tells nothing about the false positive and negative rates. In fact, in a real-world deployment of any classifier, it is often true that the error rate of one class of data comes at a higher cost than misclassifying other types of data. For the problem of malicious web page detection, a false negative is potentially much more harmful than a false-positive since it could result in an infected system [the machine learning task is a binary classification task identifying a malicious uniform resource locator].
An ROC curve shows the false-positive and false-negative rates on the X and Y axes respectively. Therefore, the curve represents the predictive quality of the classifier independent of error costs (and class imbalance in the training data) [16]. In order to choose our classifier we examined the ROC Area- Under-the-Curve value (AUC) and accuracy rate of each type of classifier. Table III shows J48 outperforming its counterparts in AUC of its ROC. Furthermore, it achieved the highest overall classification accuracy for the n-gram model [the evaluation metric is an area under curve metric].
Page 6 III. Results, Table VI shows the confusion matrix for the J48 model with 10-fold cross validation on the entire data set. Each row represents the instances of an actual class, each column represents the instances of the predicted class by J48.
PNG
media_image3.png
123
455
media_image3.png
Greyscale
[the machine learning task is a binary classification task identifying a malicious uniform resource locator;] (i.e. the system classifies URLs as Malicious or Benign)).
Bavly and Darling are considered to be analogous to the claim invention because they are in the same field of distributed machine learning. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to have modified Bavly in view of Darling to disclose using area under the curve as an evaluation metric. Doing so to select the best classifier based on the best predictive quality (Darling E. Classification Algorithms Page 199 para 4, An ROC curve shows the false-positive and false-negative rates on the X and Y axes respectively. Therefore, the curve represents the predictive quality of the classifier independent of error costs (and class imbalance in the training data) [16]. In order to choose our classifier we examined the ROC Area Under-the-Curve value (AUC) and accuracy rate of each type of classifier. Table III shows J48 outperforming its counterparts in AUC of its ROC. Furthermore, it achieved the highest overall classification accuracy for the n-gram model.).
Regarding claim 21, Bavly in view of Polleri teach the method of claim 1.
Bavly and Polleri are combined in the same rational as set forth above with respect to claim 1 and analogous claims 8 and 15.
Bavly and Polleri are combined in the same rational as set forth above with respect to claim 7 and analogous claims 14.
Darlin teaches wherein:
the evaluation metric is selected based on a machine learning task associated with the input dataset;
the machine learning task is a binary classification task identifying a uniform resource locator as being malicious;
the first feature set and the second feature set include subdivisions of uniform resource locators;
and each of the plurality of corresponding machine learning models is configured to execute the binary classification task identifying the uniform resource locator as being malicious ((Darling page 198 Table II,
PNG
media_image4.png
197
375
media_image4.png
Greyscale
[the first feature set and the second feature set include subdivisions of uniform resource locators;]
page 198-199 E. Classification Algorithms, In this study we chose to explore several classification methods. As our baseline we built a linear classifier using regularized logistic regression. Logistic regression is a parametric model for binary classification where examples are classified by their distance from a decision boundary [the machine learning task is a binary classification]. We implemented the L1 regularized logistic regression model from the LibLinear package as was done by Ma et al. [1]. For the implementation of the classifier based on n-gram modeling we used the J48 algorithm (an implementation of the C4.5 decision tree). J48 is one of many classification algorithms available in the popular Weka machine learning suite [8]. We chose J48 due to its reputation of success in this domain [15] and after evaluating its performance in comparison to Bayesian Logistic Regression, Logistic Regression, Naive Bayes, and K-Nearest Neighbors classifiers. In order to evaluate performance of each algorithm, we compared their overall accuracy and receiver operating characteristic (ROC) curve [the evaluation metric is selected based on a machine learning task associated with the input dataset].
A classifier’s accuracy is not the only necessary metric to evaluate its performance. The ambiguity lies in the fact that classification accuracy will have equal misclassification costs: an accuracy of 99% tells nothing about the false positive and negative rates. In fact, in a real-world deployment of any classifier, it is often true that the error rate of one class of data comes at a higher cost than misclassifying other types of data. For the problem of malicious web page detection, a false negative is potentially much more harmful than a false-positive since it could result in an infected system [the machine learning task is a binary classification task identifying a uniform resource locator as being malicious;]
Page 199 Table III,
PNG
media_image5.png
185
347
media_image5.png
Greyscale
[and each of the plurality of corresponding machine learning models is configured to execute the binary classification task identifying the uniform resource locator as being malicious]).
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALFREDO CAMPOS whose telephone number is (571)272-4504. The examiner can normally be reached 7:00 - 4:00 pm M - F.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michael J. Huntley can be reached at (303) 297-4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ALFREDO CAMPOS/Examiner, Art Unit 2129
/MICHAEL J HUNTLEY/Supervisory Patent Examiner, Art Unit 2129