DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 05/07/2024. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Objections
Claims 1, 7-8, 14-15, and 20 are objected to because of the following informalities:
As to Claim 1, 8 and 15
clustering pipelines responsive to the clustering datasets should read as clustering pipelines responsive to the set of clustering datasets
As to Claim 7,14 and claim 20
The limitation “Combining the encoded trained clustering pipe with the internal scores.” should read as “Combining the encoded trained clustering pipelines with the internal scores.”
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefore, subject to the conditions and requirements of this title.
Claim 1-20 is rejected under 35 USC § 101 because claimed invention is directed to the abstract idea without significantly more.
As to claim 1:
Step 1 Analysis: Is the claim to a process, machine, manufacture or composition of matter? See MPEP § 2106.03.
Claim 1 is a method claim, therefore it falls under one of four categories of statutory subject matter.
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation “creating a set of clustering datasets using the classification datasets, and creating a plurality of unsupervised clustering pipelines; generating trained unsupervised clustering pipelines by training the set of unsupervised clustering pipelines responsive to the clustering datasets;” ” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
The limitation “processing the trained unsupervised clustering pipelines to generate internal scores and external scores for the set of clustering datasets; creating an encoded clustering pipeline by encoding the trained clustering pipeline using the external score as a label” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
The limitation “and generating a trained supervised machine learning model by combining the internal scores and the encoded clustering pipelines.” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
Step 2A Prong Two Analysis: Does the claim recite additional elements that integrate the judicial exception into a practical application? See MPEP § 2106.04(d).
“A method of creating a machine learning clustering meta learning model for use in solving a machine learning clustering problem, the method comprising” is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §§ 2106.04(d), 2106.05(f)(2).
The limitation “obtaining a plurality of information related to the machine learning clustering problem, wherein the plurality of information includes classification datasets, machine learning transformers and clustering estimators” is an additional element that amounts to adding insignificant extra-solution activity of mere data gathering to the judicial exception. The claim recites generating an output in the form of instruction for an activity. See MPEP §§ 2106.04(d), 2106.05(g).
Step 2B Analysis: Does the claim recite additional elements that amount to significantly more than the judicial exception? See MPEP § 2106.05.
“A method of creating a machine learning clustering meta learning model for use in solving a machine learning clustering problem, the method comprising” is an additional element to adding insignificant extra-solution activity of mere data gathering to the judicial exception. See MPEP § 2106.05(g). Furthermore, the additional element is directed to a method Performed by computational unit, which the courts have recognized as well‐understood, routine, and conventional when they are claimed in a generic manner. See MPEP § 2106.05(d)(II).
The limitation “obtaining a plurality of information related to the machine learning clustering problem, wherein the plurality of information includes classification datasets, machine learning transformers and clustering estimators” amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP § 2106.05(g). Furthermore, the additional element is directed at receiving data from different domains i.e. ML transformers, classification datasets[...] which the courts have recognized as well‐understood, routine, and conventional when they are claimed in a generic manner. See MPEP § 2106.05(d)(II).
Therefore, in examining elements as recited by the limitations individually and as an ordered combination, as a whole the independent claim limitations do not recite what have the courts have identified as “significantly more”.
As to claim 8
Step 1 Analysis: Is the claim to a process, machine, manufacture or composition of matter? See MPEP § 2106.03.
Claim 8 is drawn to the Hardware components. i.e. (machine), therefore claim 8 falls under one of four categories of statutory subject matter (machine/products/apparatus, process/method, manufactures and compositions of matter.
Step 2A Prong Two Analysis: Does the claim recite additional elements that integrate the judicial exception into a practical application? See MPEP § 2106.04(d).
The limitation “A computing system, comprising: a processor configured to perform operations for creating a machine learning clustering meta learning model for use in solving a machine learning clustering problem,” is additional element is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §§ 2106.04(d), 2106.05(f)(2)
Step 2B Analysis: Does the claim recite additional elements that amount to significantly more than the judicial exception? See MPEP § 2106.05.
The limitation ““A computing system, comprising: a processor configured to perform operations for creating a machine learning clustering meta learning model for use in solving a machine learning clustering problem” ” is the additional claim elements of one or more processors; executed by at least one of the processors, are not sufficient to amount to significantly more than the judicial exception since these additional claim elements are recited at a high level of generality (i.e. using a generic processor).
And for all other claim elements of claim 8 they are rejected using the PEG analysis of claim 1 since they are analogous claims.
As to claim 15
Step 1 Analysis: Is the claim to a process, machine, manufacture or composition of matter? See MPEP § 2106.03.
Claim 15 is drawn to a computer product (i.e. product), therefore claim 15 falls under one of four categories of statutory subject matter (machine/products/apparatus, process/method, manufactures and compositions of matter.
Step 2A Prong Two Analysis: Does the claim recite additional elements that integrate the judicial exception into a practical application? See MPEP § 2106.04(d).
The limitation “A computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a processor to cause the processor to perform operations for” is additional element is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §§ 2106.04(d), 2106.05(f)(2)
Step 2B Analysis: Does the claim recite additional elements that amount to significantly more than the judicial exception? See MPEP § 2106.05.
The limitation “A computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a processor to cause the processor to perform operations for” ” is the additional claim elements of one or more processors; a memory coupled to at least one of the processors; a stored instructions stored in the memory and executed by at least one of the processors, are not sufficient to amount to significantly more than the judicial exception since these additional claim elements are recited at a high level of generality (i.e. using a generic processor and generic memory.
And for all other claim elements of claim 15 they are rejected using the PEG analysis of claim 1 since they are analogous claims.
As to claim 2
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation “wherein the classification datasets include unseen datasets.” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
As to claim 3
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation “wherein creating a set of clustering datasets includes creating a repository of labeled datasets, wherein the labeled datasets are used to match a clustering pipeline with a particular labeled dataset” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
As to claim 4
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation “wherein processing includes generating a plurality of selected top k clustering pipelines by identifying a plurality of top k clustering pipelines from a plurality of k clustering pipelines.” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
As to claim 5
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation “wherein processing includes executing the plurality of selected top k clustering pipelines to obtain the internal scores, wherein the selected top k clustering pipelines are encoded and combined with the internal scores to generate an input to the machine learning clustering meta learning model.” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
As to claim 6
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation “wherein the internal scores include a silhouette_score, a calinski_harabasz_score and a davies_bouldin_score, and wherein the external scores include a normalized_mutual_info_score, a fowlkes_mallows_score and an adjusted_rand_score.” is the abstract idea of a mental process and mathematical concept that can practically be performed in the human mind, with or without the use of physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
As to claim 7
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation “wherein generating a trained supervised machine learning model includes generating an encoded trained clustering pipeline by encoding the trained clustering pipeline with the external scores as a label and combining the encoded trained clustering pipe with the internal scores.” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
As to claim 9
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation “wherein the classification datasets include unseen datasets.” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
As to claim 10
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation “wherein creating a set of clustering datasets includes creating a repository of labeled datasets, wherein the labeled datasets are used to match a clustering pipeline with a particular labeled dataset” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
As to claim 11
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation “wherein processing includes generating a plurality of selected top k clustering pipelines by identifying a plurality of top k clustering pipelines from a plurality of k clustering pipelines.” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
As to claim 12
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation “wherein processing includes executing the plurality of selected top k clustering pipelines to obtain the internal scores, wherein the selected top k clustering pipelines are encoded and combined with the internal scores to generate an input to the machine learning clustering meta learning model.” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
As to claim 13
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation “wherein the internal scores include a silhouette_score, a calinski_harabasz_score and a davies_bouldin_score, and wherein the external scores include a normalized_mutual_info_score, a fowlkes_mallows_score and an adjusted_rand_score.” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
As to claim 14
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation “wherein generating a trained supervised machine learning model includes generating an encoded trained clustering pipeline by encoding the trained clustering pipeline with the external scores as a label and combining the encoded trained clustering pipe with the internal scores.” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
As to claim 16
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation “wherein the classification datasets include unseen datasets.” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
The limitation “wherein creating a set of clustering datasets includes creating a repository of labeled datasets, wherein the labeled datasets are used to match a clustering pipeline with a particular labeled dataset” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
As to claim 17
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation “wherein processing includes generating a plurality of selected top k clustering pipelines by identifying a plurality of top k clustering pipelines from a plurality of k clustering pipelines.” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
As to claim 18
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation “wherein processing includes executing the plurality of selected top k clustering pipelines to obtain the internal scores, wherein the selected top k clustering pipelines are encoded and combined with the internal scores to generate an input to the machine learning clustering meta learning model.” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
As to claim 19
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation “wherein the internal scores include a silhouette_score, a calinski_harabasz_score and a davies_bouldin_score, and wherein the external scores include a normalized_mutual_info_score, a fowlkes_mallows_score and an adjusted_rand_score.” is the abstract idea of a mental process and mathematical concept that can practically be performed in the human mind, with or without the use of physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
As to claim 20
Step 2A Prong One Analysis: Does the claim recite an abstract idea, law of nature, or natural phenomenon? See MPEP § 2106.04(II)(A)(1).
The limitation “wherein generating a trained supervised machine learning model includes generating an encoded trained clustering pipeline by encoding the trained clustering pipeline with the external scores as a label and combining the encoded trained clustering pipe with the internal scores.” is the abstract idea of a mental process that can practically be performed in the human mind, with or without the use of physical aid such as pen and paper (including an observation, evaluation, judgment, opinion). See MPEP § 2106.04(a)(2)(III).
Claim Rejections – 35 USC § 103
The following is a quotation of 35 U.S.C. 103, which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Claim 1-3 ,7-10 14-16, 20 are rejected under 35 U.S.C. 103 as being unpatentable over Garg et. al. “Meta-Unsupervised-Learning: A supervised approach to unsupervised learning” (“Garg”) in view of Shapira et. al “Automatic selection of clustering algorithms using supervised graph embedding” (“Shapira”) and in view of Vikash et. al. “Supervising Unsupervised Learning” (“Vikash”)
As to Claim 1
Garg teaches “A method of creating a machine learning clustering meta learning model for use in solving a machine learning clustering problem, the method comprising:” (Garg,” [abs] We introduce a new paradigm. [A method] to investigate unsupervised learning, reducing unsupervised learning to supervised learning. Specifically, we mitigate the subjectivity in unsupervised decision-making by leveraging knowledge acquired from prior, possibly heterogeneous, supervised learning tasks. We demonstrate the versatility of our framework via comprehensive expositions and detailed experiments on several unsupervised problems such as (a) clustering, [creating a machine learning clustering meta learning model for use in solving a machine learning clustering problem, the method comprising]
2. obtaining a plurality of information related to the machine learning clustering problem, wherein the plurality of information includes classification datasets, machine learning transformers and clustering estimators. (Garg,” [pg-10] We downloaded all classification datasets from OpenML5 that had at most 10,000 instances, 500 features,10 classes, and no missing data to obtain a corpus of 339 datasets. [obtaining a plurality of information related to the machine learning clustering problem wherein the plurality of information includes classification datasets] [pg-10] The ten base clustering algorithms were chosen to be five clustering algorithms from scikit-learn (K-Means, Spectral, Agglomerative Single Linkage, Complete Linkage, and Ward) [machine learning transformers and clustering estimators]
Examiner notes: Under BRI, in light of specification, “[0062] where the plurality of information may include classification data sets, machine learning transformers and clustering estimators. The method 200 includes creating a set of clustering datasets using the classification datasets, as shown in operational block 204, and creating a set of unsupervised clustering pipelines using the machine learning transformers and clustering estimators, as shown in operational block 206. “Selecting these algorithms from scikit-learn details information being obtained from clustering estimators and transformers.
3. creating a set of clustering datasets using the classification datasets, (Garg [pg-3] Ignoring labels, each dataset can be viewed as an unsupervised learning problem.[pg-10] We downloaded all classification datasets from OpenML5 that had at most 10,000 instances, 500 features, 10 classes, [using the classification datasets], and no missing data to obtain a corpus of 339 datasets. For our purposes, we extracted the numeric features from all these datasets ignoring the categorical features. [creating a set of clustering datasets]
Examiner notes: creating a set of clustering datasets is interpreted as datasets created from Ignoring categorical features from the classification datasets.
4. and creating a plurality of unsupervised clustering pipelines. (Grag –“[pg-5] Meta-unsupervised learning, [unsupervised] which is henceforth the focus of this paper, simply refers to case where μ is a meta-distribution over datasets X 2 X and ground truth labeling Y 2 Y, and classifiers C are unsupervised learning algorithms that take an entire dataset X as input, such as clustering algorithms. [and creating a plurality of clustering pipelines]
5. processing the trained unsupervised clustering pipelines to generate internal scores and external scores for the set of clustering datasets. (Garg [Pg-13] We computed the Silhouette scores [generate internal scores] and Adjusted Rand Index (ARI) scores [and external scores] for each of the 339 datasets from k = 2 to 10. We conducted 10 independent experiments for each dataset to account for statistical significance and thus obtained two 90-dimensional vectors per dataset for Silhouette and ARI scores. [pg-11] First, one can run each of the algorithms on the repository and see which algorithm has the lowest average error. Error is calculated with respect to the ground truth labels by the ARI (see Section 3). We compare algorithms on the 250 openml binary classification datasets with at most 2000 instances. The ten base clustering algorithms were chosen to be five clustering algorithms from [processing the trained unsupervised clustering pipelines] scikit-learn (K-Means, Spectral, Agglomerative Single Linkage, Complete Linkage, and Ward) [set of clustering datasets.]
Examiner notes: Each Clustering Algorithm details about clustering datasets.
Garg does not explicitly teach:
generating trained unsupervised clustering pipelines by training the set of unsupervised clustering pipelines responsive to the clustering datasets.
creating an encoded clustering pipeline by encoding the trained clustering pipeline using the external score as a label.
and generating a trained supervised machine learning model by combining the internal scores and the encoded clustering pipelines
Shapira teaches “generating trained unsupervised clustering pipelines by training the set of
unsupervised clustering pipelines responsive to the clustering datasets. Shapira [pg-9] Given a collection of datasets D, a set of clustering algorithms A, and internal index m, we quantify the performance of all combinations of d ∈ D and a ∈ A using m. [by training the set of unsupervised clustering pipelines responsive to the clustering datasets]. We denote the result of this evaluation as Pa,d,m. Based on the performance, we rank the clustering algorithms for each dataset, such that the algorithm with the best performance occupies the first rank position, and the algorithm with the worst performance holds the last position. The ranking of each algorithm is de_ned as Ra;d;m.[…] With regard to step one, it is also important to mentioning the following two points. First, to deal with the non-deterministic nature of some algorithms in A, we repeat this assessment 10 times and compute the average results; thus, the result Pa;d;m [generating trained unsupervised clustering pipelines] is computed based on the average performance score of algorithm a on d, using m.
2. “and generating a trained supervised machine learning model by combining the internal scores and the encoded clustering pipelines.” (Shapira, [pg-13] we use the ranking version of the XGBoost algorithm as a meta-learner, since prior research [8] showed that XGBoost is well suited for producing a list of promising candidates. Once the meta-model has been trained, [and generating a trained supervised machine learning model] predictions for previously unseen datasets can be made.
PNG
media_image1.png
392
714
media_image1.png
Greyscale
Examiner notes: Under The BRI, In light of specification, ([paragraph “[0038] This may be accomplished by performing feature selection on X to train an XGB model (where an XGB model is a supervised learning algorithm that is used to make predictions on continuous numerical data) […] , generating a supervised machine learning model is interpreted as generating a XGBoost meta model. Additionally In Algorithm 2, encoded clustering pipelines is interpreted as Step 8, Ma and Step9 is interpreted as internal score, and step 10 is interpreted as combining internal scores and encoded clustering pipelines.
Shapira and Grag are related to the same field of endeavor (Meta learning). In view of the teachings of Shapira it would have been obvious for a person of ordinary skill in the art to apply the teachings of Shapira to Grag before the effective filing date of the claimed invention in order to optimize the meta model obtained through encoding to accurately recommend the best performing algorithm for unseen datasets. (Shapira” [abs] Using the embedding representations obtained, MARCOGE trains a ranking meta-model capable of accurately recommending top-performing algorithms for a new dataset and clustering evaluation measure.”)
Shapira in view of Grag does not teach:
creating an encoded clustering pipeline by encoding the trained clustering pipeline using the external score as a label.
Vikash teaches “creating an encoded clustering pipeline by encoding the trained clustering pipeline using the external score as a label.” (Vikash, pg-6 pg-7 “[Pg-6] We illustrate the main ideas with k = 2 clusters. First, one can run each of the algorithms on the repository and see which algorithm has the lowest average error. Error is calculated with respect to the ground truth labels [as a label] by the ARI. [using the external scores] [Pg-7] That is for each clustering algorithm Cj, we fit ARI (Yi, Cj(Xi)) from features Φ(Xi, Cj(Xi)) ∈ R5 [by encoding the trained clustering pipeline] over problems Xi, Yi using ν-SVR regression, with default parameters as implemented by scikit-learn. Call this estimator aˆj(X, Cj(X)). [creating an encoded clustering pipeline] To cluster a new dataset X ∈ Rd×m, the meta-algorithm then chooses Cj(X) for the j with greatest accuracy estimate aˆj(X, Cj(X)). The 250 problems were with train and test sets of varying sizes.
Vikash and Grag are related to the same field of endeavor (Meta learning). In view of the teachings of Vikash it would have been obvious for a person of ordinary skill in the art to apply the teachings of Vikash to Grag before the effective filing date of the claimed invention in order to optimize the performance of models and choose the best possible clustering algorithm which improvs efficiency and accuracy in meta learning. (Vikash “[pg- We show how a repository of multiple datasets annotated with ground truth labels can be used to improve average performance on several unsupervised tasks, even with simple algorithms. Theoretically, this enables us to make UL problems, such as clustering, well-defined. Prior datasets may prove useful for a variety of reasons, from simple to complex. They may help one choose the best clustering algorithm or parameter settings or transfer shared features that can be identified as useful.”)
As to Claim 8
Grag teaches “A computing system, comprising: a processor configured to perform operations” [Pg-6] We downloaded all classification datasets from OpenML (http://www.openml.org) that had at most 10,000 instances, 500 features, 10 classes, and no missing data to obtain a corpus of 339 datasets…. First, one can run each of the algorithms on the repository and see which algorithm has the lowest average error.
Examiner notes: Under BRI,Downloading Classification sets necessarily requires device/ computer and running an algorithm necessarily requires a processor to perform operations.
And for all the other limitations of Claim 8, it is rejected on the same basis as Claim 1. As the Claim are Analogous.
As to Claim 15
Grag teaches “computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a processor to cause the processor to perform operations” (Grag [Pg-6] We downloaded all classification datasets from OpenML (http://www.openml.org) that had at most 10,000 instances, 500 features, 10 classes, and no missing data to obtain a corpus of 339 datasets…. First, one can run each of the algorithms on the repository and see which algorithm has the lowest average error.
Examiner notes: Downloading Classification sets necessarily requires device/ computer containing computer storage medium within and running an algorithm necessarily requires a processor to perform operations.
And for all the other limitations of Claim 15, it is rejected on the same basis as Claim 1. As the Claim are Analogous.
As to Claim 2 and Analogous claim 9
Grag in view of Shapira and in view of Vikash teaches the method of claim 1.
Grag further teaches “wherein the classification datasets include unseen datasets.” (Grag[pg-10] We downloaded all classification datasets from OpenML5 that had at most 10,000 instances, 500 features, 10 classes, and no missing data to obtain a corpus of 339 datasets. For our purposes, we extracted numeric features from all these datasets ignoring the categorical features. […] [pg-13] The training and test sets were obtained using the following procedure. Each of the 339 datasets were designated [classification dataset] to be either a training set or test set. The number of training sets was varied over a wide range to take values in the set {140, 160, . . ., 280}. For each such split size, the training examples from the different sets [include unseen datasets] together formed the meta-training set, and the remaining sets formed the meta-test set. Thus, 8 such (training, test) partitions were obtained corresponding to these sizes.
Grag, Saphira, and Vikash are combinable for the same rationale as set forth above with respect to Claim 1
As to Claim 3 and Analogous claim 10
Grag in view of Shapira and in view of Vikash teaches the method of claim 1.
Vikash further wherein creating a set of clustering datasets includes creating a repository of labeled datasets, wherein the labeled datasets are used to match a clustering pipeline with a particular labeled dataset. (Vikash, abs, pg-2, “[abs] We introduce a framework to leverage knowledge acquired from a repository of (heterogeneous) supervised datasets to new unsupervised datasets. Pg-2 we suppose that we have a repository of datasets, annotated with ground truth labels [a repository of labeled datasets], that is drawn from a meta-distribution µ over problems, and that the given data X was drawn from this same distribution (though without labels) [wherein creating a set of clustering datasets included.] From this collection, one could, at a minimum, learn which type of algorithm works best, or, even better, which type of algorithm works best for which type of data. [wherein the labeled datasets are used to match a clustering pipeline with a particular labeled dataset.]”
Grag, Saphira, and Vikash are combinable for the same rationale as set forth above with respect to Claim 1
As to Claim 7 and Analogous claim 14 and claim 20
Grag in view of Shapira and in view of Vikash teaches the method of claim 1.
Shapira further teaches “wherein generating a trained supervised machine learning model includes” (Shapira “[pg-40] By modeling the interactions of the dataset’s instances as a graph and extracting an embedding representation that serves the same function as meta-features, we were able to develop a meta-learning model capable of effectively recommending top-performing algorithms for previously unseen datasets. [wherein generating a trained supervised machine learning model includes]”).
and combining the encoded trained clustering pipe with the internal scores. (Shapira, [pg-13] we use the ranking version of the XGBoost algorithm as a meta-learner, since prior research [8] showed that XGBoost is well suited for producing a list of promising candidates. Once the meta-model has been trained, predictions for previously unseen datasets can be made.
PNG
media_image1.png
392
714
media_image1.png
Greyscale
Examiner notes: In Algorithm 2, encoded trained clustering pipe is interpreted as Step 8, Ma and Step9 is interpreted as internal score, and step 10 is interpreted as combining the encoded trained clustering pipe with the internal scores.
Vikash further teaches “generating an encoded trained clustering pipeline by encoding the trained clustering pipeline with the external scores as a label. “(Vikash [Pg-6] We illustrate the main ideas with k = 2 clusters. First, one can run each of the algorithms on the repository and see which algorithm has the lowest average error. Error is calculated with respect to the ground truth labels [as a label] by the ARI. [with the external scores] [Pg-7] That is for each clustering algorithm Cj, [the trained clustering pipeline] we fit ARI (Yi, Cj(Xi)) from features Φ(Xi, Cj(Xi)) ∈ R5 [by encoding the trained clustering pipeline] over problems Xi, Yi using ν-SVR regression, with default parameters as implemented by scikit-learn. Call this estimator aˆj(X, Cj(X)). [generating an encoded trained clustering pipeline] To cluster a new dataset X ∈ Rd×m, the meta-algorithm then chooses Cj(X) for the j with greatest accuracy estimate aˆj(X, Cj(X)). The 250 problems were with train and test sets of varying sizes.
Grag, Saphira, and Vikash are combinable for the same rationale as set forth above with respect to Claim 1
As to Claim 16
Grag in view of Shapira and in view of Vikash teaches product of claim 15.
Grag further teaches “wherein the classification datasets include unseen datasets.” (Grag[pg-10] We downloaded all classification datasets from OpenML5 that had at most 10,000 instances, 500 features, 10 classes and no missing data to obtain a corpus of 339 datasets. For our purposes, we extracted numeric features from all these datasets ignoring the categorical features. […] [pg-13] The training and test sets were obtained using the following procedure. Each of the 339 datasets were designated [classification dataset] to be either a training set or test set. The number of training sets was varied over a wide range to take values in the set {140, 160, . . ., 280}. For each such split size, the training examples from the different sets [include unseen datasets] together formed the meta-training set, and the remaining sets formed the meta-test set. Thus, 8 such (training, test) partitions were obtained corresponding to these sizes.
Vikash further teaches “wherein creating a set of clustering datasets includes creating a repository of labeled datasets, wherein the labeled datasets are used to match a clustering pipeline with a particular labeled dataset” (Vikash, abs, pg-2, “[abs] We introduce a framework to leverage knowledge acquired from a repository of (heterogeneous) supervised datasets to new unsupervised datasets. Pg-2 we suppose that we have a repository of datasets, annotated with ground truth labels [a repository of labeled datasets], that is drawn from a meta-distribution µ over problems, and that the given data X was drawn from this same distribution (though without labels) [wherein creating a set of clustering datasets included.] From this collection, one could, at a minimum, learn which type of algorithm works best, or, even better, which type of algorithm works best for which type of data. [wherein the labeled datasets are used to match a clustering pipeline with a particular labeled dataset.]”
Grag, Saphira, and Vikash are combinable for the same rationale as set forth above with respect to Claim 1
Claim 4 -5 and 11-12 are rejected under 35 U.S.C. 103 as being unpatentable over (“Garg”) in view of (“Shapira”) and in view of (“Vikash”) and in further view of Laadan et. al, “Rankl: Meta Learning-Based Approach for Pre-Ranking Machine Learning Pipelines” (Laadan”)
As to claim 4 and Analogous claim 11 and claim 17
Grag in view of Shapira and in view of Vikash teaches the method of claim 1.
Grag in view of Shapira and in view of Vikash does not teach:
wherein processing includes generating a plurality of selected top k clustering pipelines by identifying a plurality of top k clustering pipelines from a plurality of k clustering pipelines
Laadan teaches wherein processing includes generating a plurality of selected top k clustering pipelines by identifying a plurality of top k clustering pipelines from a plurality of k clustering pipelines. (Laadan, [pg.-3] Next, the top-ranked pipelines are evaluated. Finally, the actual performance is recorded and added to our knowledgebase for future use. [Pg-4] For the training of our meta-learner, we used XGBoost. More specifically, we used the XGBRanker model with the pairwise ranking objective function and shallow trees of 150 estimators. Additionally, we used the following hyper-parameters settings: learning rate of 0.1, max depth of 8 and 150 estimators. We set the number of pipelines returned by RankML to k = 10.
[Pg-5] During the test phase, for each di 2 D, we used the matching Mdi meta-model to rank all possible pipelines and produce a ranked list based on predicted performance (see the online phase in Figure1. [wherein processing includes generating a plurality of selected top k clustering pipelines] The evaluated dataset di was then split into train and test sets using a 80%/20% ratio. The K top-ranked pipelines (by Mdt) were then trained on the training set of di and evaluated on its test set. [by identifying the plurality of top k clustering pipelines from a plurality of k clustering pipelines.])
Under BRI, in light of specification (“paragraph [0039] This may be accomplished by selecting the top 20 pipelines using the XGB meta learner and, for each of these top 20 pipelines,” generating a selected top k pipeline by identifying top k pipelines is interpreted to be disclosed by Laadan pg-3-4 and 5”)
Laadan and Grag are related to the same field of endeavor (Meta learning). In view of the teachings of Laadan it would have been obvious for a person of ordinary skill in the art to apply the teachings of Laadan to Grag before the effective filing date of the claimed invention in order to identify effective pipelines and reduce computation cost and complexity. (Laadan [pg-7] “a novel meta learning based approach for ranking machine learning pipelines. By exploring the interactions between datasets and pipeline topology, we were able to train learning models capable of identifying effective pipelines without performing computationally expensive analysis. By doing so, we address one of the main shortcomings of AutoML-based systems:
long running times and computational complexity.”).
As to Claim 5 and Analogous claim 12 and claim 18
Grag in view of Shapira and in view of Vikash and in further view of Laadan teaches the method of claim 4.
Laadan teaches “wherein processing includes executing the plurality of selected top k clustering pipelines” (laadan, “[pg -5] On the other hand, utilizes the meta-model to rank all the pipelines in the knowledge base with respect to the analyzed dataset and then returns its top-ranked pipelines. [selected top k clustering pipelines] These pipelines are then trained on the datasets train set and evaluated on the test set. [wherein processing includes executing the plurality of]
Shapira further teaches “[wherein processing includes executing the plurality of selected top k clustering pipelines] to obtain the internal scores” (Shapira [PG-8] Table 3 shows the number of times each algorithm held the first rank position. For example,If the selected measure is the Dunn index, then the DBSCAN algorithm gets the first rank position, followed by SL, AL, EAC, CL, MST, KM, KKM, KHM, GMF, MS, PSC, WL, FC, GMT, and GMD, respectively.
PNG
media_image2.png
232
700
media_image2.png
Greyscale
Examiner notes: Under BRI, In light of specification “(paragraph “[0032] Clustering pipelines are created in the form of [imputation, scaling, feature engineering, clustering estimator], where the clustering estimators may include Optis, DBScan[…]. Additionally, the method includes computing the clustering internal measures such as silhouette score […]” silhouette score is being obtained as internal score […].
wherein the selected top k clustering pipelines are encoded and combined with the internal scores (Shapira, [pg-13]
PNG
media_image1.png
392
714
media_image1.png
Greyscale
Examiner notes: Recall wherein the selected top k clustering pipelines are disclosed by laden at pg-5(see above). In algorithm 2, step 8 is interpreted as encoding the pipeline and combined with internal scores (Ra,d,m ) is interpreted as step 10.
“[are encoded and combined with the internal scores to] generate an input to the machine learning clustering meta learning model.(Shapira [pg-8, also see fig 2 (as explained below in subsection 3.3.1). We denote this score as am #top. […] [pg-9] we evaluate clustering algorithms on a large collection of diverse datasets. Then, we generate meta-features that are used to train a meta-model. [generate an input to the machine learning clustering meta learning model] The training phase includes four steps and is illustrated in Figure 1,
Grag, Shapira Vikash, and Laadan are combinable for the same rationale as set forth above with respect to Claim 1
Claims 6,13 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over (“Garg”) in view of (“Vikash”) and in view of (“Shapira”) and in further view of (“Laadan”) and in view of Coulter et.al pre grant num- US 20230252140 A1 (“Coulter”)
As to Claim 6 and Analogous claim 13 and claim 19
Grag in view of Shapira and in view of Vikash teaches the method of claim 1.
Grag further teaches wherein the internal scores include a silhouette_score, [a, and wherein the external scores include a normalized_mutual_info_score, a fowlkes_mallows_score] and an adjusted_rand_score.
(Garg [Pg-13] We computed the Silhouette scores [internal scores include a silhouette_score] and Adjusted Rand Index (ARI) scores [wherein the external scores include an adjusted_rand_score] for each of the 339 datasets from k = 2 to 10.
Grag in view of Shapira and in view of Vikash does not teach:
a calinski_harabasz_score and a davies_bouldin_score, normalized_mutual_info_score, a fowlkes_mallows_score
Coutler teaches calinski_harabasz_score and a davies_bouldin_score, normalized_mutual_info_score, a fowlkes_mallows_score ( Coutler [0042] In some implementations, re-assigning cohorts can include an elbow method, classification accuracy, rand index, Fowlkes-Mallow’s index, [a fowlkes_mallows_score] adjusted mutual information, normalized mutual information [normalized_mutual_info_score],, silhouette score, Davies-Bouldin index [a davies_bouldin_score],, Calinski-Harabasz index, [a calinski_harabasz_score] etc.
Coutler and Grag are related to the same field of endeavor (Machine learning model). In view of the teachings of Laadan it would have been obvious for a person of ordinary skill in the art to apply the teachings of Coutler to Grag before the effective filing date of the claimed invention in order to use these scores to evaluate the quality of clustering process while performing robust optimization and validation pipeline in meta-learning clustering. (Coulter “paragraph [0051] the scores can be translated to a single normalized scale (e.g. 0-100). The scores can be a metric that can be used to determine how valuable events are for present and/or future scenarios. Value can be associated with common or uncommon events, which can either have or not have a use in current or future scenarios.”).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NIROJ KOIRALA whose telephone number is (571)270-0748. The examiner can normally be reached Monday -Friday 8am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, MICHAEL HUNTLEY can be reached on (303) 297-4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/N.K./Examiner, Art Unit 2129
/MICHAEL J HUNTLEY/Supervisory Patent Examiner, Art Unit 2129