Detailed Action
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1-17, and 25-29 are pending for examination. Claims 1 and 25 are independent.
Election/Restrictions
Applicant’s election without traverse of claims 1-17 and 25-29 in the reply filed on 06/15/2026 is acknowledged.
Drawings
The drawings are objected to as failing to comply with 37 CFR 1.84(p)(4) because reference characters "artificial intelligence application 1210" and "artificial intelligence application 1206" have both been used to designate the artificial intelligence application. Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
The drawings are objected to as failing to comply with 37 CFR 1.84(p)(4) because reference character “1206” has been used to designate both the substantial dataset 1206 and the artificial intelligence application 1206. Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Specification
The disclosure is objected to because of the following informalities:
Para 0003 Ine 6 states "use an infinite number of attributes to organize and eventually data". It is unclear what "eventually data" meant to state.
Para 0022 first two lines mislabels with the wrong numbers the items inside the Legend box, such as services 116, which shows service 126 instead in Fig 1.
Para 0043 lines 1-2 states “semantic type 1310” which appears to be mislabeled.
Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-17, and 25-29 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
The term “substantial” in claims 1, 5-7, 9-11, 14, 25, and 27 is a relative term which renders the claim indefinite. The term “substantial” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention.
The term “at least one of a maximally similar and a maximally diverse data objects to be labeled.” in claim 4 is a relative term which renders the claim indefinite. The term “maximally” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention.
The term “wherein the substantial dataset is an unstructured data that is so voluminous that traditional data processing methods are unable to discreetly organize the data in a structured schema” in claim 10 is a relative term which renders the claim indefinite. The term “so voluminous” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention.
Claim 2 states “at least one of an ensemble of single machine learning model in the ensemble”. It is unclear if the claim is describing a single machine learning model or multiple machine learning models. For purpose of examination, examiner interprets the limitation as describing a single model or multiple models.
Claim 25 recites the limitation "the substantial dataset" in line 3. There is insufficient antecedent basis for this limitation in the claim.
Dependent claims are also rejected under 112(b) for not resolving the 112(b) issues for the independent claims and other claims they depend on.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-17, and 25-29 rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1
According to the first part of the analysis, in the instant case, claims 1-17 are directed to a method, and claims 25-29 are also directed to a method. Thus, each of the claims falls within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter).
Regarding Claim 1:
2A Prong 1:
A method of
organizing a substantial dataset in a manner through which it can serve as an input to at least one of a machine learning model and an artificial-intelligence application, (This step for organizing a dataset is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., judgment).)
(This step for creating attributes is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., evaluation).) and
efficiently generating a data representation for at least one of the machine learning model and the artificial-intelligence application that is usable on the substantial dataset. (This step for generating a data representation is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., evaluation).)
2A Prong 2: This judicial exception is not integrated into a practical application.
Additional elements:
wherein the machine learning model is focused on at least one of a classification, a prediction, a pattern search, a trend search, a data cluster search, a data mining, and a knowledge discovery; (The specification of data to be stored is understood to be a field of use limitation. The limitation further specifies the machine learning model - See MPEP 2106.05(h).)
automatically (This suggests a computer to perform the step for creating and is understood to be mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f).)
The additional elements as disclosed above alone or in combination do not integrate the judicial exception into practical application as they are field of use in combination of generic computer functions that are implemented to perform the disclosed abstract idea above.
2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
wherein the machine learning model is focused on at least one of a classification, a prediction, a pattern search, a trend search, a data cluster search, a data mining, and a knowledge discovery; (The specification of data to be stored is understood to be a field of use limitation. The limitation further specifies the machine learning model - See MPEP 2106.05(h).)
automatically (This suggests a computer to perform the step and is understood to be mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f).)
The additional elements as disclosed above in combination of the abstract idea
are not sufficient to amount to significantly more than the judicial exception as they are
field of use in combination of generic computer functions that are implemented to perform the disclosed abstract idea above.
Regarding Claim 25
2A Prong 1:
A method of
(This step for creating attributes is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., evaluation).)
efficiently generating a data representation for the at least one of the machine learning model and the artificial-intelligence application that is usable on the substantial dataset; (This step for generating a data representation is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., evaluation).)
using an information pipeline module to compute a set of data transformations and perform calculations on the attributes; (This step for computing transformations and calculations is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., evaluation) and mathematical calculations.) and
applying the attribute specifications from the information library on the substantial data (This step for applying attribute specifications is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., evaluation).)
2A Prong 2: This judicial exception is not integrated into a practical application.
Additional elements:
Automatically (This suggests a computer to perform the step and is understood to be mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f).)
collecting a set of parameters that can be applied to the substantial data in an information library; (This step is directed to receiving information, which is understood to be insignificant extra-solution activity and data gathering. See MPEP 2106.05(g).)
and storing a result data as in the machine learning model. (This step directed to storing information, is understood to be insignificant extra- solution activity and data gathering. See MPEP 2106.05(g).)
The additional elements as disclosed above alone or in combination do not integrate the judicial exception into practical application as they are insignificant extra solution activity in combination of generic computer functions that are implemented to perform the disclosed abstract idea above.
2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
Automatically (This suggests a computer to perform the step and is understood to be mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f).)
collecting a set of parameters that can be applied to the substantial data in an information library; (This step is directed to receiving information, which is understood to be insignificant extra-solution activity and is well understood, routine and conventional activity of transmitting and receiving data as identified by the court (MPEP2106.05(d)(ll)(i))))
and storing a result data as in the machine learning model. (This step is directed to storing information, which is understood to be insignificant extra-solution activity and is well understood, routine and conventional activity as identified by the court (MPEP 2106.05(d)(ll)(IV)))))
The additional elements as disclosed above in combination of the abstract idea
are not sufficient to amount to significantly more than the judicial exception as they are
well, understood, routine and conventional activity as disclosed in combination
of generic computer functions that are implemented to perform the disclosed abstract idea above.
Regarding Claim 2
2A Prong 1:
creating additional sets of attributes to be used by at least one of an ensemble of single machine learning model in the ensemble. (This step is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., evaluation).)
2A Prong 2 & 2B: The claim does not recite any additional elements.
Regarding Claim 3
2A Prong 1:
comparing data objects by means of whether they have same combinations of attribute values on each single set of attributes; (This step is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., judgment ).)
computing a similarity score between the objects as an amount of the sets of attributes for which the corresponding combinations of attribute values are the same; (This step is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., evaluation).)
forming a query-by-example artificial-intelligence application through the comparing data objects by means of whether they have the same combinations of attribute values on each single set of attributes; (This step is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., evaluation).) and
computing the similarity score between the objects as the amount of the sets of attributes for which the corresponding combinations of attribute values are the same. (This step is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., evaluation).)
2A Prong 2 & 2B: The claim does not recite any additional elements.
Regarding Claim 4
2A Prong 1: The claim does not recite any Abstract idea.
2A Prong 2 & 2B:
using the query-by-example artificial intelligence application as a component of at least one of a diagnostic application of the machine learning model which searches for similar historical data objects to a currently diagnosed data objects and a data labeling application which searches for at least one of a maximally similar and a maximally diverse data objects to be labeled. (This step is adding the words “apply it” (or an equivalent) with the judicial exception, or merely applying an artificial intelligence as a tool to perform the abstract idea - see MPEP 2106.05(f).)
Regarding Claim 5
2A Prong 1: The claim does not recite any Abstract idea.
2A Prong 2:
collecting a set of parameters that can be applied to the substantial data in an information library, wherein the information library is a repository of attribute specifications, and wherein the attribute specifications have hyperparameters that are usable to compute various alternative versions of the attributes; these versions may have different practical properties that make some of them more useful for any one of training the accurate machine learning models and designing an efficient artificial intelligence application. (This step is directed to receiving information, which is understood to be insignificant extra-solution activity and data gathering. See MPEP 2106.05(g).)
2B:
collecting a set of parameters that can be applied to the substantial data in an information library, wherein the information library is a repository of attribute specifications, and wherein the attribute specifications have hyperparameters that are usable to compute various alternative versions of the attributes; these versions may have different practical properties that make some of them more useful for any one of training the accurate machine learning models and designing an efficient artificial intelligence application. (This step is directed to receiving information, which is understood to be insignificant extra-solution activity and is well understood, routine and conventional activity of transmitting and receiving data as identified by the court (MPEP2106.05(d)(ll)(i))))
Regarding Claim 6
2A Prong 1:
using an information pipeline module to compute a set of data transformations and perform calculations of the attribute values; (This step is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., evaluation).) and
applying the attribute specifications from the information library on the substantial data (This step is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., evaluation).)
2A Prong 2:
storing result data as in an efficient data storage, wherein computed attribute values are storable in a compressed format to decrease required storage size. (This step directed to storing information, is understood to be insignificant extra- solution activity and data gathering. See MPEP 2106.05(g).)
2B:
storing result data as in an efficient data storage, wherein computed attribute values are storable in a compressed format to decrease required storage size. (This step is directed to storing information, which is understood to be insignificant extra-solution activity and is well understood, routine and conventional activity as identified by the court (MPEP 2106.05(d)(ll)(IV)))))
Regarding Claim 7
2A Prong 1:
interpreting a description of the hyper parameter space when applying the method to create the automatic representation of data suitable for any of the machine learning models, intelligent algorithms, and artificial intelligence applications; (This step is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., judgment/evaluation).)
determining which type of data transformation to use when identifying the space of parameters based on the semantic type when the description is interpreted. (This step is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., judgment/evaluation).)
2A Prong 2:
interacting with a user of the machine learning model through an information-finder module, wherein the user is asked to define at least one of: a type of the substantial data, a location where the substantial data and attribution specifications are stored, an upload method of the substantial data, and a semantic type of the substantial data; (This step is directed to transmitting or receiving information, which is understood to be insignificant extra-solution activity and data gathering. See MPEP 2106.05(g).)
2B:
interacting with a user of the machine learning model through an information-finder module, wherein the user is asked to define at least one of: a type of the substantial data, a location where the substantial data and attribution specifications are stored, an upload method of the substantial data, and a semantic type of the substantial data; (This step is directed to transmitting or receiving information, which is understood to be insignificant extra-solution activity and is well understood, routine and conventional activity of transmitting and receiving data as identified by the court (MPEP2106.05(d)(ll)(i))))
Regarding Claim 8
2A Prong 1:
optimizing the automatic representation in a compute-efficient manner that is suitable for the machine learning model based on the semantic type using the description of the hyper parameter space. (This step is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., judgment/evaluation).)
2A Prong 2 & 2B: The claim does not recite any additional elements.
Regarding Claim 9
2A Prong 1:
estimating whether a particular attribute specification is useful for a particular substantial dataset, based on the semantic type and using the description of the hyper parameter space; (This step is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., judgment/evaluation).)
learning a set of attributes in the substantial data that are important by applying an attribute selection/extraction method; (This step is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., judgment/evaluation).) and
using the set of identified important attributes to make predictions on new data. (This step is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., judgment/evaluation).)
2A Prong 2 & 2B: The claim does not recite any additional elements.
Regarding Claim 10
2A Prong 1: The claim does not recite any Abstract idea.
2A Prong 2 & 2B:
wherein the substantial dataset is an unstructured data that is so voluminous that traditional data processing methods are unable to discreetly organize the data in a structured schema. (The specification of data to be stored is understood to be a field of use limitation. The limitation further specifies the substantial dataset. - See MPEP 2106.05(h).)
Regarding Claim 11
2A Prong 1:
analyzing the substantial dataset computationally to reveal at least one pattern, trend, and association relating to at least one of a human behavior and a human- computer interaction. (This step is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., judgment/evaluation).)
2A Prong 2 & 2B: The claim does not recite any additional elements.
Regarding Claim 12
2A Prong 1:
sampling randomly a space of possible hyper parameter settings whereby the sampled setting becomes a means to compute corresponding attributes and evaluate them through a random-sampling application. (This step is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., judgment/evaluation).)
2A Prong 2 & 2B: The claim does not recite any additional elements.
Regarding Claim 13
2A Prong 1:
applying intelligence to the random sampling through at least one of a stratified sampling technique, a systematic sampling, a cluster sampling, a probability proportional to size sampling, and an adaptive sampling method. (This step is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., judgment/evaluation).)
2A Prong 2 & 2B: The claim does not recite any additional elements.
Regarding Claim 14
2A Prong 1: The claim does not recite any Abstract idea.
2A Prong 2:
forming a data representation in a form of a tensor; (The specification of data to be stored is understood to be a field of use limitation. The limitation further specifies data representation. - See MPEP 2106.05(h).)
storing the tensor in a single cell of the substantial database; wherein the tensor is at least one of a multi-dimensional array and a mathematical object to generalize scalars, vectors, and matrices; (This step directed to storing information, is understood to be insignificant extra- solution activity and data gathering. See MPEP 2106.05(g).)
conveniently using the tensor in a machine learning method by retrieving the data from the single cell of the substantial database, wherein the substantial dataset includes a set of columns comprising any one of integers floats, characters, strings, and categories; (This step is directed to transmitting or receiving information, which is understood to be insignificant extra-solution activity and data gathering. See MPEP 2106.05(g).) and
fitting the machine learning model using the tensor. (This step is adding the words “apply it” (or an equivalent) with the judicial exception, or merely applying machine learning as a tool to perform the abstract idea - see MPEP 2106.05(f).)
2B:
forming a data representation in a form of a tensor; (The specification of data to be stored is understood to be a field of use limitation. The limitation further specifies data representation. - See MPEP 2106.05(h).)
storing the tensor in a single cell of the substantial database; wherein the tensor is at least one of a multi-dimensional array and a mathematical object to generalize scalars, vectors, and matrices; (This step is directed to storing information, which is understood to be insignificant extra-solution activity and is well understood, routine and conventional activity as identified by the court (MPEP 2106.05(d)(ll)(IV)))))
conveniently using the tensor in a machine learning method by retrieving the data from the single cell of the substantial database, wherein the substantial dataset includes a set of columns comprising any one of integers floats, characters, strings, and categories; (This step is directed to transmitting or receiving information, which is understood to be insignificant extra-solution activity and is well understood, routine and conventional activity of transmitting and receiving data as identified by the court (MPEP2106.05(d)(ll)(i)))) and
fitting the machine learning model using the tensor. (This step is adding the words “apply it” (or an equivalent) with the judicial exception, or merely applying machine learning as a tool to perform the abstract idea - see MPEP 2106.05(f).)
Regarding Claim 15
2A Prong 1:
evaluating a chosen attribute using a heuristic evaluation approach, wherein the heuristic approaches comprise at least one of an entropy of a decision class conditioned by the chosen attribute to be evaluated, an accuracy of the machine learning model trained on the representation extended by the chosen attribute; (This step is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., judgment/evaluation).) and
creating a score for the chosen attribute; (This step is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., judgment/evaluation).)
permitting an external domain expert to assist in fine-tuning the method of automatic representation of data. (This step is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., judgment/evaluation).)
2A Prong 2 & 2B: The claim does not recite any additional elements.
Regarding Claim 16
2A Prong 1:
assuring a diversity of hyper parameter settings corresponding to the chosen attributes while constructing the set of attributes; (This step is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., judgment/evaluation).)
defining a metric over a space of hyperparameters to indicate which settings any of closer to each other or distant from each other; (This step is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., judgment/evaluation).)and
ensuring a complementary mix of sources of information required for calculating the set of attributes when a multimodal data comprising any of a set of images, videos, audio, and text data. (This step is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., judgment/evaluation).)
2A Prong 2 & 2B: The claim does not recite any additional elements.
Regarding Claim 17
2A Prong 1:
maintaining a heatmap of hyper parameter settings for which corresponding attributes have been already examined, whereby that heatmap registers which attributes were evaluated as good and which as bad during an iterative process of selecting attributes and adding best ones to the set of attributes, wherein iteration of the heatmap is additionally used during a process of random selection of the hyperparameters settings whereby the settings which are closer based on a metric over the space of possible hyper parameter settings to the good ones and more distant from the bad ones examined before. (This step is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., judgment/evaluation).)
2A Prong 2 & 2B: The claim does not recite any additional elements.
Regarding Claim 26
2A Prong 1:
interpreting a description of the hyper parameter when applying the method to create the automatic representation of data to apply an accurate semantics. (This step is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., judgment/evaluation).)
2A Prong 2 & 2B: The claim does not recite any additional elements.
Regarding Claim 27
2A Prong 1: The claim does not recite any Abstract idea.
2A Prong 2:
communicating with a user of the machine learning model through an information-finder module; wherein the user to define at least one of: a type of the substantial data, a location where the substantial data and attribution specifications are stored, an upload method of the substantial data, and a semantic type of the substantial data. (This step is directed to transmitting or receiving information, which is understood to be insignificant extra-solution activity and data gathering. See MPEP 2106.05(g).)
2B:
communicating with a user of the machine learning model through an information-finder module; wherein the user to define at least one of: a type of the substantial data, a location where the substantial data and attribution specifications are stored, an upload method of the substantial data, and a semantic type of the substantial data. (This step is directed to transmitting or receiving information, which is understood to be insignificant extra-solution activity and is well understood, routine and conventional activity of transmitting and receiving data as identified by the court (MPEP2106.05(d)(ll)(i))))
Regarding Claim 28
2A Prong 1:
determining which type of data transformation to use when identifying the set of parameters based on the semantic type when the description is interpreted. (This step is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., judgment/evaluation).)
2A Prong 2 & 2B: The claim does not recite any additional elements.
Regarding Claim 29
2A Prong 1:
optimizing the automatic representation in a compute efficient manner that is suitable for the machine learning model based on the semantic type using the description of the hyper parameter. (This step is practically performable in the human mind and is understood to be a recitation of a mental process (i.e., judgment/evaluation).)
2A Prong 2 & 2B: The claim does not recite any additional elements.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1-2, 5-6, 10-13, and 25 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Zhang et al. (US 20230132064 A1).
Regarding Claim 1
Zhang discloses: A method of automatic representation of data ([Abstract] describes an automated machine learning (AutoM) framework.), comprising:
organizing a substantial dataset in a manner through which it can serve as an input to at
least one of a machine learning model and an artificial-intelligence application, ([Para 0030-0031, 0039, 0074 and Fig 2-3] describes a data collection and solution/cleaning configuration for data that serves as input for machine learning models.)
wherein the machine learning model is focused on at least one of a classification, a prediction, a pattern search, a trend search, a data cluster search, a data mining, and a knowledge discovery; ([Para 0033-0035, 0038, 0090, Fig 2, and Fig 4] describe machine learning model libraries with models that classify/predict.)
automatically creating a set of attributes for at least one of the machine learning model and the artificial-intelligence application that is usable on the substantial dataset; ([Para 0029, 0034, 0038-0039, 0091-0092, Fig 2 and Fig 4-5] describes automatic feature extraction and construction for machine learning models.) and
efficiently generating a data representation for at least one of the machine learning model and the artificial-intelligence application that is usable on the substantial dataset. ([Para 0029, 0034, 0038-0039, 0077-0079, 0084 Fig 2 and Fig 4-5] describes data preprocessing and features processing/engineering which produces a machine learning ready representation.)
Regarding Claim 2
Zhang discloses: The method of claim 1 further comprising:
creating additional sets of attributes to be used by at least one of an ensemble of single machine learning model in the ensemble. ([Para 0029, 0034, 0038-0039, 0091-0092, Fig 2 and Fig 4-5] describes automatic feature extraction and construction for machine learning model and library.)
Regarding Claim 5
Zhang discloses: The method of claim 1 further comprising:
collecting a set of parameters that can be applied to the substantial data in an information library, wherein the information library is a repository of attribute specifications, and wherein the attribute specifications have hyperparameters that are usable to compute various alternative versions of the attributes; these versions may have different practical properties that make some of them more useful for any one of training the accurate machine learning models and designing an efficient artificial intelligence application. ([Para 0030-0031, 0039, 0067-0070, 0074, 0079 and Fig 2-3] describes a data solution configuration/parameter for data and hyperparameters as well.)
Regarding Claim 6
Zhang discloses: The method of claim 5 further comprising:
using an information pipeline module to compute a set of data transformations and perform calculations of the attribute values; ([Para 0029, 0034, 0038-0039, 0077-0079, 0084, Fig 2 and Fig 4-5] describes data preprocessing and features processing/engineering which produces a representation. The solution generator functions as a pipeline.) and
applying the attribute specifications from the information library on the substantial data and storing result data as in an efficient data storage, wherein computed attribute values are storable in a compressed format to decrease required storage size. ([Para 0038-0039, 0075, 0091, Fig 2 and Fig 4-5] describe feature specifications that define feature processing parameters. [Para 0031 0042-0043, 0072 0083, 0097, and 0121] describes a historical database that stores solution data. [Para 0074, 0131] describe a compressed format)
Regarding Claim 10
Zhang discloses: The method of claim 1: wherein the substantial dataset is an unstructured data that is so voluminous that traditional data processing methods are unable to discreetly organize the data in a structured schema.([Para 0031 0074 0111 and Fig 3-4], Zhang describes the data as coming potentially from sensors or images that requires data cleaning/pre-processing )
Regarding Claim 11
Zhang discloses: The method of claim 10 wherein the method further comprising:
analyzing the substantial dataset computationally to reveal at least one pattern, trend, and association relating to at least one of a human behavior and a human- computer interaction. ([Para 0029, 0034, 0038-0039, 0091-0092, Fig 2 and Fig 4-5] describes feature extraction and machine learning model to analyze the dataset, which reveal patterns.)
Regarding Claim 12
Zhang discloses: The method of claim 11 wherein the method further comprising:
sampling randomly a space of possible hyper parameter settings whereby the sampled setting becomes a means to compute corresponding attributes and evaluate them through a random-sampling application. ([Para 0092, 0094, and 0099] describes random search for hyperparameters optimization.)
Regarding Claim 13
Zhang discloses: The method of claim 12 further comprising:
applying intelligence to the random sampling through at least one of a stratified sampling technique, a systematic sampling, a cluster sampling, a probability proportional to size sampling, and an adaptive sampling method. ([Para 0092-0094 and Fig 5], Zhang describes clustering algorithms.)
Regarding Claim 25
Zhang discloses: A method of automatic representation of data ([Abstract] describes an automated machine learning (AutoM) framework.), comprising:
automatically creating a set of attributes for at least one of a machine learning model and an artificial-intelligence application that is usable on the substantial dataset; ([Para 0029, 0034, 0038-0039, 0091-0092, Fig 2 and Fig 4-5] describes automatic feature extraction and construction for machine learning models.)
efficiently generating a data representation for the at least one of the machine learning model and the artificial-intelligence application that is usable on the substantial dataset; ([Para 0029, 0034, 0038-0039, 0077-0079, 0084 Fig 2 and Fig 4-5] describes data preprocessing and features processing/engineering which produces a machine learning ready representation.)
collecting a set of parameters that can be applied to the substantial data in an information library; ([Para 0030-0031, 0039, 0067-0070, 0074, 0079 and Fig 2-3] describes a data solution configuration/parameter for data and hyperparameters as well.)
using an information pipeline module to compute a set of data transformations and perform calculations on the attributes; ([Para 0029, 0034, 0038-0039, 0077-0079, 0084, Fig 2 and Fig 4-5] describes data preprocessing and features processing/engineering which produces a representation. The solution generator functions as a pipeline.) and
applying the attribute specifications from the information library on the substantial data and storing a result data as in the machine learning model. ([Para 0038-0039, 0075, 0091, Fig 2 and Fig 4-5] describe feature specifications that define feature processing parameters. [Para 0031 0042-0043, 0072 0083, 0097, and 0121] describes a historical database that stores solution data.)
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 3-4 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang in view of Riply (US 2005/0055345 A1, hereinafter "Riply").
Regarding Claim 3
Zhang discloses: The method of claim 2 further comprising:
Zhang does not explicitly disclose: comparing data objects by means of whether they have same combinations of attribute values on each single set of attributes; computing a similarity score between the objects as an amount of the sets of attributes for which the corresponding combinations of attribute values are the same; forming a query-by-example artificial-intelligence application through the comparing data objects by means of whether they have the same combinations of attribute values on each single set of attributes; and computing the similarity score between the objects as the amount of the sets of attributes for which the corresponding combinations of attribute values are the same
However, Ripley discloses in the same field of endeavor: comparing data objects by means of whether they have same combinations of attribute values on each single set of attributes; ([Para 0004, 0013, and 0114-0123] describes a similarity search engine that compares objects/records using multiple attributes.)
computing a similarity score between the objects as an amount of the sets of attributes for which the corresponding combinations of attribute values are the same; ([Para 0004, 0013, and 0114-0123] describes calculating similarity score for each attribute and aggregating.)
forming a query-by-example artificial-intelligence application through the comparing data objects by means of whether they have the same combinations of attribute values on each single set of attributes; ([Para 0013-0015, 0066, 00114, Fig 4 and Fig 37-39] describes Query object format as Query by Example (QBE).) and
computing the similarity score between the objects as the amount of the sets of attributes for which the corresponding combinations of attribute values are the same. ([Para 0004, 0013, and 0114-0123] describes calculating similarity score for each attribute and aggregating.)
It would have been obvious to a person of ordinary skill in art before the effective filling date of the invention to implement the function of Similarity Engine disclosed by Ripley into the method of Automated Machine learning disclosed by Zhang to compute similarity scores. The modification would have been obvious because one of the ordinary skills of the art would be motivated to utilize the feature of Similarity Engine disclosed by Ripley as all the references are in the field of machine learning. A person of ordinary skill of the art would have been motivated to perform the combination for being able to enable similarity determinations between records and target data base records.
Regarding Claim 4
Zhang in view of Ripley discloses: The method of claim 3 further comprising:
using the query-by-example artificial intelligence application as a component of at least one of a diagnostic application of the machine learning model which searches for similar historical data objects to a currently diagnosed data objects and a data labeling ([Para 0031 0042-0043, 0072 0083, 0097, and 0121], Zhang describes a historical database.) application which searches for at least one of a maximally similar and a maximally diverse data objects to be labeled. ([Para 0013-0015, 0066, 00114, Fig 4 and Fig 37-39], Ripley describes similarity search and Query by Example (QBE).)
Claim(s) 7-9 and 26-29 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang in view of Marques et al. (US 20200090003 A1, hereinafter "Marques").
Regarding Claim 7
Zhang discloses: The method of claim 6 further comprising: interpreting a description of the hyper parameter space when applying the method to create the automatic representation of data suitable for any of the machine learning models, intelligent algorithms, and artificial intelligence applications; ([Para 0030-0039 Fig 2 and Fig 6] describes solutions selection and hyperparameter optimization.)
interacting with a user of the machine learning model through an information-finder module, wherein the user is asked to define at least one of: a type of the substantial data, a location where the substantial data and attribution specifications are stored, an upload method of the substantial data, and a semantic type of the substantial data; ([Para 0024, 0038, 0074, 0106-0112] describes a user interface for uploading data solution configurations.) and
determining which type of data transformation to use when identifying the space of parameters based on the ([Para 0030-0039 Fig 2 and Fig 6] describes solutions selection, metadata, and hyperparameter optimization.)
Zhang does not explicitly disclose: based on the semantic type when the description is interpreted.
However, Marques discloses in the same field of endeavor: determining which type of data transformation to use when identifying the space of parameters based on the semantic type when the description is interpreted. ([Para 0040-0042, 0074-0081, and Fig 2-6] describe determining semantic information associated with data/features. Semantic information is used for automated feature engineering processes to select/guide feature transformation.)
It would have been obvious to a person of ordinary skill in art before the effective filling date of the invention to implement the function of Semantic aware feature engineering disclosed by Marques into the method of Automated Machine learning disclosed by Zhang to apply semantics. The modification would have been obvious because one of the ordinary skills of the art would be motivated to utilize the feature of Semantic aware feature engineering disclosed by Marques as all the references are in the field of machine learning. A person of ordinary skill of the art would have been motivated to perform the combination for being able to extend data types with semantic meaning and embeds domain knowledge into data transformations and generate features.
Regarding Claim 8
Zhang in view of Marques discloses: The method of claim 7 further comprising: optimizing the automatic representation in a compute-efficient manner that is suitable for the machine learning model based on the semantic type using the description of the hyper parameter space. ([Para 0030-0039 Fig 2 and Fig 6], Zhang describes solutions selection and hyperparameter optimization. [Para 0040-0042, 0074-0081, and Fig 2-6], Marques describes semantic information is used for automated feature engineering processes to select/guide feature transformation.)
Regarding Claim 9
Zhang in view of Marques discloses: The method of claim 1:
estimating whether a particular attribute specification is useful for a particular substantial dataset, based on the semantic type and using the description of the hyper parameter space; ([Para 0030-0039 Fig 2 and Fig 6], Zhang describes solutions selection and hyperparameter optimization. [Para 0040-0042, 0074-0081, and Fig 2-6], Marques describes semantic information is used for automated feature engineering processes to select/guide feature transformation.)
learning a set of attributes in the substantial data that are important by applying an attribute selection/extraction method; ([Para 0029, 0034, 0038-0039, 0091-0092, Fig 2 and Fig 4-5], Zhang describes automatic feature extraction and construction for machine learning models.) and
using the set of identified important attributes to make predictions on new data. ([Para 0029, 0034, 0038-0039, 0091-0092, Fig 2 and Fig 4-5], Zhang describes feature extraction for machine learning models to make predictions.)
Regarding Claim 26
Zhang discloses: The method of claim 25 further comprising: interpreting a description of the hyper parameter when applying the method to create the automatic representation of data ([Para 0030-0039 Fig 2 and Fig 6] describes solutions selection, metadata, and hyperparameter optimization.)
Zhang does not explicitly disclose: apply an accurate semantics.
However, Marques discloses in the same field of endeavor: interpreting a description of the hyper parameter when applying the method to create the automatic representation of data to apply an accurate semantics. ([Para 0040-0042, 0074-0081, and Fig 2-6] describe determining semantic information associated with data/features. Semantic information is used for automated feature engineering processes to select/guide feature transformation.)
It would have been obvious to a person of ordinary skill in art before the effective filling date of the invention to implement the function of Semantic aware feature engineering disclosed by Marques into the method of Automated Machine learning disclosed by Zhang to apply semantics. The modification would have been obvious because one of the ordinary skills of the art would be motivated to utilize the feature of Semantic aware feature engineering disclosed by Marques as all the references are in the field of machine learning. A person of ordinary skill of the art would have been motivated to perform the combination for being able to extend data types with semantic meaning and embed domain knowledge into data transformations and generate features.
Regarding Claim 27
Zhang in view of Marques discloses: The method of claim 26 further comprising: communicating with a user of the machine learning model through an information-finder module; wherein the user to define at least one of: a type of the substantial data, a location where the substantial data and attribution specifications are stored, an upload method of the substantial data, and a semantic type of the substantial data. ([Para 0024, 0038, 0074, 0106-0112], Zhang describes a user interface for uploading data solution configurations.)
Regarding Claim 28
Zhang in view of Marques discloses: The method of claim 27 further comprising: determining which type of data transformation to use when identifying the set of parameters based on the semantic type when the description is interpreted. ([Para 0040-0042, 0074-0081, and Fig 2-6], Marques describe determining semantic information associated with data/features. Semantic information is used for automated feature engineering processes to select/guide feature transformation.)
Regarding Claim 29
Zhang in view of Marques discloses: The method of claim 28 further comprising: optimizing the automatic representation in a compute efficient manner that is suitable for the machine learning model based on the semantic type using the description of the hyper parameter. ([Para 0030-0039 Fig 2 and Fig 6], Zhang describes solutions selection and hyperparameter optimization. [Para 0040-0042, 0074-0081, and Fig 2-6], Marques describes semantic information is used for automated feature engineering processes to select/guide feature transformation.)
Claim(s) 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang in view of Borje et al. (US 20190164096 A1, hereinafter "Borje").
Regarding Claim 14
Zhang discloses: The method of claim 1 further comprising:
forming a data representation ([Para 0029, 0034, 0038-0039, 0077-0079, 0084 Fig 2 and Fig 4-5] describes data preprocessing and features processing/engineering which produces a machine learning ready representation.)
conveniently using the tensor in a machine learning method by retrieving the data from the single cell of the substantial database, wherein the substantial dataset includes a set of columns comprising any one of integers floats, characters, strings, and categories; ([Para 0029, 0031-0034, 0038-0039, 0072 0083, 0097, Fig 2 and Fig 4-5] describes retrieving datasets and solution information from databases) and
fitting the machine learning model using the tensor. ([Para 0030-0032, 0100 Fig 2 and Fig 7] describes applying ML models to the datasets and performing model training/evaluation.)
Zhang does not explicitly disclose: forming a data representation in a form of a tensor; storing the tensor in a single cell of the substantial database; wherein the tensor is at least one of a multi-dimensional array and a mathematical object to generalize scalars, vectors, and matrices; conveniently using the tensor in a machine learning method by retrieving the data from the single cell of the substantial database, wherein the substantial dataset includes a set of columns comprising any one of integers floats, characters, strings, and categories; and fitting the machine learning model using the tensor.
However, Marques discloses in the same field of endeavor: forming a data representation in a form of a tensor; ([Para 0031-0040, 0065, 0094] describes tensor vectors.)
storing the tensor in a single cell of the substantial database; wherein the tensor is at least one of a multi-dimensional array and a mathematical object to generalize scalars, vectors, and matrices; ([Para 0020-0025, 0040, 0098-0100, Fig 3 and Fig 5] describes storing tensor vectors in cells.)
conveniently using the tensor in a machine learning method by retrieving the data from the single cell of the substantial database, wherein the substantial dataset includes a set of columns comprising any one of integers floats, characters, strings, and categories; ([Para 0021-0023, 0067-0076, Fig 7] describes dataset with columns comprising categories/strings.) and
fitting the machine learning model using the tensor. ([Para 0019, 0054, 0093, Fig 5 and Fig 8] describe training a machine learning model.)
It would have been obvious to a person of ordinary skill in art before the effective filling date of the invention to implement the function of Tensor vectors disclosed by Borje into the method of Automated Machine learning disclosed by Zhang to apply tensor vectors. The modification would have been obvious because one of the ordinary skills of the art would be motivated to utilize the feature of Tensor vectors disclosed by Borje as all the references are in the field of machine learning. A person of ordinary skill of the art would have been motivated to perform the combination for being able to perform machine learning operations with tensor vectors.
Claim(s) 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang in view of Moghadam et al. (US 2020/0334569 A1, hereinafter "Moghadam").
Regarding Claim 15
Zhang discloses: The method of claim 1 further comprising:
evaluating a chosen attribute using a heuristic evaluation approach, wherein the heuristic approaches comprise at least one of an entropy of a decision class conditioned by the chosen attribute to be evaluated, an accuracy of the machine learning model trained on the representation extended by the chosen attribute; ([Para 0030-0032, 0077-0079, 0084, 0100, Fig 2 and Fig 7] describes feature/solution evaluation.) and
Zhang does not explicitly disclose: creating a score for the chosen attribute; permitting an external domain expert to assist in fine-tuning the method of automatic representation of data.
However, Moghadam discloses in the same field of endeavor: evaluating a chosen attribute using a heuristic evaluation approach, wherein the heuristic approaches comprise at least one of an entropy of a decision class conditioned by the chosen attribute to be evaluated, an accuracy of the machine learning model trained on the representation extended by the chosen attribute; ([Para 0057-0060, 0063, 0068 and Fig 3] describes evaluating accuracy of models with features.)
creating a score for the chosen attribute; ([Para 0057-0060, 0063, 0068 and Fig 3] describes RML generates suitability/performance scores.)
permitting an external domain expert to assist in fine-tuning the method of automatic representation of data. ([Para 0052, 0090-0092] describes utilizing human experts for training.)
It would have been obvious to a person of ordinary skill in art before the effective filling date of the invention to implement the function of Hyperparameter Predictors disclosed by Moghadam into the method of Automated Machine learning disclosed by Zhang to apply feature scoring and user expertise. The modification would have been obvious because one of the ordinary skills of the art would be motivated to utilize the feature of Hyperparameter Predictors disclosed by Moghadam as all the references are in the field of machine learning. A person of ordinary skill of the art would have been motivated to perform the combination for being able to improve the accuracy of automatically selecting models.
Claim(s) 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang in view of Moghadam and Adams et al. (US 2016/0292129 A1, hereinafter "Adams").
Regarding Claim 16
Zhang in view of Moghadam discloses: The method of claim 15 further comprising:
assuring a diversity of hyper parameter settings corresponding to the chosen attributes while constructing the set of attributes; ([Para 0030-0034 0098 and Fig 6], Zhang describes hyperparameter optimizations such as s grid search and random search, and Bayesian optimization.)
Zhang in view of Moghadam does not explicitly disclose: defining a metric over a space of hyperparameters to indicate which settings any of closer to each other or distant from each other; and ensuring a complementary mix of sources of information required for calculating the set of attributes when a multimodal data comprising any of a set of images, videos, audio, and text data.
However, Adams discloses in the same field of endeavor: defining a metric over a space of hyperparameters to indicate which settings any of closer to each other or distant from each other; ([Para 0007, 0061, 0093-0094 and 0136] describes relating values of hyperparameters of a machine learning system using distance.) and
ensuring a complementary mix of sources of information required for calculating the set of attributes when a multimodal data comprising any of a set of images, videos, audio, and text data. ([Para 0076-0077] discloses multimodal data.)
It would have been obvious to a person of ordinary skill in art before the effective filling date of the invention to implement the function of Bayesian Optimization disclosed by Adams into the method of Zhang in view of Moghadam to define a metric over a hyperparameter space. The modification would have been obvious because one of the ordinary skills of the art would be motivated to utilize the feature of Bayesian Optimization disclosed by Adams as all the references are in the field of machine learning. A person of ordinary skill of the art would have been motivated to perform the combination for being able to find the best hyperparameters for a machine learning model.
Claim(s) 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang in view of Moghadam, Adams, and Kubota et al. (US 2022/0147829 A1, hereinafter "Kubota").
Regarding Claim 17
Zhang in view of Moghadam and Adams discloses: The method of claim 16 further comprising:
maintaining a ([Para 0034, 0039, 0068-0072, Fig 2, and Fig 5-6], Zhang describes maintaining solution file examined and using random search/selection for generating solutions ) whereby the settings which are closer based on a metric over the space of possible hyper parameter settings to the good ones and more distant from the bad ones examined before. ([Para 0007, 0061, 0093-0094 and 0136], Adams describes relating values of hyperparameters of a machine learning system using distance.)
Zhang in view of Moghadam and Adams does not explicitly disclose: heatmap
However, Kubota discloses in the same field of endeavor: heatmap of hyper parameter settings ([Para 0062-0063 and Fig 9] disclose a heatmap of hyper parameter settings.)
It would have been obvious to a person of ordinary skill in art before the effective filling date of the invention to implement the function of hyperparameter heatmaps disclosed by Kubota into the method of Zhang in view of Moghadam and Adams to provide a heatmap of hyper parameter settings. The modification would have been obvious because one of the ordinary skills of the art would be motivated to utilize the feature of hyperparameter heatmaps disclosed by Kubota as all the references are in the field of machine learning. A person of ordinary skill of the art would have been motivated to perform the combination for being able to explore and visualize different hyperparameter settings.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Arcot Desai et al. (US 20200272857 A1) describes labeling large datasets. Skerry-Ryan et al. (US 20240104394 A1) describes automatic machine learning production.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to TEWODROS E MENGISTU whose telephone number is (571)270-7714. The examiner can normally be reached Mon-Fri 9:30-5:30.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, ABDULLAH KAWSAR can be reached at (571)270-3169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/TEWODROS E MENGISTU/ Primary Examiner, Art Unit 2127