DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-8 are rejected under 35 U.S.C. § 101 because the claimed invention is directed to a judicial exception (an abstract idea) without significantly more.
Step 1
The claims are directed to a statutory category because they recite a computer-implemented method.
Step 2A, Prong One
Claim 1 recites limitations that describe a process of selecting information, analyzing information, and mathematically evaluating information for training a machine learning model, including:
retrieving a subset of data files from a dataset;
dividing the subset into training and evaluation datasets;
training a machine learning model;
determining predictive accuracy;
determining a number of least learned data files by a targeted sampling process;
generating a second training dataset including the least learned data files; and
retraining the machine learning model using the second training dataset.
These limitations recite concepts involving observation, evaluation, ranking, selection, and organization of information. The claimed determination of predictive accuracy, predictive confidence, targeted sampling, and selection of least learned data files are mathematical evaluations performed on data and constitute mathematical concepts and mental processes that fall within the abstract idea grouping identified in the 2019 Revised Patent Subject Matter Eligibility Guidance (84 Fed. Reg. 50) and MPEP §2106.04(a)(2).
The dependent claims likewise recite additional mathematical and analytical operations, including:
establishing a benchmark accuracy (claims 2-3),
arranging data files according to predictive confidence (claim 4),
generating a quartile plot and identifying a first quartile (claim 5),
limiting the number of selected files (claim 6), and
iteratively repeating the sampling and retraining process until a benchmark is satisfied (claims 7-8).
These additional limitations likewise recite mathematical analysis, data evaluation, and organization of information, which remain within the judicial exception.
Accordingly, claims 1-8 recite an abstract idea.
Step 2A, Prong Two
The claims do not integrate the judicial exception into a practical application.
The additional elements beyond the abstract idea merely recite execution of the abstract idea on generic computer technology, including a generic computer and a generic machine learning model.
The claims do not:
improve the functioning of a computer or another technology;
improve the operation of the machine learning model itself through a specific technological architecture;
improve memory management, processor utilization, communication protocols, or computer resource allocation;
effect a transformation of a particular article;
employ a particular machine in a meaningful way beyond serving as a tool to perform the abstract idea; or
otherwise apply the abstract idea in a manner that imposes a meaningful limit on the judicial exception.
Instead, the machine learning model merely serves as a tool for carrying out the abstract process of selecting training data, evaluating prediction performance, and retraining the model.
Accordingly, the claims do not integrate the judicial exception into a practical application.
Step 2B
The additional claim elements, individually and as an ordered combination, do not amount to significantly more than the abstract idea.
The claims merely require conventional computer functions including:
retrieving stored data;
dividing data into datasets;
executing machine learning training;
evaluating prediction accuracy;
ranking data according to confidence;
generating datasets; and
repeating the process iteratively.
These functions represent well-understood, routine, and conventional computer operations that merely automate the abstract idea using generic computer components.
Considering the elements as an ordered combination does not alter this conclusion. The ordered sequence simply performs conventional iterative machine learning training in which additional data are selected based upon evaluation results and used for subsequent training. The sequence does not produce any technological improvement to computer functionality but merely improves the quality of the machine learning model itself through additional training data, which is an improvement in the abstract idea rather than an improvement in computer technology.
Accordingly, the claims do not recite significantly more than the judicial exception and are therefore ineligible under 35 U.S.C. §101.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1 and 4 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by SHIM et al. [US 2021/0383158 A1].
Regarding claim 1, SHIM teaches “A computer-implemented method comprising: retrieving a subset of data files from an original dataset, wherein data files not included in the subset of data files are a remaining dataset;” as “selecting a subset of the first candidate points by aggregating the determined KNN-SVs of the first candidate points” [¶0023]
“dividing the subset of data files into an initial training dataset and an evaluation dataset;” as “wherein the evaluation set and the first candidate set comprise different data points,” [¶0023]
“executing an in initial training pass on a machine learning model, wherein the initial training pass trains the machine learning model using the initial training dataset;” as “As a result, an artificial intelligence based object identifying apparatus trains the artificial neural network using a machine learning algorithm or requests a trained artificial neural network from the AI server 120 to receive the trained artificial neural network from the AI server 120. ” [¶0049]
“determining, after the initial training pass, a predictive accuracy of the machine learning model using the evaluation dataset;” as “ when every possible subset of data points is considered, s(i) measures the average marginal improvement of utility given by the sample i. By setting the utility as test accuracy in ML classification tasks, the SV can discover how much of the test accuracy is attributed to a training instance.” [¶0096]
“determining, by a targeted sampling process, a number of least learned data files from the remaining dataset; generating a second training dataset including the number of least learned data files, wherein the second training dataset includes the initial training dataset; and” as “obtaining an evaluation set for a first type of training data and a second type of training data from a first class-balanced random subset of the training data samples from the memory and a first candidate set from a second class-balanced random subset of the training data samples from the memory excluding any training data included in the second type of training data” [¶0023]
“executing a second training pass on the machine learning model, wherein the second training pass trains the machine learning model using the second training dataset.” as “obtaining an evaluation set for a third type of training data from the first class-balanced random subset of the training data samples from the memory and a second candidate set from a randomly selected subset of the training data samples from the memory and the new images from the input batch, wherein a size of the second candidate set corresponds to a number of the new images in addition to a number of a size of the randomly selected training data samples from the memory,” [¶0023]
Regarding claim 4, SHIM teaches “wherein the targeted sampling process comprises: arranging data files in the remaining dataset in increasing order of predictive confidence based on the initial training pass.” as “MIR 403 chooses replay samples whose loss most increases after a current task update. MIR403 is a recently proposed method aiming to improve the MemoryRetrieval strategy. ” [¶0081]
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 2-3 and 7-8 is/are rejected under 35 U.S.C. 103 as being unpatentable over SHIM et al. [US 2021/0383158 A1] in view of Jin et al. [US 2020/0311557 A1].
Claim 2 is rejected over SHIM and Jin.
SHIM does not explicitly teach setting a benchmark accuracy, wherein the benchmark accuracy is a desired level of predictive accuracy of the machine learning model.
However, Jin teaches “setting a benchmark accuracy, wherein the benchmark accuracy is a desired level of predictive accuracy of the machine learning model.” as “The predefined acceptability criterion can include for example, a predefined benchmark or threshold correspondence value. For instance, the predefined acceptability criterion can include a minimum Z-score, a maximum degree of deviation, a minimum percentage of correspondence value, a maximum degree of separation, or the like. In accordance with these embodiments, the target data acceptability component 108 can determine whether the target data set 126 is within the scope of the training data set 124 or otherwise exhibits an acceptable degree of correspondence with the training data set 124 based on whether the degree of correspondence measurement data meets the acceptability criterion.” [¶0045]
SHIM and Jin are analogous arts because they teach training models and machine learning.
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of SHIM and Jin before him/her, to modify the teachings of SHIM to include the teachings of Jin with the motivation of facilitate improving the accuracy of results generated based on application of DNN models (e.g., target neural network model) in the field while minimizing errors and downstream effects of inaccurate inferences generated based on DNN model application to new data sets. [Jin, ¶0035]
Claim 3 is rejected over SHIM and Jin.
SHIM does not explicitly teach determining, after the initial training pass, that the predictive accuracy of the machine learning model is less than the benchmark accuracy.
However, Jin teaches “determining, after the initial training pass, that the predictive accuracy of the machine learning model is less than the benchmark accuracy.” as “if the data evaluation process results in a determination that the target data set is associated with a low confidence score/level that fails to exceed a minimum threshold such that the target data set can be considered outside the scope of the training data set, the disclosed techniques can declare the target data set as inapplicable to the DNN model and forgo proceeding to the model scope evaluation step.” [¶0027]
Claim 7 is rejected over SHIM and Jin.
SHIM teaches “executing the targeted sampling process a number of additional iterations, wherein each iteration of the number of additional iterations determines a new number of least learned documents, wherein each new number of least learned documents is included in a new training dataset, and wherein each new training set is used in an additional training pass to train the machine learning model.” as “ hyperparameters refer to parameters which are set before learning in a machine learning algorithm, and include a learning rate, a number of iterations, a mini-batch size, an initialization function, and the like.” [¶0045]
Claim 8 is rejected over SHIM and Jin.
SHIM does not explicitly teach wherein the number of additional iterations cause the predictive accuracy of the machine learning model to equal or exceed the benchmark accuracy.
However, Jin teaches “wherein the number of additional iterations cause the predictive accuracy of the machine learning model to equal or exceed the benchmark accuracy.” as “if the data evaluation process results in a determination that the target data set is associated with a confidence score/level exceeding a minimum threshold such that the target data set can be considered within the scope of the training data set, the disclosed techniques can further proceed with the model scope evaluation.” [¶0027]
Claim(s) 5 is/are rejected under 35 U.S.C. 103 as being unpatentable over SHIM et al. [US 2021/0383158 A1] in view of Li et al. [US 2024/0412868 A1].
Claim 5 is rejected over SHIM and Li.
SHIM does not explicitly teach wherein the targeted sampling process comprises: generating a quartile plot, wherein the quartile plot plots data points that represent the data files in the remaining dataset, wherein the quartile plot defines a first quartile, and wherein datapoints in the first quartile represent the number of least learned data files.
However, Li teaches “wherein the targeted sampling process comprises: generating a quartile plot, wherein the quartile plot plots data points that represent the data files in the remaining dataset, wherein the quartile plot defines a first quartile, and wherein datapoints in the first quartile represent the number of least learned data files.” as “A box plot displays the quartile range of the data, and observations that exceed the upper quartile plus 1.5 times the interquartile range can be considered as outliers' cutoff value. [0182] 2) Modified Z-Score” [¶0180]
SHIM and Li are analogous arts because they teach training models and machine learning.
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of SHIM and Li before him/her, to modify the teachings of SHIM to include the teachings of Li with the motivation of significant advantages including their non-invasive nature, automation, and relatively low cost compared with many other clinical detection methods. [Li, ¶0059]
Claim(s) 6 is/are rejected under 35 U.S.C. 103 as being unpatentable over SHIM et al. [US 2021/0383158 A1] in view of Li et al. [US 2024/0412868 A1] and in further view of Jin et al. [US 2020/0311557 A1].
Claim 5 is rejected over SHIM, Li and Jin.
The combination of SHIM and Li does not explicitly teach wherein the number of least learned data files is limited to one of a predetermined number of data files and a percentage of the number of data files in the remaining dataset.
However, Jin teaches “wherein the number of least learned data files is limited to one of a predetermined number of data files and a percentage of the number of data files in the remaining dataset.” as “ the second confidence score can increase as the second degree of correspondence increases. In one or more embodiments, based a determination that the target data set 126 is outside the scope of the target neural network model 128 and/or association of the target data set 126 with an unacceptable confidence score (e.g., relative to a minimum confidence score), the model acceptability component 118 can prevent application of the target neural network model 128 to the target data set 126 or otherwise facilitate weighting results generated based on application of the target neural network model 128 to the target data set 126 accordingly.” [¶0061]
SHIM, Li and Jin are analogous arts because they teach training models and machine learning.
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of SHIM, Li and Jin before him/her, to modify the teachings of combination of SHIM and Li to include the teachings of Jin with the motivation of facilitate improving the accuracy of results generated based on application of DNN models (e.g., target neural network model) in the field while minimizing errors and downstream effects of inaccurate inferences generated based on DNN model application to new data sets. [Jin, ¶0035]
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MASUD K KHAN whose telephone number is (571)270-0606. The examiner can normally be reached Monday-Friday (8am-5pm).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Hosain Alam can be reached at (571) 272-3978. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MASUD K KHAN/ Primary Examiner, Art Unit 2132