DETAILED ACTION
This communication is in response to the amendment filed 6/5/26 in which claim 1 was amended, and claims 18-20 were newly presented. Claims 1-3 and 6-20 are pending. Claims 4-5 were previously canceled.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 6/5/26 has been entered.
Response to Arguments
Applicant argues:
The Examiner relies on Xu paragraph 64 to assert that the determination of a mislabeled state is taught by the prior art. However, a substantive review of Xu reveals a fundamental mismatch in the scale of analysis:
Xu (Macro-Level): Xu's logic focuses on the global statistical reliability of an entire dataset. Specifically, Xu teaches evaluating whether a "scoring split" of the original dataset meets a threshold to determine if the dataset as a whole is suitable for model updates. Xu focuses on "reject rates" to manage the quality of the through-the-door population, not the validity of any specific, individual data point.
Amended claim 1 (Micro-Level): Claim 1, as amended, requires the processor to count misidentifications for each of the sample data. The claimed invention tracks the behavioral consistency of a specific, individual sample across multiple iterations. Only when an individual sample consistently results in an estimated label that contradicts its original label is that specific sample determined to be in a mislabeled state.
Under the Broadest Reasonable Interpretation (BRI) standard, it is unreasonable to equate the global dataset evaluation of Xu with the per-sample diagnostic mechanism of the present invention. Combining Yu's model-level evaluation with Xu's dataset-level statistics would not lead to a system that diagnoses the labeling state of an individual sample based on its specific accumulated count.
Applicant’s arguments with respect to claim 1 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or non-obviousness.
Claims 1, 7, 9, 18, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Yu (US 9,607,272 B1; patented Mar. 28, 2017) in view of Brodley, Carla E., and Mark A. Friedl. "Identifying mislabeled training data." Journal of artificial intelligence research 11 (1999): 131-167 (“Brodley”).
Regarding claim 1, Yu discloses [a] data analysis device that constructs a machine learning model based on pieces of labeled teacher data for a plurality of sample data and identifies and labels an unknown sample data using the machine learning model, the data analysis device comprising: (Yu 3:42-57 (“The training documents may then be used to train a classification model for the predictive coding system. Once the classification model has been trained (e.g., to generate a first trained classification model), the effectiveness of the predictive coding system can be determined for a set of validation documents selected from the corpus of electronic discovery documents. The effectiveness of the predictive coding system can be based on the quality of the trained classification model once the classification model has been trained, and can be determined by comparing a predictive coding system classification for each validation document and a user classification for each validation document. Therefore, the quality of the training documents is crucial to the quality and effectiveness of the trained classification model and the effectiveness of the predictive coding system that uses the trained classification model.”))
a memory; and a processor, (Yu 16:10-18 (“The exemplary computer system 500 includes a processing device (processor) 502, a main memory 504 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) such as synchronous DRAM (SDRAM), double data rate (DDR SDRAM), or DRAM (RDRAM), etc.), a static memory 506 (e.g., flash memory, static random access memory (SRAM), etc.), and a data storage device 518, which communicate with each other via a bus 530.”))
wherein the memory stores pieces of teacher data to which a correct label is labeled, the pieces of teacher data including first model construction data, second model construction data, and model verification data, the model verification data being labeled with a first label, (Yu 4:42-54 (“The data in the electronic discovery documents data repository 120 can include a corpus of electronic discovery documents that need to be reviewed and classified. Examples of electronic discovery documents can include, and are not limited to, electronic discovery documents which have been divided into a set of training documents that have been selected by an administrator (document reviewer, etc.), a set of validation documents that have been selected by an administrator (document reviewer, etc.), an unlabeled remainder of electronic discovery documents that need to be classified or labeled, and any other electronically stored information that can be associated with electronic discovery documents, etc.”))
wherein the processor performs:
a first process of constructing the machine learning model using the first model construction data; (Yu 5:7-36 (“During operation of system 100, a predictive coding system 110 can train a (untrained) classification model 140, to generate a (first) trained classification model 145. To train the classification model 140, an initial training set of documents is needed by the predictive coding system 110. To generate the initial training set, the predictive coding system defines a set of query/search terms based on a topic of interest. The topic of interest can be provided by a user or administrator of the predictive coding system 110. The predictive coding system 110 can perform a search (keyword search and/or concept search) with the (stemmed) terms on the electronic discovery documents in electronic discovery documents data repository 120 and can return documents based on the search. In one embodiment, the predictive coding system 110 selects all documents returned by the search as training documents. In an alternate embodiment, the predictive coding system 110 selects a predetermined number of documents returned by the search as the training documents. For example, the predictive coding system 110 can select 1000 random documents from the documents returned by the search as training documents.
The predictive coding system 110 can cause a user interface to be presented to an administrator or reviewer via client device 102A-102N. The user interface can present the training documents to the administrator or reviewer and request one or more inputs from the administrator or reviewer on the client device 102A-102N over network 104, such as a label or a classification for each training document (e.g., confidential, not confidential, relevant, not relevant, privileged, not privileged, responsive, not responsive, etc.).”))
a second process of applying the machine learning model to the model verification data to label an estimated label; (Yu 5:45-52 (“The validation documents 160 can include a set of validation documents used to validate the trained classification model 145 in the predictive coding system 110. The trained classification model 145 can classify the set of validation documents in the predictive coding system 110. The predictive coding system 110 can further present the set of validation documents to an administrator or reviewer via client device 102A-102N over network 104.”))
a third process of retraining the machine learning model using the second model construction data; and (Yu 5:66-6:4 (“In one embodiment, the predictive coding system 110 includes a training data generation module 130. The training data generation module 130 can incrementally enhance the training set for predictive coding in one or more iterations, and retrain the classification model with the enhanced training set in each iteration.”); Yu 6:58-62 (“In one embodiment, the predictive coding system 110 retrains the classification model 140 based on only the set of updated training documents in training documents 150, and may not be based on any previous training of the classification model 140. For example, if the classification model was previously trained using documents A1, . . . , A1000, and the updated set of training documents contains documents A1, . . . , A1000 and B1, . . . , B100, the classification model is retrained using documents A1, . . . , A1000 and B1, . . . , B100, without the use of any previous version of the trained classification model that was built using documents A1, . . . , A1000.”))
a fourth process of repeating the second process after the third process,
wherein a combination of the third process and the fourth process is performed at least once, (Yu 17-36 (“If the training data generation module 130 determines that the classification model 140 should be retrained, the training data generation module 130 can generate training data from a subset of the unlabeled documents in unlabeled documents 170 and provide the training data to the predictive coding system 110 to cause the classification model 140 to be retrained by generating a new trained classification model 145 (e.g., a second trained classification model) that has an improved effectiveness than the previous trained classification model 145 (e.g., first trained classification model). In each iteration, the training data generation module 130 can generate the training data by selecting a predetermined number of additional documents from unlabeled documents 170 as training data. The training data generation module 130 can select each of the additional documents by randomly choosing a group of unlabeled documents from unlabeled documents 170, calculating a score for each document in the chosen group of unlabeled documents, and selecting the document in the chosen group of unlabeled documents with the lowest score.”))
wherein the processor counts a number of misidentifications in which the estimated label does not coincide with the first label for each of the plurality of sample data, after the combination of the third process and the fourth process is performed at least once, and (Yu 6:4-8 (“The training data generation module 130 can determine an effectiveness of the trained classification model 145 and determine whether the classification model 140 (e.g., untrained classification model) should be retrained.”); Yu 10:40-43 (“For example, the effectiveness measure can be a precision of the trained classification model, a recall of the trained classification model, an F-measure of the trained classification model, etc.”); Yu 10:56-63 (“The precision for the trained classification model can be defined as: precision=TP/(TP+FP), where TP is the number of true positives in the set of validation documents, and FP is the number of false positives in the set of validation documents.”)).
Yu does not expressly disclose wherein the processor determines that one of the plurality of the sample data is in a mislabeled state when the number of misidentifications obtained from the sample data, or a misidentification rate derived therefrom, is equal to or higher than a threshold (but see Brodley pp. 136-137 (Section 3.2) (“In filtering, an ensemble classifier detects mislabeled instances by constructing a set of base-level detectors (classifiers) and then using their classification errors to identify mislabeled instances. The general approach is to tag an instance as mislabeled if x of the m base-level classifiers cannot classify it correctly. In this work we examine both majority and consensus filters. A majority vote filter tags an instance as mislabeled if more than half of the m base level classifiers classify it incorrectly.”)).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Yu to incorporate the teachings of Brodley to identify mislabeled documents by majority vote of a set of base level classifiers, at least because using multiple models (e.g., a particular model trained on different training datasets) will provide a better method for detecting mislabeled instances than collecting information from a single model. Brodley p. 136.
Regarding claim 7, Yu, in view of Brodley, discloses the invention of claim 1 as discussed above. Yu further discloses wherein the processor uses a support vector machine as a machine learning technique (Yu 5:37-44 (“The predictive coding system 110 can add each of the documents labeled by the user to a set of training documents, such as training documents 150 in the electronic discovery documents data repository 120. The predictive coding system 110 can train an untrained classification model, such as classification model 140 (e.g., an SVM model) using the set of training documents in training documents 150 to generate a trained classification model 145.”)).
Regarding claim 9, Yu, in view of Brodley, discloses the invention of claim 1 as discussed above. Yu further discloses wherein the processor uses a linear discrimination method as a machine learning technique (Yu 5:37-44 (“The predictive coding system 110 can add each of the documents labeled by the user to a set of training documents, such as training documents 150 in the electronic discovery documents data repository 120. The predictive coding system 110 can train an untrained classification model, such as classification model 140 (e.g., an SVM model) using the set of training documents in training documents 150 to generate a trained classification model 145.”)).
Regarding claim 18, Yu, in view of Brodley, discloses the invention of claim 1 as discussed above. Yu does not expressly disclose wherein the pieces of teacher data are partitioned into the first model construction data, the second model construction data, and the model verification data, and wherein the processor performs the first process through the fourth process after the memory stores the pieces of teacher data (but see Brodley p. 136 (“This section describes a general procedure for identifying mislabeled instances in a training set. The first step is to identify candidate instances by using m learning algorithms (called filter algorithms) to tag instances as correctly or incorrectly labeled. To this end, a n fold cross-validation is performed over the training data. For each of the n parts, the m algorithms are trained on the other n 1 parts. The m resulting classifiers are then used to tag each instance in the excluded part as either correct or mislabeled.”)).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Yu to incorporate the teachings of Brodley to perform the n-fold cross-validation on n folds of the teacher data, at least because the trained models resulting from the n fold cross validation provide better results than a single trained model trained on a single dataset.
Regarding claim 20, Yu, in view of Brodley, discloses the invention of claim 1 as discussed above. Yu does not expressly disclose wherein upon determining that the one of the plurality of sample data is mislabeled, the processor either (i) removes the one of the plurality of sample data that is mislabeled from the pieces of teacher data or (ii) relabels the preexisting first label of the any one of the plurality of sample data that is mislabeled with a corrected label (but see Brodley p. 136 (“At the end of the n-fold cross-validation each instance in the training data has been tagged. Using this information, the second step is to form a classifier using a new version of the training data for which all of the instances identified as mislabeled are removed. Filtering can be based on one or more of the m base level classifiers' tags. The filtered set of training instances is provided as input to the final learning algorithm.”)).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Yu to incorporate the teachings of Brodley to perform the n-fold cross-validation on n folds of the teacher data, at least because the trained models resulting from the n fold cross validation provide better results than a single trained model trained on a single dataset.
Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Yu and Brodley as applied to claim 1 above, and further in view of Nicholson, Bryce, et al. "Label noise correction methods." 2015 IEEE International conference on data science and advanced analytics (DSAA). IEEE, 2015 (“Nicholson”).
Regarding claim 6, Yu, in view of Brodley, discloses the invention of claim 1 as discussed above. Yu does not expressly disclose wherein the processor uses random forest as a machine learning technique (but see Nicholson I Introduction (“Another group of methods that implicitly handle label noise, using classification methods that are robust to mislabeled data, are bagging, boosting, and random forests [8], Bayesian approaches—e.g., [9]—among many others”)).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Yu to incorporate the teachings of Nicholson to use random forests as the classification model, at least because random forests are robust to mislabeled data.
Claims 8 and 10 are rejected under 35 U.S.C. 103 as being unpatentable over Yu and Brodley as applied to claim 1 above, and further in view of Ben-Hur (US 2010/0205124 A1; published Aug. 12, 2010).
Regarding claim 8, Yu, in view of Brodley, discloses the invention of claim 1 as discussed above. Yu does not expressly disclose wherein the processor uses a neural network as a machine learning technique (but see Ben-Hur ¶ 4 (“Machine-learning approaches, which include neural networks, hidden Markov models, belief networks and kernel-based classifiers such as support vector machines, are ideally suited for domains characterized by the existence of large amounts of data, noisy patterns and the absence of general theories.”)).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Yu to incorporate the teachings of Ben-Hur to use a neural network as the classification model, at least because a neural network is a type of classification model like SVM.
Regarding claim 10, Yu, in view of Brodley, discloses the invention of claim 1 as discussed above. Yu does not expressly disclose wherein the processor uses a non-linear discrimination method as a machine learning technique (but see Ben-Hur ¶ 4 (“Machine-learning approaches, which include neural networks, hidden Markov models, belief networks and kernel-based classifiers such as support vector machines, are ideally suited for domains characterized by the existence of large amounts of data, noisy patterns and the absence of general theories.”)).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Yu to incorporate the teachings of Ben-Hur to use a neural network as the classification model, at least because a neural network is a type of classification model like SVM.
Claim 2 is rejected under 35 U.S.C. 103 as being unpatentable over Yu and Brodley as applied to claim 1 above, and further in view of Karlov (US 2003/0065535 A1; published Apr. 3, 2003).
Regarding claim 2, Yu, in view of Brodley, discloses the invention of claim 1 as discussed above. Yu does not expressly disclose wherein the processor
removes the one of the plurality of sample data determined to be in the mislabeled state from the pieces of teacher data to generate an updated pieces of teacher data, and performs the combination of the third process and the fourth process using the updated pieces of teacher data (but see Karlov ¶ 10 (“In another aspect of the invention, a method is provided for identifying a patient disease diagnosis that appears to be mislabeled. Each patient who has contributed a data record to the clinical data, including one or more test results, will be associated with a clinical disease diagnosis. The possibility of a mislabeling is identified when a data analysis such as described above is performed and a set of probability density functions (pdf) are produced that can provide a hypothesized disease diagnosis for each patient, as well as for new patients. This analysis can identify a patient to whom one or more of the tests was administered, but for whom the disease diagnosis predicted by the inventive method is different from the clinical diagnosis assigned to that patient. If a clinical diagnosis is determined to be mislabeled, that patient's data record can be removed from consideration in performing a future iteration of the estimation technique in accordance with the invention.”)).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Yu and Xu to incorporate the teachings of Karlov to remove documents from the training data set if it is determined that the documents were mislabeled, at least because doing so would improve the accuracy of the retrained model.
Claims 11-15 are rejected under 35 U.S.C. 103 as being unpatentable over Yu and Brodley as applied to claim 1 above, and further in view of Miranda, André LB, et al. "Use of classification algorithms in noise detection and elimination" Hybrid Artificial Intelligence Systems: 4th International Conference, HAIS 2009, Salamanca, Spain, June 10-12, 2009, Proceedings 4 Springer Berlin Heidelberg, 2009 (“Miranda”).
Regarding claim 11, Yu, in view of Brodley, discloses the invention of claim 1 as discussed above. Yu does not expressly disclose wherein the processor determines that one of the pieces of teacher data having a highest misidentification rate is incorrect (but see Miranda pp. 317-319 (an ensemble method for noise elimination in classification problems in which an instance is removed from a training set if it cannot be classified correctly by all (i.e., 100%), or the majority of, the classifiers built on parts of the training set)).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Yu and Xu to incorporate the teachings of Miranda to identify mislabeled instances based on the false positive (or false negative) rate, at least because doing so would most improve the accuracy of the classification model.
Regarding claim 12, Yu, in view of Brodley, discloses the invention of claim 1 as discussed above. Yu does not expressly disclose wherein the processor determines that pieces of the teacher data as many as a number specified by a user in descending order of a misidentification rate are incorrect (but see Miranda pp. 317-319 (an ensemble method for noise elimination in classification problems in which an instance is removed from a training set if it cannot be classified correctly by all (i.e., 100%), or the majority of, the classifiers built on parts of the training set)).
Yu is combinable with Miranda for the same reasons as set forth above.
Regarding claim 13, Yu, in view of Brodley, discloses the invention of claim 1 as discussed above. Yu does not expressly disclose wherein the processor determines that one of the pieces of teacher data having a misidentification rate of 100% is incorrect (but see Miranda pp. 317-319 (an ensemble method for noise elimination in classification problems in which an instance is removed from a training set if it cannot be classified correctly by all (i.e., 100%), or the majority of, the classifiers built on parts of the training set)).
Yu is combinable with Miranda for the same reasons as set forth above.
Regarding claim 14, Yu, in view of Brodley, discloses the invention of claim 1 as discussed above. Yu does not expressly disclose wherein the processor determines that one of the pieces of teacher data whose misidentification rate is equal to or higher than a threshold set by a user is incorrect (but see Miranda pp. 317-319 (an ensemble method for noise elimination in classification problems in which an instance is removed from a training set if it cannot be classified correctly by all (i.e., 100%), or the majority of, the classifiers built on parts of the training set)).
Yu is combinable with Miranda for the same reasons as set forth above.
Regarding claim 15, Yu, in view of Brodley and Miranda, discloses the invention of claim 2 as discussed above. Yu does not expressly disclose wherein generating the updated teacher data is repeated until a misidentification rate becomes equal to or lower than a predetermined threshold (but see Miranda p. 420 (New classifiers are then trained using the new training data sets, and their accuracies are again evaluated using the validation folds. If the new performance recorded is better than that obtained previously, the preprocessing cycle is repeated. Pre-processing stops when a performance degradation occurs.)).
Yu is combinable with Miranda for the same reasons as set forth above.
Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over Yu and Brodley as applied to claim 1 above, and further in view of Zhang, Y., Noise Tolerant Data Mining (2008) (Ph.D. dissertation, The University of Vermont), available at https://citeseerx.ist.psu.edu/document?repid=rep1&type=pdf&doi=56948dc4ff87bdc4605d5eb4d06085be86aba12b (“Zhang”).
Regarding claim 3, Yu, in view of Brodley, discloses the invention of claim 1 as discussed above. Yu does not expressly disclose wherein the first model construction data, the second model construction data, and the model verification data are obtained by dividing the pieces of teacher data in random (but see Zhang pp. 16-17 (a well-known method for handling noise is to detect and discard the instances which are subject to noise according to certain evaluation methods: “The essential idea is using m learning algorithms to filter out the instances that are prone to labeling errors. The method first splits the training data into n parts, like an n-fold cross-validation. For each of the n parts, the m filtering algorithms are trained on the other n−1 parts. The m resulting classifiers are then used to predict on instances in the excluded part, and finally decide whether the instances are correctly labeled or not. All of the instances identified as mislabeled are removed, and the filtered set of training instances is provided as the input to the final learning algorithm.”)).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Yu to incorporate the teachings of Zhang to prune the mislabeled data and retrain/retest the model using the pruned data, at least because doing so would improve the accuracy of the classifier.
Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over Yu and Brodley as applied to claim 1 above, and further in view of Drucker (US 2011/0307422 A1; published Dec. 15, 2011).
Regarding claim 16, Yu, in view of Brodley, discloses the invention of claim 1 as discussed above. Yu does not expressly disclose wherein the processor creates a table or a graph based on an identification result and displays the table or graph on a display (but see Drucker ¶ 48 (“FIG. 6 illustrates a graph 600 plotting incorrectness versus entropy using a plurality of data points. It can be seen from FIG. 6 that some of the data falls within the canonical area 510, some data falls within the unsure area 520, and some data falls within the confused area 530.”)).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Yu to incorporate the teachings of Drucker to create a graph plotting incorrectness versus entropy of the data points, at least because doing so would enable a user to obtain information from the graph based on where the data falls on the graph.
Claim 17 is rejected under 35 U.S.C. 103 as being unpatentable over Yu and Brodley as applied to claim 1 above, and further in view of Lam (US 2002/0107712 A1; published Aug. 8, 2002).
Regarding claim 17, Yu, in view of Brodley, discloses the invention of claim 1 as discussed above. Yu does not expressly disclose wherein
upon determining that any one of the plurality of sample data is mislabeled, the processor either (i) removes corresponding model verification data from the pieces of teacher data or (ii) relabels the first label (but see Lam ¶ 117 (“When validation shows that results don't meet expectations, there are three immediate options for correction. The simple option, but possibly adequate, is to get more data. Another option is to fix mislabeled data on hand, retaining the existing categorization scheme. The option that takes more cognitive effort is to rethink the categorization scheme in the light of the technology, keeping in mind the feasibility of relabeling the data.”)).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Yu and Xu to incorporate the teachings of Lam to fix the mislabeled data, at least because doing so would ensure the adequacy of the classification model in light of the latest training data. See Lam ¶ 32.
Claim 19 is rejected under 35 U.S.C. 103 as being unpatentable over Yu and Brodley as applied to claim 1 above, and further in view of Bouziane, Hafida, Belhadri Messabih, and Abdallah Chouarfia. "Profiles and majority voting-based ensemble method for protein secondary structure prediction." Evolutionary Bioinformatics 7 (2011): EBO-S7931 (“Bouziane”).
Regarding claim 19, Yu, in view of Brodley, discloses the invention of claim 1 as discussed above. Yu does not expressly disclose wherein the plurality of sample data are re-partitioned into different combinations of the first model construction data, the second model construction data, and the model verification data each time the combination of the third process and the fourth process is repeated (but see Bouziane p. 174 second column (“In order to estimate the generalization error, methods are typically tested using k-fold cross-validation, where a dataset is split into k subsets. In each step of the cross-validation, k − 1 of them are used for training and the remaining one for testing. The process is repeated k times, until all the k subsets are used once for testing. The prediction accuracy is estimated by calculating the average accuracy across all the k steps. In this study, we used the seven-fold cross-validation on both RS126 and CB513 datasets.”)).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Yu to incorporate the teachings of Bouziane to perform an iterative n fold cross-validation over n-1 parts training data, at least because doing so yields significant performance gains when compared with a single best classifier. See Bouziane Abstract.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHAHID KHAN whose telephone number is (571)270-0419. The examiner can normally be reached M-F, 9-5 est.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Usmaan Saeed can be reached at (571)272-4046. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SHAHID K KHAN/Primary Examiner, Art Unit 2146