The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1 and 9 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Moniri (Analytical Chemistry 2020, hereinafter called Moniri ‘20). In the paper Moniri ’20 teaches performing data-driven multiplexing in a single fluorescent channel using machine learning methods, by virtue of the information in the amplification curve. This new approach, referred to as amplification curve analysis (ACA), was shown using an intercalating dye (EvaGreen), reducing the cost and complexity of the assay and enabling the use of melting curve analysis for validation. As a case study, 3 carbapenem-resistant genes were multiplexed to show the impact of this approach on global challenges such as antimicrobial resistance. Figures 2A), 2C), 3B) and 4A) show the plotting of a fluorescent signal over time (the amplification curve) from a substance. The paragraph bridging pages 13139-13140 discusses data-driven multiplexing using supervised machine learning of the amplification curve. After establishing that information exists in the amplification curve using unsupervised methods, supervised learning methods can be used to exploit this information to perform multiplexing. Several machine learning algorithms exist for classification tasks such as k-nearest neighbors, support vector machines, and deep neural networks. The paper used the k-nearest neighbors algorithm, which is an intuitive nonparametric method for detecting single targets as well as single targets in the presence of multiple targets. Thus claim 1 is anticipated. Because the training of the machine learning algorithm required multiple plots, claim 9 is also anticipated.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Moniri ‘20 as applied to claim 1 above. With respect to claim 7 figure 1 shows amplification curves without a scale. Figures 2A), 2C), 3B) and 4A) show amplification curves with a scale. Figures 2A), 2C) and 3B) show a normalized fluorescence scale while figure 4A shows a scale associated with the raw amplification curve. The last full paragraph on page 13139 discusses the information available in the amplification curve data with an unsupervised machine learning algorithm (see figure 5 for visualization of the data using the unsupervised machine learning). In figure 5, the different targets fall in a different region and can therefore be distinguished automatically using statistical machine learning. Therefore, they demonstrated that even after normalizing for fluorescent intensity, the kinetic information that is encoded in the amplification curve can provide sufficient information to perform data-driven multiplexing. Moreover, it is interesting to observe that the region indicated within the dashed red circle of figure 5 shows amplification curves that do not fully plateau, and therefore are similar across the 3 targets. This suggests that the entire curve is necessary to extract sufficient kinetic information. The fact that changing the scale through normalization does not erase the kinetic information of the amplification curve and that Moniri '20 uses different scales or the lack thereof in the various displays in the figures of the paper would point to the display of a scale as a non-critical feature as long as the information being extracted to distinguish the targets remains in the amplification curve data presented to the statistical machine learning algorithm. Thus although each of the amplification curves in the above figures includes a set of axes with either no defined scale or different defined scales used in the figures, it would have been an obvious modification at the time the application was filed to remove a defined scale from the display of the amplification curve as shown in figure 1 because as shown by Moniri ’20, display of a scale is not a noncritical feature of the machine learning algorithm being able to extract the kinetic data from amplification curve fluorescence.
Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Moniri ‘20 as applied to claim 1 above, and further in view of Cascio (Applied Sciences 2019) or Hu (Analytical Methods 2019). With respect to claim 8, the paragraph bridging pages 13139-13140 of Moniri ’20 teaches that several machine learning algorithms exist for classification tasks such as k-nearest neighbors, support vector machines, and deep neural networks. Moniri '20 did not teach the use of a convolutional neural network for that purpose.
In the paper Cascio teaches an automatic system for fluorescence intensity classification to support the autoimmune diagnostics in HEp-2 image analysis. The system is based on the use of a pre-trained convolutional neural network (CNN) to extract features and a support vector machine (SVM) classifier for the positive or negative association. Page 2 of the paper describes different prior techniques used to characterize fluorescence intensity images. The second to last paragraph on the page describes recent scientific research on pattern recognition and teaches that Convolutional Neural Networks (CNNs) have been proven to be efficient and reliable models to achieve remarkable performance for image classification and object detection tasks. Moreover, it has been demonstrated that pre-trained CNN architectures can play an important role in terms of features extractors, and allow high classification performance. Page 7 of the paper teaches that the problem of automatic classification of fluorescence intensity in HEp-2 images was been addressed. To that end, CNN pre-trained networks were analyzed as feature pullers, combined with the traditional SVM (support vector machine) classifier. In particular, several more well-known pre-trained CNN architectures were used. For each architecture used, the different layers were analyzed in order to find the most discriminating one for the classification problem addressed. In order to identify the set of features having the highest classification power for the characterization of the fluorescence intensity, the analysis carried out for the different configurations allowed the identification, in terms of network and layer, of the best-performing solution. A comparison of the performances was presented with other recent state-of-the-art methods that highlight the quality of the proposed system and the very promising capabilities in discriminating the positive and negative images of the IIF test. The pre-trained network number and the relative layers used as features extractors, allowed them to obtain better fluorescence intensity classification performance than those obtained from other recent state-of-the-art methods.
In the paper Hu teaches a method based on a Mask R-convolutional neural network (CNN) model for processing digital polymerase chain reaction (dPCR) images. A digital polymerase chain reaction (dPCR) using fluorescence images for collecting quantitative information needs efficient software tools to automate the image analysis process. However, due to the broad range of fluorescence image characteristics, such as the uneven fluorescence intensity distribution, irregular structures of microarrays, and diverse and unpredictable morphologies of microchambers, existing tools fail to extract signals from these images properly, thus posing challenges for the improvement of detection accuracy. In this paper, a deep learning method based on the Mask R-CNN model was used for image processing to achieve more accurate quantification of nucleic acids in both microarray and droplet dPCR. This Mask R-CNN based method uses massive dPCR fluorescence image data to train a model that has the ability to recognize target signals in dPCR images precisely and automatically, regardless of the non-uniform luminosity or spot impurities appearing in dPCR images. When the Mask R-CNN model is used to process images with non-uniform luminosity, the true positive rate of this model can reach 97.56%. By contrast, the true positive rate of threshold segmentation is only 68.29%. As for dealing with images with spot impurities, which caused a 7.25% fault of the positive points in the threshold segmentation method, the error rate decreased to zero using the Mask R-CNN model. In addition, the labeling of annotated pictures is time-consuming. Therefore, an iteration method was used in the annotation stage to reduce time spent on labeling. With proper modifications, this method has great potential to serve as an alternative to achieve a highly efficient fluorescence image process for dPCR or other digital assays. The last paragraph of the left column on page 3411 teaches that they trained two different models to process chip-based and droplet PCR images. Once the training was completed, it took only a few seconds for each test to discover all the positive microchambers exactly using the same type of pictures. As a proof of concept this classification method was conducted to deal with non-uniform illumination of droplets or chip-based images and the performance of the method was much better than that of the threshold segmentation method. As for the possible impurities retained on fluorescence images that may affect the final quantification results, the method can filter out such interfering signal efficiently without the need for coordinate delimitation. In addition, a similar iteration method is employed to reduce the labeling time and this process can be finished more quickly. Finally, once the optimized parameters are settled, the model can be applied to process most images of the same type of dPCR chip without any manual modifications, thus ensuring the authenticity and reproducibility of the experimental results. Therefore, this classification can simplify the analysis of dPCR fluorescence images and improve the accuracy of quantification results.
It would have been obvious to one of ordinary skill in the art at the time the application was filed to modify the method of Moniri ’20 to use a convolutional neural network as the machine learning model as taught by Cascio or Hu because of its ability to handle fluorescent image analysis in general as taught by Cascio or to handle dPCR fluorescent image analysis as taught by Hu in a manner that was better than other methods working on/with the same problem(s).
Claims 2, 6 and 10 are rejected under 35 U.S.C. 103 as being unpatentable over Moniri ‘20 as applied to claim 1 above, and further in view of Vess (US 2005/0118620). With respect to claims 2, 6 and 10, the second through fourth paragraphs of page 13136 teach that the PCR and real-time dPCR amplification was performed with a SsoFast EvaGreen Supermix with Low ROX on either a LightCycler 96 Real-Time PCR System for the PCR or qdPCR 37K digital chips and Fluidigm’s Biomark HD system to perform the dPCR experiments. These PCR/dPCR instruments/systems would have a memory and processor in electronic communication with the memory to control the instrument/system and/or produce the amplification plots shown in figures 1, 2A), 2C), 3B) and 4A) based on software stored in the memory. The data analysis including the software that was developed and/or used to process the data was also described. Moniri ’20 does not teach generating a second plot of a second fluorescence signal measured from the substance, wherein detecting the target molecule is further based on the second plot.
In the patent publication Vess teaches methods, apparatus, and systems, including computer program products, implement techniques for determining an amount of target nucleic acid ("target") in a sample. Signal data is received for a plurality of cycles of an amplification experiment performed on the target and a standard nucleic acid ("standard"). The signal data includes a series of signal values indicating a quantity of standard present during cycles of the standard amplification, and a series of signal values indicating a quantity of target present during cycles of the target amplification. A target growth curve value is defined using the target signal values and a standard growth curve value is defined using the standard signal values. An initial amount of the target is calculated according to a calibration equation using an initial amount of the standard, and the target and standard growth curve values, where the calibration equation is a nonlinear equation (see at least paragraph [0014]). Figure 2G shows the plotting of 2 growth curves. Paragraph [0017] teaches that the calibration equation can be derived from target growth curve values and standard growth curve values from a series of amplification experiments performed on known quantities of the target nucleic acid and a known quantity of the standard nucleic acid. The calibration equation can relate the initial amount of the target nucleic acid to the known quantity of the standard nucleic acid using the target and standard growth curve values. The calibration equation can be defined by plotting a correlate of the initial amount of the target nucleic acid against a function of the known quantity of the standard nucleic acid and the target and standard growth curve values to produce a calibration plot, and fitting a curve to the calibration plot. The function of the known quantity of the standard nucleic acid and the target and standard growth curve values can be a function of the difference between the target and standard growth curve values. Paragraphs [0051]-[0053] teach that monitoring the signal indicative of the concentration of target nucleic acid produces a set of signal data for the target (120), and monitoring the signal indicative of the concentration of standard nucleic acid produces a set of signal data for the standard (130). The target data and the standard data can be stored, such as in a database. Each set of signal data (120, 130) can represent a growth curve for the amplification of a nucleic acid, and is processed separately but similarly. The signal data 120, 130 can be plotted as a function of the amplification cycle number to generate a growth curve. Amplification of a nucleic acid by PCR is an exponential process, so a PCR growth curve is typically and generally exponential in form, at least during the initial stages of a PCR. Figures 2A-J illustrate examples of growth curves that can be obtained in PCR experiments according to the method of figure 1. With specific reference to claim 10, paragraph [0146] teaches that the invention can be implemented in digital electronic circuitry, or in computer hardware, firmware, software, or in combinations of them. Apparatus of the invention can be implemented in a computer program product tangibly embodied in a machine-readable storage device for execution by a programmable processor; and method steps of the invention can be performed by a programmable processor executing a program of instructions to perform functions of the invention by operating on input data and generating output. The invention can be implemented advantageously in one or more computer programs that are executable on a programmable system including at least one programmable processor coupled to receive data and instructions from, and to transmit data and instructions to, a data storage system, at least one input device, and at least one output device. Suitable processors include, by way of example, both general and special purpose microprocessors. Generally, a processor will receive instructions and data from a read-only memory and/or a random access memory. The essential elements of a computer are a processor for executing instructions and a memory. Generally, a computer will include one or more mass storage devices for storing data files; such devices include magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and optical disks. Storage devices suitable for tangibly embodying computer program instructions and data include all forms of non-volatile memory, including by way of example semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM disks.
With respect to claims 2 and 6, it would have been obvious to one of ordinary skill in the art at the time the application was filed to modify the Moniri ’20 method by providing a second standard nucleic acid as taught by Vess and monitor/plot its fluorescent growth curve and use it to form a calibration curve to analyze the target nucleic acid as taught by Vess because of the ability to quantitate the amount of target in a sample as taught by Vess. With respect to claim 10, it would have been obvious to one of ordinary skill in the art at the time the application was filed to implement the method of Moniri ’20 on a coupled memory and processor as taught by Vess because of its use in PCR instruments to both control the instrument and carry out data analysis as taught by Vess.
Claims 3-5 and 11-12 are rejected under 35 U.S.C. 103 as being unpatentable over Moniri ‘20 in view of Vess as applied to claims 1-2 and 10 above, and further in view of Harada (US 2018/0163260). With respect to these claims Moniri ’20 does not teach that the two plats are shifted to spatially separate regions.
In the patent publication Harada teaches a gene polymorphism analysis device for determining the allele mating type of a gene polymorphism on the basis of the first fluorescence intensity change over time and the second fluorescence intensity change over time measured on the first allele and second allele that constitute a gene polymorphism of a target DNA. The determination of the allele mating type of the gene polymorphism is performed on samples arranged in M rows and N columns, and the determination results and graphs are displayed on a display device. The graphs show the first fluorescence intensity change over time and the second fluorescence intensity change over time of each sample. The determination results and the graphs are arranged in M rows and N columns so as to correspond to the layout including M rows and N columns of the samples. Figure 2 is a figure showing an example of a display screen showing 96 determination results obtained by the device of figure 1. Paragraph [0027] teaches that the data display control unit 11c of figure 1 controls the display of the following display units on the display device 13: SNP plotted display unit (SNP image) for showing the fluorescence signal intensity of the first allele by the horizontal axis and the fluorescence signal intensity derived from the second allele by the vertical axis, a waveform data display unit (reaction curve image) for displaying in piles a plurality of waveform data (reaction curves) illustrating the fluorescence signal intensity change over time after the start of measurement, and a display unit (graph image) in which the determination result and the waveform data are individually arranged (spatially separated) so as to become the same 8 rows 12 columns as that of the sample arrangement of a plurality of wells (the spatial separation is both vertical and horizonal). Figure 2 is a figure showing an example of a screen on which 32 determination results are displayed by the gene polymorphism analysis device 1. On the right half screen (graph image), 96 cells are arranged and displayed so as to become the same 8 rows and 12 columns as that of the sample arrangement of a plurality of wells. And one waveform data is displayed in a 4×8 cell into which the measurement information has been registered, respectively. As for the waveform data, the fluorescence signal intensity change over time derived from the first allele is shown by a solid line on the same coordinates, and the fluorescence signal intensity change over time derived from the second allele is shown by a dashed line. The vertical axis of the coordinates shows the fluorescence signal intensity (arbitrary unit), and the horizontal axis shows the elapsed time (the number of measurement cycles) after the start of measurement. The minimum and maximum values of the vertical axis and the horizontal axis are the same in all wells and in the initial display state of the waveform data display unit (reaction curve image) on the lower left screen. At the upper left of the waveform data, number “1” is displayed when the determination result is “allele 1,” number “12” is displayed when the determination result is “allele 1 & 2,” number “2” is displayed when the determination result is “allele 2,” “ND” is displayed when the determination result is “unknown,” and “NC” is displayed when the determination result is “NC.”
It would have been obvious to one of ordinary skill in the art at the time the application was filed to display the fluorescence data of Moniri ’20 in spatially different locations as shown by Harada because of the ability to view each time dependent change in the fluorescence by itself either separated vertically and/or horizontally rather than as a composite graph of the fluorescence change to separate them in a manner representing the structure of the device in which the amplifications were performed as taught by Harada or the ability or to separate them based on the different reactions being measure by each time dependent graph as resulted through the display described by Harada.
Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over Moniri ’20 as described above in view of Ullerich (Journal of Laboratory Medicine 2017). While Moniri ’20 teaches the use of a LightCycler 96 Real-Time PCR System for the PCR or qdPCR 37K digital chips and Fluidigm’s Biomark HD system to perform the dPCR experiments, Moniri ’20 does not teach that they operate to produce pulse-controlled amplification.
In the paper Ullerich presents exemplary approaches on how to improve the most time-limiting part of polymerase chain reactions, the heating and cooling steps. A new technology, Laser PCR®, promises to deliver a solution fast enough for point-of-care applications. In the section bridging pages 239-240, Ullerich teaches that polymerase chain reaction (PCR) with its amplification, detection, and if required, quantification of the target DNA and/or RNA creates a challenge for point-of-care testing (POCT). The challenge with PCR lies in its more complex handling as compared with dipstick assays, extensive costs (primarily for the additional work-flow steps and the costly enzyme needed) and lengthy turn-around-times of typically an hour or more. Especially for POCT, it is a necessity to have fast and easy-to-use tests that show very accurate results in <15 min. Integral to PCR are the applied 30–40 cycles of heating and cooling of the reaction liquid. These thermal cycling steps allow for DNA denaturation, primer annealing, and DNA elongation by the DNA polymerase. Performing temperature ramps between high and low temperatures is the most time-limiting aspect of a PCR. Besides accounting for the cycling of the temperature of the bulk reaction liquid, the thermalization of the heating and cooling element and the contacted walls of the reaction vessel also drag the reaction time. At that time, only a few instruments were available for near-patient molecular testing, and these mostly still did not deliver on ease of use, time-to-result and/or sensitivity. However, an increasing number of technological innovations were pushing the boundaries of PCR in terms of speed and usability for POCT. The paragraph bridging pages 240-241 describes/lists several isothermal amplification technologies that did not rely on the time-consuming heating and cooling steps. Page 242 teaches that laser PCR developed by GNA Biosolutions GmbH operates on the principle of pulse controlled amplification (PCA) of nucleic acids on functionalized nanoparticles. Laser-activated PCA triggers nearly instantaneous heating and cooling ramps locally, while the bulk of the reaction is held at a constant temperature. Laser irradiation of colloidal nanoparticles (preferably made of gold) enables the local and ultra-fast heating of the homogeneously dispersed nanoparticles, given that the laser wavelength and the plasmonic absorption properties of the gold nanoparticles match. The laser heats up the nanoparticles (in picomolar concentrations) in short pulses (e.g. microseconds), thereby keeping the bulk reaction temperature largely unchanged during nucleic acid amplification. By effectively restricting thermal cycling to the nanoparticle, with the primers attached to the nanoparticle’s surface, Laser PCR can achieve heating and cooling cycles a million times faster than conventional PCR (see Figure 1). Increased speed should not of course come at the expense sensitivity. Laser PCR can generate pulse controlled amplification of as few as 10 spiked target DNA copies in under 10 min in its current state of development (see Figure 2 showing curves/graphs similar to those of Moniri '20). Laser PCR was combined with reporter probes and real-time fluorescence detection. Existing commercially available hydrolysis probe assays, and other probe-based assay formats, can be easily combined with Laser PCR. Importantly, the sensitivity observed in target-spiked Laser PCR experiments can be seen also in large-volume reactions (see Figure 3). For example, as few as 10 target copies of genomic DNA purified from MRSA were amplified and detected in real time in a reaction volume of 100 μL (see Figure 3). This is especially important in the POC setting where samples may require dilution and flexibility of processing. Further optimization regarding base line corrections, and means for quantification, will be implemented in the future. Since the technology uses probe-based detection, multiplexing is possible by adding more fluorescence channels, but not shown at this stage of development.
It would have been obvious to one of ordinary skill in the art at the time the application was filed to adapt the non-transitory tangible computer-readable medium comprising instructions when executed cause a processor of an electronic device to determine a plot of a fluorescence signal measured from a pulse-controlled amplification (PCA) procedure as taught by Ullerich and execute a machine learning model to detect a target nucleic acid strand in a nucleic acid sample based on the plot because Ullerich shows that the fluorescence plot from a pulse-controlled amplification (PCA) procedure is substantially similar to a plot of a fluorescence signal measured for other PCR based instruments so that there would have been an expectation that the machine learning model of Moniri ’20 would have been able to perform the detection on a fluorescence plot from a pulse-controlled amplification procedure.
Claim 14 is rejected under 35 U.S.C. 103 as being unpatentable over Moniri ’20 in view of Ullerich as applied to claim 13 above, and further in view of Cascio or Hu both as described above. Moniri ‘20 did not teach the use of a convolutional neural network with a plurality of convolution layers and a fully connected classifier layer.
It would have been obvious to one of ordinary skill in the art at the time the application was filed to modify the method of Moniri ’20 to use a convolutional neural network as the machine learning model as taught by Cascio or Hu because of its ability to handle fluorescent image analysis in general as taught by Cascio or to handle dPCR fluorescent image analysis as taught by Hu in a manner that was better than other methods working on/with the same problem(s).
Claim 15 is rejected under 35 U.S.C. 103 as being unpatentable over Moniri ’20 in view of Ullerich as applied to claim 13 above, and further in view of Kumar (PeerJ 2020). Figure 6 and the paragraph bridging pages 13140-13141 of Moniri ’20 teach that the machine learning model was trained, but do not teach that the machine learning model is trained based on minority oversampling.
In the paper Kumar teaches that machine learning techniques are increasingly used in the analysis of high throughput genome sequencing data to better understand the disease process and design of therapeutic modalities. In the current study, they applied state of the art machine learning (ML) algorithms (Random Forest (RF), Support Vector Machine Radial Kernel (svmR), Adaptive Boost (AdaBoost), averaged Neural Network (avNNet), and Gradient Boosting Machine (GBM)) to stratify the HNSCC patients in early and late clinical stages (TNM) and to predict the risk using miRNAs expression profiles. A six miRNA signature was identified that can stratify patients in the early and late stages. The mean accuracy, sensitivity, specificity, and area under the curve (AUC) was found to be 0.84, 0.87, 0.78, and 0.82, respectively indicating the robust performance of the generated model. The prognostic signature of eight miRNAs was identified using LASSO (least absolute shrinkage and selection operator) penalized regression. These miRNAs were found to be significantly associated with overall survival of the patients. The pathway and functional enrichment analysis of the identified biomarkers revealed their involvement in important cancer pathways such as GP6 signalling, Wnt signalling, p53 signalling, granulocyte adhesion, and dipedesis. in the paragraph bridging pages 6-7 of the paper Kumar discusses data balancing. A challenge in biomedical studies is working with imbalanced data sets i.e., unequal number of normal and disease samples. Unbalanced predictive variable ratio do not meet the assumptions of the machine learning models and its predict biased. Therefore, SMOTE algorithm was used to balance the data using the 'DMwR' library package. The SMOTE is a popular algorithm to deal with unbalanced data. There are two types of sampling strategies commonly employed in machine learning viz. oversampling and under-sampling for unbiased classification. In under-sampling, the samples are reduced based on k nearest neighbor (kNN) clustering centroid distance while in oversampling minority classes are amplified to balance both predictors. In their data set, the number of late stage patient samples was reduced so as to match the number of early stage patient samples (N = 101) using the under-sampling method and in oversampling, early stage data was amplified (N = 352) to balanced early and later-stage samples. Thereafter, the datasets (under and over sampled) were systematically distributed into training (70%) and test set (30%) separately. In under-sampled data, 141 and 61 samples were used as training and test set respectively whereas in oversampled data 478 and 206 samples were used as training and test set respectively. First, the model was trained using only training set the test set was kept aside for independent external cross-validation. A 10-fold repeated internal cross-validation with 10-time iterations to randomized data set was done to avoid generation of over/under fitting model. The cost functions were optimized (100 to 3,000 with 100 steps per iteration) to achieve accurate classification. The performance of the generated model was examined based on sensitivity, specificity, accuracy, and AUC (ROC), precision and MCC by external independent data. The last full paragraph on page 18 of the paper teaches that the small cohort and imbalanced data highlight the major challenges in this study. The SMOTE algorithm was therefore utilized to generate a balanced dataset for classification of stages. It was interesting to note that the models generated with oversampling of the data showed better accuracy as compared to the undersampled data. Perhaps, the undersampling of the data might have resulted in loss of important information and thus reduced accuracy.
It would have been obvious to one of ordinary skill in the art at the time the application was filed to train the Moniri ’20 machine learning model on balance data based on over sampling or minority data as taught by Kumar because of the greater accuracy of that method compared to the other methods studied as taught by Kumar.
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. The additionally cited art is related to various amplification methods and machine learning models.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Arlen Soderquist whose telephone number is (571)272-1265. The examiner can normally be reached 1st week Monday-Thursday, 2nd week Monday-Friday.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Lyle Alexander can be reached at (571)272-1254. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ARLEN SODERQUIST/ Primary Examiner, Art Unit 1797