Prosecution Insights
Last updated: August 17, 2026
Application No. 17/853,044

AUTOMATIC PARKINSONS DISEASE DETECTION BASED ON THE COMBINATION OF LONG-TERM ACOUSTIC FEATURES AND MEL FREQUENCY COEFFICIENTS (MFCCs)

Final Rejection §103
Filed
Jun 29, 2022
Examiner
LEE, MICHAEL CHRISTOPHER
Art Unit
2128
Tech Center
2100 — Computer Architecture & Software
Assignee
Imam Abdulrahman Bin Faisal University
OA Round
4 (Final)
62%
Grant Probability
Moderate
5-6
OA Rounds
0m
Est. Remaining
88%
With Interview

Examiner Intelligence

Grants 62% of resolved cases
62%
Career Allowance Rate
95 granted / 153 resolved
+7.1% vs TC avg
Strong +26% interview lift
Without
With
+26.4%
Interview Lift
resolved cases with interview
Typical timeline
3y 3m
Avg Prosecution
53 currently pending
Career history
197
Total Applications
across all art units

Statute-Specific Performance

§101
30.1%
-9.9% vs TC avg
§103
45.2%
+5.2% vs TC avg
§102
10.5%
-29.5% vs TC avg
§112
12.8%
-27.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 153 resolved cases

Office Action

§103
DETAILED ACTION Notice of AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 2/17/2026 has been entered. Response to Amendment Applicant’s Amendment and remarks dated 2/17/2026 have been considered. Claims 4, 11, and 18 are cancelled. Claims 1-3, 5-10, 12-17, and 19-20 are pending. Response to Arguments On page 13 of Applicant’s 2/17/2026 Amendment and remarks, Applicant asserts that no new matter has been added by the claim amendments. The examiner agrees that at least Fig. 3, as explained on pp. 21-23 of the instant specification, provides sufficient written description support for the claim amendments. On pages 13-17 of Applicant’s 2/17/2026 Amendment and remarks, with respect to the rejections of all claims under 35 U.S.C. 101, Applicant makes several arguments. The examiner finds Applicant’s arguments on pages 15-16, with respect to Step 2A, Prong 2, to be persuasive. In particular, Applicant’s argument that it would not be reasonable for a “human to divide voice signals into overlapping frames, perform windowing on the overlapping frames,...” to practically be performed in the human mind, particularly in view of the new Hanning window limitation, and explanation. Therefore, this particular claimed signal processing technique appears to be an improvement in technologies with respect to extracting acoustic features for the purpose of predicting diseases such as Parkinson’s disease, and is therefore considered to be subject matter eligible. All rejections under 35 U.S.C. 101 are hereby withdrawn. The examiner respectfully submits that Applicant’s remaining arguments with respect to 35 U.S.C. 101 are moot in view of the withdrawal of all such rejections. On pages 17-21 of Applicant’s 2/17/2026 Amendment and remarks, with respect to the rejections of claim 1 under 35 U.S.C. 103, Applicant argues that the combination of SOLANA-LAVALLE and KRNETA does not make obvious the “a set C, set C comprising backward stepwise selection of the set B of long-term acoustic features combined with the set A of short-term acoustic features” limitation. The examiner respectfully submits that Applicant’s arguments are moot because KRNETA is no longer utilized in any rejections. The new rejections rely on the ZHUANG reference for this limitation. On pages 21-22 of Applicant’s 2/17/2026 Amendment and remarks, with respect to the rejections of the remaining claims under 35 U.S.C. 103, Applicant argues that such claims should be allowed for the same reasons explained with respect to claim 1. The examiner respectfully disagrees for the same reasons explained above with respect to claim 1. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1-3, 6-10, 13-17, and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Solana-Lavalle, et al. "Automatic Parkinson disease detection at early stages as a pre-diagnosis tool by using classifiers and a small set of vocal features." Biocybernetics and Biomedical Engineering 40.1 (2020): pp. 505-516, hereinafter referenced as SOLANA-LAVALLE, in view of Zhuang, Juntang, et al. "Prediction of pivotal response treatment outcome with task fMRI using random forest and variable selection." 2018 IEEE 15th International Symposium on Biomedical Imaging (ISBI 2018). IEEE, 2018, hereinafter referenced as ZHUANG, and further in view of US 20150265205 A1, hereinafter referenced as ROSENBEK, and further in view of Kumar, Ashwin Nair Anil, et al. "Text dependent voice recognition system using MFCC and VQ for security applications." 2017 International conference of Electronics, Communication and Aerospace Technology (ICECA). Vol. 2. IEEE, 2017, pp. 130-136, hereinafter referenced as KUMAR, and further in view of Firmansyah, Muhammad Raafi'U., et al. "Comparison of windowing function on feature extraction using MFCC for speaker identification." 2021 international conference on intelligent cybernetics technology & applications (ICICYTA). IEEE, 2021, hereinafter referenced as FIRMANSYAH. Regarding Claim 1 SOLANA-LAVALLE teaches: A machine-learning method to differentiate between patients with neurodegenerative disease and healthy patients, the method comprising: (SOLANA-LAVALLE, p. 506, section 2: “The proposed method for vocal-based PD detection consists of different stages.” SOLANA-LAVALLE, p. 509, section 2.4: “In this work, we are using four classification techniques; kNN, SVM, MLP and Random Forest.”; Examiner’s Note (EN): Page 505, section 1, defines PD to be “Parkinson Disease”, and the proposed method determines if a patient has PD or not (e.g., is healthy), where kNN, SVM, MPL, and Random Forest are each types of machine learning algorithms; pursuant to MPEP 2131.01 II, to “explain the meaning of a term used in the primary reference”, the examiner cites to “Classification with Machine Learning”, available at https://web.archive.org/web/20210801170139/https://apmonitor.com/do/index.php/Main/MachineLearningClassifier (archived on August 1, 2021), to establish that the meaning of “kNN” (or k-Nearest Neighbors), “SVM” (or support vector machine), “MLP” (or multi-layer perceptron), and “Random Forest” are each types of machine learning classification algorithms) obtaining, via at least one microphone, a first plurality of recordings of voice signals from known healthy humans and known neurogenerative diseased humans; (SOLANA-LAVALLE, p. 507, section 2.1: “The dataset for this study is the Parkinson's Disease Classification dataset, which is found within the Machine Learning Repository of the University of California Irvine. This dataset was generated by the Cerrahpsa Faculty of Medicine at the Department of Neurology, Istanbul University, from 188 PD patients (107 men and 81 women) with ages ranging from 33 to 87 years old (65.1 years old +/- 10.9 years), and from 64 healthy individuals (23 men and 41 women) with ages from 41 to 82 years old (61.1 years old +/- 8.9 years). ... The voice samples were recorded after the subjects under study were diagnosed as being healthy or having Parkinson disease.”; SOLANA-LAVALLE, p. 507, section 2.3: “During generation of the dataset, the sustained phonation of vowel /a/ was recorded from each individual. The recording time for each recorded instance was 220 s through the use of a microphone with a sampling frequency of 44 kHz”; (EN): this dataset, obtained from the Univ. of California, Irvine, includes recordings of voice samples from known healthy patients and patients with Parkinson Disease, and SOLANA-LAVALLE further discloses using a microphone to capture such recordings) extracting one or more long-term acoustic features from the recordings of the first plurality of voice signals; (SOLANA-LAVALLE, p. 507, section 2.3: “A total of 754 vocal features were extracted from each recorded instance”; SOLANA-LAVALLE, pp. 507-508, section 2.3.1: “The baseline feature set has been one of the most commonly used sets in speech analysis. For the case of the proposed method, only two baseline features were used, jitter and detrended fluctuation analysis (DFA).”; (EN): jitter and DFA correspond to the recited “long-term acoustic features”; the examiner notes that the instant specification at page 12, lines 11-14 specifically identifies jitter and DFA as examples of long-term acoustic features) extracting Mel frequency coefficients (MFCCs) from the recording of each of the first plurality of voice signals; (SOLANA-LAVALLE, p. 508, section 2.3.2: “The Mel frequency cepstral coefficients (MFCC) are obtained by applying the discrete cosine transform (DCT)”) creating a set A of short-term acoustic features based on the MFCCs; (SOLANA-LAVALLE, p. 508, section 2.3.2: “There are features related to the change in the cepstral coefficients over time, which are called delta coefficients. Typical MFCC-related features are the 12 MFCC coefficients (the mean of the sixth coefficient was selected by Wrappers for the case of the SVM classifier), 12 delta (or velocity) MFCC features (the mean of the third delta was selected by Wrappers for the MLP, the standard deviations of the seventh delta and ninth delta were selected by Wrappers for the SVM), 12 double-delta (or acceleration) MFCC features (the standard deviations of the first and ninth double-deltas were selected by Wrappers for the RF classifier), 1 energy feature,1 delta energy feature, 1 double-delta energy feature (selected by Wrappers for the case of the kNN, MLP and SVM).”; SOLANA-LAVALLE, p. 512, Table 2: “Entropy 3rd coef., entropy-log 8th coef., 1st coef., mean 4th coef” in the “Baseline features” row; Examiner’s Note: SOLANA-LAVALLE discloses deriving different features from the MFCCs, and that a sub-set of these features (corresponding to recited “set A”) is identified in Table 2). performing a backward stepwise selection of the long-term acoustic features to obtain a set B of long-term acoustic features and (SOLANA-LAVALLE, p. 507, section 2.2: “the Wrappers feature subset selection is a computationally efficient alternative because it considers a reduced number of these subsets. It combines forward stepwise selection and backward stepwise selection. Forward stepwise selection begins with a model with no features, and it adds features to the model, one at a time. At each step, the feature, that gives the best improvement, is added to the model. Backward stepwise selection begins with all n features, and it iteratively removes the least useful feature, one at a time. Wrappers adds a new feature while it also checks the relevance of already added features. If it finds an insignificant feature, then it removes that particular feature. The steps involved are: ... c) Perform backward elimination of any previously added feature” SOLANA-LAVALLE, p. 512, Table 2: “DFA” and “jitter” in the “MFCCs” row; (EN): the Wrappers feature subset selection utilizes backward stepwise selection to remove less useful features from the Baseline features (e.g., for the RF model, the reduced set of Baseline features is limited to jitter, corresponding to recited set (B)); Table 2, which is the combined version of 4 categories of features (Baseline, MFCCs, TQWT, and Number of features) corresponds to recited “set C”)) configuring a random forest classification model with the features of sets A, B, (SOLANA-LAVALLE, p. 510, section 2.4: “In a random forest algorithm, a majority of the classifications is not considered at each split in the tree by forcing each split to consider only a subset of the features. The RF implementation consisted of a set of 87–120 decision trees by reaching a tree depth of 150.”; SOLANA-LAVALLE, p. 512, Table 2 – see “RF” column; SOLANA-LAVALLE, p. 513, section 4: “The RF classifier reaches the highest PD detection performance when it is tested with a feature set selected by running Wrappers with RF.” (EN): Table 2 shows that the features that were used to train the Random Forest classifier, including at least Entropy 3rd coef. (corresponding to a short-term feature of recited “Set A”) and hitter (corresponding to a long-term feature of recited “Set B”)) applying the second plurality of voice signals against the random forest classification model in order to determine which patients in the second plurality of voice signals are healthy patients and which are neurodegenerative disease patients. (SOLANA-LAVALLE, p. 511, section 2.6: “Thus, the set is divided into ten folds. One fold is picked for testing of a classifier and the other nine folds are left for training. This process, of choosing one fold for testing and the rest for training, is repeated ten times (ten-fold cross-validation).”; SOLANA-LAVALLE, p. 511, section 3: “Table 3 shows the PD detection performance in four different classifiers, (kNN, MLP, SVM and RF) when they are tested with a feature set selected by running the Wrappers algorithm with a kNN classifier. The best results are highlighted in boldface. Tables 4–6 show the PD detection performance of the four classifiers, when they are tested with a feature set selected by running the Wrappers algorithm with MLP, SVM and RF, respectively. These results were obtained by using 10-fold cross-validation.”; (EN): the dataset was split into a training set and testing set, where the testing set corresponds to recited “second plurality of voice signals”, and this testing set is applied the RF classifier, with results shown in Tables 3-6 on p. 512) Wherein the Mel frequency coefficients are extracted by a method comprising: (SOLANA-LAVALLE, p. 508, section 2.3.2: “The Mel frequency cepstral coefficients (MFCC) are obtained by applying the discrete cosine transform (DCT)”) However, SOLANA-LAVALLE fails to explicitly teach: a set C, set C comprising backward stepwise selection of the set B of long-term acoustic features combined with the set A of short-term acoustic features; set C obtaining a second plurality of voice signals from humans of undetermined health status dividing the voice signals into overlapping frames, each frame containing a plurality of samples wherein the overlap is between 30% and 50% of the frame; windowing the overlapping frames wherein the window is of length 20-40ms, wherein each frame is multiplied by a Hanning window of a length equal to a number of the plurality of samples; applying a Fast Fourier Transform (FFT) to convert the voice signal to a frequency domain; calculating logarithm of an average value of a spectral power density in each of the frames to model the voice signal in a cepstral domain; creating Mel filterbanks within the cepstral domain; and performing a discrete cosine transformation (DCT) on the Mel filterbanks to determine the Mel frequency coefficients. However, in a related field of endeavor (feature selection in the medical field, see p. 97, section 1), ZHUANG teaches: a set C, set C comprising backward stepwise selection of the set B of long-term acoustic features combined with the set A of short-term acoustic features; (ZHUANG, p. 97, section 1: “We propose a two-step feature selection method to solve this problem. The first step selects all relevant variables [11], and the second step performs stepwise variable selection to find the minimal optimal set of predictor variables.”; ZHUANG, p. 98, section 2.3.3: “Number of variables after candidate variable selection is still large compared to dataset size, and the dataset may contain many correlated variables. Therefore, we select variables by stepwise building of tree ensembles. We tried both forward and backward stepwise variable selection.”; Examiner’s Note: ZHUANG teaches a 2-step process for feature selection, where the first step selects all relevant variables (corresponding to combining the features of sets A and B), and the second step uses backward stepwise selection to determine the minimal optimal set of predictor variables (corresponding to recited “set C”); the SOLANA-LAVALLE-ZHUANG combination now creates a set C as explained by ZHUANG, by using backward stepwise selection (as in ZHUANG) to the combined sets of A and B, in order to determine the minimal optimal set of predictor variables as disclosed by ZHUANG) configuring a random forest classification model with the features of sets A, B, and C in order to classify healthy patients and neurodegenerative disease patients; (ZHUANG, p. 97, section 1: “We propose a two-step feature selection method to solve this problem. The first step selects all relevant variables [11], and the second step performs stepwise variable selection to find the minimal optimal set of predictor variables.”; ZHUANG, p. 98, section 2.3.3: “Number of variables after candidate variable selection is still large compared to dataset size, and the dataset may contain many correlated variables. Therefore, we select variables by stepwise building of tree ensembles. We tried both forward and backward stepwise variable selection.”; Examiner’s Note: the SOLANA-LAVALLE-ZHUANG combination now configures a random forest model (as in SOLANA-LAVALLE) using the set C created using the backward stepwise selection teachings of ZHUANG, where set C includes features of sets A and B that were selected during backwards stepwise selection) Before the effective filing date of the present application, it would have been obvious to one of ordinary skill in the art to combine the teachings of SOLANA-LAVALLE with the teachings of ZHUANG as explained above. As disclosed by ZHUANG, one of ordinary skill would have been motivated to do so in order to use ZHUANG’s 2-step process to lower the “risk of eliminating truly predictive variables.” (pp. 97-98, section 1). However, SOLANA-LAVALLE and ZHUANG fail to explicitly teach: obtaining a second plurality of voice signals from humans of undetermined health status dividing the voice signals into overlapping frames, each frame containing a plurality of samples wherein the overlap is between 30% and 50% of the frame; windowing the overlapping frames wherein the window is of length 20-40ms, wherein each frame is multiplied by a Hanning window of a length equal to a number of the plurality of samples; applying a Fast Fourier Transform (FFT) to convert the voice signal to a frequency domain; calculating logarithm of an average value of a spectral power density in each of the frames to model the voice signal in a cepstral domain; creating Mel filterbanks within the cepstral domain; and performing a discrete cosine transformation (DCT) on the Mel filterbanks to determine the Mel frequency coefficients. However, in a related field of endeavor (screening for neurological diseases, such as PD, using speech characteristics as a biomarker, see para. 0003), ROSENBEK teaches: obtaining a second plurality of voice signals from humans of undetermined health status (ROSENBEK, para. 0049: “the identification device 200 can be located at the testing site of a patient. In one such embodiment, the identification device 200 can be part of a computer or mobile device such as a smartphone. The interface 201 can include a user interface such as a graphical user interface (GUI) provided on a screen or display. An input to the identification device 200 can include a microphone, which is connected to the device in such a manner that a speech sample can be recorded into the device 200. Alternately, a speech sample can be recorded on another medium and copied (or otherwise transmitted) to the device 200. Once the speech sample is input to the device 200, the processor of the computer or mobile device can provide the processor 202 of the device 200 and perform the identification procedures to determine the health state of the subject.”; (EN): the SOLANA-LAVALLE-ROSENBEK combination now collects speech samples from other humans to be tested using the RF PD classifier of SOLANA-LAVALLE; the examiner notes that one of ordinary skill would understand that samples can be collected prior to diagnosis, such that the voice samples are collected based on the recited “undetermined health status” for later determination) Before the effective filing date of the present application, it would have been obvious to one of ordinary skill in the art to combine the teachings of SOLANA-LAVALLE with the teachings of ZHUANG and ROSENBEK as explained above. As disclosed by ROSENBEK, one of ordinary skill would have been motivated to do so in order to monitor a patient over time to “predict the likelihood of one or more neurological/neurodegenerative or other disease, such as infectious and/or respiratory disease, condition(s).” (para. 0029). However, SOLANA-LAVALLE, ZHUANG, and ROSENBEK fail to explicitly teach: dividing the voice signals into overlapping frames, each frame containing a plurality of samples wherein the overlap is between 30% and 50% of the frame; windowing the overlapping frames wherein the window is of length 20-40ms, wherein each frame is multiplied by a Hanning window of a length equal to a number of the plurality of samples; applying a Fast Fourier Transform (FFT) to convert the voice signal to a frequency domain; calculating logarithm of an average value of a spectral power density in each of the frames to model the voice signal in a cepstral domain; creating Mel filterbanks within the cepstral domain; and performing a discrete cosine transformation (DCT) on the Mel filterbanks to determine the Mel frequency coefficients. However, in a related field of endeavor (feature extraction of voice signals, p. 130, section I), KUMAR teaches: dividing the voice signals into overlapping frames, each frame containing a plurality of samples wherein the overlap is between 30% and 50% of the frame; (KUMAR, p. 131, section II.B: “Therefore, in order to analyze the signal, it must be split into frames of few milliseconds. Another advantage of frame based analysis is that it improves the efficiency of the system by analyzing groups of samples (in a frame) as opposed to analyzing each sample separately. Each frame overlaps with the previous frame so as to ensure a smooth transition of signal from one frame to another i.e. less discontinuities. Large overlapping creates smoother transition of signals between frames but results in a smaller time shift in the signal which requires higher processing power. ... Referring previous studies [1,4], typical number of samples for frame length (N) and overlap (M) are 256 and 100 respectively.”; (EN): 100/256 is about 39%; the examiner notes that MPEP 2131.03 I.A explains that a specific example in the prior art which is within a claimed range anticipates the range, and 39% overlap is specifically within the claimed 30-50% overlap range) windowing the overlapping frames wherein the window is of length 20-40ms, wherein each frame is multiplied by a samples; (KUMAR, p. 131, section II.B: “The ideal frame length lies in the range of 20ms – 40ms”; KUMAR, p. 131, section II.C: “Since each frame consists of 256 samples, a 256-point Hamming window as illustrated in Fig. 3 was utilized for this paper.” (EN): KUMAR discloses that the window length equals the frame length, and that the frame length is in the range of 20-40 ms, which matches the claimed range precisely, and further discloses that a Hamming window of 256 points is used, where such Hamming window is based on the 256 samples per frame) applying a Fast Fourier Transform (FFT) to convert the voice signal to a frequency domain; (KUMAR, pp. 132-133, section III.A: “The first step in the feature extraction phase is to convert the speech signal into the frequency domain using the Fourier transform. ... The Fast Fourier Transform (FFT) is used as it is a fast algorithm for implementing the Discrete Fourier Transform (DFT) to obtain the frequency spectrum.”) calculating logarithm of an average value of a spectral power density in each of the frames to model the voice signal in a cepstral domain; (KUMAR, p. 133, section III.C: “A cepstrum is obtained by taking the inverse Fourier transform of the logarithm of an estimated spectrum of a signal. The only remaining step after obtaining the log mel spectrum is to convert the spectrum back into the time domain. Since the mel spectrum coefficients (and their logarithm) are real numbers, the time domain conversion can be achieved by utilizing the even property of the real cepstrum whereby it is possible to express the inverse DFT in terms”) creating Mel filterbanks within the cepstral domain; and (KUMAR, p. 133, section III.B: “The triangular mel filter bank is then designed per the new bins and are applied to the frequency spectrum to produce the mel spectrum. An example of the mel filter bank consisting of 24 filters used by Davis and Mermelstein [9] is shown in Fig. 6.”) performing a discrete cosine transformation (DCT) on the Mel filterbanks to determine the Mel frequency coefficients. (KUMAR, p. 131, section III.C: “Applying the DCT to the log mel spectrum converts the mel spectrum into the mel cepstrum. The coefficients of this resulting mel cepstrum are called the Mel Frequency Cepstral Coefficients (MFCCs) which are the feature vectors used in this paper to characterize an individual.”; (EN): the SOLANA-LAVALLE-ZHUANG-ROSENBEK-KUMAR combination now calculates MFCCs from the audio samples of SOLANA-LAVELLE using the process of KUMAR) Before the effective filing date of the present application, it would have been obvious to one of ordinary skill in the art to combine the teachings of SOLANA-LAVALLE with the teachings of ZHUANG, ROSENBEK, and KUMAR as explained above. As disclosed by KUMAR, one of ordinary skill would have been motivated to do so in order to because the “MFCCs are based on the mel scale which is a scale that accurately approximates the human hearing system.” (p. 135, section V). However, SOLANA-LAVALLE, ZHUANG, ROSENBEK, and KUMAR fail to explicitly teach: Hanning window However, in a related field of endeavor (MFCC extraction with respect to voice signals, see p. 1, section I), FIRMANSYAH teaches and makes obvious: wherein each frame is multiplied by a Hanning window of a length equal to a number of the plurality of samples (FIRMANSYAH, p. 1, section I: “This article focuses on the windowing process by experimenting with several other window functions such as hanning, bartlett, blackman, kaiser, and gaussian.” FIRMANSYAH, p. 4, section III: “Afterward, seen in Fig.5, the hanning and blackman window functions can reduce the ends of the signal towards zero better than the hamming and bartlett functions”; Examiner’s Note: the SOLANA-LAVALLE-ZHUANG-ROSENBEK-KUMAR-FIRMANSYAH combination now calculates MFCCs from the audio samples of SOLANA-LAVELLE using the process of KUMAR, but replacing the Hamming window of KUMAR with the Hanning window of FIRMANSYAH) Before the effective filing date of the present application, it would have been obvious to one of ordinary skill in the art to combine the teachings of SOLANA-LAVALLE with the teachings of ZHUANG, ROSENBEK, KUMAR, and FIRMANSYAH as explained above. As disclosed by FIRMANSYAH, one of ordinary skill would have been motivated to do so in order to because FIRMANSYAH experimentally found situations where a hanning window can “reduce the ends of the signal towards zero better than the hamming” window function. (p. 4, section III). Regarding Claim 2 SOLANA-LAVELLE, ZHUANG, ROSENBEK, KUMAR, and FIRMANSYAH disclose the method of claim 1. SOLANA-LAVELLE further teaches: determining an accuracy, a specificity, and a sensitivity of the random forest classification model wherein the accuracy is calculated by: PNG media_image1.png 240 390 media_image1.png Greyscale where TP is true positive (TP) indicates a number of correctly classified diseased patients and true negative (TN) expresses a number of correctly classified healthy patients and false positive (FP) indicates a number of incorrectly classified healthy subjects, and false negative (FN) expresses a number of incorrectly classified diseased patients. (SOLANA-LAVALLE, p. 510, section 2.5: PNG media_image2.png 130 476 media_image2.png Greyscale PNG media_image3.png 218 494 media_image3.png Greyscale Examiner’s Note: SOLANA-LAVALLE teaches the identical equations for accuracy, specificity, and sensitivity, where Table 6 on page 512 has the results for the RF classifier) Regarding Claim 3 SOLANA-LAVELLE, ZHUANG, ROSENBEK, KUMAR, and FIRMANSYAH disclose the method of claim 1. SOLANA-LAVELLE further teaches: wherein the random forest classification model is created by: (SOLANA-LAVALLE, p. 510, section 2.4: “In a random forest algorithm, a majority of the classifications is not considered at each split in the tree by forcing each split to consider only a subset of the features. The RF implementation consisted of a set of 87–120 decision trees by reaching a tree depth of 150.”) dividing the first plurality of voice signals into a training set and a test set of voice signals; (SOLANA-LAVALLE, p. 511, section 2.6: “Feature vectors, from PD or healthy subjects, are stored into two sets: C1 for patients with PD, C2 for healthy subjects. Each set, C1 and C2, is separated into ten fragments, .... Then, a fragment c1;i (from C1) and a corresponding fragment c2;i (from C2) are randomly combined into ci. The result of these random mixings is ten fragments or folds ..., where each fold contains instances form PD and healthy subjects. Thus, the set is divided into ten folds. One fold is picked for testing of a classifier and the other nine folds are left for training. This process, of choosing one fold for testing and the rest for training, is repeated ten times (ten-fold cross-validation).”; (EN): from the first set of voice signals, 9/10 are used for training and 1/10 is used for testing) building the random forest classification model using bootstrap sampling of the training set wherein the model comprises multiple decision trees produced by multiple training subsets; and (SOLANA-LAVALLE, p. 510, section 2.4: “Decision trees have high variance, which implies that region splitting could be quite different for different partitions of the training set. Bootstrap or bagging is a procedure for reducing variance, where multiple decision trees are built using different bootstrapped training sets, and the tree classifications are averaged. For a given observation, the assigned class is the most occurring class among multiple tree classifications. In a random forest algorithm, a majority of the classifications is not considered at each split in the tree by forcing each split to consider only a subset of the features. The RF implementation consisted of a set of 87–120 decision trees by reaching a tree depth of 150.”; (EN): Boostrap sampling is used on the 87-120 different decision trees, which are created using “different bootstrapped training sets” (corresponding to recited “multiple training subsets”) testing the model with the test set of voice signals. (SOLANA-LAVALLE, p. 511, section 3: “Table 3 shows the PD detection performance in four different classifiers, (kNN, MLP, SVM and RF) when they are tested with a feature set selected by running the Wrappers algorithm with a kNN classifier. The best results are highlighted in boldface. Tables 4–6 show the PD detection performance of the four classifiers, when they are tested with a feature set selected by running the Wrappers algorithm with MLP, SVM and RF, respectively. These results were obtained by using 10-fold cross-validation.”; (EN): the dataset was split into a training set and testing set, and testing set is applied the RF classifier, with results shown in Tables 3-6 on p. 512) Regarding Claim 6 SOLANA-LAVELLE, ZHUANG, ROSENBEK, KUMAR, and FIRMANSYAH disclose the method of claim 1. SOLANA-LAVELLE further teaches: wherein the long-term acoustic features comprise any of:a relative average perturbation; a jitter; an amplitude perturbation quotient; a shimmer; a detrended fluctuation analysis; a minimum intensity; a maximum intensity; a mean intensity; and a formant frequency. (SOLANA-LAVALLE, pp. 507-508, section 2.3.1: “The baseline feature set has been one of the most commonly used sets in speech analysis. For the case of the proposed method, only two baseline features were used, jitter and detrended fluctuation analysis (DFA).”; (EN): jitter and DFA correspond to the recited “long-term acoustic features”; the examiner notes that the instant specification at page 12, lines 11-14 specifically identifies jitter and DFA as examples of long-term acoustic features) Regarding Claim 7 SOLANA-LAVELLE, ZHUANG, ROSENBEK, KUMAR, and FIRMANSYAH disclose the method of claim 1. SOLANA-LAVELLE further teaches: wherein the neurodegenerative disease is Parkinson's disease. (SOLANA-LAVALLE, p. 506, section 2: “The proposed method for vocal-based PD detection consists of different stages.”; Examiner’s Note (EN): Page 505, section 1, defines PD to be “Parkinson Disease”) Regarding Claim 8 SOLANA-LAVELLE teaches: A medical diagnostic system, comprising: (SOLANA-LAVALLE, p. 506, section 2: “The proposed method for vocal-based PD detection consists of different stages.”) one or more processors, (SOLANA-LAVALLE, p. 511, section 3: “These sets of experiments were run in a processor Intel core i5-8250U at 1.60 GHz with 6 MB of cache memory, 4 Cores and 8 threads.”) a memory, (SOLANA-LAVALLE, p. 511, section 3: “These sets of experiments were run in a processor Intel core i5-8250U at 1.60 GHz with 6 MB of cache memory, 4 Cores and 8 threads.”) a microphone, and (SOLANA-LAVALLE, p. 507, section 2.3: “The recording time for each recorded instance was 220 s through the use of a microphone with a sampling frequency of 44 kHz.”) a circuitry configured to: , (SOLANA-LAVALLE, p. 511, section 3: “These sets of experiments were run in a processor Intel core i5-8250U at 1.60 GHz with 6 MB of cache memory, 4 Cores and 8 threads.”; (EN): cores comprise electronic circuitry that can be configured) The remaining limitations of claim 8 correspond to the method of claim 1, and therefore this claim is rejected for the same reasons explained above with respect to claim 1 in view or 35 U.S.C. 103 and the SOLANA-LAVELLE, ZHUANG, ROSENBEK, KUMAR, and FIRMANSYAH references. Claim 9 depends from claim 8 and claims a system that corresponds to the method of claim 2, and is therefore rejected for the same reasons explained with respect to claims 2 and 8. Claim 10 depends from claim 8 and claims a system that corresponds to the method of claim 3, and is therefore rejected for the same reasons explained with respect to claims 3 and 8. Claim 13 depends from claim 8 and claims a system that corresponds to the method of claim 6, and is therefore rejected for the same reasons explained with respect to claims 6 and 8. Claim 14 depends from claim 8 and claims a system that corresponds to the method of claim 7, and is therefore rejected for the same reasons explained with respect to claims 7 and 8. Regarding Claim 15 SOLANA-LAVELLE teaches: A non-transitory computer-readable storage medium storing computer-readable instructions that, when executed by one or more processors, cause the one or more processors to: (SOLANA-LAVALLE, p. 511, section 3: “These sets of experiments were run in a processor Intel core i5-8250U at 1.60 GHz with 6 MB of cache memory, 4 Cores and 8 threads.”) The remaining limitations of claim 15 correspond to the method of claim 1, and therefore this claim is rejected for the same reasons explained above with respect to claim 1 in view or 35 U.S.C. 103 and the SOLANA-LAVELLE, ZHUANG, ROSENBEK, KUMAR, and FIRMANSYAH references. The examiner notes that removing the distinction between first voice signals pertaining to “known healthy humans” and “known neurogenerative diseases humans” does not change the analysis under 35 U.S.C. 103. Claim 16 depends from claim 15 and claims a non-transitory computer-readable storage medium that corresponds to the method of claim 2, and is therefore rejected for the same reasons explained with respect to claims 2 and 15. Claim 17 depends from claim 15 and claims a non-transitory computer-readable storage medium that corresponds to the method of claim 3, and is therefore rejected for the same reasons explained with respect to claims 3 and 15. Claim 19 depends from claim 15 and claims a non-transitory computer-readable storage medium that corresponds to the method of claim 6, and is therefore rejected for the same reasons explained with respect to claims 6 and 15. Claim 20 depends from claim 15 and claims a non-transitory computer-readable storage medium that corresponds to the method of claim 7, and is therefore rejected for the same reasons explained with respect to claims 7 and 15. Claims 5 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over SOLANA-LAVELLE in view of ZHUANG, ROSENBEK, KUMAR, and FIRMANSYAH and further in view of Wood, Angela M., et al. "How should variable selection be performed with multiply imputed data?." Statistics in medicine 27.17 (2008): pp. 3227-3246, hereinafter referenced as WOOD. Regarding Claim 5 SOLANA-LAVELLE, ZHUANG, ROSENBEK, KUMAR, and FIRMANSYAH disclose the method of claim 1. SOLANA-LAVELLE further teaches: wherein the backward stepwise selection of the long-term acoustic features comprises: (SOLANA-LAVELLE, p. 507, section 2.2: “Backward stepwise selection begins with all n features, and it iteratively removes the least useful feature, one at a time.”; (EN): p. 512, table 2, shows the jitter and DFA features are selected for different models) starting with a model with a full set of long-term acoustic features; (SOLANA-LAVELLE, p. 507, section 2.2: “Backward stepwise selection begins with all n features, and it iteratively removes the least useful feature, one at a time.”) iteratively removing a particular feature that has the least significance for model accuracy; (SOLANA-LAVELLE, p. 507, section 2.2: “Backward stepwise selection begins with all n features, and it iteratively removes the least useful feature, one at a time.”; (EN): one of ordinary skill would understand that the “least useful feature” can be determined with respect to model accuracy, which is one of the performance metrics tracked as shown on p. 512, Tables 3-6) However, SOLANA-LAVELLE, ZHUANG, ROSENBEK, KUMAR, and FIRMANSYAH fail to explicitly teach: removing the particular feature from the model when removal of the particular feature from the model improves model performance wherein performance is measured by accuracy, specificity, sensitivity, or an area under a curve; returning the particular feature to the model when removing the particular feature worsens the model performance; and repeating removal of each feature in the set of long-term acoustic features until the best performance of the model is achieved as measured by accuracy, a specificity, a sensitivity, or an area under the curve. However, in a related field of endeavor (backward stepwise selection in the medical research field, see p. 3228, section 1), WOOD teaches: removing the particular feature from the model when removal of the particular feature from the model improves model performance wherein performance is measured by accuracy, specificity, sensitivity, or an area under a curve; (WOOD, p. 3230, section 2.3: “An alternate approach is backward selection. Starting with a model with all potential variables, each variable in the model is tested for exclusion from the model. ... The backward stepwise selection procedure consists of backward selection, followed by forward selection and iterates if necessary. Starting with a model with all potential variables, the backward selection process removes from the model the variable that is least statistically important (at α per cent significance).”; (EN): the SOLANA-LAVELLE-ZHUANG-ROSENBEK-KUMAR-FIRMANSYAH-WOOD combination now modifies the Random Forest PD classifier of SOLANA-LAVELLE, which teaches backward stepwise selection, to remove the “least statistically important” variable as in WOOD, where such least statistically important variable is measured with respect to accuracy, specificity, and/or sensitivity, which are each performance metrics tracked by SOLANA-LAVELLE) returning the particular feature to the model when removing the particular feature worsens the model performance; and (WOOD, p. 3230, section 2.3: “The backward stepwise selection procedure consists of backward selection, followed by forward selection and iterates if necessary. ... The forward selection process checks whether removed variables should be added back into the model”; (EN): the SOLANA-LAVELLE-ZHUANG-ROSENBEK-KUMAR-FIRMANSYAH-WOOD combination now modifies the Random Forest PD classifier of SOLANA-LAVELLE, which teaches backward stepwise selection, to return features as in WOOD) repeating removal of each feature in the set of long-term acoustic features until the best performance of the model is achieved as measured by accuracy, a specificity, a sensitivity, or an area under the curve. (WOOD, p. 3230, section 2.3: “An alternate approach is backward selection. Starting with a model with all potential variables, each variable in the model is tested for exclusion from the model. ... The backward stepwise selection procedure consists of backward selection, followed by forward selection and iterates if necessary. Starting with a model with all potential variables, the backward selection process removes from the model the variable that is least statistically important (at α per cent significance).”; (EN): the SOLANA-LAVELLE-ZHUANG-ROSENBEK-KUMAR-FIRMANSYAH-WOOD combination now modifies the Random Forest PD classifier of SOLANA-LAVELLE, which teaches backward stepwise selection, to iteratively remove the “least statistically important” variable as in WOOD, where such least statistically important variable is measured with respect to accuracy, specificity, and/or sensitivity, which are each performance metrics tracked by SOLANA-LAVELLE) Before the effective filing date of the present application, it would have been obvious to one of ordinary skill in the art to combine the teachings of SOLANA-LAVALLE, with the teachings of ZHUANG, ROSENBEK, KUMAR, FIRMANSYAH, and WOOD as explained above. As disclosed by WOOD, one of ordinary skill would have been motivated to do so because backward stepwise selection “is appealing due to its popularity and availability in statistical software.” (p. 3228, section 1). Claim 12 depends from claim 8 and claims a system that corresponds to the method of claim 5, and is therefore rejected for the same reasons explained with respect to claims 5 and 8. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHAEL C LEE whose telephone number is (571)272-4933. The examiner can normally be reached M-F 12:00 pm - 8:00 pm ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Omar Fernandez Rivas can be reached at 571-272-2589. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MICHAEL C. LEE/Examiner, Art Unit 2128
Read full office action

Prosecution Timeline

Show 1 earlier event
Jun 18, 2025
Non-Final Rejection mailed — §103
Sep 12, 2025
Response Filed
Nov 17, 2025
Final Rejection mailed — §103
Feb 17, 2026
Request for Continued Examination
Feb 25, 2026
Response after Non-Final Action
May 12, 2026
Non-Final Rejection mailed — §103
Jul 21, 2026
Response Filed
Aug 12, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12645972
Performing Property Estimation Using Quantum Gradient Operation on Quantum Computing System
3y 7m to grant Granted Jun 02, 2026
Patent 12603081
METHOD AND SERVER FOR A TEXT-TO-SPEECH PROCESSING
4y 7m to grant Granted Apr 14, 2026
Patent 12602605
QUANTUM COMPUTER ARCHITECTURE BASED ON MULTI-QUBIT GATES
3y 11m to grant Granted Apr 14, 2026
Patent 12591915
METHODS AND SYSTEMS FOR DETERMINING RECOMMENDATIONS BASED ON REAL-TIME OPTIMIZATION OF MACHINE LEARNING MODELS
5y 0m to grant Granted Mar 31, 2026
Patent 12585743
INTERFACE ACCESS PROCESSING METHOD, COMPUTER DEVICE AND STORAGE MEDIUM
1y 6m to grant Granted Mar 24, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
62%
Grant Probability
88%
With Interview (+26.4%)
3y 3m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 153 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month