DETAILED ACTION
This communication is in response to the Amendments and Arguments filed on 8/6/2026.
Claims 1-20 and 33-35 are pending and have been examined.
All previous objections / rejections not mentioned in this Office Action have been withdrawn by the examiner.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments and Amendments
Regarding the previous the rejections under 35 U.S.C. § 101, applicant has amended independent claim 1 to include “apply, by the signal processing circuitry, the one or more metrics of speech as model input to a trained machine learning (ML) model to output at least one of a classification or prediction for evaluation of cognitive function for the subject, the trained predictive model having previously been fit with the one or more metrics of speech and trained on training data labeled to separate healthy subjects from other subjects according to cognitive function”. Therefore, 101 rejection is withdrawn for claims 1-10, 33, and 34. Claims 21-32 are cancelled by applicant.
Regarding the Applicant’s arguments for the rejections under 35 U.S.C. § 102, applicant has amended independent claim 1 and 11. Since Applicant’s arguments are directed towards the new amendment, the arguments are moot in view of new grounds for rejection.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) are: “audio input circuitry configured to receive” in claim 1, “signal processing circuitry configured to” in claims 1, 6, 8, 9, 10, 16, 33, and “classifier configured to” in claims 10, 20.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-20, 34, and 35 are rejected under 35 U.S.C. 103 as being unpatentable over Zaldua et al. (U.S. PG Pub No. 20240023877), hereinafter Zaldua, in view of “Connected Language in Late Middle-Aged Adults at Risk for Alzheimer’s Disease” by Mueller et al., hereinafter Mueller.
Regarding claim 1 Zaldua teaches:
A device for evaluating cognitive function based on speech, the device comprising: (P0045, System comprising a mobile device and a data processing apparatus for carrying out the method of detecting cognitive impairment.)
audio input circuitry configured to receive an audio signal provided by a subject; and (P0049, Receiving audio data representing recorded utterances.; P0155, The mobile device may also record the utterances made by the patient P and generate audio data. For this purpose, the mobile device may be provided with a microphone or an array of microphones.)
signal processing circuitry configured to: process the input signal to detect one or more metrics of speech of the subject, the one or more metrics of speech including a measure of semantic relevancy that defines a degree of overlap between reference content units of a picture and a description of the picture from the audio signal, (P0052, The received audio data may then be processed by a speech-to-text engine to produce a text transcription of the recorded utterances. The speech-to-text engine may be implemented locally or remotely in an external server.; P0055, Next, the text transcription of the recorded utterances may be processed to calculate a plurality of test variables.; P0116, Image description test.; P0123, Number of correct F-words.; P0151, In an image description text, the predetermined visual information may include an image, and text-based instruction asking the patient to describe as many things as they can see in the image within a time limit.)
wherein to detect the measure of semantic relevancy the signal processing circuitry generates a transcript of the description from the audio signal using automated speech recognition (ASR), and (P0052, The received audio data may then be processed by a speech-to-text engine to produce a text transcription of the recorded utterances. The speech-to-text engine may be implemented locally or remotely in an external server.)
algorithmically computes a semantic relevance score by accessing a predefined set of content units associated with a picture presented in a picture description task, comparing words in the transcript to the predefined set of content units including comparing the words in the transcript to respective word families corresponding to the predefined set of content units, identifying, from the transcript, content units of the predefined set represented in the transcript, and (P0055, Next, the text transcription of the recorded utterances may be processed to calculate a plurality of test variables.; P0018, Test variables may comprise one, two, or all of: mean number of correct words, mean percentage of incorrect words, and mean average correct word closeness.; P0089, Following list of test variables are found to be suitable: [Examples P0090-P0125].; P0116, Image description test.; P0123, Number of correct F-words.; P0151, In an image description text, the predetermined visual information may include an image, and text-based instruction asking the patient to describe as many things as they can see in the image within a time limit.)
applying, by the signal processing circuitry, the one or more metrics of speech as a model input to a trained machine learning (ML) model to output at least one of a classification or prediction for evaluation of cognitive function for the subject, the trained predictive model having previously been fit with the one or more metrics of speech and and trained on training data labeled to separate healthy subjects from other subjects according to cognitive function. (P0056, The plurality of test variables may be passed on to a trained detection model. The trained detection model, taking the plurality of test variables as input, may calculate an impairment possibility indicating a likelihood that the patient suffers from the cognitive impairment.; P0146, Each of the detection models and the final detection model may be implemented using any suitable Machine Learning algorithm.; P0157, The detection models require prior training with reference data. Each of the detection models may be trained separately with its own set of reference data.; P0158, As in an actual performance of the method of detecting cognitive impairment disclosed in the present application, for the purpose of preparing training data, the neuropsychological test may also be presented to the patient P1, P2 by displaying predetermined visual information and/or providing predetermined audible information to the patient P1, P2 prompting them to make utterances.)
Zaldua does not specifically teach:
dividing a number of the identified content units by a total number of words in the transcript, and
Mueller, however, teaches:
dividing a number of the identified content units by a total number of words in the transcript, and (Semantic Unit Idea Density: total number of semantic units divided by total number of words.)
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to divide content units by total number of words. It would have been obvious to combine the references because the use of density ratio is a known technique to yield a predictable result of measuring language efficiency and conciseness. (Mueller, Introduction).
Regarding claim 11 Zaldua teaches:
A computer-implemented method for evaluating cognitive function based on speech, the method comprising: (P0045, System comprising a mobile device and a data processing apparatus for carrying out the method of detecting cognitive impairment.)
receiving, with audio input circuitry, an input signal provided by a subject; (P0049, Receiving audio data representing recorded utterances.; P0155, The mobile device may also record the utterances made by the patient P and generate audio data. For this purpose, the mobile device may be provided with a microphone or an array of microphones.)
generating, via signal processing circuitry, a transcript from the input signal; (P0052, The received audio data may then be processed by a speech-to-text engine to produce a text transcription of the recorded utterances. The speech-to-text engine may be implemented locally or remotely in an external server.; P0055, Next, the text transcription of the recorded utterances may be processed to calculate a plurality of test variables.)
automatically extracting, via the signal processing circuitry, a plurality of speech features defining digital speech-based measures from the transcript; (P0018, Test variables may comprise one, two, or all of: mean number of correct words, mean percentage of incorrect words, and mean average correct word closeness.; P0089, Following list of test variables are found to be suitable: [Examples P0090-P0125].)
the plurality of speech features including a metric of semantic relevance for a picture description task, wherein the metric of semantic relevance is determined by accessing a predefined set of content units associated with a picture presented in the picture description task, comparing words in the transcript to the predefined set of content units, and (P0055, Next, the text transcription of the recorded utterances may be processed to calculate a plurality of test variables.; P0018, Test variables may comprise one, two, or all of: mean number of correct words, mean percentage of incorrect words, and mean average correct word closeness.; P0089, Following list of test variables are found to be suitable: [Examples P0090-P0125].; P0116, Image description test.; P0123, Number of correct F-words.; P0151, In an image description text, the predetermined visual information may include an image, and text-based instruction asking the patient to describe as many things as they can see in the image within a time limit.)
selecting, via signal processing circuitry, a restricted number of the plurality of speech features as predictors for model input, the predictors previously shown to be representative of vocabulary, language processing, and ability to convey relevant picture details; and (P0089, Following list of test variables are found to be suitable: [Examples P0090-P0125].; P0088, It may be desirable to choose test variables which, individually, provides significant predictive power.; P0116, Image description test.; P0123, Number of correct F-words.; P0151, In an image description text, the predetermined visual information may include an image, and text-based instruction asking the patient to describe as many things as they can see in the image within a time limit.)
applying, by the signal processing circuitry, the model input to a trained machine learning (ML) model to output at least one of a classification or prediction for evaluation of cognitive function for the subject, the trained predictive model having previously been fit with the plurality of speech features and trained on training data labeled to separate healthy subjects from other subjects according to cognitive function. (P0056, The plurality of test variables may be passed on to a trained detection model. The trained detection model, taking the plurality of test variables as input, may calculate an impairment possibility indicating a likelihood that the patient suffers from the cognitive impairment.; P0146, Each of the detection models and the final detection model may be implemented using any suitable Machine Learning algorithm.; P0157, The detection models require prior training with reference data. Each of the detection models may be trained separately with its own set of reference data.; P0158, As in an actual performance of the method of detecting cognitive impairment disclosed in the present application, for the purpose of preparing training data, the neuropsychological test may also be presented to the patient P1, P2 by displaying predetermined visual information and/or providing predetermined audible information to the patient P1, P2 prompting them to make utterances.)
Zaldua does not specifically teach:
computing a semantic relevance score by dividing a number of words in the transcript corresponding to the predefined set of content units by a total number of words in the transcript;
Mueller, however, teaches:
computing a semantic relevance score by dividing a number of words in the transcript corresponding to the predefined set of content units by a total number of words in the transcript; (Semantic Unit Idea Density: total number of semantic units divided by total number of words.)
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to divide content units by total number of words. It would have been obvious to combine the references because the use of density ratio is a known technique to yield a predictable result of measuring language efficiency and conciseness. (Mueller, Introduction).
Regarding claim 2 and 12 Zaldua view of Mueller teach claim 1 and 11.
Zaldua further teaches:
wherein the evaluation of the cognitive function comprises detection or prediction of future cognitive decline. (P0044, Present disclosure may enable scalable and accurate screening of cognitive impairment, which may be useful for large-scale screening for early signs of cognitive impairment.; P0081, By dividing the final impairment probability into three ranges using two predetermined thresholds, the indication may indicate an absence of impairment or one of two degrees of impairment. It should be understood that a finer division of the ranges of final impairment probability may be used. That is, three or more predetermined thresholds may be defined and applied to the final impairment probability, so that the indication may indicate four or more possible outcomes corresponding to varying degrees of cognitive impairment or non-impairment.)
Regarding claim 3 and 13 Zaldua view of Mueller teach claim 1 and 11.
Zaldua further teaches:
wherein the evaluation of the cognitive function comprises a prediction or classification of normal cognition, early mild cognitive impairment, mild cognitive impairment, or dementia. (P0082, It may be useful, from a health policy point of view, to provide fast and effective screening of mild cognitive impairment or early stage dementia. In this case, it may be enough to use two predetermined thresholds resulting in three possible outcomes, namely 1) no cognitive impairment, 2) mild cognitive impairment, and 3) dementia.)
Regarding claim 4 and 14 Zaldua view of Mueller teach claim 1 and 11.
Zaldua further teaches:
wherein the one or more metrics of speech of the subject comprises a metric of semantic relevance, word count, ratio of unique words to total number of words (MATTR), pronoun-to-noun ratio, propositional density, number of pauses during an audio speech recording within the input signal, or any combination thereof. (P0089, Following list of test variables are found to be suitable: [Examples P0090-P0125].; P0118, Number of nouns.; P0119, Number of verbs.; P0120, Ratio nouns/pronouns.)
Regarding claim 5 and 15 Zaldua view of Mueller teach claim 4 and 14.
Zaldua further teaches:
wherein the metric of semantic relevance measures a degree of overlap between a content of a picture and a description of the picture detected from the speech in the input signal. (P0116, Image description test.; P0123, Number of correct F-words.; P0151, In an image description text, the predetermined visual information may include an image, and text-based instruction asking the patient to describe as many things as they can see in the image within a time limit.)
Regarding claim 6 and 16 Zaldua view of Mueller teach claim 1 and 11.
Zaldua further teaches:
display an output comprising the evaluation. (P0186, Results of a method of the invention may be displayed to a user or stored in any suitable storage medium.)
Regarding claim 7 and 17 Zaldua view of Mueller teach claim 1 and 11.
Zaldua further teaches:
generates a wherein the notification element that comprises a display. (P0155, With reference to FIG. 8, the presentation of predetermined visual information and/or predetermined audio information may be achieved by a mobile device, which may be provided with a display and/or loudspeakers.)
Regarding claim 8 Zaldua in view of Mueller teach claim 7.
Zaldua further teaches:
cause the display to prompt the subject to provide a speech sample from which the input signal is derived. (P0147, The patient may be asked to make utterances in accordance with different neuropsychological tests, which utterances may be recorded to generate the audio data used in subsequent processing. In order to prompt the patient to make these utterances, the present method may further comprise displaying predetermined visual information.)
Regarding claim 18 Zaldua in view of Mueller teaches claim 11.
Zaldua further teaches:
prompting the subject to provide a speech sample from which the input signal is derived. (P0147, The patient may be asked to make utterances in accordance with different neuropsychological tests, which utterances may be recorded to generate the audio data used in subsequent processing. In order to prompt the patient to make these utterances, the present method may further comprise displaying predetermined visual information.)
Regarding claim 9 and 19 Zaldua view of Mueller teach claim 1 and 11.
Zaldua further teaches:
utilize at least one machine learning classifier to generate the evaluation of the cognitive function of the subject. (P0145, The calculation of the final impairment probability may be performed using the trained final detection model.; P0146, Detection model may be implemented using any suitable Machine Learning algorithm.)
Regarding claim 10 Zaldua view of Mueller teach claim 9.
Zaldua further teaches:
utilize a plurality of machine learning classifiers comprising a first classifier configured to evaluate the subject for a first cognitive function or condition and a second classifier configured to evaluate the subject for a second cognitive function or condition. (P0073, The text description may be processed to calculate first and second pluralities of test variables associated with first and second neuropsychological tests, and first and second impairment probabilities may be calculated by first and second trained detection models. … The final impairment probability may be calculated based on the first, second.)
Regarding claim 20 Zaldua view of Mueller teach claim 19.
Zaldua further teaches:
the at least one machine learning classifier comprises a first classifier configured to evaluate the subject for a first cognitive function or condition and a second classifier configured to evaluate the subject for a second cognitive function or condition. (P0073, The text description may be processed to calculate first and second pluralities of test variables associated with first and second neuropsychological tests, and first and second impairment probabilities may be calculated by first and second trained detection models. … The final impairment probability may be calculated based on the first, second.)
Regarding claim 34 Zaldua in view of Mueller teach claim 1.
Zaldua further teaches:
wherein to detect the one or more metrics of speech the signal processing circuitry executes computer-implemented algorithms that process the speech audio and transcript to identify linguistic and acoustic elements and to compute quantitative measures within the audio signal that relate to vocabulary including a ratio of unique words to total number of words (MATTR), a propositional density, a type-to-token ratio (TTR), and mean word length. (P0089, Following list of test variables are found to be suitable: [Examples P0090-P0125].; P0118, Number of nouns.; P0119, Number of verbs.; P0120, Ratio nouns/pronouns.; P0006, The language features in question are speaking rate, number of pause fillers (e.g. “ums” and “ahs”), the difficulty of words, or the parts of speech of words following the pause fillers.)
Regarding claim 35 Zaldua in view of Mueller teach claim 11.
Zaldua further teaches:
wherein the predictors comprise at least one of: ratio of unique words to total number of words (MATTR), pronoun-to-noun ratio, propositional density, parse tree height, mean length of word, type-to-token ratio, proportion of relevant details correctly identified in a picture description, duration of pauses relative to total speaking duration, and word count. (P0089, Following list of test variables are found to be suitable: [Examples P0090-P0125].; P0118, Number of nouns.; P0119, Number of verbs.; P0120, Ratio nouns/pronouns.; P0006, The language features in question are speaking rate, number of pause fillers (e.g. “ums” and “ahs”), the difficulty of words, or the parts of speech of words following the pause fillers.)
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 33 is rejected under 35 U.S.C. 103 as being unpatentable over Zaldua in view of Mueller and in further view of "Cognitive impairment screening using m-health: an android implementation of the mini-mental state examination (MMSE) using speech recognition" by Devos et al., hereinafter Devos.
Regarding claim 33 Zaluda teach claim 1.
Zaluda in view of Mueller does not specifically teach:
wherein the signal processing circuitry is further configured to determine semantic relevance scores for a plurality of speech samples obtained at different timepoints and implement the scores to generate or track a longitudinal change indication.
Devos, however, teaches:
wherein the signal processing circuitry is further configured to determine semantic relevance scores for a plurality of speech samples obtained at different timepoints and implement the scores to generate or track a longitudinal change indication. (Introduction, Assessing the cognitive impairment of a person on a periodically. … A score is linked to these results and an indication of the person’s cognitive impairment (for instance, the state of dementia, if appropriate) is given.)
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to track cognitive decline over time. It would have been obvious to combine the references because assessing the cognitive impairment of a person on a periodically basis is needed to adapt the care to the level of impairment. (Devos, Introduction)
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Vairavan et al. (U.S. PG Pub No. 20200327882): System and method for detecting cognitive decline using speech analysis.
Kim et al. : (U.S. PG Pub No. 11759145) Technique for identifying dementia based on voice data.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DANIEL WONSUK CHUNG whose telephone number is (571)272-1345. The examiner can normally be reached Monday - Friday (7am-4pm)[PT].
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, PIERRE-LOUIS DESIR can be reached at (571)272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DANIEL W CHUNG/Examiner, Art Unit 2659
/EDGAR X GUERRA-ERAZO/Primary Examiner, Art Unit 2656