Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendments
Applicant’s amendments to the specification, drawings, and claims have overcome the objections and rejections under 35 USC § 112.
Response to Arguments Regarding Claim Rejections - 35 USC § 101
Applicant’s argument that the claimed invention provides an improvement to a language model is not persuasive. The claim does not recite a specific improvement to the architecture, training, or operation of the language model itself. Rather, the language model is recited generically and is used as a tool to evaluate transcribed text and generate evaluation feedback. The claim likewise does not require a model specifically trained to perform the claimed invention or identify any particular training data, training technique, model parameters, or modification that improves the model’s functioning. Instead, the claim is directed to the use of generic computing and AI components to perform the claimed evaluation and classification process. Accordingly, Applicant’s arguments do not demonstrate that the judicial exception is integrated into a practical application or that the claim recites significantly more than the abstract idea, and the rejection under 35 USC § 101 is maintained.
Response to Arguments Regarding Claim Rejections - 35 USC § 103
Regarding Claims 1-20, Applicant’s remarks filed on July 13, 2026 have been fully considered. The examiner has reviewed each of the Applicant’s arguments and respectfully responds as follows.
Applicant’s arguments primarily address the cited references individually, whereas the rejection is based
upon the combined teachings of the cited references, where each reference is relied upon only for the
teachings for which it is cited.
Applicant’s Arguments
Applicant argues that Agbavor uses GPT3 merely as a feature encoder whose output is a numerical vector, rather than an evaluation, opinion, scoring rationale, or fluency assessment.
Examiner’s Response
Applicant’s characterization of Agbavor is acknowledged. Agbavor is not relied upon, by itself, as teaching LLM generated fluency evaluation feedback. Agbavor is relied upon for converting speech into transcribed text, extracting transcript based linguistic embeddings and acoustic features, and using those features for dementia classification. Yao and Weston are relied upon for generating evaluation feedback through a generative language model. The rejection does not equate Agbavor’s embeddings with the claimed evaluation feedback, rather it combines Agbavor’s feature extraction and classification framework with the prompt based evaluation mechanisms taught by Yao and Weston.
Applicant’s Arguments
Agbavor does not require the LLM to apply evaluation criteria, compare the transcript against score anchors, or consider whether important image elements were included, assess consistency or repletion or comment on wording accuracy. By contrast the claimed LLM receives a structured fluency evaluation request, generates evaluation feedback, and linguistic features extracted from the transcript and that feedback.
Examiner’s Response
Applicant’s characterization of Agbavor individually is acknowledged. However, the rejection does not equate Agbavor’s embedding with the claimed evaluation feedback or rely on Agbavor alone. Yao teaches prompting a generative AI engine to evaluate a user’s response and provide explanations or corrections. Weston further teaches supplying a transcript and structured clinical rating criteria to a generative model to generate feedback concerning recalled key elements.
The combined system preserves the different stages. Yao and Weston generate textual evaluation feedback and Agbavor’s text embedding technique is then applied to the transcript and evaluation feedback to produce linguistic features for use with acoustic features in dementia classification. Thus, the rejection does not rely merely on corrections learned from Agbavor’s transcript embedding. It adds an intermediate LLM evaluation stage before feature extraction and classification. A person of ordinary skill in the art would have been motivated to add this intermediate evaluation stage to provide Agbavor’s classifier with higher level linguistic information not expressly identified in the raw transcript. Yao teaches that prompting an LLM produces evaluation and explanatory information concerning a user’s response, while Weston teaches that structured LLM generated clinical ratings may be processed by a downstream computer and used in a diagnostic process. Applying Agbavor’s text embedding technique to this additional textual information would predictably enrich the linguistic feature representation. Both the transcript and LLM feedback are textual information. Agbavor’s embedding process is already designed to convert text into numerical linguistic features, so applying the same process to the related textual feedback would use the technique for its established purpose.
Applicant’s Arguments
Applicant’s argues that a language model may generate embeddings in one context and natural language feedback in another, but those uses are not interchangeable. An embedding is a mathematical representation whereas a generative LLM responding to Applicant’s prompt produces feedback expressly directly to fluency or clarity.
Examiner’s Response
Applicant’s distinction is acknowledged, but the rejection does not treat the embedding and evaluation feedback as interchangeable. They perform separate, sequential functions in the combination. Yao and Weston’s generative model first produces textual evaluation feedback concerning the user’s response. Agbavor’s embedding technique then converts the transcript and that feedback into numerical linguistic features for dementia classification.
Applicant’s Arguments
Even if LLM was known to generate feedback in other contexts, the cited art does not teach using that feedback as an intermediate diagnostic feature source in a cognitive impairment classifier.
Examiner’s Response
Although no individual reference teaches the complete limitation, the combined references suggest it. Yao and Weston teach generating textual evaluation feedback concerning a user’s response, while Agbavor teaches converting textual linguistic information into embeddings used for dementia classification. It would be obvious to apply Agbavor’s embedding technique to the generated feedback, thereby using that feedback as an intermediate linguistic feature source for cognitive impairment classifier. This would predictably provide the classifier with additional diagnostic information and improve the assessment.
Applicant’s Arguments
Applicant argues that Agbavor predicts dementia using acoustic features and transcript embeddings but does not teach generating fluency evaluation feedback from a prompt containing the recited evaluation criteria. Applicant further argues that Agbavor does not extract linguistic features from that feedback and use those features with acoustic features to classify the user into dementia, mild cognitive impairment or normal groups. Also, they mention Agbavor’s transcript embeddings are not equivalent to generating evaluation feedback from a structured prompt and using that feedback as a classification feature source.
Examiner’s response
Applicant’s characterized of Agbavor individually is acknowledged. But the rejection relies on the combined references. Yao teaches prompting an LLM to evaluate a user’s response and generate explanations or corrections, while Weston teaches supplying a transcript and structured clinical evaluation criteria to a generative model to obtain evaluation results. Agbavor teaches converting textual linguistic information into embeddings and using linguistic and acoustic features together for dementia classification. This combination would create feedback derived linguistic features that can be used with Agbavor’s acoustic features in its cognitive impairment classifier while preserving the distance evaluation, feature extraction, and classification stages.
Applicant’s Arguments
Yao’s feedback is pedagogical, not diagnostic. Yao concerns language teaching and generates corrections intended to improve second language performance, rather than diagnostic feedback used in a cognitive impairment pipeline.
Examiner’s Response
The claim does not require the LLM generated feedback itself to constitute a diagnosis or clinical opinion. It requires or states “evaluation feedback related to fluency” and subsequently requires cognitive impairment classification based on acoustic and linguistic features. Yao is relied upon for the narrower function of prompting an generative AI engine to evaluate a user’s spoken response, identify language mistakes, and provide explanations. These functions are reasonably pertinent to the problem of automatically evaluating linguistic characteristic in a spoken response.
Weston supplies the clinical assessment context and structured evaluation criteria, while Agbavor supplies the extraction of textual linguistic features and dementia classification framework. That fact that Yao applies its technique in a pedagogical context does not negate its relevance to automatically evaluation linguistic characteristics of a transcribed spoken response.
Applicant’s Arguments
Applicant argues that Yao’s identification of language learning mistakes is different than from the claimed fluency evaluation feedback. Yao does not teach a structured prompt containing diagnostic evaluation criteria, score ranges, score units, or scored transcripts examples, nor does Yao transform the resulting feedback into classifier input features.
Examiners response
Yao is not relied upon alone for the detailed evaluation criteria or subsequent feature extraction. Weston teaches providing a transcript and structured clinical rating instructions to a generative model, including scoring guidance, story elements, examples responses, and evaluation that may use a numerical scale 1-10 . Yao teaches the broader prompt based generation of language evaluation feedback, while Agbavor teaches converting textual linguistic information into classifier input embeddings. Additionally, the claim requires at least one evaluation criterion text selected from the listed alternatives, not every listed criterion. It would have been obvious to apply Agbavor’s embedding technique to the transcript and the structured textual feedback generated according to Yao and Weston, thereby producing linguistic features for use in Agbavor’s dementia classifier. Thus rejection relies on the combined teachings rather than Yao alone.
Applicant’s Arguments
Applicant argues that a person or ordinary skill would not have imported Yao’s language learning correction mechanism into Agbover’s dementia prediction system in the claimed manner. Agbavor represents transcripts as embeddings for machine learning classification, whereas Yao maintains an instructional dialogue and improves engagement in learning secondary language. These references address different problems and use LLM outputs for different purposes. Even if combined, applicant contends that the results would be a speech processing or chatbot system providing language learning corrections, not a system that treats LLM feedback as diagnostic information, converts the feedback into linguistic features, and combines those features with acoustic features for dementia classification.
Examiners response
The proposed combination does not substitute Yao’s educational chatbot for Agbavor’s dementia prediction system or import Yao’s educational purpose. Agbavor remains the primary system, including its transcription, acoustic feature extraction, linguistic embeddings, and dementia classifier. Yao contributes the narrower technical technique of prompting the LLM to evaluate a transcribed spoken response and generate textual feedback. Weston supplies the clinical assessment connection by teaching that a patient transcript and structured clinical rating instructions may be provided to a generative model to produce evaluation used in diagnostic problem. Weston’s rating criteria include scoring guidance and story elements and its output identifies whether key elements were recalled. Because the claim requires feedback concerning at least one listed characteristic, Weston evaluation of key element inclusion is sufficient. One would have been motivated to add Yao and Weston prompt based evaluation stage to Agbavor to obtain higher level linguistic information concerning mistakes, disfluencies, and omitted elements that is not expressed in raw text. Applying Agbavors’ embedding process to the transcript and textual evaluation feedback would generate additional linguistic feature for combination with Agbavor’s acoustic features . The resulting system therefore would remain Agbavor’s dementia classification system with intermediate LLM evaluation stage, rather than becoming a learning chatbot.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim [1, 2, 4, 7, 8, 9, 11, 14, 15, 17, 20] rejected under 35 U.S.C. 101 because the claimed invention is directed to a mental process without significantly more.
Regarding claim 1, an electronic apparatus comprising: a communication circuit configured to acquire utterance voice of a user; a memory storing at least one instruction related to an artificial intelligence model; and one or more processors functionally connected to the communication circuit and the memory, the one or more processors being configured to execute the at least one instruction to:
[A person can listen to the user speaking]
convert the utterance voice into transcribed text;
[A person can listen to the speech and write down the spoken words]
generate a first prompt related to a fluency evaluation request based on the transcribed text corresponding to the utterance voice of the user, the first prompt including the transcribed text and at least one evaluation criterion text selected from among an explanatory phrase related to the transcribed text, a score range and a score unit of the fluency evaluation, and example transcribed text associated with a plurality of scores within the score range;
[A person can formulate a question or instruction for evaluating the fluency of the transcribed speech, A person can review the transcribed together with written evaluation, a person can read an explanatory phrase describing or relating to the subject of the transcript, a person can use a predetermined scoring scale and score unit when evaluating fluency, A person can compare the user’s transcript with example transcripts that have already been assigned different scores];
input the first prompt to a large language model and acquire, from the large language model, evaluation feedback related to fluency, the evaluation feedback being related to at least one of inclusion of key elements, consistency, term repetition, and wording accuracy in the transcribed text;
[A person can review the transcript under the stated criteria and provide feedback concerning fluency. A person can read the transcript and identify whether key information is present, whether it is consistent, whether terms are repeated and whether the wording is accurate. ];
extract acoustic features based on the utterance voice;
[A person can listen to the voice and identify audible characteristics such as pauses, pitch, etc.].
extract linguistic features based on transcribed text corresponding to the utterance voice and the evaluation feedback; and
[A person can read the transcript and feedback and identify linguistic features such as word choice, repetition, sentence structure, and fluency].
classify the user into one of a dementia group, a mild cognitive impairment group, or a normal group based on the acoustic features and the linguistic features.
[A person can consider the identified acoustic and linguistic information and categorize the subject according to the specified groups].
As described above, these limitations can be carried out as a series of mental steps.
This judicial exception is not integrated into a practical application because the only additional
elements recited are a large language model, an interface module, first extraction unit, second
extraction unit, classification module, a memory, a processor, and a computer program executing on a
device, and these additional elements are nothing more than instructions to apply the mental process
using a general-purpose software model and general-purpose hardware.
These claims do not include additional elements that are sufficient to amount to significantly
more than the judicial exception because, as described above, the only additional elements recited are a
large language model, an interface module, first extraction unit, second extraction unit, classification
module, a memory, a processor, and a computer program executing on device, and these additional
elements are nothing more than instructions to apply the mental process using general-purpose
software and hardware.
Regarding claim 2 and 9 recite a system and a method wherein the utterance voice includes a
voice of the user that describes a painting or photo. [person can speak without a system and describe a
painting or photo].
These additional limitations do not prevent the process from being carried out as a mental
process.
Regarding claim 4 and 11 recite electronic apparatus of claim 1 and 8 , wherein the plurality of scores include a lowest score and a highest score within the score range [person can determine a score
whether it should be rated as high or low].
These limitations can be carried out as a series of mental steps.
This judicial exception is not integrated into a practical application because the only additional
elements recited are an electronic system device, and these additional elements are
nothing more than instructions to apply the mental process using the hardware.
The claims do not include additional elements that are sufficient to amount to significantly more
than the judicial exception because, as described above, the only additional elements recited are an
electronic system executing on a device, and these additional elements are nothing
more than instructions to apply the mental process using the hardware.
Regarding claim 7, the electronic apparatus of claim 1, wherein the one or more processors are further configured to generate a test opinion that summarizes a result of the classification on the cognitive impairment and the evaluation feedback through the large language model.
[This summary of the results can be done by a person].
As described above, these limitations can be carried out as a series of mental steps.
This judicial exception is not integrated into a practical application because the only additional
elements recited is a large language model, and this additional element is nothing more than instructions to apply the mental process using the hardware and software.
These claims do not include additional elements that are sufficient to amount to significantly
more than the judicial exception because, as described above, the only additional elements recited is
a large language model and this additional element is nothing more than instructions to apply the mental process using general-purpose software and hardware.
Regarding claim 8, a computer-implemented method of classifying cognitive impairment, the method comprising: acquiring, by an electronic apparatus, utterance voice of a user;
[A person can listen to the user speaking]
converting, by one or more processors of the electronic apparatus, the utterance voice into transcribed text;
[A person can listen to the speech and write down the spoken words]
generating, by the one or more processors, a first prompt related to a fluency evaluation request based on transcribed text corresponding to utterance voice of a user, the first prompt including the transcribed text and at least one evaluation criterion text selected from among an explanatory phrase related to the transcribed text, a score range and a score unit of the fluency evaluation, and example transcribed text rated with a plurality of scores within the score range;
[A person can formulate a question or instruction for evaluating the fluency of the transcribed speech, A person can review the transcribed together with written evaluation, a person can read an explanatory phrase describing or relating to the subject of the transcript, a person can use a predetermined scoring scale and score unit when evaluating fluency, A person can compare the user’s transcript with example transcripts that have already been assigned different scores];
acquiring, by inputting the first prompt to a large language model, evaluation feedback related to fluency in response to the first prompt through [[a]] the large language model, the evaluation feedback being related to at least one of inclusion of key elements, consistency, term repetition, and wording accuracy in the transcribed text;
[A person can review the transcript under the stated criteria and provide feedback concerning fluency. A person can read the transcript and identify whether key information is present, whether it is consistent, whether terms are repeated and whether the wording is accurate. ];
extracting, by the one or more processors, acoustic features and linguistic features based on the utterance voice, the transcribed text, and the evaluation feedback;
[A person can listen to the voice and identify audible characteristics such as pauses, pitch, etc.].
[A person can read the transcript and feedback and identify linguistic features such as word choice, repetition, sentence structure, and fluency].
and classifying, by the one or more processors, the user into one of a dementia group, a mild cognitive impairment group, or a normal group based on the acoustic features and the linguistic features.
[A person can consider the identified acoustic and linguistic information and categorize the subject according to the specified groups].
As described above, these limitations can be carried out as a series of mental steps.
This judicial exception is not integrated into a practical application because the only additional
elements recited are a large language model, an interface module, first extraction unit, second
extraction unit, classification module, a memory, a processor, and a computer program executing on a
device, and these additional elements are nothing more than instructions to apply the mental process
using a general-purpose software model and general-purpose hardware.
These claims do not include additional elements that are sufficient to amount to significantly
more than the judicial exception because, as described above, the only additional elements recited are a
large language model, an interface module, first extraction unit, second extraction unit, classification
module, a memory, a processor, and a computer program executing on device, and these additional
elements are nothing more than instructions to apply the mental process using general-purpose
software and hardware.
Regarding claim 14, the method of claim 8, further comprising: summarizing, by the one or more processors, a result of the classification on the cognitive impairment and the evaluation feedback through the large language model to generate a test opinion.
[This summary of the results can be done by a person].
As described above, these limitations can be carried out as a series of mental steps.
This judicial exception is not integrated into a practical application because the only additional
elements recited is a large language model, and this additional element is nothing more than instructions to apply the mental process using the hardware and software.
These claims do not include additional elements that are sufficient to amount to significantly
more than the judicial exception because, as described above, the only additional elements recited is
a large language model and this additional element is nothing more than instructions to apply the mental process using general-purpose software and hardware.
Regarding claim 15, an electronic apparatus comprising: a memory in which at least one instruction related to an artificial intelligence model is stored; and a processor functionally connected to the memory, the processor executing the at least one instruction to: acquire utterance voice of a user;
[A person can listen to the user speaking]
convert the utterance voice into transcribed text;
[A person can listen to the speech and write down the spoken words]
generate a first prompt related to a fluency evaluation request based on transcribed text corresponding to utterance voice of a user, the first prompt including the transcribed text and at least one evaluation criterion text selected from among an explanatory phrase related to the transcribed text, a score range and a score unit of the fluency evaluation, and example transcribed text rated with a plurality of scores within the score range;
[A person can formulate a question or instruction for evaluating the fluency of the transcribed speech, A person can review the transcribed together with written evaluation, a person can read an explanatory phrase describing or relating to the subject of the transcript, a person can use a predetermined scoring scale and score unit when evaluating fluency, A person can compare the user’s transcript with example transcripts that have already been assigned different scores];
input the first prompt to a large language model, and acquire evaluation feedback from the large language model, the evaluation feedback being related to at least one of inclusion of key elements, consistency, term repetition, and wording accuracy in the transcribed text;
[A person can review the transcript under the stated criteria and provide feedback concerning fluency. A person can read the transcript and identify whether key information is present, whether it is consistent, whether terms are repeated and whether the wording is accurate. ];
extract acoustic features and linguistic features based on the utterance voice, the evaluation feedback, and the transcribed text;
[A person can listen to the voice and identify audible characteristics such as pauses, pitch, etc.].
[A person can read the transcript and feedback and identify linguistic features such as word choice, repetition, sentence structure, and fluency].
and classify a cognitive impairment group to which the user belongs the user into one of a dementia group, a mild cognitive impairment group, or a normal group based on the acoustic features and the linguistic features.
[A person can consider the identified acoustic and linguistic information and categorize the subject according to the specified groups].
As described above, these limitations can be carried out as a series of mental steps.
This judicial exception is not integrated into a practical application because the only additional
elements recited are a large language model, an interface module, first extraction unit, second
extraction unit, classification module, a memory, a processor, and a computer program executing on a
device, and these additional elements are nothing more than instructions to apply the mental process
using a general-purpose software model and general-purpose hardware.
These claims do not include additional elements that are sufficient to amount to significantly
more than the judicial exception because, as described above, the only additional elements recited are a
large language model, an interface module, first extraction unit, second extraction unit, classification
module, a memory, a processor, and a computer program executing on device, and these additional
elements are nothing more than instructions to apply the mental process using general-purpose
software and hardware.
Regarding claim 17, the electronic apparatus of claim 15, further comprising a communication module, wherein the processor executes the at least one instruction to: acquire the utterance voice from an external electronic apparatus through the communication module and provide the test opinion to the external electronic apparatus through the communication module, to the external electronic apparatus through the communication module, a test opinion generated by summarizing a result of the classification on the cognitive impairment and the evaluation feedback through the large language model.
[This summary of the results can be done by a person and an opinion can be expressed].
As described above, these limitations can be carried out as a series of mental steps.
This judicial exception is not integrated into a practical application because the only additional
elements recited are an communication module and a large language model, and these additional elements are nothing more than instructions to apply the mental process using the hardware and software.
These claims do not include additional elements that are sufficient to amount to significantly
more than the judicial exception because, as described above, the only additional elements recited are
an communication module and a large language model and these additional elements are nothing more than instructions to apply the mental process using general-purpose software and hardware.
Regarding claim 20, the electronic apparatus of claim 15, wherein the processor executes the at least one instruction to: summarize a result of the classification on the cognitive impairment and the evaluation feedback through the large language model to generate a test opinion.
[This summary of the results can be done by a person].
As described above, these limitations can be carried out as a series of mental steps.
This judicial exception is not integrated into a practical application because the only additional
elements recited is a large language model, and this additional element is nothing more than instructions to apply the mental process using the hardware and software.
These claims do not include additional elements that are sufficient to amount to significantly
more than the judicial exception because, as described above, the only additional elements recited is
a large language model and this additional element is nothing more than instructions to apply the mental process using general-purpose software and hardware.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims [1, 8, 15] are rejected under 35 U.S.C. 103 as being unpatentable
over Agbavor( Agbavor, F. and Liang, H., 2022. Predicting dementia from spontaneous speech
using large language models. PLOS digital health, 1(12), p.e0000168.) in view of Yao (US
.2024/0321131 A) in view of and in further view of Weston (US2025/0132036A1).
Regarding claim 1,
Agbavor teaches
An electronic apparatus comprising:
convert the utterance voice into transcribed text;
[Page 10, lines 1-2 “GPT-3 based text embeddings are afterwards derived from the transcribed text obtained via wav2vec2.”]; [Page 9 Section : Text embeddings from GPT-3 “The waveform is
then tokenized using Wav2Vec2Tokenizer and if necessary, divided them into smaller chunks
(with the maximum size of 100,000 in our case) to fit into memory, which is subsequently fed
into the Wav2Vec2ForCTC (a wav2vec model for speech recognition) and decoded as text
transcripts.”];
extract acoustic features based on the utterance voice;
[ Page 10, lines21- 24 “In this work, acoustic features are extracted directly from speech using OpenSMILE (open-source Speech and Music Interpretation by Large-space Extraction), a widely used open- source toolkit for audio feature extraction and classification of speech and music signals [39]”];
extract linguistic features based on transcribed text and the
[Page 3, Fig1. Schematic showing two different feature representations that are derived from speech- A.
The acoustic features are engineered to capture the acoustic characteristics of speech and
therefore the pathological speech behavior. B. The linguistic features, represented as text
embeddings, are derived from the transcribed text. Central to our proposed approach is the
GPT-3 based text embeddings (shaded), which entail meaningful vector representations that can
capture lexical, syntactic, and semantic properties for dementia classification] ;
classify the user into one of a dementia group, a mild cognitive impairment group, or a normal group based on the acoustic features and the linguistic features.
[Page 11, lines 5- 9 “The model may use acoustic features from speech, linguistic features
(embeddings) from transcribed speech, or both. As such, we use (1) the acoustic features
extracted from speech audio data, (2) the text embeddings from each GPT-3 base model
(Babbage or Ada), and (3) the combination of both as inputs for three different kinds of
commonly used machine learning models, including Support Vector Classifier (SVC), Random
Forest (RF), and Logistic Regression (LR)”]; [Page 3, lines 1-10 “We report the results from two
tasks which include AD(Alzheimer’s Dementia) vs non-AD classification and AD severity prediction using a subject’s MMSE score. For the classification task, either the acoustic features or GPT-3 embeddings (Adaand Babbage) or both are fed into a machine-learning model such as support vector classifier
(SVC), logistic regression (LR) or random forest (RF). As a comparison, we further perform
finetuning on the GPT-3 model to see if there is any advantage over the GPT-3 embedding”];
[Page 3, lines 12-17 “In this section we present the AD classification results between AD and
non-AD (or healthy control) subjects based on different features: our proposed GPT-3 based text
embeddings, the acoustic features, and their combination. We also benchmark the GPT-3 based
text embeddings against the mainstream fine-tuning approach. We show that the GPT-3 based
text embeddings considerably outperform both the acoustic feature-based approach and the
fine-tuned model”]; [ Page 5, lines 1-4 “To evaluate whether the acoustic features and the text
embeddings can provide complementary information to augment the AD classification, we
combine the acoustic features from speech audio data and the GPT-3 based text embeddings by
simply concatenating them”]; [ Page 11, line 18- 23 “MMSE is perhaps the most common measure for assessing the severity of AD. We perform regression analysis using both the acoustic features and text embeddings from GPT-3 (Ada and Babbage) to predict the MMSE score. The scores normally range from 0 to 30, with scores of 26 or higher being considered normal [3]. A score of 20 to 24 suggests mild
dementia, 13 to 20 suggests moderate dementia, and less than 12 indicates severe dementia. As
such, the prediction is clipped to a range between 0 and 30”].
However, Agbavor does not teach
a communication circuit configured to acquire utterance voice of a user;
a memory storing at least one instruction related to an artificial intelligence model; and one or more processors functionally connected to the communication circuit and the memory,
the one or more processors being configured to execute the at least one instruction to:
generate a first prompt related to a fluency evaluation request based on the transcribed text corresponding to the utterance voice of the user, the first prompt including the transcribed text and at least one evaluation criterion text selected from among an explanatory phrase related to the transcribed text, a score range and a score unit of the fluency evaluation, and example transcribed text associated with a plurality of scores within the score range;
input the first prompt to a large language model and acquire, from the large language model, evaluation feedback related to fluency, the evaluation feedback being related to at least one of inclusion of key elements, consistency, term repetition, and wording accuracy in the transcribed text;
But Yao teaches
a communication circuit configured to acquire utterance voice of a user;
[0032 “Audio input device 112 can receive user-provided audio messages and provide these messages to audio-text conversion engine 106.”];
a memory storing at least one instruction related to an artificial intelligence model;
[0140 “FIG. 8 presents an exemplary computing system that facilitates an AI-based language learning partner chatbot, in accordance with an aspect of the present disclosure. In this example, a computing system 800 can include a processor 802, a memory device 804, and a storage device 808.”]
[0142 “Storage device 808 can store data 230 as well as computer-executable instructions which when executed by processor 802 can cause processor 802 to implement a number of functions and features.”];
and one or more processors functionally connected to the communication circuit and the memory, the one or more processors being configured to execute the at least one instruction to:
[0142 “Storage device 808 can store data 230 as well as computer-executable instructions which when executed by processor 802 can cause processor 802 to implement a number of functions and features.”];
generate a first prompt related to a fluency evaluation request based on the transcribed text corresponding to the utterance voice of the user, the first prompt including the transcribed text
[0030 “In one embodiment, generative AI engine 102 can include an LLM. This LLM can generate coherent and contextually relevant text, which can form the basis of the conversation content for chatbot system 100, and answer questions based on a given context or passage of text. In general, generic LLMs can answer or attempt to answer any question or request (which are referred to as "prompt"). LLMs are typically not entirely intuitive about the exact nuances or specifics the user might be interested in, such as the dialogue carried out by a trained language learning partner”];
[0031 “In order to induce the LLM in generative AI engine 102 to provide the desired response, or to initiate a conversation with the appropriate messages, one aspect of the present disclosure uses a prompt engine 104 to generate the appropriate prompts to the LLM, such that the LLM can provide a desired output. Specifically, prompt engine 104 can generate prompts to cause generative AI engine 102 to initiate a dialogue session or to respond to a user-provided message”];
[0033 “During operation, user 120 can start the language learning process by saying a command to audio sub-system 108, which relays the audio signal to audio-text conversion engine 106. Audio-text conversion engine 106 in turn converts the voice command to a text command and passes the text command to prompt engine 104. Subsequently, prompt engine 104 can generate a prompt and transmit the prompt to generative AI engine 102, which responds with the desired message to be played back to user 120.”];
[0037 “In some aspects, the prompt engine is responsible for creating, based on the user audio messages, prompts to the generative AI engine to induce the AI engine to provide desired responses”]; [0070 “Note that if the user pronounces any of those words incorrectly, the audio-text conversion engine can recognize and convert such mispronounced word to an incorrectly spelled word, which is then included in the prompt sent to the generative AI engine. For example, if the user pronounces
"Gracias" as "Graziaz," the prompt engine can include this mispronounced (which results in mis-
spelling) word in the prompt to the generative AI engine. As a result, the generative AI engine
can provide a message to help the user correct their mistake”];
input the first prompt to a large language model and acquire, from the large language model, evaluation feedback
[0069 “As a result, the generative AI engine can then evaluate the user's response and provide explanation and correction if necessary.”] [0070 “Note that if the user pronounces any of those words incorrectly, the audio-text conversion engine can recognize and convert such mispronounced word to an incorrectly spelled word, which is then included in the prompt sent to the generative AI engine. For example, if the user pronounces “Gracias” as “Graziaz,” the prompt engine can include this mispronounced (which results in mis-spelling) word in the prompt to the generative AI engine. As a result, the generative AI engine can provide a message to help the user correct their mistake.” where identified mistakes, explanations, and corrections reasonably constitute evaluation feedback];
[0120 “One of the key features of the present system is that the chatbot can evaluate the user's response and provide specific feedback to help the user correct any potential mistakes. To do so, the prompt engine can generate optimized prompts based on the user's response, which can cause the generative AI engine to provide specific feedback for the user”].
It would have been obvious to one of ordinary skill in the art before the effective filing
date of the claimed invention to combine the teachings Agbavor with the teaching of Yao because doing so would enable Agbavor’s system to obtain additional textual information identifying mistakes and explaining linguistic characteristics of the user’s response. This would thereby produce linguistic features based on both sources for use with Agbavor’s acoustic features. This would predictably enrich the classifier’s linguistic representation and improve the diagnostic assessment.
Agbavor in view of Yao do not teach
generate a first prompt
input the first prompt to a large language model and acquire, from the large language model,
But Weston teaches
generate a first prompt, the first prompt including the transcribed text and at least one evaluation criterion text selected from among an explanatory phrase related to the transcribed text,
[“0133 “In step S506, the SOP encoding 506 and the rating SOP encoding 507 are
provided as inputs to a generative ML model 508 along with a transcript 510]; [0136 “As before,
these inputs can be provided as inputs to the model in a prompt optionally with instructions to
format the output in a structured form, such as JSON or XML. The generative ML model 508 may
be prompted to provide the output in the form of a "rating sheet", although the output could be
in various forms of natural language or structured Text"]; [0111 "Segmentation" in this context
refers to the grouping or mapping of different parts of a transcription to corresponding
questions and answers in template ( e.g., SOP) for performing the clinical assessment”] ; [0113
“The method 400 uses an SOP 402, which is a template for administering a clinical assessment
comprising a plurality of sections corresponding to different clinical questions. A first section
comprises the "count to five" test described previously. A second section comprises a story
recall test, wherein the interviewer tells a patient a story then asks the patient to recall some of
the story, or as many aspects of the story as possible”]; [0128” The method 500 uses a similar
approach to the methods 200, 300 and 400, wherein an SOP 502 is optionally fed into the SOP
encoder 204 to provide an SOP encoding 506 in steps S502 and S504, respectively, as described
previously. The SOP 502 is identical to the SOP 402 and provides instructions for administering
the story recall test”]; [0129 “In the method 500, a rating SOP 503 is also optionally encoded by
the SOP encoder 204 into a rating SOP encoding 507 in steps S502 and S504, respectively”];
[0130 “As shown, the rating SOP encoding 507 comprises an instructions sub-section and a text
string entry providing instructions for scoring the task. The instructions can include guidance
such as "allow paraphrases". A "story elements" subsection is provided that lists each aspect of the story. The example above has been limited to a single story element ("Allison") for
simplicity, though it would be appreciated that further elements may be provided in practice.
Guidance for scoring the particular story element is also provided in the form of a "scoring
guidance" subsection. An "output_ schema" subsection is included to provide a format for the
story element. The output schema conditions the model to provide its output in a particular
format. In this case, 'rating' the story recall means producing a report that includes the
element index and whether or not that element was recalled. A schema, or template, has
children (a type and description) that tells you (a) what kind of data should be populated in
this field, and (b) a description of what the field means. A "recalled" section provides a
description for the evaluation("whether this element was recalled") and a type of the
evaluation ("boo!"). In other examples, the evaluation could be measured in other ways, such
using a scale of 1 to 10 which may be reflected in the "type" subsection"]; [0142 “The rating SOP
encoding 507 can include instructions on how to interpret or adjust an evaluation based on
disfluencies of the patient”].
input the first prompt to a large language model and acquire, from the large language model,
[0133 “In step S506, the SOP encoding 506 and the rating SOP encoding 507 are provided as inputs to a generative ML model 508 along with a transcript 510.”];
[0137 “In step S508, the generative ML model analyses the provided inputs and provides an output rating sheet 512.”];
[0142 “The rating SOP encoding 507 can include instructions on how to interpret or adjust an evaluation based on disfluencies of the patient. Alternatively, or in addition, the generative ML model 508 can be trained with positive training data corresponding to patients diagnosed with the health condition in question, in order assess the relevance of disfluencies (or other speech features) to each criteria in the SOP 502.”];
It would have been obvious to one of ordinary skill in the art before the effective filing
date of the claimed invention to combine the teachings Agbavor in view of Yao with Weston’s structured clinical rating prompt because Weston teaches providing a patient transcript and evaluation criteria to an LLM to generate a clinical rating identifying recalled key elements and accounting for disfluencies. This combination would enable Yao’s prompt based feedback mechanism to produce clinically relevant and consistency structured evaluation information. Applying Agbavor’s text embedding technique to that feedback would predictably provide additional linguistic features for its dementia classifier, thereby improving interpretability and diagnostic.
Regarding claim 8,
Agbavor teaches
A computer-implemented method of classifying cognitive impairment, the method comprising:
acquiring, by an electronic apparatus, utterance voice of a user;
[Page 9, section Text embeddings from GPT3 “To our knowledge, this is the first application of GPT-3 to predicting dementia from speech. In our approach (Fig 1B), we first convert voice to text using Wav2Vec 2.0 pretrained model [36], a state-of-the-art model for automatic speech recognition.”];
[Page 4, Paragraph 4, “Our AI model could be deployed as a web application or even a voice-powered app used at the doctor’s office to aid clinicians in AD screening and early diagnosis.”];
converting, by one or more processors of the electronic apparatus, the utterance voice into transcribed text;
[Page 10, lines 1-2 “GPT-3 based text embeddings are afterwards derived from the transcribed text obtained via wav2vec2.”]; [Page 9 Section : Text embeddings from GPT-3 “The waveform is
then tokenized using Wav2Vec2Tokenizer and if necessary, divided them into smaller chunks
(with the maximum size of 100,000 in our case) to fit into memory, which is subsequently fed
into the Wav2Vec2ForCTC (a wav2vec model for speech recognition) and decoded as text
transcripts.”];
extracting, by the one or more processors, acoustic features and linguistic features based on the utterance voice, the transcribed text,
[ Page 10, lines21- 24 “In this work, acoustic features are extracted directly from speech using OpenSMILE (open-source Speech and Music Interpretation by Large-space Extraction), a widely used open- source toolkit for audio feature extraction and classification of speech and music signals [39]”];
[Page 3, Fig1. Schematic showing two different feature representations that are derived from speech- A.
The acoustic features are engineered to capture the acoustic characteristics of speech and
therefore the pathological speech behavior. B. The linguistic features, represented as text
embeddings, are derived from the transcribed text. Central to our proposed approach is the
GPT-3 based text embeddings (shaded), which entail meaningful vector representations that can
capture lexical, syntactic, and semantic properties for dementia classification] ;
classifying, by the one or more processors, the user into one of a dementia group, a mild cognitive impairment group, or a normal group based on the acoustic features and the linguistic features.
[Page 11, lines 5- 9 “The model may use acoustic features from speech, linguistic features
(embeddings) from transcribed speech, or both. As such, we use (1) the acoustic features
extracted from speech audio data, (2) the text embeddings from each GPT-3 base model
(Babbage or Ada), and (3) the combination of both as inputs for three different kinds of
commonly used machine learning models, including Support Vector Classifier (SVC), Random
Forest (RF), and Logistic Regression (LR)”]; [Page 3, lines 1-10 “We report the results from two
tasks which include AD(Alzheimer’s Dementia) vs non-AD classification and AD severity prediction using a subject’s MMSE score. For the classification task, either the acoustic features or GPT-3 embeddings (Adaand Babbage) or both are fed into a machine-learning model such as support vector classifier
(SVC), logistic regression (LR) or random forest (RF). As a comparison, we further perform
finetuning on the GPT-3 model to see if there is any advantage over the GPT-3 embedding”];
[Page 3, lines 12-17 “In this section we present the AD classification results between AD and
non-AD (or healthy control) subjects based on different features: our proposed GPT-3 based text
embeddings, the acoustic features, and their combination. We also benchmark the GPT-3 based
text embeddings against the mainstream fine-tuning approach. We show that the GPT-3 based
text embeddings considerably outperform both the acoustic feature-based approach and the
fine-tuned model”]; [ Page 5, lines 1-4 “To evaluate whether the acoustic features and the text
embeddings can provide complementary information to augment the AD classification, we
combine the acoustic features from speech audio data and the GPT-3 based text embeddings by
simply concatenating them”]; [ Page 11, line 18- 23 “MMSE is perhaps the most common measure for assessing the severity of AD. We perform regression analysis using both the acoustic features and text embeddings from GPT-3 (Ada and Babbage) to predict the MMSE score. The scores normally range from 0 to 30, with scores of 26 or higher being considered normal [3]. A score of 20 to 24 suggests mild
dementia, 13 to 20 suggests moderate dementia, and less than 12 indicates severe dementia. As
such, the prediction is clipped to a range between 0 and 30”].
However, Agbavor does not teach
generating, by the one or more processors, a first prompt related to a fluency evaluation request based on transcribed text corresponding to utterance voice of a user, the first prompt including the transcribed text and at least one evaluation criterion text selected from among an explanatory phrase related to the transcribed text, a score range and a score unit of the fluency evaluation, and example transcribed text rated with a plurality of scores within the score range;
acquiring, by inputting the first prompt to a large language model, evaluation feedback related to fluency in response to the first prompt through the large language model, the evaluation feedback being related to at least one of inclusion of key elements, consistency, term repetition, and wording accuracy in the transcribed text;
But Yao teaches
Generating, by the one or more processors, a first prompt related to a fluency evaluation request based on the transcribed text corresponding to the utterance voice of the user, the first prompt including the transcribed text
[0030 “In one embodiment, generative AI engine 102 can include an LLM. This LLM can generate coherent and contextually relevant text, which can form the basis of the conversation content for chatbot system 100, and answer questions based on a given context or passage of text. In general, generic LLMs can answer or attempt to answer any question or request (which are referred to as "prompt"). LLMs are typically not entirely intuitive about the exact nuances or specifics the user might be interested in, such as the dialogue carried out by a trained language learning partner”];
[0031 “In order to induce the LLM in generative AI engine 102 to provide the desired response, or to initiate a conversation with the appropriate messages, one aspect of the present disclosure uses a prompt engine 104 to generate the appropriate prompts to the LLM, such that the LLM can provide a desired output. Specifically, prompt engine 104 can generate prompts to cause generative AI engine 102 to initiate a dialogue session or to respond to a user-provided message”];
[0033 “During operation, user 120 can start the language learning process by saying a command to audio sub-system 108, which relays the audio signal to audio-text conversion engine 106. Audio-text conversion engine 106 in turn converts the voice command to a text command and passes the text command to prompt engine 104. Subsequently, prompt engine 104 can generate a prompt and transmit the prompt to generative AI engine 102, which responds with the desired message to be played back to user 120.”];
[0037 “In some aspects, the prompt engine is responsible for creating, based on the user audio messages, prompts to the generative AI engine to induce the AI engine to provide desired responses”]; [0070 “Note that if the user pronounces any of those words incorrectly, the audio-text conversion engine can recognize and convert such mispronounced word to an incorrectly spelled word, which is then included in the prompt sent to the generative AI engine. For example, if the user pronounces
"Gracias" as "Graziaz," the prompt engine can include this mispronounced (which results in mis-
spelling) word in the prompt to the generative AI engine. As a result, the generative AI engine
can provide a message to help the user correct their mistake”];
acquiring, by inputting the first prompt to a large language model, evaluation feedback
[0069 “As a result, the generative AI engine can then evaluate the user's response and provide explanation and correction if necessary.”] [0070 “Note that if the user pronounces any of those words incorrectly, the audio-text conversion engine can recognize and convert such mispronounced word to an incorrectly spelled word, which is then included in the prompt sent to the generative AI engine. For example, if the user pronounces “Gracias” as “Graziaz,” the prompt engine can include this mispronounced (which results in mis-spelling) word in the prompt to the generative AI engine. As a result, the generative AI engine can provide a message to help the user correct their mistake.” where identified mistakes, explanations, and corrections reasonably constitute evaluation feedback];
[0120 “One of the key features of the present system is that the chatbot can evaluate the user's response and provide specific feedback to help the user correct any potential mistakes. To do so, the prompt engine can generate optimized prompts based on the user's response, which can cause the generative AI engine to provide specific feedback for the user”].
It would have been obvious to one of ordinary skill in the art before the effective filing
date of the claimed invention to combine the teachings Agbavor with the teaching of Yao because doing so would enable Agbavor’s system to obtain additional textual information identifying mistakes and explaining linguistic characteristics of the user’s response. This would thereby produce linguistic features based on both sources for use with Agbavor’s acoustic features. This would predictably enrich the classifier’s linguistic representation and improve the diagnostic assessment.
Agbavor in view of Yao do not teach
Generating, by one or more processors, a first prompt
Acquiring, by inputting the first prompt to a large language model, ,
But Weston teaches
Generating, by one or more processors, a first prompt, the first prompt including the transcribed text and at least one evaluation criterion text selected from among an explanatory phrase related to the transcribed text,
[“0133 “In step S506, the SOP encoding 506 and the rating SOP encoding 507 are
provided as inputs to a generative ML model 508 along with a transcript 510]; [0136 “As before,
these inputs can be provided as inputs to the model in a prompt optionally with instructions to
format the output in a structured form, such as JSON or XML. The generative ML model 508 may
be prompted to provide the output in the form of a "rating sheet", although the output could be
in various forms of natural language or structured Text"]; [0111 "Segmentation" in this context
refers to the grouping or mapping of different parts of a transcription to corresponding
questions and answers in template ( e.g., SOP) for performing the clinical assessment”] ; [0113
“The method 400 uses an SOP 402, which is a template for administering a clinical assessment
comprising a plurality of sections corresponding to different clinical questions. A first section
comprises the "count to five" test described previously. A second section comprises a story
recall test, wherein the interviewer tells a patient a story then asks the patient to recall some of
the story, or as many aspects of the story as possible”]; [0128” The method 500 uses a similar
approach to the methods 200, 300 and 400, wherein an SOP 502 is optionally fed into the SOP
encoder 204 to provide an SOP encoding 506 in steps S502 and S504, respectively, as described
previously. The SOP 502 is identical to the SOP 402 and provides instructions for administering
the story recall test”]; [0129 “In the method 500, a rating SOP 503 is also optionally encoded by
the SOP encoder 204 into a rating SOP encoding 507 in steps S502 and S504, respectively”];
[0130 “As shown, the rating SOP encoding 507 comprises an instructions sub-section and a text
string entry providing instructions for scoring the task. The instructions can include guidance
such as "allow paraphrases". A "story elements" subsection is provided that lists each aspect of the story. The example above has been limited to a single story element ("Allison") for
simplicity, though it would be appreciated that further elements may be provided in practice.
Guidance for scoring the particular story element is also provided in the form of a "scoring
guidance" subsection. An "output_ schema" subsection is included to provide a format for the
story element. The output schema conditions the model to provide its output in a particular
format. In this case, 'rating' the story recall means producing a report that includes the
element index and whether or not that element was recalled. A schema, or template, has
children (a type and description) that tells you (a) what kind of data should be populated in
this field, and (b) a description of what the field means. A "recalled" section provides a
description for the evaluation("whether this element was recalled") and a type of the
evaluation ("boo!"). In other examples, the evaluation could be measured in other ways, such
using a scale of 1 to 10 which may be reflected in the "type" subsection"]; [0142 “The rating SOP
encoding 507 can include instructions on how to interpret or adjust an evaluation based on
disfluencies of the patient”].
Acquiring, by inputting the first prompt to a large language model,
[0133 “In step S506, the SOP encoding 506 and the rating SOP encoding 507 are provided as inputs to a generative ML model 508 along with a transcript 510.”];
[0137 “In step S508, the generative ML model analyses the provided inputs and provides an output rating sheet 512.”];
[0142 “The rating SOP encoding 507 can include instructions on how to interpret or adjust an evaluation based on disfluencies of the patient. Alternatively, or in addition, the generative ML model 508 can be trained with positive training data corresponding to patients diagnosed with the health condition in question, in order assess the relevance of disfluencies (or other speech features) to each criteria in the SOP 502.”];
It would have been obvious to one of ordinary skill in the art before the effective filing
date of the claimed invention to combine the teachings Agbavor in view of Yao with Weston’s structured clinical rating prompt because Weston teaches providing a patient transcript and evaluation criteria to an LLM to generate a clinical rating identifying recalled key elements and accounting for disfluencies. This combination would enable Yao’s prompt based feedback mechanism to produce clinically relevant and consistency structured evaluation information. Applying Agbavor’s text embedding technique to that feedback would predictably provide additional linguistic features for its dementia classifier, thereby improving interpretability and diagnostic.
Regarding claim 15,
Agbavor teaches
An electronic apparatus comprising:
acquire utterance voice of a user
[Page 9, section Text embeddings from GPT3 “To our knowledge, this is the first application of GPT-3 to predicting dementia from speech. In our approach (Fig 1B), we first convert voice to text using Wav2Vec 2.0 pretrained model [36], a state-of-the-art model for automatic speech recognition.”];
[Page 4, Paragraph 4, “Our AI model could be deployed as a web application or even a voice-powered app used at the doctor’s office to aid clinicians in AD screening and early diagnosis.”];
convert the utterance voice into transcribed text;
[Page 10, lines 1-2 “GPT-3 based text embeddings are afterwards derived from the transcribed text obtained via wav2vec2.”]; [Page 9 Section : Text embeddings from GPT-3 “The waveform is
then tokenized using Wav2Vec2Tokenizer and if necessary, divided them into smaller chunks
(with the maximum size of 100,000 in our case) to fit into memory, which is subsequently fed
into the Wav2Vec2ForCTC (a wav2vec model for speech recognition) and decoded as text
transcripts.”];
extract acoustic features and linguistic features based on the utterance voice, the
[ Page 10, lines21- 24 “In this work, acoustic features are extracted directly from speech using OpenSMILE (open-source Speech and Music Interpretation by Large-space Extraction), a widely used open- source toolkit for audio feature extraction and classification of speech and music signals [39]”];
[Page 3, Fig1. Schematic showing two different feature representations that are derived from speech- A.
The acoustic features are engineered to capture the acoustic characteristics of speech and
therefore the pathological speech behavior. B. The linguistic features, represented as text
embeddings, are derived from the transcribed text. Central to our proposed approach is the
GPT-3 based text embeddings (shaded), which entail meaningful vector representations that can
capture lexical, syntactic, and semantic properties for dementia classification] ;
classify the user into one of a dementia group, a mild cognitive impairment group, or a normal group based on the acoustic features and the linguistic features.
[Page 11, lines 5- 9 “The model may use acoustic features from speech, linguistic features
(embeddings) from transcribed speech, or both. As such, we use (1) the acoustic features
extracted from speech audio data, (2) the text embeddings from each GPT-3 base model
(Babbage or Ada), and (3) the combination of both as inputs for three different kinds of
commonly used machine learning models, including Support Vector Classifier (SVC), Random
Forest (RF), and Logistic Regression (LR)”]; [Page 3, lines 1-10 “We report the results from two
tasks which include AD(Alzheimer’s Dementia) vs non-AD classification and AD severity prediction using a subject’s MMSE score. For the classification task, either the acoustic features or GPT-3 embeddings (Adaand Babbage) or both are fed into a machine-learning model such as support vector classifier
(SVC), logistic regression (LR) or random forest (RF). As a comparison, we further perform
finetuning on the GPT-3 model to see if there is any advantage over the GPT-3 embedding”];
[Page 3, lines 12-17 “In this section we present the AD classification results between AD and
non-AD (or healthy control) subjects based on different features: our proposed GPT-3 based text
embeddings, the acoustic features, and their combination. We also benchmark the GPT-3 based
text embeddings against the mainstream fine-tuning approach. We show that the GPT-3 based
text embeddings considerably outperform both the acoustic feature-based approach and the
fine-tuned model”]; [ Page 5, lines 1-4 “To evaluate whether the acoustic features and the text
embeddings can provide complementary information to augment the AD classification, we
combine the acoustic features from speech audio data and the GPT-3 based text embeddings by
simply concatenating them”]; [ Page 11, line 18- 23 “MMSE is perhaps the most common measure for assessing the severity of AD. We perform regression analysis using both the acoustic features and text embeddings from GPT-3 (Ada and Babbage) to predict the MMSE score. The scores normally range from 0 to 30, with scores of 26 or higher being considered normal [3]. A score of 20 to 24 suggests mild
dementia, 13 to 20 suggests moderate dementia, and less than 12 indicates severe dementia. As
such, the prediction is clipped to a range between 0 and 30”].
However, Agbavor does not teach
Generate a first prompt related to a fluency evaluation request based on transcribed text corresponding to utterance voice of a user, the first prompt including the transcribed text and at least one evaluation criterion text selected from among an explanatory phrase related to the transcribed text, a score range and a score unit of the fluency evaluation, and example transcribed text rated with a plurality of scores within the score range;
Input the first prompt to a large language model, and acquire evaluation feedback from the large language model, the evaluation feedback being related to at least one of inclusion of key elements, consistency, term repetition, and wording accuracy in the transcribed text;
But Yao teaches
Generate a first prompt related to a fluency evaluation request based on the transcribed text corresponding to the utterance voice of the user, the first prompt including the transcribed text
[0030 “In one embodiment, generative AI engine 102 can include an LLM. This LLM can generate coherent and contextually relevant text, which can form the basis of the conversation content for chatbot system 100, and answer questions based on a given context or passage of text. In general, generic LLMs can answer or attempt to answer any question or request (which are referred to as "prompt"). LLMs are typically not entirely intuitive about the exact nuances or specifics the user might be interested in, such as the dialogue carried out by a trained language learning partner”];
[0031 “In order to induce the LLM in generative AI engine 102 to provide the desired response, or to initiate a conversation with the appropriate messages, one aspect of the present disclosure uses a prompt engine 104 to generate the appropriate prompts to the LLM, such that the LLM can provide a desired output. Specifically, prompt engine 104 can generate prompts to cause generative AI engine 102 to initiate a dialogue session or to respond to a user-provided message”];
[0033 “During operation, user 120 can start the language learning process by saying a command to audio sub-system 108, which relays the audio signal to audio-text conversion engine 106. Audio-text conversion engine 106 in turn converts the voice command to a text command and passes the text command to prompt engine 104. Subsequently, prompt engine 104 can generate a prompt and transmit the prompt to generative AI engine 102, which responds with the desired message to be played back to user 120.”];
[0037 “In some aspects, the prompt engine is responsible for creating, based on the user audio messages, prompts to the generative AI engine to induce the AI engine to provide desired responses”]; [0070 “Note that if the user pronounces any of those words incorrectly, the audio-text conversion engine can recognize and convert such mispronounced word to an incorrectly spelled word, which is then included in the prompt sent to the generative AI engine. For example, if the user pronounces
"Gracias" as "Graziaz," the prompt engine can include this mispronounced (which results in mis-
spelling) word in the prompt to the generative AI engine. As a result, the generative AI engine
can provide a message to help the user correct their mistake”];
Input the first prompt to a large language model, and acquire evaluation feedback from the large language model, evaluation feedback
[0069 “As a result, the generative AI engine can then evaluate the user's response and provide explanation and correction if necessary.”] [0070 “Note that if the user pronounces any of those words incorrectly, the audio-text conversion engine can recognize and convert such mispronounced word to an incorrectly spelled word, which is then included in the prompt sent to the generative AI engine. For example, if the user pronounces “Gracias” as “Graziaz,” the prompt engine can include this mispronounced (which results in mis-spelling) word in the prompt to the generative AI engine. As a result, the generative AI engine can provide a message to help the user correct their mistake.” where identified mistakes, explanations, and corrections reasonably constitute evaluation feedback];
[0120 “One of the key features of the present system is that the chatbot can evaluate the user's response and provide specific feedback to help the user correct any potential mistakes. To do so, the prompt engine can generate optimized prompts based on the user's response, which can cause the generative AI engine to provide specific feedback for the user”].
It would have been obvious to one of ordinary skill in the art before the effective filing
date of the claimed invention to combine the teachings Agbavor with the teaching of Yao because doing so would enable Agbavor’s system to obtain additional textual information identifying mistakes and explaining linguistic characteristics of the user’s response. This would thereby produce linguistic features based on both sources for use with Agbavor’s acoustic features. This would predictably enrich the classifier’s linguistic representation and improve the diagnostic assessment.
Agbavor in view of Yao do not teach
Generate a first prompt
input the first prompt to a large language model, and acquire evaluation feedback from the large language model, the
But Weston teaches
Generate, a first prompt, the first prompt including the transcribed text and at least one evaluation criterion text selected from among an explanatory phrase related to the transcribed text,
[“0133 “In step S506, the SOP encoding 506 and the rating SOP encoding 507 are
provided as inputs to a generative ML model 508 along with a transcript 510]; [0136 “As before,
these inputs can be provided as inputs to the model in a prompt optionally with instructions to
format the output in a structured form, such as JSON or XML. The generative ML model 508 may
be prompted to provide the output in the form of a "rating sheet", although the output could be
in various forms of natural language or structured Text"]; [0111 "Segmentation" in this context
refers to the grouping or mapping of different parts of a transcription to corresponding
questions and answers in template ( e.g., SOP) for performing the clinical assessment”] ; [0113
“The method 400 uses an SOP 402, which is a template for administering a clinical assessment
comprising a plurality of sections corresponding to different clinical questions. A first section
comprises the "count to five" test described previously. A second section comprises a story
recall test, wherein the interviewer tells a patient a story then asks the patient to recall some of
the story, or as many aspects of the story as possible”]; [0128” The method 500 uses a similar
approach to the methods 200, 300 and 400, wherein an SOP 502 is optionally fed into the SOP
encoder 204 to provide an SOP encoding 506 in steps S502 and S504, respectively, as described
previously. The SOP 502 is identical to the SOP 402 and provides instructions for administering
the story recall test”]; [0129 “In the method 500, a rating SOP 503 is also optionally encoded by
the SOP encoder 204 into a rating SOP encoding 507 in steps S502 and S504, respectively”];
[0130 “As shown, the rating SOP encoding 507 comprises an instructions sub-section and a text
string entry providing instructions for scoring the task. The instructions can include guidance
such as "allow paraphrases". A "story elements" subsection is provided that lists each aspect of the story. The example above has been limited to a single story element ("Allison") for
simplicity, though it would be appreciated that further elements may be provided in practice.
Guidance for scoring the particular story element is also provided in the form of a "scoring
guidance" subsection. An "output_ schema" subsection is included to provide a format for the
story element. The output schema conditions the model to provide its output in a particular
format. In this case, 'rating' the story recall means producing a report that includes the
element index and whether or not that element was recalled. A schema, or template, has
children (a type and description) that tells you (a) what kind of data should be populated in
this field, and (b) a description of what the field means. A "recalled" section provides a
description for the evaluation("whether this element was recalled") and a type of the
evaluation ("boo!"). In other examples, the evaluation could be measured in other ways, such
using a scale of 1 to 10 which may be reflected in the "type" subsection"]; [0142 “The rating SOP
encoding 507 can include instructions on how to interpret or adjust an evaluation based on
disfluencies of the patient”].
input the first prompt to a large language model, acquire evaluation feedback from the large language model,
[0133 “In step S506, the SOP encoding 506 and the rating SOP encoding 507 are provided as inputs to a generative ML model 508 along with a transcript 510.”];
[0137 “In step S508, the generative ML model analyses the provided inputs and provides an output rating sheet 512.”];
[0142 “The rating SOP encoding 507 can include instructions on how to interpret or adjust an evaluation based on disfluencies of the patient. Alternatively, or in addition, the generative ML model 508 can be trained with positive training data corresponding to patients diagnosed with the health condition in question, in order assess the relevance of disfluencies (or other speech features) to each criteria in the SOP 502.”];
It would have been obvious to one of ordinary skill in the art before the effective filing
date of the claimed invention to combine the teachings Agbavor in view of Yao with Weston’s structured clinical rating prompt because Weston teaches providing a patient transcript and evaluation criteria to an LLM to generate a clinical rating identifying recalled key elements and accounting for disfluencies. This combination would enable Yao’s prompt based feedback mechanism to produce clinically relevant and consistency structured evaluation information. Applying Agbavor’s text embedding technique to that feedback would predictably provide additional linguistic features for its dementia classifier, thereby improving interpretability and diagnostic.
Claims [2, 7, 9 ,14, 17, 20] are rejected under 35 U.S.C. 103 as being unpatentable over
Agbavor(Agbavor, F. and Liang, H., 2022. Predicting dementia from spontaneous speech using
large language models. PLOS digital health, 1(12), p.e0000168. ) in view of Yao (US
2024/0321131 A) and in view of Weston (US2025/0132036A1) and in further view of Rentoumi (US. Patent No. 11,114,113 B2).
Regarding claim 2,
Agbavor in view of Yao in view of Weston does not teach the electronic apparatus of Claim 1, wherein the utterance voice includes a voice of the user that describes a painting or photo.
However, Rentoumi teaches electronic apparatus of claim 1, wherein the utterance
voice includes a voice of the user that describes a painting or photo .
[“Column 3, lines 31-44“According to the present disclosure, a system for early detection of Alzheimer's disease is provided. The system generates a prediction of whether a patient has Alzheimer's disease,
another form of dementia, or other neurodegenerative disease based on an occurrence of
speech, such as the patient's speech in response to a task or free form speech. For example, the
patient may be shown a picture and asked to describe the picture, be asked to retell a popular
short story or fairy tale, or be asked to describe how to perform a specific task. The patient's
speech is recorded and a transcript is generated of the speech for analysis. The transcript may
be generated using available speech to text applications or, in some implementations, may be
generated by hand”].
It would have been obvious to one of ordinary skill in the art before the effective filing
date of the claimed invention to combine the teachings Agbavor in view of Yao in view of Weston with the teachings of Rentoumi because it would automatically analyze a user’s spoken description and
generating contextual prompts or feedback based on spoken description. The visual stimulus
would help generate intelligent prompts or feedback, thereby improving classification accuracy.
Regarding Claim 7, Agbavor in view of Yao in view of Weston teach the electronic apparatus of claim 1,
Agbavor teaches classification on the cognitive impairment – [Page 5, lines 1-5 “To evaluate whether the acoustic features and the text embeddings can provide complementary information to augment the AD classification, we combine the acoustic features from speech audio data and the GPT-3 based text embeddings by simply concatenating them. Table 4 shows the results for both the 10-fold CV and evaluation on the test set for different machine learning models”]; [Page 11, lines 40-42 “We also report the averaged AUC scores, along with the corresponding standard deviations over the 10-fold CV when comparing the different models using acoustic features, GPT-3 embeddings (both Ada and Babbage) for AD classification”]; [Page 3, lines 1-11 “We report the results from two tasks which include AD vs non-AD classification and AD severity prediction using a subject’s MMSE score. For the classification task, either the acoustic features or GPT-3
embeddings (Ada and Babbage) or both are fed into a machine-learning model such as support
vector classifier (SVC), logistic regression (LR) or random forest (RF). As a comparison, we further
perform finetuning on the GPT-3 model to see if there is any advantage over the GPT-3
embedding. For the AD severity prediction, we perform the regression analysis based on both
the acoustic features and GPT-3 embeddings to estimate a subject’s MMSE score using three
regression models, i.e., support vector regressor (SVR), ridge regression (Ridge) and random
forest regressor (RFR)”].
However, Agbavor does not teach the evaluation feedback through the LLM
But Yao teaches evaluation feedback through the LLM
[0069 “As a result, the generative AI engine can then evaluate the user's response and provide explanation and correction if necessary.”] [0070 “Note that if the user pronounces any of those words incorrectly, the audio-text conversion engine can recognize and convert such mispronounced word to an incorrectly spelled word, which is then included in the prompt sent to the generative AI engine. For example, if the user pronounces “Gracias” as “Graziaz,” the prompt engine can include this mispronounced (which results in mis-spelling) word in the prompt to the generative AI engine. As a result, the generative AI engine can provide a message to help the user correct their mistake.” where identified mistakes, explanations, and corrections reasonably constitute evaluation feedback];
[0120 “One of the key features of the present system is that the chatbot can evaluate the user's response and provide specific feedback to help the user correct any potential mistakes. To do so, the prompt engine can generate optimized prompts based on the user's response, which can cause the generative AI engine to provide specific feedback for the user”].
It would have been obvious to one of ordinary skill in the art before the effective filing
date of the claimed invention to combine the teachings Agbavor with the teaching of Yao because doing so would enable Agbavor’s system to obtain additional textual information identifying mistakes and explaining linguistic characteristics of the user’s response. This would thereby produce linguistic features based on both sources for use with Agbavor’s acoustic features. This would predictably enrich the classifier’s linguistic representation and improve the diagnostic assessment.
However, Agbavor in view of Yao view of Weston do not teach the one or more processors are
further configured to generate a test opinion.
But Rentoumi teaches one or more processors are further configured to generate a test opinion-
[Page 2, lines 29-37 “The example system includes a prediction module including a trained
classification model, wherein the trained classification model is trained to generate a prediction
of the disease state for a patient based on the speech using the plurality of lingual features
extracted from the speech. A communication interface is configured to return the prediction of
the disease state and one or more analytics regarding the speech and the lingual features to a
user device for display to a user”].[Column 16, lines 19-25 “ The technology described herein may be implemented as logical operations and/or modules in one or more systems. The logical operations may be implemented as a sequence of processor-implemented steps directed by software programs executing in one or more computer systems and as interconnected machine or circuit modules within one or more computer systems, or as a combination of both.”];
It would have been obvious to one of ordinary skill in the art before the effective filing
date of the claimed invention to combine the teachings Agbavor in view of Yao in view of Weston with the teachings Rentoumi because generating the prediction or opinion of the summarized
classification results obtained from the transcribed text would allow the language model to
show the opinion/prediction to the user. This combination represents predictable use of known
techniques to transform classification outputs into user understandable feedback or opinion,
thereby improving interpretability, usability, and practical value of the system by enabling
meaningful opinion/prediction summaries of the cognitive assessment results.
Regarding claim 9, the method of claim 8, wherein the utterance voice includes a voice of
the user that describes a painting or photo.
Claim 9 is rejected for the same reasons as claim 2.
Regarding claim 14, the method of claim 8, further comprising: summarizing, by one or more processors, a result of the classification on the cognitive impairment and the evaluation feedback through the large language model to generate a test opinion.
Claim 14 is rejected for the same reasons as claim 7.
Regarding claim 17, Agbavor in view of Yao does teach the electronic apparatus of
claim 15, further comprising a communication module, wherein the processor executes the at
least one instruction to:
Agbavor teaches result of the classification on the cognitive impairment
[Page 5, lines 1-5 “To evaluate whether the acoustic features and the text embeddings can provide complementary information to augment the AD classification, we combine the acoustic features from speech audio data and the GPT-3 based text embeddings by simply concatenating them. Table 4 shows the results for both the 10-fold CV and evaluation on the test set for different machine learning models”]; [Page 11, lines 40-42 “We also report the averaged AUC scores, along with the corresponding standard deviations over the 10-fold CV when comparing the different models using acoustic features, GPT-3 embeddings (both Ada and Babbage) for AD classification”]; [Page 3, lines 1-11 “We report the results from two tasks which include AD vs non-AD classification and AD severity prediction using a subject’s MMSE score. For the classification task, either the acoustic features or GPT-3
embeddings (Ada and Babbage) or both are fed into a machine-learning model such as support
vector classifier (SVC), logistic regression (LR) or random forest (RF). As a comparison, we further
perform finetuning on the GPT-3 model to see if there is any advantage over the GPT-3
embedding. For the AD severity prediction, we perform the regression analysis based on both
the acoustic features and GPT-3 embeddings to estimate a subject’s MMSE score using three
regression models, i.e., support vector regressor (SVR), ridge regression (Ridge) and random
forest regressor (RFR)”].
However, Agbavor does not teach
acquire the utterance voice from an external electronic apparatus through the communication module
and the evaluation feedback through the large language model.
But Yao teaches acquire the utterance voice from an external electronic apparatus through the communication module
[0032 “Audio input device 112 can receive user-provided audio messages and provide these messages to audio-text conversion engine 106.”];
the evaluation feedback through the large language model
[0069 “As a result, the generative AI engine can then evaluate the user's response and provide explanation and correction if necessary.”] [0070 “Note that if the user pronounces any of those words incorrectly, the audio-text conversion engine can recognize and convert such mispronounced word to an incorrectly spelled word, which is then included in the prompt sent to the generative AI engine. For example, if the user pronounces “Gracias” as “Graziaz,” the prompt engine can include this mispronounced (which results in mis-spelling) word in the prompt to the generative AI engine. As a result, the generative AI engine can provide a message to help the user correct their mistake.” where identified mistakes, explanations, and corrections reasonably constitute evaluation feedback];
[0120 “One of the key features of the present system is that the chatbot can evaluate the user's response and provide specific feedback to help the user correct any potential mistakes. To do so, the prompt engine can generate optimized prompts based on the user's response, which can cause the generative AI engine to provide specific feedback for the user”].
It would have been obvious to one of ordinary skill in the art before the effective filing
date of the claimed invention to combine the teachings Agbavor with the teaching of Yao because doing so would enable Agbavor’s system to obtain additional textual information identifying mistakes and explaining linguistic characteristics of the user’s response. This would thereby produce linguistic features based on both sources for use with Agbavor’s acoustic features. This would predictably enrich the classifier’s linguistic representation and improve the diagnostic assessment.
Agbavor in view of Yao does not teach provide to the external electronic apparatus through the communication module, a test opinion generated by summarizing a result
However, Rentoumi teaches the to the external electronic apparatus through the communication module, a test opinion.
[Page 2, lines 29-37 “The example system includes a prediction module including a
trained classification model, wherein the trained classification model is trained to generate a prediction of the disease state for a patient based on the speech using the plurality of lingual
features extracted from the speech. A communication interface is configured to return the
prediction of the disease state and one or more analytics regarding the speech and the lingual
features to a user device for display to a user”].
It would have been obvious to one of ordinary skill in the art before the effective filing
date of the claimed invention to combine the teachings Agbavor in view of Yao with the
teachings Rentoumi because this would allow the generating of a prediction or opinion and for
the results to be displayed to a user via a communication module. This combination represents
predictable use of known techniques to transform outputs into user understandable feedback or
opinion, thereby improving interpretability, usability, and practical value of the system by
enabling meaningful opinion/prediction results.
Regarding claim 20, The electronic apparatus of claim 15, wherein the processor
executes the at least one instruction to: summarize a result of the classification on the cognitive
impairment and the evaluation feedback through the large language model to generate a test
opinion.
Claim 20 is rejected for the same reasons as claim 7.
Claims [4, 11 ] are rejected under 35 U.S.C. 103 as being unpatentable over
Agbavor(Agbavor, F. and Liang, H., 2022. Predicting dementia from spontaneous speech using
large language models. PLOS digital health, 1(12), p.e0000168. ) in view of Yao (US
2024/0321131 A) and in further view of Weston (US2025/0132036A1) and in further view of
Kurlowicz “The Mini Mental State Examination (MMSE)” (https://cgatoolkit.ca/Uploads/ContentDocuments/MMSE.pdf).
Regarding claim 4,
Agbavor in view of Yao does not teach electronic apparatus of claim 1,
wherein the plurality of scores include a lowest score and a highest score within the score range.
But Weston teaches scores include a lowest score and a highest score within the score
range - [0111 "Segmentation" in this context refers to the grouping or mapping of different
parts of a transcription to corresponding questions and answers in template ( e.g., SOP) for
performing the clinical assessment”]; [0130 “As shown, the rating SOP encoding 507 comprises
an instructions sub-section and a text string entry providing instructions for scoring the task. The
instructions can include guidance such as "allow paraphrases". A "story elements" subsection is
provided that lists each aspect of the story. The example above has been limited to a single
story element ("Allison") for simplicity, though it would be appreciated that further elements
may be provided in practice. Guidance for scoring the particular story element is also provided
in the form of a "scoring guidance" subsection. An "output_ schema" subsection is included to
provide a format for the story element. The output schema conditions the model to provide its
output in a particular format. In this case, 'rating' the story recall means producing a report that
includes the element index and whether or not that element was recalled. A schema, or
template, has children (a type and description) that tells you (a) what kind of data should be
populated in this field, and (b) a description of what the field means. A "recalled" section
provides a description for the evaluation("whether this element was recalled") and a type of the
evaluation ("boo!"). In other examples, the evaluation could be measured in other ways, such
using a scale of 1 to 10 which may be reflected in the "type" subsection"]; [0142 “The rating SOP
encoding 507 can include instructions on how to interpret or adjust an evaluation based on
disfluencies of the patient”];[ [0004 “Once administered, clinical assessments also need
to be evaluated or scored, which is known as rating].
It would have been obvious to one of ordinary skill in the art before the effective filing
date of the claimed invention to combine the teachings of Agbavor in view of Yao with the
teachings of Weston because scoring the transcribed text used to evaluate will enhance the quality
and reliably of the transcribed text which in turn will help the language model generate outputs
that are more consistent which then improves accuracy of the classification results.
However, Weston doesn’t explicitly say scores include a lowest score and a highest
score within the score range.
But, Kurlowicz teaches scores include lowest and a highest score – [Page 1 lines 8-10
“The Mini Mental State Examination (MMSE) is a tool that can be used to systematically and
thoroughly assess mental status. It is an 11-question measure that tests five areas of cognitive
function: orientation, registration, attention and calculation, recall, and language. The maximum
score is 30. A score of 23 or lower is indicative of cognitive impairment. The MMSE takes only 5-
10 minutes to administer and is therefore practical to use repeatedly and routinely”].
It would have been obvious to one of ordinary skill in the art before the effective filing
date of the claimed invention to combine the teachings of Agbavor in view of Yao and in view of
Weston with the teachings Kurlowicz because having a plurality of scoring ranging from high to
low scores with the transcribed text used will enhance the quality and reliably of the transcribed
text which in turn will help the language model generate outputs that are more consistent ,
which then improves accuracy of the classification results.
Regarding claim 11, the method of claim 8, wherein the plurality of scores include a
lowest score and a highest score within the score range.
Claim 11 is rejected for the same reasons as claim 4.
Claims [16] is rejected under 35 U.S.C. 103 as being unpatentable over Agbavor(
Agbavor, F. and Liang, H., 2022. Predicting dementia from spontaneous speech using large
language models. PLOS digital health, 1(12), p.e0000168.) in view of Yao (US 2024/0321131 A)
and in view of Weston (US20250132036A1) and in further view of Luz (Detecting cognitive decline using speech only: The ADReSSoChallenge, arxiv.org/pdf/2104.09356).
Regarding claim 16, Agbavor teaches the electronic apparatus of claim 15, wherein the
processor executes the at least one instruction to: request training data of the artificial
intelligence model in relation to an image, utterance voice, and transcribed text of a cognitively
impaired patient from the large language model; and construct the artificial intelligence model
based on the training data – [Page 9 , lines 7-15 “The dataset used in this study is derived from
the ADReSSo Challenge [21], which consists of set of speech recordings of picture descriptions
produced by cognitively normal subjects and patients with an AD diagnosis, who were asked to
describe the Cookie Theft picture from the Boston Diagnostic Aphasia Examination [6,35]. There are totally 237 speech recordings, with 70/30 split balanced for demographics, resulting in 166
and 71 in the training set and the test set, respectively. In the training set, there are 87 samples
from AD subjects and 79 from non-AD (or healthy control) subjects. The datasets were matched
so as to avoid potential biases often overlooked in assessment of AD detection methods,
including incidences of repetitive speech from the same individual, variations in speech quality,
and imbalanced distribution of gender and age. The detailed procedures to match the data
demographically according to propensity scores were described in Luz et al. [21]”]; [ Page 10,
lines 43-46 , Page 11, 1-2 “To fine tune our own custom GPT-3 models, we use the OpenAI
command-line interface, which is released to the public. We simply follow the instructions about
fine-tuning, provided by OpenAI, to prepare the training data that consists of 166 paragraphs,
totaling 19,123 words that are used to fine tune one of the base models (Babbage and Ada in
our case) with speech transcripts. Tokens used to train a model are relatively cheaper, as billed
at 50% of the base prices”].
However, Agbavor in view of Yao in view of Weston does not explicitly teach request training data of the artificial intelligence model in relation to a transcribed text of a cognitively impaired patient from the large language model.
But Luz teaches that dataset that is derived from the ADReSSo Challenge also includes
transcribed text – [Page 1, column 2, lines 5-11 “The ADReSSo Challenge provides a forum for
researchers working on approaches to cognitive decline detection based on speech data to test
their existing methods or develop novel approaches on a new shared standardized dataset. The
approaches that performed best on last year’s dataset [4] employed features extracted from
manual transcripts which were provided along with the audio data [6, 7]”].
It would have been obvious to one of ordinary skill in the art before the effective filing
date of the claimed invention to combine the teachings Agbavor in view Yao in view of Weston with Luz because incorporating training data feedback in relation to image, voice, and transcribed text of impaired patients and using it to train language model enables data driven answer generation were the
language models leverages patterns learned from training data to produce more accurate
responses thereby improving diagnostic outputs.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHEZA ABDUL AZIZ whose telephone number is (571)272-9610. The examiner can normally be reached Monday-Friday 7:30am-5pm Alternate Fridays off.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at (571) 272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SHEZA ABDUL AZIZ/Examiner, Art Unit 2657
/DANIEL C WASHBURN/Supervisory Patent Examiner, Art Unit 2657