Prosecution Insights
Last updated: October 02, 2026
Application No. 19/063,817

ADAPTIVE, INDIVIDUALIZED, AND CONTEXTUALIZED TEXT-TO-SPEECH SYSTEMS AND METHODS

Non-Final OA §DP
Filed
Feb 26, 2025
Priority
Dec 07, 2022 — continuation of 12/266,340
Examiner
MANOHARAN, SHASHIDHAR SHANKAR
Art Unit
Tech Center
Assignee
Truist Bank
OA Round
1 (Non-Final)
80%
Grant Probability
Favorable
1-2
OA Rounds
7m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 80% — above average
80%
Career Allowance Rate
4 granted / 5 resolved
+20.0% vs TC avg
Strong +33% interview lift
Without
With
+33.3%
Interview Lift
resolved cases with interview
Fast prosecutor
2y 2m
Avg Prosecution
26 currently pending
Career history
33
Total Applications
across all art units

Statute-Specific Performance

§101
17.2%
-22.8% vs TC avg
§103
64.8%
+24.8% vs TC avg
§102
4.7%
-35.3% vs TC avg
§112
9.4%
-30.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 5 resolved cases

Office Action

§DP
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13. The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer. Claims 1-20 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-20 of U.S. Patent No. US 12266340 B2. Although the claims at issue are not identical, they are not patentably distinct from each other as laid out in the chart below. Instant Application US Patent: US 12266340 B2 Claim 1: A computing system for adaptive, individualized, and contextualized text-to- speech, the system comprising: Claim 1: A computing system for adaptive, individualized, and contextualized text-to- speech, the system comprising: A memory; A memory; one or more processors in communication with the memory; one or more processors in communication with the memory; And program instructions executable by the one or more processors via the memory to: And program instructions executable by the one or more processors via the memory to: iteratively train, using training data, at least one neural network to categorize communication elements from audio data, the neural network being trained to perform natural language processing on the audio data to generate transcribed data from the audio data, parse the generated transcribed data, and perform a reduction analysis on the transcribed data in order to interpret the transcribed data to derive the communication elements for categorization, the communication elements including intent, context, emotion, and other circumstantial factors, the training including: iteratively train, using training data, at least one recurrent neural network (RNN) to categorize communication elements from audio data, the RNN being trained to perform natural language processing on the audio data to generate transcribed data from the audio data, parse the generated transcribed data, and perform a reduction analysis on the transcribed data in order to interpret the transcribed data to derive the communication elements for categorization, the communication elements including intent, context, emotion, and other circumstantial factors, the training including: inserting the training data into an iterative training and testing loop to predict a target variable; inserting the training data into an iterative training and testing loop to predict a target variable; And repeatedly predicting the target variable during each iteration of the training and testing loop, wherein each iteration of the training and testing loop has differing weights applied to one or more nodes of the neural network, each of the differing weights being updated with each iteration of the training and testing loop to reduce error in predicting the target variable and improve predictability of the neural network; And repeatedly predicting the target variable during each iteration of the training and testing loop, wherein each iteration of the training and testing loop has differing weights applied to one or more nodes of the RNN, each of the differing weights being updated with each iteration of the training and testing loop to reduce error in predicting the target variable and improve predictability of the RNN: deploy the at least one trained neural network to facilitate categorization of the communication elements; deplov the at least one trained RNN to facilitate categorization of the communication elements: receive, in real-time from a user device, input audio data of a user via a telephonic communication means, the input audio data comprising a plurality of communication elements; receive, in real-time from a user via a user device, input audio data of a user via a telephonic communication means, the input audio data comprising a plurality of communication elements; apply the input audio data to the at least one trained neural network to categorize one or more communication elements of the plurality of communication elements, the categorizing including assigning at least one contextual category to a communication element of the plurality of communication elements; apply the input audio data to the at least one trained RNN model to categorize one or more communication elements of the plurality of communication elements, the categorizing including assigning at least one contextual category to a communication element of the plurality of communication elements; generate text comprising a response to one or more of the plurality of communication elements, the response including one or more individualized and contextualized qualities predicted to provide an optimal outcome based at least in part on (i) the assigned at least one contextual category and (ii) the user from which the input audio data is received; generate text comprising a response to one or more of the plurality of communication elements, the response including one or more individualized and contextualized qualities predicted to provide an optimal outcome based at least in part on (i) the assigned at least one contextual category and (ii) the user from which the input audio data is received; implement text-to-speech processing of the generated text, the text-to- speech processing producing an audio output comprising (a) the response and (b) a speech pattern predicted to facilitate the optimal outcome, the speech pattern including at least one prosody element, the at least one prosody element including a timbre intended to elicit certain emotions of the user associated with a desired outcome, the timbre being based on the at least one contextual category assigned to the communication element of the plurality of communication elements; implement text-to-speech processing of the generated text, the text-to-speech processing producing an audio output comprising (a) the response and (b) a speech pattern predicted to facilitate the optimal outcome, the speech pattern including at least one prosody element, the at least one prosody element includin2 a timbre intended to elicit certain emotions of the user associated with a desired outcome, the timbre being based on the at least one contextual category assigned to the communication element of the plurality of communication elements; and provide, to the user device, the audio output and based thereon measure a reaction of the user in response to the speech pattern and the timbre of the audio output according to a quantifiable quality score and using the quantifiable quality score to modify future iterations of the text-to-speech processing in order to provide one or more future audio outputs comprising a revised speech pattern. and provide, to the user via the user device, the audio output and based thereon measure a reaction of the user in response to the speech pattern and the timbre of the audio output according to a quantifiable quality score and using the quantifiable quality score to modify future iterations of the text-to-speech processing in order to provide one or more future audio outputs comprising a revised speech pattern. Claim 2: The computing system for adaptive, individualized, and contextualized text-to- speech according to claim 1, wherein the input audio data includes multiple speech utterances and based thereon performing the receiving, performing, and applying after each speech utterance of the multiple speech utterances, and wherein the assigning is adapted after each speech utterance in context with the multiple speech utterances received thus far from the user. Claim 2: The computing system for adaptive, individualized, and contextualized text-to- speech according to claim 1, wherein the input audio data includes multiple speech utterances and based thereon performing the receiving, performing, and applying after each speech utterance of the multiple speech utterances, and wherein the assigning is adapted after each speech utterance in context with the multiple speech utterances received thus far from the user. Claim 3: The computing system for adaptive, individualized, and contextualized text-to- speech according to claim 1, wherein the program instructions are further executable to detect a pause in the input audio data being received and, based at least in part on the detecting, to determine whether the user has completed a current stream of speech utterances, wherein the providing the audio output is based on the determining that the user has completed the current stream of speech utterances. Claim 3: The computing system for adaptive, individualized, and contextualized text-to- speech according to claim 1, wherein the program instructions are further executable to detect a pause in the input audio data being received and, based at least in part on the detecting, to determine whether the user has completed a current stream of speech utterances, wherein the providing the audio output is based on the determining that the user has completed the current stream of speech utterances. Claim 4: The computing system for adaptive, individualized, and contextualized text-to- speech according to claim 1, wherein the program instructions are further executable to store the input audio data to a user profile of the user, wherein the user profile comprises stored user data that includes any prior interaction data of the user associated with prior interactions between the user and the computing system. Claim 4: The computing system for adaptive, individualized, and contextualized text-to- speech according to claim 1, wherein the program instructions are further executable to store the input audio data to a user profile of the user, wherein the user profile comprises stored user data that includes any prior interaction data of the user associated with prior interactions between the user and the computing system. Claim 5: The computing system for adaptive, individualized, and contextualized text-to- speech according to claim 4, wherein the program instructions are further executable to identify an existing user profile of the user and based thereon: access and analyze the stored user data; identify any stored communication elements that contributed to any positive resolutions from the prior interaction data of the user; determine that the identified stored communication elements would contribute to the optimal outcome; and incorporate the identified stored communication elements into the response. Claim 5: The computing system for adaptive, individualized, and contextualized text-to- speech according to claim 5, wherein the program instructions are further executable to identify an existing user profile of the user and based thereon: access and analyze the stored user data; identify any stored communication elements that contributed to any positive resolutions from the prior interaction data of the user; determine that the identified stored communication elements would contribute to the optimal outcome; and incorporate the identified stored communication elements into the response. Claim 6: The computing system for adaptive, individualized, and contextualized text-to- speech according to claim 5, wherein the user profile includes customer feedback provided by the user, the customer feedback identifying the stored communication elements that contributed to a positive resolution. Claim 6: The computing system for adaptive, individualized, and contextualized text-to- speech according to claim 6, wherein the user profile includes customer feedback provided by the user, the customer feedback identifying the stored communication elements that contributed to a positive resolution. Claim 7: The computing system for adaptive, individualized, and contextualized text-to- speech according to claim 5, wherein the user profile includes agent feedback provided by an agent's previous communication with the user, the agent feedback identifying the stored communication elements that contributed to a positive resolution. Claim 7: The computing system for adaptive, individualized, and contextualized text-to- speech according to claim 6, wherein the user profile includes agent feedback provided by an agent's previous communication with the user, the agent feedback identifying the stored communication elements that contributed to a positive resolution. Claim 8: The computing system for adaptive, individualized, and contextualized text-to- speech according to claim 1, wherein the at least one prosody element further includes a prosody element selected from the group consisting of tone, cadence, pitch, volume, pronunciation, speaking rate, articulation, fluency, intensity, inflection, and resonance, that is predicted to be appeasing to the user in accordance with the assigned at least one contextual category. Claim 8: The computing system for adaptive, individualized, and contextualized text-to-speech according to claim 1, wherein the at least one prosody element further includes a prosody element selected from the group consisting of tone, cadence, pitch, volume, pronunciation, speaking rate, articulation, fluency, intensity, inflection, and resonance, that is predicted to be appeasing to the user in accordance with the assigned at least one contextual category. Claim 9: The computing system for adaptive, individualized, and contextualized text-to- speech according to claim 1, wherein the timbre selected is predicted to be appeasing to the user. Claim 9: The computing system for adaptive, individualized, and contextualized text-to-speech according to claim 1, wherein the timbre selected is predicted to be appeasing to the user. Claim 10: The computing system for adaptive, individualized, and contextualized text-to- speech according to claim 1, wherein the at least one trained neural network categorizes the at least one communication element of the plurality of communication elements according to an identified speech pattern, wherein the identified speech pattern includes a regional dialect and based thereon selecting the speech pattern based on the regional dialect. Claim 10: The computing system for adaptive, individualized, and contextualized text-to-speech according to claim 1, wherein the at least one trained RNN categorizes the at least one communication element of the plurality of communication elements according to an identified speech pattern, wherein the identified speech pattern includes a regional dialect and based thereon selecting the speech pattern based on the regional dialect. Claim 11: The computing system for adaptive, individualized, and contextualized text-to- speech according to claim 1, wherein the text-to-speech processing utilizes a deployed text-to- speech model, and wherein the program instructions are further executable to continuously modify a prosodic configuration of the deployed text-to-speech model, based on receiving future input audio data, providing the one or more future audio outputs, and identifying future reactions of the user, with each future reaction of the future reactions having a respective measured quantifiable quality score corresponding to the one or more future audio outputs. Claim 11: The computing system for adaptive, individualized, and contextualized text-to- speech according to claim 1, wherein the text-to-speech processing utilizes a deployed text-to- speech model, and wherein the program instructions are further executable to continuously modify a prosodic configuration of the deployed text-to-speech model, based on receiving future input audio data, providing the one or more future audio outputs, and identifying future reactions of the user, with each future reaction of the future reactions having a respective measured quantifiable quality score corresponding to the one or more future audio outputs. Claim 12: The computing system for adaptive, individualized, and contextualized text-to- speech according to claim 1, wherein the quantifiable quality score is further incorporated into modifying a prosodic configuration of the a natural language generation model that is used to perform the generating of the text, the prosodic configuration using the quantifiable quality score to incorporate one or more speech synthesis markup language (SSML) tags to facilitate producing the revised speech pattern. Claim 12: The computing system for adaptive, individualized, and contextualized text-to-speech according to claim 1, wherein the quantifiable quality score is further incorporated into modifying a prosodic configuration of the a natural language generation model that is used to perform the generating of the text, the prosodic configuration using the quantifiable quality score to incorporate one or more speech synthesis markup language (SSML) tags to facilitate producing the revised speech pattern. Claim 13: The computing system for adaptive, individualized, and contextualized text-to- speech according to claim 1, wherein the user device comprises an external device that is external to the computing system, and wherein the program instructions are further executable to transmit the audio output via a telephonic communication means. Claim 13: The computing system for adaptive, individualized, and contextualized text-to-speech according to claim 1, wherein the user device comprises an external device that is external to the computing system, and wherein the program instructions are further executable to transmit the audio output via a telephonic communication means. Claim 14: A computing system for adaptive, individualized, and contextualized text-to- speech, the system comprising: a memory; one or more processors in communication with the memory; and program instructions executable by the one or more processors via the memory to: iteratively train, using training data, at least neural network to categorize communication elements from audio data, the neural network being trained to perform natural language processing on the audio data to generate transcribed data from the audio data, parse the generated transcribed data, and perform a reduction analysis on the transcribed data in order to interpret the transcribed data to derive the communication elements for categorization, the communication elements including intent, context, emotion, and other circumstantial factors, the training including: inserting the training data into an iterative training and testing loop to predict a target variable; and repeatedly predicting the target variable during each iteration of the training and testing loop, wherein each iteration of the training and testing loop has differing weights applied to one or more nodes of the neural network, each of the differing weights being updated with each iteration of the training and testing loop to reduce error in predicting the target variable and improve predictability of the neural network; deploy the at least one trained neural network to facilitate categorization of the communication elements; receive, during a single interaction between the user and the computing system and via a telephonic communication means, an initial audio input provided via a user device by a user; classify a context of the audio input, the context indicating a purpose for receiving the audio input; identify a saved profile associated with a speech pattern, the saved profile corresponding to a highest level of positive resolutions resulting from previously received audio having similarly classified contexts to the context of the received audio input; provide, using text-to-speech processing, an audible output to the user via the user device, the audible output incorporating one or more speech pattern qualities of the saved profile into a speech pattern, the speech pattern qualities including at least one prosody element, the at least one prosody element including a timbre intended to elicit certain emotions of the user associated with a desired outcome, the timbre being based on the context assigned to the audio input; and modify the speech pattern in response to subsequent audio inputs provided by the user, the modifying occurring in real time during the single interaction between the user and the computing system. Claim 14: A computing system for adaptive, individualized, and contextualized text-to-speech, the system comprising: a memory; one or more processors in communication with the memory; and program instructions executable by the one or more processors via the memory to: iteratively train, using training data, at least recurrent neural network (RNN) to categorize communication elements from audio data, the RNN being trained to perform natural language processing on the audio data to generate transcribed data from the audio data, parse the generated transcribed data, and perform a reduction analysis on the transcribed data in order to interpret the transcribed data to derive the communication elements for categorization, the communication elements including intent, context, emotion, and other circumstantial factors, the training including: inserting the training data into an iterative training and testing loop to predict a target variable; and repeatedly predicting the target variable during each iteration of the training and testing loop, wherein each iteration of the training and testing loop has differing weights applied to one or more nodes of the RNN, each of the differing weights being updated with each iteration of the training and testing loop to reduce error in predicting the target variable and improve predictability of the RNN; deploy the at least one trained RNN to facilitate categorization of the communication elements; receive, during a single interaction between the user and the computing system and via a telephonic communication means, an initial audio input provided via a user device by a user; classify a context of the audio input, the context indicating a purpose for receiving the audio input; identify a saved profile associated with a speech pattern, the saved profile corresponding to a highest level of positive resolutions resulting from previously received audio having similarly classified contexts to the context of the received audio input; provide, using text-to-speech processing, an audible output to the user via the user device, the audible output incorporating one or more speech pattern qualities of the saved profile into a speech pattern, the speech pattern qualities including at least one prosody element, the at least one prosody element including a timbre intended to elicit certain emotions of the user associated with a desired outcome, the timbre being based on the context assigned to the audio input; and modify the speech pattern in response to subsequent audio inputs provided by the user, the modifying occurring in real time during the single interaction between the user and the computing system. Claim 15: The computing system for adaptive, individualized, and contextualized text-to- speech according to claim 14, wherein the program instructions are further executable to access one or more saved profiles within a prosody category selected from one or more prosody categories, each prosody category being associated with different prosody element, the one or more saved profiles comprising the identified saved profile, and wherein the one or more saved profiles are contextually classified based on one or more purposes for user interactions between users and the computing system. Claim 15: The computing system for adaptive, individualized, and contextualized text-to-speech according to claim 14, wherein the program instructions are further executable to access one or more saved profiles within a prosody category selected from one or more prosody categories, each prosody category being associated with different prosody element, the one or more saved profiles comprising the identified saved profile, and wherein the one or more saved profiles are contextually classified based on one or more purposes for user interactions between users and the computing system. Claim 16: The computing system for adaptive, individualized, and contextualized text-to- speech according to claim 14, wherein the context of the audio input is provided by the user based on the user providing an input indicating a reason for initiating the single interaction. Claim 16: The computing system for adaptive, individualized, and contextualized text-to-speech according to claim 14, wherein the context of the audio input is provided by the user based on the user providing an input indicating a reason for initiating the single interaction. Claim 17: A computer-implemented method for adaptive, individualized, and contextualized text-to-speech, the computer-implemented method comprising: iteratively training, using training data, at least one neural network to categorize communication elements from audio data, the neural network being trained to perform natural language processing on the audio data to generate transcribed data from the audio data, parse the generated transcribed data, and perform a reduction analysis on the transcribed data in order to interpret the transcribed data to derive the communication elements for categorization, the communication elements including intent, context, emotion, and other circumstantial factors, the training including: inserting the training data into an iterative training and testing loop to predict a target variable; and repeatedly predicting the target variable during each iteration of the training and testing loop, wherein each iteration of the training and testing loop has differing weights applied to one or more nodes of the neural network, each of the differing weights being updated with each iteration of the training and testing loop to reduce error in predicting the target variable and improve predictability of the neural network; deploying the at least one trained neural network to facilitate categorization of the communication elements; receiving, in real-time from a user device, input audio data of a user, the input audio data comprising a plurality of communication elements; applying the input audio data to the at least one trained neural network to categorize one or more communication elements of the plurality of communication elements, the categorizing including assigning at least one contextual category to a communication element of the plurality of communication elements; generating text comprising a response to one or more of the plurality of communication elements, the response including one or more individualized and contextualized qualities predicted to provide an optimal outcome based at least in part on (i) the assigned at least one contextual category and (ii) the user from which the input audio data is received; implementing text-to-speech processing of the generated text, the text-to-speech processing producing an audio output comprising (a) the response and (b) a speech pattern predicted to facilitate the optimal outcome, the speech pattern including at least one prosody element, the at least one prosody element including a timbre intended to elicit certain emotions of the user associated with a desired outcome, the timbre being based on the at least one contextual category assigned to the communication element of the plurality of communication elements; and providing, to the user device, the audio output and based thereon measure a reaction of the user in response to the speech pattern and the timbre of the audio output according to a quantifiable quality score and using the quantifiable quality score to modify future iterations of the text-to-speech processing in order to provide one or more future audio outputs comprising a revised speech pattern. Claim 17: A computer-implemented method for adaptive, individualized, and contextualized text-to-speech, the computer-implemented method comprising: iteratively training, using training data, at least one recurrent neural network (RNN) to categorize communication elements from audio data, the RNN being trained to perform natural language processing on the audio data to generate transcribed data from the audio data, parse the generated transcribed data, and perform a reduction analysis on the transcribed data in order to interpret the transcribed data to derive the communication elements for categorization, the communication elements including intent, context, emotion, and other circumstantial factors, the training including: inserting the training data into an iterative training and testing loop to predict a target variable; and repeatedly predicting the target variable during each iteration of the training and testing loop, wherein each iteration of the training and testing loop has differing weights applied to one or more nodes of the RNN, each of the differing weights being updated with each iteration of the training and testing loop to reduce error in predicting the target variable and improve predictability of the RNN; deploying the at least one trained RNN to facilitate categorization of the communication elements; receiving, in real-time from a user device, input audio data of a user, the input audio data comprising a plurality of communication elements; applying the input audio data to the at least one trained RNN to categorize one or more communication elements of the plurality of communication elements, the categorizing including assigning at least one contextual category to a communication element of the plurality of communication elements; generating text comprising a response to one or more of the plurality of communication elements, the response including one or more individualized and contextualized qualities predicted to provide an optimal outcome based at least in part on (i) the assigned at least one contextual category and (ii) the user from which the input audio data is received; implementing text-to-speech processing of the generated text, the text-to-speech processing producing an audio output comprising (a) the response and (b) a speech pattern predicted to facilitate the optimal outcome, the speech pattern including at least one prosody element, the at least one prosody element including a timbre intended to elicit certain emotions of the user associated with a desired outcome, the timbre being based on the at least one contextual category assigned to the communication element of the plurality of communication elements; and providing, to the user device, the audio output and based thereon measure a reaction of the user in response to the speech pattern and the timbre of the audio output according to a quantifiable quality score and using the quantifiable quality score to modify future iterations of the text-to-speech processing in order to provide one or more future audio outputs comprising a revised speech pattern. Claim 18: The computer-implemented method for adaptive, individualized, and contextualized text-to-speech according to claim 17, wherein the input audio data includes multiple speech utterances and based thereon performing the receiving, performing, and applying after each speech utterance of the multiple speech utterances, and wherein the assigning is adapted after each speech utterance in context with the multiple speech utterances received thus far from the user. Claim 18: The computer-implemented method for adaptive, individualized, and contextualized text-to-speech according to claim 17, wherein the input audio data includes multiple speech utterances and based thereon performing the receiving, performing, and applying after each speech utterance of the multiple speech utterances, and wherein the assigning is adapted after each speech utterance in context with the multiple speech utterances received thus far from the user. Claim 19: The computer-implemented method for adaptive, individualized, and contextualized text-to-speech according to claim 17, wherein the quantifiable quality score is further incorporated into modifying a prosodic configuration of the a natural language generation model that is used to perform the generating of the text, the prosodic configuration using the quantifiable quality score to incorporate one or more speech synthesis markup language (SSML) tags to facilitate producing the revised speech pattern. Claim 12: The computing system for adaptive, individualized, and contextualized text-to-speech according to claim 1, wherein the quantifiable quality score is further incorporated into modifying a prosodic configuration of the a natural language generation model that is used to perform the generating of the text, the prosodic configuration using the quantifiable quality score to incorporate one or more speech synthesis markup language (SSML) tags to facilitate producing the revised speech pattern. Claim 20: The computer-implemented method for adaptive, individualized, and contextualized text-to-speech according to claim 17, wherein the user device comprises an external device that is external to the computing system, and wherein the program instructions are further executable to transmit the audio output via a telephonic communication means. Claim 13: The computing system for adaptive, individualized, and contextualized text-to-speech according to claim 1, wherein the user device comprises an external device that is external to the computing system, and wherein the program instructions are further executable to transmit the audio output via a telephonic communication means. Allowable Subject Matter Claims 1-20 would be allowable if rewritten or amended or a terminal disclaimer filed to overcome the rejection(s) under non-statutory double patenting set forth in this Office action. More specifically the limitation of independent claims 1, 14, and 17 regarding: “implement text-to-speech processing of the generated text, the text-to- speech processing producing an audio output comprising (a) the response and (b) a speech pattern predicted to facilitate the optimal outcome, the speech pattern including at least one prosody element, the at least one prosody element including a timbre intended to elicit certain emotions of the user associated with a desired outcome, the timbre being based on the at least one contextual category assigned to the communication element of the plurality of communication elements.” are not mentioned in any prior art found either alone or in an obvious combination thereof. Dependent claims 2-13, 15-16, 18-20 are rejected by virtue of their dependencies to the independent claims. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHASHIDHAR S MANOHARAN whose telephone number is (571)272-6772. The examiner can normally be reached M-F 8:00-4:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Flanders can be reached at 571-272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SHASHIDHAR SHANKAR MANOHARAN/Examiner, Art Unit 2655 /ANDREW C FLANDERS/Supervisory Patent Examiner, Art Unit 2655
Read full office action

Prosecution Timeline

Feb 26, 2025
Application Filed
Aug 27, 2026
Non-Final Rejection mailed — §DP (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737537
GOVERNANCE AND CONFIDENCE ASSESSMENT OF LLM
2y 6m to grant Granted Sep 15, 2026
Patent 12682890
MASK-CONFORMER AUGMENTING CONFORMER WITH MASK-PREDICT DECODER UNIFYING SPEECH RECOGNITION AND RESCORING
2y 4m to grant Granted Jul 14, 2026
Patent 12682173
MODULAR FRAMEWORK FOR EVALUATING LANGUAGE MODELS
2y 4m to grant Granted Jul 14, 2026
Study what changed to get past this examiner. Based on 3 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
80%
Grant Probability
99%
With Interview (+33.3%)
2y 2m (~7m remaining)
Median Time to Grant
Low
PTA Risk
Based on 5 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month