Prosecution Insights
Last updated: August 17, 2026
Application No. 18/599,431

RECIPIENT-SPECIFIC VOICE TONE ADJUSTMENT IN TELEPHONY

Non-Final OA §101§103
Filed
Mar 08, 2024
Examiner
WOZNIAK, JAMES S
Art Unit
2655
Tech Center
2600 — Communications
Assignee
International Business Machines Corporation
OA Round
3 (Non-Final)
59%
Grant Probability
Moderate
3-4
OA Rounds
1y 2m
Est. Remaining
98%
With Interview

Examiner Intelligence

Grants 59% of resolved cases
59%
Career Allowance Rate
237 granted / 403 resolved
-3.2% vs TC avg
Strong +40% interview lift
Without
With
+39.5%
Interview Lift
resolved cases with interview
Typical timeline
3y 7m
Avg Prosecution
18 currently pending
Career history
431
Total Applications
across all art units

Statute-Specific Performance

§101
19.4%
-20.6% vs TC avg
§103
42.9%
+2.9% vs TC avg
§102
16.1%
-23.9% vs TC avg
§112
16.8%
-23.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 403 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment In response to the Final Office Action from 3/3/2026, Applicant has filed a Request for Continued Examination (RCE) on 4/11/2026. In this reply, Applicant has amended independent claims 1, 7, and 15 to further recite that the extracting of voice tone data comprises analyzing audio data of the plurality of voice samples using a machine learning model trained to identify characteristics from the plurality of voice samples to derive one or more tone-related characteristics and storing the one or more tone-related characteristics as the voice tone data and that the speech output generated during a voice communication comprises audio generated from the text using a text-to-speech model, and wherein the text-to-speech model generates the speech output with a voice tone corresponding to the voice tone data by producing output audio data having one or more tone-related characteristics corresponding to the voice tone data. Applicant has also argued that the prior art of record fails to teach the limitations added via the instant amendment (Remarks, Pages 20-21). These arguments have been fully considered, however, are moot with respect to the new grounds of rejection responsive to the amended claims and further in view of Li, et al. (U.S. PG Publication: 2025/0349282 A1). Response to Arguments Patent Subject Matter Eligibility Rejections under 35 U.S.C. 101: Applicant traverses the rejection of claims 1-20 under 35 U.S.C. 101 relying upon multiple arguments. First, Applicant contends that under the 2019 Patent Subject Matter Eligibility Guidelines of Step 2A prong 1, the rejection relies only upon conclusions that do not rely upon evidence showing how a human could perform the claimed operations as recited. Applicant contends that the characterization as listening, writing and speaking does not address the recited analysis of audio data or the storage and reuse of derived tone-related characteristics (Remarks, Page 11). In response, it is noted that the proceeding 35 U.S.C. 101 provides an explanation for how these steps can be practically performed by a human being listening and then mentally noting and remembering characteristics of a person's tone that can be reproduced in an imitation. For example, a person could remember the particular tone or emphasis placed on certain words, accents, and general volume levels used by a target speaker. The use of a machine learning model (MLM) to perform such processing is not addressed as part of the analysis under Step 2A prong 1. Instead, the use of a generic MLM merely stands in for/automates otherwise human activity with a computer and like the speech-to-text and text-to-speech models was also not invented or improved by Applicant per their own admission (Paragraph 0043- "examples of presently available techniques to extract voice tone data are a voice pitch analyzer which determines voice pitch (i.e., frequency) and variations in pitch, and machine learning models trained to identify emotions such as happiness, sadness, anger, or excitement within a plurality of voice samples"). Accordingly, Applicant arguments directed towards step 2A prong 1 are not found to be persuasive. With respect to step 2A prong 2, Applicant traverses the alleged Office Action position that speech-to-text and text-to-speech models are generic and presently available and therefore do not provide an invention concept on the grounds that the claims require a specific use of these components to analysis audio data to determine and store tone characteristics (Remarks, Page 11). In response, it is noted that Applicant arguments mischaracterize the position of record under step 2A prong 2. The use of the claimed models acts as mere computer automations of processes that can otherwise be performed by a human under the broadest reasonable interpretation (BRI). See, for example, Mere automation of manual processes, such as using a generic computer to process an application for financing a purchase, Credit Acceptance Corp. v. Westlake Services, 859 F.3d 1044, 1055, 123 USPQ2d 1100, 1108-09 (Fed. Cir. 2017) and Recentive Analytics, Inc. v. Fox Corp. (Fed. Cir. April 18, 2025)- “Machine learning is a burgeoning and increasingly important field and may lead to patent-eligible improvements in technology. Today, we hold only that patents that do no more than claim the application of generic machine learning to new data environments, without disclosing improvements to the machine learning models to be applied, are patent ineligible under § 101." The statement that these models are admitted by Applicant to be presently available pertains to the analysis in step 2B where it is found that these models are not directed towards an inventive concept to direct the claimed invention towards significantly more than an abstract idea under the BRI. Moving to the 2B analysis under the 2019 PEG, Applicant next traverses the alleged Office Action position that the claimed elements do not amount to a generic automation of a human process and argues that this position does not address the specific manner in which the claims analyze audio to determine tone-related characteristics for use during audio data generation (Remarks, Pages 11-12). In response, it is noted that Applicant remarks related to 2B more accurately pertain to step 2A prong 2. Step 2B/Berkheimer asks whether the remaining elements after the abstract idea analysis of Step 2A prong 1 are well-known, routine, and conventional. In the case of the instant claims, the answer here is "yes" as Applicant has admitted that all of the models utilized in a combination are "presently available." As per the claims requiring some "specific manner" of processing, it is unclear what limitations provide the specifics of how audio analysis and synthesis is being carried out in a technical manner so as to preclude performance. The claimed invention also places no limits on how any of the "presently available" models are being used in a specific manner outside of generic/high level "using". As such, the claims are drafted at a high level of abstraction so as to read on a human mental process under the BRI. Accordingly, Applicant arguments directed towards step 2B are not found to be persuasive. In response to Applicant arguments directed towards the precedential Ex parte Desjardins decision (Remarks, Pages 12-13), while it is agreed that the instant amendments do introduce a machine learning model for voice tone analysis and extraction, the Desjardins decision relates to improvements in artificial intelligence. It is unclear how Applicant's claimed machine learning model constitutes an improvement when it was admitted by Applicant to be "presently available" (Paragraph 0043). Moreover, similar to the Recentive Analytics decision, the claims do no more than claim "the application of generic machine learning" without even a new data environment where no improvements to the machine learning model are neither claimed nor disclosed. Accordingly, these arguments are not found to be persuasive. Applicant eligibility arguments found on pages 13-18 are largely similar to those already considered (see Final Office Action Remarks, pages 3-6) with the exception that the recited claims have been updated to reflect the amended claim language. In particular regards to the amended claim language, how a human can practically perform such processes under the BRI, and how the remaining elements/models have been addressed under step 2A prong 2 and 2B, Applicant is directed towards the preceding remarks. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1, 3-9, 11-15, and 17-20 are rejected under 35 U.S.C. 101 because under the broadest reasonable interpretation (BRI), the claimed invention is directed to an abstract idea without significantly more. Independent Claims 1, 7, and 15 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claims regard a process that, as drafted under its BRI, covers performance of the limitations as a mental process, but for the recitation of generic computer components/models. In regards to the process/functionality of claims 1, 7, and 15 the claimed functionality could be practiced as a mental process in the following manner: extracting, from a plurality of voice samples of a user, voice tone data corresponding to a specified voice tone of the user wherein converting, generating, during a voice communication, a speech output corresponding to the text, wherein the speech output comprises audio generated from the text (a human can read out the written text and vocally reproduce a voice tone (e.g., by raising their pitch, volume, cadence, or tone) by recalling the tone analysis). This judicial exception is not integrated into a practical application. Outside of the identified abstract idea, the claimed invention only recites processors and storage media which amount to no more than mere instructions to implement an otherwise abstract idea using generic computer components and a mention of generic speech-to-text, text-to-speech models, and machine learning models that are a mere machine/software automations recited at a high level of abstraction of otherwise practical human processes of mental voice characteristic analysis, transcription, and reading/speaking processes. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The above identified additional generic computer components are no more than mere instructions to apply the exception using generic computer components that are well-known, routine, and conventional as is evidenced by Bancorp Services v. Sun Life (Fed. Cir. 2012) and Alice Corp. v. CLS Bank (2014). As for evidence that the speech-to-text and text-to-speech models are well-known, routine, and conventional activity that does not direct patent ineligible subject matter to significantly more than the abstract idea, see the following prior art: Subramanian, et al. (U.S. PG Publication: 2007/0208569 A1- Paragraph 0051-0052- speech-to-text is "known" and Paragraph 0082- text-to-speech synthesis is "known"), Patel, et al. (U.S. PG Publication: 2016/0379622 A1- Paragraph 0063- text-to-speech has many "known techniques"), and Jaroker (U.S. PG Publication: 2005/0010407 A1- Paragraph 0043- speech recognition is "generally known"). Moreover, see Applicant admission that the models used are “presently available”- machine learning model (see Paragraphs 0014 and 0043), speech-to-text (see Paragraphs 0016 and 0045), text-to-speech (see Paragraphs 0048 and 0050). Accordingly, independent claims 1, 7, and 15 under the BRI are not patent eligible under 35 U.S.C. 101. The remaining dependent claims do not add patent eligible subject matter to their respective parent claims and have also been rejected under 35 U.S.C. 101: Claims 3, 6, 11, 14, 17, and 20 further limit the data being processed in the independent claims that can be understood and analyzed by a human. Claims 4, 12, and 18 regard a human mentally deciding on a recipient for their communication. Claims 5, 13, and 19 regard a human mentally deciding upon a tone that they used in the past for a particular participant. Claim 8 regards generic computer structures as addressed in claim 1 and a network transfer that does not patentably limit the recited computer-program process (nor would such transfer make the claim eligible if claimed as part of the program instructions). Claim 9 regards a human deciding on how much a conversion serviced was accessed and mentally calculating a bill at an established rate (e.g., per use, time-based). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 3-8, 11-15, and 17-20 are rejected under 35 U.S.C. 103 as being unpatentable over Subramanian, et al. (U.S. PG Publication: 2007/0208569 A1) in view of Li, et al. (U.S. PG Publication: 2025/0349282 A1). With respect to Claim 1, Subramanian discloses: extracting, from a plurality of voice samples of a user , voice tone data corresponding to a specified voice tone of the user (speech features that are extracted from speech communication samples of the speaker/user are used to extract various specified types of voice tone information (e.g., pitch, tone, cadence, amplitude, etc., Paragraphs 0051-0052; such voice patters are labeled and used to populate a database for synthesis, Paragraphs 0065, 0071, 0076, 0090, 0095, and 0110) and storing one or more tone-related characteristics as the voice tone data (dictionary/database stored for a specific user (e.g., a profile), Paragraphs 0071, 0081, and 0094-0095); converting, using a speech-to-text model, a speech input to corresponding text (speech recognition model that converts a speech communication into text, Paragraph 0035, 0049, and 0051-0052); and generating, during a voice communication, a speech output corresponding to the text, wherein the speech output comprises audio generated from the text using a text-to-speech model and wherein the text-to-speech model generates the speech output with a voice tone corresponding to the voice tone data by producing output audio data having one or more tone-related characteristics corresponding to the voice tone data (generation and playback relying on a text to speech model that is modulated using the voice patterns including tone information, Paragraphs 0047, 0076, 0082-0084, 0114, and 0117; telephony/voice communication application, Paragraphs 0006 and 0105). Subramanian does not teach the more specific extraction process involving analyzing audio data of the plurality of voice samples using a machine learning model trained to identify characteristics from the plurality of voice samples to derive one or more tone-related characteristics. Li, however, discloses the extracting comprises analyzing audio data of the plurality of voice samples using a machine learning model trained to identify characteristics from the plurality of voice samples to derive one or more tone-related characteristics and storing the one or more tone-related characteristics as the voice tone data (feature extraction module (Fig. 2, Element 202) that extracts prosodic voice tone features from a "reference speaker" (reference speech, Fig. 2, Element 208) audio for text-to-speech (TTS) synthesis in a "personalized voice that captures the identity of the speaker, as well as the acoustic and prosodic characteristics of the natural speaking voice of the target speaker, Paragraphs 0063, 0065; note that the feature extractor may be trained by machine learning as disclosed in Paragraph 0057). Note also that the prosodic/tone information extracted from a target speaker in Li is used with a text-to-speech model in a voice chat application that generates the speech output with a voice tone corresponding to the voice tone data by producing output audio data having one or more tone-related characteristics corresponding to the voice tone data (see Fig. 6, Element 206 and Paragraph 0065-0066). Subramanian and Li are analogous art because they are from a similar field of endeavor in personalized text-to-speech synthesis. Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date to utilize the machine-learning based approach for prosody/tone extraction for voice cloning in text-to-speech synthesis as taught by Li in the user tone extraction taught by Subramanian to provide a predictable result of implanting a model that can be trained to learn and extract optimal features for speech synthesis to capture personalized prosodic characteristics of a target voice. With respect to Claim 3, Subramanian further discloses: The computer-implemented method of claim 1, wherein the voice tone data is maintained in a user-specific voice tone repository (dictionary/database for a specific user (e.g., a profile), Paragraphs 0071, 0081, and 0094-0095). With respect to Claim 4, Subramanian further discloses: The computer-implemented method of claim 1, further comprising: selecting, for use in the voice communication with a communication recipient, the voice tone data (the voice tone information is selected based upon particular scenarios such as speaker and "audience," Paragraphs 0042-0045 and 0047-0048). With respect to Claim 5, Subramanian further discloses: The computer-implemented method of claim 4, wherein the voice tone data was previously selected for use in a previous voice communication with the communication recipient (user profiles that include voice tone data used/frequently used in past communications and learned, Paragraphs 0044-0045, 0047, 0049, 0063, and 0071; Fig. 4). With respect to Claim 6, Subramanian further discloses: The computer-implemented method of claim 4, wherein the voice tone data is default voice tone data ("generic or default profile" including voice tone data, Paragraphs 0049 and 0117; see the last line of the rightmost column in the audience profiles shown in Fig. 4). Claim 7 is directed towards an embodiment that implements the method of claim 1 as one or more computer readable storage media storing processor-executable program instructions, and thus, is rejected under similar rationale. Moreover, Subramanian teaches method implementation as a computer-readable medium storing processor-executable instructions (Paragraph 0021). With respect to Claim 8, Subramanian further discloses: The computer program product of claim 7, wherein the stored program instructions are stored in a computer readable storage device in a data processing system (processor embodied in a data processing system/computer also having a memory storing program code, Paragraph 0024), and wherein the stored program instructions are transferred over a network from a remote data processing system (this wherein clause described an implementation environment that does not patentably limit the claimed product claim (i.e., one or more computer readable storage media, and program instructions collectively stored on the one or more computer readable storage media) because it does not limit or modify the structure of the claimed computer readable medium, and thus, need not be addressed with prior art to render the added subject matter of claim 8 unpatentable. It should be noted, however, that Paragraph 0023-0024 does teach communication of program instructions through a network). Claims 11-14 contain subject matter respectively similar to Claims 3-6, and thus, are rejected under similar rationale. Claim 15 is directed towards an embodiment that implements the method of claim 1 as a computer system comprising a processor and one or more computer readable storage media storing processor-executable program instructions, and thus, is rejected under similar rationale. Moreover, Subramanian teaches method implementation as a computer system comprising one or more processors and a computer-readable medium storing program code (Paragraphs 0021 and 0024). Claims 17-20 contain subject matter respectively similar to Claims 3-6, and thus, are rejected under similar rationale. Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Subramanian, et al. in view of Li, et al. and further in view of Yu, et al. (U.S. PG Publication: 2021/0256958 A1). With respect to Claim 9, Subramanian in view of Li teaches the computer program product of claim 7 comprising one or more computer readable storage media and program instructions stored thereupon. Claim 9 relates to the stored program instructions also being stored in a "stored in a computer readable storage device in a server data processing system, and wherein the stored program instructions are downloaded in response to a request over a network to a remote data processing system for use in a computer readable storage device associated with the remote data processing system." Parent claim 7, however, is not directed towards a network computing system (i.e., such a system is outside the scope of the claimed invention) nor does the wherein clause modify the recited "one or more computer-readable storage media" or add to the instructions stored on that media. As such, the wherein clause is not patentably limiting. Claim 9 does include further program instructions comprising: program instructions to meter use of the program instructions associated with the request; and program instructions to generate an invoice based on the metered use. These program instructions while not taught by Subramanian in view of Li are taught by Yu. Specifically, Yu discloses software resources for metering that tracks the usage of resources and "billing or invoicing" for such consumption of these resources (Paragraph 0064). Subramanian, Li, and Yu are analogous art because they are from a similar field of endeavor in voice conversion. Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date to modify the teachings of Subramanian in view of Li to include the metering/billing instructions taught by Yu to provide a predictable result of allowing a developer to profit off of and recover development costs from implementing a voice service. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Zhou, et al. (U.S. PG Publication: 2020/0410976 A1)- discloses a neural network-based classifier that is used in a content extraction block to extract pitch contours and intonation that is then used to train a vocal model neural network for speech synthesis (Paragraphs 0039-0041; Fig. 1). Any inquiry concerning this communication or earlier communications from the examiner should be directed to JAMES S WOZNIAK whose telephone number is (571)272-7632. The examiner can normally be reached 7-3, off alternate Fridays. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant may use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Flanders can be reached at (571)272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. JAMES S. WOZNIAK Primary Examiner Art Unit 2655 /JAMES S WOZNIAK/Primary Examiner, Art Unit 2655
Read full office action

Prosecution Timeline

Show 3 earlier events
Dec 11, 2025
Applicant Interview (Telephonic)
Dec 22, 2025
Response Filed
Mar 03, 2026
Final Rejection mailed — §101, §103
Mar 26, 2026
Examiner Interview Summary
Mar 26, 2026
Applicant Interview (Telephonic)
Apr 11, 2026
Request for Continued Examination
Apr 13, 2026
Response after Non-Final Action
Jun 09, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12700401
Speech Wakeup Method and Apparatus, Device, Storage Medium, and Program Product
2y 5m to grant Granted Aug 04, 2026
Patent 12694222
USER INTERFACE AUTOMATION USING NATURAL LANGUAGE
2y 3m to grant Granted Jul 28, 2026
Patent 12682889
METHOD, DEVICE, AND PROGRAM FOR PROVIDING MATCHING INFORMATION THROUGH ANALYSIS OF SOUND INFORMATION
2y 8m to grant Granted Jul 14, 2026
Patent 12670901
SYSTEMS AND METHODS FOR USING CONTEXTUAL INTERIM RESPONSES IN CONVERSATIONS MANAGED BY A VIRTUAL ASSISTANT SERVER
2y 7m to grant Granted Jun 30, 2026
Patent 12664980
METHOD AND SYSTEM FOR CONVERTING PLANT PROCEDURES TO VOICE INTERFACING SMART PROCEDURES
2y 9m to grant Granted Jun 23, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
59%
Grant Probability
98%
With Interview (+39.5%)
3y 7m (~1y 2m remaining)
Median Time to Grant
High
PTA Risk
Based on 403 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month