DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
In response to the Non-Final Office Action from 4/27/2025, Applicant has filed an amendment on 7/17/2026. In this reply, Applicant has amended independent claim 1 to specify that the audio segment on interest is analyzed to determine a prosody of the audio segment of interest that includes stress, rhythm, and intonation along with significant amendments to the text string conversion resulting in an audio response "having a characteristic voice of a non-human social agent that is independent of a voice of the human speaker, wherein a rendering of the audio segment of interest in the audio response is modified, using the determined prosody of the audio segment of interest, to replicate the pronunciation of the audio segment of interest by the human speaker, and wherein portions of the audio response other than the audio segment of interest retain the characteristic voice of the non-human social agent unmodified by the prosody of the audio segment of interest." Independent claims 8 and 15 were similarly amended. New claims 27-30 have been added while claims 24-25 have been cancelled.
Applicant has also argued that the prior art of record fails to teach the limitations added via the instant amendment because the secondary reference reproduces the same voice attributes of recorded speaker for all words where the amended claim confines a known speaker's attributes to a single segment of interest while other segments in the same response remain in a separate, independent voice (Remarks, Pages 12-14).
These arguments have been fully considered, however, are not found to be persuasive for the reasons noted in the Response to Arguments section.
Response to Arguments
Applicant argues that the prior art combination of Peddinti, et al. (U.S. PG Publication: 2022/0284882 A1) and Tischer (U.S. PG Publication: 2004/0111271 A1) fails to teach the limitation added to the independent claims regarding "text string conversion resulting in an audio response "having a characteristic voice of a non-human social agent that is independent of a voice of the human speaker, wherein a rendering of the audio segment of interest in the audio response is modified, using the determined prosody of the audio segment of interest, to replicate the pronunciation of the audio segment of interest by the human speaker, and wherein portions of the audio response other than the audio segment of interest retain the characteristic voice of the non-human social agent unmodified by the prosody of the audio segment of interest." In particular, Applicant examines the teachings of Tischer in detail and argues that Tischer is concerned with emphasis, rhythm and intonation belong to the same voice of the recoded speaker and so contains "no teaching or suggestion of confining the use of a known speaker's recorded attributes to a single segment of a response while other portions of that same response remain in a separate, independent voice (Remarks, Pages 13-14).
In response, it is noted that these arguments amount to a piecemeal deconstruction of a prior art rejection based upon a combination under 35 U.S.C. 103. It is noted that one cannot show nonobviousness by attacking references individually where the rejections are based on combinations of references. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981); In re Merck & Co., 800 F.2d 1091, 231 USPQ 375 (Fed. Cir. 1986). The argued limitation here is addressed with the combined teachings of Peddinti and Tischer.
First, Applicant is directed to consider the discussion of Peddinti regarding text-to-speech (TTS) audio produced by a "digital assistant" that includes a preferred user pronunciation for a "particular word" among other speech reproduced by a digital assistant (Paragraph 0053). It should also be noted that the TTS pronunciation utilizes "prosody and/or speech characteristics" (Paragraph 0040). Peddinti, thus teaches use of a speaker's/user's pronunciation/prosody for a "particular" or "unique" word among other digital assistant TTS audio and addresses the Applicant's piecemeal analysis that Tischer does not address confining prosodic information to a "single segment of a response while other portions of that same response remain in a separate, independent voice."
Peddinti only lacks details of the prosody as specifically including "stress, rhythm, and intonation as claimed though it should be noted that the ordinary and customary meaning of the term includes "the rhythmic and intonational aspect of language" (see https://www.merriam-webster.com/dictionary/prosody). Tischer resolves this minor deficiency in Peddinti by disclosing prosodic information in the form of emphasis/stress, rhythm, and intonations (see at least cited Paragraph 0037) to enable recognizable clarity in the unique words provided in Peddinti using well-known specifics of prosodic information. Accordingly, when taken in combination, Peddinti and Tischer results in a preferred pronunciation of a unique or particular word by a digital assistant among other digital assistant speech using prosody where the specific types of prosodic information are established by Tischer and the Applicant's piecemeal arguments directed towards Tischer are not found to be persuasive.
Applicant next argues that the relied upon combination is not supported by an adequate motivation to combine. In particular, Applicant contends that Tischer does not teach that the benefits mentioned in the Office Action would result from applying prosodic attributes to a non-human voice and that Tischer teaches away from that purpose because the proposed combination would frustrate the purpose for which Tischer's recorded prosodic attributes are used (Remarks, Pages 14-15).
In response, Applicant is again directed towards the previously cited teachings of Peddinti that describe "particular" or "unique" words pronounced in the style of a user utilizing prosody. Tischer provides the specific prosodic information for insertion into rendering the preferred pronunciation of Peddinti in the form of rhythm, stress, and intonation as described above. Tischer's specific prosodic information provides specific type of information to provide recognizable pronunciations of these words with clarity in a modification to Peddinti since specifics of prosody are added via the teachings of Tischer. Applicant's teaching away arguments seem to argue a modification of Tischer or Tischer teachins away from Peddinti that does not reflect the grounds of rejection since both references rely upon prosodic information to generate synthetic speech in a manner of a user/target speaker. Accordingly, due to the link between prosody and Peddinti's focus on particular or unique words, Tischer does aid Peddinti in increasing clarity of those words and these arguments are not found to be persuasive.
The prior art rejections of the remaining dependent claims have been traversed for reasons similar to the independent claims (see Remarks, Pages 16-20). In response, Applicant is directed to the preceding reply directed towards the independent claims.
Applicant provides some additional arguments directed towards the subject matter of claims 5, 12, and 21 by arguing that the rejection conflates non-verbal and non-vocal where the instant application "expressly distinguishes" between these expressions and the laugh or crying taught by Ingel, et al. (U.S. PG Publication: 2025/0006182 A1) is not a non-vocal sound as claimed.
In response, it is pointed out that the specification's "expressly" distinguishing argued by Applicant is not "clear and unmistakable" definition/disavowal of scope for these terms, but a description of embodiments (see MPEP 2111(IV)). Thus, non-vocal sounds excluding laughter for example is not disavowed given the description of embodiments argued by Applicant. In particular, Applicant is directed to consider the types of laughter detailed in Ingel as including manually produced sounds such as a "snicker" or a "snort" that are non-vocal sounds considered as laughs by Ingel (Paragraphs 0402). Thus, since the specification does not disavow all laughs from the scope of "non-vocal" and Ingel teaches non-vocal mimicked sounds including laughs, this argument has not been found to be persuasive.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 4, 8, 11, 15, 18, and 30 are rejected under 35 U.S.C. 103 as being unpatentable over Peddinti, et al. (U.S. PG Publication: 2022/0284882 A1) in view of Tischer (U.S. PG Publication: 2004/0111271 A1).
With respect to Claim 1, Peddinti discloses:
A system comprising:
a microphone (microphone for audio capture, Paragraphs 0026 and 0029);
an audio speaker (speaker for communicating an audible audio signal, Paragraphs 0026 and 0035);
an automatic speech recognition (ASR) sensor (ASR system/sensor relating acoustics to word confidence scores including a neural network and/or confidence models, Paragraphs 0029, 0049, 0057; Fig. 1, Element 140);
a text-to-speech (TTS) module (text-to-speech (TTS) system/software, Paragraphs 0027 and 0035); and
a speech-to-text (STT) module (speech-to-text software within the ASR system that uses confidence data to generate/assemble a transcription of the spoken user query, Paragraphs 0025, 0029, 0049, and 0057);
a computing platform having a hardware processor and a system memory storing a software code and a natural language understanding (NLU) model; the hardware processor configured to execute the software code to (computing device having a processor and memory that stores software code and an NLU model, Paragraphs 0021, 0029, and 0066-0067):
receive, via the microphone, an audio input, the audio input including speech by a human speaker (input speech data from a user is received at a digital assistant via a microphone, Paragraphs 0026 and 0029; Fig. 1, Element 12);
produce, using the ASR sensor and the STT module, a machine generated a text transcription of the audio input (Paragraph 0029- “generate, as output, a transcription 142 of the query 12”; utilizing the ASR system to generate the word confidence scores and STT software to generate the transcription, Paragraphs 0025, 0029, 0049, and 0057; note that the ASR system and STT software a machine based thus the transcription is “machine generated” as claimed);
identify, using the NLU model and the machine generated text transcription, an audio segment of interest (identification of (e.g., "the particular word" such as the name of a city referenced in a spoken user query) audio data in the spoken input using the NLU module/model, Paragraph 0021 and 0029-0031);
analyze one or more audio characteristics of the feature of the audio segment of interest including a prosody of the segment of interest (obtaining audio related features in the form of "pronunciation-related features" for the particular “unique” term, Paragraphs 0031, 0040 (prosody as a characteristic), and 0045);
generate, using the machine generated text transcription, a text string corresponding to the audio segment of interest (generating “a textual representation of the response to the query” with respect to the spoken user query audio data term of interest, Paragraphs 0029-0031);
convert the text string, using the TTS module, to produce an audio response having a characteristic voice of a non-human social agent that is independent of a voice of the human speaker, wherein a rendering of the audio segment of interest in the audio response is modified, using the determined prosody of the audio segment of interest, to replicate the pronunciation of the audio segment of interest by the human speaker, and wherein portions of the audio response other than the audio segment of interest retain the characteristic voice of the non-human social agent unmodified by the prosody of the audio segment of interest (use of the text-to-speech (TTS) system of the digital assistant relying upon a user pronunciation of the particular word in the query differs from that of the TTS pronunciation and modifies the response such that the digital assistant (i.e., non-human social agent) utters "the synthesized speech that approximates how a human would pronounce words formed by the sequence of graphemes/characters defining the TTS input 152 including the textual representation of the response to the query 12," Paragraphs 0027, 0031-0032 (i.e., the human speaker pronunciation is “replicated” because it is an approximation of the human pronunciation); see also paragraph 0027- "the TTS audio 154 pronounces the particular word using the one of the user pronunciation 202"; see also Paragraphs 0053-0054 describing that the user pronunciation characteristics are relied upon for the “particular word” among other digital assistant speech); and
play, using the audio speaker, the audio response having the characteristic voice of the non-human social agent while replicating the pronunciation of the audio segment of interest of the human speaker (audibly outputting the synthesized speech of the digital assistant via the speaker response to the user query relying upon the user pronunciation of the audio of interest, Paragraphs 0027, 0031-0032, and 0035).
While Peddinti teaches analyzing an audio segment of interest (e.g., “particular” or “unique” word with respect to "pronunciation-related features” such as prosody in generating a TTS output replicating a user's pronunciation by a digital assistant (see Paragraphs 0031-0032 and 0040), Peddinti does not teach that prosodic analysis involves determining stress, rhythm, and intonation of the audio segment of interest. Tischer, however, discloses that speech samples recorded of interest from a person's "own voice file" are analyzed and then used for "customizing...text to speech" with parameters including "emphasis" (i.e., stress), "rhythm", and "intonations" in the voice input samples of words (Paragraphs 0009, 0034-0035, 0037,0041, 0053, 0055, and 0061 (discussing the use of a person's "own voice file")).
Peddinti and Tischer are analogous art because they are from a similar field of endeavor in text-to-speech synthesis. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to utilize the prosodic features taught by Tischer in the word-specific pronunciations taught by Peddinti to provide a predictable result of providing synthesized voices that are recognizable and with greater clarity (Tischer, Paragraph 0023).
With respect to Claim 4, Peddinti further discloses:
The system of claim 1, wherein the audio segment of interest comprises a first name, a surname, a nickname, a name of a pet, a place name, a brand name, or a company name (contact name or musical artist first and last names, Paragraph 0022 and 0057; name of a place (city/restaurant examples), Paragraphs 0030 and 0046).
Claim 8 is a method embodiment carrying out the functionality of the system set forth in claim 1, and thus, is rejected under similar rationale.
Claim 11 contains subject matter similar to claim 4, and thus, are rejected under similar rationale.
Claim 15 is a "computer-readable non-transitory medium" embodiment storing processor instructions for carrying out the functionality of system claim 1, and thus, is rejected under similar rationale. Moreover, Peddinti further recites method implementation as a program stored on a non-transitory computer-readable medium (Paragraphs 0071-0072).
Claim 18 contains subject matter similar to claim 4, and thus, are rejected under similar rationale.
With respect to Claim 30, Peddinti discloses:
The system of claim 1, wherein the characteristic voice of the non- human social agent is a voice independent of any voice of the human speaker ("digital voice assistant" with a voice that is altered only with respect to "unique" or "particular" words such as names, Paragraphs 0024, 0042, 0053-0054, 0057; Fig. 1, Element 115).
Claims 5, 12, and 21 are rejected under 35 U.S.C. 103 as being unpatentable over Peddinti, et al. in view of Tischer and further in view of Ingel, et al. (U.S. PG Publication: 2025/0006182 A1).
With respect to Claim 5, Peddinti in view of Tischer teaches the system for outputting a response to a spoken user input relying upon a user pronunciation as applied to Claim 1. Although Peddinti models TTS speech pronunciations of particular terms after user pronunciations, the particular terms discussed in Peddinti in view of Tischer do not refer to "a non-vocal sound produced by the human speaker" as set forth in claim 5. Ingel, however, teaches a voice assistant voice clone that clones/mimics non-vocal sounds from a user such as laughter (Paragraphs 0061, 0174, 0405, and 0420).
Peddinti, Tischer, and Ingel are analogous art because they are from a similar field of endeavor in text-to-speech synthesis. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to include non-speech/verbal vocalizations taught by Ingel in the user pronunciations that are adapted into TTS taught by Peddinti in view of Tischer to provide a predictable result of producing a synthetic voice that better approximates a user by also including distinct non-verbal sounds like crying and laughter.
Claims 12 and 21 contain subject matter similar to claim 5, and thus, is rejected under similar rationale.
Claims 6, 13, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Peddinti, et al. in view of Tischer and further in view of Schaaf, et al. (U.S. Patent: 9,405,741 B1).
With respect to Claim 6, Peddinti in view of Tischer teaches the system for outputting a response to a spoken user input relying upon a user pronunciation as applied to Claim 1. Peddinti in view of Tischer does not teach the substitute language replacement procedure for prohibited words in a reply as set forth in claim 6. Schaaf, however, discloses:
a language database stored in the system memory, the language database including a list of prohibited words and a plurality of generic responses (database of terms that are prohibited/offensive and corresponding generic replacements (e.g., “Playing the requested song” instead of actually using the prohibited song title), Col. 3, Line 35- Col. 4, Line 17; Col. 9, Lines 7-14; Col. 15, Line 56- Col. 16, Line 30), wherein the hardware processor is further configured to execute the software code to:
determine whether the text string comprises a word in the list of prohibited words; select, based on the machine generated text transcription, a substitute response from among the plurality of generic responses; and replace the text string with the substitute response such that the audio response is produced using the substitute response in lieu of the audio segment of interest (spoken request is transcribed with natural language understanding and when a word that is on the prohibited/offensive list is encountered, a generic substitution in the audio output response is provided (e.g., “Playing the requested song” instead of actually using the prohibited song title in a similar response), Col. 11, Lines 1-25; Col. 15, Line 56- Col. 16, Line 30; note that Peddinti discloses a particular word such as a name that is otherwise produced in accordance with a user pronunciation as applied to claim 1).
Peddinti, Tischer, and Schaaf are analogous art because they are from a similar field of endeavor in text-to-speech synthesis in digital assistants. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to enable the digital assistant of Peddinti in view of Tischer to use the prohibited term substitution taught by Schaaf in response generation to provide a predictable result of improving the user experience by removing a output that the user may find to be inappropriate (Schaaf, Col. 2, Lines 26-34).
Claims 13 and 19 contain subject matter similar to claim 6, and thus, are rejected under similar rationale.
Claims 7, 14, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Peddinti, et al. in view of Tischer in view of Winter, et al. (U.S. PG Publication: 2019/0096387 A1; cited in the PTO-892 from 7/21/2025).
With respect to Claim 7, Peddinti in view of Tischer teaches the system for outputting a response to a spoken user input relying upon a user pronunciation as applied to Claim 1. Peddinti in view of Tischer does not teach the detection and removal of impediments in a spoken input of interest for a system response as set forth in claim 7. Winter, however, discloses:
The system of claim 1, wherein the audio segment of interest includes a speech impediment element (detection of user audio self-repairs/impediments such as a stutter or repetition in an utterance, Paragraphs 0034-0036), and wherein the hardware processor is further configured to execute the software code to: remove, before playing the audio response, the speech impediment from the audio response (performing natural language processing to "remove" the speech impediment including the stutter/repetition from the TTS output, Paragraphs 0030 and 0034-0035).
Peddinti, Tischer, and Winter are analogous art because they are from a similar field of endeavor in interactive systems utilizing text-to-speech synthesis. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to employ the filler word removal taught by Winter in the response generation taught by Peddinti in view of Tischer to provide for a more accurate read back of the audio of interest (Winter, Paragraph 0034).
Claims 14 and 20 contain subject matter similar to claim 7, and thus, are rejected under similar rationale.
Claims 23 and 26 are rejected under 35 U.S.C. 103 as being unpatentable over Peddinti, et al. in view of Tischer in view of Hantrakul, et al. (U.S. PG Publication: 2023/0377591 A1).
With respect to Claim 23, Peddinti in view of Tischer teaches the system for outputting a non-human digital assistant response to a spoken user input relying upon a user pronunciation as applied to Claim 1. Although Peddinti deals with computer systems and neural networks (see Paragraph 0035) that likely could provide a digital agent response to a received audio input in near real-time on the order of one hundred milliseconds or less, Peddinti in view of Tischer do not explicitly describe such near-real time processing. Hantrakul, however, discloses speech synthesis to generate a "real time" reply with a perception of an immediate response on the order of "within few milliseconds." Note that the ordinary and customary meaning of "few" is a "small number" (https://www.dictionary.com/browse/few). Thus, a small number would be within the range disclosed by the applicant and obvious to choose on the order of 50 ms or less in order to provide a response that is "virtually immediate when observed by a user" (Paragraph 0029).
Peddinti, Tischer, and Hantrakul are analogous art because they are from a similar field of endeavor in speech synthesis. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to provide the responses of Peddinti in view of Tischer in real time as taught by Hantrakul to provide a predictable result of providing a response that is "virtually immediate when observed by a user" (Paragraph 0029) and implement continuous interaction between the user and the digital assistant in Peddinti.
Claim 26 contains subject matter similar to claim 23, and thus, are rejected under similar rationale.
Claims 27- 29 are rejected under 35 U.S.C. 103 as being unpatentable over Peddinti, et al. in view of Tischer in view of Takano, et al. (U.S. PG Publication: 2020/0218781 A1).
With respect to Claim 27, Peddinti in view of Tischer teaches the system for outputting a non-human digital assistant response to a spoken user input relying upon a user pronunciation as applied to Claim 1. Peddinti in view of Tischer does not teach wherein the non-human social agent comprises a robot, and the hardware processor is further configured to execute the software code to control the one or more mechanical actuators to produce a facial expression or articulate a limb of the robot. Takano, however, discloses a non-human social agent that comprises a robot (Fig. 2, Element 120; Paragraph 0023) that has “mechanical actuators” in, e.g., limbs, hands, head, that are controlled to produce facial gestures/expressions or articulate a limb of the robot via hand gestures (Paragraphs 0044, 0050, 0054, and 0078-0079).
Peddinti, Tischer, and Takano are analogous art because they are from a similar field of endeavor systems utilizing speech synthesis. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to realize the digital assistant of Peddinti in view of Tischer in a mechanical robot as taught by Takano to provide a predictable result of a more-human like agent/assistant interaction including both voice and gestures.
With respect to Claim 28, Takano further discloses:
The system of claim 1, further comprising a camera and a facial recognition sensor, the facial recognition sensor being configured to recognize a face using image data captured by the camera (video input device such as camera, Paragraphs 0022 and 0148, acting as an input of video data images to a facial feature recognition sensor, Paragraphs 0037, 0044, and 0050, for a predictable result of better understanding an interacting user).
With respect to Claim 29, Takano further discloses:
The system of claim 1, wherein the non-human social agent comprises one of a robot, a digital character rendered on a display, or an agent exhibiting characteristics of a fictional character (a robot, Fig. 2, Element 120 and Paragraph 0023; see also rendering of a 3D humanoid model representing a virtual agent in paragraph 0041 corresponding to the claimed digital character on a display).
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Marzinzik, et al. (U.S. PG Publication: 2024/0281706)- teaches the use of a "scheme" for mimicking a user by a chatbot (Paragraphs 0052 and 0086).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JAMES S WOZNIAK whose telephone number is (571)272-7632. The examiner can normally be reached 7-3, off alternate Fridays.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant may use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Flanders can be reached at (571)272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
JAMES S. WOZNIAK
Primary Examiner
Art Unit 2655
/JAMES S WOZNIAK/Primary Examiner, Art Unit 2655