DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This Office Action is in response to the amendment filed April 17, 2026. Claims 1, 4, and 9 have been amended. Claims 1-19 remain pending.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 2-3 and 10-11 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 2 (and by dependency claim 3) recites “a situational attribute” in line 2. As claim 1 recites “a situational attribute” at line 4, it is unclear if the recitation at claim 2 refers to the same situational attribute previously recited, or if it refers to a new/different attribute.
Claim 10 (and by dependency claim 11) recites “a situational attribute” in lines 1-2. As claim 9 recites “a situational attribute” at line 3, it is unclear if the recitation at claim 10 refers to the same situational attribute previously recited, or if it refers to a new/different attribute
Claim Rejections - 35 USC § 103
The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
Claims 1-19 are rejected under 35 U.S.C. 103 as being unpatentable over Kirsch et al (US Patent Application Publication No. 2010/0057465), hereinafter Kirsch, in view of Flores et al (US Patent Application Publication No. 2017/0244834), hereinafter Flores.
Kirsch discloses variable text-to-speech for automotive application. Regarding claim 1, Kirsch teaches a method of speech synthesis [Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031], the method comprising: processing a sensor signal of a vehicle and determining, based on the sensor signal, a situational attribute [312/320; para 0025-0026-- one or more TTS tuning modules 322, 324 that classify vehicle states and relate them to human reactions and associated countermeasures that can be brought into effect through the adjustment of TTS speech synthesis parameters, the parameters can be updated or modified based on the one or more TTS tuning models 322, 324 that take into account the current state of the vehicle and its environment. The TTS control engine 202 then exports parameters from the TTS parameter module 326 to the TTS speech synthesizer 310 to cause the TTS audio stream played to the driver to be modified according to the identified countermeasures, where the system can provide tuning modules relating TTS voice speed and TTS voice volume to vehicle speed];
producing a TTS prosody parameter by applying a model that transforms the situational attribute into TTS prosody parameters [TTS parameters & TTS Tuning Model -- Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031]; synthesizing digital audio samples of speech, an attribute of which depends upon the TTS prosody parameter [TTS Speech Synthesizer/ tuning modules relating TTS voice speed and TTS voice volume to vehicle speed -- Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031]; and driving a speaker to produce audio as represented by the digital audio samples, wherein prosody parameter can change at run time for more dynamic effects [p0025 – TTS audio stream played to the driver based on current state of vehicle/environment-- Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031]. Kirsch fails to teach processing data relating to at least one of an attribute of a listener and a profile for the listener or that the prosody parameter is based on the processed data and wherein the TTS prosody parameters computed by the model comprise a machine learning algorithm trained using historical listener attribute and behavior data. In a similar field of endeavor, Flores [para 0022; 0039; 0044-0054] teaches virtual voice response agent individually configured for a user, in which the voice response agent 136 can be customized that result in TTS engine 122 providing synthesized speech having characteristics [from data gathered in the user profile and user interactions] that are tailored to the user 150 and the user's sentiments on the present call to be customized to present a speech personality matching the user's personality traits and speech patterns, and appropriate for the user's current sentiment [“listener’s responses”]; can provide processes to identify speech traits can be based on pattern recognition, including hidden Markov models, neural networks, pattern matching, frequency estimation, mixed models and deep learning [para 0040]; provides for a custom agent manager 124 can collect and analyze data pertaining to VIVR agent features 134 selected by various users 150-154. Based on such analysis, the custom agent manager 124 can learn how different users select different VIVR agent features 134. From time to time, the custom agent manager 124 can automatically update one or more baseline VIVR agent profiles 132 to implement VIVR agent features [para 0065] and specifically teaches the system’s user interaction analytics can identify virtual intelligent agent features shown to be more effective in satisfying users [para 0052]. Therefore, one having ordinary skill in the art at the time of the invention would have recognized the advantages of implementing the user interaction analytics to generate user tailored speech synthesis as suggested by Flores, in the system of Kirsch, and the results would have been predictable and would provide speech synthesis that is more effective in satisfying the user, as suggested by Flores and subsequently modify the vehicle sensor based synthesized speech output so as to ensure the speech is intelligible for the user and the user is able to ascertain the important content of the speech.
Regarding claim 2, the combination of Kirsch and Flores teaches processing the sensor signal determines a value of a situational attribute [sensor interface -- Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031; Flores’ user interaction analytics], and producing the TTS prosody parameter is in dependence upon the value of the situational attribute [TTS parameters & TTS Tuning Module -- Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031; Flores’ user interaction analytics].
Regarding claim 3, the combination of Kirsch and Flores teaches the dependence upon the value of the situational attribute is programmable using text rules [TTS parameters & TTS Tuning Modules -- Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031; Flores’ user interaction analytics].
Regarding claim 4, Kirsch teaches a method of speech synthesis [Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031], the method comprising: processing a sensor signal of a vehicle to determine a value of a situational attribute [312; 320; para 0025-0026-- one or more TTS tuning modules 322, 324 that classify vehicle states and relate them to human reactions and associated countermeasures that can be brought into effect through the adjustment of TTS speech synthesis parameters, the parameters can be updated or modified based on the one or more TTS tuning models 322, 324 that take into account the current state of the vehicle and its environment. The TTS control engine 202 then exports parameters from the TTS parameter module 326 to the TTS speech synthesizer 310 to cause the TTS audio stream played to the driver to be modified according to the identified countermeasures, where the system can provide tuning modules relating TTS voice speed and TTS voice volume to vehicle speed]; producing a TTS parameter according to a model in dependence upon a value of the situational attribute [TTS parameters & TTS Tuning Model -- Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031]; synthesizing digital audio samples of speech based on input text such that an attribute of the digital audio samples depends upon the TTS parameter [TTS Speech Synthesizer -- Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031]; and driving a speaker to produce audio as represented by the digital audio samples [p0025 – TTS audio stream played to the driver based on current state of vehicle/environment-- Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031]. Kirsch fails to teach processing data relating to at least one of an attribute of a listener and a profile for the listener to, in part, determine the value of the situational attribute, wherein the model utilizes two or more listener profile features selected from the group consisting of age, gender, emotional state and linguistic background or detecting a user behavior in response to the audio and updating the model based on the detected user behavior. In a similar field of endeavor, Flores [para 0022; 0039; 0044-0054] teaches virtual voice response agent individually configured for a user, in which the voice response agent 136 can be customized that result in TTS engine 122 providing synthesized speech having characteristics [from data gathered in the user profile and user interactions – including a language spoken by the user 150, a dialect spoken by the user 150, a particular accent of the user's speech, vocabulary, language and colloquialisms used by the user 150, the user's speech rate or speech tempo, a gender corresponding to the user's tone of voice, a sentiment of the user 150] that are tailored to the user 150 and the user's sentiments [“user behavior in response to the audio”] on the present call to be customized to present a speech personality matching the user's personality traits and speech patterns, and appropriate for the user's current sentiment [“listener’s responses”]; can provide processes to identify speech traits can be based on pattern recognition, including hidden Markov models, neural networks, pattern matching, frequency estimation, mixed models and deep learning [para 0040]; provides for a custom agent manager 124 can collect and analyze data pertaining to VIVR agent features 134 selected by various users 150-154. Based on such analysis, the custom agent manager 124 can learn how different users select different VIVR agent features 134. From time to time, the custom agent manager 124 can automatically update one or more baseline VIVR agent profiles 132 to implement VIVR agent features [para 0065] and specifically teaches the system’s user interaction analytics can identify virtual intelligent agent features shown to be more effective in satisfying users [para 0052]. Therefore, one having ordinary skill in the art at the time of the invention would have recognized the advantages of implementing the user interaction analytics to generate user tailored speech synthesis as suggested by Flores, in the TTS tuning module system of Kirsch, and the results would have been predictable and would provide speech synthesis that is more effective in satisfying the user, as suggested by Flores.
Regarding claim 5, the combination of Kirsch and Flores teaches the dependence upon the value of the situational attribute is programmable using text rules [TTS parameters & TTS Tuning Modules -- Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031].
Regarding claim 6, the combination of Kirsch and Flores teaches the TTS parameter represents a prosody attribute [TTS parameters & TTS Tuning Modules…pitch, speed, volume -- Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031; Flores’ user interaction analytics].
Regarding claim 7, the combination of Kirsch and Flores teaches prosody can be changed at run time for more dynamic effects [TTS speed and volume changes with changing vehicle speed -- Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031; Flores’ real time user interaction analytics].
Regarding claim 8, the combination of Kirsch and Flores teaches synthesizing the digital audio samples of speech such that prosody attribute is further based on markup in the input text [TTS parameters & TTS Tuning Modules -- Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031].
Regarding claim 9, Kirsch teaches a method of speech synthesis [Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031], the method comprising: processing a sensor signal of a vehicle [312; 320] and determining, based on the sensor signal, a situational attribute [para 0025-0026-- one or more TTS tuning modules 322, 324 that classify vehicle states and relate them to human reactions and associated countermeasures that can be brought into effect through the adjustment of TTS speech synthesis parameters, the parameters can be updated or modified based on the one or more TTS tuning models 322, 324 that take into account the current state of the vehicle and its environment. The TTS control engine 202 then exports parameters from the TTS parameter module 326 to the TTS speech synthesizer 310 to cause the TTS audio stream played to the driver to be modified according to the identified countermeasures, where the system can provide tuning modules relating TTS voice speed and TTS voice volume to vehicle speed]; producing a TTS parameter by applying a model that transforms at least the situational attribute into TTS parameters [TTS parameters & TTS Tuning Model -- Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026 -- one or more TTS tuning modules 322, 324 that classify vehicle states and relate them to human reactions and associated countermeasures that can be brought into effect through the adjustment of TTS speech synthesis parameters, the parameters can be updated or modified based on the one or more TTS tuning models 322, 324 that take into account the current state of the vehicle and its environment. The TTS control engine 202 then exports parameters from the TTS parameter module 326 to the TTS speech synthesizer 310 to cause the TTS audio stream played to the driver to be modified according to the identified countermeasures, where the system can provide tuning modules relating TTS voice speed and TTS voice volume to vehicle speed; 0027-0031]; synthesizing digital audio samples of speech, an attribute of which depends upon the TTS parameter [TTS Speech Synthesizer -- Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031]; and driving a speaker to produce audio as represented by the digital audio samples [p0025 – TTS audio stream played to the driver based on current state of vehicle/environment-- Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031]. Kirsch fails to teach processing data relating to at least one of an attribute of a listener and a profile for the listener or that the prosody parameter is based on the processed data, wherein the TTS prosody parameters computed by the model comprise a machine learning algorithm trained using historical listener attribute and behavior data. In a similar field of endeavor, Flores [para 0022; 0039; 0044-0054] teaches virtual voice response agent individually configured for a user, in which the voice response agent 136 can be customized that result in TTS engine 122 providing synthesized speech having characteristics that are tailored to the user 150 and the user's sentiments on the present call to be customized to present a speech personality matching the user's personality traits and speech patterns, and appropriate for the user's current sentiment [“listener’s responses”] and specifically teaches the system’s user interaction analytics can identify virtual intelligent agent features shown to be more effective in satisfying users [para 0052]. Therefore, one having ordinary skill in the art at the time of the invention would have recognized the advantages of implementing the user interaction analytics to generate user tailored speech synthesis as suggested by Flores, in the system of Kirsch, and the results would have been predictable and would provide speech synthesis that is more effective in satisfying the user, as suggested by Flores and subsequently modify the vehicle sensor based synthesized speech output so as to ensure the speech is intelligible for the user and the user is able to ascertain the important content of the speech.
Regarding claim 10, the combination of Kirsch and Flores teaches processing the sensor signal determines a value of a situational attribute [sensor interface -- Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031; Flores’s sentiment processing], and producing the TTS parameter is in dependence upon the value of the situational attribute [TTS parameters & TTS Tuning Module -- Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031; Flores’ sentiment processing].
Regarding claim 11, the combination of Kirsch and Flores teaches the dependence upon the value of the situational attribute is programmable using text rules [TTS parameters & TTS Tuning Modules -- Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031].
Regarding claim 12, the combination of Kirsch and Flores teaches the TTS parameter represents a prosody attribute [TTS parameters & TTS Tuning Modules…pitch, speed, volume -- Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031; Flores’ user interaction analytics].
Regarding claim 13, the combination of Kirsch and Flores teaches prosody can change at run time for more dynamic effects [TTS speed and volume changes with changing vehicle speed -- Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031; Flores’ real-time synthesis parameter updates].
Regarding claim 14, Kirsch teaches a text-to-speech system [Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031], comprising a computer processor programmed to: perform machine learned parametric speech synthesis using a TTS voice parameter [TTS speech synthesis based on TTS parameters & TTS Tuning Modules -- Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031] and produce the TTS voice parameter by a function that transforms at least one voice attribute and at least one situational attribute according to a model [TTS parameters & TTS Tuning Model -- Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031] wherein the model has a specified rule-based algorithm coded with a user text rule [TTS parameters & TTS Tuning Modules for synthesis based on driver preferences-- Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031]. Kirsch fails to specifically teach measuring how listeners respond to the TTS voice parameter and improving the machine learned parametric speech synthesis based on how listeners respond to the TTS voice parameter. In a similar field of endeavor, Flores [para 0022; 0039; 0044-0054] teaches virtual voice response agent individually configured for a user, in which the voice response agent 136 can be customized that result in TTS engine 122 providing synthesized speech having characteristics that are tailored to the user 150 and the user's sentiments on the present call to be customized to present a speech personality matching the user's personality traits and speech patterns, and appropriate for the user's current sentiment [“listener’s responses”] and specifically teaches the system’s user interaction analytics can identify virtual intelligent agent features shown to be more effective in satisfying users [para 0052]. Therefore, one having ordinary skill in the art at the time of the invention would have recognized the advantages of implementing the user interaction analytics to generate user tailored speech synthesis as suggested by Flores, in the system of Kirsch, and the results would have been predictable and would provide speech synthesis that is more effective in satisfying the user, as suggested by Flores.
Regarding claim 15, the combination of Kirsch and Flores teaches the user text rule depends on situational attributes [TTS parameters & TTS Tuning Modules for synthesis based on driver preferences-- Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031; Flores’ sentiment processing].
Regarding claim 16, the combination of Kirsch and Flores teaches a situational attribute is noise level [Interior noise sensor (208); Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031].
Regarding claim 17, the combination of Kirsch and Flores teaches the user text rule depends on noise level [Interior noise sensor (208); Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031].
Regarding claim 18, the combination of Kirsch and Flores teaches the TTS voice parameter is related to volume [TTS parameters & TTS Tuning Modules…pitch, speed, volume -- Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031].
Regarding claim 19, the combination of Kirsch and Flores teaches the at least one situational attribute is a noise level [Interior noise sensor (208); Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031; Flores’ sentiment processing]; the TTS voice parameter is related to volume [TTS parameters & TTS Tuning Modules…pitch, speed, volume -- Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031]; the user text rule depends on a value of the at least one situational attribute [TTS parameters & TTS Tuning Modules for synthesis based on driver preferences-- Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031]; and the value of the at least one situational attribute is received from a sensor in a vehicle [sensor interface -- Fig 2; Fig 3; Fig 7; Fig 8; para 0009-0012; 0022-0026; 0027-0031].
Response to Arguments
Applicant's arguments filed April 17, 2026 have been fully considered but they are not persuasive.
At pages 6-7, with respect to claims 1 and 9, Applicant argues Kirsch does not disclose determining a situational attribute as an intermediate abstraction and then applying a model that transforms that attribute together with listener-related data to produce a TTS parameter as provided by the amended claims. Applicant further argues Flores operates in a different context and does not address vehicle sensor signals or the derivation of situational attributes from such signals and also argues nor does Flores disclose a model that jointly processes situational attributes and listener-related data as required by the amended claims. The examiner notes that these arguments are toward the newly presented amended claim language. These additional elements have been mapped to Kirsch/Flores references as indicated in the rejection above.
Applicant argues neither reference teaches or suggests a model that receives multiple distinct attribute types and performs a transformation to generate TTS parameters. In response to applicant's arguments against the references individually, one cannot show nonobviousness by attacking references individually where the rejections are based on combinations of references. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981); In re Merck & Co., 800 F.2d 1091, 231 USPQ 375 (Fed. Cir. 1986).
Applicant argues the combination proposed by the Examiner therefore requires impermissible hindsight reconstruction, as it selectively extracts unrelated aspects of each reference without any teaching or suggestion that such features should be combined in the claimed manner. In response to applicant's argument that the examiner's conclusion of obviousness is based upon improper hindsight reasoning, it must be recognized that any judgment on obviousness is in a sense necessarily a reconstruction based upon hindsight reasoning. But so long as it takes into account only knowledge which was within the level of ordinary skill at the time the claimed invention was made, and does not include knowledge gleaned only from the applicant's disclosure, such a reconstruction is proper. See In re McLaughlin, 443 F.2d 1392, 170 USPQ 209 (CCPA 1971).
At pages 7-8, with respect to claim 4, Applicant argues Kirsch is directed to adjusting speech output based on vehicle conditions, but does not monitor user behavior in response to the synthesized speech, nor does it update any model based on such behavior; Flores discusses analyzing user interactions, but does not disclose updating a model used to generate TTS parameters based on detected user behavior in response to synthesized audio, particularly in the context of vehicle-based situational attributes; and further argues neither Kirsch nor Flores discloses or suggests such a closed-loop system for updating a model used to generate TTS parameters. The examiner notes that these arguments are toward the newly presented amended claim language. These additional elements have been mapped to Kirsch/Flores references as indicated in the rejection above. Additionally, in response to applicant's arguments against the references individually, one cannot show nonobviousness by attacking references individually where the rejections are based on combinations of references. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981); In re Merck & Co., 800 F.2d 1091, 231 USPQ 375 (Fed. Cir. 1986).
Applicant argues the combination of Kirsch and Flores also lacks any teaching, suggestion, or motivation to provide a system in which a model that generates TTS parameters is updated based on user behavior resulting from the synthesized speech. In response to applicant’s argument that there is no teaching, suggestion, or motivation to combine the references, the examiner recognizes that obviousness may be established by combining or modifying the teachings of the prior art to produce the claimed invention where there is some teaching, suggestion, or motivation to do so found either in the references themselves or in the knowledge generally available to one of ordinary skill in the art. See In re Fine, 837 F.2d 1071, 5 USPQ2d 1596 (Fed. Cir. 1988), In re Jones, 958 F.2d 347, 21 USPQ2d 1941 (Fed. Cir. 1992), and KSR International Co. v. Teleflex, Inc., 550 U.S. 398, 82 USPQ2d 1385 (2007). In this case, Flores specifically teaches the system’s user interaction analytics can identify virtual intelligent agent features shown to be more effective in satisfying users [para 0052].
Applicant argues implementing such a feedback loop would fundamentally change the operation of Kirsch's system from a static or rule-based adjustment mechanism to a dynamic, learning system. The Examiner respectfully disagrees. The Examiner notes, the system of Kirsch operates as a dynamic system as disclosed, since the TTS parameters can be updated or modified based on the one or more TTS tuning models that take into account the current state of the vehicle and its environment. The TTS control engine then exports parameters from the TTS parameter module to the TTS speech synthesizer to cause the TTS audio stream played to the driver to be modified according to the identified countermeasures, where the system can provide tuning modules relating TTS voice speed and TTS voice volume to vehicle speed.
At page 9, with respect to claim 14, Applicant argues Kirsch does not disclose or suggest using voice attributes as an input to a function that generates TTS parameters. The Examiner respectfully disagrees. Kirsch’s TTS control system adjusts TTS synthesized voice parameters based on measurements of the state of vehicle to increase the level of intelligibility. TTS audio parameters that may be adjusted include the TTS voice volume, the TTS voice speed, the TTS voice pitch, and the speakers to which the TTS voice is directed. Additionally, other characteristics of a TTS voice, such as the gender of the voice, the language, or a particular regional accent may also be adjusted according to sensor inputs, such as a microphone that samples the driver's voice [para 0010].
Applicant argues neither Kirsch nor Flores discloses or suggests a model defined or controlled by a user text rule. The Examiner notes, applying the parameters to implement the TTS system necessarily requires applying text rules, and since Kirsch specifically teaches adjusting the parameters based on the driver’s voice or driver preferences [para 0010], the TTS of Kirsch provides a form of user text rules.
Applicant argues Flores does not disclose updating a model that generates TTS parameters based on listener response to synthesized speech. The Examiner respectfully disagrees. Flores specifically teaches [para 0022; 0039; 0044-0054] a virtual voice response agent individually configured for a user, in which the voice response agent 136 can be customized that result in TTS engine 122 providing synthesized speech having characteristics that are tailored to the user 150 and the user's sentiments on the present call to be customized to present a speech personality matching the user's personality traits and speech patterns, and appropriate for the user's current sentiment, where a customized speech personality that matches a user’s current personality/sentiment requires monitoring the user’s sentiment and updating the speech personality so that it matches or reflects the user’s current personality.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ANGELA A ARMSTRONG whose telephone number is (571)272-7598. The examiner can normally be reached M,T,TH,F 11:30-8:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre Desir can be reached at 571-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
ANGELA A. ARMSTRONG
Primary Examiner
Art Unit 2659
/ANGELA A ARMSTRONG/Primary Examiner, Art Unit 2659