Prosecution Insights
Last updated: October 01, 2026
Application No. 18/040,142

HYBRID TEXT TO SPEECH

Final Rejection §103
Filed
Jan 31, 2023
Priority
Apr 26, 2021 — nonprovisional of PCTCN2021089825
Examiner
CAUDLE, PENNY LOUISE
Art Unit
2657
Tech Center
2600 — Communications
Assignee
Microsoft Technology Licensing, LLC
OA Round
2 (Final)
70%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
85%
With Interview

Examiner Intelligence

Grants 70% — above average
70%
Career Allowance Rate
59 granted / 84 resolved
+8.2% vs TC avg
Moderate +14% lift
Without
With
+14.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
18 currently pending
Career history
97
Total Applications
across all art units

Statute-Specific Performance

§101
21.3%
-18.7% vs TC avg
§103
47.1%
+7.1% vs TC avg
§102
15.0%
-25.0% vs TC avg
§112
15.9%
-24.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 84 resolved cases

Office Action

§103
DETAILED ACTION This examination is in response to the communication filed on 07/17/2026. Claims 16-35 are currently pending, where claims 16, 24, 25, 33 and 34 have been amended. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) submitted on 06/15/2026 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Response to Amendment/Arguments Applicant’s arguments with respect to claims 16-35 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or non-obviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 16-35 are rejected under 35 U.S.C. 103 as being unpatentable over Osowski et al. (US 2014/0200894 A1; herein “Osowski”) further in view of Bakis et al. (US 2003/0046077 A1; herein “Bakis”) still further in view of Peng et al. (US 2021/0118440 A1; herein “Peng”). Regarding claim 16, Osowski teaches a device (Fig. 2, TTS device 202), comprising: a processor (Fig. 2, controller/processor 208); and a memory (Fig. 2, memory 210) comprising a cache (Fig. 2, TTS Storage and Fig. 1 Local TTS Unit Database) and computer program code, the memory and the computer program code configured to, with the processor, cause the processor to: receive textual data from a user application (Fig. 5 Step 504 “Receive text for TTS processing” ); send the received textual data to both a remote text to speech (TTS) engine and to a TTS engine in the device (Fig. 5, step 508 “Perform speech synthesis with units in local database” and step 510 “Obtain units from remote TTS device”. In addition, ¶[0040] teaches “local TTS processing may also be combined with distributed TTS processing. Where a portion of text to be converted uses units available in a local database, that portion of text may be processed locally. Where a portion of text to be converted uses units not available in a local database, the local device may obtain the units from a remote device. The units from the remote device may then concatenated with the local units for construction of the audio speech for output to a user” ); receive speech data from both the remote TTS engine and the TTS engine in the device, the received speech data corresponding to the received textual data (¶[0040] teaches “Where a portion of text to be converted uses units available in a local database, that portion of text may be processed locally. Where a portion of text to be converted uses units not available in a local database, the local device may obtain the units from a remote device. The units from the remote device may then concatenated with the local units for construction of the audio speech for output to a user” ); select, based on a selection policy, the received speech data from the remote TTS engine, the TTS engine in the device, or both (¶[0040] teaches “Where a portion of text to be converted uses units available in a local database, that portion of text may be processed locally. Where a portion of text to be converted uses units not available in a local database, the local device may obtain the units from a remote device. The units from the remote device may then concatenated with the local units for construction of the audio speech for output to a user” and ¶[0041] teaches “selection of units from input text may be performed by a remote device where the remote device is aware of what units are available on a local device. The remote device may determine the desired units to use in synthesizing the text and send the local device the unit sequence, along with the unit speech segments that are unavailable on the local device. The local device may then take the unit speech segments sent to it by the remote device, along with the unit speech segments that are available locally, and perform unit concatenation and complete the speech synthesis based on those unit segments and the unit sequence sent from the remote device”); and transmit the selected speech data to the user application (¶[0040] teaches “…The units from the remote device may then concatenated with the local units for construction of the audio speech for output to a user”). Osowski fails to disclose the device is configured to determine that the received textual data is missing from the cache. Bakis teaches a method and system for text-to-speech (TTS) caching that includes inter alia,“ receiving a text input and comparing the received text input to at least one entry in a text-to-speech cache memory…In the text input matches one of the entries in the text-to-speech cache memory, the cached speech output specified by the matching entry can be provided” More specifically, Bakis teaches determining that the received textual data is missing from the cache (Fig. 2, steps 220 and 230 and ¶[025] teaches “In step 230, a determination can be made as to whether a match for the received text input was located within the TTS cache…If not, the method can continue to step 250.”) Osowski differs from the claimed invention, as defined in claims 1, 25 and 34, in that Osowski fails to disclose or suggest TTS caching, e.g., determining whether or not a TTS output has been previously cached. TTS caching is known in the art as evidenced by Bakis. Therefore, it would have been obvious to one having ordinary skill in the art to have modified the TTS system of Osowski to further include TTS caching in order to increase TTS efficiency but reducing redundant processing. (Bakis, ¶[0005]). The combination of Osowski and Bakis fails to disclose sending all of the received textual data to both the remote text to speck (TTS) engine and to a TTS engine in the device as recited in amended claim 16. Peng teaches an assistant system which includes both a local TTS engine (Fig. 2, TTS 240) and a remote TTS engine (FIG. 2, TTS 238). In addition, Peng further teaches sending all of the received textual data to both the remote text to speech (TTS) engine and to a TTS engine (¶[0063] teaches “…the output of the response execution module 232 on the server-side may be sent to a remote text-to-speech (TTS) module 238. Similarly, the output of the response expansion module 236 on the client-side may be sent to a local TTS module 240. Both TTS modules may convert a response to audio signal. In particular embodiments, the output from …the TTS modules on both sides, may be finally sent to a local renders output module 242”). The combination of Osowski and Bakis differs from the claimed invention, as defined in claim 16, in that the combination fails to explicitly disclose sending the all or same received text data to both a remote and local TTS engine. Utilizing both a remote and local TTS engine for text to speech conversion is known in the art as evidenced by Peng. Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the invention to have modified the system taught by the combination of Osowski and Bakis to include sending the all of or the same received text to both the local and remote TTS engines as taught by Peng as it merely constitutes the combination of known processes to achieve the predictable result of allowing the system to generate a response suitable for the client system based on both the local and remote TTS engine outputs (Peng, ¶[0063]). Regarding claim 25, Osowski teaches a computer-implemented method comprising: receiving textual data from a user application (Fig. 5 Step 504 “Receive text for TTS processing” ); sending received textual data to both a remote text to speech (TTS) engine and to a TTS engine in the device (Fig. 5, step 508 “Perform speech synthesis with units in local database” and step 510 “Obtain units from remote TTS device”. In addition, ¶[0040] teaches “local TTS processing may also be combined with distributed TTS processing. Where a portion of text to be converted uses units available in a local database, that portion of text may be processed locally. Where a portion of text to be converted uses units not available in a local database, the local device may obtain the units from a remote device. The units from the remote device may then concatenated with the local units for construction of the audio speech for output to a user” ); receiving speech data from both the remote TTS engine and the TTS engine in the device, the received speech data corresponding to the received textual data (¶[0040] teaches “Where a portion of text to be converted uses units available in a local database, that portion of text may be processed locally. Where a portion of text to be converted uses units not available in a local database, the local device may obtain the units from a remote device. The units from the remote device may then concatenated with the local units for construction of the audio speech for output to a user” ); selecting, based on a selection policy, the received speech data from the remote TTS engine, the TTS engine in the device, or both (¶[0040] teaches “Where a portion of text to be converted uses units available in a local database, that portion of text may be processed locally. Where a portion of text to be converted uses units not available in a local database, the local device may obtain the units from a remote device. The units from the remote device may then concatenated with the local units for construction of the audio speech for output to a user” and ¶[0041] teaches “selection of units from input text may be performed by a remote device where the remote device is aware of what units are available on a local device. The remote device may determine the desired units to use in synthesizing the text and send the local device the unit sequence, along with the unit speech segments that are unavailable on the local device. The local device may then take the unit speech segments sent to it by the remote device, along with the unit speech segments that are available locally, and perform unit concatenation and complete the speech synthesis based on those unit segments and the unit sequence sent from the remote device”); and transmitting the selected speech data to the user application (¶[0040] teaches “…The units from the remote device may then concatenated with the local units for construction of the audio speech for output to a user”). Osowski fails to disclose determining that the received textual data is missing from the cache. Bakis teaches a method and system for text-to-speech (TTS) caching that includes inter alia,“ receiving a text input and comparing the received text input to at least one entry in a text-to-speech cache memory…In the text input matches one of the entries in the text-to-speech cache memory, the cached speech output specified by the matching entry can be provided” More specifically, Bakis teaches determining that the received textual data is missing from the cache (Fig. 2, steps 220 and 230 and ¶[025] teaches “In step 230, a determination can be made as to whether a match for the received text input was located within the TTS cache…If not, the method can continue to step 250.”) Osowski differs from the claimed invention, as defined in claim 25, in that Osowski fails to disclose or suggest TTS caching, e.g., determining whether or not a TTS output has been previously cached. TTS caching is known in the art as evidenced by Bakis. Therefore, it would have been obvious to one having ordinary skill in the art to have modified the TTS system of Osowski to further include TTS caching in order to increase TTS efficiency but reducing redundant processing. (Bakis, ¶[0005]). The combination of Osowski and Bakis fails to disclose sending the same received textual data to both the remote text to speech (TTS) engine and to a TTS engine in the device as recited in amended claim 25. Peng teaches an assistant system which includes both a local TTS engine (Fig. 2, TTS 240) and a remote TTS engine (FIG. 2, TTS 238). In addition, Peng further teaches sending the same received textual data to both the remote text to speech (TTS) engine and to a TTS engine (¶[0063] teaches “…the output of the response execution module 232 on the server-sed may be sent to a remote text-to-speech (TTS) module 238. Similarly, the output of the response expansion module 236 on the client-side may be sent to a local TTS module 240. Both TTS modules may convert a response to audio signal. In particular embodiments, the output from …the TTS modules on both sides, may be finally sent to a local renders output module 242”). The combination of Osowski and Bakis differs from the claimed invention, as defined in claim 25, in that the combination fails to explicitly disclose sending all or the same received text data to both a remote and local TTS engine. Utilizing both a remote and local TTS engine for text to speech conversion is known in the art as evidenced by Peng. Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the invention to have modified the system taught by the combination of Osowski and Bakis to include sending the all of or the same received text to both the local and remote TTS engines as taught by Peng as it merely constitutes the combination of known processes to achieve the predictable result of allowing the system to generate a response suitable for the client system based on both the local and remote TTS engine outputs (Peng, ¶[0063]). Regarding claim 34, Osowski teaches a computer storage medium comprising a plurality of instructions, when executed by a processor (Fig. 2, controller/processor 208), cause the processor to: receive textual data from a user application (Fig. 5 Step 504 “Receive text for TTS processing” ); send the received textual data to both a remote text to speech (TTS) engine and to a TTS engine in the device (Fig. 5, step 508 “Perform speech synthesis with units in local database” and step 510 “Obtain units from remote TTS device”. In addition, ¶[0040] teaches “local TTS processing may also be combined with distributed TTS processing. Where a portion of text to be converted uses units available in a local database, that portion of text may be processed locally. Where a portion of text to be converted uses units not available in a local database, the local device may obtain the units from a remote device. The units from the remote device may then concatenated with the local units for construction of the audio speech for output to a user” ); receive speech data from both the remote TTS engine and the TTS engine in the device, the received speech data corresponding to the received textual data (¶[0040] teaches “Where a portion of text to be converted uses units available in a local database, that portion of text may be processed locally. Where a portion of text to be converted uses units not available in a local database, the local device may obtain the units from a remote device. The units from the remote device may then concatenated with the local units for construction of the audio speech for output to a user” ); select, based on a selection policy, the received speech data from the remote TTS engine, the TTS engine in the device, or both (¶[0040] teaches “Where a portion of text to be converted uses units available in a local database, that portion of text may be processed locally. Where a portion of text to be converted uses units not available in a local database, the local device may obtain the units from a remote device. The units from the remote device may then concatenated with the local units for construction of the audio speech for output to a user” and ¶[0041] teaches “selection of units from input text may be performed by a remote device where the remote device is aware of what units are available on a local device. The remote device may determine the desired units to use in synthesizing the text and send the local device the unit sequence, along with the unit speech segments that are unavailable on the local device. The local device may then take the unit speech segments sent to it by the remote device, along with the unit speech segments that are available locally, and perform unit concatenation and complete the speech synthesis based on those unit segments and the unit sequence sent from the remote device”); and transmit the selected speech data to the user application (¶[0040] teaches “…The units from the remote device may then concatenated with the local units for construction of the audio speech for output to a user”). Osowski fails to disclose the device is configured to determine that the received textual data is missing from the cache. Bakis teaches a method and system for text-to-speech (TTS) caching that includes inter alia,“ receiving a text input and comparing the received text input to at least one entry in a text-to-speech cache memory…In the text input matches one of the entries in the text-to-speech cache memory, the cached speech output specified by the matching entry can be provided” More specifically, Bakis teaches determining that the received textual data is missing from the cache (Fig. 2, steps 220 and 230 and ¶[025] teaches “In step 230, a determination can be made as to whether a match for the received text input was located within the TTS cache…If not, the method can continue to step 250.”) Osowski differs from the claimed invention, as defined in claims 1, 25 and 34, in that Osowski fails to disclose or suggest TTS caching, e.g., determining whether or not a TTS output has been previously cached. TTS caching is known in the art as evidenced by Bakis. Therefore, it would have been obvious to one having ordinary skill in the art to have modified the TTS system of Osowski to further include TTS caching in order to increase TTS efficiency but reducing redundant processing. (Bakis, ¶[0005]). The combination of Osowski and Bakis fails to disclose at least a portion of the received textual data sent to the remote TTS engine being equivalent to a least a portion of the received textual data sent to the TTS engine in the device as recited in amended claim 34. Peng teaches an assistant system which includes both a local TTS engine (Fig. 2, TTS 240) and a remote TTS engine (FIG. 2, TTS 238). In addition, Peng further teaches at least a portion of the received textual data sent to the remote TTS engine being equivalent to a least a portion of the received textual data sent to the TTS engine in the device (¶[0063] teaches “…the output of the response execution module 232 on the server-sed may be sent to a remote text-to-speech (TTS) module 238. Similarly, the output of the response expansion module 236 on the client-side may be sent to a local TTS module 240. Both TTS modules may convert a response to audio signal. In particular embodiments, the output from …the TTS modules on both sides, may be finally sent to a local renders output module 242”). The combination of Osowski and Bakis differs from the claimed invention, as defined in claim 34, in that the combination fails to explicitly disclose sending the same portion of the received text data to both a remote and local TTS engine. Utilizing both a remote and local TTS engine for text to speech conversion is known in the art as evidenced by Peng. Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the invention to have modified the system taught by the combination of Osowski and Bakis to include sending the all of or the same received text to both the local and remote TTS engines as taught by Peng as it merely constitutes the combination of known processes to achieve the predictable result of allowing the system to generate a response suitable for the client system based on both the local and remote TTS engine outputs (Peng, ¶[0063]). Regarding claims 17 and 26, the combination of Osowski, Bakis and Peng teaches all of the elements of claims 16 and 25 (see detailed element mapping above). In addition, Osowski further teaches the selection policy includes rules to prioritize at least one of a cognition-driven policy, a performance-driven policy, or a quality- driven policy (¶[0043] teaches “The local device may be configured to adjust the local unit database based on a variety of factors such as available local storage, network connection quality, bandwidth used in communicating with the network, frequency of TTS requests, desired TTS quality or domain of use…”; ¶[0044] teaches “the local unit database may also be configured based on a user of the device…to provide faster TTS output…to include frequently used speech units”; and ¶[0045] teaches “the local device may cache a certain local unit database configured for a particular application…” and ¶[0047] teaches “…the size of a unit database is variable, and ultimately may be a result of a design choice between resources (i.e., storage or bandwidth) consumption and desired TTS quality. The choice may be configured by a device manufacturer, application designer, or user”. Furthermore, ¶[0051] teaches “ In certain aspects, it may be desirable to have the local device be aware of which speech synthesis units it can handle to provide a sufficient quality result (such as those units that are locally stored with a number of robust examples), and which units are better handled by a remote TTS device (such as those units that are locally stored with a limited number of examples). For example, if a local unit database includes at least one example of each speech synthesis unit, the local device may not be aware when it should use its locally stored units for synthesis and when it should turn to the remote device. In this aspect, the local device may also be configured with a list of units and their corresponding acoustic features that are available at a remote TTS device and whose audio files should be retrieved from the remote device for speech synthesis.” The teachings of ¶[0051] in combination with the teachings in ¶[0046]-[0047] would lead one skilled in the art to configured the system to determine whether to utilize the local database or the remoted database based on resource consumption, i.e., performance, desired quality, and/or user preference). Regarding claims 18 and 27, the combination of Osowski, Bakis and Peng teaches all of the elements of claims 16 and 25 (see detailed element mapping above). In addition, Osowski further teaches the selection policy is a reactive selection policy (the “or” language makes this limitation optional ) or a proactive selection policy (¶[0011] teaches the selection process is proactive in as much as either the local device or the remote device determines whether the local database has sufficient information to complete the speech process, thereby making the selection of local or remote data proactive). Regarding claims 19 and 28, the combination of Osowski, Bakis and Peng teaches all of the elements of claims 16 and 25 (see detailed element mapping above). In addition, Osowski further teaches the processor is further configured to: select, based on the selection policy, the speech data generated from both the remote TTS engine and the TTS engine in the device (¶[0011] teaches “…When the text is received by the local device 104 for TTS processing, the local device checks to see if the local unit database 108 include the units to perform speech synthesis for the input text If the units are available locally, the local device 104 performs TTS processing, as shown in block 112. If the units are not available locally, the local device contacts the remote device 116 over the network 114.The remote device 116 then sends the units to the local device 104 which completes the speech synthesis using the units obtained from the remote device 116. In another aspect the text may be sent directly to the remote device 116 which may then perform the check 110 to see if the local unit database 108 includes the units to perform speech synthesis for the input text. The remote device 116 may then send the local device 104 the unit sequence from the text along with any units not stored in the local TTS unit database 108 and instructions to concatenate the locally available units with the received units to complete the speech synthesis”); combine the selected speech data into comprehensive speech data (¶[0011] teaches “…In another aspect the text may be sent directly to the remote device 116 which may then perform the check 110 to see if the local unit database 108 includes the units to perform speech synthesis for the input text. The remote device 116 may then send the local device 104 the unit sequence from the text along with any units not stored in the local TTS unit database 108 and instructions to concatenate the locally available units with the received units to complete the speech synthesis”), wherein the comprehensive speech data includes at least a portion of the speech data generated from the remote TTS engine and at least a portion of the speech data generated from the TTS engine in the device (¶[0011] teaches “…and instructions to concatenate the locally available units with the received units to complete the speech synthesis ); and transmit the comprehensive speech data (¶[0034] teaches “Audio waveforms including the speech output from the TTS module 214 may be sent to an audio output device 204 for playback to a user or may be sent to the input/output device 206 for transmission to another device, such as another TTS device 202, for further processing or output to a user.”). Regarding claims 20 and 29, the combination of Osowski, Bakis and Peng teaches all of the elements of claims 16 and 25 (see detailed element mapping above). In addition, Osowski further teaches the processor is further configured to: determine to send the received textual data to the remote TTS engine and the TTS engine in the device based on a transmission policy and wherein the transmission policy is based at least in part on the selection policy (The broadest reasonable interpretation of this limitation is that selection of the which data to use and such the need to transmit the text information to both the local and remote devices includes selection based on latency, bandwidth, and processing load, such an interpretation is supported by ¶[0053] of the application as filed; ¶[0043] teaches “The local device may be configured to adjust the local unit database based on a variety of factors such as available local storage, network connection quality, bandwidth used in communicating with the network, frequency of TTS requests, desired TTS quality or domain of use…”)¶[0043] teaches “The local device may be configured to adjust the local unit database based on a variety of factors such as available local storage, network connection quality, bandwidth used in communicating with the network, frequency of TTS requests, desired TTS quality or domain of use…”). Regarding claims 21 and 30, the combination of Osowski, Bakis and Peng teaches all of the elements of claims 20 and 29 (see detailed element mapping above). In addition, Osowski further teaches the remote TTS engine is a TTS engine executed and stored in a cloud (Fig. 1, TTS device 116 and ¶[0017] teaches that “the TTS device 202 may connect to a network, such as the Internet or private network, which may include a distributed computing environment”). Regarding claims 22 and 31, the combination of Osowski, Bakis and Peng teaches all of the elements of claims 16 and 25 (see detailed element mapping above). In addition, Osowski further teaches the selected speech data is an audio version of the received textual data (¶[0018] teaches “includes a TTS module 214 for processing textual data into audio waveforms including speech…The textual data may originate from an internal component of the TTS device 202 or may be received by the TTS device 202 from an input device such as a keyboard or may be sent to the TTS device 202 over a network connection” Thus, the output corresponding to the selected speech data is an audio version of the received textual data). Regarding claims 23 and 32, the combination of Osowski, Bakis and Peng teaches all of the elements of claims 16 and 25 (see detailed element mapping above). In addition, Bakis further teaches to determine whether the received textual data is stored in the cache, the processor is further configured to identify whether the received textual data matches a keyword stored in the cache (¶[0025] teaches “In Step 220, the received text can be compared to the entries within the TTS cache to determine whether a match exists… in addition to comparing text, whether plain text or annotated text, the attributes of received text inputs can be compared to the attributes of the TTS cache entries.” Comparing plain or annotated text is interpreted as matching keywords, i.e., text). Osowski differs from the claimed invention, as defined in claims 1, 25 and 34, in that Osowski fails to disclose or suggest TTS caching, e.g., determining whether or not a TTS output has been previously cached. TTS caching is known in the art as evidenced by Bakis. Therefore, it would have been obvious to one having ordinary skill in the art to have modified the TTS system of Osowski to further include TTS caching in order to increase TTS efficiency but reducing redundant processing. (Bakis, ¶[0005]) Regarding claims 24 and 33, the combination of Osowski, Bakis and Peng teaches all of the elements of claims 16 and 25 (see detailed element mapping above). In addition, Bakis further teaches the processor is further configured to, in response to identifying the received textual data is stored in the cache: identify corresponding speech data to the received textual data identified in the cache (¶[0026] teaches “In step 240, when a match exists within the TTS cache, the TTS System can retrieve the spoken output specified by the matched entry…”); and bypass the remote TTS engine and the TTS engine in the device and transmit the corresponding speech data to the user application (¶[0026] teaches “In step 240, when a match exists…rather than continuing to process the received text input to construct a spoken output, the Spoken output Specified by the matched TTS cache entry can be provided as an output.”). Osowski differs from the claimed invention, as defined in claims 1, 25 and 34, in that Osowski fails to disclose or suggest TTS caching, e.g., determining whether or not a TTS output has been previously cached. TTS caching is known in the art as evidenced by Bakis. Therefore, it would have been obvious to one having ordinary skill in the art to have modified the TTS system of Osowski to further include TTS caching in order to increase TTS efficiency but reducing redundant processing. (Bakis, ¶[0005]) Regarding claim 35, the combination of Osowski, Bakis and Peng teaches all of the elements of claim 34 (see detailed element mapping above). In addition, XXX further teaches the selection policy includes rules to prioritize at least one of a cognition-driven policy, a performance-driven policy, or a quality-driven policy, the rules are selected by a user (¶[0043] teaches “The local device may be configured to adjust the local unit database based on a variety of factors such as available local storage, network connection quality, bandwidth used in communicating with the network, frequency of TTS requests, desired TTS quality or domain of use…”; ¶[0044] teaches “the local unit database may also be configured based on a user of the device…to provide faster TTS output…to include frequently used speech units”; ¶[0045] teaches “the local device may cache a certain local unit database configured for a particular application…” and ¶[0047] teaches “…the size of a unit database is variable, and ultimately may be a result of a design choice between resources (i.e., storage or bandwidth) consumption and desired TTS quality. The choice may be configured by a device manufacturer, application designer, or user”. Furthermore, ¶[0051] teaches “ In certain aspects, it may be desirable to have the local device be aware of which speech synthesis units it can handle to provide a sufficient quality result (such as those units that are locally stored with a number of robust examples), and which units are better handled by a remote TTS device (such as those units that are locally stored with a limited number of examples). For example, if a local unit database includes at least one example of each speech synthesis unit, the local device may not be aware when it should use its locally stored units for synthesis and when it should turn to the remote device. In this aspect, the local device may also be configured with a list of units and their corresponding acoustic features that are available at a remote TTS device and whose audio files should be retrieved from the remote device for speech synthesis.” The teachings of ¶[0051] in combination with the teachings in ¶[0046]-[0047] would lead one skilled in the art to configured the system to determine whether to utilize the local database or the remoted database based on resource consumption, i.e., performance, desired quality, and/or user preference), and the selection policy is a reactive selection policy (the “OR” language make this limitation optional) or a proactive selection policy (¶[0011] teaches the selection process is proactive in as much as either the local device or the remote device determines whether the local database has sufficient information to complete the speech process, thereby making the selection of local or remote data proactive ). Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to PENNY L CAUDLE whose telephone number is (703)756-1432. The examiner can normally be reached M-Th 8:00 am to 5:00 pm eastern. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at 571-272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /PENNY L CAUDLE/Examiner, Art Unit 2657 /DANIEL C WASHBURN/Supervisory Patent Examiner, Art Unit 2657
Read full office action

Prosecution Timeline

Jan 31, 2023
Application Filed
Jun 10, 2026
Non-Final Rejection mailed — §103
Jul 22, 2026
Interview Requested
Jul 28, 2026
Applicant Interview (Telephonic)
Jul 28, 2026
Examiner Interview Summary
Aug 17, 2026
Response Filed
Sep 01, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12731587
VISUAL SPEECH RECOGNITION FOR DIGITAL VIDEOS UTILIZING GENERATIVE ADVERSARIAL LEARNING
4y 7m to grant Granted Sep 08, 2026
Patent 12724812
AUTOMATIC GENERATION OF HANDOUTS FROM MULTI-MODAL DOCUMENTS
2y 8m to grant Granted Sep 01, 2026
Patent 12718829
SYSTEMS AND METHODS FOR VOICE RECEPTION AND DETECTION
3y 10m to grant Granted Aug 25, 2026
Patent 12711313
HUMAN-MACHINE COLLABORATIVE CONVERSATION INTERACTION SYSTEM AND METHOD
3y 2m to grant Granted Aug 18, 2026
Patent 12706091
PRONUNCIATION-AWARE EMBEDDING GENERATION FOR CONVERSATIONAL AI SYSTEMS AND APPLICATIONS
2y 6m to grant Granted Aug 11, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
70%
Grant Probability
85%
With Interview (+14.5%)
2y 11m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 84 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month