Prosecution Insights
Last updated: September 17, 2026
Application No. 18/533,507

METHOD AND SYSTEM FOR GENERATING SYNTHESIS VOICE USING STYLE TAG REPRESENTED BY NATURAL LANGUAGE

Non-Final OA §103
Filed
Dec 08, 2023
Priority
Jun 08, 2021 — RE 10-2021-0074436 +2 more
Examiner
SIRJANI, FARIBA
Art Unit
2659
Tech Center
2600 — Communications
Assignee
Neosapience Inc.
OA Round
3 (Non-Final)
75%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 75% — above average
75%
Career Allowance Rate
430 granted / 570 resolved
+13.4% vs TC avg
Strong +32% interview lift
Without
With
+31.7%
Interview Lift
resolved cases with interview
Typical timeline
2y 9m
Avg Prosecution
20 currently pending
Career history
589
Total Applications
across all art units

Statute-Specific Performance

§101
16.0%
-24.0% vs TC avg
§103
51.6%
+11.6% vs TC avg
§102
12.7%
-27.3% vs TC avg
§112
11.7%
-28.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 570 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Now DETAILED ACTION Claims 1 and 8-20 were submitted. Claims 1, 14, and 20 are independent and are all method Claims. Claims 13 and 14-20 are directed to visual TTS and were withdrawn from consideration by an election on 12/18/2025. Claims 1-11 and 13 were under examination of which Claim 1 is independent. Claims 3-7 were canceled. Claims 1-2, 8-11, and 13 are now under examination of which Claim 1 is independent. Claims are amended. This Application was published as U.S. 20240105160. Apparent priority: 8 June 2021 (earlier of two KR priorities). Applicant’s amendments and arguments are considered but are either unpersuasive or moot in view of the new grounds of rejection. Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 7/15/2026 has been entered. Response to Amendments and Arguments Arguments are moot in view of the new grounds of rejection. Claim 1 is amended as follows: 1. A method for generating a synthesis voice for text, the method being performed by one or more processors and comprising: acquiring a text-to-speech synthesis model trained to generate a synthesis voice for a training text, wherein the text-to-speech synthesis model is configured to generate and output the synthesis voice based on reference voice data and a training style tag represented by natural language; receiving a target text associated with the first speaker; receiving a selection for a preset, wherein the preset includes a style represented by natural language, the style tag having been previously used to generate a synthesis voice of a second speaker different from the first speaker; acquiring the style tag included in the preset as a style tag for the target text; inputting the style tag acquired from the preset and the target text associated with the first speaker into the text-to-speech synthesis model; and acquiring a synthesis voice for the target text, the synthesis voice for the target text generated and output by the text-to-speech synthesis model, wherein the synthesis voice for the target text is a synthesis voice of the first speaker and comprises voice style features related to the style tag acquired from the preset. Applicant has pointed to Fig. 16 and Description for support: [0178] FIG. 16 is a diagram illustrating an example of reusing a style tag. The processor may provide a user interface to store the style tags as presets. [0179] The processor may store a specific style tag as a preset according to a user input. In FIG. 16, it is illustrated that a style tag 1610 corresponding to “in a tone with regret and irritation, as if screaming” is stored as a preset1 1620. For example, a preset menu or preset shortcut key to separately store the style tag 1610 may be predefined in the user interface, and the user may separately store the style tag as a preset using the preset menu or preset shortcut key. [0180] A plurality of different style tags may be stored as a preset list. For example, a first style tag may be stored in the preset1, a second style tag may be stored in a preset2, and a third style tag may be stored in a preset3. [0181] The style tags stored as the presets may be reused as the style tags for the other target texts. For example, if the processor receives a selection input for a preset, the processor may acquire a style tag included in the preset as a style tag for the target text. In FIG. 16, it is illustrated that the style tag corresponding to a preset1 1630 is applied to a sentence corresponding to “Ah, hitting the goal post! It's hard to score a goal.” [0182] Meanwhile, to facilitate at least some style adjustment (or modification) in the synthesis voice generated based on the style tag, the system for generating a synthesis voice may output a user interface that visually represents the embedding vector representing the corresponding voice style features. PNG media_image1.png 354 546 media_image1.png Greyscale See also: [0115] A second encoder 520 may receive a style tag 558 represented by natural language, extract voice style features 560 from the style tag 558, and provide the extracted voice style features 560 to the decoder 540. The style tag 558 may be written between delimiters. For example, the delimiter for the style tag 558 may be a symbol such as parentheses, a slash, a backslash, a less-than sign (<), and a greater-than sign (>). In addition, the voice style features 560 may be provided to the decoder 540 in vector form (e.g., embedding vector). The voice style feature 560 may be a text domain-based vector for the style tag 558 represented by natural language. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-2 and 8-9, 11 are rejected under 35 U.S.C. 103 as being unpatentable over Kim (U.S. 20210142783) in view of Chicote (U.S. 20210097976). PNG media_image2.png 332 326 media_image2.png Greyscale PNG media_image3.png 520 650 media_image3.png Greyscale PNG media_image4.png 490 716 media_image4.png Greyscale Regarding Claim 1, Kim teaches: 1. A method for generating a synthesis voice for text, the method being performed by one or more processors and comprising: [Kim, Figure 2, “synthetic speech generation system 230.” Figure 3 showing the hardware including memory 312/332 and processor 314/334.] acquiring a text-to-speech synthesis model trained to generate a synthesis voice for a training text, wherein the text-to-speech synthesis model is configured to generate and output the synthesis voice based on reference voice data and a training style tag represented by natural language; [Kim, Figure 8, ““[0134] The artificial neural network-based text-to-speech synthesis device may be trained using a large database existing as the text-speech signal pair. …to finally obtain a single artificial neural network text-to-speech synthesis model that outputs a desired speech when any text is input.” “[0126] … For example, the encoder 810 may use a machine learning model (e.g., a probability model, an artificial neural network, or the like) that has already been trained, to obtain the character embeddings on the basis of the divided input text. Furthermore, the encoder 810 may update the machine learning model while performing machine learning. ….”] receiving a target text associated with the first speaker; [Kim, Figure 6, S610. See Figures 9-10 when three speakers are participating: Jin-Hyuk and Beom-Su and Sun-Young. Text of speech for each is received.] receiving a selection for a preset, wherein the preset includes a style represented by natural language, [Kim, Figure 6, S620. Figure 9, 920 and 930 show the setting of the style for the TTS. The style tags are “Preset.” But they are not taught to be copied from another speaker. “[0141] According to an embodiment, when one is selected from among the plurality of sentences 910 received through the user interface and an icon 912 associated with the speech style is clicked, a speech style setting interface 920 may be displayed. According to the user's selection of one from among a plurality of speech styles included in the speech style setting interface 920, the speech style selected for the given sentence may be determined….” “[0118] Then, at S620, the speech style characteristics for the received one or more sentences may be determined. According to an embodiment, in response to a user input through one or more user interfaces, at least a part of the one or more sentences outputted through the user interfaces may be selected, and the speech style characteristics for at least a part selected sentences may be determined. …”] the style tag having been previously used to generate a synthesis voice of a second speaker different from the first speaker; [Kim, includes a recommendation feature that comes close to this limitation: “[0118] … In another embodiment, the synthetic speech generation system may recommend or determine the speech style characteristics for one or more sentences and provide them to the user terminal, and the user terminal may determine the speech style characteristics for the corresponding sentences based on the received speech style characteristics.” The recommendation of Kim is close to the “preset” of the Claim but not express.] acquiring the style tag included in the preset as a style tag for the target text; [Kim, Figure 10, 920 and 930 are style tags and appear as tags #3 and #5 in the text that is to be converted to speech: “[0141] … For example, when the user selects a sentence 922 “I am the CEO.” and clicks an icon 912 associated with the speech style list, the speech style setting interface 920 may be displayed. When the user selects a portion corresponding to “3” in the speech style setting interface 920, “#3” may be determined as the setting information in the sentence 922 “I am the CEO.” In addition, the speech style characteristic for the sentence 922 “I am the CEO.” may be determined or set as the speech style characteristic “awkwardly” which is a predetermined speech style characteristic corresponding to “#3”….”] inputting the style tag acquired from the preset and the target text associated with the first speaker into the text-to-speech synthesis model; and [Kim, Figures 9 and 10 show that the speech style and style tags (#3, #5, slow, fast) are appearing adjacent the text that is input to the TTS.] acquiring a synthesis voice for the target text, the synthesis voice for the target text generated and output by the text-to-speech synthesis model, [Kim, Figure 6, S 630. “[0119] Next, at S750 (sic), the synthetic speeches for the one or more sentences reflecting the speech style characteristics may be outputted. Here, the one or more sentences and the speech style characteristic may be inputted to the artificial neural network text-to-speech synthesis model and the synthetic speech may be generated based on the speech data outputted from the artificial neural network text-to-speech synthesis model. For example, the synthetic speech may be included in the user terminal or may be outputted through a connected speaker.”] wherein the synthesis voice for the target text is a synthesis voice of the first speaker and comprises voice style features related to the style tag acquired from the preset. [Kim, Figures 9 and 10 and the description teach the generation of speech of each speaker according to a different and user-selected style tag.] The only part that Kim does not teach is that the style tag is selected from a set of tags previously used to generate speech for another speaker Chicote teaches: receiving a selection for a preset, wherein the preset includes a style represented by natural language, the style tag having been previously used to generate a synthesis voice of a second speaker different from the first speaker; . [Chicote, Figure 4, the “vocal characteristic description 410” includes voice characteristics/styles of previously synthesized speeches and “voice sample 412” includes voice of a second speaker different from the first speaker: “[0016] …. The speech model may be trained to generate audio output that resembles the speaking style, tone, accent, or other vocal characteristic(s) of a particular speaker using training data from one or more human speakers. …” “[0018] The present disclosure relates to systems and methods for synthesizing speech from first input data (e.g., text data) using one or more vocal characteristics represented in second input data (e.g., audio recordings of people speaking, a textual description of a style of speaking, and/or a picture of a particular speaker). …” “[0047] FIG. 4 illustrates a vocal characteristic component 402, which may include a vocal characteristic model 406 for generating at least a portion of the vocal characteristic data 120 based on a received vocal characteristic description 410, (e.g., text data describing one or more vocal characteristics), a received voice sample 412 (e.g., a representation of speech exhibiting one or more vocal characteristics),…” “[0048] … For example, if the vocal characteristic description 410 includes the phrase “sound like a professor,” the NLU component 404 may determine that the vocal characteristic data 120 includes terms (or values representing the terms) “distinguished” and “received pronunciation accent.”….” Both “distinguished” and “accent” are similar to the tags shown in Figures of the instant Application. “[0049] ….For example, if the vocal characteristic description 410 corresponds to “sound like Elmo from Sesame Street,” the NLU component 404 may determine an intent that the vocal characteristic description 410 sounds like that television character and that the vocal characteristic data 120 includes indications such as “childlike.”…”] PNG media_image5.png 266 626 media_image5.png Greyscale Kim and Chicote pertain to generation of emotive and expressive speech and it would have been obvious to combine the method of selection of style from Chicote with the system of Kim as one way of setting the TTS style. This combination falls under combining prior art elements according to known methods to yield predictable results or use of known technique to improve similar devices (methods, or products) in the same way. See MPEP 2141, KSR, 550 U.S. at 418, 82 USPQ2d at 1396. Regarding Claim 2, Kim teaches: 2. The method according to claim 1, wherein the receiving the selection includes: providing a user interface to select the preset; and [Kim in Figures 9-10 shows that the style tags are presented to the user on the UI. “[0144] FIG. 10 is a diagram illustrating an exemplary screen 1000 of the user interface for providing a speech synthesis service according to an embodiment of the present disclosure. The speech style characteristics may be determined for the received one or more sentences 1010. The speech style characteristics may be determined or changed based on the setting information for visual representation of at least a part of one or more sentences….”] receiving the preset through the user interface. [Kim, Figures 9-10, user is selecting the style tags such as #3, #5 or slow, fast.] Regarding Claim 8, Kim teaches: 8. The method according to claim 1, wherein the text-to-speech synthesis model is configured to generate the synthesis voice for the target text reflecting the voice style features, based on features of reference voice data related to the style tag acquired from the preset. [Kim, Figures 6 and 9-13.] Regarding Claim 9, Kim teaches: 9. The method according to claim 1, wherein the text-to-speech synthesis model is configured to acquire embedding features for the style tag acquired from the preset, and generate the synthesis voice for the target text reflecting the voice style features based on the acquired embedding features. [Kim, Figure 8 showing the speech synthesizer as a series of encoder 810 and decoder 820 which operate based on embeddings including character embeddings and articulatory feature embedding vector which is the style embedding. “Speech data of the speaker 821” is another input which will be in the form of an embedding and determines style. “[0038] FIG. 8 is a diagram illustrating a configuration of an artificial neural network-based text-to-speech synthesis device, and a network for extracting an embedding vector that can distinguish each of a plurality of speakers according to an embodiment of the present disclosure.”] Regarding Claim 11, Kim teaches: 11. The method according to claim 1, wherein the text-to-speech synthesis model is configured to extract sequential prosodic features from the style tag acquired from the preset and generate the synthesis voice for the target text reflecting the sequential prosodic features as the voice style features. [Kim: “[0156] In this embodiment, the speech style characteristic may include a sequential prosody characteristic including prosody information corresponding to at least one unit of a frame, a phoneme, a letter, a syllable, a word, or a sentence in chronological order. In an example, the prosody information may include at least one of information on the volume of the sound, information on the pitch of the sound, information on the length of the sound, information on the pause duration of the sound, or information on the speed of the sound. In addition, the style of the sound may include any form, manner, or nuance that the sound or speech expresses, and may include, for example, tone, intonation, emotion, and the like inherent in the sound or speech. Further, the sequential prosody characteristic may be represented by a plurality of embedding vectors, and each of the plurality of embedding vectors may correspond to the prosody information included in chronological order.” “[0158] According to another embodiment, in order to change the speech style characteristic of at least a part of the given sentence, the user may provide the speech of the user reading the given sentence in a manner desired by the user to the synthetic speech generation system through the user interface. The synthetic speech generation system may input the received speech to an artificial neural network configured to infer the input speech as the sequential prosody characteristic, and output the sequential prosody characteristics corresponding to the received speech. Here, the outputted sequential prosody characteristics may be expressed by one or more embedding vectors. These one or more embedding vectors may be reflected in the graph provided through the interface 1320.”] Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Kim and Chicote in view of Shekhar (U.S. 20220028367). Regarding Claim 10, Kim teaches the use of a loss function and during the process of training loss function is minimized: “[0134] The artificial neural network-based text-to-speech synthesis device may be trained using a large database existing as the text-speech signal pair. A loss function may be defined by comparing the output to the text that is entered as the input, with the corresponding target speech signal. The text-to-speech synthesis device may learn the loss function through the error back propagation algorithm to finally obtain a single artificial neural network text-to-speech synthesis model that outputs a desired speech when any text is input.” Minimizing loss function is well-known and common in training but not express in Kim. Chicote includes training but does not discuss loss. Shekhar teaches: 10. The method according to claim 1, wherein the text-to-speech synthesis model is trained to minimize a loss between a first style feature extracted from the voice data and a second style feature extracted from the training style tag. [Shekhar, Figure 5, training of a neural network model involves minimizing/optimizing a loss function. “[0099] The expressive audio generation system 106 can utilize the loss function 508 to determine the loss (i.e., error) resulting from the expressive speech neural network 504 by comparing the predicted context-based speech map 506 with a ground truth 510 (e.g., a ground truth context-based speech map). The expressive audio generation system 106 can back propagate the determined loss to the expressive speech neural network 504 (as shown by the dashed line 512) to optimize the model by updating its parameters/weights. In particular, the expressive audio generation system 106 can back propagate the determined loss to each channel of the expressive speech neural network 504 (e.g., the character-level channel, the word-level channel, and, in some instances, the speaker identification channel) as well as the decoder of the expressive speech neural network 504 to update the respective parameters/weights of that channel. In some embodiments, the expressive audio generation system 106 back propagates the determined loss to each component of the expressive speech neural network (e.g., the character-level encoder, the word-level encoder, the location-sensitive attention mechanism, etc.) update the parameters/weights of that component individually. Consequently, with each iteration of training, the expressive audio generation system 106 gradually improves the accuracy with which the expressive speech neural network 504 can generate context-based speech maps for input texts (e.g., by lowering the resulting loss value). As shown, the expressive audio generation system 106 can thus generate the trained expressive speech neural network 514.”] Kim/Chicote and Shakeri pertain to generation of expressive speech and part of the expression is included in prosody and therefore it would have been obvious to combine the express teachings of Shakeri wrt to minimizing loss which is a well-known aspect of model training with the system of combination as a routine method of training. This combination falls under combining prior art elements according to known methods to yield predictable results or use of known technique to improve similar devices (methods, or products) in the same way. See MPEP 2141, KSR, 550 U.S. at 418, 82 USPQ2d at 1396. Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over Kim and Chicote in view of Bocchieri (U.S. 20120130709). Regarding Claim 13, Kim and Chicote do not mention an API call. Bocchieri teaches: 13. The method according to claim 1, wherein the receiving of the target text is via an API call. [Bocchieri is directed to generation of LM and AM models that can be used in speech recognition and therefore includes input of text corresponding to speech and teaches that the communications are performed via API calls. “[0011] Disclosed are systems, methods, and non-transitory computer-readable storage media for generating speech models from the perspective of a server and from a client device. The server receives a standard feature stream and/or an optional proprietary feature stream, transcriptions, and parameter values as inputs from a network client independent of and/or without access to or specific knowledge of internal operations of the automatic speech recognition system. However, the network client may have general knowledge of the available tools and functionalities via the API. The system processes the inputs to train an acoustic model and a language model. Then the server transmits the acoustic model and the language model to the network client. The client communicates with the server via an API. The client provides text, such as transcriptions of the input speech. The input speech can be recorded live from a user or can be selected from a database of previously recorded speech. The client device extracts features from the input speech and the input text based on configuration parameters and transmits, via an API call, the features, the input speech, the input text, and configuration parameter values to the server. Later, the client receives from the server an acoustic model and a language model generated based on the features or the input speech, the input text, and the configuration parameter values.” “[0033] The system then extracts features from the input speech and the input text based on configuration parameters (404). The system transmits, via an API call, the features, the input speech, the input text, and configuration parameter values to the server (406). The configuration parameter values can indicate one or more specific task, application, or desired use for the requested models. The ASR server can process the input speech, text, features, and so forth for a significant amount of time. While many API calls in other applications may result in a near-instantaneous response, the server may take several hours, days, or longer to generate an AM and LM in response to the API call. Thus, the system may wait for a long time for the server to respond to the API call. In one variation, the client does not keep a constant communication channel open with the server, such as an HTTP or HTTPS session or other persistent session, while waiting for the response to the API call from the server.”] Kim/Chicote and Bocchieri pertain to speech processing and it would have been obvious to combine the API call of Boccieri which is used for receiving all types of input including text with the system of combination in order to have a variation in the implementation. This combination falls under combining prior art elements according to known methods to yield predictable results or use of known technique to improve similar devices (methods, or products) in the same way. See MPEP 2141, KSR, 550 U.S. at 418, 82 USPQ2d at 1396. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Kim (U.S. 2020082807): (same inventor) PNG media_image6.png 296 750 media_image6.png Greyscale PNG media_image7.png 410 370 media_image7.png Greyscale Kim (U.S. 20210142783): (same inventor) [0053] As used herein, “setting information” may include visually recognizable information for distinguishing speech style characteristics that are set for one or more sentences through the user interface. For example, it may mean information such as a font, a font style, a font color, a font size, a font effect, an underline, an underline style, and the like that is applied to one or more sentences. As another example, the setting information such as “#3”, “slow”, and “1.5 s” indicative of speech style, sound effect, or silence may be displayed through the user interface. PNG media_image8.png 544 796 media_image8.png Greyscale [0063] The speech style characteristics may be determined for the received one or more sentences. These speech style characteristics may be determined or changed based on the setting information for one or more sentences. In an embodiment, such setting information may be determined or changed according to a user input. For example, the user may input or change setting information through a plurality of icons 136 located on a lower-left side of the user interface screen 100. According to another embodiment, the synthetic speech generation system may analyze one or more sentences to automatically determine the setting information for one or more sentences. For example, as shown in the user interface screen 100, the setting information 116 (“#3”) may be determined and displayed in the sentence “I am the CEO.”, and the speech style characteristic of the sentence “I am the CEO.” may be determined as the speech style characteristic of “awkwardly” corresponding to the setting information 116 (“#3”). As another example, the setting information 118 (“slow”) may be determined and displayed in the sentence of “I am glad to meet you.”, and the speech style characteristic for the sentence “I am glad to meet you” may be determined to be the slow speed style characteristic. [0072] The user terminals 210_1, 210_2, and 210_3 may determine or change the setting information for at least a part of the one or more sentences. According to an embodiment, the user terminal may select a sentence for at least a part of the one or more sentences outputted through the user interface, and designate a predetermined value and/or term indicative of a specific speech style for the selected sentence to thus determine or change the setting information for the selected sentence. The determination or change of the setting information may be performed in response to a user input. According to another embodiment, the user terminals 210_1, 210_2, and 210_3 may change the setting information (e.g., a font, a font style, a font color, a font size, a font effect, an underline, an underline style, or the like) for visual representation of at least a part of the outputted one or more sentences. For example, the user terminals 210_1, 210_2, and 210_3 may change the font size for at least a part of the outputted one or more sentences from 10 to 12, to thus change the setting information for at least a part of the outputted sentences. As another example, the user terminals 210_1, 210_2, and 210_3 may change the font color for at least a part of the outputted one or more sentences from black to red, to thus change the setting information for at least a part of the outputted sentences. [0089] According to an embodiment, the processor 314 may receive an input to determine or change the setting information for at least a part of one or more sentences through the input device 320. For example, the processor 314 may receive an input to change the setting information for the speech style or speech speed. As another example, the processor 314 may receive an input to change the setting information for visual representation, such as a font, a font style, a font color, a font size, a font effect, an underline or underline style, for the part of one or more sentences. As still another example, the processor 314 may receive an input to select at least a part of the one or more speech style characteristic candidates received from the synthetic speech generation system 230. As another example, the processor 314 may receive an input to change a value indicative of the speech style characteristic through an interface for changing the speech style characteristic for at least a part of one or more sentences. Based on the received input, the processor 314 may determine or change the setting information for at least a part of one or more sentences. Alternatively, the processor 314 may provide the received input to the synthetic speech generation system 230 through the communication module 316, and receive the speech style characteristic determined or changed according to setting information from the synthetic speech generation system 230. [0090] According to another embodiment, the processor 314 may receive an input to add a visual representation indicative of the characteristics of an effect to be inserted between a plurality of sentences through the input device 320. For example, the processor 314 may receive an input to add a visual representation indicative of sound effects to be inserted between a plurality of sentences. As another example, the processor 314 may receive an input to add a visual representation indicative of a time period of silence to be inserted between a plurality of sentences. The processor 314 may provide the input to add a visual representation indicative of the sound effect to the synthetic speech generation system 230 through the communication module 316, and receive a synthetic speech including or reflecting the sound effect from the synthetic speech generation system 230. PNG media_image9.png 550 756 media_image9.png Greyscale [0143] FIG. 9 shows an operation in which the speech style characteristic is determined according to an input through the user interface, but embodiment is not limited thereto, and in the synthetic speech generation system, the speech style characteristics may be automatically determined according to the analyzed result using the natural language processing or the like. For example, the synthetic speech generation system may recognize a sentence “Well . . . ” and determine the speech style characteristic “hesitantly” for the next sentence, that is, the sentence 932 “I'm glad to meet you”. In this case, unlike FIG. 9, “hesitantly” may be displayed in front of the sentence 932 “I am glad to meet you.” PNG media_image10.png 500 792 media_image10.png Greyscale [0144] FIG. 10 is a diagram illustrating an exemplary screen 1000 of the user interface for providing a speech synthesis service according to an embodiment of the present disclosure. The speech style characteristics may be determined for the received one or more sentences 1010. The speech style characteristics may be determined or changed based on the setting information for visual representation of at least a part of one or more sentences. In this case, the setting information for visual representation may include a font, a font style, a font color, a font size, a font effect, an underline, an underline style, or the like. In an embodiment, the setting information for visual representation may be determined or changed according to a user input. According to another embodiment, the synthetic speech generation system may analyze one or more sentences and automatically determine the setting information for visual representation of the one or more sentences. For example, as shown in the user interface screen 100, a font thickness of a sentence 1014 “emotion to text” may be determined in bold, and the speech style characteristic for the sentence 1014 “emotion to text” may be determined to be a bold speech style characteristic. As another example, an underline may be added to the sentence 1016 “artificial intelligence speech actor service”, and the speech style characteristic for the sentence 1016 of the “artificial intelligence voice actor service” may be determined to be an emphasizing speech style characteristic. As another example, the space between letters in the sentence 1018 “I am glad to meet you” may be determined to be wide, and the speech style characteristic of the sentence 1018 “I am glad to meet you” may be determined to be a slow-speed style characteristic. As another example, the sentence 1022 “What is this service?” may be determined to be tilted, and the speech style characteristic of the sentence 1022 “What is this service?” may be determined to be a sharp-tone speech style characteristic. As another example, the font of the sentence 1024 “We are constantly improving and upgrading the sound quality for better quality.” may be determined to be in an archetype, and the speech style characteristic for the sentence 1024 may be determined to be a sincere speech style characteristic. [0145] A silence may be inserted between the plurality of received sentences 1010. The time of silence to be inserted may be determined or changed based on the visual representation indicative of a time period of silence added between a plurality of received sentences. In this case, the visual representation indicative of the time period of silence may mean a space between two sentences among a plurality of sentences. For example, as shown, a space 1020 between the sentences “If you have any questions, please raise your hand and ask a question” and “Yes, lady in the front, ask a question, please.” may be determined to be wide, and the silence for a time corresponding to the space 1020 may be added between the two sentences. Kim (U.S. 11514887): (Same inventor) PNG media_image11.png 442 496 media_image11.png Greyscale “The text-to-speech synthesis terminal 100 may be configured to output speech data for an input text reflecting an articulatory feature of the designated speaker. For example, as shown in FIG. 1, when output speech data for the input text “How are you” is generated, the output speech data may reflect an articulatory feature of the selected “Person 1.” Here, an articulatory feature of a specific speaker may simulate the speaker's voice and also may include at least one of a variety of factors, such as style, prosody, emotion, tone, pitch, etc. included in the articulation. In order to generate the output speech data, the text-to-speech synthesis terminal 100 may provide an input text and a designated speaker to a text-to-speech synthesis apparatus and receive synthesized speech data (e.g., speech data “How are you” reflecting the articulatory feature of “Person 1”) from the text-to-speech synthesis apparatus. The text-to-speech synthesis apparatus will be described in detail below with reference to FIG. 2. The text-to-speech synthesis terminal 100 may output the synthesized speech data to the user 110. Unlike this, the text-to-speech synthesis terminal 100 may include the text-to-speech synthesis apparatus.” 5:55-6:10. Eide (U.S. 20060229872) A technique for producing speech output in a text-to-speech system is provided. A message is created for communication to a user in a natural language generator of the text-to-speech system. The message is annotated in the natural language generator with a synthetic speech output style. The message is conveyed to the user through a speech synthesis system in communication with the natural language generator, wherein the message is conveyed in accordance with the synthetic speech output style. Abstract [0023] If the message is annotated automatically, in block 206, the message is annotated in accordance with a defined set of rules that instruct as to when and where to provide a reminder of the synthetic nature of the system during communication with the caller. This built-in mechanism decides which sentences should contain a synthetic speech output style and what those synthetic speech output styles should be. A simple example of such a rule would be "on the first sentence and every 10 sentences thereafter, vary the speed on the central word of the utterance." Alternatively, the system could randomly assign certain sentences to contain a synthetic speech output style, and randomly choose which synthetic speech output style to include. Stephen (U.S. 20100042410) [0039] In step 230, selected models are applied to the input textual content. This application produces prosody annotations associated with textual units in the textual content. These textual units include but are not limited to syllables, words, phrases, sentences, and larger textual units. As discussed above in reference to FIG. 1, the annotations cover a board range of prosody characteristics, and their syntaxes and semantics vary with embodiments and their prosody models. Zhao (U.S. 20210097976): PNG media_image12.png 588 630 media_image12.png Greyscale [0052] The acoustic feature annotator 308 annotates the speech input 300 with acoustic features to generate acoustic feature annotations 312. The acoustic feature annotator 308 may, for example, generate tags that identify the prosody, pitch, energy, and harmonics at various points in the speech input 300. In some aspects, the acoustic feature annotator 308 operates in parallel with the speech-to-text system 302 (i.e., the speech-to-text system 302 generates the text result 310 at the same time as the acoustic feature annotator 308 generates the acoustic feature annotations 312). Chicote (U.S. 20210097976): [0018] The present disclosure relates to systems and methods for synthesizing speech from first input data (e.g., text data) using one or more vocal characteristics represented in second input data (e.g., audio recordings of people speaking, a textual description of a style of speaking, and/or a picture of a particular speaker). In various embodiments, the description of a voice is received; this voice description may include text data, speech data, or other data that describes vocal characteristics of a desired voice. These vocal characteristics may include a desired age, gender, accent, accent, and/or emotion. For example, the system may receive input data that represents the request: “Generate speech that sounds like a 40-year-old news anchor.” In other embodiments, a sample of a voice is received (from, e.g., a user); this voice sample may include speech data representing speech that has vocal characteristics resembling those desired by the user. For example, the system may receive speech data that represents the request, “Generate speech that sounds like this: ‘Hasta la vista, baby,’ in which the phrase “Hasta la vista, baby' is spoken in an Arnold Schwarzenegger-sounding accent. The system may receive image data that includes an image of Arnold Schwarzenegger and determine based thereon that the vocal characteristics include an Arnold Schwarzenegger-sounding accent. [0021] The speech model 100 may receive vocal characteristic data 120. The vocal characteristic data 120 may be one or more numbers arranged in a vector. Each number in the vector may represent a value denoting a particular vocal characteristic. A value may be a binary value to denote a yes/no vocal characteristic, such as “male” or “female.” A value may be an integer value to denote other vocal characteristics, such as age. Other values may be one number in a range of values to represent a degree of a vocal characteristic, such as “coarseness” or “speed.” In some embodiments, the vocal characteristic data 120 includes natural-language descriptions of vocal characteristics in lieu of numbers representing them. The vocal characteristic data 120 may be a fixed size (e.g., a vector having a dimension of 20 elements) or may vary in size. The size of the vocal characteristic data 120, and the vocal characteristics it represents, may be the same for all received text data 110 or may vary with different text data 110. For example, first vocal characteristic data 120 may correspond to “male, old, angry” while second vocal characteristic data 120 may correspond to “female, French accent.” PNG media_image13.png 212 528 media_image13.png Greyscale [0028] The TTS front end may also process context data 215, such as text tags or text metadata, that may indicate, for example, how specific words should be pronounced, for example by indicating the desired output speech quality in tags formatted according to the speech synthesis markup language (SSML) or in some other form. For example, a first text tag may be included with text marking the beginning of when text should be whispered (e.g., <begin whisper>) and a second tag may be included with text marking the end of when text should be whispered (e.g., <end whisper>). The tags may be included in the input text data 110 and/or the text for a TTS request may be accompanied by separate metadata indicating what text should be whispered (or have some other indicated audio characteristic). The speech synthesis engine 218 may compare the annotated phonetic units models and information stored in the TTS unit storage 272 and/or TTS parametric storage 280 for converting the input text into speech. The TTS front end and speech synthesis engine 218 may include their own controller(s)/processor(s) and memory or they may use the controller/processor and memory of server system, user device, or other device, for example. Similarly, the instructions for operating the TTS front end and speech synthesis engine 218 may be located within the TTS component 295, within the memory and/or storage of the system, user device, or within an external device. Farbrizio (U.S. 10839033): In some implementations, once the model has been trained, the model can generalize items which may not be in the original training data and can generate sequences of words that are semantically equivalent referring expressions for an item. For example, in some implementations, the model can include an encoder-decoder recurrent neural network (RNN). The RNN model can include a sequence-to-sequence predictive model trained to generate semantically equivalent referring expressions for items. The sequence-to-sequence predictive model can utilize an encoder to process item names and a decoder to process referring expressions. The vectors generated by the encoder and the decoder can be hidden layer outputs generate by the model respectively over time. This provides the benefit of enhanced contextual relevancy between item names and referring expressions that can be utilized to configure conversational system with conversational models capable of responding accurately to unseen expressions referring to an item which can be provided by a user in a search query or other dialog input. User Selection: Yu (U.S. 11741984): Please refer to FIG. 2, which illustrates a display window of an acoustic scene conversion in accordance with an embodiment of the present invention. The acoustic scene conversion application 200 may be configured to test the two phases of the acoustic scene conversion, respectively. The application 200 may be installed in a mobile phone. In the display window as shown in FIG. 2, several items may be inputted by user. In order to more precisely retrieving voice sounds, a gender selection 210 is used to specify user's gender, male or female. Although the gender selection 210 only shows two options including male or female, a third option, child, may be included in accordance with an embodiment of the present invention, since user's gender and/or age is decisive for his/her voice. In order to specify the acoustic scene, the scene selection 220 includes a scene selection button. In other embodiments, the scene selection 220 may include an AI model selection button. After the button is pressed, a dialog window of the application would be popped up for displaying of available scenes, i.e., AI models trained according to the scenes. After one of the available scenes is selected, name of the selected scene would be displayed on the window. A new scene selection 230 as shown in FIG. 2 includes a new scene selection button. After the button is pressed, a dialog window of the application would be popped up for displaying of available acoustic scenes. Although the embodiment as shown in FIG. 2 shows three selection inputs, it is already discussed that predetermined setting values may be used as default values so as it may free the user from inputting his/her selections each time. The inputs may be automatically filled according to dynamic recognitions. 5:59-6:23. PNG media_image14.png 540 512 media_image14.png Greyscale Church (U.S. 20190370283): [0051] As previously mentioned, methods presented herein may be designed to receive any number of parameters, such as user-selected target parameters that may define an event or a sound in a recording. In embodiments, in order to directly or indirectly remove certain substantive content, a user may select recording times and durations, keywords that may be used as wake-up words to initiate a recording or flag certain substantive content for removal or non-removal, and any number of other user preferences, e.g., parameters that may be associated with certain content types, such as a speaker, a time, a location, an environmental parameter, and the like, that may be used to identify to-be-removed content or content to be kept. Li (U.S. 12032922): PNG media_image15.png 670 834 media_image15.png Greyscale PNG media_image16.png 682 882 media_image16.png Greyscale PNG media_image17.png 660 686 media_image17.png Greyscale Poirier (U.S. RE44248) PNG media_image18.png 724 298 media_image18.png Greyscale Hager (U.S. 20090276215): [0081] As shown in FIG. 10, the navigation area 274 can also include an inbox selection mechanism 282, a my history selection mechanism 284, a settings selection mechanism 286, a help selection mechanism 288, and/or a log off selection mechanism 290. A user can select the inbox selection mechanism 282 in order to view the main page 272. The user can select the my history selection mechanism 284 in order to access previously corrected transcriptions. In some embodiments, if a user selects the my history selection mechanism 284, the correction interface 260 displays a history page (not shown) similar to the main page 272 that lists previously corrected transcriptions. Alternatively or in addition to displaying the information displayed in the main page 272 (e.g., file name, checked out by, checked in by, creation date, priority), the history page can display correction date(s) for each transcription. Mahyar (U.S. 10930263): Model selection. Ju (U.S. 20100145694): An automated "Voice Search Message Service" provides a voice-based user interface for generating text messages from an arbitrary speech input. Specifically, the Voice Search Message Service provides a voice-search information retrieval process that evaluates user speech inputs to select one or more probabilistic matches from a database of pre-defined or user-defined text messages. These probabilistic matches are also optionally sorted in terms of relevancy. A single text message from the probabilistic matches is then selected and automatically transmitted to one or more intended recipients. Optionally, one or more of the probabilistic matches are presented to the user for confirmation or selection prior to transmission. Correction or recovery of speech recognition errors avoided since the probabilistic matches are intended to paraphrase the user speech input rather than exactly reproduce that speech, though exact matches are possible. Consequently, potential distractions to the user are significantly reduced relative to conventional speech recognition techniques. Abstract. Pan (US 20220293091): PNG media_image19.png 270 952 media_image19.png Greyscale Systems are configured for generating spectrogram data characterized by a voice timbre of a target speaker and a prosody style of source speaker by converting a waveform of source speaker data to phonetic posterior gram (PPG) data, extracting additional prosody features from the source speaker data, and generating a spectrogram based on the PPG data and the extracted prosody features. The systems are configured to utilize/train a machine learning model for generating spectrogram data and for training a neural text-to-speech model with the generated spectrogram data. Abstract. TTS and Model: Fang (U.S. 20230298564): “Embodiments of this application provide a speech synthesis method performed by an electronic device. The method includes: acquiring a target text to be synthesized into a speech; generating hidden layer features and prosodic features of the target text, and predicting pronunciation duration of characters in the target text using an acoustic model corresponding to the target text; generating acoustic features corresponding to the target text based on the hidden layer features, the prosodic features and the pronunciation duration; and synthesizing a target speech corresponding to the target text according to the acoustic features. Using the solution provided by the embodiments of this application is beneficial to reducing the difficulty of speech synthesis.” Abstract. Oh (U.S. 20070106514): “[0017] Another aspect of the present invention provides a speech synthesis method for adjusting a speech style, comprising the steps of: (a) receiving a sentence with a marked friendliness level; (b) selecting a prosodic model based on the marked friendliness level of the sentence; and (c) generating a synthesized speech of the sentence with the marked friendliness level by obtaining speech segments from a synthesis unit database on the basis of the selected prosodic model, the synthesis unit database storing speech segments for each friendliness level.” Gururani (U.S. 20210151029): “[0001] Text-to-speech systems are systems that emulate human speech by processing text and outputting a synthesized utterance of the text. However, conventional text to speech systems may produce unrealistic, artificial sounding speech output and also may not capture the wide variation of human speech. Techniques have been developed to produce more expressive text-to-speech systems, however many of these systems do not enable fine-grained control of the expressivity by a user. In addition, many systems for expressive text-to-speech use large, complex models requiring a significant number of training examples and/or high-dimensional features for training.” “[0013] … In addition, by being trained using low-dimensional representations, models described in this specification can better disentangle between different aspects of speech style, thus providing a user with more control when generating expressive speech from text.” Yang (U.S. 20210366462), Figure 9 determines emotion from the text of a received message. “[0172] In addition, the TTS device 100 according to an embodiment of the present invention may further include a semantic analysis module, a context analysis module, and an emotion determination module, and the semantic analysis module,…”Yang, Figure 10 shows an emotion vector that includes the plurality of emotions (neutral, love, happy, anger, sad) that is detected in this input text. “[0176] For example, FIG. 10 is a diagram for explaining an emotion vector according to an embodiment of the present invention. FIG. 10, if a received message M1 is “Love you”, the TTS device 100 calculates a first emotion vector through the syntactic analysis module, and the first emotion vector may be a vector in which weights are set to a plurality of emotion elements EA1, EA2, EA3, EA4, and EA5). For example, the first emotion vector is an emotion vector containing emotion classification information that is999 generated by assigning a weight “0” to neutral EA1, anger EA4 and sad EA5, a weight “0.9” to love EA2, and a weight “0.1” to happy EA3 with respect to the message M1 of “Love you”. Yang, Figures 17a and 17b show the output of the suggested/recommended emotions E to the user for his selection. “[0215] FIGS. 17A and 17B shows an example of transmitting a message with emotion classification information set therein according to an embodiment of the present invention.” “[0216] Referring to FIG. 17A, the transmitting device 12 may execute an application for writing a text message and the text message may be input. In this case, an emotion setting menu Es may be provided. The emotion setting menu Es may be displayed as an icon. When the emotion setting menu ES is selected after the text message is input, the transmitting device 12 may display at least one emotion item on a display so that a user can select emotion classification information with respect to the input text message. According to an embodiment, the at least one emotion item may be provided as a pop-up window. According to an embodiment, at least one candidate emotion item frequently used by the user may be recommended as the at least one emotion item” Shakeri (U.S. 11790884), Figure 1, “speech audio generator 110” is a TTS and “speech script 108” is the text that is converted to audio of a particular style: “… In some examples, the speech audio generator comprises a text-to-speech module configured to receive text data representing speech content, player identifier data for the player, and optionally, speech style features, and to output speech audio in the voice of the player.” 5:27-33. Figure 2, “synthesizer 202” and Figure 3, “speech content data 301” indicate TTS. “The data representing speech content may comprise text data. The text data may be any digital data representing text. …” 6:1-2. “Methods and systems described herein also enable the learning of an accurate representation of a speaker's voice in the form of a speaker embedding. By learning a suitable speaker embedding for each speaker, and inputting this along with source acoustic features into the voice convertor module, the performance of the source speech audio (e.g. prosody) may be retained while realistically transforming the voice of the source speech audio into that of the player.” 4:47-55. Shakeri (U.S. 11790884) Figures 5, 6, and 7 are directed to model training and each include a “model trainer 508/608/708.” Minimizing a loss function between the training data and input data is a feature of training. The loss is the error and must be optimized/minimized: “Model trainer 608 receives the predicted target acoustic features 607 and the “ground-truth” target acoustic features 602, and updates the parameters of acoustic feature encoder 605 and acoustic feature decoder 606 in order to optimize an objective function. The objective function comprises a loss in dependence on the predicted target acoustic features 607 and the ground-truth target acoustic features 602. For example, the loss may measure a mean-squared error between the predicted target acoustic features 607 and the ground-truth target acoustic features 602….” 16:1-10. Any inquiry concerning this communication or earlier communications from the examiner should be directed to FARIBA SIRJANI whose telephone number is (571)270-1499. The examiner can normally be reached 9 to 5, M-F. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre Desir can be reached at 571-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Fariba Sirjani/ Primary Examiner, Art Unit 2659
Read full office action

Prosecution Timeline

Dec 08, 2023
Application Filed
Dec 30, 2025
Non-Final Rejection mailed — §103
Mar 18, 2026
Response Filed
Apr 16, 2026
Final Rejection mailed — §103
Jul 15, 2026
Request for Continued Examination
Jul 20, 2026
Response after Non-Final Action
Sep 10, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12731572
INTERACTION SERVICE PROVIDING SYSTEM, INFORMATION PROCESSING APPARATUS, INTERACTION SERVICE PROVIDING METHOD, AND RECORDING MEDIUM
2y 9m to grant Granted Sep 08, 2026
Patent 12725626
APPARATUS FOR OUTPUTTING AN AUDIO SIGNAL IN A VEHICLE CABIN
2y 2m to grant Granted Sep 01, 2026
Patent 12711310
LEARNING APPARATUS, TEXT GENERATION APPARATUS, LEARNING METHOD, TEXT GENERATION METHOD AND PROGRAM
4y 0m to grant Granted Aug 18, 2026
Patent 12705266
METHOD AND APPARATUS FOR AI-ASSISTED VIRTUAL ASSISTANT FOR SME AGENT
3y 0m to grant Granted Aug 11, 2026
Patent 12694866
SPEECH RECOGNITION METHOD AND APPARATUS, DEVICE, AND STORAGE MEDIUM
3y 9m to grant Granted Jul 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
75%
Grant Probability
99%
With Interview (+31.7%)
2y 9m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 570 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month