DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s arguments, see Remarks, filed 07/02/2026, with respect to the rejection of claims 5, 7, 8, 14, and 15 under 35 U.S.C. § 112(b) have been fully considered and are persuasive. The rejection of claims 5, 7, 8, 14, and 15 has been withdrawn.
Applicant’s arguments, see Remarks, filed 07/02/2026, with respect to the rejection of claims 1-5, 7-12, 14, and 15 under 35 U.S.C. § 102(a)(1) have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new grounds of rejection is made under 35 U.S.C. § 103 in view of Garman and Richards for claims 1-5, 7-12, 14, and 15-16, and in view of Garman, Richards and Lin for claims 6 and 13. The new grounds of rejection is provided in view of a further search conducted by the examiner in view of the new the language of the independent claims 1 and 9, that contain language not previously recited in neither one of claims 1-15.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-5, 7-12, 14 and 15-16 are rejected under 35 U.S.C. 103 as being unpatentable over Garman (US PG Pub 20210256961) in view of Richards (US PG Pub 20210390944).
As per claims 1 and 9, Garman discloses: A method of providing voice synthesis service and an artificial intelligence-based synthesis service system, comprising: an artificial intelligence device (Garman; Fig. 10; p. 0070 - In the example shown in FIG. 10, exemplary memory contents are shown for both a training system (AI/computing device) and a speech synthesis system (AI/computing device). While in some cases, these may both be implemented in the same computer system, typically, they are implemented in different computer systems. For example, the training system may typically be implemented in a server system or systems, while the speech synthesis system may typically be implemented in an end user device, such as a smartphone, tablet, personal computer, etc); and a computing device configure to exchanges data with the artificial intelligence device (Garman; Fig. 10; p. 0070 - In the example shown in FIG. 10, exemplary memory contents are shown for both a training system (AI/computing device) and a speech synthesis system (AI/computing device). While in some cases, these may both be implemented in the same computer system, typically, they are implemented in different computer systems. For example, the training system may typically be implemented in a server system or systems, while the speech synthesis system may typically be implemented in an end user device, such as a smartphone, tablet, personal computer, etc), wherein the computing device includes: a processor (Garman; Fig. 10, items 1002A-1002N; p. 0066 - Computer system 1000 may include one or more processors (CPUs) 1002A-1002N) configured to: receive sound source data for synthesizing a speaker's voice for a plurality of predefined first texts through a voice synthesis service platform that provides a development toolkit (Garman; Fig. 8, item 104; p. 0044 - the end-user can optionally add his or her voice to the inventory of speakers. To add a voice, the user may read and record a series of prompts in his or her native or fluent language, for example…; see also p. 0024 - Examples of such uses may include presenting voice instructions, reading texts of books, magazines, etc. aloud, etc…; see also p. 0062); learn tone conversion for sound source data of a speaker using a pre-generated tone conversion base model (Garman; p. 0041 - …When trained, DRNN 228 may be a single universal (base) speech model and may encode all the information necessary to produce speech for all of the trained languages and speakers. The system learns, in the sense of “deep learning”, the phonemes for the various languages, as well as the voice characteristics of each of the speakers. The resulting DRNN 228 may be a single model of speech that contains all of the phonemes and prosodic patterns (tone conversion) for each of the languages and the voice characteristics of each of the included speakers; see also p. 0036; see also p. 0037-0040 – Every utterance (speaker’s sound source data) used for training may be encoded using a feature extraction function 230…; see also p. 0027 - In embodiments, the language 208 for each utterance is presented to the DRNN via embedding, while all other layers are language-independent. This allows the ASR 212 to employ transfer learning from one language to the next, resulting in an ASR 212 that gets more robust with each language that is added. For example, training on 8 different languages may produce an accuracy of about 80% at frame level phoneme identification. In addition to an embedding for each language, there is an embedding for a “universal” language. This “universal” language may be trained with a small percentage of data from all languages, and is useful for doing recognition on an “unknown” language that is not already trained, such as a “new” language); generate a first voice synthesis model for the speaker through learning the tone conversion (Garman; p. 0073 - Model training routines 1016 may include software routines to train the model using, for example, a sequence of embedded phonemes, prosodic values, language identifiers, and speaker identifiers as input, along with acoustic features, to generate trained model data 1018). Garman, however, fails to disclose when second text is being inputted, generate a second voice synthesis model through voice synthesis inference based on the first voice synthesis model for the speaker and the second text; and generate a synthesized voice using the second voice synthesis model, wherein when the second voice synthesis model is generated as a learning result, a tone conversion learning module generates a model similar to the first voice synthesis model generated according to a request or setting, or the tone conversion learning module is combined with another user's previously generated voice synthesis model for a corresponding user to generate a new voice synthesis model depending on a parasitic-generated voice synthesis model of the user, and various new voice synthesis models are combined and generated, wherein the combined and generated voice synthesis models are configured to link or map to each other by assigning identifiers, or are stored together, and wherein when the tone conversion learning module completes learning, the tone conversion learning module saves learning completion status information in a model management of the user. Richards does teach when second text is being inputted (Richards; p. 0066 - configurable neural speech synthesis inference model 310 is operable to receive input text (the input text being text received after pre-training the baseline speech synthesis model of p. 0075, using transcriptions (first text) & 0077) and one or more voice property values and generate synthesized speech audio as an output; see also p. 0078 - a configurable speech synthesis model 704 takes as input a voice property value and text), generate a second voice synthesis model through voice synthesis inference based on the first voice synthesis model for the speaker and the second text (Richards; p. 0077 - A pre-trained baseline speech synthesis model generates a particular voice for the speech that it synthesizes. For example, a target voice with a general accent, middle to young age, and neutral sounding gender may be preferred. After having pre-trained a baseline speech synthesis model, it is possible to perform transfer training (generate second voice synthesis model through voice synthesis inference) by training an improved speech synthesis model that has one or more additional input nodes to the neural network (based on the first voice synthesis model for the speaker and the second text), the nodes indicating voice property values); and generate a synthesized voice using the second voice synthesis model (Richards; p. 0066 - configurable neural speech synthesis inference model 310 is operable to receive input text and one or more voice property values and generate synthesized speech audio as an output) wherein when the second voice synthesis model is generated as a learning result (Richards; p. 0066 - Configurable neural speech synthesis inference model 310 is capable of inferring probabilities of properties of certain input samples and is both a part of training and a result of training a configurable neural speech synthesis model), a tone conversion learning module generates a model similar to the first voice synthesis model generated according to a request or setting, or the tone conversion learning module is combined with another user's previously generated voice synthesis model for a corresponding user to generate a new voice synthesis model depending on a parasitic-generated voice synthesis model of the user (Richards; p. 0077 - After having pre-trained a baseline speech synthesis model, it is possible to perform transfer training by training an improved speech synthesis model that has one or more additional input nodes to the neural network (generates a model similar to the first voice synthesis model), the nodes indicating voice property values; see also p. 0090 - A request for synthesized speech is received (according to a request or setting)), and various new voice synthesis models are combined and generated, wherein the combined and generated voice synthesis models are configured to link or map to each other by assigning identifiers (Richards; p. 0067 - In an embodiment, some neural speech synthesis models may use more than one internal neural network. For example, one may be trained to produce an audio spectrogram, and another uses the spectrogram to produce a waveform. Other ways of dividing the work of speech synthesis between different neural and expert-designed models are possible. FIG. 4A illustrates an exemplary embodiment 400 of the custom voice system 222 showing additional components in accordance with various embodiments. In this example, custom voice system 222 represents an example two-piece inference model for configurating neural speech synthesis and includes feature model 402 and vocoder 404. In an embodiment, high-level feature model 402 takes as input text to be converted to speech audio and one or more voice property values. It produces a spectrogram of speech as output. A vocoder 404 takes as input the spectrogram and produces synthesized speech audio as an output that can be stored in synthesized audio data store 312 or other appropriate data store, and/or otherwise utilized. FIG. 4B illustrates example 420 of spectrogram 422 of speech audio produced by high-level feature model 402 and used as input to a vocoder 404 (examples of combinations of linked models)), or are stored together (Richards; p. 0059-0060 – models stored together), and wherein when the tone conversion learning module completes learning, the tone conversion learning module saves learning completion status information in a model management of the user (Richards; p. 0051 - …allows a user to quickly and easily try different voice sounds and thereby find a voice that meets the needs of their product or use. Further, it allows for saving the property values and comparing them to others to ensure that they are different enough that different products' voices will be distinct (saving property values of model once transfer training is completed, based on the configuration property values for tone conversion set by the user)…). Therefore, it would have been obvious to one of ordinary skill in the art to modify the method and system of Garman to include when second text is being inputted, generate a second voice synthesis model through voice synthesis inference based on the first voice synthesis model for the speaker and the second text; and generate a synthesized voice using the second voice synthesis model, wherein when the second voice synthesis model is generated as a learning result, a tone conversion learning module generates a model similar to the first voice synthesis model generated according to a request or setting, or the tone conversion learning module is combined with another user's previously generated voice synthesis model for a corresponding user to generate a new voice synthesis model depending on a parasitic-generated voice synthesis model of the user, and various new voice synthesis models are combined and generated, wherein the combined and generated voice synthesis models are configured to link or map to each other by assigning identifiers, or are stored together, and wherein when the tone conversion learning module completes learning, the tone conversion learning module saves learning completion status information in a model management of the user, as taught by Richards, in order to enable the speech synthesis model to learn how to adapt the sound of the synthesized voice according to the voice property values (Richards; p. 0077) and thus provide a configurability that has the benefit of enabling rapid experimentation and testing of voices (tone conversion) that can affect the perception and relatability of machines that employ speech synthesis as configured (Richards; p. 0031). As per claims 2 and 10, Garman in view of Richards discloses: The method and system of claims 1 and 9, wherein the receiving the sound source data for synthesizing the speaker's voice for the plurality of predefined first texts includes: receiving the speaker's sound source multiple times for each first text; and generating sound source data for synthesizing the speaker's voice based on the speaker's sound source received multiple times (Garman; Fig. 8, item 104; p. 0044 - the end-user can optionally add his or her voice to the inventory of speakers. To add a voice, the user may read and record a series of prompts in his or her native or fluent language, for example…; see also p. 0024 - Examples of such uses may include presenting voice instructions, reading texts of books, magazines, etc. aloud, etc…; see also p. 0062).
As per claims 3 and 11, Garman in view of Richards discloses: The method and system of claims 2 and 10, wherein the sound source data for voice synthesis of the speaker is an average value of the speaker's sound source received multiple times (Garman; p. 0032 - In embodiments, the variables of prosody used by the system may include, for example, pitch, duration, and loudness. For each of these, the prosodic value may be relative to an average or predictable value for that variable).
As per claims 4 and 12, Garman in view of Richards discloses: The method and system of claims 3 and 11, wherein the learning the tone conversion includes performing speaker transfer learning based on the tone conversion base model (Garman; p. 0027 - In embodiments, the language 208 for each utterance is presented to the DRNN via embedding, while all other layers are language-independent. This allows the ASR 212 to employ transfer learning from one language to the next, resulting in an ASR 212 that gets more robust with each language that is added. For example, training on 8 different languages may produce an accuracy of about 80% at frame level phoneme identification. In addition to an embedding for each language, there is an embedding for a “universal” language. This “universal” language may be trained with a small percentage of data from all languages, and is useful for doing recognition on an “unknown” language that is not already trained, such as a “new” language).
As per claims 5, Garman in view of Richards discloses: The method of claim 1, wherein a plurality of the first voice synthesis models is generated for the speaker (Garman; p. 0073 - Model training routines 1016 may include software routines to train the model using, for example, a sequence of embedded phonemes, prosodic values, language identifiers, and speaker identifiers as input, along with acoustic features, to generate trained model data 1018; see also p. 0052 - The inputs for the phonemes 316, accents 308, and speakers 306 may be fed into an Embedding layer to generated embeddings 320, 324, 326. The prosodic inputs may be fed into an Embedding Bag laver to generate embeddings 322. These inputs may include, but are not limited to, stress, tone, focus, syllable position, punctuation type, part-of-speech; generating the speaker embeddings (voice synthesis inference) using the prosodic inputs (tone conversion)).
As per claims 7 and 14, Garman in view of Richards discloses: The method and system of claims 1 and 9, further comprising: receiving a speaker ID and third text (Garman; p. 0045 - …the inputs may include: the text to be spoken 302, the language of the text 304, identification of the speaker 306, and the output accent 308 (which may be the same as the language of the text)…; see also p. 0034 - each speaker 210 has a unique identifier that may, for example, be derived from the name of the corpus that contains them and the identifier within that corpus); calling the generated second voice synthesis model for the speaker corresponding to the speaker ID (Garman; p. 0052 - A speaker may be chosen. This may be a built-in speaker or an enrolled speaker. An output language may be chosen. This may be the language of the text or some other language. Typically, the language of the text may be chosen, producing accent-free speech in the target language. However, if desired, any accent can be introduced intentionally. The series of phonemes may be converted to speech using the output language and the selected speaker's voice characteristics); synthesizing voice for the third text based on the called second voice synthesis model (Garman; p. 0053 - The output of DRNN 228 for each frame represents the acoustic features of speech for that frame. Decoder 312 may take the acoustic features and generate a speech signal); and generating a synthesized voice for the third text (Garman; p. 0053 - The output of DRNN 228 for each frame represents the acoustic features of speech for that frame. Decoder 312 may take the acoustic features and generate a speech signal).
As per claims 8 and 15, Garman in view of Richards discloses: The method and system of claims 7 and 14, further comprising: receiving an input for at least one of volume level, pitch, and speed for the generated synthesized voice; and adjusting one of a volume level, pitch, and speed for the generated synthesized voice based on the received input (Garman; p. 0051 - Finally, the duration of each phoneme, in frames, may be determined 406. For this step, speaker 306 may be used, as speech rate varies between speakers. For a “fast” speaker, phonemes will have, on average, shorter duration. Note also that phoneme duration varies between languages. The text conversion system may handle this interaction of speaker and language to produce a speech rate that is consistent with both the speaker and language. In addition to duration, the other prosodic elements, pitch, and loudness, are also postulated during text conversion).
As per claim 16, Garman in view of Richards disclose:
The method of claim 1, upon which claim 16 depends. And further, Richards teaches wherein a similar model to the first voice synthesis model is a model in which some predefined parts of the first voice synthesis model have been arbitrarily modified and changed (Richards; p. 0077 - After having pre-trained a baseline speech synthesis model, it is possible to perform transfer training by training an improved speech synthesis model that has one or more additional input nodes to the neural network (generates a model similar to the first voice synthesis model), the nodes indicating voice property values).
Claims 6 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Garman in view of Richards and further in view of Lin (US PG Pub 20200058288).
As per claims 6, Garman in view of Richards discloses: The method of claim 1, upon which claims 6. Garman in view of Richards, however, fails to disclose wherein only the first text selected from the plurality of predefined first text is used for the voice synthesis. Lin does teach wherein only the first text selected from the plurality of predefined first text is used for the voice synthesis (Lin; p. 0041 - the processing apparatus 170 may select the text script 153 for model training (step S230). The text script 153 for model training may be the same or different from the indicating text in step S210, or may be other text materials designed to facilitate subsequent training of the timbre transformation model (for example, sentences including all finals or vowels)). Therefore, it would have been obvious to one of ordinary skill in the art to modify the method and system of Garman in view of Richards to include wherein only the first text selected from the plurality of predefined first text is used for the voice synthesis, as taught by Lin, because although a specific text article can be converted into synthetic human voice through text-to-speech (TTS) technology, there is no related existing products that provides a friendly operation interface for the user to select the voice timbre of a specific person that the user intends to listen to (Lin; p. 0004). As per claim 13, the claim recites language similar to the combination of claims 5 and 6, and therefore claim 13 is rejected similarly in view of Garman, Richards and Lin.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. The prior art made of record and not relied upon includes: Fernandez Guajardo (US PG Pub 20220392428) discloses: An online system receives, from a client device of a posting user, a script for a voice-based content item. The online system retrieves a voice synthesis model stored in the user profile of the posting user and generates a synthetic audio stream using the retrieved voice synthesis model and based on the received script. The online system presents the generated synthetic audio stream to the posting user and receives instructions for modifying the synthetic audio stream. The online system generates a second audio stream based on the received instructions and composes the voice-based tent item based on the generated second audio stream (Fernandez Guajardo; Abstract).
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Rodrigo A Chavez whose telephone number is (571)270-0139. The examiner can normally be reached Monday - Friday 9-6 ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached at 5712727602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/RODRIGO A CHAVEZ/Examiner, Art Unit 2658
/RICHEMOND DORVIL/Supervisory Patent Examiner, Art Unit 2658