DETAILED ACTION
Notice of AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 08/24/2026 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Response to Amendment
3. Amendment filed 07/08/2026 has been considered by Examiner. Claims 1, 8-10 and 17-19 have been amended. Claims 1-20 are pending, and likewise Claims 1- 20 have been examined.
Response to Arguments
Applicant’s amendments and arguments filed 07/08/2026, with respect to claim(s) 1-20 have been fully considered.
35 U.S.C 101 rejections of Claims 1-20 have been withdrawn in view of the amended claims filed on 07/08/2026. Applicant amended independent claim 1, 10 and 19 by adding limitations, “receiving a phonetic text phrase via a phonetic text entry field of the interactive user interface, the phonetic text phrase corresponding to the at least one candidate word; providing the phonetic text phrase to a text-to-speech module; synthesizing, by the text-to-speech module, the phonetic text phrase into synthesized audio data; and outputting the synthesized audio data for playback via a speaker”, which shows a step by step procedure of generating synthesized audio data by using text to speech module and tying the generated synthesized audio data with a practical application of playing back the audio data via speaker to the user, thereby overcoming the 35 U.S.C. 101 rejection.
Applicant’s arguments filed 07/08/2026, with respect to claim(s) 1-20, under 35 U.S.C. 103 have been fully considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 8, 10, 17 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Legat et al. ( US 20140222415 A1), hereinafter referenced as Legat, in view of Mika et al. (EP 1081589 A2), hereinafter referenced as Mika, further in view of Lacoss-Arnold et al. ( US 20180108344 A1), hereinafter referenced as Lacoss-Arnold.
Regarding Claim 1, Legat teaches a method comprising:
processing words in a text document to identify a first plurality of candidate words that are predicted to be mispronounced during automated text-to-speech processing of the text document ( Legat: Para. [0089]-[0092], Fig. 2, the primary text-to-speech synthesizer 210-1 receives text sample 205-1. The processing resource 240-1 first analyzes the text sample 205-1 for occurrence of out-of-vocabulary words ( first plurality of candidate words that are predicted to be mispronounced) by using applying morpho-syntactic or other suitable analysis such as lexicon lookup algorithm 110-1) ;
filtering the first plurality of candidate words to remove one or more candidate words of the first plurality of candidate words and obtain a second plurality of candidate words that have fewer candidate words than the first plurality of candidate words ( Legat: Para.[0093], Fig. 2, the processing resource 240-1 uses the trained classification model 175-1 ( filtering) to estimate a probability that a respective detected out-of-vocabulary words will be mispronounced ( second plurality of candidate words) using the a locally available source such as grapheme- to-phoneme algorithm 111-1 during text-to-speech synthesis);
Legat while teaching the method of claim 1, fails to explicitly teach the claimed, annotating the text document to obtain an annotated text document that identifies the second plurality of candidate words; [[and]] outputting, for display via an interactive user interface, at least a portion of the annotated text document that identifies at least one candidate word of the second plurality of candidate words; receiving a phonetic text phrase via a phonetic text entry field of the interactive user interface, the phonetic text phrase corresponding to the at least one candidate word; providing the phonetic text phrase to a text-to-speech module; synthesizing, by the text-to-speech module, the phonetic text phrase into synthesized audio data; and outputting the synthesized audio data for playback via a speaker.
However, Mika does teach the claimed, annotating the text document to obtain an annotated text document that identifies the second plurality of candidate words ( Mika: Para.[0065], Fig. 6 illustrates highlighting ( annotating) of displayed text which are identified as problematic/difficult to pronounce for the text-to-speech engine 16);
[[and]] outputting, for display via an [interactive] user interface, at least a portion of the annotated text document that identifies at least one candidate word of the second plurality of candidate words ( Mika: Para.[0065], Fig. 6 illustrates displaying of highlighted ( annotated) texts/words which are identified as problematic/difficult to pronounce for the text-to-speech engine 16, such as word “e-mail”, “Bientot” are highlighted in the display).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Mika’s teaching of a user interface having a display for displaying text and a speech synthesizer including a loudspeaker, arranged to convert an input, dependent upon a text, to an audio output representative of a person reading the text , into the system and method of text-to-speech synthesizing which implements a lexicon lookup algorithm, taught by Legat, because, this would improve the accuracy of text to speech conversion and the level of comprehension of a user by identifying difficult words to pronounce.(Mika, Para.[0003]-[0005]).
Legat in view of Mika, while teaching the method of claim 1, fail to explicitly teach the claimed, receiving a phonetic text phrase via a phonetic text entry field of the interactive user interface, the phonetic text phrase corresponding to the at least one candidate word; providing the phonetic text phrase to a text-to-speech module; synthesizing, by the text-to-speech module, the phonetic text phrase into synthesized audio data; and outputting the synthesized audio data for playback via a speaker.
However, Lacoss-Arnold does teach the claimed, receiving a phonetic text phrase via a phonetic text entry field of the interactive user interface, the phonetic text phrase corresponding to the at least one candidate word ( Lacoss-Arnold: Para.[0018],[0035], TTS computing component 104 may receive text data and generate a machine pronunciation of the text data according to at least one phonetic rule stored in the TTS database 108 and provide the machine pronunciation to the user via user interface component 106, which plays back the machine pronunciation such that the user may hear the text data instead of reading the text data. If the machine pronunciation from TTS computing component 104 is phonetically incorrect to the user, then the user may provide a pronunciation correction to user interface component 106 ( interactive user interface). For example, user may provide “Piasa Street" to the text field, which is a street located in Alton, Ill. "Piasa" is a Native American word pronounced "PIE-uh-saw," however, based on the phonetic rules of the TTS computing device, the generated machine pronunciation may be "pee-AH-zah." In such cases, the user may provide a pronunciation correction of "PIE-uh-saw" ( phonetic text phrase) to the TTS computing device);
providing the phonetic text phrase to a text-to-speech module ( Lacoss-Arnold: Para.[0034],[0035], Fig. 1, TTS computing component 104 of TTS computing device 102 may receive text data);
synthesizing, by the text-to-speech module, the phonetic text phrase into synthesized audio data ( Lacoss-Arnold: Para.[0034],[0035], Fig. 1, TTS computing component 104 of TTS computing device 102 may receive text data and generate a machine pronunciation of the text data according to at least one phonetic rule stored in the TTS database 108.);
and outputting the synthesized audio data for playback via a speaker ( Lacoss-Arnold: Para.[0049],[0053], Fig. 3, TTS computing device 102 may include at least one media output component 208 for presenting information to user 202 and may generates a machine pronunciation for the text data that is output to user 202, by an audio output device such as a speaker of the media output component 208).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Lacross-Arnold’s teaching of systems and methods for correcting text-to speech pronunciation by generating a machine pronunciation of a text data according to at least one phonetic rule, and provide the machine pronunciation to a user interface, into the system and method, taught by Legat in view of Mika, because, this would provide the technical advantages of convenient and efficient correction of the TIS system's machine translations, increasing regional dialect capacity of the TTS systems, and increased user satisfaction and interaction with the TTS system.( Lacross-Arnold, Para.[0031]).
Claim 10 is a computing device claim comprising: a memory configured to store a text document ( Legat: Para.[0153], Fig. 5, Computer readable storage medium 312 can be any suitable device such as memory, optical storage, hard drive, floppy disk configured to store instructions and/ or data) ;one or more processors configured to ( Legat: Para.[0151], Fig. 5, processor 313): performing the steps in method claim 1 above and as such, claim 10 is similar in scope and content to claim 1 and therefore, claim 10 is rejected under similar rationale as presented against claim 1 above.
Claim 19 is a non-transitory computer readable storage medium claim having instructions stored thereon that, when executed, cause one or more processors to ( Legat: Para.[0151],[0153],[0156], Fig. 5, processor 313 accesses computer readable storage media 312 via the use of interconnect 311 in order to launch, run, execute, interpret or otherwise perform the instructions of application stored on computer readable storage medium 312): performing the steps in method claim 1 above and as such, claim 19 is similar in scope and content to claim 1 and therefore, claim 19 is rejected under similar rationale as presented against claim 1 above.
Regarding Claim 8, Legat in view of Mika, further in view of Lacoss-Arnold teach the method of claim 1. Mika further teaches, further comprising: receiving an input via the [interactive] user interface selecting the at least one candidate word of the second plurality of candidate words ( Mika: Para.[0024], [0037], Fig. 2, controller 14 accesses text from memory and display on the user interface which includes a display 4);
obtaining, responsive to receiving the input pronunciation audio data representative of a verbal pronunciation of the at least one candidate word of the second plurality of candidate words ( Mika: Para.[0027], Fig.2, The text-to-speech engine 16 receives a text input 18 from the controller and converts the text input to a synthetic speech output 22 which is transduced by the speaker 6 to sound waves);
and outputting the pronunciation audio data for playback via the [[a]] speaker ( Mika: Para.[0027], Fig.2, The text-to-speech engine 16 drives the loudspeaker 6. It receives a text input 18 from the controller and converts the text input to a synthetic speech output 22 which is transduced by the speaker 6 to sound waves. The speech output may be one word at a time, one phrase at a time or one sentence at a time).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Mika’s teaching of a user interface having a display for displaying text and a speech synthesizer including a loudspeaker, arranged to convert an input, dependent upon a text, to an audio output representative of a person reading the text , into the system and method, taught by Legat in view of Lacoss-Arnold, because, this would improve the accuracy of text to speech conversion and the level of comprehension of a user by identifying difficult words to pronounce.(Mika, Para.[0003]-[0005]).
Lacoss-Arnold further teaches the claimed, interactive user interface ( Lacoss-Arnold: Para.[0035], TTS computing component 104 may receive text data and generate a machine pronunciation of the text data according to at least one phonetic rule stored in the TTS database 108 and provide the machine pronunciation to the user via user interface component 106, which plays back the machine pronunciation such that the user may hear the text data instead of reading the text data. If the machine pronunciation from TTS computing component 104 is phonetically incorrect to the user, then the user may provide a pronunciation correction to user interface component 106 ( interactive user interface)).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Lacross-Arnold’s teaching of systems and methods for correcting text-to speech pronunciation by generating a machine pronunciation of a text data according to at least one phonetic rule, and provide the machine pronunciation to a user interface, into the system and method, taught by Legat in view of Mika, because, this would provide the technical advantages of convenient and efficient correction of the TIS system's machine translations, increasing regional dialect capacity of the TTS systems, and increased user satisfaction and interaction with the TTS system.( Lacross-Arnold, Para.[0031]).
Claim 17 is a computing device claim performing the steps in method claim 8 above and as such, claim 17 is similar in scope and content to claim 8 and therefore, claim 17 is rejected under similar rationale as presented against claim 8 above.
Claims 2, and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Legat et al. ( US 20140222415 A1), hereinafter referenced as Legat, in view of Mika et al. (EP 1081589 A2), hereinafter referenced as Mika, further in view of Lacoss-Arnold et al. ( US 20180108344 A1), hereinafter referenced as Lacoss-Arnold, further in view of Fliedner et al. ( US 8825620 B2), hereinafter referenced as Fliedner.
Regarding Claim 2, Legat in view of Mika, further in view of Lacoss-Arnold teach the method of claim 1. Legat in view of Mika, further in view of Lacoss-Arnold fail to teach the claimed, wherein filtering the first plurality of candidate words includes: identifying the one or more candidate words of the first plurality of candidate words that is a stop word; and removing the one or more candidate words of the first plurality of candidate words that are the stop words to obtain the second plurality of candidate words.
However, Fliedner does teach the claimed, wherein filtering the first plurality of candidate words includes: identifying the one or more candidate words of the first plurality of candidate words that is a stop word ( Fliedner: Column 8, lines 4-16, Fig. 3, The text processing unit 306 can perform any of a number of text processing procedures. Once the sentences ( or other groupings of text) are tokenized, the tokens can be analyzed for stop words, in order to filter out words that might cause problems with, or otherwise negatively impact, an indexing and/or search procedure. Stop words can include, for example, "a," "an," "of," "the," and other such words as known in the art);
and removing the one or more candidate words of the first plurality of candidate words that are the stop words to obtain the second plurality of candidate words ( Fliedner: Column 17, lines 5-8, removing stop words to normalize).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Fliedner’s teaching of behavioral word segmentation for use in processing search queries, into the system and method, taught by Legat in view of Mika, further in view of Lacoss-Arnold, because, by segmenting words of search query a refined list can be achieved for a better result. (Fliedner, Column 11, lines 1-14).
Claim 11 is a computing device claim performing the steps in method claim 2 above and as such, claim 11 is similar in scope and content to claim 2 and therefore, claim 11 is rejected under similar rationale as presented against claim 2 above.
Claims 3, 6, 12 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Legat et al. ( US 20140222415 A1), hereinafter referenced as Legat, in view of Mika et al. (EP 1081589 A2), hereinafter referenced as Mika, further in view of Lacoss-Arnold et al. ( US 20180108344 A1), hereinafter referenced as Lacoss-Arnold, further in view of Tiwari et al. (US 20210158804 A1), hereinafter referenced as Tiwari.
Regarding Claim 3, Legat in view of Mika, further in view of Lacoss-Arnold teach the method of claim 1. Legat in view of Mika, further in view of Lacoss-Arnold fail to explicitly teach the claimed, wherein filtering the first plurality of candidate words includes: identifying a candidate word count for the first plurality of candidate words that indicates a number of times each candidate word appears in the first plurality of candidate words; and removing the one or more candidate words from the first plurality of candidate words having the candidate word count that exceeds a threshold.
However, Tiwari does teach the claimed, wherein filtering the first plurality of candidate words includes: identifying a candidate word count for the first plurality of candidate words that indicates a number of times each candidate word appears in the first plurality of candidate words ( Tiwari: Para.[0021],Fig. 1A, The Confusion Management Logic 12 receives (or reads) data or words/phrases from a plurality of data sources, including an uncommon word List 20 and a corpus data set 22 and receives a unigram score (or word frequency) for each word. Para.[0039], word frequency or likelihood may be determined based on word count or word count ratio);
and removing the one or more candidate words from the first plurality of candidate words having the candidate word count that exceeds a threshold ( Tiwari: Para.[0042], if the probability (or word frequency or likelihood) of a corpus word is less than a probability or likelihood threshold (Tp), and the corpus word is not on the uncommon word list then it is not an important corpus word and removes the word from the Lexicon and saves the updated Lexicon).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Tiwari’s teaching of systems and methods to improve the performance of an automatic speech recognition (ASR) system using a confusion index indicative of the amount of confusion between words , into the system and method, taught by Legat in view of Mika, further in view of Lacoss-Arnold, because, by taking the language into account, the accuracy of the word conflict resolution can be improved. (Tiwari, Para.[0073], [0074]).
Claim 12 is a computing device claim performing the steps in method claim 3 above and as such, claim 12 is similar in scope and content to claim 3 and therefore, claim 12 is rejected under similar rationale as presented against claim 3 above.
Regarding Claim 6, Legat in view of Mika, further in view of Lacoss-Arnold teach the method of claim 1. Legat in view of Mika, further in view of Lacoss-Arnold fail to explicitly teach the claimed, wherein filtering the first plurality of candidate words includes: applying a language model to the first plurality of candidate words to determine a perplexity of each candidate word of the first plurality of candidate words; and removing, based on the perplexity of each candidate word of the first plurality of candidate words, the one or more candidate words of the first plurality of candidate words.
However, Tiwari does teach the claimed, wherein filtering the first plurality of candidate words includes: applying a language model to the first plurality of candidate words to determine a perplexity of each candidate word of the first plurality of candidate words ( Tiwari: Para.[0021],Fig. 1A, The Confusion Management Logic 12 receives (or reads) data or words/phrases from a plurality of data sources, including an uncommon word List 20 and a corpus data set 22, respectively, and calculates a Confusion Index (CI) or Confusion Score, which is an indication of the amount of confusion between words ( perplexity), using parameters received from a Language Model 26 and an Acoustic Score Tool 28, respectively and a weighting factor received from a CI Parameter table);
and removing, based on the perplexity of each candidate word of the first plurality of candidate words, the one or more candidate words of the first plurality of candidate words ( Tiwari: Para.[0076], the confusion index may be used to identify corpus words that are likely to cause confusion, and the system (or a user) can then decide to keep or remove words from the lexicon depending upon their application and the level of potential confusion).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Tiwari’s teaching of systems and methods to improve the performance of an automatic speech recognition (ASR) system using a confusion index indicative of the amount of confusion between words , into the system and method, taught by Legat in view of Mika, further in view of Lacoss-Arnold, because, by taking the language into account, the accuracy of the word conflict resolution can be improved. (Tiwari, Para.[0073], [0074]).
Claim 15 is a computing device claim performing the steps in method claim 6 above and as such, claim 15 is similar in scope and content to claim 6 and therefore, claim 15 is rejected under similar rationale as presented against claim 6 above.
Claims 4, 7, 13, 16 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Legat et al. ( US 20140222415 A1), hereinafter referenced as Legat, in view of Mika et al. (EP 1081589 A2), hereinafter referenced as Mika, further in view of Lacoss-Arnold et al. ( US 20180108344 A1), hereinafter referenced as Lacoss-Arnold, further in view of Wan et al. (US 20230035947 A1), hereinafter referenced as Wan.
Regarding Claim 4, Legat in view of Mika, further in view of Lacoss-Arnold teach the method of claim 1. Legat in view of Mika, further in view of Lacoss-Arnold fail to explicitly teach the claimed, wherein filtering the first plurality of candidate words includes: identifying the one or more candidate words of the first plurality of candidate words that are homographs not specified in a common homograph list, and removing the one or more candidate words of the first plurality of candidate words that are identified as homographs not specified in the common homograph list.
However, Wan does teach the claimed, wherein filtering the first plurality of candidate words includes: identifying the one or more candidate words of the first plurality of candidate words that are homographs not specified in a common homograph list ( Wan: Para.[0075],[0111],[0112], identifying homonym phrases ( another type of homograph, with same spelling and pronunciation, different meaning) in the hot word list),
and removing the one or more candidate words of the first plurality of candidate words that are identified as homographs not specified in the common homograph list ( Wan; Para.[0117], Keywords are filtered by calculating language model scores of sentences including homonyms of the keywords. The homonym phrases associated with lower language model score is removed from the hot word list.
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Wan’s teaching of speech recognition with a customized language model , into the system and method, taught by Legat in view of Mika, further in view of Lacoss-Arnold, because, this would improve the accuracy of speech recognition by acquiring and updating the word list with the help of customized language model.(Wan, Para.[0018]-[0022]).
Claim 13 is a computing device claim performing the steps in method claim 4 above and as such, claim 13 is similar in scope and content to claim 4 and therefore, claim 13 is rejected under similar rationale as presented against claim 4 above.
Claim 20 is a non-transitory computer readable storage medium claim performing the steps in method claim 4 above and as such, claim 20 is similar in scope and content to claim 4 and therefore, claim 20 is rejected under similar rationale as presented against claim 4 above.
Regarding Claim 7, Legat in view of Mika, further in view of Lacoss-Arnold teach the method of claim 1. Legat in view of Mika, further in view of Lacoss-Arnold fail to explicitly teach the claimed, wherein filtering the first plurality of candidate words includes: applying a learning model to the first plurality of candidate words to determine a confidence score for each candidate word of the first plurality of candidate words that are homographs; and removing, based on the confidence score for each candidate word of the first plurality of candidate words that are homographs, the one or more candidate words of the first plurality of candidate words.
However, Wan does teach the claimed, wherein filtering the first plurality of candidate words includes: applying a learning model to the first plurality of candidate words to determine a confidence score for each candidate word of the first plurality of candidate words that are homographs ( Wan: Para.[0111],[0115], customized language model ( learning model) calculates score for the homonym phrases in the list);
and removing, based on the confidence score for each candidate word of the first plurality of candidate words that are homographs, the one or more candidate words of the first plurality of candidate words ( Wan; Para.[0117], Keywords are filtered by calculating language model scores of sentences including homonyms of the keywords. The homonym phrases associated with lower language model score is removed from the hot word list.
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Wan’s teaching of speech recognition with a customized language model , into the system and method, taught by Legat in view of Mika, further in view of Lacoss-Arnold, because, this would improve the accuracy of speech recognition by acquiring and updating the word list with the help of customized language model.(Wan, Para.[0018]-[0022]).
Claim 16 is a computing device claim performing the steps in method claim 7 above and as such, claim 16 is similar in scope and content to claim 7 and therefore, claim 16 is rejected under similar rationale as presented against claim 7 above.
Claims 5 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Legat et al. ( US 20140222415 A1), hereinafter referenced as Legat, in view of Mika et al. (EP 1081589 A2), hereinafter referenced as Mika, further in view of Lacoss-Arnold et al. ( US 20180108344 A1), hereinafter referenced as Lacoss-Arnold, further in view of Maergner et al. (US 20170287474 A1), hereinafter referenced as Maergner.
Regarding Claim 5, Legat in view of Mika, further in view of Lacoss-Arnold teach the method of claim 1. Legat in view of Mika, further in view of Lacoss-Arnold fail to explicitly teach the claimed, wherein filtering the first plurality of candidate words includes: identifying the one or more candidate words of the first plurality of candidate words that are named entities specified in a common named entities list; and removing the one or more candidate words of the first plurality of candidate words that are identified as named entities specified in the common named entities list.
However, Maergner does teach the claimed, wherein filtering the first plurality of candidate words includes: identifying the one or more candidate words of the first plurality of candidate words that are named entities specified in a common named entities list (Maergner: Para.[0040], Fig .2, the plurality of named entities 218 may be stored as a list comprising the named entities 218) ;
and removing the one or more candidate words of the first plurality of candidate words that are identified as named entities specified in the common named entities list (Maergner: Para.[0047], Fig .2, the list comprising the plurality of named entities 218 may be pre-modified for use in the speech recognition system (e.g., system 200). The system may access the plurality of named entities 218 and determine a subset of named entities that match at least one predefined criterion. The system may then remove the subset of named entities from the list comprising the plurality of named entities 218).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Maergner’s teaching of methods and systems for improving speech recognition of multilingual named entities , into the system and method, taught by Legat in view of Mika, further in view of Lacoss-Arnold, because, this would improve the accuracy of speech recognition with increased accuracy in recognizing named entities in speech.( Maergner, Para.[0002]-[0004]).
Claim 14 is a computing device claim performing the steps in method claim 5 above and as such, claim 14 is similar in scope and content to claim 5 and therefore, claim 14 is rejected under similar rationale as presented against claim 5 above.
Claims 9 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Legat et al. ( US 20140222415 A1), hereinafter referenced as Legat, in view of Mika et al. (EP 1081589 A2), hereinafter referenced as Mika, further in view of Lacoss-Arnold et al. ( US 20180108344 A1), hereinafter referenced as Lacoss-Arnold, further in view of Naik et al. (US 9620104 B2), hereinafter referenced as Naik.
Regarding Claim 9, Legat in view of Mika, further in view of Lacoss-Arnold teach the method of claim 1. Mika further teaches, further comprising: receiving an input via the [interactive] user interface selecting the at least one candidate word of the second plurality of candidate words ( Mika: Para.[0024], [0037], Fig. 2, controller 14 accesses text from memory and display on the user interface which includes a display 4);
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Mika’s teaching of a user interface having a display for displaying text and a speech synthesizer including a loudspeaker, arranged to convert an input, dependent upon a text, to an audio output representative of a person reading the text , into the system and method, taught by Legat in view of Lacoss-Arnold, because, this would improve the accuracy of text to speech conversion and the level of comprehension of a user by identifying difficult words to pronounce.(Mika, Para.[0003]-[0005]).
Legat in view of Mika further in view of Lacoss-Arnold, while teaching the method of claim 9, fail to explicitly teach the claimed, interactive user interface, receiving a verbal pronunciation of the at least one candidate word of the second plurality of candidate words; identifying, based on the verbal pronunciation, a potential pronunciation from a plurality of potential pronunciations; and associating the potential pronunciation to the at least one candidate word of the second plurality of candidate words.
However, Naik does teach the claimed, interactive user interface ( Naik: Column 10, lines 48-64, the digital assistant system 300 represents the server portion of a digital assistant implementation, and interacts with the user through a client-side portion residing on a user device ( interactive user interface)).
receiving a verbal pronunciation of the at least one candidate word of the second plurality of candidate words ( Naik: Column 13, lines 16-23, receiving an utterance of a word “ tomato”) ;
identifying, based on the verbal pronunciation, a potential pronunciation from a plurality of potential pronunciations ( Naik: Column 13, lines 10-23, receiving an utterance of a word “ tomato”. Identifies the sequence of phonemes “tuh-may-doe” and “tuh-mah-doe” for “tomato” by using language model ( potential pronunciation)) ;
and associating the potential pronunciation to the at least one candidate word of the second plurality of candidate words ( Naik: Column 12, lines 63-67, column 13, lines 1-15, associating one of the candidate pronunciations as a predicted pronunciation (e.g., the most likely pronunciation)).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Naik’s teaching of a system and method where digital assistants that make use of user-specified pronunciations of words for speech synthesis and recognition, into the system and method, taught by Legat in view of Mika, further in view of Lacoss-Arnold, because, this would improve the accuracy of speech recognition and synthesis, by using digital assistant's ability to fulfill a user's request by interacting with user to process difficult pronunciations.(Naik, Column 1, lines 25-67, column 2, lines 1-6 ).
Claim 18 is a computing device claim performing the steps in method claim 9 above and as such, claim 18 is similar in scope and content to claim 9 and therefore, claim 18 is rejected under similar rationale as presented against claim 9 above.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NADIRA SULTANA whose telephone number is (571)272-4048. The examiner can normally be reached M-F,7:30 am-5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Paras D. Shah can be reached on (571) 270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/NADIRA SULTANA/Examiner, Art Unit 2653
/Paras D Shah/Supervisory Patent Examiner, Art Unit 2653
09/11/2026