DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1-20 are pending and have been examined. Claims 1, 10, and 19 are independent.
This Application was published as U.S. 20250166599.
Priority
Acknowledgment is made of applicant's claim for foreign priority under 35 U.S.C. 119(a)-(d) to Korean Application No. 10-2023-0161153, filed November 20, 2023. Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55, retrieved electronically under 37 CFR 1.55(i) on December 30, 2024.
Claim Objections
Claims 1, 10, and 19 are objected to because of the following informalities:
a. In claims 1 and 10, "at least one or more symbols" should apparently read "at least one symbol."
b. In claim 19, "instructions executed in at least one processor" should apparently read "instructions executable in at least one processor."
Appropriate correction is required.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 6-8, 10, 15-17, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Goel et al. (U.S. 20220269870; filed December 30, 2021; published August 25, 2022), hereinafter Goel, in view of Oh (U.S. 6141642, Samsung Electronics Co., Ltd.; filed October 16, 1998; issued October 31, 2000), hereinafter Oh.
Goel was published on August 25, 2022, before the November 20, 2023 effective filing date of the claimed invention, and qualifies as prior art under 35 U.S.C. 102(a)(1).
Oh was published on October 31, 2000, before the November 20, 2023 effective filing date of the claimed invention, and qualifies as prior art under 35 U.S.C. 102(a)(1).
Regarding claim 1, Goel discloses:
1. A method of outputting text input as sound from an electronic device, the method comprising: (Goel discloses a method of reading out the text content of a message as sound at a device of the recipient: "the assistant system may provide a user with audio readout of a communication content (e.g., a message) when the communication content includes non-Latin script content items" (Goel, paras. [0008] and [0124]-[0125]).)
receiving, by at least one processor, a text input including characters (Goel discloses receiving content directed to the recipient for rendering: "the assistant system 140 may access a communication content comprising zero or more Latin script text strings and one or more non-Latin script content items" (Goel, paras. [0124]-[0125] and [0128]; see also claim 2).)
and at least one or more symbols; (Goel discloses that the received content contains emoji and symbol items: "the one or more non-Latin script content items may comprise one or more of an emoji or a symbol. Emoji and symbols may be parsable by the assistant system 140" (Goel, para. [0130]; see also para. [0128] and claim 2). The specification defines a symbol as a Unicode emoji (Spec., para. [0041]), and the emoji and symbol items of Goel are those symbols.)
generating, by the at least one processor, a plurality of ordered groups by sequencing, wherein the plurality of ordered groups includes (Goel discloses processing the content into ordered units for rendering: "determine a readout of the communication content based on one or more parsing rules. In particular embodiments, the one or more parsing rules may specify one or more formats for the" (Goel, para. [0125]; see also Fig. 9, step 920).)
a symbol group including symbols, by segmenting the text input sequentially (Goel discloses identifying the emoji and symbol items within the text and grouping them into units, the number of items in a unit being counted against a threshold: "the one or more attributes may comprise one or more of a threshold requirement for the one or more non-Latin script content items or a description difficulty associated with each of the one or more non-Latin script content items" (Goel, paras. [0130]-[0131]).)
by symbol; (Goel discloses that the division is made on the symbol items themselves: "the one or more non-Latin script content items may comprise one or more of an emoji or a symbol. Emoji and symbols may be parsable by the assistant system 140" (Goel, para. [0130]).)
displaying, by the at least one processor, the symbols of the symbol group on a display (Goel discloses splitting the rendering into an audio readout and an on-screen visual component carrying the symbol items: "the assistant system may provide a user with audio readout of a communication content (e.g., a message) when the communication content includes non-Latin script content items" (Goel, paras. [0008] and [0124]), the message being displayed on a screen while the readout plays: "FIG. 8 illustrates an example readout of a communication content comprising non-Latin script text strings" (Goel, para. [0150]; see also Fig. 8).)
Goel does not teach “from at least two languages”; “character groups including characters from a common language, and”; “by language and”; “generating, by the at least one processor, sound segments corresponding to each of the character groups by use of a plurality of text-to-speech (TTS) engines supporting different languages;”; “generating, by the at least one processor, an output sound by merging the sound segments according to the sequencing of the character groups within the plurality of ordered groups; and”; and “and outputting the output sound by use of a speaker.”
Oh discloses:
from at least two languages (Oh's starting point is input text containing more than one language: "a multiple language processing portion receiving multiple language text and dividing the input text into sub-texts according to language" (Oh, col. 1, l. 60 - col. 2, l. 2; worked example at col. 6, ll. 17-58).)
character groups including characters from a common language, and (Oh divides the input text into sub-texts, each sub-text consisting of the characters of one language: "a multiple language processing portion receiving multiple language text and dividing the input text into sub-texts according to language" (Oh, col. 1, l. 60 - col. 2, l. 2; see also col. 4, ll. 25-32).)
by language and (Oh makes the division on language boundaries and processes the sub-texts in the order in which they occur: "(a) checking characters of an input multiple language text one by one until a character of a different language from the character under process is found" (Oh, steps (a)-(d) at col. 2, ll. 6-19).)
generating, by the at least one processor, sound segments corresponding to each of the character groups by use of a plurality of text-to-speech (TTS) engines supporting different languages; (Oh provides a plurality of text-to-speech engines, one for each supported language, and assigns each sub-text to the engine for its language: "a text-to-speech engine portion having a plurality of text-to-speech engines, one for each language, for converting the sub-texts divided by the multiple language processing portion into audio wave data" (Oh, col. 1, ll. 63-67; six-language example at col. 4, ll. 55-64; Korean engine 214 and English engine 212 at col. 4, l. 65 - col. 5, l. 2).)
generating, by the at least one processor, an output sound by merging the sound segments according to the sequencing of the character groups within the plurality of ordered groups; and (Oh outputs the synthesized sub-texts one after another in the order of the sub-texts through a single audio processing path: "repeating the steps (a) through (c) while replacing the current processed language by the different language found in the step (a), if there are more characters to be converted in the input text" (Oh, col. 2, ll. 6-19; col. 5, ll. 12-21; worked example at col. 6, ll. 17-58). Sequential output of the sound segments in the order of their character groups through the single audio path is a merging of the sound segments according to that sequencing.)
and outputting the output sound by use of a speaker. (Oh outputs the synthesized speech through the audio processor to the speaker: "a speaker for converting the analog audio signal converted by the audio processor into sound and outputting the sound" (Oh, col. 5, ll. 12-21).)
Goel and Oh pertain to converting text into audible speech for output by a device. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the readout of Goel such that the text is divided into sub-texts by language, each sub-text is synthesized by the text-to-speech engine supporting that language, and the resulting sound segments are output in the order of the sub-texts through the single audio path of Oh, the symbol items remaining routed to the display as Goel provides. One of ordinary skill in the art would have been motivated to make this modification in order to render text containing more than one language intelligibly in each of its languages, as Oh expressly teaches (Oh, col. 1, l. 60 - col. 2, l. 2).
Regarding claim 6, Goel in view of Oh teaches the claim from which it depends as set forth above, and Goel further teaches:
6. The method of claim 1, wherein the characters and the symbols in the text input include characters and symbols represented by Unicode. (Goel bases the individual readouts on the Unicode descriptions associated with the corresponding emoji or symbols: "The description of the one or more emojis or symbols may comprise individual readouts for one or more of the emojis or symbols. In particular embodiments, the individual readouts are based on Unicode descriptions" (Goel, para. [0130]; see also claim 17).)
No further modification of the combination is required; the rationale for combining Goel and Oh is the one provided for claim 1.
Regarding claim 7, Goel in view of Oh teaches the claim from which it depends as set forth above, and Goel further teaches:
7. The method of claim 1, wherein the text input is provided in response to a query from a user of the electronic device, or provided as a mobile message received on a user terminal communicatively linked to the electronic device. (The claim recites alternatives, of which the second is taught. Goel discloses content sent by a sender over a network and received at the client system associated with the recipient: "the assistant system 140 may receive the communication content from a sender. The communication content may be directed to one or more recipients" (Goel, para. [0128]; see also claim 2 and para. [0134] with Fig. 6A), that client system being linked to the device on which the readout occurs: "The rendering device 137 may be configured to render outputs generated by the assistant system 140 to the user" (Goel, para. [0033]).)
No further modification of the combination is required; the rationale for combining Goel and Oh is the one provided for claim 1.
Regarding claim 8, Goel in view of Oh teaches the claim from which it depends as set forth above, and Goel further teaches:
8. The method of claim 1, wherein the plurality of TTS engines includes: TTS engines embedded in the electronic device, TTS engines provided by a server communicatively linked to the electronic device, or a combination of the TTS engines in the electronic device and the TTS engines by the server. (Goel places the text-to-speech component within the server-side delivery system: "The delivery system 230 may comprise a CU composer 370, a response generation component 380, a dialog state writing component 382, and a text-to-speech (TTS) component 390" (Goel, para. [0102]), and discloses that the readout is performed on the device, at the server, or in a blended mode using both: "the assistant system 140 may assist a user via an architecture built upon client-side processes and server-side processes which may operate in various operational modes" (Goel, paras. [0047]-[0048]; see also paras. [0008] and [0124]).)
No further modification of the combination is required; the rationale for combining Goel and Oh is the one provided for claim 1.
Regarding claim 10, Goel discloses:
10. An electronic apparatus for outputting text input as sound, the apparatus comprising: (Goel discloses a client system that reads out the text content of a message as sound: "the assistant system may provide a user with audio readout of a communication content (e.g., a message) when the communication content includes non-Latin script content items" (Goel, paras. [0008] and [0124]-[0125]).)
a memory configured to store instructions; and (Goel discloses a non-transitory memory holding the instructions that carry out the readout: "memory 1004 includes main memory for storing instructions for processor 1002 to execute or data for processor 1002 to operate on" (Goel, para. [0181]; see also paras. [0179] and [0182] and claim 20).)
at least one processor configured to execute the instructions, wherein the at least one processor is configured for, by executing the instructions: (Goel discloses one or more processors that execute the instructions loaded from the memory: "computer system 1000 includes a processor 1002, memory 1004, storage 1006, an input/output (I/O) interface 1008, a communication interface 1010, and a bus 1012" (Goel, paras. [0179]-[0180]; see also claim 20).)
receiving the text input including characters (Goel discloses receiving content directed to the recipient for rendering: "the assistant system 140 may access a communication content comprising zero or more Latin script text strings and one or more non-Latin script content items" (Goel, paras. [0124]-[0125] and [0128]).)
and at least one or more symbols; (Goel discloses that the received content contains emoji and symbol items: "the one or more non-Latin script content items may comprise one or more of an emoji or a symbol. Emoji and symbols may be parsable by the assistant system 140" (Goel, para. [0130]; see also para. [0128] and claim 2).)
generating a plurality of ordered groups by sequencing, wherein the plurality of ordered groups includes (Goel discloses processing the content into ordered units for rendering: "determine a readout of the communication content based on one or more parsing rules. In particular embodiments, the one or more parsing rules may specify one or more formats for the" (Goel, para. [0125]; see also Fig. 9, step 920).)
a symbol group including symbols, by segmenting the text input sequentially (Goel discloses identifying the emoji and symbol items and grouping them into units: "the one or more attributes may comprise one or more of a threshold requirement for the one or more non-Latin script content items or a description difficulty associated with each of the one or more non-Latin script content items" (Goel, paras. [0130]-[0131]).)
by symbol; (Goel discloses that the division is made on the symbol items themselves: "the one or more non-Latin script content items may comprise one or more of an emoji or a symbol. Emoji and symbols may be parsable by the assistant system 140" (Goel, para. [0130]).)
displaying the symbols of the symbol group on a display (Goel discloses an on-screen visual component carrying the symbol items alongside the audio readout: "the assistant system may provide a user with audio readout of a communication content (e.g., a message) when the communication content includes non-Latin script content items" (Goel, paras. [0008] and [0124]; para. [0150] and Fig. 8).)
Goel does not teach “from at least two languages”; “character groups including characters from a common language, and”; “by language and”; “generating sound segments corresponding to each of the character groups by use of a plurality of text-to-speech (TTS) engines supporting different languages;”; “generating an output sound by merging the sound segments according to the sequencing of the character groups within the plurality of ordered groups; and”; and “and outputting the output sound by use of a speaker.”
Oh discloses:
from at least two languages (Oh's starting point is input text containing more than one language: "a multiple language processing portion receiving multiple language text and dividing the input text into sub-texts according to language" (Oh, col. 1, l. 60 - col. 2, l. 2).)
character groups including characters from a common language, and (Oh divides the input text into sub-texts, each consisting of the characters of one language: "a multiple language processing portion receiving multiple language text and dividing the input text into sub-texts according to language" (Oh, col. 1, l. 60 - col. 2, l. 2; col. 4, ll. 25-32).)
by language and (Oh makes the division on language boundaries and processes the sub-texts in order: "(a) checking characters of an input multiple language text one by one until a character of a different language from the character under process is found" (Oh, col. 2, ll. 6-19).)
generating sound segments corresponding to each of the character groups by use of a plurality of text-to-speech (TTS) engines supporting different languages; (Oh provides one engine per supported language and assigns each sub-text to the engine for its language: "a text-to-speech engine portion having a plurality of text-to-speech engines, one for each language, for converting the sub-texts divided by the multiple language processing portion into audio wave data" (Oh, col. 1, ll. 63-67; col. 4, ll. 55-64; col. 4, l. 65 - col. 5, l. 2).)
generating an output sound by merging the sound segments according to the sequencing of the character groups within the plurality of ordered groups; and (Oh outputs the synthesized sub-texts one after another in order through a single audio processing path: "repeating the steps (a) through (c) while replacing the current processed language by the different language found in the step (a), if there are more characters to be converted in the input text" (Oh, col. 2, ll. 6-19; col. 5, ll. 12-21).)
and outputting the output sound by use of a speaker. (Oh outputs the synthesized speech through the audio processor to the speaker: "a speaker for converting the analog audio signal converted by the audio processor into sound and outputting the sound" (Oh, col. 5, ll. 12-21).)
Goel and Oh pertain to converting text into audible speech for output by a device. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the readout of Goel such that the text is divided into sub-texts by language, each sub-text is synthesized by the text-to-speech engine supporting that language, and the resulting sound segments are output in the order of the sub-texts through the single audio path of Oh, the symbol items remaining routed to the display as Goel provides. One of ordinary skill in the art would have been motivated to make this modification in order to render text containing more than one language intelligibly in each of its languages, as Oh expressly teaches (Oh, col. 1, l. 60 - col. 2, l. 2).
Regarding claim 15, Goel in view of Oh teaches the claim from which it depends as set forth above, and Goel further teaches:
15. The electronic apparatus of claim 10, wherein the characters and the symbols in the text input include characters and symbols represented by Unicode. (Goel bases the individual readouts on the Unicode descriptions associated with the corresponding emoji or symbols: "The description of the one or more emojis or symbols may comprise individual readouts for one or more of the emojis or symbols. In particular embodiments, the individual readouts are based on Unicode descriptions" (Goel, para. [0130]; see also claim 17).)
No further modification of the combination is required; the rationale for combining Goel and Oh is the one provided for claim 1.
Regarding claim 16, Goel in view of Oh teaches the claim from which it depends as set forth above, and Goel further teaches:
16. The electronic apparatus of claim 10, wherein the text input is provided in response to a query from a user of the electronic apparatus, or provided as a mobile message received on a user terminal communicatively linked to the electronic apparatus. (The claim recites alternatives, of which the second is taught. Goel discloses content sent by a sender and received at the client system associated with the recipient: "the assistant system 140 may receive the communication content from a sender. The communication content may be directed to one or more recipients" (Goel, para. [0128]; see also claim 2 and para. [0134] with Fig. 6A), that client system being linked to the device on which the readout occurs: "The rendering device 137 may be configured to render outputs generated by the assistant system 140 to the user" (Goel, para. [0033]).)
No further modification of the combination is required; the rationale for combining Goel and Oh is the one provided for claim 1.
Regarding claim 17, Goel in view of Oh teaches the claim from which it depends as set forth above, and Goel further teaches:
17. The electronic apparatus of claim 10, wherein the plurality of TTS engines includes: TTS engines embedded in the electronic apparatus, TTS engines provided by a server communicatively linked to the electronic apparatus, or a combination of the TTS engines in the electronic apparatus and the TTS engines by the server. (Goel places the text-to-speech component within the server-side delivery system: "The delivery system 230 may comprise a CU composer 370, a response generation component 380, a dialog state writing component 382, and a text-to-speech (TTS) component 390" (Goel, para. [0102]), and discloses that the readout is performed on the device, at the server, or in a blended mode using both: "the assistant system 140 may assist a user via an architecture built upon client-side processes and server-side processes which may operate in various operational modes" (Goel, paras. [0047]-[0048]).)
No further modification of the combination is required; the rationale for combining Goel and Oh is the one provided for claim 1.
Regarding claim 19, Goel discloses:
19. A non-transitory computer-readable recording medium storing instructions executed in at least one processor for causing the at least one processor to perform: (Goel discloses non-transitory computer-readable storage media embodying software that, when executed, performs the readout: "One or more computer-readable non-transitory storage media embodying software that is operable when executed to: access a communication content" (Goel, para. [0186]; see also claim 19).)
receiving a text input including characters (Goel discloses receiving content directed to the recipient for rendering: "the assistant system 140 may access a communication content comprising zero or more Latin script text strings and one or more non-Latin script content items" (Goel, paras. [0124]-[0125] and [0128]).)
and at least one symbol; (Goel discloses that the received content contains emoji and symbol items: "the one or more non-Latin script content items may comprise one or more of an emoji or a symbol. Emoji and symbols may be parsable by the assistant system 140" (Goel, para. [0130]; see also para. [0128] and claim 2).)
generating a plurality of ordered groups by sequencing, wherein the plurality of ordered groups includes (Goel discloses processing the content into ordered units for rendering: "determine a readout of the communication content based on one or more parsing rules. In particular embodiments, the one or more parsing rules may specify one or more formats for the" (Goel, para. [0125]; see also Fig. 9, step 920).)
a symbol group including symbols, by segmenting the text input sequentially (Goel discloses identifying the emoji and symbol items and grouping them into units: "the one or more attributes may comprise one or more of a threshold requirement for the one or more non-Latin script content items or a description difficulty associated with each of the one or more non-Latin script content items" (Goel, paras. [0130]-[0131]).)
by symbol; (Goel discloses that the division is made on the symbol items themselves: "the one or more non-Latin script content items may comprise one or more of an emoji or a symbol. Emoji and symbols may be parsable by the assistant system 140" (Goel, para. [0130]).)
displaying the symbols of the symbol group on a display (Goel discloses an on-screen visual component carrying the symbol items alongside the audio readout: "the assistant system may provide a user with audio readout of a communication content (e.g., a message) when the communication content includes non-Latin script content items" (Goel, paras. [0008] and [0124]; para. [0150] and Fig. 8).)
Goel does not teach “from at least two languages”; “character groups including characters from a common language, and”; “by language and”; “generating sound segments corresponding to each of the character groups by use of a plurality of text-to-speech (TTS) engines supporting different languages;”; “generating an output sound by merging the sound segments according to the sequencing of the character groups within the plurality of ordered groups; and”; and “and outputting the output sound by use of a speaker.”
Oh discloses:
from at least two languages (Oh's starting point is input text containing more than one language: "a multiple language processing portion receiving multiple language text and dividing the input text into sub-texts according to language" (Oh, col. 1, l. 60 - col. 2, l. 2).)
character groups including characters from a common language, and (Oh divides the input text into sub-texts, each consisting of the characters of one language: "a multiple language processing portion receiving multiple language text and dividing the input text into sub-texts according to language" (Oh, col. 1, l. 60 - col. 2, l. 2; col. 4, ll. 25-32).)
by language and (Oh makes the division on language boundaries and processes the sub-texts in order: "(a) checking characters of an input multiple language text one by one until a character of a different language from the character under process is found" (Oh, col. 2, ll. 6-19).)
generating sound segments corresponding to each of the character groups by use of a plurality of text-to-speech (TTS) engines supporting different languages; (Oh provides one engine per supported language: "the multiple language processing portion can simultaneously include an English processor, Korean processor, Japanese processor, French processor, German processor, and a Mandarin Chinese processor" (Oh, col. 1, ll. 63-67; col. 4, ll. 55-64).)
generating an output sound by merging the sound segments according to the sequencing of the character groups within the plurality of ordered groups; and (Oh outputs the synthesized sub-texts one after another in order through a single audio path: "repeating the steps (a) through (c) while replacing the current processed language by the different language found in the step (a), if there are more characters to be converted in the input text" (Oh, col. 2, ll. 6-19; col. 5, ll. 12-21).)
and outputting the output sound by use of a speaker. (Oh outputs the synthesized speech through the audio processor to the speaker: "a speaker for converting the analog audio signal converted by the audio processor into sound and outputting the sound" (Oh, col. 5, ll. 12-21).)
Goel and Oh pertain to converting text into audible speech for output by a device. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the readout of Goel such that the text is divided into sub-texts by language, each sub-text is synthesized by the text-to-speech engine supporting that language, and the resulting sound segments are output in the order of the sub-texts through the single audio path of Oh, the symbol items remaining routed to the display as Goel provides. One of ordinary skill in the art would have been motivated to make this modification in order to render text containing more than one language intelligibly in each of its languages, as Oh expressly teaches (Oh, col. 1, l. 60 - col. 2, l. 2).
Claims 2, 11, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Goel in view of Oh as applied to claims 1, 10, and 19 above, and further in view of Iso-Sipila (U.S. 20050144003, Nokia Corporation; filed December 8, 2003; published June 30, 2005), hereinafter Iso-Sipila.
Iso-Sipila was published on June 30, 2005, before the November 20, 2023 effective filing date of the claimed invention, and qualifies as prior art under 35 U.S.C. 102(a)(1).
Regarding claim 2, Goel in view of Oh teaches the claim from which it depends as set forth above. The combination of Goel and Oh does not disclose the following, which Iso-Sipila teaches:
2. The method of claim 1, further including: replacing, by the at least one processor, characters of a character group expressed in a language not supported by the plurality of TTS engines with a phonetic representation according to a first language among the different languages supported by the plurality of TTS engines; and generating, by the at least one processor, using a TTS engine supporting the first language among the plurality of TTS engines, a sound segment for the phonetic representation. (Iso-Sipila generates the audio output of a word in a first language with a speech synthesizing engine that does not support that language, by mapping the phonemes of the word onto the phonemes of a second language the engine does support and synthesizing with that second language's engine and prosody models: "an audio output of a word in a first language can be generated by a speech synthesizing engine not having actual support for this language. Instead, the pronunciation phonemes of the word are mapped onto phonemes of at least one second language, for which the speech synthesizing engine does have support" (Iso-Sipila, para. [0007]), the target language being selected in dependence on the first language: "The at least one second language is advantageously selected based on the first language" (Iso-Sipila, para. [0011]; see also paras. [0028]-[0031] and [0036] with Fig. 3, steps S1 to S6, the German to US English example at paras. [0038]-[0041], and claims 1, 2, 6, and 10). The phoneme sequence of the supported second language onto which the word is mapped is the phonetic representation according to a first language among the languages the engines support, and the engine for that language generates the sound segment.)
Goel, Oh, and Iso-Sipila pertain to generating audible speech from written text by text-to-speech synthesis. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify the combination such that a character group identified by the language division of Oh as being in a language none of the plurality of text-to-speech engines supports is replaced with the phonetic representation of a supported language and synthesized by the engine for that language, in the manner of Iso-Sipila. One of ordinary skill in the art would have been motivated to make this modification in order that such a character group is still output as sound, as Iso-Sipila expressly teaches that the number of languages to be supported exceeds the number of existing engines (Iso-Sipila, para. [0009]).
Regarding claim 11, Goel in view of Oh teaches the claim from which it depends as set forth above. The combination of Goel and Oh does not disclose the following, which Iso-Sipila teaches:
11. The electronic apparatus of claim 10, wherein the at least one processor is configured to further perform: replacing characters of a character group expressed in a language not supported by the plurality of TTS engines with a phonetic representation according to a first language among the different languages supported by the plurality of TTS engines; and generating, using a TTS engine supporting the first language among the plurality of TTS engines, a sound segment for the phonetic representation. (Iso-Sipila generates the audio output of a word in a first language with a speech synthesizing engine that does not support that language, by mapping the phonemes of the word onto the phonemes of a second language the engine does support and synthesizing with that second language's engine and prosody models: "an audio output of a word in a first language can be generated by a speech synthesizing engine not having actual support for this language. Instead, the pronunciation phonemes of the word are mapped onto phonemes of at least one second language, for which the speech synthesizing engine does have support" (Iso-Sipila, para. [0007]), the target language being selected in dependence on the first language: "The at least one second language is advantageously selected based on the first language" (Iso-Sipila, para. [0011]; see also paras. [0028]-[0031] and [0036] with Fig. 3, steps S1 to S6, the German to US English example at paras. [0038]-[0041], and claims 1, 2, 6, and 10). The phoneme sequence of the supported second language onto which the word is mapped is the phonetic representation according to a first language among the languages the engines support, and the engine for that language generates the sound segment.)
Goel, Oh, and Iso-Sipila pertain to generating audible speech from written text by text-to-speech synthesis. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify the combination such that a character group identified by the language division of Oh as being in a language none of the plurality of text-to-speech engines supports is replaced with the phonetic representation of a supported language and synthesized by the engine for that language, in the manner of Iso-Sipila. One of ordinary skill in the art would have been motivated to make this modification in order that such a character group is still output as sound, as Iso-Sipila expressly teaches that the number of languages to be supported exceeds the number of existing engines (Iso-Sipila, para. [0009]).
Regarding claim 20, Goel in view of Oh teaches the claim from which it depends as set forth above. The combination of Goel and Oh does not disclose the following, which Iso-Sipila teaches:
20. The non-transitory computer-readable recording medium of claim 19, wherein the at least one processor is configured to further perform: replacing characters of a character group expressed in a language not supported by the plurality of TTS engines with a phonetic representation according to a first language among the different languages supported by the plurality of TTS engines; and generating, using a TTS engine supporting the first language among the plurality of TTS engines, a sound segment for the phonetic representation. (Iso-Sipila generates the audio output of a word in a first language with a speech synthesizing engine that does not support that language, by mapping the phonemes of the word onto the phonemes of a second language the engine does support and synthesizing with that second language's engine and prosody models: "an audio output of a word in a first language can be generated by a speech synthesizing engine not having actual support for this language. Instead, the pronunciation phonemes of the word are mapped onto phonemes of at least one second language, for which the speech synthesizing engine does have support" (Iso-Sipila, para. [0007]), the target language being selected in dependence on the first language: "The at least one second language is advantageously selected based on the first language" (Iso-Sipila, para. [0011]; see also paras. [0028]-[0031] and [0036] with Fig. 3, steps S1 to S6, the German to US English example at paras. [0038]-[0041], and claims 1, 2, 6, and 10). The phoneme sequence of the supported second language onto which the word is mapped is the phonetic representation according to a first language among the languages the engines support, and the engine for that language generates the sound segment.)
Goel, Oh, and Iso-Sipila pertain to generating audible speech from written text by text-to-speech synthesis. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify the combination such that a character group identified by the language division of Oh as being in a language none of the plurality of text-to-speech engines supports is replaced with the phonetic representation of a supported language and synthesized by the engine for that language, in the manner of Iso-Sipila. One of ordinary skill in the art would have been motivated to make this modification in order that such a character group is still output as sound, as Iso-Sipila expressly teaches that the number of languages to be supported exceeds the number of existing engines (Iso-Sipila, para. [0009]).
Claims 3 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Goel in view of Oh as applied to claims 1 and 10 above, and further in view of Story et al. (U.S. 9213705, Audible, Inc.; filed December 19, 2011; issued December 15, 2015), hereinafter Story.
Story was published on December 15, 2015, before the November 20, 2023 effective filing date of the claimed invention, and qualifies as prior art under 35 U.S.C. 102(a)(1).
Regarding claim 3, Goel in view of Oh teaches the claim from which it depends as set forth above. The combination of Goel and Oh does not disclose the following, which Story teaches:
3. The method of claim 1, further including: stopping, by the at least one processor, the displaying of the symbols of the symbol group on the display upon termination of an output of a sound segment corresponding to a character group preceding the symbol group. (Story associates each retrieved content item with a beginning time mark and/or an ending time mark within the audio content, the marks indicating the portion of the audio content during which the item is presented: "the presentation module 216 may associate each retrieved image or other content item with a beginning time mark and/or ending time mark within the audio content" (Story, col. 10, ll. 45-57; see also Fig. 4, block 412), and presents each item for display during playback of its corresponding audio portion only: "each of the one or more images is presented for display during playback of a corresponding portion of the audio content" (Story, col. 16, ll. 33-42; see also claim 1 at col. 16, ll. 5-42). The narration audio of Story expressly includes text-to-speech audio content: "narration audio (such as spoken word recordings of magazine or newspaper articles, podcasts, text-to-speech audio content, etc.)" (Story, col. 1, ll. 8-10). In the combination the symbol group is not itself spoken, so the audio portion the symbol group accompanies is the sound segment of the preceding character group, and bounding the presentation by the ending time mark of that portion stops the displaying of the symbols upon termination of the output of that sound segment.)
Goel, Oh, and Story pertain to presenting written content to a user as audio together with a synchronized visual display. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify the combination such that the display of the symbol group is bounded by the ending time mark of the sound segment it accompanies, in the manner of Story. One of ordinary skill in the art would have been motivated to make this modification in order to keep the content shown on the display matched to the audio then being output, as Story expressly teaches (Story, col. 10, ll. 45-57).
Regarding claim 12, Goel in view of Oh teaches the claim from which it depends as set forth above. The combination of Goel and Oh does not disclose the following, which Story teaches:
12. The electronic apparatus of claim 10, wherein the at least one processor is configured to further perform: stopping the displaying of the symbols of the symbol group on the display upon termination of an output of a sound segment corresponding to a character group preceding the symbol group. (Story associates each retrieved content item with a beginning time mark and/or an ending time mark within the audio content, the marks indicating the portion of the audio content during which the item is presented: "the presentation module 216 may associate each retrieved image or other content item with a beginning time mark and/or ending time mark within the audio content" (Story, col. 10, ll. 45-57; see also Fig. 4, block 412), and presents each item for display during playback of its corresponding audio portion only: "each of the one or more images is presented for display during playback of a corresponding portion of the audio content" (Story, col. 16, ll. 33-42; see also claim 1 at col. 16, ll. 5-42). The narration audio of Story expressly includes text-to-speech audio content: "narration audio (such as spoken word recordings of magazine or newspaper articles, podcasts, text-to-speech audio content, etc.)" (Story, col. 1, ll. 8-10). In the combination the symbol group is not itself spoken, so the audio portion the symbol group accompanies is the sound segment of the preceding character group, and bounding the presentation by the ending time mark of that portion stops the displaying of the symbols upon termination of the output of that sound segment.)
Goel, Oh, and Story pertain to presenting written content to a user as audio together with a synchronized visual display. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify the combination such that the display of the symbol group is bounded by the ending time mark of the sound segment it accompanies, in the manner of Story. One of ordinary skill in the art would have been motivated to make this modification in order to keep the content shown on the display matched to the audio then being output, as Story expressly teaches (Story, col. 10, ll. 45-57).
Claims 4 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Goel in view of Oh as applied to claims 1 and 10 above, and further in view of Kalisky et al. (U.S. 10649726; filed August 16, 2017; issued May 12, 2020), hereinafter Kalisky.
Kalisky was published on May 12, 2020, before the November 20, 2023 effective filing date of the claimed invention, and qualifies as prior art under 35 U.S.C. 102(a)(1).
Regarding claim 4, Goel in view of Oh teaches the claim from which it depends as set forth above. The combination of Goel and Oh does not disclose the following, which Kalisky teaches:
4. The method of claim 1, wherein the displaying of the symbols of the symbol group on the display is timed based on a position of a character group preceding the symbol group and a number of characters in the character group preceding the symbol group. (Kalisky computes display parameters comprising the position of the next text unit within the text portion and the point in time at which that unit is to be displayed, that point in time being obtained by multiplying an average per-character playback time by the number of characters: "calculating the entire text reading time 801 is done by multiplying the average time of reading a single character multiplied by the total number of characters of the entire text" (Kalisky, claim 1(c)(1)-(2); see also col. 10, step 801, and col. 11, step 1005 with Fig. 10). The display time of a unit therefore follows from its position in the text and from the character count of the text preceding it.)
Goel, Oh, and Kalisky pertain to displaying written content in time with the audio playback of that content. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify the combination such that the time at which the symbol group is displayed is computed from the position of the preceding character group and the number of characters in it, in the manner of Kalisky. One of ordinary skill in the art would have been motivated to make this modification in order to place the display of each unit at the point in the playback at which the preceding text has been spoken, as Kalisky expressly teaches (Kalisky, claim 1(c)(1)-(2)).
Regarding claim 13, Goel in view of Oh teaches the claim from which it depends as set forth above. The combination of Goel and Oh does not disclose the following, which Kalisky teaches:
13. The electronic apparatus of claim 10, wherein the displaying of the symbols of the symbol group on the display is timed based on a position of a character group preceding the symbol group and a number of characters in the character group preceding the symbol group. (Kalisky computes display parameters comprising the position of the next text unit within the text portion and the point in time at which that unit is to be displayed, that point in time being obtained by multiplying an average per-character playback time by the number of characters: "calculating the entire text reading time 801 is done by multiplying the average time of reading a single character multiplied by the total number of characters of the entire text" (Kalisky, claim 1(c)(1)-(2); see also col. 10, step 801, and col. 11, step 1005 with Fig. 10). The display time of a unit therefore follows from its position in the text and from the character count of the text preceding it.)
Goel, Oh, and Kalisky pertain to displaying written content in time with the audio playback of that content. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify the combination such that the time at which the symbol group is displayed is computed from the position of the preceding character group and the number of characters in it, in the manner of Kalisky. One of ordinary skill in the art would have been motivated to make this modification in order to place the display of each unit at the point in the playback at which the preceding text has been spoken, as Kalisky expressly teaches (Kalisky, claim 1(c)(1)-(2)).
Claims 5 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Goel in view of Oh as applied to claims 1 and 10 above, and further in view of Ebersman et al. (U.S. 20150222586, Facebook, Inc.; filed February 5, 2014; published August 6, 2015), hereinafter Ebersman.
Ebersman was published on August 6, 2015, before the November 20, 2023 effective filing date of the claimed invention, and qualifies as prior art under 35 U.S.C. 102(a)(1).
Regarding claim 5, Goel in view of Oh teaches the claim from which it depends as set forth above, and Goel further teaches:
5. The method of claim 1, further including: identifying, by the at least one processor, among the character groups, a character group including a number of characters above a preset threshold; (Goel counts the items of a group against a preset threshold and takes a different rendering action according to whether the count exceeds it: "the one or more attributes may comprise one or more of a threshold requirement for the one or more non-Latin script content items or a description difficulty associated with each of the one or more non-Latin script content items" (Goel, paras. [0127] and [0131]; see also paras. [0010], [0129], and [0133] and claims 6, 10, and 12).)
Goel does not teach “determining, by the at least one processor, a symbol corresponding to an emotional state of characters in the character group including the number of characters above the preset threshold; and replacing, by the at least one processor, the character group including the number of characters above the preset threshold with a symbol group including the symbol corresponding to the emotional state.”, which Ebersman teaches:
determining, by the at least one processor, a symbol corresponding to an emotional state of characters in the character group including the number of characters above the preset threshold; and replacing, by the at least one processor, the character group including the number of characters above the preset threshold with a symbol group including the symbol corresponding to the emotional state. (Ebersman conducts sentiment analysis on a portion of the text the author has input and supplies ideograms each corresponding to an identified sentiment: "the messaging platform may suggest certain ideograms to the author of a message, based at least partly on the text content that the author has input" (Ebersman, paras. [0058]-[0062] and [0069] with Fig. 5; see also para. [0071] with Fig. 6A), the selected ideogram being inserted in substitution for the analyzed text: "the author will select one or more of the ideograms to be inserted into the message at step 590, substituting the analyzed text" (Ebersman, para. [0006]; see also para. [0068] and claim 1). The sentiment identified for the text portion is the emotional state of the characters of the group, and the ideogram substituted for that portion is the symbol group replacing it.)
Goel, Oh, and Ebersman pertain to the rendering of written messages and the symbols carried in them. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify the combination such that a character group whose character count exceeds the threshold of Goel is subjected to the sentiment analysis of Ebersman and replaced with the ideogram corresponding to the identified sentiment. One of ordinary skill in the art would have been motivated to make this modification in order to shorten the readout of a long character group while preserving the sentiment it carries, as Ebersman expressly teaches that the suggested ideogram may substitute for the analyzed text (Ebersman, para. [0006]).
Regarding claim 14, Goel in view of Oh teaches the claim from which it depends as set forth above, and Goel further teaches:
14. The electronic apparatus of claim 10, wherein the at least one processor is configured to further perform: identifying, among the character groups, a character group including a number of characters above a preset threshold; (Goel counts the items of a group against a preset threshold and takes a different rendering action according to whether the count exceeds it: "the one or more attributes may comprise one or more of a threshold requirement for the one or more non-Latin script content items or a description difficulty associated with each of the one or more non-Latin script content items" (Goel, paras. [0127] and [0131]; see also paras. [0010], [0129], and [0133] and claims 6, 10, and 12).)
Goel does not teach “determining a symbol corresponding to an emotional state of characters in the character group including the number of characters above the preset threshold; and replacing the character group including the number of characters above the preset threshold with a symbol group including the symbol corresponding to the emotional state.”, which Ebersman teaches:
determining a symbol corresponding to an emotional state of characters in the character group including the number of characters above the preset threshold; and replacing the character group including the number of characters above the preset threshold with a symbol group including the symbol corresponding to the emotional state. (Ebersman conducts sentiment analysis on a portion of the text the author has input and supplies ideograms each corresponding to an identified sentiment: "the messaging platform may suggest certain ideograms to the author of a message, based at least partly on the text content that the author has input" (Ebersman, paras. [0058]-[0062] and [0069] with Fig. 5; see also para. [0071] with Fig. 6A), the selected ideogram being inserted in substitution for the analyzed text: "the author will select one or more of the ideograms to be inserted into the message at step 590, substituting the analyzed text" (Ebersman, para. [0006]; see also para. [0068] and claim 1). The sentiment identified for the text portion is the emotional state of the characters of the group, and the ideogram substituted for that portion is the symbol group replacing it.)
Goel, Oh, and Ebersman pertain to the rendering of written messages and the symbols carried in them. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify the combination such that a character group whose character count exceeds the threshold of Goel is subjected to the sentiment analysis of Ebersman and replaced with the ideogram corresponding to the identified sentiment. One of ordinary skill in the art would have been motivated to make this modification in order to shorten the readout of a long character group while preserving the sentiment it carries, as Ebersman expressly teaches that the suggested ideogram may substitute for the analyzed text (Ebersman, para. [0006]).
Claims 9 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Goel in view of Oh as applied to claims 1 and 10 above, and further in view of Blanco et al. (U.S. 20070194902, Microsoft Corporation; filed February 17, 2006; published August 23, 2007), hereinafter Blanco.
Blanco was published on August 23, 2007, before the November 20, 2023 effective filing date of the claimed invention, and qualifies as prior art under 35 U.S.C. 102(a)(1).
Regarding claim 9, Goel in view of Oh teaches the claim from which it depends as set forth above. The combination of Goel and Oh does not disclose the following, which Blanco teaches:
9. The method of claim 1, wherein the display includes a heads-up display of a vehicle provided with the electronic device. (Blanco presents its adaptive user interface on the heads-up display 202 of an automobile: "The automobile computing system 200 provides an adaptive user interface that may be displayed on a HUD 202 of an automobile" (Blanco, para. [0024]; see also Fig. 2 and claims 1, 10, and 17). The automobile computing system on which that interface runs is itself provided with a text-to-speech capability that presents information audibly through the automobile's speakers: "a text-to-speech capability may be provided within the automobile computing system 200 to allow audible presentation of information, for instance, via the automobile's speakers 210" (Blanco, para. [0031]; see also para. [0029] with Fig. 2 for speakers 218, and para. [0052] with Fig. 3K).)
Goel, Oh, and Blanco pertain to presenting message and document content to a user both visually and as synthesized speech. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify the combination such that the display on which the symbols are shown is the heads-up display of the automobile of Blanco, the sound being output through the automobile's speakers. One of ordinary skill in the art would have been motivated to make this modification in order to present the content to a driver without requiring the driver to look away from the road, as Blanco expressly teaches (Blanco, paras. [0024] and [0031]).
Regarding claim 18, Goel in view of Oh teaches the claim from which it depends as set forth above. The combination of Goel and Oh does not disclose the following, which Blanco teaches:
18. The electronic apparatus of claim 10, wherein the display includes a heads-up display of a vehicle provided with the electronic apparatus. (Blanco presents its adaptive user interface on the heads-up display 202 of an automobile: "The automobile computing system 200 provides an adaptive user interface that may be displayed on a HUD 202 of an automobile" (Blanco, para. [0024]; see also Fig. 2 and claims 1, 10, and 17). The automobile computing system on which that interface runs is itself provided with a text-to-speech capability that presents information audibly through the automobile's speakers: "a text-to-speech capability may be provided within the automobile computing system 200 to allow audible presentation of information, for instance, via the automobile's speakers 210" (Blanco, para. [0031]; see also para. [0029] with Fig. 2 for speakers 218, and para. [0052] with Fig. 3K).)
Goel, Oh, and Blanco pertain to presenting message and document content to a user both visually and as synthesized speech. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify the combination such that the display on which the symbols are shown is the heads-up display of the automobile of Blanco, the sound being output through the automobile's speakers. One of ordinary skill in the art would have been motivated to make this modification in order to present the content to a driver without requiring the driver to look away from the road, as Blanco expressly teaches (Blanco, paras. [0024] and [0031]).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
a. Li et al. (U.S. 8898066): multi-lingual text-to-speech synthesis in which a first and a second acoustic-prosodic model are selected and merged under a controllable accent weighting parameter, with a phonetic unit transformation table, to synthesize speech carrying a controlled degree of accent.
b. Gu et al. (U.S. 10685186): emoji input method that performs word segmentation on input text, classifies the text with an emotion classification model to determine an emotion label, and obtains, from each of several themes, an emoji corresponding to that label.
c. Ricci (U.S. 20200067786): dynamic reconfiguration of a vehicle heads-up display, in which a configuration area of a graphical user interface presents a visual representation of a virtual dash display and gesture input rearranges the dash information, readouts, instruments, indicators, or controls shown on the unit.
d. Ostermann et al. (U.S. 9536544): multi-media message delivered by an animated entity whose voice is generated by text-to-speech conversion, in which emoticons in the message text are translated into corresponding facial expressions, the position of an emoticon in the text determining when the expression is executed during delivery.
e. Radebaugh (U.S. 9767789): identification of emoticons within a source text and tagging of the text with corresponding mood indicators, so that the emotional characteristics of the synthesized speech follow the emoticons found in the text.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to YUVAL H. LEVENTAL whose telephone number is (571) 270-3130. The examiner can normally be reached Monday-Friday, 8:00 AM - 5:00 PM.
If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, PIERRE-LOUIS DESIR, can be reached at (571) 272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/YUVAL HAIM LEVENTAL/
Examiner, Art Unit 2659
/PIERRE LOUIS DESIR/Supervisory Patent Examiner, Art Unit 2659