DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on February 17th, 2026 has been entered.
Response to Amendment
The amendment filed February 17th, 2026 has been entered. Claims 1, 4, 5, 14, 15, 17, 18, and 28-30 have been amended. Claims 31-41 have been added. Claims 1-41 are pending and have been examined.
Response to Arguments
Applicant’s arguments with respect to claim(s) 1-30 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-3, 14-18, 29-31, 33-36, and 38-41 is/are rejected under 35 U.S.C. 103 as being unpatentable over Blumenfeld et al. (WO Pub. No. 2011/011224 A1 hereinafter Blumenfeld), in view of Bowers et al. (US Pat. Pub. No. 2023/0046658 A1 hereinafter Bowers).
Regarding claim 1, Blumenfeld discloses an electronic device, comprising: one or more processors (Blumenfeld, [0013]: "Additional features that may be included in exemplary SGD embodiments concern computer readable medium that, when executed by a processor, instruct the SGD to perform specific functionality."); a memory (Blumenfeld, [0013]: "instructions stored in SGD memory may adapt the device to perform integrated web access for vocabulary generation in which content is selected from the Internet, and useful terms, phrases and/or images are extracted and available for immediate use in aided communication via the SGD."); and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for: receiving a first set of one or more user inputs that includes an input requesting a text-to-speech service (Blumenfeld, [0072]: "In accordance with general functionality of a speech generation device, a user provides text, symbols corresponding to text, and/or related or additional information in a "Message Window" which may then be interpreted by a text-to-speech engine and provided as audio output via the speakers 714."); receiving, via the text-to-speech user interface, first text; in response to receiving the first text, displaying the first text in the text-to-speech user interface (Blumenfeld, [0072]: "In accordance with general functionality of a speech generation device, a user provides text, symbols corresponding to text, and/or related or additional information in a "Message Window" which may then be interpreted by a text-to-speech engine and provided as audio output via the speakers 714."); detecting a second user input requesting output of the first text (Blumenfeld, Fig. 21; [00114]: "After input defining one or more speech styles and/or speech effects is received for a given portion of text (e.g., text displayed in a message window), the marked-up text is relayed to an interpretation engine such that the marked text can be analyzed by software rules in step 2014 to determine an optimum speech-to-text engine for speaking the marked-up selected text and associated speech styles/effects."); and in response to detecting the second user input: generating, using a personalized voice model associated with a first user, a first synthesized speech output of the first text; and initiating output of the first synthesized speech output (Blumenfeld, [0014]: "Another feature that may be provided in exemplary SGD embodiments corresponds to a voice inflection mark up feature that can add inflection or other speech style to the speech output of the SGD or that can add a predefined speech effect. Exemplary speech styles for modifying text may be defined in accordance with a variety of parameters, including but not limited to voice, volume, pitch, speech rate, spell mode, emphasis and context. Exemplary speech effects that may be inserted into a communicated message may include but are not limited to coughing, laughing, crying, kissing, sniffing, swallowing, sighing, whistling, yawning, or making noises such as "eh", "oh", "mmm", etc. The selected speech styles and effects can help communicate different meanings just based on the different inflections imposed on different syllables of the words being spoken by the SGD. In addition, such speech modifiers provide more interactive and enjoyable communication by an SGD user."). However, Blumenfeld fails to expressly recite in response to receiving the first set of one or more user inputs, displaying a text-to- speech user interface including with a candidate text library interface, wherein displaying the text- to-speech user interface with the candidate text library interface includes concurrently displaying: a plurality of phrase affordances respectively corresponding to a plurality of candidate phrases, wherein the plurality of candidate phrases is associated with a respective phrase category; and a plurality of phrase category affordances that includes a respective phrase category affordance corresponding to the respective phrase category; and generating, using a personalized voice model associated with a first user, a first synthesized speech output of the first text that simulates a speech characteristic of the first user.
Bowers teaches in response to receiving the first set of one or more user inputs, displaying a text-to- speech user interface including with a candidate text library interface (Bowers, Fig. 4B; [0094]: “In some implementations, a prosodic properties user interface 482 can optionally be visually rendered via the user interface 480 of the client device 410 (e.g., as indicated in FIGS. 4B-4D as dashed lines) when the additional user spoken input is detected at the client device 410. In other implementations, the prosodic properties user interface 482 can be visually rendered via the user interface 480 of the client device 410 in response to detecting user interface input (e.g., touch input or spoken input) directed to the graphical elements 462R and 462J, respectively in FIGS. 4B and 4C.”), wherein displaying the text- to-speech user interface with the candidate text library interface includes concurrently displaying: a plurality of phrase affordances respectively corresponding to a plurality of candidate phrases, wherein the plurality of candidate phrases is associated with a respective phrase category (Bowers, Fig. 4B, 456B1-456B4; [0096]: “In some implementations, the client device 410 can generate one or more suggestions 456B1-456B4 that each include a corresponding candidate textual segment that is responsive to the further additional user spoken input from Randy. The one or more suggestions 456B1-456B4 can be visually rendered on the client device 410 via the textual segment user interface 481. The client device 410 can generate the one or more suggestions using an automatic suggestion engine as described in more detail herein (e.g., with respect to automatic suggestion engine 150 of FIG. 1).); and a plurality of phrase category affordances that includes a respective phrase category affordance corresponding to the respective phrase category (Bowers, Fig. 4B, 482; [0094]: “The prosodic properties user interface 482 can include, for example, a scale 442 having an indicator 444 that indicates how “informal,” “neutral,” or “formal” any synthesized speech generated by the client device 410 is with respect to the additional participant… In some versions of those implementations, the indicator 444 may be slidable along the scale 442 in response to user interface input directed to the indicator 444. When the indicator 444 is adjusted along the scale 442, at least one prosodic property in the set of prosodic properties is adjusted to reflect the adjustment. For example, if the indicator 444 is adjusted on the scale 442 to reflect more informal speech, then any synthesized speech subsequently generated by the client device 410 that is directed to Randy include a more excited tone (as compared to a monotonous tone)”; Here, the slidable scale is seen as a plurality of phrase category affordances for the phrase categories formal, neutral, and informal.); and generating, using a personalized voice model associated with a first user, a first synthesized speech output of the first text that simulates a speech characteristic of the first user (Bowers, [0088]: “By establishing Tim's speaker embedding, the client device 410 can model Tim's voice by generating synthesized speech that represents Tim's speech based on Tim's speaker embedding.”; [0096]: “Upon receiving further user interface input at the client device 410 directed to one of suggestions 456B1-456B4, the client device 410 can generate synthesized speech, using Tim's speech embedding established in FIG. 4A, that includes the candidate textual segment of the selected suggestion, and that is synthesized with a set of prosodic properties.”).
Blumenfeld and Bowers are analogous arts because they both belong to the same field of text and audio processing. It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to have modified the speech generation device of Blumenfeld to incorporate the teachings of Bowers to concurrently display phrase categories with corresponding phrases that can be output using a voice model based on the user’s voice. These features can improve conversations for speech-impaired users without significant increases to system resources (Bowers, [0005]). As such, the system provides a better user experience for a variety of users.
Regarding claim 2, the rejection of claim 1 is incorporated. Blumenfeld, in view of Bowers, discloses all of the elements of the current invention as stated above. Bowers further teaches the one or more programs further including instructions for: training the personalized voice model at least in part on audio inputs received from the first user (Bowers, [0086]: “Referring initially to FIG. 4A, a speaker embedding for the given user 301 of the client device 410 can be established. In some implementations, the speaker embedding for the given user 301 can be generated as described herein (e.g., with respect to identification engine 112 of FIG. 1), and can be utilized in generating synthesized speech that represents speech of the given user 301. In some versions of those implementations, an automated assistant executing at least in part on the client device 410 can cause a plurality of prompts to be rendered audibly via speaker(s) of the client device 410 and/or of an additional computing device (e.g., computing device 310A of FIG. 3) and/or visually via the user interface 480 of the client device 410. The plurality of prompts can be part of a conversation between the given user 301 and the automated assistant, and each of the prompts can solicit spoken input from the given user 301 of the client device 410. Further, the spoken input from the given user 301 that is responsive to each of the prompts can be detected via one or more microphones of the client device 410… the automated assistant can continuously cause the prompts to be rendered at the client device 410 until the speaker embedding for the given user 301 of the client device 410 is established.”). The same motivation for claim 1 applies equally to claim 2.
Regarding claim 3, the rejection of claim 1 is incorporated. Blumenfeld, in view of Bowers, discloses all of the elements of the current invention as stated above. Bowers further teaches wherein the personalized voice model is trained to generate synthesized speech outputs that simulate speech characteristics of the first user (Bowers, [0088]: “By establishing Tim's speaker embedding, the client device 410 can model Tim's voice by generating synthesized speech that represents Tim's speech based on Tim's speaker embedding.”). The same motivation for claim 1 applies equally to claim 3.
Regarding claim 14, the rejection of claim 1 is incorporated. Blumenfeld, in view of Bowers, discloses all of the elements of the current invention as stated above. Blumenfeld further discloses wherein the text-to-speech user interface includes a candidate text library affordance, and wherein the plurality of phrase affordances and the plurality of phrase category affordances are displayed in response to detecting a selection of the candidate text library affordance (Blumenfeld, [00101]: "software instructions may be configured to display an embedded vocabulary list in a similar fashion (i.e., as an icon) as any other item in a given vocabulary list box. The example in Fig. 17 shows a user interface displaying a vocabulary list box with the vocabulary list "Animals" having the "Animal Sounds" vocabulary list embedded within it. When an embedded vocabulary list is selected by a user, its contents will then be displayed in the vocabulary list box. All the words in the original vocabulary list "Animals" (e.g., aardvark, alligator, animal, ant, anteater, antelope, ape, armadillo, etc.) as well as the words in the embedded vocabulary list "Animal Sounds" (e.g., bark, bleat, buzz, chirp, cluck, croak, gobble, growl, hiss, etc.) are then available for a user to use in a message.").
Regarding claim 15, the rejection of claim 1 is incorporated. Blumenfeld, in view of Bowers, discloses all of the elements of the current invention as stated above. Blumenfeld further discloses the one or more programs further including instructions for: detecting a selection of a first phrase category affordance of the plurality of phrase category affordances, wherein the first phrase category affordance corresponds to a first phrase category; and in response to detecting a selection of the first phrase category affordance: displaying at least a first phrase affordance of the plurality of phrase affordances, wherein the candidate phrase corresponding to the first phrase affordance is associated with the first phrase category; and foregoing displaying at least a second phrase affordance of the plurality of phrase affordances, wherein the candidate phrase corresponding to the second phrase affordance is not associated with the first phrase category (Blumenfeld, [00101]: "software instructions may be configured to display an embedded vocabulary list in a similar fashion (i.e., as an icon) as any other item in a given vocabulary list box. The example in Fig. 17 shows a user interface displaying a vocabulary list box with the vocabulary list "Animals" having the "Animal Sounds" vocabulary list embedded within it. When an embedded vocabulary list is selected by a user, its contents will then be displayed in the vocabulary list box. All the words in the original vocabulary list "Animals" (e.g., aardvark, alligator, animal, ant, anteater, antelope, ape, armadillo, etc.) as well as the words in the embedded vocabulary list "Animal Sounds" (e.g., bark, bleat, buzz, chirp, cluck, croak, gobble, growl, hiss, etc.) are then available for a user to use in a message.").
Regarding claim 16, the rejection of claim 1 is incorporated. Blumenfeld, in view of Bowers, discloses all of the elements of the current invention as stated above. Blumenfeld further discloses the one or more programs further including instructions for: displaying a category addition affordance; detecting a selection of the category addition affordance; and in response to detecting the selection of the category addition affordance: receiving a fifth user input indicating a new phrase category; receiving a sixth user input indicating at least one candidate phrase; and associating the at least one candidate phrase with the new phrase category (Blumenfeld, [00102]: "suppose that a user adds twelve words to an SGD dictionary and marks them all with the "Clothing" category. The next time a user opens a page with a vocabulary list containing a vocabulary search on "Clothing", those twelve new words will automatically appear in the vocabulary list box. By inserting a search itself to a vocabulary list, a search can be performed dynamically. As such, when other items are added to a dictionary that matches a user's search criteria, a user does not have to go back and separately add the new words to different preexisting vocabulary lists. The results of an embedded vocabulary search can be intermixed with the other vocabulary items within the vocabulary list. They will appear side by side with the individual vocabulary items, and they can be sorted based on the settings of the vocabulary list box {alphabetically, as is, or by frequency of use).").
Regarding claim 17, the rejection of claim 1 is incorporated. Blumenfeld, in view of Bowers, discloses all of the elements of the current invention as stated above. Blumenfeld further discloses wherein the candidate text library interface includes a phrase addition affordance, and the one or more programs further including instructions for: detecting a selection of the phrase addition affordance (Blumenfeld, [0091]: "interface menus may be provided that enable a user to incorporate key words located on the Internet (i.e., "web words") into a user- customized vocabulary list."); and in response to detecting a selection of the phrase addition affordance: receiving second text; and in response to receiving the second text, adding a phrase affordance corresponding to the second text to the plurality of phrase affordances (Blumenfeld, [0091]: "Such identified web words may then be incorporated into a user's custom message by selecting such words from a vocabulary list.").
Regarding claim 18, the rejection of claim 1 is incorporated. Blumenfeld, in view of Bowers, discloses all of the elements of the current invention as stated above. Blumenfeld further discloses wherein receiving the first text includes detecting a selection of a third phrase affordance of the plurality of phrase affordances, wherein the third phrase affordance corresponds to at least a portion of the first text (Blumenfeld, [00106]: "The first slot is associated with the "breakfast" concept and the second slot is associated with the "fruit" concept. The slots allow a user to create dynamic messages with a reduced number of selections. By selecting the slots and changing the filler text, the example phrase "I want oatmeal and a banana for breakfast" can quickly and easily be changed to read as follows: "I want toast and a nectarine for breakfast".").
Regarding claim 29, Blumenfeld discloses a non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a first electronic device, cause the first electronic device to (Blumenfeld, [0013]: “Additional features that may be included in exemplary SGD embodiments concern computer readable medium that, when executed by a processor, instruct the SGD to perform specific functionality.”): receive a first user input requesting a text-to-speech service (Blumenfeld, [0072]: "In accordance with general functionality of a speech generation device, a user provides text, symbols corresponding to text, and/or related or additional information in a "Message Window" which may then be interpreted by a text-to-speech engine and provided as audio output via the speakers 714."); receive, via the text-to-speech user interface, first text; in response to receiving the first text, display the first text in the text-to-speech user interface (Blumenfeld, [0072]: "In accordance with general functionality of a speech generation device, a user provides text, symbols corresponding to text, and/or related or additional information in a "Message Window" which may then be interpreted by a text-to-speech engine and provided as audio output via the speakers 714."); detect a second user input requesting output of the first text (Blumenfeld, Fig. 21; [00114]: "After input defining one or more speech styles and/or speech effects is received for a given portion of text (e.g., text displayed in a message window), the marked-up text is relayed to an interpretation engine such that the marked text can be analyzed by software rules in step 2014 to determine an optimum speech-to-text engine for speaking the marked-up selected text and associated speech styles/effects."); and in response to detecting the second user input: generate, using a personalized voice model associated with a first user, a first synthesized speech output of the first text; and initiate output of the first synthesized speech output (Blumenfeld, [0014]: "Another feature that may be provided in exemplary SGD embodiments corresponds to a voice inflection mark up feature that can add inflection or other speech style to the speech output of the SGD or that can add a predefined speech effect. Exemplary speech styles for modifying text may be defined in accordance with a variety of parameters, including but not limited to voice, volume, pitch, speech rate, spell mode, emphasis and context. Exemplary speech effects that may be inserted into a communicated message may include but are not limited to coughing, laughing, crying, kissing, sniffing, swallowing, sighing, whistling, yawning, or making noises such as "eh", "oh", "mmm", etc. The selected speech styles and effects can help communicate different meanings just based on the different inflections imposed on different syllables of the words being spoken by the SGD. In addition, such speech modifiers provide more interactive and enjoyable communication by an SGD user."). However, Blumenfeld fails to expressly recite in response to receiving the first set of one or more user inputs, display a text-to- speech user interface including with a candidate text library interface, wherein displaying the text- to-speech user interface with the candidate text library interface includes concurrently displaying: a plurality of phrase affordances respectively corresponding to a plurality of candidate phrases, wherein the plurality of candidate phrases is associated with a respective phrase category; and a plurality of phrase category affordances that includes a respective phrase category affordance corresponding to the respective phrase category; and generate, using a personalized voice model associated with a first user, a first synthesized speech output of the first text that simulates a speech characteristic of the first user.
Bowers teaches in response to receiving the first set of one or more user inputs, display a text-to- speech user interface including with a candidate text library interface (Bowers, Fig. 4B; [0094]: “In some implementations, a prosodic properties user interface 482 can optionally be visually rendered via the user interface 480 of the client device 410 (e.g., as indicated in FIGS. 4B-4D as dashed lines) when the additional user spoken input is detected at the client device 410. In other implementations, the prosodic properties user interface 482 can be visually rendered via the user interface 480 of the client device 410 in response to detecting user interface input (e.g., touch input or spoken input) directed to the graphical elements 462R and 462J, respectively in FIGS. 4B and 4C.”), wherein displaying the text- to-speech user interface with the candidate text library interface includes concurrently displaying: a plurality of phrase affordances respectively corresponding to a plurality of candidate phrases, wherein the plurality of candidate phrases is associated with a respective phrase category (Bowers, Fig. 4B, 456B1-456B4; [0096]: “In some implementations, the client device 410 can generate one or more suggestions 456B1-456B4 that each include a corresponding candidate textual segment that is responsive to the further additional user spoken input from Randy. The one or more suggestions 456B1-456B4 can be visually rendered on the client device 410 via the textual segment user interface 481. The client device 410 can generate the one or more suggestions using an automatic suggestion engine as described in more detail herein (e.g., with respect to automatic suggestion engine 150 of FIG. 1).); and a plurality of phrase category affordances that includes a respective phrase category affordance corresponding to the respective phrase category (Bowers, Fig. 4B, 482; [0094]: “The prosodic properties user interface 482 can include, for example, a scale 442 having an indicator 444 that indicates how “informal,” “neutral,” or “formal” any synthesized speech generated by the client device 410 is with respect to the additional participant… In some versions of those implementations, the indicator 444 may be slidable along the scale 442 in response to user interface input directed to the indicator 444. When the indicator 444 is adjusted along the scale 442, at least one prosodic property in the set of prosodic properties is adjusted to reflect the adjustment. For example, if the indicator 444 is adjusted on the scale 442 to reflect more informal speech, then any synthesized speech subsequently generated by the client device 410 that is directed to Randy include a more excited tone (as compared to a monotonous tone)”; Here, the slidable scale is seen as a plurality of phrase category affordances for the phrase categories formal, neutral, and informal.); and generate, using a personalized voice model associated with a first user, a first synthesized speech output of the first text that simulates a speech characteristic of the first user (Bowers, [0088]: “By establishing Tim's speaker embedding, the client device 410 can model Tim's voice by generating synthesized speech that represents Tim's speech based on Tim's speaker embedding.”; [0096]: “Upon receiving further user interface input at the client device 410 directed to one of suggestions 456B1-456B4, the client device 410 can generate synthesized speech, using Tim's speech embedding established in FIG. 4A, that includes the candidate textual segment of the selected suggestion, and that is synthesized with a set of prosodic properties.”).
Blumenfeld and Bowers are analogous arts because they both belong to the same field of text and audio processing. It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to have modified the speech generation device of Blumenfeld to incorporate the teachings of Bowers to concurrently display phrase categories with corresponding phrases that can be output using a voice model based on the user’s voice. These features can improve conversations for speech-impaired users without significant increases to system resources (Bowers, [0005]). As such, the system provides a better user experience for a variety of users.
Regarding claim 30, Blumenfeld discloses a method, comprising: at an electronic device with a display (Blumenfeld, [0008]: “A variety of physical input devices and software interface features may be provided to facilitate the capture of user input to define what information should be displayed in a message window and ultimately communicated to others as spoken output, text message, phone call, e-mail or other outgoing”), one or more processors (Blumenfeld, [0013]: "Additional features that may be included in exemplary SGD embodiments concern computer readable medium that, when executed by a processor, instruct the SGD to perform specific functionality."), and memory (Blumenfeld, [0013]: "instructions stored in SGD memory may adapt the device to perform integrated web access for vocabulary generation in which content is selected from the Internet, and useful terms, phrases and/or images are extracted and available for immediate use in aided communication via the SGD."): receiving a first user input requesting a text-to-speech service (Blumenfeld, [0072]: "In accordance with general functionality of a speech generation device, a user provides text, symbols corresponding to text, and/or related or additional information in a "Message Window" which may then be interpreted by a text-to-speech engine and provided as audio output via the speakers 714."); receiving, via the text-to-speech user interface, first text; in response to receiving the first text, displaying the first text in the text-to-speech user interface (Blumenfeld, [0072]: "In accordance with general functionality of a speech generation device, a user provides text, symbols corresponding to text, and/or related or additional information in a "Message Window" which may then be interpreted by a text-to-speech engine and provided as audio output via the speakers 714."); detecting a second user input requesting output of the first text (Blumenfeld, Fig. 21; [00114]: "After input defining one or more speech styles and/or speech effects is received for a given portion of text (e.g., text displayed in a message window), the marked-up text is relayed to an interpretation engine such that the marked text can be analyzed by software rules in step 2014 to determine an optimum speech-to-text engine for speaking the marked-up selected text and associated speech styles/effects."); and in response to detecting the second user input: generating, using a personalized voice model associated with a first user, a first synthesized speech output of the first text; and initiating output of the first synthesized speech output (Blumenfeld, [0014]: "Another feature that may be provided in exemplary SGD embodiments corresponds to a voice inflection mark up feature that can add inflection or other speech style to the speech output of the SGD or that can add a predefined speech effect. Exemplary speech styles for modifying text may be defined in accordance with a variety of parameters, including but not limited to voice, volume, pitch, speech rate, spell mode, emphasis and context. Exemplary speech effects that may be inserted into a communicated message may include but are not limited to coughing, laughing, crying, kissing, sniffing, swallowing, sighing, whistling, yawning, or making noises such as "eh", "oh", "mmm", etc. The selected speech styles and effects can help communicate different meanings just based on the different inflections imposed on different syllables of the words being spoken by the SGD. In addition, such speech modifiers provide more interactive and enjoyable communication by an SGD user."). However, Blumenfeld fails to expressly recite in response to receiving the first set of one or more user inputs, displaying a text-to- speech user interface including with a candidate text library interface, wherein displaying the text- to-speech user interface with the candidate text library interface includes concurrently displaying: a plurality of phrase affordances respectively corresponding to a plurality of candidate phrases, wherein the plurality of candidate phrases is associated with a respective phrase category; and a plurality of phrase category affordances that includes a respective phrase category affordance corresponding to the respective phrase category; and generating, using a personalized voice model associated with a first user, a first synthesized speech output of the first text that simulates a speech characteristic of the first user.
Bowers teaches in response to receiving the first set of one or more user inputs, displaying a text-to- speech user interface including with a candidate text library interface (Bowers, Fig. 4B; [0094]: “In some implementations, a prosodic properties user interface 482 can optionally be visually rendered via the user interface 480 of the client device 410 (e.g., as indicated in FIGS. 4B-4D as dashed lines) when the additional user spoken input is detected at the client device 410. In other implementations, the prosodic properties user interface 482 can be visually rendered via the user interface 480 of the client device 410 in response to detecting user interface input (e.g., touch input or spoken input) directed to the graphical elements 462R and 462J, respectively in FIGS. 4B and 4C.”), wherein displaying the text- to-speech user interface with the candidate text library interface includes concurrently displaying: a plurality of phrase affordances respectively corresponding to a plurality of candidate phrases, wherein the plurality of candidate phrases is associated with a respective phrase category (Bowers, Fig. 4B, 456B1-456B4; [0096]: “In some implementations, the client device 410 can generate one or more suggestions 456B1-456B4 that each include a corresponding candidate textual segment that is responsive to the further additional user spoken input from Randy. The one or more suggestions 456B1-456B4 can be visually rendered on the client device 410 via the textual segment user interface 481. The client device 410 can generate the one or more suggestions using an automatic suggestion engine as described in more detail herein (e.g., with respect to automatic suggestion engine 150 of FIG. 1).); and a plurality of phrase category affordances that includes a respective phrase category affordance corresponding to the respective phrase category (Bowers, Fig. 4B, 482; [0094]: “The prosodic properties user interface 482 can include, for example, a scale 442 having an indicator 444 that indicates how “informal,” “neutral,” or “formal” any synthesized speech generated by the client device 410 is with respect to the additional participant… In some versions of those implementations, the indicator 444 may be slidable along the scale 442 in response to user interface input directed to the indicator 444. When the indicator 444 is adjusted along the scale 442, at least one prosodic property in the set of prosodic properties is adjusted to reflect the adjustment. For example, if the indicator 444 is adjusted on the scale 442 to reflect more informal speech, then any synthesized speech subsequently generated by the client device 410 that is directed to Randy include a more excited tone (as compared to a monotonous tone)”; Here, the slidable scale is seen as a plurality of phrase category affordances for the phrase categories formal, neutral, and informal.); and generating, using a personalized voice model associated with a first user, a first synthesized speech output of the first text that simulates a speech characteristic of the first user (Bowers, [0088]: “By establishing Tim's speaker embedding, the client device 410 can model Tim's voice by generating synthesized speech that represents Tim's speech based on Tim's speaker embedding.”; [0096]: “Upon receiving further user interface input at the client device 410 directed to one of suggestions 456B1-456B4, the client device 410 can generate synthesized speech, using Tim's speech embedding established in FIG. 4A, that includes the candidate textual segment of the selected suggestion, and that is synthesized with a set of prosodic properties.”).
Blumenfeld and Bowers are analogous arts because they both belong to the same field of text and audio processing. It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to have modified the speech generation device of Blumenfeld to incorporate the teachings of Bowers to concurrently display phrase categories with corresponding phrases that can be output using a voice model based on the user’s voice. These features can improve conversations for speech-impaired users without significant increases to system resources (Bowers, [0005]). As such, the system provides a better user experience for a variety of users.
Regarding claim 31, the rejection of claim 1 is incorporated. Blumenfeld, in view of Bowers, discloses all of the elements of the current invention as stated above. Bowers further teaches the one or more programs further including instructions for: while displaying the text-to-speech user interface with the candidate text library interface (Bowers, Fig. 4B; [0094]: “In some implementations, a prosodic properties user interface 482 can optionally be visually rendered via the user interface 480 of the client device 410 (e.g., as indicated in FIGS. 4B-4D as dashed lines) when the additional user spoken input is detected at the client device 410. In other implementations, the prosodic properties user interface 482 can be visually rendered via the user interface 480 of the client device 410 in response to detecting user interface input (e.g., touch input or spoken input) directed to the graphical elements 462R and 462J, respectively in FIGS. 4B and 4C.”; [0096]: “In some implementations, the client device 410 can generate one or more suggestions 456B1-456B4 that each include a corresponding candidate textual segment that is responsive to the further additional user spoken input from Randy. The one or more suggestions 456B1-456B4 can be visually rendered on the client device 410 via the textual segment user interface 481. The client device 410 can generate the one or more suggestions using an automatic suggestion engine as described in more detail herein (e.g., with respect to automatic suggestion engine 150 of FIG. 1).), receiving an input selecting a respective phrase affordance from the plurality of phrase affordances, wherein the respective phrase affordance corresponds to a respective candidate phrase; and in response to receiving the input selecting the respective phrase affordance from the plurality of phrase affordances: generating, using the personalized voice model associated with the first user, a second synthesized speech output of the respective candidate phrase that simulates the speech characteristic of the first user; and initiating output of the second synthesized speech output of the respective candidate phrase (Bowers, [0088]: “By establishing Tim's speaker embedding, the client device 410 can model Tim's voice by generating synthesized speech that represents Tim's speech based on Tim's speaker embedding.”; [0096]: “Upon receiving further user interface input at the client device 410 directed to one of suggestions 456B1-456B4, the client device 410 can generate synthesized speech, using Tim's speech embedding established in FIG. 4A, that includes the candidate textual segment of the selected suggestion, and that is synthesized with a set of prosodic properties.”). The same motivation for claim 1 applies equally to claim 31.
Regarding claim 33, the rejection of claim 29 is incorporated. Blumenfeld, in view of Bowers, discloses all of the elements of the current invention as stated above. Blumenfeld further discloses the one or more programs further including instructions for: detecting a selection of a first phrase category affordance of the plurality of phrase category affordances, wherein the first phrase category affordance corresponds to a first phrase category; and in response to detecting a selection of the first phrase category affordance: displaying at least a first phrase affordance of the plurality of phrase affordances, wherein the candidate phrase corresponding to the first phrase affordance is associated with the first phrase category; and foregoing displaying at least a second phrase affordance of the plurality of phrase affordances, wherein the candidate phrase corresponding to the second phrase affordance is not associated with the first phrase category (Blumenfeld, [00101]: "software instructions may be configured to display an embedded vocabulary list in a similar fashion (i.e., as an icon) as any other item in a given vocabulary list box. The example in Fig. 17 shows a user interface displaying a vocabulary list box with the vocabulary list "Animals" having the "Animal Sounds" vocabulary list embedded within it. When an embedded vocabulary list is selected by a user, its contents will then be displayed in the vocabulary list box. All the words in the original vocabulary list "Animals" (e.g., aardvark, alligator, animal, ant, anteater, antelope, ape, armadillo, etc.) as well as the words in the embedded vocabulary list "Animal Sounds" (e.g., bark, bleat, buzz, chirp, cluck, croak, gobble, growl, hiss, etc.) are then available for a user to use in a message.").
Regarding claim 34, the rejection of claim 29 is incorporated. Blumenfeld, in view of Bowers, discloses all of the elements of the current invention as stated above. Blumenfeld further discloses the one or more programs further including instructions for: displaying a category addition affordance; detecting a selection of the category addition affordance; and in response to detecting the selection of the category addition affordance: receiving a fifth user input indicating a new phrase category; receiving a sixth user input indicating at least one candidate phrase; and associating the at least one candidate phrase with the new phrase category (Blumenfeld, [00102]: "suppose that a user adds twelve words to an SGD dictionary and marks them all with the "Clothing" category. The next time a user opens a page with a vocabulary list containing a vocabulary search on "Clothing", those twelve new words will automatically appear in the vocabulary list box. By inserting a search itself to a vocabulary list, a search can be performed dynamically. As such, when other items are added to a dictionary that matches a user's search criteria, a user does not have to go back and separately add the new words to different preexisting vocabulary lists. The results of an embedded vocabulary search can be intermixed with the other vocabulary items within the vocabulary list. They will appear side by side with the individual vocabulary items, and they can be sorted based on the settings of the vocabulary list box {alphabetically, as is, or by frequency of use).").
Regarding claim 35, the rejection of claim 29 is incorporated. Blumenfeld, in view of Bowers, discloses all of the elements of the current invention as stated above. Blumenfeld further discloses wherein the candidate text library interface includes a phrase addition affordance, and the one or more programs further including instructions for: detecting a selection of the phrase addition affordance (Blumenfeld, [0091]: "interface menus may be provided that enable a user to incorporate key words located on the Internet (i.e., "web words") into a user- customized vocabulary list."); and in response to detecting a selection of the phrase addition affordance: receiving second text; and in response to receiving the second text, adding a phrase affordance corresponding to the second text to the plurality of phrase affordances (Blumenfeld, [0091]: "Such identified web words may then be incorporated into a user's custom message by selecting such words from a vocabulary list.").
Regarding claim 36, the rejection of claim 29 is incorporated. Blumenfeld, in view of Bowers, discloses all of the elements of the current invention as stated above. Bowers further teaches the one or more programs further including instructions for: while displaying the text-to-speech user interface with the candidate text library interface (Bowers, Fig. 4B; [0094]: “In some implementations, a prosodic properties user interface 482 can optionally be visually rendered via the user interface 480 of the client device 410 (e.g., as indicated in FIGS. 4B-4D as dashed lines) when the additional user spoken input is detected at the client device 410. In other implementations, the prosodic properties user interface 482 can be visually rendered via the user interface 480 of the client device 410 in response to detecting user interface input (e.g., touch input or spoken input) directed to the graphical elements 462R and 462J, respectively in FIGS. 4B and 4C.”; [0096]: “In some implementations, the client device 410 can generate one or more suggestions 456B1-456B4 that each include a corresponding candidate textual segment that is responsive to the further additional user spoken input from Randy. The one or more suggestions 456B1-456B4 can be visually rendered on the client device 410 via the textual segment user interface 481. The client device 410 can generate the one or more suggestions using an automatic suggestion engine as described in more detail herein (e.g., with respect to automatic suggestion engine 150 of FIG. 1).), receiving an input selecting a respective phrase affordance from the plurality of phrase affordances, wherein the respective phrase affordance corresponds to a respective candidate phrase; and in response to receiving the input selecting the respective phrase affordance from the plurality of phrase affordances: generating, using the personalized voice model associated with the first user, a second synthesized speech output of the respective candidate phrase that simulates the speech characteristic of the first user; and initiating output of the second synthesized speech output of the respective candidate phrase (Bowers, [0088]: “By establishing Tim's speaker embedding, the client device 410 can model Tim's voice by generating synthesized speech that represents Tim's speech based on Tim's speaker embedding.”; [0096]: “Upon receiving further user interface input at the client device 410 directed to one of suggestions 456B1-456B4, the client device 410 can generate synthesized speech, using Tim's speech embedding established in FIG. 4A, that includes the candidate textual segment of the selected suggestion, and that is synthesized with a set of prosodic properties.”). The same motivation for claim 1 applies equally to claim 31.
Regarding claim 38, the rejection of claim 30 is incorporated. Blumenfeld, in view of Bowers, discloses all of the elements of the current invention as stated above. Blumenfeld further discloses detecting a selection of a first phrase category affordance of the plurality of phrase category affordances, wherein the first phrase category affordance corresponds to a first phrase category; and in response to detecting a selection of the first phrase category affordance: displaying at least a first phrase affordance of the plurality of phrase affordances, wherein the candidate phrase corresponding to the first phrase affordance is associated with the first phrase category; and foregoing displaying at least a second phrase affordance of the plurality of phrase affordances, wherein the candidate phrase corresponding to the second phrase affordance is not associated with the first phrase category (Blumenfeld, [00101]: "software instructions may be configured to display an embedded vocabulary list in a similar fashion (i.e., as an icon) as any other item in a given vocabulary list box. The example in Fig. 17 shows a user interface displaying a vocabulary list box with the vocabulary list "Animals" having the "Animal Sounds" vocabulary list embedded within it. When an embedded vocabulary list is selected by a user, its contents will then be displayed in the vocabulary list box. All the words in the original vocabulary list "Animals" (e.g., aardvark, alligator, animal, ant, anteater, antelope, ape, armadillo, etc.) as well as the words in the embedded vocabulary list "Animal Sounds" (e.g., bark, bleat, buzz, chirp, cluck, croak, gobble, growl, hiss, etc.) are then available for a user to use in a message.").
Regarding claim 39, the rejection of claim 30 is incorporated. Blumenfeld, in view of Bowers, discloses all of the elements of the current invention as stated above. Blumenfeld further discloses displaying a category addition affordance; detecting a selection of the category addition affordance; and in response to detecting the selection of the category addition affordance: receiving a fifth user input indicating a new phrase category; receiving a sixth user input indicating at least one candidate phrase; and associating the at least one candidate phrase with the new phrase category (Blumenfeld, [00102]: "suppose that a user adds twelve words to an SGD dictionary and marks them all with the "Clothing" category. The next time a user opens a page with a vocabulary list containing a vocabulary search on "Clothing", those twelve new words will automatically appear in the vocabulary list box. By inserting a search itself to a vocabulary list, a search can be performed dynamically. As such, when other items are added to a dictionary that matches a user's search criteria, a user does not have to go back and separately add the new words to different preexisting vocabulary lists. The results of an embedded vocabulary search can be intermixed with the other vocabulary items within the vocabulary list. They will appear side by side with the individual vocabulary items, and they can be sorted based on the settings of the vocabulary list box {alphabetically, as is, or by frequency of use).").
Regarding claim 40, the rejection of claim 30 is incorporated. Blumenfeld, in view of Bowers, discloses all of the elements of the current invention as stated above. Blumenfeld further discloses wherein the candidate text library interface includes a phrase addition affordance, and the method further comprises: detecting a selection of the phrase addition affordance (Blumenfeld, [0091]: "interface menus may be provided that enable a user to incorporate key words located on the Internet (i.e., "web words") into a user- customized vocabulary list."); and in response to detecting a selection of the phrase addition affordance: receiving second text; and in response to receiving the second text, adding a phrase affordance corresponding to the second text to the plurality of phrase affordances (Blumenfeld, [0091]: "Such identified web words may then be incorporated into a user's custom message by selecting such words from a vocabulary list.").
Regarding claim 41, the rejection of claim 30 is incorporated. Blumenfeld, in view of Bowers, discloses all of the elements of the current invention as stated above. Bowers further teaches while displaying the text-to-speech user interface with the candidate text library interface (Bowers, Fig. 4B; [0094]: “In some implementations, a prosodic properties user interface 482 can optionally be visually rendered via the user interface 480 of the client device 410 (e.g., as indicated in FIGS. 4B-4D as dashed lines) when the additional user spoken input is detected at the client device 410. In other implementations, the prosodic properties user interface 482 can be visually rendered via the user interface 480 of the client device 410 in response to detecting user interface input (e.g., touch input or spoken input) directed to the graphical elements 462R and 462J, respectively in FIGS. 4B and 4C.”; [0096]: “In some implementations, the client device 410 can generate one or more suggestions 456B1-456B4 that each include a corresponding candidate textual segment that is responsive to the further additional user spoken input from Randy. The one or more suggestions 456B1-456B4 can be visually rendered on the client device 410 via the textual segment user interface 481. The client device 410 can generate the one or more suggestions using an automatic suggestion engine as described in more detail herein (e.g., with respect to automatic suggestion engine 150 of FIG. 1).), receiving an input selecting a respective phrase affordance from the plurality of phrase affordances, wherein the respective phrase affordance corresponds to a respective candidate phrase; and in response to receiving the input selecting the respective phrase affordance from the plurality of phrase affordances: generating, using the personalized voice model associated with the first user, a second synthesized speech output of the respective candidate phrase that simulates the speech characteristic of the first user; and initiating output of the second synthesized speech output of the respective candidate phrase (Bowers, [0088]: “By establishing Tim's speaker embedding, the client device 410 can model Tim's voice by generating synthesized speech that represents Tim's speech based on Tim's speaker embedding.”; [0096]: “Upon receiving further user interface input at the client device 410 directed to one of suggestions 456B1-456B4, the client device 410 can generate synthesized speech, using Tim's speech embedding established in FIG. 4A, that includes the candidate textual segment of the selected suggestion, and that is synthesized with a set of prosodic properties.”). The same motivation for claim 1 applies equally to claim 31.
Claim(s) 4, 5, 24, 32, and 37 is/are rejected under 35 U.S.C. 103 as being unpatentable over Blumenfeld, in view of Bowers, as applied to claims 1-3, 14-18, 29, and 30 above, and further in view of Roth et al. (US Pat. Pub. No. 2005/0038657 A1 hereinafter Roth).
Regarding claim 4, the rejection of claim 1 is incorporated. Blumenfeld, in view of Bowers, discloses all of the elements of the current invention as stated above. However, Blumenfeld, in view of Bowers, fails to expressly recite wherein the input requesting a text-to-speech service includes a hardware button input.
Roth teaches wherein the input requesting a text-to-speech service includes a hardware button input (Roth, [0503]: "If the user presses the "4" key when in the edit options menu, a text-to-speech or TTS menu will be displayed. In this menu, the "4" key toggles TTS play on or off.").
Blumenfeld, Bowers, and Roth are analogous arts because they each belong to the same field of text and audio processing. It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to have modified the speech generation device of Blumenfeld, as modified by the synthesized speech audio generator of Bowers, to incorporate the teachings of Roth to include specific hardware button inputs. This allows for the use of physical buttons to navigate menus and perform functions (Roth, [0502]). Including physical buttons in the system allows for more various inputs that can be used by people who may not be able to or may not want to use other input methods.
Regarding claim 5, the rejection of claim 1 is incorporated. Blumenfeld, in view of Bowers, discloses all of the elements of the current invention as stated above. However, Blumenfeld, in view of Bowers, fails to expressly recite wherein the input requesting a text-to-speech service includes an input selecting a text-to-speech affordance.
Roth teaches wherein the input requesting a text-to-speech service includes an input selecting a text-to-speech affordance (Roth, [0503]: "If the user presses the "4" key when in the edit options menu, a text-to-speech or TTS menu will be displayed. In this menu, the "4" key toggles TTS play on or off."). The same motivation for claim 4 applies equally to claim 5.
Regarding claim 24, the rejection of claim 1 is incorporated. Blumenfeld, in view of Bowers, discloses all of the elements of the current invention as stated above. However, Blumenfeld, in view of Bowers, fails to expressly recite wherein the second user input requesting the output of the first text includes a hardware button press.
Roth teaches wherein the second user input requesting the output of the first text includes a hardware button press (Roth, [0504]: "The TTS submenu also includes a choice, selected by pressing the "5" key, that allows the user to play the current selection whenever he or she desires to do so, as indicated by functions 8924 and 8926."). The same motivation for claim 4 applies equally to claim 24.
Regarding claim 32, the rejection of claim 29 is incorporated. Blumenfeld, in view of Bowers, discloses all of the elements of the current invention as stated above. However, Blumenfeld, in view of Bowers, fails to expressly recite wherein the input requesting a text-to-speech service includes a hardware button input.
Roth teaches wherein the input requesting a text-to-speech service includes a hardware button input (Roth, [0503]: "If the user presses the "4" key when in the edit options menu, a text-to-speech or TTS menu will be displayed. In this menu, the "4" key toggles TTS play on or off.").
Blumenfeld, Bowers, and Roth are analogous arts because they each belong to the same field of text and audio processing. It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to have modified the speech generation device of Blumenfeld, as modified by the synthesized speech audio generator of Bowers, to incorporate the teachings of Roth to include specific hardware button inputs. This allows for the use of physical buttons to navigate menus and perform functions (Roth, [0502]). Including physical buttons in the system allows for more various inputs that can be used by people who may not be able to or may not want to use other input methods.
Regarding claim 37, the rejection of claim 30 is incorporated. Blumenfeld, in view of Bowers, discloses all of the elements of the current invention as stated above. However, Blumenfeld, in view of Bowers, fails to expressly recite wherein the input requesting a text-to-speech service includes a hardware button input.
Roth teaches wherein the input requesting a text-to-speech service includes a hardware button input (Roth, [0503]: "If the user presses the "4" key when in the edit options menu, a text-to-speech or TTS menu will be displayed. In this menu, the "4" key toggles TTS play on or off.").
Blumenfeld, Bowers, and Roth are analogous arts because they each belong to the same field of text and audio processing. It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to have modified the speech generation device of Blumenfeld, as modified by the synthesized speech audio generator of Bowers, to incorporate the teachings of Roth to include specific hardware button inputs. This allows for the use of physical buttons to navigate menus and perform functions (Roth, [0502]). Including physical buttons in the system allows for more various inputs that can be used by people who may not be able to or may not want to use other input methods.
Claim(s) 6, 9, 10, 13, 20-23, and 28 is/are rejected under 35 U.S.C. 103 as being unpatentable over Blumenfeld, in view of Bowers, as applied to claims 1-3, 14-18, 29, and 30 above, and further in view of Cha et al. (US Pat. Pub. No. 2023/0239401 A1 hereinafter Cha).
Regarding claim 6, the rejection of claim 1 is incorporated. Blumenfeld, in view of Bowers, discloses all of the elements of the current invention as stated above. However, Blumenfeld, in view of Bowers, fails to expressly recite wherein the text-to-speech user interface includes a text input field, and wherein displaying the first text in the text-to-speech user interface includes inserting the first text in the text input field.
Cha teaches wherein the text-to-speech user interface includes a text input field, and wherein displaying the first text in the text-to-speech user interface includes inserting the first text in the text input field (Cha, [0030]: "The text input field (or text input window) 14 is configured to display the text data which the CTS user types").
Blumenfeld, Bowers, and Cha are analogous arts because they each belong to the same field of text and audio processing. It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to have modified the speech generation device of Blumenfeld, as modified by the synthesized speech audio generator of Bowers, to incorporate the teachings of Cha to include a text input field and virtual keyboard. Including these features allows for additional user friendly input methods (Cha, [0030]-[0031]). Creating a more user friendly system improves the overall user experience of the system.
Regarding claim 9, the rejection of claim 6 is incorporated. Blumenfeld, in view of Bowers and Cha, discloses all of the elements of the current invention as stated above. Cha further teaches wherein receiving the first text includes receiving a third user input to the text input field (Blumenfeld, [00106]: "The first slot is associated with the "breakfast" concept and the second slot is associated with the "fruit" concept. The slots allow a user to create dynamic messages with a reduced number of selections. By selecting the slots and changing the filler text, the example phrase "I want oatmeal and a banana for breakfast" can quickly and easily be changed to read as follows: "I want toast and a nectarine for breakfast"."). The same motivation for claim 6 applies equally to claim 9.
Regarding claim 10, the rejection of claim 1 is incorporated. Blumenfeld, in view of Bowers, discloses all of the elements of the current invention as stated above. However, Blumenfeld, in view of Bowers, fails to expressly recite the one or more programs further including instructions for: displaying a virtual keyboard.
Cha teaches the one or more programs further including instructions for: displaying a virtual keyboard (Cha, [0029]: "the CTS user device 10 includes a user interface which has a chat window (or caption display window) 12, a text input field (or text input window) 14, an answer suggestion display field 16, and optionally a virtual keyboard 18."). The same motivation for claim 6 applies equally to claim 10.
Regarding claim 13, the rejection of claim 10 is incorporated. Blumenfeld, in view of Bowers and Cha, discloses all of the elements of the current invention as stated above. Cha further teaches wherein receiving the first text includes receiving a fourth user input via the virtual keyboard (Cha, [0029]: "the CTS user device 10 includes a user interface which has a chat window (or caption display window) 12, a text input field (or text input window) 14, an answer suggestion display field 16, and optionally a virtual keyboard 18."). The same motivation for claim 6 applies equally to claim 13.
Regarding claim 20, the rejection of claim 1 is incorporated. Blumenfeld, in view of Bowers, discloses all of the elements of the current invention as stated above. However, Blumenfeld, in view of Bowers, fails to expressly recite wherein the second user input requesting the output of the first text includes a selection of a confirmation affordance.
Cha teaches wherein the second user input requesting the output of the first text includes a selection of a confirmation affordance (Cha, [0030]: "The text input field 14 may have a “send” button, and the CTS user device 10 sends the text data to the TTS system 60 when the CTS user touches or presses the “send” button or “enter” button on the virtual keyboard 18."). The same motivation for claim 6 applies equally to claim 20.
Regarding claim 21, the rejection of claim 20 is incorporated. Blumenfeld, in view of Bowers and Cha, discloses all of the elements of the current invention as stated above. Cha further teaches wherein the confirmation affordance is displayed included in a text entry field of the text-to-speech user interface (Cha, [0030]: "The text input field 14 may have a “send” button, and the CTS user device 10 sends the text data to the TTS system 60 when the CTS user touches or presses the “send” button or “enter” button on the virtual keyboard 18."). The same motivation for claim 6 applies equally to claim 21.
Regarding claim 22, the rejection of claim 20 is incorporated. Blumenfeld, in view of Bowers and Cha, discloses all of the elements of the current invention as stated above. Cha further teaches the one or more programs further including instructions for: in response to receiving the first text, displaying the confirmation affordance (Cha, [0031]: "Once an answer is selected, the CTS user device 10 may send the answer to the TTS system 60 or display the answer in the text input field 14 so that the CTS user can edit the answer and send the edited answer to the TTS system 60 by touching or pressing the “send” button” or “enter” button on the virtual keyboard 18."). The same motivation for claim 6 applies equally to claim 22.
Regarding claim 23, the rejection of claim 20 is incorporated. Blumenfeld, in view of Bowers and Cha, discloses all of the elements of the current invention as stated above. Cha further teaches wherein the confirmation affordance is displayed included in a virtual keyboard of the text-to-speech user interface (Cha, [0030]: "The text input field 14 may have a “send” button, and the CTS user device 10 sends the text data to the TTS system 60 when the CTS user touches or presses the “send” button or “enter” button on the virtual keyboard 18."). The same motivation for claim 6 applies equally to claim 23.
Regarding claim 28, the rejection of claim 1 is incorporated. Blumenfeld, in view of Bowers, discloses all of the elements of the current invention as stated above. However, Blumenfeld, in view of Bowers, fails to expressly recite wherein the first text includes at least a first word and a second word, and the one or more programs further including instructions for: while outputting the first text as the first synthesized speech output: displaying the first text; while outputting a first portion of the first synthesized speech output corresponding to the first word, visually emphasizing the first word in the displayed first text; and while outputting a second portion of the first synthesized speech output corresponding to the second word, visually emphasizing the second word in the displayed first text.
Cha teaches wherein the first text includes at least a first word and a second word, and the one or more programs further including instructions for: while outputting the first text as the first synthesized speech output: displaying the first text (Cha, [0033]: "the text data may be animated in gradually changing the color, font, style, or size of the letters of the text data to show the current point in time of playing the speech."); while outputting a first portion of the first synthesized speech output corresponding to the first word, visually emphasizing the first word in the displayed first text; and while outputting a second portion of the first synthesized speech output corresponding to the second word, visually emphasizing the second word in the displayed first text (Cha, [0033]: "the text data may be animated in gradually changing the color, font, style, or size of the letters of the text data to show the current point in time of playing the speech."). The same motivation for claim 6 applies equally to claim 28.
Claim(s) 7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Blumenfeld, in view of Bowers and Cha, as applied to claims 6, 9, 10, 13, 20-23, and 28 above, and further in view of Ghassabian, Benjamin (US Pat. Pub. No. 2016/0041965 A1 hereinafter Ghassabian).
Regarding claim 7, the rejection of claim 6 is incorporated. Blumenfeld, in view of Bowers and Cha, discloses all of the elements of the current invention as stated above. However, Blumenfeld, in view of Bowers and Cha, fails to expressly recite the one or more programs further including instructions for: displaying the text input field with first dimensions; and while receiving the first text, modifying the first dimensions based on the first text.
Ghassabian teaches the one or more programs further including instructions for: displaying the text input field with first dimensions; and while receiving the first text, modifying the first dimensions based on the first text (Ghassabian, [0852]: "The size of the text box may be dynamically adjusted based on the length of the text being typed.").
Blumenfeld, Bowers, Cha, and Ghassabian are analogous arts because they all belong to the same field of data processing. It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to have modified the speech generation device of Blumenfeld, as modified by the synthesized speech audio generator of Bowers and the text-to-speech system of Cha, to incorporate the teachings of Ghassabian to dynamically change the dimensions of the text input field. This ensures that all of the input text is displayed (Ghassabian, [0852]). Properly displaying input improves the user experience by providing feedback to the user.
Claim(s) 8 is/are rejected under 35 U.S.C. 103 as being unpatentable over Blumenfeld, in view of Bowers and Cha, as applied to claims 6, 9, 10, 13, 20-23, and 28 above, and further in view of DeMattei, Mark (US Pat. Pub. No. 2019/0124021 A1 hereinafter DeMattei).
Regarding claim 8, the rejection of claim 6 is incorporated. Blumenfeld, in view of Bowers and Cha, discloses all of the elements of the current invention as stated above. However, Blumenfeld, in view of Bowers and Cha, fails to expressly recite the one or more programs further including instructions for: in response to receiving the first text, displaying a deletion affordance included in the text input field.
DeMattei teaches the one or more programs further including instructions for: in response to receiving the first text, displaying a deletion affordance included in the text input field (DeMattei, [0159]: "the backspace button can operate to delete text in the text entry mode").
Blumenfeld, Bowers, Cha, and DeMattei are analogous arts because they all belong to the same field of text and video processing. It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to have modified the speech generation device of Blumenfeld, as modified by the synthesized speech audio generator of Bowers and the text-to-speech system of Cha, to incorporate the teachings of DeMattei to include a text deletion button that appears in response to input. This allows a user to delete text, but only in a text entry mode (DeMattei, [0159]). This ensures that a user can edit their input if necessary.
Claim(s) 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Blumenfeld, in view of Bowers and Cha, as applied to claims 6, 9, 10, 13, 20-23, and 28 above, and further in view of Threlkeld et al. (US Pat. Pub. No. 2017/0285813 A1 hereinafter Threlkeld).
Regarding claim 11, the rejection of claim 10 is incorporated. Blumenfeld, in view of Bowers and Cha, discloses all of the elements of the current invention as stated above. However, Blumenfeld, in view of Bowers and Cha, fails to expressly recite wherein displaying the virtual keyboard is performed in accordance with a determination that a hardware keyboard is not detected by the electronic device.
Threlkeld teaches wherein displaying the virtual keyboard is performed in accordance with a determination that a hardware keyboard is not detected by the electronic device (Threlkeld, [0034]: " In one or more implementations, external display module 108 is configured to detect whether a hardware keyboard is coupled to touch-capable display device 104. If a hardware keyboard is not coupled to the touch-capable display device 104, then output module 112 may cause display of an “on-screen” keyboard on touch-capable display device 104 that enables the user to type by touching locations on the touch-capable display device 104 that correspond to keys of the on-screen keyboard.").
Blumenfeld, Bowers, Cha, and Threlkeld are analogous arts because they all belong to the same field of data processing. It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to have modified the speech generation device of Blumenfeld, as modified by the synthesized speech audio generator of Bowers and the text-to-speech system of Cha, to incorporate the teachings of Threlkeld to display a virtual keyboard if no hardware keyboard is detected. This enables multiple forms of user input (Threlkeld, [0034]). Enabling additional forms of input improves the accessibility of the system.
Claim(s) 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Blumenfeld, in view of Bowers and Cha, as applied to claims 6, 9, 10, 13, 20-23, and 28 above, and further in view of Lee et al. (KR Pub. No. 2019/0021103 A hereinafter Lee).
Regarding claim 12, the rejection of claim 10 is incorporated. Blumenfeld, in view of Bowers and Cha, discloses all of the elements of the current invention as stated above. However, Blumenfeld, in view of Bowers and Cha, fails to expressly recite wherein displaying the virtual keyboard is performed in response to detecting a selection of a keyboard affordance included in the text-to-speech user interface.
Lee teaches wherein displaying the virtual keyboard is performed in response to detecting a selection of a keyboard affordance included in the text-to-speech user interface (Lee, Pg. 9, paragraph 3: "The first terminal 11 may further display icons providing various functions. For example, an icon 75 for providing a handwriting mode for inputting text by directly writing a character using a touch screen or an electronic pen, and a keyboard (both a physical keyboard or a keyboard displayed on the touch screen) An icon 76 that provides a typing mode for typing text, and a call termination icon 77 may be provided.").
Blumenfeld, Bowers, Cha, and Lee are analogous arts because they all belong to the same field of text and video processing. It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to have modified the speech generation device of Blumenfeld, as modified by the synthesized speech audio generator of Bowers and the text-to-speech system of Cha, to incorporate the teachings of Lee to display a virtual keyboard if a corresponding icon is selected. This allows the user to activate the virtual keyboard if they want to enter text (Lee, Pg. 9, paragraph 3). This ensures that the virtual keyboard can be activated manually if it does not otherwise activate.
Claim(s) 19 and 25-27 is/are rejected under 35 U.S.C. 103 as being unpatentable over Blumenfeld, in view of Bowers, as applied to claims 1, 14-18, 29, and 30 above, and further in view of Davies et al. (US Pat. Pub. No. 2020/0035218 A1 hereinafter Davies).
Regarding claim 19, the rejection of claim 1 is incorporated. Blumenfeld, in view of Bowers, discloses all of the elements of the current invention as stated above. However, Blumenfeld, in view of Bowers, fails to expressly recite wherein the text-to-speech user interface includes a cancel affordance that, when selected, causes displaying the text-to-speech user interface to cease.
Davies teaches wherein the text-to-speech user interface includes a cancel affordance that, when selected, causes displaying the text-to-speech user interface to cease (Davies, [0103]: "TTS system 116 can display a message (e.g., “Drag here to close”) at location 224 to indicate to the user that releasing TTS selector 204 at location 224 will deactivate TTS system 116.").
Blumenfeld, Bowers, and Davies are analogous arts because they each belong to the same field of text and audio processing. It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to have modified the speech generation device of Blumenfeld, as modified by the synthesized speech audio generator of Bowers, to incorporate the teachings of Davies to include cancel and pause functions in the text-to-speech system. This allows the user to pause and/or close the program when wanted (Davies, [0098]-[0100]). This increases interactivity of the system, thus improving the user experience.
Regarding claim 25, the rejection of claim 1 is incorporated. Blumenfeld, in view of Bowers, discloses all of the elements of the current invention as stated above. However, Blumenfeld, in view of Bowers, fails to expressly recite the one or more programs further including instructions for: in response to initiating the output of the first text as the first synthesized speech output, displaying a pause affordance.
Davies teaches the one or more programs further including instructions for: in response to initiating the output of the first text as the first synthesized speech output, displaying a pause affordance (Davies, [0098]: "during playback of the speech data, TTS system 116 can display TTS controller 214 to provide the user with options to control the playback. TTS controller 214 can be associated with TTS interface 140 of TTS system 116. TTS system 116 can display TTS controller 214 with an “X” icon, and the user can activate TTS controller 214 (e.g., by tapping the “X” icon) to pause or stop playback."). The same motivation for claim 19 applies equally to claim 25.
Regarding claim 26, the rejection of claim 25 is incorporated. Blumenfeld, in view of Bowers and Davies, discloses all of the elements of the current invention as stated above. Davies further teaches the one or more programs further including instructions for: detecting a selection of the pause affordance; and in response to detecting the selection of the pause affordance: cease the output of the first text as the first synthesized speech output; cease displaying the pause affordance; and displaying a play affordance (Davies, [0098]: "during playback of the speech data, TTS system 116 can display TTS controller 214 to provide the user with options to control the playback. TTS controller 214 can be associated with TTS interface 140 of TTS system 116. TTS system 116 can display TTS controller 214 with an “X” icon, and the user can activate TTS controller 214 (e.g., by tapping the “X” icon) to pause or stop playback."). The same motivation for claim 19 applies equally to claim 26.
Regarding claim 27, the rejection of claim 25 is incorporated. Blumenfeld, in view of Bowers and Davies, discloses all of the elements of the current invention as stated above. Davies further teaches the one or more programs further including instructions for: after outputting the first text as the first synthesized speech output, ceasing displaying the pause affordance (Davies, [0101]: "when TTS system 116 completes playback of the speech data, or if the user interrupts playback of the speech data via TTS controller 214, then TTS system 116 can display TTS selector 204 returning from the second position to the first position to indicate that TTS system 116 is ready to receive another TTS request from the user."). The same motivation for claim 19 applies equally to claim 27.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Haley, Mark R. (US Pat. Pub. No. 2003/0009342 A1) discloses software that converts text-to-speech in any language.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to TYLER J BECKER whose telephone number is (703)756-1271. The examiner can normally be reached M-Th, 7:15am-5:45pm PT.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at (571) 272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/TYLER BECKER/ Examiner, Art Unit 2657
/DANIEL C WASHBURN/ Supervisory Patent Examiner, Art Unit 2657