Prosecution Insights
Last updated: October 01, 2026
Application No. 19/064,810

CONVERSATION SUPPORT DEVICE, CONVERSATION SUPPORT SYSTEM, CONVERSATION SUPPORT METHOD, AND STORAGE MEDIUM

Non-Final OA §103
Filed
Feb 27, 2025
Priority
Mar 06, 2024 — JP 2024-034031
Examiner
CRESPO FEBLES, HECTOR J
Art Unit
Tech Center
Assignee
Honda Motor Co., Ltd.
OA Round
1 (Non-Final)
100%
Grant Probability
Favorable
1-2
OA Rounds
6m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 100% — above average
100%
Career Allowance Rate
1 granted / 1 resolved
+40.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Fast prosecutor
2y 1m
Avg Prosecution
16 currently pending
Career history
13
Total Applications
across all art units

Statute-Specific Performance

§101
4.7%
-35.3% vs TC avg
§103
75.0%
+35.0% vs TC avg
§102
15.6%
-24.4% vs TC avg
§112
3.1%
-36.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claim(s) 1, 2, 6 and 7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Subramanian; Balan et al. (US 20070208569 A1), hereinafter SUBRAMANIAN, in view of Harazi; Liron et al. (US 20230252972 A1), hereinafter HARAZI. Regarding claim 1, SUBRAMANIAN teaches: a voice conversion unit configured to convert the vocal sound, the voice conversion unit recognizes a portion of the text displayed on the display unit, which is selected by a user, as specified text, when an emotion for the specified text is input, SUBRAMANIAN [0005] “A text and emotion markup abstraction for a voice communication in a source language is translated into a target language and then voice synthesized and adjusted for emotion. The emotion metadata is translated into emotion metadata for a target language using emotion translation definitions for the target language. The text is translated into a text for the target language using text translation definitions. Additionally, the translated emotion metadata is used to emotion mine words that have an emotion connotation in the culture of the target language. The emotion words are than substituted for corresponding words in the target language text. The translated text and emotion words are modulated into a synthesized voice. The delivery of the synthesized voice can be adjusted for emotion using the translated emotion metadata. Modifications to the synthesized voice patterns are derived by emotion mining an emotion-to-voice pattern dictionary for emotion voice patterns, which are used to modify the delivery of the modulated voice.” SUBRAMANIAN [0122] “Using the present invention as described immediately above, a user could receive an abstraction of a voice communication, translate the textual and emotion content of the abstraction and hear the communication in the user's language with emotion consistent with the user's culture. In one example, a speaker creates an audio message for a recipient who speaks a different language. The speech communication is received at PC 1012 with integrated emotion communication architecture 200. Using the dictionary definitions appropriate for the speaker, the voice communication is converted into text which preserves the emotion of the speech with emotion markup metadata and is transmitted to the recipient. The text with emotion markup is received at the recipients device, for instance at laptop 1026 with emotion communication architecture 200 integrated thereon. Using the dictionary definitions for the recipient's language and culture, the text and emotion are translated and emotion words included in the text that are consistent with the recipient's culture. The text is then voice synthesized and the synthesized delivery is adjusted for the emotion. Of course, the user of PC 1012 can designate which portions of text to adjust with the voice synthesized using the emotion metadata.” SUBRAMANIAN [0036] “Selection commands may also be received at emotion markup component 210, issued by a user, for specifying particular words, phrases, sentences and passages in the communication for emotion analysis. These commands may also designate which type of analysis, text pattern analysis (text mining), or voice analysis, to use for extracting emotion from the selected portion of the communication.” converts the text into a vocal sound so that the vocal sound of the text becomes a vocal sound corresponding to the selected emotion, SUBRAMANIAN [0109]”In accordance with one exemplary embodiment of the present invention, text with emotion markup metadata is converted to voice communication, with or without language translation. This aspect of the invention will be discussed with regard to instant messaging (IM). A user of a PC, laptop, PDA, cell phone, telephone or other network appliance creates a textual message that includes emotion inferences, for instance using one of PCs 1012 or 1016, one of laptops 1018, 1026, 1047 or 1067, one of PDAs 1020 or 1058, one of cell phones 1056 or 1059, or even using one of telephones 1046, 1048, or 1049. The emotion inferences may include emoticons, highlighting, punctuation or some other emphasis indicative of emotion. In accordance with one exemplary embodiment of the present invention, the device that creates the message may or may not be configured with emotion markup component 210 for marking up the text. In any case, the text message with emotion markup is transmitted to a device that includes emotion translation component 250, either separately, or in emotion communication architecture 200, such as laptop 1026. The emotion markup should be in a standard format or contain standard markup metadata that can be recognized as emotion content by emotion translation component 250. If it is not recognizable, the text and nonstandard emotion markup can be processed into standardized emotion markup metadata by any device that includes emotion markup component 210, using the sender's profile information (see FIG. 4).” and outputs the converted vocal sound corresponding to the emotion from the voice output unit. SUBRAMANIAN [0110] “Once the text and emotion markup metadata are received at emotion translation component 250, the recipient can choose between content delivery modes, e.g., text or voice. The recipient of the text message may also specify a language for content delivery. The language selection is used for populating text-to-text dictionary 253 with the appropriate text definitions for translating the text to the selected language. The language selection is also used for populating emotion-to-emotion dictionary 255 with the appropriate emotion definitions for translating the emotion to the culture of the selected language, and for populating emotion-to-voice pattern dictionary 222 with the appropriate voice pattern definitions for adjusting the synthesized audio voice for emotion. The language selection also dictates which word and phrase definitions are appropriate for populating emotion-to-phrase dictionary 220, used for emotion mining for emotion charged words that are particular to the culture of the selected language.” SUBRAMANIAN [0109] “In accordance with one exemplary embodiment of the present invention, text with emotion markup metadata is converted to voice communication, with or without language translation. ...” SUBRAMANIAN does not teach, but HARAZI teaches: A conversation support device comprising: a display unit configured to display input text; HARAZI [0130] “The graphical user interface 700 can receive input that selects a text input option 730. In response, the graphical user interface 700 allows the user to modify or add more text to the text region 720. The graphical user interface 700 displays a speaker selection option 740. In response to receiving a user selection of the speaker selection option 740, the graphical user interface 700 presents a list of available speakers (speakers for which one or more embeddings associated with different emotions are stored and available). The text to speech system 230 processes the words of the text region 720 (e.g., as the words are input or after a command to process is received). “ a voice output unit configured to output a vocal sound into which the text has been converted; HARAZI [0131] “The text to speech system 230 generates a new embedding for a speaker in response to determining that the currently stored embeddings for the selected speaker fail to include the determined emotion. The text to speech system 230 applies the new embedding to the words in the text region 720 to render or generate an audio stream 722 in which the selected speaker audibly (verbally) speaks the words in the text region 720 with the determined emotion (e.g., a level 8 joy and a level 2 fear).” It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of SUBRAMANIAN a device with a display, explicitly. The benefit and motivation of such modification is discussed by HARAZI in the following portion: HARAZI [0016] “The disclosed techniques improve the efficiency of using the electronic device by providing a text to speech (TTS) system that dynamically generates an audio file in which a speaker's voice is used to virtually speak words given as text using one or more emotions derived from the input. …” Regarding claim 2, SUBRAMANIAN does not teach, but HARAZI teaches: The conversation support device according to claim 1, wherein the voice conversion unit recognizes all input text as the specified text when a user does not select a part of the text. HARAZI [0130] “The graphical user interface 700 can receive input that selects a text input option 730. In response, the graphical user interface 700 allows the user to modify or add more text to the text region 720. The graphical user interface 700 displays a speaker selection option 740. In response to receiving a user selection of the speaker selection option 740, the graphical user interface 700 presents a list of available speakers (speakers for which one or more embeddings associated with different emotions are stored and available). The text to speech system 230 processes the words of the text region 720 (e.g., as the words are input or after a command to process is received). The text to speech system 230 determines an emotion associated with the words (e.g., by way of an emotion selected by the user responsive to selection of the speaker selection option 740 or by automatically computing an intensity or level of emotion for one or more words in the text region 720). In some examples, the graphical user interface 700 can include an automatic mode option. In response to receiving input that selects the automatic mode option, the text to speech system 230 processes words of the text string and automatically determines or selects a particular voice with an automatically determined emotion and/or level of emotion.” It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of SUBRAMANIAN the capability to process a full message in the emotion text-to-speech program instead of just selecting a portion. The benefit and motivation of such modification is discussed by HARAZI in the following portion: HARAZI [0022] “…The messaging client 104 can generalize the style components of the identified speakers corresponding to the specified emotion and can combine the generalized style components with the voice component of the specified speaker. This results in a new embedding for the specified speaker that is associated with the specified emotion. This new embedding can be used to generate the audio in which the specified speaker virtually speaks the one or more words of the text string with the specified emotion. In some examples, the new embedding—can be generated using a machine learning technique (e.g., a neural network).” Regarding claim 6, arguments analogous to claim 1 are applicable, furthermore SUBRAMANIAN teaches: A conversation support method SUBRAMANIAN [0128] “The flowchart and block diagrams in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present invention. …” Regarding claim 7, arguments analogous to claim 1 are applicable, furthermore SUBRAMANIAN does not teach, but HARAZI teaches: A computer-readable non-transitory storage medium that stores a program causing a computer of a conversation support device to execute: HARAZI [0132] “FIG. 8 is a flowchart illustrating example operations of the messaging client 104 in performing process 800, according to example examples. The process 800 may be embodied in computer-readable instructions for execution by one or more processors such that the operations of the process 800 may be performed in part or in whole by the functional components of the messaging server system 108; accordingly, the process 800 is described below by way of example with reference thereto. However, in other examples at least some of the operations of the process 800 may be deployed on various other hardware configurations. The operations in the process 800 can be performed in any order, in parallel, or may be entirely skipped and omitted.” It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of SUBRAMANIAN a device with a computer-readable non-transitory storage medium, explicitly. The benefit and motivation of such modification is discussed by HARAZI in the following portion: HARAZI [0016] “The disclosed techniques improve the efficiency of using the electronic device by providing a text to speech (TTS) system that dynamically generates an audio file in which a speaker's voice is used to virtually speak words given as text using one or more emotions derived from the input. …” Claim(s) 3 and 4 is/are rejected under 35 U.S.C. 103 as being unpatentable over SUBRAMANIAN, in view of HARAZI in further view of Bruns; Miroslawa et al. (US 20190028416 A1), hereinafter BRUNS. Regarding claim 3, SUBRAMANIAN does not teach, but HARAZI teaches: an emotion addition button image for issuing an instruction to convert the vocal sound, HARAZI [0130] “The graphical user interface 700 can receive input that selects a text input option 730. In response, the graphical user interface 700 allows the user to modify or add more text to the text region 720. The graphical user interface 700 displays a speaker selection option 740. In response to receiving a user selection of the speaker selection option 740, the graphical user interface 700 presents a list of available speakers (speakers for which one or more embeddings associated with different emotions are stored and available). The text to speech system 230 processes the words of the text region 720 (e.g., as the words are input or after a command to process is received). The text to speech system 230 determines an emotion associated with the words (e.g., by way of an emotion selected by the user responsive to selection of the speaker selection option 740 or by automatically computing an intensity or level of emotion for one or more words in the text region 720). In some examples, the graphical user interface 700 can include an automatic mode option. In response to receiving input that selects the automatic mode option, the text to speech system 230 processes words of the text string and automatically determines or selects a particular voice with an automatically determined emotion and/or level of emotion.” (See Figure 1) PNG media_image1.png 669 838 media_image1.png Greyscale Figure 1. GUI example from HARAZI It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of SUBRAMANIAN a graphical user interface capable to execute the command claimed when pressed. The benefit and motivation of such modification is discussed by HARAZI in the following portion: HARAZI [0024] “The messaging server system 108 supports various services and operations that are provided to the messaging client 104. Such operations include transmitting data to, receiving data from, and processing data generated by the messaging client 104. This data may include message content, client device information, geolocation information, media augmentation and overlays, message content persistence conditions, social network information, watermarks (combined indications of messages and reactions being read or presented to a user of a client device 102) and live event information, as examples. Data exchanges within the messaging system 100 are invoked and controlled through functions available via user interfaces (UIs) of the messaging client 104.” SUBRAMANIAN in view of HARAZI does not teach but BRUNS teaches: The conversation support device according to claim 1, wherein the display unit displays an image in which speech other than that of the user is converted into text through voice recognition, BRUNS [0075] “FIG. 4 shows one preferred method of providing a user interface for communicating 400 in accordance with the present invention. A box 404 includes the title of the page “Friend #1.” This page shows all the communications between the user and Friend #1. A box 402 includes the title “You have a new message from Friend #1. The box 402 further includes a box 406 entitled “Play Message” and a box 408 entitled “Show Original Text Message.” A box 412 shows the date (Date 1) when the message was delivered. When the user taps on the box 406, the user will be directed to another page entitled “Play Friend #1 Message,” see FIG. 6 discussed in more details below. When the user taps on the box 408, the user will be directed to another page entitled “Show Friend #1 Original Text Message.” A box 410 shows the animation character of Friend #1.” (See Figure 2) PNG media_image2.png 669 518 media_image2.png Greyscale Figure 2. GUI example from BRUNS 2 a text input area for inputting the text, BRUNS [0080]” FIG. 5 shows one preferred method of providing a user interface for communicating 500 in accordance with the present invention. A box 502 includes the title of the page “Send Message.” This page is presented for the user to compose his/her electronic message. A box 504 shows the user's animation character, in this case the image of the Avatar character in the movie “Avatar.” In one instance, when the user is directed to this page, the use is presented with the moving images of the animation character Avatar uttering the words “Please Send Your Message.” A box 506 includes the title “Type your message here” and when tapped by the user, the user device, such as the device 102, provides a means such as a keyboard for the user to input the text. A box 508 includes the title “Reply” and when pressed causes the message to be transmitted. As discussed above, the steps of electronic message to speech conversion and generation of moving images of the animation character may occur by the user's device, the server, or recipient device. As such, if the user device is so designated, tapping on the box 508 performs those steps and transmits the speech and moving images. A box 510 provides some of the above options and when tapped on a specific option it will direct the user to another page with the corresponding title.” (See Figure 3) PNG media_image3.png 651 539 media_image3.png Greyscale Figure 3. GUI example from BRUNS BRUNS [0089] “FIG. 10 shows a flow diagram 1000 of one preferred method of communicating in accordance with the present invention which may be implemented utilizing the computer network system depicted in FIG. 1. According, to this embodiment, the method comprises composing an electronic message, such as a text message, via the device 102, at 1002. The method further comprises selecting an animation character, via the device 102, at 1006. The method further comprises transmitting the electronic message and animation character, via the device 102, at 1010. The method further comprises receiving the electronic message and animation character, via the device 104, at 1014. The method further comprises converting, the electronic message into speech, via the device 104, at 1018. The method further comprises generating moving images of the animation character, via the device 104, at 1022. The method further comprises transmitting the speech and moving images, via the device 104, at 1026. The method further comprises receiving the speech and moving images, via the device 106, at 1030. The method further comprises outputting the speech, via the device 106, at 1034. The method further comprises displaying the moving images, via the device 106, at 1038.” and an output button image for outputting the converted vocal sound. BRUNS [0080] “FIG. 5 shows one preferred method of providing a user interface for communicating 500 in accordance with the present invention. A box 502 includes the title of the page “Send Message.” This page is presented for the user to compose his/her electronic message. A box 504 shows the user's animation character, in this case the image of the Avatar character in the movie “Avatar.” In one instance, when the user is directed to this page, the use is presented with the moving images of the animation character Avatar uttering the words “Please Send Your Message.” A box 506 includes the title “Type your message here” and when tapped by the user, the user device, such as the device 102, provides a means such as a keyboard for the user to input the text. A box 508 includes the title “Reply” and when pressed causes the message to be transmitted. As discussed above, the steps of electronic message to speech conversion and generation of moving images of the animation character may occur by the user's device, the server, or recipient device. As such, if the user device is so designated, tapping on the box 508 performs those steps and transmits the speech and moving images. A box 510 provides some of the above options and when tapped on a specific option it will direct the user to another page with the corresponding title.” It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of SUBRAMANIAN in view of HARAZI a graphical user interface capable to execute the command claimed when pressed and display information. The benefit and motivation of such modification is discussed by BRUNS in the following portion: BRUNS [0072] “An App, such as the App 110, may be further configured to include methods of providing a user interface to facilitate a visual representation of the methods of communication discussed herein. The user interface may be implemented on one or more devices such as the devices, 102, 104, and 106, separately or in combination. …” Regarding claim 4, SUBRAMANIAN teaches: until the voice output unit outputs the converted vocal sound corresponding to the emotion. SUBRAMANIAN [0005] “A text and emotion markup abstraction for a voice communication in a source language is translated into a target language and then voice synthesized and adjusted for emotion. The emotion metadata is translated into emotion metadata for a target language using emotion translation definitions for the target language. The text is translated into a text for the target language using text translation definitions. Additionally, the translated emotion metadata is used to emotion mine words that have an emotion connotation in the culture of the target language. The emotion words are than substituted for corresponding words in the target language text. The translated text and emotion words are modulated into a synthesized voice. The delivery of the synthesized voice can be adjusted for emotion using the translated emotion metadata. Modifications to the synthesized voice patterns are derived by emotion mining an emotion-to-voice pattern dictionary for emotion voice patterns, which are used to modify the delivery of the modulated voice.” SUBRAMANIAN [0101] “Returning to step 808, if the text and emotion markup are to be translated, the text to text dictionary is populated with translation from the original language of the text and markup, to the language of the user (step 810). Next, the text with emotion markup is received (step 813) and the emotion information is parsed (step 815). The text is translated from the original language to the language of the user with the text to text dictionary (step 818). The process then continues by checking if the text is marked for emotion adjustment (step 820), and the emotion metadata is translated to the user's cultural using the definitions in emotion to emotion dictionary (step 822). The emotion to word/phrase dictionary is emotion mined for words that convey the emotion consistent with the culture of the user (step 824). And a check is made to determine whether to synthesize the text into audio (step 826). If not, the translated text (with the translated emotion) is output (step 836). Otherwise, the text is modulated (step 828) the modulated voice is adjusted for emotion by altering the tone, camber and frequency of synthesized voice (step 830). The synthesized voice with emotion is the output (step 836). The process reiterates from step 813 until all the text has been output as audio and the process ends.” SUBRAMANIAN in view of HARAZI does not teach but BRUNS teaches: The conversation support device according to claim 3, wherein the display unit does not display the input text in a display area of an image in which speech other than that of the user is converted into text through voice recognition, BRUNS [0080] “FIG. 5 shows one preferred method of providing a user interface for communicating 500 in accordance with the present invention. A box 502 includes the title of the page “Send Message.” This page is presented for the user to compose his/her electronic message. A box 504 shows the user's animation character, in this case the image of the Avatar character in the movie “Avatar.” In one instance, when the user is directed to this page, the use is presented with the moving images of the animation character Avatar uttering the words “Please Send Your Message.” A box 506 includes the title “Type your message here” and when tapped by the user, the user device, such as the device 102, provides a means such as a keyboard for the user to input the text. A box 508 includes the title “Reply” and when pressed causes the message to be transmitted. As discussed above, the steps of electronic message to speech conversion and generation of moving images of the animation character may occur by the user's device, the server, or recipient device. As such, if the user device is so designated, tapping on the box 508 performs those steps and transmits the speech and moving images. A box 510 provides some of the above options and when tapped on a specific option it will direct the user to another page with the corresponding title.” (See Figure 3 wherein the message is displayed with the rest of the conversation only once it has been sent.) It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of SUBRAMANIAN in view of HARAZI a graphical user interface capable to execute the command claimed when pressed and display information. The benefit and motivation of such modification is discussed by BRUNS in the following portion: BRUNS [0072] “An App, such as the App 110, may be further configured to include methods of providing a user interface to facilitate a visual representation of the methods of communication discussed herein. The user interface may be implemented on one or more devices such as the devices, 102, 104, and 106, separately or in combination. …” Claim(s) 5 is/are rejected under 35 U.S.C. 103 as being unpatentable over SUBRAMANIAN in view of HARAZI, in further view of Maurer; Aaron Josephus et al. (US 20240176960 A1), hereinafter MAURER. Regarding claim 5, SUBRAMANIAN teaches: a voice output unit for outputting a vocal sound into which the text has been converted, SUBRAMANIAN [0107] “... It is expected that some or all of these devices will be configured with internal or external audio input/output components (microphones and speakers), for instance PC 1012 is shown with external microphone 1014 and external speaker(s) 1013.” SUBRAMANIAN [0037] “Emotion translation component 250 receives a communication, typically text with emotion markup metadata, and parses the emotion content. Emotion translation component 250 synthesizes the text into a natural language and adjusts the tone, cadence and amplitude of the voice delivery for emotion based the emotion metadata accompanying the text. Alternatively, prior to modulating the communication stream, emotion translation component 250 may translate the text and emotion metadata into the language of the listener.” and a voice conversion unit for converting the vocal sound, the voice conversion unit recognizes a portion of the text displayed on the display unit, which is selected by a user as specified text, when an emotion for the specified text is input, SUBRAMANIAN [0005] “A text and emotion markup abstraction for a voice communication in a source language is translated into a target language and then voice synthesized and adjusted for emotion. The emotion metadata is translated into emotion metadata for a target language using emotion translation definitions for the target language. The text is translated into a text for the target language using text translation definitions. Additionally, the translated emotion metadata is used to emotion mine words that have an emotion connotation in the culture of the target language. The emotion words are than substituted for corresponding words in the target language text. The translated text and emotion words are modulated into a synthesized voice. The delivery of the synthesized voice can be adjusted for emotion using the translated emotion metadata. Modifications to the synthesized voice patterns are derived by emotion mining an emotion-to-voice pattern dictionary for emotion voice patterns, which are used to modify the delivery of the modulated voice.” SUBRAMANIAN [0122] “Using the present invention as described immediately above, a user could receive an abstraction of a voice communication, translate the textual and emotion content of the abstraction and hear the communication in the user's language with emotion consistent with the user's culture. In one example, a speaker creates an audio message for a recipient who speaks a different language. The speech communication is received at PC 1012 with integrated emotion communication architecture 200. Using the dictionary definitions appropriate for the speaker, the voice communication is converted into text which preserves the emotion of the speech with emotion markup metadata and is transmitted to the recipient. The text with emotion markup is received at the recipients device, for instance at laptop 1026 with emotion communication architecture 200 integrated thereon. Using the dictionary definitions for the recipient's language and culture, the text and emotion are translated and emotion words included in the text that are consistent with the recipient's culture. The text is then voice synthesized and the synthesized delivery is adjusted for the emotion. Of course, the user of PC 1012 can designate which portions of text to adjust with the voice synthesized using the emotion metadata.” SUBRAMANIAN [0036] “Selection commands may also be received at emotion markup component 210, issued by a user, for specifying particular words, phrases, sentences and passages in the communication for emotion analysis. These commands may also designate which type of analysis, text pattern analysis (text mining), or voice analysis, to use for extracting emotion from the selected portion of the communication.” converts the text into a vocal sound so that the vocal sound of the text becomes a vocal sound corresponding to the selected emotion, outputs the converted vocal sound corresponding to the emotion from the voice output unit, SUBRAMANIAN [0122] “Using the present invention as described immediately above, a user could receive an abstraction of a voice communication, translate the textual and emotion content of the abstraction and hear the communication in the user's language with emotion consistent with the user's culture. In one example, a speaker creates an audio message for a recipient who speaks a different language. The speech communication is received at PC 1012 with integrated emotion communication architecture 200. Using the dictionary definitions appropriate for the speaker, the voice communication is converted into text which preserves the emotion of the speech with emotion markup metadata and is transmitted to the recipient. The text with emotion markup is received at the recipients device, for instance at laptop 1026 with emotion communication architecture 200 integrated thereon. Using the dictionary definitions for the recipient's language and culture, the text and emotion are translated and emotion words included in the text that are consistent with the recipient's culture. The text is then voice synthesized and the synthesized delivery is adjusted for the emotion. Of course, the user of PC 1012 can designate which portions of text to adjust with the voice synthesized using the emotion metadata.” SUBRAMANIAN [0036] “Selection commands may also be received at emotion markup component 210, issued by a user, for specifying particular words, phrases, sentences and passages in the communication for emotion analysis. These commands may also designate which type of analysis, text pattern analysis (text mining), or voice analysis, to use for extracting emotion from the selected portion of the communication.” transmits the input text to the conference support device, SUBRAMANIAN [0109] “In accordance with one exemplary embodiment of the present invention, text with emotion markup metadata is converted to voice communication, with or without language translation. This aspect of the invention will be discussed with regard to instant messaging (IM). A user of a PC, laptop, PDA, cell phone, telephone or other network appliance creates a textual message that includes emotion inferences, for instance using one of PCs 1012 or 1016, one of laptops 1018, 1026, 1047 or 1067, one of PDAs 1020 or 1058, one of cell phones 1056 or 1059, or even using one of telephones 1046, 1048, or 1049. The emotion inferences may include emoticons, highlighting, punctuation or some other emphasis indicative of emotion. In accordance with one exemplary embodiment of the present invention, the device that creates the message may or may not be configured with emotion markup component 210 for marking up the text. In any case, the text message with emotion markup is transmitted to a device that includes emotion translation component 250, either separately, or in emotion communication architecture 200, such as laptop 1026. The emotion markup should be in a standard format or contain standard markup metadata that can be recognized as emotion content by emotion translation component 250. If it is not recognizable, the text and nonstandard emotion markup can be processed into standardized emotion markup metadata by any device that includes emotion markup component 210, using the sender's profile information (see FIG. 4).” SUBRAMANIAN does not teach, but HARAZI teaches: wherein the terminal includes a display unit for displaying input text, HARAZI [0130] “The graphical user interface 700 can receive input that selects a text input option 730. In response, the graphical user interface 700 allows the user to modify or add more text to the text region 720. The graphical user interface 700 displays a speaker selection option 740. In response to receiving a user selection of the speaker selection option 740, the graphical user interface 700 presents a list of available speakers (speakers for which one or more embeddings associated with different emotions are stored and available). The text to speech system 230 processes the words of the text region 720 (e.g., as the words are input or after a command to process is received). “ It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of SUBRAMANIAN a graphical user interface capable to execute the command claimed when pressed. The benefit and motivation of such modification is discussed by HARAZI in the following portion: HARAZI [0024] “The messaging server system 108 supports various services and operations that are provided to the messaging client 104. Such operations include transmitting data to, receiving data from, and processing data generated by the messaging client 104. This data may include message content, client device information, geolocation information, media augmentation and overlays, message content persistence conditions, social network information, watermarks (combined indications of messages and reactions being read or presented to a user of a client device 102) and live event information, as examples. Data exchanges within the messaging system 100 are invoked and controlled through functions available via user interfaces (UIs) of the messaging client 104.” It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of SUBRAMANIAN a device with a display, explicitly. The benefit and motivation of such modification is discussed by HARAZI in the following portion: HARAZI [0016] “The disclosed techniques improve the efficiency of using the electronic device by providing a text to speech (TTS) system that dynamically generates an audio file in which a speaker's voice is used to virtually speak words given as text using one or more emotions derived from the input. …” SUBRAMANIAN in view of HARAZI does not teach, but MAURER teaches: A conversation support system comprising: a terminal; MAURER [0021] “In at least one example, the server(s) 102 can communicate with a user computing device 104 via one or more network(s) 106. That is, the server(s) 102 and the user computing device 104 can transmit, receive, and/or store data (e.g., content, information, or the like) using the network(s) 106, as described herein. The user computing device 104 can be any suitable type of computing device, e.g., portable, semi-portable, semi-stationary, or stationary. Some examples of the user computing device 104 can include a tablet computing device, a smart phone, a mobile communication device, a laptop, a netbook, a desktop computing device, a terminal computing device, a wearable computing device, an augmented reality device, an Internet of Things (IOT) device, or any other computing device capable of sending communications and performing the functions according to the techniques described herein. While a single user computing device 104 is shown, in practice, the example environment 100 can include multiple (e.g., tens of, hundreds of, thousands of, millions of) user computing devices. In at least one example, user computing devices, such as the user computing device 104, can be operable by users to, among other things, access communication services via the communication platform. A user can be an individual, a group of individuals, an employer, an enterprise, an organization, and/or the like.” and a conference support device, MAURER [0134] “These various types of functions 336 may in turn integrate with APIs 344. In some examples, APIs 344 are associated with third-party services that functions 336 employ to provide a custom integration between a particular third-party service and a group-based communication system. Examples of third-party service integrations include video conferencing, sales, marketing, customer service, project management, and engineering application integration. In such an example, one of the triggers 318 would be a slash command 326 that is used to trigger a hosted-code function 342, which makes an API call to a third-party video conferencing provider by way of one of the APIs 344. As shown in FIG. 3B, the APIs 344 may themselves also become a source of any number of triggers 318 or events 328. Continuing the above example, successful completion of a video conference would trigger one of the functions 336 that sends a message initiating a further API call to the third-party video conference provider to download and archive a recording of the video conference and store it in a group-based communication system channel.” MAURER [0146] “In some examples, the process 500 may begin at operation 502, which may include receiving teleconferencing meeting data associated with a channel of the group-based communication platform. In some examples, the teleconferencing meeting data may be the ambient data associated with a synchronous multimedia collaboration session that is occurring within the channel of the group-based communication platform. Thus, in some examples, the teleconferencing meeting data may include raw audio-visual data and user reaction data. Further, the user reaction data of some examples includes one or more of an emoji selected by a user, a user's perceived physical expressions (e.g., a gesture detected using machine vision techniques from video data), messages or text input by the user (e.g., during the meeting), or a thread of messages input by a plurality of users (e.g., associated with a channel or virtual space). In some examples, the user reaction data comprises messages or text input to a user interface proximate video data substantially simultaneously during generation of the audio-visual data.” and when a display image is acquired from the conference support device, displays the acquired display image on the display unit, MAURER[0063] “The user computing device 104 can further be equipped with various input/output devices 136 (e.g., I/O devices). Such I/O devices 136 can include a display, various user interface controls (e.g., buttons, joystick, keyboard, mouse, touch screen, etc.), audio speakers, connection ports and so forth.” MAURER [0034] “In some examples, the audio/video component 118 can be configured to generate a transcript of the conversation, and further can store the transcript in association with the audio and/or video data. The transcript can include a textual representation of the audio and/or video data. In at least one example, the audio/video component 118 can use known speech recognition techniques to generate the transcript. In some examples, the audio/video component 118 can generate the transcript concurrently or substantially concurrently with the conversation. That is, in some examples, the audio/video component 118 can be configured to generate a textual representation of the conversation while it is being conducted. ...” and the conference support device, after the input text is acquired from the terminal, transmits the text to the terminal as the display image. MAURER [0084] “For purposes of this discussion, a “message” can refer to any electronically generated digital object provided by a user using the user computing device 104 and that is configured for display within a communication channel and/or other virtual space for facilitating communications (e.g., a virtual space associated with direct message communication(s), etc.) as described herein. A message may include any text, image, video, audio, or combination thereof provided by a user (using a user computing device). For instance, the user may provide a message that includes text, as well as an image and a video, within the message as message contents. ...” MAURER [0076] “Additionally, or in the alternative, in some examples, a virtual space can be associated with one or more canvases with which the user is associated. In at least one example, the canvas can include a flexible canvas for curating, organizing, and sharing collections of information between users. That is, the canvas can be configured to be accessed and/or modified by two or more users with appropriate permissions. In at least one example, the canvas can be configured to enable sharing of text, images, videos, GIFs, drawings (e.g., user-generated drawing via a canvas interface), gaming content (e.g., users manipulating gaming controls synchronously or asynchronously), and/or the like. In at least one example, modifications to a canvas can include adding, deleting, and/or modifying previously shared (e.g., transmitted, presented) data. In some examples, content associated with a canvas can be shareable via another virtual space, such that data associated with the canvas is accessible to and/or rendered interactable for members of the virtual space.” It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of SUBRAMANIAN in view of HARAZI a terminal device and a conference device capable of connecting to each other and display an image of the text. The benefit and motivation of such modification is discussed by MAURER in the following portion: MAURER [0065] “FIG. 2A illustrates a user interface 200 of a group-based communication system, which will be useful in illustrating the operation of various examples discussed herein. The group-based communication system may include communication data such as messages, queries, files, mentions, users or user profiles, interactions, tickets, channels, applications integrated into one or more channels, conversations, workspaces, or other data generated by or shared between users of the group-based communication system. In some instances, the communication data may comprise data associated with a user, such as a user identifier, channels to which the user has been granted access, groups with which the user is associated, permissions, and other user-specific information.” Summary of references cited SUBRAMANIAN is main reference for emotion text to speech, it teaches selecting portions of text to apply emotion, converting text message from user and emotion information into voice message modulated by the emotion information. HARAZI teaches the graphical user interface (GUI) of a device, it also teaches emotion-based text to speech and a no-selection automated emotion-based text to speech of a message, that provides an emotion for the entire message instead of a section. BRUNS discloses another GUI more focus on displaying text/audio conversation between the main user and a second user as well as display text box and send/reply button and functionality, even though BRUNS also teach a modified text to speech, the rejection does not rely upon their text-to-speech because of their use for a goal different than the current application. Finally, MAURER teaches a device/system focused on a conference/teleconferencing as well as a terminal device interacting with the device. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to HECTOR J. CRESPO FEBLES whose telephone number is (571)272-4512 and email is hcrespofebles@uspto.gov. The examiner can normally be reached Mon - Fri 7:30 - 5:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at (571) 272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /HECTOR J. CRESPO FEBLES/ Examiner, Art Unit 2657 /DANIEL C WASHBURN/ Supervisory Patent Examiner, Art Unit 2657
Read full office action

Prosecution Timeline

Feb 27, 2025
Application Filed
Aug 26, 2026
Non-Final Rejection mailed — §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
100%
Grant Probability
99%
With Interview (+0.0%)
2y 1m (~6m remaining)
Median Time to Grant
Low
PTA Risk
Based on 1 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month