Prosecution Insights
Last updated: October 04, 2026
Application No. 17/531,861

SYSTEM WITH SPEAKER REPRESENTATION, ELECTRONIC DEVICE AND RELATED METHODS

Final Rejection §103
Filed
Nov 22, 2021
Priority
Nov 27, 2020 — DK PA202070795
Examiner
ZHU, RICHARD Z
Art Unit
2654
Tech Center
2600 — Communications
Assignee
GN Audio A/S
OA Round
8 (Final)
69%
Grant Probability
Favorable
9-10
OA Rounds
0m
Est. Remaining
85%
With Interview

Examiner Intelligence

Grants 69% — above average
69%
Career Allowance Rate
509 granted / 734 resolved
+7.3% vs TC avg
Strong +16% interview lift
Without
With
+15.7%
Interview Lift
resolved cases with interview
Typical timeline
3y 3m
Avg Prosecution
25 currently pending
Career history
765
Total Applications
across all art units

Statute-Specific Performance

§101
13.1%
-26.9% vs TC avg
§103
59.7%
+19.7% vs TC avg
§102
20.5%
-19.5% vs TC avg
§112
4.4%
-35.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 734 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Acknowledgement Acknowledgement is made of applicant’s amendment made on 08/05/2026. Applicant’s submission filed has been entered and made of record. Status of the Claims Claims 1-3, 6-12, 14, 16-20, and 22 are pending. Response to Applicant’s Arguments In response to “Applicant further respectfully submits that, because Costello nowhere discloses determination of a speaker's appearance based on analysis of an audio signal, and instead discloses determination of solely speaker sentiment, emotional state and mental state based on analysis of an audio signal” and “Applicant further respectfully submits that, because Costello, for at least the reasons given above, neither teaches nor suggests the feature of currently amended claim 1 of "determining a first primary appearance metric", it necessarily follows that Costello neither teaches nor suggests the features of currently amended claim 1 of "determining a first speaker representation based on the first primary sentiment metric and the first primary appearance metric" (underline added for emphasis) and of "determining a first speaker representation" and of "outputting a first speaker representation" (underline added for emphasis)”. Costello discloses a visualization program 285 that generates an icon 316a-316c representing the facial expression of the individual and whether that individual is speaking at a given time based on analytics generated by the trained neuromorphic processor, which were analyzed by a sentiment analysis program 275 (¶32). Note that the neuromorphic processor contextualize specific sentiments with conversation parameters such as keywords, key phrases, conversation tone (¶34). Therefore, it appears that Costello uses at least a speaker tone feature / conversation tone to determine individual’s facial expression (i.e., speaker’s appearance). In response to “Applicant respectfully notes that claim 1, as currently amended, following the suggestion of the Examiner in the July 23, 2026 interview, is further distinguishable from Costello for at least the reason that Costello, in disclosing icons and word labels representative of the sentiment, emotional state or mental state of a speaker, as shown in Fig. 3, reproduced above, neither teaches nor suggests the feature of currently amended claim 1 of "wherein determining a first speaker representation comprises generating a first avatar" (underline added for emphasis)” and “Applicant respectfully submits that, because Costello discloses, e.g., as shown in Fig. 3, reproduced above, only use of a generic icon to display a generic facial expression associated with a given sentiment of a speaker and "whether that individual is speaking at a given time", or use of a word label to represent a given sentiment of a speaker, see, e.g., Fig. 3 and para. [0032], and because neither an icon nor a word label corresponds to an avatar, which would be commonly understood to mean a graphic representation of the appearance of a specific individual, Costello neither teaches, let alone suggests, the feature of currently amended claim 1 of "wherein determining a first speaker representation comprises generating a first avatar" (underline added for emphasis)”. According to Wikipedia: “Avatars can be two-dimensional icons in Internet forums and other online communities, where they are also known as profile pictures (pfps), userpics, or formerly picons (personal icons, or possibly "picture icons"). Alternatively, an avatar can take the form of a three-dimensional model, as used in online worlds and video games, or an imaginary character with no graphical appearance”.1 Therefore, the plain meaning of “avatar” may include icons 316a-316c generated for the customers in Fig. 3 of Costello. Perhaps the applicant can further distinguish how avatar generated by the invention differ from the picture icons of Costello. In response to “Applicant respectfully notes that claim 1, as currently amended, following the suggestion of the Examiner in the July 23, 2026 interview, is further distinguishable from Costello for at least the reason that Costello, in disclosing determination of solely sentiments, emotional states and mental states based on analysis of an audio signal, neither teaches nor suggests the feature of currently amended claim 1 that "the first primary appearance metric [indicative of a primary appearance of the first speaker] is a gender, weight, height, age, language, language capability, hearing capability, dialect, health, personality, or understanding capability metric" (underline added for emphasis)” and “Applicant notes that, though Dawson discloses identification of the sentiment of a plurality of participants within a physical area (para. [0004]) using "sentiment identifier 56" based on input "from audio sensor 76" (see para. [0053]), for at least the reason that Dawson does not disclose determination and representation of speaker appearance based on analysis of an audio signal, Dawson does not teach, let alone suggest, the features of currently amended claim 1 that are neither taught nor suggested by Costello: "determining one or more appearance metrics indicative of an appearance of a first speaker", including "a first primary appearance metric indicative of a primary appearance of the first speaker", where "determining one or more first appearance metrics comprises extracting one or more speaker features from the first audio signal" (underline added for emphasis); "the first primary appearance metric [indicative of a primary appearance of the first speaker] is a gender, weight, height, age, language, language capability, hearing capability, dialect, health, personality, or understanding capability metric"; and, "wherein determining a first speaker representation comprises generating a first avatar"”. In view of such amendment to claims 1, 14, and 20, rejection under Costello and Dawson has been withdrawn. Upon further search and consideration, please see details of a new combination of references set forth below. Claim Rejections - 35 USC § 103 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 103 that form the basis for the rejections under this section made in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-3, 10-11, 14, 16-17, 19-20, and 22 are rejected under 35 USC 103(a) as being unpatentable over Costello et al. (US 2018/0122368 A1) in view of Dawson et al. (US 2020/0159487 A1) and Lie et al., (SCREAM: SCREen-based NavigAtion in voice Messages). Regarding Claims 1 and 14, Costello discloses electronic device (Fig. 2 and see ¶27, device 240 being an example of smartphone 140) comprising a processor (¶24, smartphone 140 includes processing circuit; ¶27, device 240 with neuromorphic processor), a memory (¶24, smartphone 140 having internal memory), and an interface (¶27, network connection for device 240 / smartphone 140), wherein the processor is configured to: obtain one or more audio signals including a first audio signal (¶25 and ¶30, receive audio and video data in a multiparty communication), wherein the first audio signal is obtained from a call or conversation between a user of the electronic device and a caller, wherein the user is an agent (¶¶37-38, in a call center, guiding a user / speaker based on analyzing speech of the customer to make recommendations to the user / speaker); determine, substantially in real time with a delay of less than 10 seconds (¶22, “real time” means within a portion of a second or within a few seconds; ¶25 and ¶32, real-time speech recognition capabilities), one or more first sentiment metrics indicative of a first speaker state based on the first audio signal, the one or more first sentiment metrics including a first primary sentiment metric indicative of a primary sentiment state of the first speaker (¶30, provide pattern analysis of the audio and the video to produce analytics for determining sentiments of each speaker of the communication), wherein the first speaker is the caller (¶38, analyzing the speech of the customer), wherein to determine one or more first sentiment metrics comprises to extract one or more speaker features from the first audio signal (¶34, neuromorphic processor utilizing machine learning and deep learning to recognize participant’s sentiments through speech analysis and to contextualize specific sentiments with conversation parameters including but not limited to keywords, key phrases, and specific conversation phrases); determine, substantially in real time with a delay of less than 10 seconds (¶22, “real time” means within a portion of a second or within a few seconds; ¶25 and ¶32, real-time speech recognition capabilities), one or more first appearance metrics indicative of an appearance of the first speaker based on the first audio signal, the one or more first appearance metrics including a first primary appearance metric indicative of a primary appearance of the first speaker, wherein to determine one or more first appearance metrics indicative of an appearance of the first speaker comprises to extract one or more speaker appearance features from the first audio signal (¶34, apply machine learning and deep learning to recognize participants’ sentiments through speech analysis and to contextualize the detected sentiments with the content of the conversation on the multiparty communication; ¶32, visualization program 285 generated an icon representing facial expression of the individual and whether that individual is speaking at a given time based on sentiment analysis), and wherein the one or more speaker appearance features comprise one or more of: a speaker tone feature (¶34, contextualize specific sentiments with conversation parameters including conversation tone to provide real-time insights into the emotional state of other participants), a speaker intonation feature, a speaker power feature, a speaker pitch feature, a speaker voice quality feature, a speaker rate feature, an acoustic feature, and a speaker spectral band energy feature; determine, substantially in real time with a delay of less than 10 seconds (¶22, “real time” means within a portion of a second or within a few seconds; ¶23 and ¶34, provide conference participants with real-time insights into the emotional state of other participants; ¶36, collect and analyze audio in the multiparty communication and ultimately provide updates to the visualization in real time), a first speaker representation based on the first primary sentiment metric and the first primary appearance metric (¶21, determine the sentiments of one or more participants based on the analytics and generate a visual representation of the sentiments; ¶32 and Fig. 3, visualization program 285 generates icon 316a-316c representing facial expression of individual), wherein determining a first speaker representation comprises generating a first avatar (¶32, visualization program 285 generated an icon representing facial expression of the individual and whether that individual is speaking at a given time based on sentiment analysis);2 and output, substantially in real time with a delay of less than 10 seconds (¶22, “real time” means within a portion of a second or within a few seconds; ¶23 and ¶34, provide conference participants with real-time insights into the emotional state of other participants), via the interface, the first speaker representation (¶21 and ¶36, display the sentiments during the course of the communication; Fig. 3 and ¶¶30-31, generate a graphical representation of the sentiments and display the representation on a device GUI), wherein the first speaker representation corresponds to the caller (¶38, call center provides guidance and analysis of a caller’s sentiments and mental state to the speaker / agent to de-escalate a complaint). Costello does not disclose wherein the one or more speaker features comprise paralinguistic features, and wherein the one or more first sentiment metrics are based on the paralinguistic features. Dawson teaches an electronic device comprising a processor, a memory, and an interface (¶21, computer system 12 comprising processor 16, memory 28, interface 22 such as a hand held or laptop device) configured to implement a sentiment identifier using vocal recognition techniques to analyze an audio feed from audio sensor to extract one or more paralinguistic speaker features from the audio feed and determine one or more sentiment metrics based on the one or more paralinguistic speaker features (¶53, sentiment identifier 56 uses vocal recognition techniques to analyze audio feed to identify paralanguage elements expressed by participants and to assign element values indicative of a participant’s mood). It would’ve been obvious to one ordinarily skilled in the art before the effective filing date of the invention to determine the one or more first sentiment metrics based on extracting paralinguistic features by weighing particular words or phrases spoken by the caller (Costello, ¶34, conversation parameters comprising keywords and key phrases for contextualizing sentiments) based on recognized paralanguage behavior accompanying a word or phrase (Dawson, ¶53) in order to use vocal recognition techniques on the first audio signal to identify paralanguage elements expressed by the caller and to assign indications of the caller’s mood (Dawson, ¶53). Costello does not disclose wherein the first primary appearance metric is a gender, weight, height. age, language, language capability, hearing capability, dialect, health, personality, or understanding capability metric. Lie teaches an electronic device obtaining a first audio signal from a caller (Fig. 2, speech to audio sampler): PNG media_image1.png 431 317 media_image1.png Greyscale determining one or more first sentiment metrics indicative of a first speaker state based on the first audio signal, the one or more first sentiment metrics including a first primary sentiment metric indicative of a primary sentiment state of a first speaker, wherein the first speaker is the caller, wherein determining one or more first sentiment metrics comprises extracting one or more speaker features from the first audio signal (p. 403, “However, monotone speech is often perceived as sad, while the pitch of the excited caller typically will vary. The same is assumed to be true for the energy of the audio signal. Calculating the excitement factor from these variations is then straightforward”); determining one or more first appearance metrics indicative of an appearance of the first speaker based on the first audio signal, the one or more first appearance metrics including a first primary appearance metric indicative of a primary appearance of the first speaker, wherein the first primary appearance metric is a gender (p. 403, “For a machine, the easier one to guess is probably gender / age, i.e., categorizing the caller as man, woman, or child”), weight, height, age (p. 403, 3 Caller characteristics, “For a machine, the easier one to guess is probably gender / age, i.e., categorizing the caller as man, woman, or child”), language, language capability, hearing capability, dialect, health, personality, or understanding capability metric, wherein the determining one or more first appearance metrics indicative of an appearance of the first speaker comprises extracting one or more speaker appearance features from the first audio signal, and wherein the one or more speaker appearance features comprise one or more of: a speaker tone feature, a speaker intonation feature, a speaker power feature, a speaker pitch feature (Fig. 2 shows speech analysis comprising energy estimator and pitch detector to detect pitch and energy for detecting caller characteristics), a speaker voice quality feature, a speaker rate feature, an acoustic feature, and a speaker spectral band energy feature (p. 402, 2. Speech analysis, analyze speech by transforming speech waveform (Fig. 3a) into the frequency domain as spectrogram, which indicates how much energy the signal contains in the various frequencies (i.e., spectral bands) as a function of time); determining a first speaker representation based on the first primary sentiment metric (Fig. 5: The wide open eye and mouth express excitement) and the first primary appearance metric (Fig. 5: the chroma of the frame determines the gender / age category), wherein determining a first speaker representation comprises generating a first avatar (p. 403, 4 Rendering images, Fig. 5, rendering module in Fig. 2 visualizes sound to generate an image of the caller); and outputting via the interface of the electronic device, the first speaker representation, wherein the first speaker representation corresponds to the caller (p. 404, Fig. 6; image of the caller rendered by rendering module of Fig. 2). It would’ve been obvious to one ordinarily skilled in the art before the effective filing date of the invention to determine first primary appearance metric comprising at least gender and age by extracting speaker appearance features from the first audio signal comprising at least a speaker pitch feature and a speaker spectral band energy feature in order to analyze incoming voice messages to extract certain caller characteristics for a system to render and display images that help the recipient navigate the messages (Lie, Abstract). Regarding Claim 2, Costello discloses wherein the one or more first sentiment metrics includes a first secondary sentiment metric indicative of a secondary sentiment state of the first speaker (Costello ¶34, apply machine learning speech analysis and to contextualize specific sentiments with conversation parameters to recognize participants’ sentiments; compare Dawsob ¶51, sentiment identifier 56 captures paralanguage sentiments such as mood, speaking style, rhythm, stress). Regarding Claim 3, Costello discloses wherein the one or more first appearance metrics includes a first secondary appearance metric indicative of a secondary appearance of the first speaker (¶34, conversation parameters not limited to conversation tone). Regarding Claim 10, Costello discloses wherein obtaining one or more audio signals comprises obtaining a second audio signal (¶30, receive audio and video data, perform pattern analysis of the audio and video data to produce analytics, utilize the analytics to determine the sentiment of each speaker in a multiparty communications session; i.e., this requires obtaining audio data of at least a second speaker); the method comprising: determining one or more second sentiment metrics indicative of a second speaker state based on the second audio signal, the one or more second sentiment metrics including a second primary sentiment metric indicative of a primary sentiment state of a second speaker (¶34, apply machine learning and deep learning to recognize participants’ sentiments through speech analysis and to contextualize the detected sentiments with the content of the conversation on the multiparty communication); obtaining one or more second appearance metrics indicative of an appearance of the second speaker, the one or more second appearance metrics including a second primary appearance metric indicative of a primary appearance of the second speaker (¶34, apply machine learning and deep learning to recognize participants’ sentiments through speech analysis and to contextualize the detected sentiments with the content of the conversation on the multiparty communication); determining a second speaker representation based on the second primary sentiment metric and the second appearance metric (¶21, determine the sentiments of one or more participants based on the analytics and generate a visual representation of the sentiments; ¶32 and Fig. 3, visualization program 285 generates icon 316a-316c representing facial expression of individual); and outputting, via the interface of the electronic device, the second speaker representation (¶21 and ¶36, display the sentiments during the course of the communication; Fig. 3 and ¶¶30-31, generate a graphical representation of the sentiments and display the representation on a device GUI). Regarding Claim 11, Costello discloses wherein the second speaker representation is an agent representation (¶30, determine sentiments of each speaker; ¶38, the speaker at the call center de-escalating a complaint from caller / customer). Regarding Claim 16, Costello discloses wherein the electronic device is selected from the group consisting of a mobile phone, a laptop computer, and a table computer (¶24, smartphone 140). Regarding Claim 17, Costello discloses wherein the interface comprises a display (¶30, device 240 GUI). Regarding Claim 19, Costello discloses wherein to obtain the one or more audio signals comprises to generate the one or more audio signals (¶42, receive communication data comprising audio data generated by a multiparty communication). Regarding Claim 20, Costello discloses system (Fig. 7) comprising: a server device (Fig. 7, cloud computing node 10 being server 12 of Fig. 6); and an electronic device in communication with the server device (Fig. 7, computers 54 communicating with cloud computing node 10, computer 54 being cellular telephone 54 corresponding to the smartphone 140 / device 240 described in ¶24 and ¶27), the electronic device comprising a processor, a memory, and an interface (¶24 and ¶27, e.g., smartphone 140 including processing circuit (e.g., an example device 240 with neuromorphic processor) executing program code stored on an internal memory with network connection), wherein the processor is configured to: obtain one or more audio signals including a first audio signal (¶25 and ¶30, receive audio and video data in a multiparty communication), wherein the first audio signal is obtained from a call or conversation between a user of the electronic device and a caller, wherein the user is an agent (¶¶37-38, in a call center, guiding a user / speaker based on analyzing speech of the customer to make recommendations to the user / speaker); determine, substantially in real time with a delay of less than 10 seconds (¶22, “real time” means within a portion of a second or within a few seconds; ¶25 and ¶32, real-time speech recognition capabilities), one or more first sentiment metrics indicative of a first speaker state based on the first audio signal, the one or more first sentiment metrics including a first primary sentiment metric indicative of a primary sentiment state of the first speaker (¶30, provide pattern analysis of the audio and the video to produce analytics for determining sentiments of each speaker of the communication), wherein the first speaker is the caller (¶38, analyzing the speech of the customer), wherein to determine one or more first sentiment metrics comprises to extract one or more speaker features from the first audio signal (¶34, neuromorphic processor utilizing machine learning and deep learning to recognize participant’s sentiments through speech analysis and to contextualize specific sentiments with conversation parameters including but not limited to keywords, key phrases, and specific conversation phrases); determine, substantially in real time with a delay of less than 10 seconds (¶22, “real time” means within a portion of a second or within a few seconds; ¶25 and ¶32, real-time speech recognition capabilities), one or more first appearance metrics indicative of an appearance of the first speaker based on the first audio signal, the one or more first appearance metrics including a first primary appearance metric indicative of a primary appearance of the first speaker, wherein to determine one or more first appearance metrics indicative of an appearance of the first speaker comprises to extract one or more speaker appearance features from the first audio signal (¶34, apply machine learning and deep learning to recognize participants’ sentiments through speech analysis and to contextualize the detected sentiments with the content of the conversation on the multiparty communication; ¶32, visualization program 285 generated an icon representing facial expression of the individual and whether that individual is speaking at a given time based on sentiment analysis), and wherein the one or more speaker appearance features comprise one or more of: a speaker tone feature (¶34, contextualize specific sentiments with conversation parameters including conversation tone to provide real-time insights into the emotional state of other participants), a speaker intonation feature, a speaker power feature, a speaker pitch feature, a speaker voice quality feature, a speaker rate feature, a linguistic feature, an acoustic feature, and a speaker spectral band energy feature; determine, substantially in real time with a delay of less than 10 seconds (¶22, “real time” means within a portion of a second or within a few seconds; ¶23 and ¶34, provide conference participants with real-time insights into the emotional state of other participants; ¶36, collect and analyze audio in the multiparty communication and ultimately provide updates to the visualization in real time), a first speaker representation based on the first primary sentiment metric and the first primary appearance metric (¶21, determine the sentiments of one or more participants based on the analytics and generate a visual representation of the sentiments; ¶32 and Fig. 3, visualization program 285 generates icon 316a-316c representing facial expression of individual), wherein determining a first speaker representation comprises generating a first avatar (¶32, visualization program 285 generated an icon representing facial expression of the individual and whether that individual is speaking at a given time based on sentiment analysis);3 and output, substantially in real time with a delay of less than 10 seconds (¶22, “real time” means within a portion of a second or within a few seconds; ¶23 and ¶34, provide conference participants with real-time insights into the emotional state of other participants), via the interface, the first speaker representation (¶21 and ¶36, display the sentiments during the course of the communication; Fig. 3 and ¶¶30-31, generate a graphical representation of the sentiments and display the representation on a device GUI), wherein the first speaker representation corresponds to the caller (¶38, call center provides guidance and analysis of a caller’s sentiments and mental state to the speaker / agent to de-escalate a complaint). Costello does not disclose wherein the one or more speaker features comprise paralinguistic features, and wherein the one or more first sentiment metrics are based on the paralinguistic features. Dawson teaches an electronic device comprising a processor, a memory, and an interface (¶21, computer system 12 comprising processor 16, memory 28, interface 22 such as a hand held or laptop device) configured to implement a sentiment identifier using vocal recognition techniques to analyze an audio feed from audio sensor to extract one or more paralinguistic speaker features from the audio feed and determine one or more sentiment metrics based on the one or more paralinguistic speaker features (¶53, sentiment identifier 56 uses vocal recognition techniques to analyze audio feed to identify paralanguage elements expressed by participants and to assign element values indicative of a participant’s mood). It would’ve been obvious to one ordinarily skilled in the art before the effective filing date of the invention to determine the one or more first sentiment metrics based on extracting paralinguistic features by weighing particular words or phrases spoken by the caller (Costello, ¶34, conversation parameters comprising keywords and key phrases for contextualizing sentiments) based on recognized paralanguage behavior accompanying a word or phrase (Dawson, ¶53) in order to use vocal recognition techniques on the first audio signal to identify paralanguage elements expressed by the caller and to assign indications of the caller’s mood (Dawson, ¶53). Costello does not disclose wherein the first primary appearance metric is a gender, weight, height. age, language, language capability, hearing capability, dialect, health, personality, or understanding capability metric. Lie teaches an electronic device obtaining a first audio signal from a caller (Fig. 2, speech to audio sampler): PNG media_image1.png 431 317 media_image1.png Greyscale determining one or more first sentiment metrics indicative of a first speaker state based on the first audio signal, the one or more first sentiment metrics including a first primary sentiment metric indicative of a primary sentiment state of a first speaker, wherein the first speaker is the caller, wherein determining one or more first sentiment metrics comprises extracting one or more speaker features from the first audio signal (p. 403, “However, monotone speech is often perceived as sad, while the pitch of the excited caller typically will vary. The same is assumed to be true for the energy of the audio signal. Calculating the excitement factor from these variations is then straightforward”); determining one or more first appearance metrics indicative of an appearance of the first speaker based on the first audio signal, the one or more first appearance metrics including a first primary appearance metric indicative of a primary appearance of the first speaker, wherein the first primary appearance metric is a gender (p. 403, “For a machine, the easier one to guess is probably gender / age, i.e., categorizing the caller as man, woman, or child”), weight, height, age (p. 403, 3 Caller characteristics, “For a machine, the easier one to guess is probably gender / age, i.e., categorizing the caller as man, woman, or child”), language, language capability, hearing capability, dialect, health, personality, or understanding capability metric, wherein the determining one or more first appearance metrics indicative of an appearance of the first speaker comprises extracting one or more speaker appearance features from the first audio signal, and wherein the one or more speaker appearance features comprise one or more of: a speaker tone feature, a speaker intonation feature, a speaker power feature, a speaker pitch feature (Fig. 2 shows speech analysis comprising energy estimator and pitch detector to detect pitch and energy for detecting caller characteristics), a speaker voice quality feature, a speaker rate feature, an acoustic feature, and a speaker spectral band energy feature (p. 402, 2. Speech analysis, analyze speech by transforming speech waveform (Fig. 3a) into the frequency domain as spectrogram, which indicates how much energy the signal contains in the various frequencies (i.e., spectral bands) as a function of time); determining a first speaker representation based on the first primary sentiment metric (Fig. 5: The wide open eye and mouth express excitement) and the first primary appearance metric (Fig. 5: the chroma of the frame determines the gender / age category), wherein determining a first speaker representation comprises generating a first avatar (p. 403, 4 Rendering images, Fig. 5, rendering module in Fig. 2 visualizes sound to generate an image of the caller); and outputting via the interface of the electronic device, the first speaker representation, wherein the first speaker representation corresponds to the caller (p. 404, Fig. 6; image of the caller rendered by rendering module of Fig. 2). It would’ve been obvious to one ordinarily skilled in the art before the effective filing date of the invention to determine first primary appearance metric comprising at least gender and age by extracting speaker appearance features from the first audio signal comprising at least a speaker pitch feature and a speaker spectral band energy feature in order to analyze incoming voice messages to extract certain caller characteristics for a system to render and display images that help the recipient navigate the messages (Lie, Abstract). Regarding Claim 22, Costello discloses wherein the server device is a cloud server (¶44, cloud computing node 10). Claims 6-9 and 18 are rejected under 35 USC 103(a) as being unpatentable over Costello et al. (US 2018/0122368 A1) in view of Dawson et al. (US 2020/0159487 A1) and Lie et al., (SCREAM: SCREen-based NavigAtion in voice Messages) as applied to claims 1 and 14, in further view of Chen et al. (WO 2017092194 A1, see IP.com translation). Regarding claims 6-7 and 18, Costello does not disclose wherein determining the first speaker representation comprises determining a first primary feature of a first avatar based on the first primary sentiment metric, and wherein the first speaker representation comprises the first avatar. Chen discloses electronic device comprising a processor, a memory (p. 8, “Preferably, the device further comprises a storage unit…”), and an interface (p. 7, “an apparatus for animating a communication interface during communication, the apparatus being located in the terminal device of both users of the communication relationship”) configured to: obtain audio signals from a call or conversation between a user of the electronic device and a caller (p. 9, “Step 11: Receive content generated by the user based on the communication during the communication. In this step, first, a communication relationship is established between the users, and the communication relationship may be established by performing a dial-up call or by using an instant messaging software for voice, text, or video communication”; p. 10, “if the user communicates using a voice call, the voice recognition unit in the device identifies the voice content generated by the user during the call”), determine one or more first sentiment metrics indicative of a first speaker state based on the first audio signal, the one or more first sentiment metrics including a first primary sentiment metric indicative of a primary sentiment state of the first speaker (p. 10, “if the user communicates using a voice call, the voice recognition unit in the device identifies the voice content generated by the user during the call, and parses the semantics of the voice content, and then according to the semantics The corresponding state change instruction is generated”; e.g., p. 11, “the avatar on the communication interface changes the mouth shape, motion, shape, dress, body type, sound and/or expression of the avatar according to the state change instruction, for example, the user said during the voice call: " I am so happy. The expression of the avatar corresponding to the user on the interface starts to become a laugh according to the content of the call”; i.e., determine user’s sentiment is “happy” based on semantic analysis of the voice content), determine one or more first appearance metrics indicative of an appearance of the first speaker based on the first audio signal, the one or more first appearance metrics including a first primary appearance metric indicative of a primary appearance of the first speaker (p. 11, “The first way is: the avatar on the communication interface changes the mouth shape, motion, shape, dress, body type, sound and/or expression of the avatar according to the state change instruction…” and “For example, the user said in the process of voice call: "Now I suddenly feel It’s so cold. The user’s avatar dress changes on the interface. Maybe the avatar was originally dressed in summer clothes, and now it’s thick clothes in winter”; i.e., determine user’s avatar appearance with thick winter cloth based on semantic analysis of the voice content); determine, during the conversation, a first speaker representation based on the first primary sentiment metric and the first primary appearance metric (p. 7, “Preferably, the modifying the state of the user avatar on the communication interface according to the state change instruction comprises: The state of the mouth shape, motion, shape, dress, body type, sound, and/or expression of the user avatar is modified according to the state change command” and “The user avatar and/or virtual scene is an avatar and/or a virtual scene created based on the real image of the user and/or the real scene in which the user is located”); wherein determining the first speaker representation comprises determining a first primary feature of a first avatar based on the first primary sentiment metric (p. 7, “Preferably, the user avatar and/or the virtual scene on the communication interface comprises an avatar and/or a virtual scene of the communication parties, the communication content being the communication content of the first party of the communication parties, according to the state The change instruction modifies the state of the user avatar and/or the virtual scene on the communication interface to cause the communication interface to produce an animation effect”), and wherein the first speaker representation comprises the first avatar (p. 7, “And changing a state of the first party's avatar and/or the virtual scene on the communication interface according to the state change instruction”); wherein the first primary feature is selected from a mouth feature (p. 7, “The state of the mouth shape, motion, shape, dress, body type, sound, and/or expression of the user avatar is modified according to the state change command”), an eye feature (Fig. 5-2, relevant animation), a nose feature (Fig. 5-2, relevant animation), a forehead feature (Fig. 5-2, relevant animation), an eyebrow feature (Fig. 5-2, relevant animation), a hair feature (Fig. 5-2, relevant animation), an ear feature (Fig. 5-2, relevant animation), a beard feature, a gender feature (Fig. 5-2, relevant animation), a cheek feature (Fig. 5-2, relevant animation), an accessory feature (p. 11, “For example, the user said in the process of voice call: "Now I suddenly feel It’s so cold. The user’s avatar dress changes on the interface. Maybe the avatar was originally dressed in summer clothes, and now it’s thick clothes in winter”), a skin feature (Fig. 5-2, relevant animation), a body feature (p. 7, “The state of the mouth shape, motion, shape, dress, body type, sound, and/or expression of the user avatar is modified according to the state change command”), and a head dimension feature (Fig. 5-2, relevant animation). It would’ve been obvious to one ordinarily skilled in the art before the effective filing date of the invention to determine the first speaker representation comprising a first avatar with a first primary feature based on the first primary sentiment metric in order to enable a corresponding animation effect according to content of communication between users (Chen, Abstract). Regarding Claims 8-9, Costello does not disclose wherein determining the first speaker representation comprises determining a first secondary feature of a first avatar based on the first primary appearance metric. Chen discloses wherein determining the first speaker representation comprises determining a first secondary feature of a first avatar based on the first primary appearance metric (pp. 11-12, “For example, the user said in the process of voice call: "Now I suddenly feel It’s so cold. The user’s avatar dress changes on the interface. Maybe the avatar was originally dressed in summer clothes, and now it’s thick clothes in winter. The avatar action may become shivering”); wherein the first secondary feature is different from a first primary feature of the first avatar (pp. 11-12, “For example, the user said in the process of voice call: "Now I suddenly feel It’s so cold. The user’s avatar dress changes on the interface. Maybe the avatar was originally dressed in summer clothes, and now it’s thick clothes in winter. The avatar action may become shivering”), wherein the first primary feature is based on the first primary sentiment metric (p. 11, “for example, the user said during the voice call: "I am so happy. The expression of the avatar corresponding to the user on the interface starts to become a laugh according to the content of the call”), and wherein the first secondary feature is selected from a mouth feature (p. 7, “The state of the mouth shape, motion, shape, dress, body type, sound, and/or expression of the user avatar is modified according to the state change command”), an eye feature (Fig. 5-2, relevant animation), a nose feature (Fig. 5-2, relevant animation), a forehead feature (Fig. 5-2, relevant animation), an eyebrow feature (Fig. 5-2, relevant animation), a hair feature (Fig. 5-2, relevant animation), an ear feature (Fig. 5-2, relevant animation), a beard feature, a gender feature (Fig. 5-2, relevant animation), a cheek feature (Fig. 5-2, relevant animation), an accessory feature (p. 7, “The state of the mouth shape, motion, shape, dress, body type, sound, and/or expression of the user avatar is modified according to the state change command”), a skin feature (Fig. 5-2, relevant animation), a body feature (p. 7, “The state of the mouth shape, motion, shape, dress, body type, sound, and/or expression of the user avatar is modified according to the state change command”), and a head dimension feature (Fig. 5-2, relevant animation). It would’ve been obvious to one ordinarily skilled in the art before the effective filing date of the invention to determine the first speaker representation comprises determining a first secondary feature of a first avatar based on the first primary appearance metric in order to enable a corresponding animation effect according to content of communication between users (Chen, Abstract). Claim 12 is rejected under 35 USC 103(a) as being unpatentable over Costello et al. (US 2018/0122368 A1) in view of Dawson et al. (US 2020/0159487 A1) and Lie et al., (SCREAM: SCREen-based NavigAtion in voice Messages) as applied to claim 1, in further view of Kawakami et al. (US 2022/0232191 A1). Regarding Claim 12, Costello does not disclose detecting a termination of speech, and in accordance with detecting the termination of speech, storing a speaker record in the memory and/or transmitting a speaker record to a server device of the system, the speaker record comprising a first speaker record indicative of one or more of first appearance metric data and first sentiment metric data of the first speaker. Kawakami teaches a system generating avatar / speaker record for users engaging in conversation in virtual space (¶41 and ¶43), detecting a termination of speech / conversation (¶48, a trigger generator 162 generates a conversation end trigger corresponding to an end of started conversation by a user), and storing a speaker record / avatar data corresponding to a speaking user in a memory and transmitting the speaker record to a server device of the system (¶64, avatar arrangement information includes a display mode of the avatar; ¶69, in response to detection of conversation end trigger, avatar controller 333a updates avatar arrangement information and stores the arrangement data in storage 32; per Fig. 4 and ¶58, storage 32 being at a content distribution server 3). It would’ve been obvious to one ordinarily skilled in the art before the effective filing date of the invention to detect a termination of speech and storing a speaker record indicative of one or more of first appearance metric data and first sentiment metric data of the first speaker (compare Costello, Fig. 3 and ¶32, display mode / icon arrangement information includes facial expression icons and sentiments of the participants) in memory / transmitting the speaker record to server / cloud computing node 10 in accordance with detecting the termination of speech in order to faithfully reflect the operations of users (Kawakami ¶69). Conclusion Applicant's amendment necessitated the new grounds of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to examiner Richard Z. Zhu whose telephone number is 571-270-1587 or examiner’s supervisor Hai Phan whose telephone number is 571-272-6338. Examiner Richard Zhu can normally be reached on M-Th, 0730:1700. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /RICHARD Z ZHU/Primary Examiner, Art Unit 2654 08/21/2026 1 en.wikipedia.org/wiki/Avatar_(computing) 2 en.wikipedia.org/wiki/Avatar_(computing): “Avatars can be two-dimensional icons in Internet forums and other online communities, where they are also known as profile pictures (pfps), userpics, or formerly picons (personal icons, or possibly "picture icons")”. 3 en.wikipedia.org/wiki/Avatar_(computing): “Avatars can be two-dimensional icons in Internet forums and other online communities, where they are also known as profile pictures (pfps), userpics, or formerly picons (personal icons, or possibly "picture icons")”.
Read full office action

Prosecution Timeline

Show 18 earlier events
Jan 27, 2026
Response after Non-Final Action
Apr 24, 2026
Request for Continued Examination
Apr 26, 2026
Response after Non-Final Action
May 05, 2026
Non-Final Rejection mailed — §103
Jul 09, 2026
Interview Requested
Jul 23, 2026
Examiner Interview Summary
Aug 05, 2026
Response Filed
Aug 25, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12738263
Method and System for Constructing Speech Recognition Model and Speech Processing
2y 10m to grant Granted Sep 15, 2026
Patent 12699692
TECHNIQUES FOR EFFICIENT ENCODING IN NEURAL SEMANTIC PARSING SYSTEMS
2y 6m to grant Granted Aug 04, 2026
Patent 12694875
METHOD, DEVICE AND SYSTEM OF CONTEXTUAL AND USER LOCATION BASED OPERATIONAL EXECUTION VIA A GENERATIVE ARTIFICIAL INTELLIGENCE (AI) COMPUTING PLATFORM IN RESPONSE TO USER INTERACTION THEREWITH
2y 7m to grant Granted Jul 28, 2026
Patent 12657538
AUDIO SIGNAL PROCESSING AND DYNAMIC NATURAL LANGUAGE UNDERSTANDING
3y 5m to grant Granted Jun 16, 2026
Patent 12633287
DOMAIN MODEL DRIVEN PROCESSING OF DIALOG INCLUDING AMBIGUOUS INTENTS
2y 5m to grant Granted May 19, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

9-10
Expected OA Rounds
69%
Grant Probability
85%
With Interview (+15.7%)
3y 3m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 734 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month