Prosecution Insights
Last updated: October 01, 2026
Application No. 17/345,429

Situationally Aware Social Agent

Final Rejection §103§112
Filed
Jun 11, 2021
Examiner
HICKS, AUSTIN JAMES
Art Unit
2142
Tech Center
2100 — Computer Architecture & Software
Assignee
Disney Enterprises Inc.
OA Round
7 (Final)
75%
Grant Probability
Favorable
8-9
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 75% — above average
75%
Career Allowance Rate
315 granted / 420 resolved
+20.0% vs TC avg
Strong +26% interview lift
Without
With
+25.8%
Interview Lift
resolved cases with interview
Typical timeline
3y 2m
Avg Prosecution
49 currently pending
Career history
469
Total Applications
across all art units

Statute-Specific Performance

§101
13.3%
-26.7% vs TC avg
§103
54.0%
+14.0% vs TC avg
§102
15.8%
-24.2% vs TC avg
§112
14.4%
-25.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 420 resolved cases

Office Action

§103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant's arguments filed 6/22/2026 have been fully considered but they are not persuasive. Applicant argues, “(1) None of Scott, Peconsti, or King discloses, teaches, or suggest the limitation ‘audio-based venue data corresponding to an environment surrounding the user, the environment including a source of sound other than the user,’ as recited by independent claims 1 and 14.” Remarks 1. This is taught by Scott para 21 “In another example, a microphone may be employed to measure loudness of the environment, and change the system behavior prior to receiving any voice input, such as by showing a prompt on the screen that changes when someone walks closer to a reference point.” This loudness is inside of the broadest reasonable interpretation of the claim in light of the specification paragraph 55, “the situationally aware social agent may be configured to distinguish between user relevant environmental cues and background noise, and to advantageously filter out the latter to enhance sensitivity to the former…” Further, Scott teaches this in para 48-49 “when motion information (e.g., angle of arrival information) is available (e.g., from radar information), a beamforming estimate can be used to enhance speech recognition…. [0049] Also, the system (e.g., the client device 102) can disambiguate between multiple sound sources, such as by filtering out the position of a known noise-producing device (e.g., a television) or background noise. …” Applicant argues, “(2) Peconsti fails to disclose ‘at least one directional microphone,’ as recited by independent claims 1, 14 and 20, or such a directional microphone being ‘configured to collect audio data and identify an angle of arrival of the audio data,’ as recited by independent claim 1.” Remarks 2. Directional microphone is not a term of art, and it is not mentioned in the specification. Peconsti teaches determining the angle of arrival second paragraph “Thus, it is need to estimate the angle of arrival of a given sound. To achieve that, an electronic circuit to acquire signals from 2 microphones into the arduino was built.”, which is the only defining characteristic of Applicant’s claimed directional microphone. Applicant argues, “(3) The reliance by the Final Office Action on paragraph [0082] of King for ‘correlate the radar-based location data and the audio-based location data with the map of the venue,’ as recited by independent claim 1, and for ‘correlating the radar-based venue data and the audio-based venue data with the map of the venue,’ as recited by independent claim 20, is in error, because paragraph [0082] of King discloses no such correlation.” Remarks 3. The specification paragraph 55 says that this correlation may be, “the situationally aware social agent may be configured to distinguish between user relevant environmental cues and background noise, and to advantageously filter out the latter to enhance sensitivity to the former…” Scott teaches this in para 48-49 “when motion information (e.g., angle of arrival information) is available (e.g., from radar information), a beamforming estimate can be used to enhance speech recognition…. [0049] Also, the system (e.g., the client device 102) can disambiguate between multiple sound sources, such as by filtering out the position of a known noise-producing device (e.g., a television) or background noise. …” Applicant argues, “(4) None of the cited references discloses ‘the interactive expression includes context from the portion of the audio data collected from the source of sound,’ as recited by independent claims 1 and 14.” Remarks 4. This claim is not a term of art and it is not defined in the specification. The specification paragraph 55 says that this interactive expression may be filtering background noise and only responding to the user, “the situationally aware social agent may be configured to distinguish between user relevant environmental cues and background noise, and to advantageously filter out the latter to enhance sensitivity to the former…” Scott teaches this in para 48-49 “when motion information (e.g., angle of arrival information) is available (e.g., from radar information), a beamforming estimate can be used to enhance speech recognition…. [0049] Also, the system (e.g., the client device 102) can disambiguate between multiple sound sources, such as by filtering out the position of a known noise-producing device (e.g., a television) or background noise. …” Scott paragraph 21 also teaches this, “In another example, a microphone may be employed to measure loudness of the environment, and change the system behavior prior to receiving any voice input, such as by showing a prompt on the screen that changes when someone walks closer to a reference point.” Applicant argues, “(5) None of the cited references discloses ‘processing the radar data to obtain radar-based venue data corresponding to the environment surrounding the user’ and ‘correlating the radar-based venue data and the audio-based venue data,’ as recited by independent claim 14.” Remarks 14. A prima facie case for this amended element was not made in the Final rejection 5/21/2026. Although the subject matter of this amended claim element was presented and properly rejected in claim 13 and 26. Therefore, this claim 14 would have been properly finally rejected in the Final rejection 5/21/2026 had the Final rejection specifically addressed the amended claim element. The specification paragraph 55 says that this interactive expression may be filtering background noise and only responding to the user, “the situationally aware social agent may be configured to distinguish between user relevant environmental cues and background noise, and to advantageously filter out the latter to enhance sensitivity to the former…” Scott teaches this in para 48-49 “when motion information (e.g., angle of arrival information) is available (e.g., from radar information), a beamforming estimate can be used to enhance speech recognition…. [0049] Also, the system (e.g., the client device 102) can disambiguate between multiple sound sources, such as by filtering out the position of a known noise-producing device (e.g., a television) or background noise. …” Claim Rejections - 35 USC § 112 The following is a quotation of the first paragraph of 35 U.S.C. 112(a): (a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention. Claims 1, 4-11, 13, 14, 17-24 and 26-31 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention. In claims 1, 14 and 20, Applicant claims a “directional microphone”. This is not described in the specification. In claims 1 and 14, Applicant claims, “wherein the interactive expression includes context from the portion of the audio data collected from the source of sound.” In the specification, Applicant discusses context of the neural network, which does not describe this claim element. Specification paragraph 41 states, “ such correlation of the radar-based location data with the audio-based location data further advantageously enables system 100 to disregard or make us of additional context from any background sounds in the environment(s) of user(s) 128 a/128 b, as well to use the correlation as a metric for aiming microphone(s) 154 a/154 b or microphone(s) 235 or radar 152 a/152 b or radar detector 231, or for re-engagement with the conversation at a later time.” Specification paragraph 55 states, “the situationally aware social agent may be configured to distinguish between user relevant environmental cues and background noise, and to advantageously filter out the latter to enhance sensitivity to the former, thereby enabling interactions with the users that are natural, engaging, and attentive to the environmental context or “situation” in which the interaction is occurring.” None of these sections of the specification describe what is now claimed. Therefore, the claim element lacks written description in the specification. Claims 1, 4-11, 13, 14, 17-19, 21-24 and 26-31 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, because the specification, while being enabling for “the situationally aware social agent may be configured to distinguish between user relevant environmental cues and background noise, and to advantageously filter out the latter to enhance sensitivity to the former, thereby enabling interactions with the users that are natural, engaging, and attentive to the environmental context or “situation” in which the interaction is occurring” (spec. 55), does not reasonably provide enablement for “wherein the interactive expression includes context from the portion of the audio data collected from the source of sound” (claim 14), and “the interactive expression includes context from the portion of the audio data collected from the source of sound”. Claim 1. The specification does not enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make the invention commensurate in scope with these claims. Claim 20 does not include this limitation. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 4, 10, 11, 13 and 26-29 are rejected under 35 U.S.C. 103 as being unpatentable over US20170289766A1 to Scott et al, Yet another arduino sound localizer (2 microphones, angle of arrival) by Peconsti and US20180227694A1 to King. Claims 5 and 6 are rejected under 35 U.S.C. 103 as being unpatentable over US20170289766A1 to Scott et al, Yet another arduino sound localizer (2 microphones, angle of arrival) by Peconsti, US20180227694A1 to King and US20200320427A1 to Kennedy et al. Claims 7 and 9 are rejected under 35 U.S.C. 103 as being unpatentable over US20170289766A1 to Scott et al, Yet another arduino sound localizer (2 microphones, angle of arrival) by Peconsti, US20180227694A1 to King and US20200184306A1 to Buhman et al. Claims 8 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over US20170289766A1 to Scott et al, Yet another arduino sound localizer (2 microphones, angle of arrival) by Peconsti, US20180227694A1 to King, US20200184306A1 to Buhman et al and US20190198006A1 to Guo et al. Claims 14, 17, 23, 24, 30 and 31 are rejected under 35 U.S.C. 103 as being unpatentable over US20170289766A1 to Scott et al and Yet another arduino sound localizer (2 microphones, angle of arrival) by Peconsti. Claims 18 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over US20170289766A1 to Scott et al, Yet another arduino sound localizer (2 microphones, angle of arrival) by Peconsti and US20200320427A1 to Kennedy et al. Claims 21 and 22 are rejected under 35 U.S.C. 103 as being unpatentable over US20170289766A1 to Scott et al, Yet another arduino sound localizer (2 microphones, angle of arrival) by Peconsti and US20200184306A1 to Buhman et al. Scott teaches Claim 1. A system comprising: a processing hardware; (Scott fig. 1 [104]) a memory storing a software code, (Scott fig. 1 [106]) a radio detection and ranging (radar) or a radar detector configured to collect radar data; (Scott para 43 “The physical presence of people (i.e. people nearby the system) may be detected using sensors 132 like…microwave radar…”) at least one (Scott, Paragraph [0043], “The physical presence of people (i.e. people nearby the system) may be detected using sensors 132 like…microwave radar, microphones or cameras,… modalities like radar can provide more fine-grained information, that can include a positioning element (e.g. x/y/z position relative to the PC)”) a social agent instantiated as a robot, a virtual character, or a tabletop or wall-mounted device, the social agent comprising an output unit configured to effectuate an interactive expression of the social agent, (Scott fig. 1 client device includes several tabletop device with screens and “first digital assistant” (Scott para 106) that are equivalent to Applicant’s social agent. The interactive expression is taught in the transition between experiences in Scott para 106, “the first digital assistant experience is adapted to generate a second digital assistant experience at the client device that is based on a difference between the first detected distance and the second detected distance (block 706).”) the processing hardware configured to execute the software code to: process the radar data to obtain radar-based location data corresponding to a location of a user within a venue; (Scott para 50 “Position: As noted above, radar or camera-based sensors 132 may provide a position for one or multiple users. The position is then used to infer context, e.g. approaching the client device 102, moving away from the client device 102, presence in a different room than the client device 102, and so forth.”) process the audio data (Scott para 47 “The physical presence of people (i.e. people nearby the system) may be detected using… microphones…” Scott para 21 “In another example, a microphone may be employed to measure loudness of the environment, and change the system behavior prior to receiving any voice input, such as by showing a prompt on the screen that changes when someone walks closer to a reference point.”) correlate the radar-based location data and the audio-based location data with the (Scott para 47 “he physical presence of people (i.e. people nearby the system) may be detected using sensors 132 … microphones or cameras, and using techniques such as Doppler radar …”)_ identify, based on the location of the user, the determined environment and the audio-based venue data, an interactive expression for use by the social agent to interact with the user, wherein the interactive expression is identified using a portion of the audio data collected from the source of sound other than the user; and (Scott para 21 “In another example, a microphone may be employed to measure loudness of the environment, and change the system behavior prior to receiving any voice input, such as by showing a prompt on the screen that changes when someone walks closer to a reference point.” The interactive expression is the prompt. The loudness of the environment is the portion of the audio collected from a source other than the user.) execute, using the output unit, the interactive expression used by the social agent, wherein the interactive expression includes context from the portion of the audio data collected from the source of sound. (Scott para 21 “In another example, a microphone may be employed to measure loudness of the environment, and change the system behavior prior to receiving any voice input, such as by showing a prompt on the screen that changes when someone walks closer to a reference point.” Showing the prompt on the screen is executing the identified expression.) Scott doesn’t teach directional microphone configured to collect audio data and identify an angle of arrival of the audio data. However, Peconsti teaches directional microphone configured to collect audio data and identify an angle of arrival of the audio data. (Peconsti second paragraph “Thus, it is need to estimate the angle of arrival of a given sound. To achieve that, an electronic circuit to acquire signals from 2 microphones into the arduino was built.” The two microphones are one directional microphone.) Scott, the claims and peonsti are all directed to microphone applications. It would have been obvious to a person having ordinary skill in the art, at the time of filing, to receive angle of arrival data “in order to achieve proper monitoring, the monitoring system should be able to attribute sounds to the specific [person]…” Peconsti first paragraph. Scott doesn’t teach a map of the venue. However, King teaches a memory storing a software code, and a map of a venue; (King para 55 “The database may be used by algorithms to present a display of a seating map of a specific venue…” King para 15 “obtaining spatial reference data for a specific venue. The method also includes creating a digital model of the specific venue.”) correlate the radar-based location data and the audio-based location data with the map of the venue to determine the environment surrounding the user within the venue; (King para 82 “navigation matrix of panoramic video and audio viewports that in a particular geographic location or venue;…” correlating radar data, audio and map data is taught by the navigation matrix.) King, Scott and the claims deal with detecting audio in different environments. It would have been obvious to a person having ordinary skill in the art, at the time of filing, to include spatial reference data of a specific venue in Scott because “the representation of the specific venue may also include a representation of a specific stage or other performance venue may be superimposed with graphical depiction of historical data related to the venue. In some embodiments such a representation may aid in a process of designing audio capture locations for a future spectator event.” King para 65. Improving audio quality based on the venue will improve the robots accuracy in collecting audio data. Scott teaches Claims 4 and 17 (Original): The system of claim 1, wherein the interactive expression comprises one of speech or text. (Scott para 21 “In another example, a microphone may be employed to measure loudness of the environment, and change the system behavior prior to receiving any voice input, such as by showing a prompt on the screen that changes when someone walks closer to a reference point.” Scott para 67 “very large characters can be displayed that provide simple messages and/or prompts, such as “Hello!,” “May I Help You?,” and so forth.” Scott para 65 “digital assistant 126 outputs an audio prompt…”) Scott teaches Claims 5 and 18 (Previously Presented): The system of claim 1, wherein the social agent is instantiated (Scott para 67 “very large characters can be displayed that provide simple messages and/or prompts, such as “Hello!,” “May I Help You?,” and so forth.”) Scott’s characters are different than applicant’s virtual character. However, Kennedy teaches the virtual character. (Kennedy para 16 “social agent may take the form of a virtual character rendered on a display…”) Kennedy, Scott and the claims are all directed to social agents. It would have been obvious to a person having ordinary skill in the art, at the time of filing, to have a virtual character relay Scott’s messages because devices without virtual characters “tend to lack character and naturalness…” Kennedy para 14. Scott teaches Claims 6 and 19 (Previously Presented): The system of claim 4, wherein the interactive expression (Scott para 67 displays a text expression.) Scott doesn’t teach a gesture etc. However, Kennedy teaches that expression comprises at least one of a gesture, a facial expression, or a posture. (Kennedy para 71 “where personality profile 646 b of the character assumed by interactive social agent 116 a or 116 b is that of an evil villain, the expression smile-smile-smile might be remapped to modified expression (sneer-sneer-sneer) 648.”) Kennedy, Scott and the claims are all directed to social agents. It would have been obvious to a person having ordinary skill in the art, at the time of filing, to have a virtual character relay Scott’s messages with a facial expression because devices without virtual characters “tend to lack character and naturalness…” Kennedy para 14. Scott teaches Claims 7 and 21 (Previously Presented): The system of claim 1, wherein the processing hardware is further configured to execute the software code to: recognize, (Scott para 52 ” Identity recognition can employ camera-based face recognition or more coarse-grained recognition techniques that approximate the identity of a user.” Scott para 49 “When the identity of a user is known (such as discussed below), it is possible to apply a different speech recognition model that actually fits the user's accent, language, acoustic speech frequencies, and demographic.”) Scott doesn’t teach the anonymous user history. However, Buhman teaches how to recognize, using the audio data, the user as an anonymous user with whom the social agent has previously interacted. (“virtual agent 150/350 is typically able to distinguish one anonymous human guest with whom a previous character interaction has occurred from another, as well as from anonymous human guests having no previous interaction experience with the character…” and “the presence of guest 126a/126b or guest object(s) 148 can be detected based on sensor data received from input module 130/230. That sensor data may also be used to reference interaction history database 108 to identify guest 126a/126b” where input module 130/230 contains a microphone to utilize audio data in (Buhmann, Drawings, Figure 2, 238).) The claims Buhmann as well as Scott are directed towards the field of interactive agents. It would have been obvious to a person having ordinary skill in the art, at the time of filing, to recognize anonymous users “in order to improve the performance of virtual agent…” Buhmann para 63. Scott teaches Claim 8 (Currently Amended): The system of claim 7, further comprising: a wherein the processing hardware is configured to execute the software code (Scott para 52 ” Identity recognition can employ camera-based face recognition or more coarse-grained recognition techniques that approximate the identity of a user.”) The Scott/Buhmann combination fails to explicitly teach recognizing a user using a trained machine learning model. However, Guo teaches a trained machine learning model. (Guo para 92 “these augmented features are then used to assess the probability that a particular word, phoneme, or phone was heard. In more modern systems, this computation is performed by a specially trained deep neural network.”) The claims, Guo, Scott and Buhman are all directed to interactive agents. It would have been obvious to a person having ordinary skill in the art, at the time of filing, to use a trained NN to recognize the user because the “neural networks or GMMs may optionally be trained for a specific individual to give improved results.” Guo para 92. Scott teaches Claims 9 and 22 (Previously Presented): The system of claim 1, wherein the memory further stores an interaction history interaction history of the social agent with the user, and wherein the processing hardware is configured to execute the software code to identify, further using the interaction history, the interactive expression for use by the social agent to interact with the user. (Scott, Paragraph [0095], “Different contextual factors are detailed throughout this discussion, and include information such as…interaction history with the digital assistant, and so forth.” And (Scott, [0018]) “the digital assistant can respond to queries, provide appropriate information, offer suggestions, adapt UI visualizations, and takes actions to assist the user depending on the context and sensor data…”) Scott doesn’t teach a database for holding the interaction history. However, Buhmann teaches an interaction history database. (Buhmann para 54 “identification of guest 126 a/126 b or guest object(s) 148 may be performed by software code 110, executed by hardware processor 104, and using input module 130/230 and interaction history database 108.”) The claims, Scott and Buhmann are directed towards the field of interactive agents. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of Buhmann with the teachings of Scott by performing interactions that take into account previous interactions with a user that are stored in a database. Buhmann provides as additional motivation for combination (Buhmann, Background, " conventional conversational agents…omit many of the properties that a real human would offer in an interaction that make that interaction not only informative but also enjoyable or entertaining. For example, an interaction between two humans may be influenced or shaped by the personalities of those human participants, as well as the history or storyline of their previous interactions. Thus, there remains a need in the art for a virtual agent capable of engaging in an extended interaction…"). Scott teaches Claims 10 and 23 (Previously Presented): The system of claim 1, wherein the processing hardware is further configured to execute the software code to: recognize, using at least one of the radar-based location data or the audio-based location data, a relocation of the user relative to the social agent. (Scott, Paragraph [0043], “The physical presence of people (i.e. people nearby the system) may be detected using sensors 132 like…microwave radar, microphones or cameras,… modalities like radar can provide more fine-grained information, that can include a positioning element (e.g. x/y/z position relative to the PC)” and (Scott, [0052]) “As the person approaches the system, indications such as icons, animations, and/or audible alerts may be output to signal that different types of interaction are active…”). Scott teaches Claims 11 and 24 (Previously Presented): The system of claim 1, wherein the processing hardware is further configured to execute the software code to enhance, using at least one of sound produced by or a data input received from the source of sound, a signal-to-noise ratio of the audio data. (Scott, Paragraph [0045], “the system (e.g., the client device 102) can disambiguate between multiple sound sources, such as by filtering out the position of a known noise- producing device (e.g., a television) or background noise…”) Scott teaches Claims 13 and 26 (Currently Amended): The system of claim 1, wherein the processing hardware is further configured to execute the software code to: process the radar data to obtain radar-based venue data corresponding to the environment surrounding the user, and recognize, using the radar-based venue data, the source of sound as an inanimate source of sound; or (Scott para 47 “The physical presence of people (i.e. people nearby the system) may be detected using sensors 132 like… microwave radar, microphones…” Scott para 49 “the system (e.g., the client device 102) can disambiguate between multiple sound sources, such as by filtering out the position of a known noise-producing device (e.g., a television) or background noise.”) recognize, using the (Scott para 49 “the system (e.g., the client device 102) can disambiguate between multiple sound sources, such as by filtering out the position of a known noise-producing device (e.g., a television) or background noise.”) Scott doesn’t teach a map of the venue. However, King teaches recognize, using the map of the venue, the source of sound (King para 82 “navigation matrix of panoramic video and audio viewports that in a particular geographic location or venue;…” correlating radar data, audio and map data is taught by the navigation matrix.) King, Scott and the claims deal with detecting audio in different environments. It would have been obvious to a person having ordinary skill in the art, at the time of filing, to include spatial reference data of a specific venue in Scott because “the representation of the specific venue may also include a representation of a specific stage or other performance venue may be superimposed with graphical depiction of historical data related to the venue. In some embodiments such a representation may aid in a process of designing audio capture locations for a future spectator event.” King para 65. Improving audio quality based on the venue will improve the robots accuracy in collecting audio data. Claim 14 (Currently Amended): A method for use by a system including a processing hardware and a memory storing a software code, the method comprising collecting, using a radio detection and ranging (radar) or a radar detector, radar data; (Scott para 47 “The physical presence of people (i.e. people nearby the system) may be detected using sensors 132 like… microwave radar, microphones…”) collecting, using at least one directional microphone, audio data; (Scott para 47 “The physical presence of people (i.e. people nearby the system) may be detected using sensors 132 like… microwave radar, microphones…”) identifying, using the at least one (Scott, Paragraph [0043], “The physical presence of people (i.e. people nearby the system) may be detected using sensors 132 like…microwave radar, microphones or cameras,… modalities like radar can provide more fine-grained information, that can include a positioning element (e.g. x/y/z position relative to the PC)”) processing, by the software code executed by the processing hardware, the radar data to obtain radar-based location data corresponding to a location of a user within a venue; (Scott para 50 “Position: As noted above, radar or camera-based sensors 132 may provide a position for one or multiple users. The position is then used to infer context, e.g. approaching the client device 102, moving away from the client device 102, presence in a different room than the client device 102, and so forth.”) processing, by the software code executed by the processing hardware, the audio data and (Scott para 21 “In another example, a microphone may be employed to measure loudness of the environment, and change the system behavior prior to receiving any voice input, such as by showing a prompt on the screen that changes when someone walks closer to a reference point.” The loudness of the environment is the portion of the audio collected from a source other than the user.) processing, by the software code executed by the processing hardware, the radar data to obtain radar-based venue data corresponding to the environment surrounding the user; (Scott para 48-49 “when motion information (e.g., angle of arrival information) is available (e.g., from radar information), a beamforming estimate can be used to enhance speech recognition…. [0049] Also, the system (e.g., the client device 102) can disambiguate between multiple sound sources, such as by filtering out the position of a known noise-producing device (e.g., a television) or background noise. …” Applicant’s spec. 55 says, “situationally aware social agent may be configured to distinguish between user relevant environmental cues and background noise, and to advantageously filter out the latter to enhance sensitivity to the former…” The definition of this claim element in spec. 55, is taught by Scott para 48-49.) correlating, by the software code executed by the processing hardware, the radar-based venue data and the audio-based venue data; (Scott para 48-49 “when motion information (e.g., angle of arrival information) is available (e.g., from radar information), a beamforming estimate can be used to enhance speech recognition…. [0049] Also, the system (e.g., the client device 102) can disambiguate between multiple sound sources, such as by filtering out the position of a known noise-producing device (e.g., a television) or background noise. …”) identifying, by the software code executed by the processing hardware based on the location of the user and the correlated radar-based venue data and audio-based venue data, an interactive expression for use by the social agent to interact with the user, wherein the interactive expression is identified using a portion of the audio data collected from the source of sound other than the user; and (Scott para 21 “In another example, a microphone may be employed to measure loudness of the environment, and change the system behavior prior to receiving any voice input, such as by showing a prompt on the screen that changes when someone walks closer to a reference point.” The interactive expression is the prompt. The loudness of the environment is the portion of the audio collected from a source other than the user. The filtering is based on an identification that a portion of the audio is from a source other than the user, Scott para 48-49 “when motion information (e.g., angle of arrival information) is available (e.g., from radar information), a beamforming estimate can be used to enhance speech recognition…. [0049] Also, the system (e.g., the client device 102) can disambiguate between multiple sound sources, such as by filtering out the position of a known noise-producing device (e.g., a television) or background noise. …”) executing, using an output unit of the social agent, the interactive expression used by the social agent, wherein the social agent is instantiated as a robot, a virtual character, or a tabletop or wall-mounted device, (Scott fig. 1 client device includes several tabletop device with screens and “first digital assistant” (Scott para 106) that are equivalent to Applicant’s social agent. The interactive expression is taught in the transition between experiences in Scott para 106, “the first digital assistant experience is adapted to generate a second digital assistant experience at the client device that is based on a difference between the first detected distance and the second detected distance (block 706).”) and wherein the interactive expression includes context from the portion of the audio data collected from the source of sound. (Scott para 21 “In another example, a microphone may be employed to measure loudness of the environment, and change the system behavior prior to receiving any voice input, such as by showing a prompt on the screen that changes when someone walks closer to a reference point.” The interactive expression is the prompt. The loudness of the environment is the portion of the audio collected from a source other than the user. The filtering is based on an identification that a portion of the audio is from a source other than the user, Scott para 48-49 “when motion information (e.g., angle of arrival information) is available (e.g., from radar information), a beamforming estimate can be used to enhance speech recognition…. [0049] Also, the system (e.g., the client device 102) can disambiguate between multiple sound sources, such as by filtering out the position of a known noise-producing device (e.g., a television) or background noise. …”) Scott doesn’t teach using the at least one directional microphone, an angle of arrival of the audio data; However, Peconsti teaches using the at least one directional microphone, an angle of arrival of the audio data. (Peconsti second paragraph “Thus, it is need to estimate the angle of arrival of a given sound. To achieve that, an electronic circuit to acquire signals from 2 microphones into the arduino was built.” The two microphones are one directional microphone.) Scott, the claims and peonsti are all directed to microphone applications. It would have been obvious to a person having ordinary skill in the art, at the time of filing, to receive angle of arrival data “in order to achieve proper monitoring, the monitoring system should be able to attribute sounds to the specific [person]…” Peconsti first paragraph. Claims 27 and 30 (Previously presented): The system of claim 1, wherein: the radar data and the audio data are timestamped, and the processing hardware is further configured to execute the software code to correlate the radar-based location data and the audio-based location data to determine the location of the user at a given point in time. (Scott para 18 “Aspects of digital assistant experience based on presence sensing include using presence sensing… and adapt the visual experience based on … context information such as the time of day…”) Scott teaches Claims 28 and 31 (Previously presented): The system of claim 1, wherein: the source of sound comprises an entertainment system or a person different from the user, and the interactive expression incorporates a subject matter of an output of the entertainment system or speech of the person. (Scott para 49 “the system (e.g., the client device 102) can disambiguate between multiple sound sources, such as by filtering out the position of a known noise-producing device (e.g., a television) or background noise.”) Scott teaches Claim 29 (Previously presented): The system of claim 1, wherein the output unit comprises at least one of a Text-To-Speech (TTS) module, a speaker, a display, a mechanical actuator, or a haptic actuator configured to effectuate the interactive expression. (Scott para 73 “if the system has access to multiple speakers, different speakers can be chosen for output to Bob and Alice, and respective volume levels at the different speakers can be optimized for Bob and Alice.”) Scott teaches Claim 20 (Previously Presented): The method of claim 14, further comprising: A method for use by a system including a processing hardware and a memory storing (Scott fig. 6) receiving, by the software code executed by the processing hardware, radar data and audio data collected by at least one (Scott, Paragraph [0043], “The physical presence of people (i.e. people nearby the system) may be detected using sensors 132 like…microwave radar, microphones or cameras,… modalities like radar can provide more fine-grained information, that can include a positioning element (e.g. x/y/z position relative to the PC)”) identifying, by the software code executed by the processing hardware, (Scott, Paragraph [0043], “The physical presence of people (i.e. people nearby the system) may be detected using sensors 132 like…microwave radar, microphones…” processing, by the software code executed by the processing hardware, the radar data, the audio data, (Scott para 50 “Position: As noted above, radar or camera-based sensors 132 may provide a position for one or multiple users. The position is then used to infer context, e.g. approaching the client device 102, moving away from the client device 102, presence in a different room than the client device 102, and so forth.” The venue is the place where the client device is.) recognizing, by the software code executed by the processing hardware, using the (Scott para 52 ” Identity recognition can employ camera-based face recognition or more coarse-grained recognition techniques that approximate the identity of a user.” Scott para 49 “When the identity of a user is known (such as discussed below), it is possible to apply a different speech recognition model that actually fits the user's accent, language, acoustic speech frequencies, and demographic.”) processing, by the software code executed by the processing hardware, the radar data, the audio data, (Scott para 21 “In another example, a microphone may be employed to measure loudness of the environment, and change the system behavior prior to receiving any voice input, such as by showing a prompt on the screen that changes when someone walks closer to a reference point.” The interactive expression is the prompt.) correlating, by the software code executed by the processing hardware, the radar-based venue data and the audio-based venue data (Scott para 47 “he physical presence of people (i.e. people nearby the system) may be detected using sensors 132 … microphones or cameras, and using techniques such as Doppler radar …”)_ identifying, by the software code executed by the processing hardware based on the location of the at least one user and the determined environment, an interactive expression for use by the social agent to interact with the at least one user; and (Scott para 21 “In another example, a microphone may be employed to measure loudness of the environment, and change the system behavior prior to receiving any voice input, such as by showing a prompt on the screen that changes when someone walks closer to a reference point.” The interactive expression is the prompt.) executing, by the software code executed by the processing hardware, the interactive expression used by the social agent. (Scott para 21 “In another example, a microphone may be employed to measure loudness of the environment, and change the system behavior prior to receiving any voice input, such as by showing a prompt on the screen that changes when someone walks closer to a reference point.” Showing the prompt on the screen is executing the identified expression.) Scott doesn’t teach directional microphone configured to collect audio data and identify an angle of arrival of the audio data. However, Peconsti teaches directional microphone configured to collect audio data and identify an angle of arrival of the audio data. (Peconsti second paragraph “Thus, it is need to estimate the angle of arrival of a given sound. To achieve that, an electronic circuit to acquire signals from 2 microphones into the arduino was built.” The two microphones are one directional microphone.) Scott, the claims and peonsti are all directed to microphone applications. It would have been obvious to a person having ordinary skill in the art, at the time of filing, to receive angle of arrival data “in order to achieve proper monitoring, the monitoring system should be able to attribute sounds to the specific [person]…” Peconsti first paragraph. Scott doesn’t teach the anonymous user history. However, Buhman teaches how to recognizing, by the software code executed by the processing hardware and using the (“virtual agent 150/350 is typically able to distinguish one anonymous human guest with whom a previous character interaction has occurred from another, as well as from anonymous human guests having no previous interaction experience with the character…” and “the presence of guest 126a/126b or guest object(s) 148 can be detected based on sensor data received from input module 130/230. That sensor data may also be used to reference interaction history database 108 to identify guest 126a/126b” where input module 130/230 contains a microphone to utilize audio data in (Buhmann, Drawings, Figure 2, 238).) The claims Buhmann as well as Scott are directed towards the field of interactive agents. It would have been obvious to a person having ordinary skill in the art, at the time of filing, to recognize anonymous users “in order to improve the performance of virtual agent…” Buhmann para 63. Buhmann and Scott don’t teach a trained NN. However, Guo teaches a trained NN. (Guo para 92 “hese augmented features are then used to assess the probability that a particular word, phoneme, or phone was heard. In more modern systems, this computation is performed by a specially trained deep neural network.”) The claims, Guo, Scott and Buhman are all directed to interactive agents. It would have been obvious to a person having ordinary skill in the art, at the time of filing, to use a trained NN to recognize the user because the “neural networks or GMMs may optionally be trained for a specific individual to give improved results.” Guo para 92. Scott doesn’t teach a map of the venue. However, King teaches a memory storing a software code, and a map of a venue; (King para 55 “The database may be used by algorithms to present a display of a seating map of a specific venue…” King para 15 “obtaining spatial reference data for a specific venue. The method also includes creating a digital model of the specific venue.”) correlating, by the software code executed by the processing hardware, the radar-based venue data and the audio-based venue data with the map of the venue to determine the environment surrounding the at least one user within the venue; (King para 82 “navigation matrix of panoramic video and audio viewports that in a particular geographic location or venue;…” correlating radar data, audio and map data is taught by the navigation matrix.) King, Scott and the claims deal with detecting audio in different environments. It would have been obvious to a person having ordinary skill in the art, at the time of filing, to include spatial reference data of a specific venue in Scott because “the representation of the specific venue may also include a representation of a specific stage or other performance venue may be superimposed with graphical depiction of historical data related to the venue. In some embodiments such a representation may aid in a process of designing audio capture locations for a future spectator event.” King para 65. Improving audio quality based on the venue will improve the robots accuracy in collecting audio data. Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Austin Hicks whose telephone number is (571)270-3377. The examiner can normally be reached Monday - Thursday 8-4 PST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Mariela Reyes can be reached at (571) 270-1006. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /AUSTIN HICKS/Primary Examiner, Art Unit 2142
Read full office action

Prosecution Timeline

Show 12 earlier events
Jan 26, 2026
Response after Non-Final Action
Feb 02, 2026
Non-Final Rejection mailed — §103, §112
Apr 30, 2026
Response Filed
May 21, 2026
Final Rejection mailed — §103, §112
Jun 22, 2026
Response after Non-Final Action
Jun 22, 2026
Notice of Allowance
Aug 04, 2026
Response after Non-Final Action
Aug 11, 2026
Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748999
EXPECTATION VALUE ESTIMATION METHOD AND APPARATUS IN QUANTUM SYSTEM, DEVICE, AND SYSTEM
3y 10m to grant Granted Sep 29, 2026
Patent 12718943
FACILITATING INTERPRETABILITY OF CLASSIFICATION MODEL
4y 6m to grant Granted Aug 25, 2026
Patent 12711096
OPTICAL CO-PROCESSOR ARCHITECTURE USING ARRAY OF WEAK OPTICAL PERCEPTRON
3y 10m to grant Granted Aug 18, 2026
Patent 12711355
SYSTEM AND METHOD FOR AN ADJUSTABLE NEURAL NETWORK
3y 4m to grant Granted Aug 18, 2026
Patent 12705474
REDUCED POWER CONSUMPTION ANALOG OR HYBRID MAC NEURAL NETWORK
4y 6m to grant Granted Aug 11, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

8-9
Expected OA Rounds
75%
Grant Probability
99%
With Interview (+25.8%)
3y 2m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 420 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month