Prosecution Insights
Last updated: August 17, 2026
Application No. 19/059,827

APPARATUS AND METHOD FOR RECOGNIZING CONVERSATIONAL CONTEXT IN ROBOT

Non-Final OA §101§102§103§112
Filed
Feb 21, 2025
Priority
Mar 29, 2024 — RE 10-2024-0043346
Examiner
LAM, PHILIP HUNG FAI
Art Unit
Tech Center
Assignee
Electronics and Telecommunications Research Institute
OA Round
1 (Non-Final)
85%
Grant Probability
Favorable
1-2
OA Rounds
1y 0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 85% — above average
85%
Career Allowance Rate
127 granted / 150 resolved
+24.7% vs TC avg
Strong +48% interview lift
Without
With
+48.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 6m
Avg Prosecution
29 currently pending
Career history
170
Total Applications
across all art units

Statute-Specific Performance

§101
24.0%
-16.0% vs TC avg
§103
55.7%
+15.7% vs TC avg
§102
11.1%
-28.9% vs TC avg
§112
4.1%
-35.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 150 resolved cases

Office Action

§101 §102 §103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION Introduction This office action is in response to Applicant’s response to submission filed on 2/21/2025. Claims 1-20 are pending of which claims 1, 11 and 20 are independent. As such, claims 1-20 have been examined. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1, 2, 6-10, 11, and 15-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Claim 1 recites an apparatus that, under the broadest reasonable interpretation, claims limitations that cover performance of the limitations in the human mind with the assistance of physical aids (e.g., pen and paper), but for the recitation of generic or well-known or conventional computer components. That is, other than reciting “a robot”, “a memory configured to store at least one program”, “and a processor configured to execute the program”, and “a robot's Point-of-View (POV) video”, nothing in these claim limitations precludes the steps from practically being performed in the mind and/or a part of human social interaction. As a whole, claim 1 pertains to interacting with another person, which is a mental process and/or human gathering activities that a human can do. Individually, each of the limitations also pertains to a mental process and/or insignificant extra solution activity, for example: recognizing a speaker and an intended recipient from a robot's Point-of-View (POV) video; (e.g., look who is talking and looking at who the recipient is.) and classifying a social interaction state of the robot as one of predefined social interaction states depending on the recognized speaker and the recognized intended recipient. (e.g., mentally or using pen to label or identify what is happening between parties, what is the social interaction about.) The judicial exception is not integrated into a practical application. In particular, the claims only recites generic computing components. Such generic computing components are recited at a high-level of generality (i.e., as a generic processor performing a generic computer function of receiving, determining, or outputting information) such that they amount to no more than mere instructions to apply the exception using generic computer components. Accordingly, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea. Claim 1 does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional limitations of using generic computer components amount to no more than mere instructions to apply the exception using generic computer components. Mere instructions to apply an exception using generic computer components cannot provide an inventive concept. Claim 1 is not patent eligible. The examiner further notes that the use of claimed generic computer components (“a robot”, “a memory configured to store at least one program”, “and a processor configured to execute the program”, and “a robot's Point-of-View (POV) video”) to obtain, extract, and/or generate data invokes such generic computer components “merely as a tool to perform an existing process”. MPEP 2106.05(f). MPEP 2106.05(f) further explains: Use of a computer or other machinery in its ordinary capacity for economic or other tasks (e.g., to receive, store, or transmit data) or simply adding a general purpose computer or computer components after the fact to an abstract idea (e.g., a fundamental economic practice or mathematical equation) does not integrate a judicial exception into a practical application or provide significantly more. See Affinity Labs v. DirecTV, 838 F.3d 1253, 1262, 120 USPQ2d 1201, 1207 (Fed. Cir. 2016) (cellular telephone); TLI Communications LLC v. AV Auto, LLC, 823 F.3d 607, 613, 118 USPQ2d 1744, 1748 (Fed. Cir. 2016) (computer server and telephone unit). Similarly, "claiming the improved speed or efficiency inherent with applying the abstract idea on a computer" does not integrate a judicial exception into a practical application or provide an inventive concept. Intellectual Ventures I LLC v. Capital One Bank (USA), 792 F.3d 1363, 1367, 115 USPQ2d 1636, 1639 (Fed. Cir. 2015). Claim 1 recites generic computer components (““a robot”, “a memory configured to store at least one program”, “and a processor configured to execute the program”, and “a robot's Point-of-View (POV) video”), with respect to performing tasks. MPEP 2106.05(d) and (f) further provides examples of court decisions where the courts found generic computing components to be mere instructions to apply a judicial exception, and further explains “increased speed” (e.g., using a computer to increase the speed of an otherwise mental process) does not provide an inventive concept. For example: A commonplace business method or mathematical algorithm being applied on a general purpose computer, Alice Corp. Pty. Ltd. V. CLS Bank Int’l, 573 U.S. 208, 223, 110 USPQ2d 1976, 1983 (2014); Gottschalk v. Benson, 409 U.S. 63, 64, 175 USPQ 673, 674 (1972); Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015). A process for monitoring audit log data that is executed on a general-purpose computer where the increased speed in the process comes solely from the capabilities of the general-purpose computer, FairWarning IP, LLC v. Iatric Sys., 839 F.3d 1089, 1095, 120 USPQ2d 1293, 1296 (Fed. Cir. 2016) (emphasis added). Performing repetitive calculations. Bancorp Services v. Sun Life, 687 F.3d 1266, 1278, 103 USPQ2d 1425, 1433 (Fed. Cir. 2012) ("The computer required by some of Bancorp’s claims is employed only for its most basic function, the performance of repetitive calculations, and as such does not impose meaningful limits on the scope of those claims.") Claim 11 recites a method claim that corresponds to the apparatus of claim 1 and is therefore rejected under the same grounds as claim 1 above. Claim 11 is not patent eligible. Claim 20 is a slightly lengthier version of claim 1, and the analysis is similarly provided below. Claim 20 recites: recognizing a speaker and an intended recipient based on multimodal recognition of audio and video from a robot's Point-of-View (POV) video; (e.g., listening and look at who is talking and look at who the recipient is.) classifying a social interaction state of the robot as one of predefined social interaction states depending on the recognized speaker and the recognized intended recipient; (e.g., mentally or using pen to label or identify what is happening between parties, what is the social interaction about.) and generating a response of the robot corresponding to a result classified as one of the predefined social interaction states, wherein, when the social interaction state is a first state in which in the speaker is a main user and the intended recipient is the robot, the response of the robot is generated as an interaction with the main user, wherein, when the social interaction state is a second state in which the speaker is the main user and the intended recipient is a surrounding user, the response of the robot is generated as waiting or intervention based on a result of determining context information, and wherein, when the social interaction state is a third state in which the speaker is the surrounding user and the intended recipient is unclear, the response of the robot is generated as waiting. (e.g., observe what is going on, is the speaker the primary user interacting with the robot, or another person, and providing an response or waiting for moment to intervene as necessary, or if someone not the primary user is speaking, but its not clear who they are talking to, then observe and wait.) Claim 20 is not patent eligible. Claims 2, 6-10, and 15-19 depend from independent claims 1, and 11 respectively, do not remedy any of the deficiencies of claims 1 and 11, and therefore are rejected on the same grounds as claim 1, and 11 from above. Claim 2 further comprising: wherein the program is configured to, in the recognizing, recognize the speaker and the intended recipient based on multimodal recognition of audio and video from the robot's POV video. (e.g., figure out who is talking based on senses, hearing and eyesight.) Claim 6 further recite: wherein the predefined social interaction states include at least one of a first state in which the speaker is a main user and the intended recipient is the robot, a second state in which the speaker is the main user and the intended recipient is a surrounding user, or a third state in which the speaker is the surrounding user and the intended recipient is unclear, or a combination thereof. (e.g., observe what is going on, is the speaker the primary user interacting with the robot, or another person, and providing a response or waiting for moment to intervene as necessary, or if someone not the primary user is speaking, but it’s not clear who they are talking to, then observe and wait.) Claim 7 further comprising: wherein the program is configured to further perform: generating a response of the robot corresponding to a result classified as one of the predefined social interaction states. (e.g., figure out what the interaction is about, then providing a response.) Claim 8 further recites: wherein the program is configured to, when the social interaction state is the first state, generate the response of the robot as an interaction with the main user. (e.g., providing a response if it has been determined the speaker is talking directly to the operator or assistant.) Claim 9 further recites: wherein the program is configured to, when the social interaction state is the second state, generate the response of the robot as waiting or intervention based on a result of determining context information. (e.g., when determine that the speaker is talking to someone else nearby, wait or intervene when the situation calls for it.) Claim 10 further recites: wherein the program is configured to, when the social interaction state is the third state, generate the response of the robot as waiting. (e.g., when it is determined that a nearby person is talking, but not sure who that person is addressing, the operator or assistant just wait and standsby.) The analysis of Claims 15-19 corresponds to claims 6-10, and therefore similar rationale of rejection is applied to these claims respectively. In sum, claims 2, 6-10, and 15-19 depend from claims 1, and 11 respectively, and further recite mental processes as explained above. None of the additional limitations recited in claims 2, 6-10, and 15-19 amount to anything more than the same or a similar abstract idea as recited in claims 1. Nor do any limitations in claims 2, 6-10, and 15-19: (a) integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea or (b) amount to significantly more than the judicial exception because the additional limitations of using generic computer components amounts to no more than mere instructions to apply the exception using generic computer components. Claims 2, 6-10, and 15-19 are not patent eligible. Dependent claims 3-5 and 12-14 recites a specific implementation using a robot to recognize the context of interaction between a user and their surrounding environment, therefore are deemed patent eligible. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim 13 and 14 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 13 recites the limitation “the combined feature vector” in lines 3-4. There is insufficient antecedent basis for this limitation in the claim. For the sake of compact prosecution, Examiner will interpret claim 13 as depended from claim 12 instead of claim 11, because claim 12 discloses combined feature vector. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 1-2, 6-8, 11, and 15-17 are rejected under 35 U.S.C. 102 (a)(1) as being anticipated by Lang, S., Kleinehagenbrock, M., Hohenner, S., Fritsch, J., Fink, G. A., & Sagerer, G. (2003, November). Providing the basis for human-robot-interaction: A multi-modal attention system for a mobile robot. In Proceedings of the 5th international conference on Multimodal interfaces (pp. 28-35). Regarding Claim 1, Lang discloses: 1. An apparatus for recognizing a conversational context in a robot, ([sect 1.] On a mobile robot that operates in an environment where several people are moving around, it is not always obvious for the robot which of the surrounding persons wants to interact with it. Therefore, it is necessary to develop techniques that allow a mobile robot to automatically recognize when and how long a user’s attention is directed towards it for communication.); comprising: a memory configured to store at least one program; ([categories and subject descriptors] image processing and computer vision, it would be implied the use of memory) and a processor configured to execute the program, wherein the program is configured to perform: ([categories and subject descriptors] image processing and computer vision, it would be implied the use of a processor) recognizing a speaker and an intended recipient from a robot's Point-of-View (POV) video; ([sect 4, multi-modal person tracking] discuss using audio and vision tracking from Robot’s sensors to see who is talking to the Robot.) Also see figs.2 and 3. and classifying a social interaction state of the robot as one of predefined social interaction states depending on the recognized speaker and the recognized intended recipient. ([Sect 6 Focus the attention] discuss how the robot decides who to look at during a conversation, especially focus on speaker turns, communication partners, and shifting of attention between people, waiting for the speaker speak again, or if the speaker is talking to the robot) Also see fig. 8 reproduced below for viewing convenience, which describes the robot categorizing interaction states based on the speaker and the recipients. PNG media_image1.png 271 452 media_image1.png Greyscale Regarding Claim 2, Lang discloses all the elements of claim 1, Lang further discloses: wherein the program is configured to, in the recognizing, recognize the speaker and the intended recipient based on multimodal recognition of audio and video from the robot's POV video. ([sect 8 Summary] In this paper we presented a multi-modal attention system for a mobile robot. The system is able to observe several persons in the vicinity of the robot and to decide based on a combination of acoustic and visual cues whether one of these is willing to engage in a communication with the robot. This attentional behavior is realized by combining an approach for multi-modal person tracking with the localization of sound sources and the detection of head orientation derived from a face recognition system. Note that due to the integration of cues from multiple modalities it is possible to verify the position of a speech source in 3D space using a singlepair of microphones only. Persons that are observed by the robot and are also talking are considered persons of interest. If a person of interest is also facing the robot it will become the current communication partner. Otherwise the robot assumes that the speech was addressed to another person present.) Also see fig. 3 (reproduced below for viewing convenience), and section 4. PNG media_image2.png 368 437 media_image2.png Greyscale wherein the program is configured to perform, in the recognizing, encoding an audio feature extracted from the robot's POV video; encoding a face region in a frame extracted from the robot's POV video; encoding the frame extracted from the robot's POV video; generating a combined feature vector from information obtained by encoding at least one of the audio feature, the face region or the frame, or a combination thereof; and recognizing the speaker and the intended recipient from the combined feature vector. Regarding Claim 6, Lang discloses all the elements of claim 1, Lang further discloses: wherein the predefined social interaction states include at least one of a first state in which the speaker is a main user and the intended recipient is the robot, a second state in which the speaker is the main user and the intended recipient is a surrounding user, ([Sect 6 Focus the attention] discuss how the robot decides who to look at during a conversation, especially focus on speaker turns, communication partners, and shifting of attention between people, waiting for the speaker speak again, or if the speaker is talking to the robot) In addition to the attention system described so far, which enables the robot to detect the person of interest and to maintain its attention during interaction, the robot decides whether the person of interest is addressing the robot and, therefore, is considered as communication partner. This decision is based on the orientation of the person’s head, as it is assumed that humans face their addressees for most of the time while they are talking to them. Whether a tracked person faces the robot or not is derived from the face recognition system. If the face of the person of interest is detected for more than 20 % of the time the person is speaking, this person is considered to be the communication partner. Also see fig. 8 reproduced earlier in the mapping for claim 1 for viewing convenience, which describes the robot categorizing interaction states based on the speaker and the recipients.) Regarding Claim 7, Lang discloses all the elements of claim 6, Lang further discloses: wherein the program is configured to further perform: generating a response of the robot corresponding to a result classified as one of the predefined social interaction states. ([Sect 6 Focus the attention] discuss how the robot decides who to look at during a conversation, especially focus on speaker turns, communication partners, and shifting of attention between people, waiting for the speaker speak again, or if the speaker is talking to the robot) Also see fig. 8 reproduced earlier in the mapping for claim 1 for viewing convenience, which describes the robot categorizing interaction states based on the speaker and the recipients. Regarding Claim 8, Lang discloses all the elements of claim 7, Lang further discloses: wherein the program is configured to, when the social interaction state is the first state, generate the response of the robot as an interaction with the main user. ([Sect 6 Focus the attention] discuss how the robot decides who to look at during a conversation, especially focus on speaker turns, communication partners, and shifting of attention between people, waiting for the speaker speak again, or if the speaker is talking to the robot) Also see fig. 8 reproduced earlier in the mapping for claim 1 for viewing convenience, which describes the robot categorizing interaction states based on the speaker and the recipients.) Claim 11 recites a method claim that corresponds to the apparatus of claim 1 and is therefore rejected under the same grounds as claim 1 above. Claim 15-17 are method claims with limitations corresponding to the limitations of Claims 6-8 respectively and are rejected under similar rationale. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 3 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Lang, in view of Chang (US 20240265917). Regarding Claim 3, Lang discloses all of claim 2, Lang further discloses: encoding a face region in a frame extracted from the robot's POV video; ([sect 5.1 Face Recognition] and recognizing the speaker and the intended recipient from the combined feature (section 4 and summary and other sections throughout discuss using multimodal feature to detect speaker and intended recipient) Lang does not disclose the following feature, Chang from the related art discloses: wherein the program is configured to perform, in the recognizing, encoding an audio feature extracted from the robot's POV video; ([0034] an audio encoder 242.) encoding the frame extracted from the robot's POV video; ([0034] visual higher-order feature representations generated from the video frames 222) generating a combined feature vector from information obtained by encoding at least one of the audio feature, the face region or the frame, or a combination thereof; ([0034] fusing acoustic higher-order feature representations 246, 246a-n generated by the audio encoder 242 with visual higher-order feature representations generated from the video frames 222 to generate audiovisual higher-order feature representations 248, 248a-n.) Lang and Chang are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Lang to combine the teaching of Chang, because in noisy environment, video of speaker can help the system understand what the speaker is trying to say (Chang, [Background]). Claim 12 a method claim with limitations similar to the limitations of Claim 3 and is rejected under similar rationale. Claims 4-5 and 13-14 are rejected under 35 U.S.C. 103 as being unpatentable over Lang, in view of Chang (US 20240265917), and further in view of Zhang (US 11681364). Regarding Claim 4, Lang and Chang disclose all of claim 3, Lang and Chang do not appear to disclose the feature recited below. wherein the program is configured to further perform, in the recognizing, recognizing and encoding a gaze point from the frame extracted from the robot's POV video; ([col. 8, lines 41-59] the encoder-decoder network 250 may be a CNN-based encoder-decoder, where the encoder may consist of two layers of 3×3 convolution layer with a stride 2, followed by a decoder that may consist of two 3×3 transposed-convolution layers also with a stride 2. A final layer of the decoder may have a channel number of 1, and all other layers may have a channel layer of the image feature map 133 (e.g., 256). In some implementations, the encoder-decoder network 250 may have more or fewer convolution layers and/or transposed-convolution layers.) generating a gaze point feature vector based on encoded information and the combined feature vector; ([col. 8, lines 41-59] the encoder-decoder network 250 may be a CNN-based encoder-decoder, where the encoder may consist of two layers of 3×3 convolution layer with a stride 2, followed by a decoder that may consist of two 3×3 transposed-convolution layers also with a stride 2. A final layer of the decoder may have a channel number of 1, and all other layers may have a channel layer of the image feature map 133 (e.g., 256). In some implementations, the encoder-decoder network 250 may have more or fewer convolution layers and/or transposed-convolution layers. The final output of the decoder may be the gaze direction probability map 134, which may have the same w′×h′ of the combined feature map 245; for example, 7×7, where the value at each element represents a probability that the POI has a gaze fixed in the direction of the person/object at that position, with zero values at elements not corresponding to a detected person or object. The gaze direction probability map 134 can be used in the gaze prediction 140 to predict gaze prediction data 141.) decoding the gaze point feature vector; ([col. 8, lines 41-59] …followed by a decoder that may consist of two 3×3 transposed-convolution layers also with a stride 2.) and recognizing the gaze point based on decoded information. ([col. 8, lines 41-59] The final output of the decoder may be the gaze direction probability map 134, which may have the same w′×h′ of the combined feature map 245; for example, 7×7, where the value at each element represents a probability that the POI has a gaze fixed in the direction of the person/object at that position, with zero values at elements not corresponding to a detected person or object. The gaze direction probability map 134 can be used in the gaze prediction 140 to predict gaze prediction data 141.) Lang/Chang/Zhang are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Lang and Chang to combine the teaching of Zhang, because the system can leverage the gaze pattern information to improve processes such as speech processing in a multi-user environment (Zhang, [Col. 3, lines 51-53]). Regarding Claim 5, Lang and Chang disclose all of claim 4, Zhang further discloses: wherein the program is configured to recognize the speaker and the intended recipient based on the combined feature vector and decoded information of the gaze point feature vector in the recognizing. ([col. 6, lines 44-53] The gaze prediction 140 may combine these feature vectors—for example, by concatenating them—to determine a combined feature vector. The gaze prediction 140 may include performing additional processing on the combined feature vector, such as processing the combined feature vector using an LSTM, to determine the gaze prediction data 141. The system 115 may send the gaze prediction data 141 to downstream processes 150, such as NLU, entity resolution, reranking, or for use by a skill 1090.) Where the rationale for the combination would be similar to the one provided earlier. Claim 13-14 are method claims with limitations corresponding to the limitations of Claims 4 and 5 respectively and are rejected under similar rationale. Claims 9 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Lang, in view of Nguyen (US 20220310094). Regarding Claim 9, Lang discloses all of claim 7, Although Lang discloses determining the second state (where main user is interacting with other nearby user or person), Lang do not appear to disclose responding by waiting or intervention based on the result of determining context information. Nguyen in the related arti discloses: wherein the program is configured to, when the social interaction state is the second state, generate the response of the robot as waiting or intervention based on a result of determining context information. ([0068] whether and/or which automated assistant functions are performed may depend on both sets of interaction confidence levels, in the case where a given user is assigned both. Thus, automated assistant 120 may only perform the one or more automated assistant functions or may only perform certain automated assistant functions based on the two types of interaction confidence levels and/or the relationship between the two types of interaction confidence levels. For example, automated assistant 120 may only perform “background operations” (e.g., automated assistant functions that are determined to be unlikely to disturb the interaction between the users) when, for instance, a given user has an interaction confidence level with another user that satisfies one threshold but an interaction confidence level with automated assistant 120 that fails to satisfy another threshold (or vice versa), when the difference between the interaction confidence levels fails to satisfy yet another threshold, and/or when the interaction confidence level with another user is higher than the interaction confidence level with automated assistant 120. As another example, automated assistant 120 may not perform any responsive automated assistant functions based on detecting a given user's speech and/or gestures when their interaction confidence level with another user is higher than their interaction confidence level with automated assistant 120.) Lang and Nguyen are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Lang to combine the teaching of Nguyen, because the system techniques described herein to determine whether they should trigger responsive action by the automated assistant, or should be ignored or discarded which leads to better user experience (Nguyen, [0006]). Claim 18 a method claim with limitations similar to the limitations of Claim 9 and is rejected under similar rationale. Claims 10 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Lang, in view of Miyazaki (US 20210210096). Regarding Claim 10, Lang discloses all of claim 7, Lang does not appear to disclose the following feature. Miyazaki in the related art discloses: wherein the program is configured to, when the social interaction state is the third state, generate the response of the robot as waiting. ([0113] FIG. 10(a) illustrates the same state as that of the above FIG. 8(a), and the detailed description thereof is omitted here. FIG. 10(b) illustrates a case where, for a reason that the person E is an angry person, it has been determined that the integration of the person E into the conversation group G2 is inappropriate, and no speech output (screen display) for promoting the integration of the person E into the conversation group G2 is performed from the agent system 100.) [the system determines that the angry person (person E) should not be responded to, although the system could determine that the person is upset their line of sight is blocked, but its not clear if person E is talking to themselves or to person C or D or to both] fig. 10B is reproduced below for viewing convenience. PNG media_image3.png 512 408 media_image3.png Greyscale Lang and Miyazaki are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Lang to combine the teaching of Miyazaki, because the configuration makes it possible to determine that integration of conversation groups that is not unintended by users is inappropriate (Miyazaki, [0011]). Claim 19 a method claim with limitations similar to the limitations of Claim 10 and is rejected under similar rationale. Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Lang, in view of Nguyen (US 20220310094), and further in view of Miyazaki (US 20210210096). 20. A method for recognizing a conversational context in a robot, comprising: recognizing a speaker and an intended recipient based on multimodal recognition of audio and video from a robot's Point-of-View (POV) video; ([sect 4, multi-modal person tracking] discuss using audio and vision tracking from Robot’s sensors to see who is talking to the Robot.) Also see figs.2 and 3. classifying a social interaction state of the robot as one of predefined social interaction states depending on the recognized speaker and the recognized intended recipient; ([Sect 6 Focus the attention] discuss how the robot decides who to look at during a conversation, especially focus on speaker turns, communication partners, and shifting of attention between people, waiting for the speaker speak again, or if the speaker is talking to the robot) Also see fig. 8 reproduced early (see above) for viewing convenience, which describes the robot categorizing interaction states based on the speaker and the recipients. and generating a response of the robot corresponding to a result classified as one of the predefined social interaction states, ([Sect 6 Focus the attention] discuss how the robot decides who to look at during a conversation, especially focus on speaker turns, communication partners, and shifting of attention between people, waiting for the speaker speak again, or if the speaker is talking to the robot) Also see fig. 8 reproduced earlier in the mapping for claim 1 for viewing convenience, which describes the robot categorizing interaction states based on the speaker and the recipients.) wherein, when the social interaction state is a first state in which in the speaker is a main user and the intended recipient is the robot, the response of the robot is generated as an interaction with the main user, ([Sect 6 Focus the attention] discuss how the robot decides who to look at during a conversation, especially focus on speaker turns, communication partners, and shifting of attention between people, waiting for the speaker speak again, or if the speaker is talking to the robot) Also see fig. 8 reproduced earlier in the mapping for claim 1 for viewing convenience, which describes the robot categorizing interaction states based on the speaker and the recipients.) Although Lang discloses determining the second state (where main user is interacting with other nearby user or person), Lang do not appear to disclose responding by waiting or intervention based on the result of determining context information. Nguyen in the related art discloses: wherein, when the social interaction state is a second state in which the speaker is the main user and the intended recipient is a surrounding user, the response of the robot is generated as waiting or intervention based on a result of determining context information, ([0068] whether and/or which automated assistant functions are performed may depend on both sets of interaction confidence levels, in the case where a given user is assigned both. Thus, automated assistant 120 may only perform the one or more automated assistant functions or may only perform certain automated assistant functions based on the two types of interaction confidence levels and/or the relationship between the two types of interaction confidence levels. For example, automated assistant 120 may only perform “background operations” (e.g., automated assistant functions that are determined to be unlikely to disturb the interaction between the users) when, for instance, a given user has an interaction confidence level with another user that satisfies one threshold but an interaction confidence level with automated assistant 120 that fails to satisfy another threshold (or vice versa), when the difference between the interaction confidence levels fails to satisfy yet another threshold, and/or when the interaction confidence level with another user is higher than the interaction confidence level with automated assistant 120. As another example, automated assistant 120 may not perform any responsive automated assistant functions based on detecting a given user's speech and/or gestures when their interaction confidence level with another user is higher than their interaction confidence level with automated assistant 120.) Lang and Nguyen are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Lang to combine the teaching of Nguyen, because the system techniques described herein to determine whether they should trigger responsive action by the automated assistant, or should be ignored or discarded which leads to better user experience (Nguyen, [0006]). Lang and Nguyen do not appear to disclose the following feature. Miyazaki in the related art discloses: and wherein, when the social interaction state is a third state in which the speaker is the surrounding user and the intended recipient is unclear, the response of the robot is generated as waiting. ([0113] FIG. 10(a) illustrates the same state as that of the above FIG. 8(a), and the detailed description thereof is omitted here. FIG. 10(b) illustrates a case where, for a reason that the person E is an angry person, it has been determined that the integration of the person E into the conversation group G2 is inappropriate, and no speech output (screen display) for promoting the integration of the person E into the conversation group G2 is performed from the agent system 100.) [the system determines that the angry person (person E) should not be responded to, although the system could determine that the person is upset their line of sight is blocked, but its not clear if person E is talking to themselves or to person C or D or to both] Lang, Nguyen and Miyazaki are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Lang and Nguyen to combine the teaching of Miyazaki, because the configuration makes it possible to determine that integration of conversation groups that is not unintended by users is inappropriate (Miyazaki, [0011]). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Liu US 20230008363 – discloses method/device in which Robot interacts with multiparty using audio and video data. See Abstract, para 0214 and figs. 1-6, 8A, 8D, 9D, and 9E for additional details. Foster, M. E., Gaschler, A., & Giuliani, M. (2017). Automatically classifying user engagement for dynamic multi-party human–robot interaction. International Journal of Social Robotics, 9(5), 659-674. – discloses Robot interaction with multi-person in a bar setting, and disclose robot classifying each interaction as engagement or nonengagement according to context using multimodal approach. See Abstract, sections 3-9, and fig. 2 for additional details. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Philip H Lam whose telephone number is (571)272-1721. The examiner can normally be reached 9 AM-3 PM Pacific time. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bhavesh Mehta can be reached on 571-272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /PHILIP H LAM/ Examiner, Art Unit 2656
Read full office action

Prosecution Timeline

Feb 21, 2025
Application Filed
Jul 30, 2026
Non-Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12688847
ERROR-CORRECTION AND EXTRACTION IN REQUEST DIALOGS
4y 1m to grant Granted Jul 21, 2026
Patent 12682164
CHAT SUPPORT PLATFORM HAVING AUTOMATIC KEYWORD CORRECTION
3y 3m to grant Granted Jul 14, 2026
Patent 12670519
CONTENT RECOMMENDATION USING RETRIEVAL AUGMENTED ARTIFICIAL INTELLIGENCE
3y 2m to grant Granted Jun 30, 2026
Patent 12657395
METHODS AND SYSTEMS FOR AVOIDING OFFENSIVE LANGUAGE BASED ON PERSONAS
2y 9m to grant Granted Jun 16, 2026
Patent 12639529
ENHANCING LARGE LANGUAGE MODELS USING IN-CONTEXT LEARNING AND ONLINE KNOWLEDGE
2y 6m to grant Granted May 26, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
85%
Grant Probability
99%
With Interview (+48.0%)
2y 6m (~1y 0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 150 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month