Prosecution Insights
Last updated: October 01, 2026
Application No. 19/203,991

DEVICES AND METHODS FOR INVOKING DIGITAL ASSISTANTS

Non-Final OA §103§DOUBLEPATENT
Filed
May 09, 2025
Priority
Mar 12, 2021 — provisional 63/160,404 +1 more
Examiner
KIM, JONATHAN C
Art Unit
Tech Center
Assignee
Apple Inc.
OA Round
1 (Non-Final)
74%
Grant Probability
Favorable
1-2
OA Rounds
1y 0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 74% — above average
74%
Career Allowance Rate
271 granted / 368 resolved
+13.6% vs TC avg
Strong +39% interview lift
Without
With
+38.7%
Interview Lift
resolved cases with interview
Typical timeline
2y 5m
Avg Prosecution
17 currently pending
Career history
392
Total Applications
across all art units

Statute-Specific Performance

§101
19.9%
-20.1% vs TC avg
§103
50.8%
+10.8% vs TC avg
§102
11.5%
-28.5% vs TC avg
§112
10.4%
-29.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 368 resolved cases

Office Action

§103 §DOUBLEPATENT
DETAILED ACTION This Office Action is in response to the correspondence filed by the applicant on 5/9/2025. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the claims at issue are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); and In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on a nonstatutory double patenting ground provided the reference application or patent either is shown to be commonly owned with this application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The USPTO internet Web site contains terminal disclaimer forms which may be used. Please visit http://www.uspto.gov/forms/. The filing date of the application will determine what form should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to http://www.uspto.gov/patents/process/file/efs/guidance/eTD-info-I.jsp. Claims 1-19 are rejected on the ground of nonstatutory double patenting as being unpatentable over Claims 1-42 of US PAT 12,327,554. Although the claims, at issue are not identical, they are not patentably distinct from each other because the claims of the instant application are rejected as being unpatentable over the claims of the US PAT. Please see below for the mapping in the table, where the bolded limitations indicate the corresponding limitations between the US PAT and instant application. Instant application: 19/203,991 US PAT 12,327,554 1. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device having a display, cause the electronic device to: display, on the display, a virtual reality (“VR”) environment, wherein the VR environment includes a first avatar representing a first domain-specific digital assistant, and wherein the first domain-specific digital assistant corresponds to a first domain; and while displaying the first avatar within the VR environment: receive a user voice input; identify one or more domains corresponding to the user voice input; determine, based on the identified one or more domains, whether the user voice input represents an intent to invoke the first domain-specific digital assistant or an intent to invoke a second domain-specific digital assistant; and in accordance with a determination that the user voice input represents an intent to invoke the second domain-specific digital assistant: invoke the second domain-specific digital assistant; determine, based on the user voice input, a first digital assistant response using the second domain-specific digital assistant; and output, via a speaker, the first digital assistant response. 1. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device having a display, cause the electronic device to: display, on the display, a virtual reality (“VR”) environment, wherein the VR environment includes a first avatar representing a domain-specific digital assistant, wherein the domain-specific digital assistant corresponds to a first domain, and wherein the domain-specific digital assistant corresponds to a single software application stored on the electronic device; and while displaying the first avatar within the VR environment: receive a user voice input; determine whether the user voice input represents an intent to invoke the domain-specific digital assistant or an intent to invoke a second digital assistant; and in accordance with a determination that the user voice input represents an intent to invoke the second digital assistant: invoke the second digital assistant, wherein virtual objects within a threshold distance of a user are suspended during a first period of time while the second digital assistant is invoked; determine, based on the user voice input, a first digital assistant response using the second digital assistant; and output, via a speaker, the first digital assistant response. Other independent claims 18 and 19 are also similar to the independent claims 17 and 30 of the US PAT. With respect to the dependent claims, each of the claims maps to a corresponding dependent claim of the US PAT or are found within the scope of the independent claim. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-3, 6-9, 11-12, 14-19 are rejected under 35 U.S.C. 103 as being unpatentable over SUGIHARA (US 2022/0005470 A1), and in further view of KEPHART (US 2022/0028366 A1). REGARDING CLAIM 1, SUGIHARA discloses a non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device (Fig.1 – “Agent Device”) having a display (Fig. 1 – “Display”), cause the electronic device to: display, on the display, a [virtual reality ("VR")] interaction environment (Fig. 9; Par 98 – “The master agent function element 152 causes the agent image EIA having a dialogue among the agent images EIA to EIC corresponding to the sub-agent function elements 154A to 154C generated by the display controller 156 to be displayed in front of the other agents EIB and EIC when viewed from the occupant. Also, the master agent function element 152 may adjust the size of each agent image in accordance with a positional relationship between the agent images EIA to EIC on the image space.”), wherein the [VR] environment includes a first avatar representing a first domain-specific digital assistant (Fig. 9 EIA; Par 99 – “As shown in FIG. 9, the occupant can be allowed to easily ascertain that there are a plurality of agents by displaying the agent images EIA to EIC corresponding to the agents capable of having a dialogue with the occupant on the display 132.”), and wherein the first domain-specific digital assistant corresponds to a first domain (Fig. 4 Function Information; Agent Information; Par 67 – “For example, when the recognized command is “activation of the air conditioner,” the master agent function element 152 determines an agent capable of executing in-vehicle equipment control, which is control corresponding to the command, with reference to the function information table 172. In the example of FIG. 4, the master agent function element 152 acquires that the agent capable of starting the air conditioner is agent A and determines agent A that has a dialogue with the occupant.”); and while displaying the first avatar within the VR environment: receive a user voice input (Par 64 – “For example, when the meaning such as “Turn on the air conditioner” or “Please turn on the air conditioner” is recognized as the recognition result, the master agent function element 152 generates a command obtained through the replacement with the standard text information “activation of the air conditioner.””; Par 67 – “For example, when the recognized command is “activation of the air conditioner,” the master agent function element 152 determines an agent capable of executing in-vehicle equipment control, ….”); identify one or more domains corresponding to the user voice input (Fig. 4 – “Function Information; Agent A; Agent B; Agent C”; Par 109 – “In the third scene, the master agent function element 152 determines agent (agent B in the example of FIG. 11) capable of implementing a function corresponding to the command as an agent that has a dialogue with the occupant with reference to command information of the function information table 172 on the basis of a command corresponding to a request recognized from the speech sound of the occupant.”); determine, based on the identified one or more domains (Fig. 4 – “Function Information; Agent A; Agent B; Agent C”; Par 109 – “In the third scene, the master agent function element 152 determines agent (agent B in the example of FIG. 11) capable of implementing a function corresponding to the command as an agent that has a dialogue with the occupant with reference to command information of the function information table 172 on the basis of a command corresponding to a request recognized from the speech sound of the occupant.”), whether the user voice input represents an intent to invoke the first domain-specific digital assistant or an intent to invoke a second domain-specific digital assistant (Par 67 – “In the example of FIG. 4, the master agent function element 152 acquires that the agent capable of starting the air conditioner is agent A and determines agent A that has a dialogue with the occupant. Also, in the case of a function that can be executed by a plurality of agents, such as the shop/store search, the master agent function element 152 may determine an agent on the basis of a priority predetermined for each function.”); and in accordance with a determination that the user voice input represents an intent to invoke the second domain-specific digital assistant (Par 70 – “Thereby, for example, the master agent function element 152 can ascertain content of the dialogue between the sub-agent function element 154A and the occupant, select another sub-agent function element 154 from which a more appropriate answer is likely to be obtained, and perform control such as switching to the other sub-agent function element 154 that has been selected.”): invoke the second domain-specific digital assistant (Fig. 11; Par 110 – “For example, as shown in FIG. 11, when the main agent having the dialogue with the occupant is switched from agent A to agent B, the master agent function element 152 causes an agent speech sound such as “The request is answered by agent B.” to be output from agent A and causes an agent speech sound such as “I will respond.” to be output from agent B.”); determine, based on the user voice input, a first digital assistant response using the second domain-specific digital assistant (Par 110 – “For example, as shown in FIG. 11, when the main agent having the dialogue with the occupant is switched from agent A to agent B, the master agent function element 152 causes an agent speech sound such as “The request is answered by agent B.” to be output from agent A and causes an agent speech sound such as “I will respond.” to be output from agent B.”; Par 111 – “Also, the master agent function element 152 switches the speech sound input collected by the microphone 124B from the speech sound input interface 154Aa of the sub-agent function element 154A to the speech sound input interface 154Ba of the sub-agent function element 154B. Thereby, a dialogue or the like can be implemented between agent B and the occupant.”); and output, via a speaker, the first digital assistant response (Par 19 – “and causing, by the computer, a speaker configured to output a speech sound inside the vehicle cabin to output the generated speech sound of the agent”; Fig. 11 – “I will respond”; Pars 110 and 111 as above; Par 93 – “Also, when the command is “What is the distance to station A?,” the dialogue generator 230 acquires control content of speech sound control for outputting a speech sound “*** [km] from here.” and control content of display control for displaying a route image to station A.”). SUGIHARA does not explicitly teach the [square-bracketed] limitations and teaches the underlined features instead. In other words, SUGIHARA teaches a voice/speech interaction between a human user and digital assistants in a vehicle environment, but does not explicitly teach the environment is a [virtual reality] environment. SUGIHARA teaches multiple sub-agent function element for each domain (e.g., Agent A for in-vehicle equipment control, Agent B for Radio control, etc.). The domain-specific agent (e.g., Agent A) correspond to sub-agent function element (e.g., application for controlling a navigator, radio, etc.,) stored in the agent device (100), but SUGIHARA does not explicitly teach the domain-specific digital assistant corresponds to [a single software] application. KEPHART discloses the [square-bracketed] limitations. KEPHART discloses a method/system utilizing speech signals for virtual assistants comprising: display, on the display, a [virtual reality ("VR")] environment (KEPHART Par 63 – “Reference should now be had to FIG. 3. Environment 301 is an exemplary embodiment that includes a physical environment in which one or more human negotiators Hi, H2 (numbered 303-1 and 303-2) are situated. The physical environment can range from a laptop computer to a conference room with one or two flat screen displays to a fully immersive environment (e.g. virtual reality (VR) or augmented reality (AR)).”), wherein the [VR] environment includes a first avatar representing a first domain-specific digital assistant (KEPHART Par 66 – “Thus, in the physical environment, natural input is collected from humans, and the response of one or more agents is also rendered to humans in a relatively natural manner (for example, via cartoon avatars 309-1, 309-2, 309-3 that have a reasonable resemblance to an actual human).”), and wherein the first domain-specific digital assistant corresponds to a first domain (KEPHART Par 195 – “The prior art includes an educational scenario in which students practice the Mandarin Chinese language and culture through spoken role-play with embodied AI agents in an immersive environment. Initial studies of AI-assisted language education had shown that immersion has a beneficial impact. In the Mandarin education scenario, the agents play various roles, including shopkeepers who compete with each other for the student's business.”; Par 198 – “The main character also needs a metal helmet, metal breastplates, shoulder guards, sword belts, leather boots, axes, bows, arrows and swords. There are several villages in which he can buy food, clothes, armor sets, and the like, but he needs a specialized currency for that.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method/system of SUGIHARA to include an agent interacting with a human user in a virtual reality environment, as taught by KEPHART. One of ordinary skill would have been motivated to include an agent interacting with a human user in a virtual reality environment, in order to provide a more natural interaction so that a user can obtain language skills (Par 195). REGARDING CLAIM 2, SUGIHARA in view of KEPHART discloses the non-transitory computer-readable storage medium of claim 1, wherein determining whether the user voice input represents an intent to invoke the first domain-specific digital assistant or an intent to invoke a second domain-specific digital assistant (SUGIHARA Par 67 – “In the example of FIG. 4, the master agent function element 152 acquires that the agent capable of starting the air conditioner is agent A and determines agent A that has a dialogue with the occupant. Also, in the case of a function that can be executed by a plurality of agents, such as the shop/store search, the master agent function element 152 may determine an agent on the basis of a priority predetermined for each function.”) includes: determining whether the user voice input includes a first domain-specific digital assistant trigger corresponding to the domain-specific digital assistant or a second domain-specific digital assistant trigger corresponding to the second digital assistant (SUGIHARA Par 63 – “For example, the master agent function element 152 recognizes a word for calling any agent (interactive agent) such as an input speech sound “Hey!” or “Hello!” or recognizes a word (for example, a wake word) for designating and calling an agent implemented by each of the sub-agent function elements 154A to 154C.”; Par 68 – “Also, when a wake word for calling a specific agent has been recognized, the master agent function element 152 may determine an agent that has a dialogue with the occupant on the basis of the wake word.”). REGARDING CLAIM 3, SUGIHARA in view of KEPHART discloses the non-transitory computer-readable storage medium of claim 2. SUGIHARA further discloses the method/system, wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the electronic device to: in accordance with a determination that the user voice input does not include the first digital assistant trigger or the second digital assistant trigger (Par 63 – “For example, the master agent function element 152 recognizes a word for calling any agent (interactive agent) such as an input speech sound “Hey!” or “Hello!” or recognizes a word (for example, a wake word) for designating and calling an agent implemented by each of the sub-agent function elements 154A to 154C.”; Par 68 – “Also, when a wake word for calling a specific agent has been recognized, the master agent function element 152 may determine an agent that has a dialogue with the occupant on the basis of the wake word.”), [determine whether a gaze of a user of the electronic device is directed at the first avatar when the user voice input is received; and in response to determining that the gaze of the user is directed at the first avatar when the user voice input is received], perform natural language processing based on the user voice input to determine a domain associated with the user voice input (Par 62 –“The master agent function element 152 executes natural language processing on the text information obtained through the conversion and recognizes the meaning of the text information.”; Par 64 –“For example, when the meaning such as “Turn on the air conditioner” or “Please turn on the air conditioner” is recognized as the recognition result, the master agent function element 152 generates a command obtained through the replacement with the standard text information “activation of the air conditioner.””; Par 67 – “For example, when the recognized command is “activation of the air conditioner,” the master agent function element 152 determines an agent capable of executing in-vehicle equipment control, which is control corresponding to the command, with reference to the function information table 172. In the example of FIG. 4, the master agent function element 152 acquires that the agent capable of starting the air conditioner is agent A and determines agent A that has a dialogue with the occupant.”; Fig. 4 – “Function Information: In-vehicle equipment control, Shop/store search, Route Guidance ….”). SUGIHARA does not explicitly teach the [square-bracketed] limitations. In other words, SUGIHARA teaches different methods for invoking a digital assistant: 1) by using a wake word for calling a specific agent and 2) by matching the recognized command (i.e., a domain associated with the user voice input) with the function information associated with each digital assistant. Thus, when a wake word for calling specific digital assistant is not included in the user voice command, the method/system matches the domain associated with the user voice input to the function information (e.g., is the command related to in-vehicle equipment control, shop/store search, route guidance, etc. ?) KEPHART discloses the [square-bracketed] limitations. KEPHART discloses a method/system utilizing speech signals for virtual assistants comprising: in accordance with a determination that the user voice input does not include the first digital assistant trigger or the second digital assistant trigger (KEPHART Par 179 – “Specifically regarding the first challenge, one approach to distinguishing which agent is being addressed is to require the human participant to use names associated with each agent. However, while this “wake-word” approach is acceptable for one-shot interactions with an agent (as is the case for Amazon's “Alexa®” agent) (registered mark of AMAZON TECHNOLOGIES, INC. NORTH SEATTLE WASH.), it proves tedious and unnatural in extended dialogs such as negotiations. Various prior-art approaches have sought techniques for determining the addressee without resorting to a wake word, via various multimodal cues such as intonation, pitch, head-gaze, vocal energy, and the like.”), [determine whether a gaze of a user of the electronic device is directed at the first avatar when the user voice input is received (KEPHART Par 179 – “Various prior-art approaches have sought techniques for determining the addressee without resorting to a wake word, via various multimodal cues such as intonation, pitch, head-gaze, vocal energy, and the like. This has been done to determine the addressee in human-kiosk, human-robot, human-human, and human-human-agent conversations, and in the human robot interaction field, using approaches such as identifying visual focus of attention or moving the robot's head to signify turns. Given that it is common for people to look at the AI agent that they are speaking to, especially when the AI agent is embodied as an animated avatar, an embodiment of this invention uses a simple approach based on head pose coupled with semiotics of inferred user attention. Specifically, at any given moment in time, the agent to which a user is paying attention is inferred by using an algorithm to determine the user's head orientation, projecting that orientation onto the display, and identifying the closest avatar lying within a specified angular or linear distance (if any). Then, to determine the addressee for a given utterance, the amount of time that the user was looking at each agent during that utterance is computed, and the agent that was being looked at the most during that utterance is identified as the addressee (in some instances, only when the amount of time exceeds a specified threshold).”); and in response to determining that the gaze of the user is directed at the first avatar when the user voice input is received] (KEPHART Par 179 – “Specifically, at any given moment in time, the agent to which a user is paying attention is inferred by using an algorithm to determine the user's head orientation, projecting that orientation onto the display, and identifying the closest avatar lying within a specified angular or linear distance (if any). Then, to determine the addressee for a given utterance, the amount of time that the user was looking at each agent during that utterance is computed, and the agent that was being looked at the most during that utterance is identified as the addressee (in some instances, only when the amount of time exceeds a specified threshold).”), perform natural language processing based on the user voice input to determine a domain associated with the user voice input (KEPHART Par 179 – “Then, to determine the addressee for a given utterance, the amount of time that the user was looking at each agent during that utterance is computed, and the agent that was being looked at the most during that utterance is identified as the addressee (in some instances, only when the amount of time exceeds a specified threshold).”; Par 68 – “In the non-limiting example of FIG. 5, element 387 is a Natural Language Understanding module such as IBM's Watson™ NLU that assists with interpreting human utterances as negotiation or other speech acts; element 388 is a conversation agent such as IBM's Watson™ Assistant that can be used in conjunction with element 387 to help classify human utterances into various categories of negotiation or other speech act; and element 389 is a representation of the conversational context, which can be maintained by the conversation agent internal to the service (as it is for Watson™ Assistant) or externally, as shown in the example.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method/system of SUGIHARA to include invoking a digital assistant using gaze detection, as taught by KEPHART. One of ordinary skill would have been motivated to include invoking a digital assistant using gaze detection, in order to provide more natural interaction between a user and a digital assistant. REGARDING CLAIM 6, SUGIHARA in view of KEPHART discloses the non-transitory computer-readable storage medium of claim 3, wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the electronic device to: determine whether the determined domain associated with the user voice input matches the first domain corresponding to the first domain-specific digital assistant (SUGIHARA Par 67 –“Par 64 –“For example, when the meaning such as “Turn on the air conditioner” or “Please turn on the air conditioner” is recognized as the recognition result, the master agent function element 152 generates a command obtained through the replacement with the standard text information “activation of the air conditioner.””; Par 67 – “For example, when the recognized command is “activation of the air conditioner,” the master agent function element 152 determines an agent capable of executing in-vehicle equipment control, which is control corresponding to the command, with reference to the function information table 172.”; Fig. 4 – “Function Information: In-vehicle equipment control, Shop/store search, Route Guidance ….”). REGARDING CLAIM 7, SUGIHARA in view of KEPHART discloses the non-transitory computer-readable storage medium of claim 6, wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the electronic device to: in accordance with a determination that the determined domain matches the first domain (SUGIHARA Par 64 –“For example, when the meaning such as “Turn on the air conditioner” or “Please turn on the air conditioner” is recognized as the recognition result, the master agent function element 152 generates a command obtained through the replacement with the standard text information “activation of the air conditioner.””; Par 67 – “For example, when the recognized command is “activation of the air conditioner,” the master agent function element 152 determines an agent capable of executing in-vehicle equipment control, which is control corresponding to the command, with reference to the function information table 172.”), determine, based on the user voice input, a second digital assistant response using the first domain-specific digital assistant (SUGIHARA Par 92 – “The dialogue generator 230 acquires the control content associated with the command For example, when the command is “activation of the conditioner,” the dialogue generator 230 acquires control content of the equipment control for activating the air conditioner installed in the vehicle M and control content of speech sound control for outputting the speech sound “The air conditioner has been activated,” and control content of display control for displaying a temperature within the vehicle cabin and a set temperature.”); and output, via the speaker, the second digital assistant response (SUGIHARA Fig. 10 – “Air Conditioner has been Activated”; Par 19 – “and causing, by the computer, a speaker configured to output a speech sound inside the vehicle cabin to output the generated speech sound of the agent”). REGARDING CLAIM 8, SUGIHARA in view of KEPHART discloses the non-transitory computer-readable storage medium of claim 6, wherein the electronic device invokes the second digital assistant in response to determining that the determined domain does not match the first domain (SUGIHARA Par 109 – “In the third scene, the master agent function element 152 determines agent (agent B in the example of FIG. 11) capable of implementing a function corresponding to the command as an agent that has a dialogue with the occupant with reference to command information of the function information table 172 on the basis of a command corresponding to a request recognized from the speech sound of the occupant. At this time, the master agent function element 152 changes a display form so that the agent image EIB is displayed in front of the other agent images EIA and EIC at the timing when a main agent that has a dialogue with the occupant is switched from the sub-agent function element 154A to the sub-agent function element 154B.”). REGARDING CLAIM 9, SUGIHARA in view of KEPHART discloses the non-transitory computer-readable storage medium of claim 8, wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the electronic device to: in response to determining that the determined domain does not match the first domain (SUGIHARA Par 109 – “In the third scene, the master agent function element 152 determines agent (agent B in the example of FIG. 11) capable of implementing a function corresponding to the command as an agent that has a dialogue with the occupant with reference to command information of the function information table 172 on the basis of a command corresponding to a request recognized from the speech sound of the occupant. At this time, the master agent function element 152 changes a display form so that the agent image EIB is displayed in front of the other agent images EIA and EIC at the timing when a main agent that has a dialogue with the occupant is switched from the sub-agent function element 154A to the sub-agent function element 154B.”) and prior to invoking the second domain-specific digital assistant (SUGIHARA Fig. 11 – “The Request is answered by Agent B” -> “I will Respond”; Par 110 – “For example, as shown in FIG. 11, when the main agent having the dialogue with the occupant is switched from agent A to agent B, the master agent function element 152 causes an agent speech sound such as “The request is answered by agent B.” to be output from agent A and causes an agent speech sound such as “I will respond.” to be output from agent B.”: output a first audio output that informs the user of the electronic device that the first domain-specific digital assistant is not capable of providing a digital assistant response based on the user voice input (SUGIHARA Fig. 11 – “The Request is answered by Agent B”; Par 120 – “Also, when the function for the request cannot be executed, the master agent function element 152 determines another sub-agent function element capable of executing the function for the request among the plurality of sub-agent function elements 154 (step S112).”; Par 110 – “For example, as shown in FIG. 11, when the main agent having the dialogue with the occupant is switched from agent A to agent B, the master agent function element 152 causes an agent speech sound such as “The request is answered by agent B.” to be output from agent A and causes an agent speech sound such as “I will respond.” to be output from agent B.”; Par 111 – “Also, the master agent function element 152 switches the speech sound input collected by the microphone 124B from the speech sound input interface 154Aa of the sub-agent function element 154A to the speech sound input interface 154Ba of the sub-agent function element 154B. Thereby, a dialogue or the like can be implemented between agent B and the occupant.”). REGARDING CLAIM 11, SUGIHARA in view of KEPHART discloses the non-transitory computer-readable storage medium of claim 1, wherein invoking the second domain-specific digital assistant includes displaying, within the VR environment (Note that KEPHART already teaches the VR environment as explained the rejection of claim 1.), a second avatar representing the second domain-specific digital assistant (SUGIHARA Fig. 11; Par 109 – “At this time, the master agent function element 152 changes a display form so that the agent image EIB is displayed in front of the other agent images EIA and EIC at the timing when a main agent that has a dialogue with the occupant is switched from the sub-agent function element 154A to the sub-agent function element 154B.”). REGARDING CLAIM 12, SUGIHARA in view of KEPHART discloses the non-transitory computer-readable storage medium of claim 11, wherein the electronic device outputs the first digital assistant response such that the second avatar appears to be speaking the first digital assistant response (SUGIHARA Par 110 – “For example, as shown in FIG. 11, when the main agent having the dialogue with the occupant is switched from agent A to agent B, the master agent function element 152 causes an agent speech sound such as “The request is answered by agent B.” to be output from agent A and causes an agent speech sound such as “I will respond.” to be output from agent B. In this case, the master agent function element 152 causes a sound image position MPA of the agent speech sound for agent A to be localized near the display position of the agent image EIA and causes a sound image position MPB of the agent speech sound for agent B to be localized near the display position of the agent image EIB. Thereby, the occupant can be allowed to sense that smooth cooperation is being performed between the agents.”). REGARDING CLAIM 14, SUGIHARA in view of KEPHART discloses the non-transitory computer-readable storage medium of claim 1, wherein the domain-specific digital assistant determines digital assistant responses only for user inputs associated with the first domain (SUGIHARA Fig. 4 – “Function Information”; Par 66 – “FIG. 4 is a diagram showing an example of content of the function information table 172. In the function information table 172, the agent identification information is associated with the function information. The function information includes, for example, in-vehicle equipment control, a shop/store search, route guidance, traffic information notification, radio control, household equipment control, and product ordering. Also, the agent information includes, for example, agents A to C implemented by the sub-agent function elements 154A to 154C. Also, although “1” is stored in the function that can be implemented by the agent and “0” is stored in the function that cannot be implemented in the example of FIG. 4, other identification information may be used.”; In other words, some agents are capable for executing commands related to more functions (i.e., domains) than other agents. For example, Agent B can execute commands related to 4 domains while Agent A can execute commands related to 3 domains, etc. Thus, there is an agent (say Agent X), which is only capable for executing commands related to a single domain.). REGARDING CLAIM 15, SUGIHARA in view of KEPHART discloses the non-transitory computer-readable storage medium of claim 1, wherein the second domain-specific digital assistant corresponding to a second domain that is different from the first domain (SUGIHARA Fig. 4 – “In-Vehicle Equipment Control” -> “Traffic Information Notification”; Par 109 – “In the third scene, the master agent function element 152 determines agent (agent B in the example of FIG. 11) capable of implementing a function corresponding to the command as an agent that has a dialogue with the occupant with reference to command information of the function information table 172 on the basis of a command corresponding to a request recognized from the speech sound of the occupant. At this time, the master agent function element 152 changes a display form so that the agent image EIB is displayed in front of the other agent images EIA and EIC at the timing when a main agent that has a dialogue with the occupant is switched from the sub-agent function element 154A to the sub-agent function element 154B.”; Note Fig. 4 shows different domains (i.e., function information) associated with different Agents.), and wherein the second domain-specific digital assistant determines digital assistant responses only for user inputs associated with the second domain (SUGIHARA Par 56 – “Also, the map information 142 may include road information, traffic regulation information, address information (an address and a postal code), facility information, telephone number information, and the like. The map information 142 may be updated at any time by the communicator 110 communicating with another device.”; In other words, Agent A is capable for executing commands related to “In-vehicle Equipment Control” domain, whereas Agent B is capable for executing commands related to “Traffic Information Notification” Domain. Since some agents are capable for executing commands related to more functions (i.e., domains) than other agents, Agent B can be assigned with a single domain (e.g., Traffic Information Notification).). REGARDING CLAIM 16, SUGIHARA in view of KEPHART discloses the non-transitory computer-readable storage medium of claim 1, wherein the first domain-specific digital assistant corresponds to a first application and the second domain-specific digital assistant corresponds to a second application (SUGIHARA Par 57 – “Also, the navigator 140 may be implemented by, for example, the function of a terminal device such as a smartphone or a tablet terminal possessed by the occupant. Also, the navigator 140 may transmit a current position and a destination to the server 200 or the navigation server via the communicator 110 and acquire a route equivalent to the route on the map from the server 200 or the navigation server. Also, the navigator 140 may implement the above-described function of the navigator 140 according to a function of a navigation application which is executed by the agent controller 150.”; Par 75 – “Also, the sub-agent function element 154 may receive the input of the speech sound from the microphone 124B or the dialogue information and the output control information obtained from the communicator 110 through an application programming interface (API), select function elements for executing a process based on the received input (the display controller 156, the speech sound controller 158, and the equipment controller 160), and cause the selected function element to execute the process via the API.”), different from the first application (SUGIHARA Fig. 4 – “Function Information; In-vehicle Equipment control; shop/store search; route guidance; traffic information notification; radio control; household equipment control; product ordering”). REGARDING CLAIM 17, SUGIHARA in view of KEPHART discloses the non-transitory computer-readable storage medium of claim 1, wherein the first domain-specific digital assistant corresponds to a knowledge base unavailable to a general digital assistant (SUGIHARA Fig. 1 – “Server 200A; 200B; 200C”; Fig. 4 – “Function Information”; Fig. 9 EIA: MPA – “Are there any request?”; Par 98 –“For example, when the word for calling any agent has been recognized, the master agent function element 152 determines agent A of the sub-agent function element 154A designated in advance as an agent that has a dialogue with the occupant. The master agent function element 152 causes the agent image EIA having a dialogue among the agent images EIA to EIC corresponding to the sub-agent function elements 154A to 154C generated by the display controller 156 to be displayed in front of the other agents EIB and EIC when viewed from the occupant. Also, the master agent function element 152 may adjust the size of each agent image in accordance with a positional relationship between the agent images EIA to EIC on the image space.”; Par 66 – “FIG. 4 is a diagram showing an example of content of the function information table 172. In the function information table 172, the agent identification information is associated with the function information. The function information includes, for example, in-vehicle equipment control, a shop/store search, route guidance, traffic information notification, radio control, household equipment control, and product ordering. Also, the agent information includes, for example, agents A to C implemented by the sub-agent function elements 154A to 154C. Also, although “1” is stored in the function that can be implemented by the agent and “0” is stored in the function that cannot be implemented in the example of FIG. 4, other identification information may be used.”; Par 91 – “FIG. 8 is a diagram showing an example of content of the answer information DB 244B provided in the server 200B. In the answer information DB 244B, for example, the command information is associated with the control content to be executed by the sub-agent function element 154B.” In other words, Agent A can be designated to have a dialogue with the occupant in advance. In this case, Agent A acts as a general digital assistant because it will be a first agent to respond to the occupant’s calling. When the occupant issues a specific command (e.g., controlling radio radio), Agent B will be invoked to respond to the command, because Agent A cannot execute the function. In other words, Agent A does not have an access to the radio controlling function and its corresponding knowledge base 200B that is associated with Agent B’s functions.). REGARDING CLAIM 18, SUGIHARA in view of KEPHART discloses an electronic device comprising: a display (SUGIHARA Fig. 1 – “Display”); one or more processors (SUGIHARA Fig. 1 – “Agent Controller”); a memory (SUGIHARA Fig. 1 – “Storage”); and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions (SUGIHARA Par 58 – “CPU executing a program”) for: performing the steps of Claim 1; thus, it is rejected under the same rationale. REGARDING CLAIM 19, SUGIHARA in view of KEPHART discloses a method, comprising: at an electronic device (SUGIHARA Fig. 1 – “Agent Device”) having one or more processors (SUGIHARA Fig. 1 – “Agent Controller”), memory (Fig. 1 – “Storage”), and a display (SUGIHARA Fig. 1 – “Display”): performing the steps of Claim 1; thus, it is rejected under the same rationale. Claims 4-5 are rejected under 35 U.S.C. 103 as being unpatentable over SUGIHARA (US 2022/0005470 A1) in view of KEPHART (US 2022/0028366 A1), and in further view of KOLL (US 2023/0120370 A1). REGARDING CLAIM 4, SUGIHARA in view of KEPHART discloses the non-transitory computer-readable storage medium of claim 3. SUGIHARA in view of KEPHART teaches performing natural language processing on the user voice input to determine a domain associated with the user voice input when the gaze is detected, but does not explicitly teach detecting a gaze after the voice input is received. KOLL discloses a method/system for tracking user attention in conversational agent systems, wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the electronic device to: in accordance with a determination that the gaze of the user is not directed at the first avatar when the user voice input is received, determine whether the gaze of the user is directed at the first avatar within a predetermined period of time after the user voice input is received (KOLL Fig. 2B; Par 29 – “Referring ahead to FIG. 2B, a flow diagram depicts one embodiment of the method 200 expanding upon FIG. 2A, 206 , in which the results processing component determines that the attention tracking component received the non-audio input within a first time interval proximate to a second time interval during which the audio input was received. As depicted in FIG. 2B, the method 200 may include receiving, by the results processing component 112, from the attention tracking component 102, an indication of a first time interval during which the attention tracking component 102 received the non-speech input 104 (206 a). The method 200 may include determining, by the results processing component, that the first time interval followed a second time interval during which the speech processing component received at least a portion of the audio input (206 b). The method 200 may include storing, by the results processing component, the at least the portion of the audio input received during the second time interval (206 c). Therefore, the method 200 may include storing speech that occurs during a time before the attention tracking component 102 determines that a user has an intention to interact with the system and that speech may be processed and the system may direct the execution of actionable commands within that speech even though the user does not indicate an intention to interact with the system until after the commands were spoken.”); and in response to determining that the gaze of the user is directed at the first avatar within the predetermined period of time, perform natural language processing based on the user voice input to determine a domain associated with the user voice input (KOLL Fig. 2B; Par 29 – “The method 200 may include storing, by the results processing component, the at least the portion of the audio input received during the second time interval (206 c). Therefore, the method 200 may include storing speech that occurs during a time before the attention tracking component 102 determines that a user has an intention to interact with the system and that speech may be processed and the system may direct the execution of actionable commands within that speech even though the user does not indicate an intention to interact with the system until after the commands were spoken.”; In other words, the system checks whether the non-audio input (i.e., gaze) is received within the first time period after the second time interval (i.e., the period of receiving audio input). If so, the system concludes that user intends to interact with the system, and the system process the user input to execute an actionable command. If the non-audio input (i.e., gaze) is not received within the time period after the second time interval (i.e., the period of receiving audio input), the system concludes the user does not intend to interact with the system, and does not process the user input.). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method/system of SUGIHARA in view of KEPHART to include checking the gaze of the user before/after receiving a user speech input, as taught by KOLL. One of ordinary skill would have been motivated to include checking the gaze of the user before/after receiving a user speech input, in order to accurately determine user’s intention to interact with a virtual assistant. REGARDING CLAIM 5, SUGIHARA in view of KEPHART and KOLL discloses the non-transitory computer-readable storage medium of claim 4, wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the electronic device to: in response to determining that the gaze of the user is not directed at the first avatar within the predetermined period of time, forgo further processing of the user voice input (KOLL Par 25 – “The method 200 may include determining, by the speech processing component, that the received audio input includes at least one actionable command. In some embodiments, the speech processing component 108 may receive all captured audio input but determine whether or not to process (e.g., apply a speech recognition process to the audio input) any portion of the speech in the audio input based on an instruction from the results processing component; ….”; Par 27 – “The results processing component 112 may determine whether the non-audio input was received within the first time interval proximate to the receiving of the audio input to determine whether or not the system should conclude that the user intended to interact with the system to have at least one actionable command executed.”; In other words, the system checks whether the non-audio input (i.e., gaze) is received within the first time period. If so, the system concludes that user intends to interact with the system, and the system process the user input to execute an actionable command. If the non-audio input (i.e., gaze) is not received within the time period, the system concludes the user does not intend to interact with the system, and does not process the user input.). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method/system of SUGIHARA in view of KEPHART to include checking the gaze of the user before/after receiving a user speech input, as taught by KOLL. One of ordinary skill would have been motivated to include checking the gaze of the user before/after receiving a user speech input, in order to accurately determine user’s intention to interact with a virtual assistant. Claims 10 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over SUGIHARA (US 2022/0005470 A1) in view of KEPHART (US 2022/0028366 A1), and in further view of BROWN (US 2015/0186156 A1). REGARDING CLAIM 10, SUGIHARA in view of KEPHART discloses the non-transitory computer-readable storage medium of claim 9. SUGIHARA in view of KEPHART does not explicitly teach requesting user approval to invoke the second digital assistant. BROWN discloses a method/system for interacting with multiple digital assistants, wherein the first audio output includes a user query that requests user approval to invoke the second domain-specific digital assistant (BROWN Fig. 12 -- “Would you like to talk to a sports virtual assistant?”; Par 35 – “Each virtual assistant of the virtual assistant team 102 may be configured for multi-modal input/output (e.g., receive and/or respond in audio or speech, text, touch, gesture, etc.), multi-language communication (e.g., receive and/or respond according to any type of human language) … ”), and wherein the electronic device invokes the second domain-specific digital assistant in response to receiving a user input representing a user approval for invoking the second domain-specific digital assistant (BROWN Par 142 – “Here, a suggestion to switch to a sports virtual assistant 1202 is made by rotating an icon 1204 (e.g., ribbon) to present the sports virtual assistant 1202 in a center of the conversation user interface 1200. In other examples, the suggestion may be made in other ways. In this example, the user has accepted the suggestion by selecting the sports virtual assistant 1202, and the sports virtual assistant 1202 is enabled to interact with the user, as illustrated by an icon 1206. The sports virtual assistant 1202 may then communicate with the user.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method/system of SUGIHARA in view of KEPHART to include requesting user approval for invoking a digital assistant, as taught by BROWN. One of ordinary skill would have been motivated to include requesting user approval for invoking a digital assistant, in order to prevent unintended invocation of a service. REGARDING CLAIM 13, SUGIHARA in view of KEPHART discloses the non-transitory computer-readable storage medium of claim 1. SUGIHARA in view of KEPHART does not explicitly teach a predefined voice setting associated with a digital assistant. BROWN discloses a method/system for interacting with multiple digital assistants, wherein the electronic device outputs the first digital assistant response based on at least one of a predefined voice setting associated with the second domain-specific digital assistant (BROWN Par 99 – “An audible manner of output—how a virtual assistant speaks to a user. This may include an accent of the virtual assistant (e.g., English, Australian, etc.), a fluctuation in the virtual assistants speech (e.g., pronouncing a first word of a sentence different than other words of a sentence), how fast words are spoken, and so on.”) and a predefined language setting associated with the second domain-specific digital assistant (BROWN Par 100 – “A language in which a virtual assistant communicates (e.g., Spanish, German, French, English, etc.). This may include a language that is understood by the virtual assistant and/or a language that is spoken or otherwise used to output information by the virtual assistant.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method/system of SUGIHARA in view of KEPHART to include a predefined voice setting for a digital assistant, as taught by BROWN. One of ordinary skill would have been motivated to include a predefined voice setting for a digital assistant, in order to interact with a user according the user’s preference. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to JONATHAN C KIM whose telephone number is (571)272-3327. The examiner can normally be reached Monday to Friday 8:00 AM thru 4:00 PM EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew C Flanders can be reached at 571-272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JONATHAN C KIM/Primary Examiner, Art Unit 2655
Read full office action

Prosecution Timeline

May 09, 2025
Application Filed
Jan 02, 2026
Response after Non-Final Action
Sep 22, 2026
Non-Final Rejection mailed — §103, §DOUBLEPATENT (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12744034
ELECTRONIC DEVICE AND CONTROL METHOD THEREOF
2y 8m to grant Granted Sep 22, 2026
Patent 12744038
SYSTEMS AND TECHNIQUES FOR USING A DIGITAL ASSISTANT WITH AN ENHANCED ENDPOINTER
2y 6m to grant Granted Sep 22, 2026
Patent 12711982
AUDIO PROCESSING
3y 3m to grant Granted Aug 18, 2026
Patent 12700403
METHOD, APPARATUS, ELECTRONIC DEVICE AND STORAGE MEDIUM FOR TEXT CONTENT MATCHING
2y 5m to grant Granted Aug 04, 2026
Patent 12688850
ELECTRONIC DEVICE AND CONTROL METHOD THEREFOR
2y 5m to grant Granted Jul 21, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
74%
Grant Probability
99%
With Interview (+38.7%)
2y 5m (~1y 0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 368 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month