Prosecution Insights
Last updated: October 01, 2026
Application No. 19/179,613

VIRTUAL ASSISTANT INTERACTIONS IN A 3D ENVIRONMENT

Non-Final OA §103
Filed
Apr 15, 2025
Priority
May 09, 2024 — provisional 63/644,809
Examiner
CHEN, YU
Art Unit
Tech Center
Assignee
Apple Inc.
OA Round
1 (Non-Final)
68%
Grant Probability
Favorable
1-2
OA Rounds
1y 4m
Est. Remaining
98%
With Interview

Examiner Intelligence

Grants 68% — above average
68%
Career Allowance Rate
738 granted / 1087 resolved
+7.9% vs TC avg
Strong +30% interview lift
Without
With
+29.6%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
85 currently pending
Career history
1184
Total Applications
across all art units

Statute-Specific Performance

§101
2.3%
-37.7% vs TC avg
§103
46.4%
+6.4% vs TC avg
§102
23.6%
-16.4% vs TC avg
§112
22.4%
-17.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1087 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1-7, 12-20 are rejected under 35 U.S.C. 103 as being unpatentable over Lai et al. (US Pub 2023/0316594 A1) in view of Klingler et al. (US Pub 2025/0181207 A1). As to claim 1, Lai discloses a method comprising: at an electronic device having a processor, a display, and one or more sensors (Fig .2A): receiving data corresponding to a first user activity in the 3D coordinate system for a first period of time (¶0044, “The user input may include text (e.g., online chat), especially in an instant messaging application or other applications, voice, eye-tracking, user motion such as gestures or running, or a combination of them.” ¶0045, “the virtual assistant application 130 may be accessed via a web browser 135. In some instances, the virtual assistant application 130 passively listens to and watches interactions of the user in the real-world, and processes what it hears and sees (e.g., explicit input such as audio commands or interface commands, contextual awareness derived from audio or physical actions of the user, objects in the real-world, environmental triggers such as weather or time, and the like) in order to interact with the user in an intuitive manner.”); identifying a user interaction event associated with the virtual assistant in the 3D environment based on the data corresponding to the user activity (¶0044, “The virtual assistant may perform concierge-type services (e.g., making dinner reservations, purchasing event tickets, making travel arrangements, and the like), provide information (e.g., reminders, information concerning an object in an environment, information concerning a task or interaction, answers to questions, training regarding a task or activity, and the like), goal assisted services (e.g., generating and implementing an exercise regimen to achieve a certain level of fitness or weight loss, implementing electronic devices such as lights, heating, ventilation, and air conditioning systems, coffee maker, television, etc. generate and execute a morning routine such as wake up, get ready for work, make breakfast, and travel to work, and the like), or combinations thereof.”); providing a graphical indication corresponding to one or more attributes associated with the virtual assistant based on identifying the user interaction event (¶0046, “The virtual assistant application 130 may present the response to the user at the client system 130 (e.g., rendering virtual content overlaid on a real-world object within the display). The presented responses may be based on different modalities such as audio, text, image, and video. As an example, and not by way of limitation, context concerning activity of a user in the physical world may be analyzed and determined to initiate an interaction for completing an immediate task or goal” “The virtual assistant application 110 may then present the traffic information to the user as text (e.g., as virtual content overlaid on the physical environment such as real-world object) or audio (e.g., spoken to the user in natural language through a speaker associated with the client system 105).” ¶0148, “the virtual assistant generates a graph of objects, attributes, and relationships between objects extracted from the input data.”); and in accordance with receiving data corresponding to a second user activity for a second period of time, generating one or more user interface elements that are positioned at 3D positions based on the 3D coordinate system associated with the 3D environment (¶0096, “The virtual assistant 500 incorporates elements of interactive responses (e.g., voice or text) and context awareness to assist, e.g., deliver information and services, users via one or more interactions.” Fig. 8, ¶0132, “in response to request by the user for a user interface, the client system generates and renders the user interface in the extended reality environment displayed to the user. The user interface includes one or more user interface elements for interacting with the virtual assistant. In some instances, the user interface is rendered at a position locked relative to a physical or virtual object. For example, the client system may generate and render a user interface including one or more user interface elements (e.g., virtual buttons) on the surface of a physical object.” ¶0138, “In order to make the modification, the user can request that the client system render a user interface 930 to interact with the virtual assistant and communicate the modification.”). Lai does not explicitly disclose presenting a view of a three-dimensional (3D) environment, wherein a virtual assistant is positioned at a 3D position based on a 3D coordinate system associated with the 3D environment. Klingler discloses presenting a view of a three-dimensional (3D) environment, wherein a virtual assistant is positioned at a 3D position based on a 3D coordinate system associated with the 3D environment (Klingler, Fig. 13A to F, ¶0008, “implement an interactive agent (e.g., bot, avatar, digital human, or robot). For example, systems and methods are disclosed that implement or support an interaction modeling language and/or interaction modeling application programming interface (API) that uses a standardized interaction categorization schema, multimodal human-machine interactions, backchanneling, an event-driven architecture, management of interaction flows, deployment using one or more large language models, sensory processing and action execution, interactive visual content, interactive agent (e.g., bot) animations, expectations actions and signaling, and/or other features.” ¶0010, “an interactive avatar (e.g., an animated digital character) or other bot may support any number of simultaneous interaction modalities and corresponding interaction channels to engage with the user, such as channels for character or bot actions (e.g., speech, gestures, postures, movement, vocal bursts, etc.), scene actions (e.g., two-dimensional (2D) GUI overlays, 3D scene interactions, visual effects, music, etc.), and user actions (e.g., speech, gesture, posture, movement, etc.). Actions based on different modalities may occur sequentially or in parallel (e.g., waving and saying hello).” ¶0277, “interactive agents such as digital avatars may be developed and/or deployed for various applications, such as customer service, virtual assistants, interactive entertainment or gaming, digital twins (e.g., for video conferencing participants), education or training, health care, virtual or augmented reality experiences, social media interactions, marketing and advertising, and/or other applications.”). Lai and Klingler are considered to be analogous art because all pertain to interactive systems. It would have been obvious before the effective filing date of the claimed invention to have modified Lai with the features of “presenting a view of a three-dimensional (3D) environment, wherein a virtual assistant is positioned at a 3D position based on a 3D coordinate system associated with the 3D environment” as taught by Klingler. The suggestion/motivation would have been in order to develop and/or deploy interactive agents (Klingler, ¶0019). As to claim 2, claim 1 is incorporated and the combination of Lai and Klingler discloses the one or more user interface elements are customized based on a large language model (LLM) associated with the virtual assistant (Lai, ¶0048, “The virtual assistant engine 110 may use artificial intelligence systems 140 (e.g., rule based systems or machine-learning based systems such as natural-language understanding models) to analyze the input based on a user's profile and other relevant information. The result of the analysis may comprise different interactions associated with a task or goal of the user.” “The virtual assistant engine 110 may generate virtual content for the user using various techniques such as natural-language generating, virtual object rendering, and the like.”. Klingler, ¶0054, “deployment using one or more language models (e.g., LLMs, VLMs, multimodal language models, etc.),”). As to claim 3, claim 1 is incorporated and the combination of Lai and Klingler discloses the one or more attributes associated with the virtual assistant are based on adjustable settings that comprises at least one of: a type of large language model (LLM); a type of personality; a response style; a temperature style; and a pedagogical approach selection (Lai, ¶0039, “generating a graph of objects, attributes, and relationships between objects extracted from the input data; determining one or more interactions to be presented, initiated, or executed based on the graph and a profile associated with the user; determining virtual content data to be used for rendering virtual content based on the one or more interactions; and rendering the virtual content in the extended reality environment displayed to the user based on the virtual content data. The virtual content is used to present, initiate, or execute the one or more interactions for the user.” ¶0100, “identify user information 537 within the data 540 based on the analysis, and recommend or select the goals 520 and action spaces 527 for the user profile 515 based on the analysis. In some instances, a user can make adjustments or modifying the goals 525, action spaces 527 (including the rules, decisions trees, or vectors), and user information 537 within the user profile 515 using the virtual assistant application 505 or a separate application accessed via the client system.” ¶0120-¶0123, Fig. 7A -7 C.) As to claim 4, claim 1 is incorporated and the combination of Lai and Klingler discloses generating the one or more user interface elements comprises determining one or more candidate representations based on a determined context of one or more utterances (Lai, ¶0003, “for an individual based on voice or text utterances (e.g., commands or questions).”, “As the number of utterances increase, the software learns over time what users want when they provide various utterances.” ¶0006, “analyzing information and contextual clues from an extended reality environment and based on the analysis, intuitively superimposing and integrating customized digital information into the artificial reality environment via a virtual assistant to recommend and lead the user into suggested action. The information and contextual clues analyzed for providing the interaction may include various inputs such as eye-tracking, user gestures, environmental sensor input, or input obtainable from remote devices.”). As to claim 5, claim 4 is incorporated and the combination of Lai and Klingler discloses the one or more candidate representations are updated based on data corresponding to user activity for a second period of time (Lai, ¶0123, “The virtual assistant continues to execute process 600 (e.g., continuously, semi-continuously, or periodically) through-out execution of the yoga regimen to ensure the user remains engaged and is following the tasks defined for the workflow or to determine whether a new interaction is trigged by the context of the new data.” ¶0158, “the virtual assistant can guide the user meet-up with the friend in accordance with the tasks and virtual content data associated with the workflow (e.g., displaying a route to meet the friend). The virtual assistant continues to execute process 1200 (e.g., continuously, semi-continuously, or periodically) through-out execution of the meet a friend interaction to ensure the user remains engaged and is following the tasks defined for the workflow or to determine whether a new interaction is trigged by the context of the new data.”). As to claim 6, claim 4 is incorporated and the combination of Lai and Klingler discloses the one or more candidate representations comprises at least one of: a candidate text representation; a candidate audio representation; a candidate image representation; a candidate video representation; and a candidate virtual object representation (Lai, ¶0123, “The virtual assistant continues to execute process 600 (e.g., continuously, semi-continuously, or periodically) through-out execution of the yoga regimen to ensure the user remains engaged and is following the tasks defined for the workflow or to determine whether a new interaction is trigged by the context of the new data.” ¶0158, “the virtual assistant can guide the user meet-up with the friend in accordance with the tasks and virtual content data associated with the workflow (e.g., displaying a route to meet the friend). The virtual assistant continues to execute process 1200 (e.g., continuously, semi-continuously, or periodically) through-out execution of the meet a friend interaction to ensure the user remains engaged and is following the tasks defined for the workflow or to determine whether a new interaction is trigged by the context of the new data.” ¶0046, “The virtual assistant application 110 may then present the traffic information to the user as text (e.g., as virtual content overlaid on the physical environment such as real-world object) or audio (e.g., spoken to the user in natural language through a speaker associated with the client system 105).”). As to claim 7, claim 6 is incorporated and the combination of Lai and Klingler discloses the candidate virtual object representation comprises a 3D interactive model (Lai, ¶0051, “content objects may include virtual objects such as virtual interfaces, 2D or 3D graphics, media content, or other suitable virtual objects.” ¶0062, “stereo video that produces a three-dimensional (3D) effect to the viewer”). As to claim 12, claim 1 is incorporated and the combination of Lai and Klingler discloses determining a context of a user based on at least one of the first user activity, the second user activity, and one or more physiological cues of the user; and updating the graphical indication of the virtual assistant is based on the determined context (Lai, ¶0045, “the virtual assistant application 130 passively listens to and watches interactions of the user in the real-world, and processes what it hears and sees (e.g., explicit input such as audio commands or interface commands, contextual awareness derived from audio or physical actions of the user, objects in the real-world, environmental triggers such as weather or time, and the like) in order to interact with the user in an intuitive manner.” ¶0046, “The presented responses may be based on different modalities such as audio, text, image, and video. As an example, and not by way of limitation, context concerning activity of a user in the physical world may be analyzed and determined to initiate an interaction for completing an immediate task or goal, which may include the virtual assistant application 130 retrieving traffic information (e.g., via a remote system 115).” ¶0056, “the extended reality application determines interaction information to be presented for the frame of reference of extended reality system 205 and, in accordance with the current context of the user 220, renders the extended reality content 225.”). As to claim 13, claim 1 is incorporated and the combination of Lai and Klingler discloses the data corresponding to the first user activity or the data corresponding to the second user activity is obtained via the one or more sensors on the device (Lai, ¶0056, “data from any external sensors, such as third-party information or device, to capture information within the real world, physical environment, such as motion by user 220 and/or feature tracking information with respect to user 220. Based on the sensed data, the extended reality application determines interaction information to be presented for the frame of reference of extended reality system 205 and, in accordance with the current context of the user 220, renders the extended reality content 225.” ¶0065, “augmented reality system 300 may include one or more sensors, such as sensor 320. Sensor 320 may generate measurement signals in response to motion of augmented reality system 300 and may be located on substantially any portion of frame 310. Sensor 320 may represent one or more of a variety of different sensing mechanisms, such as a position sensor, an inertial measurement unit (IMU), a depth camera assembly, a structured light emitter and/or detector, or any combination thereof. In some embodiments, augmented reality system 300 may or may not include sensor 320 or may include more than one sensor. In embodiments in which sensor 320 includes an IMU, the IMU may generate calibration data based on measurement signals from sensor 320. Examples of sensor 320 may include, without limitation, accelerometers, gyroscopes, magnetometers, other suitable types of sensors that detect motion, sensors used for error correction of the IMU, or some combination thereof.”). As to claim 14, claim 1 is incorporated and the combination of Lai and Klingler discloses the data corresponding to the first user activity or the data corresponding to the second user activity comprises gaze data comprising a stream of gaze vectors corresponding to gaze directions over time during use of the electronic device (Lai, ¶0113, “the virtual content module 580 may trigger generation and rendering of virtual content 543 by the client system (including virtual assistant application 505 and I/O interfaces 545) based on a current field of view of user, as may be determined by real-time gaze tracking of the user, or other conditions.” ¶0121, ¶0155). As to claim 15, claim 1 is incorporated and the combination of Lai and Klingler discloses the data corresponding to the first user activity or the data corresponding to the second user activity comprises an audio stream that includes one or more utterances or instructions received via an input device (Lai, ¶0067, “using higher numbers of acoustic transducers 325 may increase the amount of audio information collected and/or the sensitivity and accuracy of the audio information. In contrast, using a lower number of acoustic transducers 325 may decrease the computing power required by an associated controller 335 to process the collected audio information. In addition, the position of each acoustic transducer 325 of the microphone array may vary. For example, the position of an acoustic transducer 325 may include a defined position on the user, a defined coordinate on frame 310, an orientation associated with each acoustic transducer 325” ¶0107, “If the user input is based on an audio modality (e.g., the user may speak to the virtual assistant application 505 or send a video including speech to the virtual assistant application 505), the virtual assistant engine 510 may process it using an automatic speech recognition (ASR) module 552 to convert the user input into text and use the messaging platform 550 to extract the specific information such as identifying named entities within the text.”). As to claim 16, claim 1 is incorporated and the combination of Lai and Klingler discloses the data corresponding to the first user activity or the data corresponding to the second user activity comprises hands data that includes a hand pose skeleton of multiple joints for each of multiple instants in time during use of the electronic device (Lai, ¶0061, “The client system 200 may detect user interface gestures and other gestures using an inside-out or outside-in tracking system of image capture devices and or external cameras.” ¶0113, “the client system performs object recognition within image data captured by the image capture devices of HMD to identify objects in the physical environment such as the user, the user's hand, and/or physical objects. Further, the client system tracks the position, orientation, and configuration of the objects in the physical environment over a sliding window of time.” ¶0165, “The user can use a gesture or combination of gestures to navigate through the user interface 1430 and view and/or select suggestions 1435.”. Klingler, ¶0150, “the user may make a gesture or pose detected using computer vision” detecting skeleton of multiple joints are obvious and well known in the art when detecting gesture (OFFICIAL NOTICE).). As to claim 17, claim 1 is incorporated and the combination of Lai and Klingler discloses the data corresponding to the first user activity or the data corresponding to the second user activity comprises at least one of hands data, controller data, gaze data, and head movement data (Lai, ¶0061, “The client system 200 may detect user interface gestures and other gestures using an inside-out or outside-in tracking system of image capture devices and or external cameras.” ¶0063, “Examples of such external devices include handheld controllers, mobile devices, desktop computers, devices worn by a user, devices worn by one or more other users, and/or any other suitable external system.” ¶0113, “the client system performs object recognition within image data captured by the image capture devices of HMD to identify objects in the physical environment such as the user, the user's hand, and/or physical objects. Further, the client system tracks the position, orientation, and configuration of the objects in the physical environment over a sliding window of time.” “the virtual content module 580 may trigger generation and rendering of virtual content 543 by the client system (including virtual assistant application 505 and I/O interfaces 545) based on a current field of view of user, as may be determined by real-time gaze tracking of the user, or other conditions.” ¶0165, “The user can use a gesture or combination of gestures to navigate through the user interface 1430 and view and/or select suggestions 1435.”. Klingler, ¶0150, “the user may make a gesture or pose detected using computer vision”.) As to claim 18, claim 1 is incorporated and the combination of Lai and Klingler discloses the electronic device comprises a head- mounted device (HMD) (Lai, ¶0092, “HMD 465 generally represents any type or form of virtual reality system, such as virtual reality system 350 in FIG. 3B.” ¶0101.). As to claim 19, the combination of Lai and Klingler discloses a device comprising: a non-transitory computer-readable storage medium; and one or more processors coupled to the non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium comprises program instructions that, when executed on the one or more processors, cause the one or more processors to perform operations comprising: presenting a view of a three-dimensional (3D) environment, wherein a virtual assistant is positioned at a 3D position based on a 3D coordinate system associated with the 3D environment; receiving data corresponding to a first user activity in the 3D coordinate system for a first period of time; identifying a user interaction event associated with the virtual assistant in the 3D environment based on the data corresponding to the user activity; providing a graphical indication corresponding to one or more attributes associated with the virtual assistant based on identifying the user interaction event; and in accordance with receiving data corresponding to a second user activity for a second period of time, generating one or more user interface elements that are positioned at 3D positions based on the 3D coordinate system associated with the 3D environment (See claim 1 for detailed analysis.). As to claim 20, the combination of Lai and Klingler discloses a non-transitory computer-readable storage medium, storing program instructions executable on a device to perform operations comprising: presenting a view of a three-dimensional (3D) environment, wherein a virtual assistant is positioned at a 3D position based on a 3D coordinate system associated with the 3D environment; receiving data corresponding to a first user activity in the 3D coordinate system for a first period of time; identifying a user interaction event associated with the virtual assistant in the 3D environment based on the data corresponding to the user activity; providing a graphical indication corresponding to one or more attributes associated with the virtual assistant based on identifying the user interaction event; and in accordance with receiving data corresponding to a second user activity for a second period of time, generating one or more user interface elements that are positioned at 3D positions based on the 3D coordinate system associated with the 3D environment (See claim 1 for detailed analysis.). Claims 8-10 are rejected under 35 U.S.C. 103 as being unpatentable over Lai et al. (US Pub 2023/0316594 A1) in view of Klingler et al. (US Pub 2025/0181207 A1) and Treat et al. (US Pub 2025/0324188 A1). As to claim 8, claim 6 is incorporated and the combination of Lai and Klingler does not disclose providing a stream of spatialized audio at a 3D position within the 3D coordinate system associated with the 3D environment, wherein the 3D position of the stream of spatialized audio corresponds to the 3D position of the virtual assistant. Treat discloses providing a stream of spatialized audio at a 3D position within the 3D coordinate system associated with the 3D environment, wherein the 3D position of the stream of spatialized audio corresponds to the 3D position of the virtual assistant (Treat, ¶0403, “Certain user devices may implement spatial audio, which may also be referred to as spatial sound, three-dimensional (3D) audio or 3D sound. For such user devices, the auditory operating system shell may implement a spatial audio framework for agents. The auditory operating system shell may grant each agent a distinct virtual position relative to the user. Through advanced audio rendering, the auditory operating system shell may precisely simulate direction and distance cues for each agent's voice. The placement of each agent may be configurable and may be continuously updated based on factors such as interaction context, user movements, head orientation, environmental factors, and agent identity and role.” ¶0404-0406. ¶0409, “an agent may also be referred to as an artificial intelligence agent, a synthetic intelligence agent, a digital assistant, a voice assistant, an artificial agent, or as an assistant.”) Lai, Klingler and Treat are considered to be analogous art because all pertain to interactive systems. It would have been obvious before the effective filing date of the claimed invention to have modified Lai with the features of “providing a stream of spatialized audio at a 3D position within the 3D coordinate system associated with the 3D environment, wherein the 3D position of the stream of spatialized audio corresponds to the 3D position of the virtual assistant” as taught by Treat. The suggestion/motivation would have been in order to precisely simulate direction and distance cues for each agent's voice (Treat, ¶0403.). As to claim 9, claim 8 is incorporated and the combination of Lai, Klingler and Treat discloses utterances associated with the stream of spatialized audio are correlated to utterances associated with the candidate text representation (Treat, ¶0450, “The ML or AI system 4108 may provide text 4110 to the auditory operating system 3816. The auditory operating system 3816 may utilize the text to speech system 4112 to convert the text into speech. The auditory operating system 3816 may then provide the speech audio stream to various entities, such as the user via the user device 3804 or the interface device 3802 or the auditory operating system shell 3818.”). As to claim 10, claim 8 is incorporated and the combination of Lai, Klingler and Treat discloses utterances associated with the stream of spatialized audio are different than the candidate text representation (Lai, ¶0046, “The virtual assistant application 110 may then present the traffic information to the user as text (e.g., as virtual content overlaid on the physical environment such as real-world object) or audio (e.g., spoken to the user in natural language through a speaker associated with the client system 105).” Klingler, ¶0010, “an interactive avatar (e.g., an animated digital character) or other bot may support any number of simultaneous interaction modalities and corresponding interaction channels to engage with the user, such as channels for character or bot actions (e.g., speech, gestures, postures, movement, vocal bursts, etc.), scene actions (e.g., two-dimensional (2D) GUI overlays, 3D scene interactions, visual effects, music, etc.), and user actions (e.g., speech, gesture, posture, movement, etc.). Actions based on different modalities may occur sequentially or in parallel (e.g., waving and saying hello). As such, the interactive agent may execute any number of flows that specify a sequence of multimodal actions (e.g., different types of bot or user actions) using any number of supported interaction modalities and corresponding interaction channels.” ¶0056,” when designing an interactive avatar experience, a designer may want to support many different output interaction modalities, or ways of interacting with a user. A designer may want their avatar to talk, make gestures, show something in a GUI, make sounds, or interact in other ways. Likewise, a designer may want to support different types of input interaction modalities, or ways for a user to interact with the system. For example, a designer may want to support detecting and responding when a user provides an answer to a question verbally, by selecting an item on a screen, or making a gesture like a thumbs up to confirm a choice. One possible implication of multimodality is that a designer may want flexibility in how interactions are temporarily aligned. For example, a designer may want an avatar to say something while performing a gesture, or may want to initiate a gesture at a specific moment when the avatar says something in particular. As such, it may be desirable to support different types of independently controllable interaction modalities.” Due to Klingler’s multimodalities, the output of audio and the output of text can be different due to different output models. Treat, ¶0451, “The audio mixed reality devices and systems described herein may render immersive spatial audio scenes that can be layered over, selectively mixed with, or used to modify real-world acoustic inputs.”). Claims 11 is rejected under 35 U.S.C. 103 as being unpatentable over Lai et al. (US Pub 2023/0316594 A1) in view of Klingler et al. (US Pub 2025/0181207 A1) and Fu et al. (US Pub 2025/0005317 A1) As to claim 11, claim 1 is incorporated and the combination of Lai and Klingler does not disclose the graphical indication is a virtual effect corresponding to an eye or pair of eyes associated with the virtual assistant. Fu teaches disclose the graphical indication is a virtual effect corresponding to an eye or pair of eyes associated with the virtual assistant (Fu, ¶0024, “the visual companion may be configured to have eye contact with the one or more users. In some embodiments, the visual companion may be configured to differentiate and identify the one or more users.” ¶0025, “the visual companion may be configured to switch eye contact when interacting with multiple users. In some embodiments, the visual companion may be configured to switch to different personalities based on the analysis of one user who the visual companion has eye contact with and has communicated with.”). Lai, Klingler and Fu are considered to be analogous art because all pertain to interactive systems. It would have been obvious before the effective filing date of the claimed invention to have modified Lai with the features of “the graphical indication is a virtual effect corresponding to an eye or pair of eyes associated with the virtual assistant” as taught by Fu. The suggestion/motivation would have been in order to switch eye contact when interacting with multiple users (Fu, ¶0025.) Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Schliemann et al. (US Pub 2022/0137724 A1) teaches hand tracking of a user in a mixed reality environment. Jonker et al. (US Pub 2024/0071378 A1) teaches defining and modifying behavior in an extended reality environment based on natural language processing and/or user demonstrations. Prasad et al. (US Pub 2025/0004544 A1) teaches gaze initiated actions. Kaplan (US Pub 2023/0108256 A1) teaches adds context into the user speech to aid interpreting the speech. Any inquiry concerning this communication or earlier communications from the examiner should be directed to YU CHEN whose telephone number is (571)270-7951. The examiner can normally be reached on M-F 8-5 PST Mid-day flex. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Xiao Wu can be reached on 571-272-7761. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /YU CHEN/Primary Examiner, Art Unit 2613
Read full office action

Prosecution Timeline

Apr 15, 2025
Application Filed
Sep 02, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743818
DISCOVERING AND MITIGATING BIASES IN LARGE PRE-TRAINED MULTIMODAL BASED IMAGE EDITING
2y 8m to grant Granted Sep 22, 2026
Patent 12745524
DISPLAY DEVICE AND MOBILE ELECTRONIC DEVICE INCLUDING THE SAME
2y 7m to grant Granted Sep 22, 2026
Patent 12737930
FEATURE LAYERS FOR RENDERING OF DESIGN OPTIONS
3y 2m to grant Granted Sep 15, 2026
Patent 12737898
RESPIRATION FEATURE EXTRACTION METHOD BASED ON BODY SURFACE SIGNIFICANCE ANALYSIS
3y 0m to grant Granted Sep 15, 2026
Patent 12732646
RENDERING A MODELED SCENE
2y 4m to grant Granted Sep 08, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
68%
Grant Probability
98%
With Interview (+29.6%)
2y 10m (~1y 4m remaining)
Median Time to Grant
Low
PTA Risk
Based on 1087 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month