DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 07/16/2026 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 2, 5, 8-10, 13, 16-18 are rejected under 35 U.S.C. 103 as being unpatentable over DELEEUW (US 2022/0083136 A1) in view of MULLINS et al. (US 2016/0343168 A1).
RE claim 1, Deleeuw is made of record as teaching technologies for natural language interactions with a virtual assistant [abstract]. The virtual personal assistant (208) responds to spoken user commands and displays an avatar on the display (128) to provide information on the status of the virtual personal assistant (208) [0027].
Deleeuw teaches a method performed by one or more computing systems (see Fig. 1, Fig. 2), the method comprising:
(a)
rendering, on an extended-reality (XR) display device comprising a headset configured to be worn by a user, a first output image of an XR assistant avatar having a first form,
Fig. 6A (blocks 602-606), computing device (100) of Deleeuw may execute method (600) for natural interaction with the virtual personal assistant (208) [0052]. The avatar of the virtual personal assistant (208) is displayed in a disengaged state (said a first output image of an XR assistant avatar having a first form) on the display (128) [0052]. When in the disengaged state, the avatar may be semitransparent, allowing background applications on the computing device (100) to shine through the avatar [0053].The disengaged state may also display the avatar in a smaller size in a corner of the display (128) [0053]. Computing device (100) includes a display (128) [0019]. Deleeuw teaches the display (128) may be embodied as any type of display capable of displaying digital information [0020].
Mullins is made of record as teaching a head mounted device that displays a virtual fictional character based on a combination of the user-based context, ambient-based context, and the application context [abstract]. Mullins teaches different virtual characters may be displayed based on the task, conditions of the user, and conditions ambient to the HMD. The conditions of the user may identify how the user feels physically and mentally while performing the task by looking at user-based sensor data (i.e., user-based context) [0022]. The virtual character may be an avatar [0024]. Fig. 1, the network environment (100) includes an HMD (101) (said headset) [0037, 0039].
It would have been obvious before the effective filing date of the claimed invention to utilize the HMD of Mullins as the display for Deleeuw. Deleeuw discloses the display (128) may be embodied as any type of display capable of displaying digital information [0020]. Furthermore, it would be beneficial to provide a hands free display, such as the headset of Mullins, in order for the user to accomplish physical tasks with both hands available (e.g., installing an appliance [Mullins: 0023]).
(b)
wherein the XR assistant avatar is configured to provide access by the user to an assistant system through a natural-language dialog;
Fig. 6A (block 608), the computing device (100) monitors for activation of the virtual personal assistant (208) by the user [0054]. The computing device (100) interprets the audio input to determine whether the user has uttered a code word for activating the virtual personal assistant (208) [0054]. The computing device (100) may analyze the audio input to determine whether the user is addressing the virtual personal assistant (208). The device (100) may perform speech recognition on the audio input [0057]. The speech recognition module (202) is configured to perform speech recognition on audio input data received from the audio input module (204) [0025]. The speech recognition module (202) may use a speech recognition grammar provided by an application such as the virtual personal assistant (208) to rank and filter speech recognition results. The speech recognition module (202) may recognize speech in a dictation or free speech mode. The dictation or free speech mode may use a full natural language vocabulary and grammar to recognize results, and thus may produce additional likely speech recognition results [0025].
(c)
detecting an action of the user associated with the XR display device;
Fig. 6A (block 608), the computing device (100) monitors for activation of the virtual personal assistant (208) by the user [0054]. The computing device (100) interprets the audio input to determine whether the user has uttered a code word for activating the virtual personal assistant (208) (said action of the user) [0054].
(d)
determining, based at least in part on the detected action, a change in sentiment of the user;
The engagement module (214) determines the user’s level of engagement (said sentiment) with the virtual personal assistant (208) based on eye tracking sensor data (132) (said detected action) and/or audio sensor data (130) (said detected action) [0028]. The engagement module (214) provides the engagement level to the virtual personal assistant (208), allowing the virtual personal assistant (208) to modify the avatar accordingly [0028]. Fig. 6A (block 614) device (100) determines that the user’s gaze has focused on the avatar for a length of time longer than a certain threshold, or when the code word has been detected [0055]. Thus, the engagement module (214) can detect the user was initially not looking at the avatar and then changes their gaze to look at the avatar (said change in sentiment).
(e)
determining, based at least in part on the change in sentiment of the user, a second form of the XR assistant avatar, wherein the second form is different from the first form; and
Fig. 6A (block 614) device (100) determines that the user’s gaze has focused on the avatar for a length of time longer than a certain threshold, or when the code word has been detected [0055]. Fig. 6A (block 616) displays the avatar in a ready state, such as rendering the avatar as making eye contact with the user and decreasing the transparency so that it appears more solid (said second form, different from first form) [0056].
(f)
rendering, on the XR display device, a second output image of the XR assistant avatar having the second form.
Fig. 6A (block 616) displays the avatar in a ready state, such as rendering the avatar as making eye contact with the user and decreasing the transparency so that it appears more solid (said rendering a second output image having a second form) [0056].
RE claim 2, Deleeuw teaches wherein
(a)
the detected action is associated with a behavior of the user, and
The engagement module (214) of Deleeuw determines the user’s level of engagement (said sentiment) with the virtual personal assistant (208) based on eye tracking sensor data (132) (said detected action) and/or audio sensor data (130) (said detected action) [0028]. The engagement module (214) provides the engagement level to the virtual personal assistant (208), allowing the virtual personal assistant (208) to modify the avatar accordingly [0028]. Fig. 6A (block 614) device (100) determines that the user’s gaze (said behavior of the user) has focused on the avatar for a length of time longer than a certain threshold, or when the code word has been detected [0055].
(b)
the change in sentiment is determined based at least in part on the behavior.
As discussed in claim 2(a), the user’s level of engagement (said sentiment) with the virtual personal assistant is determined [0028]. Thus, the engagement module (214) can detect the user was initially not looking at the avatar and then changes their gaze to look at the avatar (said change in sentiment) as detected by the focus length compared to a threshold [0055].
RE claim 5, Deleeuw teaches further comprising:
(a)
changing, based at least in part on the change in sentiment of the user, a set of actions executable by the assistant system.
As discussed in the rationale of claim 1(d), the engagement module (214) of Deleeuw determines the user’s level of engagement (said sentiment) with the virtual personal assistant (208) [0028]. The avatar can change to the ready state when it is determined the user is now gazing at the avatar (said change in sentiment). Fig. 6A (block 630), Deleeuw teaches the user is engaged with the avatar based on engagement level [0058]. Fig. 6B (block 632), device (632) displays the avatar in an engaged state [0059]. The engaged state indicates to the user that the virtual personal assistant (208) is actively interpreting commands issued by the user. Fig. 6B (642) device (100) determine whether or not a command has been received, when yes, avatar is displayed in a working state which indicates the virtual personal assistant (208) is currently executing a task (said changing set of actions executable by the assistant system) [0065]. The working state includes a representation of the task being performed, e.g., an application icon or a representation of the avatar performing a task [0065-0066].
RE claim 8, Deleeuw teaches wherein the second form is different from the first form in one or more of
(a)
emotion, appearance, size, shape, movement, gesture, color, shading, outline, brightness or luminescence.
As discussed in the rationale of claim 1(e), when the avatar changes to the ready state (said second form), the avatar can be displayed with a decreased transparency so that it appears more solid (said appearance) [0056]. When the system further determines an engaged state [0058], the device (100) may decrease the transparence of the avatar or be displayed as full opaque (said appearance) [0059]. The device (100) may adjust the size (said size) and/or position of the avatar (Fig. 6A (638)). The avatar may be rendered close to or in front of the currently-active application on the display (128), or the avatar may be increased in size (said size) [0059].
RE claim 9, claim 9 recites similar limitations as claim 1 but in manufacture form. Therefore, the same rationale used for claim 1 is applied. Furthermore, Deleeuw teaches implementing the method as instructions carried by or stored on non-transitory computer-readable storage medium [0014].
RE claim 10, claim 10 recites similar limitations as claim 2 but in manufacture form. Therefore, the same rationale used for claim 2 is applied.
RE claim 13, claim 13 recites similar limitations as claim 5 but in manufacture form. Therefore, the same rationale used for claim 5 is applied.
RE claim 16, claim 16 recites similar limitations as claim 8 but in manufacture form. Therefore, the same rationale used for claim 8 is applied.
RE claim 17, claim 17 recites similar limitations as claim 1 but in system form. Therefore, the same rationale used for claim 1 is applied. Furthermore, Deleeuw in view of Mullins teaches an extended-reality (XR) display device comprising:
(I)
a headset configured to be worn by a user;
Deleeuw teaches the display (128) may be embodied as any type of display capable of displaying digital information [0020].
Mullins is made of record as teaching a head mounted device that displays a virtual fictional character based on a combination of the user-based context, ambient-based context, and the application context [abstract]. Fig. 1, the network environment (100) includes an HMD (101) (said headset) [0037, 0039].
It would have been obvious before the effective filing date of the claimed invention to utilize the HMD of Mullins as the display for Deleeuw. Deleeuw discloses the display (128) may be embodied as any type of display capable of displaying digital information [0020]. Furthermore, it would be beneficial to provide a hands free display, such as the headset of Mullins, in order for the user to accomplish physical tasks with both hands available (e.g., installing an appliance [Mullins: 0023]).
(ii)
a display;
Deleeuw teaches the display (128) may be embodied as any type of display capable of displaying digital information [0020]. Mullins teaches an HMD (101) (said headset) [0037, 0039, Fig. 1].
(iii)
one or more processors; and
Deleeuw teaches the device (100) includes processor (120) [0018].
(iv)
a memory coupled to the one or more processors, the memory comprising instructions which, when executed by the one or more processors, cause the XR display device to execute the steps of claim 1.
Fig. 1 of Deleeuw, memory (124) [0018].
RE claim 18, claim 18 recites similar limitations as claim 2 but in system form. Therefore, the same rationale used for claim 2 is applied.
Claims 3-4, 11-12, 19 are rejected under 35 U.S.C. 103 as being unpatentable over DELEEUW (US 2022/0083136 A1) in view of MULLINS et al. (US 2016/0343168 A1) as applied to claim 1, and in further view of KIM et al. (US 12,205,5077 B1).
RE claim 3, Deleeuw in view of Mullins teaches the limitations of claim 3 with the exception of discussing a machine-learning model. Kim is made of record as teaching a virtual conversation companion [abstract]. Kim teaches wherein
(a)
the second form is further determined based at least in part on a machine-learning model associated with the XR assistant avatar.
Kim teaches intent classifier component (520) that may implement machine learning (ML) model(s) [16:41-44]. Various ML techniques may be used to train and operate ML models [16:44-17:4]. The response management component (540) is configured to generate output data to the user (5) via device (110) [17:65-66]. The response management component (540) may generate output data including video data of a dynamic avatar (said second form), image data, and audio data corresponding to synthesized speech. Inputs to the response management component (540) may include the spoken natural language input data (505), and ASR output data representing a spoken natural language input [18:1-7].
It would have been obvious before the effective filing date of the claimed invention to use the intent classifier component of Kim with the teachings of Deleeuw in view of Mullins because it allows the digital assistant to understand natural human speech, learn from user behavior and improve accuracy over time.
RE claim 4, in further view of Kim, Kim teaches wherein
(a)
The machine-learning model is trained on a plurality of prior reactions associated with a plurality of morphings of the XR assistant avatar.
Various ML techniques may be used to train and operate ML models [16:44-17:4]. In order to apply the machine learning techniques, the ML processes themselves need to be trained. Training a machine learning component such as, one of the first or second models, requires establishing a “ground truth” for the training examples. In machine learning, the term “ground truth” refers to the accuracy of a training set’s classification for supervised learning techniques. Various techniques may be used to train the models including backpropagation, statistical learning, supervised learning, semi-supervised learning, stochastic learning, or other know techniques [17:5-15]. As taught in the rationale of claim 3, the response management component (540) may generate output data including video data of a dynamic avatar (said morphings), image data, and audio data corresponding to synthesized speech. Inputs to the response management component (540) may include the spoken natural language input data (505), and ASR output data representing a spoken natural language input [18:1-7].
The same motivation to combine as discussed in the rationale of claim 3 is incorporated herein.
RE claim 11, claim 11 recites similar limitations as claim 3 but in manufacture form. Therefore, the same rationale used for claim 3 is applied.
RE claim 12, claim 12 recites similar limitations as claim 4 but in manufacture form. Therefore, the same rationale used for claim 4 is applied.
RE claim 19, claim 19 recites similar limitations as claims 3 and 4 but in system form. Therefore, the same rationale used for claims 3 and 4 is applied.
Allowable Subject Matter
Claims 6-7, 14-15, 20 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
The following is an examiner’s statement of reasons for allowance:
Reference CASE, Jr. et al. (US 7,254,516 B2) is made of record as teaching an athletic performing monitoring system that provides real time information [abstract]. GPS data may be used for controlling an audio, video, or other display device during the athletic performance based on data obtained via the GPS [6:17-28]. Different audio can be output pending on the GPS information related to time, distance, type of terrain, elevation or altitude etc. [19:16-35]. Case provides the example of playing slow or relaxing songs under certain conditions, such as when the athlete’s hear rate or pulse rate exceeds 150 per minute [19:55-65].
However, the cited prior art does not disclose or render obvious the combination of elements recited in the claims as whole. Specifically, the cited prior art fails to disclose or render obvious the limitations: detecting the speed of movement of the user and altering the sentiment of the assistant avatar in relation to the speed.
Any comments considered necessary by applicant must be submitted no later than the payment of the issue fee and, to avoid processing delays, should preferably accompany the issue fee. Such submissions should be clearly labeled “Comments on Statement of Reasons for Allowance.”
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHELLE L SAMS:
direct telephone number:
(571) 272-7661
email:
michelle.sams@uspto.gov
personal fax number:
(571)273-7661
The examiner is currently part time and can be reached Mon.-Fri. 5:30am-9:30am.
Examiner interviews are available via telephone and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kee M. Tung can be reached on (571)272-7794. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MICHELLE L SAMS/
Primary Examiner, Art Unit 2611
11 August 2026