DETAILED ACTION
This action is in response to communications: Amendment filed May 21, 2026.
Claims 1-15 are pending in this case. Claims 1, 3-6, 12, and 14 have been newly amended. No claims have been newly added or cancelled. This action is made FINAL.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-8, and 12-15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Steptoe et al. (US 11,468,616) in view of Tomlin et al. (US 10,529,359).
As to claim 1, Steptoe et al. further disclose a non-transitory computer-readable data storage medium (Figure 1, system 100, further illustrated as system 200 in Figure 2, with memory 120) storing program code (e.g. storing modules 102 as well as data and/or computer-readable instructions, column 3, lines 48-59) executable by a processor (e.g. physical processor 130, column 3, lines 60 thru column 4, line 10) to perform processing (process of Figure 3) comprising: detecting speech using a microphone (step 330, detecting that a user produced a sound, where column 16, lines 31-56 notes user device 202 (illustrated as computing device 202) may include a microphone or other audio sensor, where detecting module 108 may receive sound 220, which may include one or more phonemes articulated by the user, via the microphone or other audio sensor, detecting module 108 may then identify the one or more phonemes included in sound 220 using any suitable speech recognition technique) of a head-mountable display (HMD)(column 6, lines 12-18 notes user device 202 may be a wearable device, with lines 19-28 further noting the user device 202 including one or more sensors, e.g. a microphone and other audio sensors, where column 5, lines 10-19 notes system 100 may encompass all of system 200, e.g. each of user device 202, server 206, and target device 208, implementing one or more module 102, where column 28, lines 64 thru column 29, lines 22 notes system may be implemented in conjunction with an artificial reality system, which may include a virtual-reality (VR), augmented reality (AR), mixed reality (MR), hybrid reality, or some combination and/or derivative thereof, the artificial reality system provides artificial reality content that may be implemented on various platforms, including a head-mounted display (HMD) connected to a host computer system, a standalone HMD, a mobile device or computing system, or any other hardware platform capable of providing artificial reality content to one or more viewers, where column 17, lines 14-21 further notes the computer-generated avatar may include an augmented reality avatar presented within an augmented reality environment, a virtual reality avatar presented within a virtual reality environment, and/or a video conferencing avatar presented within a video conferencing application, and thus may be considered the user device 202 may be considered a head-mounted display (HMD), the user of user device 202 being a wearer of the HMD), the speech including a phoneme (e.g. column 16, lines 49-51 notes sound 220 received may include one or more phonemes articulated by the user); determining whether a wearer of the HMD uttered the speech (step 330, detecting that a user produced a sound, where column 16, lines 31-56 further notes based on identifying the one or more phonemes, detecting module 108 may detect that the user has produced sound 220, where column 16, lines 57 thru column 17, lines 2 gives illustration of multiple users of different devices, where the detecting module may pick up the sound 220 specifically from user device 202 and detect that the sound 220 came from user of user device 202 (NOTE: determining whether a wearer uttered the speech may additionally include steps 310 and 320 of Figure 3, see claims below)); and in response to determining that the wearer uttered the speech (e.g. in response to each of steps 310, 320, and 330, detecting the user has produced sound 220), rendering an avatar representing the wearer to have a viseme corresponding to the phoneme (step 340, directing a computer-generated avatar that represents the user to produce the viseme in accordance with the set of action unit parameters associated with each unit in response to detecting that the user has produced the sound, where column 17, lines 3 thru column 18, lines 22 notes computer-generated avatar 238 may be configured to produce one or more visemes via one or more action units (AUs) and/or AU parameters, where Figure 4 illustrates a table 400 that shows various examples to visemes, associated phonemes, examples of words that, when pronounced by a human, may cause the human’s face to produce the visemes, and visual depictions of a human face producing associated visemes, where Figure 5 further illustrates a table 500 that shows various examples of AUs, AU names, and visual depictions of a human face when the AU is at a maximum intensity, where column 8, lines 42-52 notes an “action unit” may include any information that describes actions of individual muscles and/or group of muscles, e.g. an AU may be associated with one or more facial muscles that may be engaged by a user to produce a viseme).
As noted above, Steptoe et al. describe its system 100/200 for performing the operations as described, where user device 202, e.g. computing device 202, as part of system 200, which may include wearable devices, further comprising various sensors, including imaging devices and a microphone. Steptoe et al. further disclose its system may be implemented in conjunction with an artificial reality system, which may include a virtual-reality (VR), augmented reality (AR), mixed reality (MR), hybrid reality, or some combination and/or derivative thereof, the artificial reality system provides artificial reality content that may be implemented on various platforms, including a head-mounted display (HMD) connected to a host computer system, a standalone HMD, a mobile device or computing system, or any other hardware platform capable of providing artificial reality content, e.g. the computer-generated avatar, to one or more viewers. Therefore, it would have been obvious to one of ordinary skill in the art at the time of the invention to implement the system 100/200 of Steptoe et al., e.g. user device 202, as a wearable device including an head-mounted display (HMD), such that the computer-generated avatar may be rendered and displayed as described, thus yielding predictable results, without changing the scope of the invention. Additionally, as noted above, Steptoe et al. describe multiple users of different devices, e.g. user of user device 202 (e.g. “source user”) and a user of target device 208 (e.g. “target user”), where the detecting module 108 may pick up the sound 220 specifically from the user of user device 202 and detect that the sound 220 came from that user of user device 202, thus disclose “determining whether a wearer of the HMD uttered the speech.” However, Steptoe et al. do not explicitly disclose “…determining whether a wearer of the HMD or another user uttered the speech…”
Tomlin et al. disclose determining whether a wearer of the HMD (e.g. wearer 102 of head-mounted display (HMD) device 104) or another user (e.g. other person 110) uttered the speech (e.g. is speaking)(Figure 11, column 9, lines 36 thru column 11, lines 17 notes HMD device 1100 (e.g. similar to HMD device 104) may include various sensors, e.g. a microphone array comprising six microphones, microphones 1112 and 1114 may be omnidirectional to capture sound in the general area/direction in front of the HMD (e.g. capture sound from other person), microphones 1116 and 1118 may be forward facing and aimed downward to capture sound emitted from the wearer’s mouth, and microphones 1120 and 1122 may be omnidirectional to capture sound in the general area/direction on each side of the HMD (e.g. capture sound from other person), the captured sound, e.g. as an audio data stream, may be further analyzed by controller 1104 to determine a source location of human speech segments, e.g. the wearer of the HMD or other person; one or more image sensors, e.g. outward facing image sensors 1108, for capturing visual data from the physical environment in which the HMD device 1100 is located, e.g. to detect movements performed by the wearer or by a person or physical object within a field of view of the display 1102, e.g. detect a user speaking to the wearer of the HMD device; and an inertial measurement unit (IMU) for providing position and/or orientation data of the HMD device 1100 to the controller, e.g. to determine a direction of a user that has engaged the wearer of the HMD device in conversation, where Figure 9 and associated text, e.g. column 7, lines 19-55 further notes image sensors, e.g. one or more outward facing image sensors noted above, for receiving image data, e.g. images of one or both speakers potentially engaged in a conversation (e.g. images of a user from the perspective of the wearer of the HMD device or images of both users from the perspective of a sensor device), the image data may be fed into a feature recognition stage to analyze images to determine whether a user’s mouth is moving and/or fed into a user identification stage to recognize a user that is speaking, where a user identification stage may output the identity of a speaker to the conversation detector stage, which may use the speaker identity to classify segments of human speech as being spoken by particular identified users).
It would have been obvious to one of ordinary skill in the at the time of the invention to modify Steptoe et al.’s system and method of detecting speech, determining whether a wearer of a head-mounted display (HMD) device uttered the speech, and rendering an avatar with Tomlin et al.’s method of detecting conversation, including determining whether a wearer of the HMD and/or another user uttered speech to further modify visual content displayed based on the conversation and/or the person speaking (see at least column 1, lines 51 thru column 2, lines 38 of Tomlin et al.).
As to claim 2, Steptoe et al. modified with Tomlin et al. disclose the processing comprises: displaying the rendered avatar representing the wearer of the HMD (Steptoe, e.g. displaying computer-generated avatar representing the user that produced the sound, e.g. wearer of user device 202 as the HMD)(Steptoe, column 2, lines 56 thru column 3, lines 3 notes the computer-generated avatar 238 may represent a user within an artificial environment, such as a VR and/or AR environment, to accurately and realistically reproduce, in real-time, facial expressions and/or facial motions associated with a series of phonemes produced by the user (e.g., words spoken by the user) and/or body actions executed by the user).
As to claim 3, Steptoe et al. modified with Tomlin et al. disclose the processing comprises: in response to determining that the other user uttered the speech, not rendering the avatar to have the viseme corresponding to the phoneme (Steptoe, modified with Tomlin, e.g. as noted in claim 1, step 340 is performed “in response to” each of steps 310, 320, and 330, detecting that the user produced the sound, thus if the output of any of steps 310, 320, and 330 does not render the user uttered the speech, including previous step 330, it is not detected that the user produced the sound (as modified with Tomlin, e.g. as noted in claim 1, as it is identified that the other person 110 is speaking and not the wearer of the HMD), it would be obvious that step 340 is not performed, thus the computer-generated avatar is not rendered).
As to claim 4, Steptoe et al. modified with Tomlin et al. disclose determining whether the wearer of the HMD uttered the speech (Steptoe, Figure 3) comprises: detecting, using a camera of the HMD (Steptoe, e.g. via one or more sensors and/or imaging devices, where column 6, lines 19-28 notes user device 202 (e.g. HMD) may include one or more sensors that may gather data associated with the environment that includes the user device 202, e.g. an imaging device configured to capture one or more portions of an electromagnetic spectrum (e.g. a visible-light camera, an infrared camera, an ultraviolet camera, etc.); modified with Tomlin, e.g. via one or more image sensors, e.g. one or more outward facing image sensors 1108 of HMD 1100), whether mouth movement of the wearer occurred while the speech was detected (Steptoe, column 11, lines 52-67 further notes identifying module 104 may include one or more sensors, e.g. imaging device and/or an audio recording device, for capturing a set of images 232, e.g. a series of images, a video file, a set of frames of a video file, and recording audio 236, e.g. a recording of the user producing phonemes 234 and/or responding to prompts included in phonemes 234, as the user speaks sets of phonemes 234; modified with Tomlin, e.g. whether mouth movement of the wearer occurred while the speech was detected)(Steptoe, step 310, column 7, lines 33-57 notes identifying a set of AUs associated with a face of a user, each AU associated with at least one muscle group engaged by the user to produce a viseme associated with a sound produced by the user, where a “viseme” may include a visual representation of a configuration (shape) of a face of a person as the person produces an associated set of phonemes, each viseme associated with a mouth shape for a specific set of phonemes (see Figures 4 and 5), where column 12, lines 1-23 further notes identifying module 104 may identify viseme 218 from the set of images 232 via image recognition and may employ any suitable face tracking and/or identification system to identify a set of AUs 210 associate with face 212, e.g. from the set of captured images, and associate a feature of the face of the user, e.g. a portion of the user’s face that may generally correspond to a set of muscle groups associated with the set of AUs, with the set of AUs based on the identification of the viseme, the set of images, and recorded audio, and step 320, column 12, lines 31 thru column 13, lines 14 further notes, for each AU in the set of AUs, determining a set of AU parameters associated with the AU and the viseme, the set of AU parameters may include (1) an onset curve associated with the viseme, and (2) a falloff curve associated with the viseme, which may specify a maximum velocity and/or activation/deactivation velocity of a given muscle group, thus as a person strains his or her lips, such as with rapid queuing or chaining of mouth shapes, e.g. visemes, the person may distort a shape of his or her mouth in varying ways to make way for a next mouth shape, then at step 330, column 16, lines 31 thru column 17, lines 2 notes detecting that the user has produced the sound; modified with Tomlin, column 7, lines 19-35 notes image data stream 924 received from image sensor 906, which may include images of one or both speakers potentially engaged in conversation, the image data stream 924 may be fed to a feature recognition stage 926, which may analyze the images to determine whether a user’s mouth is moving and/or fed into a user identification stage to recognize a user that is speaking, where a user identification stage may output the identity of a speaker to the conversation detector stage, which may use the speaker identity to classify segments of human speech as being spoken by particular identified users); in response to detecting that the mouth movement of the wearer occurred while the speech was detected, determining that the wearer uttered the speech (Steptoe, e.g. determining the user of user device 202, e.g. wearer of HMD, uttered the speech based on the determinations at each of steps 310, 320, and 330 as noted above, which includes determining mouth (and/or lip) movements (steps 310, 320) while detecting sound (step 330); modified with Tomlin, as noted above, based on the analysis of at least the image data stream 924, e.g. determining a user’s mouth moving, and further determining the identity of the person speaking may be determined, e.g. the wearer 102 of the HMD 104); and in response to not detecting that the mouth movement of the wearer occurred while the speech was detected, determining that the other user uttered the speech (Steptoe, e.g. determining the user of user device 202, e.g. wearer of HMD, did not utter the speech based on the determinations at each of steps 310, 320, and 330 as noted, which includes determining mouth (and/or lip) movements (steps 310, 320) and/or detecting sound (step 330); modified with Tomlin, as noted above, based on the analysis of at least the image data stream 924, e.g. determining a user’s mouth moving, and further determining the identity of the person speaking may be determined, e.g. the other person 110).
As to claim 5, Steptoe et al. modified with Tomlin et al. disclose determining whether the wearer of the HMD uttered the speech (Steptoe, Figure 3) comprises: detecting, using a sensor of the HMD other than the microphone (Steptoe, e.g. via one or more sensors and/or imaging devices, where column 6, lines 19-28 notes user device 202 (e.g. HMD) may include one or more sensors that may gather data associated with the environment that includes the user device 202, e.g. an imaging device configured to capture one or more portions of an electromagnetic spectrum (e.g. a visible-light camera, an infrared camera, an ultraviolet camera, etc.); modified with Tomlin, one or more image sensors, e.g. via one or more outward facing image sensors 1108 of HMD 1100), whether mouth movement of the wearer occurred while the speech was detected (Steptoe, column 11, lines 52-67 further notes identifying module 104 may include one or more sensors, e.g. imaging device and/or an audio recording device, for capturing a set of images 232, e.g. a series of images, a video file, a set of frames of a video file, and recording audio 236, e.g. a recording of the user producing phonemes 234 and/or responding to prompts included in phonemes 234, as the user speaks sets of phonemes 234; modified with Tomlin, column 7, lines 19-35 notes image data stream 924 received from image sensor 906, which may include images of one or both speakers potentially engaged in conversation, the image data stream 924 may be fed to a feature recognition stage 926, which may analyze the images to determine whether a user’s mouth is moving and/or fed into a user identification stage to recognize a user that is speaking, where a user identification stage may output the identity of a speaker to the conversation detector stage, which may use the speaker identity to classify segments of human speech as being spoken by particular identified users); in response to detecting that the mouth movement of the wearer occurred while the speech was detected, determining that the wearer uttered the speech; and in response to not detecting that the mouth movement of the wearer occurred while the speech was detected, determining that the other user uttered the speech (Steptoe, modified with Tomlin, see details claim 4 above).
As to claim 6, Steptoe et al. modified with Tomlin et al. disclose the microphone comprises a microphone array (Steptoe, e.g. as noted in claim 1, a microphone or other audio sensor; modified with Tomlin, e.g. as noted in claim 1, microphone array of six microphones, e.g. microphones 1112, 1114, 1116, 1118, 1120, and 1122), and determining whether the wearer of the HMD uttered the speech comprises: detecting, using the microphone array, whether the speech was uttered from a direction of a mouth of the wearer (Steptoe, e.g. as noted in claim 1, step 330, detecting that a user produced a sound, where column 16, lines 31-56 notes user device 202 (illustrated as computing device 202) may include a microphone or other audio sensor, where detecting module 108 may receive sound 220, which may include one or more phonemes articulated by the user, via the microphone or other audio sensor; modified with Tomlin, column 9, lines 41 thru column 10, lines 20 notes microphone array comprising six microphones to capture sound, e.g. an audio data stream, where at least microphones 1116 and 1118 may be forward facing and aimed downward to capture sound emitted from the wearer’s mouth, and the audio data stream may be further analyzed to determine a source location of human speech segments, e.g. from the wearer of the HMD); in response to detecting that the speech was uttered from the direction of the mouth of the wearer, determining that the wearer uttered the speech (Steptoe, e.g. in response to each of steps 310, 320, and 330, detecting the user has produced sound 220, e.g. via microphone or other audio sensor; modified with Tomlin, as noted above, microphones 1116 and 1118 capture sound in the direction of the wearer’s mouth, thus may be determined the source location as the wearer of the HMD); and in response to detecting that the speech was not uttered from the direction of the mouth of the wearer, determining that the other user uttered the speech (Steptoe, e.g. in response to determining the user of user device 202, e.g. wearer of HMD, did not utter the speech based on the determinations at each of steps 310, 320, and 330 as noted, e.g. detecting sound (step 330); modified with Tomlin, e.g. as noted above, microphone array comprising other microphones, e.g. microphones 1112 and 1114 may be omnidirectional to capture sound in the general area/direction in front of the HMD (e.g. capture sound from other person), and microphones 1120 and 1122 may be omnidirectional to capture sound in the general area/direction on each side of the HMD (e.g. capture sound from other person), the captured sound, e.g. as an audio data stream, may be analyzed by controller 1104 to determine a source location of human speech segments, e.g. from the other person).
As to claim 7, Steptoe et al. modified with Tomlin et al. disclose the wearer is determined as having uttered the speech (Steptoe, e.g. based on the determination of steps 310, 320, and 330 of Figure 3), wherein the processing further comprises: capturing facial images of the wearer while the speech is detected, using a camera of the HMD (Steptoe, e.g. via one or more sensors and/or imaging devices, where column 6, lines 19-28 notes user device 202 (e.g. HMD) may include one or more sensors that may gather data associated with the environment that includes the user device 202, e.g. an imaging device configured to capture one or more portions of an electromagnetic spectrum (Steptoe, e.g. a visible-light camera, an infrared camera, an ultraviolet camera, etc.)), the facial images comprising the viseme corresponding to the phoneme (Steptoe, column 11, lines 52-67 further notes identifying module 104 may include one or more sensors, e.g. imaging device and/or an audio recording device, for capturing a set of images 232, e.g. a series of images, a video file, a set of frames of a video file, and recording audio 236, e.g. a recording of the user producing phonemes 234 and/or responding to prompts included in phonemes 234, as the user speaks sets of phonemes 234)(Steptoe, see details of claim 4 regarding steps 310 and 320 for identifying visemes corresponding to the phoneme, e.g. from the set of captured images 232), and wherein the avatar is rendered to have the viseme corresponding to the phoneme based on both the phoneme within the detected speech and the captured facial images including the viseme (Steptoe, e.g. as noted in claim 1, step 340, directing a computer-generated avatar that represents the user to produce the viseme in accordance with the set of action unit parameters associated with each unit in response to detecting that the user has produced the sound, where column 17, lines 3 thru column 18, lines 22 notes computer-generated avatar 238 may be configured to produce one or more visemes via one or more action units (AUs) and/or AU parameters).
As to claim 8, Steptoe et al. modified with Tomlin et al. disclose the processing further comprises: capturing sensor data while the speech is detected, using one or multiple sensors other than the camera and the microphone of the HMD (Steptoe, column 6, lines 19-28 notes user device 202 (e.g. HMD) may include one or more sensors that may gather data associated with the environment that includes the user device 202, e.g. an imaging device configured to capture one or more portions of an electromagnetic spectrum (e.g. a visible-light camera, an infrared camera, an ultraviolet camera, etc.), an inertial measurement unit (IMU), an accelerometer, a global positioning system device, a thermometer, a barometer, an altimeter, and other audio sensors, where one type of camera may be considered different from another as well as other audio sensors may be different from the microphone, thus “other than the camera and microphone”), and wherein the avatar is rendered to have the viseme corresponding to the phoneme further based on the captured sensor data (Steptoe, e.g. as noted in claim 4, steps 310 and 320, the set of captured images 232 may be from any one of the imaging devices and step 330 notes detecting that the user produced the sound with the microphone or other audio sensors, then at step 340, directing a computer-generated avatar that represents the user to produce the viseme in accordance with the set of action unit parameters associated with each unit in response to detecting that the user has produced the sound (e.g. using the other audio sensors), where column 17, lines 3 thru column 18, lines 22 notes computer-generated avatar 238 may be configured to produce one or more visemes via one or more action units (AUs) and/or AU parameters (identified via other capturing devices)).
As to claim 12, Steptoe et al. modified with Tomlin et al. disclose a method comprising: detecting, by a processor using a microphone of an electronic device, speech including a phoneme (see details of claim 1); determining, by the processor, whether a user of the electronic device uttered the speech or whether another user uttered the speech (see details of claim 1); in response to determining that the user of the electronic device uttered the speech, rendering, by the processor, an avatar representing the user of the electronic device to have a viseme corresponding to the phoneme (see details of claim 1); and displaying, by the processor, the avatar representing the user of the electronic device (see details of claim 2 above). Claim 12 is similar in scope to claims 1 and 2 combined, and is therefore rejected under similar rationale. Please see the rejection and rationale of claims 1 and 2 above.
As to claim 13, Steptoe et al. modified with Tomlin et al. disclose the user is determined as having uttered the speech, wherein the method further comprises: capturing facial images of the user while the speech is detected, by a processor using a camera, the facial images comprising the viseme corresponding to the phoneme, and wherein the avatar is rendered to have the viseme corresponding to the phoneme based on both the phoneme within the detected speech and the captured facial images including the viseme (see details of claim 7 above).
As to claim 14, Steptoe et al. modified with Tomlin et al. disclose a head-mountable display (HMD) (Steptoe, e.g. Figure 2, user device, where details of claim 1 notes user device 202 may be an head-mounted display (HMD); modified with Tomlin, head-mounted display (HMD) device 104, further illustrated as HMD 1100 of Figure 11) comprising: a microphone to detect speech including a phoneme (Steptoe, e.g. column 6, lines 19-28 notes user device 202 (e.g. HMD) may include one or more sensors that may gather data associated with the environment that includes the user device 202, including a microphone, where step 330, where column 16, lines 47-56 notes microphone for receiving sound which may include one or more phonemes articulated by the user; modified with Tomlin, column 9, lines 36 thru column 10, lines 20 notes HMD device 1100 may include various sensors including a microphone array comprising six microphones, microphones 1112 and 1114 may be omnidirectional to capture sound in the general area/direction in front of the HMD (e.g. capture sound from other person), microphones 1116 and 1118 may be forward facing and aimed downward to capture sound emitted from the wearer’s mouth, and microphones 1120 and 1122 may be omnidirectional to capture sound in the general area/direction on each side of the HMD (e.g. capture sound from other person), the captured sound, e.g. as an audio data stream, may be further analyzed by controller 1104 to determine a source location of human speech segments, e.g. the wearer of the HMD or other person); a camera to capture facial images of a wearer of the HMD while the speech is detected (Steptoe, e.g. one or more sensors as an imaging devices, where column 6, lines 19-28 notes user device 202 (e.g. HMD) may include one or more sensors that may gather data associated with the environment that includes the user device 202, e.g. an imaging device configured to capture one or more portions of an electromagnetic spectrum (e.g. a visible-light camera, an infrared camera, an ultraviolet camera, etc.), where column 11, lines 52-67 further notes identifying module 104 may include one or more sensors, e.g. imaging device and/or an audio recording device, for capturing a set of images 232, e.g. a series of images, a video file, a set of frames of a video file, and recording audio 236, e.g. a recording of the user producing phonemes 234 and/or responding to prompts included in phonemes 234, as the user speaks sets of phonemes 234; modified with Tomlin, column 9, lines 36-40 and column 10, lines 26-41 notes one or more image sensors, e.g. outward facing image sensors 1108, for capturing visual data from the physical environment in which the HMD device 1100 is located, e.g. to detect movements performed by the wearer or by a person or physical object within a field of view of the display 1102, e.g. detect a user speaking to the wearer of the HMD device, where Figure 9 and associated text, e.g. column 7, lines 19-55 further notes image sensors, e.g. one or more outward facing image sensors noted above, for receiving image data, e.g. images of one or both speakers potentially engaged in a conversation (e.g. images of a user from the perspective of the wearer of the HMD device or images of both users from the perspective of a sensor device), e.g. including facial images, e.g. capturing mouth moving); and a processor to: detect whether the wearer of the HMD uttered the speech (e.g. detecting whether the wearer of the HMD (or another user) uttered the speech, see the details of claim 1); detect whether mouth movement of the wearer occurred while the speech was detected, from the captured facial images (e.g. detecting mouth movements of the wearer of the HMD, see the details of claims 1, 4, and 5); in response to detecting that the mouth movement of the wearer occurred while the speech was detected, render an avatar representing the wearer to have a viseme corresponding to the phoneme (e.g. in response to determining the wearer of the HMD uttered the speech based on mouth movements, rendering an avatar, see details to claims 1, 4, and 5); and in response to determining that another user uttered the speech, not rendering the avatar to have the viseme (e.g. in response to determining the wearer of the HMD did not utter the speech based on mouth movements, e.g. the other person uttered the speech as determined in Tomlin, not rendering an avatar, see details to claims 1 and 3-5). Claim 14 is similar in scope to claims 1 and 3-5 combined, and is therefore rejected under similar rationale. Please see the rejections and rationale regarding the claims above.
As to claim 15, Steptoe et al. modified with Tomlin et al. disclose the captured facial images comprise the viseme corresponding to the phoneme, and wherein the avatar is rendered to have the viseme corresponding to the phoneme based on both the phoneme within the detected speech and the captured facial images including the viseme (see details of claim 7 above).
Claim(s) 9 and 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Steptoe et al. (US 11,468,616) in view of Tomlin et al. (US 10,529,359) as applied to claim 7 above, and further in view of Xiao et al. (US 11,113,859).
As to claim 9, Steptoe et al. modified with Tomlin et al. disclose rendering the avatar (Steptoe, e.g. Figure 3, step 340), but do not disclose, but Xiao et al. disclose rendering the avatar (Figure 4) comprises: applying a model to the captured facial images to generate blendshape weights corresponding to a facial expression of the wearer while the wearer uttered the speech (step 410, column 10, lines 56-64 notes obtaining an audio stream and image data, the audio stream may include or capture a vocal output of a person and image data may include or capture a face image of the person, step 440, column 11, lines 16-45 notes determining blendshapes with corresponding weights to form a 3D model of an avatar, a blendshape can include, correspond to, or be indicative of a structure, shape, and/or profile of a face part (e.g. eye, node, eyebrow, lips, etc.) of an avatar, where a corresponding weight may indicate an amount of emphasis (e.g. on the structure, shape, and/or profile of the face part), the system may obtain, detect, determine, or extract landmarks of a face of the person in the image data, landmarks can include, correspond to, or be indicative of data indicating locations or shapes of different body parts, and according to the landmarks, the system may determine, identify, and/or select a number of blendshapes); identifying the phoneme within the detected speech (step 420, column 10, lines 65 thru column 11, lines 7 notes the system predicts phonemes of vocal output from the audio stream); modifying the generated blendshape weights based on the identified phoneme (step 430, column 11, lines 8-15 notes the system translates the predicted phonemes into visemes, a viseme can include, correspond to, or be indicative of a model or data indicating a mouth shape associated with a particular sound or a phoneme, where as noted above, step 440, determining blendshapes with corresponding weights to form a 3D model of an avatar, where column 11, lines 43-45 notes this step may be performed after one or both of step 420 and 430, thus may render blendshapes determined or generated with respect to the predicted phonemes translated into visemes); and rendering the avatar from the modified blendshape weights (step 450, column 11, lines 46 thru column 12, lines 4 notes the system combining the visemes with the 3D model of the avatar, the system generates or constructs the 3D model of the avatar such that face parts of the 3D model are formed or shaped as indicated by the blendshapes, the system also combines the visemes with the 3D model of the avatar to form a 3D representation of the avatar, by synchronizing the visemes with the 3D model of the avatar in time).
It would have been obvious to one of ordinary skill in the art at the time of the invention to further modify Steptoe et al. modified with Tomlin et al.’s system and method of rendering an avatar with Xiao et al.’s method of using blendshapes and corresponding weights to allow generating the 3D model or representation of the avatar with realistic expressions and/or facial movements in sync with the vocal output of the person to provide improved artificial reality (e.g. including augmented and virtual reality) experiences (column 5, lines 51-55 of Xiao et al.).
As to claim 11, Steptoe et al. modified with Tomlin et al. disclose rendering the avatar (Steptoe, e.g. Figure 3, step 340), but do not disclose, but Xiao et al. disclose rendering the avatar (Figure 4) comprises: applying a model to the captured facial image and to the detected speech to generate blendshape weights corresponding to a facial expression of the wearer while the wearer uttered the speech (step 410, column 10, lines 56-64 notes obtaining an audio stream and image data, the audio stream may include or capture a vocal output of a person and image data may include or capture a face image of the person, step 440, column 11, lines 16-45 notes determining blendshapes with corresponding weights to form a 3D model of an avatar, a blendshape can include, correspond to, or be indicative of a structure, shape, and/or profile of a face part (e.g. eye, node, eyebrow, lips, etc.) of an avatar, where a corresponding weight may indicate an amount of emphasis (e.g. on the structure, shape, and/or profile of the face part), the system may obtain, detect, determine, or extract landmarks of a face of the person in the image data, landmarks can include, correspond to, or be indicative of data indicating locations or shapes of different body parts, and according to the landmarks, the system may determine, identify, and/or select a number of blendshapes, see additional details of claim 9); and rendering the avatar from the blendshape weights (step 450, column 11, lines 46 thru column 12, lines 4 notes the system combining the visemes with the 3D model of the avatar, the system generates or constructs the 3D model of the avatar such that face parts of the 3D model are formed or shaped as indicated by the blendshapes, the system also combines the visemes with the 3D model of the avatar to form a 3D representation of the avatar, by synchronizing the visemes with the 3D model of the avatar in time).
It would have been obvious to one of ordinary skill in the art at the time of the invention to further modify Steptoe et al. modified with Tomlin et al.’s system and method of rendering an avatar with Xiao et al.’s method of using blendshapes and corresponding weights to allow generating the 3D model or representation of the avatar with realistic expressions and/or facial movements in sync with the vocal output of the person to provide improved artificial reality (e.g. including augmented and virtual reality) experiences (column 5, lines 51-55 of Xiao et al.).
Allowable Subject Matter
Claim 10 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including ALL of the limitations of the base claim AND any intervening claims.
The following is a statement of reasons for the indication of allowable subject matter: The prior art of record fails to teach or suggest, singly or combined, the limitations of the claim as recited (see Applicant’s Remarks on pages 10 and 11 of the Amendment filed regarding claim 10).
Response to Arguments
Applicant's arguments filed May 21, 2026 have been fully considered but they are not persuasive. Applicant amends independent claim 1 to recite, “…determining whether a wearer of the HMD or another user uttered the speech…” Applicant argues on pages 7 and 8 of the Amendment filed that the prior art of record, e.g. Steptoe, fails to teach or suggest the limitations of the claim as now amended.
In reply, in light of the amendments of independent claim 1, Steptoe is now modified with newly found reference, Tomlin et al. (US 10,529,359), for teaching the limitations of the claim as now amended. Please see the rejection and rationale of claim 1 above.
Applicant amends independent claim 12 to recite, “…determining, by the processor, whether a user of the electronic device uttered the speech or whether another user uttered the speech…” Applicant argues on page 9 of the Amendment filed that the prior art of record, e.g. Steptoe, fails to teach or suggest the limitations of the claim as now amended.
In reply, independent claim 12 is similar in scope to claims 1 and 2 combined. In light of the amendments of independent claim 12, Steptoe is now modified with newly found reference, Tomlin et al. (US 10,529,359), for teaching the limitations of the claim as now amended. Please see the rejection and rationale of claim 12 (claims 1 and 2) above.
Applicant amends independent claim 14 to recite, “…detect whether the wearer of the HMD uttered the speech…and in response to determining that another user uttered the speech, not rendering the avatar to have the viseme…” Applicant argues on pages 9 and 10 of the Amendment filed that the prior art of record, e.g. Steptoe, fails to teach or suggest the limitations of the claim as now amended.
In reply, independent claim 14 is similar in scope to claims 1 and 3-5 combined. In light of the amendments of independent claim 14, Steptoe is now modified with newly found reference, Tomlin et al. (US 10,529,359), for teaching the limitations of the claim as now amended. Please see the rejection and rationale of claim 14 (claims 1 and 3-5) above.
Applicant amends dependent claim 3 to recite, “…in response to determining that the other user uttered the speech, not rendering the avatar to have the viseme corresponding to the phoneme.” Applicant further argues on page 10 of the Amendment filed regarding dependent claim 3, that “…The Office alleges that “it is not detected that the user produced the sound, it would be obvious that step 340 is not performed, thus the computer-generated avatar is not rendered”… Applicant respectfully disagrees. As noted above, Steptoe assumes that if there is a sound, then the user of the electronic device is the source of the sound. Further, while Steptoe mentions not rendering an avatar to have a viseme when a user is not speaking, Steptoe fails to contemplate a situation where someone else makes a sound or speaks.”
In reply, in light of the amendments of claim 3, the Steptoe is now modified with newly found reference, Tomlin et al. (US 10,529,359), for teaching the limitations of the claim as now amended. Please see the rejection and rationale of claim 3 above.
Applicant further argues claims 6 and 9-11 depend from claim 1, thus are allowable for similar reasons as claim 1 as Karakotsios and Xiao fail to solve the deficiencies of Steptoe.
In reply, in light of the amendments of claim 1, Steptoe is now modified with newly found reference, Tomlin et al. (US 10,529,359), for teaching the limitations of the claim as now amended. Tomlin further teaches portions of the limitations of claim 6, thus Karakotsios is no longer used in the rejection of claim 6. Claims 9 and 11 stand rejected, now in view of Steptoe, Tomlin and Xiao as noted above.
Applicant’s arguments, see pages 10 and 11, filed May 21, 2026, with respect to dependent claim 10 have been fully considered and are persuasive. Applicant argues that Xiao does not teach “…applying a first model to the captured facial image…” and “…applying a second model to the detected speech…” thus does not teach two distinct models where each distinct model is applied to a different element. The Examiner finds this argument persuasive. Therefore, the 35 U.S.C. 103 claim rejection of claim 10 has been withdrawn.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JACINTA M CRAWFORD whose telephone number is (571)270-1539. The examiner can normally be reached 8:30a.m. to 4:30p.m.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, King Y. Poon can be reached at (571)272-7440. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JACINTA M CRAWFORD/Primary Examiner, Art Unit 2617