DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant's arguments filed 05/08/2026 have been fully considered but they are not persuasive. The applicant’s arguments are with regards to the independent claims which are now rejected in view of Poore eta al as detailed in this action.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1, 4, 6, 7, 8, 11, 12, 13, 16, 17, 19 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Poore et al. (US 2021/0055367)(Hereinafter referred to as Poore).
Regarding claim 1, Poore teaches A method (A head-mountable device can include multiple microphones for directional audio detection. The head-mountable device can also include a speaker for audio output and/or a display for visual output. The head-mountable device can be configured to provide visual outputs based on audio inputs by displaying an indicator on a display based on a location of a source of a sound. The head-mountable device can be configured to audio outputs based on audio inputs by modifying an audio output of the speaker based on a detected sound and a target characteristic. Such characteristics can be based on a direction of a gaze of the user, as detected by an eye sensor. See abstract) comprising:
detecting an ambient sound from audio data generated by a plurality of microphones on a wearable device (In particular, a head-mountable device can be provided with multiple microphones for capturing audio information (e.g., sounds) from multiple sources that are located in different directions with respect to the head-mountable device. Multiple microphones distributed across the head-mountable device can provide directional audio detection. The head-mountable device can use the data collected by the microphones to provide visual and/or audio outputs to the user. For example, the detected audio inputs can be rendered with visual outputs by providing indicators directing the user to the source of the sound. This can allow the user to correctly and readily identify the location of the source, even when the user is not readily able to hear the sound independently of the head-mountable device. By further example, the detected audio inputs can be rendered with audio outputs that emphasize (e.g., amplify) certain sounds over others to help the user distinguish between different sounds. See paragraph [0016]);
determining that the ambient sound includes a non-speech sound based on an audio segment of the audio data that includes the ambient sound (In operation 802, a head-mountable device determines a target characteristic of a sound to be detected. The target characteristic can be based on a user input. For example, the target characteristic can be selected (e.g., from a menu) and/or input by a user. The target characteristic can be based on a user input in which the user selects a previously recorded sound to form the basis for analysis of subsequently detected sounds. The target characteristic can be a frequency, volume (e.g., amplitude), location, type of source, and/or range of one or more of the above. For example, a sound of a particular type can be targeted, so that the audio output of the head-mountable device is focused on sounds having such a target characteristic. See paragraph [0066]);
determining a location of a sound source that generated the non-speech sound using the audio segment (As shown in FIG. 4, the display 190 can identify a source 20 of a detected sound as having a particular location (e.g., direction of origin) with respect to the head-mountable device 100. Such determinations can be performed by an array of microphones, as discussed herein. Upon determination of the location of the source 20, the corresponding location on the display 190 can also be determined based on a known spatial relationship between the microphones and the display 190 of the head-mountable device 100. As further shown in FIG. 4, an indicator 300 can be visually output by the display 190 to indicate the location of the source 20. Such an output can help the user visually identify the location of the source 20 even when the user is unable to directly identify the location-based on the user's own detection of the sound. See paragraph [0056]);
generating textual data about at least one attribute of the sound source that generated the non-speech sound using the audio segment (The indicator 300 can include an icon, symbol, graphic, text, word, number, character, picture, or other visible feature that can be displayed at, on, and/or near the source 20 as displayed on the display 190. For example, the indicator 300 can correspond to a known characteristic (e.g., identity, name, color, etc.) of the source 20. Additionally or alternatively, the indicator 300 can include visual features such as color, highlighting, glowing, outlines, shadows, or other contrasting features that allow portions thereof to be more distinctly visible when displayed along with the view to the external environment and/or objects therein. The indicator 300 can move across the display 190 as the user moves the head-mountable device to change the field-of-view being captured and/or displayed. For example, the indicator 300 can maintain its position with respect to the source 20 as the source 20 moves within the display 190 due to the user's movement. See paragraph [0057]); and
displaying the textual data at a display position in a display that is based on the location of the sound source that generated the non-speech sound (The indicator 300 can include an icon, symbol, graphic, text, word, number, character, picture, or other visible feature that can be displayed at, on, and/or near the source 20 as displayed on the display 190. For example, the indicator 300 can correspond to a known characteristic (e.g., identity, name, color, etc.) of the source 20. Additionally or alternatively, the indicator 300 can include visual features such as color, highlighting, glowing, outlines, shadows, or other contrasting features that allow portions thereof to be more distinctly visible when displayed along with the view to the external environment and/or objects therein. The indicator 300 can move across the display 190 as the user moves the head-mountable device to change the field-of-view being captured and/or displayed. For example, the indicator 300 can maintain its position with respect to the source 20 as the source 20 moves within the display 190 due to the user's movement. See paragraph [0057])(See figure 4).
Regarding claim 4, Poore teaches The method of claim 1, wherein the display position includes a three-dimensional position (Referring now to FIG. 3, a user can wear and/or operate a head-mountable device that provides visual outputs based on audio inputs. As shown in FIG. 3, a user 10 can wear the head-mountable device 100, which provides a field-of-view 90 and external environment. A source 20 of a sound 30 can be located within the field-of-view 90. Other sources of sounds can also be located within the field-of-view 90 and/or outside the field-of-view 90. As each of the sounds are received by the user, the head-mountable device 100 can provide visual outputs that guide the user's attention to particular sources of sound. See paragraph [0054])( As shown in FIG. 4, the display 190 can identify a source 20 of a detected sound as having a particular location (e.g., direction of origin) with respect to the head-mountable device 100. Such determinations can be performed by an array of microphones, as discussed herein. Upon determination of the location of the source 20, the corresponding location on the display 190 can also be determined based on a known spatial relationship between the microphones and the display 190 of the head-mountable device 100. As further shown in FIG. 4, an indicator 300 can be visually output by the display 190 to indicate the location of the source 20. Such an output can help the user visually identify the location of the source 20 even when the user is unable to directly identify the location-based on the user's own detection of the sound. See paragraph [0056]), the method further comprising:
computing a direction and a distance of the non-speech sound relative to the wearable device based on the audio segment (As shown in FIG. 4, the display 190 can identify a source 20 of a detected sound as having a particular location (e.g., direction of origin) with respect to the head-mountable device 100. Such determinations can be performed by an array of microphones, as discussed herein. Upon determination of the location of the source 20, the corresponding location on the display 190 can also be determined based on a known spatial relationship between the microphones and the display 190 of the head-mountable device 100. As further shown in FIG. 4, an indicator 300 can be visually output by the display 190 to indicate the location of the source 20. Such an output can help the user visually identify the location of the source 20 even when the user is unable to directly identify the location-based on the user's own detection of the sound. See paragraph [0056]);
determining the location of the sound source that generated the non-speech soundAs shown in FIG. 4, the display 190 can identify a source 20 of a detected sound as having a particular location (e.g., direction of origin) with respect to the head-mountable device 100. Such determinations can be performed by an array of microphones, as discussed herein. Upon determination of the location of the source 20, the corresponding location on the display 190 can also be determined based on a known spatial relationship between the microphones and the display 190 of the head-mountable device 100. As further shown in FIG. 4, an indicator 300 can be visually output by the display 190 to indicate the location of the source 20. Such an output can help the user visually identify the location of the source 20 even when the user is unable to directly identify the location-based on the user's own detection of the sound. See paragraph [0056]); and
displaying the textual data at the three-dimensional position that corresponds to the location of the sound source that generated the non-speech sound (As shown in FIG. 4, the display 190 can identify a source 20 of a detected sound as having a particular location (e.g., direction of origin) with respect to the head-mountable device 100. Such determinations can be performed by an array of microphones, as discussed herein. Upon determination of the location of the source 20, the corresponding location on the display 190 can also be determined based on a known spatial relationship between the microphones and the display 190 of the head-mountable device 100. As further shown in FIG. 4, an indicator 300 can be visually output by the display 190 to indicate the location of the source 20. Such an output can help the user visually identify the location of the source 20 even when the user is unable to directly identify the location-based on the user's own detection of the sound. See paragraph [0056]) (The indicator 300 can include an icon, symbol, graphic, text, word, number, character, picture, or other visible feature that can be displayed at, on, and/or near the source 20 as displayed on the display 190. For example, the indicator 300 can correspond to a known characteristic (e.g., identity, name, color, etc.) of the source 20. Additionally or alternatively, the indicator 300 can include visual features such as color, highlighting, glowing, outlines, shadows, or other contrasting features that allow portions thereof to be more distinctly visible when displayed along with the view to the external environment and/or objects therein. The indicator 300 can move across the display 190 as the user moves the head-mountable device to change the field-of-view being captured and/or displayed. For example, the indicator 300 can maintain its position with respect to the source 20 as the source 20 moves within the display 190 due to the user's movement. See paragraph [0057])(See figure 4).
Regarding claim 6, Poore teaches The method of claim 4, further comprising:
determining an updated value for at least one of the direction or the distance as the sound source moves relative to the wearable device (see figures 5 and 6, outside field of view); and
adjusting the three-dimensional position of the textual data based on the updated value (See figure 6, 300) (The indicator 300 can include an icon, symbol, graphic, text, word, number, character, picture, or other visible feature that can be displayed at, on, and/or near the source 20 as displayed on the display 190. For example, the indicator 300 can correspond to a known characteristic (e.g., identity, name, color, etc.) of the source 20. Additionally or alternatively, the indicator 300 can include visual features such as color, highlighting, glowing, outlines, shadows, or other contrasting features that allow portions thereof to be more distinctly visible when displayed along with the view to the external environment and/or objects therein. The indicator 300 can move across the display 190 as the user moves the head-mountable device to change the field-of-view being captured and/or displayed. For example, the indicator 300 can maintain its position with respect to the source 20 as the source 20 moves within the display 190 due to the user's movement. See paragraph [0057]).
Regarding claim 7, Poore teaches The method of claim 1, further comprising:
generating a directional cue about the location of the sound source that generated the non- speech sound, the directional cue including a textual description that describes a spatial orientation of the sound source with respect to the wearable device; and
displaying the directional cue in the display (See figure 4 and figure 6)(The indicator 300 can include an icon, symbol, graphic, text, word, number, character, picture, or other visible feature that can be displayed at, on, and/or near the source 20 as displayed on the display 190. For example, the indicator 300 can correspond to a known characteristic (e.g., identity, name, color, etc.) of the source 20. Additionally or alternatively, the indicator 300 can include visual features such as color, highlighting, glowing, outlines, shadows, or other contrasting features that allow portions thereof to be more distinctly visible when displayed along with the view to the external environment and/or objects therein. The indicator 300 can move across the display 190 as the user moves the head-mountable device to change the field-of-view being captured and/or displayed. For example, the indicator 300 can maintain its position with respect to the source 20 as the source 20 moves within the display 190 due to the user's movement. See paragraph [0057]).
Regarding claim 8, Poore teaches A wearable device (A head-mountable device can include multiple microphones for directional audio detection. The head-mountable device can also include a speaker for audio output and/or a display for visual output. The head-mountable device can be configured to provide visual outputs based on audio inputs by displaying an indicator on a display based on a location of a source of a sound. The head-mountable device can be configured to audio outputs based on audio inputs by modifying an audio output of the speaker based on a detected sound and a target characteristic. Such characteristics can be based on a direction of a gaze of the user, as detected by an eye sensor. See abstract) comprising:
at least one processor (As shown in FIG. 2, the head-mountable device 100 can include a controller 270 with one or more processing units that include or are configured to access a memory 218 having instructions stored thereon. See paragraph [0041]); and
a non-transitory computer-readable medium storing executable instructions that cause the at least one processor (As shown in FIG. 2, the head-mountable device 100 can include a controller 270 with one or more processing units that include or are configured to access a memory 218 having instructions stored thereon. See paragraph [0041]) to:
detect an ambient sound from audio data generated by a plurality of microphones on the wearable device (In particular, a head-mountable device can be provided with multiple microphones for capturing audio information (e.g., sounds) from multiple sources that are located in different directions with respect to the head-mountable device. Multiple microphones distributed across the head-mountable device can provide directional audio detection. The head-mountable device can use the data collected by the microphones to provide visual and/or audio outputs to the user. For example, the detected audio inputs can be rendered with visual outputs by providing indicators directing the user to the source of the sound. This can allow the user to correctly and readily identify the location of the source, even when the user is not readily able to hear the sound independently of the head-mountable device. By further example, the detected audio inputs can be rendered with audio outputs that emphasize (e.g., amplify) certain sounds over others to help the user distinguish between different sounds. See paragraph [0016]);
determine that the ambient sound includes a non-speech sound based on an audio segment of the audio data that includes the ambient sound (In operation 802, a head-mountable device determines a target characteristic of a sound to be detected. The target characteristic can be based on a user input. For example, the target characteristic can be selected (e.g., from a menu) and/or input by a user. The target characteristic can be based on a user input in which the user selects a previously recorded sound to form the basis for analysis of subsequently detected sounds. The target characteristic can be a frequency, volume (e.g., amplitude), location, type of source, and/or range of one or more of the above. For example, a sound of a particular type can be targeted, so that the audio output of the head-mountable device is focused on sounds having such a target characteristic. See paragraph [0066]);
determine a location of a sound source that generated the non-speech sound using the audio segment (As shown in FIG. 4, the display 190 can identify a source 20 of a detected sound as having a particular location (e.g., direction of origin) with respect to the head-mountable device 100. Such determinations can be performed by an array of microphones, as discussed herein. Upon determination of the location of the source 20, the corresponding location on the display 190 can also be determined based on a known spatial relationship between the microphones and the display 190 of the head-mountable device 100. As further shown in FIG. 4, an indicator 300 can be visually output by the display 190 to indicate the location of the source 20. Such an output can help the user visually identify the location of the source 20 even when the user is unable to directly identify the location-based on the user's own detection of the sound. See paragraph [0056]);
generate textual data about at least one attribute of the sound source that generated the non-speech sound using the audio segment (The indicator 300 can include an icon, symbol, graphic, text, word, number, character, picture, or other visible feature that can be displayed at, on, and/or near the source 20 as displayed on the display 190. For example, the indicator 300 can correspond to a known characteristic (e.g., identity, name, color, etc.) of the source 20. Additionally or alternatively, the indicator 300 can include visual features such as color, highlighting, glowing, outlines, shadows, or other contrasting features that allow portions thereof to be more distinctly visible when displayed along with the view to the external environment and/or objects therein. The indicator 300 can move across the display 190 as the user moves the head-mountable device to change the field-of-view being captured and/or displayed. For example, the indicator 300 can maintain its position with respect to the source 20 as the source 20 moves within the display 190 due to the user's movement. See paragraph [0057]); and
display the textual data at a display position in a display that is based on the location of the sound source that generated the non-speech sound (The indicator 300 can include an icon, symbol, graphic, text, word, number, character, picture, or other visible feature that can be displayed at, on, and/or near the source 20 as displayed on the display 190. For example, the indicator 300 can correspond to a known characteristic (e.g., identity, name, color, etc.) of the source 20. Additionally or alternatively, the indicator 300 can include visual features such as color, highlighting, glowing, outlines, shadows, or other contrasting features that allow portions thereof to be more distinctly visible when displayed along with the view to the external environment and/or objects therein. The indicator 300 can move across the display 190 as the user moves the head-mountable device to change the field-of-view being captured and/or displayed. For example, the indicator 300 can maintain its position with respect to the source 20 as the source 20 moves within the display 190 due to the user's movement. See paragraph [0057])(See figure 4).
Regarding claim 11, Poore teaches the wearable device of claim 8, wherein the display position includes a three-dimensional position (Referring now to FIG. 3, a user can wear and/or operate a head-mountable device that provides visual outputs based on audio inputs. As shown in FIG. 3, a user 10 can wear the head-mountable device 100, which provides a field-of-view 90 and external environment. A source 20 of a sound 30 can be located within the field-of-view 90. Other sources of sounds can also be located within the field-of-view 90 and/or outside the field-of-view 90. As each of the sounds are received by the user, the head-mountable device 100 can provide visual outputs that guide the user's attention to particular sources of sound. See paragraph [0054])( As shown in FIG. 4, the display 190 can identify a source 20 of a detected sound as having a particular location (e.g., direction of origin) with respect to the head-mountable device 100. Such determinations can be performed by an array of microphones, as discussed herein. Upon determination of the location of the source 20, the corresponding location on the display 190 can also be determined based on a known spatial relationship between the microphones and the display 190 of the head-mountable device 100. As further shown in FIG. 4, an indicator 300 can be visually output by the display 190 to indicate the location of the source 20. Such an output can help the user visually identify the location of the source 20 even when the user is unable to directly identify the location-based on the user's own detection of the sound. See paragraph [0056]), wherein the executable instructions include instructions that cause the at least one processor to:
compute a direction and a distance of the non-speech sound relative to the wearable device using the audio segment (As shown in FIG. 4, the display 190 can identify a source 20 of a detected sound as having a particular location (e.g., direction of origin) with respect to the head-mountable device 100. Such determinations can be performed by an array of microphones, as discussed herein. Upon determination of the location of the source 20, the corresponding location on the display 190 can also be determined based on a known spatial relationship between the microphones and the display 190 of the head-mountable device 100. As further shown in FIG. 4, an indicator 300 can be visually output by the display 190 to indicate the location of the source 20. Such an output can help the user visually identify the location of the source 20 even when the user is unable to directly identify the location-based on the user's own detection of the sound. See paragraph [0056]);
determine the location of the sound source that generated the non-speech sound based on the direction and the distance of the non-speech sound (As shown in FIG. 4, the display 190 can identify a source 20 of a detected sound as having a particular location (e.g., direction of origin) with respect to the head-mountable device 100. Such determinations can be performed by an array of microphones, as discussed herein. Upon determination of the location of the source 20, the corresponding location on the display 190 can also be determined based on a known spatial relationship between the microphones and the display 190 of the head-mountable device 100. As further shown in FIG. 4, an indicator 300 can be visually output by the display 190 to indicate the location of the source 20. Such an output can help the user visually identify the location of the source 20 even when the user is unable to directly identify the location-based on the user's own detection of the sound. See paragraph [0056]); and
display the textual data at the three-dimensional position in the display based on the location of the sound source of the non-speech sound (As shown in FIG. 4, the display 190 can identify a source 20 of a detected sound as having a particular location (e.g., direction of origin) with respect to the head-mountable device 100. Such determinations can be performed by an array of microphones, as discussed herein. Upon determination of the location of the source 20, the corresponding location on the display 190 can also be determined based on a known spatial relationship between the microphones and the display 190 of the head-mountable device 100. As further shown in FIG. 4, an indicator 300 can be visually output by the display 190 to indicate the location of the source 20. Such an output can help the user visually identify the location of the source 20 even when the user is unable to directly identify the location-based on the user's own detection of the sound. See paragraph [0056]) (The indicator 300 can include an icon, symbol, graphic, text, word, number, character, picture, or other visible feature that can be displayed at, on, and/or near the source 20 as displayed on the display 190. For example, the indicator 300 can correspond to a known characteristic (e.g., identity, name, color, etc.) of the source 20. Additionally or alternatively, the indicator 300 can include visual features such as color, highlighting, glowing, outlines, shadows, or other contrasting features that allow portions thereof to be more distinctly visible when displayed along with the view to the external environment and/or objects therein. The indicator 300 can move across the display 190 as the user moves the head-mountable device to change the field-of-view being captured and/or displayed. For example, the indicator 300 can maintain its position with respect to the source 20 as the source 20 moves within the display 190 due to the user's movement. See paragraph [0057])(See figure 4).
Regarding claim 12, Poore teaches The wearable device of claim 11, wherein the executable instructions include instructions that cause the at least one processor to:
determine an updated value for at least one of the direction or the distance as the sound source moves relative to the wearable device (see figures 5 and 6, outside field of view); and
adjust the three-dimensional position of the textual data based on the updated value (See figure 6, 300) (The indicator 300 can include an icon, symbol, graphic, text, word, number, character, picture, or other visible feature that can be displayed at, on, and/or near the source 20 as displayed on the display 190. For example, the indicator 300 can correspond to a known characteristic (e.g., identity, name, color, etc.) of the source 20. Additionally or alternatively, the indicator 300 can include visual features such as color, highlighting, glowing, outlines, shadows, or other contrasting features that allow portions thereof to be more distinctly visible when displayed along with the view to the external environment and/or objects therein. The indicator 300 can move across the display 190 as the user moves the head-mountable device to change the field-of-view being captured and/or displayed. For example, the indicator 300 can maintain its position with respect to the source 20 as the source 20 moves within the display 190 due to the user's movement. See paragraph [0057]).
Regarding claim 13, Poore teaches The wearable device of claim 8, wherein the executable instructions include instructions that cause the at least one processor to:
determine that the location of the sound source that generated the non-speech sound is outside a field of view of a user of the wearable device (Referring now to FIG. 5, a source 20 of a sound 30 can be located outside of the field-of-view 90 provided by the head-mountable device 100. Other sources of sounds can also be located within the field-of-view 90 and/or outside the field-of-view 90. As each of the sounds are received by the user, the head-mountable device 100 can provide visual outputs that guide the user's attention to particular sources of sound, even when such sources are outside of the field-of-view. See paragraph [0058]);
in response to the location of the sound source being determined as outside the field of view, generate a directional cue about the location of the sound source, the directional cue including a textual description that describes a spatial orientation of the sound source with respect to the wearable device (Upon determination of the location of the source 20, it can be further determined that the location of the source is not within a field-of-view provided by the display 190. Such a determination can be made based on a known spatial relationship between the microphones and the display 190 of the head-mountable device 100. As further shown in FIG. 6, the indicator 300 can be visually output by the display 190 to indicate the location of the source even when the source is not displayed within the field-of-view of the display 190. As such, the indicator 300 can suggest to the user the direction in which the user may change its position and/or orientation to capture a view of the source. Such an output can help the user visually identify the location of the source even when the user is unable to directly identify the location-based on the user's own detection of the sound. See paragraph [0060])( The indicator 300 can include an icon, symbol, graphic, text, word, number, character, picture, or other visible feature that can be displayed at, on, and/or near the portion of the display 190 that most closely corresponds to the location of the source. By further example, the indicator 300 can correspond to a known characteristic (e.g., identity, name, color, etc.) of the source. Additionally or alternatively, the indicator 300 can include visual features such as color, highlighting, glowing, outlines, shadows, or other contrasting features that allow portions thereof to be more distinctly visible when displayed along with the view to the external environment and/or objects therein. See paragraph [0061]); and
display the directional cue in the display ( The indicator 300 can include an icon, symbol, graphic, text, word, number, character, picture, or other visible feature that can be displayed at, on, and/or near the portion of the display 190 that most closely corresponds to the location of the source. By further example, the indicator 300 can correspond to a known characteristic (e.g., identity, name, color, etc.) of the source. Additionally or alternatively, the indicator 300 can include visual features such as color, highlighting, glowing, outlines, shadows, or other contrasting features that allow portions thereof to be more distinctly visible when displayed along with the view to the external environment and/or objects therein. See paragraph [0061])(See figure 6).
Regarding claim 16, Poore teaches A non-transitory computer-readable medium storing executable instructions that when executed by at least one processor cause the at least one processor to execute operations (As shown in FIG. 2, the head-mountable device 100 can include a controller 270 with one or more processing units that include or are configured to access a memory 218 having instructions stored thereon. See paragraph [0041]), the operations comprising:
detecting an ambient sound from audio data generated by a plurality of microphones on a wearable device (In particular, a head-mountable device can be provided with multiple microphones for capturing audio information (e.g., sounds) from multiple sources that are located in different directions with respect to the head-mountable device. Multiple microphones distributed across the head-mountable device can provide directional audio detection. The head-mountable device can use the data collected by the microphones to provide visual and/or audio outputs to the user. For example, the detected audio inputs can be rendered with visual outputs by providing indicators directing the user to the source of the sound. This can allow the user to correctly and readily identify the location of the source, even when the user is not readily able to hear the sound independently of the head-mountable device. By further example, the detected audio inputs can be rendered with audio outputs that emphasize (e.g., amplify) certain sounds over others to help the user distinguish between different sounds. See paragraph [0016]);
determining that the ambient sound includes a non-speech sound based on an audio segment of the audio data that includes the ambient sound(In operation 802, a head-mountable device determines a target characteristic of a sound to be detected. The target characteristic can be based on a user input. For example, the target characteristic can be selected (e.g., from a menu) and/or input by a user. The target characteristic can be based on a user input in which the user selects a previously recorded sound to form the basis for analysis of subsequently detected sounds. The target characteristic can be a frequency, volume (e.g., amplitude), location, type of source, and/or range of one or more of the above. For example, a sound of a particular type can be targeted, so that the audio output of the head-mountable device is focused on sounds having such a target characteristic. See paragraph [0066]);
determining, using the audio segment, that a location of a sound source that generated the non-speech sound is outside a field of view of a user of the wearable device (Referring now to FIG. 5, a source 20 of a sound 30 can be located outside of the field-of-view 90 provided by the head-mountable device 100. Other sources of sounds can also be located within the field-of-view 90 and/or outside the field-of-view 90. As each of the sounds are received by the user, the head-mountable device 100 can provide visual outputs that guide the user's attention to particular sources of sound, even when such sources are outside of the field-of-view. See paragraph [0058]);
generating textual data about at least one attribute of the sound source that generated the non-speech sound using the audio segment (The indicator 300 can include an icon, symbol, graphic, text, word, number, character, picture, or other visible feature that can be displayed at, on, and/or near the source 20 as displayed on the display 190. For example, the indicator 300 can correspond to a known characteristic (e.g., identity, name, color, etc.) of the source 20. Additionally or alternatively, the indicator 300 can include visual features such as color, highlighting, glowing, outlines, shadows, or other contrasting features that allow portions thereof to be more distinctly visible when displayed along with the view to the external environment and/or objects therein. The indicator 300 can move across the display 190 as the user moves the head-mountable device to change the field-of-view being captured and/or displayed. For example, the indicator 300 can maintain its position with respect to the source 20 as the source 20 moves within the display 190 due to the user's movement. See paragraph [0057]);
generating a directional cue using the audio segment, the directional cue including a textual description that describes a spatial orientation of the sound source that generated the non- speech sound with respect to the wearable device (Upon determination of the location of the source 20, it can be further determined that the location of the source is not within a field-of-view provided by the display 190. Such a determination can be made based on a known spatial relationship between the microphones and the display 190 of the head-mountable device 100. As further shown in FIG. 6, the indicator 300 can be visually output by the display 190 to indicate the location of the source even when the source is not displayed within the field-of-view of the display 190. As such, the indicator 300 can suggest to the user the direction in which the user may change its position and/or orientation to capture a view of the source. Such an output can help the user visually identify the location of the source even when the user is unable to directly identify the location-based on the user's own detection of the sound. See paragraph [0060])( The indicator 300 can include an icon, symbol, graphic, text, word, number, character, picture, or other visible feature that can be displayed at, on, and/or near the portion of the display 190 that most closely corresponds to the location of the source. By further example, the indicator 300 can correspond to a known characteristic (e.g., identity, name, color, etc.) of the source. Additionally or alternatively, the indicator 300 can include visual features such as color, highlighting, glowing, outlines, shadows, or other contrasting features that allow portions thereof to be more distinctly visible when displayed along with the view to the external environment and/or objects therein. See paragraph [0061]); and
displaying the textual data and the directional cue at a display position in a display of the wearable device ( The indicator 300 can include an icon, symbol, graphic, text, word, number, character, picture, or other visible feature that can be displayed at, on, and/or near the portion of the display 190 that most closely corresponds to the location of the source. By further example, the indicator 300 can correspond to a known characteristic (e.g., identity, name, color, etc.) of the source. Additionally or alternatively, the indicator 300 can include visual features such as color, highlighting, glowing, outlines, shadows, or other contrasting features that allow portions thereof to be more distinctly visible when displayed along with the view to the external environment and/or objects therein. See paragraph [0061])(See figure 6).
Regarding claim 17, Poore teaches The non-transitory computer-readable medium of claim 16, wherein the textual data includes information identifying a type of the non-speech sound and information identifying a category of a physical object generating the non-speech sound, the textual data also identifying a characteristic about the physical object that is different from the category of the physical object (In operation 802, a head-mountable device determines a target characteristic of a sound to be detected. The target characteristic can be based on a user input. For example, the target characteristic can be selected (e.g., from a menu) and/or input by a user. The target characteristic can be based on a user input in which the user selects a previously recorded sound to form the basis for analysis of subsequently detected sounds. The target characteristic can be a frequency, volume (e.g., amplitude), location, type of source, and/or range of one or more of the above. For example, a sound of a particular type can be targeted, so that the audio output of the head-mountable device is focused on sounds having such a target characteristic. See paragraph [0066]) (The indicator 300 can include an icon, symbol, graphic, text, word, number, character, picture, or other visible feature that can be displayed at, on, and/or near the source 20 as displayed on the display 190. For example, the indicator 300 can correspond to a known characteristic (e.g., identity, name, color, etc.) of the source 20. Additionally or alternatively, the indicator 300 can include visual features such as color, highlighting, glowing, outlines, shadows, or other contrasting features that allow portions thereof to be more distinctly visible when displayed along with the view to the external environment and/or objects therein. The indicator 300 can move across the display 190 as the user moves the head-mountable device to change the field-of-view being captured and/or displayed. For example, the indicator 300 can maintain its position with respect to the source 20 as the source 20 moves within the display 190 due to the user's movement. See paragraph [0057]).
Regarding claim 19, Poore teaches The non-transitory computer-readable medium of claim 16, wherein the display position is a first display position, wherein the operations further comprise:
determining that the sound source of the non-speech sound has moved into the field of view of the user of the wearable device (As shown in FIG. 4, the display 190 can identify a source 20 of a detected sound as having a particular location (e.g., direction of origin) with respect to the head-mountable device 100. Such determinations can be performed by an array of microphones, as discussed herein. Upon determination of the location of the source 20, the corresponding location on the display 190 can also be determined based on a known spatial relationship between the microphones and the display 190 of the head-mountable device 100. As further shown in FIG. 4, an indicator 300 can be visually output by the display 190 to indicate the location of the source 20. Such an output can help the user visually identify the location of the source 20 even when the user is unable to directly identify the location-based on the user's own detection of the sound. See paragraph [0056])( Referring now to FIG. 3, a user can wear and/or operate a head-mountable device that provides visual outputs based on audio inputs. As shown in FIG. 3, a user 10 can wear the head-mountable device 100, which provides a field-of-view 90 and external environment. A source 20 of a sound 30 can be located within the field-of-view 90. Other sources of sounds can also be located within the field-of-view 90 and/or outside the field-of-view 90. As each of the sounds are received by the user, the head-mountable device 100 can provide visual outputs that guide the user's attention to particular sources of sound. See paragraph [0054]);
determining spatial coordinates of the sound source of the non-speech sound based on the audio data (As shown in FIG. 4, the display 190 can identify a source 20 of a detected sound as having a particular location (e.g., direction of origin) with respect to the head-mountable device 100. Such determinations can be performed by an array of microphones, as discussed herein. Upon determination of the location of the source 20, the corresponding location on the display 190 can also be determined based on a known spatial relationship between the microphones and the display 190 of the head-mountable device 100. As further shown in FIG. 4, an indicator 300 can be visually output by the display 190 to indicate the location of the source 20. Such an output can help the user visually identify the location of the source 20 even when the user is unable to directly identify the location-based on the user's own detection of the sound. See paragraph [0056]); and
displaying the textual data at a second display position, the second display position being based on the spatial coordinates of the sound source of the non-speech sound (The indicator 300 can include an icon, symbol, graphic, text, word, number, character, picture, or other visible feature that can be displayed at, on, and/or near the source 20 as displayed on the display 190. For example, the indicator 300 can correspond to a known characteristic (e.g., identity, name, color, etc.) of the source 20. Additionally or alternatively, the indicator 300 can include visual features such as color, highlighting, glowing, outlines, shadows, or other contrasting features that allow portions thereof to be more distinctly visible when displayed along with the view to the external environment and/or objects therein. The indicator 300 can move across the display 190 as the user moves the head-mountable device to change the field-of-view being captured and/or displayed. For example, the indicator 300 can maintain its position with respect to the source 20 as the source 20 moves within the display 190 due to the user's movement. See paragraph [0057]).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claim(s) 2, 5, 9 are rejected under 35 U.S.C. 103 as being unpatentable over Poore et al. (US 2021/0055367)(Hereinafter referred to as Poore) in view of Takumi Asakura, (“Augmented-Reality Presentation of Household Sounds for Deaf and Hard-of-Hearing People”. 2023)(Hereinafter referred to as Asakura).
Regarding claim 2, Poore teaches the method of claim 1, but is silent to wherein generating the textual data includes inputting the audio segment into a machine-learning model, and generating, by the machine-learning model, the textual data based on the audio segment, the machine-learning model being stored on the wearable device.
Asakura teaches classifying audio utilizing a machine learning model and outputting a description for a hearing-impaired user (Normal-hearing people use sound as a cue to recognize various events that occur in their surrounding environment; however, this is not possible for deaf and hearing of hard (DHH) people, and in such a context they may not be able to freely detect their surrounding environment. Therefore, there is an opportunity to create a convenient device that can detect sounds occurring in daily life and present them visually instead of auditorily. Additionally, it is of great importance to appropriately evaluate how such a supporting device would change the lives of DHH people. The current study proposes an augmented-reality-based system for presenting household sounds to DHH people as visual information. We examined the effect of displaying both the icons indicating sounds classified by machine learning and a dynamic spectrogram indicating the real-time time frequency characteristics of the environmental sounds. First, the issues that DHH people perceive as problems in their daily lives were investigated through a survey, suggesting that DHH people need to visualize their surrounding sound environment. Then, after the accuracy of the machine-learning based classifier installed in the proposed system was validated, the subjective impression of how the proposed system increased the comfort of daily life was obtained through a field experiment in a real residence. The results confirmed that the comfort of daily life in household spaces can be improved by combining not only the classification results of machine learning but also the real-time display of spectrograms. See abstract).
Poore and Asakura teach of presenting information based on audio data and Asakura teaches that the audio can be classified by a machine learning model, therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the system of Poore with the classification techniques of Asakura such that the system could classify a wider range of sounds.
Regarding claim 5, Poore teaches The method of claim 1, but is silent to further comprising:
generating a filtered audio segment by applying a filter to the audio segment to reduce background noise in the audio segment; and
inputting the filtered audio segment to a machine-learning model to generate the textual data.
Asakura teaches setting a threshold pressure level to remove background noise and classifying audio utilizing a machine learning model and outputting a description for a hearing-impaired user (Normal-hearing people use sound as a cue to recognize various events that occur in their surrounding environment; however, this is not possible for deaf and hearing of hard (DHH) people, and in such a context they may not be able to freely detect their surrounding environment. Therefore, there is an opportunity to create a convenient device that can detect sounds occurring in daily life and present them visually instead of auditorily. Additionally, it is of great importance to appropriately evaluate how such a supporting device would change the lives of DHH people. The current study proposes an augmented-reality-based system for presenting household sounds to DHH people as visual information. We examined the effect of displaying both the icons indicating sounds classified by machine learning and a dynamic spectrogram indicating the real-time time frequency characteristics of the environmental sounds. First, the issues that DHH people perceive as problems in their daily lives were investigated through a survey, suggesting that DHH people need to visualize their surrounding sound environment. Then, after the accuracy of the machine-learning based classifier installed in the proposed system was validated, the subjective impression of how the proposed system increased the comfort of daily life was obtained through a field experiment in a real residence. The results confirmed that the comfort of daily life in household spaces can be improved by combining not only the classification results of machine learning but also the real-time display of spectrograms. See abstract) (Household sounds, such as the sound of running water in a distant kitchen, have very low sound pressure levels, and show almost the same level of sound pressure as background noise, but their low sound pressure levels do not mean that they are less important. In the present study, the sound pressure level was set to the level of the background noise plus 5 dB, which eliminated the possibility that meaningless background noise would be recognized as an environmental sound. See page 11, second to last paragraph).
Poore and Asakura teach of presenting information based on audio data and Asakura teaches that the audio can remove noise and then be classified by a machine learning model, therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the system of Poore with the noise filtering and classification techniques of Asakura such that the system could classify a wider range of sounds while ignoring irrelevant sounds.
Regarding claim 9, Poore teaches the wearable device of claim 8, but is silent to wherein the executable instructions include instructions that cause the at least one processor to:
input the audio segment into a machine-learning model; and
generate, by the machine-learning model, the textual data using the audio segment, the textual data includes information identifying a type of the non-speech sound and information identifying a category of a physical object that produced the non-speech sound.
Asakura teaches classifying audio utilizing a machine learning model and outputting a description for a hearing-impaired user (Normal-hearing people use sound as a cue to recognize various events that occur in their surrounding environment; however, this is not possible for deaf and hearing of hard (DHH) people, and in such a context they may not be able to freely detect their surrounding environment. Therefore, there is an opportunity to create a convenient device that can detect sounds occurring in daily life and present them visually instead of auditorily. Additionally, it is of great importance to appropriately evaluate how such a supporting device would change the lives of DHH people. The current study proposes an augmented-reality-based system for presenting household sounds to DHH people as visual information. We examined the effect of displaying both the icons indicating sounds classified by machine learning and a dynamic spectrogram indicating the real-time time frequency characteristics of the environmental sounds. First, the issues that DHH people perceive as problems in their daily lives were investigated through a survey, suggesting that DHH people need to visualize their surrounding sound environment. Then, after the accuracy of the machine-learning based classifier installed in the proposed system was validated, the subjective impression of how the proposed system increased the comfort of daily life was obtained through a field experiment in a real residence. The results confirmed that the comfort of daily life in household spaces can be improved by combining not only the classification results of machine learning but also the real-time display of spectrograms. See abstract).
Poore and Asakura teach of presenting information based on audio data and Asakura teaches that the audio can be classified by a machine learning model, therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the system of Poore with the classification techniques of Asakura such that the system could classify a wider range of sounds.
Claim(s) 3, 10, 14, 15, 18, 22 are rejected under 35 U.S.C. 103 as being unpatentable over Poore et al. (US 2021/0055367)(Hereinafter referred to as Poore) in view of McCulloch et al. (US 2014/0337023)(Hereinafter referred to as McCulloch).
Regarding claim 3, Poore teaches the method of claim 1, wherein the ambient sound is a first ambient sound, the sound source is a first sound source, the display position is a first display position, the audio segment is a first audio segment, and the textual data is first textual data (In particular, a head-mountable device can be provided with multiple microphones for capturing audio information (e.g., sounds) from multiple sources that are located in different directions with respect to the head-mountable device. Multiple microphones distributed across the head-mountable device can provide directional audio detection. The head-mountable device can use the data collected by the microphones to provide visual and/or audio outputs to the user. For example, the detected audio inputs can be rendered with visual outputs by providing indicators directing the user to the source of the sound. This can allow the user to correctly and readily identify the location of the source, even when the user is not readily able to hear the sound independently of the head-mountable device. By further example, the detected audio inputs can be rendered with audio outputs that emphasize (e.g., amplify) certain sounds over others to help the user distinguish between different sounds. See paragraph [0016])( Other sources of sounds can also be located within the field-of-view 90 and/or outside the field-of-view 90. As each of the sounds are received by the user, the head-mountable device 100 can provide visual outputs that guide the user's attention to particular sources of sound. See paragraph [0054]), the method further comprising:
detecting a second ambient sound from the audio data, the second ambient sound at least partially overlapping with the first ambient sound (In particular, a head-mountable device can be provided with multiple microphones for capturing audio information (e.g., sounds) from multiple sources that are located in different directions with respect to the head-mountable device. Multiple microphones distributed across the head-mountable device can provide directional audio detection. The head-mountable device can use the data collected by the microphones to provide visual and/or audio outputs to the user. For example, the detected audio inputs can be rendered with visual outputs by providing indicators directing the user to the source of the sound. This can allow the user to correctly and readily identify the location of the source, even when the user is not readily able to hear the sound independently of the head-mountable device. By further example, the detected audio inputs can be rendered with audio outputs that emphasize (e.g., amplify) certain sounds over others to help the user distinguish between different sounds. See paragraph [0016]), but is silent to
determining that the second ambient sound includes speech based on a second audio segment of the audio data that includes the second ambient sound;
generating second textual data about at least one attribute of a second sound source that generated the speech using the second audio segment;
initiating conversion of the speech to a textual translation; and
displaying the textual translation and the second textual data at a second display position of the display that is based on a location of the second sound source that generated the speech.
McCulloch teaches detecting speech and utilizing beamforming to locate and translate the speech into a text overlay. (Embodiments that relate to converting audio inputs from an environment into text are disclosed. For example, in one disclosed embodiment a speech conversion program receives audio inputs from a microphone array of a head-mounted display device. Image data is captured from the environment, and one or more possible faces are detected from image data. Eye-tracking data is used to determine a target face on which a user is focused. A beamforming technique is applied to at least a portion of the audio inputs to identify target audio inputs that are associated with the target face. The target audio inputs are converted into text that is displayed via a transparent display of the head-mounted display device. See abstract)(See figure 3).
Poore and McCulloch teach of presenting indicators over audible objects and McCulloch teaches that by using a beamforming and audio translation technique text can be located above the speaker, therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the system of Poore with the beamforming and audio translation to text of McCulloch such that a user could visualize responses from multiple individuals.
Regarding claim 10, Poore teaches The wearable device of claim 8, wherein the ambient sound is a first ambient sound, the sound source is a first sound source, the display position is a first display position, the audio segment is a first audio segment, and the textual data is first textual data, wherein the executable instructions include instructions that cause the at least one processor (In particular, a head-mountable device can be provided with multiple microphones for capturing audio information (e.g., sounds) from multiple sources that are located in different directions with respect to the head-mountable device. Multiple microphones distributed across the head-mountable device can provide directional audio detection. The head-mountable device can use the data collected by the microphones to provide visual and/or audio outputs to the user. For example, the detected audio inputs can be rendered with visual outputs by providing indicators directing the user to the source of the sound. This can allow the user to correctly and readily identify the location of the source, even when the user is not readily able to hear the sound independently of the head-mountable device. By further example, the detected audio inputs can be rendered with audio outputs that emphasize (e.g., amplify) certain sounds over others to help the user distinguish between different sounds. See paragraph [0016])( Other sources of sounds can also be located within the field-of-view 90 and/or outside the field-of-view 90. As each of the sounds are received by the user, the head-mountable device 100 can provide visual outputs that guide the user's attention to particular sources of sound. See paragraph [0054]) to:
detect a second ambient sound from the audio data (In particular, a head-mountable device can be provided with multiple microphones for capturing audio information (e.g., sounds) from multiple sources that are located in different directions with respect to the head-mountable device. Multiple microphones distributed across the head-mountable device can provide directional audio detection. The head-mountable device can use the data collected by the microphones to provide visual and/or audio outputs to the user. For example, the detected audio inputs can be rendered with visual outputs by providing indicators directing the user to the source of the sound. This can allow the user to correctly and readily identify the location of the source, even when the user is not readily able to hear the sound independently of the head-mountable device. By further example, the detected audio inputs can be rendered with audio outputs that emphasize (e.g., amplify) certain sounds over others to help the user distinguish between different sounds. See paragraph [0016]), but is silent to
determine that the second ambient sound includes speech based on a second audio segment of the audio data that includes the second ambient sound;
generate second textual data about a second sound source that generated the speech using the second audio segment, the second textual data including a textual description of a physical characteristic of the second sound source;
initiate conversion of the speech to a textual translation; and
display the textual translation and the second textual data at a second display position in the display that is based on a location of the second sound source that generated the speech.
McCulloch teaches detecting speech and utilizing beamforming to locate and translate the speech into a text overlay. (Embodiments that relate to converting audio inputs from an environment into text are disclosed. For example, in one disclosed embodiment a speech conversion program receives audio inputs from a microphone array of a head-mounted display device. Image data is captured from the environment, and one or more possible faces are detected from image data. Eye-tracking data is used to determine a target face on which a user is focused. A beamforming technique is applied to at least a portion of the audio inputs to identify target audio inputs that are associated with the target face. The target audio inputs are converted into text that is displayed via a transparent display of the head-mounted display device. See abstract)(See figure 3).
Poore and McCulloch teach of presenting indicators over audible objects and McCulloch teaches that by using a beamforming and audio translation technique text can be located above the speaker, therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the system of Poore with the beamforming and audio translation to text of McCulloch such that a user could visualize responses from multiple individuals.
Regarding claim 14, Poore teaches The wearable device of claim 8, but is silent to wherein the plurality of microphones are arranged as a beamforming array, wherein the executable instructions include instructions that cause the at least one processor to:
receive an audio signal from the beamforming array; and
compute a direction of arrival of the ambient sound relative to the wearable device using the audio signal.
McCulloch teaches detecting speech and utilizing beamforming to locate audio target associated with a particular face and translate the speech into a text overlay. (Embodiments that relate to converting audio inputs from an environment into text are disclosed. For example, in one disclosed embodiment a speech conversion program receives audio inputs from a microphone array of a head-mounted display device. Image data is captured from the environment, and one or more possible faces are detected from image data. Eye-tracking data is used to determine a target face on which a user is focused. A beamforming technique is applied to at least a portion of the audio inputs to identify target audio inputs that are associated with the target face. The target audio inputs are converted into text that is displayed via a transparent display of the head-mounted display device. See abstract)(See figure 3).
Poore and McCulloch teach of presenting indicators over audible objects and McCulloch teaches that by using a beamforming and audio translation technique text can be located above the speaker, therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the system of Poore with the beamforming and audio translation to text of McCulloch such that a user could visualize responses from multiple individuals.
Regarding claim 15, Poore in view of McCulloch teaches the wearable device of claim 10, wherein the executable instructions include instructions that cause the at least one processor to:
detect the second ambient sound at least partially in parallel with detecting the first ambient sound (Poore; In particular, a head-mountable device can be provided with multiple microphones for capturing audio information (e.g., sounds) from multiple sources that are located in different directions with respect to the head-mountable device. Multiple microphones distributed across the head-mountable device can provide directional audio detection. The head-mountable device can use the data collected by the microphones to provide visual and/or audio outputs to the user. For example, the detected audio inputs can be rendered with visual outputs by providing indicators directing the user to the source of the sound. This can allow the user to correctly and readily identify the location of the source, even when the user is not readily able to hear the sound independently of the head-mountable device. By further example, the detected audio inputs can be rendered with audio outputs that emphasize (e.g., amplify) certain sounds over others to help the user distinguish between different sounds. See paragraph [0016]).
Regarding claim 18, Poore teaches the non-transitory computer-readable medium of claim 16, wherein the ambient sound is a first ambient sound, the sound source is a first sound source, the audio segment is a first audio segment, and the textual data is first textual data (In particular, a head-mountable device can be provided with multiple microphones for capturing audio information (e.g., sounds) from multiple sources that are located in different directions with respect to the head-mountable device. Multiple microphones distributed across the head-mountable device can provide directional audio detection. The head-mountable device can use the data collected by the microphones to provide visual and/or audio outputs to the user. For example, the detected audio inputs can be rendered with visual outputs by providing indicators directing the user to the source of the sound. This can allow the user to correctly and readily identify the location of the source, even when the user is not readily able to hear the sound independently of the head-mountable device. By further example, the detected audio inputs can be rendered with audio outputs that emphasize (e.g., amplify) certain sounds over others to help the user distinguish between different sounds. See paragraph [0016])( Other sources of sounds can also be located within the field-of-view 90 and/or outside the field-of-view 90. As each of the sounds are received by the user, the head-mountable device 100 can provide visual outputs that guide the user's attention to particular sources of sound. See paragraph [0054]), wherein the operations further comprise:
detecting a second ambient sound from the audio data at least partially in parallel with detecting the first ambient sound from the audio data (In particular, a head-mountable device can be provided with multiple microphones for capturing audio information (e.g., sounds) from multiple sources that are located in different directions with respect to the head-mountable device. Multiple microphones distributed across the head-mountable device can provide directional audio detection. The head-mountable device can use the data collected by the microphones to provide visual and/or audio outputs to the user. For example, the detected audio inputs can be rendered with visual outputs by providing indicators directing the user to the source of the sound. This can allow the user to correctly and readily identify the location of the source, even when the user is not readily able to hear the sound independently of the head-mountable device. By further example, the detected audio inputs can be rendered with audio outputs that emphasize (e.g., amplify) certain sounds over others to help the user distinguish between different sounds. See paragraph [0016]), but is silent to
determining that the second ambient sound includes speech based on a second audio segment of the audio data that includes the second ambient sound;
generating second textual data about at least one attribute of a second sound source that generated the speech using the second audio segment, the second textual data including a textual description of a physical characteristic of the second sound source;
initiating conversion of the speech to a textual translation; and
displaying the textual translation and the second textual data at a display position in the display that is based on a location of the second sound source that generated the speech.
McCulloch teaches detecting speech and utilizing beamforming to locate and translate the speech into a text overlay. (Embodiments that relate to converting audio inputs from an environment into text are disclosed. For example, in one disclosed embodiment a speech conversion program receives audio inputs from a microphone array of a head-mounted display device. Image data is captured from the environment, and one or more possible faces are detected from image data. Eye-tracking data is used to determine a target face on which a user is focused. A beamforming technique is applied to at least a portion of the audio inputs to identify target audio inputs that are associated with the target face. The target audio inputs are converted into text that is displayed via a transparent display of the head-mounted display device. See abstract)(See figure 3).
Poore and McCulloch teach of presenting indicators over audible objects and McCulloch teaches that by using a beamforming and audio translation technique text can be located above the speaker, therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the system of Poore with the beamforming and audio translation to text of McCulloch such that a user could visualize responses from multiple individuals.
Regarding claim 22, Poore teaches The method of claim 1, wherein the ambient sound is a first ambient sound classified as ambient noise and the audio segment is a first audio segment (In particular, a head-mountable device can be provided with multiple microphones for capturing audio information (e.g., sounds) from multiple sources that are located in different directions with respect to the head-mountable device. Multiple microphones distributed across the head-mountable device can provide directional audio detection. The head-mountable device can use the data collected by the microphones to provide visual and/or audio outputs to the user. For example, the detected audio inputs can be rendered with visual outputs by providing indicators directing the user to the source of the sound. This can allow the user to correctly and readily identify the location of the source, even when the user is not readily able to hear the sound independently of the head-mountable device. By further example, the detected audio inputs can be rendered with audio outputs that emphasize (e.g., amplify) certain sounds over others to help the user distinguish between different sounds. See paragraph [0016])( Other sources of sounds can also be located within the field-of-view 90 and/or outside the field-of-view 90. As each of the sounds are received by the user, the head-mountable device 100 can provide visual outputs that guide the user's attention to particular sources of sound. See paragraph [0054]), the method further comprising:
detecting a second ambient sound from the audio data that temporally overlaps with the first ambient sound (In particular, a head-mountable device can be provided with multiple microphones for capturing audio information (e.g., sounds) from multiple sources that are located in different directions with respect to the head-mountable device. Multiple microphones distributed across the head-mountable device can provide directional audio detection. The head-mountable device can use the data collected by the microphones to provide visual and/or audio outputs to the user. For example, the detected audio inputs can be rendered with visual outputs by providing indicators directing the user to the source of the sound. This can allow the user to correctly and readily identify the location of the source, even when the user is not readily able to hear the sound independently of the head-mountable device. By further example, the detected audio inputs can be rendered with audio outputs that emphasize (e.g., amplify) certain sounds over others to help the user distinguish between different sounds. See paragraph [0016]), but is silent to wherein the second ambient sound is classified as speech;
extracting a second audio segment of the audio data that includes the second ambient sound;
converting the second audio segment into a textual transcription of the speech;
displaying the textual data at a first location on the display corresponding to the sound source of the ambient noise; and
displaying the textual transcription of the speech at a second location on the display corresponding to a sound source of the speech.
McCulloch teaches detecting speech and utilizing beamforming to locate and translate the speech into a text overlay. (Embodiments that relate to converting audio inputs from an environment into text are disclosed. For example, in one disclosed embodiment a speech conversion program receives audio inputs from a microphone array of a head-mounted display device. Image data is captured from the environment, and one or more possible faces are detected from image data. Eye-tracking data is used to determine a target face on which a user is focused. A beamforming technique is applied to at least a portion of the audio inputs to identify target audio inputs that are associated with the target face. The target audio inputs are converted into text that is displayed via a transparent display of the head-mounted display device. See abstract)(See figure 3).
Poore and McCulloch teach of presenting indicators over audible objects and McCulloch teaches that by using a beamforming and audio translation technique text can be located above the speaker, therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the system of Poore with the beamforming and audio translation to text of McCulloch such that a user could visualize responses from multiple individuals.
Claim(s) 21 is/are rejected under 35 U.S.C. 103 as being unpatentable over Poore et al. (US 2021/0055367)(Hereinafter referred to as Poore) in view of Alter et al. (US 2009/0293012)(Hereinafter referred to as Alter).
Regarding claim 21, Poore teaches The method of claim 1, further comprising:
determining that the location of the sound source is initially outside a field of view of a user of the wearable device (See figure 5);
displaying a directional sound cue on the display indicating a relative direction to turn toward the sound source (See figure 6);
detecting a movement of the wearable device such that the location of the sound source enters the field of view of the user of the wearable device (Transition to inside field of view from outside field of view, transition from figures 5, 6 location to figures 3 and 4), but is silent to
in response to the location of the sound source entering the field of view, removing the directional sound cue from the display and displaying the textual data anchored at a three- dimensional position in the display corresponding to the location of the sound source.
Alter teaches one indicator reticle when the object is within the field of view and an arrow indicating to rotate the device such that the object is within the field of view (FIG. 11 shows a synthetic vision device used to visualize a direct real world view from a user's actual position. A reticle symbol is used to designate an object in the view. In the figure synthetic vision device 210 is pointed toward a house 225 and a tree 230. Images 1110 and 1120 of the house and tree respectively appear on screen 215. Also shown on the screen is a reticle symbol 1130. The symbol is used to designate an object in the view. In this case a user has designated the door by placing the reticle symbol over the image of the door 1110. The device 210 remembers where the reticle is based on its three dimensional position coordinates. See paragraph [0093])( FIG. 12 shows a synthetic vision device used to visualize a direct real world view from a user's actual position. An arrow symbol is used to designate the direction in which the device must be pointed to see the reticle which was visible in the view shown in FIG. 11. In FIG. 12 the device 210 has been turned to the right compared to the view in FIG. 11. In FIG. 12 only the image 1120 of the tree 230 appears on screen. The house 225 is off screen to the left because the device is not pointed at it. However, arrow 1210 indicates the direction in which the device must be rotated to bring reticle symbol 1130 into view. See paragraph [0094]).
Poore and Alter teach of indicating the location of objects with augmentations and Alter teaches that different augmentations can be utilized to indicate that the object is within the view or that the user needs to move such that the object is within the field of view, therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the system of Poore with the plurality of object indicators based on location of the object such that the user could have desired directions as to how to locate the object.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Fragoso et al. (“TranslatAR: A Mobile Augmented Reality Translator”, IEEE. 2010.)(Hereinafter referred to as Fragoso), generally teaches augmented reality translation.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NICHOLAS R WILSON whose telephone number is (571)272-0936. The examiner can normally be reached M-F 7:30-5:00PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kee Tung can be reached at (572)-272-7794. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/NICHOLAS R WILSON/Primary Examiner, Art Unit 2611