DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
Claim Rejections - 35 USC § 102
The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
Claim(s) 1-6, 9, 11-14, 17-19, 25, and 27-30 is/are rejected under 35 U.S.C. 102(a)(1)/(a)(2) as being anticipated by Li et al., US 2021/0321212 A1 (previously cited and hereinafter Li).
Regarding claim 1, Li anticipates “A device comprising: a processor” (see Li, ¶ 0076-0077, 0097, 0114, and 0117, figure 8, unit 800, and figure 14, units 1400 and 1402),
“configured to: perform signal enhancement of an input audio signal to generate an enhanced mono audio signal” by teaching that the processor of the computer system is configured by instructions (e.g., software) (see Li, ¶ 0006, 0076, 0114, and 0117-0118, figure 8, unit 800, and figure 14, units 1400, 1402, 1416, 1424, and 1426), to perform a signal enhancement of an input audio signal, such as processing the input audio signal received from a microphone array and separating the sounds coming from different directions into multiple audio tracks that each correspond, one to one, with a sound coming from a specific direction (see Li, ¶ 0077-0079, and figure 8, units 802, 806, and 808); Li teaches the generated enhanced mono audio signal, such as the output track(s) of the acoustic beamformer is an enhanced single channel (i.e., monophonic or mono) audio signal, such that each output track corresponds to sound from one direction or source, and the generated enhanced mono audio signal is the output of the acoustic beamformer because the output mono audio signals have improved signal to noise ratio (SNR) compared to the input from the microphone array (see Li, ¶ 0079 and 0084-0085, figure 8, unit 806, and figures 9A-9B), and additionally, Li teaches the audio tracks output from the acoustic beamformer are processed by a noise reduction unit to reduce the background noise in the mono audio signals (see Li, ¶ 0079 and figure 8, unit 808);
“generate at least one directional audio signal based on the input audio signal” by teaching the multiple track output of the acoustic beamformer, which is based on the input audio signal from the microphone array (see Li, ¶ 0077-0079 and 0085, figure 8, units 802 and 806, and figure 9B, units 904 and 906), and by teaching the tracks are used to generate output to groups of loudspeakers or to generate the output of HRTF filters based on the spatial information of the separated sound sources (see Li, ¶ 0079-0081 and 0085); and
“mix the enhanced mono audio signal and a second audio signal to generate a stereo audio signal, the second audio signal based on the at least one directional audio signal” by teaching that the system mixes the enhanced mono audio signal, such as mixing, with a mixer, one of the multiple output tracks corresponding to one of the separated directional sound sources with the other output tracks of the acoustic beamformer and/or HRTF filters to generate a binaural output or stereo output (see Li, ¶ 0038-0044 and 0081, figures 2A-B, and figure 8, units 800, 810, and 812), where the binaural output or stereo output corresponds to the generated stereo audio signal, the at least one of the other output tracks of the noise reduction unit and/or the acoustic echo canceller (AEC) correspond to a second audio signal based on at least one directional audio signal, and any single one of the output tracks of the acoustic beamformer corresponds to the enhanced mono audio signal.
Regarding claim 2, see the preceding rejection with respect to claim 1 above. Li anticipates the “device of claim 1, wherein the second audio signal is associated with a context of the input audio signal” by teaching these features in use with videoconference software to determine which attendee is speaking and process the input audio signals based on the location of the speaker (i.e., sound source direction) to output binaural output to another attendee of the videoconference (see Li, ¶ 0086-0088 and 0090-0091, figure 8, unit 802, figure 9B, unit 904, and figure 10).
Regarding claim 3, see the preceding rejection with respect to claim 1 above. Li anticipates the “device of claim 1, wherein the processor is configured to use a neural network to perform the signal enhancement” by teaching these features in use with videoconference software to determine which attendee is speaking and process the input audio signals based on the location of the speaker (i.e., sound source direction) to output binaural output to another attendee of the videoconference (see Li, ¶ 0086-0088 and 0090-0091, figure 8, unit 802, figure 9B, unit 904, and figure 10) and by teaching that the determination of the speaker is performed by a trained neural network (e.g., NN, CNN, DNN, RNN, etc.) (see Li, ¶ 0090 and 0052-0053, and figure 10, unit 1016).
Regarding claim 4, see the preceding rejection with respect to claim 1 above. Li anticipates the “device of claim 1, wherein the input audio signal is based on microphone output of one or more microphones” because the input audio signal is based on the output of a microphone array (see Li, ¶ 0077-0078 and figure 13, step ).
Regarding claim 5, see the preceding rejection with respect to claim 1 above. Li anticipates the “device of claim 1, wherein the processor is configured to decode encoded audio data to generate the input audio signal” because Li teaches that the microphone array comprises digital microphones where the processor would decode the audio data and spatial information from the digital microphones to further process the input audio signal (see Li, ¶ 0077-0078).
Regarding claim 6, see the preceding rejection with respect to claim 1 above. Li anticipates the “device of claim 1, wherein the signal enhancement includes at least one of noise suppression, audio zoom, beamforming, dereverberation, bass adjustment, or equalization” because Li teaches the signal enhancement includes at least noise suppression and/or beamforming (see Li, ¶ 0079), and additionally teaches equalization (see Li, claims 1 and 12).
Regarding claim 9, see the preceding rejection with respect to claim 1 above. Li anticipates the “device of claim 1, wherein the processor is configured to:
perform signal enhancement of a second input audio signal to generate a second enhanced mono audio signal” by teaching that the performed signal enhancement of an input audio signal generates multiple audio tracks that each correspond, one to one, with a sound coming from a specific direction (see Li, ¶ 0077-0079, and figure 8, units 802, 806, and 808), and therefore, Li teaches a second generated enhanced mono audio signal, such as the output track corresponding to a different direction or source from the first enhanced mono audio signal (see Li, ¶ 0079 and 0084-0085, figure 8, unit 806, and figures 9A-9B), and additionally, Li teaches the audio tracks output from the acoustic beamformer are processed by a noise reduction unit to reduce the background noise in the mono audio signals (see Li, ¶ 0079 and figure 8, unit 808); and
“generate the stereo audio signal based on mixing the enhanced mono audio signal, the second audio signal, and a third audio signal, the third audio signal based on the second enhanced mono audio signal” by teaching that the system mixes the enhanced mono audio signal, such as mixing, with a mixer, one of the multiple output tracks corresponding to one of the separated directional sound sources with the other output tracks of the acoustic beamformer and/or HRTF filters to generate a binaural output or stereo output (see Li, ¶ 0038-0044 and 0081, figures 2A-B, and figure 8, units 800, 810, and 812), where the binaural output or stereo output corresponds to the generated stereo audio signal, the at least one of the other output tracks of the noise reduction unit and/or the acoustic echo canceller (AEC) correspond to a second audio signal based on at least one directional audio signal, and any single one of the output tracks of the acoustic beamformer corresponds to the enhanced mono audio signal.
Regarding claim 11, see the preceding rejection with respect to claim 1 above. Li teaches the generation of binaural output using head related transfer function (HRTF) filters, where a pair of HRTF filters are convolved with a sound track (e.g., a sound source corresponds to a monophonic, or single, channel) to generate binaural signals for left and right ears (i.e., a monophonic sound track is convolved with one filter to generate the left ear output and the same track is convolved with the other filter to generate the right ear output) (see Li, ¶ 0039-0041). Li also teaches that HRTF filters encode how a human listener perceives sound arriving from a specific 3D location, and the listener estimates, or perceives, the location of a sound source based on binaural cues, such as interaural time differences and interaural intensity differences, where a sound source arriving from a specific direction with respect to the listener has varying time differences of arrival at each ear (i.e., the sound source is delayed with respect to one ear or the other depending on the location of said sound source) (see Li, ¶ 0039-0040).
Therefore, Li anticipates the “device of claim 1, wherein the processor is configured to apply a delay to a directional audio signal of the at least one directional audio signal to generate a delayed audio signal, wherein the second audio signal is based on the delayed audio signal” because Li teaches using HRTF filters to generate the binaural output, where the second audio signal is delayed depending on the selected HRTF filter pair, which are characterized by a delay to create the time difference of arrival at each ear for a specific location associated with a sound source (see Li, ¶ 0039-0041 and 0079-0081, figure 2B, figure 8, units 800, 810 and 812).
Regarding claim 12, see the preceding rejection with respect to claim 11 above. Li anticipates the “device of claim 11, wherein the processor is configured to pan, based on a visual context, the delayed audio signal to generate the second audio signal” because Li teaches using HRTF filters to generate the binaural output, Li teaches that the second audio signal is attenuated or amplified with respect to the first audio signal depending on the selected HRTF filter pair, which are characterized by an interaural intensity difference to create the sound pressure difference heard at each ear for a specific location associated with a sound source (see Li, ¶ 0039-0040), and Li teaches that a user can, for example, use a touchscreen interface of a smartphone to configure a sound source location, such that a sound source location is panned in amplitude between left and right binaural outputs based on the visual touchscreen interface (see Li, ¶ 0032 and figures 7A and 7D).
Regarding claim 13, see the preceding rejection with respect to claim 1 above. Li anticipates the “device of claim 1, wherein the processor is configured to use a first neural network to process the input audio signal to generate a first neural network output that includes the at least one directional audio signal” because Li teaches these features in use with videoconference software to determine which attendee is speaking and process the input audio signals based on the location of the speaker (i.e., sound source direction) to output binaural output to another attendee of the videoconference (see Li, ¶ 0086-0088 and 0090-0091, figure 8, unit 802, figure 9B, unit 904, and figure 10) and by teaching that the determination of the speaker is performed by a trained neural network (e.g., NN, CNN, DNN, RNN, etc.) (see Li, ¶ 0090 and 0052-0053, and figure 10, unit 1016), such that at least one directional audio signal is generated using the trained neural network and beamforming.
Regarding claim 14, see the preceding rejection with respect to claim 13 above. Li anticipates the “device of claim 13, wherein the processor is configured to use a second neural network to perform signal enhancement of the input audio signal to generate a second neural network output that includes the enhanced mono audio signal” because Li teaches that sound separation is also performed by a trained neural network, such as a trained neural network to separate different types of audio from an input multi-channel signal (see Li, ¶ 0051-0063 and claims 1-2, and figures 4-5E).
Regarding claim 17, see the preceding rejection with respect to claim 1 above. Li teaches separating sound into various categories including an unidentified sound type or category, which reads on a background signal (see Li, ¶ 0032).
Therefore, Li anticipates the “device of claim 1, wherein the processor is configured to generate a background audio signal from an input audio signal, wherein the second audio signal is based at least in part on the background audio signal” (see Li, ¶ 0032).
Regarding claim 18, see the preceding rejection with respect to claim 17 above. Li anticipates the “device of claim 17, wherein the processor is configured to:
apply a delay to the background audio signal to generate a delayed background audio signal” because Li teaches using HRTF filters to generate the binaural output, where the background signal is delayed depending on the selected HRTF filter pair, which are characterized by a delay to create the time difference of arrival at each ear for a specific location associated with a sound source (see Li, ¶ 0032 and 0039-0040); and
“attenuate the delayed background audio signal to generate the second audio signal” because Li teaches that a delayed background signal is attenuated (or amplified) with respect to the first audio signal depending on the selected HRTF filter pair, which are characterized by an interaural intensity difference to create the sound pressure difference heard at each ear for a specific location associated with a sound source (see Li, ¶ 0032 and 0039-0040).
Regarding claim 19, see the preceding rejection with respect to claim 18 above. Li anticipates the “device of claim 18, wherein the processor is configured to attenuate the delayed background audio signal based on a visual context to generate the second audio signal” because Li teaches that a delayed background signal is attenuated (or amplified) with respect to the first audio signal depending on the selected HRTF filter pair, which is selected via a touchscreen interface of a smartphone to configure a sound source location (see Li, ¶ 0032 and figures 7A and 7D).
Regarding claim 25, see the preceding rejection with respect to claim 1 above. Li anticipates the device of claim 1, and likewise anticipates:
“A method comprising:
performing, at a device, signal enhancement of an input audio signal to generate an enhanced mono audio signal” by teaching that the computer system is configured by instructions (e.g., software) (see Li, ¶ 0006, 0076, 0114, and 0117-0118, figure 8, unit 800, and figure 14, units 1400, 1402, 1416, 1424, and 1426), to perform a signal enhancement of an input audio signal, such as processing the input audio signal received from a microphone array and separating the sounds coming from different directions into multiple audio tracks that each correspond, one to one, with a sound coming from a specific direction (see Li, ¶ 0077-0079, and figure 8, units 802, 806, and 808); Li teaches the generated enhanced mono audio signal, such as the output track(s) of the acoustic beamformer is an enhanced single channel (i.e., monophonic or mono) audio signal, such that each output track corresponds to sound from one direction or source, and the generated enhanced mono audio signal is the output of the acoustic beamformer because the output mono audio signals have improved signal to noise ratio (SNR) compared to the input from the microphone array (see Li, ¶ 0079 and 0084-0085, figure 8, unit 806, and figures 9A-9B), and additionally, Li teaches the audio tracks output from the acoustic beamformer are processed by a noise reduction unit to reduce the background noise in the mono audio signals (see Li, ¶ 0079 and figure 8, unit 808);
“generate, at the device, at least one directional audio signal based on the input audio signal” by teaching the multiple track output of the acoustic beamformer, which is based on the input audio signal from the microphone array (see Li, ¶ 0077-0079 and 0085, figure 8, units 802 and 806, and figure 9B, units 904 and 906), and by teaching the tracks are used to generate output to groups of loudspeakers or to generate the output of HRTF filters based on the spatial information of the separated sound sources (see Li, ¶ 0079-0081 and 0085); and
“mixing, at the device, the enhanced mono audio signal and a second audio signal to generate a stereo audio signal, the second audio signal based on the at least one directional audio signal” by teaching that the system mixes the enhanced mono audio signal, such as mixing, with a mixer, one of the multiple output tracks corresponding to one of the separated directional sound sources with the other output tracks of the acoustic beamformer and/or HRTF filters to generate a binaural output or stereo output (see Li, ¶ 0038-0044 and 0081, figures 2A-B, and figure 8, units 800, 810, and 812), where the binaural output or stereo output corresponds to the generated stereo audio signal, the at least one of the other output tracks of the noise reduction unit and/or the acoustic echo canceller (AEC) correspond to a second audio signal based on at least one directional audio signal, and any single one of the output tracks of the acoustic beamformer corresponds to the enhanced mono audio signal.
Regarding claim 27, see the preceding rejection with respect to claim 1 above. Li anticipates the device of claim 1, and likewise anticipates:
“A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to:
perform signal enhancement of an input audio signal to generate an enhanced mono audio signal” by teaching that the computer system is configured by instructions (e.g., software) (see Li, ¶ 0006, 0076, 0114, and 0117-0118, figure 8, unit 800, and figure 14, units 1400, 1402, 1416, 1424, and 1426), to perform a signal enhancement of an input audio signal, such as processing the input audio signal received from a microphone array and separating the sounds coming from different directions into multiple audio tracks that each correspond, one to one, with a sound coming from a specific direction (see Li, ¶ 0077-0079, and figure 8, units 802, 806, and 808); Li teaches the generated enhanced mono audio signal, such as the output track(s) of the acoustic beamformer is an enhanced single channel (i.e., monophonic or mono) audio signal, such that each output track corresponds to sound from one direction or source, and the generated enhanced mono audio signal is the output of the acoustic beamformer because the output mono audio signals have improved signal to noise ratio (SNR) compared to the input from the microphone array (see Li, ¶ 0079 and 0084-0085, figure 8, unit 806, and figures 9A-9B), and additionally, Li teaches the audio tracks output from the acoustic beamformer are processed by a noise reduction unit to reduce the background noise in the mono audio signals (see Li, ¶ 0079 and figure 8, unit 808);
“generate at least one directional audio signal based on the input audio signal” by teaching the multiple track output of the acoustic beamformer, which is based on the input audio signal from the microphone array (see Li, ¶ 0077-0079 and 0085, figure 8, units 802 and 806, and figure 9B, units 904 and 906), and by teaching the tracks are used to generate output to groups of loudspeakers or to generate the output of HRTF filters based on the spatial information of the separated sound sources (see Li, ¶ 0079-0081 and 0085); and
“mix the enhanced mono audio signal and a second audio signal to generate a stereo audio signal, the second audio signal based on the at least one directional audio signal” by teaching that the system mixes the enhanced mono audio signal, such as mixing, with a mixer, one of the multiple output tracks corresponding to one of the separated directional sound sources with the other output tracks of the acoustic beamformer and/or HRTF filters to generate a binaural output or stereo output (see Li, ¶ 0038-0044 and 0081, figures 2A-B, and figure 8, units 800, 810, and 812), where the binaural output or stereo output corresponds to the generated stereo audio signal, the at least one of the other output tracks of the noise reduction unit and/or the acoustic echo canceller (AEC) correspond to a second audio signal based on at least one directional audio signal, and any single one of the output tracks of the acoustic beamformer corresponds to the enhanced mono audio signal.
Regarding claim 28, see the preceding rejection with respect to claim 27 above. Li anticipates the “non-transitory computer-readable medium of claim 27, wherein the signal enhancement is based at least in part on a configuration setting, a user input, or both” where the signal enhancement is based on at least a configuration setting (see Li, ¶ 0081 and figure 8, unit 814).
Regarding claim 29, see the preceding rejection with respect to claim 1 above. Li anticipates the device of claim 1, and likewise anticipates:
“An apparatus comprising:
means for performing signal enhancement of an input audio signal to generate an enhanced mono audio signal” by teaching a microphone array for performing a signal enhancement of an input audio signal by separating an input audio signal into different channels and processing the different channels with noise reduction and/or acoustic echo cancellation, such as an acoustic beamformer receiving an input audio signal from a 3D microphone array, separating different sound sources from different directions, and processing the separated sound source tracks with a noise reduction unit to generate at least one enhanced single channel audio signal (see Li, ¶ 0006, 0076-0079, 0084-0085, 0090, and 0097, figure 8, units 800, 802, and 806, figure 9A, and figure 9B);
“means for generating at least one directional audio signal based on the input audio signal” by teaching the acoustic beamformer that generates the multiple track output, which is based on the input audio signal from the microphone array (see Li, ¶ 0077-0079 and 0085, figure 8, units 802 and 806, and figure 9B, units 904 and 906), and by teaching the multiple tracks are used to generate output to groups of loudspeakers or to generate the output of HRTF filters based on the spatial information of the separated sound sources (see Li, ¶ 0079-0081 and 0085); and
“means for mixing the enhanced mono audio signal and a second audio signal to generate a stereo audio signal, the second audio signal based on the at least one directional audio signal” by teaching that the system mixes the enhanced mono audio signal, such as mixing, with a mixer, one of the multiple output tracks corresponding to one of the separated directional sound sources with the other output tracks of the acoustic beamformer and/or HRTF filters to generate a binaural output or stereo output (see Li, ¶ 0038-0044 and 0081, figures 2A-B, and figure 8, units 800, 810, and 812), where the binaural output or stereo output corresponds to the generated stereo audio signal, the at least one of the other output tracks of the noise reduction unit and/or the acoustic echo canceller (AEC) correspond to a second audio signal based on at least one directional audio signal, and any single one of the output tracks of the acoustic beamformer corresponds to the enhanced mono audio signal.
Regarding claim 30, see the preceding rejection with respect to claim 29 above. Li anticipates the “apparatus of claim 29, wherein the means for performing the signal enhancement and the means for mixing the enhanced mono audio signal and the second audio signal are integrated into at least one of a smart speaker, a speaker bar, a computer, a tablet, a display device, a television, a gaming console, a music player, a radio, a digital video player, a camera, a navigation device, a vehicle, a headset, an augmented reality headset, a mixed reality headset, a virtual reality headset, an aerial vehicle, a home automation system, a voice-activated device, a wireless speaker and voice activated device, a portable electronic device, a communication device, an internet-of-things (IoT) device, a virtual reality (VR) device, a base station, or a mobile device” by teaching that the means for performing the signal enhancement and the means for mixing (i.e., the 3D microphone system with the acoustic beamformer and mixer) are integrated into at least one of a computer (see Li, ¶ 0076-0079, 0081, 0097, 0114, and 0117-0118, figure 8, units 800, 802, and 806, and figure 14, units 1400, 1402, 1404, 1406, and 1426), and additional teaches the features integrated into other similar devices (see Li, ¶ 0115-0116).
Claim Rejections - 35 USC § 103
The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claim(s) 7 and 8 is/are rejected under 35 U.S.C. 103 as being unpatentable over Li as applied to claim 1 above, and further in view of Seamans, US 9,967,693 B1 (previously cited).
Regarding claim 7, see the preceding rejection with respect to claim 1 above. Li anticipates the device of claim 1, where the device is configured to use a neural network to separate different sound sources (see Li, ¶ 0046-0053 and 0090, and figure 3). However, Li does not appear to teach that “the processor is configured to use a neural network to mix the enhanced mono audio signal and the second audio signal to generate the stereo audio signal”.
Seamans teaches an advanced binaural sound imaging system, where a neural network is used to eliminate or reduce dead zones and/or speaker crosstalk (see Seamans, abstract and column 1, line 52 – column 2, line 7). It would have been obvious to one of ordinary skill in the art at the time of the effective filing date to modify Li with the teachings of Seamans for the purpose of eliminate or reduce dead zones and/or speaker crosstalk for reproduction of binaural audio via loudspeakers (see Seamans, abstract and column 1, line 63 – column 2, line 7).
Therefore, the combination of Li and Seamans makes obvious the “device of claim 1, wherein the processor is configured to use a neural network to mix the enhanced mono audio signal and the second audio signal to generate the stereo audio signal” because Seamans makes it obvious to use a neural network to enhance the mixed binaural output, where a trained neural network performs a first function of separating sound sources and a second function of mapping sound sources into binaural left and right channels, such that the mapping sound sources to the binaural output reads on a mix created by a neural network (see Seamans, column 8, lines 4-37 and figure 6).
Regarding claim 8, see the preceding rejection with respect to claim 1 above. Li anticipates the “device of claim 1, wherein the processor is configured to:
use a first neural network to perform signal enhancement of an input audio signal to generate the enhanced mono audio signal” because the device is configured to use a neural network to separate different sound sources (see Li, ¶ 0046-0053 and 0090, and figure 3).
However, Li does not appear to teach that “a second neural network to mix the enhanced mono audio signal and the second audio signal”.
Seamans teaches an advanced binaural sound imaging system, where a neural network is used to eliminate or reduce dead zones and/or speaker crosstalk (see Seamans, abstract and column 1, line 52 – column 2, line 7). It would have been obvious to one of ordinary skill in the art at the time of the effective filing date to modify Li with the teachings of Seamans for the purpose of eliminate or reduce dead zones and/or speaker crosstalk for reproduction of binaural audio via loudspeakers (see Seamans, abstract and column 1, line 63 – column 2, line 7).
Therefore, the combination of Li and Seamans makes obvious the “device of claim 1, wherein the processor is configured to use a neural network to mix the enhanced mono audio signal and the second audio signal to generate the stereo audio signal” because Seamans makes it obvious to use a neural network to enhance the mixed binaural output, where a trained neural network performs a first function of separating sound sources and a second function of mapping sound sources into binaural left and right channels, such that the mapping sound sources to the binaural output makes obvious a second neural network to create the mixed binaural output (see Seamans, column 8, lines 4-37 and figure 6).
Claim(s) 20-23 and 26 is/are rejected under 35 U.S.C. 103 as being unpatentable over Li as applied to claims 1 and 17 above, and further in view of Sarkar, US 2019/0373395 A1 (previously cited).
Regarding claim 20, see the preceding rejection with respect to claim 17 above. Li anticipates the device of claim 17 where the device separates sound sources based on their direction with respect to the 3D microphone, and the binaural output is based on the actual spatial locations of each separated sound source, such as a second audio source arriving from a different direction of other audio sources (see Li, ¶ 0076-0079 and 0081, and figures 8 and 9A-9B).
However, Li does not appear to teach a reverberation model to process a background audio signal.
Sarkar teaches an augmented reality (AR) device that adjusts audio characteristics for AR based on a generated 3D map of a location (see Sarkar, abstract). Sarkar teaches that audio for an AR application includes simulating reverberation characteristics of a room (see Sarkar, ¶ 0003), such that an AR device that modifies the audio signal based on image data captured from an image sensor (see Sarkar, ¶ 0018-0020, 0022-0023, and 0030, and figure 1, units 102 and 108). It would have been obvious to one of ordinary skill in the art at the time of the effective filing date to modify Li with the teachings of Sarkar for the purpose of providing an augmented reality where the user can perceive differences between different locations for sounds in the augmented reality (see Li, ¶ 0021-0022 in view of Sarkar, ¶ 0014).
Therefore, the combination of Li and Sarkar makes obvious the “device of claim 17, wherein the processor is configured to use a reverberation model to process the background audio signal, the at least one directional audio signal, or a combination thereof, to generate a reverberation signal, wherein the second audio signal includes the reverberation signal” because Sarkar makes it obvious to generate a reverberation model to process the sound objects, including the background sound signal, such that the second audio signal includes a reverberation signal (see Sarkar, ¶ 0031-0034).
Regarding claim 21, see the preceding rejection with respect to claim 1 above. Li anticipates the device of claim 1, but does not appear to teach a reverberation model to process a background audio signal.
Sarkar teaches an augmented reality (AR) device that adjusts audio characteristics for AR based on a generated 3D map of a location (see Sarkar, abstract). Sarkar teaches that audio for an AR application includes simulating reverberation characteristics of a room (see Sarkar, ¶ 0003), such that an AR device that modifies the audio signal based on image data captured from an image sensor (see Sarkar, ¶ 0018-0020, 0022-0023, and 0030, and figure 1, units 102 and 108). It would have been obvious to one of ordinary skill in the art at the time of the effective filing date to modify Li with the teachings of Sarkar for the purpose of providing an augmented reality where the user can perceive differences between different locations for sounds in the augmented reality (see Li, ¶ 0021-0022 in view of Sarkar, ¶ 0014).
Therefore, the combination of Li and Sarkar makes obvious the “device of claim 1, wherein the processor is configured to:
determine, based on image data, a visual context of the input audio signal, the image data representing a visual scene associated with an audio source of the input audio signal” where image data captured from an image sensor is used to determine a reverberation model (see Sarkar, ¶ 0018-0020, 0022-0023, and 0030, and figure 1, units 102 and 108); and
“use a reverberation model to generate a synthesized reverberation signal corresponding to the visual context, wherein the second audio signal includes the synthesized reverberation signal” because Sarkar makes it obvious to generate a reverberation model to process the sound objects, such that the second audio signal includes a synthesized reverberation signal (see Sarkar, ¶ 0031-0034).
Regarding claim 22, see the preceding rejection with respect to claim 21 above. The combination makes obvious the “device of claim 21, wherein the visual context is based on surfaces of an acoustic environment, room geometry, or both” because the image data captured from an image sensor, used to determine a reverberation model, determines dimensions of the environment and materials associated with the room (see Sarkar, ¶ 0023 and 0029-0030, and figure 1, units 102 and 108).
Regarding claim 23, see the preceding rejection with respect to claim 21 above. The combination makes obvious the “device of claim 21, wherein the image data is based on at least one of camera output, a graphic visual stream, decoded image data, or stored image data” because the image date is from a camera (see Sarkar, ¶ 0020 and figure 1, unit 208).
Regarding claim 26, see the preceding rejection with respect to claim 25 above. Li anticipates the “method of claim 25, further comprising:
determining a location context based on location data” by separating sounds coming from different directions with an acoustic beamformer and a mixer outputs binaural output based on the actual spatial locations of the sound sources, such that the location context is based on the look directions of the acoustic beamformer (see Li, ¶ 0041 and 0079-0081, figure 2B, figure 8, units 800, 810 and 812).
Li does not appear to teach a reverberation model to process a background audio signal.
Sarkar teaches an augmented reality (AR) device that adjusts audio characteristics for AR based on a generated 3D map of a location (see Sarkar, abstract). Sarkar teaches that audio for an AR application includes simulating reverberation characteristics of a room (see Sarkar, ¶ 0003), such that an AR device that modifies the audio signal based on image data captured from an image sensor (see Sarkar, ¶ 0018-0020, 0022-0023, and 0030, and figure 1, units 102 and 108). It would have been obvious to one of ordinary skill in the art at the time of the effective filing date to modify Li with the teachings of Sarkar for the purpose of providing an augmented reality where the user can perceive differences between different locations for sounds in the augmented reality (see Li, ¶ 0021-0022 in view of Sarkar, ¶ 0014).
Therefore, the combination of Li and Sarkar makes obvious the method of claim 25 also comprising “using a reverberation model to generate a synthesized reverberation signal corresponding to the location context, wherein the second audio signal includes the synthesized reverberation signal” because Sarkar makes it obvious to generate a reverberation model to process the sound objects, such that the second audio signal includes a synthesized reverberation signal (see Sarkar, ¶ 0031-0034).
Claim(s) 24 is/are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Li and Sarkar as applied to claim 21 above, and further in view of Adsumilli et al., US 2017/0366896 A1 (previously cited and hereinafter Adsumilli).
Regarding claim 24, see the preceding rejection with respect to claim 21 above. The combination of Li and Sarkar makes obvious the device of claim 21, where the device determines a visual context. However the combination does not appear to teach the feature “to determine the visual context based at least in part on performing face detection on the image data”
Adsumilli teaches a system and method for generating a model of geometric relationships between various audio sources. In particular, Adsumilli teaches detecting a face as a visual object and matching the detected face to an audio source (see Adsumilli, ¶ 0061).
It would have been obvious to one of ordinary skill in the art at the time of the effective filing date to modify the combination of Li and Sarkar with the teachings of Adsumilli for the purpose of improving the tracking of audio sources as they move around the environment (see Li, ¶ 0086-0090 in view of Adsumilli, ¶ 0002 and 0014-0016).
Therefore, the combination of Li, Sarkar, and Adsumilli makes obvious the “device of claim 21, wherein the processor is configured to determine the visual context based at least in part on performing face detection on the image data” where it is obvious to track different people in the system using face detection to improve localization of different speaker’s voices (see Li, ¶ 0090-0091 in view of Adsumilli, ¶ 0061).
Response to Arguments
Applicant’s arguments with respect to the 35 U.S.C. 102(a)(2) rejections of claim(s) 1-6, 9, 25, and 27-30 over Morsy (US 2023/0335091 A1) have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Applicant's arguments filed 06/25/2024 with respect to the 35 U.S.C. 102(a)(1)/(a)(2) rejections of claim(s) 1, 11-14, and 17-19 over Li (US 2021/0321212 A1) have been fully considered but they are not persuasive. Examiner Li teaches the features as stated above, where Li teaches features that anticipates the broadest reasonable interpretation (BRI) of the claim language. For instance, the instant specification discloses that an “enhanced mono audio signal” comprises one of a noise suppressed speech signal, an audio zoomed signal, a beamformed signal, a dereverberated signal (i.e., a signal where reverberation, or echoes, are partially or completely removed), a source separated signal, a bass adjusted signal, and an equalized signal, (as disclosed by instant specification, pp. 8-10, ¶ 0053-0059). The BRI further encompasses the plain meaning of an “enhanced mono” audio signal, where enhanced refers to a desired change in the audio signal, and mono refers to a monophonic audio signal, such as a single audio track, channel, and/or signal. Li teaches these features as shown above with respect to the 35 USC 102 rejections of claims 1-6, 9, 11-19, 25, and 27-30.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Wung et al., US 2019/0172476 A1 (previously cited and hereinafter Wung), discloses a deep learning driven multi-channel filtering for speech enhancement (see Wung, abstract and figures 1 and 4-8);
Eubank et al., US 2021/0035597 A1 (previously cited and hereinafter Eubank), discloses an electronic device that performs bandwidth-reduction operations for transmitting less data to another electronic device, where the electronic device processes audio, obtained from a microphone array, to separate a speech signal from one or more ambient signals (see Eubank, abstract and ¶ 0002), where the controller (i.e., processor) performs signal enhancement of the input microphone array signals to generate a separated voice signal (i.e., an enhanced mono audio signal) (see Eubank, ¶ 0038-0039, 0041-0042 and figure 1, units 1-3, 5, 7, and 16) and the mixing is performed at a receiving device and/or the same device (see Eubank, ¶ 0067 and figure 4, unit 30); and
Hantrakul et al., US 2023/0154451 A1 (previously cited and hereinafter Hantrakul), discloses a differential wavetable synthesizer that uses a generative machine learning model to synthesize sounds (see Hantrakul, abstract and figures 4-6).
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Daniel R Sellers whose telephone number is (571)272-7528. The examiner can normally be reached Mon - Fri 10:00-4:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Fan S Tsang can be reached on (571)272-7547. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Daniel R Sellers/Primary Examiner, Art Unit 2694