DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
Response to Arguments
Applicant’s arguments with respect to claim(s) 1-9, 11-14, and 17-30 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Claim Rejections - 35 USC § 102
The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
Claim(s) 1-4, 6-7, 9, 11-14, 17-19, 25, and 27-30 is/are rejected under 35 U.S.C. 102(a)(1)/(a)(2) as being anticipated by Eubank et al. (US 2021/0035597 A1, previously cited as pertinent, and hereinafter Eubank).
Regarding claim 1, Eubank anticipates “A device comprising: a processor” (see Eubank, ¶ 0035 and 0039, and figure 1, units 1 and 5),
“configured to: perform signal enhancement of an input audio signal to generate a first directional enhanced mono audio signal associated with speech from a first audio source in a soundscape” by teaching a speech and ambient separator that receives an input audio signal from a microphone array and generates a speech signal that is enhanced by spectral shaping or another method to improve the signal-to-noise ratio (SNR) (see Eubank, ¶ 0041-0042 and 0044-0045, and figure 1, units 1-3, 5, 7, and 16);
“perform signal enhancement of the input audio signal to generate a second directional enhanced mono audio signal associated with a second audio source in the soundscape” by teaching that the speech and ambient separator also generates another speech signal from another talker and the speech and ambient separator uses noise suppression operations to improve the SNR for the separated speech and ambient sounds (see Eubank, ¶ 0042 and 0044-0045, and figure 1, units 2-3 and 7); and
“generate a stereo audio signal based at least on the first directional enhanced mono audio signal, the second directional enhanced mono audio signal, and a context signal indicative of an audio context of the first audio source relative to the soundscape” by first teaching a context signal, such as separated ambient sounds, diffuse background sounds, and/or wind noise, because these separated sounds provide audio context of the first audio source relative to the acoustic environment, then teaching that the system determines the sound object’s positions within the environment, and teaching that a receiver device receives the digital audio data and spatial metadata to output a set of mixed signals that the user perceives as being emitted from a particular location within an acoustic space (see Eubank, ¶ 0042-0044, 0058-0060, 0067, 0069-0070, 0074, and 0080, figure 1, units 2-3, 7, and 17-18, and figure 4, units 20-22, 25, and 30).
Regarding claim 2, see the preceding rejection with respect to claim 1 above. Eubank anticipates the “device of claim 1, wherein the context signal is based at least on a background audio signal corresponding to a diffuse noise associated with the soundscape” by teaching a context signal, such as the sound object and sound bed identifier identifies an ambient or diffuse background sound as the sound bed of the acoustic environment (see Eubank, ¶ 0042-0043 and 0058-0060).
Regarding claim 3, see the preceding rejection with respect to claim 1 above. Eubank anticipates the “device of claim 1, wherein the processor is configured to use a neural network to perform the signal enhancement” because the sound source separation uses machine learning, such as neural networks, in order to determine a sound object from its features or characteristics (see Eubank, ¶ 0051-0052, 0054-0055, 0093-0094, and 0110).
Regarding claim 4, see the preceding rejection with respect to claim 1 above. Eubank anticipates the “device of claim 1, wherein the input audio signal is based on microphone output of one or more microphones” because the input audio signal is based on the output of a microphone array (see Eubank, ¶ 0038 and 0042, and figure 1, units 2-3).
Regarding claim 6, see the preceding rejection with respect to claim 1 above. Eubank anticipates the “device of claim 1, wherein the signal enhancement includes at least one of noise suppression, audio zoom, beamforming, dereverberation, bass adjustment, or equalization” because Eubank teaches the signal enhancement includes at least noise suppression and/or beamforming (see Eubank, ¶ 0042, 0045, 0051, and 0053-0054, figure 1, units 5 and 10, and figure 2, unit 71).
Regarding claim 7, see the preceding rejection with respect to claim 1 above. Eubank anticipates the “device of claim 1, wherein the processor is configured to use a neural network to generate the stereo audio signal” by teaching machine learning algorithms, such as neural networks, to better perform sound source separation and therefore generate the stereo audio signal from the separated sounds (see Eubank, ¶ 0051-0052, 0054-0055, 0093-0094, and 0110).
Regarding claim 9, see the preceding rejection with respect to claim 1 above. Eubank anticipates the “device of claim 1, wherein the processor is configured to:
perform signal enhancement of a second input audio signal to generate a third directional enhanced mono audio signal” by teaching that the speech and ambient separator also generates other sound objects, such as directional ambient sounds, and the speech and ambient separator uses noise suppression operations to improve the SNR for the ambient sounds (see Eubank, ¶ 0042 and 0044-0045, and figure 1, units 2-3 and 7) ; and
“generate the stereo audio signal based on the first directional enhanced mono audio signal, the second directional enhanced mono audio signal, and the third directional enhanced mono audio signal” by teaching that the system determines the sound object’s positions within the environment, and teaching that a receiver device receives the digital audio data and spatial metadata to output a set of mixed signals that the user perceives as being emitted from a particular location within an acoustic space (see Eubank, ¶ 0042-0044, 0058-0060, 0067, 0069-0070, 0074, and 0080, figure 1, units 2-3, 7, and 17-18, and figure 4, units 20-22, 25, and 30).
Regarding claim 11, see the preceding rejection with respect to claim 1 above. Eubank anticipates the “device of claim 2, wherein the processor is configured to
apply a delay to the background audio signal to generate a delayed audio signal, wherein the context signal is based on the delayed audio signal” by teaching the beamformer to separate the ambient or background sounds, where the beamformer applies delays to two or more of the microphone signals to steer the beam pattern in the direction of one or more of the background sounds (see Eubank, ¶ 0044-0045 and 0054).
Regarding claim 12, see the preceding rejection with respect to claim 11 above. Eubank anticipates the “device of claim 11, wherein the processor is configured to pan, based on a visual context, the delayed audio signal to generate the context signal” by teaching a camera and object recognition to better identify sound objects and improve beam-steering to separate the sound objects (see Eubank, ¶ 0049, 0051-0054, and 0110, figure 1, units 4-5 and 10, and figure 2, units 4 and 71).
Regarding claim 13, see the preceding rejection with respect to claim 11 above. Eubank anticipates the “device of claim 11, wherein the processor is configured to attenuate the delayed audio signal to generate the context signal” by teaching the beamformer to separate the ambient or background sounds, where the beamformer applies delays to two or more of the microphone signals and sums the signals with associated weights (e.g., attenuation and/or amplification values) to steer the beam pattern in the direction of one or more of the background sounds (see Eubank, ¶ 0044-0045 and 0054).
Regarding claim 14, see the preceding rejection with respect to claim 13 above. Eubank teaches the “device of claim 1, wherein the processor is configured to use a first neural network to perform signal enhancement of the input audio signal to generate a first neural network output that includes the first directional enhanced mono audio signal or the second directional enhanced mono audio signal” by teaching machine learning algorithms, such as neural networks, to better perform sound source separation and therefore generate the directional enhanced mono audio signals (see Eubank, ¶ 0051-0052, 0054-0055, 0093-0094, and 0110).
Regarding claim 17, see the preceding rejection with respect to claim 1 above. Eubank anticipates the “device of claim 1, wherein the second audio source is noise in a direction of the first audio source” by teaching that the sound source separator performs separation of speech and ambient signals by suppressing speech signals in the ambient signal and suppressing ambient signals in the speech signal, which indicates the noise source is in the direction of the first audio source (see Eubank, ¶ 0042-0043).
Regarding claim 18, see the preceding rejection with respect to claim 1 above. Eubank anticipates the “device of claim 1, wherein the second audio source is speech” by teaching that other detected sounds include a different talker (see Eubank, ¶ 0042).
Regarding claim 19, see the preceding rejection with respect to claim 13 above. Eubank anticipates the “device of claim 13, wherein the processor is configured to attenuate the delayed audio signal based on a visual context to generate the context signal” by teaching the beamformer to separate the ambient or background sounds, where the beamformer applies delays to two or more of the microphone signals and sums the signals with associated weights (e.g., attenuation and/or amplification values) to steer the beam pattern in the direction of one or more of the background sounds and by teaching a camera and object recognition to better identify sound objects and improve beam-steering to separate the sound objects (see Eubank, ¶ 0044-0045, 0049, 0051-0054, and 0110, figure 1, units 4-5 and 10, and figure 2, units 4 and 71).
Regarding claim 25, Eubank anticipates “A method comprising:
performing, at a device, signal enhancement of an input audio signal to generate a first directional enhanced mono audio signal associated with speech from a first audio source in a soundscape” by teaching a speech and ambient separator that receives an input audio signal from a microphone array and generates a speech signal that is enhanced by spectral shaping or another method to improve the signal-to-noise ratio (SNR) (see Eubank, ¶ 0041-0042 and 0044-0045, and figure 1, units 1-3, 5, 7, and 16)
“performing, at the device, signal enhancement of the input audio signal to generate a second directional enhanced mono audio signal associated with a second audio source in the soundscape” by teaching that the speech and ambient separator also generates another speech signal from another talker and the speech and ambient separator uses noise suppression operations to improve the SNR for the separated speech and ambient sounds (see Eubank, ¶ 0042 and 0044-0045, and figure 1, units 2-3 and 7)
“generating, at the device, a stereo audio signal based at least on the first directional enhanced mono audio signal, the second directional enhanced mono audio signal, and a context signal indicative of an audio context of the first audio source relative to the soundscape” by first teaching a context signal, such as separated ambient sounds, diffuse background sounds, and/or wind noise, because these separated sounds provide audio context of the first audio source relative to the acoustic environment, then teaching that the system determines the sound object’s positions within the environment, and teaching that a receiver device receives the digital audio data and spatial metadata to output a set of mixed signals that the user perceives as being emitted from a particular location within an acoustic space (see Eubank, ¶ 0042-0044, 0058-0060, 0067, 0069-0070, 0074, and 0080, figure 1, units 2-3, 7, and 17-18, and figure 4, units 20-22, 25, and 30).
Regarding claim 27, Eubank anticipates “A non-transitory computer-readable medium storing instructions that” (see Eubank, ¶ 0035, 0039, and 0112, and figure 1, units 1 and 5), “when executed by one or more processors, cause the one or more processors to:
perform signal enhancement of an input audio signal to generate a first directional enhanced mono audio signal associated with speech from a first audio source in a soundscape” by teaching a speech and ambient separator that receives an input audio signal from a microphone array and generates a speech signal that is enhanced by spectral shaping or another method to improve the signal-to-noise ratio (SNR) (see Eubank, ¶ 0041-0042 and 0044-0045, and figure 1, units 1-3, 5, 7, and 16);
“perform signal enhancement of the input audio signal to generate a second directional enhanced mono audio signal associated with a second audio source in the soundscape” by teaching that the speech and ambient separator also generates another speech signal from another talker and the speech and ambient separator uses noise suppression operations to improve the SNR for the separated speech and ambient sounds (see Eubank, ¶ 0042 and 0044-0045, and figure 1, units 2-3 and 7); and
“generate a stereo audio signal based at least on the first directional enhanced mono audio signal, the second directional enhanced mono audio signal, and a context signal indicative of an audio context of the first audio source relative to the soundscape” by first teaching a context signal, such as separated ambient sounds, diffuse background sounds, and/or wind noise, because these separated sounds provide audio context of the first audio source relative to the acoustic environment, then teaching that the system determines the sound object’s positions within the environment, and teaching that a receiver device receives the digital audio data and spatial metadata to output a set of mixed signals that the user perceives as being emitted from a particular location within an acoustic space (see Eubank, ¶ 0042-0044, 0058-0060, 0067, 0069-0070, 0074, and 0080, figure 1, units 2-3, 7, and 17-18, and figure 4, units 20-22, 25, and 30).
Regarding claim 28, see the preceding rejection with respect to claim 27 above. Eubank anticipates the “non-transitory computer-readable medium of claim 27, wherein the signal enhancement is based at least in part on a configuration setting, a user input, or both” where the signal enhancement, such as noise reduction is configured by the device and/or by teaching user input for indication of spatial rendering of sound objects (see Eubank, ¶ 0042, 0054, and 0111).
Regarding claim 29, Eubank anticipates “An apparatus” (see Eubank, ¶ 0035 and 0039, and figure 1, units 1 and 5) “comprising:
means for performing signal enhancement of an input audio signal to generate a first directional enhanced mono audio signal associated with speech from a first audio source in a soundscape” by teaching a speech and ambient separator that receives an input audio signal from a microphone array and generates a speech signal that is enhanced by spectral shaping or another method to improve the signal-to-noise ratio (SNR) (see Eubank, ¶ 0041-0042 and 0044-0045, and figure 1, units 1-3, 5, 7, and 16);
“means for performing signal enhancement of the input audio signal to generate a second directional enhanced mono audio signal associated with a second audio source in the soundscape” by teaching that the speech and ambient separator also generates another speech signal from another talker and the speech and ambient separator uses noise suppression operations to improve the SNR for the separated speech and ambient sounds (see Eubank, ¶ 0042 and 0044-0045, and figure 1, units 2-3 and 7); and
“means for generating a stereo audio signal based at least on the first directional enhanced mono audio signal, the second directional enhanced mono audio signal, and a context signal indicative of an audio context of the first audio source relative to the soundscape” by first teaching a context signal, such as separated ambient sounds, diffuse background sounds, and/or wind noise, because these separated sounds provide audio context of the first audio source relative to the acoustic environment, then teaching that the system determines the sound object’s positions within the environment, and teaching that a receiver device receives the digital audio data and spatial metadata to output a set of mixed signals that the user perceives as being emitted from a particular location within an acoustic space (see Eubank, ¶ 0042-0044, 0058-0060, 0067, 0069-0070, 0074, and 0080, figure 1, units 2-3, 7, and 17-18, and figure 4, units 20-22, 25, and 30).
Regarding claim 30, see the preceding rejection with respect to claim 29 above. Eubank anticipates the “apparatus of claim 29, wherein the means for performing the signal enhancement and the means for generating the stereo audio signal are integrated into at least one of a smart speaker, a speaker bar, a computer, a tablet, a display device, a television, a gaming console, a music player, a radio, a digital video player, a camera, a navigation device, a vehicle, a headset, an augmented reality headset, a mixed reality headset, a virtual reality headset, an aerial vehicle, a home automation system, a voice-activated device, a wireless speaker and voice activated device, a portable electronic device, a communication device, an internet-of-things (IoT) device, a virtual reality (VR) device, a base station, or a mobile device” by teaching the features are integrated into a head-mounted display (HMD) (see Eubank, ¶ 0035, 0067, and 0069).
Claim Rejections - 35 USC § 103
The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claim(s) 5 is/are rejected under 35 U.S.C. 103 as being unpatentable over Eubank as applied to claim 1 above, and further in view of Li et al., US 2021/0321212 A1 (previously cited and hereinafter Li).
Regarding claim 5, see the preceding rejection with respect to claim 1 above. Eubank teaches the device of claim 1, however does not appear to teach the features “wherein the processor is configured to decode encoded audio data to generate the input audio signal”.
Li teaches that the microphone array comprises digital microphones where the processor would decode the audio data and spatial information from the digital microphones to further process the input audio signal (see Li, ¶ 0077-0078). It would have been obvious to one of ordinary skill in the art at the time of the effective filing date to modify Eubank with the teachings of Li for the purpose of improving audio rendering for virtual conferencing (see Li, ¶ 0086-0087 and figure 10).
Therefore the combination of Eubank and Li makes obvious the “device of claim 1, wherein the processor is configured to decode encoded audio data to generate the input audio signal” because Li makes obvious a microphone array comprising digital microphones where the processor would decode the audio data and spatial information from the digital microphones to further process the input audio signal (see Eubank, ¶ 0038 and 0041-0042 in view of Li, ¶ 0077-0078).
Claim(s) 8 is/are rejected under 35 U.S.C. 103 as being unpatentable over Eubank as applied to claim 1 above, and further in view of Seamans, US 9,967,693 B1 (previously cited).
Regarding claim 8, see the preceding rejection with respect to claim 1 above. Eubank teaches the “device of claim 1, wherein the processor is configured to:
use a first neural network to perform signal enhancement of an input audio signal to generate the first directional enhanced mono audio signal or the second directional enhanced” mono “audio signal” by teaching machine learning algorithms, such as neural networks, to better perform sound source separation (see Eubank, ¶ 0051-0052, 0054-0055, and 0093-0094). However, Eubank does not appear to teach the use of “a second neural network to generate the stereo audio signal”.
Seamans teaches an advanced binaural sound imaging system, where a neural network is used to eliminate or reduce dead zones and/or speaker crosstalk (see Seamans, abstract and column 1, line 52 – column 2, line 7). It would have been obvious to one of ordinary skill in the art at the time of the effective filing date to modify Eubank with the teachings of Seamans for the purpose of eliminate or reduce dead zones and/or speaker crosstalk for reproduction of binaural audio via loudspeakers (see Seamans, abstract and column 1, line 63 – column 2, line 7).
Therefore, the combination of Eubank and Seamans makes obvious the “use a second neural network to generate the stereo audio signal” because Seamans makes it obvious to use a neural network to enhance the mixed binaural output, where a trained neural network performs a first function of separating sound sources and a second function of mapping sound sources into binaural left and right channels, such that the mapping sound sources to the binaural output reads on a mix created by a neural network (see Seamans, column 8, lines 4-37 and figure 6)
Claim(s) 20-23 and 26 is/are rejected under 35 U.S.C. 103 as being unpatentable over Eubank as applied to claims 1-2 and 25 above, and further in view of Sarkar, US 2019/0373395 A1 (previously cited).
Regarding claim 20, see the preceding rejection with respect to claim 2 above. Eubank teaches the device of claim 2, where the audio signal is processed by reverberation filters (see Eubank, ¶ 0080-0081). However, Eubank does not appear to teach a reverberation model for processing the background audio signal.
Sarkar teaches an augmented reality (AR) device that adjusts audio characteristics for AR based on a generated 3D map of a location (see Sarkar, abstract). Sarkar teaches that audio for an AR application includes simulating reverberation characteristics of a room (see Sarkar, ¶ 0003), such that an AR device that modifies the audio signal based on image data captured from an image sensor (see Sarkar, ¶ 0018-0020, 0022-0023, and 0030, and figure 1, units 102 and 108). It would have been obvious to one of ordinary skill in the art at the time of the effective filing date to modify Eubank with the teachings of Sarkar for the purpose of providing an augmented reality where the user can perceive differences between different locations for sounds in the augmented reality (see Eubank, ¶ 0080-0081 in view of Sarkar, ¶ 0014).
Therefore, the combination of Eubank and Sarkar makes obvious the “device of claim 2, wherein the processor is configured to use a reverberation model to process the background audio signal, the first directional enhanced mono audio signal, the second directional enhanced mono audio signal, or a combination thereof, to generate a reverberation signal, wherein the context signal includes the reverberation signal” because Sarkar makes it obvious to generate a reverberation model to process the sound objects, including the background sound signal, such that the second audio signal includes a reverberation signal (see Eubank, ¶ 0081 in view of Sarkar, ¶ 0031-0034).
Regarding claim 21, see the preceding rejection with respect to claim 1 above. Eubank teaches the “device of claim 1, wherein the processor is configured to:
determine, based on image data, a visual context of the input audio signal, the image data representing a visual scene associated with the first audio source” by teaching a camera and object recognition to better identify sound objects and improve beam-steering to separate the sound objects (see Eubank, ¶ 0049, 0051-0054, and 0110, figure 1, units 4-5 and 10, and figure 2, units 4 and 71). However, Eubank does not appear to teach a reverberation model to generate a synthesized reverberation signal according to the visual context.
Sarkar teaches an augmented reality (AR) device that adjusts audio characteristics for AR based on a generated 3D map of a location (see Sarkar, abstract). Sarkar teaches that audio for an AR application includes simulating reverberation characteristics of a room (see Sarkar, ¶ 0003), such that an AR device that modifies the audio signal based on image data captured from an image sensor (see Sarkar, ¶ 0018-0020, 0022-0023, and 0030, and figure 1, units 102 and 108). It would have been obvious to one of ordinary skill in the art at the time of the effective filing date to modify Eubank with the teachings of Sarkar for the purpose of providing an augmented reality where the user can perceive differences between different locations for sounds in the augmented reality (see Eubank, ¶ 0080-0081 in view of Sarkar, ¶ 0014).
Therefore, the combination of Eubank and Sarkar makes obvious the feature to “use a reverberation model to generate a synthesized reverberation signal corresponding to the visual context, wherein the context audio signal further includes the synthesized reverberation signal” because Sarkar makes it obvious to generate a reverberation model to process the sound objects, such that the second audio signal includes a synthesized reverberation signal (see Sarkar, ¶ 0031-0034).
Regarding claim 22, see the preceding rejection with respect to claim 21 above. The combination makes obvious the “device of claim 21, wherein the visual context is based on surfaces of an acoustic environment, room geometry, or both” because the image data captured from an image sensor, used to determine a reverberation model, determines dimensions of the environment and materials associated with the room (see Sarkar, ¶ 0023 and 0029-0030, and figure 1, units 102 and 108).
Regarding claim 23, see the preceding rejection with respect to claim 21 above. The combination makes obvious the “device of claim 21, wherein the image data is based on at least one of camera output, a graphic visual stream, decoded image data, or stored image data” because the image date is from a camera (see Sarkar, ¶ 0020 and figure 1, unit 208).
Regarding claim 26, see the preceding rejection with respect to claim 25 above. Eubank teaches the “method of claim 25, further comprising:
determining a location context based on location data” by teaching a camera and object recognition to better identify sound objects and improve beam-steering to separate the sound objects (see Eubank, ¶ 0049, 0051-0054, and 0110, figure 1, units 4-5 and 10, and figure 2, units 4 and 71). However, Eubank does not appear to teach a reverberation model to generate a synthesized reverberation signal according to the location context.
Sarkar teaches an augmented reality (AR) device that adjusts audio characteristics for AR based on a generated 3D map of a location (see Sarkar, abstract). Sarkar teaches that audio for an AR application includes simulating reverberation characteristics of a room (see Sarkar, ¶ 0003), such that an AR device that modifies the audio signal based on image data captured from an image sensor (see Sarkar, ¶ 0018-0020, 0022-0023, and 0030, and figure 1, units 102 and 108). It would have been obvious to one of ordinary skill in the art at the time of the effective filing date to modify Eubank with the teachings of Sarkar for the purpose of providing an augmented reality where the user can perceive differences between different locations for sounds in the augmented reality (see Eubank, ¶ 0080-0081 in view of Sarkar, ¶ 0014).
Therefore, the combination of Eubank and Sarkar makes obvious the method of claim 25 also comprising “using a reverberation model to generate a synthesized reverberation signal corresponding to the location context, wherein the context signal includes the synthesized reverberation signal” because Sarkar makes it obvious to generate a reverberation model to process the sound objects, such that the second audio signal includes a synthesized reverberation signal (see Sarkar, ¶ 0031-0034).
Claim(s) 24 is/are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Eubank and Sarkar as applied to claim 21 above, and further in view of Adsumilli et al., US 2017/0366896 A1 (previously cited and hereinafter Adsumilli).
Regarding claim 24, see the preceding rejection with respect to claim 21 above. The combination of Eubank and Sarkar makes obvious the device of claim 21, where the device determines a visual context. However the combination does not appear to teach the feature “to determine the visual context based at least in part on performing face detection on the image data”.
Adsumilli teaches a system and method for generating a model of geometric relationships between various audio sources. In particular, Adsumilli teaches detecting a face as a visual object and matching the detected face to an audio source (see Adsumilli, ¶ 0061).
It would have been obvious to one of ordinary skill in the art at the time of the effective filing date to modify the combination of Eubank and Sarkar with the teachings of Adsumilli for the purpose of improving the tracking of audio sources as they move around the environment (see Eubank, ¶ 0049, 0051, and 0110 in view of Adsumilli, ¶ 0002 and 0014-0016).
Therefore, the combination of Eubank, Sarkar, and Adsumilli makes obvious the “device of claim 21, wherein the processor is configured to determine the visual context based at least in part on performing face detection on the image data” where it is obvious to track different people in the system using face detection to improve localization of different speaker’s voices (see Eubank, ¶ 0049, 0051, and 0110 in view of Adsumilli, ¶ 0061).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Wung et al., US 2019/0172476 A1 (previously cited and hereinafter Wung), discloses a deep learning driven multi-channel filtering for speech enhancement (see Wung, abstract and figures 1 and 4-8);
Hantrakul et al., US 2023/0154451 A1 (previously cited and hereinafter Hantrakul), discloses a differential wavetable synthesizer that uses a generative machine learning model to synthesize sounds (see Hantrakul, abstract and figures 4-6); and
Nesta et al. (US 2020/0184985 A1 and hereinafter Nesta), teaches a multi-stream target-speech detection and channel fusion audio processing system, which uses N different speech enhancement modules each using a different enhancement and providing a different number of output streams (see Nesta, abstract, ¶ 0020-0021, 0035-0038, 0040-0042, 0044, and 0047-0049, figure 1, units 102 and 110, figure 2, unit 202f, figure 3, units 300, 305a-305n, 320, 322, and 324, and figure 4, unit 400 and 405a-405n).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Daniel R Sellers whose telephone number is (571)272-7528. The examiner can normally be reached Mon - Fri 10:00-4:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vivian C Chin can be reached at (571)272-7848. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Daniel R Sellers/ Primary Examiner, Art Unit 2694