DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file
Information Disclosure Statement
No IDS
Drawings
The drawings submitted on 8/6/2024 have been considered and accepted.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 1, 3, 4, and 5, is/are rejected under 35 U.S.C. 103 as being unpatentable over Tisch_ Rosenbaum et al. US 9807530 B1 in view of Sensarn et al. US 11302342 B1 and further in view of US Li et al. 11172285 B1 and further in view of Mortensen et al. US 10360926 B2.
Regarding Claim 1, Tisch_ Rosenbaum teaches:
1. An image capture device, comprising: a first microphone; a second microphone; and a processor configured to: Tisch_ Rosenbaum teach (“(22) In an embodiment, the microphone selection control 130, the audio combiner 135, and/or the audio encoder 140 are implemented as a processor and a non-transitory computer-readable storage medium storing instructions that when executed by the processor carry out the functions attributed to the microphone selection controller 130, the audio combiner 135, and/or audio encoder 140 described herein. The microphone selection controller 130, audio combiner 135, and audio encoder 140 may be implemented using a common processor or separate processors. In other embodiments, the microphone selection controller 130, audio combiner 135, and/or audio encoder 140 may be implemented in hardware, (e.g., with an FPGA or ASIC), firmware, or a combination of hardware, firmware and software.” Col. 5, lines 7-20) by Tisch_ Rosenbaum et al. US 9807530 B1
obtain a first microphone signal from the first microphone; Tisch_ Rosenbaum teaches (“(24) FIG. 2 is a flowchart illustrating an embodiment of a process for generating an audio output audio signal from audio signals captured from multiple different microphones. Audio signals are received 202 from at least two microphones, which may include the drainage microphone 110, the first reference microphone 122, the second reference microphone 124 or any combination thereof. ...” col. 5, lines 29-35) Tisch_ Rosenbaum_ al. US 9807530 B1
obtain a second microphone signal from the second microphone; Tisch_ Rosenbaum teaches (“(24) FIG. 2 is a flowchart illustrating an embodiment of a process for generating an audio output audio signal from audio signals captured from multiple different microphones. Audio signals are received 202 from at least two microphones, which may include the drainage microphone 110, the first reference microphone 122, the second reference microphone 124 or any combination thereof. ...” col. 5, lines 29-35) Tisch_ Rosenbaum al. US 9807530 B1
Tisch_ Rosenbaum further teaches:
select non-voice sub-band frequency bins from the first microphone signal and the second microphone signal based on a lowest energy value of each respective non-voice sub-band frequency bin; (“(20) The audio combiner 135 combines the blocks or portions thereof (e.g., particular frequency sub-bands) of audio received from the microphone selection controller 130 to generate a combined audio signal. This combining may include combining blocks or portions thereof (e.g., particular frequency sub-bands) received from the different microphones 110, 122, 124.” Col. 4, lines 64-67) (“(26) … The uncorrelated processing algorithm may select, for each frequency band, a frequency component of an audio signal having the lowest uncorrelated noise and combine these frequency components together to create the combined audio signal. An example embodiment of an uncorrelated audio signal processing algorithm is described in further detail below with respect to FIG. 3. The combined audio signal is then encoded 218.” Col. 6, lines 30-38) (“(27) FIG. 3 is a flowchart illustrating an embodiment of an uncorrelated audio signal processing algorithm. This algorithm may be applied, for example, when wind noise or other uncorrelated noise is detected in the audio signals. For each sub-band, the audio signal having the lowest wind noise is selected 302 for that time interval. For example, in one embodiment, the audio signal having the lowest root-mean-square signal level in a given frequency sub-band is selected as the audio signal with the lowest wind noise for that frequency sub-band. In each sub-band, noise suppression processing is then applied 304 on the selected audio signal based on the sub-band correlation metric for that sub-band to further reduce the noise. The original (non-offset) audio signals for each of the selected sub-bands with noise suppression processing applied are then combined 306 to generate the output audio signal.” Col. 6, lines 30-55) by Tisch_ Rosenbaum al. US 9807530 B1
Tisch_ Rosenbaum further teaches:
output a composite signal that comprises the selected non-voice sub-band frequency bins and the selected voice sub-band frequency bins.
Tisch_ Rosenbaum_ Rosenbaumteaches (“(26) … The uncorrelated processing algorithm may select, for each frequency band, a frequency component of an audio signal having the lowest uncorrelated noise and combine these frequency components together to create the combined audio signal. An example embodiment of an uncorrelated audio signal processing algorithm is described in further detail below with respect to FIG. 3. The combined audio signal is then encoded 218.” Col. 6, line 30-38) (“(27) FIG. 3 is a flowchart illustrating an embodiment of an uncorrelated audio signal processing algorithm. This algorithm may be applied, for example, when wind noise or other uncorrelated noise is detected in the audio signals. For each sub-band, the audio signal having the lowest wind noise is selected 302 for that time interval. For example, in one embodiment, the audio signal having the lowest root-mean-square signal level in a given frequency sub-band is selected as the audio signal with the lowest wind noise for that frequency sub-band. In each sub-band, noise suppression processing is then applied 304 on the selected audio signal based on the sub-band correlation metric for that sub-band to further reduce the noise. The original (non-offset) audio signals for each of the selected sub-bands with noise suppression processing applied are then combined 306 to generate the output audio signal.” Col. 6, lines 30-55) (“(20) The audio combiner 135 combines the blocks or portions thereof (e.g., particular frequency sub-bands) of audio received from the microphone selection controller 130 to generate a combined audio signal. This combining may include combining blocks or portions thereof (e.g., particular frequency sub-bands) received from the different microphones 110, 122, 124.” col.4, lines 64-67) Tisch_ Rosenbaum al. US 9807530 B1
Tisch_ Rosenbaum does not explicitly teach determine coherence values between microphones.
Sensarn teaches:
determine coherence values between the first microphone signal and the second microphone signal across a frequency band, Sensarn teaches (“(205) … a coherence value (COH) representing a coherence between the first microphone and the second microphone,”) by Sensarn et al. US 11302342 B1
wherein the frequency band comprises a voice sub-band and non-voice sub-bands, and Sensarn teaches (“(124) FIG. 13 is a flowchart conceptually illustrating an example method for performing tap detection when wind is present according to embodiments of the present disclosure. As illustrated in FIG. 13, the device 110 may receive (1310) audio data corresponding to microphones (e.g., first audio data corresponding to a first microphone and second audio data corresponding to a second microphone), may convert (1312) the audio data from a time domain to a subband domain, and may determine (1314) if wind is detected in the audio data. For example, the device 110 may compare energy values in different frequency ranges to determine that wind is represented.” Col. 19, lines 48-58) (“(125) If the device 110 determines that wind is not detected in the audio data, the device 110 may select (1318) frequencies in a first frequency range (e.g., lower than 500 Hz) to perform tap detection processing. If the device 110 determines that wind is detected in the audio data, the device 110 may select (1320) frequencies in a second frequency range (e.g., higher than 3 kHz) to perform tap detection processing.” Col. 19, lines 60-67) by Sensarn et al. US 11302342 B1
Sensarn further teaches:
wherein the voice sub-band and the non-voice sub-bands each comprise frequency bins and Sensarn teaches where each amplitude value is for a different tone or “bin.” So, for example, if the sound wave consisted solely of a pure sinusoidal 1 kHz tone, then the frequency domain representation would consist of a discrete amplitude spike in the bin containing 1 kHz, with the other bins at zero. In other words, each tone “k” is a frequency index (e.g., frequency bin). (“(46) As used herein, a frequency band corresponds to a frequency range having a starting frequency and an ending frequency. Thus, the total frequency range may be divided into a fixed number (e.g., 256, 512, etc.) of frequency ranges, with each frequency range referred to as a frequency band and corresponding to a uniform size. …”) (“(55) Using a Fourier transform, a sound wave such as music or human speech can be broken down into its component “tones” of different frequencies, each tone represented by a sine wave of a different amplitude and phase. Whereas a time-domain sound wave (e.g., a sinusoid) would ordinarily be represented by the amplitude of the wave over time, a frequency domain representation of that same waveform comprises a plurality of discrete amplitude values, where each amplitude value is for a different tone or “bin.” So, for example, if the sound wave consisted solely of a pure sinusoidal 1 kHz tone, then the frequency domain representation would consist of a discrete amplitude spike in the bin containing 1 kHz, with the other bins at zero. In other words, each tone “k” is a frequency index (e.g., frequency bin).” Col 7, lines 55-67) (“(56) FIG. 2A illustrates an example of time indexes 216 (e.g., microphone audio data x(t) 210) and frame indexes 218 (e.g., microphone audio data x(n) 212 in the time domain and microphone audio data X(n, k) 216 in the frequency domain). For example, the system 100 may apply FFT processing to the time-domain microphone audio data x(n) 212, producing the frequency-domain microphone audio data X(n,k) 214, where the tone index “k” (e.g., frequency index) ranges from 0 to K and “n” is a frame index ranging from 0 to N. As illustrated in FIG. 2A, the history of the values across iterations is provided by the frame index “n”, which ranges from 1 to N and represents a series of samples over time.” Col. 8, lines 3-15) (“(125) If the device 110 determines that wind is not detected in the audio data, the device 110 may select (1318) frequencies in a first frequency range (e.g., lower than 500 Hz) to perform tap detection processing. If the device 110 determines that wind is detected in the audio data, the device 110 may select (1320) frequencies in a second frequency range (e.g., higher than 3 kHz) to perform tap detection processing.”) (“(126) Thus, the device 110 may perform tap detection processing using different frequency ranges depending on whether wind is detected in the audio data, as wind may cause relatively large ILD that may be incorrectly interpreted as a tap event. However, wind noise typically occurs at low frequencies below 1 kHz, whereas tap events cause ILD values across all frequencies, so the device 110 may perform tap detection processing using higher frequencies without departing from the disclosure. While FIG. 13 illustrates examples of the first frequency range and the second frequency range, the disclosure is not limited thereto. Instead, the first frequency range and the second frequency range may vary without departing from the disclosure.”) by Sensarn et al. US 11302342 B1
determine that wind is present based on the determined coherence values for each frequency bin; Sensarn teaches (“(128) … the tap detection component 1420 may determine that wind conditions are present and set an ILD threshold value to the wind threshold value (e.g., 30 dB).”) (“If the signal detection component 1510 determines that a signal is present (e.g., signal present), the signal detection component 1510 may output the signal to the inter-level difference (ILD) calculation component 1520 and/or the wind noise detection component 1530. …”) (“(141) The wind noise detection component 1530 may perform wind detection processing by determining a coherence between two input audio signals (e.g., two microphone signals). Two-channel coherence is defined as ratio of the cross power spectral density (PSD) and product of auto power spectral densities. Therefore, the wind noise detection component 1530 may use coherence calculation 1660 illustrated in FIG. 16C and shown below: …” Col. 22, lines 20-27) (“(141) The wind noise detection component 1530 may perform wind detection processing by determining a coherence between two input audio signals (e.g., two microphone signals). Two-channel coherence is defined as ratio of the cross power spectral density (PSD) and product of auto power spectral densities. Therefore, the wind noise detection component 1530 may use coherence calculation 1660 illustrated in FIG. 16C and shown below: … … where the PSDs are computed using smoothed periodogram (e.g., Power Spectral Density (PSD) calculation 1670) shown below: … … (144) However, wind noise is a low-frequency, non-stationary signal that is uncorrelated at different channels. The metric to be used for wind noise detection is magnitude coherence averaged over low frequencies [0-300] Hz,” col. 22, lines 20-30) by Sensarn et al. US 11302342 B1
Sensarn is considered to be analogous to the claimed invention because it relates electronic devices are commonly used to capture and process audio data.
Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify Tisch_ Rosenbaum to incorporate the teachings of Sensarn in order to include powerful microphone signal.
One could have been motivated to do so because system will improve performance with microphones tap detection. (“ (102) …. The device 110 may perform tap detection using the isolated microphone signals to improve performance of the tap detection component 810, when the microphones are not positioned symmetrically relative to the loudspeaker(s) of the device 110, and/or for other reasons.“) by Sensarn et al. US 11302342 B1
The combination does not explicitly teach a coherence value is determined for each frequency bin.
Li teaches:
a coherence value is determined for each frequency bin; (“(42) … In addition to performing one or more of these smoothing functions to the initial coherence values, the coherence-determination component 128 may calculate a coherence values for a first set of the frequency bins to determine coherence values for a remainder of the frequency bins. …” Col. Lines) by Li et al. US 11172285 B1
Li is considered to be analogous to the claimed invention because it relates to process audio signals to lessen the impact that wind and/or other environmental noise has upon the resulting quality of these audio signals.
Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify Tisch_ Rosenbaum and Sensarn to incorporate the teachings of Li in order to include a coherence value is determined for each frequency bin.
One could have been motivated to do so because system can process one or more audio signals based at least in part on these values to lessen the impact of wind and/or other unwanted environmental noise from the resulting signals. (“(24) (44) After determining the final coherence values for the number of N frequency ranges, the output-audio-signal component 132 may process one or more audio signals based at least in part on these values to lessen the impact of wind and/or other unwanted environmental noise from the resulting signals. …).” Col. 12, lines 10-15).“) by Li et al. US 11172285 B1
The combination does not explicitly teach select voice sub-band frequency bins from a predetermined microphone signal;
Mortensen teaches
select voice sub-band frequency bins from a predetermined microphone signal; and Mortensen teaches (“(46) Based on the insights on formant frequencies of vowels, a low-complexity and low-power voice activity detector can be used to determine whether a voice is present or not by providing filters that can detect presence of voice activity resulting from a reasonable number of vowels. A voice activity detector can include a first channel for processing a first audio stream and detecting activity in a first frequency band and a second channel for processing the first audio stream and detecting activity in a second frequency band. It is important to note that the first frequency band and the second frequency band are not just any frequency band, but they are selected carefully to allow the voice activity detector to detect the presence of voice, i.e., vowels. Accordingly, the first frequency band includes a first group of formant frequencies characteristic of vowels and the second frequency band includes a second group of formant frequencies characteristic of vowels. Furthermore, the voice activity detector includes a first decision module for observing the first channel and the second channel to determine whether voice activity is present in the first audio stream.” Col. 6, lines 10-29) (“(137) FIG. 12 shows an exemplary waveform of human speech having voice activity and ambient noise, according to some embodiments of the disclosure. In FIG. 12, the waveform shows a period of ambient noise, a period of voice onset (when someone is just about to utter a word), a period of voice activity (shown by the high activity in the waveform), and back to another period of ambient noise. When the voice activity detector detects the presence of voice and triggers a process, such as voice command detection, the samples in the sample storage is flushed and forwarded to the process for further processing. Due to a delay of the voice activity detector, the samples stored in sample storage to be provided to the process for processing corresponds to the period marked by “VAD DELAY/BUFFER”. The samples would include some voice onset and some voice activity.” Col. 24, lines 1-15) (“(148) In one example, the VAD has three channels, e.g., with two channels emphasizing male and female “OH” formants (as in “[OH]kay Bobby”), and third channel being used to reduce false alarms. False alarms can be triggered by noise or audio activity with wide band energy, which would trigger energy being detected for the Formant VAD channels. To detect such false alarms, it is possible to add an additional channel that detects energy outside of formant bands. If CH0 detects energy for male “OH” formants and CH1 detects energy for female “OH” formants, CH2 can be added to detect energy in out of formant bands, and the outputs of the three channels CH0, CH1, and CH2 can be combined like this: OUT=(CH0 or CH1) and not (CH2).” Col. 25, lines 36-48) by Mortensen et al. US 10360926 B2
Mortensen is considered to be analogous to the claimed invention because it relates to the field of audio signal processing, in particular, to voice activity detection for trigging a process in a processing system.
Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify Tisch_ Rosenbaum, Sensarn, Li, and to incorporate the teachings of Mortensen in order to include voice sub-band frequency bin.
One could have been motivated to do so because system will have high accuracy VAD algorithms can be perform to detect voice. (“(33) Voice activity detection (VAD) (also known as speech activity detection or speech detection) involves determining whether one or more voices is present or not present in an audio stream. In many cases, an audio stream can be noisy, which can make it difficult for a system to detect voices. Noise can come from sources not associated with speech, e.g., low-frequency sounds from a fan, a refrigerator, a helicopter, a motor, a loud bang, sounds from a keyboard, etc. Besides issues from noise, voices in an audio stream can be imperfect simply due to the way the audio was captured and/or transmitted. For at least these reasons, some VAD algorithms can be quite complicated, especially if the VAD algorithm is expected to perform with high accuracy.” col. 3, lines lines 35-47) by Mortensen et al. US 10360926 B2
Regarding Claim 3, the combination teaches the device claim 1 as identified above.
Sensarn further teaches
3. The image capture device of claim 1, wherein the predetermined microphone signal is the first microphone signal. Sensarn teaches a first acoustic impulse response (AIR) (i.e. minimum duration) associated with the first microphone (i.e. predetermined microphone). (“(74) … which can be approximated as a ratio between a first acoustic impulse response (AIR) associated with the first microphone (e.g., |A.sub.1[f]|.sup.2) and a second acoustic impulse response (AIR) associated with the second microphone (e.g., |A.sub.2[f]|.sup.2). …” col. 12, lines 10-15) by Sensarn et al. US 11302342 B1
Sensarn is considered to be analogous to the claimed invention because it relates electronic devices are commonly used to capture and process audio data.
Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify Tisch_ Rosenbaum to incorporate the teachings of Sensarn in order to include powerful microphone signal.
One could have been motivated to do so because system will improve performance with microphones tap detection. (“ (102) …. The device 110 may perform tap detection using the isolated microphone signals to improve performance of the tap detection component 810, when the microphones are not positioned symmetrically relative to the loudspeaker(s) of the device 110, and/or for other reasons.“) by Sensarn et al. US 11302342 B1
Regarding Claim 4, the combination teaches the device claim 1 as identified above.
Sensarn further teaches:
4. The image capture device of claim 1, wherein the predetermined microphone signal is the second microphone signal. Sensarn teaches first microphone closer (i.e. predetermined first/second microphone). Sensarn teaches therefore, it is expected that a first microphone closer to the tap location always receives a more powerful signal than a second microphone that is farther from the tap location. (“(78) In a near-field case, the source to microphone distances r.sub.i are comparable to the inter-microphone distance d. The attenuation factors 1/(√{square root over (4π)}r.sub.i) are distinct for different r.sub.i, which results in different power levels in different microphones. Thus, a tap event corresponds to a tap sound and is considered a near-field source as taps near the mics satisfy r.sub.i≤5 cm for d=2.6 cm mic spacing. Therefore, it is expected that a first microphone closer to the tap location always receives a more powerful signal than a second microphone that is farther from the tap location. ….“ col. 12, lines 46-60) by Sensarn et al. US 11302342 B1
Sensarn is considered to be analogous to the claimed invention because it relates electronic devices are commonly used to capture and process audio data.
Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify Tisch_ Rosenbaum to incorporate the teachings of Sensarn in order to include predetermined microphone signal.
One could have been motivated to do so because system will improve performance with microphones tap detection. (“ (102) …. The device 110 may perform tap detection using the isolated microphone signals to improve performance of the tap detection component 810, when the microphones are not positioned symmetrically relative to the loudspeaker(s) of the device 110, and/or for other reasons.“) by Sensarn et al. US 11302342 B1
Regarding Claim 5, the combination teaches the device claim 2 as identified above
5. The image capture device of claim 1, wherein coherence values of a subset of the frequency bins are averaged to create a wind meter value that indicates a presence of wind. Sensarn teaches (“(144) However, wind noise is a low-frequency, non-stationary signal that is uncorrelated at different channels. The metric to be used for wind noise detection is magnitude coherence averaged over low frequencies [0-300] Hz, illustrated in FIG. 16C as magnitude coherence calculation 1680 and shown below: …” col. 22, lines 45-55) (“(128) … the tap detection component 1420 may determine that wind conditions are present and set an ILD threshold value to the wind threshold value (e.g., 30 dB).”) (“If the signal detection component 1510 determines that a signal is present (e.g., signal present), the signal detection component 1510 may output the signal to the inter-level difference (ILD) calculation component 1520 and/or the wind noise detection component 1530. …”) (“(141) The wind noise detection component 1530 may perform wind detection processing by determining a coherence between two input audio signals (e.g., two microphone signals). Two-channel coherence is defined as ratio of the cross power spectral density (PSD) and product of auto power spectral densities. Therefore, the wind noise detection component 1530 may use coherence calculation 1660 illustrated in FIG. 16C and shown below: …” Col. 22, lines 20-27) (“(141) The wind noise detection component 1530 may perform wind detection processing by determining a coherence between two input audio signals (e.g., two microphone signals). Two-channel coherence is defined as ratio of the cross power spectral density (PSD) and product of auto power spectral densities. Therefore, the wind noise detection component 1530 may use coherence calculation 1660 illustrated in FIG. 16C and shown below: … … where the PSDs are computed using smoothed periodogram (e.g., Power Spectral Density (PSD) calculation 1670) shown below: … … (144) However, wind noise is a low-frequency, non-stationary signal that is uncorrelated at different channels. The metric to be used for wind noise detection is magnitude coherence averaged over low frequencies [0-300] Hz,” col. 22, lines 20-30) by Sensarn et al. US 11302342 B1
Claim 2, is/are rejected under 35 U.S.C. 103 as being unpatentable over Tisch_ Rosenbaum, Sensarn, Li, and Mortensen in view of Babu et al. US 20230253010 A1.
Regarding Claim 2, the combination teaches the device claim 1 as identified above
The combination does not explicitly teach voice sub-band ranges from 300 Hz to 8000 Hz.
Babu teaches:
2. The image capture device of claim 1, wherein the voice sub-band ranges from 300 Hz to 8000 Hz. Babu teaches However, as discussed above, the spikes in the duration bounded by duration 310 are not speech events but instead the result of the car hitting a speed breaker. (“[0048] FIG. 3A, FIG. 3B, and FIG. 3C … In FIG. 3C, a dispersion determination is made (e.g., a standard deviation of spectrum in the voice frequencies (300 Hz-4 kHz)) for at least one of a frame or series of audio frames. Because speech may have energy in a wider frequency range compared to engine noise, a standard deviation of spectrum in regions containing speech may be one or more of significantly or noticeably different. … … For example, the highlighted region inside the duration 310 shows non-speech noise that registers a spike in the standard deviation plot at FIG. 3C. ”) (“[0049] … FIG. 3C may represent the output of the standard deviation block in the first stage 130. The spikes in FIG. 3C may represent speech events. However, as discussed above, the spikes in the duration bounded by duration 310 are not speech events but instead the result of the car hitting a speed breaker. …”) (“[0080] … as shown in FIG. 4C, where the upper region of the cepstrum captures the pitch information in speech (e.g., from about 3.3 milliseconds to about 10 milliseconds, or bins 80-240 corresponding to a sample rate of 24 kilohertz). However, not all sounds associated with speech carry pitch information. Pitch is generally an unambiguous indicator of speech, however it is present only in certain regions of speech. …”) by Babu et al. US 20230253010 A1
Babu is considered to be analogous to the claimed invention because it relates voice activity detection.
Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify over Tisch_ Rosenbaum, Sensarn, Li, and Mortensen to incorporate the teachings of Babu in order to include a voice sub-band ranges from 300 Hz to 8000 Hz.
One could have been motivated to do so because pitch indicator a helpful indicator of speech. (“[0044] Speech may include certain pitch characteristics that can be detected in the upper region of the cepstrum, which may make the pitch indicator 244 a helpful indicator of speech. …”) by Babu et al. US 20230253010 A1
Claim 6 is/are rejected under 35 U.S.C. 103 as being unpatentable over Tisch_ Rosenbaum, Sensarn, Li, and Mortensen further in view of Hetherington et al. US 20140376742 A1.
Regarding Claim 6, the combination teaches the device claim 1 as identified above
The combination does not explicitly teach the lowest energy value corresponds to a high coherence value.
Hetherington teaches
6. The image capture device of claim 1, wherein the lowest energy value corresponds to a high coherence value. Hetherington teaches (“[0051] … The CohLR value calculated by the coherence calculator 302 may also be used to increase the gain of the lower amplitude microphone signal 118 when the spectral coherence is relatively high. …”) by Hetherington et al. US 20140376742 A1
Hetherington is considered to be analogous to the claimed invention because it relates to the field of processing sound fields.
Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify over Tisch_ Rosenbaum, Sensarn, Li, and Mortensen to incorporate the teachings of Hetherington in order to include lowest energy value corresponds to a high coherence value.
One could have been motivated to do so because the spectral coherence CohLR value may improve the fidelity . (“[0050] Further processing of the spectral coherence CohLR value may improve the fidelity. …”) by Hetherington et al. US 20140376742 A1
Claim 7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Tisch_ Rosenbaum, Sensarn, Li, and Mortensen in view of Miyasaka et al. US 10564927 B1.
Regarding Claim 7, the combination teaches the device claim 1 as identified above.
The combination does not explicitly wherein each frequency bin is 93.75 Hz.
Miyasaka teaches:
7. The image capture device of claim 1, wherein each frequency bin is 93.75 Hz. Miyasaka teaches the frequency resolution (i.e. each frequency bin) is 93.75 Hz (“(56) FIG. 5 illustrates the state where the time axis signal with a sampling frequency of 48 kHz is converted into a 256-point complex Fourier series, and is stored in the memory 11. Since the sampling frequency is 48 kHz, the frequency resolution is 93.75 Hz (=24000/256). The 256-point complex frequency signal is stored in the memory 11 with this frequency resolution up to 24 kHz, which is the Nyquist frequency.” Col. 10, lines 7-15) by Miyasaka et al. US 10564927 B1
Miyasaka is considered to be analogous to the claimed invention because it relates to a signal processing apparatus.
Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify over Tisch_ Rosenbaum, Sensarn, Li, and Mortensen to incorporate the teachings of Miyasaka in order to include each frequency bin is 93.75 Hz.
One could have been motivated to do so because system can accurately find a portion of the frequency signal of the sound. (101) According to this, it is possible to accurately find a portion of the frequency signal of the sound where high frequency components are insufficient, and to complement this insufficient portion with frequency components. Accordingly, it is possible to achieve the high-quality sound of the reproduction signal when reproducing the sound signal.“ col. 16, lines 55-60) by Miyasaka et al. US 10564927 B1
Claim 8, 12, and 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Tisch_ Rosenbaum et al. US 9807530 B1 in view of Sensarn et al. US 11302342 B1 and further in view of US Li et al. 11172285 B1 and further in view of Mortensen et al. US 10360926 B2 and further in view of ZHOU et al. CN 110648678 A
Claim 8 is a device claim with a limitation similar to the limitation of device Claim 1 and is rejected under similar rationale. Additionally,
Regarding Claim 8, Tisch_ Rosenbaum further teaches
8. An image capture device, comprising: a first microphone; a second microphone; and a processor configured to: Tisch_ Rosenbaum teach (“(22) In an embodiment, the microphone selection control 130, the audio combiner 135, and/or the audio encoder 140 are implemented as a processor and a non-transitory computer-readable storage medium storing instructions that when executed by the processor carry out the functions attributed to the microphone selection controller 130, the audio combiner 135, and/or audio encoder 140 described herein. The microphone selection controller 130, audio combiner 135, and audio encoder 140 may be implemented using a common processor or separate processors. In other embodiments, the microphone selection controller 130, audio combiner 135, and/or audio encoder 140 may be implemented in hardware, (e.g., with an FPGA or ASIC), firmware, or a combination of hardware, firmware and software.” Col. 5, lines 7-20) by Tisch_ Rosenbaum et al. US 9807530 B1
The combination does not explicitly teach select voice sub-band frequency bins from the first microphone signal based on an average energy per microphone in the voice sub-band.
ZHOU teaches:
select voice sub-band frequency bins from the first microphone signal based on an average energy per microphone in the voice sub-band; and ZHOU teaches (“the average value of the In some embodiments, the method further comprises: frequency of the high frequency sound energy is selected [4kHz, 8kHz] range, high-frequency sound energy average value is obtained by calculating the current frame and the history of several frames of high frequency voice energy value is obtained.” Page 5, first para) (“7. the average value of the range according to claim 3 used for the scene recognition method with multiple microphone conference, wherein the high-frequency voice energy of the frequency selected from [4kHz, 8kHz], the high frequency sound energy average value is obtained by calculating the current frame and the history of several frames of high frequency voice energy value is obtained.” Page 19, claim 7) by ZHOU et al. CN 110648678 A
ZHOU is considered to be analogous to the claimed invention because it relates to sound processing field, specifically relates to a scene recognition method and a system with multiple microphone conference.
Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify Tisch_ Rosenbaum, Sensarn, Li, and Mortensen to incorporate the teachings of ZHOU in order to include average energy value per microphone.
One could have been motivated to do so because system can eliminate signal noise and noise and improve the selection quality sound microphone voice output channel probability. (“using voice detection benefit to eliminate signal noise and noise and improve the selection quality sound microphone voice output channel probability.” Page 8, PARA 2) by ZHOU et al. CN 110648678 A
Regarding Claim 12, the combination teaches the device claim 8 as identified above.
The combination does not explicitly teach the non-voice sub-band range is below 300 Hz.
Sensarn further teaches
12. The image capture device of claim 8, wherein the non-voice sub-band range is below 300 Hz. Sensarn teaches (“(144) However, wind noise is a low-frequency, non-stationary signal that is uncorrelated at different channels. The metric to be used for wind noise detection is magnitude coherence averaged over low frequencies [0-300] Hz, illustrated in FIG. 16C as magnitude coherence calculation 1680 and shown below:…” col. 22, lines 46-51) (“(205) … a first microphone value associated with a first microphone and a second microphone value associated with a second microphone, a power value (POW) representing a root-mean-squared (RMS) power value, a coherence value (COH) representing a coherence between the first microphone and the second microphone, …” col. 36, lines 18-35) by Sensarn et al. US 11302342 B1
Sensarn is considered to be analogous to the claimed invention because it relates electronic devices are commonly used to capture and process audio data.
Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify Tisch_ Rosenbaum to incorporate the teachings of Sensarn in order to include powerful microphone signal.
One could have been motivated to do so because system will improve performance with microphones tap detection. (“ (102) …. The device 110 may perform tap detection using the isolated microphone signals to improve performance of the tap detection component 810, when the microphones are not positioned symmetrically relative to the loudspeaker(s) of the device 110, and/or for other reasons.“) by Sensarn et al. US 11302342 B1
Claim 13 is a device claim with a limitation similar to the limitation of device Claim 5 and is rejected under similar rationale.
Claim 9 is/are rejected under 35 U.S.C. 103 as being unpatentable over Tisch_ Rosenbaum, Sensarn, Li, Mortensen, and ZHOU and further in view of Stokes et al. US 20090214048 A1and further in view of Ramakrishnan et al. US 8571231 B2.
Regarding Claim 9, the combination teaches the device claim 8 as identified above
Stokes further teaches:
9. The image capture device of claim 8, wherein the processor is further configured to: select voice sub-band frequency bins from the second microphone signal based on the average energy per microphone in the voice sub-band; and Stokes teaches (“[0045] Given the foregoing, one way of implementing the HDRES parameter adaptation is using the following process. Referring to FIGS. 4A through 4D, the process entails first selecting a previously unselected frequency sub-band from the set of pre-defined sub-bands within the prescribed overall frequency range (400). The average power of the speaker signal segment corresponding to the current AEC segment is then computed for the selected sub-band (402). It is next determined if the speaker signal segment's average power for the selected sub-band exceeds a prescribed speaker signal power threshold (404). If the speaker signal power threshold is not exceeded, the HDRES parameters associated with the selected sub-band are not adapted and the process skips to action (432). In one embodiment, the HDRES parameters that are associated with a selected sub-band are all those that correspond to frequencies falling within the prescribed frequency range surrounding one of the aforementioned fundamental frequency bands or their harmonics in which the selected sub-band also falls. If the speaker signal power threshold is exceeded, however, the average power of the near-end microphone signal segment corresponding to the current AEC segment is computed for the selected sub-band (406). It is then determined if the near-end microphone segment's average power exceeds a prescribed microphone signal power threshold (408). If the microphone signal power threshold is not exceeded, the HDRES parameters associated with the selected sub-band are not adapted and the process skips to action (432). If the microphone signal power threshold is exceeded, the average power of the estimated residual echo component of the current AEC segment is computed for the selected sub-band (410). It is then determined if the estimated residual echo component's average power exceeds a prescribed residual echo power threshold (412). If the residual echo power threshold is not exceeded, the HDRES parameters associated with the selected sub-band are not adapted and the process skips to action (432).”) by Stokes et al. US 20090214048 A1
Stokes is considered to be analogous to the claimed invention because it relates to Harmonic distortion residual echo suppression (HDRES).
Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify Tisch_ Rosenbaum, Sensarn, Li, Mortensen and ZHOU to incorporate the teachings of Stokes in order to include average energy value per microphone.
One could have been motivated to do so because voice can be echo free. (“[0001] … An Acoustic Echo Canceller (AEC) suppresses the component of the near-end microphone signal corresponding to the captured speaker audio signal, thereby reducing the perceived echo effect at the far-end location.”) by Stokes et al. US 20090214048 A1
Ramakrishnan teaches:
apply a smoothing algorithm to the voice sub-band. Ramakrishnan teaches (“(69) … An electronic device 102 may obtain 802 an audio signal. As discussed above, an electronic device 102 may obtain 802 an audio signal by capturing an audio signal using a microphone or by receiving an audio signal (e.g., from another electronic device). …”) col.12, Lines 20-25) (“(71) … For example, the electronic device 102 may compute a running average of the magnitude or power of the frequency domain audio signal using different smoothing or averaging factors during VAD active periods (e.g., when voice or speech is detected) compared to VAD inactive periods (e.g., when voice or speech is not detected). More specifically, the smoothing factor may be larger when voice is detected than when voice is not detected using the VAD.” Col. 12, lines 42-53) by Ramakrishnan et al. US 8571231 B2
Ramakrishnan is considered to be analogous to the claimed invention because it relates suppressing noise in an audio signal.
Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify Tisch_ Rosenbaum, Sensarn, Li, Mortensen, Zhou, and Stokes to incorporate the teachings of Ramakrishnan in order to include a lowest energy value audio segment.
One could have been motivated to do so because system will improve the quality of audio signals. (“(80) The noise suppression module 910 employs frequency domain noise suppression techniques to improve the quality of audio signals 904. ...”) by Ramakrishnan et al. US 8571231 B2
Claim 10 and 11, is/are rejected under 35 U.S.C. 103 as being unpatentable over Tisch_ Rosenbaum, Sensarn, Li, Mortensen and Zhou in view of Babu et al. US 20230253010 A1.
Regarding Claim 10, the combination teaches the device claim 8 as identified above.
10. The image capture device of claim 8, wherein the voice sub-band frequency bins of the first microphone signal are selected for a minimum duration. Babu teaches where the upper region of the cepstrum captures the pitch information in speech (e.g., from about 3.3 milliseconds to about 10 milliseconds, or bins 80-240 corresponding to a sample rate of 24 kilohertz). 3.3 milliseconds (i.e minimum duration) (“[0080] … as shown in FIG. 4C, where the upper region of the cepstrum captures the pitch information in speech (e.g., from about 3.3 milliseconds to about 10 milliseconds, or bins 80-240 corresponding to a sample rate of 24 kilohertz). However, not all sounds associated with speech carry pitch information. Pitch is generally an unambiguous indicator of speech, however it is present only in certain regions of speech. …”) by Babu et al. US 20230253010 A1
Babu is considered to be analogous to the claimed invention because it relates voice activity detection.
Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify over Tisch_ Rosenbaum, Sensarn, Li, Mortensen and Zhou to incorporate the teachings of Babu in order to include a voice sub-band ranges from 300 Hz to 8000 Hz.
One could have been motivated to do so because pitch indicator a helpful indicator of speech. (“[0044] Speech may include certain pitch characteristics that can be detected in the upper region of the cepstrum, which may make the pitch indicator 244 a helpful indicator of speech. …”) by Babu et al. US 20230253010 A1
Regarding Claim 11, the combination teaches the device claim 10 as identified above.
11. The image capture device of claim 10, wherein the minimum duration is 5 milliseconds. Babu teaches (“[0080] … as shown in FIG. 4C, where the upper region of the cepstrum captures the pitch information in speech (e.g., from about 3.3 milliseconds to about 10 milliseconds, or bins 80-240 corresponding to a sample rate of 24 kilohertz). However, not all sounds associated with speech carry pitch information. Pitch is generally an unambiguous indicator of speech, however it is present only in certain regions of speech. …”) by Babu et al. US 20230253010 A1
Babu is considered to be analogous to the claimed invention because it relates voice activity detection.
Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify over Tisch_ Rosenbaum, Sensarn, Li, Mortensen and Zhou to incorporate the teachings of Babu in order to include a voice sub-band ranges from 300 Hz to 8000 Hz.
One could have been motivated to do so because pitch indicator a helpful indicator of speech. (“[0044] Speech may include certain pitch characteristics that can be detected in the upper region of the cepstrum, which may make the pitch indicator 244 a helpful indicator of speech. …”) by Babu et al. US 20230253010 A1
Claim 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Tisch_ Rosenbaum, Sensarn, Li, Mortensen, Zhou and further in view of Hetherington et al. US 20140376742 A1.
Claim 14 is a device claim with a limitation similar to the limitation of device Claim 6 and is rejected under similar rationale.
Regarding Claim 14, the combination teaches the device claim 8 as identified above
14. The image capture device of claim 8, wherein the lowest energy value corresponds to a high coherence value. Hetherington teaches (“[0051] … The CohLR value calculated by the coherence calculator 302 may also be used to increase the gain of the lower amplitude microphone signal 118 when the spectral coherence is relatively high. …”) by Hetherington et al. US 20140376742 A1
Hetherington is considered to be analogous to the claimed invention because it relates to the field of processing sound fields.
Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify over Tisch_ Rosenbaum, Sensarn, Li, and Mortensen to incorporate the teachings of Hetherington in order to include lowest energy value corresponds to a high coherence value.
One could have been motivated to do so because the spectral coherence CohLR value may improve the fidelity . (“[0050] Further processing of the spectral coherence CohLR value may improve the fidelity. …”) by Hetherington et al. US 20140376742 A1
Claim 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Tisch_ Rosenbaum, Sensarn, Li, Mortensen, Zhou and further in view of Miyasaka et al. US 10564927 B1.
Regarding Claim 15, the combination teaches the device claim 8 as identified above.
The combination does not explicitly teach wherein each frequency bin is 93.75 Hz.
15. The image capture device of claim 8, wherein each frequency bin is 93.75 Hz. Miyasaka teaches the frequency resolution (i.e. each frequency bin) is 93.75 Hz (“(56) FIG. 5 illustrates the state where the time axis signal with a sampling frequency of 48 kHz is converted into a 256-point complex Fourier series, and is stored in the memory 11. Since the sampling frequency is 48 kHz, the frequency resolution is 93.75 Hz (=24000/256). The 256-point complex frequency signal is stored in the memory 11 with this frequency resolution up to 24 kHz, which is the Nyquist frequency.” Col. 10, lines 7-15) by Miyasaka et al. US 10564927 B1
Miyasaka is considered to be analogous to the claimed invention because it relates to a signal processing apparatus.
Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify over Tisch_ Rosenbaum, Sensarn, Li, Mortensen, and Zhou to incorporate the teachings of Miyasaka in order to include each frequency bin is 93.75 Hz.
One could have been motivated to do so because system can accurately find a portion of the frequency signal of the sound. (101) According to this, it is possible to accurately find a portion of the frequency signal of the sound where high frequency components are insufficient, and to complement this insufficient portion with frequency components. Accordingly, it is possible to achieve the high-quality sound of the reproduction signal when reproducing the sound signal.” col. 16, lines 55-60) by Miyasaka et al. US 10564927 B1
Claim 16, is/are rejected under 35 U.S.C. 103 as being unpatentable over Tisch_ Rosenbaum et al. US 9807530 B1 in view of Sensarn et al. US 11302342 B1 and further in view of US Li et al. 11172285 B1 and further in view of Mortensen et al. US 10360926 B2 and further in view of Thyssen et al. US 20130216057 A1
Claim 16 is a method claim with a limitation similar to the limitation of device Claim 1 and is rejected under similar rationale. Additionally,
The combination does not explicitly teach selecting voice sub-band frequency bins from the first microphone signal based on a lowest coherence value;
Thyssen teaches:
selecting voice sub-band frequency bins from the first microphone signal based on a lowest coherence value; and Thyssen teaches (“[0050] Generally speaking, if the measure of coherence for a given frequency bin is low, then desired speech is likely being received via the microphone with little echo being present. However, if the measure of coherence is high then there is likely to be significant acoustic echo. In accordance with certain embodiments, an aggressive tracker is utilized that maps a high measure of coherence to a fast update rate for the estimated statistics and maps a low measure of coherence to a low update rate for the estimated statistics, which may include not updating at all. In an embodiment in which the statistics are estimated by calculating a running mean, the aforementioned mapping may be achieved by controlling the weight attributed to the current instantaneous statistics when calculating the mean. Thus, to achieve a slow update rate, little or no weight may be assigned to the current instantaneous statistics, but to achieve a fast update rate, more significant weight may be assigned current instantaneous statistics.”) by Thyssen et al. US 20130216057 A1
Thyssen is considered to be analogous to the claimed invention because it relates to performing echo cancellation in an audio communication system.
Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify Tisch_ Rosenbaum, Sensarn, Li, and Mortensen to incorporate the teachings of Thyssen in order to include selecting voice sub-band frequency bins from the first microphone signal based on a lowest coherence value.
One could have been motivated to do so because system have a longer impulse response. (“[0040] … Effectively, the time direction filters in individual frequency bins increase the frequency resolution by providing a non-flat frequency response within a bin. This approach also enables the acoustic echo canceller to have a longer tail length than that otherwise provided by the size of the FFT, which is useful for environments having high reverberation and thus a longer impulse response”). by Thyssen et al. US 20130216057 A1
Claim 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Tisch_ Rosenbaum, Sensarn, Li, Mortensen and Thyssen and further in view of Chen et al. US 20140278397 A1 and further in view of Ramakrishnan et al. US 8571231 B2
Regarding Claim 17, the combination teaches the method claim 16 as identified above.
Chen teaches:
17. The method of claim 16, further comprising: selecting voice sub-band frequency bins from the second microphone signal based on the lowest coherence value; and Chen teaches (“[0079] ... If the measure of coherence for a given frequency bin is low, then desired speech is likely being received via the microphone with little echo being present. However, if the measure of coherence is high, then there is likely to be significant acoustic echo. In accordance with certain embodiments disclosed therein, a high measure of coherence is mapped to a fast update rate for the estimated statistics and a low measure of coherence is mapped to a slow update rate for the estimated statistics, which may include not updating at all.”) by Chen et al. US 20140278397 A1
Chen is considered to be analogous to the claimed invention because it relates to speech processing algorithms.
Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify Tisch_ Rosenbaum, Sensarn, Li, Mortensen and Thyssen to incorporate the teachings of Chen in order to include selecting voice sub-band frequency bins from the first microphone signal based on a lowest coherence value.
One could have been motivated to do so because speaker identification (SID) can improve the identification of when the far-end speaker is talking and when the near-end speaker is talking. (“[0075] … The parameters are only updated during a far-end single talk condition because in such a condition far-end speech signal 314 is strong and there is no near-end speech signal to interfere with proper parameter adaptation. speaker identification (SID) can improve the identification of when the far-end speaker is talking and when the near-end speaker is talking.”). by Chen et al. US 20140278397 A1
The combination does not explicitly teach applying a smoothing algorithm to the voice sub-band.
Ramakrishnan teaches:
applying a smoothing algorithm to the voice sub-band. Ramakrishnan teaches (“(69) … An electronic device 102 may obtain 802 an audio signal. As discussed above, an electronic device 102 may obtain 802 an audio signal by capturing an audio signal using a microphone or by receiving an audio signal (e.g., from another electronic device). …”) col.12, Lines 20-25) (“(71) … For example, the electronic device 102 may compute a running average of the magnitude or power of the frequency domain audio signal using different smoothing or averaging factors during VAD active periods (e.g., when voice or speech is detected) compared to VAD inactive periods (e.g., when voice or speech is not detected). More specifically, the smoothing factor may be larger when voice is detected than when voice is not detected using the VAD.” Col. 12, lines 42-53) by Ramakrishnan et al. US 8571231 B2
Ramakrishnan is considered to be analogous to the claimed invention because it relates identifying speech using energy features, entropy features, and frequency features of audio segments.
Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify Tisch_ Rosenbaum, Sensarn, Li, Mortensen, Thyssen, and Chen to incorporate the teachings of Ramakrishnan in order to include a lowest energy value audio segment.
One could have been motivated to do so because system will and improving the quality of the desired signal. (“(22) The systems and methods disclosed herein pertain generally to the field of signal processing solutions used for improving voice quality of electronic devices (e.g., wireless communication devices). More specifically, the systems and methods disclosed herein focus on suppressing noise (e.g., ambient noise, background noise) and improving the quality of the desired signal.”) by Ramakrishnan et al. US 8571231 B2
Claim 18 and 19, is/are rejected under 35 U.S.C. 103 as being unpatentable over Tisch_ Rosenbaum, Sensarn, Li, Mortensen and Thyssen and further in view of Babu et al. US 20230253010 A1
Regarding Claim 18, the combination teaches the method claim 16 as identified above.
The combination does not explicitly teach the voice sub-band frequency bins of the first microphone signal are selected for a minimum duration.
Babu teaches:
18. The method of claim 16, wherein the voice sub-band frequency bins of the first microphone signal are selected for a minimum duration. Babu teaches where the upper region of the cepstrum captures the pitch information in speech (e.g., from about 3.3 milliseconds to about 10 milliseconds, or bins 80-240 corresponding to a sample rate of 24 kilohertz). 3.3 milliseconds (i.e minimum duration) (“[0080] … as shown in FIG. 4C, where the upper region of the cepstrum captures the pitch information in speech (e.g., from about 3.3 milliseconds to about 10 milliseconds, or bins 80-240 corresponding to a sample rate of 24 kilohertz). However, not all sounds associated with speech carry pitch information. Pitch is generally an unambiguous indicator of speech, however it is present only in certain regions of speech. …”) by Babu et al. US 20230253010 A1
Babu is considered to be analogous to the claimed invention because it relates voice activity detection.
Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify over Tisch_ Rosenbaum, Sensarn, Li, Mortensen and Thyssen to incorporate the teachings of Babu in order to include a voice sub-band ranges from 300 Hz to 8000 Hz.
One could have been motivated to do so because pitch indicator a helpful indicator of speech. (“[0044] Speech may include certain pitch characteristics that can be detected in the upper region of the cepstrum, which may make the pitch indicator 244 a helpful indicator of speech. …”) by Babu et al. US 20230253010 A1
Regarding Claim 19, the combination teaches the method claim 18 as identified above.
Babu further teaches:
19. The method of claim 18, wherein the minimum duration is 5 milliseconds. Babu teaches (“[0080] … as shown in FIG. 4C, where the upper region of the cepstrum captures the pitch information in speech (e.g., from about 3.3 milliseconds to about 10 milliseconds, or bins 80-240 corresponding to a sample rate of 24 kilohertz). However, not all sounds associated with speech carry pitch information. Pitch is generally an unambiguous indicator of speech, however it is present only in certain regions of speech. …”) by Babu et al. US 20230253010 A1
Babu is considered to be analogous to the claimed invention because it relates voice activity detection.
Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify over Tisch_ Rosenbaum, Sensarn, Li, Mortensen and Zhou to incorporate the teachings of Babu in order to include a voice sub-band ranges from 300 Hz to 8000 Hz.
One could have been motivated to do so because pitch indicator a helpful indicator of speech. (“[0044] Speech may include certain pitch characteristics that can be detected in the upper region of the cepstrum, which may make the pitch indicator 244 a helpful indicator of speech. …”) by Babu et al. US 20230253010 A1
Claim 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Tisch_Rosenbaum, Sensarn, Li, , Mortensen, Thyssen, and further in view of Yan et al. US 20230300544 A1.
Regarding Claim 20, the combination teaches the method claim 16 as identified above.
Yan teaches
20. The method of claim 16, wherein the non-voice sub-band range is above 8000 Hz.
Yan teaches (“[0080] In some embodiments, the hearing level of the wearer of the bone conduction hearing aid may be different at the same sound level in different frequency bands. For example, at a sound level of 20 dBC, the hearing level of the wearer in a high frequency band (e.g., 8000 Hz-12000 Hz) may be equal to a certain value within a range of 41 dBHL-60 dBHL; and the hearing level of the wearer in a low frequency band may all be equal to a certain value within a range of 26 dBHL-40 dBHL.”) by Yan et al. US 20230300544 A1
Yan is considered to be analogous to the claimed invention because it relates the field of bone conduction hearing aids.
Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify Tisch_Rosenbaum, Sensarn, Li, Mortensen, and Thyssen to incorporate the teachings of Yan in order to include a non-voice sub-band range is above 8000 Hz.
One could have been motivated to do so because system may modify or enhance the sound signal. (“[0065] …. The one or more sound signal processing algorithms may modify or enhance the sound signal. ...”) by Yan et al. US 20230300544 A1
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Emori et al. (U.S. PG Pub No. US 20110071825 A1): VAD and the sub-band power is compared from one microphone to another to select the microphone (speaker) with the maximum power value.
Goodwin, (U.S. PG Pub No. US 8781137 B1) -Sub-bands containing wind noise and removing wind noise.
PEDERSEN et al. (US PG Pub No. 20240284127 A1): Hearing aid with two or more microphones and a wind noise controller configured to, recurringly, determine a first wind noise level.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to FOUZIA HYE SOLAIMAN whose telephone number is (571)270-5656. The examiner can normally be reached M-F (8-5)AM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Paras D. Shah can be reached at (571) 270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/F.H.S./Examiner, Art Unit 2653
/Paras D Shah/Supervisory Patent Examiner, Art Unit 2653
07/29/2026