Prosecution Insights
Last updated: October 04, 2026
Application No. 18/904,906

ECHO CANCELLATION AND AUDIO PROCESSING FOR EMBEDDED SYSTEMS

Non-Final OA §102
Filed
Oct 02, 2024
Priority
Jun 07, 2024 — CN 202410733049.1
Examiner
TESHALE, AKELAW
Art Unit
2694
Tech Center
2600 — Communications
Assignee
Beken Corporation
OA Round
1 (Non-Final)
82%
Grant Probability
Favorable
1-2
OA Rounds
9m
Est. Remaining
98%
With Interview

Examiner Intelligence

Grants 82% — above average
82%
Career Allowance Rate
709 granted / 864 resolved
+20.1% vs TC avg
Strong +16% interview lift
Without
With
+16.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
19 currently pending
Career history
879
Total Applications
across all art units

Statute-Specific Performance

§101
7.4%
-32.6% vs TC avg
§103
45.7%
+5.7% vs TC avg
§102
34.1%
-5.9% vs TC avg
§112
5.6%
-34.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 864 resolved cases

Office Action

§102
DETAILED ACTION Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action. Claims 1-20 are rejected under 35 U.S.C. 102 (a) (1) as being anticipated by U.S Pub. No. 2014/0126745 A1 to Dickins et al. (hereinafter “Dickins”). Regarding claim 1, Dickins teaches a method at an embedded system for signal processing (please see Fig.1), the method comprising: receiving a reference sound signal transmitted from a client device; capturing, by a microphone of the embedded system, a microphone sound signal to be transmitted to the client device (paragraphs [0091] and [0098]; system 100 that include echo suppression include a reference signal input processor 111 to accept one or more reference signals, a transformer 113 and a spectral banding element 115 to form a banded frequency domain amplitude metric representation 116 of the one or more reference signals. Such versions of system 100 include a predictor 117 of a banded frequency domain amplitude metric representation of the echo 118 based on adaptively determined filter coefficients); determining an energy ratio between the reference sound signal and the microphone sound signal (paragraphs [0144] and [00223]; a signal to noise ratio of at least 3 dB is required for a contribution to the signal level parameter S. If the current signal level is large relative to the noise and echo estimate, the summation term has a maximum of 1 for each band); obtaining a preliminary echo cancelling coefficient (paragraphs [0039], [0045] and [0091]; the means for predicting includes means for adaptively determining echo filter coefficients coupled to means for determining an estimate of the banded spectral amplitude metric of the noise, means for voice-activity detecting (VAD) using the estimate of the banded spectral amplitude metric of the mixed-down signal, and means for updating the filter coefficients based on the estimates of the banded spectral amplitude metric of the mixed-down signal and of the noise, and the previously predicted echo spectral content); determining a preliminary output signal based on the reference sound signal, the microphone sound signal, and the preliminary echo cancelling coefficient (Fig.1, paragraphs [0091], [0098] and [0106]; a plurality of input signals, e.g., signals from a plurality of spatially separated microphones; and, for echo suppression, (b) one or more reference signals, e.g., signals from or to be rendered by one or more loudspeakers and that can cause echoes); iteratively updating the preliminary output signal to generate a first output signal, the iterative updating of the preliminary output signal (paragraphs [0039], [0045] and [0091] and [0098]; the means for predicting 117, 123, 125, 127 includes means for adaptively determining 125, 127 echo filter coefficients 128 coupled to means for determining 123 an estimate of the banded spectral amplitude metric of the noise 124, means for voice-activity detecting (VAD) using the estimate of the banded spectral amplitude metric of the mixed-down signal 122, and means for updating 127 the filter coefficients 128. The output of the VAD is coupled to means for updating and determined if the means for updating updates the filter coefficients. The filter coefficients are updated based on the estimates of the banded spectral amplitude metric of the mixed-down signal 122 and of the noise 124, and the previously predicted echo spectral content 118) including: updating an updating index based on a predetermined step size, the reference sound signal, and an output signal determined in a preceding iteration (Fig.1, paragraphs [0039], [0129]- [0131]; overlapping and combining is dependent on the frame size, transform size and window functions, and should be designed to achieve a accurate reconstruction of the input signal in the absence of any processing or modification of the signal, X.sub.n, in the frequency domain); updating the preliminary echo cancelling coefficient based on the updated updating index, the energy ratio, and a predetermined suppression depth; and updating the preliminary output signal based on the reference sound signal, the microphone sound signal, and the updated preliminary echo cancelling coefficient (Abstract, paragraphs [0039], [0045], [0052] and [0091]; the banded signal 110 is a sufficiently accurate estimate of the banded spectral amplitude metric of the mixed-down signal 122, so that signal spectral estimator 121 is not used. The results of the VAD 125 are used by an adaptive filter updater 127 to determine whether to update the filter coefficients 128 based on the estimates of the banded spectral amplitude metric of the mixed-down signal 122 (or 110) and of the noise 124, and the previously predicted echo spectral content 118); generating a second output signal based on the first output signal and a noise estimate (Abstract, paragraphs [0039], [0040], [0058] and [0081]- [0092]; adaptively determine the filter coefficients, a noise estimator determines an estimate of the banded spectral amplitude metric of the noise. A voice-activity detector (VAD) uses the banded spectral amplitude metric of the noise, an estimate of the banded spectral amplitude metric of the mixed-down signal determined by a signal spectral estimator, and previously predicted echo spectral content to ascertain whether there is voice or not. In some embodiments, the banded signal is a sufficiently accurate estimate of the banded spectral amplitude metric of the mixed-down signal, so that signal spectral estimator is not used. The output of the VAD is used by an adaptive filter updater to determine whether or not to update the filter coefficients, the updating based on the estimates of the banded spectral amplitude metric of the mixed-down signal and of the noise, and the previously predicted echo spectral content); and transmitting the second output signal to the client device in replacement of the microphone sound signal (Fig.1, paragraphs [0039], [0095]- [0098], [0111] and [0140]; output of the VAD is used by an adaptive filter updater to determine whether or not to update the filter coefficients, the updating based on the estimates of the banded spectral amplitude metric of the mixed-down signal and of the noise, and the previously predicted echo spectral content). Regarding claim 2, Dickins teaches the method of claim 1, further comprising: determining that the energy ratio between the microphone sound signal and the reference sound signal is less than a ratio threshold and the reference sound signal is greater than an echo threshold; and increasing a frequency of the iteratively updating of the preliminary output signal (paragraphs [0083], [0156] and [0177]- [0178]; suppression in order to facilitate suppression of residual noise to a constant power across frequency value relative to the hearing threshold. One suggested approach for normalization of the bands is to scale such that the 1 kHz band has unity energy gain from the input, and the other bands are scaled such that a noise source having a relative spectrum matching the threshold of hearing would be white or constant power across the bands. In some sense, this is a pre-emphasis filter on the bands prior to analysis which causes a drop in sensitivity in the lower and higher bands. This normalization is useful, since if the residual noise is controlled to be constant across the bands, this achieves a perceptually white noise when close to the hearing threshold). Regarding claim 3, Dickins teaches the method of claim 1, further comprising: determining that the energy ratio between the microphone sound signal and the reference sound signal is less than a ratio threshold and the reference sound signal is greater than an echo threshold (paragraphs [0177]- [0178] and [0226]- [0227]; suppression in order to facilitate suppression of residual noise to a constant power across frequency value relative to the hearing threshold. One suggested approach for normalization of the bands is to scale such that the 1 kHz band has unity energy gain from the input, and the other bands are scaled such that a noise source having a relative spectrum matching the threshold of hearing would be white or constant power across the bands. In some sense, this is a pre-emphasis filter on the bands prior to analysis which causes a drop in sensitivity in the lower and higher bands. This normalization is useful, since if the residual noise is controlled to be constant across the bands, this achieves a perceptually white noise when close to the hearing threshold); comparing the output signal determined in the preceding iteration with a threshold (paragraphs [0148], [0201] and [0211]; echo may cause an over-estimation of the noise component. For this reason, one embodiment of the invention includes echo-gated noise estimation: updating the noise estimate N.sub.b', and stopping the update of the noise estimate when the predicted echo level is significant compared with the previous noise estimate. That is, that noise estimator 123 provides an estimate which is gated when the predicted echo spectral content is significant compared to the previously estimated noise spectral content); in response to a comparison result that the output signal determined in the preceding iteration is greater than the threshold, increasing the updating index based on the predetermined step size (paragraphs [0083] and [0177]; one suggested approach for normalization of the bands is to scale such that the 1 kHz band has unity energy gain from the input, and the other bands are scaled such that a noise source having a relative spectrum matching the threshold of hearing would be white or constant power across the bands. In some sense, this is a pre-emphasis filter on the bands prior to analysis which causes a drop in sensitivity in the lower and higher bands. This normalization is useful, since if the residual noise is controlled to be constant across the bands, this achieves a perceptually white noise when close to the hearing threshold); and in response to a comparison result that the output signal determined in the preceding iteration is less than the threshold, decreasing the updating index based on the predetermined step size (paragraphs [0119] and [0129]-[0131]; overlapping and combining is dependent on the frame size, transform size and window functions, and should be designed to achieve a accurate reconstruction of the input signal in the absence of any processing or modification of the signal, X.sub.n, in the frequency domain). Regarding claim 4, Dickins teaches the method of claim 1, wherein the generating of the second output signal based on the first output signal and the noise estimate comprises: determining the noise estimate based on the first output signal and at least one predetermined smoothing factor (paragraphs [0040] and [0046]; ensuring smoothness by carrying out time smoothing and, in some embodiments, band-to-band smoothing. In some embodiments that include post-processing, the means for post-processing includes means for spatially-selective voice activity detecting using two or more of the spatial features to generate a signal classification, such that the post-processing is according to the signal classification); determining a filter coefficient based on the first output signal and the noise estimate (paragraphs [0039]-[0042]; a noise suppression probability indicator, e.g., noise suppression gain determined using an estimate of noise spectral content. In some embodiments, the estimate of noise spectral content is a spatially-selective estimate of noise spectral content. In some embodiments that include echo suppression, the noise suppression probability indicator, e.g., suppression gain includes echo suppression); filtering the first output signal based on the filter coefficient to generate a filtered first output signal (paragraphs [0039]- [0040]; the system include a predictor of a banded frequency domain amplitude metric representation of the echo based on adaptively determined filter coefficients. To adaptively determine the filter coefficients, a noise estimator determines an estimate of the banded spectral amplitude metric of the noise); adjusting the at least one predetermined smoothing factor based on the filtered first output signal (paragraphs [0039]-[0040]; he system include a predictor of a banded frequency domain amplitude metric representation of the echo based on adaptively determined filter coefficients. To adaptively determine the filter coefficients, a noise estimator determines an estimate of the banded spectral amplitude metric of the noise); updating the noise estimate based on the filtered first output signal and the adjusted at least one smoothing factor (paragraphs [0039]-[0040]; he system include a predictor of a banded frequency domain amplitude metric representation of the echo based on adaptively determined filter coefficients. To adaptively determine the filter coefficients, a noise estimator determines an estimate of the banded spectral amplitude metric of the noise); updating the filter coefficient based on the updated noise estimate and the filtered first output signal; and filtering the filtered first output signal based on the updated filter coefficient to generate the second output signal (paragraphs [0039]- [0040]; the system include a predictor of a banded frequency domain amplitude metric representation of the echo based on adaptively determined filter coefficients. To adaptively determine the filter coefficients, a noise estimator determines an estimate of the banded spectral amplitude metric of the noise). Regarding claim 5, Dickins teaches the method of claim 1, further comprising: attenuating frequency bands of second output signal below a cut-off frequency to generate a third output signal (paragraph [0168]; It can be seen that this has a unity gain across the spectrum with a high pass characteristic having a cut-off frequency around 100 Hz. The high frequency shelf and banding are not essential components of the embodiments presented herein, but are suggested features for use on typical microphone input signals for the case of the signal of interest being a voice input). Regarding claim 6, Dickins teaches the method of claim 1, further comprising: adjusting a plurality of frequency bands in the second output signal based on a plurality of parameters across the plurality of frequency bands to generate a third output signal (paragraphs [0045], [0049] –[0050] and [0083]; echo suppression include means for accepting one or more reference signals and for forming a banded frequency domain amplitude metric representation of the one or more reference signals, and means for predicting a banded frequency domain amplitude metric representation of the echo). Regarding claim 7, Dickins teaches the method of claim 1, further comprising: injecting a noise signal to the second output signal to generate a third output signal (paragraphs [00061]-[0062]; determining banded spatial features from the plurality of sampled input signals; calculating a first set of suppression probability indicators, including an out-of-location suppression probability indicator determined using two or more of the spatial features, and a noise suppression probability indicator determined using an estimate of noise spectral content; combining the first set of probability indicators to determine a first combined gain for each band. The first combined gain, after post-processing if post-processing is included, forms a final gain for each band; and applying an interpolated final gain determined from the first combined gain. Interpolating the final gain produces final bin gains to apply to bin data of the mixed-down signal to form suppressed signal data). Regarding claim 8, Dickins teaches the method of claim 7, wherein the noise signal is a pink noise that decreases in amplitude as frequency increases (paragraphs [0039], [0045], [0052] and [0091]; the results of the VAD 125 are used by an adaptive filter updater 127 to determine whether to update the filter coefficients 128 based on the estimates of the banded spectral amplitude metric of the mixed-down signal 122 (or 110) and of the noise 124, and the previously predicted echo spectral content). Regarding claim 9, Dickins teaches the method of claim 1, further comprising: dynamically calculating a gain coefficient for each frame of the second output signal based on a maximum amplitude within the frame; and apply the gain coefficient on the each frame of the second output signal to generate a third output signal (paragraphs [0040]-[0042]; gain calculator further calculates an additional echo suppression probability indicator, e.g., an echo suppression gain). Regarding claim 10, Dickins teaches the method of claim 1, further comprising: providing a graphical user interface; receiving user configurations via the graphical user interface; and adjusting parameters of the embedded system based on the user configurations (paragraphs [0062] and [0172]; a set of parameters and using: an estimate of noise spectral content, the banded frequency domain amplitude metric representation of the echo, and the banded spatial features, the set of parameters including whether the estimate of noise spectral content is spatially selective or not, which indication of voice activity an instantiation determines being controlled by a selection of the parameters, voice activity). Regarding claim 11, Dickins teaches the method of claim 1, further comprising: determining a size of data buffered in the embedded system (paragraphs [0123] and [0127]; discrete finite length Fourier transform, such as implemented by the FFT, is often referred to as a circulant transform due to the implicit assumption that the signal in the transform window is in some way periodic or repetitive. Most general forms of circulant transforms can be represented by buffering, a window, a twist (real value to complex value transformation) and a DFT, e.g., FFT. An optional complex twist after the DFT can be used to adjust the frequency domain representation to match specific transform definitions); estimating a delay time based on the size of data buffered in the embedded system; and shifting the reference sound signal by the delay time to align the reference sound signal with the microphone sound signal (paragraphs [0139] and [0149]; a microphone inputs to effect appropriate time delay to be used in the mixing down). Regarding claim 12, Dickins teaches an embedded system for signal processing, the embedded system (please see Fig.1) comprising: a processor; and a memory storing instructions that, when executed by the processor, configure the embedded system to: receive a reference sound signal transmitted from a client device (paragraphs [0091] and [0098]; system 100 that include echo suppression include a reference signal input processor 111 to accept one or more reference signals, a transformer 113 and a spectral banding element 115 to form a banded frequency domain amplitude metric representation 116 of the one or more reference signals. Such versions of system 100 include a predictor 117 of a banded frequency domain amplitude metric representation of the echo 118 based on adaptively determined filter coefficients); capture, by a microphone of the embedded system, a microphone sound signal to be transmitted to the client device (paragraphs [0091] and [0098]; system 100 that include echo suppression include a reference signal input processor 111 to accept one or more reference signals, a transformer 113 and a spectral banding element 115 to form a banded frequency domain amplitude metric representation 116 of the one or more reference signals. Such versions of system 100 include a predictor 117 of a banded frequency domain amplitude metric representation of the echo 118 based on adaptively determined filter coefficients); generate a first output signal based on the reference sound signal, the microphone sound signal, and a preliminary echo cancelling coefficient (paragraphs [0039], [0045] and [0091]; the means for predicting includes means for adaptively determining echo filter coefficients coupled to means for determining an estimate of the banded spectral amplitude metric of the noise, means for voice-activity detecting (VAD) using the estimate of the banded spectral amplitude metric of the mixed-down signal, and means for updating the filter coefficients based on the estimates of the banded spectral amplitude metric of the mixed-down signal and of the noise, and the previously predicted echo spectral content); determine a noise estimate based on the first output signal and at least one predetermined smoothing factor (paragraphs [0040] and [0046]; ensuring smoothness by carrying out time smoothing and, in some embodiments, band-to-band smoothing. In some embodiments that include post-processing, the means for post-processing includes means for spatially-selective voice activity detecting using two or more of the spatial features to generate a signal classification, such that the post-processing is according to the signal classification); determine a filter coefficient based on the first output signal and the noise estimate; filter the first output signal based on the filter coefficient to generate a filtered first output signal (paragraphs [0039]-[0042]; a noise suppression probability indicator, e.g., noise suppression gain determined using an estimate of noise spectral content. In some embodiments, the estimate of noise spectral content is a spatially-selective estimate of noise spectral content. In some embodiments that include echo suppression, the noise suppression probability indicator, e.g., suppression gain includes echo suppression); adjust the at least one predetermined smoothing factor based on the filtered first output signal (paragraphs [0039]-[0040]; he system include a predictor of a banded frequency domain amplitude metric representation of the echo based on adaptively determined filter coefficients. To adaptively determine the filter coefficients, a noise estimator determines an estimate of the banded spectral amplitude metric of the noise); update the noise estimate based on the filtered first output signal and the adjusted at least one smoothing factor; update the filter coefficient based on the updated noise estimate and the filtered first output signal (paragraphs [0039]-[0040]; he system include a predictor of a banded frequency domain amplitude metric representation of the echo based on adaptively determined filter coefficients. To adaptively determine the filter coefficients, a noise estimator determines an estimate of the banded spectral amplitude metric of the noise); filter the filtered first output signal based on the updated filter coefficient to generate a second output signal; and transmit the second output signal to the client device in replacement of the microphone sound signal (paragraphs [0039]- [0040]; the system include a predictor of a banded frequency domain amplitude metric representation of the echo based on adaptively determined filter coefficients. To adaptively determine the filter coefficients, a noise estimator determines an estimate of the banded spectral amplitude metric of the noise). Regarding claim 13, Dickins teaches the embedded system of claim 12, wherein to generate the first output signal based on the reference sound signal, the microphone sound signal, and the preliminary echo cancelling coefficient, the instructions configure the embedded system to: determine an energy ratio between the reference sound signal and the microphone sound signal (paragraphs [0144] and [00223]; a signal to noise ratio of at least 3 dB is required for a contribution to the signal level parameter S. If the current signal level is large relative to the noise and echo estimate, the summation term has a maximum of 1 for each band); obtain the preliminary echo cancelling coefficient (paragraphs [0039], [0045] and [0091]; the means for predicting includes means for adaptively determining echo filter coefficients coupled to means for determining an estimate of the banded spectral amplitude metric of the noise, means for voice-activity detecting (VAD) using the estimate of the banded spectral amplitude metric of the mixed-down signal, and means for updating the filter coefficients based on the estimates of the banded spectral amplitude metric of the mixed-down signal and of the noise, and the previously predicted echo spectral content); determine a preliminary output signal based on the reference sound signal, the microphone sound signal, the preliminary echo cancelling coefficient (Fig.1, paragraphs [0091], [0098] and [0106]; a plurality of input signals, e.g., signals from a plurality of spatially separated microphones; and, for echo suppression, (b) one or more reference signals, e.g., signals from or to be rendered by one or more loudspeakers and that can cause echoes); and iteratively update the preliminary output signal to generate the first output signal, the iterative updating of the preliminary output signal (paragraphs [0039], [0045] and [0091] and [0098]; the means for predicting 117, 123, 125, 127 includes means for adaptively determining 125, 127 echo filter coefficients 128 coupled to means for determining 123 an estimate of the banded spectral amplitude metric of the noise 124, means for voice-activity detecting (VAD) using the estimate of the banded spectral amplitude metric of the mixed-down signal 122, and means for updating 127 the filter coefficients 128. The output of the VAD is coupled to means for updating and determined if the means for updating updates the filter coefficients. The filter coefficients are updated based on the estimates of the banded spectral amplitude metric of the mixed-down signal 122 and of the noise 124, and the previously predicted echo spectral content 118) including: update an updating index based on a predetermined step size, the reference sound signal, and an output signal determined in a preceding iteration (Fig.1, paragraphs [0039], [0129]- [0131]; overlapping and combining is dependent on the frame size, transform size and window functions, and should be designed to achieve a accurate reconstruction of the input signal in the absence of any processing or modification of the signal, X.sub.n, in the frequency domain); update the preliminary echo cancelling coefficient based on the updated updating index, the energy ratio, and a predetermined suppression depth (Abstract, paragraphs [0039], [0045], [0052] and [0091]; the banded signal 110 is a sufficiently accurate estimate of the banded spectral amplitude metric of the mixed-down signal 122, so that signal spectral estimator 121 is not used. The results of the VAD 125 are used by an adaptive filter updater 127 to determine whether to update the filter coefficients 128 based on the estimates of the banded spectral amplitude metric of the mixed-down signal 122 (or 110) and of the noise 124, and the previously predicted echo spectral content 118); and update the preliminary output signal based on the reference sound signal, the microphone sound signal, and the updated preliminary echo cancelling coefficient (Fig.1, paragraphs [0039], [0095]- [0098], [0111] and [0140]; output of the VAD is used by an adaptive filter updater to determine whether or not to update the filter coefficients, the updating based on the estimates of the banded spectral amplitude metric of the mixed-down signal and of the noise, and the previously predicted echo spectral content). Regarding claim 14, Dickins teaches the embedded system of claim 13, wherein the instructions further configure the embedded system to: determine that the energy ratio between the microphone sound signal and the reference sound signal is less than a ratio threshold and the reference sound signal is greater than an echo threshold; and increase a frequency of the iteratively updating of the preliminary output signal (paragraphs [0083], [0156] and [0177]- [0178]; suppression in order to facilitate suppression of residual noise to a constant power across frequency value relative to the hearing threshold. One suggested approach for normalization of the bands is to scale such that the 1 kHz band has unity energy gain from the input, and the other bands are scaled such that a noise source having a relative spectrum matching the threshold of hearing would be white or constant power across the bands. In some sense, this is a pre-emphasis filter on the bands prior to analysis which causes a drop in sensitivity in the lower and higher bands. This normalization is useful, since if the residual noise is controlled to be constant across the bands, this achieves a perceptually white noise when close to the hearing threshold). Regarding claim 15, Dickins teaches the embedded system of claim 13, wherein the instructions further configure the embedded system to: determine that the energy ratio between the microphone sound signal and the reference sound signal is less than a ratio threshold and the reference sound signal is greater than an echo threshold (paragraphs [0177]- [0178] and [0226]- [0227]; suppression in order to facilitate suppression of residual noise to a constant power across frequency value relative to the hearing threshold. One suggested approach for normalization of the bands is to scale such that the 1 kHz band has unity energy gain from the input, and the other bands are scaled such that a noise source having a relative spectrum matching the threshold of hearing would be white or constant power across the bands. In some sense, this is a pre-emphasis filter on the bands prior to analysis which causes a drop in sensitivity in the lower and higher bands. This normalization is useful, since if the residual noise is controlled to be constant across the bands, this achieves a perceptually white noise when close to the hearing threshold); compare the output signal determined in the preceding iteration with a threshold (paragraphs [0148], [0201] and [0211]; echo may cause an over-estimation of the noise component. For this reason, one embodiment of the invention includes echo-gated noise estimation: updating the noise estimate N.sub.b', and stopping the update of the noise estimate when the predicted echo level is significant compared with the previous noise estimate. That is, that noise estimator 123 provides an estimate which is gated when the predicted echo spectral content is significant compared to the previously estimated noise spectral content); in response to a comparison result that the output signal determined in the preceding iteration is greater than the threshold, increase the updating index based on the predetermined step size (paragraphs [0083] and [0177]; one suggested approach for normalization of the bands is to scale such that the 1 kHz band has unity energy gain from the input, and the other bands are scaled such that a noise source having a relative spectrum matching the threshold of hearing would be white or constant power across the bands. In some sense, this is a pre-emphasis filter on the bands prior to analysis which causes a drop in sensitivity in the lower and higher bands. This normalization is useful, since if the residual noise is controlled to be constant across the bands, this achieves a perceptually white noise when close to the hearing threshold); and in response to a comparison result that the output signal determined in the preceding iteration is less than the threshold, decrease the updating index based on the predetermined step size (paragraphs [0119] and [0129]-[0131]; overlapping and combining is dependent on the frame size, transform size and window functions, and should be designed to achieve an accurate reconstruction of the input signal in the absence of any processing or modification of the signal, X.sub.n, in the frequency domain). Regarding claim 16, Dickins teaches the embedded system of claim 12, wherein the instructions further configure the embedded system to: inject a noise signal to the second output signal to generate a third output signal (paragraphs [00061]-[0062]; determining banded spatial features from the plurality of sampled input signals; calculating a first set of suppression probability indicators, including an out-of-location suppression probability indicator determined using two or more of the spatial features, and a noise suppression probability indicator determined using an estimate of noise spectral content; combining the first set of probability indicators to determine a first combined gain for each band. The first combined gain, after post-processing if post-processing is included, forms a final gain for each band; and applying an interpolated final gain determined from the first combined gain. Interpolating the final gain produces final bin gains to apply to bin data of the mixed-down signal to form suppressed signal data). Regarding claim 17, Dickins teaches the embedded system of claim 12, wherein the instructions further configure the embedded system to: dynamically calculate a gain coefficient for each frame of the second output signal based on a maximum amplitude within the frame (Abstract and paragraphs [0040]-[0042]; gain calculator further calculates an additional echo suppression probability indicator, e.g., an echo suppression gain).; and apply the gain coefficient on the each frame of the second output signal to generate a third output signal (Abstract and paragraphs [0040]-[0042]; gain calculator further calculates an additional echo suppression probability indicator, e.g., an echo suppression gain). Regarding claim 18, Dickins teaches the embedded system of claim 12, wherein the instructions further configure the embedded system to: provide a graphical user interface; receive user configurations via the graphical user interface; and adjust parameters of the embedded system based on the user configurations (paragraphs [0062] and [0172]; a set of parameters and using: an estimate of noise spectral content, the banded frequency domain amplitude metric representation of the echo, and the banded spatial features, the set of parameters including whether the estimate of noise spectral content is spatially selective or not, which indication of voice activity an instantiation determines being controlled by a selection of the parameters, voice activity). Regarding claim 19, Dickins teaches the embedded system of claim 12, wherein the instructions further configure the embedded system to: determine a size of data buffered in the embedded system (paragraphs [0123] and [0127]; discrete finite length Fourier transform, such as implemented by the FFT, is often referred to as a circulant transform due to the implicit assumption that the signal in the transform window is in some way periodic or repetitive. Most general forms of circulant transforms can be represented by buffering, a window, a twist (real value to complex value transformation) and a DFT, e.g., FFT. An optional complex twist after the DFT can be used to adjust the frequency domain representation to match specific transform definitions); estimate a delay time based on the size of data buffered in the embedded system; and shift the reference sound signal by the delay time to align the reference sound signal with the microphone sound signal (paragraphs [0139] and [0149]; a microphone inputs to effect appropriate time delay to be used in the mixing down). Claim 20 is a non-transitory computer-readable storage medium claim correspond to method claim 1. Therefore, claim 20 has been analyzed and rejected based on method claim 1. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. U.S Pub. No. 2023/0094054 A1 to Li et al. discloses performing an echo cancellation process on a voice input signal according to an echo reference signal to obtain an echo cancellation signal; performing a FFT on the echo reference signal to obtain a reference spectrum signal for each frame; performing the FFT on the echo cancellation signal to obtain a speech spectrum signal for each frame; using the reference spectrum signal and the speech spectrum signal of a current frame to obtain a priori signal-to-noise ratio of the current frame according to a principle of additive noise; filtering the speech spectrum signal of the current frame by a Wiener filter coefficient of the current frame determined by the prior signal-to-noise ratio of the current frame to obtain a target spectrum signal of each frame; performing an IFFT on the target spectrum signal of each frame to obtain a target voice signal (Abstract). Any inquiry concerning this communication or earlier communications from the examiner should be directed to AKELAW A TESHALE whose telephone number is (571)270-5302. The examiner can normally be reached 9 am -6pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, AHMAD MATAR can be reached at (571)272-7488. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. AKELAW TESHALE Primary Examiner Art Unit 2694 /AKELAW TESHALE/Primary Examiner, Art Unit 2694
Read full office action

Prosecution Timeline

Oct 02, 2024
Application Filed
Aug 20, 2026
Non-Final Rejection mailed — §102 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12745022
DEVICE FOR DETECTION OF A CLIPPED RINGING SIGNAL
2y 2m to grant Granted Sep 22, 2026
Patent 12726565
DERIVING UPDATES TO AN EMERGENCY USER PROFILE FROM COMMUNICATIONS ASSOCIATED WITH AN EMERGENCY INCIDENT
2y 8m to grant Granted Sep 01, 2026
Patent 12719982
FORECASTING CALL QUALITY USING MACHINE LEARNING TECHNIQUES
2y 7m to grant Granted Aug 25, 2026
Patent 12707009
GENERATIVE AND ADAPTIVE MEDIATOR FOR REAL-TIME INTERACTIONS WITH CONVERSATIONAL AGENTS
2y 5m to grant Granted Aug 11, 2026
Patent 12707010
SYSTEM AND METHOD FOR DUAL-DEVICE COMMUNICATION SYNCHRONIZATION
2y 3m to grant Granted Aug 11, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
82%
Grant Probability
98%
With Interview (+16.0%)
2y 10m (~9m remaining)
Median Time to Grant
Low
PTA Risk
Based on 864 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month