Prosecution Insights
Last updated: August 17, 2026
Application No. 18/494,302

METHOD AND APPARATUS FOR NEURAL NETWORK AUGMENTED KALMAN FILTER FOR ACOUSTIC HOWLING SUPPRESSION

Non-Final OA §103§112
Filed
Oct 25, 2023
Examiner
GERMICK, JOHNATHAN R
Art Unit
4100
Tech Center
4100
Assignee
Tencent Technology (Shenzhen) Company Limited
OA Round
1 (Non-Final)
46%
Grant Probability
Moderate
1-2
OA Rounds
1y 9m
Est. Remaining
76%
With Interview

Examiner Intelligence

Grants 46% of resolved cases
46%
Career Allowance Rate
46 granted / 101 resolved
-14.5% vs TC avg
Strong +30% interview lift
Without
With
+30.1%
Interview Lift
resolved cases with interview
Typical timeline
4y 7m
Avg Prosecution
28 currently pending
Career history
124
Total Applications
across all art units

Statute-Specific Performance

§101
28.5%
-11.5% vs TC avg
§103
39.2%
-0.8% vs TC avg
§102
17.1%
-22.9% vs TC avg
§112
14.3%
-25.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 101 resolved cases

Office Action

§103 §112
DETAILED ACTION This action is responsive to the Claims filed 10/25/2023. Claims 1-20 are pending in the case. Claims 1, 11 and 20 are independent claims. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 5 and 15 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claims 5 and 15 recite the limitation "the observation matrix estimation". There is insufficient antecedent basis for this limitation in the claim. Examiner notes the parent claims recite “the observations covariance estimation”. For the purposes of compact prosecution “observation matrix estimation” is understood to refer to “the observations covariance estimation”. Claim Rejections - 35 U.S.C. § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. §§ 102 and 103 (or as subject to pre-AIA 35 U.S.C. §§ 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. § 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1, 8-10, 11, 18-20 are rejected under 35 U.S.C. § 103 as being unpatentable over Wan et al. “Acoustic Howling Suppression with Neural Networks using Real-world Data”, further in view of Yu et al. “A Deep Neural Network Based Kalman Filter for Time Domain Speech Enhancement” Claim 1/11/20 Wan teaches, A method performed by at least one processor of an acoustic howling suppression (AHS) system, the method comprising: [from claim 11] An acoustic howling suppression (AHS) system, comprising: at least one memory configured to store program code; and at least one processor configured to read the program code and operate as instructed by the program code, the program code including: [from claim 20] A non-transitory computer readable medium having instructions stored therein, which when executed by a processor in an acoustic howling suppression (AHS) system cause the processor to execute a method comprising: (pg 1 “In this work, to conduct howling suppression using deep learning, we first collected 36 hours of howling sound as the training and test sets…” pg 3 “The structure diagram of Wave-U-Net used in this work is provided in Figure 4… with each successive level running at half the time resolution of the previous level, which reduces memory consumption.” PHOSITA would understand such a model oriented to reducing memory consumption is run on a processor with memory.) receiving, from an input source device, an audio signal; (pg 2 Section A “We use the SP3874AN sound as the howling source, and the artificial mouth is sited in front of the microphone,”)filtering the audio signal …to reduce acoustic howling included in the audio signal (pg 1 “we first collected 36 hours of howling sound as the training and test sets …After that, several representing networks of CRN, Wave-U-NET, and DCCRN are utilized to suppress the howling” neural network are used to suppress the howling in the sound or audio, corresponding to the claim requiring filtering to reduce acoustic howling.) Wan does not explicitly teach, refining one or more parameters of a Kalman filter based on one or more neural networks; and… using the Kalman filter with the one or more refined parameters of the Kalman filter Yu however when addressing refining Kalman filter parameters with neural networks for audio enhancement teaches, refining one or more parameters of a Kalman filter based on one or more neural networks; and… using the Kalman filter with the one or more refined parameters of the Kalman filter (pg 1 abstract “In this paper, we present a novel deep neural network (DNN) based Kalman filter (KF) algorithm for speech enhancement, where DNN is applied for estimating key parameters in the KF, namely, the linear prediction coefficients (LPCs).” Section 3 pg 2 “The overall block diagram of our DNN based SE system with KF is depicted in Fig.1…. In the enhancement stage, a KF with the the DNN-based estimated parameters is applied to the noisy speech to obtain the enhanced speech.”) Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the audio filtering based neural network of Wan to comprise a Kalman based filter informed by neural network parameter estimates as described by Yu. One would have been motivated to make such a combination because Wan and Yu are directed to audio filtering and enhancement using neural networks. Further, Yu notes “our proposed DNN KF algorithm is able to estimate LPCs from noisy speech more accurately and robustly, leading to an improved performance” Claim 8/18 Wan/Yu teaches claim 1/11 Further Wan teaches, determining whether acoustic howling is detected in the audio signal… based on a determination that the acoustic howling is detected, stopping training of the one or more neural networks. ( pg 2 “To train the network, we create a totally 7812, 2232, 1116 howling-clean pairs in the ratio of 7:2:1 for training, validation, and test, respectively, with signal-to-howling ratio (SHR) ranging from-5 dB to 20 dB. To test the model, six different SHRs are used for model evaluation, which are {−5,0,5,10,15,20} dB, respectively” pg 4 Conclusion “we collected the howling dataset, and three different representing networks were utilized for training. The results show that the deep learning-based method can effectively suppress the howling” the neural network detects the suppression of howling, the training is stopped based in part on degree of howling detection (i.e the SHR) during the training process.) Claim 9/19 Wan/Yu teaches claim 1/11 Further Wan when combined with Yu teaches, herein the acoustic howling is detected based on a determination that an amplitude of an output of the Kalman filter signal exceeds an amplitude threshold for a predetermined period. (pg 2 section 2 “we create a totally 7812, 2232, 1116 howling-clean pairs in the ratio of 7:2:1 for training, validation, and test, respectively, with signal-to-howling ratio (SHR) ranging from-5 dB to 20 dB. To test the model, six different SHRs are used for model evaluation, which are {−5,0,5,10,15,20} dB, respectively.” Pg 3 “The loss function of DCCRN is SI-SNR, which has been widely used to replace the mean square error(MSE).SI-SNR is defined as: PNG media_image1.png 87 304 media_image1.png Greyscale ” the output S, which when combined with Yu is the output of the Kalman filters (i.e the cleaned audio signal), compared to the target signal (via the difference function), which is a detection of amplitude of the howling/noise. The magnitude of e_noise is the degree to which the amplitude of the noise exceeds the threshold target signal over a sampling period) Claim 10 Wan/Yu teaches claim 1 Wan teaches, wherein the input source device is a microphone that receives first audio from a user of the microphone and second audio from an amplifier. (pg 2 Section A “We use the SP3874AN sound as the howling source, and the artificial mouth is sited in front of the microphone,” Also see Figure 1. The experimenters are the users of the microphone, the sound source is an amplifier.) Claim(s) 2, 4, 6, 12, 14 and 16 are rejected under 35 U.S.C. § 103 as being unpatentable over Wan/Yu, further in view of Chen et al “DynaNet: Neural Kalman Dynamical Model for Motion Estimation and Prediction” Claim 2/12 Wan/Yu teaches claim 1/11 Wan/Yu does not explicitly teach, wherein the one or more parameters comprises a reference signal of the Kalman filter, and wherein the one or more neural networks includes a first neural network that refines the reference signal of the Kalman filter based on the audio signal to generate a learned reference signal. Chen when addressing a Kalman net conditioned with a neural network system teaches and combined with Wan/Yu, wherein the one or more parameters comprises a reference signal of the Kalman filter (Figure 2 pg 4 “ PNG media_image2.png 378 712 media_image2.png Greyscale caption “DynaNet framework consists of the neural observation model to extract latent states z, the neural transition model to generate the evolving relation A, and a recursive Kalman filter to infer” the neural transition model in part outputs a transition matrix which conditions the Kalman filter, therefore considered a reference signal. In fact, each input to the Kalman filter is a reference signal of some kind as claimed it conditions the Kalman filter.) and wherein the one or more neural networks includes a first neural network that refines the reference signal of the Kalman filter based on the audio signal to generate a learned reference signal. (Figure 3 pg 4-5 “In this deterministic transition model, the dependence of the transition matrix on historical latent states is specified by an RNN. This RNN recursively processes previous hidden states(zt−1,ht−1) of the dynamic model and LSTM” the deterministic transition model generates the transition matrix A which is the reference signal of the Kalman filter, this network generates the learned output signal A, when combined with the neural network for audio processing of Wan/Yu is understood to describe an audio signal.) Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the audio filtering based neural network of Wan/Yu to comprise a set of neural networks for generating Kalman network parameters to improve the Kalman filtering as described by Chen. One would have been motivated to make such a combination because Wan/Yu and Chen are directed to filtering with neural network conditioned Kalman networks. Further, Chen notes “Through deeply coupled DNNs and SSMs, DynaNet can scale to high-dimensional data as well as model very complex motion dynamics in real world. By using the Kalman filter on feature space, DynaNet is able to reason about latent system states, allowing reliable inference and predictions even with missing observations.” (Conclusion pg 11) Claim 4/14 Wan/Yu/Chen teaches claim 2/12 Further Chen teaches, wherein the one or more parameters includes an observation covariance matrix estimation, and wherein the one or more neural networks includes a second neural network that refines the observation covariance matrix estimation using the learned reference signal of the Kalman filter. ( pg 3 “In our neural emission model, an encoder f_encoder is used to extract both features at and an estimation of uncertainty σa t from the observations xt at timestep… The coupled uncertainties σ represent the measurement belief that is transformed into the observation noise R in a Kalman filter… a and σ are further used in the update stage of a differentiable Kalman filter” Figure 2 caption “and Q and R are process noise matrix and observation noise matrix, respectively.” Here R is generated in part by the second neural network, the encoder, it is a matrix estimation of the covariance of observation noise. Using learned reference signals of the Kalman filter, here the inputs to the Kalman filter are considered the learned reference signals. Further, Examiner notes that, mathematically, to PHOSITA, the noise parameters of a Kalman filter are defined as covariance matrices. Just as neural networks definitionally input and output matrices. ) Claim 6/16 Wan/Yu/Chen teaches claim 4/14 Further Chen teaches, wherein the one or more parameters includes a noise covariance estimation, and wherein the one or more neural networks includes a third neural network that refines the noise covariance estimation using the learned reference signal of the Kalman filter. ( pg 6 “In contrast to hand-tuning process noise Q and measurement noise R in a conventional KF, these two terms are jointly learned by our proposed neural dynamical model”, Figure 2 caption “and Q and R are process noise matrix and observation noise matrix, respectively.” As shown in the figure the process noise matrix which is the estimated covariance of noise using a learned reference signal is generated by the neural transition component of the model for estimated Q. The portion of the joint neural network responsible for estimating P is considered the third neural network.) Claim(s) 3, 5, 7, 13, 15, 17 are rejected under 35 U.S.C. § 103 as being unpatentable over Wan/Yu/Chen, further in view of, Lei et al. “A hybrid model based on deep LSTM for predicting high-dimensional chaotic systems” Claim 3/13 Wan/Yu/Chen teaches claim 2/12 Further Chen teaches, wherein the learned reference signal is generated based on a … long short-term memory (LSTM) network applied to the audio signal and the reference signal of the Kalman filter. (pg 4 Section 3 B “The deterministic transition and resampled transition are two individual strategies to generate a transition matrix. The parameters of long short-term memory (LSTM) in these two modules are separately learned from data.” When combined with Wan/Yu the transition matrix is a references signal of the applied audio signal, the LSTM neural network is applied to the audio signal and reference signal. As noted previously this output applied to the Kalman filter, thus of the Kalman filter as claimed. ) Wan/Yu/Chen does not explicitly teach, two layer long short-term memory Lei teaches when addressing multi layer LSTMs, two layer long short-term memory (pg 6 Conclusion “a hybrid method based on the deep LSTM with multi-layers to predict”) Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify LSTM networks of Wan/Yu/Chen to comprise utilize multi layer LSTMs. One would have been motivated to make such a combination because as noted by Lei “Under the evaluation criteria, the prediction performance can be sorted as: single-layer LSTM < multi-layer LSTM < hybrid single-layer LSTM < hybrid multi-layer LSTM… increasing the depth of LSTM model can better the prediction performance to some extent” (pg 6 Lei) Claim 5/15 Wan/Yu/Chen teaches claim 4/14 Further Chen teaches, wherein the observation matrix estimation is generated based on a … network applied to an output of the Kalman filter. ( pg 3 “In our neural emission model, an encoder f_encoder is used to extract both features at and an estimation of uncertainty σa t from the observations xt at timestep” as previously noted the observations generated by the encoder network corresponding to the claims, where the observations are applied to the output step of the Kalman filters shown in figure 2) Wan/Yu/Chen does not explicitly teach, two layer long short-term memory Lei teaches when addressing multi layer LSTMs, two layer long short-term memory (pg 6 Conclusion “a hybrid method based on the deep LSTM with multi-layers to predict”) Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify LSTM networks of Wan/Yu/Chen to comprise utilize multi layer LSTMs. One would have been motivated to make such a combination because as noted by Lei “Under the evaluation criteria, the prediction performance can be sorted as: single-layer LSTM < multi-layer LSTM < hybrid single-layer LSTM < hybrid multi-layer LSTM… increasing the depth of LSTM model can better the prediction performance to some extent” (pg 6 Lei). Further Lei notes “Reserver Computing (RC), which uses the advantages of data and system dynamic structure, and pointed out that this method is also applicable to other machine learning models1. In fact, both the RC and the LSTM model are derived from general recurrent neural network …In practice, we can use all of the resources, data and prerequisite knowledge, and combine them to achieve an optimal prediction result” Claim 7/17 Wan/Yu/Chen teaches claim 4/14 Further Chen teaches, wherein noise covariance estimation is generated based on a … short-term memory (LSTM) network … pg 4 Section 3 B “The deterministic transition and resampled transition are two individual strategies to generate a transition matrix. The parameters of long short-term memory (LSTM) in these two modules are separately learned from data.”… PNG media_image3.png 396 496 media_image3.png Greyscale as shown in the figure the process noise estimate, corresponding to noise covariance estimation, is based on a LSTM network module of learned transitions) Further Wan teaches, applied to an echo path. ( pg 1 “Due to the closed loop between the microphone and loudspeaker, howling occurs in scenarios such as meetings, KTV, and smart classroom…. Because of the positive feedback between the microphone and loudspeaker, howling happens” pg 2 “This section introduces the howling dataset collection method and the related audio equipment used to record the howling dataset. PNG media_image4.png 306 375 media_image4.png Greyscale ” PHOSITA would recognize that the recording system created to create the howling data set is an echo path. positive audio feedback is when a microphone picks up a singal which is output by an amplifier and picked up again, otherwise known as echo, as described in specification paragraph 0002.) Wan/Yu/Chen does not explicitly teach, two layer long short-term memory Lei teaches when addressing multi layer LSTMs, two layer long short-term memory (pg 6 Conclusion “a hybrid method based on the deep LSTM with multi-layers to predict”) Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify LSTM networks of Wan/Yu/Chen to comprise utilize multi layer LSTMs. One would have been motivated to make such a combination because as noted by Lei “Under the evaluation criteria, the prediction performance can be sorted as: single-layer LSTM < multi-layer LSTM < hybrid single-layer LSTM < hybrid multi-layer LSTM… increasing the depth of LSTM model can better the prediction performance to some extent” (pg 6 Lei). Further Lei notes “Reserver Computing (RC), which uses the advantages of data and system dynamic structure, and pointed out that this method is also applicable to other machine learning models1. In fact, both the RC and the LSTM model are derived from general recurrent neural network …In practice, we can use all of the resources, data and prerequisite knowledge, and combine them to achieve an optimal prediction result” Conclusion Prior art not relied upon: Yixuan Zhang, Hao Zhang, Meng Yu, Dong Yu “neural network augmented Kalman filter for robust acoustic howling suppression” discusses the claimed multi neural network Kalman Gain determination for howling suppression. Any inquiry concerning this communication or earlier communications from the examiner should be directed to JOHNATHAN R GERMICK whose telephone number is (571)272-8363. The examiner can normally be reached M-F 9:30-4:30. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kakali Chaki can be reached on 571-272-3719. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /J.R.G./ Examiner, Art Unit 2122 /KAKALI CHAKI/Supervisory Patent Examiner, Art Unit 2122
Read full office action

Prosecution Timeline

Oct 25, 2023
Application Filed
Jul 15, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12699893
SELF-SUPERVISED REPRESENTATION LEARNING USING BOOTSTRAPPED LATENT REPRESENTATIONS
5y 2m to grant Granted Aug 04, 2026
Patent 12694266
CORRELATION RECURRENT UNIT FOR IMPROVING PREDICTION PERFORMANCE OF TIME-SERIES DATA AND CORRELATION RECURRENT NEURAL NETWORK
3y 7m to grant Granted Jul 28, 2026
Patent 12670369
CENTRAL SCHEDULER AND INSTRUCTION DISPATCHER FOR A NEURAL INFERENCE PROCESSOR
8y 2m to grant Granted Jun 30, 2026
Patent 12664441
METHOD AND SYSTEM FOR IDENTIFYING RELEVANT VARIABLES
5y 5m to grant Granted Jun 23, 2026
Patent 12645917
NEURAL ARCHITECTURE FOR SELF SUPERVISED EVENT LEARNING AND ANOMALY DETECTION
6y 11m to grant Granted Jun 02, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
46%
Grant Probability
76%
With Interview (+30.1%)
4y 7m (~1y 9m remaining)
Median Time to Grant
Low
PTA Risk
Based on 101 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month