Prosecution Insights
Last updated: August 17, 2026
Application No. 18/955,791

AUDIO NOISE DETERMINATION USING ONE OR MORE NEURAL NETWORKS

Non-Final OA §103
Filed
Nov 21, 2024
Priority
May 14, 2020 — continuation of 11/678,120 +1 more
Examiner
PATEL, YOGESHKUMAR G
Art Unit
2691
Tech Center
2600 — Communications
Assignee
NVIDIA Corporation
OA Round
1 (Non-Final)
83%
Grant Probability
Favorable
1-2
OA Rounds
6m
Est. Remaining
87%
With Interview

Examiner Intelligence

Grants 83% — above average
83%
Career Allowance Rate
563 granted / 675 resolved
+21.4% vs TC avg
Minimal +3% lift
Without
With
+3.2%
Interview Lift
resolved cases with interview
Typical timeline
2y 3m
Avg Prosecution
12 currently pending
Career history
683
Total Applications
across all art units

Statute-Specific Performance

§101
5.2%
-34.8% vs TC avg
§103
68.6%
+28.6% vs TC avg
§102
12.2%
-27.8% vs TC avg
§112
11.6%
-28.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 675 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims 1-20 are cancelled. Claims 21-40 are new and pending. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 21-30, 32-37, and 40 is/are rejected under 35 U.S.C. 103 as being unpatentable over Calle et al. (US PGPUB #2018/0358003) in view of Lee et al. (US PGPUB #2018/0190268). Regarding Claim 21, Calle discloses one or more processors (title, abstract, figs. 1-12; ¶0024), comprising: circuitry to use one or more neural networks (Calle ¶0016, fig. 7) to generate one or more denoised audio signals (Calle ¶0044 discloses the UE 310 can include a noise filter/ suppression, beam-forming component 312 that filters or suppresses noise and performs beam forming on the speech signal picked up by one or more microphones of the UE 310. ¶0069 discloses a high-fidelity speech model for multiple voices [e.g., voice biometrics] can be learned to increase speech quality, a deep learning-based voice discriminator and a voice activity detector can be learned to detect and discriminate a voice signal [e.g., in low signal-to-noise ratio (SNR)], a directional beam former function can be learned to localize each voice of a plurality of multiple voices, a neural network can be trained to recover the accurate speech signal output by reducing room echo and channel problems [e.g., transcoding problems]. ¶0075 discloses the voice discriminator is a voice activity detector that detects and discriminates voice signal in low SNR, triggers on voice/speech, and rejects detected environmental noise. ¶0077 discloses a talker's voice embeddings [or voice model] can be captured, learned, and updated on-device. The system can be robust from noise effects in the local environment) based, at least in part, on using one or more convolutional portions of the one or more neural networks (Calle ¶0074 discloses a generative temporal convolutional auto-encoder neural network can be used to learn and then generate high fidelity voice. In one configuration, a temporal network can be substituted with a 3D neural network, clockwork network [or RNN] implementation. In one configuration, a temporal network can be used to handle voice aging and temporal envelope effects of different voices. ¶0078: fig. 9) to identify one or more first portions of one or more audio signals (Calle ¶0074 discloses raw speech from the transmit side can be detected and captured, and cleaner high fidelity speech output can be reconstructed [either optimized for human listening fidelity, or optimized for speech recognition fidelity]). Calle may not explicitly disclose using one or more recurrent portions of the one or more neural networks, parallel to the one or more convolutional portions, to identify one or more second portions of the one or more audio signals. However, Lee (title, abstract, figs. 1-11) teaches using one or more recurrent portions of the one or more neural networks, parallel to the one or more convolutional portions, to identify one or more second portions of the one or more audio signals (Lee ¶0090 discloses in the illustrated neural network example of fig. 4, the weight determiner can be implemented through at least one neural network layer of the illustrated neural network that can have a separate input layer from an input layer that receives the speech signal Vt input or can be implemented through layers of the neural network that are configured in parallel with connections between the input of the neural network and the example layer 415. ¶0114 discloses operations of fig. 9 can be performed in parallel or simultaneously. ¶0106 disclose there can also be convolutional connections, bias connections, and other neural network contextual or memory connections, as only examples. figs. 4-5 and 7-8). Calle and Lee are analogous art as they pertain to determining audio noise using neural network. Therefore it would have been obvious to someone of ordinary skill in the art before the effective filing date of the invention was made to modify improving speech (as taught by Calle) to simultaneously input a plurality of speech frames to the speech recognizing model to determine an attention weight of each of the speech frames input to the speech recognizing model and apply the determined attention weights to each of the speech frames corresponding to the determined attention weights (as taught by Lee, ¶0085) to implement through a neural network that can have a separate input layer from an input layer that receives the speech signal input or to implement through layers of the neural network that are configured in parallel (Lee, ¶0090). Regarding Claim 22, Calle in view of Lee discloses the one or more processors of claim 21, wherein the one or more neural networks further comprise one or more second recurrent portions in series with the one or more convolutional portions and the one or more recurrent portions (Calle ¶0026 discloses a recurrent architecture can be helpful in recognizing patterns that span more than one of the input data chunks that are delivered to the neural network in a sequence). Regarding Claim 23, Calle in view of Lee discloses the one or more processors of claim 21, wherein the one or more neural networks are to generate the one or more denoised audio signals (Calle fig. 3-4: 312, 412: noise filter/suppression. beam forming) based, at least in part, on one or more audio spectrograms representing the one or more audio signals (Calle ¶0047 discloses the ASR can be constructed a neural network with convolutional layers acting on speech features, including MFCC [Mel-Frequency Cepstral Coefficients], spectrogram and gammatone features, or conceivably on the audio signal itself, given sufficient processing power. In addition, the ASR can contain various RNN layers including bi-direction RNN. Examples of specialized RNNs include LSTM [Long Short-Term Memory] units and GRU [Gated Recurrent Units], which can further be configured to process incoming data front-to-back, or in the case of buffered data, both front-to-back and back-to-front, creating a so called bidirectional RNN networks that is known to improve accuracy; fig. 2). Regarding Claim 25, Calle in view of Lee discloses the one or more processors of claim 21. But Calle may not explicitly disclose wherein the one or more neural networks are to generate the one or more denoised audio signals based, at least in part, on generating an audio mask based, at least in part, on the identified one or more first portions and one or more second portions. However, Lee (title, abstract, figs. 1-11) teaches wherein the one or more neural networks are to generate the one or more denoised audio signals based, at least in part, on generating an audio mask based, at least in part, on the identified one or more first portions and one or more second portions (Lee ¶0069 discloses speech recognizing model implemented by the speech recognizing apparatus 110 configured as the neural network may dynamically implement spectral masking by receiving a feedback on a result calculated by the neural network at the previous time. When the spectral masking is performed, feature values for each frequency band can selectively not be used in full as originally determined/captured, but rather, a result of a respective adjusting of the magnitudes of all or select feature values for all or select frequency bands, e.g., according to the dynamically implemented spectral masking, can be used for or within speech recognition. Also, for example, such a spectral masking scheme can be dynamically implemented to intensively recognize a speech of a person other than noise from a captured speech signal and/or to intensively recognize a speech of a particular or select speaker to be recognized when plural speeches of a plurality of speakers are present in the captured speech signal). Calle and Lee are analogous art as they pertain to determining audio noise using neural network. Therefore it would have been obvious to someone of ordinary skill in the art before the effective filing date of the invention was made to modify improving speech (as taught by Calle) to configure the neural network to dynamically implement spectral masking by receiving a feedback on a result calculated by the neural network at the previous time (as taught by Lee, ¶0069) to implement through a neural network that can have a separate input layer from an input layer that receives the speech signal input or to implement through layers of the neural network that are configured in parallel (Lee, ¶0090). Regarding Claim 26, Calle in view of Lee discloses the one or more processors of claim 21, wherein the one or more convolutional portions of the one or more neural networks are to identify at least one or more spatial patterns of the one or more audio signals (Calle ¶0027 discloses a convolutional network 106 can be locally connected, and is further configured such that the connection strengths associated with the inputs for each neuron in the second layer are shared [e.g., connection strength 108]. More generally, a locally connected layer of a network may be configured so that each neuron in a layer will have the same or a similar connectivity pattern, but with connections strengths that may have different values [e.g., 110, 112, 114, and 116]. The locally connected connectivity pattern can give rise to spatially distinct receptive fields in a higher layer, because the higher layer neurons in a given region may receive inputs that are tuned through training to the properties of a restricted portion of the total input to the network). Regarding Claim 27, Calle in view of Lee discloses the one or more processors of claim 21, wherein the one or more recurrent portions of the one or more neural networks are to identify at least one or more temporal patterns of the one or more audio signals (Calle ¶0069 discloses temporal network can be substituted with clockwork network [or recurrent neural network (RNN)] to handle voice aging and temporal effects of different voices. ¶0080 disclose the CNN can learn the temporal features [time-based features] distributed spatially in the CNN. In one configuration, instead of the temporal CNN, an RNN or clockwork CNN may be used to reconstruct the voice). Claims 28-30 and 32-37 are rejected for the same reasons as set forth in Claims 21-23 and 25-27. Regarding Claim 40, Calle in view of Lee discloses the system of claim 35, wherein the one or more processors are comprised in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations (Calle fig. 8: deep learning based VAD/SVA 804 [0037]. ¶0011 discloses fig. 2 is a block diagram illustrating an exemplary Deep Convolutional Network [DCN]. ¶0033 discloses DCNs are networks of convolutional networks, configured with additional pooling and normalization layers. DCNs can achieve state-of-the-art performance on many tasks. DCNs can be trained using supervised learning in which both the input and output targets are known for many exemplars and are used to modify the weights of the network by use of gradient descent methods); a system for performing remote operations (Calle ¶0057 discloses at 502 the UE can receive a first voice stream from a remote UE [e.g., the UE 310 or 410]; figs. 3-5); a system for performing real-time streaming (Calle fig. 5: 506 construct, by using a neural network, a second voice stream based on the first voice stream. ¶0060 discloses the UE can identify [e.g., through classification] in real time the voice of the user speaking in the first voice stream. That way the method can pull up appropriate user models based on who is speaking. The classification technique can be based on a neural network that detects the particular voice features. For example, a first person is talking on the phone, the first person can put a second person on the phone, and the voice model switches to the second person's voice. Also refer to fig. 11: voice stream); a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for performing conversational AI operations; a system implementing one or more multi-model language models (MMLM); a system implementing one or more large language models (LLMs); a system implementing one or more small language models (SLMs); a system implementing one or more vision language models (VLMs); a system for generating synthetic data; a system for generating synthetic data using AI; a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources (Calle fig. 4: 402; cloud service). Claims 24 and 31 is/are rejected under 35 U.S.C. 103 as being unpatentable over Calle et al. (US PGPUB #2018/0358003) in view of Lee et al. (US PGPUB #2018/0190268) further in view of Korjani (US Patent #10614827). Regarding Claim 24, Calle in view of Lee discloses the one or more processors of claim 21, but may not explicitly disclose wherein the one or more neural networks further comprise one or more portions to concatenate the one or more first portions and the one or more second portions of the one or more audio signals. However, Korjani (title, abstract, figs. 1-3) teaches wherein the one or more neural networks further comprise one or more portions to concatenate the one or more first portions and the one or more second portions of the one or more audio signals (Korjani col. 3, lines 43-49 discloses the frames of user speech, after removal of the noise, are concatenated and converted by the waveform reconstruction module 158 as needed into an audio file or audio stream that may then be played back to the user via a speaker 160 on the user phone. col. 4, lines 38-46 discloses once noise is isolated and removed from all the frames, the frames of clean speech are assembled or otherwise reconstructed 350 into a complete waveform). Calle, Lee, and Korjani are analogous art as they pertain to determining audio noise using neural network. Therefore it would have been obvious to someone of ordinary skill in the art before the effective filing date of the invention was made to modify teachings of Calle in view of Lee in light of the teachings of Korjani to train the neural network only to isolate noise, and the noise profile that is estimated for each frame consists of the noise estimate alone (as taught by Korjani, col. 1, lines 65-67) since the neural network is not trained to estimate speech in the speech data, the noise filter accurately removes noise independent of the language in which the user utters the speech data (Korjani, col. 1, line 67 to col. 2, line 3). Claim 31 is rejected for the same reasons as set forth in Claim 24. Claims 38-39 is/are rejected under 35 U.S.C. 103 as being unpatentable over Calle et al. (US PGPUB #2018/0358003) in view of Lee et al. (US PGPUB #2018/0190268) further in view of Tommy et al. (US #2021/0118462). Regarding Claim 38, Calle in view of Lee discloses the system of claim 35, wherein the one or more convolutional portions of the one or more neural networks are to identify at least one or more of: fundamental frequencies, or speech, harmonics (Calle ¶0028 discloses in a spectral image certain neurons may focus on fundamental frequencies of human voice other neurons may learn the relationship between harmonics). Calle in view of Lee may not explicitly disclose wherein the one or more convolutional portions of the one or more neural networks are to identify at least one or more of: noise patterns. However, Tommy (title, abstract, figs. 1-3) teaches wherein the one or more convolutional portions of the one or more neural networks are to identify at least one or more of: noise patterns (Tommy ¶0033 discloses the form of the noise signal pattern generated cannot be determined using signal processing techniques. The DNN is trained to understand these behaviors of noises, their patterned behavior or their momentary nature with time and their spectral frequency deviations or differences. ¶0037 discloses the DNN is trained to learn about the behavior pattern, and accordingly the DNN adjusts the weights of neuron to suppress the given noises. This exhibits the intelligent behavior of the DNN, which allows it to provide degree to which it should be reduced, based on the noise impact or involvement). Calle, Lee, and Tommy are analogous art as they pertain to determining audio noise using neural network. Therefore it would have been obvious to someone of ordinary skill in the art before the effective filing date of the invention was made to modify teachings of Calle in view of Lee in light of the teachings of Tommy to understand the behaviors of noises, their patterned behavior or their momentary nature with time and their spectral frequency deviations or differences (Tommy, ¶0033) identify a plurality of noises or their combination to suppress them and enhance the deteriorated input signal in a dynamic manner (Tommy, ¶0002). Regarding Claim 39, Calle in view of Lee discloses the system of claim 35, wherein the one or more recurrent portions of the one or more neural networks are to identify at least one or more of: fundamental frequencies, speech, or harmonics (Calle ¶0028 discloses in a spectral image certain neurons may focus on fundamental frequencies of human voice other neurons may learn the relationship between harmonics). Calle in view of Lee may not explicitly disclose wherein the one or more recurrent portions of the one or more neural networks are to identify at least one or more of: noise patterns. However, Tommy (title, abstract, figs. 1-3) teaches wherein the one or more recurrent portions of the one or more neural networks are to identify at least one or more of: noise patterns (Tommy ¶0033 discloses the form of the noise signal pattern generated cannot be determined using signal processing techniques. The DNN is trained to understand these behaviors of noises, their patterned behavior or their momentary nature with time and their spectral frequency deviations or differences. ¶0037 discloses the DNN is trained to learn about the behavior pattern, and accordingly the DNN adjusts the weights of neuron to suppress the given noises. This exhibits the intelligent behavior of the DNN, which allows it to provide degree to which it should be reduced, based on the noise impact or involvement). Calle, Lee, and Tommy are analogous art as they pertain to determining audio noise using neural network. Therefore it would have been obvious to someone of ordinary skill in the art before the effective filing date of the invention was made to modify teachings of Calle in view of Lee in light of the teachings of Tommy to understand the behaviors of noises, their patterned behavior or their momentary nature with time and their spectral frequency deviations or differences (Tommy, ¶0033) identify a plurality of noises or their combination to suppress them and enhance the deteriorated input signal in a dynamic manner (Tommy, ¶0002). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to YOGESHKUMAR G PATEL whose telephone number is (571)272-3957. The examiner can normally be reached 7:30 AM-4 PM PST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Duc Nguyen can be reached at (571) 272-7503. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /YOGESHKUMAR PATEL/Primary Examiner, Art Unit 2691
Read full office action

Prosecution Timeline

Nov 21, 2024
Application Filed
Aug 04, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12706578
Zone Volume Control
2y 5m to grant Granted Aug 11, 2026
Patent 12701355
APPARATUS AND METHOD FOR NARROWBAND DIRECTION-OF-ARRIVAL ESTIMATION
2y 4m to grant Granted Aug 04, 2026
Patent 12701356
ELECTRONIC DEVICE AND OPERATION METHOD THEREOF
2y 4m to grant Granted Aug 04, 2026
Patent 12682910
PARTIALLY ADAPTIVE AUDIO BEAMFORMING SYSTEMS AND METHODS
2y 5m to grant Granted Jul 14, 2026
Patent 12676165
METHOD FOR MIXING MICROPHONE INPUTS, APPARATUS, AND COMPUTER PROGRAM PRODUCT
2y 4m to grant Granted Jul 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
83%
Grant Probability
87%
With Interview (+3.2%)
2y 3m (~6m remaining)
Median Time to Grant
Low
PTA Risk
Based on 675 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month