Prosecution Insights
Last updated: October 02, 2026
Application No. 18/068,187

EFFICIENT FREQUENCY-BASED AUDIO RESAMPLING FOR USING NEURAL NETWORKS

Final Rejection §103§112
Filed
Dec 19, 2022
Examiner
SUBRAMANI, NANDINI
Art Unit
2656
Tech Center
2600 — Communications
Assignee
NVIDIA Corporation
OA Round
4 (Final)
65%
Grant Probability
Moderate
5-6
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 65% of resolved cases
65%
Career Allowance Rate
64 granted / 99 resolved
+2.6% vs TC avg
Strong +48% interview lift
Without
With
+47.5%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
15 currently pending
Career history
114
Total Applications
across all art units

Statute-Specific Performance

§101
12.9%
-27.1% vs TC avg
§103
65.3%
+25.3% vs TC avg
§102
11.1%
-28.9% vs TC avg
§112
8.5%
-31.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 99 resolved cases

Office Action

§103 §112
DETAILED ACTION Introduction Applicant's submission filed on 06/08/2026 has been entered. Claims 1-6, 9-13 and 16-20 are pending in the application and have been examined. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Examiner Recommendation It is the suggestion of the Examiner, to overcome the Rejections under 35 USC 112(a), the claims be amended to clarify the second frequency being higher than the lower frequency representation versus the first frequency representation as shown to be supported in the specifications as indicated further in this Office Action. Further, it is recommended to amend as the trained neural network used to generate the second audio at second frequency. Response to Amendment The response filed on 06/08/2026 has been correspondingly accepted and considered in this Office Action. Claims 1-6, 9-13 and 16-20 have been examined. Claims 7-8, 14-15 have been cancelled. Applicant' s amendments to claim 1, 9 and 17, training of the neural network as included in cancelled claim 8 overcome the 35 U.S.C 101 rejections previously set forth in the Non-Final Office Action mailed 01/12/2026. The dependent claims 2-6, 10-13, 16 and 18-20 overcome the 35 U.S.C 101 rejections previously set forth in the Non-Final Office Action mailed 01/12/2026 based on their dependency to the amended claims 1, 9 and 17 respectively. Therefore, the above referenced rejections under 35 U.S.C. 101 are withdrawn. Response to Arguments Applicant's arguments filed 06/08/2026 have been fully considered as follows: Applicant’s arguments with respect to claim 1 (also representative of claims 9 and 17) state that “e Office Action rejects claims 1-20 as allegedly failing to comply with the written description requirement, specifically regarding the language "the first audio data at a second frequency that is higher than the low frequency representation of the first frequency." Applicant respectfully disagrees and points the Office's attention to at least paragraphs [0027]-[0028] and [0033] of the Specification. For example, paragraph [0028] discloses: … This may include, for example, entering zero values for any higher frequency bands where data is not otherwise present, such as any frequency bands above 16 kHz for a 48 kHz target spectrogram. …. Furthermore, paragraph [0033] discloses: The super-resolution neural network 214 in this example, however, is designed to accept higher frequency data, such as a spectrogram at 48 kHz. Accordingly, this process uses a re-sampler 210 to convert the clipped frequency spectrogram 208 into a high frequency spectrogram at the frequency required for the neural network 214, or at least for which the network was designed. In this example, an FFT resampler takes in the clipped frequency spectrogram and generates a corresponding resampled spectrogram 212 that includes higher frequency bands. In this example, the resampler 210 inserts padding values, such as zero or other values, in the additional higher frequency bands (e.g., those above 16 kHz). …. ” The examiner respectfully disagrees, the claims 1, 9 and 17 cite “resample the low frequency representation to generate a second frequency-based representation of the first audio data at a second frequency that is higher than the first frequency” or the equivalent; however the cited paragraphs in the specifications ( bolded above) specifically refer to the resampled spectrogram as a higher frequency than clipped frequency spectrogram which is the low frequency of the filtered first frequency-based representation of the first audio data at a first frequency. The rejections of claims 1, 9 and 17 under 35 U.S.C. 112(a) are sustained and updated accordingly in this Office Action under Claim Rejections under 35 USC 112. Applicant’s arguments with respect to claim 1 (also representative of claims 9 and 17) state that “As discussed during the Examiner Interview, the Office alleges that Eskimez discloses "performing filtering of the first frequency-based representation to generate a low frequency representation of the first frequency-based representation," citing to Eskimez section IV, D. See Non-Final Office Action, p. 12. … For at least these reasons, Eskimez does not disclose the elements of claim 1 as discussed herein.” Applicant’s arguments above with respect to claim 1 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. In response to the art rejection(s) of the remainder of dependent claims are rejected under 35 U.S.C 103, in case said claims are correspondingly discussed and/or argued for at least the same rationale presented in Remarks filed 06/08/2026, Examiner respectfully notes as follows. For completeness, should the mentioned claims be likewise traversed for similar reasons to independent claims 1, 9 and 17 correspondingly, Examiner respectfully directs Applicant to the same previous supra reasons provided in the response directed towards claims 1, 9 and 17 correspondingly discussed above. For at least the same supra provided reasons, Examiner likewise respectfully disagrees, and Applicant's arguments have been fully considered but they are not persuasive. Claim Rejections - 35 USC § 112 The following is a quotation of the first paragraph of 35 U.S.C. 112(a): (a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention. The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112: The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention. Claims 1, 9 and 17 and the dependent claims 2-6, 10-16 and 18-20 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention. The support for the amendments in independent claims 1, 9 and 17 line 6 regarding “the first audio data at a second frequency that is higher than different from the low frequency representation the first frequency” is not supported by the specifications. Upon review of the Specification, the Examiner was uncertain where there is support in the specification for the newly amended limitations and request that in a next response, Applicant indicate the respective supporting portions of the specification. The amendments to claim 1 is inferred as being based on Specifications [0028, 0033-0034] regarding the resampled spectrogram as a higher frequency than clipped frequency spectrogram, further the super resolution neural network is trained by comparing the high frequency spectrogram with the resampled spectrogram, where the resampled spectrogram is resampled at the designed frequency of the super-resolution neural network 214. The dependent claims 2-8, 10-16 and 18-20 are also rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement due to the lack of support to the amendments in the independent claims 1, 9 and 17 respectively. Claim Rejections - 35 USC § 103 The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action. Claims 1-5, 9-12 and 16-20 are rejected under 35 U.S.C. 103 as being unpatentable over Li, X., et.al. (2019). Speech audio super-resolution for speech recognition. In Proc. Interspeech 2019 (pp. 3416-3420) further in view of Kumar et. al. US PgPub. 2022/0101872. Regarding claim 1, Li teaches a method comprising: converting a first audio data, in a time domain, to a first frequency-based representation of the first audio data at a first frequency (Li, sect 3, 3.4 discusses STFT computation of the signal to process with Narrow Band portion); performing filtering of the first frequency-based representation to generate a low frequency representation of the first frequency-based representation (Li, sect 5 discusses the low pass filter when generating LB from WB audio for the testing); resampling the low frequency representation to generate a second frequency-based representation of the first audio data at a second frequency that is higher than the first frequency and includes padded values for one or more higher frequency audio bands (Li, sect 3, discusses the generation of WB(higher frequency than NB) from NB signal with zero spectral content(zero padding) in HB portion ; second frequency that is higher than the first frequency is interpreted as second frequency is higher than the low frequency representation of the first frequency-based representation as indicated in the specifications [0028, 0032,0038]); adjusting one or more network parameters of one or more neural networks based on at least a loss value calculated by comparing the differences between a first enhanced waveform generated for the first frequency-based representation and a second enhanced waveform generated for the second frequency-based representation (see Li, sect 3.2, The discriminator serves the purpose of distinguishing the generated audio(second enhanced waveform) from true WB audio(first enhanced waveform). The adversarial loss expresses the similarity between generated samples and real samples estimated by the discriminator. This is usually combined with additional loss functions designed for task-specific purposes as shown in equation (3)); and generating, using one or more neural networks and based at least on the second frequency- based representation of the first audio data, second audio data at the second frequency (Li, fig. 1 Bottom, our end-to-end approach using cGAN ). Li teaches generating, using one or more neural networks and based at least on the second frequency- based representation of the first audio data, second audio data at the second frequency, to further compact prosecution, Kumar is used to further teach generating, using one or more neural networks and based at least on the second frequency-based representation of the first audio data, second audio data at the second frequency (see Kumar [0102] the model can be trained to transform magnitude spectrograms to a given sampling rate (here, Sampling Rate B). By applying an inverse transform to the second magnitude spectrogram and phase associated with the first magnitude spectrogram, the media production platform can obtain a second discrete audio signal (“Discrete Audio Signal B”) that has the higher sampling rate ). Li and Kumar are considered to be analogous to the claimed invention because they relate to processing audio data using neural network based super resolution models. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Li on training the audio super resolution models with the upsampling discrete audio signals to higher sampling rates teachings and trainings of super resolution model teachings of Kumar to improve audio quality( see Kumar, [0025-0026]). PNG media_image1.png 290 252 media_image1.png Greyscale Regarding claim 2, Li in view of Kumar teaches the method of claim 1. Kumar further teaches wherein the first frequency-based representation is a frequency-based spectrogram generated using a short-time Fourier transform (STFT) operation (see Kumar, [0051] Assume, for example, that the media production platform 210 acquires input indicative of a selection of an audio file to be upsampled. Generally, the audio file will be in the form of a discrete audio signal with a frequency less than 40,000 Hz. For example, the discrete audio signal may have a sampling rate of 8,000 Hz, 16,000 Hz, 22,050 Hz, or 24,000 Hz depending on the computing device used for recording. In such a scenario, the upsampling module 216 can apply a Fourier transform (e.g., an STFT) to the audio file to produce a first magnitude spectrum). The same motivation to combine as claim 1 applies here. Regarding claim 3, Li in view of Kumar teaches the method of claim 1. Li teaches wherein the resampling is performed using a fast Fourier transform (FFT) resampler to add one or more padded entries to one or more frequency bands in the second frequency-based representation that do not contain data values in the first frequency- based representation( Li, sect 3, discuss NB signal as HB portion with noise and processing using STFT to both HB and noise to remove the HB portion( the NB signal x as the WB signal with zero spectral con tent in the HB portion) ). Regarding claim 4, Li in view of Kumar teaches the method of claim 1. Li further teaches wherein the one or more neural networks infer one or more non-zero audio data values for one or more frequency bands from the second frequency- based representation (see Li, 3.1, design an end-to-end generator that takes the LB 1D audio signal x as input and generates WB 1D audio signal ˜x. (Given NB audio x is equivalent to the ˜x that has same number of samples as WB signal ˜x, we propose to recover the WB signal ˜x by using a cGAN that removes noise n from ˜x.)). Kumar further teaches wherein the one or more neural networks infer one or more non-zero audio data values for one or more frequency bands from the second frequency- based representation (see Kumar, [0102] The media production platform can then apply a model to the first magnitude spectrogram to produce a second magnitude spectrogram (“Magnitude Spectrogram B”). As discussed above, the model can be trained to transform magnitude spectrograms to a given sampling rate (here, Sampling Rate B) and non zero values as shown in Fig. 8 ) The same motivation to combine as claim 1 applies here. Regarding claim 5, Li in view of Kumar teaches the method of claim 1. Kumar teaches wherein the generating of the first frequency-based representation of the first audio data and the resampling are performed using at least one graphics processing unit (GPU) (see Kumar, [0045] The processor 202 can have generic characteristics similar to general-purpose processors, or the processor 202 may be an application-specific integrated circuit (ASIC) that provides control functions to the computing device 200(GPU). As shown in FIG. 2, the processor 202 can be coupled to all components of the computing device 200, either directly or indirectly, for communication purposes). The same motivation to combine as claim 1 applies here. Regarding claim 9, is directed to a processor claim corresponding to the method claim presented in claim 1 and is rejected under the same grounds stated above regarding claim 1. Kumar further teaches cause output of the second audio data using one or more output devices(see Kumar, [0045] discusses display mechanisms, Fig. 9, 918). Regarding claim 10, is directed to a processor claim corresponding to the method claim presented in claim 2 and is rejected under the same grounds stated above regarding claim 2. Regarding claim 11, is directed to a processor claim corresponding to the method claim presented in claim 3 and is rejected under the same grounds stated above regarding claim 3. Regarding claim 12, is directed to a processor claim corresponding to the method claim presented in claim 4 and is rejected under the same grounds stated above regarding claim 4. Regarding claim 16, Li in view of Kumar teaches the processor of claim 9. Kumar further teaches wherein the processor is comprised in at least one of: a system for rendering graphical output(see Kumar, Fig. 9, 918); a system for performing deep learning operations(see Kumar, [0050] At a high level, the model is a machine learning framework that is comprised of one or more algorithms adapted to upsample an audio signal that is provided as input, thereby producing another audio signal having a different sampling rate. ); a system implemented at least partially using cloud computing resources(see Kumar, [0108], Fig. 9, 912/914). Regarding claim 17, Li teaches using one or more neural networks, a resampled frequency-based representation of audio data, being in a time domain, the resampled frequency-based representation generated based at least on resampling of a low frequency representation of an initial frequency-based representation of the audio data to pad one or more values of one or more higher frequency audio bands(see Li, Fig. 1 Bottom, our end-to-end approach using cGAN; Li sect 3 We consider the NB signal x as the WB signal with zero spectral con tent in the HB portion ¯x, which can be further considered as the original WB signal ˜x mixed with noise n that cancels all the HB energy in phase), wherein the resampled audio data is at a second frequency that is higher than a first frequency of the initial frequency-based representation and includes padded values for the one or more higher frequency audio bands(see Li, sect 3, discusses the generation of WB(higher frequency than NB) from NB signal with zero spectral content(zero padding) in HB portion ; second frequency that is higher than the first frequency is interpreted as second frequency is higher than the low frequency representation of the first frequency-based representation as indicated in the specifications [0028, 0032,0038]), and wherein the resampled audio data is generated by neural networks including one or more network parameters adjusted based on at least a loss value calculated by comparing differences between a first enhanced waveform generated for the first frequency-based representation and a second enhanced waveform generated for the second frequency-based representation(see Li, sect 3.2, The discriminator serves the purpose of distinguishing the generated audio(second enhanced waveform) from true WB audio(first enhanced waveform). The adversarial loss expresses the similarity between generated samples and real samples estimated by the discriminator. This is usually combined with additional loss functions designed for task-specific purposes as shown in equation (3)). Kumar further teaches a system, comprising: one or more processors to generate resampled audio data based at least on processing, using one or more neural networks, a resampled frequency-based representation of audio data, being in a time domain (see Kumar, Fig. 8, Fig. 9).The motivation to combine is similar to claim 1 applies here. Regarding claim 18, is directed to a system claim corresponding to the method claim presented in claim 2 and is rejected under the same grounds stated above regarding claim 2. Regarding claim 19, is directed to a system claim corresponding to the method claim presented in claim 3 and is rejected under the same grounds stated above regarding claim 3. Regarding claim 20, is directed to a system claim corresponding to the processor claim presented in claim 16 and is rejected under the same grounds stated above regarding claim 16. Claims 6 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Li, X., et.al. (2019). Speech audio super-resolution for speech recognition. In Proc. Interspeech 2019 (pp. 3416-3420) in view of Kumar et. al. US PgPub. 2022/0101872 further in view of Zeyu et. al, US PgPub. 2023/0162725. Regarding claim 6, Li in view of Kumar teaches the method of claim 5. However, Li in view of Kumar fail to teach wherein the first audio data is received in a first audio stream, and wherein a batch of audio streams including the first audio stream is to be processed in parallel using one or more GPUs. However, Zeyu further teaches wherein the first audio data is received in a first audio stream, and wherein a batch of audio streams including the first audio stream is to be processed in parallel using one or more GPUs (see Zeyu, [0097] the processor(s) 1202 may include one or more central processing units (CPUs), graphics processing units (GPUs)). See Zeyu, [0074] As illustrated in FIG. 10, the method 1000 includes an act 1006 providing the upsampled audio data to an audio super resolution model, the audio super resolution model trained to perform bandwidth expansion from narrow-band to wide-band. In some embodiments, the audio may be processed in parallel by dividing the audio into batches. Each batch may then be processed by dividing the batch into sub-batches, with each sub-batch being processed in parallel; Zeyu [0103]). Li, Kumar and Zeyu are considered to be analogous to the claimed invention because they relate to processing audio data using neural network based super resolution models. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Li in view of Kumar on using the audio super resolution models with batch processing of audio data teachings and trainings of audio super resolution model teachings of Zeyu to improve processing speed ( see Zeyu, [0043]). Regarding claim 13, is directed to a processor claim corresponding to the method claim presented in claim 6 and is rejected under the same grounds stated above regarding claim 6. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Schmit et. al US PgPub. 2020/0243102 teaches a method for generating a bandwidth-enhanced audio signal for an input audio signal using spectral vectors used for the purpose of raw signal generation and used for the purpose of raw signal processing using the parametric representation output by the neural network processor (see Schmit, Fig. 2e). Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to NANDINI SUBRAMANI whose telephone number is (571)272-3916. The examiner can normally be reached Monday - Friday 12:00pm - 5:00 pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bhavesh M Mehta can be reached at (571)272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /NANDINI SUBRAMANI/ Examiner, Art Unit 2656 /BHAVESH M MEHTA/ Supervisory Patent Examiner, Art Unit 2656
Read full office action

Prosecution Timeline

Show 5 earlier events
Jul 18, 2025
Final Rejection mailed — §103, §112
Oct 20, 2025
Request for Continued Examination
Oct 27, 2025
Response after Non-Final Action
Jan 12, 2026
Non-Final Rejection mailed — §103, §112
May 14, 2026
Applicant Interview (Telephonic)
May 19, 2026
Examiner Interview Summary
Jun 08, 2026
Response Filed
Sep 11, 2026
Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749498
QUALITY ESTIMATION MODEL FOR PACKET LOSS CONCEALMENT
3y 9m to grant Granted Sep 29, 2026
Patent 12731575
Attention-Based Joint Acoustic and Text On-Device End-to-End Model
3y 7m to grant Granted Sep 08, 2026
Patent 12700416
Audio Transcoding Method and Apparatus, Audio Transcoder, Device, and Storage Medium
3y 9m to grant Granted Aug 04, 2026
Patent 12688863
EMOTIONALLY-AWARE VOICE RESPONSE GENERATION METHOD AND APPARATUS
4y 7m to grant Granted Jul 21, 2026
Patent 12670912
ATTENTIVE SCORING FUNCTION FOR SPEAKER IDENTIFICATION
2y 9m to grant Granted Jun 30, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
65%
Grant Probability
99%
With Interview (+47.5%)
3y 0m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 99 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month