Prosecution Insights
Last updated: October 02, 2026
Application No. 17/921,858

METHOD, APPARATUS AND SYSTEM FOR ENHANCING MULTI-CHANNEL AUDIO IN A DYNAMIC RANGE REDUCED DOMAIN

Final Rejection §103§112
Filed
Oct 27, 2022
Priority
Apr 30, 2020 — provisional 63/018,282 +2 more
Examiner
ZHANG, LESHUI
Art Unit
2695
Tech Center
2600 — Communications
Assignee
Dolby International AB
OA Round
4 (Final)
78%
Grant Probability
Favorable
5-6
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 78% — above average
78%
Career Allowance Rate
743 granted / 954 resolved
+15.9% vs TC avg
Strong +35% interview lift
Without
With
+35.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 9m
Avg Prosecution
25 currently pending
Career history
986
Total Applications
across all art units

Statute-Specific Performance

§101
5.8%
-34.2% vs TC avg
§103
44.8%
+4.8% vs TC avg
§102
14.4%
-25.6% vs TC avg
§112
29.1%
-10.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 954 resolved cases

Office Action

§103 §112
DETAILED ACTION The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This office action is in response to claim amendment filed on June 26, 2026 and wherein claims 1-52, 59 previously canceled and remained cancelation status, claims 66-67 previously withdrawn and remain withdrawn status, claim 73 amended and claim 74 newly added. In virtue of this communication, claims 53-58, 60-74 are currently pending in this Office Action. With respect to the rejection of claim 73 under 35 USC §112(a), as set forth in the previous Office Action, the claim amendment, and argument, see paragraph 1 of page 11 in Remarks filed on June 26, 2026, have been fully considered and the argument is persuasive. Therefore, the rejection of claim 73 under 35 USC § 112(a), as set forth in the previous Office Action, has been withdrawn. The Office appreciates the explanation of the amendment and analyses of the prior arts, and however, although the claims are interpreted in light of the specification, limitations from the specification are not read into the claims. See In re Van Geuns, 988 F.2d 1181, 26 USPQ2d 1057 (Fed. Cir. 1993) and MPEP 2145. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(a): (a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention. Claim 74 is rejected under 35 U.S.C. 112(a), as failing to comply with the written description requirement. The claims contain subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor(s) or joint inventor, or for pre-AIA the inventor(s), at the time the application was filed, had possession of the claimed invention. Claim 74 recites “wherein the decoder stage outputs the two or more enhanced channels simultaneously” which is not supported by the original disclosure, including original claims and drawings. For example, the specification reads “processing/enhancing, which is performed simultaneously on the two or more channels of the multi-channel audio signal. In this case, jointly refers to the simultaneous enhancement of the two or more channels of the dynamic range reduced raw multi-channel audio signal by the multi-channel Generator (USPGPub 20230178084 A1, para 82)”, etc., i.e., “simultaneously” or “at the same time” is referred to processing of the “dynamic range reduced raw multi-channel audio signal” by “Generator”, which would not be considered as support of claimed limitation above, because “simultaneously” or “at the same time” has nothing to do with “outputs” by “decoder”. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 53, 55, 59-60, 63-64, 68, 71, 73 are rejected under 35 U.S.C. 103 as being unpatentable over Biswas (US 20180358028 A1) and in view of references Blyth et al (US 20150008962 A1, hereinafter Blyth) and Pascual et al. (“SEGAN: Speech Enhancement Generative Adversarial Network”, [1703.09452] SEGAN: Speech Enhancement Generative Adversarial Network, Computer Science, Machine Learning, v3, 2017, hereinafter Pascual). Claim 53: Biswas teaches a method (title and abstract, ln 1-12 and a decoder in fig. 5 and method steps in fig. 3B) of generating, in a dynamic range reduced domain (the audio signal and processing from a compression 104 at an audio encoder side to expansion 114 at an audio decoder side is in dynamic range reduced domain due to compression 104, para 34), an enhanced multi-channel audio signal (audio out 512 in fig. 5 or 112 in fig. 1, having ordinal dynamic range with coding noise reduced, para 34) from an audio bitstream (received bitstream 501 in fig. 5) including a multi-channel audio signal (multiple channels handled by duplicating the operation separately on each channel, or grouped channels having similarities, para 59, e.g., including stereo or multi-channel signals, para 35, 60), the method comprising: receiving the audio bitstream (via a part of core decoder 502 for receiving the bitstream in fig. 5); core decoding the audio bitstream (the bitstream decoded to generate decoded portion of bitstream into QMF analysis 504 via other part of the core decoder 502 in fig. 5) and obtaining a dynamic range reduced raw multi-channel audio signal based on the received audio bitstream (output from the core decoder 502 in fig. 5), wherein the dynamic range reduced raw multi-channel audio signal comprises two or more channels (including stereo two channels, para 68 and two or more channels, para 59); inputting two or more channels (stereo as multi-channel signals, para 35, at observed stereo-panned transient signals, para 60) of the dynamic range reduced raw multi-channel audio signal (output signal from core decoder 502 in fig. 5) together into a multi-channel generator (including QMF and/or parametric stereo PS tool, A-CPL tool set, para 67-68, for reconstructing full bandwidth prior to expanding, e.g., parametric stereo tool and A-CPL toolset, and the reconstructing includes using side information and reconstructing stereo channel from the mono, para 67-68, and the processed channel can be group channel from multiple channels having similarities, para 60, i.e., receiving the stereo signals at the same time as claimed together for avoiding audible stereo image artifacts, para 60); enhancing, with the multi-channel generator, the two or more channels (by using QMF and/or parametric stereo PS tool, A-CPL tool set, para 67-68 and grouped channel to be processed, and the discussion above) in the dynamic range reduced domain (the processing above is performed prior to the expanding by the expander 506 in fig. 5) to output an enhanced dynamic range reduced multi-channel audio signal (reconstructed full bandwidth compressed signal, para 67) for subsequent expansion of the dynamic range (expander 506 to be used for expanding the compressed signals in fig. 5), wherein the enhanced dynamic range reduced multi-channel audio signal comprises two or more enhanced channels that are enhanced relative to the two or more channels of the dynamic range reduced raw multi-channel audio signal (stereo and two or more channels, the discussion above) and wherein the multi-channel generator is configured (QMF analysis filterbank, as part of the multi-channel generator discussed above, is pre-configured to implement band separation through bandpass filters, para 61, and opposite operation at decoder of fig. 5, para 66; A-SPX tool as a part of the multi-channel generator is above, is pre-configured so that the high frequency band is shaped using side information received by A-SPX in fig. 5, para 64). However, Biswas does not explicitly teach that simultaneously performs disclosed enhancement of the two or more channels and does not explicitly teach wherein the multi-channel generator is trained in the dynamic range reduced domain in a Generative Adversarial Network GAN setting. Blyth teaches an analogous field of endeavor by disclosing a method of generating, in a dynamic range reduced domain, an enhanced multi-channel audio signal (title and abstract, ln 1-18 and in fig. 3 and the method for multi-channel in fig. 10) and wherein a multi-channel generator is disclosed (fig. 10, including INTERP 1001, 1002, ENVR 1011, CTRLR 1012, element 1010 for right input audio channel channel DinR and similar elements for left input audio channel DinL, para 125, 130-135) and wherein inputting two or more channels (stereo DinR and DinL in fig. 10, and also applied to more than two channels, e.g., in 5.1 output configurations, para 129) of the dynamic range reduced raw multi-channel audio signal (dynamic range enhancement DRE in fig. 10, para 114 and including the received signal having a low signal levels compared to noise condition, para 130, i.e., the claimed dynamic range reduced) together into the multi-channel generator (for evaluating envelope for both stereo signals, and using the DAR/DAL, DBR/DBL at the same time, para 129); simultaneously enhancing, with the multi-channel generator, the two or more channels (by evaluating digital gain GDGR/GDGL and applied GDIGR/GDIGL to the interpolated audio signals 64fs at left and right channel paths through a multiplier 1003 for the left and the right channels, respectively and specifically, all four sample inputs applied for evaluating the digital gain values, para 129) to output an enhanced dynamic range reduced multi-channel audio signal (output from the multiplier 1003 for the left and the right channels in fig. 10) for subsequent expansion of the dynamic range (through analog amplifier 1005 for the left and the right channels, for compensation of the digital gains GDIGR/GDIGL, para expander 506 to be used for expanding the compressed signals in fig. 5), wherein the enhanced dynamic range reduced multi-channel audio signal comprises two or more enhanced channels that are enhanced relative to the two or more channels of the dynamic range reduced raw multi-channel audio signal (stereo two channels, outputted from the multiplier 1003 and compared to the input left and the input right channels DinR, DinL in fig. 10) for benefits of enhancing the dynamic range processing for multi-channel audio signals (by improving noise performance at low signal levels, para 10, by accurately measuring the dynamic range of the audio signals, para 120, in a cost-saving manner, para 10). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have applied wherein simultaneously enhancing, with the multi-channel generator, the two or more channels in the dynamic range reduced domain to output an enhanced dynamic range reduced multi-channel audio signal for subsequent expansion of the dynamic range, as taught by Blyth, to enhancing, with the multi-channel generator, the two or more channels in the dynamic range reduced domain in the method, as taught by Biswas, for the benefits discussed above. However, the combination of Biswas and Blyth does not explicitly teach wherein the multi-channel generator is trained in the dynamic range reduced domain in a Generative Adversarial Network GAN setting. Pascual teaches an analogous field of endeavor by disclosing a method (title and abstract, ln 1-19 and training process in fig. 1 and an architecture for speech enhancement network in fig. 2) and wherein a multi-channel generator is disclosed (including encoder and decoder in fig. 2, represented in G, and for two input channels of 16384 samples, p.3, col 2, para 4, e.g., two speakers with 20 alternative noise conditions, or 28 speakers and 40 different noise conditions into the same model, abstract, and the same model trained by 28 from 30 speakers from the Voice Bank corpus and 2 speeches used for test, Session 4.1 Data Set, para 1-2) is trained in a dynamic range reduced domain (spectral domain, abstract, and compressed through a number of strided convolutional layers, p.2, col 2, para 2-3) in a Generative Adversarial Network GAN setting (setting to carry out GAN training processing in fig. 1) for benefits of enhancing intelligibility and quality of generated speech (via introducing multiple and complex noisy conditions, abstract, session 2 Generative Adversarial Networks, p.2, col 1, para 3 and session 3. Speech Enhancement GAN, p.2) through improved training performance (simplifying training process and saving training time, session 3 Speech Enhancement GAN, p.2, col 2). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have applied the multi-channel generator that is trained in the dynamic range reduced domain in a Generative Adversarial Network GAN setting, as taught by Pascual, to the configuration of the multi-channel Generator in the dynamic range reduced domain in the method, as taught by the combination of Biwas and Blyth, for the benefits discussed above. Claim 68 has been analyzed and rejected according to claim 53 above and the combination of Biswas, Blyth, and Pascual further teaches an apparatus that includes a receiver (part of core decoder 502 to receive bitstream 501 in fig. 5), a core decoder (other part of core decoder 502 in fig. 5), and a multi-channel generator (including QMF and A-SPX processing unit 508, etc., in fig. 5), for implementing the method of claim 53. Claim 55: the combination of Biswas, Blyth, and Pascual further teaches, according to claim 53 above, wherein the method further includes a step of expanding the enhanced dynamic range reduced multi-channel audio signal to an expanded dynamic range domain by performing an expansion operation on the two or more channels (Biswas, the spectral magnitudes for emphasizing a weak spectral content, para 41, 47-53 and Blyth, through analog portion of the DRE in fig. 10, para 127-129), wherein the expansion operation is a companding operation based on a p-norm of spectral magnitude for calculating respective gain values (Biswas, p-norm of the spectral magnitudes used for emphasizing weak spectral content, para 41 and details in para 47-53 and Blyth, through analog portion of the DRE in fig. 10, via ENVG 1008, CTLG 1009, PSU 1007, etc., for providing maximum to support dynamic range enhancement, para 127-129 and Pascual, through the decoder to compand in fig. 2). Claim 71 has been analyzed and rejected according to claims 68, 55 above. Claim 59: the combination of Biswas and Blyth teaches all the elements of claim 59, according to claim 53 above, including the multi-channel generator in the dynamic range reduced domain (Biswas, and Blyth, the discussion in claim 53 above), except explicitly teaching wherein the multi-channel generator is a generator trained in the dynamic range reduced domain in a Generative Adversarial Network GAN setting. Pascual teaches an analogous field of endeavor by disclosing a method (title and abstract, ln 1-19 and training process in fig. 1 and an architecture for speech enhancement network in fig. 2) and wherein a generator is disclosed (including encoder and decoder in fig. 2) to be trained in a spectral domain (spectral domain, abstract, session Introduction, p.1 or waveform domain, abstract) in a generative adversarial network GAN setting (setting in fig. 1, mapping and discriminating processing defined by the formula 1, and improved formula 2, p.2, D-discriminator, G-generator, Ẍ - real speech sample, session 2 Generative Adversarial Networks, p.1-2) for benefits of enhancing intelligibility and quality of generated speech (via introducing multiple and complex noisy conditions, abstract, session 2 Generative Adversarial Networks, p.2, col 1, para 3 and session 3. Speech Enhancement GAN, p.2) with improved training performance (simplifying training process and saving training time, session 3 Speech Enhancement GAN, p.2, col 2). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have applied the generator trained in the generative adversarial network GAN setting, as taught by Pascual, to the multi-channel Generator in the dynamic range reduced domain in the method, as taught by the combination of Biwas and Blyth, for the benefits discussed above. Claim 60: the combination of Biwas, Blyth, and Pascual further teaches, according to claims 53, 59 above, wherein the multi-channel generator (Biwas, discussed in claim 53 above, and Pascual, generator G and the discussion in claim 53, 59 above) includes an encoder stage and a decoder stage (Biwas, discussed in claim 53 above, and Pascual, encoder-decoder architecture for speech enhancement through G network in fig. 2) arranged in a mirror symmetric manner (Pascual, G has symmetric architecture between encoder or convolutions and decoder or deconvolutions, which are connected through a latent representation layer z in fig. 2, session 3. Speech Enhancement GAN, p.2), wherein the encoder stage and the decoder stage each include L layers with N filters in each layer (Pascual, strided convolution layers and each of the layers has N steps of fitlers for outputting from the layer, session 3, p.2), wherein L is a natural number ≥ 1 and wherein N is a natural number ≥ 1 (Pascual, output from each layer, i.e., from the layer filter, are inputted to the next layer as input of the next layer in fig. 2, e.g., N=2, session 4.2 SEGAN Setup, p.3), and wherein a size of the N filters in each layer of the encoder stage and the decoder stage is the same (Pascual, the encoding process is reversed in the decoding stage and each encoding layer connected to its homologous decoding layer by passing compression performed in the middle of the model, session 3 Speech Enhancement GAN, p.2, col 2) and each of the N filters in the encoder stage and the decoder stage operates with a stride of > 1 (Pascual, L/2, L/4, L/8, …, in fig. 2, session 3, p.2). Claim 63: the combination of Biwas, Blyth, and Pascual further teaches, according to claims 60 above, wherein one or more skip connections exist between respective homologous layers of the multi-channel generator (Pascual, skip connections used for connecting each encoding layer to its homologous decoding layer, p.2, session 3. Speech Enhancement GAN, col 2, para 3). Claim 64: the combination of Biwas, Blyth, and Pascual further teaches, according to claims 60 above, wherein the multi-channel generator includes, between the encoder stage and the decoder stage, a stage for modifying multi-channel audio in the dynamic range reduced domain based at least on a dynamic range reduced coded multi-channel audio feature space (Biwas, such as side information used in parametric stereo tool for reconstructing original channels, para 68-69, and Pascual, a latent vector z and a thought vector c layers concatenated each other and between the encoder and the decoder layers in fig. 2, session 3 Speech Enhancement GAN, p.2, col 2, para 2). Claims 54, 56-58, 69-70, 72 are rejected under 35 U.S.C. 103 as being unpatentable over Biwas (above) and in view of references Blyth (above), Pascual (above), and Riedmiller et al (US 20150104021 A1, hereinafter Riedmiller). Claim 69: the combination of Biswas, Blyth, and Pascual further teaches, according to claim 68 above, wherein the received audio bitstream includes metadata (Biswas, including side information used for replicating higher frequency components, para 64, companding control per channel transmitted from the encoder in fig. 4), wherein the metadata include one or more items of companding control data (Biswas, companding control per channel and placed into the bitstream 414 and including companding mode control in fig. 4), wherein the companding control data include information on a companding mode among one or more companding modes (Biswas, companding mode selected among on/off/average modes, para 35, switching between individual companding and jointly companding of channels, para 60, different tools used before expanding, e.g., parametric stereo PS tool or A-CPL toolset, etc., para 67-68) that had been used for encoding the multi-channel audio signal (Biswas, the companding control information passed to compressor 406 at the encoder in fig. 4, para 89-91). However, the combination of Biswas, Blyth, and Pascual does not explicitly teach a demultiplexer for demultiplexing the received audio bitstream. Riedmiller teaches an analogous field of endeavor by disclosing an apparatus (title and abstract, ln 1-17, a decoding system in fig. 2c) for generating, in a dynamic range reduced domain (the multi-channel audio signal is encoded and subject to being dynamic range controlled through the decoder in fig. 2c, DRC1, DRC2, DRC3 upon the coding mode, parametric coding mode for low bitrate and discrete coding mode for high bitrate, para 33), an enhanced multi-channel audio signal (reconstructed channel signals X in fig. 2c) from an audio bitstream (bitstream P received by demultiplex 60 or 70 in fig. 2c) including a multi-channel audio signal (n-channel audio signal X and m-channel core signal Y in fig. 2c, abstract), wherein A demultiplexer is disclosed for demultiplexing the received audio bitstream (demultiplexing encoded audio signal Ẋ, Ẏ and DRC2, DRC3, DRC1, and α in fig. 2c) for benefits of improving performance in audio encoding and audio decoding with compatibility for legacy system (by improving bandwidth efficiency, computational efficiency, error resilience, para 4). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have applied the demultiplexer for demultiplexing the received audio bitstream, as taught by Riedmiller, to the receiver for receiving the audio bitstream in the apparatus, as taught by the combination of Biwas, Blyth, and Pascual, for the benefits discussed above. Claim 70: the combination of Biwas, Blyth, Pascual and Riedmiller further teaches, according to claim 69 above, wherein the multi-channel Generator is configured to jointly enhance the two or more channels of the dynamic range reduced raw multi-channel audio signal in the dynamic range reduced domain (Biswas, discussion in claim 68 above) depending on the companding mode indicated by the companding control data (Biswas, jointly companding of the channels based on detected similarity between channels, para 60; different tools used before expanding, e.g., parametric stereo PS tool or A-CPL toolset, etc., para 67-68 and discussion above, and Riedmiller, companding modes defined by usage of DRC1, DRC2, and DRC3 and combination of thereof at decoder in fig. 2c and upon either parametric decoding mode as a low bitrate mode or discrete decoding mode as a high bitrate mode, para 33). Claim 72: the combination of Biwas, Blyth, Pascual and Riedmiller further teaches, according to claim 68 above, wherein the apparatus further includes a dynamic range reduction unit configured to perform a dynamic range reduction operation after core decoding the audio bitstream to obtain the dynamic range reduced raw multi- channel audio signal (Riedmiller, part of decoding unit 71, compared to a late DRC processing by element 74 in fig. 2c, and the decoding unit 71 further performs DRC by using DRC3 applied on the decoded core signal, para 64, and DRC is used for range compression, para 30, i.e., dynamic range reduction operation). Claim 54: the combination of Biswas, Blyth, and Pascual further teaches, according to claim 68 above, core decoding the audio bitstream (Biswas, via core decoder 502 in fig. 5), except wherein performing a dynamic range reduction operation to obtain the dynamic range reduced raw multi- channel audio signal after core decoding the audio bitstream. Riedmiller teaches an analogous field of endeavor by disclosing an apparatus (title and abstract, ln 1-17, a decoding system in fig. 2c) for generating, in a dynamic range reduced domain (the multi-channel audio signal is encoded and subject to being dynamic range controlled through the decoder in fig. 2c, DRC1, DRC2, DRC3 upon the coding mode, parametric coding mode for low bitrate and discrete coding mode for high bitrate, para 33), an enhanced multi-channel audio signal (reconstructed channel signals X in fig. 2c) from an audio bitstream (bitstream P received by demultiplex 60 or 70 in fig. 2c) including a multi-channel audio signal (n-channel audio signal X and m-channel core signal Y in fig. 2c, abstract), wherein performing a dynamic range reduction operation to obtain the dynamic range reduced raw multi- channel audio signal after core decoding the audio bitstream (a part of decoding unit 71, compared to a late DRC processing by element 74 in fig. 2c, and the decoding unit 71 further performs DRC by using DRC3 applied on the decoded core signal, para 64, and DRC is used for range compression, para 30, i.e., dynamic range reduction operation), for the benefits discussed in claim 72 above. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have applied performing the dynamic range reduction operation to obtain the dynamic range reduced raw multi- channel audio signal after core decoding the audio bitstream, as taught by Riedmiller, to the core decoding of the audio bitstream in the method, as taught by the combination of Biwas, Blyth, and Pascual, for the benefits discussed above. Claim 56 has been analyzed and rejected according to claims 53, 69 above. Claim 57 has been analyzed and rejected according to claims 56, 69 above. Claim 58 has been analyzed and rejected according to claims 57, 70 above. Claims 61-62, 65 are rejected under 35 U.S.C. 103 as being unpatentable over Biwas (above) and in view of reference Blyth (above), Pascual (above) and Nesta et al (US 20200349965 A1, hereinafter Nesta). Claim 61: the combination of Biwas, Blyth, and Pascual teaches all the elements of claim 61, according to claims 60 above, including the multi-channel Generator (Biwas, discussed in claim 53 above, and Pascual, the discussion in claims 60 above) comprising an input layer and an output layer (Pascual, first layer to receive noisy input and last layer to output enhanced output in fig. 2), except wherein the multi-channel Generator further includes a non-strided convolutional layer as an input layer prepending the encoder stage. Nesta teaches an analogous field of endeavor by a method for enhanced multi-channel audio signal (title and abstract, ln 1- 18 and method step implementation in a noise reduction system in figs. 1A/1B/1C, applied to including multichannel audio signals, para 4) and wherein a multi-channel generator is disclosed (a deep neural network DNN 100 in figs. 1A and autoencoder neural network 120 in fig. 1B, para 19-20) to include a non-strided convolutional layer as an input layer prepending the encoder stage (input layer 104 or 126 by receiving input samples 102, 122 to prepending to the hidden layers 110 or embedding vectors 130 in figs. 1A/1B, respectively, with no strided layer, para 19-20) and a non-strided transposed convolutional layer as an output layer subsequently following the decoder stage (following the noise embedding layer and speech embedding layer, a output layer 106 for noise embedding and 128 for speech embedding in figs. 1A/1B, respectively, para 19-20) for benefits of enhancing speech (by improving denoising and enhancing target by priorly providing noise signal and target signal in the deep embeddings, para 8). Therefore, it would have been obvious to one of ordinary skill in the art at the time the invention was made to have applied wherein the multi-channel Generator further includes the non-strided convolutional layer as the input layer prepending the encoder stage, as taught by Nesta, to the input layer of the multichannel Generator in the method, as taught by the combination of Biwas, Blyth, and Pascual, for the benefits discussed above. Claim 62 has been analyzed and rejected according to claims 60-61 above. Claim 65: the combination of Biwas, Blyth, Pascual, and Nesta further teaches, according to claims 64, 61 above, wherein a random noise vector z is used in the dynamic range reduced coded multi-channel audio feature space for modifying multi-channel audio in the dynamic range reduced domain (Nesta, the noise reduction neural network NR-NN 150 is trained with random speech and noise sequences to produce an enhanced signal 160, para 21), wherein the use of the random noise vector z is conditioned on a bit rate of the audio bitstream and/or on a number of channels of the multi-channel audio signal (Biwas, channel grouping information is transmitted through the bitstream, para 60 and used for reconstructing the number of original channels, para 34, and Pascual, discrete decoding for high bitrate and parametric decoding for low bitrate, para 33). Claim 74 is rejected under 35 U.S.C. 103 as being unpatentable over Biswas (above) and in view of reference Pascual (above). Claim 74: Biswas teaches a method (title and abstract, ln 1-12 and a decoder in fig. 5 and method steps in fig. 3B) of generating, in a dynamic range reduced domain (the audio signal and processing from a compression 104 at an audio encoder side to expansion 114 at an audio decoder side is in dynamic range reduced domain due to compression 104, para 34), an enhanced multi-channel audio signal (audio out 512 in fig. 5 or 112 in fig. 1, having ordinal dynamic range with coding noise reduced, para 34, and multi-channel audio signals, including stereo audio signals, para 60) from an audio bitstream (received bitstream 501 in fig. 5) including a multi-channel audio signal (multiple channels handled by duplicating the operation separately on each channel, or grouped channels having similarities, para 59, e.g., including stereo multi-channel signals, para 35, 60), the method comprising: receiving the audio bitstream (via a part of core decoder 502 for receiving the bitstream in fig. 5); core decoding the audio bitstream (the bitstream decoded to generate decoded portion of bitstream into QMF analysis 504 via other part of the core decoder 502 in fig. 5) and obtaining a dynamic range reduced raw multi-channel audio signal based on the received audio bitstream (output from the core decoder 502 in fig. 5), wherein the dynamic range reduced raw multi-channel audio signal comprises two or more channels (including stereo two channels, para 60, 68 and two or more channels, para 59); inputting the two or more channels (including stereo as multi-channel signals, para 35, at observed stereo-panned transient signals, para 60) of the dynamic range reduced raw multi-channel audio signal (output signal from core decoder 502 in fig. 5) together as a multi-channel input to a multi-channel generator (including QMF and/or parametric stereo PS tool, A-CPL tool set, para 67-68, for reconstructing full bandwidth prior to expanding, e.g., parametric stereo tool and A-CPL toolset, and the reconstructing includes using side information and reconstructing stereo channel from the mono, para 67-68, and the processed channel can be group channel from multiple channels having similarities, para 60, i.e., receiving the stereo signals at the same time as claimed together for avoiding audible stereo image artifacts, para 60); simultaneously enhancing, with the multi-channel generator, the two or more channels (by using QMF and/or parametric stereo PS tool, A-CPL tool set, para 67-68 and grouped channel to be processed, and including stereo audio signals and the discussed above and it is inherency for the stereo audio signals to be processed and inputted at the same time for supporting video frame synchronous audio coding, para 73, and for synchronization manner in compression at encoding and companding at decoding, para 66, and e.g., audible stereo image artifacts would be caused by independently companding of the individual channels, i.e., preferably processed at the same time or simultaneously, para 60) in the dynamic range reduced domain (the processing above is performed prior to the expanding by the expander 506 in fig. 5) to output an enhanced dynamic range reduced multi-channel audio signal (reconstructed full bandwidth compressed signal, para 67) for subsequent expansion of the dynamic range (expander 506 to be used for expanding the compressed signals in fig. 5), wherein the enhanced dynamic range reduced multi-channel audio signal comprises two or more enhanced channels that are enhanced relative to the two or more channels of the dynamic range reduced raw multi-channel audio signal (stereo and two or more channels, the discussion above) and wherein the multi-channel generator is configured (QMF analysis filterbank, as part of the multi-channel generator discussed above, is pre-configured to implement band separation through bandpass filters, para 61, and opposite operation at decoder of fig. 5, para 66; A-SPX tool as a part of the multi-channel generator is above, is pre-configured so that the high frequency band is shaped using side information received by A-SPX in fig. 5, para 64) and outputting the two or more enhanced channels simultaneously (outputting the stereo audio signals at the same time to avoid audible stereo image artifacts, para 60). However, Biswas does not explicitly teach an encoder stage of the multi-channel generator and wherein the multi-channel is inputted to the encoder stage of the multi-channel generator and does not explicitly teach wherein the multi-channel generator is trained in the dynamic range reduced domain in a Generative Adversarial Network GAN setting, wherein the multi-channel generator comprises an encoder stage and a decoder stage arranged in a mirror symmetric manner, with one or more skip connections between respective homologous layers of the encoder stage and the decoder stage, and wherein it is the decoder stage to perform disclosed simultaneous output of the two or more enhanced channels. Pascual teaches an analogous field of endeavor by disclosing a method (title and abstract, ln 1-19 and training process in fig. 1 and an architecture for speech enhancement network in fig. 2) and wherein a multi-channel generator is disclosed (including encoder and decoder in fig. 2, represented in G, and for two input utterances from two speakers or two channels of 16384 samples, p.3, col 2, para 4, e.g., two speakers with 20 alternative noise conditions, or 28 speakers and 40 different noise conditions into the same model, abstract, and the same model trained by 28 from 30 speakers from the Voice Bank corpus and 2 speeches used for test, Session 4.1 Data Set, para 1-2) is trained in a dynamic range reduced domain (spectral domain, abstract, and compressed through a number of strided convolutional layers, p.2, col 2, para 2-3) in a Generative Adversarial Network GAN setting (setting to carry out GAN training processing in fig. 1) and wherein the multi-channel generator comprises an encoder stage (an encoder in G, low portion of the Encoder-decoder architecture in fig. 2) to receive the multi-channel input (the speech from two speakers with 20 alternative noises in the test set, abstract) and a decoder stage (the upper mirrored portion of the Encoder-Decoder architecture in fig. 2) arranged in a mirror symmetric manner (the upper and the lower portions are mirrored each other in fig. 2, p.3, Session SEGAN Setup, col 2, para 3), with one or more skip connections between respective homologous layers of the encoder stage and the decoder stage (the arrows between encoder and decoder blocks denote skip connections, session 3. Speech Enhancement GAN, p.2, col 2, para 3), and wherein it is the decoder stage to perform the simultaneous output of the two or more enhanced channels (the clear utterance from at least two speakers in the test is outputted as an output signal in fig. 2, and two speakers in the test, and 28 speakers in the training the same model, abstract) for benefits of enhancing intelligibility and quality of generated speech (via introducing multiple and complex noisy conditions, in multiple channels of speakers, abstract, session 2 Generative Adversarial Networks, p.2, col 1, para 3 and session 3. Speech Enhancement GAN, p.2) through improved training performance (simplifying training process and saving training time, session 3 Speech Enhancement GAN, p.2, col 2). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have applied the encoder stage of the multi-channel generator and wherein the multi-channel generator is trained in the dynamic range reduced domain in the Generative Adversarial Network GAN setting, wherein the multi-channel generator is trained in the dynamic range reduced domain in the Generative Adversarial Network GAN setting, wherein the multi-channel generator comprises the encoder stage to receive the multi-channel input and the decoder stage arranged in the mirror symmetric manner, with the one or more skip connections between the respective homologous layers of the encoder stage and the decoder stage, and wherein it is the decoder stage to perform the simultaneous output of the two or more enhanced channels, as taught by Pascual, to the multi-channel Generator, the configuration of the multi-channel Generator in the dynamic range reduced domain, and simultaneously outputting the two or more enhanced channels, in the method, respectively as taught by Biswas, for the benefits discussed above. Allowable Subject Matter Claim 73 would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Response to Remarks Applicant's arguments filed on June 6, 2026 have been fully considered and but are not persuasive. Although a new claim 74 is added and claim 73 depending on claim 53 has been amended, a response is considered necessary for several of applicant’s arguments since reference Biswas and Blyth will continue to be used to meet several claimed limitations. With respect to prior art rejection of claim 53, applicant argued “Biswas does not disclose or suggest a multi-channel generator trained in a GAN setting that simultaneously enhances multiple channels together … Applicant’s specification explicitly defines ‘jointly’ as processing that is performed simultaneously on the two or more channels … wherein the two or more channels … are input into the multi-channel Generator at the same time”, etc., as asserted in paragraphs 3-4 of page 12 in Remarks filed on June 26, 2026 and “Biswas teaches the opposite approach: “multiple channels could be handled by repeating the approach individually on each channel” and “Pascual’s reference to “two input channels in Section 4.2 describes the discriminator D receiving a noise/clean signal pair for …” as asserted in paragraph 2 of page 13 in Remarks filed on June 26, 2026. In response to argument above, the Office respectfully disagrees because (1) the Office Action clearly wrote that the prior art rejection is under 35 U.S.C. 103(a), i.e., the primary prior art Biswas does not have to teach every limitations claim 53 recited, (2) in light of the application specification about the reference of “simultaneously” to the specification “jointly processing” (application USPGPub 20230178084 A1, para 82), etc., Biswis clearly teaches the “jointly companding the channels” of multi-channel audio signals including stereo audio signals or companding on the groups (para 60) which anticipated the “jointly processing” as the interpretation of claimed term “simultaneously”, but applicant is in silence, (3) Biswis clearly teaches that audio signals are multiple channel audio signals, including stereo audio signal inputting and processing (stereo-panned transient signals, para 60), and it is well-known in the art that stereo audio signals would have been inputted and processed simultaneously, for balanced sound effects, preventing from artifacts due to non-synchronized pair of stereo audio channels, and the consideration of accuracy of the dynamic range measurement, etc., and this is further proved by secondary prior art Blyth that teaches, stereo audio signal, dynamic range enhancement simultaneously and applied to stereo audio signals (fig. 10, shared with ENVG 1008, CTRLG 1009, PSU 1007 by taking the stereo type of audio signals as inputs, i.e., jointly processing), as discussed in office action above (office action above, p.7), etc., (4) claim 53 broadly recited “bitstream including a multi-channel audio signal” and “comprises two or more channels”, etc., with no recitation of what multiple channels are or comprising and thus, in its BRI, Biswas teaches “multi-channel” audio signals, including stereo audio signals, or repeatedly generating multiple signals from an audio source, and this is undoubtedly that Biswis’ “stereo audio signals” or “audio signals” are audio signals in at least two channels, which are not opposite to the broadly claimed “multi-channel audio signals” or “multiple channels”, and applicant provides no evidence that Biswas’ “stereo” or repeatedly generated “audio signals” would not be “multiple channels” or is the opposite to the “multi-channel audio signals” or “multiple channels” as recited in claim 53. Similarly, Pascual clearly teaches speech enhancement GAN and the speech is a type of audio signals with and without noise, i.e., multi-channel audio signals. In addition, claim 53 recited that “multi-channel generator is trained in …” without recitation of any requirement for the training, whether the training data has to be a single channel or multiple channels, and without recitation of how it is trained, but applicant is also in silence. Based on the analyses and evidence above, the prior art rejection of claim 53 under 35 U.S.C. 103(a) maintained. Applicant further argued “Blyth’s envelope detection and gain allocation for DRE is not a neural network-based enhancement system”, asserted in paragraph 3 of page 13 in Remarks filed on June 26, 2026, and “Pascual’s reference to two input channels of 16384 samples refers to the discriminator network D, not the generator G, and these two channels are noisy signal and clean signal pair for the discriminator’s classification task, not simultaneous multi-channel audio processing. Pascual’s SEGAN is designed for single-channel speech enhancement, processing one audio channel at a time to remove noise, not for simultaneously enhancing multiple audio channels together as a unified multi-channel input”, as asserted in paragraph 4 of page 13 and paragraph 1 of page 14 in Remarks filed on June 26, 2026 and further argued “”. In response to the argument above, the Office further disagrees because (1) claim 53 failed to recite the argued “neural network-based enhancement system”, but merely a method of generating, in a dynamic range reduced domain, an enhanced multi-channel audio signal from …a multi-channel audio signal”, etc., (2) with respect to the argued “multi-channel audio signals”, Pascual does not teach or limit Pascual’s GAN to be only for a single channel speech application, and instead, Pascual clearly teaches multiple speeches with multiple noises by using the same model, as discussed in the office action (p.7-8 of the office action above, Pascual, test sets having 2 speakers and 20 alternative noise conditions, i.e., multiple channels in GAN application or speech as input, session 4.1 Data Set, p.3), although the training is performed by using “speech” and “noisy speech”. However, applicant is also in silence about Pascual’s disclosure above, and (3) it is further emphasized that the prior art rejection of claim 53 is under 35 U.S.C. 103(a), and thus, the second and/or the third prior arts do not have to teach the limitations the primary prior art has taught, e.g., the combination of Biswas and Blyth has taught the limitation “simultaneously enhancing, …, two or more channels in …”, and thus, Pascual do not need to disclose the limitation of “simultaneously enhancing …” as argued above. Therefore, based on the analyses and evidences above, the prior art rejection of claim 53 under 35 U.S.C. 103(a) maintained. For the at least similar reasons described above, the prior art rejection of other independent claim 68 under 35 U.S.C. 103(a) maintained. Similarly, the prior art rejection of dependent claims 54-58, 60-65, 69-72 maintained. It is further noted that claims 66-67 are currently withdrawn and it is suggested to cancel claims 66-67 because claims 66-67 recited training related subject matters that may have confliction with limitation of “the multi-channel generator is trained in the dynamic range reduced domain in a GAN setting” as recited in claim 53. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to LESHUI ZHANG whose telephone number is (571)270-5589. The examiner can normally be reached Monday-Friday 6:30amp-4:00pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vivian Chin can be reached at 571-272-7848. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /LESHUI ZHANG/ Primary Examiner, Art Unit 2695
Read full office action

Prosecution Timeline

Show 12 earlier events
Dec 29, 2025
Non-Final Rejection mailed — §103, §112
Mar 20, 2026
Applicant Interview (Telephonic)
Mar 20, 2026
Examiner Interview Summary
Jun 17, 2026
Interview Requested
Jun 23, 2026
Examiner Interview Summary
Jun 23, 2026
Applicant Interview (Telephonic)
Jun 26, 2026
Response Filed
Sep 09, 2026
Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749492
METHOD FOR QUANTIZING LINE SPECTRAL FREQUENCIES
1y 11m to grant Granted Sep 29, 2026
Patent 12740755
Digital Auscultation Device
2y 5m to grant Granted Sep 22, 2026
Patent 12744040
METHOD FOR PROCESSING MISRECOGNIZED AUDIO SIGNALS, AND DEVICE THEREFOR
2y 3m to grant Granted Sep 22, 2026
Patent 12739591
APPARATUS AND METHOD FOR GENERATING CONTROL SIGNALS FOR A LOUDSPEAKER SYSTEM WITH SPECTRAL INTERLACING IN THE LOWER FREQUENCY RANGE
2y 5m to grant Granted Sep 15, 2026
Patent 12731592
COMBINING SPATIAL AUDIO STREAMS
2y 11m to grant Granted Sep 08, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
78%
Grant Probability
99%
With Interview (+35.3%)
2y 9m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 954 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month