DETAILED ACTION
This communication is in response to the Application filed on November 27, 2024.
Claims 1 - 20 are pending and have been examined.
Claims 1, 11 and 20 are independent.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on July 25, 2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Drawings
The drawings filed on November 27, 2024 have been accepted and considered by the Examiner.
Double Patenting Note
The Examiner notes that previously published patents U.S. Patent Application Publication 2025/0140265 was analyzed for Double Patenting. However, based on the current claim scope no Double patenting was found.
Claim Rejections - 35 USC § 102
The following is a quotation of pre-AIA 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1, 11 and 20 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Brendel et al., (WO 2026041746 A1), hereinafter referred to as Brendel.
Regarding Claims 1, 11 and 20 Brendel teaches:
1. A method performed by at least one processor, the method comprising, 11. an apparatus comprising, and 20. A non-transitory computer readable medium having instructions stored therein:
at least one memory configured to store program code; and [Brendel, “The implementation can be performed using a digital storage medium, for example a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable (i.e., the claimed “program code”) computer system such that the respective method is performed. Therefore, the digital storage medium may be computer readable.” Pg. 13, Ln. 29-35 - Pg. 14, Ln.1-4]
at least one processor configured to read the program code and operate as instructed by the program code, the program code including: [Brendel, “According to further embodiments, the audio processor may comprise a downstream task module configured to process the enhanced (audio) signal to obtain an output (audio) signal or, in general output data. For audio coding (i.e., the claimed “program code”) chosen as the downstream task, the output is a bitstream (if sender and receiver is treated together, the output signal is indeed an audio signal), for speech recognition the output is text, for scene classification the output of the downstream task are labels.” Pg. 2, Ln. 1-6]
which when executed by a processor cause the processor to execute a method comprising: [Brendel, “According to further embodiments, the audio processor may comprise a downstream task module configured to process the enhanced (audio) signal to obtain an output (audio) signal or, in general output data.” Pg. 2, Ln. 1-6]
PNG
media_image1.png
365
449
media_image1.png
Greyscale
inputting a mixture signal comprising speech, background noise, and reverberation into a denoising stage that generates a denoised mixture signal with reduced noise compared to the mixture signal; [Brendel clearly shows in Figure 3 a progressive pipeline with two sequential stages (1st stage: speech enhancement/denoise and 2nd stage: codec) to handle both noise and reverberation: “An embodiment provides an audio processor comprising an audio enhancement module, like a speech enhancement module, configured to process an input audio signal (i.e., the claimed “speech, background noise, and reverberation”) to obtain an enhanced audio signal (i.e., the claimed “denoised mixture signal with reduced noise”).” Pg. 1, Ln. 30-32; “In other words, this means that according to embodiments the audio processor primarily consists of a speech enhancer (i.e., the claimed “denoising stage”) and a neural codec, wherein the speech enhancer may reduce interfering sources, e.g., background noise, but also, depending on the application, other undesired signal components such as reverberation etc. (i.e., the claimed “mixture signal comprising speech, background noise, and reverberation”).” Pg. 3, Ln. 8-11; “The entity 12 may be an audio enhancement module or speech enhancement module which can be configured to reduce interferences from the input audio signal IS to output the enhanced signal ES without or with reduced interferences. In other words, this means that the entity 12 may be configured to reduce interfering sources, e.g., background noise, but also reverberation, etc. (i.e., the claimed “mixture signal comprising speech, background noise, and reverberation”)” Pg. 9, Ln. 4-8]
fusing the denoised mixture signal with the mixture signal to generate a fused signal; and [Brendel Figure 3 “18 combination”; Brendel clearly shows in Figure 3 fusing/combination: Brendel, “The result output by the combiner is a combination or weighted combination (i.e., the claimed “fusing”) of ES and IS (i.e., the claimed “fusing the denoised mixture signal with the mixture signal to generate a fused signal”).” Pg. 11, Ln. 31-32]
inputting the fused signal into a dereverberation and restoration (DR) stage to generate an output signal with reduced reverberation compared to the mixture signal. [Brendel clearly shows in Figure 3 inputting fused signal/combination signal into codec/dereverberation and restoration (DR) stage: “In other words, this means that according to embodiments the audio processor primarily consists of a speech enhancer and a neural codec (i.e., the claimed “dereverberation and restoration (DR) stage”), wherein the speech enhancer may reduce interfering sources, e.g., background noise, but also, depending on the application, other undesired signal components such as reverberation etc..” Pg. 3, Ln. 8-11]
Claim Rejections - 35 USC § 103
The following is a quotation of pre-AIA 35 U.S.C. 103(a) which forms the basis for all obviousness rejections set forth in this Office action:
(a) A patent may not be obtained though the invention is not identically disclosed or described as set forth in section 102 of this title, if the differences between the subject matter sought to be patented and the prior art are such that the subject matter as a whole would have been obvious at the time the invention was made to a person having ordinary skill in the art to which said subject matter pertains. Patentability shall not be negatived by the manner in which the invention was made.
Claims 2 and 12 are rejected under 35 U.S.C. 103(a) as being unpatentable over Brendel as applied in claim 1 above in view of Xiong et al., (CN114242043A), hereinafter referred to as Xiong.
Regarding Claims 2 and 12, Brendel has been discussed above. Brendel further teaches:
wherein inputting the mixture signal into the denoising stage further comprises: [Brendel, see mapping applied to claim 1]
inputting the mixture signal into a long short-term memory (LSTM) neural network that generates the denoised mixture signal. [Brendel, see mapping applied to claim 1]
Brendel fails to teach long short-term memory (LSTM) neural network.
However, Xiong teaches:
inputting the mixture signal into a long short-term memory (LSTM) neural network that generates the denoised mixture signal. [Xiong, “Long Short-Term Memory (LSTM) networks are a type of recurrent neural network that can learn long term dependencies.” Par. n0059]
Brendel and Xiong pertain to voice processing systems and are analogous to the instant application. Accordingly, it would have been obvious to one of ordinary skill in the voice processing systems art to modify Brendel’s teachings of “speech enhancer (i.e., the claimed “denoising stage”) and a neural codec” (Brendel, Pg. 3, Ln. 8-11) with the teachings of “Long Short-Term Memory (LSTM) networks” (Xiong, Par. n0059) taught by Xiong in order to “achieve both noise reduction and reverberation reduction” (Xiong, Par. n0003).
Claims 6 - 8 and 16 - 18 are rejected under 35 U.S.C. 103(a) as being unpatentable over Brendel as applied in claim 1 above, and in further view of Pia et al., (WO2023175198A1), hereinafter referred to as Pia.
Regarding Claims 6 and 16, Brendel has been discussed above. Brendel further teaches:
wherein the inputting the fused signal into the DR stage further comprises: [Brendel, see mapping applied to claim 1]
inputting the fused signal into an encoder to generate an encoded fused signal; [Brendel, see mapping applied to claim 1; “The denoised signal is then further processed by the neural codec encoder 16, so as to obtain the bitstream BS (i.e., the claimed “encoded fused signal”).” Pg. 10, Line 3-4; “For example, the downstream task model may comprise a neural audio/speech codec or encoder, or a system involving a neural audio/speech encoder and decoder.” Pg. 2, Ln. 7-9]
inputting the quantized signal into a decoder to generate the output signal. [Brendel, see mapping applied to claim 1; “For example, the downstream task model may comprise a neural audio/speech codec or encoder, or a system involving a neural audio/speech encoder and decoder.” Pg. 2, Ln. 7-9]
While bitstream output taught by Brendel is quantized, Brendel fails to teach explicitly teach quantizer.
However, Pia teaches:
inputting the encoded fused signal into a quantizer to generate a quantized signal; and [Pia, “a learnable quantizer to associate, to each frame of the first multi-dimensional or a processed version of the first multi-dimensional audio signal representation of the input audio signal, indexes of at least one codebook, so as to generate the bitstream.” Pg. 2, Ln.32-33 - Pg. 3, Ln. 1-2]
inputting the quantized signal into a decoder to generate the output signal. [Pia, “There are presented vocoder techniques and more in general techniques for generating an audio signal representation (e.g., a bitstream) and for generating an audio signal (i.e. the claimed “output signal”) (e.g., at a decoder).” Pg. 1, Ln. 4-5]
Brendel and Pia pertain to voice processing systems and are analogous to the instant application. Accordingly, it would have been obvious to one of ordinary skill in the voice processing systems art to modify Brendel’s teachings of “speech enhancer (i.e., the claimed “denoising stage”) and a neural codec” (Brendel, Pg. 3, Ln. 8-11) with the teachings of “quantizer” (Pia, Pg. 2, Ln.32-33) taught by Pia in order to “achieve robustness against a wide range of different types of noises and reverberation. (Pia, Pg. 52, Ln. 16-17).
Regarding Claims 7 and 17, Brendel in view of Pia has been discussed above. The combination further teaches:
wherein the quantizer quantizes the encoded fused signal in accordance with scalar quantization. [Brendel, see mapping applied to claims 1, 6; Pia, see mapping applied to claim 6; Pia, “It is noted that that the quantizer 300 may be inputted with a scalar, a vector, or more in general a tensor.” Pg. 33, Ln. 30-31]
Regarding Claims 8 and 18, Brendel in view of Pia has been discussed above. The combination further teaches:
wherein the quantizer quantizes the encoded fused signal in accordance with vector quantization. [Brendel, see mapping applied to claims 1, 6; Pia, see mapping applied to claim 6; Pia, “It is noted that that the quantizer 300 may be inputted with a scalar, a vector, or more in general a tensor.” Pg. 33, Ln. 30-31]
Claims 9 and 19 are rejected under 35 U.S.C. 103(a) as being unpatentable over Brendel in view of Xiong as applied in claim 5 above, and in further view of Pia et al., (WO2023175198A1), hereinafter referred to as Pia.
Regarding Claims 9 and 19, Brendel in view of Xiong has been discussed above. The combination fails to teach quantizer quantizes the encoded fused signal in accordance with a combination of scalar quantization and vector quantization.
However, Pia teaches:
wherein the quantizer quantizes the encoded fused signal in accordance with a combination of scalar quantization and vector quantization. [Brendel, see mapping applied to claims 1, 6; Pia, see mapping applied to claim 6; Pia, “Therefore, at the quantizer 300 there may be a multiplicity (i.e. the claimed “combination”) of residual codebooks, so that: the second residual codebook qe associates to indexes to be encoded in the audio signal representation, codes (e.g., scalar (i.e., the claimed “scalar quantization”, vector (i.e., the claimed “vector quantization”),” Pg. 36, Ln. 5-8]
Brendel, Xiong, Pia and pertain to voice processing systems and are analogous to the instant application. Accordingly, it would have been obvious to one of ordinary skill in the voice processing systems art to modify Brendel’s teachings of “speech enhancer (i.e., the claimed “denoising stage”) and a neural codec” (Brendel, Pg. 3, Ln. 8-11) with the teachings of “Long Short-Term Memory (LSTM) networks” (Xiong, Par. n0059) taught by Xiong and the teachings of “quantizer” (Pia, Pg. 2, Ln.32-33) taught by Pia in order to “achieve both noise reduction and reverberation reduction” (Xiong, Par. n0003) and “achieve robustness against a wide range of different types of noises and reverberation (Pia, Pg. 52, Ln. 16-17).
Allowable Subject Matter
Claims 3 - 5, 10, and 13 - 15 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Regarding Claim 3, although Brendel teaches “speech enhancer (i.e., the claimed “denoising stage”) and a neural codec” (Brendel, Pg. 3, Ln. 8-11) and Xiong teaches “Long Short-Term Memory (LSTM) networks” (Xiong, Par. n0059), none of the combination teach “LSTM is pre-trained in accordance with a loss function determined in accordance with (i) a scale-invariant signal-to-distortion ratio (SI-SDR) of the denoised mixture signal and a target reverberant signal and (ii) a determined loss on a spectrum magnitude difference between the denoised mixture signal and the target reverberant signal.”
Claim 13 is recited similar to Claim 3 and also contain similar allowable subject matter.
Claims 4 - 5 and 14 - 15 depend on Claims 3 and 13, respectively, and therefore are allowable by virtue of their dependency.
Regarding Claim 10, although Brendel teaches “speech enhancer (i.e., the claimed “denoising stage”) and a neural codec” (Brendel, Pg. 3, Ln. 8-11) and Pia teaches “quantizer” (Pia, Pg. 2, Ln.32-33), none of the combination teach “wherein the DR stage is trained in accordance with a (i) a first loss between a target output signal and the output signal, (ii) a second loss using a hinge loss based on one or more multi-scale short-time Fourier transform (STFT) discriminators, (iii) a third loss in accordance with an L1 feature loss, and (iv) a loss between the quantized signal and the encoded fused signal.”
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Ji et al., (U.S. Patent Application Publication 2022/0122597) teaches speech enhancement with simultaneous dereverberation and denoising.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to EUNICE LEE whose telephone number is 571-272-1886. The examiner can normally be reached M-F 8:00 AM - 5:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bhavesh Mehta can be reached on 571-272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/EUNICE LEE/Examiner, Art Unit 2656
/BHAVESH M MEHTA/Supervisory Patent Examiner, Art Unit 2656