DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55. Accordingly, the claims are given the benefit of the earlier filing date of GB 2319935.9, filed December 22nd, 2023.
Information Disclosure Statement
The information disclosure statements (IDS) submitted on April 14th, 2025 and June 3rd, 2026 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statements are being considered by the examiner.
Drawings
The drawings are objected to under 37 CFR 1.83(a) because they fail to show element names or functions as described in the specification. Figures 1 and 5 show interrelated shapes and reference numbers, but no further naming or labeling to illustrate the nature or purpose of the shapes. Figure 10 shows an unlabeled plane with numbers, having no icons or names, plotted about the plane. Any structural detail that is essential for a proper understanding of the disclosed invention should be shown in the drawing. MPEP § 608.02(d).
Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as “amended.” If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d).
Specification
The disclosure is objected to because of the following informalities: 37 CFR 1.52(b)(6) reads as follows:
Other than in a reissue application or reexamination or supplemental examination proceeding, the paragraphs of the specification, other than in the claims or abstract, may be numbered at the time the application is filed, and should be individually and consecutively numbered using Arabic numerals, so as to unambiguously identify each paragraph. The number should consist of at least four numerals enclosed in square brackets, including leading zeros (e.g., [0001]). The numbers and enclosing brackets should appear to the right of the left margin as the first item in each paragraph, before the first word of the paragraph, and should be highlighted in bold. A gap, equivalent to approximately four spaces, should follow the number. Nontext elements (e.g., tables, mathematical or chemical formulae, chemical structures, and sequence data) are considered part of the numbered paragraph around or above the elements, and should not be independently numbered. If a nontext element extends to the left margin, it should not be numbered as a separate and independent paragraph. A list is also treated as part of the paragraph around or above the list, and should not be independently numbered. Paragraph or section headers (titles), whether abutting the left margin or centered on the page, are not considered paragraphs and should not be numbered.
The paragraphs of the Specification are not numbered. Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claims recite a mental process that can be performed in the human mind or with the aid of pen and paper. This judicial exception is not integrated into a practical application because a computer is invoked merely as a tool to execute an abstract idea. The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because an abstract idea is merely applied on a generic computer without any element that would otherwise preclude performance of the abstrac.
Regarding claim 1, the claim recites “An apparatus for noise suppression comprises:at least one processor; and at least one memory including computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to:obtain at least one audio signal for a current frame or one or more previous frames, based on at least two microphone signals for the current frame or one or more previous frames;use a program code to predict an output signal for a future frame based, at least in part, on the at least one audio signal for the current frame or one or more previous frames; anduse the output signal for processing the future frame of the at least two microphone signals in a first audio signal process and using the output signal for processing the future frame of an output of the first audio signal process in a second audio signal process to enable noise suppression.”
The limitations of “obtain at least one audio signal…” “predict an output signal…” and “use… a first audio signal process and… a second audio signal process to enable noise suppression” as drafted cover broad activities which can be performed in the mind or with the aid of pen and paper. Taken individually, or as a whole, these limitations describe acts which are equivalent to human mental work of reading numerical data and predicting patterns.
The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the steps of the claimed invention can be performed mentally, and no additional features in the claims would preclude them from being performed as such. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claim 2, the claim depends from claim 1, and thus recites the limitations of claim 1, “wherein the at least one audio signal comprises an output of the first audio signal process.”
Taken individually, or as a whole with claim 1, these limitations describe methods of organizing human activity. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claim 3, the claim depends from claim 1, and thus recites the limitations of claim 1, “wherein the first audio signal process and the second audio signal process are consecutive processes in which the output of the first audio signal process is provided as an input to the second audio signal process.”
Taken individually, or as a whole with claim 1, these limitations describe methods of organizing human activity. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claim 4, the claim depends from claim 1, and thus recites the limitations of claim 1, “wherein the first audio signal process comprises a beamforming process, and wherein the beamforming process comprises processing the future frame of the at least two microphone signals using the output signal.”
The limitation of “a beamforming process” as drafted covers technology that is well-understood and routine in the art. At the time of filing, beamforming was well-understood and readily available to a person having ordinary skill in the art of digital signal processing and noise suppression. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claim 5, the claim depends from claim 4, and thus recites the limitations of claims 1 and 4, “wherein the output signal comprises a gain to be applied to at least one of the at least two microphone signals of the beamforming process.”
Taken individually, or as a whole with the preceding claims, these limitations describe methods of organizing human activity. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claim 6, the claim depends from claim 4, and thus recites the limitations of claims 1 and 4, “wherein the output signal comprises an amplitude to be used for at least one of the at least two microphone signals of the beamforming process.”
Taken individually, or as a whole with the preceding claims, these limitations describe methods of organizing human activity. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claim 7, the claim depends from claim 1, and thus recites the limitations of claim 1, “wherein the second audio signal process comprises a spectral noise suppression process.”
Taken individually, or as a whole with claim 1, these limitations describe methods of organizing human activity. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claim 8, the claim depends from claim 7, and thus recites the limitations of claims 1 and 7, “wherein the output signal comprises a gain to be applied to the input of the spectral noise suppression process.”
Taken individually, or as a whole with the preceding claims, these limitations describe methods of organizing human activity. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claim 9, the claim depends from claim 7, and thus recites the limitations of claims 1 and 7, “wherein the output signal comprises an amplitude to be used for the spectral noise suppression process.”
Taken individually, or as a whole with the preceding claims, these limitations describe methods of organizing human activity. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claim 10, the claim depends from claim 1, and thus recites the limitations of claim 1, “wherein the output signal is applied to the future frame of the respective audio signal processes in a frequency domain.”
Taken individually, or as a whole with claim 1, these limitations describe methods of organizing human activity. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claim 11, the claim depends from claim 1, and thus recites the limitations of claim 1, “wherein the program code receives a single input and provides a single output.”
Taken individually, or as a whole with claim 1, these limitations describe methods of organizing human activity. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claim 12, the claim depends from claim 1, and thus recites the limitations of claim 1, “wherein the program code comprises a machine learning program.”
The limitation of “a machine learning program” as drafted covers technology that is well-understood and routine in the art. At the time of filing, machine learning was well-understood and readily available to a person having ordinary skill in the art of digital signal processing and noise suppression. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claim 13, the claim depends from claim 12, and thus recites the limitations of claims 1 and 12, “wherein the machine learning program comprises a neural network circuit.”
The limitation of “a neural network circuit” as drafted covers technology that is well-understood and routine in the art. At the time of filing, neural networks were well-understood and readily available to a person having ordinary skill in the art of digital signal processing and noise suppression. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claim 14, the claim depends from claim 1, and thus recites the limitations of claim 1, “wherein the same output signal is applied to future frames of multiple audio signal processes.”
Taken individually, or as a whole with claim 1, these limitations describe methods of organizing human activity. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claim 15, the claim depends from claim 1, and thus recites the limitations of claim 1, “wherein the number of current or previous frames in the obtained audio signal that are used to predict the output signal for a future frame are selected based, at least in part, on latency requirements.”
Taken individually, or as a whole with claim 1, these limitations describe methods of organizing human activity. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claim 16, the claim depends from claim 1, and thus recites the limitations of claim 1, “wherein the apparatus is for use in an audio communication setting.”
Taken individually, or as a whole with claim 1, these limitations describe methods of organizing human activity. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claim 17, the claim depends from claim 16, and thus recites the limitations of claims 1 and 16, “wherein the audio communication setting is at least one of;a one-way communication setting; ora two-way communication setting.”
Taken individually, or as a whole with the preceding claims, these limitations describe methods of organizing human activity. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claim 18, the claim depends from claim 1, and thus recites the limitations of claim 1, “wherein the apparatus is at least one of:a telephone;a camera;a computing device;a teleconferencing device;a television;a virtual reality device; oran augmented reality device.”
At the time of filing, each of the listed embodiments was well-understood and readily available to a person having ordinary skill in the art of digital signal processing and noise suppression. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claim 19, apparatus claim 19 and method claim 1 are related as a method and system of using the same, with each system element’s function corresponding to the method step. Accordingly, claim 19 are similarly rejected under the same rationale as applied to claim 1.
Regarding claim 20, computer-readable medium claim 20 and appartus claim 1 are related as system and computer-readable medium for performing the same, with each computer-readable medium element’s function corresponding to the method step. Accordingly, claim 20 are similarly rejected under the same rationale as applied to claim 1.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-3, 7-14 and 16-20 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by U.S. Patent Application Publication 2018/0040333 to Wung et al. (hereinafter, “Wung”).
Regarding claims 1, 19 and 20, Wung teaches a method, system and computer-readable medium for noise suppression comprising: at least one processor; and at least one memory including computer program code, the at least one memory and the computer program code (paragraph [0039], "Keeping the above points in mind, FIG. 8 is a block diagram illustrating components that may be present in one such electronic device 10, and which may allow the device 10 to function in accordance with the techniques discussed herein. The various functional blocks shown in FIG. 8 may include hardware elements (including circuitry), software elements (including computer code stored on a computer-readable medium, such as a hard drive or system memory), or a combination of both hardware and software elements… For example, in the illustrated embodiment, these components may include a display 12, input/output (I/O) ports 14, input structures 16, one or more processors 18, memory device(s) 20, non-volatile storage 22, expansion card(s) 24, RF circuitry 26, and power source 28.") configured to, with the at least one processor, cause the apparatus at least to:
obtain at least one audio signal for a current frame or one or more previous frames, based on at least two microphone signals for the current frame or one or more previous frames (paragraph [0017], "FIG. 2 illustrates a block diagram of a system 200 for performing speech enhancement using a Deep Neural Network (DNN)-based signal according to one embodiment of the invention. System 200 may be included in the electronic device 10 and comprises a microphone 120 and a loudspeaker 130. While the system 200 in FIG. 2 includes only one microphone 120, it is understood that at least one of the microphones in the electronic device 10 may be included in the system 200. Accordingly, a plurality of microphone 120 may be included in the system 200. It is further understood that the at least one microphone 120 may be included in a headset used with the electronic device 10.");
use a program code to predict an output signal for a future frame based, at least in part, on the at least one audio signal for the current frame or one or more previous frames (paragraph [0023], "Once the DNN 170 is trained offline, the DNN 170 in FIG. 2 receives the microphone signal, the reference signal, the AEC echo-cancelled signal, and an estimated loudspeaker signal in the frequency domain from the time-frequency transformer 160. In the embodiment in FIG. 2, the DNN 170 generates a clean speech signal in the frequency domain," and paragraph [0033], "Referring back to FIG. 5, each of the feature buffers 3501-3504 receives the outputs of the first and second normalization units 6301, 6302 from each of the feature processors 4101-4104. Each of the feature buffers 3501-3504 may stack (or buffer) the extracted features, respectively, with a number of past or future frames."); and
use the output signal for processing the future frame of the at least two microphone signals in a first audio signal process and using the output signal for processing the future frame of an output of the first audio signal process in a second audio signal process to enable noise suppression (paragraph [0028], "As shown in FIG. 3, the DNN 370 transmits the speech reference signal to a noise suppressor 390. In one embodiment, the noise suppressor 390 may also receive the AEC echo-cancelled signal in the frequency domain from the time-frequency transformer 160. The noise suppressor 390 suppresses the noise or residual echo in the AEC echo-cancelled signal based on the speech reference and outputs a clean speech signal in the frequency domain to the frequency-time transformer 180. As in FIG. 2, the frequency-time transformer 180 in FIG. 3 transforms the clean speech signal in the frequency domain to a clean speech signal in the time domain.").
Regarding claim 2, Wung further teaches a system wherein the at least one audio signal comprises an output of the first audio signal process (paragraph [0028], "As shown in FIG. 3, the DNN 370 transmits the speech reference signal to a noise suppressor 390. In one embodiment, the noise suppressor 390 may also receive the AEC echo-cancelled signal in the frequency domain from the time-frequency transformer 160. The noise suppressor 390 suppresses the noise or residual echo in the AEC echo-cancelled signal based on the speech reference and outputs a clean speech signal in the frequency domain to the frequency-time transformer 180. As in FIG. 2, the frequency-time transformer 180 in FIG. 3 transforms the clean speech signal in the frequency domain to a clean speech signal in the time domain.").
Regarding claim 3, Wung further teaches a system wherein the first audio signal process and the second audio signal process are consecutive processes in which the output of the first audio signal process is provided as an input to the second audio signal process (paragraph [0028], "As shown in FIG. 3, the DNN 370 transmits the speech reference signal to a noise suppressor 390. In one embodiment, the noise suppressor 390 may also receive the AEC echo-cancelled signal in the frequency domain from the time-frequency transformer 160. The noise suppressor 390 suppresses the noise or residual echo in the AEC echo-cancelled signal based on the speech reference and outputs a clean speech signal in the frequency domain to the frequency-time transformer 180. As in FIG. 2, the frequency-time transformer 180 in FIG. 3 transforms the clean speech signal in the frequency domain to a clean speech signal in the time domain.").
Regarding claim 7, Wung further teaches a system wherein the second audio signal process comprises a spectral noise suppression process (paragraph [0028], "As shown in FIG. 3, the DNN 370 transmits the speech reference signal to a noise suppressor 390. In one embodiment, the noise suppressor 390 may also receive the AEC echo-cancelled signal in the frequency domain from the time-frequency transformer 160. The noise suppressor 390 suppresses the noise or residual echo in the AEC echo-cancelled signal based on the speech reference and outputs a clean speech signal in the frequency domain to the frequency-time transformer 180. As in FIG. 2, the frequency-time transformer 180 in FIG. 3 transforms the clean speech signal in the frequency domain to a clean speech signal in the time domain.").
Regarding claim 8, Wung further teaches a system wherein the output signal comprises a gain to be applied to the input of the spectral noise suppression process (paragraph [0022], "In another embodiment, the output of the DNN 170 may be a training gain function (e.g., an oracle gain function or an signal approximation of the gain function) to be applied to the noise speech signal instead of a signal approximation of the clean speech signal.").
Regarding claim 9, Wung further teaches a system wherein the output signal comprises an amplitude to be used for the spectral noise suppression process (paragraph [0022], "In another embodiment, the output of the DNN 170 may be a training gain function (e.g., an oracle gain function or an signal approximation of the gain function) to be applied to the noise speech signal instead of a signal approximation of the clean speech signal.").
In audio signals, gain and amplitude are constituents of the same process; by adjusting a gain value, a system would intrinsically adjust the output amplitude. Given that the outcomes are functionally identical, claims 8 and 9, under their broadest reasonable interpretation, teach indistinct methods, and thus Wung teaches the limitations of each claim.
Regarding claim 10, Wung further teaches a system wherein the output signal is applied to the future frame of the respective audio signal processes in a frequency domain (paragraph [0033], "Referring back to FIG. 5, each of the feature buffers 3501-3504 receives the outputs of the first and second normalization units 6301, 6302 from each of the feature processors 4101-4104. Each of the feature buffers 3501-3504 may stack (or buffer) the extracted features, respectively, with a number of past or future frames.").
Regarding claim 11, Wung further teaches a system wherein the program code receives a single input and provides a single output (paragraph [0017], "While the system 200 in FIG. 2 includes only one microphone 120, it is understood that at least one of the microphones in the electronic device 10 may be included in the system 200," and paragraph [0028], "As shown in FIG. 3, the DNN 370 transmits the speech reference signal to a noise suppressor 390.").
Regarding claim 12, Wung further teaches a system wherein the program code comprises a machine learning program (paragraph [0022], "The DNN 170 in FIG. 2 is trained offline by exciting the at least one microphone using a target training signal that includes a signal approximation of clean speech.").
Regarding claim 13, Wung further teaches a system wherein the machine learning program comprises a neural network circuit (paragraph [0001], "An embodiment of the invention relate generally to a system and method for performing speech enhancement using a deep neural network-based signal.").
Regarding claim 14, Wung further teaches a system wherein the same output signal is applied to future frames of multiple audio signal processes (paragraph [0033], "Referring back to FIG. 5, each of the feature buffers 3501-3504 receives the outputs of the first and second normalization units 6301, 6302 from each of the feature processors 4101-4104. Each of the feature buffers 3501-3504 may stack (or buffer) the extracted features, respectively, with a number of past or future frames.").
Regarding claim 16, Wung further teaches a system wherein the apparatus is for use in an audio communication setting (paragraph [0015], "FIG. 1 depicts near-end user and a far-end user using an exemplary electronic device in which an embodiment of the invention may be implemented. The electronic device 10 may be a mobile communications handset device such as a smart phone or a multi-function cellular phone.").
Regarding claim 17, Wung further teaches a system wherein the audio communication setting is at least one of; a one-way communication setting; or a two-way communication setting (paragraph [0015], "In the embodiment in FIG. 1, the near-end user is in the process of a call with a far-end user who is using another communications device 4. The term 'call' is used here generically to refer to any two-way real-time or live audio communications session with a far-end user (including a video call which allows simultaneous audio).").
Regarding claim 18, Wung further teaches a system wherein the apparatus is at least one of: a telephone; a camera; a computing device; a teleconferencing device; a television; a virtual reality device; or an augmented reality device (paragraph [0041], "The electronic device 10 may also take the form of other types of devices, such as mobile telephones, media players, personal data organizers, handheld game platforms, cameras, and/or combinations of such devices.").
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 4-6 are rejected under 35 U.S.C. 103 as being unpatentable over Wung as applied to claim 1 in view of U.S. Patent Application Publication 2024/0371389 to Mosayyebpour Kaskari et al. (hereinafter, “Kaskari”).
Regarding claim 4, Wung does not explicitly teach the use of beamforming, and thus, Kaskari is referenced to teach a system wherein the first audio signal process comprises a beamforming process, and wherein the beamforming process comprises processing the future frame of the at least two microphone signals using the output signal (paragraph [0023], "Various aspects relate generally to audio signal processing, and more particularly, to speech enhancement techniques that combine statistical signal processing with neural network inferencing. In some aspects, a speech enhancement system may include a linear filter, a deep neural network (DNN), and a nonlinear post-filter... The DNN infers a probability of speech in the denoised audio signal and produces a speech signal and a noise signal (representing a speech component and a noise component, respectively, of the audio signal) based on the inferred probability of speech. In some implementations, the probability of speech also may be used to update various parameters of the linear filter (such as a vector of weights associated with a multi-frame beamformer)," and paragraph [0033], "The series of input audio frames X(l,k) includes the current audio frame to be processed (X0(l,k)), a number (c) of future audio frames (Xfuture(l,k)) that follow the current audio frame X0(l,k) in time, and a number (d) of past audio frames (Xpast(l,k)) that precede the current audio frame X0(l,k) in time,").
Wung and Kaskari are considered analogous because they are each concerned with speech enhancement and noise suppression. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified Wung with the beamforming of Kaskari for the purpose of improving noise suppression performance. Given that all the claimed elements were known in the prior art, one skilled in the art could have combined the elements by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Regarding claim 5, Wung teaches a system wherein the output signal comprises a gain to be applied to at least one of the at least two microphone signals of the beamforming process (paragraph [0022], "In another embodiment, the output of the DNN 170 may be a training gain function (e.g., an oracle gain function or an signal approximation of the gain function) to be applied to the noise speech signal instead of a signal approximation of the clean speech signal.").
Wung and Kaskari are considered analogous because they are each concerned with speech enhancement and noise suppression. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have further modified Wung and Kaskari with the gain function of Wung for the purpose of improving noise suppression performance. Given that all the claimed elements were known in the prior art, one skilled in the art could have combined the elements by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Regarding claim 6, Wung teaches wherein the output signal comprises an amplitude to be used for at least one of the at least two microphone signals of the beamforming process (paragraph [0022], "In another embodiment, the output of the DNN 170 may be a training gain function (e.g., an oracle gain function or an signal approximation of the gain function) to be applied to the noise speech signal instead of a signal approximation of the clean speech signal.").
In audio signals, gain and amplitude are constituents of the same process; by adjusting a gain value, a system would intrinsically adjust the output amplitude. Given that the outcomes are functionally identical, claims 5 and 6, under their broadest reasonable interpretation, teach indistinct methods, and thus Wung teaches the limitations of each claim.
Wung and Kaskari are considered analogous because they are each concerned with speech enhancement and noise suppression. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have further modified Wung and Kaskari with the gain function of Wung for the purpose of improving noise suppression performance. Given that all the claimed elements were known in the prior art, one skilled in the art could have combined the elements by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Claim 15 is rejected under 35 U.S.C. 103 as being unpatentable over Wung as applied to claim 1 in view of U.S. Patent 12,073,847 to Fazeli et al. (hereinafter, “Fazeli”).
Regarding claim 15, Wung does not address process latency, and thus, Fazeli is referenced to teach a system wherein the number of current or previous frames in the obtained audio signal that are used to predict the output signal for a future frame are selected based, at least in part, on latency requirements (column 22, line 63, "As noted above, while the attention mechanism may be applied to both past frames and future frames, in the interest of avoiding or reducing latency, some embodiments of the present disclosure use only past frames.").
Wung and Fazeli are considered analogous because they are each concerned with speech enhancement and noise suppression. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified Wung with the latency considerations of Fazeli for the purpose of improving noise suppression performance. Given that all the claimed elements were known in the prior art, one skilled in the art could have combined the elements by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
U.S. Patent 9,640,194 to Nemala et al. teaches speech processing and noise suppression using machine learning techniques to apply gain masks.
U.S. Patent 11,398,241 to Mansour et al. teaches microphone-based noise suppression with beamforming.
U.S. Patent 12,562,176 to Mosayyebpour Kaskari et al. teaches audio signal processing using masking for noise suppression.
U.S. Patent Application Publication 2025/0157480 to Dubey et al. teaches noise suppression using DNN models to generate spectrogram masks.
U.S. Patent Application Publication 2023/0326475 to Tsiaflakis et al. teaches noise suppression using machine learning programs to create gain and spectrogram modifications.
U.S. Patent Application Publication 2015/0110284 to Niemisto et al. teaches noise suppression based on an array of microphones and a plurality of processed audio signals.
WIPO Publication 2019/199501 to Jelev et al. teaches audio frame processing using estimated masking with a trained neural network.
Great Britain Patent Application Publication 2398913 to Kelleher et al. teaches continually adaptable spectral filtering for noise suppression.
“STFT-Domain Neural Speech Enhancement with Very Low Algorithmic Latency” by Wang et al. as cited by Applicant.
“Deep Learning-Based Joint Control of Acoustic Echo Cancellation, Beamforming and Postfiltering” by Haubner et al. as cited by Applicant.
“Improved MVDR beamforming using single-channel mask prediction networks” by Erdogan et al. as cited by Applicant.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SEAN T SMITH whose telephone number is (571)272-6643. The examiner can normally be reached Monday - Friday 8:00am - 5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, PIERRE-LOUIS DESIR can be reached at (571) 272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SEAN THOMAS SMITH/Examiner, Art Unit 2659
/PIERRE LOUIS DESIR/Supervisory Patent Examiner, Art Unit 2659