Prosecution Insights
Last updated: August 17, 2026
Application No. 18/526,712

METHODS AND APPARATUSES FOR SPEECH ENHANCEMENT

Non-Final OA §101§102§103
Filed
Dec 01, 2023
Examiner
HUTCHESON, CODY DOUGLAS
Art Unit
2659
Tech Center
2600 — Communications
Assignee
Comcast Cable Communications LLC
OA Round
3 (Non-Final)
63%
Grant Probability
Moderate
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 63% of resolved cases
63%
Career Allowance Rate
19 granted / 30 resolved
+1.3% vs TC avg
Strong +42% interview lift
Without
With
+42.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 8m
Avg Prosecution
28 currently pending
Career history
65
Total Applications
across all art units

Statute-Specific Performance

§101
33.8%
-6.2% vs TC avg
§103
41.1%
+1.1% vs TC avg
§102
14.7%
-25.3% vs TC avg
§112
6.7%
-33.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 30 resolved cases

Office Action

§101 §102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 07/23/2026 has been entered. Response to Arguments 1. Regarding the rejection under 35 U.S.C. 101, Applicant's arguments filed 07/23/2026 have been fully considered but they are not persuasive. Step 2A Prong 1: Applicant first argues on pgs. 10-11 that the claimed inventions do not recite abstract ideas under Step 2A Prong 1 analysis. Specifically, Applicant argues that the amended limitations of “converting the input signal from a time domain to a set of time-frequency (TF) samples in a frequency domain” and “generating, based on the one or more TF losses applied to the one or more TF samples, an output signal by converting the TF samples from the frequency domain to the time domain, wherein the output signal comprises less non-speech than the input signal” are not performable as mental processes, and further argues that the step of determining a speech probability estimate has been mischaracterized as a mental process. The Examiner respectfully disagrees with these arguments. The amended claims still recite abstract ideas under Step 2A Prong 2. First, the amended limitations reciting the specific conversion of TF samples between time and frequency domains, while not mental processes, read as mathematical calculations, which also fall under the category of abstract ideas. Furthermore, the determining step for the speech probability estimate is recited at a high level of generality. This step provides no specific technical steps/computations that would preclude a person from being able to perform this step mentally with the aid of pen and paper, and thus this step can be grouped as a mental process in the form of an observation, evaluation, judgement, or opinion. Therefore, the amended claims still recite abstract ideas. Step 2A Prong 2: Applicant further argues on pgs. 11-14 that the amended claims integrate the judicial exception into a practical application under Step 2A Prong 2. Specifically, it is argued that the claimed invention reflects a technical improvement to noise reduction via speech probability estimates and TF losses which “enable selective suppression of non-stationary interference while preserving speech content, resulting in an output signal with reduced non speech.” (see pg. 13, section 3 ii) of Remarks). Furthermore, it is argued that the claims reflect a particular application of machine learning in a specific technological context, constituting a practical application (see pg. 14, section 5, para. 2 of Remarks). The Examiner respectfully disagrees with these arguments. The claims as currently amended do not contain any additional elements which when viewed with the claims as a whole integrate the judicial exception into a practical application via a technical improvement. The only additional limitations in the independent claims amount to generic computer components (e.g., computing device in claim 1, apparatus, processor, and memory in claim 30, CRM in claim 41), which do not integrate the judicial exception into a practical application as they do not impose any meaningful limits on practicing the abstract idea. The claimed step of determining speech probability estimates which is argued by Applicant as leading to the improvement of “selective suppression of non-stationary interference while preserving speech content” is not recited at a technical level (e.g. no specific models/architectures/technical steps are recited to show how this step is being performed), and thus falls under the mental process grouping and cannot integrate the judicial exception into a practical application. Further technical details in the claims would be needed which show how the claimed invention is a technical improvement. Hence, Applicant’s arguments are not persuasive. 2. Regarding the rejection under 35 U.S.C. 102 and 103, Applicant’s arguments have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. 3. Claims 1-11 and 30-51 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Regarding claims 1, 30, and 41, “A method”, “An apparatus”, and “One or more non-transitory computer-readable media” are recited, which is directed to one of the four statutory categories of invention (process, machine, and article of manufacture) (Step 1: YES). However, the claims limitations, under their broadest reasonable interpretation, recite mental processes or mathematical concepts which fall into the category of abstract idea (Step 2A Prong 1: YES). The following limitations recited in claim 1, and analogous claims recited in independent claims 30 and 41, under their broadest reasonable interpretation, recite mental processes or mathematical concepts: receiving, …, an input signal comprising speech and non-speech: a person listens to audio comprising speech and non-speech sounds converting the input signal from a time domain to a set of time-frequency (TF) samples in a frequency domain: converting from time to time-frequency (TF) samples of the input signal is a mathematical concept determining, based on the set of TF samples, a speech probability estimate that speech is present for each TF sample of the set of TF samples: a person analyzes the TF samples and determines a probability that there is speech in the sample determining, based on the speech probability estimate for each TF sample of the set of TF samples, one or more losses to be applied to one or more TF samples of the set of TF samples: determining losses to apply amounts to a mathematical concept. and generating, based on the one or more TF losses applied to the one or more TF samples, an output signal by converting the TF samples from the frequency domain to the time domain, wherein the output signal comprises less non-speech than the input signal: applying the losses to the samples to generate an output signal amounts to a mathematical concept; conversion from TF to time domain amounts to further a mathematical concept. Claims 1, 30, and 41 do not contain any additional elements which integrate the judicial exception into a practical application (Step 2A Prong 2: NO). The only additional limitations are “…by a computing device” (claim 1), “An apparatus comprising: one or more processors; and a memory storing processor-executable instructions that, when executed by the one or more processors, cause the apparatus to” (claim 30), and “One or more non-transitory computer-readable media storing processor-executable instructions that, when executed by at least one processor, cause the at least one processor to” (claim 41), which amount to mere instructions to implement the judicial exception using a generic computer. Even when viewed in combination, the mere instructions to implement the judicial exception using a generic computer do not integrate the judicial exception into a practical application as they do not impose any meaningful limits on practicing the abstract idea. Therefore, claims 1, 30, and 41 are directed to an abstract idea (Step 2A: YES). Claims 1, 30 and 41 do not contain any additional elements which amount to significantly more than the judicial exception (Step 2B: NO). As discussed above, the only additional limitation amount to mere instructions to implement the judicial exception using a generic computer. Even when viewed in combination, the mere instructions to implement the judicial exception using a generic computer do not amount to significantly more than the judicial exception as they do not provide an inventive concept. Therefore, claims 1, 30 and 41 are not patent eligible. Regarding dependent claims 2-11, 31-40, and 42-51, “The method”, “The apparatus”, and “The one or more non-transitory computer-readable media” are recited, which is directed to one of the four statutory categories of invention (process, machine, and article of manufacture) (Step 1: YES). However, the claims limitations, under their broadest reasonable interpretation, recite further mental processes or mathematical concepts which fall into the category of abstract idea (Step 2A Prong 1: YES). The following limitations recited in claims 2-11, and the analogous limitations recited in claims 31-40 and 42-51, under their broadest reasonable interpretation, recite mental processes or mathematical concepts: Claim 2, 31, and 42: wherein the speech probability estimate for the each TF sample of the set of TF samples is indicative of the speech being present in the each TF sample: a person determines the estimate by determining how likely speech is present in each sample. Claims 2, 31, and 42 contain no additional elements. Claim 3, 32, and 43: wherein the non-speech comprises stationary noise and non-stationary noise: a person listens to audio which has stationary noise (e.g. white noise), and non-stationary noise (e.g. wind) Claims 3, 32, and 43 contain no additional elements. Claim 4, 33, and 44: wherein the speech probability estimate further distinguishes the speech from the stationary noise and the non-stationary noise: a person determines the estimate to decide which samples contain speech and which samples contain noise. Claims 4, 33, and 44 contain no additional elements. Claim 5, 34, and 45: generating a labelled data set that comprises one or more input features and one or more indications indicative of the speech being present: a person writes down a data set of features and indications of speech being present. providing the labelled data set…to determine the speech probability estimate: a person uses the data to learn how to determine the estimate. Claim 5, 34, and 45 contain the additional limitation “to a machine learning mode, wherein the machine learning model is configured to…”, which amounts to mere instructions to implement the judicial exception using a generic computer. Claims 6, 35, and 46: receiving…a set of speech samples; applying, based on each of the set of speech samples, a speech weight to the each of the set of speech samples; receiving,…a set of non-speech samples; applying, based on each of the set of non-speech sample, a non-speech weight to the each of the set of non-speech samples; generating, by combining the speech weighted set of speech samples and the non-speech weighted set of non-speech samples, a speech augmented set; and extracting one or more input features from the speech augmented set: generating a combined speech weighted set by combining the speech weighted set and non-speech weighted set, and extracting input features from the speech augmented set amounts to mathematical concepts. Claims 6, 35, and 46 contain the additional limitation “by the computing device”, which amounts to mere instructions to implement the judicial exception using a generic computer. Claim 7, 36, and 47: wherein the one or more extracted input features comprise Mel Frequency Cepstrum Coefficients (MFCCs), Phonemes, Senones, and Mel Spectrogram: a person can write down phonemes they hear. Determining MFCCs, Senones, and Mel Spectrogram amount to mathematical concepts. Claims 7, 36, and 47 contain no additional elements. Claims 8, 37, and 48: receiving…a set of speech samples; applying, based on each of the set of speech samples, a speech weight to the each of the set of speech samples; receiving,…a set of non-speech samples; applying, based on each of the set of non-speech sample, a non-speech weight to the each of the set of non-speech samples; determining, based on the speech weighted set of speech samples and the non-speech weighted set of non-speech sample, the one or more indications indicative of the speech present: weighting a set of speech samples and a set of non-speech samples and using this to determine the one or more indications amounts to a mathematical concept. Claims 8, 37, and 48 contain the additional limitation “by the computing device”, which amounts to mere instructions to implement the judicial exception using a generic computer. Claim 9, 38, and 49: determining, based on at least one of a priori signal to noise (SNR) ratio, the speech probability estimate, or a posteriori SNR the one or more TF losses to be applied to the each of the set of TF samples: determining a prior/a posteriori SNR’s to determine TF losses amounts to a mathematical concept. Claims 9, 38, and 49 contain no additional elements. Claims 10, 39, and 50: wherein each of the set of time-frequency samples comprises a frequency bin narrowly filtered based on a frequency domain: determining samples comprising a frequency bin narrowly filtered amounts to a mathematical concept. Claims 10, 39, and 50 contain no additional elements. Claim 11, 40, and 51: wherein the input signal comprises one or more pulse code modulation (PCM) signals: using PCM signals as the input signal amounts to a mathematical concept. Claims 11, 40, and 51 contains no additional elements. Claims 2-11, 31-40, and 42-51 do not contain any additional elements which integrate the judicial exception into a practical application (Step 2A Prong 2: NO). As discussed above, the only additional limitations amount to mere instructions to implement the judicial exception using a generic computer. Even when viewed in combination, the mere instructions to implement the judicial exception using a generic computer do not integrate the judicial exception into a practical application as they do not impose any meaningful limits on practicing the abstract idea. Therefore, claims 2-11, 31-40, and 42-51 are directed to an abstract idea (Step 2A: YES). Claims 2-11, 31-40, and 42-51 do not contain any additional elements which amount to significantly more than the judicial exception (Step 2B: NO). As discussed above, the only additional limitations amount to mere instructions to implement the judicial exception using a generic computer. Even when viewed in combination, the mere instructions to implement the judicial exception using a generic computer do not amount to significantly more than the judicial exception as they do not provide an inventive concept. Therefore, claims 2-11, 31-40, and 42-51 are not patent eligible. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. 4. Claims 1-2, 9-10, 30-31, 38-39, 41-42, and 49-50 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Kaskari (US 2024/0304204 A1). Regarding claim 1, Kaskari discloses A method, comprising: receiving, by a computing device (Fig. 4, microphone 412), an input signal comprising speech and non-speech (Fig. 5, 510; para. 0027 “In some implementations, the sound waves 101 may include speech from the environment and/or user speech mixed with other environmental sounds, background noise, or interference (such as reverberant noise from a headset enclosure). Thus, the audio signal 102 may include a speech component and a noise component.”); converting the input signal from a time domain to a set of time-frequency (TF) samples in a frequency domain (Fig. 5, 520 and 530; para. 0060 “The speech enhancement system may transform the B*N time-domain samples into B*N first frequency-domain samples based on an N-point fast Fourier transform (FFT) (520). The speech enhancement system may transform the B*N first frequency-domain samples into B*N second frequency-domain samples based on a B-point FFT (530)…”); determining, based on the set of TF samples, a speech probability estimate that speech is present for each TF sample of the set of TF samples (Fig. 5, 540; para. 0060 “Further, the speech enhancement system may determine a probability of speech in the input signal based at least in part on the B*N second frequency-domain samples (540).”; para. 0024 “The DNN is configured to receive an input audio signal and infer a probability of a speech component (also referred to as “probability of speech”) in the input audio signal based on a neural network model. …The DNN module may generate a probability of speech in the fine frequency domain based on the further transformed signal.”; Fig. 2, 218); determining, based on the speech probability estimate for each TF sample of the set of TF samples, one or more TF losses to be applied to one or more TF samples of the set of TF samples (para. 0028 “In some implementations, the speech enhancement component 120 may determine a spectral suppression gain to be applied to the audio signal 102 based, at least in part, on a deep neural network (DNN) 122. For example, the DNN 122 may be trained to infer a likelihood or probability of speech in the time-frequency domain.”; para. 0029 “During the inferencing phase, the DNN 122 may determine a probability of speech in each frame of the audio signal 102, at each frequency index associated with the time-frequency domain, based on the classification results. The DNN 122 may further convert the probability of speech determined for each frequency index into a spectral suppression gain that can be used to suppress the noise component of the corresponding frame of the audio signal 102. For example, if there is a low probability of speech in a given frame of the audio signal 102 at a particular frequency index, the DNN 122 may apply a lower gain to reduce the power at that frequency index of the corresponding audio frame. As a result, the DNN 122 may dynamically attenuate the noise component of the audio signal 102 in the time-frequency domain.”); generating, based on the one or more TF losses applied to the one or more TF samples, an output signal by converting the TF samples from the frequency domain to the time domain, wherein the output signal comprises less non-speech than the input signal (Fig. 2, steps 220-226; para. 0044 “For example, the subband synthesis module 226 may transform a number N of time-frequency domain samples of Y(l,f) (or Z(l,f)), representing a frame of Y(l,f) (or Z(l,f)) in the time-frequency domain), to N time-domain samples representing an enhanced audio frame 204 of the enhanced audio signal y(t). In some implementations, the subband synthesis module 226 may perform the transformation from the time-frequency domain to the time domain using an inverse Fourier transform, such as an inverse FFT. The enhanced audio signal y(t) may be processed further and/or output into sound waves via a speaker.”; para. 0027 “The spectral suppression gain attenuates the power of the noise component of the audio signal 102, in a time-frequency domain, to produce an enhanced speech signal 104. Thus, the enhanced speech signal 104 may have a higher SNR than the audio signal 102.”). Regarding claim 2, Kaskari discloses wherein the speech probability estimate for the each TF sample of the set of TF samples is indicative of the speech being present in the each TF sample (para. 0029 “During the inferencing phase, the DNN 122 may determine a probability of speech in each frame of the audio signal 102, at each frequency index associated with the time-frequency domain, based on the classification results. The DNN 122 may further convert the probability of speech determined for each frequency index into a spectral suppression gain that can be used to suppress the noise component of the corresponding frame of the audio signal 102. For example, if there is a low probability of speech in a given frame of the audio signal 102 at a particular frequency index, the DNN 122 may apply a lower gain to reduce the power at that frequency index of the corresponding audio frame. As a result, the DNN 122 may dynamically attenuate the noise component of the audio signal 102 in the time-frequency domain.”). Regarding claim 9, Kaskari discloses determining, based on at least one of a priori signal to noise (SNR) ratio, the speech probability estimate, or a posteriori SNR, the one or more TF losses to be applied to the each of the set of TF samples (para. 0028 “In some implementations, the speech enhancement component 120 may determine a spectral suppression gain to be applied to the audio signal 102 based, at least in part, on a deep neural network (DNN) 122. For example, the DNN 122 may be trained to infer a likelihood or probability of speech in the time-frequency domain.”; para. 0029 “During the inferencing phase, the DNN 122 may determine a probability of speech in each frame of the audio signal 102, at each frequency index associated with the time-frequency domain, based on the classification results. The DNN 122 may further convert the probability of speech determined for each frequency index into a spectral suppression gain that can be used to suppress the noise component of the corresponding frame of the audio signal 102. For example, if there is a low probability of speech in a given frame of the audio signal 102 at a particular frequency index, the DNN 122 may apply a lower gain to reduce the power at that frequency index of the corresponding audio frame. As a result, the DNN 122 may dynamically attenuate the noise component of the audio signal 102 in the time-frequency domain.”). Regarding claim 10, Kaskri discloses wherein each of the set of TF samples comprises a frequency bin narrowly filtered based on a frequency domain (para. 0037 “The fine domain mapping module 214 is configured to map frames of the time-frequency domain signal X(l,f) (in the coarse domain) to a respective signal X(l,f,q) in the fine domain, where q is a sub-bin index indicating a number of sub-bins associated with a frequency bin (e.g., a given value of the frequency index f). In some implementations, the fine domain mapping module 214 may perform the mapping using an FFT of size B, and the number of sub-bins q per frequency bin is equal to B (e.g., q=1, . . . , B). Accordingly, if B=8, then the fine domain mapping module 214 applies an 8-point FFT to the time-frequency domain signal X(l,f) for each value of frequency bin index f, resulting in a signal X(l,f,q), where q=1, . . . , 8 for each value of f. More generally, the fine domain mapping module 214 maps B frames of the time-frequency domain signal X(l,f), obtained from the buffer 212, to the signal X(l,f,q) by applying a B-point FFT to the B frames of X(l,f) per value of f.”). Regarding claim 30, claim 30 is an apparatus claim with limitations similar to those in method claim 1, and is thus rejected under similar rationale. Additionally, Kaskari discloses An apparatus (Fig. 4, 400) comprising: one or more processors (Fig. 4, 420; para. 0058 “The processing system 420 may include any suitable one or more processors capable of executing scripts or instructions of one or more software programs stored in the speech enhancement system 400 (such as in the memory 430).”); and a memory (Fig. 4, 430; para. 0058 “The processing system 420 may include any suitable one or more processors capable of executing scripts or instructions of one or more software programs stored in the speech enhancement system 400 (such as in the memory 430).”) storing processor-executable instructions that, when executed by the one or more processors, cause the apparatus to (para. 0058 “The processing system 420 may include any suitable one or more processors capable of executing scripts or instructions of one or more software programs stored in the speech enhancement system 400 (such as in the memory 430).”). Regarding claims 31, 38, and 39, these claims are rejected under analogous reasons to claims 2, 9, and 10, respectively. Regarding claim 41, claim 41 is a non-transitory computer-readable media claim with limitations similar to those in method claim 1, and is thus rejected under similar rationale. Additionally, Kaskari discloses One or more non-transitory computer-readable media storing processor-executable instructions that, when executed by at least one processor, cause the at least one processor to (para. 0053 “The memory 430 also may include a non-transitory computer-readable medium (including one or more nonvolatile memory elements, such as EPROM, EEPROM, Flash memory, or a hard drive, among other examples) that may store at least the following software (SW) modules:”; para. 0058 “The processing system 420 may include any suitable one or more processors capable of executing scripts or instructions of one or more software programs stored in the speech enhancement system 400 (such as in the memory 430).”). Regarding claims 42, 49, and 50, these claims are rejected for analogous reasons to claims 2, 9, and 10, respectively. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 5. Claims 3-4, 32-33, and 43-44 are rejected under 35 U.S.C. 103 as being unpatentable over Kaskari in view of Thyssen & Borgstrom (US 2015/0071461 A1, hereinafter Thyssen). Regarding claim 3, Kaskari does not specifically disclose wherein the non-speech comprises stationary noise and non-stationary noise. Thyssen teaches wherein the non-speech comprises stationary noise and non-stationary noise (para. 0108 “Back-end SCS component 300 is configured to suppress multiple types of interfering sources (e.g., stationary noise, non-stationary noise, residual echo, etc.) present in a first signal 340.”). Kaskari and Thyssen are considered to be analogous to the claimed invention as they both are in the same field of speech enhancement. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Kaskari to incorporate the teachings of Thyssen in order to specifically analyze audio with non-speech comprising both stationary noise and non-stationary noise. Doing so would be beneficial, as this would allow for analysis and suppression of multiple different types of additive noise using different suppression branches specific to each noise type (Thyssen, para. 0035), which would allow for a wider variety of noise types to be suppressed to achieve an enhanced signal. Regarding claim 4, Kaskari does not specifically disclose wherein the speech probability estimate further distinguishes the speech from the stationary noise and the non-stationary noise. Thyssen teaches wherein the speech probability estimate further distinguishes the speech from the stationary noise and the non-stationary noise (para. 0121 “For example, as shown in FIG. 3E, plot 347 represents a time domain input waveform representing first signal 340 (which includes both speech and car noise), plot 349 represents a time-frequency plot of first signal 340”; para. 0204 “As shown in FIG. 5, the method of flowchart 500 begins at step 502, where one or more first characteristics associated with a first type of interfering source in an audio signal are determined. In accordance with an embodiment, the first type of interfering source is stationary noise. In accordance with such an embodiment, the first characteristic(s) include an SNR regarding the stationary noise with respect to the audio signal and a first measure of probability indicative of a probability that the audio signal is from a desired source with respect to the stationary noise.”; para. 0206 “At step 504, one or more second characteristics associated with a second type of interfering source in an audio signal are determined. In accordance with an embodiment, the second type of interfering source is non-stationary noise. In accordance with such an embodiment, the second characteristic(s) include an SNR regarding the non-stationary noise with respect to the audio signal and a second measure of probability indicative of a probability that the audio signal is from a desired source with respect to the non-stationary noise.”). Kaskari and Thyssen are considered to be analogous to the claimed invention as they both are in the same field of speech enhancement. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Kaskari to incorporate the teachings of Thyssen in order to specifically have the speech probability estimate distinguish the speech from the stationary noise and the non-stationary noise. Doing so would be beneficial, given the same rationale as for claim 3. Regarding claims 32 and 33, claims 32 and 33 are rejected for analogous reasons to claims 3 and 4. Regarding claims 43 and 44, claims 43 and 44 are rejected for analogous reasons to claims 3 and 4. 6. Claims 5-6, 8, 34-35, 37, 45-46, and 48 are rejected under 35 U.S.C. 103 as being unpatentable over Kaskari in view of Sivaraman et al. (US 2022/0084509 A1, hereinafter Sivaraman). Regarding claim 5, Kaskari discloses a machine learning model, wherein the machine learning model is configured to determine the speech probability estimate (para. 0029 “During the inferencing phase, the DNN 122 may determine a probability of speech in each frame of the audio signal 102, at each frequency index associated with the time-frequency domain, based on the classification results. The DNN 122 may further convert the probability of speech determined for each frequency index into a spectral suppression gain that can be used to suppress the noise component of the corresponding frame of the audio signal 102. For example, if there is a low probability of speech in a given frame of the audio signal 102 at a particular frequency index, the DNN 122 may apply a lower gain to reduce the power at that frequency index of the corresponding audio frame. As a result, the DNN 122 may dynamically attenuate the noise component of the audio signal 102 in the time-frequency domain.”). Kaskari does not specifically disclose generating a labelled data set that comprises one or more input features and one or more indications indicative of the speech being present; and providing the labelled data set to [a machine learning model…] Sivaraman teaches generating a labelled data set that comprises one or more input features and one or more indications indicative of the speech being present (para. 0075 “In some embodiments, the analytics server 102 employs supervised training to train the machine-learning models of the machine-learning architecture, where the analytics database 104 and/or the call center database 112 contains labels associated with the training call data or enrollment call data. The labels indicate, for example, the expected data for the training call data or enrollment call data.”; para. 0086 “The server, or certain layers of the machine-learning architecture, may perform one or more data augmentation operations on the input audio signal (e.g., training audio signal, enrollment audio signal). The data augmentation operations generate certain types of degradation for the input audio signal, thereby generating corresponding simulated audio signals from the input audio signal. ...”; para. 0093 “In step 204, the server trains the machine-learning architecture by applying the sub-architectures (e.g., speaker separation engine, noise suppression engine, speaker-embedding engine) on the training signals. The server trains the speech separation engine and noise suppression engine to extract spectro-temporal masks (e.g., speaker mask, noise mask) and generate features of an output signal (e.g., noisy target speaker signal, enhanced speaker signal).”); and providing the labelled data set to a machine learning model (para. 0093 “In step 204, the server trains the machine-learning architecture by applying the sub-architectures (e.g., speaker separation engine, noise suppression engine, speaker-embedding engine) on the training signals. The server trains the speech separation engine and noise suppression engine to extract spectro-temporal masks (e.g., speaker mask, noise mask) and generate features of an output signal (e.g., noisy target speaker signal, enhanced speaker signal).”). Kaskari and Sivaraman are considered to be analogous to the claimed invention as they both are in the same field of speech enhancement. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Kaskari to incorporate the teachings of Sivaraman in order to generate a labelled data set that comprises one or more input features and one or more indications indicative of the speech being present, and to provide the labelled data set to a machine learning model to determine the speech probability estimate. Doing so would be beneficial, as this would force the machine-learning architecture to evaluate and adjust for various types of degradation present in the original training data signals (Sivaraman, para. 0086). Regarding claim 6, Kaskari in view of Sivaraman discloses receiving, by the computing device, a set of speech samples (Sivaraman: para. 0085 “Certain steps of the method 200 include obtaining the input audio signals and/or pre-processing the input audio signals (e.g., training audio signal, enrollment audio signal, inbound audio signal) based upon the particular operational phase (e.g., training phase, enrollment phase, deployment phase).”); applying, based on each of the set of speech samples, a speech weight to the each of the set of speech samples (para. 0031 “As an example of the speech separation engine operations, the input audio signal containing a speech mixture signal x(t) may be represented as: x(t)=s.sub.tar(t)+αs.sub.interf(t)+n(t), where s.sub.tar(t) is the target speaker's signal; s.sub.interf(t) is an interfering speaker's signal; n(t) is the noise; and a is a scaling factor according to the SDR of the given training signal.”; speech weight for speech sample s.sub.tar(t) is 1); receiving, by the computing device, a set of non-speech samples (para. 0091 “The data augmentation operations are not limited to generating speech mixtures. Before or after generating the simulated signals containing the speech mixtures, the server additionally or alternatively performs the data augmentation operations for non-speech background noises.”); applying, based on each of the set of non-speech sample, a non-speech weight to the each of the set of non-speech samples (para. 0091 “For instance, the server may add background noises randomly selected from a large noise corpus to the simulated audio signal comprising the speech mixture. The server may apply these background noises to the simulated audio signal at SNRs ranging from, for example, 5 dB to 30 dB, though such range is not limiting on possible embodiments;”); generating, by combining the speech weighted set of speech samples and the non-speech weighted set of non-speech samples, a speech augmented set (para. 0086 “The data augmentation operations generate certain types of degradation for the input audio signal, thereby generating corresponding simulated audio signals from the input audio signal.”); and extracting one or more input features from the speech augmented set (para. 0093 “In step 204, the server trains the machine-learning architecture by applying the sub-architectures (e.g., speaker separation engine, noise suppression engine, speaker-embedding engine) on the training signals. The server trains the speech separation engine and noise suppression engine to extract spectro-temporal masks (e.g., speaker mask, noise mask) and generate features of an output signal (e.g., noisy target speaker signal, enhanced speaker signal).”). Kaskari and Sivaraman are considered to be analogous to the claimed invention as they both are in the same field of speech enhancement. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Kaskari to incorporate the teachings of Sivaraman in order to receive a set of speech and non-speech samples, to apply a speech and non-speech weight respectively to the samples, to combine the speech weighted set of speech samples and the non-speech weighted set of non-speech samples to generate a speech augmented set, and to extract one or more input features from the speech augmented set. Doing so would be beneficial, as this would force the machine-learning architecture to evaluate and adjust for various types of degradation present in the original training data signals (Sivaraman, para. 0086). Regarding claim 8, Kaskari in view of Sivaraman discloses receiving, by the computing device, a set of speech samples (Sivaraman: para. 0085 “Certain steps of the method 200 include obtaining the input audio signals and/or pre-processing the input audio signals (e.g., training audio signal, enrollment audio signal, inbound audio signal) based upon the particular operational phase (e.g., training phase, enrollment phase, deployment phase).”); applying, based on each of the set of speech samples, a speech weight to the each of the set of speech samples (para. 0031 “As an example of the speech separation engine operations, the input audio signal containing a speech mixture signal x(t) may be represented as: x(t)=s.sub.tar(t)+αs.sub.interf(t)+n(t), where s.sub.tar(t) is the target speaker's signal; s.sub.interf(t) is an interfering speaker's signal; n(t) is the noise; and a is a scaling factor according to the SDR of the given training signal.”; speech weight for speech sample s.sub.tar(t) is 1); receiving, by the computing device, a set of non-speech samples (para. 0091 “The data augmentation operations are not limited to generating speech mixtures. Before or after generating the simulated signals containing the speech mixtures, the server additionally or alternatively performs the data augmentation operations for non-speech background noises.”); applying, based on each of the set of non-speech sample, a non-speech weight to the each of the set of non-speech samples (para. 0091 “For instance, the server may add background noises randomly selected from a large noise corpus to the simulated audio signal comprising the speech mixture. The server may apply these background noises to the simulated audio signal at SNRs ranging from, for example, 5 dB to 30 dB, though such range is not limiting on possible embodiments;”); determining, based on the speech weighted set and the non-speech weighted set of non-speech sample, the one or more indications indicative of the speech present (para. 0093 “In step 204, the server trains the machine-learning architecture by applying the sub-architectures (e.g., speaker separation engine, noise suppression engine, speaker-embedding engine) on the training signals. The server trains the speech separation engine and noise suppression engine to extract spectro-temporal masks (e.g., speaker mask, noise mask) and generate features of an output signal (e.g., noisy target speaker signal, enhanced speaker signal).”). Kaskari and Sivaraman are considered to be analogous to the claimed invention as they both are in the same field of speech enhancement. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Kaskari to incorporate the teachings of Sivaraman in order to receive a set of speech and non-speech samples, to apply a speech and non-speech weight respectively to the samples, and to determine based on the speech weighted set of speech samples and the non-speech weighted set of non-speech samples the one or more indications indicative of the speech present. Doing so would be beneficial, as this would force the machine-learning architecture to evaluate and adjust for various types of degradation present in the original training data signals (Sivaraman, para. 0086). Regarding claims 34, 35, and 37, these claims are rejected for analogous reasons to claims 5, 6, and 8, respectively. Regarding claims 45, 46, and 48, these claims are rejected for analogous reasons to claims 5, 6, and 8, respectively. 7. Claims 7, 36, and 47 are rejected under 35 U.S.C. 103 as being unpatentable over Kaskari in view of Sivaraman and in further view of Ahoei et al. (US 11,977,816 B1, hereinafter Ahoei). Regarding claim 7, Kaskari in view of Sivaraman discloses wherein the one or more extracted input features comprise Mel Frequency Cepstrum Coefficients (MFCCs)…(Sivaraman, para. 0093 “In step 204, the server trains the machine-learning architecture by applying the sub-architectures (e.g., speaker separation engine, noise suppression engine, speaker-embedding engine) on the training signals.”; para. 0029 “The speech separation engine receives an input audio signal containing a mixture of speaker signals and one or more types of noise (e.g., additive noise, reverberation). The speech separation engine extracts low-level spectral features, such as such as mel-frequency cepstrum coefficients (MFCCs), and receives a voiceprint for a target speaker (sometimes called an “inbound voiceprint” or “target voiceprint”) generated by the speaker-embedding engine.”). Kaskari and Sivaraman are considered to be analogous to the claimed invention as they both are in the same field of speech enhancement. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Kaskari to incorporate the teachings of Sivaraman in order to have the one or more input features include MFCCs. Doing so would be beneficial, as MFCCs are a set of features which are frequently used for voice recognition (NPL Deruty, Intuitive understanding of MFCCs, pg. 1, 1st para.) which would be indicative of speech being present in an audio signal. Kaskari in view of Sivaraman does not specifically disclose [wherein the one or more extracted input features comprise…] Phonemes, Senones, and Mel Spectrogram. Ahoei teaches wherein the one or more extracted input features comprise Phonemes (Col. 28 Lines 1-7 “The preprocessing component 720 may transform the text data 715 into, for example, a symbolic linguistic representation, which may include linguistic context features such as phoneme data, punctuation data, syllable-level features, word-level features, and/or emotion, speaker, accent, or other features for processing by the TTS system 680.”), Senones (Col. 32 Lines 55-60 “The speech recognition engine 858 may use the acoustic model(s) 853 to attempt to match received audio feature vectors to words or subword acoustic units. An acoustic unit may be a senone, phoneme, phoneme in context, syllable, part of a syllable, syllable in context, or any other such portion of a word.”), and Mel Spectrogram (Col. 28 Lines 62-65 “This symbolic linguistic representation may be sent to the TTS model 780 for conversion into audio data (e.g., in the form of Mel-spectrograms or other frequency content data format).”). Kaskari, Sivaraman, and Ahoei are considered to be analogous to the claimed invention as they are all in the same field of speech processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Kaskari in view of Sivaraman to incorporate the teachings of Ahoei in order to have the one or more input features include phonemes, senones, and mel spectrogram. Utilizing phonemes would be beneficial as phonetic information can be used to guide speech the speech enhancement process to achieve better denoising performance (NPL Lu et al., Incorporating Broad Phonetic Information for Speech Enhancement, pg. 4, Conclusion). Furthermore, utilizing senones would be beneficial as senones carry higher-level information relating to human perception which aid in speech enhancement tasks (NPL Wang et al., A Cross-Task Transfer Learning Approach to Adapting Deep Speech Enhancement Models to Unseen Background Noise Using Paired Senone Classifiers, pg. 3 section 4.2 1st para., and pg. 4 Conclusion). Furthermore, utilizing mel spectrograms would be beneficial as they provide a concise snapshot of an audio signal while better reflecting how humans perceive amplitude and frequency compared to a normal spectrogram (NPL Doshi, Audio Deep Learning Made Simple (Part 2): Why Mel Spectrograms perform better, pg. 4 section “Spectrograms”; pg. 6 section “Mel Spectrograms”). Regarding claims 36 and 47, both claims are rejected for analogous reasons to claim 7. 8. Claims 11, 40, and 51 are rejected under 35 U.S.C. 103 as being unpatentable over Kaskari in view of Vilkamo et al. (US 2023/0402050 A1, hereinafter Vilkamo). Regarding claim 11, Kaskari does not specifically disclose wherein the input signal comprises one or more pulse code modulation (PCM) signals. Vilkamo teaches wherein the input signal comprises one or more pulse code modulation (PCM) signals (para. 0081 “The audio signals 205 can be provided to the processor 103 in any suitable format. In some examples the audio signals 205 can be provided to the processor 103 in a digital format. The digital format could comprise pulse code modulation (PCM) or any other suitable type of format.”). Kaskari and Vilkamo are considered to be analogous to the claimed invention as they both are in the same field of speech enhancement. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Kaskari to incorporate the teachings of Vilkamo in order to have the input signal comprise one or more pulse code modulation (PCM) signals. Doing so would be beneficial, as pulse-code modulation is a noise-resistant method for transmitting audio signals (NPL Plonus, Electronics and Communications for Scientists and Engineers, Chapter 9, pg. 370, section “Pulse Code Modulation (PCM)”). Regarding claims 40 and 51, both claims are rejected for analogous reasons to claim 11. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Song (US 2021/0233557 A1): determine probabilities of speech and noise (e.g. wind noise), obtain gain to apply to audio signal to perform wind noise reduction (Fig. 1) Any inquiry concerning this communication or earlier communications from the examiner should be directed to CODY DOUGLAS HUTCHESON whose telephone number is (703)756-1601. The examiner can normally be reached M-F 8:00AM-5:00PM EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre-Louis Desir can be reached at (571)-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /CODY DOUGLAS HUTCHESON/ Examiner, Art Unit 2659 /BHAVESH M MEHTA/ Supervisory Patent Examiner, Art Unit 2656
Read full office action

Prosecution Timeline

Show 1 earlier event
Apr 22, 2024
Response after Non-Final Action
Dec 10, 2025
Non-Final Rejection mailed — §101, §102, §103
Mar 10, 2026
Response Filed
Apr 23, 2026
Final Rejection mailed — §101, §102, §103
Jun 12, 2026
Response after Non-Final Action
Jul 23, 2026
Request for Continued Examination
Jul 27, 2026
Response after Non-Final Action
Aug 04, 2026
Non-Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12664970
SPEECH TRANSLATION WITH PERFORMANCE CHARACTERISTICS
3y 2m to grant Granted Jun 23, 2026
Patent 12626715
ROLE SEPARATION METHOD, ELECTRONIC DEVICE, AND COMPUTER STORAGE MEDIUM
3y 4m to grant Granted May 12, 2026
Patent 12614036
INTELLIGENT DETECTION OF BIAS WITHIN AN ARTIFICIAL INTELLIGENCE MODEL
2y 3m to grant Granted Apr 28, 2026
Patent 12603096
VOICE ENHANCEMENT METHODS AND SYSTEMS
2y 10m to grant Granted Apr 14, 2026
Patent 12591750
GENERATIVE LANGUAGE MODEL UNLEARNING
2y 3m to grant Granted Mar 31, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
63%
Grant Probability
99%
With Interview (+42.0%)
2y 8m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 30 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month