Prosecution Insights
Last updated: October 02, 2026
Application No. 18/857,711

PCEN-BASED MASK THRESHOLDING AND VOICE ACTIVITY DETECTION FOR TRAINING DNN-BASED SPEECH ENHANCEMENT MODELS

Final Rejection §102§103
Filed
Oct 17, 2024
Priority
Apr 20, 2022 — CN PCT/CN2022/087983 +3 more
Examiner
MONIKANG, GEORGE C
Art Unit
2653
Tech Center
2600 — Communications
Assignee
Dolby Laboratories Licensing Corporation
OA Round
2 (Final)
75%
Grant Probability
Favorable
3-4
OA Rounds
1y 1m
Est. Remaining
82%
With Interview

Examiner Intelligence

Grants 75% — above average
75%
Career Allowance Rate
739 granted / 981 resolved
+13.3% vs TC avg
Moderate +7% lift
Without
With
+7.1%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
29 currently pending
Career history
1004
Total Applications
across all art units

Statute-Specific Performance

§101
4.2%
-35.8% vs TC avg
§103
65.0%
+25.0% vs TC avg
§102
21.4%
-18.6% vs TC avg
§112
3.6%
-36.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 981 resolved cases

Office Action

§102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant's arguments filed 7/15/2026 have been fully considered but they are not persuasive. With regards to applicant’s argument that the amendment that calls for wherein the at least one mask is a ground truth mask target for a training sample used to train the DNN, the examiner disagrees. The Olvera et al reference teaches where a model is trained by minimizing the norm between the Mel spectrogram of the ground truth foreground event and the Mel spectrogram of the mixture multiplied element-wise by the estimated mask (pg. 3, right column, Section E: estimated mask is such that it minimizes distances to the ground truth). Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1, 3, 5-8 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Olvera et al, “Foreground-Background Ambient Sound Scene Separation”. (The Olvera et al reference is cited in IDS filed 10/17/2024) Re Claim 1, Olvera et al discloses a method of determining at least one mask for use in training a deep neural network, DNN, -based mask-based audio processing model (pg. 1, right column, 3rd paragraph: Deep Neural Network), the method comprising: obtaining a time-frequency representation of a target audio signal for use in the training (pg. 1, right column, 2nd paragraph, lines 1-6: time-frequency & 3rd paragraph: time-frequency representation; pg. 2, left column, section II); determining a per-channel energy normalization, PCEN, measure for the target audio signal (pg. 1, right column, 4th paragraph: PCEN; pg. 3, left-right column, section B); and determining the at least one mask based on the PCEN measure (pg. 1, right column, 3rd paragraph: time-frequency mask; pg. 2, right column, lines 4-8: time-frequency mask; fig. 1; pg. 4, left column, lines 1-4), wherein the at least one mask is a ground truth mask target for a training sample used to train the DNN (pg. 3, right column, Section E: estimated mask is such that it minimizes distances to the ground truth). Re Claim 3, Olvera et al discloses the method according to claim 1, wherein the target audio signal is for use as a ground truth for the training (pg. 3, right column, section E: ground truth foreground event). Re Claim 5, Olvera et al discloses the method according to claim 1, wherein the PCEN measure is determined based on a ratio of a time-frequency energy measure of the target audio signal and a running average of the time-frequency energy measure of the target audio signal (pg. 3, left - right column, section D: average). Re Claim 6, Olvera et al discloses the method according to claim 1, wherein the method further comprises: obtaining an audio mixture based on the target audio signal, wherein the audio mixture comprises, in addition to the target audio signal, audio artifacts (pg. 3, right column, section F – pg. 4, left column, section V: computation of IRM to obtain upper bound for the performance of the foreground-background separation task); and wherein the determination of the at least one mask based on the PCEN measure involves: determining at least one ideal ratio mask, IRM, measure based on the target audio signal and the audio mixture (pg. 3, right column, section F – pg. 4, left column, section V: computation of IRM to obtain upper bound for the performance of the foreground-background separation task); and adjusting the at least one IRM measure to obtain the at least one mask based on the PCEN measure (pg. 3, right column, section F – pg. 4, left column, section V: computation of IRM to obtain upper bound for the performance of the foreground-background separation task). Re Claim 7, Olvera et al discloses the method according to claim 6, wherein the audio artifacts comprises at least one of: noise, echo, or reverberation (pg. 4, left column, section V: signal to distortion/noise; wherein noise/distortion is selected from the Markush language). Re Claim 8, Olvera et al discloses the method according to claim 6, wherein the IRM measure is determined as a ratio of a time-frequency energy measure of the target audio signal to a time-frequency energy measure of the audio mixture including the audio artifacts (pg. 4, left column, lines 4-7: ideal ratio mask). Claims 15, 23 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Zhang et al, CN 112530451 A. Re Claim 15, Zhang et al discloses a method of determining a speech-aware loss function for use in training a deep neural network, DNN, -based audio processing model (claim 1: DNN), the method comprising: obtaining a time-frequency representation of a target audio signal for use in the training (claim 1: time & frequency); determining presence of speech in the target audio signal by using a voice activity detection, VAD, process (claim 1: voice signal is picked up by a detector); and determining the loss function, by controlling gradient of the loss function based on the determined presence of speech in the target audio signal (claim 1: speech enhancement system that minimizes loss from gradient descent method). Re Claim 23, Zhang et al discloses the method according to claim 15, wherein the determination of the loss function involves: determining respective loss function for speech and non-speech frames; and averaging the loss function for all time-frequency bands, thereby obtaining a final loss function (claim 1: iteration average result). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 2, 4, 10, 25 are rejected under 35 U.S.C. 103 as being unpatentable over Olvera et al, “Foreground-Background Ambient Sound Scene Separation”. Re Claim 2, Olvera et al discloses the method according to claim 1, but fails to explicitly disclose wherein the target audio signal comprises speech and/or music. Since Olvera et al discloses a DNN that is able to separate sources belonging to different classes such as rapidly varying Spectro-temporal features of short audio events (larger frame wise than a specified threshold) against more slowly varying features (lower fame wise that a specified threshold) of background sounds encountered in real-life environments, it would have been on obvious for one of ordinary skill in the art to modify Olvera et al where the rapidly varying spectro-temproal features matches music, speech since both typically comprise rapidly varying spectro-temporal features and the slowly varying features matches stationary noise components such as room noise or recording noise for the purpose of separating speech/music signals from stationary noises. Re Claim 4, Olvera et al discloses the method according to claim 1, but fail to explicitly disclose wherein the target audio signal comprises a clean audio component, and a stationary noise component such as recording noise and/or room noise. Since Olvera et al discloses a DNN that is able to separate sources belonging to different classes such as rapidly varying Spectro-temporal features of short audio events (larger frame wise than a specified threshold) against more slowly varying features (lower fame wise that a specified threshold) of background sounds encountered in real-life environments, it would have been on obvious for one of ordinary skill in the art to modify Olvera et al where the rapidly varying spectro-temproal features matches music, speech since both typically comprise rapidly varying spectro-temporal features and the slowly varying features matches stationary noise components such as room noise or recording noise for the purpose of separating speech/music signals from stationary noises. Re Claim 10, Olvera et al discloses the method according to claim 1, but fails to disclose wherein the method further comprises: classifying the target audio signal into speech frames or non-speech frames; and for each audio frame, determining the at least one mask further based on a classification of speech frames or non-speech frames. Since Olvera et al discloses a DNN that is able to separate sources belonging to different classes such as rapidly varying Spectro-temporal features of short audio events (larger frame wise than a specified threshold) against more slowly varying features (lower fame wise that a specified threshold) of background sounds encountered in real-life environments, it would have been on obvious for one of ordinary skill in the art to modify Olvera et al where the rapidly varying spectro-temproal features matches music, speech since both typically comprise rapidly varying spectro-temporal features and the slowly varying features matches stationary noise components such as room noise or recording noise for the purpose of separating speech/music signals from stationary noises. Claim 25 has been analyzed and rejected according to claims 1-2. Allowable Subject Matter Claims 9, 11, 13-14, 16-21 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter for claim 9: The prior art does not teach or moderately suggest the following limitations: Wherein respective values of the PCEN and the IRM measures are determined for each time-frequency band of the target audio signal; and wherein the adjustment of the IRM measure involves: for each time-frequency band, setting the respective IRM measure value for that time-frequency band to 0, if the respective PCEN measure value is less than a first predetermined threshold. Limitations such as these may be useful in combination with other limitations of claim 6 and ultimately claim 1. The following is a statement of reasons for the indication of allowable subject matter for claim 11: The prior art does not teach or moderately suggest the following limitations: Wherein the classification of the target audio signal into speech frames or non-speech frames involves, for each audio frame of the target audio signal: determining a respective frame-wise PCEN measure across all frequency bands of that audio frame; and if the determined frame-wise PCEN measure is larger than a second predetermined threshold, classifying that audio frame into a speech frame; otherwise, classifying that audio frame into a non-speech frame. Limitations such as these may be useful in combination with other limitations of claim 10 and ultimately claim 1. The following is a statement of reasons for the indication of allowable subject matter for claim 13: The prior art does not teach or moderately suggest the following limitations: Wherein the PCEN measure is determined according to, for each time-frequency band: P C E N t , f = ( s t , f ε + m ( t , f ∝   +   δ ) r   -   δ r , wherein S(t, f) is a time-frequency energy of the target audio signal, ε, α, δ, and r are predetermined constants, and M(t, f) is a running average of S(t, f) defined as: M⁡(t,f)=(1-s).Math.M⁡(t-1,f)+s.Math.S⁡(t,f), in which s is a predetermined smoothing factor. Limitations such as these may be useful in combination with other limitations of claim 1. The following is a statement of reasons for the indication of allowable subject matter for claim 14: The prior art does not teach or moderately suggest the following limitations: Wherein the DNN-based mask-based audio processing model is a multitask model configured for being trained for predicting a plurality of masks each corresponding to a respective audio processing aspect, and wherein the method involves: obtaining a respective target audio signal for each audio processing aspect; determining a respective PCEN measure for each target audio signal; and determining a respective mask based on the PCEN measure. Limitations such as these may be useful in combination with other limitations of claim 1. The following is a statement of reasons for the indication of allowable subject matter for claims 16-18: The prior art does not teach or moderately suggest the following limitations: Wherein the VAD process involves: determining frame-wise and/or band-wise presence of speech in the target audio signal; and wherein the controlling of the gradient of the loss function based on the determined presence of speech in the target audio signal involves: increasing the respective gradient of the loss function for a non-speech frame and/or band such that audio artifacts are suppressed more aggressively in the non-speech frame and/or band. Limitations such as these may be useful in combination with other limitations of claim 15. The following is a statement of reasons for the indication of allowable subject matter for claim 19: The prior art does not teach or moderately suggest the following limitations: Wherein the loss function is determined such that over-suppression of speech is penalized more than under-suppression of noise in the target audio signal. Limitations such as these may be useful in combination with other limitations of claim 15. The following is a statement of reasons for the indication of allowable subject matter for claims 20-21: The prior art does not teach or moderately suggest the following limitations: Wherein the loss function loss is defined as, for each time-frequency band: l o s s =   α d ⅈ f f - d i f f - 1 , where α is a predetermined constant, and diff indicates a difference between an ideal ratio mask, IRM, measure determined for the target audio signal and an estimated mask mask.sub.est predicted by the DNN-based audio processing model. Limitations such as these may be useful in combination with other limitations of claim 15. Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to GEORGE C MONIKANG whose telephone number is (571)270-1190. The examiner can normally be reached Mon. - Fri., 9AM-5PM, ALT. Fridays off. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Carolyn R Edwards can be reached at 571-270-7136. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /GEORGE C MONIKANG/Primary Examiner, Art Unit 2692 08/21/2026
Read full office action

Prosecution Timeline

Oct 17, 2024
Application Filed
Apr 29, 2026
Non-Final Rejection mailed — §102, §103
Jun 25, 2026
Interview Requested
Jul 15, 2026
Response Filed
Aug 26, 2026
Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12745031
SYSTEM FOR CHARGING AN EAR-WORN ELECTRONIC DEVICE
2y 1m to grant Granted Sep 22, 2026
Patent 12739579
HEARING INSTRUMENT AND ASSOCIATED BINAURAL HEARING SYSTEM
2y 4m to grant Granted Sep 15, 2026
Patent 12732147
ACOUSTIC PROCESSING DEVICE AND ACOUSTIC PROCESSING METHOD
3y 0m to grant Granted Sep 08, 2026
Patent 12725318
MULTILINGUAL TEXT-TO-IMAGE GENERATION
3y 5m to grant Granted Sep 01, 2026
Patent 12724969
ELECTRONIC COMMUNICATIONS SIGNATURE RECOGNITION FOR PRIVACY PRESERVING COMPUTER OPERATIONS
2y 2m to grant Granted Sep 01, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
75%
Grant Probability
82%
With Interview (+7.1%)
3y 0m (~1y 1m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 981 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month