DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant's arguments filed 7/15/2026 have been fully considered but they are not persuasive.
With regards to applicant’s argument that the amendment that calls for wherein the at least one mask is a ground truth mask target for a training sample used to train the DNN, the examiner disagrees. The Olvera et al reference teaches where a model is trained by minimizing the norm between the Mel spectrogram of the ground truth foreground event and the Mel spectrogram of the mixture multiplied element-wise by the estimated mask (pg. 3, right column, Section E: estimated mask is such that it minimizes distances to the ground truth).
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1, 3, 5-8 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Olvera et al, “Foreground-Background Ambient Sound Scene Separation”. (The Olvera et al reference is cited in IDS filed 10/17/2024)
Re Claim 1, Olvera et al discloses a method of determining at least one mask for use in training a deep neural network, DNN, -based mask-based audio processing model (pg. 1, right column, 3rd paragraph: Deep Neural Network), the method comprising: obtaining a time-frequency representation of a target audio signal for use in the training (pg. 1, right column, 2nd paragraph, lines 1-6: time-frequency & 3rd paragraph: time-frequency representation; pg. 2, left column, section II); determining a per-channel energy normalization, PCEN, measure for the target audio signal (pg. 1, right column, 4th paragraph: PCEN; pg. 3, left-right column, section B); and determining the at least one mask based on the PCEN measure (pg. 1, right column, 3rd paragraph: time-frequency mask; pg. 2, right column, lines 4-8: time-frequency mask; fig. 1; pg. 4, left column, lines 1-4), wherein the at least one mask is a ground truth mask target for a training sample used to train the DNN (pg. 3, right column, Section E: estimated mask is such that it minimizes distances to the ground truth).
Re Claim 3, Olvera et al discloses the method according to claim 1, wherein the target audio signal is for use as a ground truth for the training (pg. 3, right column, section E: ground truth foreground event).
Re Claim 5, Olvera et al discloses the method according to claim 1, wherein the PCEN measure is determined based on a ratio of a time-frequency energy measure of the target audio signal and a running average of the time-frequency energy measure of the target audio signal (pg. 3, left - right column, section D: average).
Re Claim 6, Olvera et al discloses the method according to claim 1, wherein the method further comprises: obtaining an audio mixture based on the target audio signal, wherein the audio mixture comprises, in addition to the target audio signal, audio artifacts (pg. 3, right column, section F – pg. 4, left column, section V: computation of IRM to obtain upper bound for the performance of the foreground-background separation task); and wherein the determination of the at least one mask based on the PCEN measure involves: determining at least one ideal ratio mask, IRM, measure based on the target audio signal and the audio mixture (pg. 3, right column, section F – pg. 4, left column, section V: computation of IRM to obtain upper bound for the performance of the foreground-background separation task); and adjusting the at least one IRM measure to obtain the at least one mask based on the PCEN measure (pg. 3, right column, section F – pg. 4, left column, section V: computation of IRM to obtain upper bound for the performance of the foreground-background separation task).
Re Claim 7, Olvera et al discloses the method according to claim 6, wherein the audio artifacts comprises at least one of: noise, echo, or reverberation (pg. 4, left column, section V: signal to distortion/noise; wherein noise/distortion is selected from the Markush language).
Re Claim 8, Olvera et al discloses the method according to claim 6, wherein the IRM measure is determined as a ratio of a time-frequency energy measure of the target audio signal to a time-frequency energy measure of the audio mixture including the audio artifacts (pg. 4, left column, lines 4-7: ideal ratio mask).
Claims 15, 23 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Zhang et al, CN 112530451 A.
Re Claim 15, Zhang et al discloses a method of determining a speech-aware loss function for use in training a deep neural network, DNN, -based audio processing model (claim 1: DNN), the method comprising: obtaining a time-frequency representation of a target audio signal for use in the training (claim 1: time & frequency); determining presence of speech in the target audio signal by using a voice activity detection, VAD, process (claim 1: voice signal is picked up by a detector); and determining the loss function, by controlling gradient of the loss function based on the determined presence of speech in the target audio signal (claim 1: speech enhancement system that minimizes loss from gradient descent method).
Re Claim 23, Zhang et al discloses the method according to claim 15, wherein the determination of the loss function involves: determining respective loss function for speech and non-speech frames; and averaging the loss function for all time-frequency bands, thereby obtaining a final loss function (claim 1: iteration average result).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 2, 4, 10, 25 are rejected under 35 U.S.C. 103 as being unpatentable over Olvera et al, “Foreground-Background Ambient Sound Scene Separation”.
Re Claim 2, Olvera et al discloses the method according to claim 1, but fails to explicitly disclose wherein the target audio signal comprises speech and/or music. Since Olvera et al discloses a DNN that is able to separate sources belonging to different classes such as rapidly varying Spectro-temporal features of short audio events (larger frame wise than a specified threshold) against more slowly varying features (lower fame wise that a specified threshold) of background sounds encountered in real-life environments, it would have been on obvious for one of ordinary skill in the art to modify Olvera et al where the rapidly varying spectro-temproal features matches music, speech since both typically comprise rapidly varying spectro-temporal features and the slowly varying features matches stationary noise components such as room noise or recording noise for the purpose of separating speech/music signals from stationary noises.
Re Claim 4, Olvera et al discloses the method according to claim 1, but fail to explicitly disclose wherein the target audio signal comprises a clean audio component, and a stationary noise component such as recording noise and/or room noise. Since Olvera et al discloses a DNN that is able to separate sources belonging to different classes such as rapidly varying Spectro-temporal features of short audio events (larger frame wise than a specified threshold) against more slowly varying features (lower fame wise that a specified threshold) of background sounds encountered in real-life environments, it would have been on obvious for one of ordinary skill in the art to modify Olvera et al where the rapidly varying spectro-temproal features matches music, speech since both typically comprise rapidly varying spectro-temporal features and the slowly varying features matches stationary noise components such as room noise or recording noise for the purpose of separating speech/music signals from stationary noises.
Re Claim 10, Olvera et al discloses the method according to claim 1, but fails to disclose wherein the method further comprises: classifying the target audio signal into speech frames or non-speech frames; and for each audio frame, determining the at least one mask further based on a classification of speech frames or non-speech frames. Since Olvera et al discloses a DNN that is able to separate sources belonging to different classes such as rapidly varying Spectro-temporal features of short audio events (larger frame wise than a specified threshold) against more slowly varying features (lower fame wise that a specified threshold) of background sounds encountered in real-life environments, it would have been on obvious for one of ordinary skill in the art to modify Olvera et al where the rapidly varying spectro-temproal features matches music, speech since both typically comprise rapidly varying spectro-temporal features and the slowly varying features matches stationary noise components such as room noise or recording noise for the purpose of separating speech/music signals from stationary noises.
Claim 25 has been analyzed and rejected according to claims 1-2.
Allowable Subject Matter
Claims 9, 11, 13-14, 16-21 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
The following is a statement of reasons for the indication of allowable subject matter for claim 9: The prior art does not teach or moderately suggest the following limitations:
Wherein respective values of the PCEN and the IRM measures are determined for each time-frequency band of the target audio signal; and wherein the adjustment of the IRM measure involves: for each time-frequency band, setting the respective IRM measure value for that time-frequency band to 0, if the respective PCEN measure value is less than a first predetermined threshold.
Limitations such as these may be useful in combination with other limitations of claim 6 and ultimately claim 1.
The following is a statement of reasons for the indication of allowable subject matter for claim 11: The prior art does not teach or moderately suggest the following limitations:
Wherein the classification of the target audio signal into speech frames or non-speech frames involves, for each audio frame of the target audio signal: determining a respective frame-wise PCEN measure across all frequency bands of that audio frame; and if the determined frame-wise PCEN measure is larger than a second predetermined threshold, classifying that audio frame into a speech frame; otherwise, classifying that audio frame into a non-speech frame.
Limitations such as these may be useful in combination with other limitations of claim 10 and ultimately claim 1.
The following is a statement of reasons for the indication of allowable subject matter for claim 13: The prior art does not teach or moderately suggest the following limitations:
Wherein the PCEN measure is determined according to, for each time-frequency band:
P
C
E
N
t
,
f
=
(
s
t
,
f
ε
+
m
(
t
,
f
∝
+
δ
)
r
-
δ
r
, wherein S(t, f) is a time-frequency energy of the target audio signal, ε, α, δ, and r are predetermined constants, and M(t, f) is a running average of S(t, f) defined as: M(t,f)=(1-s).Math.M(t-1,f)+s.Math.S(t,f), in which s is a predetermined smoothing factor.
Limitations such as these may be useful in combination with other limitations of claim 1.
The following is a statement of reasons for the indication of allowable subject matter for claim 14: The prior art does not teach or moderately suggest the following limitations:
Wherein the DNN-based mask-based audio processing model is a multitask model configured for being trained for predicting a plurality of masks each corresponding to a respective audio processing aspect, and wherein the method involves: obtaining a respective target audio signal for each audio processing aspect; determining a respective PCEN measure for each target audio signal; and determining a respective mask based on the PCEN measure.
Limitations such as these may be useful in combination with other limitations of claim 1.
The following is a statement of reasons for the indication of allowable subject matter for claims 16-18: The prior art does not teach or moderately suggest the following limitations:
Wherein the VAD process involves: determining frame-wise and/or band-wise presence of speech in the target audio signal; and wherein the controlling of the gradient of the loss function based on the determined presence of speech in the target audio signal involves: increasing the respective gradient of the loss function for a non-speech frame and/or band such that audio artifacts are suppressed more aggressively in the non-speech frame and/or band.
Limitations such as these may be useful in combination with other limitations of claim 15.
The following is a statement of reasons for the indication of allowable subject matter for claim 19: The prior art does not teach or moderately suggest the following limitations:
Wherein the loss function is determined such that over-suppression of speech is penalized more than under-suppression of noise in the target audio signal.
Limitations such as these may be useful in combination with other limitations of claim 15.
The following is a statement of reasons for the indication of allowable subject matter for claims 20-21: The prior art does not teach or moderately suggest the following limitations:
Wherein the loss function loss is defined as, for each time-frequency band:
l
o
s
s
=
α
d
ⅈ
f
f
-
d
i
f
f
-
1
, where α is a predetermined constant, and diff indicates a difference between an ideal ratio mask, IRM, measure determined for the target audio signal and an estimated mask mask.sub.est predicted by the DNN-based audio processing model.
Limitations such as these may be useful in combination with other limitations of claim 15.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to GEORGE C MONIKANG whose telephone number is (571)270-1190. The examiner can normally be reached Mon. - Fri., 9AM-5PM, ALT. Fridays off.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Carolyn R Edwards can be reached at 571-270-7136. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/GEORGE C MONIKANG/Primary Examiner, Art Unit 2692 08/21/2026