Last updated: April 19, 2026

Application No. 18/658,964

SPEECH PROCESSING

Non-Final OA §103

Filed

May 08, 2024

Examiner

WOO, STELLA L

Art Unit

2693

Tech Center

2600 — Communications

Assignee

Tencent Technology (Shenzhen) Company Limited

OA Round

1 (Non-Final)

Interview Optional

— +13.2% interview lift. This examiner has a relatively high allow rate; a written response may suffice.

Based on 1007 resolved cases, 2023–2026

Examiner Intelligence

WOO, STELLA L View full profile →

Grants 80% — above average

Career Allow Rate

801 granted / 1007 resolved

+17.5% vs TC avg

Moderate +13% lift

Without

With

+13.2%

Interview Lift

resolved cases with interview

Typical timeline

2y 9m

Avg Prosecution

21 currently pending

Career history

1028

Total Applications

across all art units

Statute-Specific Performance

§101

3.3%

-36.7% vs TC avg

§103

42.4%

+2.4% vs TC avg

§102

27.9%

-12.1% vs TC avg

§112

11.4%

-28.6% vs TC avg

Black line = Tech Center average estimate • Based on career data from 1007 resolved cases

Office Action

§103

DETAILED ACTION
Notice of Pre-AIA  or AIA  Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Allowable Subject Matter
Claims 4-10, 12-13, 16, 19 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA  35 U.S.C. 102 and 103 (or as subject to pre-AIA  35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA  to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.  
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.

Claim(s) 1, 14, 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Borgstrom et al. (US 2021/0074282 A1, “Borgstrom”) in view of Le Roux et al. (US 2019/0318754 A1, “Le Roux”).
As to claims 1, 14, 17, Borgstrom discloses an audio processing method comprising: 
obtaining an initial audio feature of initial audio data (initial spectrum 110 is received as an input; para. 0040); 
inputting the initial audio feature to an audio enhancement model, the audio enhancement model being iteratively trained based on a deep clustering loss function and a mask inference loss function (initial spectrum 110 is input to a deep neural network 120, noise estimator 130, SNR estimator 140 and gain mask estimator 150; para. 0044-0045; Fig. 1B); 
calculating, by processing circuitry, target audio data with reduced noise and reverberation according to a target audio feature, the target audio feature being generated by the audio enhancement model based on the initial audio feature (calculating enhanced speech with jointly suppressed noise and reverberation; para. 0007, 0045); and 
outputting the target audio data (output processor 160 outputs enhanced spectrum speech 199; para. 0044-0045; Fig. 1B).
Borgstrom differs from claim 1 in that it does not disclose the above underlined limitations.  Le Roux teaches transforming an input audio signal using a combination of deep clustering loss function and mask inference loss function (para. 0084-0087).  It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Borgstrom with the above teaching of Le Roux in order to use a known Chimera++ network, as taught by Le Roux, in order to yield significant improvement over individual models, as taught by Le Roux (para. 0084).
Claim(s) 2-3, 15, 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Bergstrom in view of Le Roux, as applied to claim 1 above, and further in view of Wojcicki et al. (US 2024/0161765 A1, “Wojcicki”).
Bergstrom in view of Le Roux differs from claims 2, 15, 18 in that it does not specifically teach: obtaining a training sample set that includes a noise audio feature, a clean audio label, a noise audio label, and a deep clustering annotation; and performing noise removal training and reverberation removal training on a preset enhancement network based on the training sample set to obtain the audio enhancement model when the preset enhancement network meets a preset condition.
Wojcicki teaches training data for a speech transformation module as including noise from a noise database, clean speech, reverberation, etc. (para. 0034-0035, 0069-0070), clustering techniques (para. 0059-0060), and periodically training the noise removal model when a preset condition is reached, e.g. at a certain time interval, at a scheduled time, after a number of new training data, after a number of new clean speech samples are collected, etc. (para. 0069).  It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Borgstrom in view of Le Roux with the above teaching of Wojcicki in order to provide a personalized noise removal model, as taught by Wojcicki (para. 0012-0016).
As to claim 3, Borgstrom in view of Le Roux and Wojcicki teaches: wherein the preset enhancement network further comprises: a hidden layer, a deep clustering layer, and a mask inference layer, the mask inference layer including an audio mask inference layer and a noise mask inference layer (Le Roux: deep neural network layers, para. 0007, 0010, 0060; mask-inference network 230 estimates a set of masks, including noisy audio and target audio, para. 0059-0060, 0132).
As to claim 11, Borgstrom in view of Le Roux and Wojcicki teaches: wherein the obtaining the training sample set further comprises: 
obtaining a first sample speech with noise and reverberation that is acquired based on a microphone (Wojcicki: machine learning model is trained with speech of the user satisfying a noise threshold and collected during one or more communication sessions, para. 0072); 
performing speech feature extraction on the first sample speech, to obtain a noise speech feature (Wojcicki: noises captured during communication sessions are detected and used for training, para. 0016); 
obtaining a second sample speech including a clean speech without noise and with reverberation and a clean speech without noise and reverberation (Wojcicki: clean speech of a user is received, and combined with reverberation; para. 0069); 
performing speech feature extraction on the second sample speech, to obtain a first clean speech label and a second clean speech label (Wojcicki: speech samples may be collected for training periodically, at a certain time interval, at a scheduled time, after a number of new training data, after a number of new clean speech samples are collected, etc.; para. 0069); and 
determining the deep clustering annotation according to the first sample speech and the second sample speech (Wojcicki: clustering techniques, para. 0059-0060).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Mandel et al. (US 2022/0358904 A1) teach the use of a combination of the deep clustering loss and mask inference loss (para. 0055).
Yang et al. (US 2025/0131941 A1) teach noise and reverberation reduction.
Chhetri et al. (US 12,272,369 B1) teach dereverberation and noise reduction.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Stella L Woo whose telephone number is (571)272-7512. The examiner can normally be reached Monday - Friday, 8 a.m. to 5 p.m.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ahmad Matar can be reached at 571-272-7488. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.

STELLA L. WOO
Primary Examiner
Art Unit 2693



/Stella L. Woo/            Primary Examiner, Art Unit 2693

Read full office action

Prosecution Timeline

May 08, 2024

Application Filed

Jan 05, 2026

Non-Final Rejection — §103

Feb 05, 2026

Examiner Interview Summary

Feb 05, 2026

Applicant Interview (Telephonic)

Precedent Cases

Applications granted by this same examiner with similar technology

18/209,475

Patent 12602416

HYBRID ARTIFICIAL INTELLIGENCE SYSTEM FOR SEMI-AUTOMATIC PATENT CLAIMS ANALYSIS

2y 5m to grant Granted Apr 14, 2026

18/454,212

Patent 12587613

System and method for documenting and controlling meetings with labels and automated operations

2y 5m to grant Granted Mar 24, 2026

18/466,814

Patent 12585681

Methods for Converting Electronic Presentations Into Autonomous Information Collection and Feedback Systems

2y 5m to grant Granted Mar 24, 2026

18/543,126

Patent 12581038

AUDIO PROCESSING IN VIDEO CONFERENCING SYSTEM USING MULTIMODAL FEATURES

2y 5m to grant Granted Mar 17, 2026

19/220,169

Patent 12568170

PRIORITIZING EMERGENCY CALLS BASED ON CALLER RESPONSE TO AUTOMATED QUERY

2y 5m to grant Granted Mar 03, 2026

Study what changed to get past this examiner. Based on 5 most recent grants.

AI Strategy Recommendation

Get an AI-powered prosecution strategy using examiner precedents, rejection analysis, and claim mapping.

Prosecution Projections

1-2

Expected OA Rounds

80%

Grant Probability

93%

With Interview (+13.2%)

2y 9m

Median Time to Grant

Low

PTA Risk

Based on 1007 resolved cases by this examiner. Grant probability derived from career allow rate.