Prosecution Insights
Last updated: September 29, 2026
Application No. 18/923,399

Method and System of Intelligent Dynamic Voice Enhancement

Final Rejection §103
Filed
Oct 22, 2024
Priority
Oct 24, 2023 — CN 202311388318.7
Examiner
GODBOLD, DOUGLAS
Art Unit
Tech Center
Assignee
Harman International Industries Incorporated
OA Round
2 (Final)
83%
Grant Probability
Favorable
3-4
OA Rounds
10m
Est. Remaining
94%
With Interview

Examiner Intelligence

Grants 83% — above average
83%
Career Allowance Rate
925 granted / 1110 resolved
+23.3% vs TC avg
Moderate +11% lift
Without
With
+10.6%
Interview Lift
resolved cases with interview
Typical timeline
2y 9m
Avg Prosecution
18 currently pending
Career history
1129
Total Applications
across all art units

Statute-Specific Performance

§101
15.3%
-24.7% vs TC avg
§103
47.8%
+7.8% vs TC avg
§102
17.8%
-22.2% vs TC avg
§112
9.3%
-30.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1110 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This Office Action is in response to correspondence filed 19 July 2026 in reference to application 18/923,399. Claims 1-18 are pending and have been examined. Response to Amendment The amendment filed 19 July 2026 has been accepted and considered in this office action. Claims 1 and 10 have been amended. Claim Objections Claim 8 is objected to because of the following informalities: It appears that claim 8 has been erroneously deleted from the claims. For purposes of examination, claim 8 will be treated as the claim 8 filed in the previous claim set. Appropriate correction is required. Response to Arguments Applicant's arguments filed 19 July 2026 have been fully considered but they are not persuasive. Applicant argues, see Remarks pages 9-10 that cited prior fails to teach the claims as amended, namely “setting a first speech enhancement gain based on a signal power strength ratio of the center channel to a sum of the other channels; and setting a second speech enhancement gain based on a system volume level; determining a final speech enhancement gain using both the first speech enhancement gain and the second speech enhancement gain.” The examiner respectfully disagrees. Applicant cites paragraph 0044 in spec to support the claims that the instant application operates by calculating separate gains, and then combining them. However the dual processing paths referred to in paragraph 0044 seems to be refereeing to the speech detection and crossover filtering in figure 1. Moreover, the claimed second gain (Gv) based on system volume level is calculated at paragraph 0035 and equation 7 using the claimed first speech enhancement gain based on power strength Ratio (Gp). Thus the combination of references as cited in the previous rejection, and in the rejection claimed supports the limitations as claimed, at least to the extent that they are supported by the specification. Claim Rejections - 35 USC § 103 The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action. Claim(s) 1-3, 6, 7, 10-12, 15, and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Vaudrey et al. (US Patent 6,442,278) in view of Irwan et al. (US PAP 2003/044032) and further in view of Kriegel et al. (US PAP 2024/0405561). Consider claim 1, Vaudrey teaches a method for intelligent dynamic speech enhancement (abstract), comprising the steps of: performing intelligent enhancement gain control on a multi-channel audio source input to determine speech enhancement gain, the multi-channel audio source input comprising a signal of a center channel and signals of other channels (figure 4, col 7 lines 40- col 8 lines 25, adjusting center channel gain of multi-channel signal to enhance speech); applying the speech enhancement gain in dynamic loudness balancing performed on the multi-channel audio source input (figure 4, col 7 lines 40- col 8 lines 25, adjusting center channel gain of multi-channel signal to enhance speech, dynamically at col 7 line 56- col 8 line 11, gain adjustment kicks in when desired ration violated); wherein the intelligent enhancement gain control comprises setting a first speech enhancement gain based on a signal power strength ratio of the center channel to a sum of the other channels (col 7 line 56- col 8 line 11, desired VRA, or voice to remaining audio). Vaudrey does not specifically teach performing speech detection. In the same field of speech enhancement on multichannel audio, Irwan teaches performing speech detection (0013-15 determining voice probabilities and enhancing based on probabilities). It would have been obvious to one of ordinary skill in the art at the time of effective filing to perform speech detection as taught by Irwan in the system of Vaudrey in order to prevent non-speech sounds from being improperly enhanced. Vaudrey and Irwan do not specifically teach setting a second speech enhancement gain based on a system volume level; determining a final speech enhancement gain using both the first speech enhancement gain and the second speech enhancement gain. In the same field of speech enhancement, Kriegel teaches setting a second speech enhancement gain based on a system volume level (0005, 0013-16, dialogue is boosted more when the system volume is low than at higher volumes); determining a final speech enhancement gain using both the first speech enhancement gain and the second speech enhancement gain (In combination with Vaudrey, the final gain would be based on both power ratio of the center channel and the system volume, as the processing could occur in series, as that is what is supported by applicant’s specification). It would have been obvious to one of ordinary skill in the art at the time of effective filing to adjust enhancement based on system volume as taught by Kriegel in the system of Vaudrey and Irwan in order to better enhance the dialogue in various playback settings (Kriegel 0003-04). Consider claim 2, Vaudrey teaches The method of claim 1, wherein setting the speech enhancement gain based on a signal power strength ratio of the center channel to a sum of the other channels comprises: setting the speech enhancement gain to be high when the signal power strength ratio of the center channel to the sum of the other channels is small (col 7 line 56- col 8 line 11, gain of center channel kicks in when ratio is too small); and setting the speech enhancement gain to be low when the signal power strength ratio of the center channel to the sum of the other channels is large (col 7 line 56- col 8 line 11, gain of center channel is not boosted when ratio is over desired VRA). Consider claim 3, Kriegel teaches The method of claim 1, wherein setting the speech enhancement gain based on a system volume level comprises: recognizing the system volume level and setting different speech enhancement gain when the recognized system volume level is within different volume ranges (0005, 0013-16, dialogue is boosted more when the system volume is low than at higher volumes); wherein the speech enhancement gain is set to be high when the system volume level is within a low range (0005, 0013-16, dialogue is boosted more when the system volume is low than at higher volumes); and wherein the speech enhancement gain is set to be low when the system volume level is within a high range (0005, 0013-16, dialogue is boosted more when the system volume is low than at higher volumes). Consider claim 6, Irwan teaches The method of claim 1, wherein performing intelligent enhancement gain control further comprises performing soft limiting processing on the set speech enhancement gain (0015, equation limits the gain applied as the signal X gets louder). Consider claim 7, Vaudrey teaches The method of claim 1, wherein the dynamic loudness balancing performed on the multi-channel audio source input comprises: enhancing the loudness of the signal of the center channel and attenuating the loudness of the signals of the other channels based on the set speech enhancement gain (col 7 line 56- col 8 line 11, gain of center channel raised and other channels lowered when ratio is too small); and performing concatenating and mixing processing on the enhanced signal of the center channel and the attenuated signals of the other channels to generate an output signal (col 8 lines 30-col 9 line 30, downmixing adjusted channels to different speaker formats). Consider claim 10, Vaudrey teaches A system for intelligent dynamic speech enhancement, comprising: a memory configured to store computer-executable instructions (col 7 lines 60-67, hardware and software combinations); and one or more processors (col 7 lines 60-67 processor) configured to execute the computer-executable instructions to implement a method for intelligent dynamic speech enhancement comprising the steps of: performing intelligent enhancement gain control on a multi-channel audio source input to determine speech enhancement gain, the multi-channel audio source input comprising a signal of a center channel and signals of other channels (figure 4, col 7 lines 40- col 8 lines 25, adjusting center channel gain of multi-channel signal to enhance speech); applying the speech enhancement gain in dynamic loudness balancing performed on the multi-channel audio source input (figure 4, col 7 lines 40- col 8 lines 25, adjusting center channel gain of multi-channel signal to enhance speech, dynamically at col 7 line 56- col 8 line 11, gain adjustment kicks in when desired ration violated); wherein the intelligent enhancement gain control comprises setting a first speech enhancement gain based on a signal power strength ratio of the center channel to a sum of the other channels (col 7 line 56- col 8 line 11, desired VRA, or voice to remaining audio). Vaudrey does not specifically teach performing speech detection. In the same field of speech enhancement on multichannel audio, Irwan teaches performing speech detection (0013-15 determining voice probabilities and enhancing based on probabilities). It would have been obvious to one of ordinary skill in the art at the time of effective filing to perform speech detection as taught by Irwan in the system of Vaudrey in order to prevent non-speech sounds from being improperly enhanced. Vaudrey and Irwan do not specifically teach setting a second speech enhancement gain based on a system volume level; determining a final speech enhancement gain using both the first speech enhancement gain and the second speech enhancement gain. In the same field of speech enhancement, Kriegel teaches setting a second speech enhancement gain based on a system volume level (0005, 0013-16, dialogue is boosted more when the system volume is low than at higher volumes); determining a final speech enhancement gain using both the first speech enhancement gain and the second speech enhancement gain (In combination with Vaudrey, the final gain would be based on both power ratio of the center channel and the system volume, as the processing could occur in series, as that is what is supported by applicant’s specification). It would have been obvious to one of ordinary skill in the art at the time of effective filing to adjust enhancement based on system volume as taught by Kriegel in the system of Vaudrey and Irwan in order to better enhance the dialogue in various playback settings (Kriegel 0003-04). Claim 11 contains similar limitations as claim 2 and is therefore rejected for the same reasons. Claim 12 contains similar limitations as claim 3 and is therefore rejected for the same reasons. Claim 15 contains similar limitations as claim 6 and is therefore rejected for the same reasons. Claim 16 contains similar limitations as claim 7 and is therefore rejected for the same reasons. Claim(s) 4, 5, 13, and 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Vaudrey and Irwan and Kriegel as applied to claims 1 above, and further in view of Yang et al. (US PAP2011/0066428). Consider claim 4, Vaudrey teach The method of claim 1, wherein performing speech detection comprises: extracting the signal of the center channel from the multi-channel audio source input (figure 4, col 7 line 56- col 8 line 11 center channel is processed separately from other channels). Vaudrey and Irwan and Kriegel do not specifically teach performing normalization on the signal of the center channel; and performing fast autocorrelation on the normalized signal of the center channel, a result of the fast autocorrelation representing a detection confidence level which is indicative of a possibility of speech being present in the signal of the center channel. In the same field of voice detection, Yang teaches performing normalization on the signal of the center channel 0054, performing smoothing and filtering); and performing fast autocorrelation on the normalized signal of the center channel, a result of the fast autocorrelation representing a detection confidence level which is indicative of a possibility of speech being present in the signal of the center channel (0089, using autocorrelation to detect speech in the media). It would have been obvious to one of ordinary skill in the art at the time of effective filing to use autocorrelation to determine voiced signals as taught by Yang in the system of Vaudrey and Irwan and Kriegel in order to more accurately. Consider claim 5, Irwan teaches the method of claim 4, wherein performing intelligent enhancement gain control further comprises: converting the detection confidence level to the speech enhancement gain (0015, using speech probability to determine speech enhancement gain using equation given); and performing smoothing processing on the speech enhancement gain (0015, given equation allows for smooth transition from high probabilities of speech to low probabilities of speech). Claim 13 contains similar limitations as claim 4 and is therefore rejected for the same reasons. Claim 14 contains similar limitations as claim 5 and is therefore rejected for the same reasons. Claim(s) 8, 9, 17, and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Vaudrey and Irwan and Kriegel as applied to claims 1 above, and further in view of Port et al. (US PAP2021/0265966). Consider claim 8, Vaudrey and Irwan and Kriegel teach The method of claim 1, but does not specifically teach further comprising performing crossover filtering processing on the multi-channel audio source input prior to the dynamic loudness balancing performed on the multi-channel audio source input. In the same field of multichannel speech enhancement, Port teaches performing crossover filtering processing on the multi-channel audio source input prior to the dynamic loudness balancing performed on the multi-channel audio source input (0181-86, gains applied to speech bands, i.e. midrange signals). It would have been obvious to one of ordinary skill in the art at the time of effective filing to perform band limited enhancement as taught by Port in the system of Vaudrey and Irwan and Kriegel in order to better enhance speech intelligibility (Port 0181). Consider claim 9, Port teaches the method of claim 8, further comprising: performing the dynamic loudness balancing only on the multi-channel audio source input within a mid-frequency range (181-86, gains applied to speech bands, i.e. midrange signals); and concatenating and mixing the multi-channel audio source input within the mid-frequency range that has undergone the dynamic loudness balancing with the multi-channel audio source input within a low-frequency range and a high-frequency range to generate an output signal (181-86, gains applied to speech bands, i.e. midrange signals, and output with the other bandwidths to generate output signal for channel). Claim 17 contains similar limitations as claim 8 and is therefore rejected for the same reasons. Claim 18 contains similar limitations as claim 9 and is therefore rejected for the same reasons. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to DOUGLAS C GODBOLD whose telephone number is (571)270-1451. The examiner can normally be reached 6:30am-5pm Monday-Thursday. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Flanders can be reached at (571)272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. DOUGLAS GODBOLD Examiner Art Unit 2655 /DOUGLAS GODBOLD/Primary Examiner, Art Unit 2655
Read full office action

Prosecution Timeline

Oct 22, 2024
Application Filed
Apr 24, 2026
Non-Final Rejection mailed — §103
Jul 19, 2026
Response Filed
Aug 12, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12744045
USING MACHINE LEARNING AND DISCRETE TOKENS TO ESTIMATE DIFFERENT SOUND SOURCES FROM AUDIO MIXTURES
3y 1m to grant Granted Sep 22, 2026
Patent 12744052
Apparatus For Estimating Emotion Using Multimodal Model And Method Of Training The Same
2y 1m to grant Granted Sep 22, 2026
Patent 12738283
PROCESSOR FOR GENERATING A PREDICTION SPECTRUM BASED ON LONG-TERM PREDICTION AND/OR HARMONIC POST-FILTERING
2y 8m to grant Granted Sep 15, 2026
Patent 12730966
MACHINE LEARNING TECHNIQUES FOR PREDICTING AND RANKING TYPEAHEAD QUERY SUGGESTION KEYWORDS BASED ON USER CLICK FEEDBACK
2y 9m to grant Granted Sep 08, 2026
Patent 12730985
MACHINE TRANSLATION SYSTEMS UTILIZING CONTEXT DATA
2y 1m to grant Granted Sep 08, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
83%
Grant Probability
94%
With Interview (+10.6%)
2y 9m (~10m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 1110 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month