Prosecution Insights
Last updated: August 16, 2026
Application No. 17/813,542

INTELLIGENT SPEECH OR DIALOGUE ENHANCEMENT

Non-Final OA §103
Filed
Jul 19, 2022
Examiner
ISKENDER, ALVIN ALIK
Art Unit
2654
Tech Center
2600 — Communications
Assignee
Bose Corporation
OA Round
5 (Non-Final)
46%
Grant Probability
Moderate
5-6
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 46% of resolved cases
46%
Career Allowance Rate
12 granted / 26 resolved
-15.8% vs TC avg
Strong +55% interview lift
Without
With
+54.8%
Interview Lift
resolved cases with interview
Typical timeline
3y 3m
Avg Prosecution
11 currently pending
Career history
48
Total Applications
across all art units

Statute-Specific Performance

§101
13.9%
-26.1% vs TC avg
§103
56.5%
+16.5% vs TC avg
§102
25.5%
-14.5% vs TC avg
§112
3.7%
-36.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 26 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant’s arguments with respect to claim(s) 1-9, 11-15, 17-20, 22-23, 25-27 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-2, 5-9, 11-15, 17-20, 22-23, 25-27 is/are rejected under 35 U.S.C. 103 as being unpatentable over Bennett (US-20070027682-A1) in view of Stark (US 20210201926 A1), Rasmussen (US 20150003623 A1), and Gorlow (US 20220199101 A1). Regarding claim 1, Bennett discloses a method for audio signal processing, the method comprising: during playback of an audio signal, analyzing content of the audio signal prior to the playback of the content to determine whether one or more predefined conditions are met to indicate that the content includes speech; (Fig 3 Block 309: Voice Activity Detection; [0043]: Voice activity detection may be applied to segregate background and voice, or to identify the portions of the audio that contain voice) wherein the analyzing comprises analyzing metadata associated with the content, (Fig 2B, [0040]: AIPS handles separate language or audio tracks as needed, with DVD being given as an example. It is known to those with ordinary skill in the art that recognizing these tracks necessitates analyzing metadata associated with the content according to the format’s standards.) Bennett does not teach, but Gorlow does teach: in response to determining the one or more predefined conditions are met and based on analyzing the sound in the environment, automatically fading from applying no playback equalization or a first playback equalization to the audio signal to applying a second playback equalization to the audio signal, wherein the second in response to determining the one or more predefined conditions are not met and based on analyzing the sound in the environment, automatically from applying no playback equalization, the first playback equalization, or the second playback equalization to the audio signal to applying i) no playback equalization to the audio signal, or ii) a thirdapply dialogue enhancement to estimated dialogue components; [0104]-[0110]: apply a cross fade when switching between not applying dialogue equalization and applying dialogue equalization) Although Bennett teaches the use of a transition equalizer when switching between equalizer types such as in Figure 7, it does not teach fading between the equalizer types. Gorlow does teach cross-fading between speech equalization and no equalization, and it would have been obvious to one with ordinary skill in the art before the effective filing date of the present application to use Gorlow’s teachings when performing Bennett’s method because it enables seamless transitions that don’t degrade the user experience (see Gorlow [0104]-[0105]). Bennett and Gorlow do not teach analyzing content using a trained machine learning model; metadata including at least one of text associated with the content or genre data associated with the content, and wherein the content includes non-speech content. However, Stark does teach analyzing content using a trained machine learning model; ([0007]: machine learned algorithm to determine speech presence) metadata including at least one of text associated with the content or genre data associated with the content, and wherein the content includes non-speech content. ([0020]: analyze metadata for text flags such as speech or music) It would have been obvious to one with ordinary skill in the art before the effective filing date to analyze metadata associated with content as taught by Stark because it can be used to determine whether which type of audio processing should be used on the content (see Stark [0019]-[0020], Abstract). Bennett, Gorlow, and Stark do not teach analyzing sound in an environment in which the audio signal is to be played; based on analyzing the sound in the environment, applying to the audio signal a first playback equalization, a second playback equalization different from the first playback equalization; Rasmussen does teach analyzing sound in an environment in which the audio signal is to be played; ([0013]: analyze incoming sound from microphones) based on analyzing the sound in the environment, applying to the audio signal a first playback equalization, a second playback equalization different from the first playback equalization ([0013], [0075]: apply an equalization function depending on the results of the analysis) It would have been obvious to one with ordinary skill in the art before the effective filing date to apply equalization filters based on sound in the environment because the need for equalization can be dependent on how noisy the environment is (see Stark [0002]). Regarding claim 2, Bennett, Gorlow, Stark, and Rasmussen disclose the elements of parent claim 1 as described above. Stark further teaches analyzing the content of the audio signal prior to the playback of the content to determine whether the one or more predefined conditions are met involves determining a confidence level that the content includes the speech using the trained machine-learning model; ([0020]: use of a machine learning algorithm to make a decision of speech presence with a confidence level) the method further comprises adjusting a speed of transitioning from applying no playback equalization or the first playback equalization to the audio signal to applying the second playback equalization to the audio signal based on the confidence level that the content includes the speech. ([0006]-[0007]: ratio of equalization functions can be dependent on the confidence level; [0027] interpolation in transitional periods) Regarding claim 5, Bennett, Gorlow, Stark, and Rasmussen disclose the elements of its parent claim 1. Stark further discloses the method wherein the metadata indicates that the content includes speech. ([0020]: metadata speech flag) Regarding claim 6, Bennett, Gorlow, Stark, and Rasmussen disclose the elements of its parent claim 1. Bennett further discloses the method wherein the analyzing comprises analyzing a voice track of the audio signal, wherein the one or more predefined conditions includes the voice track exceeding a threshold value. ([0042]: cross correlation involves signal comparison; [0045]: proportional amplitude level of voice;) Regarding claim 7, Bennett, Gorlow, Stark, and Rasmussen disclose the elements of its parent claim 1. Bennett further discloses the method wherein the audio signal comprises different channels and the analyzing comprises analyzing of the content comprises the different channels. ([0031]: correlating channels to segregate signals) Regarding claim 8, Bennett, Gorlow, Stark, and Rasmussen disclose the elements of its parent claim 7. Bennett further discloses the method wherein analyzing the different channels comprises comparing correlated content between two channels of the different channels. ([0031]: correlating center channels with two or more other channels) Regarding claim 9, Bennett, Gorlow, Stark, and Rasmussen disclose the elements of its parent claim 1. Bennett further discloses the method wherein the analyzing comprises analyzing the center channel of the audio signal. ([0031]: center channel analysis) Regarding claim 11, Bennett, Gorlow, Stark, and Rasmussen disclose the elements of parent claim 1. Bennett further discloses the method wherein applying the second playback equalization to the audio signal comprises increasing a volume of the speech within the content relative to other content within the audio signal. [0045]: proportional amplitude of voice to background) Regarding claim 12, Bennett, Gorlow, Stark, and Rasmussen disclose the elements of parent claim 1. Bennett further discloses the method wherein applying the second playback equalization to the audio signal comprises decreasing a volume of the non-speech content within the audio signal. (Fig 3: Proportionate Amplitude Regulator 315) Regarding claim 13, Bennett, Gorlow, Stark, and Rasmussen disclose the elements of parent claim 12. Bennett further discloses the method further comprising increasing a volume of the speech within the content. ([0027]-[0028], [0045], Fig 3: Amplitude and volume of speech is independently controlled) Regarding claim 14, Bennett, Gorlow, Stark, and Rasmussen disclose the elements of parent claim 1. Bennett further discloses the method wherein the second playback equalization comprises at least one of low frequency enhancement or music playback enhancement. ([0046]: enhancement of concert sounds; [0053]: enhancement of specific frequencies in the spectrum, including bass) Regarding claim 15, Bennett, Gorlow, Stark, and Rasmussen disclose the elements of parent claim 1. Bennett further discloses the method wherein at least one of the first playback equalization or the one or more predefined conditions are configurable by a user. ([0027]-[0028]: user configurable equalization settings) Regarding claims 17-20, 22-23, 25, they contain limitations analogous to claims 1-2, 7-10, 14-15 as described above. They are thus rejected for the same reason and in the same manner as described above. Regarding claim 26, Bennett, Gorlow, Stark, and Rasmussen disclose the elements of its parent claim 1. Bennett further discloses the method wherein automatically transitioning gradually from applying no playback equalization or the first playback equalization to the audio signal to applying the second playback equalization to the audio signal comprises transitioning gradually from applying no playback equalization or the first playback equalization to the audio signal to applying the second playback equalization to the audio signal without any user input. ([0065]-[0066]: transitional equalization is dependent on automatically identifying a transitional period) Regarding claim 27, Bennett, Gorlow, Stark, and Rasmussen disclose the elements of its parent claim 1. Stark further discloses the method wherein during playback of the audio signal, analyzing the content of the audio signal using the trained machine-learning model prior to the playback of the content to determine whether the one or more predefined conditions are met to indicate that the content includes the speech comprises analyzing the content of the audio signal using the trained machine-learning model prior to the playback of the content to determine whether the one or more predefined conditions are met to indicate that the speech is a primary content of the content. ([0020]: using a machine learned algorithm to determine whether a source is speech or not speech) Claim(s) 3-4 is/are rejected under 35 U.S.C. 103 as being unpatentable over Bennett in view of Stark, Gorlow, and Rasmussen as applied to the claims above, and further in view of Chen et al. ("Mixed Stereo Audio Classification Using a Stereo-Input Mixed-to-Panned Level Feature"). Regarding claim 3, Bennett, Gorlow, Stark, Rasmussen, disclose parent claim 2 as described above. They do not disclose the method wherein the trained machine-learning model comprises a deep learning model that estimates energy levels of the audio signal. However, Chen et al. does disclose the method wherein the trained machine-learning model comprises a deep learning model that estimates energy levels of the audio signal (Section II: Estimating speech to music ratio, which is a ratio of the energy levels of speech and music components in a signal; Section II E: wide variety of existing machine learning algorithms can be used, understood by those with ordinary skill in the art to encompass deep learning techniques) It would have been obvious to one with ordinary skill in the art before the effective filing date of the claimed invention to use a trained machine-learning model that estimates energy levels of speech and music to analyze an audio signal for speech presence as it allows for the classification of speech and music in an audio signal to a high degree of accuracy (see Chen et al. Section I). Regarding claim 4, Bennett, Gorlow, Stark, Rasmussen, and Chen et al. disclose parent claim 3 as described above. Chen et al. further discloses the method wherein the energy levels of the audio signals include energy levels of any combination the speech, a music component of the audio signal, and a singing component of the audio signal. (Section I: SMR is described as the ratio between energy levels of speech and music components in an audio signal) Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALVIN ISKENDER whose telephone number is (703)756-4565. The examiner can normally be reached M-F. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, HAI PHAN can be reached at (571) 272-6338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ALVIN ISKENDER/Examiner, Art Unit 2654 /HAI PHAN/Supervisory Patent Examiner, Art Unit 2654
Read full office action

Prosecution Timeline

Show 13 earlier events
Aug 15, 2025
Applicant Interview (Telephonic)
Aug 19, 2025
Response Filed
Sep 08, 2025
Examiner Interview Summary
Dec 09, 2025
Final Rejection mailed — §103
Feb 12, 2026
Response after Non-Final Action
Mar 31, 2026
Request for Continued Examination
Apr 02, 2026
Response after Non-Final Action
Jul 22, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12632658
SYSTEM AND METHODS FOR KEY-PHRASE EXTRACTION
4y 3m to grant Granted May 19, 2026
Patent 12562244
COMBINING DOMAIN-SPECIFIC ONTOLOGIES FOR LANGUAGE PROCESSING
4y 12m to grant Granted Feb 24, 2026
Patent 12531078
NOISE SUPPRESSION FOR SPEECH ENHANCEMENT
3y 4m to grant Granted Jan 20, 2026
Patent 12505825
SPONTANEOUS TEXT TO SPEECH (TTS) SYNTHESIS
3y 1m to grant Granted Dec 23, 2025
Patent 12456457
ALL DEEP LEARNING MINIMUM VARIANCE DISTORTIONLESS RESPONSE BEAMFORMER FOR SPEECH SEPARATION AND ENHANCEMENT
3y 5m to grant Granted Oct 28, 2025
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
46%
Grant Probability
99%
With Interview (+54.8%)
3y 3m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 26 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month