Prosecution Insights
Last updated: October 02, 2026
Application No. 19/081,230

MIXING SPEECH AND NOISE IN AN EAR-WORN DEVICE BASED ON NOISE LEVEL

Non-Final OA §103
Filed
Mar 17, 2025
Priority
Mar 18, 2024 — provisional 63/566,434
Examiner
JOSHI, SUNITA
Art Unit
Tech Center
Assignee
Fortell Research Inc.
OA Round
1 (Non-Final)
81%
Grant Probability
Favorable
1-2
OA Rounds
8m
Est. Remaining
87%
With Interview

Examiner Intelligence

Grants 81% — above average
81%
Career Allowance Rate
923 granted / 1138 resolved
+21.1% vs TC avg
Moderate +6% lift
Without
With
+6.1%
Interview Lift
resolved cases with interview
Typical timeline
2y 2m
Avg Prosecution
23 currently pending
Career history
1149
Total Applications
across all art units

Statute-Specific Performance

§101
1.1%
-38.9% vs TC avg
§103
68.6%
+28.6% vs TC avg
§102
18.6%
-21.4% vs TC avg
§112
2.5%
-37.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1138 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobvious Ness. 1. Claims 1-5, 13, 21 and 23 are rejected under 35 U.S.C. 103 as being unpatentable over Sabin et. al. (US2022/0369047A1), hereinafter “Sabin “ in view of Sun, Jundai et al. ( WO2023/091276A1), hereinafter “ Jundai”. As to Claim 1, Sabin teaches an ear-worn device ( wearable hearing assistance devices, [0001]), comprising: noise reduction circuitry( active noise reduction (ANR) system 106, [0030], Figure 1) Regarding the following: comprising: neural network circuitry configured to determine, using a neural network, one or both of a neural network-predicted speech component of an input audio signal and a neural network-predicted noise component of the input audio signal; mixing circuitry configured to perform mixing resulting in an output signal corresponding to the neural network-predicted speech component of the input audio signal mixed with the neural network-predicted noise component of the input audio signal, Sabin teaches an audio machine learning base processing in a wearable hearing device. [0021]. Further, generating an ML-enhanced version of the incoming audio, then mix that enhanced signal with some of the original unprocessed signal using a coefficient chosen from an environmental noise assessment [0004]- [0006], [0032]- [0035]. The mixed result is output as the user-facing audio. This “back-mixing” of the original signal helps mask or remediate ML artifacts while still allowing the enhancement system to dominate the listening experience [0021], [0042]. The mixing amount can be adapted based on noise level, SNR, or a direct ML preference model [0032]-[0035], Figure 1). Sabin does not explicitly teach using a neural network, one or both of a neural network-predicted speech component of an input audio signal and a neural network-predicted noise component of the input audio signal; mixing circuitry configured to perform mixing resulting in an output signal corresponding to the neural network-predicted speech component of the input audio signal mixed with the neural network-predicted noise component of the input audio signal. However, Jundai in related field (speech processing) teaches a system with an with an audio signal that contains speech plus other sounds, such as background hiss, traffic, birds, wind, or other noises. It uses one model to pull out speech and another model to pull out stationary noise like constant hiss or hum. It then defines a third category by taking the difference between stationary noise and all non-speech content, which captures non-stationary noise such as birdsong, wind gusts, or passing cars. The separated parts are treated as independent ingredients that can be recombined with chosen weights. This lets a user reduce unwanted noise while keeping some ambient sound if desired. In some versions, the non-stationary noise is further filtered to isolate a specific noise object, such as birdsong or traffic. A classifier can also detect which kind of noise is present and automatically select filters or mixing settings. The system may be used for speech enhancement, but also for creative remixing of audio where ambience is intentionally retained. See at least abstract, [0002], [0009], [0013], Figure 3a, 3b and [0053]-[0063]. Separating speech from arbitrary audio signals may be performed accurately with a model (e.g. implemented with a neural network) trained to predict masks for separating speech content provided a representation of an audio signal. Additionally, the same mask used to extract the speech content may also be used to extract non-speech content meaning that the same trained model may be used to determine both the speech content and the non-speech content. [0017] and adaptive mixing control circuitry configured to control the mixing performed by the mixing circuitry based, at least in part, on: a level of an estimate of a stationary noise, component of the input audio signal; and/or a level of the neural network-predicted noise component of the input audio signal, Sabin on [0036] teaches “ a noise level detector’ in hearing aids and [0015] teaches an accurate trained model (e.g. implemented with a neural network) may be used to determine the stationary noise content given a representation of an audio signal. Stationary noise content may be defined precisely, and large amounts of training data are readily available, may be recorded or created synthetically which means stationary noise isolator model can be trained to be very accurate. It would have been obvious to one of ordinary skill in the art, before the effective filing date of the invention to upgrade the machine learning model of the hearing aid Sabin to improve audio processing by a neural network by separating speech, stationary noise, and non-stationary noise more flexibly so the output can be remixed for better intelligibility or better ambience. [0007]- [0009], [0013], [0018]. Further, on [0039] teaches the audio signal S.sub.in may be provided to a trained model wherein the trained model has been trained to output a mask M.sub.1, M.sub.2 for suppressing a certain type of noise wherein the mask M.sub.1, M.sub.2 is typically defined as the magnitude ratio between the desired speech S.sub.m,f and the audio signal mixture X.sub.m,.sub.f for each time frame and frequency bin. That is, the mask M is defined as As to Claim 2, Sabin in view of Jundai teaches the limitations of Claim 1, and wherein the mixing circuitry is configured to mix the neural network-predicted speech component of the input audio signal with the neural network-predicted noise component of the input audio signal, Jundai teaches on [0072] In the exemplary embodiment shown in fig. 6a the classifier 15 predicts birdsong as one noise object which is present in the audio signal and provides an indication of birdsong to the selector 16. Selector 16 accesses the database 171 and finds that filter data 172b describes a filter 13’ associated with birdsong (e.g. the filer with a passband between 3 kHz and 7 kHz as mentioned in the above) whereby the selector 16 selects filter data 172b and enables the birdsong filter 13’ to be applied to the non-stationary noise content Also, see [0017] and [0069]- [0070]. As to Claim 3, Sabin in view of Jundai teaches the limitations of Claim 1, and, wherein the mixing circuitry is configured to mix the neural network-predicted speech component of the input audio signal with the input audio signal, Jundai on [0073] and Figure 6b teaches [0073] Fig. 6b depicts another audio processing system 1 comprising a classifier 15 according to some implementations. The classifier 15 predicts the presence of at least one noise object (e.g. the presence of at least one noise object of a predetermined set of noise objects) and provides the predicted noise object(s) to a selector 16. The selector 16 accesses a database 173 of trained noise object isolation models 174a, 174b, 174c and selects at least one trained noise object isolation model 174a trained to predict a mask for isolating the at least one predicted noise object The predicted mask of the selected noise object isolation model 174a is applied to the audio signal to obtain the noise object. The noise object is in turn provided to mixing unit 14 and combined with the non-stationary noise stationary noise and speech content wherein each content type is provided with a respective weighting factor. Thus, the user or mixing engineer may set the weighting factors as desired and e.g. suppress the stationary and non-stationary noise and amplify only the noise object of the non- [AltContent: textbox ( )]stationary noise and the speech content As to Claim 4, Sabin in view of Jundai teaches the limitations of Claim 1, and Jundai further teaches wherein the mixing circuitry is configured to mix( Figure 6b, 14) the neural network-predicted noise component( noise object from classifier 15 which is a neural network, [0073], [0070]) of the input audio signal( audio signal) with the input audio signal ( speech isolator 11). As to Claim 5, Sabin in view of Jundai teaches the limitations of Claim 1, and Jundai further teaches wherein: the mixing circuitry is configured to perform the mixing such that the output signal comprises a component comprising the neural network-predicted noise component of the input audio signal multiplied by a weight; and the adaptive mixing control circuitry is configured, when controlling the mixing performed by the mixing circuitry, to control the weight, Jundai on [0075] and [0076] teaches the number of audio content types which are provided to the mixing unit 14 may change depending on the contents of the audio signal whereby the user or mixing engineer may select a desired relative signal strength for each of the components by selecting the weighting factors manually. However, as shown in fig. 6c the weighting factors may be determined automatically, e.g. selected by the selector 16 from a database 175 of weighting factor sets 176a, 176b, 176c based on which noise object(s) the classifier 15 predicts to be present in the audio signal. Each set 176a, 176b, 176c of weighting factors in the database comprising a value for at least each one of α.sub.2 , γ.sub.2 and μ.sub.2.[0076] For instance, if the classifier 15 predicts the presence of birdsong the selector 16 may select a set of weighting factors 176c which suppresses the stationary noise, amplifies the non-stationary noise and amplifies the speech content as birdsong is considered to not disturb the speech intelligibility while adding a pleasant ambiance. On the other hand, if the classifier 15 predicts the presence of wind sounds the selector 16 may select a different set of weighting factors 176a which suppresses the stationary noise and the non-stationary noise (which includes the wind sound) while amplifying the speech content as wind sounds is considered to not be an unwanted disturbance. As to Claim 13, Sabin in view of Jundai teaches the limitations of Claim 1, and wherein the adaptive mixing control circuitry is configured, when controlling the mixing performed by the mixing circuitry, to control a relative weight applied to the neural network-predicted noise component of the input audio signal versus the neural network-predicted speech component of the input audio signal in the output signal, Jundai teaches on [0076] and [0077] f the classifier 15 predicts the presence of birdsong the selector 16 may select a set of weighting factors 176c which suppresses the stationary noise, amplifies the non-stationary noise and amplifies the speech content as birdsong is considered to not disturb the speech intelligibility while adding a pleasant ambiance. On the other hand, if the classifier 15 predicts the presence of wind sounds the selector 16 may select a different set of weighting factors 176a which suppresses the stationary noise and the non-stationary noise (which includes the wind sound) while amplifying the speech content as wind sounds is considered to not be an unwanted disturbance. [0077] In this manner, the selector 16 automatically selects a suitable weighting factor set 176a, 176b, 176c for all audio signals according to a predetermined set of rules wherein a user or mixing engineer, optionally, provides some preferences to modify the rules. The preferences e.g. indicates a desire to suppress some noise objects more than others (e.g. suppress all manmade noise objects such as machine sounds and traffic sounds but keep all nature sounds such as birdsong, rain sound and thunder sound). Alternatively or additionally, the preferences e.g. indicates a desire to enhance speech intelligibility at the cost of less ambience wherein any reverberation and stationary noise is omitted entirely and any noise object is attenuated. As to Claim 21, Sabin in view of Jundai teaches the limitations of Claim 1, and Jundai teaches further comprising stationary noise reduction circuitry( Figure 3a, Stationary Noise isolator, 12) configured to generate the estimate of the stationary noise component of the input audio signal( audio signal, [0054] teaches the audio signal is provided to the stationary noise isolator model 12 trained to predict a mask M.sub.2 for separating the residual audio content from the stationary noise content By applying the mask M.sub.2 to the audio signal, e.g. in accordance with equation 7 in the above, at least the stationary noise content is determined at step S2b.) As to Claim 23, Sabin in view of Jundai teaches the limitations of Claim 1, and, Jundai teaches further comprising circuitry configured to smooth the estimate of the stationary noise component of the input audio signal or the neural network-predicted noise component of the input audio signal is smoothed, [0059] and Figure 3c teaches wherein the non-stationary noise is N.sub.NS processed with a bandpass filter 13 at step S3 prior to being fed to the mixer unit 14. Additionally, the filtered non-stationary noise may be smoothed with a smoothing kernel or smoothing filter (not shown) prior to being fed to the mixer unit 14. Allowable Subject Matter Claims 6-12, 14-20 and 22 are objected to as being dependent upon a rejected base claim but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to SUNITA JOSHI whose telephone number is (571)270-7227. The examiner can normally be reached 8-3. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Duc Nguyen can be reached at 5712727503. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SUNITA JOSHI/Primary Examiner, Art Unit 2691
Read full office action

Prosecution Timeline

Mar 17, 2025
Application Filed
Aug 28, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12745028
SET-TOP BOX OF REDUCED HEIGHT
2y 5m to grant Granted Sep 22, 2026
Patent 12745044
SOUND APPARATUS
2y 3m to grant Granted Sep 22, 2026
Patent 12745043
MICRO SPEAKER STRUCTURE INCLUDING DIAPHRAGM WITH LOCALLY THINNER REGIONS AND METHOD FOR FORMING THE SAME
2y 2m to grant Granted Sep 22, 2026
Patent 12739560
ACOUSTIC DEVICES
2y 10m to grant Granted Sep 15, 2026
Patent 12739570
SOUND APPARATUS
2y 2m to grant Granted Sep 15, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
81%
Grant Probability
87%
With Interview (+6.1%)
2y 2m (~8m remaining)
Median Time to Grant
Low
PTA Risk
Based on 1138 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month