Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 07/10/2024 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
DETAILED ACTION
Claim Rejections – 35 USC Code § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1, 5, 6, 11-13, and 15-16 are rejected under 35 U.S.C. §102(a)(1) as being anticipated by Ma et al., “ClearSpeech: Improving Voice Quality of Earbuds Using Both In-Ear and Out-Ear Microphones,” Proc. ACM Interactive, Mobile, Wearable and Ubiquitous Technologies, Vol. 7, No. 4, Article 170, December 2023 (“Ma”).
Regarding Claim 1, Ma teaches: A system (Fig. 3) comprising: a device of a user (earbuds, § 2.2, 7.1 and Fig. 8); a first sensor (out-ear microphone of earbud, § 1, 7.1, and Fig. 8) coupled to the device; a second sensor (in-ear microphone of earbud, § 1, 7.1, and Fig. 8) coupled to the device; and one or more processors (a laptop, a desktop, and a smartphone, § 7.1, and 8.9) coupled to the device, the one or more processors, individually or collectively, being configured to: receive, at the first sensor, a first audio signal (out-ear signal) with a first noise and a first distortion (the out-ear microphone captures air-conducted speech and is substantially exposed to external environmental noise); receive, at the second sensor, a second audio signal (in-ear signal) with a second noise and a second distortion (the in-ear microphone captures both bone- and air-conducted information and has significantly higher SNR because it is shielded by the ear canal), wherein the first noise is different than the second noise and the first distortion is different than the second distortion;
and determine an output audio signal (enhanced speech, Fig. 3) using at least a portion of at least one of a magnitude of the first audio signal or a magnitude of the second audio signal (out-ear, in-ear paired noise datasets, Fig. 5, and § 5.2)
and using at least a portion of at least one of a phase of the first audio signal or a phase of the second audio signal (the phase enhancement module receives noisy out-ear phase together with enhanced magnitude and determines enhanced phase, p. 170:6, and § 3).
Ma expressly establishes the claimed different noise characteristics. The out-ear microphone is critically influenced by external noise while the in-ear microphone experiences a substantially more limited impact.
Ma also establishes different distortion characteristics. The in-ear signal experiences occlusion-induced frequency distortion, including low-frequency amplification and high-frequency suppression, whereas the out-ear signal does not experience that same distortion.
Ma further teaches determining an output audio signal using magnitude information from the first and second audio signals. Figure 5 receives separate Out-ear Mag and In-ear Mag inputs and produces an Enhanced Mag output. See Ma, Fig. 5, and § 5.2.
Ma teaches phase processing as well. The Phase Enhancement Module receives noisy out-ear phase together with enhanced magnitude and determines enhanced phase. The enhanced magnitude and enhanced phase are then combined and inverse-STFT (iSTFT) is performed to recover enhanced speech. See Ma, p. 170:6, § 3.
Figure 3 graphically shows the complete path:
noisy in-ear/out-ear signals → magnitude enhancement → enhanced magnitude → phase enhancement → enhanced phase → enhanced spectrogram → iSTFT → enhanced speech.
Regarding claim 5, it recites the method counterpart of claim 1. Ma expressly performs the claimed acts by concurrently receiving out-ear and in-ear signals and processing the signals through magnitude and phase enhancement to obtain an enhanced speech waveform. See Ma, § 3, 5, and 6; Figs. 3, and 5.
Accordingly, for substantially the same reasons stated for claim 1, Ma anticipates claim 5.
Regarding claim 6, Ma expressly teaches that the first sensor is an out-ear microphone. The prototype shown in Fig. 8 includes one microphone embedded inside the earbud and another microphone located at the end of the earbud handle and facing downward to capture out-ear speech. Both microphones record their respective audio signals simultaneously. See Ma, p. 170:13, §7.1, Fig. 8.
Thus, Ma teaches the claimed microphone outside the device/ear-canal region.
Regarding claim 11, it requires using part of the magnitude of the first signal and part of the magnitude of the second signal.
Ma expressly teaches joint processing of in-ear and out-ear magnitude spectrograms. Figure 3 identifies separate Noisy Mag_Out and Noisy Mag_In inputs to the magnitude-enhancement stage.
Ma further teaches that the respective signals provide complementary information at different frequency bands and that combining the two improves speech enhancement.
The final fusion reconstructs complementary portions of the enhanced magnitude spectrum from the respective streams. Accordingly, Ma anticipates claim 11.
.
Claim 12 requires using all of the magnitude of the first audio signal.
Ma expressly evaluates an “Out Only” embodiment/variant in which only the out-ear stream is used. The final fully-connected layer contains 256 neurons and “recovers all the frequency bands.” Ma further states that for magnitude-only processing, the noisy out-ear phase is reused to reconstruct the time-domain speech. See Ma, p. 170:16, § 8.3, and Table 1.
Accordingly, Ma discloses determining an output signal using all of the magnitude information of the first/out-ear audio signal. Thus, using the entirety of the first/out-ear magnitude represents an expressly contemplated operating alternative.
Regarding claim 13, Ma teaches that the first/out-ear signal has greater noise than the second/in-ear signal. Ma explains that the in-ear microphone is isolated from external noise and has a substantially higher SNR than the out-ear microphone.
Ma further explains that the in-ear signal experiences greater frequency distortion because of the occlusion effect. The authors expressly state that the “in-ear speech is frequency-distorted,” while the out-ear signal contains comparatively more complete speech-frequency information.
Thus, when the out-ear signal is mapped to the first signal and the in-ear signal to the second signal:
first noise > second noise; and
first distortion < second distortion.
Regarding claim 15, it requires preprocessing the second audio signal.
Ma expressly preprocesses in-ear and out-ear signals before magnitude/phase enhancement, including segmentation and conversion to the time-frequency domain using STFT. The preprocessing produces magnitude and complex phase spectrograms for subsequent enhancement. See § 5.2.
Claim 16 requires the device to comprise a wearable device.
Ma repeatedly identifies the device as a wireless earbud, i.e., a wearable audio device. The Abstract describes ClearSpeech as a system “designed for wireless earbuds.”
Accordingly, Ma anticipates the wearable-device limitation of claim 16.
Claim Rejections – 35 USC Code § 103
4. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 2– 4, 8-10 and 18-20 are rejected under 35 U.S.C. §103 as being unpatentable over Ma et al., “ClearSpeech” in view of Merks (hereafter Merks, US 20240365073 A1) and further in view of Kim et al. (hereafter Kim, US 20220360891 A1).
Regarding claim 2, it requires: determining a level of the first noise in the first audio signal; and when that level is above a first threshold, using all of the phase of the second audio signal.
Ma teaches the underlying system of claim 1, including first/out-ear and second/in-ear microphone signals having different noise and distortion characteristics. Ma expressly finds that the in-ear microphone is more resilient to environmental noise than the out-ear microphone and that enhancement gains attributable to the in-ear signal increase at low out-ear SNR. See Ma §8.5.
Ma expressly evaluates first/out-ear signal quality according to SNR values of: −10 dB, −5 dB, 0 dB, 5 dB, and 10 dB. Ma concludes that the in-ear microphone supplies greater benefit at lower out-ear SNR. See Ma §8.5, Table 2.
Ma does not teach the claimed three-state phase-selection procedure based upon first-noise thresholds.
However, Merks teaches the missing noise estimation and threshold-dependent decision framework. Merks calculates directional variance 302; calculates noise-reference variance 304; determines a variance threshold/input 306; determines whether the measured relationship satisfies the threshold condition; and changes subsequent processing accordingly. He provides the actual mechanism for estimating microphone noise and testing a measured noise-related quantity against a threshold. See Fig. 3 and corresponding ¶ 0092–0106.
Figure 3 shows calculated directional variance and noise-reference variance supplied to a threshold determination block.
Merks also teaches generating estimated noise information from first and second microphone data and a constructed noise reference. See Fig. 4.
Kim further provides two independently available phase sources. Phase extractor 148A generates first phase 161A; phase extractor 148B generates second phase 161B (¶ 0037–0041).
Thus, it would have been obvious to one of ordinary skills in the art at the time of the effective filing date of the application to incorporate the noise-estimation and threshold-based control technique of Merks’s and the separate phase-processing architecture of Kim into Ma's dual-microphone speech-enhancement system.
Ma identifies which sensor is more reliable under high-noise conditions; Merks detects that high-noise condition by comparison with a threshold; and Kim provides the alternative second-sensor phase that can be used once that condition is detected. The combination therefore would have yielded the predictable result of improving reconstructed speech when the first microphone is significantly corrupted by environmental noise.
Accordingly, it would have been obvious to determine the output audio signal using all of the phase of the second audio signal when the determined first-noise level exceeds the first threshold, as recited in claim 2.
Regarding claim 3, it requires using part of the first phase and part of the second phase in an intermediate noise range.
Ma expressly teaches that the in-ear and out-ear microphones provide complementary speech information at different frequencies and that combining their information improves speech enhancement.
Ma also establishes that the relative advantage of the in-ear signal varies with SNR rather than existing as an all-or-nothing phenomenon.
Merks teaches adaptive processing according to measured microphone-noise conditions and threshold comparisons. Its Fig. 2 further shows scaling, weighting, combining, and differencing signals derived from different microphone channels, while Fig. 3 determines which adaptive state is appropriate from measured variance/noise conditions.
Kim provides the independently available first and second phase values.
Having provided: a high-noise region in which the second/in-ear phase is favored, and a low-noise region in which the first/out-ear phase remains reliable, it would have been obvious to provide an intermediate region in which contributions from both phase sources are used. Blending or weighting the two available phase sources in the transition region would predictably:
avoid an abrupt transition between phase sources;
preserve usable phase information from both sensors; and
follow the adaptive weighted multi-signal processing approach taught by Merks.
The modification represents application of Merks's known adaptive weighting/combination principle to Kim's separately available phase values in Ma's complementary two-sensor environment.
Regarding claim 4 requires using all of the first-signal phase below a lower-noise threshold.
Ma's SNR study shows that the comparative benefit from in-ear assistance is greatest at poor out-ear SNR and relatively less significant as the out-ear signal improves.
Merks provides threshold-dependent classification of microphone/noise conditions.
Kim expressly permits use of the first phase derived from the first microphone signal with the enhanced magnitude (¶ 0040–0042).
Thus, it would have been obvious to one of ordinary skills in the art at the time of the effective filing date of the application to use the first/out-ear phase when its measured noise falls below a lower threshold because, in that condition, the first phase is sufficiently reliable and substitution by the internally derived phase is unnecessary.
Thus, the combination naturally provides:
Noise condition
Phase used
high first noise
second phase
intermediate noise
part first + part second
low first noise
first phase
The arrangement provides a predictable adaptive phase-selection system based on the relative reliability of the two sensor signals.
Regarding dependent claims 8-10 and 18-20, the device of Ma as modified by Merk and Kim teaches all limitations for the similar reasons as set forth in the rejection of dependent claims 2-4 because these dependent claims recite similar claim limitations as those of dependent claims 2-4 (see rejection of claims 2-4 as set forth above).
Claim 7 is rejected under 35 U.S.C. §103 as being unpatentable over Ma.
Regarding claim 7, it recites: wherein the second sensor comprises a feedback microphone.
Ma expressly states that commercial earbuds commonly include an in-ear microphone used for ANC, and “ClearSpeech” deliberately proposes reusing that in-ear ANC microphone for speech enhancement to avoid additional hardware. See Ma §2.2.
A feedback ANC microphone is conventionally positioned acoustically inward of the earbud transducer to sense the residual acoustic field near the user's ear.
Thus, it would have been obvious to one of ordinary skills in the art at the time of the effective filing date of the application to use such an inward-facing feedback microphone as Ma's in-ear sensor because:
Ma expressly seeks to reuse an existing ANC in-ear microphone;
the feedback microphone occupies the appropriate inward-facing acoustic location; and
using the existing microphone avoids the hardware overhead that Ma expressly seeks to eliminate.
Claim 14 is rejected under 35 U.S.C. §103 as being unpatentable over Ma et al. in view of Luneau et al. (hereafter Luneau, US 20240005937 A1).
Regarding claim 14, it requires in relevant part: using a trained machine-learning model to determine a mask for the first audio signal, the mask being configured to at least partially denoise the first audio signal.
Ma provides the underlying claimed dual-sensor speech-enhancement system and expressly employs trained deep-learning processing for enhancing noisy speech. As previously mapped, Ma's magnitude-enhancement network receives noisy out-ear and in-ear magnitude spectrograms and produces enhanced magnitude information.
Luneau provides an especially relevant teaching concerning the use of a trained machine-learning model to remove or compensate for degradation in an audio signal.
Luneau explains in its Background that an internal sensor is relatively protected from environmental noise but may suffer from distortion and limited spectral bandwidth, whereas an external microphone can capture a more natural air-conducted signal. See Luneau ¶ 0003–0010, and p. 1.
Luneau then expressly proposes machine-learning processing to predict a higher-quality signal. ¶ 0015–0017 explain that the bone-conducted/internal audio signal is processed by a machine-learning model previously trained to produce a predicted air-conducted signal, thereby improving the internal signal.
More specifically, Luneau ¶ 0083–0088 explain that the ML processing may operate on frequency-domain representations and that the input signal may be transformed by FFT, DFT, DCT, or STFT before being supplied to the machine-learning model. The disclosed processing is directed to predicting frequency components missing or degraded in the sensor signal.
Luneau further provides a concrete trained neural-network architecture. Figure 7 shows:
feedforward layer 71 → LSTM layer 72 → LSTM layer 73 → feedforward layer 74 → output layer 76.
Figure 8 then shows that this trained machine-learning model processes the measured internal-sensor signal before determination of the resulting internal signal and confidence index.
Ma expressly discusses trained speech-enhancement systems such as PHASEN and FullSubNet+ that learn complex ideal ratio masks for enhancing noisy speech. Thus, Ma provides the actual mask concept, while Luneau reinforces that applying a trained neural model to degraded wearable-device microphone signals to obtain an enhanced representation was known.
Thus, it would have been obvious to one of ordinary skills in the art at the time of the effective filing date of the application to configure the trained speech-enhancement processing of Ma to determine a denoising mask, as suggested by Ma's disclosed mask-based speech-enhancement techniques, while implementing the trained model according to the machine-learning audio-enhancement teachings of Luneau. Ma and Luneau address the same recognized problem of improving degraded speech captured by sensors associated with wearable audio devices. Luneau expressly teaches training a neural-network model to transform a degraded sensor signal toward a higher-quality target signal, including frequency-domain processing, while Ma identifies mask estimation as a known mechanism for neural speech enhancement. A skilled artisan would have recognized generation and application of a time-frequency mask as a predictable implementation of the trained enhancement model for suppressing noise-dominated components while retaining speech-dominated components, thereby improving the quality of the first audio signal.
Claim 17 is rejected under 35 U.S.C. §103 as being unpatentable over Ma in view of Kim.
Regarding claim 17, it recites a non-transitory computer-readable medium containing instructions implementing the processing substantially corresponding to claim 1.
Ma teaches the underlying audio operations but does not emphasize the claimed CRM form.
Kim expressly teaches: a non-transitory computer-readable medium storing instructions that cause processors to determine a first phase from a first audio signal, determine a second phase from a second audio signal, generate an enhanced signal, and combine enhanced magnitude with the respective first and second phases.
See Kim ¶ 0301.
Thus, it would have been obvious to one of ordinary skills in the art at the time of the effective filing date of the application to store Ma's executable ClearSpeech signal-processing software on a non-transitory computer-readable medium as taught by Kim because Ma's DL/audio-processing functions necessarily execute as software on a processor, and Kim expressly teaches that implementation for closely analogous magnitude/phase audio processing.
The modification merely implements Ma's known software operations in a conventional persistent computer-readable storage form.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Pradip Podder whose telephone number is (571)272-8543. The examiner can normally be reached Monday - Friday 8:00 am- 5 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vivian Chin can be reached at 571-272-7547. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center,
https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PRADIP C. PODDER/
Examiner, Art Unit 2694
/VIVIAN C CHIN/Supervisory Patent Examiner, Art Unit 2695