DETAILED ACTION
This communication is in response to the Application filed on 03/07/2025. Claims 1-20 are pending and have been examined. Claims 1, 19 and 20 are independent. This Application was published as US Pub 2025/0299678.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 04/10/2025 was filed. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Priority
Applicant’s claims for benefit of a provisional application 63/568,384 submitted on 03/21/2024 is acknowledged.
Claim Objections
Claim 8 is objected to because of the following informalities: Claim 8 recites “…wherein the ASR is trained recognize speech in the directional audio data” where “…wherein the ASR is trained to recognize speech in the directional audio data” was perhaps intended. Appropriate correction is required.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 6-7, 9-14 and 16-20 are rejected under 35 U.S.C. 103 as being unpatentable over Mansour et al. (US Pat 10,657,981) in view of Panchapagesan et al. ("A conformer-based waveform-domain neural acoustic echo canceller optimized for ASR accuracy." arXiv preprint arXiv:2205.03481, 2022) further in view of Bhowmik et al. (US Pub 20200219515).
Regarding Claim 1,
Mansour discloses a non-transitory computer-readable storage medium storing one or more programs executable by one or more processors, the one or more programs comprising instructions (Mansour, Title, "Acoustic Echo Cancellation withy Loudspeaker Canceling Beamformer"; col.32:18-44, "…The device 110 may include one or more controllers/processors 1204... for processing data and computer-readable instructions, and a memory 1206...The computer instructions may be stored in a non-transitory manner in non-volatile memory 1206, storage 1208, or an external device...") for:
receiving, at an electronic device, multiple channels of audio data from a plurality of microphones (Fig.1, col.3:4-29, "…a device 110 configured to capture input audio data, generate loudspeaker canceling beam (LCB) audio data corresponding to a loud speaker of the device 110...the device 110 may include a microphone array 114 and one or more loudspeaker(s) 116...The device 110 may receive playback audio data...may capture input audio data using the microphone array 114..."),
wherein the multiple channels of audio data comprise speech from a user of the electronic device and speech from one or more other persons (col.3:18-22, "…In addition to capturing speech (e.g., the input audio data includes a representation of speech), the device 110 may capture a portion of the output audio generated by the loudspeaker(s) 116, which may be referred to as an "echo" or echo signal...");
receiving output audio data from one or more speakers, wherein the output audio data comprises speech generated using a text-to-speech technique (col.3:4-29, "…The device 110 may receive playback audio data and may generate output audio corresponding to the playback audio data using the one or more loudspeaker(s) 116...");
generating refined audio data by applying a multi-path acoustic echo cancellation (AEC) technique to the multiple channels of audio data using the output audio data from the one or more speakers as reference data (col.4:25-36, "…the device 110 may perform AEC processing based on a number of microphones included in the microphone array 114 (e.g., number of different output signals from the microphone array 114)..."; Fig.9B, col.24:61-col.25:34, "…To perform AEC processing with the modified playback audio data, FIG. 9B illustrates a multi-channel AEC (MC AEC) 930, which is capable of performing AEC for a single channel (e.g., one speaker) or for multiple channels (e.g., 5.1 surround sound)...the MC-AEC component 930 may perform AEC processing separately for each channel of the modified playback audio data 972...");
generating directional audio data by applying beamforming to the refined audio data (Fig.1, col.6:48-col.7:43, "…The device 110 may perform (126) acoustic echo cancellation to remove (e.g. subtract) the LCB audio data from the input audio data and generate modified input audio data...The device may then beamform (128) the modified input audio data into a plurality of beams ( e.g., perform a beamforming operation to generate beamformed audio data)...each beam may include audio data corresponding to a particular direction relative to the device 110..."; Fig.8A, col.18:43-19:37, "…microphone outputs 800 (e.g., input audio data captured by the microphone array 114) is input to one or more acoustic echo cancellation components (AECs) 810 and the AECs generate AEC outputs 815 by canceling an echo signal...After performing AEC to generate the AEC outputs 815, the AEC outputs 815 may be input to one or more fixed beamformer (FBF) units 820..."; col.22:1-5, "…In some examples, the AECs 810 may be positioned after the fixed beamformer (FBF) units 820 without departing from the disclosure.."),
wherein the directional audio data has more channels than the multiple channels of audio data (col.8:47-61, "…While the number of beams may correspond to the number of microphones, this need not be the case…the number of microphones may be more than, less than, or the same as the number of beams..."; Fig.8A, col.19:38-, lls., "…the microphone array 114 may include eight microphones whereas the device 110 may generate twelve beams...");
Mansour does not explicitly disclose output audio data generated using a text-to-speech technique. Panchapagesan, in the analogous field of acoustic echo cancellation training for ASR accuracy, discloses a text-to-speech playback as the specific reference content (Panchapagesan, Introduction, "…We consider the problem of recognizing queries spoken to a smart speaker device that is playing out audio such as text-to-speech
(TTS) responses, audiobooks or music..."; 4.1. Training Data, "…Real echoes are obtained...from an internal dataset collected for training Text-to-Speech (TTS) models on a Google Home device from multiple rooms. The latter was chosen since an important use case is to cancel TTS responses played by the device...").
Therefore, it would have been obvious to one of ordinary skill in the art, before effective filing date of the claimed invention, to have applied Panchapagesan's known TTS-playback context to Mansour's generic loudspeaker-reference AEC mechanism, with the reasonable expectation of yielding a device whose AEC reference signal corresponds to TTS-generated output.
Neither Mansour nor Panchapagesan discloses an automatic speech recognizer (ASR) that identifies the user or other people, or transcription process that excludes the user's speech.
Bhowmik, in the analogous field of wearable automatic transcription devices, teaches identifying, by inputting the directional audio data to an automatic speech recognizer (ASR), the speech from the user of the electronic device and the speech from the one or more other persons (Bhowmik, par [058], "…The system may be capable of "Own Voice" detection, where the system knows whether it was the user, the wearer of the ear wearable device, or someone else..."); and
generating a textual transcription for the speech from the one or more other persons, wherein the textual transcription does not include the speech from the user of the electronic device (Bhowmik, Fig.8, par [090], "…the system presents a transcript to the user including the text from three voice signals generated by three different speakers other than the user. The transcript shows a number 1, 2 or 3 for each speaker...The display also includes a map 1724 showing the relative locations of the speakers 1, 2 and 3 and the user represented with the letter U..."; par [086], "…the system can detect a user voice signal from a user wearing the ear-wearable device and process the input audio signal to exclude content of the user voice signal from the transcript...").
Therefore, it would have been obvious to a person of ordinary skill in the art to apply a speech recognition and transcription technique, which can exclude own-voice transcription of Bhowmik, to the refined and directed audio signal produced by AEC and Beamforming system of Mansour in view of Panchapagesan, with the reasonable expectation of success of yielding the predictable result of a device that identifies the wearer versus other speakers and transcribes only other speakers' speech using "own-voice" aware selective transcription.
Regarding Claim 6,
Mansour in view of Panchapagesan further in view of Bhowmik discloses the non-transitory computer-readable storage medium of claim 1, wherein the ASR comprises a trained AEC-aware model (Panchapagesan, Abstract, "…a neural AEC model operating on log-mel spectral features...can greatly improve Automatic Speech Recognition (ASR) accuracy when optimized with an auxiliary loss utilizing a pre-trained ASR model encoder...The model is trained by jointly optimizing Negative Scale-Invariant SNR (SISNR) and ASR losses on a large speech dataset (i.e., training routine)...").
Regarding Claim 7,
Mansour in view of Panchapagesan further in view of Bhowmik discloses the non-transitory computer-readable storage medium of claim 6, wherein the trained AEC-aware model is configured to differentiate between speech in the directional audio data and a residual echo from the multi-path AEC technique (Panchapagesan, Abstract, "…we find that cascading a linear adaptive AEC and a waveform-domain neural AEC is very effective..."; Fig.4: Cascade of Linear Adaptive AEC and Neural AEC; 4.5. Logmel-domain Neural AEC, "…this model predicts enhancement masks in the logmel domain, with ideal ratio masks as targets. The predicted masks are post-processed to trade-off residual noise and distortion…(i.e., differentiation process)").
Regarding Claim 9,
Mansour in view of Panchapagesan further in view of Bhowmik discloses the non-transitory computer-readable storage medium of claim 1, wherein the speech from the one or more other persons is in a first language and the textual transcription is in a second language (Bhowmik, par [047], "…The system may have the ability to translate the content of the audio signal to a language different than what was spoken. For example, the system may hear a voice signal in the French language and show content in English text...").
Regarding Claim 10,
Mansour in view of Panchapagesan further in view of Bhowmik discloses the non-transitory computer-readable storage medium of claim 1, wherein, for each portion of speech in the multiple channels of audio data: the ASR is configured to identify which person is speaking; and the textual transcription includes an indication of which person is speaking (Bhowmik, par [047], "…the system may differentiate different speakers or conversations in a complex environment based on direction or speaker recognition or cadence of conversation or any combination thereof...").
Regarding Claim 11,
Mansour in view of Panchapagesan further in view of Bhowmik discloses the non-transitory computer-readable storage medium of claim 1, wherein the speech from the user of the electronic device and the speech from one or more other persons correspond to conversation between the user and the one or more other persons (Bhowmik, par [046], "…the display device provides notice to a wearer, and provides a notice that a wearer can show to other participants in a conversation, that the verbal interaction will be recorded and transcribed..."; par [047], cadence of conversation).
Regarding Claim 12,
Mansour in view of Panchapagesan further in view of Bhowmik discloses the non-transitory computer-readable storage medium of claim 1, wherein the speech from the user of the electronic device comprises speech in a first language, and the speech from one or more other persons comprises speech in a second language (Bhowmik, par [047], "…The system may have the ability to translate the content of the audio signal to a language different than what was spoken. For example, the system may hear a voice signal in the French language and show content in English text...").
Regarding Claim 13,
Mansour in view of Panchapagesan further in view of Bhowmik discloses the non-transitory computer-readable storage medium of claim 1, wherein the multiple channels of audio data comprises a respective channel of audio data for each microphone in the plurality of microphones (Mansour, col.4:25-36, "…the device 110 may perform AEC processing based on a number of microphones included in the microphone array 114 (e.g., number of different output signals from the microphone array 114)...").
Regarding Claim 14,
Mansour in view of Panchapagesan further in view of Bhowmik discloses the non-transitory computer-readable storage medium of claim 1, wherein generating the directional audio data comprises splitting the multiple channels of audio data into a set number of audio channels corresponding to different regions of space around the electronic device (Mansour, col.5:63-67, "…at a first time the device 110 may perform a first beamforming operation to divide input audio data into 36 different portions, with each portion associated with a specific direction ( e.g., 10 degrees out of 360 degrees) relative to the device 110...").
Regarding Claim 16,
Mansour in view of Panchapagesan further in view of Bhowmik discloses the non-transitory computer-readable storage medium of claim 1, wherein the electronic device comprises a wearable device (Bhowmik, par [046], "…A transcript or notes from a verbal interaction can be delivered to the wearer of the ear-wearable devices, such as using an application running on the display device..."; par [056], "…the transcript could be shown on a smart glasses device, such as augmented reality glasses that show subtitles...").
Regarding Claim 17,
Mansour in view of Panchapagesan further in view of Bhowmik discloses the non-transitory computer-readable storage medium of claim 16, wherein the wearable device comprises an extended-reality headset (Bhowmik, par [056], "…the transcript could be shown on a smart glasses device, such as augmented reality glasses that show subtitles...").
Regarding Claim 18,
Mansour in view of Panchapagesan further in view of Bhowmik discloses the non-transitory computer-readable storage medium of claim 1, wherein the one or more programs further comprise instructions for presenting the textual transcription for speech from the one or more other persons on a display (Bhowmik, Fig.8, par [090], "…the system presents a transcript to the user including the text from three voice signals generated by three different speakers other than the user. The transcript shows a number 1, 2 or 3 for each speaker...The display also includes a map 1724 showing the relative locations of the speakers 1, 2 and 3 and the user represented with the letter U...").
Claim 19 is a method claim with limitations similar to the limitations of Claim 1 and is rejected under similar rationale. Rationale for combination is similar to that provided for Claim 1.
Claim 20 is an electronic device claim with limitations similar to the limitations of Claim 1 and is rejected under similar rationale. Rationale for combination is similar to that provided for Claim 1.
Claims 2-5 are rejected under 35 U.S.C. 103 as being unpatentable over Mansour in view of Panchapagesan further in view of Bhowmik further in view of Zhang et al. (US Pat 11,741,934).
Regarding Claim 2,
Mansour in view of Panchapagesan further in view of Bhowmik discloses the non-transitory computer-readable storage medium of claim 1 but does not explicitly disclose "wherein the multi-path AEC technique includes applying a linear filter to the multiple channels of audio data."
However, Zhang, in the analogous field of endeavor, discloses wherein the multi-path AEC technique includes applying a linear filter to the multiple channels of audio data (Zhang, Fig.6A, col.11:26-67, "…Components of an MB-AEC 420 to perform microphone based acoustic echo cancellation...adaptive filter 610...The adaptive filter 610 operates on audio signals corresponding to a same time/frame index...The adaptive filter 610 may utilize a short time Fourier transform-recursive least square (STFT-RLS) approach...The adaptive filter 610 may an STFT sub-band finite impulse response (FIR) filter with coefficients h..."; Fig.9, col15:44-16:45, "…signals 919B x2(n) through 319M xM(n), corresponding to multiple microphones 112B through 112M, may be used by the MB-AEC 420 to cancel from signal 417...The multiple microphone/multi-channel (Mc) recursive least square solution may be represented as...").
Therefore, it would have been obvious to a person of ordinary skill in the art to apply Zhang's known STFT-RLS adaptive filtering technique, to the generic multi-path AEC step of Mansour in view of Panchapagesan further in view of Bhowmik, with the reasonable expectation of success of yielding the predictable result of echo cancellation performed via STFT-domain, RLS-based, time-varying linear filtering to avoid distorting multi-channel audio signal because Zhang's technique performs the identical general function.
Regarding Claim 3,
Mansour in view of Panchapagesan further in view of Bhowmik further in view of Zhang discloses the non-transitory computer-readable storage medium of claim 2, wherein applying the linear filter comprises applies a short-time Fourier transform (STFT) to remove echoing from the multiple channels of audio data (Zhang, Fig.6A, col.11:26-67, "...The adaptive filter 610 may utilize a short time Fourier transform-recursive least square (STFT-RLS) approach...The adaptive filter 610 may an STFT sub-band finite impulse response (FIR) filter with coefficients h..."; Fig.9, col15:44-16:45, "…signals 919B x2(n) through 319M xM(n), corresponding to multiple microphones 112B through 112M, may be used by the MB-AEC 420 to cancel from signal 417...The multiple microphone/multi-channel (Mc) recursive least square solution may be represented as...").
Regarding Claim 4,
Mansour in view of Panchapagesan further in view of Bhowmik further in view of Zhang discloses the non-transitory computer-readable storage medium of claim 2, wherein applying the linear filter comprises applying a recursive least squares (RLS) algorithm to remove echoing from the multiple channels of audio data (Zhang, Fig.6A, col.11:26-67, "...The adaptive filter 610 may utilize a short time Fourier transform-recursive least square (STFT-RLS) approach..."; Fig.9, col15:44-16:45, "…signals 919B x2(n) through 319M xM(n), corresponding to multiple microphones 112B through 112M, may be used by the MB-AEC 420 to cancel from signal 417...The multiple microphone/multi-channel (Mc) recursive least square solution may be represented as...").
Regarding Claim 5,
Mansour in view of Panchapagesan further in view of Bhowmik further in view of Zhang discloses the non-transitory computer-readable storage medium of claim 2, wherein the linear filter comprises a single-time varying linear filter configured to prevent distortion of the multiple channels of audio data (Zhang, Fig.6A, 11:26-67, "…The adaptive filter 610 will continually adjust its coefficients to converge on values that will result in desired cancellation (i.e., time-varying)..."; Fig.5, col.11:2-19, "…cancellation filter coefficients determined for a noise context portion 510...maybe used for AEC applied to later audio, such as AEC performed on audio corresponding to the wakeword portion 520 and/or the voice query portion 530..."; col.13:42-67, "…By freezing the coefficients in this manner the device 110 may avoid having the speech impact the filter coefficients used for echo cancellation...thus avoiding any undesired cancellation of desired speech…(i.e., distortion prevention)").
Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Mansour in view of Panchapagesan further in view of Bhowmik further in view of He et al. ("Spatial attention for far-field speech recognition with deep beamforming neural networks." ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2020).
Regarding Claim 8,
Mansour in view of Panchapagesan further in view of Bhowmik discloses the non-transitory computer-readable storage medium of claim 1 but is silent on "the ASR is trained recognize speech in the directional audio data."
He, in the analogous field of endeavor, discloses wherein the ASR is trained recognize speech in the directional audio data (He, Fig.1, 2. Approach, "…We propose an end-to-end neural network for acoustic modeling from raw multi-channel signals with three components...Neural beamformer, which extracts speech features on multiple look directions...Attention-based pooling module...Back end, which predicts the sequence of target posteriors...The three components are trained jointly, so that the speech enhancement and feature extraction front-end are directly optimized to reduce the target classification error..."; i.e., the beamformer (front end) and the back end (ASR component) are trained together such that the ASR component is directly optimized against the directional audio data).
Therefore, it would have been obvious to a person of ordinary skill in the art to apply He's jointly trained beamformer and its backend ASR using a directional audio data to Mansour's directional beamformer output, with the reasonable expectation of success to predictably yield the ASR trained on directional audio for the system of Mansour in view of Panchapagesan further in view of Bhowmik.
Claim 15 is rejected under 35 U.S.C. 103 as being unpatentable over Mansour in view of Panchapagesan further in view of Bhowmik further in view of Fan (US Pub 2014/0270244).
Regarding Claim 15,
Mansour in view of Panchapagesan further in view of Bhowmik discloses the non-transitory computer-readable storage medium of claim 1 but does not explicitly teach microphones integrated at distinct positions, or beamforming accounting for relative positions of the microphones.
Fan, in the analogous field of eyewear-integrated, noise cancelling microphone array, discloses wherein microphones of the plurality of microphones are located at distinct locations on the electronic device, and wherein generating the directional audio data comprises accounting for relative positions of the microphones of the plurality of microphones (Fan, Fig.6, par [053], "…an array of microphones coupled to the eyeglasses frame, the array of microphones including at least a first microphone 602 and a second microphone 604, the first microphone coupled to the eyeglasses frame about a temple region...The second microphone is located diagonally across lens opening 606, although it can be positioned anywhere along the inner frame of the lens..."; par [067], "…microphone signal 1212 is inputted to a frequency response matching filter 1204. The frequency response matching filter 1204 adjusts gain, phase, and shapes the frequency response of the further microphone signal 1212...the frequency response matching filter 1204 can adjust the signal for the distance between the two microphones, such that an outputted reference signal 1232 representative of the further microphone signal 1212 can be processed with the main signal 1230...").
Therefore, it would have been obvious to a person of ordinary skill in the art to apply Fan's multi-microphone array integrated into an eyewear frame to a generic microphone array of Mansour in view of Panchapagesan further in view of Bhowmik, with the reasonable expectation of success to predictably yield the position-aware glasses feeding the existing pipeline, since the Fan's placement performs the identical geometry-aware capture functions in both contexts.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Govindaraju et al. (US Pat 11,792,570) discloses a method for processing audio data using two adaptive reference algorithm (ARA) paths in parallel is provided. First ARA processing performs noise cancellation using all microphones, while second ARA processing performs noise cancellation using only a portion of the microphones (Govindaraju, Abstract).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JANGWOEN LEE whose telephone number is (703)756-5597. The examiner can normally be reached Monday-Friday 8:00 am - 5:00 pm ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, BHAVESH MEHTA can be reached at (571)272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JANGWOEN LEE/ Examiner, Art Unit 2656
/BHAVESH M MEHTA/ Supervisory Patent Examiner, Art Unit 2656