Prosecution Insights
Last updated: October 02, 2026
Application No. 18/817,608

FRONT-END AUDIO PROCESSING FOR AUTOMATIC SPEECH RECOGNITION

Non-Final OA §103
Filed
Aug 28, 2024
Priority
Aug 28, 2023 — provisional 63/579,211
Examiner
LAM, PHILIP HUNG FAI
Art Unit
2656
Tech Center
2600 — Communications
Assignee
Shure Acquisition Holdings Inc.
OA Round
2 (Non-Final)
84%
Grant Probability
Favorable
2-3
OA Rounds
5m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 84% — above average
84%
Career Allowance Rate
130 granted / 155 resolved
+21.9% vs TC avg
Strong +51% interview lift
Without
With
+50.9%
Interview Lift
resolved cases with interview
Typical timeline
2y 6m
Avg Prosecution
29 currently pending
Career history
177
Total Applications
across all art units

Statute-Specific Performance

§101
24.1%
-15.9% vs TC avg
§103
54.7%
+14.7% vs TC avg
§102
10.5%
-29.5% vs TC avg
§112
4.5%
-35.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 155 resolved cases

Office Action

§103
DETAILED ACTION This office action is in response to Applicant’s Amended submission filed on 8/3/2026. Applicant has amended claims 20. Claims 1-20 are pending and have been examined. Response to Amendment and Arguments 35 U.S.C. 102/103 Rejections Applicant’s amendments and arguments are considered but are either unpersuasive or moot in view of the new grounds of rejection. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-5, 8, 10, and 16-20 are rejected under 35 U.S.C. 103 as being unpatentable over Wingate (US 20170243577), in view of Zhao (US 20260171104). Regarding Claim 1, Wingate discloses: 1. An apparatus comprising at least one processor and a memory storing instructions that are operable, when executed by the processor, to cause the apparatus to: (see fig. 3, which shows an intelligent microphone with DSP module (which includes processor and can execute instructions) and memory module.) receive audio data captured by one or more microphone array devices located within an audio environment; ([0037] The embedded intelligent microphone module 100 may receive sound waves from the audio source 110 or other audio sources (not shown) as analog audio signals at, for example, a microphone array component of the intelligent microphone module 100.) Also see fig. 5 flow chart. input the ([0088] At block 540, the digital audio signals may be processed. For example, the digital audio signals may be processed by DSP module 330 or source separation module 340. Processing may include beamforming, noise reduction, source separation, or any other suitable technique including the previously described audio processing techniques.) Also see fig. 5 flow chart. input the audio speech signal to an ASR model configured to generate textual data; ([0088) At block 550, processed audio signals may be transmitted for automated speech recognition. For example, the processed audio signals may be transmitted to a remote ASR service 140 as previously described.) Also see fig. 5 flow chart. and output the textual data to a post-processing system. ([0039] the digital audio signals (or the further enhanced digital audio signals) may optionally be communicated to ASR service 140 via connection 141. As previously described, the ASR service 140 may perform speech recognition or other value-added services such as executing search queries based on the recognized speech or other audio. In some embodiments, the ASR service 140 may record received digital audio signals for future processing or analysis. The ASR service 140 may communicate the text of the recognized speech or other information (e.g., search results) back to the device 120 via connection 141 or directly to a remote business service such as business service 150 via connection 145. In other embodiments, the ASR service 140 may be embedded within the device 120.) Also see para 0046, display result on UI. Wingate is silent on extract an audio feature set from the audio data; input the audio feature set to an audio source separation model Zhao in the related art discloses: extract an audio feature set from the audio data; ([0010] The method may include: extracting speech features of different pronunciation objects from an acquired speech information sequence to obtain a speech feature sequence,) input the audio feature set to an audio source separation model to generate an audio speech signal that is pre-processed for automatic speech recognition (ASR); ([0010] separating speech information output by the different pronunciation objects from the speech information sequence based on the speech mask information of the different pronunciation objects and the speech feature sequence; and inputting the speech information output by the different pronunciation objects into a speech recognition terminal, where the speech information is used to be recognized by the speech recognition terminal.) Wingate and Zhao are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Wingate to combine the teaching of Zhao for the above mentioned feature, because the present application provides a speech separation method to at least solve a technical problem of being unable to perform a speech separation on the speech (Zhao, [Background]). Regarding Claim 2, Wingate and Zhao discloses all the elements of claim 1, Wingate further discloses: wherein the instructions are further operable to cause the apparatus to: input the audio feature set to the ASR model to generate the textual data associated with the audio speech signal. ([0046] The display-based device 210 may be configured to display information related to the digital audio signals processed by intelligent microphone module 100. For example, intelligent microphone module 100 may receive speech input that an ASR service interprets as a query (e.g., “What is the weather today?”), and the display-based device 210 may be configured to display the text of the query (e.g., “What is the weather today?”) or the results of the query (e.g., 70 degrees Fahrenheit and sunny).) Also see para 0039 from above. Regarding Claim 3, Wingate and Zhao disclose all the element of claim 1, Wingate further discloses: wherein the instructions are further operable to cause the apparatus to: input the audio speech signal associated with the audio source separation model to an audio post-processing module configured to generate a filtered audio speech signal associated with the audio data. ([0102] Noise reduction module 604 can reduce noise in the one or more respective analog audio signals. The noise reduction module 602 can include an ambient noise removal module, which can be configured to reduce/remove ambient noise based on measured/artificial/estimated ambient noise from the audio signals. Filters and/or gain control can be implemented to reduce/remove ambient noise (e.g., a high pass filter, band pass filter, low pass filter, etc.) in the one or more respective audio signals. The noise reduction module 602 can include wind noise detector and/or removal module, which can be configured to indicate that wind noise is present and/or remove/reduce wind noise from the one or more respective audio signals. Filters and/or gain control can be implemented to modify the audio signals in the presence of wind noise. Coefficients of filters and/or gain control of the noise reduction module 602 can be tuned. These coefficients can affect the performance of the noise reduction module 604.) Also see para 0069 and 0108. Regarding Claim 4, Wingate and Zhao disclose all the element of claim 3, Wingate further discloses: wherein the instructions are further operable to cause the apparatus to: input the filtered audio speech signal to the ASR model to generate the textual data. ([0046] The display-based device 210 may be configured to display information related to the digital audio signals processed by intelligent microphone module 100. For example, intelligent microphone module 100 may receive speech input that an ASR service interprets as a query (e.g., “What is the weather today?”), and the display-based device 210 may be configured to display the text of the query (e.g., “What is the weather today?”) or the results of the query (e.g., 70 degrees Fahrenheit and sunny).) Also see para 0039 from above. Regarding Claim 5, Wingate and Zhao disclose all the element of claim 1, Wingate further discloses: wherein the post-processing system comprises a text post-processing system configured to enhance the textual data. ([0039] the digital audio signals (or the further enhanced digital audio signals) may optionally be communicated to ASR service 140 via connection 141. As previously described, the ASR service 140 may perform speech recognition or other value-added services such as executing search queries based on the recognized speech or other audio. In some embodiments, the ASR service 140 may record received digital audio signals for future processing or analysis. The ASR service 140 may communicate the text of the recognized speech or other information (e.g., search results) back to the device 120 via connection 141 or directly to a remote business service such as business service 150 via connection 145.) Regarding Claim 8, Wingate and Zhao disclose all the element of claim 1, Wingate further discloses: wherein the instructions are further operable to cause the apparatus to: optimize one or more of beamforming or beamsteering associated with the one or more microphone array devices based at least in part on ASR feedback data associated with the ASR model. ([0110] the integrated circuit includes or implements a feedback path 706 for the confidence level of the ASR module to control one or more parameters affecting a beam being formed by the beamforming module 602. For instance, a low confidence level can change the beam one or more ways to try and improve the performance of the ASR module 606. When the confidence level is high, which is an indication of a good quality audio signal, the size of beam can be made smaller to adaptively focus the beam towards the source as positive feedback for the beamforming module 602. When the confidence level is low, which is an indication of a bad quality audio signal, the feedback path 706 can request the beamforming module 602 to adapt the beam and/or initiate a search sequence for the source (e.g., changing the direction, increasing the size, changing the location of the beam). If the confidence level increases, the feedback path 706 can provide positive feedback that the beamforming module 602 has found the source (and possibly halt adaptive beam forming temporarily if the confidence level remains high). By improving the beam, the quality of the audio signal being processed by the ASR module 606 may ultimately improve, which in turn can increase the confidence level of the extracted speech information.) Regarding Claim 10, Wingate and Zhao disclose all the element of claim 1, Wingate further discloses: wherein the instructions are further operable to cause the apparatus to: input the audio speech signal associated with the audio source separation model to an audio post-processing module configured to generate a filtered audio speech signal associated with the audio data; ([0110] the integrated circuit includes or implements a feedback path 706 for the confidence level of the ASR module to control one or more parameters affecting a beam being formed by the beamforming module 602. For instance, a low confidence level can change the beam one or more ways to try and improve the performance of the ASR module 606. When the confidence level is high, which is an indication of a good quality audio signal, the size of beam can be made smaller to adaptively focus the beam towards the source as positive feedback for the beamforming module 602. When the confidence level is low, which is an indication of a bad quality audio signal, the feedback path 706 can request the beamforming module 602 to adapt the beam and/or initiate a search sequence for the source (e.g., changing the direction, increasing the size, changing the location of the beam). If the confidence level increases, the feedback path 706 can provide positive feedback that the beamforming module 602 has found the source (and possibly halt adaptive beam forming temporarily if the confidence level remains high). By improving the beam, the quality of the audio signal being processed by the ASR module 606 may ultimately improve, which in turn can increase the confidence level of the extracted speech information.) [a beam is a filtered signal - Also see para 0030, 0060, 0101 – beamforming is spatial, directional, filtering.] Also see para 0108. input the filtered audio speech signal to the ASR model to generate the textual data; ([0110] the integrated circuit includes or implements a feedback path 706 for the confidence level of the ASR module to control one or more parameters affecting a beam being formed by the beamforming module 602. For instance, a low confidence level can change the beam one or more ways to try and improve the performance of the ASR module 606. When the confidence level is high, which is an indication of a good quality audio signal, the size of beam can be made smaller to adaptively focus the beam towards the source as positive feedback for the beamforming module 602. When the confidence level is low, which is an indication of a bad quality audio signal, the feedback path 706 can request the beamforming module 602 to adapt the beam and/or initiate a search sequence for the source (e.g., changing the direction, increasing the size, changing the location of the beam). If the confidence level increases, the feedback path 706 can provide positive feedback that the beamforming module 602 has found the source (and possibly halt adaptive beam forming temporarily if the confidence level remains high). By improving the beam, the quality of the audio signal being processed by the ASR module 606 may ultimately improve, which in turn can increase the confidence level of the extracted speech information.) [a beam is a filtered signal] and adjust one or more parameters associated with the audio post-processing module based at least in part on ASR feedback data associated with the ASR model. ([0110] the integrated circuit includes or implements a feedback path 706 for the confidence level of the ASR module to control one or more parameters affecting a beam being formed by the beamforming module 602. For instance, a low confidence level can change the beam one or more ways to try and improve the performance of the ASR module 606. When the confidence level is high, which is an indication of a good quality audio signal, the size of beam can be made smaller to adaptively focus the beam towards the source as positive feedback for the beamforming module 602. When the confidence level is low, which is an indication of a bad quality audio signal, the feedback path 706 can request the beamforming module 602 to adapt the beam and/or initiate a search sequence for the source (e.g., changing the direction, increasing the size, changing the location of the beam). If the confidence level increases, the feedback path 706 can provide positive feedback that the beamforming module 602 has found the source (and possibly halt adaptive beam forming temporarily if the confidence level remains high). By improving the beam, the quality of the audio signal being processed by the ASR module 606 may ultimately improve, which in turn can increase the confidence level of the extracted speech information.) Claim 16 is a method claim that corresponds to claim 1, and the similar rationale applied in the rejection of claim 1 can also be applied. Claims 17-19 recites method claims that corresponds to the apparatus of claims 2-4 are therefore rejected under the same grounds as claims 2-4 above. Regarding claim 20, Wingate discloses: 20. A computer program product, stored on a computer readable medium, comprising instructions that, when executed by one or more processors of an apparatus, cause the one or more processors to: ([0060 The instruction (stored in non-transitory computer-readable memory in the intelligent microphone) may be configured to improve or enhance the digital audio signals to prepare the digital audio signals for further processing by other modules or by an external service, such as a remote (e.g., cloud-based) automated speech recognition (ASR) service.) As for the rest of the claim, they recite the elements of claim 1, therefor the rationale applied in rejection of claim 1 is also applicable to claim 20. Claims 6 and 7 are rejected under 35 U.S.C. 103 as being unpatentable over Wingate and Zhao, and further in view of Baeuml (US 20230074406). Regarding Claim 6, Wingate and Zhao disclose all the elements of Claim 1, However, Wingate and Zhao do not disclose: wherein the post-processing system comprises a large language model configured to generate one or more inferences with respect to the textual data. Baeuml (in the related field of using LLM in generating automated response) discloses: wherein the post-processing system comprises a large language model configured to generate one or more inferences with respect to the textual data. ([0043] Further, the LLM engine 150A1 and/or 150A2 can process the set of assistant outputs that are predicted to be responsive to the assistant query included in the spoken utterance captured in the stream of audio data processed by the ASR engine 130A1 and/or 130A2. As described herein (e.g., with respect to FIGS. 2-6), in some implementations, the LLM engine 150A1 and/or 150A2 can cause the set of assistant outputs to be modified, using one or more LLM outputs, to generate a set of modified assistant outputs.) Wingate, Zhao and Baeuml are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Wingate and Zhao to combine the teaching of Baeuml for the above-mentioned feature, because the LLM can interpret text accuracy for specified domain without the need to retrain the acoustic model and save the cost of creating a custom ASR model for various domains (Baeuml, [0043]). Regarding Claim 7, Wingate and Zhao disclose all the elements of Claim 1, However, Wingate and Zhao do not disclose: wherein the post-processing system comprises a user experience system configured to provide digital entertainment output. Baeuml discloses: wherein the post-processing system comprises a user experience system configured to provide digital entertainment output. ([0031] The client device 110 may be, for example, one or more of: a desktop computer, a laptop computer, a tablet, a mobile phone, a computing device of a vehicle (e.g., an in-vehicle communications system, an in-vehicle entertainment system, an in-vehicle navigation system), a standalone interactive speaker (optionally having a display), a smart appliance such as a smart television, and/or a wearable apparatus of the user that includes a computing device (e.g., a watch of the user having a computing device, glasses of the user having a computing device, a virtual or augmented reality computing device). Additional and/or alternative client devices may be provided.) Wingate, Zhao and Baeuml are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Wingate and Zhao to combine the teaching of Baeuml for the above-mentioned feature, because postprocessing can be used in cars and wearable device to deliver content (Baeuml, [0031]). Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Wingate and Zhao, and further in view of Watanabe (US 20180261225). Regarding Claim 9, Wingate and Zhao disclose all the elements of Claim 1, Wingate further discloses: wherein the instructions are further operable to cause the apparatus to: adjust one or more minimum variance distortionless response (data associated with the ASR model. ([0110] the integrated circuit includes or implements a feedback path 706 for the confidence level of the ASR module to control one or more parameters affecting a beam being formed by the beamforming module 602. For instance, a low confidence level can change the beam one or more ways to try and improve the performance of the ASR module 606. When the confidence level is high, which is an indication of a good quality audio signal, the size of beam can be made smaller to adaptively focus the beam towards the source as positive feedback for the beamforming module 602. When the confidence level is low, which is an indication of a bad quality audio signal, the feedback path 706 can request the beamforming module 602 to adapt the beam and/or initiate a search sequence for the source (e.g., changing the direction, increasing the size, changing the location of the beam). If the confidence level increases, the feedback path 706 can provide positive feedback that the beamforming module 602 has found the source (and possibly halt adaptive beam forming temporarily if the confidence level remains high). By improving the beam, the quality of the audio signal being processed by the ASR module 606 may ultimately improve, which in turn can increase the confidence level of the extracted speech information.) Although it can be said that Wingate teaches feedback path, and adapting the beam, and feedback path based on confidence of ASR, and parameters affecting beamformer like direction and size be adjusted. However, Wingate does not explicitly disclose: the beamformer being MVDR. Watanabe (in the related field of multichannel end to end speech recognition) discloses: adjust one or more minimum variance distortionless response (MVDR) coefficients associated with the one or more microphone array devices based at least in part on ASR feedback data associated with the ASR model. ([0059] In one embodiment, the network estimates the time-frequency masks, which are used to compute the time-invariant filter coefficients {g.sub.f,c}.sub.f=1,c=1.sup.F,C based on the MVDR formalizations. Also, mask-based beamforming approaches have achieved great performance in noisy speech recognition benchmarks. Therefore, one embodiment of the present invention uses a mask-based MVDR beamformer (mask-based MVDR beamformer network), where overall procedures are formalized as a differentiable network for the subsequent end-to-end speech recognition system.) Wingate, Zhao and Watanabe are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Wingate and Zhao to combine the teaching of Watanabe for the above mentioned feature, because mask based MVDR improves speech recognition by suppressing noise before reaching the ASR (Watanabe, [0059]). Claims 11-13 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Wingate and Zhao, and further in view of Applicant supplied reference, Hiroe (US 20160005394). Regarding Claim 11, Wingate and Zhao disclose all the elements of Claim 1, Wingate and Zhao do not explicitly disclose: wherein the instructions are further operable to cause the apparatus to: receive video data captured by the one or more microphone array devices or a video capture device located within the audio environment; Hiroe (in the same field of speech recognition) discloses: wherein the instructions are further operable to cause the apparatus to: ([0476] The voice recognition apparatus 150 in the embodiment performs tracking using a plurality of sound source direction information, i.e., sound source direction information acquired based on an analysis of the sound data acquired by the sound input unit 151 including the microphone array, and sound source direction information acquired based on the direction of the lip or the hand provided by the analysis of the acquired image by the image input unit 154.) extract a video feature set from the video data; ([0476] The voice recognition apparatus 150 in the embodiment performs tracking using a plurality of sound source direction information, i.e., sound source direction information acquired based on an analysis of the sound data acquired by the sound input unit 151 including the microphone array, and sound source direction information acquired based on the direction of the lip or the hand provided by the analysis of the acquired image by the image input unit 154.) and input the video feature set to the audio source separation model to generate the audio speech signal. ([0476] The voice recognition apparatus 150 in the embodiment performs tracking using a plurality of sound source direction information, i.e., sound source direction information acquired based on an analysis of the sound data acquired by the sound input unit 151 including the microphone array, and sound source direction information acquired based on the direction of the lip or the hand provided by the analysis of the acquired image by the image input unit 154.) [audio source separation model and generate audio speech signal already disclosed earlier in claim 1 by Wingate.] Wingate, Zhao and Hiroe are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Wingate and Zhao to combine the teaching of Hiroe for the above-mentioned feature, because the combination of audio source separation with visual cues may better isolate and/or track the speaker in noisy environment (Hiroe, [0476]). Regarding Claim 12, Wingate, Zhao and Hiroe disclose all the elements of Claim 11, Wingate further discloses: optimize one or more of beamforming or beamsteering associated with the one or more microphone array devices based at least in part on the ([0110] the integrated circuit includes or implements a feedback path 706 for the confidence level of the ASR module to control one or more parameters affecting a beam being formed by the beamforming module 602. For instance, a low confidence level can change the beam one or more ways to try and improve the performance of the ASR module 606. When the confidence level is high, which is an indication of a good quality audio signal, the size of beam can be made smaller to adaptively focus the beam towards the source as positive feedback for the beamforming module 602. When the confidence level is low, which is an indication of a bad quality audio signal, the feedback path 706 can request the beamforming module 602 to adapt the beam and/or initiate a search sequence for the source (e.g., changing the direction, increasing the size, changing the location of the beam). If the confidence level increases, the feedback path 706 can provide positive feedback that the beamforming module 602 has found the source (and possibly halt adaptive beam forming temporarily if the confidence level remains high). By improving the beam, the quality of the audio signal being processed by the ASR module 606 may ultimately improve, which in turn can increase the confidence level of the extracted speech information.) Wingate and Zhao do not explicitly disclose: wherein the instructions are further operable to cause the apparatus to: optimize one or more of beamforming or beamsteering associated with the one or more microphone array devices based at least in part on the video feature set. Hiroe further discloses: optimize one or more of beamforming or beamsteering associated with the one or more microphone array devices based at least in part on the video feature set. ([0476] The voice recognition apparatus 150 in the embodiment performs tracking using a plurality of sound source direction information, i.e., sound source direction information acquired based on an analysis of the sound data acquired by the sound input unit 151 including the microphone array, and sound source direction information acquired based on the direction of the lip or the hand provided by the analysis of the acquired image by the image input unit 154.) Where the rationale for the combination would be similar to the one already provided. Regarding Claim 13, Wingate and Zhao disclose all the elements of Claim 1, Wingate further discloses: wherein the audio feature set is a first audio feature set, and wherein the instructions are further operable to cause the apparatus to: receive one or more undesirable audio signals related to the audio environment; ([0088] FIG. 5 depicts a method 500 for audio processing using an intelligent microphone 100 in accordance with an embodiment of the present disclosure. At block 510, the method may begin. At block 520, analog audio signals may be received by one or more microphones, such as by microphone array module 310. At block 530, the analog audio signals may be converted to digital audio signals by one or more ADCs, such as by ADC module 320. At block 540, the digital audio signals may be processed. For example, the digital audio signals may be processed by DSP module 330 or source separation module 340. Processing may include beamforming, noise reduction, source separation, or any other suitable technique including the previously described audio processing techniques.) Wingate and Zhao does not explicitly disclose the following: extract a second audio feature set from the one or more undesirable audio signals; and input the second audio feature set to the audio source separation model to generate the audio speech signal. Hiroe discloses: extract a second audio feature set from the one or more undesirable audio signals; ([0221] Thus, by performing the voice segment detection corresponding to a plurality of the sound sources and the sound source extraction process at a former step of the voice recognition, even under the environments where the disturbing sound is present, there are a plurality of the target sounds for the voice recognition and both of which are overlapped and generated, it is possible to detect the individual target sounds and perform the voice recognition with high accuracy.) and input the second audio feature set to the audio source separation model to generate the audio speech signal. ([0221] Thus, by performing the voice segment detection corresponding to a plurality of the sound sources and the sound source extraction process at a former step of the voice recognition, even under the environments where the disturbing sound is present, there are a plurality of the target sounds for the voice recognition and both of which are overlapped and generated, it is possible to detect the individual target sounds and perform the voice recognition with high accuracy.) Wingate, Zhao and Hiroe are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Wingate to combine the teaching of Hiroe for the above-mentioned feature, because by analyzing separated noise separately, the model can subtract it from the target sound, leading to improvement of speech recognition (Hiroe, [0221]). Regarding Claim 15, Wingate/Zhao/Hiroe disclose all the elements of Claim 11, Wingate further discloses: wherein the instructions are further operable to cause the apparatus to: input the audio speech signal to an audio post-filter model configured to generate a processed audio speech signal that is further processed for the ASR; ([0110] the integrated circuit includes or implements a feedback path 706 for the confidence level of the ASR module to control one or more parameters affecting a beam being formed by the beamforming module 602. For instance, a low confidence level can change the beam one or more ways to try and improve the performance of the ASR module 606. When the confidence level is high, which is an indication of a good quality audio signal, the size of beam can be made smaller to adaptively focus the beam towards the source as positive feedback for the beamforming module 602. When the confidence level is low, which is an indication of a bad quality audio signal, the feedback path 706 can request the beamforming module 602 to adapt the beam and/or initiate a search sequence for the source (e.g., changing the direction, increasing the size, changing the location of the beam). If the confidence level increases, the feedback path 706 can provide positive feedback that the beamforming module 602 has found the source (and possibly halt adaptive beam forming temporarily if the confidence level remains high). By improving the beam, the quality of the audio signal being processed by the ASR module 606 may ultimately improve, which in turn can increase the confidence level of the extracted speech information.) [Beamforming module 602 combined with feedback path 706 can be interpreted as an audio post filter model as they perform speech improvement] and input the processed audio speech signal to the ASR model to generate the textual data. ([0110] the integrated circuit includes or implements a feedback path 706 for the confidence level of the ASR module to control one or more parameters affecting a beam being formed by the beamforming module 602. For instance, a low confidence level can change the beam one or more ways to try and improve the performance of the ASR module 606. When the confidence level is high, which is an indication of a good quality audio signal, the size of beam can be made smaller to adaptively focus the beam towards the source as positive feedback for the beamforming module 602. When the confidence level is low, which is an indication of a bad quality audio signal, the feedback path 706 can request the beamforming module 602 to adapt the beam and/or initiate a search sequence for the source (e.g., changing the direction, increasing the size, changing the location of the beam). If the confidence level increases, the feedback path 706 can provide positive feedback that the beamforming module 602 has found the source (and possibly halt adaptive beam forming temporarily if the confidence level remains high). By improving the beam, the quality of the audio signal being processed by the ASR module 606 may ultimately improve, which in turn can increase the confidence level of the extracted speech information.) Claim 14 is rejected under 35 U.S.C. 103 as being unpatentable over Wingate and Zhao, and further in view of Namazifar (US 20240428787). Regarding Claim 14, Wingate and Zhao disclose all the elements of Claim 1, Wingate and Zhao do not explicitly disclose: wherein the audio speech signal is associated with an audio embedding, a neural vocoder format, or a Residual Vector Quantization (RVQ) format. Namazifar (in the related field of speech processing) discloses: wherein the audio speech signal is associated with an ([0227] The vocoder 890 may convert the spectrogram data 845 generated by the TTS model 860 into an audio signal (e.g., an analog or digital time-domain waveform) suitable for amplification and output as audio. The vocoder 890 may be, for example, a universal neural vocoder based on Parallel WaveNet or related model. The vocoder 890 may take as input audio data in the form of, for example, a Mel-spectrogram with 80 coefficients and frequencies ranging from 50 Hz to 12 kHz. The synthesized speech audio data 895 may be a time-domain audio format (e.g., pulse-code modulation (PCM), waveform audio format (WAV), u-law, etc.) that may be readily converted to an analog signal for amplification and output by a loudspeaker.) [the claim only required one of the features recited] Wingate, Zhao and Namazifar are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Wingate and Zhao to combine the teaching of Namazifar for the above-mentioned feature, because neural vocoder can provide high quality audio generation (Namazifar, [0227]). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Cui (US 20240177717) – discloses method/system/device processing voice, where mixed voices are separately inputted into feature extraction layer in secondary processing model. See Abstract, and para 0122, 0130, 0137, and 0223 for additional details. Narayanan (US 20230038982) – discloses a method for ASR using joint acoustic echo cancellation, speech enhancement, and voice separation. See Abstract, para 0003, 0008 and 0028 and fig. 1 for additional details. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Philip H Lam whose telephone number is (571)272-1721. The examiner can normally be reached 9 AM-3 PM Pacific time. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bhavesh Mehta can be reached on 571-272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /PHILIP H LAM/ Examiner, Art Unit 2656
Read full office action

Prosecution Timeline

Aug 28, 2024
Application Filed
May 01, 2026
Non-Final Rejection mailed — §103
Jul 17, 2026
Examiner Interview Summary
Jul 17, 2026
Applicant Interview (Telephonic)
Aug 03, 2026
Response Filed
Sep 10, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12688847
ERROR-CORRECTION AND EXTRACTION IN REQUEST DIALOGS
4y 1m to grant Granted Jul 21, 2026
Patent 12682164
CHAT SUPPORT PLATFORM HAVING AUTOMATIC KEYWORD CORRECTION
3y 3m to grant Granted Jul 14, 2026
Patent 12670519
CONTENT RECOMMENDATION USING RETRIEVAL AUGMENTED ARTIFICIAL INTELLIGENCE
3y 2m to grant Granted Jun 30, 2026
Patent 12657395
METHODS AND SYSTEMS FOR AVOIDING OFFENSIVE LANGUAGE BASED ON PERSONAS
2y 9m to grant Granted Jun 16, 2026
Patent 12639529
ENHANCING LARGE LANGUAGE MODELS USING IN-CONTEXT LEARNING AND ONLINE KNOWLEDGE
2y 6m to grant Granted May 26, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

2-3
Expected OA Rounds
84%
Grant Probability
99%
With Interview (+50.9%)
2y 6m (~5m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 155 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month