Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
Introduction
This office action is in response to Applicant’s submission filed on 8/28/2024. As such, claims 1-20 have been examined.
Examiner Comment Regarding Patent Subject Matter Eligibility under 35 U.S.C. 101
Independent claims 1, and 16 involve an apparatus/method focus on technical solution for speech enhancement involving sound source separation is a process that could not practically be performed as an abstract idea such as a mental process under the broadest reasonable interpretation (BRI). Accordingly, the independent claims and their dependents by virtue of their dependency, are directed towards patent eligible subject matter under step 2A prong 1. Although Claim 20 recite similar elements, however claim 20 recites computer readable medium which could involve transitory waveform or signal per se, see rejection below for further details.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim 20 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claim(s) does/do not fall within at least one of the four categories of patent eligible subject matter because they recite “computer readable medium …” where the computer readable medium is not claimed to be non-transitory. Therefore, the broadest reasonable interpretation of “computer readable medium” includes signals per se, rendering claim 20 subject matter ineligible. It is noted that Applicant’s statement in para. 0081 of the originally filed specification stating “Alternatively, or in addition, the program instructions may be encoded on an artificially-generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode information/data for transmission to suitable receiver apparatus for execution by an information/data processing apparatus. A computer-readable storage medium may be, or be included in, a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them. Moreover, while a computer-readable storage medium is not a propagated signal, a computer-readable storage medium may be a source or destination of computer program instructions encoded in an artificially-generated propagated signal. The computer-readable storage medium may also be, or be included in, one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices).” is not considered to be a “special definition or disavowal” as it is merely a statement directed to claim construction itself and thus does not clearly set forth a special definition of the claim term that differs from the plain and ordinary meaning it would otherwise possess. See MPEP 2111.01(IV). Please note that the spec only mention computer-readable storage medium is not propagated signal, but is silent on computer readable medium.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 1-5, 8, 10, and 16-20 are rejected under 35 U.S.C. 102 (a)(1) and (a)(2) as being anticipated by Applicant supplied reference (US PG Publication version), Wingate US 20170243577.
Regarding Claim 1, Wingate discloses: 1. An apparatus comprising at least one processor and a memory storing instructions that are operable, when executed by the processor, to cause the apparatus to: (see fig. 3, which shows an intelligent microphone with DSP module (which includes processor and can execute instructions) and memory module.)
receive audio data captured by one or more microphone array devices located within an audio environment; ([0037] The embedded intelligent microphone module 100 may receive sound waves from the audio source 110 or other audio sources (not shown) as analog audio signals at, for example, a microphone array component of the intelligent microphone module 100.) Also see fig. 5 flow chart.
extract an audio feature set from the audio data; ([0088] At block 540, the digital audio signals may be processed. For example, the digital audio signals may be processed by DSP module 330 or source separation module 340. Processing may include beamforming, noise reduction, source separation, or any other suitable technique including the previously described audio processing techniques.) Also see fig. 5 flow chart.
input the audio feature set to an audio source separation model to generate an audio speech signal that is pre-processed for automatic speech recognition (ASR); ([0088] At block 540, the digital audio signals may be processed. For example, the digital audio signals may be processed by DSP module 330 or source separation module 340. Processing may include beamforming, noise reduction, source separation, or any other suitable technique including the previously described audio processing techniques.) Also see fig. 5 flow chart.
input the audio speech signal to an ASR model configured to generate textual data; ([0088) At block 550, processed audio signals may be transmitted for automated speech recognition. For example, the processed audio signals may be transmitted to a remote ASR service 140 as previously described.) Also see fig. 5 flow chart.
and output the textual data to a post-processing system. ([0039] the digital audio signals (or the further enhanced digital audio signals) may optionally be communicated to ASR service 140 via connection 141. As previously described, the ASR service 140 may perform speech recognition or other value-added services such as executing search queries based on the recognized speech or other audio. In some embodiments, the ASR service 140 may record received digital audio signals for future processing or analysis. The ASR service 140 may communicate the text of the recognized speech or other information (e.g., search results) back to the device 120 via connection 141 or directly to a remote business service such as business service 150 via connection 145. In other embodiments, the ASR service 140 may be embedded within the device 120.) Also see para 0046, display result on UI.
Regarding Claim 2, Wingate discloses all the elements of claim 1,
Wingate further discloses: wherein the instructions are further operable to cause the apparatus to: input the audio feature set to the ASR model to generate the textual data associated with the audio speech signal. ([0046] The display-based device 210 may be configured to display information related to the digital audio signals processed by intelligent microphone module 100. For example, intelligent microphone module 100 may receive speech input that an ASR service interprets as a query (e.g., “What is the weather today?”), and the display-based device 210 may be configured to display the text of the query (e.g., “What is the weather today?”) or the results of the query (e.g., 70 degrees Fahrenheit and sunny).) Also see para 0039 from above.
Regarding Claim 3, Wingate disclose all the element of claim 1,
Wingate further discloses: wherein the instructions are further operable to cause the apparatus to: input the audio speech signal associated with the audio source separation model to an audio post-processing module configured to generate a filtered audio speech signal associated with the audio data. ([0102] Noise reduction module 604 can reduce noise in the one or more respective analog audio signals. The noise reduction module 602 can include an ambient noise removal module, which can be configured to reduce/remove ambient noise based on measured/artificial/estimated ambient noise from the audio signals. Filters and/or gain control can be implemented to reduce/remove ambient noise (e.g., a high pass filter, band pass filter, low pass filter, etc.) in the one or more respective audio signals. The noise reduction module 602 can include wind noise detector and/or removal module, which can be configured to indicate that wind noise is present and/or remove/reduce wind noise from the one or more respective audio signals. Filters and/or gain control can be implemented to modify the audio signals in the presence of wind noise. Coefficients of filters and/or gain control of the noise reduction module 602 can be tuned. These coefficients can affect the performance of the noise reduction module 604.) Also see para 0069 and 0108.
Regarding Claim 4, Wingate disclose all the element of claim 3,
Wingate further discloses: wherein the instructions are further operable to cause the apparatus to: input the filtered audio speech signal to the ASR model to generate the textual data. ([0046] The display-based device 210 may be configured to display information related to the digital audio signals processed by intelligent microphone module 100. For example, intelligent microphone module 100 may receive speech input that an ASR service interprets as a query (e.g., “What is the weather today?”), and the display-based device 210 may be configured to display the text of the query (e.g., “What is the weather today?”) or the results of the query (e.g., 70 degrees Fahrenheit and sunny).) Also see para 0039 from above.
Regarding Claim 5, Wingate disclose all the element of claim 1,
Wingate further discloses: wherein the post-processing system comprises a text post-processing system configured to enhance the textual data. ([0039] the digital audio signals (or the further enhanced digital audio signals) may optionally be communicated to ASR service 140 via connection 141. As previously described, the ASR service 140 may perform speech recognition or other value-added services such as executing search queries based on the recognized speech or other audio. In some embodiments, the ASR service 140 may record received digital audio signals for future processing or analysis. The ASR service 140 may communicate the text of the recognized speech or other information (e.g., search results) back to the device 120 via connection 141 or directly to a remote business service such as business service 150 via connection 145.)
Regarding Claim 8, Wingate disclose all the element of claim 1,
Wingate further discloses: wherein the instructions are further operable to cause the apparatus to: optimize one or more of beamforming or beamsteering associated with the one or more microphone array devices based at least in part on ASR feedback data associated with the ASR model. ([0110] the integrated circuit includes or implements a feedback path 706 for the confidence level of the ASR module to control one or more parameters affecting a beam being formed by the beamforming module 602. For instance, a low confidence level can change the beam one or more ways to try and improve the performance of the ASR module 606. When the confidence level is high, which is an indication of a good quality audio signal, the size of beam can be made smaller to adaptively focus the beam towards the source as positive feedback for the beamforming module 602. When the confidence level is low, which is an indication of a bad quality audio signal, the feedback path 706 can request the beamforming module 602 to adapt the beam and/or initiate a search sequence for the source (e.g., changing the direction, increasing the size, changing the location of the beam). If the confidence level increases, the feedback path 706 can provide positive feedback that the beamforming module 602 has found the source (and possibly halt adaptive beam forming temporarily if the confidence level remains high). By improving the beam, the quality of the audio signal being processed by the ASR module 606 may ultimately improve, which in turn can increase the confidence level of the extracted speech information.)
Regarding Claim 10, Wingate disclose all the element of claim 1,
Wingate further discloses: wherein the instructions are further operable to cause the apparatus to: input the audio speech signal associated with the audio source separation model to an audio post-processing module configured to generate a filtered audio speech signal associated with the audio data; ([0110] the integrated circuit includes or implements a feedback path 706 for the confidence level of the ASR module to control one or more parameters affecting a beam being formed by the beamforming module 602. For instance, a low confidence level can change the beam one or more ways to try and improve the performance of the ASR module 606. When the confidence level is high, which is an indication of a good quality audio signal, the size of beam can be made smaller to adaptively focus the beam towards the source as positive feedback for the beamforming module 602. When the confidence level is low, which is an indication of a bad quality audio signal, the feedback path 706 can request the beamforming module 602 to adapt the beam and/or initiate a search sequence for the source (e.g., changing the direction, increasing the size, changing the location of the beam). If the confidence level increases, the feedback path 706 can provide positive feedback that the beamforming module 602 has found the source (and possibly halt adaptive beam forming temporarily if the confidence level remains high). By improving the beam, the quality of the audio signal being processed by the ASR module 606 may ultimately improve, which in turn can increase the confidence level of the extracted speech information.) [a beam is a filtered signal - Also see para 0030, 0060, 0101 – beamforming is spatial, directional, filtering.] Also see para 0108.
input the filtered audio speech signal to the ASR model to generate the textual data; ([0110] the integrated circuit includes or implements a feedback path 706 for the confidence level of the ASR module to control one or more parameters affecting a beam being formed by the beamforming module 602. For instance, a low confidence level can change the beam one or more ways to try and improve the performance of the ASR module 606. When the confidence level is high, which is an indication of a good quality audio signal, the size of beam can be made smaller to adaptively focus the beam towards the source as positive feedback for the beamforming module 602. When the confidence level is low, which is an indication of a bad quality audio signal, the feedback path 706 can request the beamforming module 602 to adapt the beam and/or initiate a search sequence for the source (e.g., changing the direction, increasing the size, changing the location of the beam). If the confidence level increases, the feedback path 706 can provide positive feedback that the beamforming module 602 has found the source (and possibly halt adaptive beam forming temporarily if the confidence level remains high). By improving the beam, the quality of the audio signal being processed by the ASR module 606 may ultimately improve, which in turn can increase the confidence level of the extracted speech information.) [a beam is a filtered signal]
and adjust one or more parameters associated with the audio post-processing module based at least in part on ASR feedback data associated with the ASR model. ([0110] the integrated circuit includes or implements a feedback path 706 for the confidence level of the ASR module to control one or more parameters affecting a beam being formed by the beamforming module 602. For instance, a low confidence level can change the beam one or more ways to try and improve the performance of the ASR module 606. When the confidence level is high, which is an indication of a good quality audio signal, the size of beam can be made smaller to adaptively focus the beam towards the source as positive feedback for the beamforming module 602. When the confidence level is low, which is an indication of a bad quality audio signal, the feedback path 706 can request the beamforming module 602 to adapt the beam and/or initiate a search sequence for the source (e.g., changing the direction, increasing the size, changing the location of the beam). If the confidence level increases, the feedback path 706 can provide positive feedback that the beamforming module 602 has found the source (and possibly halt adaptive beam forming temporarily if the confidence level remains high). By improving the beam, the quality of the audio signal being processed by the ASR module 606 may ultimately improve, which in turn can increase the confidence level of the extracted speech information.)
Claim 16 is a method claim that corresponds to claim 1, and the similar rationale applied in the rejection of claim 1 can also be applied.
Claims 17-19 recites method claims that corresponds to the apparatus of claims 2-4 are therefore rejected under the same grounds as claims 2-4 above.
Regarding claim 20, Wingate discloses: 20. A computer program product, stored on a computer readable medium, comprising instructions that, when executed by one or more processors of an apparatus, cause the one or more processors to: ([0060 The instruction (stored in non-transitory computer-readable memory in the intelligent microphone) may be configured to improve or enhance the digital audio signals to prepare the digital audio signals for further processing by other modules or by an external service, such as a remote (e.g., cloud-based) automated speech recognition (ASR) service.)
As for the rest of the claim, they recite the elements of claim 1, therefor the rationale applied in rejection of claim 1 is also applicable to claim 20.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 6 and 7 are rejected under 35 U.S.C. 103 as being unpatentable over Wingate, in view of Baeuml (US 20230074406).
Regarding Claim 6, Wingate disclose all the elements of Claim 1,
However, Wingate does not disclose: wherein the post-processing system comprises a large language model configured to generate one or more inferences with respect to the textual data.
Baeuml (in the related field of using LLM in generating automated response) discloses: wherein the post-processing system comprises a large language model configured to generate one or more inferences with respect to the textual data. ([0043] Further, the LLM engine 150A1 and/or 150A2 can process the set of assistant outputs that are predicted to be responsive to the assistant query included in the spoken utterance captured in the stream of audio data processed by the ASR engine 130A1 and/or 130A2. As described herein (e.g., with respect to FIGS. 2-6), in some implementations, the LLM engine 150A1 and/or 150A2 can cause the set of assistant outputs to be modified, using one or more LLM outputs, to generate a set of modified assistant outputs.)
Wingate and Baeuml are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Wingate to combine the teaching of Baeuml for the above-mentioned feature, because the LLM can interpret text accuracy for specified domain without the need to retrain the acoustic model and save the cost of creating a custom ASR model for various domains (Baeuml, [0043]).
Regarding Claim 7, Wingate disclose all the elements of Claim 1,
However, Wingate does not disclose: wherein the post-processing system comprises a user experience system configured to provide digital entertainment output.
Baeuml discloses: wherein the post-processing system comprises a user experience system configured to provide digital entertainment output. ([0031] The client device 110 may be, for example, one or more of: a desktop computer, a laptop computer, a tablet, a mobile phone, a computing device of a vehicle (e.g., an in-vehicle communications system, an in-vehicle entertainment system, an in-vehicle navigation system), a standalone interactive speaker (optionally having a display), a smart appliance such as a smart television, and/or a wearable apparatus of the user that includes a computing device (e.g., a watch of the user having a computing device, glasses of the user having a computing device, a virtual or augmented reality computing device). Additional and/or alternative client devices may be provided.)
Wingate and Baeuml are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Wingate to combine the teaching of Baeuml for the above-mentioned feature, because postprocessing can be used in cars and wearable device to deliver content (Baeuml, [0031]).
Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Wingate, in view of Watanabe (US 20180261225).
Regarding Claim 9, Wingate disclose all the elements of Claim 1,
Wingate further discloses: wherein the instructions are further operable to cause the apparatus to: adjust one or more minimum variance distortionless response (([0110] the integrated circuit includes or implements a feedback path 706 for the confidence level of the ASR module to control one or more parameters affecting a beam being formed by the beamforming module 602. For instance, a low confidence level can change the beam one or more ways to try and improve the performance of the ASR module 606. When the confidence level is high, which is an indication of a good quality audio signal, the size of beam can be made smaller to adaptively focus the beam towards the source as positive feedback for the beamforming module 602. When the confidence level is low, which is an indication of a bad quality audio signal, the feedback path 706 can request the beamforming module 602 to adapt the beam and/or initiate a search sequence for the source (e.g., changing the direction, increasing the size, changing the location of the beam). If the confidence level increases, the feedback path 706 can provide positive feedback that the beamforming module 602 has found the source (and possibly halt adaptive beam forming temporarily if the confidence level remains high). By improving the beam, the quality of the audio signal being processed by the ASR module 606 may ultimately improve, which in turn can increase the confidence level of the extracted speech information.)
Although it can be said that Wingate teaches feedback path, and adapting the beam, and feedback path based on confidence of ASR, and parameters affecting beamformer like direction and size be adjusted. However, Wingate does not explicitly disclose: the beamformer being MVDR.
Watanabe (in the related field of multichannel end to end speech recognition) discloses: adjust one or more minimum variance distortionless response (MVDR) coefficients associated with the one or more microphone array devices based at least in part on ASR feedback data associated with the ASR model. ([0059] In one embodiment, the network estimates the time-frequency masks, which are used to compute the time-invariant filter coefficients {g.sub.f,c}.sub.f=1,c=1.sup.F,C based on the MVDR formalizations. Also, mask-based beamforming approaches have achieved great performance in noisy speech recognition benchmarks. Therefore, one embodiment of the present invention uses a mask-based MVDR beamformer (mask-based MVDR beamformer network), where overall procedures are formalized as a differentiable network for the subsequent end-to-end speech recognition system.)
Wingate and Watanabe are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Wingate to combine the teaching of Watanabe for the above mentioned feature, because mask based MVDR improves speech recognition by suppressing noise before reaching the ASR (Watanabe, [0059]).
Claims 11-13 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Wingate, in view of Applicant supplied reference, Hiroe (US 20160005394).
Regarding Claim 11, Wingate disclose all the elements of Claim 1,
Wingate does not explicitly disclose: wherein the instructions are further operable to cause the apparatus to: receive video data captured by the one or more microphone array devices or a video capture device located within the audio environment;
Hiroe (in the same field of speech recognition) discloses: wherein the instructions are further operable to cause the apparatus to: ([0476] The voice recognition apparatus 150 in the embodiment performs tracking using a plurality of sound source direction information, i.e., sound source direction information acquired based on an analysis of the sound data acquired by the sound input unit 151 including the microphone array, and sound source direction information acquired based on the direction of the lip or the hand provided by the analysis of the acquired image by the image input unit 154.)
extract a video feature set from the video data; ([0476] The voice recognition apparatus 150 in the embodiment performs tracking using a plurality of sound source direction information, i.e., sound source direction information acquired based on an analysis of the sound data acquired by the sound input unit 151 including the microphone array, and sound source direction information acquired based on the direction of the lip or the hand provided by the analysis of the acquired image by the image input unit 154.)
and input the video feature set to the audio source separation model to generate the audio speech signal. ([0476] The voice recognition apparatus 150 in the embodiment performs tracking using a plurality of sound source direction information, i.e., sound source direction information acquired based on an analysis of the sound data acquired by the sound input unit 151 including the microphone array, and sound source direction information acquired based on the direction of the lip or the hand provided by the analysis of the acquired image by the image input unit 154.) [audio source separation model and generate audio speech signal already disclosed earlier in claim 1 by Wingate.]
Wingate and Hiroe are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Wingate to combine the teaching of Hiroe for the above-mentioned feature, because the combination of audio source separation with visual cues may better isolate and/or track the speaker in noisy environment (Hiroe, [0476]).
Regarding Claim 12, Wingate and Hiroe disclose all the elements of Claim 11,
Wingate further discloses: optimize one or more of beamforming or beamsteering associated with the one or more microphone array devices based at least in part on the ([0110] the integrated circuit includes or implements a feedback path 706 for the confidence level of the ASR module to control one or more parameters affecting a beam being formed by the beamforming module 602. For instance, a low confidence level can change the beam one or more ways to try and improve the performance of the ASR module 606. When the confidence level is high, which is an indication of a good quality audio signal, the size of beam can be made smaller to adaptively focus the beam towards the source as positive feedback for the beamforming module 602. When the confidence level is low, which is an indication of a bad quality audio signal, the feedback path 706 can request the beamforming module 602 to adapt the beam and/or initiate a search sequence for the source (e.g., changing the direction, increasing the size, changing the location of the beam). If the confidence level increases, the feedback path 706 can provide positive feedback that the beamforming module 602 has found the source (and possibly halt adaptive beam forming temporarily if the confidence level remains high). By improving the beam, the quality of the audio signal being processed by the ASR module 606 may ultimately improve, which in turn can increase the confidence level of the extracted speech information.)
Wingate does not explicitly disclose: wherein the instructions are further operable to cause the apparatus to: optimize one or more of beamforming or beamsteering associated with the one or more microphone array devices based at least in part on the video feature set.
Hiroe further discloses: optimize one or more of beamforming or beamsteering associated with the one or more microphone array devices based at least in part on the video feature set. ([0476] The voice recognition apparatus 150 in the embodiment performs tracking using a plurality of sound source direction information, i.e., sound source direction information acquired based on an analysis of the sound data acquired by the sound input unit 151 including the microphone array, and sound source direction information acquired based on the direction of the lip or the hand provided by the analysis of the acquired image by the image input unit 154.)
Where the rationale for the combination would be similar to the one already provided.
Regarding Claim 13, Wingate disclose all the elements of Claim 1,
Wingate further discloses: wherein the audio feature set is a first audio feature set, and wherein the instructions are further operable to cause the apparatus to: receive one or more undesirable audio signals related to the audio environment; ([0088] FIG. 5 depicts a method 500 for audio processing using an intelligent microphone 100 in accordance with an embodiment of the present disclosure. At block 510, the method may begin. At block 520, analog audio signals may be received by one or more microphones, such as by microphone array module 310. At block 530, the analog audio signals may be converted to digital audio signals by one or more ADCs, such as by ADC module 320. At block 540, the digital audio signals may be processed. For example, the digital audio signals may be processed by DSP module 330 or source separation module 340. Processing may include beamforming, noise reduction, source separation, or any other suitable technique including the previously described audio processing techniques.)
Wingate does not explicitly disclose the following: extract a second audio feature set from the one or more undesirable audio signals; and input the second audio feature set to the audio source separation model to generate the audio speech signal.
Hiroe discloses: extract a second audio feature set from the one or more undesirable audio signals; ([0221] Thus, by performing the voice segment detection corresponding to a plurality of the sound sources and the sound source extraction process at a former step of the voice recognition, even under the environments where the disturbing sound is present, there are a plurality of the target sounds for the voice recognition and both of which are overlapped and generated, it is possible to detect the individual target sounds and perform the voice recognition with high accuracy.)
and input the second audio feature set to the audio source separation model to generate the audio speech signal. ([0221] Thus, by performing the voice segment detection corresponding to a plurality of the sound sources and the sound source extraction process at a former step of the voice recognition, even under the environments where the disturbing sound is present, there are a plurality of the target sounds for the voice recognition and both of which are overlapped and generated, it is possible to detect the individual target sounds and perform the voice recognition with high accuracy.)
Wingate and Hiroe are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Wingate to combine the teaching of Hiroe for the above-mentioned feature, because by analyzing separated noise separately, the model can subtract it from the target sound, leading to improvement of speech recognition (Hiroe, [0221]).
Regarding Claim 15, Wingate/Hiroe disclose all the elements of Claim 11,
Wingate further discloses: wherein the instructions are further operable to cause the apparatus to: input the audio speech signal to an audio post-filter model configured to generate a processed audio speech signal that is further processed for the ASR; ([0110] the integrated circuit includes or implements a feedback path 706 for the confidence level of the ASR module to control one or more parameters affecting a beam being formed by the beamforming module 602. For instance, a low confidence level can change the beam one or more ways to try and improve the performance of the ASR module 606. When the confidence level is high, which is an indication of a good quality audio signal, the size of beam can be made smaller to adaptively focus the beam towards the source as positive feedback for the beamforming module 602. When the confidence level is low, which is an indication of a bad quality audio signal, the feedback path 706 can request the beamforming module 602 to adapt the beam and/or initiate a search sequence for the source (e.g., changing the direction, increasing the size, changing the location of the beam). If the confidence level increases, the feedback path 706 can provide positive feedback that the beamforming module 602 has found the source (and possibly halt adaptive beam forming temporarily if the confidence level remains high). By improving the beam, the quality of the audio signal being processed by the ASR module 606 may ultimately improve, which in turn can increase the confidence level of the extracted speech information.) [Beamforming module 602 combined with feedback path 706 can be interpreted as an audio post filter model as they perform speech improvement]
and input the processed audio speech signal to the ASR model to generate the textual data. ([0110] the integrated circuit includes or implements a feedback path 706 for the confidence level of the ASR module to control one or more parameters affecting a beam being formed by the beamforming module 602. For instance, a low confidence level can change the beam one or more ways to try and improve the performance of the ASR module 606. When the confidence level is high, which is an indication of a good quality audio signal, the size of beam can be made smaller to adaptively focus the beam towards the source as positive feedback for the beamforming module 602. When the confidence level is low, which is an indication of a bad quality audio signal, the feedback path 706 can request the beamforming module 602 to adapt the beam and/or initiate a search sequence for the source (e.g., changing the direction, increasing the size, changing the location of the beam). If the confidence level increases, the feedback path 706 can provide positive feedback that the beamforming module 602 has found the source (and possibly halt adaptive beam forming temporarily if the confidence level remains high). By improving the beam, the quality of the audio signal being processed by the ASR module 606 may ultimately improve, which in turn can increase the confidence level of the extracted speech information.)
Claim 14 is rejected under 35 U.S.C. 103 as being unpatentable over Wingate, in view of Namazifar (US 20240428787).
Regarding Claim 14, Wingate disclose all the elements of Claim 1,
Wingate does not explicitly disclose: wherein the audio speech signal is associated with an audio embedding, a neural vocoder format, or a Residual Vector Quantization (RVQ) format.
Namazifar (in the related field of speech processing) discloses: wherein the audio speech signal is associated with an ([0227] The vocoder 890 may convert the spectrogram data 845 generated by the TTS model 860 into an audio signal (e.g., an analog or digital time-domain waveform) suitable for amplification and output as audio. The vocoder 890 may be, for example, a universal neural vocoder based on Parallel WaveNet or related model. The vocoder 890 may take as input audio data in the form of, for example, a Mel-spectrogram with 80 coefficients and frequencies ranging from 50 Hz to 12 kHz. The synthesized speech audio data 895 may be a time-domain audio format (e.g., pulse-code modulation (PCM), waveform audio format (WAV), u-law, etc.) that may be readily converted to an analog signal for amplification and output by a loudspeaker.) [the claim only required one of the features recited]
Wingate and Namazifar are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Wingate to combine the teaching of Namazifar for the above-mentioned feature, because neural vocoder can provide high quality audio generation (Namazifar, [0227]).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Narayanan (US 20220417659) – discloses method/system/device for audio correction, it teaches corrective feedback for speech recognition correction, and separating background noise from target speech. See para 0082 for additional details.
Higuchi, T., Ito, N., Yoshioka, T., & Nakatani, T. (2016, March). Robust MVDR beamforming using time-frequency masks for online/offline ASR in noise. In 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 5210-5214). IEEE. – discloses MVDR beamforming to enhance speech signal, that use steering vector estimation method based on time frequency mask. See Abstract and section 3-5 for additional details.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Philip H Lam whose telephone number is (571)272-1721. The examiner can normally be reached 9 AM-3 PM Pacific time.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bhavesh Mehta can be reached on 571-272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PHILIP H LAM/ Examiner, Art Unit 2656