Prosecution Insights
Last updated: October 01, 2026
Application No. 18/057,268

VOICE REINFORCEMENT IN MULTIPLE SOUND ZONE ENVIRONMENTS

Non-Final OA §103
Filed
Nov 21, 2022
Priority
Dec 30, 2021 — provisional 63/295,062
Examiner
WASHBURN, DANIEL C
Art Unit
2657
Tech Center
2600 — Communications
Assignee
Cerence Operating Company
OA Round
4 (Non-Final)
50%
Grant Probability
Moderate
4-5
OA Rounds
3m
Est. Remaining
80%
With Interview

Examiner Intelligence

Grants 50% of resolved cases
50%
Career Allowance Rate
82 granted / 163 resolved
-11.7% vs TC avg
Strong +30% interview lift
Without
With
+30.2%
Interview Lift
resolved cases with interview
Typical timeline
4y 1m
Avg Prosecution
8 currently pending
Career history
179
Total Applications
across all art units

Statute-Specific Performance

§101
12.1%
-27.9% vs TC avg
§103
53.6%
+13.6% vs TC avg
§102
15.1%
-24.9% vs TC avg
§112
11.3%
-28.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 163 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 2/9/26 has been entered. Response to Arguments Applicant's arguments filed 2/9/26 have been fully considered but they are either moot or are not persuasive. Regarding Applicant’s first argument, that neither reference teaches: “the AEC using first adaptive filters and reference signals from the audio source indicative of the audio signal played back by the loudspeakers to estimate and cancel echo of the audio signal reproduced by the loudspeakers within the environment”, as described by the independent claims, this argument is a moot point, as newly cited US Patent Song (US 11,017,799) is relied on to teach this limitation. Regarding Applicant’s second argument, that neither reference teaches: “the AFC using second adaptive filters and, as a reference, per channel voice outputs of vocal effects applied to speech reinforcement to produce the reinforced voice signal to estimate and cancel feedback resulting from application of the reinforced voice signal within the environment”, as described by the independent claims, this argument is not persuasive, as US Patent Hetherington et al. (US 11,348,595) teaches this limitation, as discussed in the updated rejections below. Regarding Applicant’s third argument, that “Applicant respectfully submits that the alleged combination can only be derived through impermissible hindsight reconstruction based on Applicant's own disclosure. Neither Hetherington nor Suzuki provides any teaching, suggestion, or motivation to modify Hetherington's single-module architecture to include the claimed sequential AEC and AFC adaptive filtering”, this argument is a moot point, as a new combination of Hetherington, Suzuki, and Song is relied on to teach the amended claims, as discussed in the updated rejections below. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1, 5, 7, 9, 13, 15, 17, 21, and 23 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hetherington et al. (US 11,348,595), herein “Hetherington”, in view of Suzuki et al. (US 10,542,154), herein “Suzuki”, and further in view of Song (US 11,017,799), herein “Song”. RE claims 1, 9, and 17, Hetherington describes a system, method, and non-transitory computer-readable medium comprising instructions for sound signal processing in a vehicle multimedia system, comprising: loudspeakers configured to reproduce, within an environment, an audio signal from an audio source and a reinforced voice signal (col. 3 lns .51-63: “In this front-to-back reinforcement in which the reinforced signal is conveyed by a single channel and the infotainment comprises music in stereo the four loudspeaker signals, x_1[n], x_4[n] can be represented as: x_1[n]=FL=music left x_2[n]=FR=music right x_3[n]=RL=music left+reinforcement signal x_4[n]=RR=music right+reinforcement signal”); at least one microphone for detection of a microphone signal, where the microphone signal includes a first voice signal component that corresponds to uttered speech, a second voice signal component that corresponds to the reinforced voice signal as reproduced by the loudspeakers, and an audio signal component corresponding to the audio signal as reproduced by the loudspeakers (col. 2 lns. 39-51: “In vehicle 200 of FIG. 2, the driver and one or more co-driver's (not shown) microphone signals are captured by microphones 202 A and B and then processed and played in rear zone 206 B of the vehicle 200 through loudspeakers 204 C and D. These loudspeakers are provided with front-to-back reinforcing signals 208 C and D. Likewise, one or more rear microphone signals can be captured by microphones 202 C and D, and thereafter processed and converted into audible sound in the front zone 206 A of the vehicle 200 through loudspeakers 204 A and B if there were rear passengers communicating in the vehicle 200. These loudspeakers are provided with back-to-front re-enforcing signals 208 A and B.” Also see col. 3 lns. 27-34: “In other alternative configurations, the entertainment signals and reinforcement signals may be rendered by additional loudspeakers, e.g., tweeters or subwoofer.”); and a voice processor system configured to receive the microphone signal from the at least one microphone (col. 4 lns. 1-5 “In FIG. 3 echo and feedback cancellation estimates the impulse response paths {h_j[n]; j=1, . . . , J} given the reference channels {x_j[n]]; j=1, . . . , J} and the microphone signal Y[n], and then subtracts the echo E[n] from microphone signal Y[n].”), perform acoustic echo cancellation (AEC) of the microphone signal to produce an echo cancelled microphone signal, the AEC using first adaptive filters to estimate and cancel echo of the audio signal reproduced by the loudspeakers within the environment (col. 2 lns. 25-30 “Due to feedback and echo that accompanies the in car environment, echo and feedback cancellation is performed at the audio processor and thereafter amplified. Here, adaptive filters model the loudspeaker-to-microphone impulse responses that are executed by the audio processor to cancel echo and feedback.” Also see col. 6 ln. 65 – col. 7 ln. 9: “At 504, the process models the acoustic environment of the vehicle by modeling the physical paths from the loudspeakers to the microphones and updates the echo canceller coefficients per each reference signal and each microphone. […] The echo canceller coefficients to be updated in 506 may be Finite Impulse Response (FIR) or Infinite Impulse Response (IIR) adaptive filter coefficients per each microphone and each loudspeaker.”), perform acoustic feedback cancellation (AFC) of the microphone signal to produce a processed microphone signal, the AFC using second adaptive filters and, as a reference, per channel voice outputs of vocal effects applied to speech reinforcement to produce the reinforced voice signals to estimate and cancel feedback resulting from application of the reinforced voice signal within the environment (col. 2 lns. 25-30 “Due to feedback and echo that accompanies the in car environment, echo and feedback cancellation is performed at the audio processor and thereafter amplified. Here, adaptive filters model the loudspeaker-to-microphone impulse responses that are executed by the audio processor to cancel echo and feedback.” Also see col. 4 lns. 18-21: “Because the signals are unique, the echo paths are optimally modeled by the echo & feedback cancellation module 314 that may comprise one or more instances of an adaptive filter, for example, before the signals are post-processed by an optional post processor 316.” Further, see col. 14 lns. 5-8: “The memory 604 and/or 1004 may store information in data structures including, for example, feedback and or echo canceller coefficients that render or estimate echo signal levels.” Additionally, regarding the newly added limitations, see col. 4 lns. 1-9: “In FIG. 3 echo and feedback cancellation estimates the impulse response paths {h_j[n]; j=1, . . . , J} given the reference channels {x_j[n]]; j=1, . . . , J} and the microphone signal Y[n], and then subtracts the echo E[n] from microphone signal Y[n]. In FIG. 3, synthesizer 312, such as a real-time sound synthesizer differentiates the signals by making a non-linear modification to the reinforcing signals and/or by adding an uncorrelated signal to each channel making each of the signals unique.” (emphasis added) Col. 4 lns. 46-58: “Other audio effects such as chorus, flange, and pitch shift may also be generated by synthesizer 312 that enhance the reinforced vocals by rendering a richer, more pleasing and professional sound. Reverberation may also be added by synthesizer 312 to render a sound that simulates an in-car talker's sound (e.g., speech, song, utterances, etc.) being reflected off of a large number of surfaces and simulating a large number of reflections that build up and then decay as if the sound were absorbed by the surfaces in a much larger and/or different space. It can provide the illusion of speaking, singing, or performing in a larger acoustic space such as a night club, concert hall, or cathedral, rather than in the small confines of the vehicle's cabin.” (emphasis added) Col. 5 lns. 15-32: “In FIG. 3, after the signals are synthesized and the echo and feedback subtracted, an optional post-processor 316 may further process the signals to enhance sound. A post-processor 316, for example, may apply equalization and/or adaptive gain. The equalization may modify the tone color or timbre and the adaptive gain adjusts (e.g., amplifies or attenuates) the level of the reinforced signals processed by echo & feedback cancellation module 314 as a result of or in response to the level of environmental noise sensed via an in-cabin noise detector or estimated in the vehicle cabin. The adapted and equalized signal is then added to the signal sourced by the stereo infotainment source 318 through the signal adder circuit 320 L and R, respectively. Thereafter, the reinforced signal is translated into analog signals by DAC 308 and transmitted into in the rear zone 206 B by the two rear loudspeakers 306 A and B. As shown, the echo and feedback cancellation module 314 includes a closed loop 322 to adjust its output.”), reinforce the uttered speech in the processed microphone signal to produce the reinforced voice signal (col. 4 lns. 1-13: “In FIG. 3 echo and feedback cancellation estimates the impulse response paths {h_j[n]; j=1, . . . , J} given the reference channels {x_j[n]]; j=1, . . . , J} and the microphone signal Y[n], and then subtracts the echo E[n] from microphone signal Y[n]. In FIG. 3, synthesizer 312, such as a real-time sound synthesizer differentiates the signals by making a non-linear modification to the reinforcing signals and/or by adding an uncorrelated signal to each channel making each of the signals unique.”), and apply the reinforced voice signal and the audio signal to the loudspeakers for reproduction in the environment (col. 4 lns. 16-18: “In FIG. 3, the signal adder circuit 320 L and R adds the echo cancelled audio processed signal to the infotainment signals.”). While Hetherington describes performing echo cancellation and feedback cancellation, as described above, Hetherington doesn’t explicitly describe performing acoustic feedback cancellation (AFC) of the echo cancelled microphone signal to produce a processed microphone signal. In other words, Hetherington doesn’t explicitly describe performing echo cancellation first, followed by feedback cancellation of the echo cancelled signal. However, Suzuki describes performing acoustic feedback cancellation (AFC) of the echo cancelled microphone signal to produce a processed microphone signal (col. 7 ln. 65 – col. 8 ln. 5: “In this exemplary embodiment, second acoustic feedback canceller 60 is a circuit for further removing the second acoustic feedback signal from an output signal of first echo and crosstalk canceller 50, in which a calculated first interference signal is removed from the output signal of first microphone 21, and for outputting a signal obtained after the removal to first loudspeaker 22”). It would have been obvious before the effective filing date of the claimed invention to include in Hetherington a system and method of performing acoustic feedback cancellation (AFC) of the echo cancelled microphone signal to produce a processed microphone signal, as taught by Suzuki, in order to implement an effective technique to remove all interfering noise and feedback signals from a reinforced voice signal, which results in a high-quality reinforced voice signal (Suzuki col. 2 lns. 22-27). The process of first removing echo from the microphone signal and then removing feedback has advantages such as isolating and removing echo in the environment and then isolating and removing feedback from the echo-cancelled signal, which enables the echo and feedback removal methods to be tailored to the removal of specific sounds in the environment, which results in a higher quality echo and feedback cancellation. Hetherington in view of Suzuki doesn’t describe the AEC using first adaptive filters and reference signals from the audio source indicative of the audio signal played back by the loudspeakers to estimate and cancel echo of the audio signal reproduced by the loudspeakers within the environment. However, Song describes a system and method including an AEC that uses first adaptive filters and reference signals from the audio source indicative of the audio signal played back by the loudspeakers to estimate and cancel echo of the audio signal reproduced by the loudspeakers within the environment (col. 4 ln. 48 – col. 5 ln. 9: “In block S140, the noisy voice and the reference audio are input to an acoustic echo canceller (AEC) module as inputted data. The AEC module is configured to perform an echo cancellation on the inputted data to obtain training data having AEC residual noise. It is to be explained that, in embodiments of the present disclosure, the AEC is used to cancel inherent noises (including music, radio broadcast, text to speech (TTS) broadcast) from a signal received by a microphone, to remain effective voice data. The AEC is an essential technical means for BargeIn disruption scenario. Implementation frames of the AEC are illustrated in FIG. 3. In embodiments of the present disclosure, the AEC module needs two inputs, one is a reference audio signal x(k) which may be music, radio broadcast, TTS broadcast audio or the like, and the other one is a signal y(k) received by the microphone (i.e., the above-mentioned noisy voice). Essence of the AEC module is to train an adaptive filter h(k) such that the adaptive filer h(k) may simulate a process of “transmitting a far-end signal to the microphone through a loudspeaker and under the interior environment of the vehicle” (that is the “echo”), and to remove the echo from the signal received by the microphone to obtain an echo-cancelled voice signal. In embodiments of the present disclosure, the adaptive filter h(k) may be trained using several algorithms, such as NLMS, frequency-domain adaptive filtering or the like. Therefore, after the noisy voice and the reference audio are inputted to the AEC module as the input data, the AEC module may perform the echo cancellation on the input data” (emphasis added)). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include in Hetherington in view of Suzuki a system and method including an AEC that uses first adaptive filters and reference signals from the audio source indicative of the audio signal played back by the loudspeakers to estimate and cancel echo of the audio signal reproduced by the loudspeakers within the environment, as taught by Song, in order to improve the quality of the echo removal by using a reference audio signal for the audio to be output by the vehicle’s loudspeakers, which enables the system to very effectively remove the audio associated with the reference audio signal, without removing any of the desired voice signal. RE claims 5, 13, and 21 Hetherington describes the system of claim 1, method of claim 9, and medium of claim 17, wherein the environment includes a plurality of sound zones, the at least one microphone includes a first microphone in a first sound zone of the plurality of sound zones and a second microphone in a second sound zone of the plurality of sound zones (col. 2 lns. 34-58: “n FIG. 2 the audio processing system is part of the vehicle 200 and provides entertainment and echo and feedback cancellation. In other systems it is an accessory or a component of a motor vehicle and in other systems part of an audio system used in a room, which may be divided into zones. In vehicle 200 of FIG. 2, the driver and one or more co-driver's (not shown) microphone signals are captured by microphones 202 A and B and then processed and played in rear zone 206 B of the vehicle 200 through loudspeakers 204 C and D. These loudspeakers are provided with front-to-back reinforcing signals 208 C and D. Likewise, one or more rear microphone signals can be captured by microphones 202 C and D, and thereafter processed and converted into audible sound in the front zone 206 A of the vehicle 200 through loudspeakers 204 A and B if there were rear passengers communicating in the vehicle 200. These loudspeakers are provided with back-to-front re-enforcing signals 208 A and B.”), and the voice processor system is further configured to: reinforce first speech received to the first microphone to produce a first aspect of the reinforced voice signal in the first sound zone (col. 10 ln. 48 – col. 11 ln. 6: “In yet another application, the entertainment post processor 704 may execute synthesis signal processing that modifies the isolated speech from the multiple zones of the vehicle—where the zones comprise a front-left (or driver zone—zone one), front-right (co-driver zone or zone two), rear left (a passenger zone behind the driver or zone three), and rear-rear right (a passenger zone behind the co-driver—zone four). In this application the synthesis signal processing modifies the isolated voices coming from the different zones or alternatively, each of the occupants and modifies the spoken utterances before rending them through selected loudspeakers. The modification may occur by pitch shifting the audio of each zone and then rendering the processed utterances in different zones or combinations of zones out of selected loudspeakers. For example, the front-right zone may be pitch shifted up a half of an octave and projected into the vehicle cabin through rear loudspeaker 306 A, the front-left zone may be pitch shifted up two tenths of an octave and projected into the vehicle cabin through rear loudspeaker 306 B, the rear right zone may be pitch shifted up eight tenths of an octave and projected into the vehicle cabin through front loudspeakers 304 A and B, and the rear-left zone may be pitch shifted up an octave and projected into the vehicle cabin through front and rear loudspeakers 304 A and B and 306 A and B to render an in car harmony”); and reinforce second speech received to the second microphone to produce a second component of the reinforced voice signal in the second sound zone (col. 10 ln. 48 – col. 11 ln. 6: “In yet another application, the entertainment post processor 704 may execute synthesis signal processing that modifies the isolated speech from the multiple zones of the vehicle—where the zones comprise a front-left (or driver zone—zone one), front-right (co-driver zone or zone two), rear left (a passenger zone behind the driver or zone three), and rear-rear right (a passenger zone behind the co-driver—zone four). In this application the synthesis signal processing modifies the isolated voices coming from the different zones or alternatively, each of the occupants and modifies the spoken utterances before rending them through selected loudspeakers. The modification may occur by pitch shifting the audio of each zone and then rendering the processed utterances in different zones or combinations of zones out of selected loudspeakers. For example, the front-right zone may be pitch shifted up a half of an octave and projected into the vehicle cabin through rear loudspeaker 306 A, the front-left zone may be pitch shifted up two tenths of an octave and projected into the vehicle cabin through rear loudspeaker 306 B, the rear right zone may be pitch shifted up eight tenths of an octave and projected into the vehicle cabin through front loudspeakers 304 A and B, and the rear-left zone may be pitch shifted up an octave and projected into the vehicle cabin through front and rear loudspeakers 304 A and B and 306 A and B to render an in car harmony”). RE claims 7, 15, and 23, Hetherington teaches the system of claim 1, method of claim 9, and medium of claim 17, wherein the voice processor system is further configured to apply vocal effects to the reinforced voice signal, the vocal effects including the addition of artificial reverberation (col. 4 lns. 46-58: “Other audio effects such as chorus, flange, and pitch shift may also be generated by synthesizer 312 that enhance the reinforced vocals by rendering a richer, more pleasing and professional sound. Reverberation may also be added by synthesizer 312 to render a sound that simulates an in-car talker's sound (e.g., speech, song, utterances, etc.) being reflected off of a large number of surfaces and simulating a large number of reflections that build up and then decay as if the sound were absorbed by the surfaces in a much larger and/or different space. It can provide the illusion of speaking, singing, or performing in a larger acoustic space such as a night club, concert hall, or cathedral, rather than in the small confines of the vehicle's cabin.”), wherein the acoustic feedback cancellation uses the per-channel voice outputs of the vocal effects as the reference (col. 4 lns. 1-9: “In FIG. 3 echo and feedback cancellation estimates the impulse response paths {h_j[n]; j=1, . . . , J} given the reference channels {x_j[n]]; j=1, . . . , J} and the microphone signal Y[n], and then subtracts the echo E[n] from microphone signal Y[n]. In FIG. 3, synthesizer 312, such as a real-time sound synthesizer differentiates the signals by making a non-linear modification to the reinforcing signals and/or by adding an uncorrelated signal to each channel making each of the signals unique.” (emphasis added) Col. 4 lns. 46-58: “Other audio effects such as chorus, flange, and pitch shift may also be generated by synthesizer 312 that enhance the reinforced vocals by rendering a richer, more pleasing and professional sound. Reverberation may also be added by synthesizer 312 to render a sound that simulates an in-car talker's sound (e.g., speech, song, utterances, etc.) being reflected off of a large number of surfaces and simulating a large number of reflections that build up and then decay as if the sound were absorbed by the surfaces in a much larger and/or different space. It can provide the illusion of speaking, singing, or performing in a larger acoustic space such as a night club, concert hall, or cathedral, rather than in the small confines of the vehicle's cabin.” (emphasis added) Col. 5 lns. 15-32: “In FIG. 3, after the signals are synthesized and the echo and feedback subtracted, an optional post-processor 316 may further process the signals to enhance sound. A post-processor 316, for example, may apply equalization and/or adaptive gain. The equalization may modify the tone color or timbre and the adaptive gain adjusts (e.g., amplifies or attenuates) the level of the reinforced signals processed by echo & feedback cancellation module 314 as a result of or in response to the level of environmental noise sensed via an in-cabin noise detector or estimated in the vehicle cabin. The adapted and equalized signal is then added to the signal sourced by the stereo infotainment source 318 through the signal adder circuit 320 L and R, respectively. Thereafter, the reinforced signal is translated into analog signals by DAC 308 and transmitted into in the rear zone 206 B by the two rear loudspeakers 306 A and B. As shown, the echo and feedback cancellation module 314 includes a closed loop 322 to adjust its output.”). Claim(s) 2, 10, and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hetherington in view of Suzuki, as applied to claims 1, 9, and 17 above, and further in view of Christoph et al. (US 2007/0110254), herein “Christoph”. RE claims 2, 10, and 18, Hetherington in view of Suzuki doesn’t describe, but Christoph describes the system of claim 1, method of claim 9, and medium of claim 17, wherein the AEC is performed using a first subset of the loudspeakers, and the AFC is performed using a second, different subset of the loudspeakers ([0010]: “A dereverberation and feedback compensation system reduces the echo of audio signals received from a first audio device while reducing feedback of speech signals received from a second audio device.” [0020]: “The communication system may also comprise an audio device 9, such as a radio, a CD player, and/or a DVD player. If a back passenger 6 is in dialog with front passenger 5, the conversation may be aided by the communication system by detecting the passengers' utterances through the use of microphones 3 close to the passengers 5 and/or 6 respectively, processing the received signals, and broadcasting the processed signals to the loudspeakers close to the passengers 5 and/or 6 respectively. Because the microphones 3 may also detect audio signals generated by audio device 9 and broadcast by loudspeakers 2, the communication system may also process these signals.” [0029]: “After estimating the impulse response h.sub.i(n) between loudspeaker 2 and microphone 3, the feedback components of an audio signal from a second audio device 15, such as a loudspeaker that transmits a passenger's verbal utterances (e.g., a speech signal), may also be estimated.” [0030]: “Where the audio signals x.sub.i(n) may be broadcasted by all of the loudspeakers, the output signals of the communication system y.sub.i(n) may be transmitted by the loudspeaker(s) close to the listening communication partner only”). It would have been obvious before the effective filing date of the claimed invention to include in Hetherington in view of Suzuki a system and method wherein the AEC is performed using a first subset of the loudspeakers, and the AFC is performed using a second, different subset of the loudspeakers, as taught by Christoph, in order to reduce the processing requirements associated with the echo cancellation and feedback cancellation processes, without degrading system performance, by only performing echo and/or feedback cancellation on the loudspeakers that are outputting audio signals and/or reinforced voice signals. Claim(s) 3, 11 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hetherington in view of Suzuki, as applied to claims 1, 9, and 17 above, and further in view of Li (US 11,211,061), herein “Li”. RE claims 3, 11, and 19, Hetherington describes the system of claim 1, method of claim 9, and medium of claim 17, wherein the voice processor system is further configured to perform automatic speech recognition (ASR) on the processed microphone signal (col. 9 ln. 34 – col. 10 ln. 47: “In FIG. 7, an entertainment post processing system 704 may deliver entertainment, services, or a grammar-based or a natural language-based automatic speech recognition (ASR). Since the in-car entertainment communication system isolates speech and/or other content delivered in the vehicle 200 a parallel architecture through a tree-based ASR structure may execute speech recognition of a limited vocabulary size through one or more processing branches (or paths) when resources are limited or through an unlimited vocabulary through a natural language vocabulary that can include a dictionary in one or more or all processing branches or a combination of ASRs.”). Hetherington in view of Suzuki doesn’t explicitly describe a system or method wherein the voice processor system is further configured to perform automatic speech recognition (ASR) on the processed microphone signal to receive commands to control the voice processor system. However, Li describes a system and method wherein the voice processor system is further configured to perform automatic speech recognition (ASR) on the processed microphone signal to receive commands to control the voice processor system (col. 16 lns. 22-47: “The ASR system 2000 uses various modules to perform tasks such as audio capture and import and to provide prompt services. ASR services may be launched through a Persistent Publish/Subscribe (PPS) service when the user activates Push-to-Talk (PPT) functionally, for examples, by touching and activating a PTT button or tab on a human-machine interface (HMI) interface displayed on the touchscreen 1836. The audio module 2010 include an audio capture module that detects speech commands, including the beginning and end of sentences, and forwards the audio stream to the speech recognition modules 2015.” […] “For example, the speech recognition module 2015 would take the utterance “search media for Hero” and create a results structure” Also see col. 17 lns. 1-7: “Applications such as Media Player and Navigation may subscribe to PPS objects for changes. For example, if the user activates PTT and says “play Arcade Fire”, the speech recognition modules 2015 parse the speech command. The Media conversation module 2020 then activates the media engine, causing tracks from the desired artist to play.”). It would have been obvious before the effective filing date of the claimed invention to include in Hetherington in view of Suzuki a system or method wherein the voice processor system is further configured to perform automatic speech recognition (ASR) on the processed microphone signal to receive commands to control the voice processor system, as taught by Li, in order to enable users to interact with and control the voice processor system using voice commands, which improves the user experience by making user input easier. Claim(s) 4, 12, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hetherington in view of Suzuki, and further in view of Li, as applied to claims 3, 11, and 19, and further in view of Thoresz et al. (US 2020/0314486), herein “Thoresz”. RE claims 4, 12, and 20, Hetherington in view of Suzuki and further in view of Li doesn’t teach, but Thoresz teaches the system of claim 3, method of claim 11, and medium of claim 19, wherein the commands include one or more of to: skip a song, repeat a song, repeat a section, adjust vocal effects and/or multichannel effects, add a user for voice reinforcement, turn off a user for voice reinforcement, turn off voice reinforcement for all users, or to request to turn on a voice processor mode to send uttered speech from one user to another of the users ([0080]: “The playback control region 133d can include selectable (e.g., via touch input and/or via a cursor or another suitable selector) icons to cause one or more playback devices in a selected playback zone or zone group to perform playback actions such as, for example, play or pause, fast forward, rewind, skip to next, skip to previous, enter/exit shuffle mode, enter/exit repeat mode, enter/exit cross fade mode, etc. The playback control region 133d may also include selectable icons to modify equalization settings, playback volume, and/or other suitable playback actions.” [0081]: “in some embodiments the control device 130a is configured as an NMD (e.g., one of the NMDs 120), receiving voice commands and other sounds via the one or more microphones 135.”). It would have been obvious before the effective filing date of the claimed invention to include in Hetherington in view of Suzuki in view of Li a system or method wherein the commands include one or more of to: skip a song, repeat a song, repeat a section, adjust vocal effects and/or multichannel effects, add a user for voice reinforcement, turn off a user for voice reinforcement, turn off voice reinforcement for all users, or to request to turn on a voice processor mode to send uttered speech from one user to another of the users, as taught by Thoresz, in order to enable the user to easily control the entertainment system via voice commands, which improves the user experience by giving a wider range of commands that that the system can recognize. Claim(s) 6, 14, and 22 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hetherington in view of Suzuki, as applied to claims 5, 13, and 21 above, and further in view of Roberge (US 2015/0255088), herein “Roberge”. RE claims 6, 14, and 21, Hetherington teaches the system of claim 5, method of claim 13, and medium of claim 21, wherein the first speech is received from a first singer, the second speech is received from a second singer (col. 10 ln. 62 – col. 11 ln. 6: “For example, the front-right zone may be pitch shifted up a half of an octave and projected into the vehicle cabin through rear loudspeaker 306 A, the front-left zone may be pitch shifted up two tenths of an octave and projected into the vehicle cabin through rear loudspeaker 306 B, the rear right zone may be pitch shifted up eight tenths of an octave and projected into the vehicle cabin through front loudspeakers 304 A and B, and the rear-left zone may be pitch shifted up an octave and projected into the vehicle cabin through front and rear loudspeakers 304 A and B and 306 A and B to render an in car harmony.” Also see col. 11 lns. 20-26: “In some in-car entertainment communication system that may include the functions shown in FIG. 7, the speech recognized lyrics are stored locally in memory or in a cloud-based storage as metadata with the original music or the processed music (with and/or without the original vocal tracks) so that the processing need only occur once a music track or segment is played. When the content is rendered, the original music or the track without vocals may be rendered in the vehicle cabin through loudspeaker 302 A and B and 306 A and B. The lyrics may be displayed on one or more heads-up display for each of the occupants or transmitted to the occupants wireless or mobile devices. In these alternatives a carpool karaoke system is rendered.”) Hetherington in view of Suzuki doesn’t teach, but Roberge teaches the voice processor system is further configured to: perform an evaluation of pitch of each of the first speech and the second speech against a reference pitch ([0003]: “There is further provided a system for scoring a singer, comprising a processing module determining notes duration and pitch of a melody of a reference song and notes duration and pitch of a melody of the singer's rendering of the reference song; and a scoring processing module comparing the notes duration and the pitch of the melody of the reference song and the notes and the pitch of the melody of the singer's rendering of the reference song.”); and identify whether the first singer or the second singer provided a performance closest to the reference pitch ([0037]: “ In FIG. 4, a fixed trip set point T.sub.a is shown. In practice, the trip set point T.sub.a is set at half the value of the energy of the first peak, so as to adapt to amplitude variations of the input signal. Hence, the envelope of a first singer singing louder than a second singer stops at the same point as the envelope of a second singer singing in a lower voice, which allows an equitable scoring between the different users.” [0041]: “Thus, the processing module 100 generates a set R of N parameters, defining the melody (notes) of a song, in terms of pitch and duration (i.e. time envelope). It serves as a reference when assessing the quality of the song as sung by a karaoke user.”). It would have been obvious before the effective filing date of the claimed invention to include in Hetherington in view of Suzuki a system and method wherein the voice processor system is further configured to: perform an evaluation of pitch of each of the first speech and the second speech against a reference pitch; and identify whether the first singer or the second singer provided a performance closest to the reference pitch, as taught by Roberge, which increases interest in the karaoke system by using it as a teaching system to improve a person’s singing skills, or by using it as a game or competition, where users compete to see who is the best singer. Claim(s) 8, 16, and 24 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hetherington in view of Suzuki, as applied to claims 7, 15, and 23 above, and further in view of Buck et al. (US 10,056,092), herein “Buck”. RE claims 8, 16, and 24, Hetherington in view of Suzuki doesn’t teach, but Buck describes the system of claim 7, method of claim 15, and medium of claim 23, wherein the voice processor system is further configured to increase a step size of adjustment of the second adaptive filters responsive to detection of reverberation in the processed microphone signal (col. 4 lns. 26-33: “controlling a step size of the adaptation AIC filter, dynamically adjusting a length of the FIR filter corresponding to a reverberation time and/or the ratio of early and late residual interference, the reference signal comprises a loudspeaker signal and the RSV estimate is applied for residual echo suppression, and/or the reference signal comprises a microphone signal and the second component corresponding to late RSV is used for dereverberation.” Also see col. 10 lns. 42-62: “The optimal step size for adaptation can be computed as: [see equation 20] Control of the step size μ(k) enables good convergence behavior. In embodiments, one aim is to get a better estimate of the residual of Φϵϵ(k) and thus, to a better convergence of the AIC filter. Generally, the dynamic step size enables the filter to adapt (and converge) quickly when Φϵϵ(k) is large (i.e., the filter is not well converged) and also ensures that the filter adapts slowly when Φϵϵ(k) is small (i.e., it prevents the filter from losing good convergence). A benefit of modelling the late reverb here is to get an estimate for the early residual PSD that is not affected by the late residual PSD. As a consequence, the AIC-step-size will be small even if there is significant late reverberant energy. This improves the convergence of the AIC-filter compared to conventional AIC control methods.” Further, see col. 11 lns. 31-46: “FIG. 7A shows an illustrative sequence of steps for residual interference suppression processing. In step 700, an AIC output signal having residual interference is received. In step 752, the residual interference is estimated including estimating a power spectral density of a first part of the residual interference corresponding to early reverberation and a second part corresponding to late reverberation. In one embodiment, the first part is estimated using a real-valued FIR filter operating on a power spectral density (PSD) of a reference signal and the second part is estimated using an exponential decay over time corresponding to a reverberation time using the PSD of the reference signal. In step 754, filter parameters can be adjusted. In step 756, a filter step size can be optimized for filter convergence. In step 758, a filter length can be adjusted based upon the reverberation time.”); and decrease the step size of the adjustment of the second adaptive filters responsive to a lack of reverberation in the processed microphone signal (col. 4 lns. 26-33: “controlling a step size of the adaptation AIC filter, dynamically adjusting a length of the FIR filter corresponding to a reverberation time and/or the ratio of early and late residual interference, the reference signal comprises a loudspeaker signal and the RSV estimate is applied for residual echo suppression, and/or the reference signal comprises a microphone signal and the second component corresponding to late RSV is used for dereverberation.” Also see col. 10 lns. 42-62: “The optimal step size for adaptation can be computed as: [see equation 20] Control of the step size μ(k) enables good convergence behavior. In embodiments, one aim is to get a better estimate of the residual of Φϵϵ(k) and thus, to a better convergence of the AIC filter. Generally, the dynamic step size enables the filter to adapt (and converge) quickly when Φϵϵ(k) is large (i.e., the filter is not well converged) and also ensures that the filter adapts slowly when Φϵϵ(k) is small (i.e., it prevents the filter from losing good convergence). A benefit of modelling the late reverb here is to get an estimate for the early residual PSD that is not affected by the late residual PSD. As a consequence, the AIC-step-size will be small even if there is significant late reverberant energy. This improves the convergence of the AIC-filter compared to conventional AIC control methods.” Further, see col. 11 lns. 31-46: “FIG. 7A shows an illustrative sequence of steps for residual interference suppression processing. In step 700, an AIC output signal having residual interference is received. In step 752, the residual interference is estimated including estimating a power spectral density of a first part of the residual interference corresponding to early reverberation and a second part corresponding to late reverberation. In one embodiment, the first part is estimated using a real-valued FIR filter operating on a power spectral density (PSD) of a reference signal and the second part is estimated using an exponential decay over time corresponding to a reverberation time using the PSD of the reference signal. In step 754, filter parameters can be adjusted. In step 756, a filter step size can be optimized for filter convergence. In step 758, a filter length can be adjusted based upon the reverberation time.”). It would have been obvious before the effective filing date of the claimed invention to include in Hetherington in view of Suzuki a system and method wherein the voice processor system is further configured to increase a step size of adjustment of the second adaptive filters responsive to detection of reverberation in the processed microphone signal; and decrease the step size of the adjustment of the second adaptive filters responsive to a lack of reverberation in the processed microphone signal, as taught by Buck, in order to improve the echo cancellation process by quickly removing reverberation when it is detected, while also maintaining good echo cancellation performance when reverberation is not detected, through the use of a dynamic step size in the filter, which results in a better user experience. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Li et al. (US 11,211,061) describes an invention that is very similar to the primary reference Hetherington, where FIG. 5 and col. 10 lns. 48-53 describe, “Once the signal is in the frequency domain, acoustic echo cancellation/suppression is applied by an AEC module 1406 to remove echo from audio playing from loudspeakers in the acoustic environment (e.g., vehicle cabin) such as music, radio or other background audio using references signals 1403.” (emphasis added). Chhetri (US 11,205,438) describes a system that includes acoustic echo cancellers, where the acoustic echo cancellers remove audio data based on data received from a reference audio data source (see FIGS. 6A-B, 7A-C). Hetherington et al. (US 20180190306) describes an invention that is very similar to the primary reference Hetherington (see figures). Every et al. (US 9,978,355) describes an acoustic management system that includes an acoustic echo canceller and an acoustic feedback canceller (see FIGS. 2, 3, and 6). Any inquiry concerning this communication or earlier communications from the examiner should be directed to Daniel C Washburn whose telephone number is (571)272-5551. The examiner can normally be reached Monday-Friday 9:00 am - 5:00 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel C Washburn can be reached at 571-272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DANIEL C WASHBURN/Supervisory Patent Examiner, Art Unit 2657
Read full office action

Prosecution Timeline

Show 1 earlier event
Jan 28, 2025
Non-Final Rejection mailed — §103
Apr 24, 2025
Response Filed
Jul 14, 2025
Non-Final Rejection mailed — §103
Oct 17, 2025
Response Filed
Nov 12, 2025
Final Rejection mailed — §103
Feb 09, 2026
Request for Continued Examination
Feb 18, 2026
Response after Non-Final Action
Jul 16, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12731590
End-to-End Speech Recognition Adapted for Multi-Speaker Applications
3y 10m to grant Granted Sep 08, 2026
Patent 12602555
METHOD FOR SEARCHING FOR TEXTS IN DIFFERENT LANGUAGES BASED ON PRONUNCIATION AND ELECTRONIC DEVICE APPLYING THE SAME
3y 3m to grant Granted Apr 14, 2026
Patent 12603084
METHOD, APPARATUS, AND COMPUTER-READABLE RECORDING MEDIUM FOR CONTROLLING RESPONSE UTTERANCE BEING REPRODUCED AND PREDICTING USER INTENTION
2y 7m to grant Granted Apr 14, 2026
Patent 12511480
Pattern Recognition Using NLP-Based Tokenizing and Clustering Models
2y 4m to grant Granted Dec 30, 2025
Patent 9614588
Smart Appliances
2y 2m to grant Granted Apr 04, 2017
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

4-5
Expected OA Rounds
50%
Grant Probability
80%
With Interview (+30.2%)
4y 1m (~3m remaining)
Median Time to Grant
High
PTA Risk
Based on 163 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month