Prosecution Insights
Last updated: August 17, 2026
Application No. 18/751,015

AUDIO-FOCUS FOR AMBIENT NOISE CANCELLATION

Final Rejection §102§103§112
Filed
Jun 21, 2024
Priority
Jun 23, 2023 — provisional 63/509,794
Examiner
HE, JIALONG
Art Unit
2659
Tech Center
2600 — Communications
Assignee
Google LLC
OA Round
2 (Final)
81%
Grant Probability
Favorable
3-4
OA Rounds
10m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 81% — above average
81%
Career Allowance Rate
755 granted / 927 resolved
+19.4% vs TC avg
Strong +33% interview lift
Without
With
+33.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
25 currently pending
Career history
945
Total Applications
across all art units

Statute-Specific Performance

§101
14.1%
-25.9% vs TC avg
§103
41.8%
+1.8% vs TC avg
§102
15.2%
-24.8% vs TC avg
§112
20.7%
-19.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 927 resolved cases

Office Action

§102 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action. Response to Amendments and Arguments Regarding twice anticipation rejections under 35 U.S.C. §102, applicant amended independent claims 1, 16 and 23 by adding different limitations to each of independent claims. Applicant argued (Remarks, page 10) that previously cited references fail to teach the newly added limitations in each of independent claims. By reviewing amended independent claims, the examiner notices that the three independent claims 1, 16 and 23 were amended by adding different limitations. The three independent claims are direct to three distinct inventions. Claim 1 was amended by adding: “a machine learning model configured based on a dataset including directional information modified by an audio response associated with the environment”. Claim 16 was amended by adding: “a machine learning model trained using a dataset including directional information modified by impulse response information”. Claim 23 was amended by adding: “a machine learning model trained using a dataset including directional information convolved with impulse response information”. By reviewing the original disclosure in light of limitations recited in dependent claim 13 or dependent claim 21, the examiner believes that the original disclosure fails to provide an adequate support for newly added limitations in each of independent claims. See more explanations in a section below (under 35 U.S.C. §112(a)). Since the original disclosure does not support new limitations in the amended independent claims, the examiner interprets claimed features based on a best understanding in light of the disclosure. The examiner combines newly discovered references (Olsson, US PG Pub. 2022/0272447 and Mobin, US Pat. 11,937,073) to reject the amended claims as unpatentable under 35 U.S.C. §103. Applicant’s arguments regarding previous anticipation rejections under 35 U.S.C. 102 have been considered. The arguments are moot because the arguments do not apply to the new ground of rejections necessitated by the amendments. Claim Objections Claims 2, 11 and 13 are objected to because of the following informalities: Claims 2, which depends from claim 1, recites “wherein the first machine learning model …”. Since claim 1 has been amended by deleting a word “first”, claim 2 must also be amended accordingly. Claims 11 and 13 have a similar issue as above explained for claim 2. Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(a): (a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), first paragraph: The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same and shall set forth the best mode contemplated by the inventor of carrying out his invention. Claims 1-24 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for pre-AIA the inventor(s), at the time the application was filed, had possession of the claimed invention. Applicant amended independent claims 1, 16 and 23 by adding different limitations related to training a machine learning model using a training dataset generated using different techniques. Claim 1 recites “a dataset including directional information modified by an audio response associated with the environment”; Claim 16 recites “a dataset including directional information modified by impulse response information”; Claim 23 recites “using a dataset including directional information convolved with impulse response information”. By carefully reviewing original disclosure in light of limitations recited in a dependent claim 13 or dependent claim 21, the training dataset is generated by convoluting an audio signal (claimed “first audio signal” / “second audio signal” in dependent claims 13 or 21) with an impulse response (e.g., DRIR dataset). The specification discloses convolution between “an audio signal with a directional room impulse response, i.e., DRIR” (Spec. [0048], [0052]). Please note, DRIR is an abbreviation of a directional room impulse response (Spec. [0045], [0060], [0064]). The DRIR is NOT an added term “directional information” in the amended claims. The specification mentions “directional information” refers to a direction of a head to torso simulator, i.e., HATS, is facing in the room (Spec. [0033]) or a direction that a computer device is facing (Spec. [0040]). In fact, the DRIR includes MHTF and DRTF, which are related to room acoustic properties, such as sound reverberation or reflection propagating in a room (Spec. [0044-0046]). The disclosure describes convoluting an audio signal with a room impulse response to obtain training dataset that captures properties of audio when propagating in a room (Spec. [0086], [0091]; Fig. 11, S1120-S1130). The specification does not disclose a training dataset including “directional information” convolved by impulse response information as recited in independent claim 23 The specification also does not disclose “directional information” modified by an audio response associated with the environment (in the amended claim 1) or modified by impulse response information (in the amended claim 16). The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1-15 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, or for pre-AIA the applicant regards as the invention. Applicant amended independent claim 1 by adding “an audio response associated with the environment”. Since antecedent limitations never mention any environment, the added limitation has insufficient antecedent basis. Dependent claims 2-15 include limitation of claim 1. Claims 2-15 also rejected. Claim Rejections - 35 USC § 103 Claims 1-2, 16 and 23 are rejected under 35 U.S.C. §103 as being unpatentable over Miller et al (US PG Pub. 2022/0130375, referred to as Miller) in view of Olsson et al. (US PG Pub. 2022/0272447, referred to as Olsson). Miller is a patent publication by the same assignee (Google Inc.) of the instant application. Miller discloses when a user issuing voice commands / queries to a speech-enabled device (Miller, Fig. 1), the system with a speech-enabled device determines user’s location. The system enhances audio signals from a user’s direction and de-emphasizes other directions using beamforming processing techniques (Miller, [0038-0039], Fig. 6). Miller further discloses using deep learning neural network techniques to determine user’s location / direction (Miller, [0041], [0054]). Independent 16 and 23 are narrower than claim 1. In the following analysis, the examiner analyzes limitations recited in a narrower independent claim 16. Olsson discloses using a machine learning model (e.g., a neural network model) to detect a sound source direction during a conference (Olsson, [0004-0005], [0034], [0038-0039]). Olsson discloses generating training data sets for training neural network-based room acoustic properties and impulse responses (Olsson, [0057-0059], [0063], [0066-0068], The impulse responses may be estimated to take into account e.g. a general reverberation time for the modelled conference room; also see [0074-0077], [0108]). Regarding claims 1-2, 16 and 23, Miller discloses a method, a non-transitory computer readable medium and an apparatus (Miller, [0059], Fig. 1, a computer implemented system for processing user’s voice request by focusing on audio signals from a user’s direction), comprising: identify an audio capture device and a target direction associated with the audio capture device (Miller, [0032-0035], a voice-enabled device is configured to detecting user’s direction when a user issues a voice request); detect first audio associated with the target direction (Miller, [0041], [0054], using deep learning neural network model to determine user’s location and enhance audios from the determined a target user direction); enhance the first audio using a machine learning model configured to detect audio associated with the target direction (Miller, [0038-0039], using a beamformer to emphasis audio signals coming from a user’s direction, de-emphasizing audio signals from other directions); detect second audio associated with a direction different from the target direction (Miller, [0039], de-emphasizing noises from other directions); and diminish the second audio using the machine learning model (Miller, [0039], [0054]). Miller discloses neural network implemented system for detecting sound source direction (Miller, [0034], [0041], [0054]). Miller does not discloses using a training dataset as defined by independent claims. Olsson discloses using a machine learning model (e.g., a neural network model) to detect a sound source direction during a conference (Olsson, [0004-0005], [0034], [0038-0039]). Olsson discloses generating training data sets for training a neural network-based room acoustic properties and impulse responses (Olsson, [0057-0059], [0063], [0066-0068], The impulse responses may be estimated to take into account e.g. a general reverberation time for the modelled conference room; also see [0074-0077], [0108]). It would have been obvious to a person having ordinary skill in the art at the time the invention was filed to combine Miller’s teaching with Olsson’s teaching to generating training dataset for training a neural network by using a room model and impulse responses of the room acoustics. One having ordinary skill in the art would have been motivated to make such a modification to improve accuracy of detecting sound source direction using a neural network and considering room dimensions and speaker’s positions (Olsson, [0007-0008]). Claims 1-10, 12, 16-20, and 23-24 are rejected under 35 U.S.C. §103 as being unpatentable over Liu et al (US Pat. 12,300,261, referred to as Liu) in view of Mobin (US Pat. 11,937,037, referred to as Mobin) Liu discloses a deep neural network (DNN) implemented audio processing system (Fig. 9A) for enhancing audio signals coming from a target user’s direction and cancelling non-speech / noises from other directions (Abstract, Col. 15, lines 21-50, Fig. 8). Fig. 8 shows an illustration of beamforming obtained from a neural network implemented system in Fig. 9A. In the following analysis, the examiner analyzes limitations of independent claim 16. Other independent claims are broader than claim 16. Mobin discloses generating synthetic audio training data for training a machine learning model. It builds a virtual room or space, places sound sources and a receiver in it, and runs many acoustic simulations. The system then samples those simulated outputs to create training examples. Some examples are labeled as desired sounds, while others are labeled as interfering sounds. The trained machine learning model can support speech enhancement or dereverberation. Regarding claims 1-2, 16 and 23, Liu discloses a method, a non-transitory computer readable medium and an apparatus (Liu, Col. 31, lines 10-25, Fig. 17, a computer implemented speech enhance system for enhancing speech signals from a desired direction and cancelling non-speech / noises from other directions), comprising: identify an audio capture device and a target direction associated with the audio capture device (Liu, Col. 3, lines 54-67, enhancing speech signals from a desired direction of a user; Fig. 1 shows devices for recording user’s voice request); detect first audio associated with the target direction (Liu, Col. 3, lines 63-67, Col. 7, lines 40-56; detecting target audio signals from a direction of a desired); enhance the first audio using a machine learning model configured to detect audio associated with the target direction (Liu, Col. 4, lines 4-15; enhance audio signal from a target direction by using a neural network model); detect second audio associated with a direction different from the target direction (Liu, Abstract, Col. 4, lines 102, Col. 15, lines 1-15, 62-65; Fig. 7 and Fig. 8, cancelling noises from other directions); and diminish the second audio using the machine learning model (Liu, Abstract, Col. 8, lines 60-67, Col. 15, lines 1-15, 62-65; Fig. 7 and Fig. 8,). Liu discloses a neural network implemented beamformer for enhancing speech (Liu, Col. 9, lines 7-15, Fig. 9A). Liu further discloses training the neural network model (Liu, Col. 9, lines 20-35). Liu does not disclose a training dataset as defined by independent claims. Mobin discloses generating training data for training a neural network speech enhancement system (Mobin, Col. 4, lines 4-30). Mobin discloses creating a virtual room and considering acoustic properties of 3D room (Mobin, Col 8, lines 20-40, 3D room reverberation, reflection, echo; Col. 7, line 1-5, Col. 8, line 25-30; room impulse responses). It would have been obvious to a person having ordinary skill in the art at the time the invention was filed to combine Liu’s teaching with Mobin’s teaching to generating training dataset by considering room acoustic properties (reverberation, echo etc.) and impulse responses. One having ordinary skill in the art would have been motivated to make such a modification to increase or improve performance of a machine learning model (Mobin, Col. 13, lines 50-55). Regarding claims 3, 17 and 24, Liu in view of Mobin further discloses diminishing the second audio includes decreasing an amplitude of at least one sound wave associated with the second audio (Liu, Col. 3, lines 63-67, Col. 8, lines 60-67; Col. 10, lines 1-3, Fig. 8, non-speech / noises from other directions, side lobe directions, are cancelled and reduced). Regarding claims 4-5 and 18, Liu in view of Mobin further discloses wherein diminishing the second audio includes decreasing an amplitude of at least one sound wave associated with the second audio (Liu, Abstract, Col. 3, lines 62-67, Col. 5, lines 1-6, Fig. 9A, Neural Sidelobe canceller for cancelling noises of other directions); diminishing the second audio includes eliminating the second audio by removing the second audio from an output of the second machine learning model (Liu, Col. 10, lines 20-32, Fig. 9A, removing remaining noises after enhancing speech from target direction and cancelling noise from other directions; neural network implemented noise canceller). Regarding claim 6, Liu in view of Mobin further discloses the first machine learning model and the second machine learning model are a same machine learning model (Liu, Col. 3, lines 54-67, Fig. 9A, enhancing speech from a target direction and cancelling noises from other directions using a neural network model). Regarding claim 7, Liu in view of Mobin further discloses enhancing the first audio includes increasing an amplitude of at least one sound wave associated with the first audio (Liu, Col. 3, lines 25-54; Fig. 8, increase gain for signals from a desired direction). Regarding claim 8, Liu in view of Mobin further discloses enhancing the first audio includes de-reverbing the first audio by removing resonant frequencies from the first audio (Liu, Col. 4, lines 4-15, removing reverberation and other noises in all sidelobe directions; Fig. 5). Regarding claim 9, Liu in view of Mobin further discloses enhancing the first audio includes de-noising the first audio by filtering the first audio (Liu, Col. 6, lines 39-50, Col. 8, lines 3-36, Fig. 5, Fig. 7). Regarding claims 10 and 19, Liu in view of Mobin further discloses the target direction is associated with a focus region (Liu, Col. 3, lines 25-67, focusing on a desired direction of a user; Fig. 8, beamforming towards to the main lobe direction). Regarding claims 12 and 20, Liu in view of Mobin further discloses the enhancing of the first audio using the first machine learning model includes: compressing the first audio using a first machine learning model (Liu, Fig. 9A, #930); and decompressing the compressed audio using a second machine learning model (Liu, Fig. 9A, #990). Claim 11 is rejected under 35 U.S.C. §103 as being unpatentable over Liu in view of Mobin, and further in view of Shah et al. (US PG Pub. 2019/0222691, referred to as Shah). Liu discloses training a neural network model for enhancing audio signals from a desired direction of a user and suppressing noises of other directions (Liu, Abstract, Col. 3, lines 54-67, Fig. 8, Fig. 9A). Liu does not explicitly mentions using “using an impulse response dataset”. Shah discloses an echo cancellation system by using impulse response dataset as training data (Shah, [0032]). Liu in view of Mobin and Shan are dealing with enhancing audio and removing noise. It would have been obvious to a person having ordinary skill in the art at the time the invention was filed to combine Liu’s teaching with Shah’s teaching to use impulse response dataset as training data. One having ordinary skill in the art would have been motivated to make such a modification to improve noise / echo cancellation performance (Shan, [0024], [0026]). Examiner Notes Claims 13-15 and 21-22 are not rejected over prior art references. These dependent claims may be allowable if overcome the rejection under 35 U.S.C. 112(a) / 112(b) set forth in this office action, and rewritten in independent form by including all of the limitations of the base claim and any intervening claims. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jialong He, whose telephone number is (571) 270-5359. The examiner can normally be reached on Monday – Friday, 8:00AM – 4:30PM, EST. If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Pierre Desir can be reached on (571) 272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JIALONG HE/Primary Examiner, Art Unit 2659
Read full office action

Prosecution Timeline

Jun 21, 2024
Application Filed
Mar 13, 2026
Non-Final Rejection mailed — §102, §103, §112
Jun 11, 2026
Examiner Interview Summary
Jun 11, 2026
Applicant Interview (Telephonic)
Jun 15, 2026
Response Filed
Jul 15, 2026
Final Rejection mailed — §102, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12694871
REAL-TIME SUMMARIZATION OF VIRTUAL CONFERENCE TRANSCRIPTS
2y 9m to grant Granted Jul 28, 2026
Patent 12688361
MACHINE LEARNING BASED QUESTION AND ANSWER (Q&A) ASSISTANT
2y 3m to grant Granted Jul 21, 2026
Patent 12682902
SPEECH-TO-TEXT PROCESSING ASSISTED WITH LANGUAGE MODELS FOR CONVERSATIONAL AI SYSTEMS AND APPLICATIONS
2y 6m to grant Granted Jul 14, 2026
Patent 12664981
PARAPHRASE AND AGGREGATE WITH LARGE LANGUAGE MODELS FOR IMPROVED DECISIONS
1y 11m to grant Granted Jun 23, 2026
Patent 12658184
VISUALIZATION INTERFACE FOR VOICE INPUT
7y 11m to grant Granted Jun 16, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
81%
Grant Probability
99%
With Interview (+33.0%)
3y 0m (~10m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 927 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month