Prosecution Insights
Last updated: October 02, 2026
Application No. 18/478,759

UNIFIED AUDIO SUPPRESSION MODEL

Final Rejection §102§103
Filed
Sep 29, 2023
Examiner
FLANDERS, ANDREW C
Art Unit
2655
Tech Center
2600 — Communications
Assignee
Amazon Technologies Inc.
OA Round
3 (Final)
74%
Grant Probability
Favorable
4-5
OA Rounds
2m
Est. Remaining
88%
With Interview

Examiner Intelligence

Grants 74% — above average
74%
Career Allowance Rate
579 granted / 781 resolved
+12.1% vs TC avg
Moderate +14% lift
Without
With
+14.1%
Interview Lift
resolved cases with interview
Typical timeline
3y 2m
Avg Prosecution
6 currently pending
Career history
788
Total Applications
across all art units

Statute-Specific Performance

§101
11.2%
-28.8% vs TC avg
§103
39.8%
-0.2% vs TC avg
§102
27.7%
-12.3% vs TC avg
§112
8.0%
-32.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 781 resolved cases

Office Action

§102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement filed 26 May 2026 fails to comply with 37 CFR 1.98(a)(2), which requires a legible copy of each cited foreign patent document; each non-patent literature publication or that portion which caused it to be listed; and all other information or that portion which caused it to be listed. It has been placed in the application file, but the information referred to therein has not been considered. Cite No. 15 under the “NON PATENT LITERATURE DOCUMENTS” heading is directed to a document to Shulen et al. titled “Speakerfilter: Deep learning-based target speaker extraction using anchor speech. No such document exists in the application file wrapper. The examiner notes another similar document titled “Speakefilter-Pro: An Improved Target Speaker Extractor Combines the Time Domain and Frequency Domain” is present in the file wrapper, but not the IDS submitted 26 May 2026. Appropriate clarification and/or correction is required. All other submitted references have been considered and their corresponding entries on the IDS submitted 26 May 2026 have been annotated. Response to Arguments Applicant’s arguments filed 26 May 2026, with respect to the claim rejections under 35 U.S.C. 112 have been fully considered and are persuasive. The rejection of the claims under these grounds has been withdrawn. Applicant's arguments filed 26 May 2026, with respect to the claim rejections under 35 U.S.C. 102 have been fully considered but they are not persuasive. Applicant alleges: However, Nighman does not teach or suggest receiving, "from any one of the plurality of users of the teleconference session, a selection corresponding to an audio selection mode," as recited in amended claim 1. Examiner respectfully disagrees. Nighman details a robust and dynamic system that allows for manual user entry of configuration and other parameters at many different stages in the processing of audio. For example, Nighman details authorized talkers issuing voice commands in order to activate any of the deployed devices and/or mixing/amplifying based on a specific command, such as “start lecture capture.” Additionally, when an instructor's issues a command to “start lecture capture,” other voices are ignored and the command causes the mixer/control engine 127 to include/amplify the separated audio source corresponding to the instructor's speech while excluding/dampening speech uttered by others talkers [0136]. The examiner submits that this configuration, namely where the lecturer is emphasized and others are de-emphasized, is effectuated through the use of a voice command, and can reasonably be construed as a section of a “mode.” Applicant’s remaining arguments regarding the claim rejections under 35 U.S.C. 102 and 35 U.S.C. 103 are not persuasive for the same reasons stated above. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 1, 2, 4 – 9, 11 – 15, and 17 – 20 is/are rejected under 35 U.S.C. 102(a)(1) and 35 U.S.C. 102(a)(2) as being anticipated by Nighman et al. (hereinafter Nighman, U.S. Patent Application Publication 2023/0115674). Regarding Claim 1, Nighman discloses: A system for enhancing teleconference application audio (e.g. teleconference system; abstract, entire doc; for improving/enhancing audio signals; [0011], [0022], [0033], [0073], [0089], [0113]), the system comprising: memory that stores computer-executable instructions (e.g. software module residing in various memories/disks or other storage [0228]); and a processor in communication with the memory (e.g. storage medium coupled to a processor; [0228]), wherein the computer-executable instructions, when executed by the processor, cause the processor (e.g. software module executed by a processor; [0228]) to: obtain a voice sample of a user (e.g. voice biometrics engine 330 can perform enrollment and verification phases to record and extract a number of features from a voice print [0118]); map the voice sample to an identifier associated with the user (e.g. voice biometrics engine 330 assigns unique acoustic speech signatures to each speaker; [0118]); receive an audio mixture of a teleconference session detected by an audio sensor (e.g. receiving audio inputs; [0074], [0098]; voice activity detector 320 identifies speech sources in the input signal; see also that the microphone(s) 140 detect sounds in the environment, convert the sounds to digital audio signals, and stream the audio signals to the processing core; [0084]; note environment in teleconference meeting rooms; [0069]) wherein the user is one of a plurality of users of the teleconference session (e.g. note Fig. 2C depicting a conferencing environment [0089], and plural “talkers,” “primary speaker,” [0089] as well as multiple user devices [0092] receive, from any one of the plurality of users of the teleconference session, a selection corresponding to an audio selection mode (e.g. a user manually entering information [0160]; also note authorized talkers issuing voice commands in order to activate any of the deployed devices and/or mixing/amplifying based on a specific command, such as “start lecture capture,” [0136]; note further the use of talker-specific personalization of audio settings; [0073] which can be customized based on a source; [0097]; and in general, some or all of the signal processing operations can be customized and/or personalized based on information; [0127]; note in the alternative the ability to manually input via user interface interaction to populate tables with configuration [“mode”] data; [0158]) modify a representation of the audio mixture to include a flag that corresponds to the audio selection mode (e.g. insert flags in the point speech source stream 322, where each flag indicates a biometrically identified talker associated with a source in the enhanced point speech source stream; [0119]; in essence, the inserted flag in this instance would be based on the “start lecture capture” command detailed in [0136] which would be flagged as speech by the instructor which would ultimately “correspond” to the command issued by virtue of the fact it’s spoken by the lecturer and will implement the specific settings or “mode,” of that user); and apply the modified representation of the audio mixture as an input into a machine learning model (e.g. identify and process flags in the stream corresponding to the talker that uttered the speech content [0121]; note that any of the engines in speech source processor can apply machine learning or AI algorithms to tune or adapt; [0124]), wherein application of the modified representation of the audio mixture as the input to the machine learning model causes the machine learning model (e.g. implementing machine learning to improve recognition of unique speakers; [0124]; note that any of the blocks, including noise suppressor 345 can implement machine learning or AI algorithms; [0132]) to one of: suppress a background noise of the audio mixture (e.g. train AI models (e.g., neural network-based models) to identify and suppress noises... suppress noise specific to the deployed environment; [0134]; see also suppression of background noises; [0177]-[0179]) or suppress all noise of the audio mixture except a voice identified by the user identifier (e.g. custom automatic gain control on a speaker-by-speaker basis; [0128]; see further the process of Fig. 7, and note output whether to include desired sources, such as one or more speech sources and to exclude or attenuate other sources such as noise; [0144]; further note, when an instructor's issues a command to “start lecture capture,” while other voices are ignored... and to cause the mixer/control engine 127 to include/amplify the separated audio source corresponding to the instructor's speech... while excluding/dampening speech uttered by others talkers [0136] Thus, in some arrangements one speech source is included/amplified and the others excluded or attenuated). Regarding Claim 2, in addition to the elements stated above regarding claim 1, Nighman further discloses: wherein the modified representation of the audio mixture includes the user identifier, the audio mixture, and the flag (e.g. voice print/acoustic speech signature for talkers; [0118]; and inserted flags in the (audio) stream; [0118]-[0121]). Regarding Claim 4, in addition to the elements stated above regarding claim 1, Nighman further discloses: wherein suppressing the background noise of the audio mixture comprises preserving a second voice of a second user from being suppressed (e.g. train AI models (e.g., neural network-based models) to identify and suppress noises... suppress noise specific to the deployed environment; [0134]; see also suppression of background noises; [0177]-[0179]; further, as noted above in some arrangements one or more speech source is included/amplified and the others excluded or attenuated). Regarding Claim 5, in addition to the elements stated above regarding claim 1, Nighman further discloses: wherein the machine learning model is trained on combined training data that comprises a first training data item (e.g. machine learning adaptively trains and tunes the algorithm to the particular deployed environment in response to training data; [0103], [0115]), wherein the first training data item includes a combination of a first type of clean speech data and a first type of background noise data (e.g. training data includes publicly available corpuses including a data set of human labeled [“clean”] sound events and curated noise samples [0103], [0115]), and wherein the first type of clean speech data is identified as a target output (e.g. any of the blocks can train on data on the fly or at a later time; [0115]; not enrolment and verification phrases to record and extract for voice prints [0118]; These are used to identify user speech to output, or “target output”). Regarding Claim 6, in addition to the elements stated above regarding claim 1, Nighman further discloses: wherein the machine learning model is trained on the voice sample of the user during an enrollment phase (e.g. any of the blocks can train on data on the fly or at a later time; [0115]; not enrolment and verification phrases to record and extract for voice prints [0118]; These are used to identify user speech to output, or “target output”). Regarding Claim 7, in addition to the elements stated above regarding claim 1, Nighman further discloses: receive, during a teleconference session in which the selection is received (e.g. note that audio processing engine automatically adjusts configuration as the nature of audio sources change; [0215], in other words, as different users are speaking, different noises appear, etc, the configuration or exclusion/attenuation of various can sources change) , a second selection via the teleconference application that identifies a second portion of the audio mixture to suppress that is different than the portion of the audio mixture; and cause the second portion of the audio mixture to be suppressed (e.g. custom automatic gain control on a speaker-by-speaker basis; [0128]; see further the process of Fig. 7, and note output whether to include desired sources, such as one or more speech sources and to exclude or attenuate other sources such as noise; [0144]; see also the portions referred to above in the rejection of claim 1 related to the selection; In essence, an administrator enters participant info through a user interface, the participant info corresponds to authorized speakers and the system will then suppress other sources of speakers for the adjusted configuration of the changing audio sources). Regarding Claim 8, Nighman discloses: A method for enhancing audio of a communication application (e.g. teleconference system; abstract, entire doc; for improving/enhancing audio signals; [0011], [0022], [0033], [0073], [0089], [0113]), the method comprising: obtaining a voice sample of a user (e.g. voice biometrics engine 330 can perform enrollment and verification phases to record and extract a number of features from a voice print [0118]); mapping the voice sample to an identifier associated with the user (e.g. voice biometrics engine 330 assigns unique acoustic speech signatures to each speaker; [0118]); receiving an audio mixture of a teleconference session detected by an audio sensor (e.g. receiving audio inputs; [0074], [0098]; voice activity detector 320 identifies speech sources in the input signal; see also that the microphone(s) 140 detect sounds in the environment, convert the sounds to digital audio signals, and stream the audio signals to the processing core; [0084] note environment in teleconference meeting rooms; [0069])), wherein the user is one of a plurality of users of the teleconference session (e.g. note Fig. 2C depicting a conferencing environment [0089], and plural “talkers,” “primary speaker,” [0089] as well as multiple user devices [0092] receiving, from any one of the plurality of users of the teleconference session, a selection corresponding to an audio selection mode (e.g. a user manually entering information [0160]; also note authorized talkers issuing voice commands in order to activate any of the deployed devices and/or mixing/amplifying based on a specific command, such as “start lecture capture,” [0136]; note further the use of talker-specific personalization of audio settings; [0073] which can be customized based on a source; [0097]; and in general, some or all of the signal processing operations can be customized and/or personalized based on information; [0127]; note in the alternative the ability to manually input via user interface interaction to populate tables with configuration [“mode”] data; [0158]); modifying a representation of the audio mixture to include a flag that corresponds to the audio selection mode (e.g. insert flags in the point speech source stream 322, where each flag indicates a biometrically identified talker associated with a source in the enhanced point speech source stream; [0119]; in essence, the inserted flag in this instance would be based on the “start lecture capture” command detailed in [0136] which would be flagged as speech by the instructor which would ultimately “correspond” to the command issued by virtue of the fact it’s spoken by the lecturer and will implement the specific settings or “mode,” of that user); and applying the modified representation of the audio mixture as an input into a machine learning model (e.g. identify and process flags in the stream corresponding to the talker that uttered the speech content [0121]; note that any of the engines in speech source processor can apply machine learning or AI algorithms to tune or adapt; [0124]), wherein application of the modified representation of the audio mixture as the input to the machine learning model causes the machine learning model (e.g. implementing machine learning to improve recognition of unique speakers; [0124]; note that any of the blocks, including noise suppressor 345 can implement machine learning or AI algorithms; [0132]) to enhance a portion of the audio mixture corresponding to the audio selection mode (e.g. custom automatic gain control on a speaker-by-speaker basis; [0128]; see further the process of Fig. 7, and note output whether to include desired sources, such as one or more speech sources and to exclude or attenuate other sources such as noise; [0144]; Thus, in some arrangements one speech source is included/amplified and the others excluded or attenuated; in the alternative, consider training AI models (e.g., neural network-based models) to identify and suppress noises... suppress noise specific to the deployed environment; [0134]; see also suppression of background noises; [0177]-[0179])). Regarding Claim 9, claim 9 is directed to the method claim that corresponds to the system claimed in claim 2 and is rejected under the same grounds. Regarding Claim 11, in addition to the elements stated above regarding claim 8, Nighman further discloses: wherein the portion of the audio mixture includes a background noise of the audio mixture or all noise of the audio mixture except a voice identified by the user identifier (e.g. Each of the input signals include a mixed source combined signal component S.sub.comb, an echo component E, and a noise component N; [0196]; note further custom automatic gain control on a speaker-by-speaker basis; [0128]; see further the process of Fig. 7, and note output whether to include desired sources, such as one or more speech sources and to exclude or attenuate other sources such as noise; [0144]; further voice commands in order to activate any of the deployed devices and/or mixing/amplifying based on a specific command, such as “start lecture capture,” [0136] Thus, in some arrangements one speech source is included/amplified and the others excluded or attenuated). Regarding Claim 12, claim 12 is directed to the method claim that corresponds to the system claimed in claim 4 and is rejected under the same grounds. Regarding Claim 13, claim 13 is directed to the method claim that corresponds to the system claimed in claim 6 and is rejected under the same grounds. Regarding Claim 14, claim 14 is directed to the computer-readable medium claim that corresponds to the system claimed in claim 6 and method claim in claim 1 and is rejected under the same grounds. Regarding Claim 15, claim 15 is directed to the computer-readable medium claim that corresponds to the system claimed in claim 2 and is rejected under the same grounds. Regarding Claim 17, claim 17 is directed to the computer-readable medium claim that corresponds to the system claimed in claim 6 and is rejected under the same grounds. Regarding Claim 18, claim 18 is directed to the computer-readable medium claim that corresponds to the method claimed in claim 11 and is rejected under the same grounds. Regarding Claim 19, in addition to the elements stated above regarding claim 14, Nighman further discloses: wherein the computer-executable instructions, when executed, further cause the computer system to suppress a second voice of a second user (e.g. custom automatic gain control on a speaker-by-speaker basis; [0128]; see further the process of Fig. 7, and note output whether to include desired sources, such as one or more speech sources and to exclude or attenuate other sources such as noise; [0144]; Thus, in some arrangements one speech source is included/amplified and the others excluded or attenuated; in the alternative, consider training AI models (e.g., neural network-based models) to identify and suppress noises... suppress noise specific to the deployed environment; [0134]; see also suppression of background noises; [0177]-[0179])). Regarding Claim 20, claim 20 is directed to the computer-readable medium claim that corresponds to the system claimed in claim 4 and is rejected under the same grounds. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claim(s) 3, 10 and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Nighman et al. (hereinafter Nighman, U.S. Patent Application Publication 2023/0115674) in view of Liu et al. (hereinafter Liu, U.S. Patent Application Publication 2022/0036907). Regarding Claim 3, in addition to the elements stated above regarding claim 1, Nighman further discloses: wherein the flag is indicating whether the selection corresponds to background noise suppression or all noise suppression except the voice identified by the user identifier (e.g. flags corresponding to the speaker; [0118]-[0121]; and custom automatic gain control on a speaker-by-speaker basis; [0128]; see further the process of Fig. 7, and note output whether to include desired sources, such as one or more speech sources and to exclude or attenuate other sources such as noise; [0144]; Thus, in some arrangements one speech source is included/amplified and the others excluded or attenuated). Nighman fails to explicitly disclose that the flag is a binary bit. In a related field of endeavor (e.g. enhancing of audio during including using noise suppression in a networked conference environment), Liu details using a neural network to process features of audio and derive and then output a binary flag to provide an indicator (see Figs. 2 and 3 and [0047]). Modifying the flags disclosed by Nighman to operate as binary flags as disclosed by Liu further makes obvious: the flag is a binary bit (e.g. Nighman’s flag, now configured to be a binary flag as disclosed by Liu’s Figs. 2 and 3 and [0047]). It would have been obvious to one of ordinary skill in the art at the time the invention was filed to apply the teachings of Liu to the system of Nighman. Doing so would have provided users of Nighman’s system with the techniques of the sound enhancement system 200 of Liu which provides the advantage of performing high-fidelity audio processing and improves sound quality for both music and voice signals, see Liu [0035]. Further, given the substantial overlap of Nighman and Liu, e.g. they’re both directed to audio conferencing applications, improving audio quality, using neural networks/machine learning, and suppressing noise, integration of the various teachings from one to the other would have been seen as predictable to one of ordinary skill in the art. Regarding Claim 10, claim 10 is directed to the method claim that corresponds to the system claimed in claim 3 and is rejected under the same grounds. Regarding Claim 16, claim 16 is directed to the computer-readable medium claim that corresponds to the system claimed in claim 3 and is rejected under the same grounds. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Andrew C Flanders whose telephone number is (571)272-7516. The examiner can normally be reached M-F 8:30-5:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ANDREW C FLANDERS/ Supervisory Patent Examiner, Art Unit 2655
Read full office action

Prosecution Timeline

Sep 29, 2023
Application Filed
Aug 20, 2025
Non-Final Rejection mailed — §102, §103
Nov 11, 2025
Response Filed
Feb 27, 2026
Non-Final Rejection mailed — §102, §103
Mar 31, 2026
Examiner Interview Summary
Mar 31, 2026
Applicant Interview (Telephonic)
May 26, 2026
Response Filed
Sep 24, 2026
Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12670919
VOICE ACTIVATION DETECTING METHOD OF EARPHONES, EARPHONES AND STORAGE MEDIUM
3y 3m to grant Granted Jun 30, 2026
Patent 12650805
Playback Session Transitions Across Different Platforms
2y 11m to grant Granted Jun 09, 2026
Patent 12562160
ARBITRATION BETWEEN AUTOMATED ASSISTANT DEVICES BASED ON INTERACTION CUES
3y 2m to grant Granted Feb 24, 2026
Patent 12547835
AUTOMATIC EXTRACTION OF SEMANTICALLY SIMILAR QUESTION TOPICS
3y 1m to grant Granted Feb 10, 2026
Patent 12512089
TESTING CASCADED DEEP LEARNING PIPELINES COMPRISING A SPEECH-TO-TEXT MODEL AND A TEXT INTENT CLASSIFIER
3y 0m to grant Granted Dec 30, 2025
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

4-5
Expected OA Rounds
74%
Grant Probability
88%
With Interview (+14.1%)
3y 2m (~2m remaining)
Median Time to Grant
High
PTA Risk
Based on 781 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month