Prosecution Insights
Last updated: October 02, 2026
Application No. 19/044,339

TARGET SPEAKER MODE

Non-Final OA §103
Filed
Feb 03, 2025
Priority
Sep 24, 2021 — CN CN202111122227.X +1 more
Examiner
SULTANA, NADIRA
Art Unit
Tech Center
Assignee
Zoom Video Communications Inc.
OA Round
1 (Non-Final)
73%
Grant Probability
Favorable
1-2
OA Rounds
1y 3m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 73% — above average
73%
Career Allowance Rate
80 granted / 110 resolved
+12.7% vs TC avg
Strong +36% interview lift
Without
With
+35.7%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
22 currently pending
Career history
135
Total Applications
across all art units

Statute-Specific Performance

§101
26.1%
-13.9% vs TC avg
§103
58.8%
+18.8% vs TC avg
§102
11.1%
-28.9% vs TC avg
§112
3.3%
-36.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 110 resolved cases

Office Action

§103
DETAILED ACTION Notice of AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed in parent Application No. CN202111122227.X, filed on 09/24/2021. It is noted, however, the attempt by the Office to electronically retrieve the priority documents has failed on 02/26/2025. Claim Objections Claims 1, 2, 11, 12 are objected to because of the following informalities: Claim 1 recites in lines 7, 8 “a trained target speaker voice-activity detection ("VAD") ML model”. Claim 2 recites in line 1, “the trained VAD ML model”. It’s not clear if the trained target speaker VAD model in claim 1 and the trained VAD model in claim 2 is same model or different. Applicant is requested to make proper correction to clear the ambiguity. Claim 11 recites in lines 12, 13 “a trained target speaker voice-activity detection ("VAD") ML model”. Claim 12 recites in line 1, “the trained VAD ML model”. It’s not clear if the trained target speaker VAD model in claim 11 and the trained VAD model in claim 12 is same model or different. Applicant is requested to make proper correction to clear the ambiguity. Double Patenting The non-statutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A non-statutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The filing of a terminal disclaimer by itself is not a complete reply to a non-statutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13. The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based e-Terminal Disclaimer may be filled out completely online using web-screens. An e-Terminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about e-Terminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer. Claims 1-3, 11, 12 are rejected on the ground of non-statutory double patenting as being unpatentable over claims 1, 2 ,15 of U.S. Patent No. 12,217, 761 B2. Although the claims at issue are not identical, they are not patentably distinct from each other because claims 1-3, 11, 12 of the instant application similar in scope and content of the patented claims 1, 2 ,15 of the patent issued to the same Applicant. It is clear that all the elements of the application claims 1-3, 11, 12 are to be found in patented claims 1, 2 ,15 (as the application claims 1-3, 11, 12 fully encompasses patented claims 1, 2 ,15). The difference between the application claims and the patent claims lies in the fact that the patented claims includes many more elements and are thus much more specific. Thus, the invention of claims 1, 2 ,15 of the patent is in effect a “species” of the “generic” invention of the application claims 1-3, 11, 12. It has been held that the generic invention is “anticipated” by the “species”. See In re Goodman, 29 USPQ2d 2010 (Fed. Cir. 1993). Since application claims 1-3, 11, 12 are anticipated by claims 1, 2 ,15 of the patent, it is not patentably distinct from of the patented claims. Instant Application: 19/044,339 Issued Patent : US 12,217,761 B2 1.A method comprising: receiving, by a target speaker extraction system, audio frames of an audio signal and a corresponding video; determining, by the target speaker extraction system using a trained multi- speaker detection machine learning ("ML") model, a presence of a voice of a single speaker within a first audio frame of the audio frames; suppressing, by the target speaker extraction system using a trained target speaker voice-activity detection ("VAD") ML model, a non-target speaker in the first audio frame based on a voiceprint of a target speaker; determining, by the target speaker extraction system using the trained multi- speaker detection ML model, a presence of voices of a plurality of speakers within a second audio frame of the audio frames; separating, by the target speaker extraction system using by a trained speech separation ML model, the voice of the target speaker from a voice mixture of the plurality of speakers in the second audio frame. 1.A computer-implemented method for target speaker extraction, comprising: receiving, by a target speaker extraction system, an audio frame of an audio signal and a corresponding video, wherein the target speaker extraction system comprises a trained multi-speaker detection machine-learning ("ML") model, a trained lip-movement- based ("LM-based") target speaker voice activity detection (VAD) ML model, and a trained speech separation ML model; responsive to determining, by the trained multi-speaker detection ML model of the target speaker extraction system, a single speaker within the audio frame: inputting, by the target speaker extraction system, the audio frame and the video to the trained LM-based target speaker VAD ML model; and suppressing, by the trained LM-based target speaker VAD ML model of the target speaker extraction system and based on the video, speech in the audio frame from a non-target speaker, wherein suppressing the speech in the audio from a non- target speaker comprises comparing the audio frame to a voiceprint of a target speaker; and responsive to determining, by the trained multi-speaker detection ML model of the target speaker extraction system, a plurality of speakers within the audio frame: inputting, by the target speaker extraction system, the audio frame to the trained speech separation ML model; and separating, by the trained speech separation ML model of the target speaker extraction system, the voice of the target speaker from a voice mixture in the audio frame. 2. The method of claim 1, wherein the trained VAD ML model comprises a lip- movement ("LM")-based target speaker VAD ML model and post-processing functionality, and wherein suppressing the non-target speaker is further based on the corresponding video. Claim 1 wherein the target speaker extraction system comprises a trained multi-speaker detection machine-learning ("ML") model, a trained lip-movement- based ("LM-based") target speaker voice activity detection (VAD) ML model, and a trained speech separation ML model; and suppressing, by the trained LM-based target speaker VAD ML model of the target speaker extraction system and based on the video, 3. The method of claim 1, wherein suppressing the non-target speaker in the first audio frame comprises suppressing speech from the non-target speaker in the first audio frame by a predetermined suppression ratio. 2. The method of claim 1, wherein suppressing, by the trained LM-based target speaker VAD ML model of the target speaker extraction system, speech in the audio frame from the non-target speaker further comprises: determining, by the trained LM-based target speaker VAD ML model of the target speaker extraction system, a suppression ratio; and suppressing, by the trained LM-based target speaker VAD ML model of the target speaker extraction system, the speech in the audio frame from the non-target speaker based on the suppression ratio. 4. The method of claim 1, wherein suppressing the non-target speaker in the first audio frame comprises: generating the first voiceprint of the first audio frame; determining a similarity between the first voiceprint and the voiceprint of the target speaker; and responsive to determining similarity does not satisfy a threshold, suppressing the non-target speaker in the first audio frame. 5. The method of claim 4, wherein determining the similarity comprises determining a cosine similarity between the first voiceprint and the voiceprint of the target speaker. 6. The method of claim 1, wherein separating the voice of the target speaker from the voice mixture comprises: decomposing the voice mixture into a plurality of speech signals, each speech signal associated with a different speaker of the plurality of speakers in the second audio frame; identifying a target speech signal corresponding to the voice of the target speaker based on the voiceprint of the target speaker; and outputting the target speech signal. 7. The method of claim 1, wherein the trained speech separation ML model or the target speaker voice-activity detection ("VAD") ML model comprise a plurality of one-dimensional convolutional neural networks. 8. The method of claim 7, wherein the trained speech separation ML model or the target speaker voice-activity detection ("VAD") ML model comprises a plurality of network blocks, each network block comprises one or more convolutional blocks, and each convolutional block comprises a one or more neural networks. 9. The method of claim 7, wherein the trained speech separation ML model or the target speaker voice-activity detection ("VAD") ML model comprises a plurality of network blocks, wherein outputs of one or more convolutional blocks in a network block are summed and input to a next network block of the plurality of network blocks. 10. The method of claim 9, wherein the sum of the outputs of the one or more convolutional blocks in the network block are fused with an embedding of the respective input audio frame prior to inputting the sum to the next network block. 11. A target speaker extraction system comprising a non-transitory computer-readable medium; and one or more processors communicatively coupled to the non-transitory computer-readable medium, the one or more processors configured to execute processor-executable instructions stored in the non-transitory computer-readable medium to: receive, by a target speaker extraction system, audio frames of an audio signal and a corresponding video; determine, by the target speaker extraction system using a trained multi-speaker detection machine learning ("ML") model, a presence of a voice of a single speaker within a first audio frame of the audio frames; suppress, by the target speaker extraction system using a trained target speaker voice-activity detection ("VAD") ML model, a non-target speaker in the first audio frame based on a voiceprint of a target speaker; determine, by the target speaker extraction system using the trained multi-speaker detection ML model, a presence of voices of a plurality of speakers within a second audio frame of the audio frames; separate, by the target speaker extraction system using by a trained speech separation ML model, the voice of the target speaker from a voice mixture of the plurality of speakers in the second audio frame. 15. A target speaker extraction system comprising: a non-transitory computer-readable medium; and one or more processors communicatively coupled to the non-transitory computer-readable medium, the one or more processors configured to execute processor-executable instructions stored in the non-transitory computer-readable medium to: receive, by a target speaker extraction system, an audio frame of an audio signal and a corresponding video, wherein the target speaker extraction system comprises a trained multi-speaker detection machine-learning (“ML”) model, a trained lip-movement-based (“LM-based”) target speaker voice activity detection (VAD) ML model, and a trained speech separation ML model; responsive to determining, by the trained multi-speaker detection ML model of the target speaker extraction system, a single speaker within the audio frame: input, by the target speaker extraction system, the audio frame and the video to the trained LM-based target speaker VAD ML model; and suppress, by the trained LM-based target speaker VAD ML model of the target speaker extraction system and based on the video, speech in the audio frame from a non-target speaker, wherein suppressing the speech in the audio from a non-target speaker comprises comparing the audio frame to a voiceprint of a target speaker; and responsive to determining, by the trained multi-speaker detection ML model of the target speaker extraction system, a plurality of speakers within the audio frame: input, by the target speaker extraction system, the audio frame to the trained speech separation ML model; and separate, by the trained speech separation ML model of the target speaker extraction system, the voice of the target speaker from a voice mixture in the audio frame. 12. The system of claim 11, wherein the trained VAD ML model comprises a lip- movement ("LM")-based target speaker VAD ML model and post-processing functionality, and wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer- readable medium to suppress the non-target speaker is further based on the corresponding video. Claim 11 wherein the target speaker extraction system comprises a trained multi-speaker detection machine-learning ("ML") model, a trained lip-movement- based ("LM-based") target speaker voice activity detection (VAD) ML model, and a trained speech separation ML model; and suppressing, by the trained LM-based target speaker VAD ML model of the target speaker extraction system and based on the video, 13. The system of claim 11, wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to: generate a first voiceprint of the first audio frame; determine a similarity between the first voiceprint and the voiceprint of the target speaker; and responsive to determining similarity does not satisfy a threshold, suppress the non-target speaker in the first audio frame. 14. The system of claim 11, wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to: decompose the voice mixture into a plurality of speech signals, each speech signal associated with a different speaker of the plurality of speakers in the second audio frame; identify a target speech signal corresponding to the voice of the target speaker based on the voiceprint of the target speaker; and output the target speech signal. 15. The system of claim 11, wherein the trained speech separation ML model or the target speaker voice-activity detection ("VAD") ML model comprise a plurality of one-dimensional convolutional neural networks. 16. The system of claim 15, wherein the trained speech separation ML model or the target speaker voice-activity detection ("VAD") ML model comprises a plurality of network blocks, wherein outputs of one or more convolutional blocks in a network block are summed and input to a next network block of the plurality of network blocks. 17. A non-transitory computer readable medium comprising processor-executable instructions configured to cause one or more processors to: receive, by a target speaker extraction system, audio frames of an audio signal and a corresponding video; determine, by the target speaker extraction system using a trained multi- speaker detection machine learning ("ML") model, a presence of a voice of a single speaker within a first audio frame of the audio frames; suppress, by the target speaker extraction system using a trained target speaker voice-activity detection ("VAD") ML model, a non-target speaker in the first audio frame based on a voiceprint of a target speaker; determine, by the target speaker extraction system using the trained multi- speaker detection ML model, a presence of voices of a plurality of speakers within a second audio frame of the audio frames; separate, by the target speaker extraction system using by a trained speech separation ML model, the voice of the target speaker from a voice mixture of the plurality of speakers in the second audio frame. 18. The non-transitory computer readable medium of claim 17, further comprising processor-executable instructions configured to cause the one or more processors to: generate a first voiceprint of the first audio frame; determine a similarity between the first voiceprint and the voiceprint of the target speaker; and responsive to determining similarity does not satisfy a threshold, suppress the non-target speaker in the first audio frame. 19. The non-transitory computer readable medium of claim 17, further comprising processor-executable instructions configured to cause the one or more processors to: decompose the voice mixture into a plurality of speech signals, each speech signal associated with a different speaker of the plurality of speakers in the second audio frame; identify a target speech signal corresponding to the voice of the target speaker based on the voiceprint of the target speaker; and output the target speech signal. 20. The non-transitory computer readable medium of claim 17, wherein the trained speech separation ML model or the target speaker voice-activity detection ("VAD") ML model comprise a plurality of one-dimensional convolutional neural networks. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-5, 11-13, 17, 18 are rejected under 35 U.S.C. 103 as being unpatentable over Li et al. ( US 20210217182 A1), hereinafter referenced as Li, in view of Sivaraman et al. (US 20220084509 A1), hereinafter referenced as Sivaraman. Regarding Claim 1, Li teaches a method comprising: receiving, by a target speaker extraction system, audio frames of an audio signal and a corresponding video ( Li: Para.[0052], Fig. 3A, at step 302 of the process 300A, the user may perform a transaction at a self-service kiosk having at least one microphone and at least one camera installed therein. The user may speak various instructions, commands, requests, responses while facing one or more image capture devices, which are captured (audio and video signal of a user) in the self-service kiosk); determining, by the target speaker extraction system using a trained multi- speaker detection machine learning ("ML") model, a presence of a voice of a single speaker within a first audio frame of the audio frames (Li: Para.[0014], [0036], output from a learning model ( could be CNN, RNN, BRNN) computing audio signals may provide classifications of time-frequency frames and bins thereof by labeling such frames and bins thereof as matching respective speakers, including a target speaker. Para.[0057]-[0060], Fig. 3A, at step 304, short-time Fourier transform (STFT) is performed on captured audio signal. The time-frequency representations of the audio signals may provide time-frequency representations of single-channel audio signals ( presence of single speaker). Para.[0028],[0095], training of learning models is generally performed using training datasets, which may be massive training datasets. A learning model may be trained on audio signals containing labeled vocal content), suppressing, by the target speaker extraction system using a trained target speaker voice-activity detection ("VAD") ML model, a non-target speaker in the first audio frame [based on a voiceprint of a target speaker] ( Li: Para.[0121], [0124], Figs. 3B, 3C, blind source separation ("BSS") learning model outputs a demixing matrix, which is in the frequency domain and when applied to a time-frequency representation of mixed-source multi-channel audio signals by an operation (such as a multiplication operation against each frame), yields an objective constituent source of the mixed-source multi-channel audio signal which is, a target speaker ( by suppressing the non-target speaker). At step 3105B, VAD outputs are detected); determining, by the target speaker extraction system using the trained multi- speaker detection ML model, a presence of voices of a plurality of speakers within a second audio frame of the audio frames (Li: Para.[0014],[0036], output from a learning model ( could be CNN, RNN, BRNN) computing audio signals may provide classifications of time-frequency frames and bins thereof by labeling such frames and bins thereof as matching respective speakers, including a target speaker. Speakers may be known speakers or unknown speakers, unknown speakers labeled in output from a learning model may, regardless, be distinguished as distinct speakers from other unknown speakers ( detecting multiple speakers)); Li while teaching the method of claim 1, fails to explicitly teach the claimed, suppressing, by the target speaker extraction system using a trained target speaker voice-activity detection ("VAD") ML model, a non-target speaker in the first audio frame based on a voiceprint of a target speaker; separating, by the target speaker extraction system using by a trained speech separation ML model, the voice of the target speaker from a voice mixture of the plurality of speakers in the second audio frame. However, Sivaraman does teach the claimed, suppressing, by the target speaker extraction system using a trained target speaker voice-activity detection ("VAD") ML model, a non-target speaker in the first audio frame based on a voiceprint of a target speaker (Sivaraman: Para. [0103], [0109], Figs. 1B, 4, the machine-learning architecture 400 performs the operations of a speaker-specific speech enhancement system and comprises a speaker-specific speech enhancement engine 402, including a speech separation engine 122 and noise suppression engine 124, and a SAD engine 404 ( speech activity detection) that identifies speech and non-speech portions of audio signals. The server applies the speech separation engine on the relevant target voiceprint and the features of the inbound audio signal to generate the speaker mask. The speech separation engine then applies a speaker mask on the features of the inbound audio signal to generate the target speaker signal, which suppresses the interfering speaker ( non-target) signals); separating, by the target speaker extraction system using by a trained speech separation ML model, the voice of the target speaker from a voice mixture of the plurality of speakers in the second audio frame (Sivaraman: Para. [0029],[0030], [0033], Fig. 1B, the speech separation engine 122 is trained to estimate the speaker separation speaker mask function. The speech separation engine receives an input audio signal containing a mixture of speaker signals and one or more types of noise. The speech separation engine extracts low-level spectral features, such as mel-frequency cepstrum coefficients (MFCCs), and receives a voiceprint for a target speaker, generated by the speaker-embedding engine. Using these two inputs, the speech separation engine generates a speaker mask for suppressing speech signals of interfering speakers. The speech separation engine applies the speaker mask on the features extracted from the input audio signal containing the mixture of speech signals, thereby suppressing the interfering speech signals and generating a target speaker signal). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Sivaraman’s teaching of a machine-learning architecture system that enhances the speech audio of a user-defined target speaker by suppressing interfering speakers, as well as background noise and reverberations , into the system and method of performing source separation on mixed source single-channel and multichannel audio signals enhanced by inputting lip motion information from captured image data , taught by Li, because, this would improve the voice quality of a target speaker on a single channel audio input containing a mixture of speaker speech signals and various types of noise.(Sivaraman, Para.[0009]-[0013]). Claim 11 is system claim comprising a non-transitory computer-readable medium; and one or more processors communicatively coupled to the non-transitory computer-readable medium, the one or more processors configured to execute processor-executable instructions stored in the non-transitory computer-readable medium to ( Li: Para.[0154],[0155],[0170],[0171], Fig. 8A, The system 800 may include one or more processors 802 and system memory 804 communicatively coupled to the processor(s) 802. The processor(s) 802 and system memory 804 may be physical or may be virtualized and/or distributed. The processor(s) 802 may execute one or more modules and/or processes to cause the processor(s) 802 to perform a variety of functions), perform the steps in method claim 1 above and as such, claim 11 is similar in scope and content to claim 1 and therefore, claim 11 is rejected under similar rationale as presented against claim 1 above. Claim 17 is non-transitory storage medium claim comprising processor-executable instructions configured to cause one or more processors to ( Li: Para.[0154],[0170],[0171], Fig. 8A, The system 800 may include one or more processors 802 and system memory 804 ( not transient computer readable media) communicatively coupled to the processor(s) 802. The processor(s) 802 may execute one or more modules and/or processes to cause the processor(s) 802 to perform a variety of functions), perform the steps in method claim 1 above and as such, claim 17 is similar in scope and content to claim 1 and therefore, claim 17 is rejected under similar rationale as presented against claim 1 above. Regarding Claim 2, Li in view of Sivaraman teach the method of claim 1. Li further teaches, wherein the trained VAD ML model comprises a lip- movement ("LM")-based target speaker VAD ML model and post-processing functionality, and wherein suppressing the non-target speaker is further based on the corresponding video ( Li: Para.[0156]-[0164], Figs. 3B, 8A, 8B illustrates a system for source separation which includes, multiple face recognition module 808, target speaker selecting module 810, facial feature extracting module 812B which further include submodule 8125B, which may be configured to output, by a VAD. Para.[0075]-[0090], Fig. 3C illustrates post processing functionality of producing VAD output from lip features, such as lip motion vectors, samples pixels. Para.[0121], [0124],[0172], Figs. 3B, 3C, blind source separation ("BSS") learning model outputs a demixing matrix, which is in the frequency domain and when applied to a time-frequency representation of mixed-source multi-channel audio signals by an operation (such as a multiplication operation against each frame), yields an objective constituent source of the mixed-source multi-channel audio signal which is, a target speaker ( by suppressing the non-target speaker). Thus, the source separation on mixed source single-channel and multi-channel audio signals enhanced by inputting lip motion information from captured image data). Claim 12 is system claim performing the steps in method claim 2 above and as such, claim 12 is similar in scope and content to claim 2 and therefore, claim 12 is rejected under similar rationale as presented against claim 2 above. Regarding Claim 3, Li in view of Sivaraman teach the method of claim 1. Sivaraman further teaches, wherein suppressing the non-target speaker in the first audio frame comprises suppressing speech from the non-target speaker in the first audio frame by a predetermined suppression ratio ( Sivaraman: Para.[0029], [0032], the speech separation engine generates a speaker mask for suppressing speech signals of interfering speakers, which is a ratio of features of the target speaker signal and features of the input audio signal containing the mixture of speaker signals ( suppression ratio)). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Sivaraman’s teaching of a machine-learning architecture system that enhances the speech audio of a user-defined target speaker by suppressing interfering speakers, as well as background noise and reverberations , into the system and method of performing source separation on mixed source single-channel and multichannel audio signals enhanced by inputting lip motion information from captured image data , taught by Li, because, this would improve the voice quality of a target speaker on a single channel audio input containing a mixture of speaker speech signals and various types of noise.(Sivaraman, Para.[0009]-[0013]). Regarding Claim 4, Li in view of Sivaraman teach the method of claim 1. Sivaraman further teaches, wherein suppressing the non-target speaker in the first audio frame comprises: generating a first voiceprint of the first audio frame ( Sivaraman: Para.[0071], Fig .1B, the speaker embedding engine 126 generates an enrolled voiceprint for an enrolled user. The speaker-embedding engine 126 extracts enrollee feature vectors from the features of the enrollment signals); determining a similarity between the first voiceprint and the voiceprint of the target speaker ( Sivaraman: Para.[0072], Fig .1B, the speaker embedding engine 126 generates a similarity score between the inbound voiceprint and the enrolled voiceprint, where the similarity score represents the comparative similarities between the inbound speaker and the enrolled speaker); and responsive to determining similarity does not satisfy a threshold, suppressing the non-target speaker in the first audio frame ( Sivaraman: Para.[0102],[0103], Fig. 2, at step 214, the server generates a deployment output by applying the machine-learning architecture on the inbound audio signal and the appropriate target voiceprint, such as a risk score representing a similarity between the target speaker's features in the enhanced audio signal and an enrolled voiceprint for the target speaker. The server applies the speech separation engine on the relevant target voiceprint and the features of the inbound audio signal to generate the speaker mask. The speech separation engine then applies a speaker mask on the features of the inbound audio signal to generate the target speaker signal, which suppresses the interfering speaker signals ). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Sivaraman’s teaching of a machine-learning architecture system that enhances the speech audio of a user-defined target speaker by suppressing interfering speakers, as well as background noise and reverberations , into the system and method of performing source separation on mixed source single-channel and multichannel audio signals enhanced by inputting lip motion information from captured image data , taught by Li, because, this would improve the voice quality of a target speaker on a single channel audio input containing a mixture of speaker speech signals and various types of noise.(Sivaraman, Para.[0009]-[0013]). Claim 13 is system claim performing the steps in method claim 4 above and as such, claim 13 is similar in scope and content to claim 4 and therefore, claim 13 is rejected under similar rationale as presented against claim 4 above. Claim 18 is non-transitory storage medium claim performing the steps in method claim 4 above and as such, claim 18 is similar in scope and content to claim 4 and therefore, claim 18 is rejected under similar rationale as presented against claim 4 above. Regarding Claim 5, Li in view of Sivaraman teach the method of claim 4. Sivaraman further teaches, wherein determining the similarity comprises determining a cosine similarity between the first voiceprint and the voiceprint of the target speaker ( Sivaraman: Para.[0072], Fig .1B, the speaker embedding engine 126 generates a similarity score between the inbound voiceprint and the enrolled voiceprint, which is a cosine similarity). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Sivaraman’s teaching of a machine-learning architecture system that enhances the speech audio of a user-defined target speaker by suppressing interfering speakers, as well as background noise and reverberations , into the system and method of performing source separation on mixed source single-channel and multichannel audio signals enhanced by inputting lip motion information from captured image data , taught by Li, because, this would improve the voice quality of a target speaker on a single channel audio input containing a mixture of speaker speech signals and various types of noise.(Sivaraman, Para.[0009]-[0013]). Claims 6, 14, 19 are rejected under 35 U.S.C. 103 as being unpatentable over Li et al. ( US 20210217182 A1), hereinafter referenced as Li, in view of Sivaraman et al. (US 20220084509 A1), hereinafter referenced as Sivaraman, further in view of Grangier et al. (US 20220375492 A1), hereinafter referenced as Grangier. Regarding Claim 6, Li in view of Sivaraman teach the method of claim 1. Sivaraman further teaches, wherein separating the voice of the target speaker from the voice mixture comprises: identifying a target speech signal corresponding to the voice of the target speaker based on the voiceprint of the target speaker ( Sivaraman: Para.[0109],[0110], Fig. 4, speech activity engine 404 may use an intermediate representation like short-time fourier transform (STFT). By applying the speech enhancement engine 402 on the input audio signal and an enrolled voiceprint ca identify target speech signal) ; and outputting the target speech signal ( Sivaraman: Para.[0109], Fig. 4, generates and outputs the enhanced audio signal ( target signal)). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Sivaraman’s teaching of a machine-learning architecture system that enhances the speech audio of a user-defined target speaker by suppressing interfering speakers, as well as background noise and reverberations , into the system and method of performing source separation on mixed source single-channel and multichannel audio signals enhanced by inputting lip motion information from captured image data , taught by Li, because, this would improve the voice quality of a target speaker on a single channel audio input containing a mixture of speaker speech signals and various types of noise.(Sivaraman, Para.[0009]-[0013]). Li in view of Sivaraman while teaching the method of claim 6, fail to explicitly teach the claimed, decomposing the voice mixture into a plurality of speech signals, each speech signal associated with a different speaker of the plurality of speakers in the second audio frame. However, Grangier does teach the claimed, decomposing the voice mixture into a plurality of speech signals, each speech signal associated with a different speaker of the plurality of speakers in the second audio frame ( Grangier: Para.[0028], Fig. 1, system 200 encodes the input audio signal 122 which corresponds to the captured utterances 120 from the multiple speakers, into a sequence of T temporal embeddings 220, 220a-t and iteratively selects a respective speaker embedding 240, 240a-n for each respective speaker); Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Grangier’s teaching of end-to-end speaker diarization via iterative speaker embedding , into the system and method, taught by Li in view of Sivaraman, because, this would produce more robust and accurate speaker diarization results in real-time scenarios.(Grangier, Para.[0024]-[0026]). Claim 14 is system claim performing the steps in method claim 6 above and as such, claim 14 is similar in scope and content to claim 6 and therefore, claim 14 is rejected under similar rationale as presented against claim 6 above. Claim 19 is non-transitory storage medium claim performing the steps in method claim 6 above and as such, claim 19 is similar in scope and content to claim 6 and therefore, claim 19 is rejected under similar rationale as presented against claim 6 above. Claims 7-10, 15, 20 are rejected under 35 U.S.C. 103 as being unpatentable over Li et al. ( US 20210217182 A1), hereinafter referenced as Li, in view of Sivaraman et al. (US 20220084509 A1), hereinafter referenced as Sivaraman, further in view of Wang et al. (US 20220301573 A1), hereinafter referenced as Wang. Regarding Claim 7, Li in view of Sivaraman teach the method of claim 1. Li in view of Sivaraman fail to explicitly teach the claimed, wherein the trained speech separation ML model or the target speaker voice-activity detection ("VAD") ML model comprise a plurality of one-dimensional convolutional neural networks. However, Wang does teach the claimed, wherein the trained speech separation ML model or the target speaker voice-activity detection ("VAD") ML model comprise a plurality of one-dimensional convolutional neural networks ( Wang: Para.[0050],[0054], [0065], Fig. 3, voice filter model 112 (trained speech separation ML model) can be trained to process a frequency representation of an audio signal as well as a speaker embedding corresponding to a human speaker to generate a predicted mask, where the frequency representation can be processed with the predicted mask to generate a revised frequency representation isolating utterance(s) of the human speaker and can be a neural network model and can include a convolutional neural network portion. convolutional neural network (CNN) portion 314 of voice filter model 112. CNN portion 314 is a one-dimensional convolutional neural network). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Wang’s teaching of isolating a human voice from a frequency representation of an audio signal by generating a predicted mask using a trained voice filter model, into the system and method, taught by Li in view of Sivaraman, because, by mitigating suppression of the target human speaker in the audio data can directly result in improved speech recognition performance by an ASR system the processes the audio data to generate the speech recognition.(Wang, Para.[0011], [0015]). Claim 15 is system claim performing the steps in method claim 7 above and as such, claim 15 is similar in scope and content to claim 7 and therefore, claim 15 is rejected under similar rationale as presented against claim 7 above. Claim 20 is non-transitory storage medium claim performing the steps in method claim 7 above and as such, claim 20 is similar in scope and content to claim 7 and therefore, claim 20 is rejected under similar rationale as presented against claim 7 above. Regarding Claim 8, Li in view of Sivaraman, further in view of Wang teach the method of claim 7. Wang further teaches, wherein the trained speech separation ML model or the target speaker voice-activity detection ("VAD") ML model comprises a plurality of network blocks, each network block comprises one or more convolutional blocks, and each convolutional block comprises a one or more neural networks ( Wang: Para.[0054], Figs 2. 3, The voice filter model 112 can be a neural network model and can include a convolutional neural network portion, a recurrent neural network portion, a fully connected feed forward neural network portion, and/or additional neural network layers). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Wang’s teaching of isolating a human voice from a frequency representation of an audio signal by generating a predicted mask using a trained voice filter model, into the system and method, taught by Li in view of Sivaraman, because, by mitigating suppression of the target human speaker in the audio data can directly result in improved speech recognition performance by an ASR system the processes the audio data to generate the speech recognition.(Wang, Para.[0011], [0015]). Regarding Claim 9, Li in view of Sivaraman, further in view of Wang teach the method of claim 7. Wang further teaches, wherein the trained speech separation ML model or the target speaker voice-activity detection ("VAD") ML model comprises a plurality of network blocks, wherein outputs of one or more convolutional blocks in a network block are summed and input to a next network block of the plurality of network blocks ( Wang: Para.[0065], Fig. 3, convolutional output generated by the CNN portion 314, can be applied as input to a recurrent neural network (RNN) portion 316 of voice filter model 112. RNN portion 316 can include unidirectional memory units (e.g., long short term memory units (LSTM), gated recurrent units (GRU), and/or additional memory unit(s)). RNN output generated by the RNN portion 316 can be applied as input to a fully connected feed-forward neural network portion 320 of voice filter model 112 to generate predicted mask 322). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Wang’s teaching of isolating a human voice from a frequency representation of an audio signal by generating a predicted mask using a trained voice filter model, into the system and method, taught by Li in view of Sivaraman, because, by mitigating suppression of the target human speaker in the audio data can directly result in improved speech recognition performance by an ASR system the processes the audio data to generate the speech recognition.(Wang, Para.[0011], [0015]). Regarding Claim 10, Li in view of Sivaraman, further in view of Wang teach the method of claim 9. Wang further teaches, wherein the sum of the outputs of the one or more convolutional blocks in the network block are fused with an embedding of the respective input audio frame prior to inputting the sum to the next network block ( Wang: Para.[0065], Fig. 3, Frequency representation 302 can be applied as input to a convolutional neural network (CNN) portion 314 of voice filter model 112. Convolutional output generated by the CNN portion 314, as well as speaker embedding 318, can be applied as input to a recurrent neural network (RNN) portion 316 of voice filter model 112). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Wang’s teaching of isolating a human voice from a frequency representation of an audio signal by generating a predicted mask using a trained voice filter model, into the system and method, taught by Li in view of Sivaraman, because, by mitigating suppression of the target human speaker in the audio data can directly result in improved speech recognition performance by an ASR system the processes the audio data to generate the speech recognition.(Wang, Para.[0011], [0015]). Conclusion Listed below are the prior arts made of record and not relied upon but are considered pertinent to applicant's disclosure. Braga et al. (US 20210118427 A1) teaches a single audio-visual automated speech recognition model for transcribing speech from audio-visual data includes an encoder frontend and a decoder. The encoder includes an attention mechanism configured to receive an audio track of the audio-visual data and a video portion of the audio-visual data. The video portion of the audio-visual data includes a plurality of video face tracks each associated with a face of a respective person. For each video face track of the plurality of video face tracks, the attention mechanism is configured to determine a confidence score indicating a likelihood that the face of the respective person associated with the video face tack includes a speaking face of the audio track. The decoder is configured to process the audio track and the video face track of the plurality of video face tracks associated with the highest confidence score to determine a speech recognition result of the audio track. Wang et al. (US 20210407510 A1) teaches a computer-implemented method which includes analyzing, by a speech detection system, a media file to detect lip movement of a speaker who is visually rendered in media content of the media file. The method additionally includes identifying, by the speech detection system, audio content within the media file, and improving accuracy of a temporal correlation of the speech detection system. The method may involve correlating the lip movement of the speaker with the audio content, and determining, based on the correlation between the lip movement of the speaker and the audio content, that the audio content comprises speech from the speaker. The method may further involve recording, based on the determination that the audio content comprises speech from the speaker, the temporal correlation between the speech and the lip movement of the speaker as metadata of the media file. Various other methods, systems, and computer-readable media are disclosed. Zhang et al. (US 20210390970 A1) teaches a method, computer program, and computer system for separating a target voice from among a plurality of speakers. Video data associated with the plurality of speakers and audio data associated with each of the one or more speakers are received. Video feature data is extracted from the received video data. The target voice is identified from among the plurality of speakers based on the received audio data and the extracted video feature data.. Khoury et al. (US 20210326421 A1 ) teaches a voice biometrics system execute machine-learning architectures capable of passive, active, continuous, or static operations, or a combination thereof. Systems passively and/or continuously, in some cases in addition to actively and/or statically, enrolling speakers as the speakers speak into or around an edge device (e.g., car, television, radio, phone). The system identifies users on the fly without requiring a new speaker to mirror prompted utterances for reconfiguring operations. The system manages speaker profiles as speakers provide utterances to the system. Machine-learning architectures implement a passive and continuous voice biometrics system, possibly without knowledge of speaker identities. The system creates identities in an unsupervised manner, sometimes passively enrolling and recognizing known or unknown speakers. The system offers personalization and security across a wide range of applications, including media content for over-the-top services and IoT devices (e.g., personal assistants, vehicles), and call centers. Any inquiry concerning this communication or earlier communications from the examiner should be directed to NADIRA SULTANA whose telephone number is (571)272-4048. The examiner can normally be reached M-F,7:30 am-5:00pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Paras D. Shah can be reached on (571) 270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /NADIRA SULTANA/Examiner, Art Unit 2653
Read full office action

Prosecution Timeline

Feb 03, 2025
Application Filed
Sep 01, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737534
SYSTEMS AND METHODS FOR LANGUAGE MODEL-BASED TEXT EDITING
2y 5m to grant Granted Sep 15, 2026
Patent 12731599
ADAPTIVE NOISE ESTIMATION
3y 6m to grant Granted Sep 08, 2026
Patent 12726544
COORDINATING A CONVERSATIONAL AGENT WITH A LARGE LANGUAGE MODEL FOR CONVERSATION REPAIR
2y 7m to grant Granted Sep 01, 2026
Patent 12681967
SYSTEM AND METHOD FOR OPTIMIZING QUERY RESOLUTION ON DIGITAL CHANNELS IN A CONTACT CENTER
2y 3m to grant Granted Jul 14, 2026
Patent 12676157
THREE-DIMENSIONAL AUDIO SIGNAL CODING METHOD AND APPARATUS, AND ENCODER
2y 7m to grant Granted Jul 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
73%
Grant Probability
99%
With Interview (+35.7%)
2y 11m (~1y 3m remaining)
Median Time to Grant
Low
PTA Risk
Based on 110 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month