Prosecution Insights
Last updated: October 02, 2026
Application No. 19/047,204

AUDIO SOURCE SEPARATION USING MULTI-MODAL AUDIO SOURCE CHANNALIZATION SYSTEM

Non-Final OA §103
Filed
Feb 06, 2025
Priority
Feb 08, 2024 — provisional 63/551,121
Examiner
SUBRAMANI, NANDINI
Art Unit
Tech Center
Assignee
Shure Acquisition Holdings Inc.
OA Round
1 (Non-Final)
65%
Grant Probability
Moderate
1-2
OA Rounds
1y 4m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 65% of resolved cases
65%
Career Allowance Rate
64 granted / 99 resolved
+4.6% vs TC avg
Strong +48% interview lift
Without
With
+47.5%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
15 currently pending
Career history
114
Total Applications
across all art units

Statute-Specific Performance

§101
12.9%
-27.1% vs TC avg
§103
65.3%
+25.3% vs TC avg
§102
11.1%
-28.9% vs TC avg
§112
8.5%
-31.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 99 resolved cases

Office Action

§103
DETAILED ACTION Introduction Applicant's submission filed on 02/06/2025 has been entered. Claims 1-10 and 63-72 are pending in the application and have been examined. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-7 and 9-10 are rejected under 35 U.S.C. 103 as being unpatentable over Shaked et. al. US PgPub. 2021/0312915 (cited in IDS) in view of Mosseri, et. al. US PgPub 2020/0335121. Regarding claim 1, Shaked teaches an audio signal processing apparatus comprising one or more processors and one or more memories storing instructions that are operable, when executed by the one or more processors, to cause the audio signal processing apparatus to: input a multi-source audio signal sample associated with at least one audio capture device to an audio feature extraction model that is configured to generate one or more isolate source audio features from the multi-source audio signal sample(see Shaked, [0028, 0029, 0040] discusses microphones 110 ( audio capture devices) connected to an echo cancellation engine 140, an audio-only separation engine 150, an audio-visual separation engine 160 ( separation engine) ( as shown in Shaked, Fig. 1)); input a video signal sample associated with at least one video capture device to a video feature extraction model that is configured to generate one or more speaker source features from the video signal sample (see Shaked, [0032,0036-0038] discusses the camera or cameras 120 may record, capture, or otherwise sample images, the audio preprocessing engine 111 is configured to isolate specific audio channels based on visual gesture detection, the video input includes detected faces( speaker source features), the engine 111 may be configured to calculate the angle and position of each face relative to the microphone or microphones , Faces may be detected using facial recognition methods ); input the one or more isolate source audio features and one or more speaker source features to a multi-modal audio source channelization model that is configured to generate a source separated channel audio sample based at least in part on the one or more isolate source audio features and on the one or more speaker source features(see Shaked, [0043] The audio-visual separation engine 160 allows for the separation of audio inputs into component channels, by audio source, based on video data); and output the source separated channel audio sample to one or more audio output devices (see Shaked, [0045] he audio-only separation engine 150 is configured to separate audio inputs into component channels, wherein each component channel reflects one audio source captured in the audio input. [0060, 0062] The audio-only separation engine 150 may be configured to separate an audio input 410 into at least one audio-only separation vocal output 420 and at least one audio-only separation remainder output 421. In a further embodiment, when the human voice is included in the audio input 410, the audio separation engine may be configured to split the audio input 410 into two channels, wherein a single channel includes a specific human voice, and where the remaining channel includes the contents of the audio input 410 without the selected specific human voice ). Shaked teaches input the one or more isolate source audio features and one or more speaker source features to a multi-modal audio source channelization model that is configured to generate a source separated channel audio sample based at least in part on the one or more isolate source audio features and on the one or more speaker source features based on location based selection, Mosseri further teaches input the one or more isolate source audio features and one or more speaker source features to a multi-modal audio source channelization model that is configured to generate a source separated channel audio sample based at least in part on the one or more isolate source audio features and on the one or more speaker source features (see Mosseri, [0051, 0053, 0056] In the visual stream, the system 100 extracts visual features 120 from the stream of frames 107 for the respective speakers 110A-B in the video 105. In the audio stream, the system 100 extracts audio features 125 from the audio soundtrack 115 for the respective speakers 110A-B in the video 105. The system 100 combines the visual features 120 and the audio features 125 of the video 105 to obtain joint audio-visual embeddings 130 of the video 105. A joint audio-visual embedding represents both audio features of a soundtrack and visual features of a face; Mosseri, Fig. 1); and output the source separated channel audio sample to one or more audio output devices(see Mosseri, [0056] The system 100 generates, from the processed audio-visual embeddings 130 of the video 105, an isolated speech signal 140A for speaker A 110A and an isolated speech signal 140B for speaker B 110B). Shaked and Mosseri are considered to be analogous to the claimed invention because both relate to multi modal audio source separation. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Shaked to audio separation based on image processing with audio-visual speech separation using neural system teachings of Mosseri to produce an isolated speech signal for each speaker, in which only the speech of the respective speaker can be heard (see Mosseri, [0007]). Regarding claim 2, Shaked in view of Mosseri teaches the audio signal processing apparatus of claim 1. Shaked further teaches wherein the multi-source audio signal sample is captured by multiple audio capture devices over multiple audio channels (see Shaked, [0028, 0036] discuss different microphones/arrays, accept one or more audio sources ). Regarding claim 3, Shaked in view of Mosseri teaches the audio signal processing apparatus of claim 1. Shaked further teaches wherein the multi-source audio signal sample is captured by one audio capture device over a single audio channel (see Shaked, [0035, 0036, 0037] discuss different microphones/arrays, accept one or more audio sources, known audio sources( single audio channel) ). Regarding claim 4, Shaked in view of Mosseri teaches the audio signal processing apparatus of claim 1. Mosseri further teaches wherein the multi-modal audio source channelization model is configured to further generate a second source separated channel audio sample based at least in part on the one or more isolate source audio features and on the one or more speaker source features, and output the second source separated channel audio sample to one or more audio output devices (see Mosseri, [0056] The system 100 generates, from the processed audio-visual embeddings 130 of the video 105, an isolated speech signal 140A for speaker A 110A and an isolated speech signal 140B for speaker B 110B). The same motivation to combine as indicated in claim 1 applies here. Regarding claim 5, Shaked in view of Mosseri teaches the audio signal processing apparatus of claim 4. Mosseri further teaches wherein the source separated channel audio sample is associated with a first target audio source and the second source separated channel audio sample is associated with a second target audio source (see Mosseri, [0056, 0090, 0091] The system 100 generates, from the processed audio-visual embeddings 130 of the video 105, an isolated speech signal 140A for speaker A 110A and an isolated speech signal 140B for speaker B 110B; he system determines, from the audio-visual embedding for the video, a respective spectrogram mask of the one or more speakers (step 312). As described above with reference to FIG. 2, the system can process the audio-visual embedding for the video through a masking neural network, where the masking neural network generates a respective spectrogram mask for the one or more speakers(first, second target audio sample).The system determines, from the respective spectrogram masks and the corresponding audio soundtrack, a respective isolated speech signal for each speaker(first, second target audio sample) (step 314).). The same motivation to combine as indicated in claim 1 applies here. Regarding claim 6, Shaked in view of Mosseri teaches the audio signal processing apparatus of claim 1. Shaked further teaches wherein the one or more speaker source features generated by the video feature extraction model comprise speaking indications associated with human facial images of the video signal sample(see Shaked, [0038, 0076, 0077, Fig. 9] discusses faces may be detected using facial recognition methods including mouth position detector ( 930) as the initiation of a person to talk is an example for gesture-based activation based on facial features). Regarding claim 7, Shaked in view of Mosseri teaches the audio signal processing apparatus of claim 1. Shaked further teaches wherein the one or more speaker source features generated by the video feature extraction model comprise speaking indications associated with human gesture images of the video signal sample (see Shaked, [0076, 0077, 0107, fig. 9] discusses the initiation of a person to talk is an example for gesture-based activation. The gesture detector 940 is configured to detect human gestures.). Regarding claim 8, Shaked in view of Mosseri teaches the audio signal processing apparatus of claim 1. Shaked further teaches wherein the audio signal processing apparatus is further configured to input the multi-source audio signal sample to a source spatialization model to generate isolate source spatial input data associated with the one or more isolate source audio features(see Shaked, [0029, 0037] discusses the audio preprocessing engine 111 may be configured to generate a beamformer, normalizing and consolidating the audio inputs from the microphones 110 to capture audio originating from specific sources, positions, or angles), and wherein the one or more isolate source audio features, the isolate source spatial input data, and the one or more speaker source features are input to the multi-modal audio source channelization model to generate the source separated channel audio sample (see Shaked, [0030, 0042, 0085, 0086, 0094, 0095], Fig. 7) discusses based on the detection and tracking, the engine 170 may be configured to direct a beamformer generated by the audio preprocessing engine 111 to the speaker and not to background noise. For example, if there are a number of passengers in a car, and a passenger in the back seat speaks, the engine 170 would direct the beamformer to the speaking passenger and not to the driver. In this way, the voice separations performed by the engines 130 through 160 may be made more accurate. The 3D microphone input and angle positions are calculated and the beamformer is used to separate channels). Regarding claim 9, Shaked in view of Mosseri teaches the audio signal processing apparatus of claim 8. Shaked further teaches wherein the source separated channel audio sample comprises isolate source spatial output data (see Shaked, [0095] At S760, channels are separated by their contents. The separation of channels may include some or all aspects of the audio-visual separation engine 160 described in FIG. 5, below, the audio-only separation engine 150 described in FIG. 4, below, other, like, separation methods or systems, and any combination thereof. In an embodiment, the separation of channels by their contents may include the isolation of sounds from their individual sources.). Regarding claim 10, Shaked in view of Mosseri teaches the audio signal processing apparatus of claim 1. Shaked further teaches wherein the video signal sample is a multi-feed video signal sample associated with two or more video capture devices (see Shaked, [0032, 0061] discusses the multiple cameras 120 types, arrangements). Regarding claim 63, is directed to a method claim corresponding to the apparatus claim presented in claim 1 and is rejected under the same grounds stated above regarding claim 1. Regarding claim 64, is directed to a method claim corresponding to the apparatus claim presented in claim 2 and is rejected under the same grounds stated above regarding claim 2. Regarding claim 65, is directed to a method claim corresponding to the apparatus claim presented in claim 3 and is rejected under the same grounds stated above regarding claim 3. Regarding claim 66, is directed to a method claim corresponding to the apparatus claim presented in claim 4 and is rejected under the same grounds stated above regarding claim 4. Regarding claim 67, is directed to a method claim corresponding to the apparatus claim presented in claim 5 and is rejected under the same grounds stated above regarding claim 5. Regarding claim 68, is directed to a method claim corresponding to the apparatus claim presented in claim 6 and is rejected under the same grounds stated above regarding claim 6. Regarding claim 69, is directed to a method claim corresponding to the apparatus claim presented in claim 7 and is rejected under the same grounds stated above regarding claim 7. Regarding claim 70, is directed to a method claim corresponding to the apparatus claim presented in claim 8 and is rejected under the same grounds stated above regarding claim 8. Regarding claim 71, is directed to a method claim corresponding to the apparatus claim presented in claim 9 and is rejected under the same grounds stated above regarding claim 9. Regarding claim 72, is directed to a method claim corresponding to the apparatus claim presented in claim 10 and is rejected under the same grounds stated above regarding claim 10. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Gao, R., et. al. (2021, June). Visualvoice: Audio-visual speech separation with cross-modal consistency. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 15490-15500) teaches leveraging the speaker’s face appearance as an additional prior to isolate the corresponding vocal qualities speakers are likely to produce (see Gao, abstract). Nefian et al US PgPub. 2005/0228673 teaches The audio and video are separated and corresponding portions of the audio are mapped to the visual features for purposes of isolating audio associated with each speaker and for purposes of filtering out noise associated with the audio (see Nefian, abstract). Any inquiry concerning this communication or earlier communications from the examiner should be directed to NANDINI SUBRAMANI whose telephone number is (571)272-3916. The examiner can normally be reached Monday - Friday 12:00pm - 5:00 pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bhavesh M Mehta can be reached at (571)272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /NANDINI SUBRAMANI/ Examiner, Art Unit 2656
Read full office action

Prosecution Timeline

Feb 06, 2025
Application Filed
Sep 04, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749498
QUALITY ESTIMATION MODEL FOR PACKET LOSS CONCEALMENT
3y 9m to grant Granted Sep 29, 2026
Patent 12731575
Attention-Based Joint Acoustic and Text On-Device End-to-End Model
3y 7m to grant Granted Sep 08, 2026
Patent 12700416
Audio Transcoding Method and Apparatus, Audio Transcoder, Device, and Storage Medium
3y 9m to grant Granted Aug 04, 2026
Patent 12688863
EMOTIONALLY-AWARE VOICE RESPONSE GENERATION METHOD AND APPARATUS
4y 7m to grant Granted Jul 21, 2026
Patent 12670912
ATTENTIVE SCORING FUNCTION FOR SPEAKER IDENTIFICATION
2y 9m to grant Granted Jun 30, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
65%
Grant Probability
99%
With Interview (+47.5%)
3y 0m (~1y 4m remaining)
Median Time to Grant
Low
PTA Risk
Based on 99 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month