Prosecution Insights
Last updated: October 02, 2026
Application No. 19/075,271

METHOD OF SEPARATING SOUND SOURCE FROM AUDIO SIGNAL AND ELECTRONIC DEVICE FOR PERFORMING THE SAME

Non-Final OA §101§102§103§112
Filed
Mar 10, 2025
Priority
Feb 26, 2024 — RE 10-2024-0027505 +3 more
Examiner
SHIN, SEONG-AH A
Art Unit
Tech Center
Assignee
Samsung Electronics Co., Ltd.
OA Round
1 (Non-Final)
78%
Grant Probability
Favorable
1-2
OA Rounds
1y 0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 78% — above average
78%
Career Allowance Rate
332 granted / 423 resolved
+18.5% vs TC avg
Strong +21% interview lift
Without
With
+21.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
23 currently pending
Career history
447
Total Applications
across all art units

Statute-Specific Performance

§101
23.7%
-16.3% vs TC avg
§103
46.5%
+6.5% vs TC avg
§102
14.0%
-26.0% vs TC avg
§112
6.7%
-33.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 423 resolved cases

Office Action

§101 §102 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Status of Claims Claims 1-15 are pending in this application. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 3 and 12 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, or for pre-AIA the applicant regards as the invention. Claims 3 and 12 recite the limitation "the clustering" in line 3. There is insufficient antecedent basis for this limitation in the claim. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefore, subject to the conditions and requirements of this title. Claims 1-15 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The independent claim 1 recites “obtaining the audio signal including sound generated by a plurality of sound sources; based on a single sound source segment including only a sound generated by one sound source from among the plurality of sound sources, obtaining an embedding corresponding to at least one primary sound source from among the plurality of sound sources; and separating the at least one primary sound source from the audio signal, based on the obtained embedding”. The limitation of “obtaining…”, “obtaining…” and “separating” is a process that, under its broadest reasonable interpretation, could be performed in the human mind and requires no more than a performing of generic computer functions (e.g. collecting data, calculating). More specifically, a human may listen to sound and make judgments based on inputs such as what he/she hears. This judicial exception is not integrated into a practical application The computer is recited at a high-level of generality (i.e., as performing a generic computer function and being used as an applying) such that it amounts no more than mere instructions to apply the exception using a generic computer. Accordingly, there additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, the additional element of using a computer amounts to no more than mere instructions to apply an exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The claim is not patent eligible. With respect to claim 10, the claim is similar to claim 1 and claim 10 recites additional element of “processor” and “memory”. The processor and memory are recited at a high-level of generality (i.e., as a generic processor performing generic computer functions and being used as an applying) such that it amounts no more than mere instructions to apply the exception using a generic computer component as well. These claims further do not remedy the judicial exception being integrated into a practical application and further fail to include additional elements that are sufficient to amount to significantly more than the judicial exception. With respect to dependent claims 2 and 11, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. With respect to dependent claims 3 and 12, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. With respect to dependent claims 4 and 13, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. With respect to dependent claim 5, the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. With respect to dependent claims 6 and 14, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. With respect to dependent claim 7, the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. With respect to dependent claim 8, the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. With respect to dependent claims 9 and 15, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. Therefore, claims 1-15 are rejected. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention. Claims 1-3 and 10-12 are rejected under 35 U.S.C. 102 (a)(1) as being anticipated by Reshef et al., (US 2019/0318743 A1). Regarding claim 1, Reshef discloses a method of separating a sound source from an audio signal, the method comprising: obtaining the audio signal including sound generated by a plurality of sound sources (Fig. 2, step 50 and Fig. 3A, [0053][0056] obtaining an audio signal 52 which is generated by a plurality of sound sources); based on a single sound source segment including only a sound generated by one sound source from among the plurality of sound sources, obtaining an embedding corresponding to at least one primary sound source from among the plurality of sound sources (Fig. 2, steps 66, 72 and Fig. 3A, [0053][0055][0056] identifying segments 54 and 56 based on the metadata as belonging to participants 30 and 33; [0031][0060][0064] extracting the acoustic features from the speech segments which is defined using embedding); and separating the at least one primary sound source from the audio signal, based on the obtained embedding (Fig. 2, steps 96,98 and Fig. 3B, separating the primary sound by removing ambiguities due to background noise or both participants speaking at once). Regarding claim 2, Reshef discloses the method of claim 1, Reshef further discloses: wherein the obtaining the embedding corresponding to the at least one primary sound source, comprises: segmenting the audio signal into a plurality of frames ([0053][0064] dividing the audio recording into a series of T time frames and extracting acoustic features from the audio signal in each time frame); determining, as the single sound source segment, each of frames in which only the sound generated by the one sound source is activated among the plurality of frames (Figs. 3A and 3B, [0053]-[0056] labeling each frame belonging to participants 30 and 33); for each single sound source segment, obtaining an embedding including features of the sound activated in the single sound source segment (Figs. 3A and 3B, extracting acoustic features from the speech segments that were labeled); and determining an embedding corresponding to each sound source by performing clustering on the embeddings (Figs. 3A and 3B, [0054] filtering the raw metadata received from the conferencing data stream to remove ambiguities and gaps and merging adjacent speaker labels and small gaps between labels). Regarding claim 3, Reshef discloses the method of claim 1, Reshef further discloses: wherein the obtaining the embedding corresponding to the at least one primary sound source, further comprises: based on a result of the clustering, identifying an active segment for each sound source; aid based on a length of the active segment, determining, as the at least one primary sound source, at least one sound source from among the plurality of sound sources (Figs. 3A and 3B, [0055] identifying as speech any segment in the audio stream and distinguish speech segments from noise). Regarding claims 10-12, Claims 10-12 are the corresponding system claims to method claims 1-3. Therefore, claims 10-12 are rejected using the same rationale as applied to claims 1-3 above. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action. The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 4-8 and 13-14 are rejected under pre-AIA 35 U.S.C. 103(a) as being unpatentable over Reshef et al., (US 2019/0318743 A1) in view of Grangier et al., (US 2022/0375492 A1). Regarding claim 4, Reshef discloses the method of claim 1. Reshef does not explicitly teach, however Grangier does explicitly teach: for each frame with a preset length, separating sound in the audio signal (Grangier, Figs. 1 and 2, [0010][0028][0033] “encoding the input audio signal into a sequence of T temporal embeddings. Each temporal embedding is associated with a corresponding time step and represents speech content extracted from the input audio signal at the corresponding time step”); obtaining an embedding corresponding to the separated sound; determining whether the embedding corresponding to the separated sound matches another embedding corresponding to the at least one primary sound source (Figs. 1 and 2, [0010][0028][0033]-[0037] “For each temporal embedding in the sequence of T temporal embeddings, the operations select the respect speaker embedding by determining a probability that the corresponding temporal embedding includes a presence of voice activity by a single new speaker for which a speaker embedding was not previously selected during a previous iteration”); and based on a result of the determining whether the embedding matches the another embedding, performing diarization on the at least one primary sound source (Fig. 1, [0010][0028][0033]-[0037] providing diarization results indicating when the voice of each speaker is active in the input audio signal). Therefore, it would have been obvious to one of ordinary skill before the effective filing date of the claimed invention to incorporate the Metadata-based diarization of multi-speaker conferences as taught by Reshef with the method of speech diarization via iterative speaker embedding as taught by Grangier to identify the respect speaker embedding by determining a probability that the corresponding temporal embedding in order to improve speech recognition on the audio data (Grangier, [0028][0029]). Regarding claim 5, Reshef in view of Grangier discloses the method of claim 4, and Grangier further discloses: wherein the diarization comprises an operation indicating an active segment of sound generated by the at least one primary sound source based on a time axis (Grangier, Fig. 1, [0028][0033]-[0037] diarization results provide time-stamped speaker labels based on the per-speaker voice activity indicators predicted at each time step). The previous motivation statement as in claim 4 is still applied. Regarding claim 6, Reshef discloses the method of claim 1. Reshef does not explicitly teach, however Grangier does explicitly teach: wherein the separating the at least one primary sound source, comprises: selecting, as a target embedding, one of embeddings corresponding to the at least one primary sound source; and separating, from the audio signal, sound matching the target embedding (Grangier, [0034] selecting the respective speaker embedding for the respective speaker as the temporal embedding in the sequence of T temporal embeddings that is associated with the highest probability for the presence of voice activity by the single new speaker). The previous motivation statement as in claim 4 is still applied. Regarding claim 7, Reshef in view of Grangier discloses the method of claim 6, and Grangier further discloses: wherein the separating, from the audio signal, the sound matching the target embedding, comprises: obtaining a segmented target frame from the audio signal; concatenating an audio signal of sound generated by a sound source corresponding to the target embedding to a front or back of an audio signal of the segmented target frame (Fig. 2, [0036] obtaining an input audio signal and encoding the input audio signal into the sequence of temporal embeddings each associated with a corresponding time step t); inputting, to a sound source separation module, the concatenated audio signal and the target embedding (Fig. 2, [0036][0037] inputting the sequence of temporal embedding into a speaker selector 230); and obtaining a separated sound from the sound source separation module (Fig. 2, [0034][0037] outputting the respective speaker embedding for each speaker detected in the input audio signal). The previous motivation statement as in claim 4 is still applied. Regarding claim 8, Reshef in view of Grangier discloses the method of claim 6, and Grangier further discloses: wherein the separating, from the audio signal, the sound matching the target embedding, comprises: obtaining a segmented target frame from the audio signal; summing an audio signal of sound generated by a sound source corresponding to the target embedding into an audio signal of the segmented target frame in a same time segment (Fig. 2, [0036] obtaining an input audio signal and encoding the input audio signal into the sequence of temporal embeddings each associated with a corresponding time step t); inputting, to a sound source separation module, the summed audio signal and the target embedding (Fig. 2, [0036][0037] inputting the sequence of temporal embedding into a speaker selector 230); and obtaining a separated sound from the sound source separation module (Fig. 2, [0034][0037] outputting the respective speaker embedding for each speaker detected in the input audio signal). The previous motivation statement as in claim 4 is still applied. Regarding claims 13-14, Claims 13-14 are the corresponding system claims to method claims 4 and 6. Therefore, claims 13-14 are rejected using the same rationale as applied to claims 4 and 6 above. Claims 9 and 15 are rejected under pre-AIA 35 U.S.C. 103(a) as being unpatentable over Reshef et al., (US 2019/0318743 A1) in view of Medalion et al., (US 2023/0260519 A1). Regarding claim 9, Reshef discloses the method of claim 1. Reshef does not explicitly teach, however Medalion does explicitly teach: wherein the at least one primary sound source comprises a speaker, and wherein the method further comprises: analyzing a lip motion of at least one person in a video segment corresponding to the single sound source segment (Fig. 8, [0051][0090]-[0094] analyzing a lip motion by applying a lip localization model and speech/lip movement identification model); and based on a result of the analyzing, selecting a person corresponding to the at least one primary sound source (Fig. 8, [0051][0090]-[0094] determining whether a portion of the transcript pertaining to the segment associated with the video fragment is associated with a particular speaker, or in the identification of a particular speaker). Therefore, it would have been obvious to one of ordinary skill before the effective filing date of the claimed invention to incorporate the Metadata-based diarization of multi-speaker conferences as taught by Reshef with the method of adapting vide analysis as taught by Medalion to increase accuracy of the diarization using lip speech identification (Medalion, [0108]). Regarding claim 15, Claim 15 are the corresponding system claim to method claim 9. Therefore, claim 15 is rejected using the same rationale as applied to claim 9 above. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Please see attached form PTO-892. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SEONG-AH A. SHIN whose telephone number is (571)272-5933. The examiner can normally be reached 9 AM-3PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre-Louis Desir can be reached at 571-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. Seong-ah A. Shin Primary Examiner Art Unit 2659 /SEONG-AH A SHIN/Primary Examiner, Art Unit 2659
Read full office action

Prosecution Timeline

Mar 10, 2025
Application Filed
Sep 14, 2026
Non-Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749479
COMMUNICATION SUPPORT SYSTEM, COMMUNICATION SUPPORT METHOD, AND NON-TRANSITORY RECORDING MEDIUM
2y 9m to grant Granted Sep 29, 2026
Patent 12725610
VOICE RECOGNITION SYSTEM, SERVER, DISPLAY APPARATUS AND CONTROL METHODS THEREOF
3y 11m to grant Granted Sep 01, 2026
Patent 12725609
Hotwording by Degree
2y 2m to grant Granted Sep 01, 2026
Patent 12694220
METHOD AND SYSTEM FOR PERSONALIZED EMBEDDING SEARCH ENGINE
3y 5m to grant Granted Jul 28, 2026
Patent 12682898
KEY PHRASE SPOTTING
2y 3m to grant Granted Jul 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
78%
Grant Probability
99%
With Interview (+21.3%)
2y 7m (~1y 0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 423 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month