Prosecution Insights
Last updated: August 17, 2026
Application No. 18/521,606

METHOD PERFORMED BY ELECTRONIC DEVICE AND APPARATUS

Non-Final OA §103
Filed
Nov 28, 2023
Priority
Nov 28, 2022 — CN 202211505381.X
Examiner
PULLIAS, JESSE SCOTT
Art Unit
2655
Tech Center
2600 — Communications
Assignee
Samsung Electronics Co., Ltd.
OA Round
3 (Non-Final)
83%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
95%
With Interview

Examiner Intelligence

Grants 83% — above average
83%
Career Allowance Rate
883 granted / 1069 resolved
+20.6% vs TC avg
Moderate +13% lift
Without
With
+12.7%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
37 currently pending
Career history
1105
Total Applications
across all art units

Statute-Specific Performance

§101
15.5%
-24.5% vs TC avg
§103
52.9%
+12.9% vs TC avg
§102
19.9%
-20.1% vs TC avg
§112
4.7%
-35.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1069 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 01/12/26 has been entered. This office action is in response to correspondence 01/12/26 regarding application 18/521,606, in which claims 1, 11-13, 16, and 18-20 were amended. Claims 1, 2, 4-13, and 15-20 are pending in the application and have been considered. Response to Arguments Amended claim 18 overcomes the objection for being dependent on a cancelled claim, and so the objection is withdrawn. Specifically, Applicant amended claim 18 so that it depends on claim 13 instead of cancelled claim 14. Applicant’s arguments on pages 12-15 regarding the 35 U.S.C. 103 rejections based on Tu and Bi have been considered but are not persuasive. Applicant first argues that the prior art does not disclose "determining a corresponding target audio segment for a first audio block by comparing speech quality between an audio segment of the first audio block and a target audio segment determined for a second audio block, wherein the second audio block precedes the first audio block; and performing speech separation on the audio signal comprising a first audio segment of the first audio block, based on a hidden state information of the corresponding target audio segment for the first audio block and a hidden state information of a second audio segment preceding the first audio segment, to obtain a plurality of separated speech signals corresponding to the plurality of speakers," as recited in claim 1, allegedly because the model switching DNN approach for speech separation in Tu does not refer to hidden state information from a previous segment or block, and allegedly because Bi does not disclose separating speech signals corresponding to a plurality of speakers using DNNs where hidden states of previous segments are used for speech separation of subsequent segments, and the header_length in Bi is not a function of hidden states used for speech separation of a subsequent utterance. In response, one cannot show nonobviousness by attacking references individually where the rejections are based on combinations of references. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981); In re Merck & Co., 800 F.2d 1091, 231 USPQ 375 (Fed. Cir. 1986). performing speech separation on the audio signal based on hidden state information of the target audio to obtain a plurality of separated speech signals corresponding to the plurality of speakers (either the positive or negative DNN is used for second-pass speech separation; Section 3.2, page 63, Section 4.2, page 64, using target features and interfering features outputs from hidden layers, Fig. 3, page 63, to yield separated signals for each of the 34 speakers in the test set, Section 4, page 63). Tu does not specifically mention: at least one audio block among the plurality of audio bocks is divided into a plurality of audio segments; determining a corresponding target audio segment for the first audio block by comparing speech quality between an audio segment of a first audio block and a target audio segment determined for a second audio block, wherein the second audio block precedes the first audio block; performing speech separation on the audio signal comprising a first audio segment of the first audio block based on the corresponding target audio segment for the first audio block and a second audio segment preceding the first audio segment. Bi discloses at least one audio block among the plurality of audio bocks is divided into a plurality of audio segments (the blocks of audio between first and second starting points and first and second ending points, Col 6 lines 25-39, Fig. 3, which are divided into frames, Col 6 lines 45-55, e.g. ten milliseconds long, Col 5 lines 19-23); determining a corresponding target audio segment for the first audio block by comparing speech quality between an audio segment of a first audio block and a target audio segment determined for a second audio block, wherein the second audio block precedes the first audio block (for each frame, comparing the current SNR to thresholds “between” the starting and ending points of the segments shown in Fig 3 to look back for the actual start of speech; this segment between START and PRE_START precedes the utterance segment, and these frames are considered a “target” audio segment” based on their instantaneous SNR values, Fig. 2, Fig 3, Col 5 lines 39-53, Col 6 lines 15-24); performing speech separation on the audio signal comprising a first audio segment of the first audio block based on the corresponding target audio segment for the first audio block and a second audio segment preceding the first audio segment (the START and END endpoints are considered to separate the speech signal from the noise signal preceding and following the speech segment, Col 6 lines 23-40, Fig. 3). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Tu such that at least one audio block among the plurality of audio bocks is divided into a plurality of audio segments, and determining a corresponding target audio segment for the first audio block by comparing speech quality between an audio segment of a first audio block and a target audio segment determined for a second audio block, wherein the second audio block precedes the first audio block; and performing speech separation on the audio signal comprising a first audio segment of the first audio block, as in Bi, based on a hidden state information as in Tu of the corresponding target audio segment of Bi for the first audio block and a hidden state information as in Tu of a second audio segment preceding the first audio segment of Bi in order to avoid missing weak speech segments, as suggested by Bi (Col 9 lines 28-34), predictably increasing accuracy of speech endpointing, as suggested by Bi (Col 9 lines 28-34). On pages 14-15, Applicant further argues that a person of skill in the art would have no reason to combine Tu and Bi because they relate to different technical problems and propose different solutions to their respective technical problems. Specifically, as Applicant argues, “For example, Tu relates to second pass speech separation whereas Bi relates to utterance detection and not speech separation between speech of multiple speakers. A person of skill in the art, in order to combine Tu and Bi, would deviate from Bi's objective of utterance detection. Furthermore, the techniques used in Tu and Bi are distinct. Bi does not disclose any techniques related speech separation based on block-based temporal processing or incorporating hidden state information to guide speech separation for subsequent speech signals.” In response, the only difference between the technical problems solved by Tu and Bi is that Tu is primarily concerned with separating speech from a target speaker from interfering speech from another speaker, while Bi is primarily concerned with endpointing speech in the presence of noise. Although Bi gives the example of road noise as a background noise, Bi describes using adaptive thresholds to deal with variable noise levels (Bi, Col 4 lines 62-65), which is non-stationary noise, analogous to interfering speech. In other words, Bi separates desired speech segments from undesired background noise segments of varying intensity in the audio signal. Tu does the same thing, except that the undesired noise segments are speech from a non-target speaker rather than e.g. road noise. Those skilled in the art of speech processing at filing time would immediately recognize that background noise often includes undesired speech (e.g. a toddler speaking in the background during a call, or a TV program with spoken dialog in the background). Combining Tu and Bi would not deviate from Bi's objective of utterance detection for the reasons above, and the alleged differences in the techniques used in Tu and Bi are complementary for the reasons above. Claim Objections In claim 11, line 8, should “signal” be “signals”? Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 2, 8-13, and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Tu et al. (“SPEECH SEPARATION BASED ON SIGNAL-NOISE-DEPENDENT DEEP NEURAL NETWORKS FOR ROBUST SPEECH RECOGNITION”. ICASSP 2015) in view of Bi et al. (US 6324509). Consider claim 1, Tu discloses a method performed by an electronic device (an electronic device inherent for implementing the DNN architectures described at section 4, page 63), comprising: obtaining an audio signal comprising a speech signal uttered by a plurality of sound sources comprising a plurality of speakers, wherein the audio signal is divided into a plurality of audio blocks (audio signal from the test set of the SSC corpus consisting of two-speaker mixtures is divided into 512 sample frames, Section 4, page 63, including 500 utterances for each speaker, i.e. “an audio signal”); performing speech separation on the audio signal based on hidden state information of the target audio to obtain a plurality of separated speech signals corresponding to the plurality of speakers (either the positive or negative DNN is used for second-pass speech separation; Section 3.2, page 63, Section 4.2, page 64, using target features and interfering features outputs from hidden layers, Fig. 3, page 63, to yield separated signals for each of the 34 speakers in the test set, Section 4, page 63). Tu does not specifically mention: at least one audio block among the plurality of audio bocks is divided into a plurality of audio segments determining a corresponding target audio segment for the first audio block by comparing speech quality between an audio segment of a first audio block and a target audio segment determined for a second audio block, wherein the second audio block precedes the first audio block; performing speech separation on the audio signal comprising a first audio segment of the first audio block based on the corresponding target audio segment for the first audio block and a second audio segment preceding the first audio segment. Bi discloses at least one audio block among the plurality of audio bocks is divided into a plurality of audio segments (the blocks of audio between first and second starting points and first and second ending points, Col 6 lines 25-39, Fig. 3, which are divided into frames, Col 6 lines 45-55, e.g. ten milliseconds long, Col 5 lines 19-23) determining a corresponding target audio segment for the first audio block by comparing speech quality between an audio segment of a first audio block and a target audio segment determined for a second audio block, wherein the second audio block precedes the first audio block (for each frame, comparing the current SNR to thresholds “between” the starting and ending points of the segments shown in Fig 3 to look back for the actual start of speech; this segment between START and PRE_START precedes the utterance segment, and these frames are considered a “target” audio segment” based on their instantaneous SNR values, Fig. 2, Fig 3, Col 5 lines 39-53, Col 6 lines 15-24); performing speech separation on the audio signal comprising a first audio segment of the first audio block based on the corresponding target audio segment for the first audio block and a second audio segment preceding the first audio segment (the START and END endpoints are considered to separate the speech signal from the noise signal preceding and following the speech segment, Col 6 lines 23-40, Fig. 3). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Tu such that at least one audio block among the plurality of audio bocks is divided into a plurality of audio segments, and determining a corresponding target audio segment for the first audio block by comparing speech quality between an audio segment of a first audio block and a target audio segment determined for a second audio block, wherein the second audio block precedes the first audio block; and performing speech separation on the audio signal comprising a first audio segment of the first audio block, as in Bi, based on a hidden state information as in Tu of the corresponding target audio segment of Bi for the first audio block and a hidden state information as in Tu of a second audio segment preceding the first audio segment of Bi in order to avoid missing weak speech segments, as suggested by Bi (Col 9 lines 28-34), predictably increasing accuracy of speech endpointing, as suggested by Bi (Col 9 lines 28-34). The references cited are analogous art in the same field of speech processing. Consider claim 12, Tu discloses an electronic device comprising: at least one memory storing computer executable instructions; and at least one processor, when executing the stored instructions (an electronic device with a memory storing instructions and a processor is inherent for implementing the DNN architectures described at section 4, page 63), is configured to: obtain an audio signal comprising a speech signal uttered by a plurality of sound sources comprising a plurality of speakers, wherein the audio signal is divided into a plurality of audio blocks (audio signal from the test set of the SSC corpus consisting of two-speaker mixtures is divided into 512 sample frames, Section 4, page 63, including 500 utterances for each speaker, i.e. “an audio signal”); perform speech separation on the audio signal based on hidden state information of the target audio to obtain a plurality of separated speech signals corresponding to the plurality of speakers (either the positive or negative DNN is used for second-pass speech separation; Section 3.2, page 63, Section 4.2, page 64, using target features and interfering features outputs from hidden layers, Fig. 3, page 63, to yield separated signals for each of the 34 speakers in the test set, Section 4, page 63). Tu does not specifically mention: at least one audio block among the plurality of audio bocks is divided into a plurality of audio segments determining a corresponding target audio segment for the first audio block by comparing speech quality between an audio segment of a first audio block and a target audio segment determined for a second audio block, wherein the second audio block precedes the first audio block; performing speech separation on the audio signal comprising a first audio segment of the first audio block based on the corresponding target audio segment for the first audio block and a second audio segment preceding the first audio segment. Bi discloses at least one audio block among the plurality of audio bocks is divided into a plurality of audio segments (the blocks of audio between first and second starting points and first and second ending points, Col 6 lines 25-39, Fig. 3, which are divided into frames, Col 6 lines 45-55, e.g. ten milliseconds long, Col 5 lines 19-23) determining a corresponding target audio segment for the first audio block by comparing speech quality between an audio segment of a first audio block and a target audio segment determined for a second audio block, wherein the second audio block precedes the first audio block (for each frame, comparing the current SNR to thresholds “between” the starting and ending points of the segments shown in Fig 3 to look back for the actual start of speech; this segment between START and PRE_START precedes the utterance segment, and these frames are considered a “target” audio segment” based on their instantaneous SNR values, Fig. 2, Fig 3, Col 5 lines 39-53, Col 6 lines 15-24); performing speech separation on the audio signal comprising a first audio segment of the first audio block based on the corresponding target audio segment for the first audio block and a second audio segment preceding the first audio segment (the START and END endpoints are considered to separate the speech signal from the noise signal preceding and following the speech segment, Col 6 lines 23-40, Fig. 3). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Tu such that at least one audio block among the plurality of audio bocks is divided into a plurality of audio segments, and determining a corresponding target audio segment for the first audio block by comparing speech quality between an audio segment of a first audio block and a target audio segment determined for a second audio block, wherein the second audio block precedes the first audio block; and performing speech separation on the audio signal comprising a first audio segment of the first audio block, as in Bi, based on a hidden state information as in Tu of the corresponding target audio segment of Bi for the first audio block and a hidden state information as in Tu of a second audio segment preceding the first audio segment of Bi for reasons similar to those for claim 1. Consider claim 13, Tu discloses a non-transitory computer readable storage medium storing instructions that, when executed by at least one processor (a memory storing instructions executed by a processor is inherent for implementing the DNN architectures described at section 4, page 63), cause the at least one processor to: obtain an audio signal comprising a speech signal uttered by a plurality of sound sources comprising a plurality of speakers, wherein the audio signal is divided into a plurality of audio blocks (audio signal from the test set of the SSC corpus consisting of two-speaker mixtures is divided into 512 sample frames, Section 4, page 63, including 500 utterances for each speaker, i.e. “an audio signal”); perform speech separation on the audio signal based on hidden state information of the target audio to obtain a plurality of separated speech signals corresponding to the plurality of speakers (either the positive or negative DNN is used for second-pass speech separation; Section 3.2, page 63, Section 4.2, page 64, using target features and interfering features outputs from hidden layers, Fig. 3, page 63, to yield separated signals for each of the 34 speakers in the test set, Section 4, page 63). Tu does not specifically mention: at least one audio block among the plurality of audio bocks is divided into a plurality of audio segments determining a corresponding target audio segment for the first audio block by comparing speech quality between an audio segment of a first audio block and a target audio segment determined for a second audio block, wherein the second audio block precedes the first audio block; performing speech separation on the audio signal comprising a first audio segment of the first audio block based on the corresponding target audio segment for the first audio block and a second audio segment preceding the first audio segment. Bi discloses at least one audio block among the plurality of audio bocks is divided into a plurality of audio segments (the blocks of audio between first and second starting points and first and second ending points, Col 6 lines 25-39, Fig. 3, which are divided into frames, Col 6 lines 45-55, e.g. ten milliseconds long, Col 5 lines 19-23) determining a corresponding target audio segment for the first audio block by comparing speech quality between an audio segment of a first audio block and a target audio segment determined for a second audio block, wherein the second audio block precedes the first audio block (for each frame, comparing the current SNR to thresholds “between” the starting and ending points of the segments shown in Fig 3 to look back for the actual start of speech; this segment between START and PRE_START precedes the utterance segment, and these frames are considered a “target” audio segment” based on their instantaneous SNR values, Fig. 2, Fig 3, Col 5 lines 39-53, Col 6 lines 15-24); performing speech separation on the audio signal comprising a first audio segment of the first audio block based on the corresponding target audio segment for the first audio block and a second audio segment preceding the first audio segment (the START and END endpoints are considered to separate the speech signal from the noise signal preceding and following the speech segment, Col 6 lines 23-40, Fig. 3). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Tu such that at least one audio block among the plurality of audio bocks is divided into a plurality of audio segments, and determining a corresponding target audio segment for the first audio block by comparing speech quality between an audio segment of a first audio block and a target audio segment determined for a second audio block, wherein the second audio block precedes the first audio block; and performing speech separation on the audio signal comprising a first audio segment of the first audio block, as in Bi, based on a hidden state information as in Tu of the corresponding target audio segment of Bi for the first audio block and a hidden state information as in Tu of a second audio segment preceding the first audio segment of Bi for reasons similar to those for claim 1. Consider claim 2, Tu discloses the speech quality is identified based on at least one of speech distortion, signal-to-noise ratio, zero crossing rate, and pitch quantity (SNR, formula 3, page 63, Section 3.2). Consider claim 8, Tu discloses the signal-to-noise ratio is determined by calculating a ratio between a separated speech signal for an audio segment and an original audio signal corresponding to the audio segment (SNR, formula 3, page 63, Section 3.2). Consider claim 9, Tu discloses the performing speech separation on the audio signal comprising the first audio segment comprises: obtaining hidden layer state information of the target audio and the second audio (Fig 3, layer h1, page 63); fusing the hidden layer state information of the target audio and the second audio segment to obtain fused hidden layer state information (Fig 3, layers h2 and h3, page 63); and performing speech separation on the first audio segment based on the fused hidden layer state information (Fig 3, target features and interfering features outputs, page 63). Tu does not specifically mention determining a target audio segment. Bi discloses determining a target audio segment (for each frame, comparing the current SNR to thresholds “between” the starting and ending points of the segments shown in Fig 3 to look back for the actual start of speech; this segment between START and PRE_START precedes the utterance segment, and these frames are considered a “target” audio segment” based on their instantaneous SNR values, Fig. 2, Fig 3, Col 5 lines 39-53, Col 6 lines 15-24). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Tu by determining a target audio segment as in Bi for reasons similar to those for claim 1. Consider claim 10, Tu discloses the hidden layer state information is obtained when speech separation is performed on the target audio and the second audio segment respectively (Fig 3, mixed feature separation performed block by block based on power spectra features, Section 4, page 63); and the hidden layer state information comprises at least one of short-term speech features, long-term speech features and context features of each sound source (the power spectra features computed in 32 msec frames, therefore considered short term speech features processed by the three hidden layers, Section 4, page 63). Tu does not specifically mention determining a target audio segment. Bi discloses determining a target audio segment (for each frame, comparing the current SNR to thresholds “between” the starting and ending points of the segments shown in Fig 3 to look back for the actual start of speech; this segment between START and PRE_START precedes the utterance segment, and these frames are considered a “target” audio segment” based on their instantaneous SNR values, Fig. 2, Fig 3, Col 5 lines 39-53, Col 6 lines 15-24). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Tu by determining a target audio segment as in Bi for reasons similar to those for claim 1. Consider claim 11, Tu discloses the first audio segment comprises a plurality of audio units (frame shifts made of 256 samples, Section 4, page 63), and wherein the performing speech separation on the current audio segment based on the fused hidden layer state information comprises: performing speech separation for each audio unit, to obtain a first separated signal of the first audio segment (computing the power spectra features, and performing speech separation, Sections 3.2 and 4, page 63); and performing speech separation on the first separated signal based on the fused hidden layer state information, to obtain the plurality of separated speech signals of the first audio segment for each sound source (Fig 3, Section 3.2, second pass speech separation, page 63). Consider claim 18, Tu discloses the performing speech separation on the audio signal comprising first audio segment from the first audio block comprises: obtaining the hidden layer state information of the target audio and the second audio (Fig 3, layer h1, page 63); fusing the hidden layer state information of the target audio and the second audio segment to obtain fused hidden layer state information (Fig 3, layers h2 and h3, page 63); and performing speech separation on the first audio segment based on the fused hidden layer state information (Fig 3, target features and interfering features outputs, page 63). Tu does not specifically mention determining a target audio segment. Bi discloses determining a target audio segment (for each frame, comparing the current SNR to thresholds “between” the starting and ending points of the segments shown in Fig 3 to look back for the actual start of speech; this segment between START and PRE_START precedes the utterance segment, and these frames are considered a “target” audio segment” based on their instantaneous SNR values, Fig. 2, Fig 3, Col 5 lines 39-53, Col 6 lines 15-24). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Tu by determining a target audio segment as in Bi for reasons similar to those for claim 1. Consider claim 19, Tu discloses the hidden layer state information is obtained when speech separation is performed on the target audio and the second audio segment respectively (Fig 3, mixed feature separation performed block by block based on power spectra features, Section 4, page 63); and the hidden layer state information comprises at least one of short-term speech features, long-term speech features and context features of each sound source (the power spectra features computed in 32 msec frames, therefore considered short term speech features processed by the three hidden layers, Section 4, page 63). Tu does not specifically mention determining a target audio segment. Bi discloses determining a target audio segment (for each frame, comparing the current SNR to thresholds “between” the starting and ending points of the segments shown in Fig 3 to look back for the actual start of speech; this segment between START and PRE_START precedes the utterance segment, and these frames are considered a “target” audio segment” based on their instantaneous SNR values, Fig. 2, Fig 3, Col 5 lines 39-53, Col 6 lines 15-24). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Tu by determining a target audio segment as in Bi for reasons similar to those for claim 1. Consider claim 20, Tu discloses the first audio segment comprises a plurality of audio units (frame shifts made of 256 samples, Section 4, page 63), and wherein the performing speech separation on the first audio segment based on the fused hidden layer state information comprises: performing speech separation for each audio unit, to obtain a first separated signal of the first audio segment (computing the power spectra features, and performing speech separation, Sections 3.2 and 4, page 63); and performing speech separation on the first separated signal based on the fused hidden layer state information, to obtain the separated speech signal of the first audio segment for each sound source (Fig 3, Section 3.2, second pass speech separation, page 63). Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Tu et al. (“SPEECH SEPARATION BASED ON SIGNAL-NOISE-DEPENDENT DEEP NEURAL NETWORKS FOR ROBUST SPEECH RECOGNITION”. ICASSP 2015) in view of Bi et al. (US 6324509), in further view of Kleijn et al. (US 20180182412). Consider claim 7, Tu discloses the speech distortion is determined by calculating a difference between a separated speech signal for an audio segment and a reference audio signal, wherein the reference audio signal is an audio signal obtained by subtracting the separated speech signal from an original audio signal corresponding to the audio segment (for supervised fine-tuning, a aim at jointly minimizing the mean squared error between the DNN output and the reference clean features of the target speakers, Section 3.1, equation 2, page 62). . Tu and Bi do not specifically mention speech distortion is determined by calculating correlation. Kleijn discloses speech distortion is determined by calculating correlation (determining the distortion measure comprises determining a correlation measure, [0005]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Tu and Bi such that speech distortion is determined by calculating correlation in order to provide more robust and low-complexity demixing of sound sources when the source mixture is relative sparse in time, as suggested by Kleijn ([0016]), predictably resulting in more useful meeting recordings, as suggested by Kleijn ([0003]). The references cited are analogous art in the same field of speech processing. Allowable Subject Matter Claims 4-6 and 15-17 are objected to as being dependent on a rejected base claim, but would be allowable if rewritten in independent form including all limitations of the base and any intervening claims. Consider claim 4, the prior art does not fairly teach or suggest: “…the comparing the speech quality between the audio segment of the first audio block and the target audio segment determined for the second audio block comprises: for each sound source that has been separated, determining whether the first audio block belongs to a target audio block based on the speech quality of the first audio block; based on the first audio block belonging to the target audio block, determining whether the speech quality of the first audio segment in the first audio block is higher than that of the second audio segment of the first audio block; based on the speech quality of the first audio segment being higher than that of the second audio segment, determining whether the speech quality of the first audio segment is higher than that of the target audio segment determined for the second audio block; and determining the target audio segment corresponding to each sound source for the first audio block based on a comparison of speech quality of the first audio segment and the target audio segment determined for the second audio block.” Consider claim 5, the prior art does not fairly teach or suggest: “…the comparing the speech quality between the audio segment of the first audio block and the target audio segment determined for the second audio block comprises: for each sound source that has been separated, determining whether the first audio block belongs to a target audio block based on the speech quality of the first audio block; based on the first audio block belonging to target audio block, determining the speech quality of each audio segment in the first audio block, and selecting a third audio segment with the highest speech quality in the first audio block; determining whether the speech quality of the third audio segment is higher than that of the target audio segment determined for the second audio block; and determining the target audio segment corresponding to each sound source for the first audio block based on a comparison between the speech quality of the third audio segment and the target audio segment determined for the second audio block.” Consider claim 6, the prior art does not fairly teach or suggest: “…the determining the target audio segment corresponding to each sound source for the first audio block comprises: based on the speech quality of the first audio segment or the third audio segment being higher than that of the target audio segment determined for the second audio block, determining the first audio segment or the third audio segment as the target audio segment; and based on the speech quality of the first audio segment or the third audio segment being lower than that of the target audio segment determined for the second audio block, determining the first audio segment or the third audio segment as the target audio segment if a difference between the speech quality of the first audio segment or the third audio segment and that of the target audio segment determined for the second audio block is less than a preset threshold and a time interval between the first audio segment or the third audio segment and the target audio segment determined for the second audio block is greater than a time threshold.” Claim 15-17 recite limitations similar to those found in claims 4-6, and are allowable over the prior art for reasons similar to those described above with regard to claims 4-6. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jesse Pullias whose telephone number is 571/270-5135. The examiner can normally be reached on M-F 8:00 AM - 4:30 PM. The examiner’s fax number is 571/270-6135. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Andrew Flanders can be reached on 571/272-7516. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Jesse S Pullias/ Primary Examiner, Art Unit 2655 04/14/26
Read full office action

Prosecution Timeline

Show 3 earlier events
Sep 08, 2025
Applicant Interview (Telephonic)
Oct 17, 2025
Response Filed
Nov 10, 2025
Final Rejection mailed — §103
Jan 12, 2026
Request for Continued Examination
Jan 26, 2026
Response after Non-Final Action
Apr 16, 2026
Non-Final Rejection mailed — §103
Jul 29, 2026
Applicant Interview (Telephonic)
Jul 29, 2026
Examiner Interview Summary

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12694234
IMAGE-BASED TEXT TRANSLATION AND PRESENTATION
3y 10m to grant Granted Jul 28, 2026
Patent 12694224
Detecting Random and/or Algorithmically-Generated Character Sequences in Domain Names
2y 2m to grant Granted Jul 28, 2026
Patent 12682171
CONTEXT DISAMBIGUATION USING DEEP NEURAL NETWORKS
2y 9m to grant Granted Jul 14, 2026
Patent 12682169
ENTITY RELATION MINING METHOD AND APPARATUS, ELECTRONIC DEVICE, AND STORAGE MEDIUM
2y 4m to grant Granted Jul 14, 2026
Patent 12675648
LANGUAGE MODEL TRAINING APPARATUS, LANGUAGE MODEL TRAINING METHOD, AND STORAGE MEDIUM
2y 7m to grant Granted Jul 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
83%
Grant Probability
95%
With Interview (+12.7%)
2y 7m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 1069 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month