Prosecution Insights
Last updated: October 04, 2026
Application No. 19/012,835

SYSTEMS AND METHODS FOR MULTIPLE SPEAKER SPEECH RECOGNITION

Non-Final OA §101§DOUBLEPATENT
Filed
Jan 07, 2025
Priority
Oct 13, 2021 — CN 202111190697.X +2 more
Examiner
PULLIAS, JESSE SCOTT
Art Unit
Tech Center
Assignee
Hithink Royalflush Information Network Co. Ltd.
OA Round
1 (Non-Final)
83%
Grant Probability
Favorable
1-2
OA Rounds
10m
Est. Remaining
95%
With Interview

Examiner Intelligence

Grants 83% — above average
83%
Career Allowance Rate
885 granted / 1072 resolved
+22.6% vs TC avg
Moderate +12% lift
Without
With
+12.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
32 currently pending
Career history
1110
Total Applications
across all art units

Statute-Specific Performance

§101
15.5%
-24.5% vs TC avg
§103
53.1%
+13.1% vs TC avg
§102
19.8%
-20.2% vs TC avg
§112
4.6%
-35.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1072 resolved cases

Office Action

§101 §DOUBLEPATENT
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION This office action is in response to application 19/012,835, which was filed 01/07/25 and is a continuation of application 17/660,407, now U.S. Patent No. 12,223,945. In a preliminary amendment 01/07/25, claims 1-20 were cancelled and new claims 21-40 were added. Claims 21-40 are pending in the application and have been considered. Foreign Priority Receipt is acknowledged of certified copies of papers submitted under 35 U.S.C. 119(a)-(d), which papers have been placed of record in the file. Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory obviousness-type double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); and In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on a nonstatutory double patenting ground provided the conflicting application or patent either is shown to be commonly owned with this application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. Effective January 1, 1994, a registered attorney or agent of record may sign a terminal disclaimer. A terminal disclaimer signed by the assignee must fully comply with 37 CFR 3.73(b). Claims 21-40 are rejected on the ground of nonstatutory obviousness-type double patenting as being unpatentable over claims 1-12 of US Patent 12,223,945. Specifically, a comparison of independent claims 21, 32, and 33 in the present application with claims 1-4, 12, and 1, 10, and 11 of US Patent 12,223,945 yields the following: (Present application) (US Patent 12,223,945) Instant claim 21: obtaining speech data and a speech recognition result of the speech data, the speech data including speech of a plurality of speakers, and the speech recognition result including a plurality of words; determining, based on the speech data and the speech recognition result, a corresponding relationship between the plurality of words and the plurality of speakers; determining, based on the corresponding relationship, at least one conversion word from the plurality of words, each of the at least one conversion word corresponding to at least two of the plurality of speakers; and for each of the at least one conversion word, determining a speaker corresponding to the conversion word based on a time interval between the conversion word and at least one adjacent word of the conversion word by: determining a target adjacent word from the at least one adjacent word of the conversion word, the time interval between the adjacent word and the conversion words being less than a preset length of time, and the target adjacent word not being the conversion word; and determining the speaker corresponding to the conversion word based on the target adjacent word. US 12,223,945 claims 1-4: obtaining speech data, the speech data including speech of a plurality of speakers; …the speech recognition result including a plurality of words determining speaking time of each of the plurality of speakers by processing the speech data; determining, based on the speaking times of the plurality of speakers and the speech recognition result, a corresponding relationship between the plurality of words and the plurality of speakers; determining, based on the corresponding relationship, at least one conversion word from the plurality of words, each of the at least one conversion word corresponding to at least two of the plurality of speakers; and for each of the at least one conversion word, determining a speaker corresponding to the conversion word based on a time interval between the conversion word and at least one adjacent word of the conversion word. determining a target adjacent word from the at least one adjacent word of the conversion word, the time interval between the adjacent word and the conversion words being less than a preset length of time, and the target adjacent word not being the conversion word; and determining the speaker corresponding to the conversion word based on the target adjacent word. Instant claim 32: obtaining speech data to be recognized; determining a plurality of candidate speech recognition results of the speech data and a confidence score of each of the plurality of candidate speech recognition results using a decoding model and a re-scoring model; determining a speech recognition result from the plurality of candidate speech recognition results based on the confidence scores of the plurality of candidate speech recognition results, wherein the re-scoring model is generated based on a process of training model, the process of training model including: obtaining a first language model and a second language model, and the first language model being obtained by cropping the second language model; generating the decoding model based on the first language model; generating a first scoring model and a second scoring model based on the first language model and the second language model; and obtaining the re-scoring model based on the first scoring model and the second scoring model, wherein the obtaining the re-scoring model based on the first scoring model and the second scoring model includes: obtaining a plurality of possible speech recognition results and a first confidence score of each of the plurality of possible speech recognition results by traversing the second scoring model; for each of the plurality of possible speech recognition results, determining a second confidence score of the possible speech recognition result based on the first scoring model; updating the second scoring model based on the first confidence scores and the second confidence scores of the plurality of possible speech recognition results; and determining the re-scoring model based on the updated second scoring model. US 12,223,945 claim 12: obtaining speech data, the speech data including speech of a plurality of speakers; obtaining a plurality of candidate speech recognition results and an initial confidence score of each of the plurality of candidate speech recognition results by processing the feature information using a decoding model; determining a correction value of the candidate speech recognition result using a re-scoring model; and determining the confidence score of the candidate speech recognition result by summing the initial confidence score and the correction value of the candidate speech recognition result; and determining a speech recognition result from the plurality of candidate speech recognition results based on the confidence scores of the plurality of candidate speech recognition results, the speech recognition result including a plurality of words, wherein the re-scoring model is generated based on a process of training model, the process of training model including: obtaining a first language model and a second language model, and the first language model being obtained by cropping the second language model; generating the decoding model based on the first language model; generating a first scoring model and a second scoring model based on the first language model and the second language model; and obtaining the re-scoring model based on the first scoring model and the second scoring model, wherein the obtaining the re-scoring model based on the first scoring model and the second scoring model includes: obtaining a plurality of possible speech recognition results and a first confidence score of each of the plurality of possible speech recognition results by traversing the second scoring model; for each of the plurality of possible speech recognition results, determining a second confidence score of the possible speech recognition result based on the first scoring model; updating the second scoring model based on the first confidence scores and the second confidence scores of the plurality of possible speech recognition results; and determining the re-scoring model based on the updated second scoring model. Instant claim 33: obtaining speech data to be recognized; obtaining a plurality of candidate speech recognition results of the speech data and an initial confidence score of each of the plurality of candidate speech recognition results using a decoding model; for each of the plurality of candidate speech recognition results, determining the confidence score of the candidate speech recognition result based on the initial confidence score and a correction value of the candidate speech recognition result, the correction value being determined using a re-scoring model; and determining a speech recognition result from the plurality of candidate speech recognition results based on the confidence scores of the plurality of candidate speech recognition results, wherein the decoding model has a directed graph structure including a plurality of arcs, each arc having a weight, each candidate speech recognition result is represented by a set of arcs in the decoding model, the confidence score of each candidate speech recognition result is determined based on the weights of the set of arcs corresponding to the candidate speech recognition result, the re-scoring model has a second directed graph structure including a plurality of second arcs, each second arc having a second weight, the determining a correction value of the candidate speech recognition result using a re-scoring model comprises: determining, from the plurality of second arcs in the re-scoring model, a second set of arcs corresponding to the set of arcs of the candidate speech recognition result; and determining the correction value of the candidate speech recognition result based on the second weights of the plurality of second arcs. US 12,223,945 claims 1, 10, and 11: obtaining speech data, the speech data including speech of a plurality of speakers; obtaining a plurality of candidate speech recognition results and an initial confidence score of each of the plurality of candidate speech recognition results by processing the feature information using a decoding model; for each of the plurality of candidate speech recognition results, determining a correction value of the candidate speech recognition result using a re-scoring model; and determining the confidence score of the candidate speech recognition result by summing the initial confidence score and the correction value of the candidate speech recognition result; and determining a speech recognition result from the plurality of candidate speech recognition results based on the confidence scores of the plurality of candidate speech recognition results, 10. The system of claim 1, wherein the decoding model has a directed graph structure including a plurality of arcs, each arc having a weight, each candidate speech recognition result is represented by a set of arcs in the decoding model, and the confidence score of each candidate speech recognition result is determined based on the weights of the set of arcs corresponding to the candidate speech recognition result. 11. The system of claim 10, wherein the re-scoring model has a second directed graph structure including a plurality of second arcs, each second arc having a second weight, the determining a correction value of the candidate speech recognition result using a re-scoring model comprises: determining, from the plurality of second arcs in the re-scoring model, a second set of arcs corresponding to the set of arcs of the candidate speech recognition result; and determining the correction value of the candidate speech recognition result based on the second weights of the plurality of second arcs. As the table above demonstrates each limitation of claim 21 of the present application is found in claims 1-4 of US Patent 12,223,945, thus claim 21 of the present application is anticipated by claims 1-4 of US Patent 12,223,945. Each limitation of claim 32 of the present application is found in claim 12 of US Patent 12,223,945, thus claim 32 of the present application is anticipated by claim 12 of US Patent 12,223,945. Each limitation of claim 33 of the present application is found in claims 1, 10, and 11 of US Patent 12,223,945, thus claim 33 of the present application is anticipated by claims 1, 10, and 11 of US Patent 12,223,945. Dependent claims 22-31 and 34-40 of the present application recite subject matter that corresponds to that found in claims 1-3, 5-8, and 9 of US Patent 12,223,945 as follows, and therefore are also anticipated. (Present application) (US Patent 12,223,945) 22. The system of claim 21, wherein the determining a corresponding relationship between the plurality of words and the plurality of speakers comprises: determining a speaking time of each of the plurality of speakers by processing the speech data; and determining, based on the speaking times of the plurality of speakers and the speech recognition result, a corresponding relationship between the plurality of words and the plurality of speakers. 2. The system of claim 1, the operations further comprising: determining speaking time of each of the plurality of speakers by processing the speech data; determining, based on the speaking times of the plurality of speakers and the speech recognition result, a corresponding relationship between the plurality of words and the plurality of speakers; determining, based on the corresponding relationship, at least one conversion word from the plurality of words, each of the at least one conversion word corresponding to at least two of the plurality of speakers; and re-determining the corresponding relationship between the plurality of words and the plurality of speakers based on the at least one conversion word. 23. The system of claim 22, wherein the determining speaking time of each of the plurality of speakers by processing the speech data includes: generating at least one effective speech segment including speech content by preprocessing the speech data; and determining the speaking time of each of the plurality of speakers by processing the at least one effective speech segment. 6. The system of claim 2, wherein the determining speaking time of each of the plurality of speakers by processing the speech data includes: generating at least one effective speech segment including speech content by preprocessing the speech data; and determining the speaking time of each of the plurality of speakers by processing the at least one effective speech segment. 24. The system of claim 23, wherein the generating at least one effective speech segment including speech content by preprocessing the speech data includes: obtaining a plurality of first speech segments by processing the speech data based on the speech recognition result; obtaining a plurality of second speech segments by amplifying the plurality of first speech segments; and obtaining the at least one effective speech segment by performing a processing operation on the plurality of second speech segments, the processing operation including at least one of a merging operation or a segmentation operation. 7. The system of claim 6, wherein the generating at least one effective speech segment including speech content by preprocessing the speech data includes: obtaining a plurality of first speech segments by processing the speech data based on the speech recognition result; obtaining a plurality of second speech segments by amplifying the plurality of first speech segments; and obtaining the at least one effective speech segment by performing a processing operation on the plurality of second speech segments, the processing operation including at least one of a merging operation or a segmentation operation. 25. The system of claim 23, wherein the determining the speaking time of each of the plurality of speakers by processing the at least one effective speech segment includes: extracting a voiceprint feature vector of each of the at least one effective speech segment; and determining the speaking time of each of the plurality of speakers by clustering the voiceprint feature vector of the at least one effective speech segment. 8. The system of claim 6, wherein the determining the speaking time of each of the plurality of speakers by processing the at least one effective speech segment includes: extracting a voiceprint feature vector of each of the at least one effective speech segment; and determining the speaking time of each of the plurality of speakers by clustering the voiceprint feature vector of the at least one effective speech segment. 26. The system of claim 21, wherein to obtain the speech recognition result of the speech data, the at least one processor is configured to cause the system to perform operations including: determining feature information of the speech data; determining, based on the feature information, a plurality of candidate speech recognition results and a confidence score of each of the plurality of candidate speech recognition results using a decoding model and a re-scoring model; and determining the speech recognition result from the plurality of candidate speech recognition results based on the confidence scores of the plurality of candidate speech recognition results. 1… determining feature information of the speech data; obtaining a plurality of candidate speech recognition results and an initial confidence score of each of the plurality of candidate speech recognition results by processing the feature information using a decoding model; determining a speech recognition result from the plurality of candidate speech recognition results based on the confidence scores of the plurality of candidate speech recognition results, 27. The system of claim 26, wherein the determining, based on the feature information, a plurality of candidate speech recognition results and a confidence score of each of the plurality of candidate speech recognition results using a decoding model and a re-scoring model includes: obtaining the plurality of candidate speech recognition results and an initial confidence score of each of the plurality of candidate speech recognition results by processing the feature information using the decoding model; determining a correction value of each of the plurality of candidate speech recognition results using the re-scoring model; and for each of the plurality of candidate speech recognition results, determining the confidence score of the candidate speech recognition result based on the corresponding initial confidence score and the corresponding correction value. 1… obtaining a plurality of candidate speech recognition results and an initial confidence score of each of the plurality of candidate speech recognition results by processing the feature information using a decoding model; for each of the plurality of candidate speech recognition results, determining a correction value of the candidate speech recognition result using a re-scoring model; and determining the confidence score of the candidate speech recognition result by summing the initial confidence score and the correction value of the candidate speech recognition result; and determining a speech recognition result from the plurality of candidate speech recognition results based on the confidence scores of the plurality of candidate speech recognition results, 28. The system of claim 26, wherein the re-scoring model is generated based on a process of training model, the process of training model including: obtaining a first language model and a second language model, and the first language model being obtained by cropping the second language model; generating the decoding model based on the first language model; generating a first scoring model and a second scoring model based on the first language model and the second language model; and obtaining the re-scoring model based on the first scoring model and the second scoring model. 1… obtaining a first language model and a second language model, and the first language model being obtained by cropping the second language model; generating the decoding model based on the first language model; generating a first scoring model and a second scoring model based on the first language model and the second language model; and obtaining the re-scoring model based on the first scoring model and the second scoring model, 29. The system of claim 28, wherein the obtaining the re-scoring model based on the first scoring model and the second scoring model includes: obtaining a plurality of possible speech recognition results and a first confidence score of each of the plurality of possible speech recognition results by traversing the second scoring model; for each of the plurality of possible speech recognition results, determining a second confidence score of the possible speech recognition result based on the first scoring model; updating the second scoring model based on the first confidence scores and the second confidence scores of the plurality of possible speech recognition results; and determining the re-scoring model based on the updated second scoring model. 1… obtaining a plurality of possible speech recognition results and a first confidence score of each of the plurality of possible speech recognition results by traversing the second scoring model; for each of the plurality of possible speech recognition results, determining a second confidence score of the possible speech recognition result based on the first scoring model; updating the second scoring model based on the first confidence scores and the second confidence scores of the plurality of possible speech recognition results; and determining the re-scoring model based on the updated second scoring model. 30. The system of claim 29, wherein the first scoring model includes a plurality of first arcs, the second scoring model includes a plurality of second arcs, and the each of the plurality of possible speech recognition results is represented by a sequence of at least one second arc, and for each of the plurality of possible speech recognition results, the determining a second confidence score of the possible speech recognition result based on the first scoring model includes: for the each of the plurality of possible speech recognition results, determining, in the first scoring model, a sequence of at least one first arc corresponding to the sequence of at least one second arc of the possible speech recognition result; and determining the second confidence score of the possible speech recognition result based on the sequence of at least one first arc. 9. The system of claim 1, wherein the first scoring model includes a plurality of first arcs, the second scoring model includes a plurality of second arcs, and the each of the plurality of possible speech recognition results is represented by a sequence of at least one second arc, and for each of the plurality of possible speech recognition results, the determining a second confidence score of the possible speech recognition result based on the first scoring model includes: for the each of the plurality of possible speech recognition results, determining, in the first scoring model, a sequence of at least one first arc corresponding to the sequence of at least one second arc of the possible speech recognition result; and determining the second confidence score of the possible speech recognition result based on the sequence of at least one first arc. 31. The system of claim 21, the operations further comprising: obtaining voiceprint feature information of each of the plurality of speakers; and for each of the at least one conversion word, determining the speaker corresponding to the conversion word based on the target adjacent word corresponding to the conversion word and the voiceprint feature information of the plurality of speakers and voiceprint feature information of a target speech segment, the target speech segment including the conversion word. 5. The system of claim 2, wherein the re-determining the corresponding relationship between the plurality of words and the plurality of speakers based on the at least one conversion word includes: obtaining voiceprint feature information of each of the plurality of speakers; and for each of the at least one conversion word, determining the speaker corresponding to the conversion word based on the voiceprint feature information of the plurality of speakers and voiceprint feature information of a target speech segment, the target speech segment including the conversion word. 34. The system of claim 33, the operations further comprising: determining speaking time of each of the plurality of speakers by processing the speech data; determining, based on the speaking times of the plurality of speakers and the speech recognition result, a corresponding relationship between the plurality of words and the plurality of speakers; determining, based on the corresponding relationship, at least one conversion word from a plurality of words of the speech recognition result, each of the at least one conversion word corresponding to at least two of the plurality of speakers; and re-determining the corresponding relationship between the plurality of words and the plurality of speakers based on the at least one conversion word. 2. The system of claim 1, the operations further comprising: determining speaking time of each of the plurality of speakers by processing the speech data; determining, based on the speaking times of the plurality of speakers and the speech recognition result, a corresponding relationship between the plurality of words and the plurality of speakers; determining, based on the corresponding relationship, at least one conversion word from the plurality of words, each of the at least one conversion word corresponding to at least two of the plurality of speakers; and re-determining the corresponding relationship between the plurality of words and the plurality of speakers based on the at least one conversion word. 35. The system of claim 34, wherein the re-determining the corresponding relationship between the plurality of words and the plurality of speakers based on the at least one conversion word includes: for each of the at least one conversion word, determining a speaker corresponding to the conversion word based on a time interval between the conversion word and at least one adjacent word of the conversion word. 3. The system of claim 2, wherein the re-determining the corresponding relationship between the plurality of words and the plurality of speakers based on the at least one conversion word includes: for each of the at least one conversion word, determining a speaker corresponding to the conversion word based on a time interval between the conversion word and at least one adjacent word of the conversion word. 36. The system of claim 34, wherein the determining speaking time of each of the plurality of speakers by processing the speech data includes: generating at least one effective speech segment including speech content by preprocessing the speech data; and determining the speaking time of each of the plurality of speakers by processing the at least one effective speech segment. 6. The system of claim 2, wherein the determining speaking time of each of the plurality of speakers by processing the speech data includes: generating at least one effective speech segment including speech content by preprocessing the speech data; and determining the speaking time of each of the plurality of speakers by processing the at least one effective speech segment. 37. The system of claim 36, wherein the generating at least one effective speech segment including speech content by preprocessing the speech data includes: obtaining a plurality of first speech segments by processing the speech data based on the speech recognition result; obtaining a plurality of second speech segments by amplifying the plurality of first speech segments; and obtaining the at least one effective speech segment by performing a processing operation on the plurality of second speech segments, the processing operation including at least one of a merging operation or a segmentation operation. 7. The system of claim 6, wherein the generating at least one effective speech segment including speech content by preprocessing the speech data includes: obtaining a plurality of first speech segments by processing the speech data based on the speech recognition result; obtaining a plurality of second speech segments by amplifying the plurality of first speech segments; and obtaining the at least one effective speech segment by performing a processing operation on the plurality of second speech segments, the processing operation including at least one of a merging operation or a segmentation operation. 38. The system of claim 33, wherein the re-scoring model is generated based on a process of training model, the process of training model including: obtaining a first language model and a second language model, and the first language model being obtained by cropping the second language model; generating the decoding model based on the first language model; generating a first scoring model and a second scoring model based on the first language model and the second language model; and obtaining the re-scoring model based on the first scoring model and the second scoring model. 1… obtaining a first language model and a second language model, and the first language model being obtained by cropping the second language model; generating the decoding model based on the first language model; generating a first scoring model and a second scoring model based on the first language model and the second language model; and obtaining the re-scoring model based on the first scoring model and the second scoring model, 39. The system of claim 38, wherein the obtaining the re-scoring model based on the first scoring model and the second scoring model includes: obtaining a plurality of possible speech recognition results and a first confidence score of each of the plurality of possible speech recognition results by traversing the second scoring model; for each of the plurality of possible speech recognition results, determining a second confidence score of the possible speech recognition result based on the first scoring model; updating the second scoring model based on the first confidence scores and the second confidence scores of the plurality of possible speech recognition results; and determining the re-scoring model based on the updated second scoring model. 1… wherein the obtaining the re-scoring model based on the first scoring model and the second scoring model includes: obtaining a plurality of possible speech recognition results and a first confidence score of each of the plurality of possible speech recognition results by traversing the second scoring model; for each of the plurality of possible speech recognition results, determining a second confidence score of the possible speech recognition result based on the first scoring model; updating the second scoring model based on the first confidence scores and the second confidence scores of the plurality of possible speech recognition results; and determining the re-scoring model based on the updated second scoring model. 40. The system of claim 39, wherein the first scoring model includes a plurality of first arcs, the second scoring model includes a plurality of second arcs, and the each of the plurality of possible speech recognition results is represented by a sequence of at least one second arc, and for each of the plurality of possible speech recognition results, the determining a second confidence score of the possible speech recognition result based on the first scoring model includes: for the each of the plurality of possible speech recognition results, determining, in the first scoring model, a sequence of at least one first arc corresponding to the sequence of at least one second arc of the possible speech recognition result; and determining the second confidence score of the possible speech recognition result based on the sequence of at least one first arc. 9. The system of claim 1, wherein the first scoring model includes a plurality of first arcs, the second scoring model includes a plurality of second arcs, and the each of the plurality of possible speech recognition results is represented by a sequence of at least one second arc, and for each of the plurality of possible speech recognition results, the determining a second confidence score of the possible speech recognition result based on the first scoring model includes: for the each of the plurality of possible speech recognition results, determining, in the first scoring model, a sequence of at least one first arc corresponding to the sequence of at least one second arc of the possible speech recognition result; and determining the second confidence score of the possible speech recognition result based on the sequence of at least one first arc. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 21, 22, and 31 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. Claim 21 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim recites “obtaining speech data and a speech recognition result of the speech data, the speech data including speech of a plurality of speakers, and the speech recognition result including a plurality of words; determining, based on the speech data and the speech recognition result, a corresponding relationship between the plurality of words and the plurality of speakers; determining, based on the corresponding relationship, at least one conversion word from the plurality of words, each of the at least one conversion word corresponding to at least two of the plurality of speakers; and for each of the at least one conversion word, determining a speaker corresponding to the conversion word based on a time interval between the conversion word and at least one adjacent word of the conversion word by: determining a target adjacent word from the at least one adjacent word of the conversion word, the time interval between the adjacent word and the conversion words being less than a preset length of time, and the target adjacent word not being the conversion word; and determining the speaker corresponding to the conversion word based on the target adjacent word”. The limitation of obtaining speech data and a speech recognition result of the speech data, the speech data including speech of a plurality of speakers, and the speech recognition result including a plurality of words, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, “obtaining speech data and a speech recognition result of the speech data, the speech data including speech of a plurality of speakers, and the speech recognition result including a plurality of words” in the context of this claim encompasses e.g. obtaining a paper transcript of a conversation between multiple speakers having words and timestamps. Similarly, the limitation of “determining, based on the speech data and the speech recognition result, a corresponding relationship between the plurality of words and the plurality of speakers”, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, “determining, based on the speech data and the speech recognition result, a corresponding relationship between the plurality of words and the plurality of speakers” in the context of this claim encompasses mentally determining, based on the speech data and the speech recognition result, a corresponding relationship between the plurality of words and the plurality of speakers, for example by mentally determining which speakers said which words based on the time stamps. Similarly, the limitation of “determining, based on the corresponding relationship, at least one conversion word from the plurality of words, each of the at least one conversion word corresponding to at least two of the plurality of speakers”, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, “determining, based on the corresponding relationship, at least one conversion word from the plurality of words, each of the at least one conversion word corresponding to at least two of the plurality of speakers” in the context of this claim encompasses mentally determining, based on the corresponding relationship, at least one conversion word from the plurality of words, each of the at least one conversion word corresponding to at least two of the plurality of speakers, by mentally determining e.g. words which could have been spoken by at least two of the speakers. Similarly, the limitation of “for each of the at least one conversion word, determining a speaker corresponding to the conversion word based on a time interval between the conversion word and at least one adjacent word of the conversion word by: determining a target adjacent word from the at least one adjacent word of the conversion word, the time interval between the adjacent word and the conversion words being less than a preset length of time, and the target adjacent word not being the conversion word”, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, “for each of the at least one conversion word, determining a speaker corresponding to the conversion word based on a time interval between the conversion word and at least one adjacent word of the conversion word by: determining a target adjacent word from the at least one adjacent word of the conversion word, the time interval between the adjacent word and the conversion words being less than a preset length of time, and the target adjacent word not being the conversion word” in the context of this claim encompasses, for each of the at least one conversion word, mentally determining a speaker corresponding to the conversion word based on a time interval between the conversion word and at least one adjacent word of the conversion word by: mentally determining a target adjacent word from the at least one adjacent word of the conversion word, the time interval between the adjacent word and the conversion words being less than a preset length of time, and the target adjacent word not being the conversion word by, e.g. scanning the transcript for adjacent words and mentally determining the time intervals based on the timestamps for find words spoken within a preset length of time of the conversion words. Similarly, the limitation of “determining the speaker corresponding to the conversion word based on the target adjacent word”, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, “determining the speaker corresponding to the conversion word based on the target adjacent word” in the context of this claim encompasses mentally determining the speaker corresponding to the conversion word based on the target adjacent word. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea. This judicial exception is not integrated into a practical application. In particular, the claim only recites three additional elements – “storage device”, “set of instructions” and “processor”. The computing elements in this step are recited at a high-level of generality (i.e., as a generic storage device, a generic set of instructions, and generic processor) such that they amount to no more than mere instructions to apply the exception using generic computer elements. Accordingly, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of using a processor executing instructions stored on a storage device to perform the obtaining and determining amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The claim is not patent eligible. Specifically with respect to Step 2A, Prong Two, of the Alice/Mayo test, the judicial exception is not integrated into a practical application. Claim 1 does not recite any limitations that are not mental steps. Specifically with respect to Step 2B of the Alice/Mayo test, “the claim as a whole does not amount to significantly more than the exception itself (there is no inventive concept in the claim)”. MPEP 2106.05 Il. There are no limitations in claim 1 outside of the judicial exception. As a whole, there does not appear to contain any inventive concept. As discussed above, claim 1 is a mental process that pertains to the mental process of determining speakers corresponding to words, which can be performed entirely by a human with physical aids. Dependent claims 22 depends from claim 21, does not remedy any of the deficiencies of claim 21, and therefore is rejected on the same grounds as claim 21 above. Claim 22 merely recites additional steps for making determinations between words and speakers, all of which could be performed mentally or by writing down relationships with a pen and paper, and do not amount to anything more than substantially the same abstract idea as explained with respect to claim 21. Specifically: Claim 22 recites “determining a corresponding relationship between the plurality of words and the plurality of speakers comprises: determining a speaking time of each of the plurality of speakers by processing the speech data; and determining, based on the speaking times of the plurality of speakers and the speech recognition result, a corresponding relationship between the plurality of words and the plurality of speakers” which could be performed by mentally determining a speaking time of each of the plurality of speakers by mentally processing the speech data; and mentally determining, based on the speaking times of the plurality of speakers and the speech recognition result, a corresponding relationship between the plurality of words and the plurality of speakers. As seen above, claim 22 depends from claim 21 and further recites mental processes as explained above. None of the additional limitations recited in claim 22 amount to anything more than the same or a similar abstract idea as recited in claim 21. Nor do any limitations in claim 22 (a) integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea or (b) amount to significantly more than the judicial exception. Claim 22 is not patent eligible. Dependent claim 31 depends from claim 21, does not remedy any of the deficiencies of claim 21, and therefore is rejected on the same grounds as claim 21 above. Claim 31 merely recites additional steps for making determinations between words and speakers, all of which could be performed mentally or by writing down relationships with a pen and paper, and do not amount to anything more than substantially the same abstract idea as explained with respect to claim 21. Specifically: Claim 31 recites “obtaining voiceprint feature information of each of the plurality of speakers; and for each of the at least one conversion word, determining the speaker corresponding to the conversion word based on the target adjacent word corresponding to the conversion word and the voiceprint feature information of the plurality of speakers and voiceprint feature information of a target speech segment, the target speech segment including the conversion word” which could be performed by obtaining “voiceprint feature information” such as a written voiceprint label from a sheet of paper of each of the plurality of speakers, and for each conversion word, mentally determining the speaker corresponding to the conversion word based on the target adjacent word corresponding to the conversion word and the voiceprint feature information of the plurality of speakers and voiceprint feature information of a target speech segment, the target speech segment including the conversion word. Notably, the claim only requires “obtaining voiceprint feature information” and does not actually require extracting a voiceprint itself from the speech audio. As seen above, claim 31 depends from claim 21 and further recites mental processes as explained above. None of the additional limitations recited in claim 31 amount to anything more than the same or a similar abstract idea as recited in claim 21. Nor do any limitations in claim 31 (a) integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea or (b) amount to significantly more than the judicial exception. Claim 31 is not patent eligible. Eligible Claims Claims 23-30 and 32-40 each recite speech processing operations which given their broadest reasonable interpretation in light of the specification, would be impractical or impossible to perform as a mental process. Specifically, in claim 23, “preprocessing” the speech data is a well known term in the art that refers to techniques for processing the audio data of the speech itself prior to further analysis. Interpretations that would allow for some sort of mental “preprocessing” the speech data would be unreasonable. Claims 24-25 depend on claim 23 and also recite additional speech processing techniques which cannot be practically performed mentally. Claim 26 requires performing speech recognition operations which cannot be performed mentally. Claims 27-30 depend on claim 26 and therefore cannot be performed mentally for at least the above reason. Claims 32 and 33 require performing speech recognition operations which cannot be performed mentally. Claims 34-40 depend on claim 33 and therefore cannot be performed mentally for at least the above reason. Allowable Subject Matter Claims 21, 22, and 31 would be allowable if a proper terminal disclaimer were to overcome the double patenting rejections, as well as amended to overcome the 35 U.S.C. 101 rejections. Claims 23-40 would be allowable if a proper terminal disclaimer were to overcome the double patenting rejections. The following is the examiner’s statement of reasons for indicating subject matter allowable over the prior art of record: The closest prior art to independent claim 21 is Jung et al. (US 20190392837). Consider claim 21, Jung discloses a speech recognition system (voice recognition by a system, [0002-0003]), comprising: at least one storage device storing a set of instructions (storage device storing instructions, [0062-0063]); and at least one processor in communication with the at least one storage device, wherein when executing the set of instructions (processor executing instructions, [0060]), the at least one processor is configured to cause the system to perform operations including: obtaining speech data and a speech recognition result of the speech data, the speech data including speech of a plurality of speakers, and the speech recognition result including a plurality of words (utterances from multiple users participating in a conversation are received, [0083], the utterances are converted into text words, [0085-0086], [0098]); determining, based on the speech data and the speech recognition result, a corresponding relationship between the plurality of words and the plurality of speakers (determining that before Lisa R is able to complete the utterance at time t6, Tim G. begins speaking, [0052]); determining, based on the corresponding relationship, at least one conversion word from the plurality of words, each of the at least one conversion word corresponding to at least two of the plurality of speakers (determining that utterance 314 is an interruption with regard to utterance 312 because two voices are detected during the same time period, [0052], and include a second set of words interrupted by a third set of words, [0089], noting that Applicant’s original specification at [0080] uses “conversion word” to refer to overlapping speech); and for each of the at least one conversion word, determining a speaker corresponding to the conversion word based on a time interval between the conversion word and at least one adjacent word of the conversion word (first utterance and second utterance spoken during a same time period, [0049], [0107]). However, Jung does not disclose determining a target adjacent word from the at least one adjacent word of the conversion word, the time interval between the adjacent word and the conversion words being less than a preset length of time, and the target adjacent word not being the conversion word; and determining the speaker corresponding to the conversion word based on the target adjacent word. The closest prior art to independent claims 32 and 33 is Jung et al. (US 20190392837) and Itoh (US 20150279353). Jung discloses a speech recognition system (voice recognition by a system, [0002-0003]), comprising: at least one storage device storing a set of instructions (storage device storing instructions, [0062-0063]); and at least one processor in communication with the at least one storage device, wherein when executing the set of instructions (processor executing instructions, [0060]), the at least one processor is configured to cause the system to perform operations including: obtaining speech data to be recognized (utterances from multiple users participating in a conversation are received, [0083]); determining a plurality of candidate speech recognition results of the speech data (the utterances are converted into text words, [0085-0086], [0098]). Itoh discloses: obtaining a plurality of candidate speech recognition results and an initial confidence score of each of the plurality of candidate speech recognition results by processing the feature information using a decoding model (producing likelihood of phonemes or text using an acoustic model inherently requires processing feature information using a decoding model, since some acoustic feature must be decoded using the acoustic model, [0049]); for each of the plurality of candidate speech recognition results, determining a correction value of the candidate speech recognition result using a re-scoring model (the language model likelihood is considered “a correction value”, [0049]); determining the confidence score of the candidate speech recognition result by summing the initial confidence score and the correction value of the candidate speech recognition result (the sum of the likelihood of acoustic model and the likelihood of language model, [0049]); and determining a speech recognition result from the plurality of candidate speech recognition results based on the confidence scores of the plurality of candidate speech recognition results, the speech recognition result including a plurality of words (text of the recognition results, [0049-0051]). However, Jung and Itoh do not disclose “the process of training model including: obtaining a first language model and a second language model, and the first language model being obtained by cropping the second language model; generating the decoding model based on the first language model; generating a first scoring model and a second scoring model based on the first language model and the second language model; and obtaining the re-scoring model based on the first scoring model and the second scoring model, wherein the obtaining the re-scoring model based on the first scoring model and the second scoring model includes: obtaining a plurality of possible speech recognition results and a first confidence score of each of the plurality of possible speech recognition results by traversing the second scoring model; for each of the plurality of possible speech recognition results, determining a second confidence score of the possible speech recognition result based on the first scoring model; updating the second scoring model based on the first confidence scores and the second confidence scores of the plurality of possible speech recognition results; and determining the re-scoring model based on the updated second scoring model” as recited by claim 32, or “…determining a speech recognition result from the plurality of candidate speech recognition results based on the confidence scores of the plurality of candidate speech recognition results, wherein the decoding model has a directed graph structure including a plurality of arcs, each arc having a weight, each candidate speech recognition result is represented by a set of arcs in the decoding model, the confidence score of each candidate speech recognition result is determined based on the weights of the set of arcs corresponding to the candidate speech recognition result, the re-scoring model has a second directed graph structure including a plurality of second arcs, each second arc having a second weight, the determining a correction value of the candidate speech recognition result using a re-scoring model comprises: determining, from the plurality of second arcs in the re-scoring model, a second set of arcs corresponding to the set of arcs of the candidate speech recognition result; and determining the correction value of the candidate speech recognition result based on the second weights of the plurality of second arcs” as recited by claim 33. Claims 22-31 and 34-40 contain allowable subject matter because they further limit the allowable subject matter in parent claims 21 and 33. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. US 11670304 Khoury discloses speaker recognition at a call center by diarization using clustering US 9368109 Colibro discloses automatic speaker-based speech clustering US 11152006 Krupka discloses voice identification enrollment using voiceprints US 11636860 Gorodetski discloses word-level blind diarization of recorded calls with arbitrary number of speakers US 10366693 Gorodetski discloses acoustic signature building for a speaker through multiple sessions Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jesse Pullias whose telephone number is 571/270-5135. The examiner can normally be reached on M-F 8:00 AM - 4:30 PM. The examiner’s fax number is 571/270-6135. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Andrew Flanders can be reached on 571/272-7516. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). /Jesse S Pullias/ Primary Examiner, Art Unit 2655 09/14/26
Read full office action

Prosecution Timeline

Jan 07, 2025
Application Filed
Sep 16, 2026
Non-Final Rejection mailed — §101, §DOUBLEPATENT (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749483
SYSTEM, APPARATUS, AND METHOD FOR PROCESSING NATURAL LANGUAGE, AND NON-TRANSITORY COMPUTER READABLE RECORDING MEDIUM
3y 2m to grant Granted Sep 29, 2026
Patent 12738286
POST-PROCESSOR, AUDIO DECODER AND RELATED METHODS FOR ENHANCING TRANSIENT PROCESSING
6y 3m to grant Granted Sep 15, 2026
Patent 12730980
REAL-TIME ADAPTATION OF MACHINE LEARNING MODELS USING LARGE LANGUAGE MODELS
2y 9m to grant Granted Sep 08, 2026
Patent 12694234
IMAGE-BASED TEXT TRANSLATION AND PRESENTATION
3y 10m to grant Granted Jul 28, 2026
Patent 12694224
Detecting Random and/or Algorithmically-Generated Character Sequences in Domain Names
2y 2m to grant Granted Jul 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
83%
Grant Probability
95%
With Interview (+12.5%)
2y 7m (~10m remaining)
Median Time to Grant
Low
PTA Risk
Based on 1072 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month