Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
This office action is in response to application 19/012,835, which was filed 01/07/25 and is a continuation of application 17/660,407, now U.S. Patent No. 12,223,945. In a preliminary amendment 01/07/25, claims 1-20 were cancelled and new claims 21-40 were added. Claims 21-40 are pending in the application and have been considered.
Foreign Priority
Receipt is acknowledged of certified copies of papers submitted under 35 U.S.C. 119(a)-(d), which papers have been placed of record in the file.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory obviousness-type double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); and In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on a nonstatutory double patenting ground provided the conflicting application or patent either is shown to be commonly owned with this application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement.
Effective January 1, 1994, a registered attorney or agent of record may sign a terminal disclaimer. A terminal disclaimer signed by the assignee must fully comply with 37 CFR 3.73(b).
Claims 21-40 are rejected on the ground of nonstatutory obviousness-type double patenting as being unpatentable over claims 1-12 of US Patent 12,223,945.
Specifically, a comparison of independent claims 21, 32, and 33 in the present application with claims 1-4, 12, and 1, 10, and 11 of US Patent 12,223,945 yields the following:
(Present application) (US Patent 12,223,945)
Instant claim 21:
obtaining speech data and a speech recognition result of the speech data, the speech data including speech of a plurality of speakers, and the speech recognition result including a plurality of words;
determining, based on the speech data and the speech recognition result, a corresponding relationship between the plurality of words and the plurality of speakers;
determining, based on the corresponding relationship, at least one conversion word from the plurality of words, each of the at least one conversion word corresponding to at least two of the plurality of speakers; and
for each of the at least one conversion word, determining a speaker corresponding to the conversion word based on a time interval between the conversion word and at least one adjacent word of the conversion word by:
determining a target adjacent word from the at least one adjacent word of the conversion word, the time interval between the adjacent word and the conversion words being less than a preset length of time, and the target adjacent word not being the conversion word; and
determining the speaker corresponding to the conversion word based on the target adjacent word.
US 12,223,945 claims 1-4:
obtaining speech data, the speech data including speech of a plurality of speakers;
…the speech recognition result including a plurality of words
determining speaking time of each of the plurality of speakers by processing the speech data;
determining, based on the speaking times of the plurality of speakers and the speech recognition result, a corresponding relationship between the plurality of words and the plurality of speakers;
determining, based on the corresponding relationship, at least one conversion word from the plurality of words, each of the at least one conversion word corresponding to at least two of the plurality of speakers; and
for each of the at least one conversion word, determining a speaker corresponding to the conversion word based on a time interval between the conversion word and at least one adjacent word of the conversion word.
determining a target adjacent word from the at least one adjacent word of the conversion word, the time interval between the adjacent word and the conversion words being less than a preset length of time, and the target adjacent word not being the conversion word; and
determining the speaker corresponding to the conversion word based on the target adjacent word.
Instant claim 32:
obtaining speech data to be recognized;
determining a plurality of candidate speech recognition results of the speech data and a confidence score of each of the plurality of candidate speech recognition results using a decoding model and a re-scoring model;
determining a speech recognition result from the plurality of candidate speech recognition results based on the confidence scores of the plurality of candidate speech recognition results, wherein the re-scoring model is generated based on a process of training model, the process of training model including:
obtaining a first language model and a second language model, and the first language model being obtained by cropping the second language model;
generating the decoding model based on the first language model;
generating a first scoring model and a second scoring model based on the first language model and the second language model; and
obtaining the re-scoring model based on the first scoring model and the second scoring model, wherein the obtaining the re-scoring model based on the first scoring model and the second scoring model includes:
obtaining a plurality of possible speech recognition results and a first confidence score of each of the plurality of possible speech recognition results by traversing the second scoring model;
for each of the plurality of possible speech recognition results, determining a second confidence score of the possible speech recognition result based on the first scoring model; updating the second scoring model based on the first confidence scores and the second confidence scores of the plurality of possible speech recognition results; and
determining the re-scoring model based on the updated second scoring model.
US 12,223,945 claim 12:
obtaining speech data, the speech data including speech of a plurality of speakers;
obtaining a plurality of candidate speech recognition results and an initial confidence score of each of the plurality of candidate speech recognition results by processing the feature information using a decoding model; determining a correction value of the candidate speech recognition result using a re-scoring model; and
determining the confidence score of the candidate speech recognition result by summing the initial confidence score and the correction value of the candidate speech recognition result; and
determining a speech recognition result from the plurality of candidate speech recognition results based on the confidence scores of the plurality of candidate speech recognition results, the speech recognition result including a plurality of words, wherein the re-scoring model is generated based on a process of training model, the process of training model including:
obtaining a first language model and a second language model, and the first language model being obtained by cropping the second language model;
generating the decoding model based on the first language model;
generating a first scoring model and a second scoring model based on the first language model and the second language model; and
obtaining the re-scoring model based on the first scoring model and the second scoring model, wherein the obtaining the re-scoring model based on the first scoring model and the second scoring model includes:
obtaining a plurality of possible speech recognition results and a first confidence score of each of the plurality of possible speech recognition results by traversing the second scoring model;
for each of the plurality of possible speech recognition results, determining a second confidence score of the possible speech recognition result based on the first scoring model;
updating the second scoring model based on the first confidence scores and the second confidence scores of the plurality of possible speech recognition results; and
determining the re-scoring model based on the updated second scoring model.
Instant claim 33:
obtaining speech data to be recognized;
obtaining a plurality of candidate speech recognition results of the speech data and an initial confidence score of each of the plurality of candidate speech recognition results using a decoding model;
for each of the plurality of candidate speech recognition results, determining the confidence score of the candidate speech recognition result based on the initial confidence score and a correction value of the candidate speech recognition result, the correction value being determined using a re-scoring model; and
determining a speech recognition result from the plurality of candidate speech recognition results based on the confidence scores of the plurality of candidate speech recognition results, wherein
the decoding model has a directed graph structure including a plurality of arcs, each arc having a weight, each candidate speech recognition result is represented by a set of arcs in the decoding model, the confidence score of each candidate speech recognition result is determined based on the weights of the set of arcs corresponding to the candidate speech recognition result,
the re-scoring model has a second directed graph structure including a plurality of second arcs, each second arc having a second weight,
the determining a correction value of the candidate speech recognition result using a re-scoring model comprises:
determining, from the plurality of second arcs in the re-scoring model, a second set of arcs corresponding to the set of arcs of the candidate speech recognition result; and
determining the correction value of the candidate speech recognition result based on the second weights of the plurality of second arcs.
US 12,223,945 claims 1, 10, and 11:
obtaining speech data, the speech data including speech of a plurality of speakers;
obtaining a plurality of candidate speech recognition results and an initial confidence score of each of the plurality of candidate speech recognition results by processing the feature information using a decoding model;
for each of the plurality of candidate speech recognition results,
determining a correction value of the candidate speech recognition result using a re-scoring model; and
determining the confidence score of the candidate speech recognition result by summing the initial confidence score and the correction value of the candidate speech recognition result; and
determining a speech recognition result from the plurality of candidate speech recognition results based on the confidence scores of the plurality of candidate speech recognition results,
10. The system of claim 1, wherein
the decoding model has a directed graph structure including a plurality of arcs, each arc having a weight,
each candidate speech recognition result is represented by a set of arcs in the decoding model, and
the confidence score of each candidate speech recognition result is determined based on the weights of the set of arcs corresponding to the candidate speech recognition result.
11. The system of claim 10, wherein
the re-scoring model has a second directed graph structure including a plurality of second arcs, each second arc having a second weight,
the determining a correction value of the candidate speech recognition result using a re-scoring model comprises:
determining, from the plurality of second arcs in the re-scoring model, a second set of arcs corresponding to the set of arcs of the candidate speech recognition result; and
determining the correction value of the candidate speech recognition result based on the second weights of the plurality of second arcs.
As the table above demonstrates each limitation of claim 21 of the present application is found in claims 1-4 of US Patent 12,223,945, thus claim 21 of the present application is anticipated by claims 1-4 of US Patent 12,223,945. Each limitation of claim 32 of the present application is found in claim 12 of US Patent 12,223,945, thus claim 32 of the present application is anticipated by claim 12 of US Patent 12,223,945. Each limitation of claim 33 of the present application is found in claims 1, 10, and 11 of US Patent 12,223,945, thus claim 33 of the present application is anticipated by claims 1, 10, and 11 of US Patent 12,223,945. Dependent claims 22-31 and 34-40 of the present application recite subject matter that corresponds to that found in claims 1-3, 5-8, and 9 of US Patent 12,223,945 as follows, and therefore are also anticipated.
(Present application) (US Patent 12,223,945)
22. The system of claim 21, wherein the determining a corresponding relationship between the plurality of words and the plurality of speakers comprises: determining a speaking time of each of the plurality of speakers by processing the speech data; and determining, based on the speaking times of the plurality of speakers and the speech recognition result, a corresponding relationship between the plurality of words and the plurality of speakers.
2. The system of claim 1, the operations further comprising:
determining speaking time of each of the plurality of speakers by processing the speech data;
determining, based on the speaking times of the plurality of speakers and the speech recognition result, a corresponding relationship between the plurality of words and the plurality of speakers;
determining, based on the corresponding relationship, at least one conversion word from the plurality of words, each of the at least one conversion word corresponding to at least two of the plurality of speakers; and
re-determining the corresponding relationship between the plurality of words and the plurality of speakers based on the at least one conversion word.
23. The system of claim 22, wherein the determining speaking time of each of the plurality of speakers by processing the speech data includes: generating at least one effective speech segment including speech content by preprocessing the speech data; and determining the speaking time of each of the plurality of speakers by processing the at least one effective speech segment.
6. The system of claim 2, wherein the determining speaking time of each of the plurality of speakers by processing the speech data includes:
generating at least one effective speech segment including speech content by preprocessing the speech data; and
determining the speaking time of each of the plurality of speakers by processing the at least one effective speech segment.
24. The system of claim 23, wherein the generating at least one effective speech segment including speech content by preprocessing the speech data includes: obtaining a plurality of first speech segments by processing the speech data based on the speech recognition result; obtaining a plurality of second speech segments by amplifying the plurality of first speech segments; and obtaining the at least one effective speech segment by performing a processing operation on the plurality of second speech segments, the processing operation including at least one of a merging operation or a segmentation operation.
7. The system of claim 6, wherein the generating at least one effective speech segment including speech content by preprocessing the speech data includes:
obtaining a plurality of first speech segments by processing the speech data based on the speech recognition result;
obtaining a plurality of second speech segments by amplifying the plurality of first speech segments; and
obtaining the at least one effective speech segment by performing a processing operation on the plurality of second speech segments, the processing operation including at least one of a merging operation or a segmentation operation.
25. The system of claim 23, wherein the determining the speaking time of each of the plurality of speakers by processing the at least one effective speech segment includes: extracting a voiceprint feature vector of each of the at least one effective speech segment; and determining the speaking time of each of the plurality of speakers by clustering the voiceprint feature vector of the at least one effective speech segment.
8. The system of claim 6, wherein the determining the speaking time of each of the plurality of speakers by processing the at least one effective speech segment includes:
extracting a voiceprint feature vector of each of the at least one effective speech segment; and
determining the speaking time of each of the plurality of speakers by clustering the voiceprint feature vector of the at least one effective speech segment.
26. The system of claim 21, wherein to obtain the speech recognition result of the speech data, the at least one processor is configured to cause the system to perform operations including: determining feature information of the speech data; determining, based on the feature information, a plurality of candidate speech recognition results and a confidence score of each of the plurality of candidate speech recognition results using a decoding model and a re-scoring model; and determining the speech recognition result from the plurality of candidate speech recognition results based on the confidence scores of the plurality of candidate speech recognition results.
1…
determining feature information of the speech data;
obtaining a plurality of candidate speech recognition results and an initial confidence score of each of the plurality of candidate speech recognition results by processing the feature information using a decoding model;
determining a speech recognition result from the plurality of candidate speech recognition results based on the confidence scores of the plurality of candidate speech recognition results,
27. The system of claim 26, wherein the determining, based on the feature information, a plurality of candidate speech recognition results and a confidence score of each of the plurality of candidate speech recognition results using a decoding model and a re-scoring model includes: obtaining the plurality of candidate speech recognition results and an initial confidence score of each of the plurality of candidate speech recognition results by processing the feature information using the decoding model; determining a correction value of each of the plurality of candidate speech recognition results using the re-scoring model; and for each of the plurality of candidate speech recognition results, determining the confidence score of the candidate speech recognition result based on the corresponding initial confidence score and the corresponding correction value.
1…
obtaining a plurality of candidate speech recognition results and an initial confidence score of each of the plurality of candidate speech recognition results by processing the feature information using a decoding model;
for each of the plurality of candidate speech recognition results,
determining a correction value of the candidate speech recognition result using a re-scoring model; and
determining the confidence score of the candidate speech recognition result by summing the initial confidence score and the correction value of the candidate speech recognition result; and
determining a speech recognition result from the plurality of candidate speech recognition results based on the confidence scores of the plurality of candidate speech recognition results,
28. The system of claim 26, wherein the re-scoring model is generated based on a process of training model, the process of training model including: obtaining a first language model and a second language model, and the first language model being obtained by cropping the second language model; generating the decoding model based on the first language model; generating a first scoring model and a second scoring model based on the first language model and the second language model; and obtaining the re-scoring model based on the first scoring model and the second scoring model.
1…
obtaining a first language model and a second language model, and the first language model being obtained by cropping the second language model;
generating the decoding model based on the first language model;
generating a first scoring model and a second scoring model based on the first language model and the second language model; and
obtaining the re-scoring model based on the first scoring model and the second scoring model,
29. The system of claim 28, wherein the obtaining the re-scoring model based on the first scoring model and the second scoring model includes: obtaining a plurality of possible speech recognition results and a first confidence score of each of the plurality of possible speech recognition results by traversing the second scoring model; for each of the plurality of possible speech recognition results, determining a second confidence score of the possible speech recognition result based on the first scoring model; updating the second scoring model based on the first confidence scores and the second confidence scores of the plurality of possible speech recognition results; and determining the re-scoring model based on the updated second scoring model.
1…
obtaining a plurality of possible speech recognition results and a first confidence score of each of the plurality of possible speech recognition results by traversing the second scoring model;
for each of the plurality of possible speech recognition results, determining a second confidence score of the possible speech recognition result based on the first scoring model;
updating the second scoring model based on the first confidence scores and the second confidence scores of the plurality of possible speech recognition results; and
determining the re-scoring model based on the updated second scoring model.
30. The system of claim 29, wherein the first scoring model includes a plurality of first arcs, the second scoring model includes a plurality of second arcs, and the each of the plurality of possible speech recognition results is represented by a sequence of at least one second arc, and for each of the plurality of possible speech recognition results, the determining a second confidence score of the possible speech recognition result based on the first scoring model includes: for the each of the plurality of possible speech recognition results, determining, in the first scoring model, a sequence of at least one first arc corresponding to the sequence of at least one second arc of the possible speech recognition result; and determining the second confidence score of the possible speech recognition result based on the sequence of at least one first arc.
9. The system of claim 1, wherein the first scoring model includes a plurality of first arcs, the second scoring model includes a plurality of second arcs, and the each of the plurality of possible speech recognition results is represented by a sequence of at least one second arc, and
for each of the plurality of possible speech recognition results, the determining a second confidence score of the possible speech recognition result based on the first scoring model includes:
for the each of the plurality of possible speech recognition results,
determining, in the first scoring model, a sequence of at least one first arc corresponding to the sequence of at least one second arc of the possible speech recognition result; and
determining the second confidence score of the possible speech recognition result based on the sequence of at least one first arc.
31. The system of claim 21, the operations further comprising: obtaining voiceprint feature information of each of the plurality of speakers; and for each of the at least one conversion word, determining the speaker corresponding to the conversion word based on the target adjacent word corresponding to the conversion word and the voiceprint feature information of the plurality of speakers and voiceprint feature information of a target speech segment, the target speech segment including the conversion word.
5. The system of claim 2, wherein the re-determining the corresponding relationship between the plurality of words and the plurality of speakers based on the at least one conversion word includes:
obtaining voiceprint feature information of each of the plurality of speakers; and
for each of the at least one conversion word, determining the speaker corresponding to the conversion word based on the voiceprint feature information of the plurality of speakers and voiceprint feature information of a target speech segment, the target speech segment including the conversion word.
34. The system of claim 33, the operations further comprising: determining speaking time of each of the plurality of speakers by processing the speech data; determining, based on the speaking times of the plurality of speakers and the speech recognition result, a corresponding relationship between the plurality of words and the plurality of speakers; determining, based on the corresponding relationship, at least one conversion word from a plurality of words of the speech recognition result, each of the at least one conversion word corresponding to at least two of the plurality of speakers; and re-determining the corresponding relationship between the plurality of words and the plurality of speakers based on the at least one conversion word.
2. The system of claim 1, the operations further comprising:
determining speaking time of each of the plurality of speakers by processing the speech data;
determining, based on the speaking times of the plurality of speakers and the speech recognition result, a corresponding relationship between the plurality of words and the plurality of speakers;
determining, based on the corresponding relationship, at least one conversion word from the plurality of words, each of the at least one conversion word corresponding to at least two of the plurality of speakers; and
re-determining the corresponding relationship between the plurality of words and the plurality of speakers based on the at least one conversion word.
35. The system of claim 34, wherein the re-determining the corresponding relationship between the plurality of words and the plurality of speakers based on the at least one conversion word includes: for each of the at least one conversion word, determining a speaker corresponding to the conversion word based on a time interval between the conversion word and at least one adjacent word of the conversion word.
3. The system of claim 2, wherein the re-determining the corresponding relationship between the plurality of words and the plurality of speakers based on the at least one conversion word includes:
for each of the at least one conversion word, determining a speaker corresponding to the conversion word based on a time interval between the conversion word and at least one adjacent word of the conversion word.
36. The system of claim 34, wherein the determining speaking time of each of the plurality of speakers by processing the speech data includes: generating at least one effective speech segment including speech content by preprocessing the speech data; and determining the speaking time of each of the plurality of speakers by processing the at least one effective speech segment.
6. The system of claim 2, wherein the determining speaking time of each of the plurality of speakers by processing the speech data includes:
generating at least one effective speech segment including speech content by preprocessing the speech data; and
determining the speaking time of each of the plurality of speakers by processing the at least one effective speech segment.
37. The system of claim 36, wherein the generating at least one effective speech segment including speech content by preprocessing the speech data includes: obtaining a plurality of first speech segments by processing the speech data based on the speech recognition result; obtaining a plurality of second speech segments by amplifying the plurality of first speech segments; and obtaining the at least one effective speech segment by performing a processing operation on the plurality of second speech segments, the processing operation including at least one of a merging operation or a segmentation operation.
7. The system of claim 6, wherein the generating at least one effective speech segment including speech content by preprocessing the speech data includes:
obtaining a plurality of first speech segments by processing the speech data based on the speech recognition result;
obtaining a plurality of second speech segments by amplifying the plurality of first speech segments; and
obtaining the at least one effective speech segment by performing a processing operation on the plurality of second speech segments, the processing operation including at least one of a merging operation or a segmentation operation.
38. The system of claim 33, wherein the re-scoring model is generated based on a process of training model, the process of training model including: obtaining a first language model and a second language model, and the first language model being obtained by cropping the second language model; generating the decoding model based on the first language model; generating a first scoring model and a second scoring model based on the first language model and the second language model; and obtaining the re-scoring model based on the first scoring model and the second scoring model.
1…
obtaining a first language model and a second language model, and the first language model being obtained by cropping the second language model;
generating the decoding model based on the first language model;
generating a first scoring model and a second scoring model based on the first language model and the second language model; and
obtaining the re-scoring model based on the first scoring model and the second scoring model,
39. The system of claim 38, wherein the obtaining the re-scoring model based on the first scoring model and the second scoring model includes: obtaining a plurality of possible speech recognition results and a first confidence score of each of the plurality of possible speech recognition results by traversing the second scoring model; for each of the plurality of possible speech recognition results, determining a second confidence score of the possible speech recognition result based on the first scoring model; updating the second scoring model based on the first confidence scores and the second confidence scores of the plurality of possible speech recognition results; and determining the re-scoring model based on the updated second scoring model.
1…
wherein the obtaining the re-scoring model based on the first scoring model and the second scoring model includes:
obtaining a plurality of possible speech recognition results and a first confidence score of each of the plurality of possible speech recognition results by traversing the second scoring model;
for each of the plurality of possible speech recognition results, determining a second confidence score of the possible speech recognition result based on the first scoring model;
updating the second scoring model based on the first confidence scores and the second confidence scores of the plurality of possible speech recognition results; and
determining the re-scoring model based on the updated second scoring model.
40. The system of claim 39, wherein the first scoring model includes a plurality of first arcs, the second scoring model includes a plurality of second arcs, and the each of the plurality of possible speech recognition results is represented by a sequence of at least one second arc, and for each of the plurality of possible speech recognition results, the determining a second confidence score of the possible speech recognition result based on the first scoring model includes: for the each of the plurality of possible speech recognition results, determining, in the first scoring model, a sequence of at least one first arc corresponding to the sequence of at least one second arc of the possible speech recognition result; and determining the second confidence score of the possible speech recognition result based on the sequence of at least one first arc.
9. The system of claim 1, wherein the first scoring model includes a plurality of first arcs, the second scoring model includes a plurality of second arcs, and the each of the plurality of possible speech recognition results is represented by a sequence of at least one second arc, and
for each of the plurality of possible speech recognition results, the determining a second confidence score of the possible speech recognition result based on the first scoring model includes:
for the each of the plurality of possible speech recognition results,
determining, in the first scoring model, a sequence of at least one first arc corresponding to the sequence of at least one second arc of the possible speech recognition result; and
determining the second confidence score of the possible speech recognition result based on the sequence of at least one first arc.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 21, 22, and 31 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter.
Claim 21 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim recites “obtaining speech data and a speech recognition result of the speech data, the speech data including speech of a plurality of speakers, and the speech recognition result including a plurality of words; determining, based on the speech data and the speech recognition result, a corresponding relationship between the plurality of words and the plurality of speakers; determining, based on the corresponding relationship, at least one conversion word from the plurality of words, each of the at least one conversion word corresponding to at least two of the plurality of speakers; and for each of the at least one conversion word, determining a speaker corresponding to the conversion word based on a time interval between the conversion word and at least one adjacent word of the conversion word by: determining a target adjacent word from the at least one adjacent word of the conversion word, the time interval between the adjacent word and the conversion words being less than a preset length of time, and the target adjacent word not being the conversion word; and determining the speaker corresponding to the conversion word based on the target adjacent word”.
The limitation of obtaining speech data and a speech recognition result of the speech data, the speech data including speech of a plurality of speakers, and the speech recognition result including a plurality of words, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, “obtaining speech data and a speech recognition result of the speech data, the speech data including speech of a plurality of speakers, and the speech recognition result including a plurality of words” in the context of this claim encompasses e.g. obtaining a paper transcript of a conversation between multiple speakers having words and timestamps.
Similarly, the limitation of “determining, based on the speech data and the speech recognition result, a corresponding relationship between the plurality of words and the plurality of speakers”, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, “determining, based on the speech data and the speech recognition result, a corresponding relationship between the plurality of words and the plurality of speakers” in the context of this claim encompasses mentally determining, based on the speech data and the speech recognition result, a corresponding relationship between the plurality of words and the plurality of speakers, for example by mentally determining which speakers said which words based on the time stamps.
Similarly, the limitation of “determining, based on the corresponding relationship, at least one conversion word from the plurality of words, each of the at least one conversion word corresponding to at least two of the plurality of speakers”, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, “determining, based on the corresponding relationship, at least one conversion word from the plurality of words, each of the at least one conversion word corresponding to at least two of the plurality of speakers” in the context of this claim encompasses mentally determining, based on the corresponding relationship, at least one conversion word from the plurality of words, each of the at least one conversion word corresponding to at least two of the plurality of speakers, by mentally determining e.g. words which could have been spoken by at least two of the speakers.
Similarly, the limitation of “for each of the at least one conversion word, determining a speaker corresponding to the conversion word based on a time interval between the conversion word and at least one adjacent word of the conversion word by: determining a target adjacent word from the at least one adjacent word of the conversion word, the time interval between the adjacent word and the conversion words being less than a preset length of time, and the target adjacent word not being the conversion word”, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, “for each of the at least one conversion word, determining a speaker corresponding to the conversion word based on a time interval between the conversion word and at least one adjacent word of the conversion word by: determining a target adjacent word from the at least one adjacent word of the conversion word, the time interval between the adjacent word and the conversion words being less than a preset length of time, and the target adjacent word not being the conversion word” in the context of this claim encompasses, for each of the at least one conversion word, mentally determining a speaker corresponding to the conversion word based on a time interval between the conversion word and at least one adjacent word of the conversion word by: mentally determining a target adjacent word from the at least one adjacent word of the conversion word, the time interval between the adjacent word and the conversion words being less than a preset length of time, and the target adjacent word not being the conversion word by, e.g. scanning the transcript for adjacent words and mentally determining the time intervals based on the timestamps for find words spoken within a preset length of time of the conversion words.
Similarly, the limitation of “determining the speaker corresponding to the conversion word based on the target adjacent word”, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, “determining the speaker corresponding to the conversion word based on the target adjacent word” in the context of this claim encompasses mentally determining the speaker corresponding to the conversion word based on the target adjacent word.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea. This judicial exception is not integrated into a practical application. In particular, the claim only recites three additional elements – “storage device”, “set of instructions” and “processor”. The computing elements in this step are recited at a high-level of generality (i.e., as a generic storage device, a generic set of instructions, and generic processor) such that they amount to no more than mere instructions to apply the exception using generic computer elements. Accordingly, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of using a processor executing instructions stored on a storage device to perform the obtaining and determining amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The claim is not patent eligible.
Specifically with respect to Step 2A, Prong Two, of the Alice/Mayo test, the judicial exception is not integrated into a practical application. Claim 1 does not recite any limitations that are not mental steps.
Specifically with respect to Step 2B of the Alice/Mayo test, “the claim as a whole does not amount to significantly more than the exception itself (there is no inventive concept in the claim)”. MPEP 2106.05 Il. There are no limitations in claim 1 outside of the judicial exception. As a whole, there does not appear to contain any inventive concept. As discussed above, claim 1 is a mental process that pertains to the mental process of determining speakers corresponding to words, which can be performed entirely by a human with physical aids.
Dependent claims 22 depends from claim 21, does not remedy any of the deficiencies of claim 21, and therefore is rejected on the same grounds as claim 21 above.
Claim 22 merely recites additional steps for making determinations between words and speakers, all of which could be performed mentally or by writing down relationships with a pen and paper, and do not amount to anything more than substantially the same abstract idea as explained with respect to claim 21. Specifically:
Claim 22 recites “determining a corresponding relationship between the plurality of words and the plurality of speakers comprises: determining a speaking time of each of the plurality of speakers by processing the speech data; and determining, based on the speaking times of the plurality of speakers and the speech recognition result, a corresponding relationship between the plurality of words and the plurality of speakers” which could be performed by mentally determining a speaking time of each of the plurality of speakers by mentally processing the speech data; and mentally determining, based on the speaking times of the plurality of speakers and the speech recognition result, a corresponding relationship between the plurality of words and the plurality of speakers.
As seen above, claim 22 depends from claim 21 and further recites mental processes as explained above. None of the additional limitations recited in claim 22 amount to anything more than the same or a similar abstract idea as recited in claim 21. Nor do any limitations in claim 22 (a) integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea or (b) amount to significantly more than the judicial exception. Claim 22 is not patent eligible.
Dependent claim 31 depends from claim 21, does not remedy any of the deficiencies of claim 21, and therefore is rejected on the same grounds as claim 21 above.
Claim 31 merely recites additional steps for making determinations between words and speakers, all of which could be performed mentally or by writing down relationships with a pen and paper, and do not amount to anything more than substantially the same abstract idea as explained with respect to claim 21. Specifically:
Claim 31 recites “obtaining voiceprint feature information of each of the plurality of speakers; and for each of the at least one conversion word, determining the speaker corresponding to the conversion word based on the target adjacent word corresponding to the conversion word and the voiceprint feature information of the plurality of speakers and voiceprint feature information of a target speech segment, the target speech segment including the conversion word” which could be performed by obtaining “voiceprint feature information” such as a written voiceprint label from a sheet of paper of each of the plurality of speakers, and for each conversion word, mentally determining the speaker corresponding to the conversion word based on the target adjacent word corresponding to the conversion word and the voiceprint feature information of the plurality of speakers and voiceprint feature information of a target speech segment, the target speech segment including the conversion word. Notably, the claim only requires “obtaining voiceprint feature information” and does not actually require extracting a voiceprint itself from the speech audio.
As seen above, claim 31 depends from claim 21 and further recites mental processes as explained above. None of the additional limitations recited in claim 31 amount to anything more than the same or a similar abstract idea as recited in claim 21. Nor do any limitations in claim 31 (a) integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea or (b) amount to significantly more than the judicial exception. Claim 31 is not patent eligible.
Eligible Claims
Claims 23-30 and 32-40 each recite speech processing operations which given their broadest reasonable interpretation in light of the specification, would be impractical or impossible to perform as a mental process.
Specifically, in claim 23, “preprocessing” the speech data is a well known term in the art that refers to techniques for processing the audio data of the speech itself prior to further analysis. Interpretations that would allow for some sort of mental “preprocessing” the speech data would be unreasonable.
Claims 24-25 depend on claim 23 and also recite additional speech processing techniques which cannot be practically performed mentally.
Claim 26 requires performing speech recognition operations which cannot be performed mentally.
Claims 27-30 depend on claim 26 and therefore cannot be performed mentally for at least the above reason.
Claims 32 and 33 require performing speech recognition operations which cannot be performed mentally.
Claims 34-40 depend on claim 33 and therefore cannot be performed mentally for at least the above reason.
Allowable Subject Matter
Claims 21, 22, and 31 would be allowable if a proper terminal disclaimer were to overcome the double patenting rejections, as well as amended to overcome the 35 U.S.C. 101 rejections.
Claims 23-40 would be allowable if a proper terminal disclaimer were to overcome the double patenting rejections.
The following is the examiner’s statement of reasons for indicating subject matter allowable over the prior art of record:
The closest prior art to independent claim 21 is Jung et al. (US 20190392837).
Consider claim 21, Jung discloses a speech recognition system (voice recognition by a system, [0002-0003]), comprising:
at least one storage device storing a set of instructions (storage device storing instructions, [0062-0063]); and
at least one processor in communication with the at least one storage device, wherein when executing the set of instructions (processor executing instructions, [0060]), the at least one processor is configured to cause the system to perform operations including:
obtaining speech data and a speech recognition result of the speech data, the speech data including speech of a plurality of speakers, and the speech recognition result including a plurality of words (utterances from multiple users participating in a conversation are received, [0083], the utterances are converted into text words, [0085-0086], [0098]);
determining, based on the speech data and the speech recognition result, a corresponding relationship between the plurality of words and the plurality of speakers (determining that before Lisa R is able to complete the utterance at time t6, Tim G. begins speaking, [0052]);
determining, based on the corresponding relationship, at least one conversion word from the plurality of words, each of the at least one conversion word corresponding to at least two of the plurality of speakers (determining that utterance 314 is an interruption with regard to utterance 312 because two voices are detected during the same time period, [0052], and include a second set of words interrupted by a third set of words, [0089], noting that Applicant’s original specification at [0080] uses “conversion word” to refer to overlapping speech); and
for each of the at least one conversion word, determining a speaker corresponding to the conversion word based on a time interval between the conversion word and at least one adjacent word of the conversion word (first utterance and second utterance spoken during a same time period, [0049], [0107]).
However, Jung does not disclose determining a target adjacent word from the at least one adjacent word of the conversion word, the time interval between the adjacent word and the conversion words being less than a preset length of time, and the target adjacent word not being the conversion word; and determining the speaker corresponding to the conversion word based on the target adjacent word.
The closest prior art to independent claims 32 and 33 is Jung et al. (US 20190392837) and Itoh (US 20150279353).
Jung discloses a speech recognition system (voice recognition by a system, [0002-0003]), comprising:
at least one storage device storing a set of instructions (storage device storing instructions, [0062-0063]); and
at least one processor in communication with the at least one storage device, wherein when executing the set of instructions (processor executing instructions, [0060]), the at least one processor is configured to cause the system to perform operations including:
obtaining speech data to be recognized (utterances from multiple users participating in a conversation are received, [0083]);
determining a plurality of candidate speech recognition results of the speech data (the utterances are converted into text words, [0085-0086], [0098]).
Itoh discloses: obtaining a plurality of candidate speech recognition results and an initial confidence score of each of the plurality of candidate speech recognition results by processing the feature information using a decoding model (producing likelihood of phonemes or text using an acoustic model inherently requires processing feature information using a decoding model, since some acoustic feature must be decoded using the acoustic model, [0049]);
for each of the plurality of candidate speech recognition results, determining a correction value of the candidate speech recognition result using a re-scoring model (the language model likelihood is considered “a correction value”, [0049]);
determining the confidence score of the candidate speech recognition result by summing the initial confidence score and the correction value of the candidate speech recognition result (the sum of the likelihood of acoustic model and the likelihood of language model, [0049]); and
determining a speech recognition result from the plurality of candidate speech recognition results based on the confidence scores of the plurality of candidate speech recognition results, the speech recognition result including a plurality of words (text of the recognition results, [0049-0051]).
However, Jung and Itoh do not disclose “the process of training model including: obtaining a first language model and a second language model, and the first language model being obtained by cropping the second language model; generating the decoding model based on the first language model; generating a first scoring model and a second scoring model based on the first language model and the second language model; and obtaining the re-scoring model based on the first scoring model and the second scoring model, wherein the obtaining the re-scoring model based on the first scoring model and the second scoring model includes: obtaining a plurality of possible speech recognition results and a first confidence score of each of the plurality of possible speech recognition results by traversing the second scoring model; for each of the plurality of possible speech recognition results, determining a second confidence score of the possible speech recognition result based on the first scoring model; updating the second scoring model based on the first confidence scores and the second confidence scores of the plurality of possible speech recognition results; and determining the re-scoring model based on the updated second scoring model” as recited by claim 32, or “…determining a speech recognition result from the plurality of candidate speech recognition results based on the confidence scores of the plurality of candidate speech recognition results, wherein the decoding model has a directed graph structure including a plurality of arcs, each arc having a weight, each candidate speech recognition result is represented by a set of arcs in the decoding model, the confidence score of each candidate speech recognition result is determined based on the weights of the set of arcs corresponding to the candidate speech recognition result, the re-scoring model has a second directed graph structure including a plurality of second arcs, each second arc having a second weight, the determining a correction value of the candidate speech recognition result using a re-scoring model comprises: determining, from the plurality of second arcs in the re-scoring model, a second set of arcs corresponding to the set of arcs of the candidate speech recognition result; and determining the correction value of the candidate speech recognition result based on the second weights of the plurality of second arcs” as recited by claim 33.
Claims 22-31 and 34-40 contain allowable subject matter because they further limit the allowable subject matter in parent claims 21 and 33.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US 11670304 Khoury discloses speaker recognition at a call center by diarization using clustering
US 9368109 Colibro discloses automatic speaker-based speech clustering
US 11152006 Krupka discloses voice identification enrollment using voiceprints
US 11636860 Gorodetski discloses word-level blind diarization of recorded calls with arbitrary number of speakers
US 10366693 Gorodetski discloses acoustic signature building for a speaker through multiple sessions
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jesse Pullias whose telephone number is 571/270-5135. The examiner can normally be reached on M-F 8:00 AM - 4:30 PM. The examiner’s fax number is 571/270-6135.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Andrew Flanders can be reached on 571/272-7516.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free).
/Jesse S Pullias/
Primary Examiner, Art Unit 2655 09/14/26