DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1 - 20 are pending and claims 1 and 11 are independent claims.
Information Disclosure Statement
The information disclosure statements (IDS) submitted on 08/19/2025 and 01/31/2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 10, 11 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Shafey et al. Pat App No. US 20220199094 A1 (Shafey) in view of Lin Pat App No. US 20230320642 A1 (Lin).
Regarding Claim 1, A computer-implemented method executed on data processing hardware that causes the data processing hardware to perform operations (Shafey, para 0067, For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions) comprising:
obtaining a series of segmented labeled training samples (Shafey, para 0048, training data that includes training input audio segment sequences), each respective segmented labeled training sample comprising one or more spoken terms spoken during a conversation by multiple speakers (Shafey, para 0048-0049, training data that includes training input audio segment sequences and, for each training input audio segment sequence, a corresponding output target.), each respective spoken term characterized by a corresponding sequence of acoustic frames and paired-with a corresponding transcription of the respective spoken term and a corresponding speaker label representing an identity of a respective speaker that spoke the respective spoken term during the conversation (Shafey, para 0052-0057, the audio segment sequence includes a plurality of audio frames. For example, each audio frame can be a d dimensional log-mel filterbank energy, where d is a fixed constant, e.g., fifty, eighty, or one hundred, or a different acoustic feature representation of the corresponding portion of the audio segment… The system then determines, from the output sequence, a transcription of the audio segment data that identifies (i) words spoken in the audio segment… FIG. 4 shows an example transcription 400 generated using the joint ASR-SD neural network);
and
for each respective segmented labeled training sample (Shafey, para 0048, training data that includes training input audio segment sequences):
obtaining a corresponding 120 on training data that includes training input audio segment sequences and, for each training input audio segment sequence… The system obtains an audio segment sequence characterizing an audio segment (step 302). The audio segment may be an entire conversation or a fixed length, e.g., ten, fifteen, or thirty second, portion of a larger conversation) the corresponding
generating, as output from a joint speech recognition and speaker diarization model, by performing cross-attention on the respective segmented labeled training sample and the corresponding dynamic audio cohort, diarization results comprising a corresponding speech recognition result comprising one or more predicted terms, each respective predicted term associated with a corresponding speaker token representing a predicted identity of a speaker that spoke the respective predicted term);
generating an updated
training the joint speech recognition and speaker diarization model based on a loss derived from the generated diarization results, the corresponding transcriptions, and the corresponding speaker labels (Shafey, para 0015-0017, FIG. 1 shows an example speech processing system 100. This system 100 generates transcriptions of audio data. In particular, the transcriptions generated by the system 100 identify the words spoken in a given audio segment and, for each of the spoken words, the speaker that spoke the word, i.e., identify the role of the speaker that spoke the word in a conversation or uniquely identify an individual speaker. More specifically, the system 100 performs joint automatic speech recognition (ASR) and speaker diarization (SD) by transducing, i.e., mapping, an input sequence of audio data 110 to an output sequence 150 of output symbols using a neural network 120. This neural network 120 is referred to in this specification as “a joint ASR—SD neural network”; Shafey, para 0049, the system can optimize an objective function that measures the conditional probability assigned to the ground truth output sequence [i.e., “performing joint automatic speech recognition (ASR) and speaker diarization (SD)” follows “training joint automatic speech recognition (ASR) and speaker diarization (SD)” comes; “the objective function” as “loss derived from the generated diarization results”]).
Shafey does not specifically disclose obtaining a corresponding dynamic audio cohort associated with an immediately prior segmented labeled training sample, the corresponding dynamic audio cohort comprising a matrix of audio speech snippets of speakers that spoke prior to the respective segmented labeled training sample, and generating an updated dynamic audio cohort based on the diarization results;
However, Lin, in the same field of endeavor, discloses:
obtaining a corresponding dynamic audio cohort associated with an immediately prior segmented labeled training sample (Lin, para 0011-0012, to dynamically manage the dialogue session in real-time by identifying, in response to the current speech segment…obtain audio data for a patient-therapist dialogue session), the corresponding dynamic audio cohort comprising a matrix of audio speech snippets of speakers that spoke prior to the respective segmented labeled training sample (Lin, para 0093, Framework 2, as described herein, may include a real-time AI system to conduct sentence-level quality assurance of conversational alignment based on speaker-diarized dialogues transcribed from automatic speech recognition of a continuous audio stream(s). In such embodiments, the framework may utilize an online registration-free speaker-diarization engine to perform separation of speech utterances of multiple speakers in the conversations);
generating an updated dynamic audio cohort based on the diarization results (Lin, para 0093, a real-time AI system to conduct sentence-level quality assurance of conversational alignment based on speaker-diarized dialogues transcribed from automatic speech recognition of a continuous audio stream(s). In such embodiments, the framework may utilize an online registration-free speaker-diarization engine to perform separation of speech utterances of multiple speakers in the conversations, that learns from user feedback. a self-supervision process assigns a pseudo-action upon which the reward mapping is updated…generates new arms by transferring the learned arm parameters for similar profiles given the user feedbacks);
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Lin in the method of Shafey because this would have the benefit of analyzing psychotherapy data that includes obtaining transcript data representative of spoken dialog in one or more psychotherapy sessions conducted between a patient and a therapist, extracting speech segments from the transcript data related to one or more of the patient or the therapist, and applying a trained machine learning topic model process to the extracted speech segments (Lin, Abstract).
Regarding Claim 11, Shafey discloses a system comprising,
data processing hardware;
memory hardware in communication with the data processing hardware, the
memory hardware storing instruction that when executed on the data processing hardware cause the data processing hardware to perform operations (Shafey, para 0067, For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions) comprising:
obtaining a series of segmented labeled training samples (Shafey, para 0048, training data that includes training input audio segment sequences), each respective segmented labeled training sample comprising one or more spoken terms spoken during a conversation by multiple speakers (Shafey, para 0048-0049, training data that includes training input audio segment sequences and, for each training input audio segment sequence, a corresponding output target), each respective spoken term characterized by a corresponding sequence of acoustic frames and paired with a corresponding transcription of the respective spoken term and a corresponding speaker label representing an identity of a respective speaker that spoke the respective spoken tem1 during the conversation (Shafey, para 0052-0057, the audio segment sequence includes a plurality of audio frames. For example, each audio frame can be a d dimensional log-mel filterbank energy, where d is a fixed constant, e.g., fifty, eighty, or one hundred, or a different acoustic feature representation of the corresponding portion of the audio segment… The system then determines, from the output sequence, a transcription of the audio segment data that identifies (i) words spoken in the audio segment… FIG. 4 shows an example transcription 400 generated using the joint ASR-SD neural network); and
for each respective segmented labeled training sample (Shafey, para 0048, training data that includes training input audio segment sequences):
obtaining a corresponding 120 on training data that includes training input audio segment sequences and, for each training input audio segment sequence… The system obtains an audio segment sequence characterizing an audio segment (step 302). The audio segment may be an entire conversation or a fixed length, e.g., ten, fifteen, or thirty second, portion of a larger conversation), the corresponding uses a sophisticated technique for determining when the speaker changes during the conversation);
generating, as output from a joint speech recognition and speaker diarization model, by performing cross-attention on the respective segmented labeled training sample and the corresponding 100 performs joint automatic speech recognition (ASR) and speaker diarization (SD) by transducing, i.e., mapping, an input sequence of audio data 110 to an output sequence 150 of output symbols using a neural network 120. This neural network 120 is referred to in this specification as “a joint ASR—SD neural network.” The system 100 is referred to as performing “joint” ASR and SD because a single output sequence 150 generated using the neural network 120 defines both the ASR output for the audio data, i.e., which words are spoken in the audio segment, and the SD output for the audio data, i.e., which speaker spoke each of the words), each respective predicted term associated with a corresponding speaker token representing a predicted identity of a speaker that spoke the respective predicted term (Shafey, para 0037-0038, the prediction neural network 220 can include an embedding layer that maps each non-blank output symbol (and the placeholder output) to a respective embedding followed by one or more uni-directional LSTM or other recurrent layers. In some cases, the last recurrent layer directly generates the prediction representation while in other cases the last recurrent layer is followed by a fully-connected layer that generates the prediction representation. The joint neural network 230 is a neural network that is configured to, at each time step, process (i) the encoded representation for the audio frame at the time step and (ii) the prediction representation for the time step to generate a set of logits l.sub.t,u that includes a respective logit for each of the output symbols in the set of output symbols. As described above, the set of output symbols includes both text symbols and speaker label symbols);
generating an updated
training the joint speech recognition and speaker diarization model based on a loss derived from the generated diarization results, the corresponding transcriptions, and the corresponding speaker labels (Shafey, para 0015-0017, FIG. 1 shows an example speech processing system 100. This system 100 generates transcriptions of audio data. In particular, the transcriptions generated by the system 100 identify the words spoken in a given audio segment and, for each of the spoken words, the speaker that spoke the word, i.e., identify the role of the speaker that spoke the word in a conversation or uniquely identify an individual speaker. More specifically, the system 100 performs joint automatic speech recognition (ASR) and speaker diarization (SD) by transducing, i.e., mapping, an input sequence of audio data 110 to an output sequence 150 of output symbols using a neural network 120. This neural network 120 is referred to in this specification as “a joint ASR—SD neural network”; ”; Shafey, para 0049, the system can optimize an objective function that measures the conditional probability assigned to the ground truth output sequence ; [i.e., “performing joint automatic speech recognition (ASR) and speaker diarization (SD)” follows “training joint automatic speech recognition (ASR) and speaker diarization (SD)” comes; “the objective function” as “loss derived from the generated diarization results”]).
Shafey does not specifically disclose obtaining a corresponding dynamic audio cohort associated with an immediately prior segmented labeled training sample, the corresponding
However, Lin, in the same field of endeavor, discloses:
obtaining a corresponding dynamic audio cohort associated with an immediately prior segmented labeled training sample (Lin, para 0011-0012, to dynamically manage the dialogue session in real-time by identifying, in response to the current speech segment…obtain audio data for a patient-therapist dialogue session), the corresponding dynamic audio cohort comprising a matrix of audio speech snippets of speakers that spoke prior to the respective segmented labeled training sample (Lin, para 0093, Framework 2, as described herein, may include a real-time AI system to conduct sentence-level quality assurance of conversational alignment based on speaker-diarized dialogues transcribed from automatic speech recognition of a continuous audio stream(s). In such embodiments, the framework may utilize an online registration-free speaker-diarization engine to perform separation of speech utterances of multiple speakers in the conversations);
generating an updated dynamic audio cohort based on the diarization results (Lin, para 0093, a real-time AI system to conduct sentence-level quality assurance of conversational alignment based on speaker-diarized dialogues transcribed from automatic speech recognition of a continuous audio stream(s). In such embodiments, the framework may utilize an online registration-free speaker-diarization engine to perform separation of speech utterances of multiple speakers in the conversations, that learns from user feedback. a self-supervision process assigns a pseudo-action upon which the reward mapping is updated…generates new arms by transferring the learned arm parameters for similar profiles given the user feedbacks);
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Lin in the method of Shafey because this would have the benefit of analyzing psychotherapy data that includes obtaining transcript data representative of spoken dialog in one or more psychotherapy sessions conducted between a patient and a therapist, extracting speech segments from the transcript data related to one or more of the patient or the therapist, and applying a trained machine learning topic model process to the extracted speech segments (Lin, Abstract).
Claims 2 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Shafey in view of Lin, and further in view of Wang et al. Pat App No. US 20230089308 A1 (Wang).
Regarding Claim 2, Shafey in view of Lin discloses the computer-implemented method of claim 1.
Shafey in view of Lin does not specifically disclose wherein the matrix of audio speech snippets comprises audio-only data.
However, Wang, in the same field of endeavor, discloses wherein the matrix of audio speech snippets comprises audio-only data (Wang, para 0044-0047, The ASR model 300 processes the input audio signal 122 corresponding to the utterances 120 spoken by the multiple speakers 10 (FIG. 1) to generate the transcriptions 220 of the utterances and the sequence of speaker turn tokens 224...the ASR model 300 processes acoustic information and/or semantic information to detect speaker turns in the input audio signal 122. .. This semantic interpretation of the transcription 220 may be used independently or in conjunction with acoustic processing of the input audio signal 122… The speaker encoder 230 receives the plurality of speaker segments 225 from the segmentation module 210 (FIG. 1) and extracts the corresponding speaker-discriminative embedding 240 for each speaker segment 225. The speaker-discriminative embeddings 240 may include speaker vectors such as d-vectors or i-vectors. The speaker encoder 230 provides the speaker-discriminative embeddings 240 associated with each speaker segment 225 to the clustering module 260… the clustering module 260 constructs a similarity graph by computing pairwise similarities a.sub.ij where A represents the affinity matrix E [AltContent: rect].sup.N×N of the similarity graph).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Wang in the method of Shafey in view of Lin because this would enable receiving an input audio signal that corresponds to utterances spoken by multiple speakers and processing the input audio to generate a transcription of the utterances and a sequence of speaker turn tokens (Wang, Abstract).
Regarding Claim 12, Shafey in view of Lin discloses the system of claim 11.
Shafey in view of Lin does not specifically disclose wherein the matrix of audio speech snippets comprises audio-only data.
However, Wang, in the same field of endeavor, discloses wherein the matrix of audio speech snippets comprises audio-only data (Wang, para 0044-0047, The ASR model 300 processes the input audio signal 122 corresponding to the utterances 120 spoken by the multiple speakers 10 (FIG. 1) to generate the transcriptions 220 of the utterances and the sequence of speaker turn tokens 224...the ASR model 300 processes acoustic information and/or semantic information to detect speaker turns in the input audio signal 122. .. This semantic interpretation of the transcription 220 may be used independently or in conjunction with acoustic processing of the input audio signal 122… The speaker encoder 230 receives the plurality of speaker segments 225 from the segmentation module 210 (FIG. 1) and extracts the corresponding speaker-discriminative embedding 240 for each speaker segment 225. The speaker-discriminative embeddings 240 may include speaker vectors such as d-vectors or i-vectors. The speaker encoder 230 provides the speaker-discriminative embeddings 240 associated with each speaker segment 225 to the clustering module 260… the clustering module 260 constructs a similarity graph by computing pairwise similarities a.sub.ij where A represents the affinity matrix E [AltContent: rect].sup.N×N of the similarity graph).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Wang in the method of Shafey in view of Lin because this would enable receiving an input audio signal that corresponds to utterances spoken by multiple speakers and processing the input audio to generate a transcription of the utterances and a sequence of speaker turn tokens (Wang, Abstract).
Claims 3-4 and 13-14 are rejected under 35 U.S.C. 103 as being unpatentable over Shafey in view of Lin, and further in view of DiMaria et al. Pat App No. US 20220223157 A1 (DiMaria).
Regarding Claim 3, Shafey in view of Lin discloses the computer-implemented method of claim 1.
Shafey in view of Lin does not specifically disclose wherein the matrix of audio speech snippets comprises a predetermined number of slots for each speaker of the multiple speakers that spoke during the conversation, each respective slot configured to store a single audio speech snippet for a respective one of the multiple speakers.
However, DiMaria, in the same field of endeavor, discloses wherein the matrix of audio speech snippets comprises a predetermined number of slots for each speaker of the multiple speakers that spoke during the conversation, each respective slot configured to store a single audio speech snippet for a respective one of the multiple speakers (DiMaria, para 0009-0010, For each audio snippet from the plurality of audio snippets, the processor generates a frequency representation of the audio snippet in a time domain, wherein the frequency representation indicate amplitudes of audio signals in the audio snippet… technology that non-deterministically splits an audio file into a plurality of audio snippets, such that a duration of each audio snippet from the plurality of audio snippets is probabilistically determined to ascertain how short each audio snippet can be to include one or more utterances).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of DiMaria in the method of Shafey in view of Lin because this would enable to identify speakers in live low-frequency audio recordings, such as live phone calls, live audio streams, and the like (DiMaria, para 0003).
Regarding Claim 4, Shafey in view of Lin and DiMaria discloses the computer-implemented method of claim 3.
Furthermore, DiMaria discloses:
wherein each respective slot of the predetermined number of slots is associated with a corresponding probability (DiMaria, para 0008, In an embodiment, a system for identifying a speaker in a multi-speaker environment comprises a processor operably coupled with a memory… The processor splits the audio file into a plurality of audio snippets based at least in part upon a probability of each audio snippet comprising one or more utterances being above a threshold percentage).
Regarding Claim 13, Shafey in view of Lin discloses the system of claim 11.
Shafey in view of Lin does not specifically disclose wherein the matrix of audio speech snippets comprises a predetermined number of slots for each speaker of the multiple speakers that spoke during the conversation, each respective slot configured to store a single audio speech snippet for a respective one of the multiple speakers.
However, DiMaria, in the same field of endeavor, discloses wherein the matrix of audio speech snippets comprises a predetermined number of slots for each speaker of the multiple speakers that spoke during the conversation, each respective slot configured to store a single audio speech snippet for a respective one of the multiple speakers (DiMaria, para 0009-0010, For each audio snippet from the plurality of audio snippets, the processor generates a frequency representation of the audio snippet in a time domain, wherein the frequency representation indicate amplitudes of audio signals in the audio snippet… technology that non-deterministically splits an audio file into a plurality of audio snippets, such that a duration of each audio snippet from the plurality of audio snippets is probabilistically determined to ascertain how short each audio snippet can be to include one or more utterances).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of DiMaria in the method of Shafey in view of Lin because this would enable to identify speakers in live low-frequency audio recordings, such as live phone calls, live audio streams, and the like (DiMaria, para 0003).
Regarding Claim 14, Shafey in view of Lin and DiMaria discloses the system of claim 13.
Furthermore, DiMaria discloses:
wherein each respective slot of the predetermined number of slots is associated with a corresponding probability (DiMaria, para 0008, In an embodiment, a system for identifying a speaker in a multi-speaker environment comprises a processor operably coupled with a memory… The processor splits the audio file into a plurality of audio snippets based at least in part upon a probability of each audio snippet comprising one or more utterances being above a threshold percentage).
Claims 5-7 and 15-17 are rejected under 35 U.S.C. 103 as being unpatentable over Shafey in view of Lin, further in view of DiMaria, and further in view of Biadsy et al. Pat App No. US 20240021190 A1 (Biadsy).
Regarding Claim 5, Shafey in view of Lin and DiMaria disclose the computer-implemented method of claim 4, wherein generating the updated
dynamic audio cohort based on the diarization results comprises.
Shafey in view of Lin and DiMaria do not specifically disclose determining that at least one of the one or more predicted terms of the corresponding speech recognition result was spoken by a new speaker that did not speak during any previous segmented labeled training sample, based on determining that the at least one of the one or more predicted terms of the corresponding speech recognition result was spoken by the new speaker, sampling the corresponding sequence of acoustic frames characterizing the at least one of the one or more predicted terms, and storing the sampled corresponding sequence of acoustic frames as a respective audio speech snippet for the new speaker at one of the predetermined number of slots for the new speaker.
However, Biadsy, in the same field of endeavor, discloses:
determining that at least one of the one or more predicted terms of the corresponding speech recognition result was spoken by a new speaker that did not speak during any previous segmented labeled training sample (Biadsy, para 0025, FIG. 6B is a schematic view of a contrastive unspoken text selection process for selecting unspoken textual utterances used for training a sub-model to bias speech recognition results);
based on determining that the at least one of the one or more predicted terms of the corresponding speech recognition result was spoken by the new speaker (Biadsy, para 0041, the ASR model 200 to make predictions that suit the source speaker 104), sampling the corresponding sequence of acoustic frames characterizing the at least one of the one or more predicted terms (Biadsy, para 0060, The ground truth transcription 563 should accurately reflect the corresponding speech sample (i.e., audio data 561) such that the ground truth transcription 563 is a target output of the sub-model 215); and
storing the sampled corresponding sequence of acoustic frames as a respective audio speech snippet for the new speaker at one of the predetermined number of slots for the new speaker (Biadsy, para 0057, a plurality of speech samples spoken by a variety of different speakers... Further, the training input 510 may be paired with a label 520 indicating a target output associated with the training input 510. In other words, the training input 510 may include a plurality of speech samples corresponding to utterance spoken by different speakers and each speech sample may be paired with a corresponding label 520 indicating a transcription of the corresponding utterance).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Biadsy in the method of Shafey in view of Lin and DiMaria because this would enable training a sub-model for contextual biasing for speech recognition (Biadsy, Abstract).
Regarding Claim 6, Shafey in view of Lin and DiMaria disclose the computer-implemented method of claim 4, wherein generating the updated dynamic audio cohort based on the diarization results comprises:
Shafey in view of Lin and DiMaria do not specifically disclose determining that at least one of the one or more predicted terms of the corresponding speech recognition result was spoken by a respective speaker that did speak during a previous segmented labeled training sample, determining that a current number of audio speech snippets stored for the respective speaker fails to satisfy a threshold of audio speech snippets, based on determining that the current number of snippets stored for the respective speaker fails to satisfy the threshold of audio speech snippets, sampling the corresponding sequence of acoustic frames characterizing the at least one of the one or more predicted terms, and storing the sampled corresponding sequence of acoustic frames as a respective audio snippet for the respective speaker at one of the predetem1ined number of slots.
However, Biadsy, in the same field of endeavor, discloses:
determining that at least one of the one or more predicted terms of the corresponding speech recognition result was spoken by a respective speaker that did speak during a previous segmented labeled training sample (Biadsy, para 0010, predicted speech recognition result and determining a supervised loss term based on the predicted speech recognition result);
determining that a current number of audio speech snippets stored for the respective speaker fails to satisfy a threshold of audio speech snippets (Biadsy, para 0029, ASR models may be trained on large sets of training data including audio samples of speech to produce a robust model for speech recognition… However, there are drawbacks to using such large models, such as applying a single model for a wide variety of users with different characteristics);
based on determining that the current number of snippets stored for the respective speaker fails to satisfy the threshold of audio speech snippets, sampling the corresponding sequence of acoustic frames characterizing the at least one of the one or more predicted terms (Biadsy, para 0063-0067, The base ASR model 200 may receive the sub-model output 569 along with the audio data 561 characterizing the training utterance 560 to produce a predicted speech recognition result 565 (i.e., biased speech recognition results 224). In some implementations, the predicted speech recognition result 565 is used by a loss function 580 to generate a supervised loss term 590. That is, the loss function 580 compares the predicted speech recognition result 565 and the ground truth transcription 563 of the corresponding training utterance 560 to generate the supervised loss term 590, where the loss 590 indicates a discrepancy between the ground truth transcription 563 (i.e., the target output) and the predicted speech recognition result 565… Training utterances 560 may be obtained in a variety of different ways. Typically, training utterances 560 are collected manually, where an audio sample of an utterance is manually transcribed. However, manually labeling training data can be tedious and difficult to collect sufficient samples of labeled data for training); and
storing the sampled corresponding sequence of acoustic frames as a respective audio snippet for the respective speaker at one of the predetermined number of slots (Biadsy, para 0035, an acoustic front-end residing on the user device 110 may convert a time-domain audio waveform of the utterance 108 captured via a microphone of the user device 110 into the input spectrograms 102 or other type or form of audio data 102. Further, the front-end device may be configured to determine or obtain data representing a contextual indicator 103 affecting the utterance 108 and/or other pertinent information corresponding to the source speaker 104).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Biadsy in the method of Shafey in view of Lin and DiMaria because this would enable training a sub-model for contextual biasing for speech recognition (Biadsy, Abstract).
Regarding Claim 7, Shafey in view of Lin and DiMaria disclose the computer-implemented method of claim 4, wherein generating the updated dynamic audio cohort based on the diarization results comprises:
DiMaria further teaches:
determining that a current number of snippets stored for respective speaker satisfies a threshold of audio speech snippets (DiMaria, para 0046, If the correspondence score 174 is more than a threshold score, e.g., 84%, the audio processing engine 120 determines that the feature vector 156 corresponds to the target vector 152. In other words, the audio processing engine 120 determines that one or more utterances 158 spoken in the audio snippet 138 are uttered by the first speaker 102a); and
based on determining that the current number of snippets stored for respective speaker satisfies the threshold of audio speech snippets, sampling a random number from a random number distribution (DiMaria, para 0053, At step 304, the audio processing engine 120 splits the audio file 136 into a plurality of audio snippets 138 based on a probability of each audio snippet 138 comprising one or more utterances 158 being above a threshold percentage. For example, the audio processing engine 120 may determine a length 204 for each audio snippet 138, such that a probability of detecting at least one utterance 158 in above a threshold number of slices 162 in a row being above a threshold percentage, e.g., 80%, similar to that described in FIGS. 1 and 2. The lengths 204 of the audio snippets 138 may vary based on speech distribution patterns 148 associated with speakers 102 in the audio file 136).
Shafey in view of Lin and DiMaria do not specifically disclose determining that at least one of the one or more predicted terms of the corresponding speech recognition result was spoken by a respective speaker that did speak during a previous segmented labeled training sample.
However, Biadsy, in the same field of endeavor, discloses determining that at least one of the one or more predicted terms of the corresponding speech recognition result was spoken by a respective speaker that did speak during a previous segmented labeled training sample (Biadsy, para 0010, predicted speech recognition result and determining a supervised loss term based on the predicted speech recognition result);
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Biadsy in the method of Shafey in view of Lin and DiMaria because this would enable training a sub-model for contextual biasing for speech recognition (Biadsy, Abstract).
Regarding Claim 15, Shafey in view of Lin and DiMaria disclose the system of claim 14, wherein generating the updated dynamic audio coh011 based on the diarization results comprises.
Shafey in view of Lin and DiMaria do not specifically disclose determining that at least one of the one or more predicted terms of the corresponding speech recognition result was spoken by a new speaker that did not speak during any previous segmented labeled training sample, based on determining that the at least one of the one or more predicted terms of the corresponding speech recognition result was spoken by the new speaker, sampling the corresponding sequence of acoustic frames characterizing the at least one of the one or more predicted terms, and storing the sampled corresponding sequence of acoustic frames as a respective audio speech snippet for the new speaker at one of the predetermined number of slots for the new speaker.
However, Biadsy, in the same field of endeavor, discloses:
determining that at least one of the one or more predicted terms of the corresponding speech recognition result was spoken by a new speaker that did not speak during any previous segmented labeled training sample (Biadsy, para 0010, predicted speech recognition result and determining a supervised loss term based on the predicted speech recognition result);
based on determining that the at least one of the one or more predicted terms of the corresponding speech recognition result was spoken by the new speaker (Biadsy, para 0041, the ASR model 200 to make predictions that suit the source speaker 104), sampling the corresponding sequence of acoustic frames characterizing the at least one of the one or more predicted terms (Biadsy, para 0060, The ground truth transcription 563 should accurately reflect the corresponding speech sample (i.e., audio data 561) such that the ground truth transcription 563 is a target output of the sub-model 215); and
storing the sampled corresponding sequence of acoustic frames as a respective audio speech snippet for the new speaker at one of the predetermined number of slots for the new speaker (Biadsy, para 0057, a plurality of speech samples spoken by a variety of different speakers... Further, the training input 510 may be paired with a label 520 indicating a target output associated with the training input 510. In other words, the training input 510 may include a plurality of speech samples corresponding to utterance spoken by different speakers and each speech sample may be paired with a corresponding label 520 indicating a transcription of the corresponding utterance).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Biadsy in the method of Shafey in view of Lin and DiMaria because this would enable training a sub-model for contextual biasing for speech recognition (Biadsy, Abstract).
Regarding Claim 16, Shafey in view of Lin and DiMaria disclose the system of claim 14.
Shafey in view of Lin and DiMaria do not specifically disclose determining that at least one of the one or more predicted terms of the corresponding speech recognition result was spoken by a respective speaker that did speak during a previous segmented labeled training sample, determining that a current number of audio speech snippets stored for the respective speaker fails to satisfy a threshold of audio speech snippets, based on determining that the current number of snippets stored for the respective speaker fails to satisfy the threshold of audio speech snippets, sampling the corresponding sequence of acoustic frames characterizing the at least one of the one or more predicted terms, and storing the sampled corresponding sequence of acoustic frames as a respective audio snippet for the respective speaker at one of the predetem1ined number of slots.
However, Biadsy, in the same field of endeavor, discloses:
determining that at least one of the one or more predicted terms of the corresponding speech recognition result was spoken by a respective speaker that did speak during a previous segmented labeled training sample (Biadsy, para 0010, predicted speech recognition result and determining a supervised loss term based on the predicted speech recognition result);
determining that a current number of audio speech snippets stored for the respective speaker fails to satisfy a threshold of audio speech snippets (Biadsy, para 0029, ASR models may be trained on large sets of training data including audio samples of speech to produce a robust model for speech recognition… However, there are drawbacks to using such large models, such as applying a single model for a wide variety of users with different characteristics);
based on determining that the current number of snippets stored for the respective speaker fails to satisfy the threshold of audio speech snippets, sampling the corresponding sequence of acoustic frames characterizing the at least one of the one or more predicted terms (Biadsy, para 0063-0067, The base ASR model 200 may receive the sub-model output 569 along with the audio data 561 characterizing the training utterance 560 to produce a predicted speech recognition result 565 (i.e., biased speech recognition results 224). In some implementations, the predicted speech recognition result 565 is used by a loss function 580 to generate a supervised loss term 590. That is, the loss function 580 compares the predicted speech recognition result 565 and the ground truth transcription 563 of the corresponding training utterance 560 to generate the supervised loss term 590, where the loss 590 indicates a discrepancy between the ground truth transcription 563 (i.e., the target output) and the predicted speech recognition result 565… Training utterances 560 may be obtained in a variety of different ways. Typically, training utterances 560 are collected manually, where an audio sample of an utterance is manually transcribed. However, manually labeling training data can be tedious and difficult to collect sufficient samples of labeled data for training); and
storing the sampled corresponding sequence of acoustic frames as a respective audio snippet for the respective speaker at one of the predetermined number of slots (Biadsy, para 0035, an acoustic front-end residing on the user device 110 may convert a time-domain audio waveform of the utterance 108 captured via a microphone of the user device 110 into the input spectrograms 102 or other type or form of audio data 102. Further, the front-end device may be configured to determine or obtain data representing a contextual indicator 103 affecting the utterance 108 and/or other pertinent information corresponding to the source speaker 104).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Biadsy in the method of Shafey in view of Lin and DiMaria because this would enable training a sub-model for contextual biasing for speech recognition (Biadsy, Abstract).
Regarding Claim 17, Shafey in view of Lin disclose the system of claim 14, wherein generating the updated dynamic audio cohort based on the diarization results comprises:
DiMaria further teaches:
determining that a current number of snippets stored for respective speaker satisfies a threshold of audio speech snippets (DiMaria, para 0046, If the correspondence score 174 is more than a threshold score, e.g., 84%, the audio processing engine 120 determines that the feature vector 156 corresponds to the target vector 152. In other words, the audio processing engine 120 determines that one or more utterances 158 spoken in the audio snippet 138 are uttered by the first speaker 102a); and
based on determining that the current number of snippets stored for respective speaker satisfies the threshold of audio speech snippets, sampling a random number from a random number distribution (DiMaria, para 0053, At step 304, the audio processing engine 120 splits the audio file 136 into a plurality of audio snippets 138 based on a probability of each audio snippet 138 comprising one or more utterances 158 being above a threshold percentage. For example, the audio processing engine 120 may determine a length 204 for each audio snippet 138, such that a probability of detecting at least one utterance 158 in above a threshold number of slices 162 in a row being above a threshold percentage, e.g., 80%, similar to that described in FIGS. 1 and 2. The lengths 204 of the audio snippets 138 may vary based on speech distribution patterns 148 associated with speakers 102 in the audio file 136).
Shafey in view of Lin and DiMaria do not specifically disclose determining that at least one of the one or more predicted terms of the corresponding speech recognition result was spoken by a respective speaker that did speak during a previous segmented labeled training sample.
However, Biadsy, in the same field of endeavor, discloses determining that at least one of the one or more predicted terms of the corresponding speech recognition result was spoken by a respective speaker that did speak during a previous segmented labeled training sample (Biadsy, para 0010, predicted speech recognition result and determining a supervised loss term based on the predicted speech recognition result).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Biadsy in the method of Shafey in view of Lin and DiMaria because this would enable training a sub-model for contextual biasing for speech recognition (Biadsy, Abstract).
Claims 8 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Shafey in view of Lin, further in view of DiMaria, further in view of Biadsy, and further in view of Jin et al. Pat No. US 10347238 B2 (Jin).
Regarding Claim 8, Shafey in view of Lin, DiMaria and Biadsy disclose the computer-implemented method of claim 7, wherein generating the updated dynamic audio cohort based on the diarization results comprises:
DiMaria further teaches:
determining that the sampled random number satisfies a random number threshold (DiMaria, para 0046, If the correspondence score 174 is more than a threshold score, e.g., 84%, the audio processing engine 120 determines that the feature vector 156 corresponds to the target vector 152. In other words, the audio processing engine 120 determines that one or more utterances 158 spoken in the audio snippet 138 are uttered by the first speaker 102a); and
based on determining that the sampled random number satisfies the random number threshold (DiMaria, para 0046, If the correspondence score 174 is more than a threshold score, e.g., 84%, the audio processing engine 120 determines that the feature vector 156 corresponds to the target vector 152. In other words, the audio processing engine 120 determines that one or more utterances 158 spoken in the audio snippet 138 are uttered by the first speaker 102a):
identifying a respective one of the predetermined number of slots associated with the respective speaker based on the sampled random number (DiMaria, para 0052-0053, identify a particular speaker 102, such as the first speaker 102a in one or more audio files 136. For example, the audio processing engine 120 may receive the request from a user to determine whether one or more audio files 136 contain speech of the first speaker 102a... the audio processing engine 120 may determine a length 204 for each audio snippet 138, such that a probability of detecting at least one utterance 158 in above a threshold number of slices 162 in a row being above a threshold percentage, e.g., 80%, similar to that described in FIGS. 1 and 2);
Biadsy further teaches:
sampling the corresponding sequence of acoustic frames characterizing the at least one of the one or more predicted terms (Biadsy, para 0010, predicted speech recognition result and determining a supervised loss term based on the predicted speech recognition result).
Shafey in view of Lin, DiMaria and Biadsy do not specifically disclose replacing a respective audio speech snippet stored at the identified respective one of the predetermined number of slots with the sampled corresponding sequence of acoustic frames as a new respective audio snippet.
However, Jin, in the same field of endeavor, discloses replacing a respective audio speech snippet stored at the identified respective one of the predetermined number of slots with the sampled corresponding sequence of acoustic frames as a new respective audio snippet (Jin, col 14, ln 25-35, audio snippets 146 may be contiguous portions of an audio waveform (i.e., a sequence of frames). Audio snippets 146 and context waveform 520 are provided to concatenative synthesis module 524, which generates edited waveform 148. According to one embodiment of the present disclosure, a snippet is the corresponding audio frames 162 for an exemplar 324 or set of exemplars 324 in the temporal domain. According to one embodiment of the present disclosure, context waveform 520 may comprise surrounding audio corresponding to the query waveform 144 to be inserted or replaced).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Jin in the method of Shafey in view of Lin, DiMaria and Biadsy because this would enable interactive text-based insertion and replacement in an audio stream or file (Jin, Abstract).
Regarding Claim 18, Shafey in view of Lin, DiMaria and Biadsy disclose the system of claim 17, wherein generating the updated dynamic audio cohort based on the diarization results comprises:
DiMaria further teaches:
determining that the sampled random number satisfies a random number threshold (DiMaria, para 0046, If the correspondence score 174 is more than a threshold score, e.g., 84%, the audio processing engine 120 determines that the feature vector 156 corresponds to the target vector 152. In other words, the audio processing engine 120 determines that one or more utterances 158 spoken in the audio snippet 138 are uttered by the first speaker 102a); and
based on determining that the sampled random number satisfies the random number threshold (DiMaria, para 0046, If the correspondence score 174 is more than a threshold score, e.g., 84%, the audio processing engine 120 determines that the feature vector 156 corresponds to the target vector 152. In other words, the audio processing engine 120 determines that one or more utterances 158 spoken in the audio snippet 138 are uttered by the first speaker 102a):
identifying a respective one of the predetermined number of slots associated with the respective speaker based on the sampled random number (DiMaria, para 0052-0053, identify a particular speaker 102, such as the first speaker 102a in one or more audio files 136. For example, the audio processing engine 120 may receive the request from a user to determine whether one or more audio files 136 contain speech of the first speaker 102a... the audio processing engine 120 may determine a length 204 for each audio snippet 138, such that a probability of detecting at least one utterance 158 in above a threshold number of slices 162 in a row being above a threshold percentage, e.g., 80%, similar to that described in FIGS. 1 and 2);
Biadsy further teaches:
sampling the corresponding sequence of acoustic frames characterizing the at least one of the one or more predicted terms (Biadsy, para 0010, predicted speech recognition result and determining a supervised loss term based on the predicted speech recognition result); and
Shafey in view of Lin, DiMaria and Biadsy do not specifically disclose replacing a respective audio speech snippet stored at the identified respective one of the predetermined number of slots with the sampled corresponding sequence of acoustic frames as a new respective audio snippet.
However, Jin, in the same field of endeavor, discloses replacing a respective audio speech snippet stored at the identified respective one of the predetermined number of slots with the sampled corresponding sequence of acoustic frames as a new respective audio snippet (Jin, col 14, ln 25-35, audio snippets 146 may be contiguous portions of an audio waveform (i.e., a sequence of frames). Audio snippets 146 and context waveform 520 are provided to concatenative synthesis module 524, which generates edited waveform 148. According to one embodiment of the present disclosure, a snippet is the corresponding audio frames 162 for an exemplar 324 or set of exemplars 324 in the temporal domain. According to one embodiment of the present disclosure, context waveform 520 may comprise surrounding audio corresponding to the query waveform 144 to be inserted or replaced).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Jin in the method of Shafey in view of Lin, DiMaria and Biadsy because this would enable interactive text-based insertion and replacement in an audio stream or file (Jin, Abstract).
Claims 9 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Shafey in view of Lin, further in view of DiMaria, further in view of Biadsy, and further in view of Fang et al. Pat App No. US 20220122612 A1 (Fang).
Regarding Claim 9, Shafey in view of Lin, DiMaria and Biadsy disclose the computer-implemented method of claim 7.
Shafey in view of Lin, DiMaria and Biadsy do not specifically disclose determining that the sampled random number fails to satisfy a random number threshold, and based on determining that the sampled random number fails to satisfy the random number threshold, determining not to replace any of the audio speech snippets currently stored in association with the respective speaker.
However, Fang, in the same field of endeavor, discloses:
determining that the sampled random number fails to satisfy a random number threshold (Fang, para 0040, the second aggregate acoustic embedding 234b fails to satisfy the distance threshold 236, the comparator 230 determines that the first audio sample 202a and the second audio sample 202b are not from the same speaker 10); and
based on determining that the sampled random number fails to satisfy the random number threshold, determining not to replace any of the audio speech snippets currently stored in association with the respective speaker (Fang, para 0040, the comparator 230 may be configured such that when the distance between the first aggregate acoustic embedding 234a and the second aggregate acoustic embedding 234b satisfies a distance threshold 236, the comparator 230 determines that the first audio sample 202a and the second audio sample 202b are from the same speaker 10. Otherwise, when the distance between the first aggregate acoustic embedding 234a and the second aggregate acoustic embedding 234b fails to satisfy the distance threshold 236, the comparator 230 determines that the first audio sample 202a and the second audio sample 202b are not from the same speaker 10. The distance threshold 236 refers to a value that is set to indicate a confidence level that the speaker 10 of the first audio sample 202a is likely the same speaker 10 as the second audio sample 202b).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Fang in the method of Shafey in view of Lin, DiMaria and Biadsy because this would enable verification and identification of a user with little or limited information (e.g., audio data) about the voice of the user (Fang, para 0002-0003).
Regarding Claim 19, Shafey in view of Lin, DiMaria and Biadsy disclose the system of claim 17.
Shafey in view of Lin, DiMaria and Biadsy do not specifically disclose determining that the sampled random number fails to satisfy a random number threshold, and based on determining that the sampled random number fails to satisfy the random number threshold, determining not to replace any of the audio speech snippets currently stored in association with the respective speaker.
However, Fang, in the same field of endeavor, discloses:
determining that the sampled random number fails to satisfy a random number threshold (Fang, para 0040, the second aggregate acoustic embedding 234b fails to satisfy the distance threshold 236, the comparator 230 determines that the first audio sample 202a and the second audio sample 202b are not from the same speaker 10); and
based on determining that the sampled random number fails to satisfy the random number threshold, determining not to replace any of the audio speech snippets currently stored in association with the respective speaker (Fang, para 0040, the comparator 230 may be configured such that when the distance between the first aggregate acoustic embedding 234a and the second aggregate acoustic embedding 234b satisfies a distance threshold 236, the comparator 230 determines that the first audio sample 202a and the second audio sample 202b are from the same speaker 10. Otherwise, when the distance between the first aggregate acoustic embedding 234a and the second aggregate acoustic embedding 234b fails to satisfy the distance threshold 236, the comparator 230 determines that the first audio sample 202a and the second audio sample 202b are not from the same speaker 10. The distance threshold 236 refers to a value that is set to indicate a confidence level that the speaker 10 of the first audio sample 202a is likely the same speaker 10 as the second audio sample 202b).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Fang in the method of Shafey in view of Lin, DiMaria and Biadsy because this would enable verification and identification of a user with little or limited information (e.g., audio data) about the voice of the user (Fang, para 0002-0003).
Claims 10 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Shafey in view of Lin, and further in view of Thomas et al. Pat App No. WO 2022084851 A1 (Thomas).
Regarding Claim 10, Shafey in view of Lin disclose the computer-implemented method of claim 1.
Shafey in view of Lin do not specifically disclose rein the operations further comprise augmenting each segmented labeled training sample.
However, Thomas, in the same field of endeavor, discloses wherein the operations further comprise augmenting each segmented labeled training sample (Thomas, para 0039-0040, For instance, physician model for physician 103B included in the collection of physician models 303 can include one or more trigger words that the physician 103B uses to indicate a dictation and that the system 100 has been trained on. As another example, a physician model in the collection of physician models 303 may include information regarding whether the physician recites punctuation when dictating. That is, the physician model for the physician 103B in the collection of physician models 303 may represent any amount of learned behaviors regarding the physician 103B and the relative impact of those learned behaviors on determinations that processed audio is or is not dictation. In some implementations, the system 100 can provide a selected physician model to an audio classifier and augmentor 301 that applies the selected physician model to the audio signal. In some implementations, the audio classifier and augmentor 301 is the first classification model 211 that uses the selected physician model as part of the first stage 201A analysis of the audio signal).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Thomas in the method of Shafey in view of Lin because this would improve the accuracy of extracted words and concepts when the system 100 uses the automatic speech recognition engine 213 to analyze the portions of speech determined to be dictation in the second stage 201B (Thomas, para 0052).
Regarding Claim 20, Shafey in view of Lin disclose the system of claim 11.
Shafey in view of Lin do not specifically disclose rein the operations further comprise augmenting each segmented labeled training sample.
However, Thomas, in the same field of endeavor, discloses wherein the operations further comprise augmenting each segmented labeled training sample (Thomas, para 0039-0040, For instance, physician model for physician 103B included in the collection of physician models 303 can include one or more trigger words that the physician 103B uses to indicate a dictation and that the system 100 has been trained on. As another example, a physician model in the collection of physician models 303 may include information regarding whether the physician recites punctuation when dictating. That is, the physician model for the physician 103B in the collection of physician models 303 may represent any amount of learned behaviors regarding the physician 103B and the relative impact of those learned behaviors on determinations that processed audio is or is not dictation. [0040] In some implementations, the system 100 can provide a selected physician model to an audio classifier and augmentor 301 that applies the selected physician model to the audio signal. In some implementations, the audio classifier and augmentor 301 is the first classification model 211 that uses the selected physician model as part of the first stage 201A analysis of the audio signal).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Thomas in the method of Shafey in view of Lin because this would improve the accuracy of extracted words and concepts when the system 100 uses the automatic speech recognition engine 213 to analyze the portions of speech determined to be dictation in the second stage 201B (Thomas, para 0052).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MULUGETA T. DUGDA whose telephone number is (703)756-1106. The examiner can normally be reached Mon - Fri, 4:30am - 7:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Paras D. Shah can be reached at 571-270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MULUGETA TUJI DUGDA/Examiner, Art Unit 2653
/DOUGLAS GODBOLD/Primary Examiner, Art Unit 2655