Prosecution Insights
Last updated: October 02, 2026
Application No. 18/930,931

SPEECH RECOGNITION WITH ACCURATE TIME ALIGNMENT OF SPEECH UNITS

Non-Final OA §102§103
Filed
Oct 29, 2024
Examiner
SERRAGUARD, SEAN ERIN
Art Unit
2657
Tech Center
2600 — Communications
Assignee
NVIDIA Corporation
OA Round
1 (Non-Final)
69%
Grant Probability
Favorable
1-2
OA Rounds
1y 1m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 69% — above average
69%
Career Allowance Rate
112 granted / 162 resolved
+7.1% vs TC avg
Strong +34% interview lift
Without
With
+34.1%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
23 currently pending
Career history
188
Total Applications
across all art units

Statute-Specific Performance

§101
8.1%
-31.9% vs TC avg
§103
50.1%
+10.1% vs TC avg
§102
19.7%
-20.3% vs TC avg
§112
20.0%
-20.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 162 resolved cases

Office Action

§102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 1-8, 15-17, and 20 is/are rejected under 35 U.S.C. 102(a)(1) and 102(a)(2) as being anticipated by Radford (U.S. Pat. App. Pub. No. 2024/0354521, hereinafter Radford) with further evidence from Non-Patent Literature to Radford et al. (Radford, A., Kim, J.W., Xu, T., Brockman, G., McLeavey, C. and Sutskever, I., 2023, July. Robust speech recognition via large-scale weak supervision. In International conference on machine learning (pp. 28492-28518). PMLR., hereinafter Radford2). Regarding claim 1, Radford discloses A method (Systems and methods described with reference to the “automatic speech recognition model”; Radford, ¶ [0050]) comprising: processing, using an automatic speech recognition (ASR) model, one or more audio frames representative of a speech (“In step 320 of process 300, the machine learning platform can obtain an audio segment,” where the audio segment converted to audio frames (e.g., “windows”) and is processed.; Radford, ¶ [0057], [0086]) to generate, for a transcription unit (TU) of the speech: a first set of likelihood values, an individual likelihood value of the first set of likelihood values characterizing a probability that the TU corresponds to a respective vocabulary token of a plurality of vocabulary tokens (As described with reference to process 300, “the machine learning platform can obtain a task configuration” which “can specify the tasks that the model is supposed to perform on the audio segment” such as “transcription” including “timestamps” where the decoder iteratively predicts “textual tokens” or “timestamp tokens” and “the output of the decoder” which corresponds to the textual tokens and the timestamp tokens, in the context of a transcription task configuration, can be mapped “to a vector of activation values (e.g., logits) for the vocabulary of model 110 (or a subset thereof, such as the text... tokens...” and “another layer (e.g., a softmax layer) can convert these activation values into probabilities”, where the probabilities for the text {vocabulary} tokens refers to the full set of probabilities for respective text tokens of a plurality of text tokens and corresponds to the first set of likelihood values; Radford, ¶ [0066], [0087]-[0088]), and a second set of likelihood values, an individual likelihood value of the second set of likelihood values characterizing a probability that the TU corresponds to a respective timestamp token of a plurality of timestamp tokens (Regarding the “timestamp tokens,” and with reference to the same process 300 described above, “the output of the decoder” which corresponds to the timestamp tokens, can be mapped “to a vector of activation values (e.g., logits) for the vocabulary of model 110 (or a subset thereof, such as the... timestamp tokens...” and “another layer (e.g., a softmax layer) can convert these activation values into probabilities”, where the probabilities for the timestamp tokens refers to the full set of probabilities for respective timestamp tokens of a plurality of timestamp tokens and corresponds to the second set of likelihood values; Radford, ¶ [0066], [0087]-[0088]); and generating, using the first set of likelihood values and the second set of likelihood values, a timed transcription of the speech (“In step 340 of process 300, the machine learning platform can generate a transcript (e.g., an output transcript) from the audio segment (e.g., an input audio segment) using the model.”; Radford, ¶ [0088]). Regarding claim 2, Radford discloses wherein the generating the timed transcription of the speech comprises: selecting, responsive to a maximum likelihood value of the first set of likelihood values and the second set of likelihood values corresponding to a vocabulary token of the plurality of vocabulary tokens (“decoder 112 can be configured to accept as input” the encoded audio input and “can be configured to predict the next token in the sequence... [and] can autoregressively predict the output sequence of tokens” where the “vector of activation values (e.g., logits) for the vocabulary of model 110” for “the text and timestamp tokens” as converted “into probabilities” can be applied by the computing device “to select the next token (e.g., next-token prediction 127) according to these probabilities.”; Radford, ¶ [0063], [0066]) [selecting], the vocabulary token as the TU (“the decoder can autoregressively generate output tokens based on (e.g., using) previously generated output tokens in the decoder input and cross-attention from the encoder” where “the model may be configured to generate textual tokens for the language”; Radford, ¶ [0094]). Regarding claim 3, Radford discloses wherein the generating the timed transcription of the speech comprises: selecting, responsive to a maximum likelihood value of the first set of likelihood values and the second set of likelihood values corresponding to a timestamp token of the plurality of timestamp tokens (“decoder 112 can be configured to accept as input” the encoded audio input and “can be configured to predict the next token in the sequence... [and] can autoregressively predict the output sequence of tokens” where the “vector of activation values (e.g., logits) for the vocabulary of model 110” for “the text and timestamp tokens” as converted “into probabilities” can be applied by the computing device “to select the next token (e.g., next-token prediction 127) according to these probabilities.”; Radford, ¶ [0063], [0066]) [selecting], the timestamp token as the TU (“the decoder can autoregressively generate output tokens based on (e.g., using) previously generated output tokens in the decoder input” where “the model may be configured to generate timestamp tokens” such as “when the decoder input is not configured with a notimestamp token (e.g., the decoder input does not include the notimestamp token)” and/or “when the decoder input is configured with a timestamp token (e.g., the decoder input includes a timestamp token)”; Radford, ¶ [0094]). Regarding claim 4, Radford discloses wherein the selected timestamp token corresponds to at least one of: a first time associated with a start of a unit of the speech, or a second time associated with an end of the unit of the speech (“timestamp tokens can predict the time associated with one or more textual tokens... before and/or after each sequence of textual tokens” and “Timestamp tokens can indicate the time relative to the start of the audio portion of the training sample.” Further, the disclosure of a time frame which includes speech, necessarily includes both a start of a unit of speech and an end of a unit of speech.; Radford, ¶ [0045], [0054]). Regarding claim 5, Radford discloses wherein the unit of the speech corresponds to a single word or a portion thereof (“The textual tokens can be generated using a BPE text tokenizer (e.g., the byte-level BPE text tokenizer used for GPT-2, or the like)” which is byte pair encoding. “The BPE text tokenizer can generate the textual tokens comprising the vocabulary based on the tokens present in the multi-language transcripts in the training dataset,” where tokens corresponds to words. Further, though common words are assigned a single dedicated token in the vocabulary {a single word}, when the BPE text tokenizer encounters complex, rare, compound, or unknown words, it breaks them down into smaller chunks {a portion thereof}; Radford, ¶ [0056]). Regarding claim 6, Radford discloses wherein the timed transcription of the speech is generated using a mapping of the plurality of timestamp tokens to the one or more audio frames (“The timestamp tokens can indicate times with a resolution chosen based on the time resolution of the whisper models (e.g., the inter-sample times, or the like)” such as “20 milliseconds (e.g., 0 ms, 20 ms, 40 ms,.., 29980 ms, or the like) or other suitable values (e.g., 10 ms, 30 ms, 50 ms, 100 ms, 200 ms, 300 ms, 500 ms, 1 s, or the like).”; Radford, ¶ [0054]). Regarding claim 7, Radford discloses wherein the plurality of timestamp tokens comprises M timestamp tokens, and wherein the mapping of the plurality of timestamp tokens to the one or more audio frames comprises a modulo M mapping. (The described system discloses a fixed vocabulary of M timestamp tokens, relies on a sliding window to process continuous audio, and “Consistent with disclosed embodiments, timestamp tokens can predict the time associated with one or more textual tokens” where the “timestamp tokens can indicate times with a resolution chosen based on the time resolution of the whisper models”. Further, examiner notes that “Robust Speech Recognition via Large-Scale Weak Supervision” by Radford et al,” is incorporated by reference. The Radford paper, at pg. 14, “Comparison with other ASR models”, further explains that “Whisper models are trained on 30-second audio chunks and cannot consume longer audio inputs at once. This is not a problem with most academic datasets comprised of short utterances but presents challenges in real-world applications which often require transcribing minutes- or hours-long audio. We developed a strategy to perform buffered transcription of long audio by consecutively transcribing 30-second segments of audio and shifting the window according to the timestamps predicted by the model,” where the 30 second window shifting limit is a modulo M mapping.; Radford, ¶ [0028], [0054]). Regarding claim 8, Radford discloses further comprises: selecting the ASR model responsive to identification of a language of the speech (Discloses “different tasks that can be performed on the same input audio signal” includes “language identification” and the system “may include model selector engine 1532 (e.g., configured to select a model from among a plurality of models, such as based on input data)”; Radford, ¶ [0051], [0162]). Regarding claim 15, Radford discloses A system comprising: one or more processors (Systems and methods described with reference to the “automatic speech recognition model” as implemented in “exemplary operating environment for implementing various aspects of this disclosure” which “may include a computing device 1402 (e.g., a general-purpose computing device)” having “various hardware components, such as one or more processors 1406”; Radford, ¶ [0050], [0145]) to: process, using an automatic speech recognition (ASR) model, one or more audio frames representative of a speech to generate, for a transcription unit (TU) of the speech (“In step 320 of process 300, the machine learning platform can obtain an audio segment,” where the audio segment converted to audio frames (e.g., “windows”) and is processed.; Radford, ¶ [0057], [0086]): a first set of likelihood values, an individual likelihood value of the first set of likelihood values characterizing a probability that the TU corresponds to a respective vocabulary token of a plurality of vocabulary tokens (As described with reference to process 300, “the machine learning platform can obtain a task configuration” which “can specify the tasks that the model is supposed to perform on the audio segment” such as “transcription” including “timestamps” where the decoder iteratively predicts “textual tokens” or “timestamp tokens” and “the output of the decoder” which corresponds to the textual tokens and the timestamp tokens, in the context of a transcription task configuration, can be mapped “to a vector of activation values (e.g., logits) for the vocabulary of model 110 (or a subset thereof, such as the text... tokens...” and “another layer (e.g., a softmax layer) can convert these activation values into probabilities”, where the probabilities for the text {vocabulary} tokens refers to the full set of probabilities for respective text tokens of a plurality of text tokens and corresponds to the first set of likelihood values; Radford, ¶ [0066], [0087]-[0088]), and a second set of likelihood values, an individual likelihood value of the second set of likelihood values characterizing a probability that the TU corresponds to a respective timestamp token of a plurality of timestamp tokens (Regarding the “timestamp tokens,” and with reference to the same process 300 described above, “the output of the decoder” which corresponds to the timestamp tokens, can be mapped “to a vector of activation values (e.g., logits) for the vocabulary of model 110 (or a subset thereof, such as the... timestamp tokens...” and “another layer (e.g., a softmax layer) can convert these activation values into probabilities”, where the probabilities for the timestamp tokens refers to the full set of probabilities for respective timestamp tokens of a plurality of timestamp tokens and corresponds to the second set of likelihood values; Radford, ¶ [0066], [0087]-[0088]); and generate, using the first set of likelihood values and the second set of likelihood values, a timed transcription of the speech (“In step 340 of process 300, the machine learning platform can generate a transcript (e.g., an output transcript) from the audio segment (e.g., an input audio segment) using the model.”; Radford, ¶ [0088]). Regarding claim 16, Radford discloses wherein to generate the timed transcription of the speech, the one or more processors are to: select, responsive to a maximum likelihood value of the first set of likelihood values and the second set of likelihood values corresponding to a timestamp token of the plurality of timestamp tokens (“decoder 112 can be configured to accept as input” the encoded audio input and “can be configured to predict the next token in the sequence... [and] can autoregressively predict the output sequence of tokens” where the “vector of activation values (e.g., logits) for the vocabulary of model 110” for “the text and timestamp tokens” as converted “into probabilities” can be applied by the computing device “to select the next token (e.g., next-token prediction 127) according to these probabilities.”; Radford, ¶ [0063], [0066]), the timestamp token as the TU (“the decoder can autoregressively generate output tokens based on (e.g., using) previously generated output tokens in the decoder input” where “the model may be configured to generate timestamp tokens” such as “when the decoder input is not configured with a notimestamp token (e.g., the decoder input does not include the notimestamp token)” and/or “when the decoder input is configured with a timestamp token (e.g., the decoder input includes a timestamp token)”; Radford, ¶ [0094]). Regarding claim 17, Radford discloses wherein the selected timestamp token corresponds to at least one of: a first time associated with a start of a unit of the speech, or a second time associated with an end of the unit of the speech (“timestamp tokens can predict the time associated with one or more textual tokens... before and/or after each sequence of textual tokens” and “Timestamp tokens can indicate the time relative to the start of the audio portion of the training sample.” Further, the disclosure of a time frame which includes speech, necessarily includes both a start of a unit of speech and an end of a unit of speech.; Radford, ¶ [0045], [0054]). Regarding claim 20, Radford discloses wherein the system is comprised in at least one of: an in-vehicle infotainment system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation; a system for performing one or more medical operations; a system for performing one or more factory operations; a system for performing one or more analytics operations; a system implementing one or more inference microservices; a system for performing light transport simulations; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system implemented using an edge device; a system for generating or presenting at least one of virtual reality content, mixed reality content, or augmented reality content; a system implemented using a robot; a system for performing one or more conversational AI operations; a system implementing one or more large language models (LLMs); a system implementing one or more vision language models (VLMs); a system implementing one or more multi-modal language models; a system implementing one or more language models; a system for performing one or more generative AI operations; a system for generating synthetic data; a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources. (“an operating environment 1400 for an exemplary embodiment includes at least one computing device 1402. The computing device 1402 may be a uniprocessor or multiprocessor computing device. An operating environment 1400 may include one or more computing devices (e.g., multiple computing devices 1402) in a given computer system, which may be clustered, part of a local area network (LAN), part of a wide area network (WAN), client-server networked, peer-to-peer networked within a cloud, or otherwise communicably linked. A computer system may include an individual machine or a group of cooperating machines. A given computing device 1402 may be configured for end-users, e.g., with applications, for administrators, as a server, as a distributed processing node, as a special-purpose processing device, or otherwise configured to train machine learning models and/or use machine learning models.”; Radford, ¶ [0146]). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 9-10 and 18-19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Radford as applied to claim 1 and 16 above, and further in view of Ward (U.S. Pat. App. Pub. No. 2025/0285612, hereinafter Ward). Regarding claim 9, the rejection of claim 1 is incorporated. Radford discloses all of the elements of the current invention as stated above. Radford further discloses wherein the ASR model is trained using training data that comprises: a training speech, and a transcription for the training speech, comprising (“the training system can obtain audio paired with transcripts, consistent with disclosed embodiments. “; Radford, ¶ [0040]): a plurality of units of the training speech, and a plurality of training timestamps (“training samples can be generated from the obtained audio and transcripts. In some embodiments, training samples can be generated by selecting segments of the obtained audio. The selected segments can be associated with corresponding portions of the transcripts. The corresponding portion of a transcripts can be the subset of the transcript that occurs within the selected segment of the obtained audio,” and the training timestamps being derived from metadata.; Radford, ¶ [0036]-[0037], [0045], [0052]). However, Radford fails to expressly recite an individual training timestamp of the plurality of training timestamps associated with at least one of: a start of an individual unit of a plurality of units of the training speech, or an end of the individual unit. Ward teaches systems and methods for “generating training data for artificial intelligence models.” (Ward, ¶ [0001]). Regarding claim 9, Ward teaches an individual training timestamp of the plurality of training timestamps associated with at least one of: a start of an individual unit of a plurality of units of the training speech, or an end of the individual unit (“An alignment module 105 can provide the chunking module 106 with timing data 110. The timing data 110 can include data, such as the beginning time and end time of words in the audio file 102” as well as “timing data of other units of speech, such as phonemes... an alignment, a matching or a pairing of units of audio in the audio file 102 and corresponding units in the transcript 104.”; Ward, ¶ [0031]). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the multi-task automatic speech recognition systems of Radford to incorporate the teachings of Ward to include an individual training timestamp of the plurality of training timestamps associated with at least one of: a start of an individual unit of a plurality of units of the training speech, or an end of the individual unit. Radford teaches a multi-task ASR architecture that explicitly requires time-aligned training data. However, Radford relies on scraped data from online data sources to provide that time alignment, which is inherently both unreliable and noisy (which is also a common problem in the art, due to limited availability of curated datasets). Ward is directed to resolving this problem through the use of a multi-stage alignment pipeline (i.e., the teacher model comprising at least the audio-to-token predictor and the alignment engine) to automatically generate precise timing labels for unaligned text. A person having ordinary skill in the art would recognize that applying Ward’s multi-stage alignment pipeline to prepare the training corpora for Radford’s multi-task ASR model would yield the predictable result of producing highly accurate target transcription units, which would reduce alignment noise (as compared to web scraped data) and improve downstream accuracy of the supervised student model, as recognized by Ward. (Ward, ¶ [0003]-[0004], [0028]). Regarding claim 10, the rejection of claim 9 is incorporated. Radford discloses all of the elements of the current invention as stated above. However, Radford fails to expressly recite wherein the plurality of training timestamps is obtained using a teacher model that comprises: a first ASR model processing the training speech to determine the plurality of units of the training speech, and a second alignment model identifying, using the plurality of units of the training speech, the plurality of training timestamps. The relevance of Ward is described above with relation to claim 9. Regarding claim 10, Ward teaches wherein the plurality of training timestamps is obtained using a teacher model that comprises (Though not expressly described as a teacher model, discloses “a diagram of an environment 100, where chunking can be used to modify raw training data for an AI trainer module 116, in order to make the raw training data more suitable for hardware and/or to provide training data that yields more efficiencies. “; Ward, ¶ [0029]): a first ASR model processing the training speech to determine the plurality of units of the training speech (Discloses an “audio to token predictor 202” which “generates predicted tokens (PT) 206 from the audio file 102” and “can also generate predicted token timings 208 for the predicted tokens 206.”; Ward, ¶ [0034]), and a second alignment model identifying, using the plurality of units of the training speech, the plurality of training timestamps (“alignment engine 212 can map the predicted tokens 206 to the ground truth tokens, finding matches between the predicted tokens 206 and the ground truth tokens 210” and “the predicted token timings 208, associated with the predicted tokens 206, can be assigned to the corresponding matched ground truth tokens 210” and generating the “first stage timings 214 are timing information for each token” and “can be used to generate the timing data 110 for other units of speech, for example words.{training timestamps}”; Ward, ¶ [0037]-[0039]). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the multi-task automatic speech recognition systems of Radford to incorporate the teachings of Ward to include wherein the plurality of training timestamps is obtained using a teacher model that comprises: a first ASR model processing the training speech to determine the plurality of units of the training speech, and a second alignment model identifying, using the plurality of units of the training speech, the plurality of training timestamps. Radford teaches a multi-task ASR architecture that explicitly requires time-aligned training data. However, Radford relies on scraped data from online data sources to provide that time alignment, which is inherently both unreliable and noisy (which is also a common problem in the art, due to limited availability of curated datasets). Ward is directed to resolving this problem through the use of a multi-stage alignment pipeline (i.e., the teacher model comprising at least the audio-to-token predictor and the alignment engine) to automatically generate precise timing labels for unaligned text. A person having ordinary skill in the art would recognize that applying Ward’s multi-stage alignment pipeline to prepare the training corpora for Radford’s multi-task ASR model would yield the predictable result of producing highly accurate target transcription units, which would reduce alignment noise (as compared to web scraped data) and improve downstream accuracy of the supervised student model, as recognized by Ward. (Ward, ¶ [0003]-[0004], [0028]). Regarding claim 18, the rejection of claim 16 is incorporated. Radford discloses all of the elements of the current invention as stated above. Radford further discloses wherein the ASR model is trained using training data that comprises: a training speech, and a transcription for the training speech (“the training system can obtain audio paired with transcripts, consistent with disclosed embodiments. “; Radford, ¶ [0040]), comprising: a plurality of units of the training speech, and a plurality of training timestamps (“training samples can be generated from the obtained audio and transcripts. In some embodiments, training samples can be generated by selecting segments of the obtained audio. The selected segments can be associated with corresponding portions of the transcripts. The corresponding portion of a transcripts can be the subset of the transcript that occurs within the selected segment of the obtained audio,” and the training timestamps being derived from metadata.; Radford, ¶ [0036]-[0037], [0045], [0052]). However, Radford fails to expressly recite an individual training timestamp of the plurality of training timestamps associated with at least one of: a start of an individual unit of a plurality of units of the training speech, or an end of the individual unit. The relevance of Ward is described above with relation to claim 9. Regarding claim 18, Ward teaches an individual training timestamp of the plurality of training timestamps associated with at least one of: a start of an individual unit of a plurality of units of the training speech, or an end of the individual unit (“An alignment module 105 can provide the chunking module 106 with timing data 110. The timing data 110 can include data, such as the beginning time and end time of words in the audio file 102” as well as “timing data of other units of speech, such as phonemes... an alignment, a matching or a pairing of units of audio in the audio file 102 and corresponding units in the transcript 104.”; Ward, ¶ [0031]). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the multi-task automatic speech recognition systems of Radford to incorporate the teachings of Ward to include an individual training timestamp of the plurality of training timestamps associated with at least one of: a start of an individual unit of a plurality of units of the training speech, or an end of the individual unit. Radford teaches a multi-task ASR architecture that explicitly requires time-aligned training data. However, Radford relies on scraped data from online data sources to provide that time alignment, which is inherently both unreliable and noisy (which is also a common problem in the art, due to limited availability of curated datasets). Ward is directed to resolving this problem through the use of a multi-stage alignment pipeline (i.e., the teacher model comprising at least the audio-to-token predictor and the alignment engine) to automatically generate precise timing labels for unaligned text. A person having ordinary skill in the art would recognize that applying Ward’s multi-stage alignment pipeline to prepare the training corpora for Radford’s multi-task ASR model would yield the predictable result of producing highly accurate target transcription units, which would reduce alignment noise (as compared to web scraped data) and improve downstream accuracy of the supervised student model, as recognized by Ward. (Ward, ¶ [0003]-[0004], [0028]). Regarding claim 19, the rejection of claim 18 is incorporated. Radford discloses all of the elements of the current invention as stated above. However, Radford fails to expressly recite wherein the plurality of training timestamps is obtained using a teacher model that comprises: a first ASR model processing the training speech to determine the plurality of units of the training speech, and a second alignment model identifying, using the plurality of units of the training speech, the plurality of training timestamps. The relevance of Ward is described above with relation to claim 9. Regarding claim 19, Ward teaches wherein the plurality of training timestamps is obtained using a teacher model (Though not expressly described as a teacher model, discloses “a diagram of an environment 100, where chunking can be used to modify raw training data for an AI trainer module 116, in order to make the raw training data more suitable for hardware and/or to provide training data that yields more efficiencies. “; Ward, ¶ [0029]) that comprises: a first ASR model processing the training speech to determine the plurality of units of the training speech (Discloses an “audio to token predictor 202” which “generates predicted tokens (PT) 206 from the audio file 102” and “can also generate predicted token timings 208 for the predicted tokens 206.”; Ward, ¶ [0034]), and a second alignment model identifying, using the plurality of units of the training speech, the plurality of training timestamps (“alignment engine 212 can map the predicted tokens 206 to the ground truth tokens, finding matches between the predicted tokens 206 and the ground truth tokens 210” and “the predicted token timings 208, associated with the predicted tokens 206, can be assigned to the corresponding matched ground truth tokens 210” and generating the “first stage timings 214 are timing information for each token” and “can be used to generate the timing data 110 for other units of speech, for example words.{training timestamps}”; Ward, ¶ [0037]-[0039]). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the multi-task automatic speech recognition systems of Radford to incorporate the teachings of Ward to include wherein the plurality of training timestamps is obtained using a teacher model that comprises: a first ASR model processing the training speech to determine the plurality of units of the training speech, and a second alignment model identifying, using the plurality of units of the training speech, the plurality of training timestamps. Radford teaches a multi-task ASR architecture that explicitly requires time-aligned training data. However, Radford relies on scraped data from online data sources to provide that time alignment, which is inherently both unreliable and noisy (which is also a common problem in the art, due to limited availability of curated datasets). Ward is directed to resolving this problem through the use of a multi-stage alignment pipeline (i.e., the teacher model comprising at least the audio-to-token predictor and the alignment engine) to automatically generate precise timing labels for unaligned text. A person having ordinary skill in the art would recognize that applying Ward’s multi-stage alignment pipeline to prepare the training corpora for Radford’s multi-task ASR model would yield the predictable result of producing highly accurate target transcription units, which would reduce alignment noise (as compared to web scraped data) and improve downstream accuracy of the supervised student model, as recognized by Ward. (Ward, ¶ [0003]-[0004], [0028]). Claim 11-14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Radford in view of Ward. Regarding claim 11, Radford discloses A method (Systems and methods described with reference to the “automatic speech recognition model”; Radford, ¶ [0050]) comprising:… training, using the plurality of target TUs, a student model to generate, for a TU of the speech (“generating a dataset for training an automatic speech recognition model (e.g., the “model”), consistent with disclosed embodiments.”; Radford, ¶ [0038]): a first set of likelihood values, an individual likelihood value of the first set of likelihood values characterizing a probability that the TU corresponds to a respective vocabulary token of a plurality of vocabulary tokens (As described with reference to process 300, “the machine learning platform can obtain a task configuration” which “can specify the tasks that the model is supposed to perform on the audio segment” such as “transcription” including “timestamps” where the decoder iteratively predicts “textual tokens” or “timestamp tokens” and “the output of the decoder” which corresponds to the textual tokens and the timestamp tokens, in the context of a transcription task configuration, can be mapped “to a vector of activation values (e.g., logits) for the vocabulary of model 110 (or a subset thereof, such as the text... tokens...” and “another layer (e.g., a softmax layer) can convert these activation values into probabilities”, where the probabilities for the text {vocabulary} tokens refers to the full set of probabilities for respective text tokens of a plurality of text tokens and corresponds to the first set of likelihood values; Radford, ¶ [0066], [0087]-[0088]), and a second set of likelihood values, an individual likelihood value of the second set of likelihood values characterizing a probability that the TU is associated with a respective timestamp of a plurality of timestamps (Regarding the “timestamp tokens,” and with reference to the same process 300 described above, “the output of the decoder” which corresponds to the timestamp tokens, can be mapped “to a vector of activation values (e.g., logits) for the vocabulary of model 110 (or a subset thereof, such as the... timestamp tokens...” and “another layer (e.g., a softmax layer) can convert these activation values into probabilities”, where the probabilities for the timestamp tokens refers to the full set of probabilities for respective timestamp tokens of a plurality of timestamp tokens and corresponds to the second set of likelihood values; Radford, ¶ [0066], [0087]-[0088]). However, Radford fails to expressly recite comprising: processing, using a teacher model, one or more audio frames representative of a speech to generate a plurality of target transcription units (TUs) of the speech, wherein the plurality of target TUs comprise: a first plurality of tokens of the speech, and a second plurality of timestamps associated with the first plurality of tokens. The relevance of Ward is described above with relation to claim 9. Regarding claim 11, Ward teaches comprising: processing, using a teacher model, one or more audio frames representative of a speech to generate a plurality of target transcription units (TUs) of the speech, (Though not expressly disclosed as a teacher model, the disclosed system including “A chunking module 106” which “can use the alignment data to break up an audio file 102 into pairs of audio file chunks (AFCs) 108 and transcript portions (L #) 114” and “An alignment module 105 can provide the chunking module 106 with timing data 110” which “can include data, such as the beginning time and end time of words in the audio file 102” as well as “timing data of other units of speech, such as phonemes... an alignment, a matching or a pairing of units of audio in the audio file 102 and corresponding units in the transcript 104.”; Ward, ¶ [0031]) wherein the plurality of target TUs comprise: a first plurality of tokens of the speech, and a second plurality of timestamps associated with the first plurality of tokens (The aligned “pairs of AFC 108 and transcript portions 114 can be provided to an AI trainer module 116” for training the ASR, and “timing data 110”; Ward, ¶ [0033], [0039]). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the multi-task automatic speech recognition systems of Radford to incorporate the teachings of Ward to include comprising: processing, using a teacher model, one or more audio frames representative of a speech to generate a plurality of target transcription units (TUs) of the speech, wherein the plurality of target TUs comprise: a first plurality of tokens of the speech, and a second plurality of timestamps associated with the first plurality of tokens. Radford teaches a multi-task ASR architecture that explicitly requires time-aligned training data. However, Radford relies on scraped data from online data sources to provide that time alignment, which is inherently both unreliable and noisy (which is also a common problem in the art, due to limited availability of curated datasets). Ward is directed to resolving this problem through the use of a multi-stage alignment pipeline (i.e., the teacher model comprising at least the audio-to-token predictor and the alignment engine) to automatically generate precise timing labels for unaligned text. A person having ordinary skill in the art would recognize that applying Ward’s multi-stage alignment pipeline to prepare the training corpora for Radford’s multi-task ASR model would yield the predictable result of producing highly accurate target transcription units, which would reduce alignment noise (as compared to web scraped data) and improve downstream accuracy of the supervised student model, as recognized by Ward. (Ward, ¶ [0003]-[0004], [0028]). Regarding claim 12, the rejection of claim 11 is incorporated. Radford and Ward disclose all of the elements of the current invention as stated above. However, Radford fails to expressly recite wherein the teacher model comprises an automatic speech recognition (ASR) model and an alignment model, and wherein the processing the one or more audio frames comprises: processing, using the ASR model, the one or more audio frames to identify the first plurality of tokens of the speech; and processing, using the alignment model, the one or more audio frames and the first plurality of tokens of the speech to determine the second plurality of timestamps associated with the first plurality of tokens. The relevance of Ward is described above with relation to claim 9. Regarding claim 12, Ward teaches wherein the teacher model comprises an automatic speech recognition (ASR) model and an alignment model (Discloses an “audio to token predictor 202” and an “alignment engine 212”; Ward, ¶ [0034]), and wherein the processing the one or more audio frames comprises: processing, using the ASR model, the one or more audio frames to identify the first plurality of tokens of the speech (Discloses an “audio to token predictor 202 {ASR}” which “generates predicted tokens (PT) 206 from the audio file 102” and “can also generate predicted token timings 208 for the predicted tokens 206.”; Ward, ¶ [0034]); and processing, using the alignment model, the one or more audio frames and the first plurality of tokens of the speech to determine the second plurality of timestamps associated with the first plurality of tokens (“alignment engine 212 can map the predicted tokens 206 to the ground truth tokens, finding matches between the predicted tokens 206 and the ground truth tokens 210” and “the predicted token timings 208, associated with the predicted tokens 206, can be assigned to the corresponding matched ground truth tokens 210” and generating the “first stage timings 214 are timing information for each token” and “can be used to generate the timing data 110 for other units of speech, for example words.{training timestamps}”; Ward, ¶ [0037]-[0039]). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the multi-task automatic speech recognition systems of Radford to incorporate the teachings of Ward to include wherein the teacher model comprises an automatic speech recognition (ASR) model and an alignment model, and wherein the processing the one or more audio frames comprises: processing, using the ASR model, the one or more audio frames to identify the first plurality of tokens of the speech; and processing, using the alignment model, the one or more audio frames and the first plurality of tokens of the speech to determine the second plurality of timestamps associated with the first plurality of tokens. Radford teaches a multi-task ASR architecture that explicitly requires time-aligned training data. However, Radford relies on scraped data from online data sources to provide that time alignment, which is inherently both unreliable and noisy (which is also a common problem in the art, due to limited availability of curated datasets). Ward is directed to resolving this problem through the use of a multi-stage alignment pipeline (i.e., the teacher model comprising at least the audio-to-token predictor and the alignment engine) to automatically generate precise timing labels for unaligned text. A person having ordinary skill in the art would recognize that applying Ward’s multi-stage alignment pipeline to prepare the training corpora for Radford’s multi-task ASR model would yield the predictable result of producing highly accurate target transcription units, which would reduce alignment noise (as compared to web scraped data) and improve downstream accuracy of the supervised student model, as recognized by Ward. (Ward, ¶ [0003]-[0004], [0028]). Regarding claim 13, the rejection of claim 11 is incorporated. Radford and Ward disclose all of the elements of the current invention as stated above. However, Radford fails to expressly recite wherein the training the student model comprises: computing, using at least the second set of likelihood values and a corresponding, to the TU, target TU of the plurality of target TUs, a loss value; and modifying, using the loss value, one or more parameters of the student model. The relevance of Ward is described above with relation to claim 9. Regarding claim 12, Ward teaches wherein the training the student model comprises: computing, using at least the second set of likelihood values and a corresponding, to the TU, target TU of the plurality of target TUs, a loss value (Though not explicitly describing the process of calculating loss and/or backpropagation, Ward does teach, regarding the background art, that “Training an AI model includes generating an output from the model, comparing it against a known output, and modifying the model parameters until the model generates an output close to the known output,” which is a structural explanation of supervised training. Generating an output and comparing against a known output, in the context of training an AI model is the generation of a loss value. This loss value is generated based on “pairs of audio files and corresponding transcripts” as well as timing and alignment data.; Ward, ¶ Abstract); and modifying, using the loss value, one or more parameters of the student model (Ward further teaches “modifying the model parameters until the model generates an output close to the known output”.; Ward, ¶ Abstract). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the multi-task automatic speech recognition systems of Radford to incorporate the teachings of Ward to include wherein the training the student model comprises: computing, using at least the second set of likelihood values and a corresponding, to the TU, target TU of the plurality of target TUs, a loss value; and modifying, using the loss value, one or more parameters of the student model. Radford teaches a multi-task ASR architecture that explicitly requires time-aligned training data. However, Radford relies on scraped data from online data sources to provide that time alignment, which is inherently both unreliable and noisy (which is also a common problem in the art, due to limited availability of curated datasets). Ward is directed to resolving this problem through the use of a multi-stage alignment pipeline (i.e., the teacher model comprising at least the audio-to-token predictor and the alignment engine) to automatically generate precise timing labels for unaligned text. A person having ordinary skill in the art would recognize that applying Ward’s multi-stage alignment pipeline to prepare the training corpora for Radford’s multi-task ASR model would yield the predictable result of producing highly accurate target transcription units, which would reduce alignment noise (as compared to web scraped data) and improve downstream accuracy of the supervised student model, as recognized by Ward. (Ward, ¶ [0003]-[0004], [0028]). Regarding claim 14, the rejection of claim 11 is incorporated. Radford and Ward disclose all of the elements of the current invention as stated above. Radford further discloses further comprising: causing the trained student model to be deployed for processing of an inference speech, (The “machine learning platform can obtain a task configuration” that “specify the tasks that the model is supposed to perform on the audio segment” as including “transcription” with “timestamps”; Radford, ¶ [0087]) wherein the processing of the inference speech comprises: determining a time-aligned transcription for the inference speech (“In step 340 of process 300, the machine learning platform can generate a transcript (e.g., an output transcript) from the audio segment (e.g., an input audio segment) using the model.” which can include predicting textual tokens and timestamp tokens.; Radford, ¶ [0088]). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Doutre (U.S. Pat. App. Pub. No. 2022/0343894) discloses systems and methods for improving streaming automatic speech recognition (ASR) with non-streaming model distillation Any inquiry concerning this communication or earlier communications from the examiner should be directed to Sean E. Serraguard whose telephone number is (313)446-6627. The examiner can normally be reached 07:00-17:00 M-F. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel C. Washburn can be reached at (571) 272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Sean E Serraguard/Primary Examiner, Art Unit 2657
Read full office action

Prosecution Timeline

Oct 29, 2024
Application Filed
Jul 01, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12750409
SIMULATED CHORAL AUDIO CHATTER
3y 11m to grant Granted Sep 29, 2026
Patent 12749484
DYNAMICALLY ADAPTING ASSISTANT RESPONSES
2y 8m to grant Granted Sep 29, 2026
Patent 12743582
DYNAMIC VOCABULARIES FOR CONDITIONING A LANGUAGE MODEL FOR TRANSFORMING NATURAL LANGUAGE TO A LOGICAL FORM
2y 8m to grant Granted Sep 22, 2026
Patent 12731158
COLLABORATIVE USER SUPPORT PORTAL
4y 6m to grant Granted Sep 08, 2026
Patent 12706081
SIMULATING CROWD NOISE FOR LIVE EVENTS THROUGH EMOTIONAL ANALYSIS OF DISTRIBUTED INPUTS
5y 2m to grant Granted Aug 11, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
69%
Grant Probability
99%
With Interview (+34.1%)
3y 0m (~1y 1m remaining)
Median Time to Grant
Low
PTA Risk
Based on 162 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month