DETAILED ACTION
This communication is in response to the Amendments and Arguments filed on 6/16/2026.
Claims 1-20 are pending and have been examined.
All previous objections / rejections not mentioned in this Office Action have been withdrawn by the examiner.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Regarding the Applicant’s arguments for the rejections under 35 U.S.C. § 103, applicant has amended independent claims 1 and 11. Hence, the Applicant’s arguments are moot in view of new grounds of rejection. More specifically, the newly added limitation to claim 1 and 11 are “obtaining, from the decoding graph, a reference forced-alignment path learned from the streaming speech recognition model, the reference forced-alignment path comprising forced-alignment frames”. The added limitations raise new grounds for rejection. Since Applicant’s arguments are directed towards the new amendment, the arguments are moot in view of new grounds for rejection. Hence, new references have been applied.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 8-11, and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Senior in view of Siohan et al. (U.S. PG Pub No. 20150379983), hereinafter Siohan.
Regarding claim 1 and 11 Senior teaches:
(Claim 1) A computer-implemented method executing on data processing hardware that causes the data processing hardware to perform operations comprising: (P0012, Systems, apparatus, and computer programs, configured to perform the actions of the methods, encoded on computer storage devices.)
(Claim 11) A system comprising: data processing hardware; and memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware causes the data processing hardware to perform operations comprising: (P0012, Systems, apparatus, and computer programs, configured to perform the actions of the methods, encoded on computer storage devices. A system of one or more computers can be so configured by virtue of software, firmware, hardware, or a combination of them installed on the system that in operation cause the system to perform the actions. One or more computer programs can be so configured by virtue of having instructions that, when executed by data processing apparatus, cause the apparatus to perform the actions.)
receiving a training data set comprising a sequence of acoustic frames corresponding to a spoken utterance paired with a ground-truth transcription of the spoken utterance; (P0038, The training state of the first neural network serves as the source of output target information for training the second neural network. For example, the first neural network has been trained to provide output values indicative of the likelihoods that different phones have been observed in input data.; P0043, For a particular training utterance in the audio data, the computing system divides the audio for the utterance into a series of frames representing a particular window or subset of the utterance.; P0049, For each frame of an utterance, the second neural network can be trained with the goal of matching the output distribution that the first neural network produced for the same frame of the same utterance.)
processing, using a streaming speech recognition model, the sequence of acoustic frames to generate a decoding graph for an output sequence of label tokens, the output sequence of label tokens comprising a speech recognition result for the utterance; (P0049, The training process can continue for many different training utterances, with the internal parameters of the second neural network being adjusted during each iteration. While the second neural network is trained to provide CTC-type output distributions.; P0056, CTC models may be trained with CD phone output labels.)
determining a speech recognition model loss based on the speech recognition result and the ground-truth transcription of the spoken utterance; (P0018, Training the second neural network using the loss function comprises training the second neural network using a loss function.; P0049, Cross-entropy training can be used to align the outputs of the second neural network with the output distributions of the first neural network.)
obtaining, from the decoding graph, a reference forced-alignment path learned from the streaming speech recognition model, the reference forced-alignment path comprising forced-alignment frames (P0075, A first neural network may be trained as described above (e.g., through forced alignment or using the CTC algorithm) with clean data as input.; P0068, In training hybrid neural network systems for speech recognition, the network may be trained with a cross-entropy loss with respect to fixed targets, which are determined by forced-alignment of a set of acoustic frames with a written transcript, transformed into the phonetic domain. Forced-alignment may find the maximum-likelihood label sequence for the acoustic frames and give labels for every frame either in {0, 1} for Viterbi alignment or in [0, 1] for Baum-Welch alignment.)
determining a training loss based on the speech recognition model loss and the forced-alignment path; and (P0051, The loss functions for multiple alignment techniques or output targets can be combined in a weighted combination used as the loss function for training. Besides the generated output distributions and CTC algorithm, other types of loss functions may be used, alone or together, including algorithms for Baum-Welch alignment and Viterbi alignment. The computing system can use a weighted combination of two or more of these loss functions as a loss function while updating the parameters of the second neural network.)
training the streaming speech recognition model based on the training loss. (P0075, A first neural network may be trained using audio signals or data as input that contain no noise or very little noise—so-called “clean data”. Noise may then be artificially added to the clean data and input to the neural network so that the neural network learns to separate the noise. In embodiments of the present disclosure, a first neural network may be trained as described above (e.g., through forced alignment or using the CTC algorithm) with clean data as input.; P0059, When starting from a written transcription and training a spoken-form model, which spoken form to use may be chosen. The training may apply Viterbi alignment to a lattice containing alternative pronunciations and allow the model to choose. CTC models may be trained using a unique alignment string, which may be derived from an alignment with a DNN model.)
Senior does not specifically teach:
obtaining, from the decoding graph, a reference forced-alignment path learned from the streaming speech recognition model, the reference forced-alignment path comprising forced-alignment frames
Siohan, however, teaches:
obtaining, from the decoding graph, a reference forced-alignment path learned from the streaming speech recognition model, the reference forced-alignment path comprising forced-alignment frames (P0052, The transcript acceptor may be composed with a lexicon transducer, a context-dependency transducer, and an HMM transducer to produce a forced-alignment decoding graph. Running Viterbi decoding may then provide a sequence of context-dependent HMM state symbols along the alignment path. A set of utterances may be described by the unigram distribution of the CD state symbols collected by running forced-alignment.)
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention to obtain a forced alignment path from a decoding graph. It would have been obvious to combine the references because the use of a decoding graph is a known technique to yield predictable results of creating an alignment path in the graph for forced alignment. (Senior P0073, Fig. 3; Siohan P0052)
Regarding claim 8 and 18 Senior in view of Siohan teach claim 1 and 11.
Senior further teaches:
wherein the data processing hardware resides on a server. (P0054, The computing system or another server system may receive audio data over a network from a user device of a user, then use the trained second neural network along with a language model and other speech recognition techniques to provide a transcription of the user's utterance to the user device.)
Regarding claim 9 and 19 Senior in view of Siohan teach claim 8 and 18.
Senior further teaches:
wherein, after training the streaming speech recognition model, the trained streaming speech recognition model is configured to execute on a user device. (P0054, the second neural network can be provided to the user device and the user device can perform speech recognition using the trained second neural network.)
Regarding claim 10 and 20 Senior in view of Siohan teach claim 1 and 11.
Senior further teaches:
wherein training the streaming speech recognition model based on the training loss comprises training the streaming speech recognition model based on the training loss without using any external aligner model to constrain alignment of the decoding graph. (P0033, Unlike many DNN, HMM, and GMM acoustic models, CTC models learn how to align phones and with audio data and are not limited to a specific forced alignment.; P0056, CTC models may be trained with CD phone output labels and a “blank” symbol. The models may be initially trained using the CTC algorithm to constantly realign with the Baum-Welch algorithm and train using a cross-entropy loss.)
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 2, 5, 12, and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Senior in view of Siohan and further view of "Transformer Transducer: A Streamable Speech Recognition Model with Transformer Encoders and RNN-T Loss" by Zhang et al., hereinafter Zhang.
Regarding claim 2 and 12 Senior teaches claim 1 and 11.
Senior in view of Siohan does not specifically teach:
generating, by an audio encoder of the streaming speech recognition model, at each of a plurality of time steps, a higher order feature representation for a corresponding acoustic frame in the sequence of acoustic frames;
receiving, as input to a label encoder of the streaming speech recognition model, a sequence of non-blank symbols output by a final softmax layer;
generating, by the label encoder, at each of the plurality of time steps, a dense representation;
receiving, as input to a joint network of the streaming speech recognition model, the higher order feature representation generated by the audio encoder at each of the plurality of time steps and the dense representation generated by the label encoder at each of the plurality of time steps; and
generating, by the joint network, at each of the plurality of time steps, a probability distribution over possible speech recognition hypotheses at the corresponding time step.
Zhang, however, teaches:
generating, by an audio encoder of the streaming speech recognition model, at each of a plurality of time steps, a higher order feature representation for a corresponding acoustic frame in the sequence of acoustic frames; (Fig. 1, Audio Encoder.; Abstract, The activations from both audio and label encoders are combined with a feed-forward layer to compute a probability distribution over the label space for every combination of acoustic frame position and label history.; 1. Introduction, The RNN-T architecture (as depicted in Figure 1) is a neural network architecture that can be trained end-to-end with the RNN-T loss to map input sequences (e.g. audio feature vectors) to target sequences (e.g. phonemes, graphemes).)
receiving, as input to a label encoder of the streaming speech recognition model, a sequence of non-blank symbols output by a final softmax layer; (2.1 RNN-T Architecture and Loss, The joint network combines the audio encoder output at ti and the label encoder output given the previous non-blank output label sequence Labels using a feedforward neural network with a softmax layer, inducing a distribution over the labels.)
generating, by the label encoder, at each of the plurality of time steps, a dense representation; (Fig. 1, Label Encoder.; Abstract, Transformer computation blocks based on self-attention are used to encode both audio and label sequences independently.; 1. Introduction, The RNN-T model gives a probability distribution over the label space at every time step.)
receiving, as input to a joint network of the streaming speech recognition model, the higher order feature representation generated by the audio encoder at each of the plurality of time steps and the dense representation generated by the label encoder at each of the plurality of time steps; and (Fig 1, Inputs to Joint Network.; 2.1 RNN-T Architecture and Loss, The joint network combines the audio encoder output at ti and the label encoder output given the previous non-blank output label sequence Labels(z1:(i-1)).)
generating, by the joint network, at each of the plurality of time steps, a probability distribution over possible speech recognition hypotheses at the corresponding time step. (Abstract, The activations from both audio and label encoders are combined with a feed-forward layer to compute a probability distribution over the label space for every combination of acoustic frame position and label history.; 1. Introduction, The RNN-T architecture1 (as depicted in Figure 1) is a neural network architecture that can be trained end-to-end with the RNN-T loss to map input sequences (e.g. audio feature vectors) to target sequences (e.g. phonemes, graphemes).)
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to utilize the audio encoder, label encoder, and joint network for streaming speech recognition model. It would have been obvious to combine the references because this model satisfies a constant computational requirement for processing each frame, making it suitable for streaming in speech recognition, given the simple architecture and parallelizable nature of self-attention computations (Zhang, 1. Introduction).
Regarding claim 5 and 15 Senior in view of Siohan and further view of Zhang teach claim 2 and 12.
Senior in view of Siohan does not specifically teach:
wherein the audio encoder comprises a plurality of multi-head attention layers.
Zhang, however, teaches:
wherein the audio encoder comprises a plurality of multi-head attention layers. (Figure 2, Transformer encoder architecture.; 2.2 Transformer, We implement AudioEncoder and LabelEncoder in Eq. (3), which are LSTMs in conventional RNN-T architectures, using the Transformers described above.; Figure 2, Masked multi-head attention with relative positional encoding.)
Claims 3, 6, 13, and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Senior in view of Siohan, in view of Zhang, and further view of "Transformer-Transducer: End-to-End Speech Recognition with Self-Attention" by Yeh et al., hereinafter Yeh.
Regarding claim 3 and 13 Senior in view of Siohan and further view of Zhang teach claim 2 and 12.
Senior in view of Siohan does not specifically teach, but Zhang teaches:
wherein the label encoder comprises a stack of transformer layers, each transformer layer comprising: (Figure 2, Transformer encoder architecture.; 2.2 Transformer, We implement AudioEncoder and LabelEncoder in Eq. (3), which are LSTMs in conventional RNN-T architectures, using the Transformers described above.)
a normalization layer; (Figure 2, Layer norm.)
a masked multi-head attention layer with relative position encoding; (Figure 2, Masked multi-head attention with relative positional encoding.)
residual connections; (Figure 2.; 2.2 Transformer, We then employ a residual connection on the normalized input and the output of the dense layer to form the final output of the multi-headed attention sublayer.)
a feedforward layer. (Figure 2, Feed forward.)
Senior in view of Zhang does not specifically teach:
a stacking/unstacking layer; and
Yeh, however, teaches:
a stacking/unstacking layer; and (5.2. Model Architectures and Details, For LSTM/BLSTM this is achieved with low frame rate in which every three consecutive frames are stacked and subsampled to form the new frame, and apply subsampling of factor 2 to the output of the second LSTM/BLSTM layer.)
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to utilize stacking/unstacking layer. It would have been obvious to combine the references because the reduction of frame rate from the stacking/unstacking layer creates efficient inference. (Yeh, Abstract)
Regarding claim 6 and 16 Senior in view of Siohan and further view of Zhang teach claim 2 and 12.
Senior in view of Siohan does not specifically teach, but Zhang teaches:
wherein the audio encoder comprises a stack of transformer layers, each transformer layer comprising: (Figure 2, Transformer encoder architecture.; 2.2 Transformer, We implement AudioEncoder and LabelEncoder in Eq. (3), which are LSTMs in conventional RNN-T architectures, using the Transformers described above.)
a normalization layer; (Figure 2, Layer norm.)
a masked multi-head attention layer with relative position encoding; (Figure 2, Masked multi-head attention with relative positional encoding.)
residual connections; (Figure 2.; 2.2 Transformer, We then employ a residual connection on the normalized input and the output of the dense layer to form the final output of the multi-headed attention sublayer.)
a feedforward layer. (Figure 2, Feed forward.)
Senior in view of Zhang does not specifically teach:
a stacking/unstacking layer; and
Yeh, however, teaches:
a stacking/unstacking layer; and (5.2. Model Architectures and Details, For LSTM/BLSTM this is achieved with low frame rate in which every three consecutive frames are stacked and subsampled to form the new frame, and apply subsampling of factor 2 to the output of the second LSTM/BLSTM layer.)
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to utilize stacking/unstacking layer. It would have been obvious to combine the references because the reduction of frame rate from the stacking/unstacking layer creates efficient inference. (Yeh, Abstract)
Claims 4 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Senior in view of Siohan, in view of Zhang, and further view of "An improved embedding matching model for Chinese word segmentation" by Deng et al., hereinafter Deng.
Regarding claim 4 and 14 Senior in view of Siohan and further view of Zhang teach claim 2 and 12.
Senior in view of Siohan and further view of Zhang does not specifically teach:
wherein the label encoder comprises a bigram embedding lookup decoder model.
Deng, however, teaches:
wherein the label encoder comprises a bigram embedding lookup decoder model. (A. General Description, The model consists of two distinguishing look-up tables used for storing and extracting unigram and bigram embeddings. The embeddings stored in the feature embeddings look-up table are d-dimensional real-valued vectors.; B. Improvement of Model Structure, Our objective is to search a path of tags with the highest score, thus we utilized the Viterbi decoding algorithm.)
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to utilize bigram embedding lookup for the label encoder. It would have been obvious to combine the references because using bigram embedding utilizing a lookup table is a known technique to yield a predictable result of encoding labels. (Deng III, Feature Extraction and Improvement)
Claims 7 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Senior in view of Siohan and further view of Zhao et al. (U.S. PG Pub No. 20210312905), hereinafter Zhao.
Regarding claim 7 and 17 Senior in view of Siohan teach claim 1 and 11.
Senior in view of Siohan does not specifically teach:
wherein the streaming speech recognition model comprises a recurrent neural network-transducer (RNN-T) model architecture.
Zhao, however, teaches:
wherein the streaming speech recognition model comprises a recurrent neural network-transducer (RNN-T) model architecture. (P0002, Analyzing the audio input using a Recurrent Neural Network-Transducer (RNN-T) to obtain textual content representing the spoken content.)
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to utilize recurrent neural network-transducer (RNN-T) model architecture. It would have been obvious to combine the references because RNN-T can support real-time streaming speech recognition. (Zhao P0036)
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DANIEL WONSUK CHUNG whose telephone number is (571)272-1345. The examiner can normally be reached Monday - Friday (7am-4pm)[PT].
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, PIERRE-LOUIS DESIR can be reached at (571)272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DANIEL W CHUNG/Examiner, Art Unit 2659
/PIERRE LOUIS DESIR/Supervisory Patent Examiner, Art Unit 2659