Prosecution Insights
Last updated: October 02, 2026
Application No. 18/741,506

SPACE EFFICIENT TRAINING FOR SEQUENCE TRANSDUCTION MACHINE LEARNING

Non-Final OA §103
Filed
Jun 12, 2024
Priority
Mar 15, 2024 — provisional 63/566,055
Examiner
CADY, MATTHEW ALAN
Art Unit
Tech Center
Assignee
Microsoft Technology Licensing, LLC
OA Round
1 (Non-Final)
0%
Grant Probability
At Risk
1-2
OA Rounds
1y 0m
Est. Remaining
0%
With Interview

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 1 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 4m
Avg Prosecution
26 currently pending
Career history
19
Total Applications
across all art units

Statute-Specific Performance

§101
10.4%
-29.6% vs TC avg
§103
68.7%
+28.7% vs TC avg
§102
11.3%
-28.7% vs TC avg
§112
9.6%
-30.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-2 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wei Kang et al. (hereinafter Kang) (“US 20230386449 A1,” 2023-11-30) in view of Jinyu Li et al. (hereinafter Li) (“IMPROVING RNN TRANSDUCER MODELING FOR END-TO-END SPEECH RECOGNITION,” 2019-09-26) further in view of Chengquan Jiang et al. (hereinafter Jiang) (US 20230359697 A1, 2023-11-09). Regarding claim 1, Kang teaches; A method for training a neural transducer, the method comprising: ([0027] training the RNN-T model) obtaining a neural transducer comprising an encoder, a decoder, and a ([0027] The RNN-T model usually consists of three parts, namely, an encoder network, a prediction network and a joiner network) … providing an encoder training batch to the encoder and a decoder training batch to the decoder; ([0029] training audio data and the text label corresponding to the training audio data are input to the encoder network and the prediction network in batch) obtaining encoding embeddings and decoding embeddings from the encoder and decoder, respectively; ([0029] the first encoding result is a vector with dimensions of (N, T, E), and the first prediction result is a vector with dimensions of (N, U, D)) providing the encoding embeddings and decoding embeddings to the ([0031] the first encoding result and the first prediction result can be input to a joint module ..., wherein the joint module can be a trivial joiner network) configuring the ([0031] joint module for a joint processing to obtain the first joint result … [0034] the first joint result is a matrix L of (T, U, V) … each element 1 (t, u, v) in the matrix L represents a probability of a character v … output by an element whose coordinate is t in the audio frame number dimension and u in the text label sequence dimension … Therefore, following definitions can be made: y(t, u)=L(t, u, y.sub.u+1) … Ø(t, u)=L(t, u, Ø) … [0035] y(t, u) represents … the probability of outputting the text label… Ø(t, u) represents … the probability of outputting the null character) computing a loss for the neural transducer ([0025] calculating the RNN-T loss); … update the parameters of the neural transducer to minimize the loss ([0054] The adjustment of the network parameters can be stopped when the first network loss value is within a preset error range); and generating a modified neural transducer (the resulting trained RNN-T … [0054] the network parameters of the encoder network, the prediction network and the joiner network can be adjusted according to the first network loss value). Kang fails to explicitly teach but Li teaches; joint network ([pg. 2] The joint network … combines the encoder network output … and the prediction network output … as zt,u) … linear ([pg. 2] zt,u is connected to the output layer with a linear transform) and softmax layer (see below); [pg. 2] PNG media_image1.png 445 428 media_image1.png Greyscale backpropagating the loss through the neural transducer [pg. 3] PNG media_image2.png 240 799 media_image2.png Greyscale OBVIOUSNESS TO COMBINE LI: Kang and Li are both analogous to the present disclosure as they pertain to RNN-T models. It would have been obvious to one of ordinary skill in the art, before the effective filing date, to implement the output processing and training of Kang’s RNN-T according to Li, such that the output of the joint network is provided to a linear transformation followed by a softmax layer and the resulting loss gradient is propagated backward through the neural transducer using the chain rule. Li teaches that such linear and softmax processing converts the join network representation into normalized output probabilities suitable for the RNN-T loss computation ([Li, pg. 2] The final posterior for each output token k is obtained after applying softmax operation P(k|t,u) = softmax(hkt,u) … loss L with respect to P(k|t,u)), and further teaches propagating the resulting loss gradient backward through the output processing during training. Kang similarly trains an RNN-T by calculating a loss and using gradient information to adjust model parameters to minimize the loss. A person of ordinary skill therefore would have had reason to employ Li’s known RNN-T output and backpropagation arrangement in Kang to provide the normalized symbol probabilities used in Kang’s loss calculation and to propagate that loss through the neural transducer for parameter optimization, with the predictable result of training Kang’s RNN-T using a known gradient-based training technique. Accordingly, Kang as modified by Li teaches; backpropagating the loss through the neural transducer (Li) to update the parameters of the neural transducer to minimize the loss (Kang); Kang and Li fail to explicitly teach but Jiang teaches; a fused linear and softmax layer ([0086] the softmax operation after Q×K may be fused with matrix multiplication) OBVIOUSNESS TO COMBINE JIANG: Jiang is analogous art to the present disclosure as it pertains to machine learning operator fusion. It would have been obvious to one of ordinary skill in the art to modify the RNN-T output processing taught by Li, as applied in Kang, to fuse Li’s linear transformation with its subsequent softmax operation in the manner taught by Jiang. Li teaches that the output of the joint network is subjected to a linear transformation followed by a softmax operation, while Jiang teaches that a matrix-multiplication operation and a following softmax operation may be fused into a common kernel to reduce accesses to intermediate memory ([Jiang, 0086] In order to further improve performance, the softmax operation after Q×K may be fused with matrix multiplication, which saves memory access operations for intermediate matrices compared to a separate kernel for the softmax operation). Because Li’s Linear transformation is implemented by matrix multiplication ([Li, pg. 2] linear transform ht,u = Wyzt,u +by), a person of ordinary skill would have recognized Jiang’s fusion technique as directly applicable to Li’s adjacent linear and softmax operations. The modification would have predictably reduced intermediate memory traffic and memory overhead during RNN-T training while preserving the same normalized output probabilities used to calculate the transducer loss. Accordingly, the combination of Kang, Li, and Jiang teaches; A fused (Jiang) joint network (Li) … wherein the fused (Jiang) joint network (Li) has been modified to comprise a fused (Jiang) linear and softmax layer (Li); Regarding claim 2, Kang teaches; the encoder training batch comprises audio data ([Abstract] encoding training audio data input to an encoder network) and the decoder training batch comprises corresponding textual transcriptions ([Abstract] predicting a text label corresponding to the training audio data input to a prediction network). ([0029] training audio data and the text label corresponding to the training audio data are input to the encoder network and the prediction network in batch) Claim(s) 3-4 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kang (“US 20230386449 A1,” 2023-11-30) in view of Li (“IMPROVING RNN TRANSDUCER MODELING FOR END-TO-END SPEECH RECOGNITION” 2019-09-26) further in view of Jiang (US 20230359697 A1, 2023-11-09) as applied to claim 1 above, further in view of Vihn Nguyen et al. (hereinafter Nguyen) (“MLPerf v1.0 Training Benchmarks: Insights into a Record-Setting NVIDIA Performance,” 2021-07-30). Regarding claim 3, Kang teaches; joint embeddings based on the encoding embeddings and decoding embeddings ([0031] the first encoding result and the first prediction result can be input to a joint module for a joint processing to obtain the first joint result, wherein the joint module can be a trivial joiner network) Kang fails to teach but Li teaches; joint network ([pg. 2] joint network) OBVIOUSNESS: Using the same reasoning from claim 1 Kang and Li fail to teach but Jiang teaches fused ([0086] softmax operation after Q×K may be fused with matrix multiplication) OBVIOUSNESS: Using the same reasoning from claim 1. Accordingly, Li and Jiang teaches; fused (Jiang) joint network (Li) Kang, Li, and Jiang fail to teach but Nguyen teaches; joint network is configured to divide a computation … into slices. ([pg. 23] RNN-T … joint net operates on a portion of the batch and loops through those subbatches one by one … a batch splitting factor of 2 is used … batch sizes of the inputs to … the joint net are … B/2) OBVIOUSNESS TO COMBINE NGUYEN: Nguyen is analogous art to the present disclosure as it discloses a method for decreasing memory usage by an RNN-T model. It would have been obvious to one of ordinary skill in the art, before the effective filing date, to configure the fused joint network of Kang as modified by Li and Jiang to divide its joint computation into slices / subbatches as taught by Nguyen. Nguyen teaches that the RNN-T joint network operates on substantially larger tensors and therefore processes portions of the batch sequentially to reduce GPU memory requirements ([Nguyen, pg. 23] RNN-T … the joint net takes much larger tensors … [to not exceed] the GPU memory capacity by having a huge tensor in the joint net, we employed a technique called batch splitting). A person of ordinary skill seeking to reduce the memory required by the joint network computation of Kang as modified by Li and Jiang therefore would have had a reason to divide the computation performed on the encoder and decoder outputs into smaller slices as taught by Nguyen, with the predictable result of reducing peak memory consumption while performing the same joint network computation. Regarding claim 4, Kang teaches; joint network ([0027] The RNN-T model usually consists of … a joiner network) Kang fails to teach but Li teaches; joint network ([pg. 2] The joint network … combines the encoder network output … and the prediction network output … as zt,u) is configured to compute a … linear ([pg. 2] zt,u is connected to the output layer with a linear transform ht,u = Wyzt,u +by) and softmax function ([pg. 2] The final posterior for each output token k is obtained after applying softmax operation P(k|t,u) = softmax(hkt,u)) for each … of the joint embeddings (zt,u). OBVIOUSNESS: Using the same reasoning from claim 1. Kang and Li fail to teach but Jiang teaches; fused linear and softmax function ([0086] the softmax operation after Q×K may be fused with matrix multiplication) OBVIOUSNESS: Using the same reasoning from claim 1. Accordingly, Li and Jiang teaches; fused (Jiang) joint network (Li) … fused (Jiang) linear and softmax function (Li) Kang, Li, and Jiang fail to teach but Nguyen teaches; joint network is configured to compute … for each slice ([pg. 23] RNN-T … joint net operates on a portion of the batch and loops through those subbatches one by one … a batch splitting factor of 2 is used … batch sizes of the inputs to … the joint net are … B/2) OBVIOUSNESS: Using the same reasoning from claim 3. Accordingly, the combination of Kang, Li, Jiang, and Nguyen teaches; the fused (Jiang) joint network (Kang, Li, and Nguyen) is configured to compute a fused (Jiang) linear and softmax function (Li) for each slice (Nguyen) of the joint embeddings (Kang and Li). Claim(s) 5 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kang (“US 20230386449 A1,” 2023-11-30) in view of Li (“IMPROVING RNN TRANSDUCER MODELING FOR END-TO-END SPEECH RECOGNITION” 2019-09-26) further in view of Jiang (US 20230359697 A1, 2023-11-09), further in view of Nguyen (“MLPerf v1.0 Training Benchmarks: Insights into a Record-Setting NVIDIA Performance,” 2021-07-30) as applied to claim 4 above, further in view of Ivan Sorokin et al. (hereinafter Sorokin) (“CUDA-Warp RNN-Transducer,” 2021-08-23). Regarding claim 5, Kang teaches; joint network ([0027] The RNN-T model usually consists of … a joiner network) and Li teaches; joint network ([pg. 2] The joint network … combines the encoder network output … and the prediction network output … as zt,u) OBVIOUSNESS: Using the same reasoning from claim 1. Kang and Li fail to teach but Jiang teaches; fused ([0086] the softmax operation after Q×K may be fused with matrix multiplication) OBVIOUSNESS: Using the same reasoning from claim 1. Accordingly, Li and Jiang teaches; fused (Jiang) joint network (Li) Kang, Li and Jiang fail to teach but Nguyen teaches; joint network … computation of each slice ([pg. 23] RNN-T … joint net operates on a portion of the batch and loops through those subbatches one by one … a batch splitting factor of 2 is used … batch sizes of the inputs to … the joint net are … B/2) OBVIOUSNESS: Using the same reasoning from claim 3. Kang, Li, Jiang and Nguyen fail to explicitly teach but Sorokin teaches; … discard (discarding can simply mean not storing, as stated by the applicant specification: [0081] discarding or otherwise refraining from storing other prediction outputs) all outputs from the computation … except for the next blank output and the next token output ([pg. 3] memory-efficient version that expects log_probs with the shape (N, T, U, 2) only for blank and labels values … trainable joint network with an output (N, T, U, V)) OBVIOUSNESS TO COMBINE SOROKIN: Sorokin is analogous art to the present disclosure as it pertains to an RNN-T training framework. It would have been obvious to one of ordinary skill in the art, before the effective filing date, to modify the RNN-T training system of Kang, as modified by Li, Jiang, and Nguyen, to retain from each sliced joint network computation only the blank output and the corresponding next target / label token output, as taught by Sorokin. Kang, Li, and Nguyen each recognize that RNN-T training produces large intermediate joint tensors that impose substantial memory requirements, while Jiang teaches fusing otherwise separate output processing operations to avoid storing and retrieving unnecessary intermediate data. Moreover, Kang teaches separately constructing only the next token and blank outputs, thereby “[Kang, 0036] avoiding the allocation of the huge four-dimensional matrix and reducing the memory and improving the speed.” Additionally, Li teaches that the RNN-T loss gradient calculation has relevant nonzero terms for the next target token and blank, providing further reason to retain those outputs while refraining from storing the other vocabulary outputs (See Li, pg. 2). Sorokin further teaches a memory efficient RNN-T loss implementation that uses a (N, T, U, 2) probability representation containing only blank and label values, rather than the conventional vocabulary joint network output representation having dimensions (N, T, U, V), thereby demonstrating that only the blank and corresponding target label probabilities need be provided for the memory efficient RNN-T loss computation. A person of ordinary skill would therefore have been motivated to incorporate Sorokin’s known blank and label only probability representation into the sliced, fused, joint processing of the combined system such that only the blank and corresponding next target token probabilities required for subsequent RNN-T loss computation are written to memory. Doing so would predictably reduce memory consumption and memory traffic while preserving the information required to compute the RNN-T loss. Claim(s) 6 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kang (“US 20230386449 A1,” 2023-11-30) in view of Li (“IMPROVING RNN TRANSDUCER MODELING FOR END-TO-END SPEECH RECOGNITION” 2019-09-26) further in view of Jiang (US 20230359697 A1, 2023-11-09), further in view of Nguyen (“MLPerf v1.0 Training Benchmarks: Insights into a Record-Setting NVIDIA Performance,” 2021-07-30), further in view of Sorokin (“CUDA-Warp RNN-Transducer,” 2021-08-23), as applied to claim 5 above, further in view of Zhifeng Chen et al. (hereinafter Chen) (US 20210042620 A1, 2021-02-11). Regarding claim 6, Kang teaches; joint embeddings ([0031] the joint module can be a trivial joiner network … joint module for a joint processing to obtain the first joint result) Kang fails to explicitly teach but Li teaches; the backpropagation of the loss through the neural transducer (using the same reasoning as in claim 1) OBVIOUSNESS: Using the same reasoning from claim 1. Kang, Li, and Jiang fail to teach but Nguyen teaches; dividing the computation of joint [network outputs] into the slices ([pg. 23] RNN-T … joint net operates on a portion of the batch and loops through those subbatches one by one … a batch splitting factor of 2 is used … batch sizes of the inputs to … the joint net are … B/2) OBVIOUSNESS: Using the same reasoning from claim 3. Kang, Li, Jiang, Nguyen, and Sorokin fail to explicitly teach but Chen teaches backpropagation of the loss through the neural [network] ([0077] compute a loss … for the neural network … the system performs backpropagation operations at each network layer to compute an output gradient of an objective function) … includes recalculating the [outputs] ([0062] outputs generated by an internal layer … recomputed … for the backpropagation function for the internal layer) OBVIOUSNESS TO COMBINE CHEN: Chen is analogous art to the present disclosure as it pertains to neural network training techniques. Kang teaches a joint network of an RNN-T which computes joint embeddings, Li teaches backpropagating loss through an RNN-T, Nguyen teaches dividing the RNN-T joint-network computation into subbatches / slices to reduce memory usage, while Chen teaches reducing neural network training memory by recomputing forward activations when those activations are required during backpropagation ([0077] The recomputation technique described above reduces peak memory requirement). Therefore, it would have been obvious to apply Chen’s known recomputation technique to Nguyen’s subbatch-wise joint network processing such that, during backpropagation, the joint network computation for each subbatch is re-executed to regenerate the corresponding joint embeddings when needed for gradient calculation, thereby avoiding the need to retain the joint embeddings from the forward pass and predictably reducing peak training memory consumption. Accordingly, the combination teaches; the backpropagation of the loss through the neural transducer (Li) includes recalculating (Chen) the joint embeddings (Kang) by dividing the computation of joint embeddings into the slices (Nguyen). Claim(s) 7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wei Kang et al. (hereinafter Kang) (“US 20230386449 A1,” 2023-11-30) in view of Jinyu Li et al. (hereinafter Li) (“IMPROVING RNN TRANSDUCER MODELING FOR END-TO-END SPEECH RECOGNITION” 2019-09-26) further in view of Chengquan Jiang et al. (hereinafter Jiang) (US 20230359697 A1, 2023-11-09) as applied to claim 1 above, further in view of Shang-Xuan Zou et al. (hereinafter Zou) (US 20180357541 A1, 2018-12-13). Regarding claim 7, Kang teaches; neural transducer Kang, Li, and Jiang fail to teach but Zou teaches; dynamically manage memory ([0049] perform the heuristic computation repeatedly until a maximum mini-batch size fit to the estimated available memory is found) storage by at least one of (a) increasing the batch size ([0049] try the mini-batch sizes with an ascending order in the heuristic computation … can try the target mini-batch size of 32, 64, 128, 256, 512, and 1024 with the ascending order in the heuristic computation) and/or (b) storing slices of transformed embeddings based on a memory storage maximum capacity and/or a predetermined threshold of memory storage. OBVIOUSNESS: Zou is analogous art to the present disclosure because it relates to memory aware optimization of neural network training on computing devices having limited memory resources. It would have been obvious to modify the RNN-T training system of Kang, as modified by Li and Jiang, to increase the training batch size according to available memory capacity, as taught by Zou. Kang and Li recognize that the large intermediate tensors generated during RNN-T training impose substantial memory requirements, and Jiang teaches reducing intermediate memory usage and memory traffic through fused processing. Zou teaches estimating available memory based on GPU memory constraints and progressively testing increasingly larger mini-batch sizes that fit within the available memory. A person of ordinary skill would therefore have been motivated to apply the known memory aware batch sizing technique of Zou to the memory efficient RNN-T system of Kang as modified by Li and Jiang so that memory made available by the RNN-T memory reduction techniques can be utilized to process larger training batches without exceeding the computing system’s memory capacity. Such a modification would have predictably improved memory utilization and training throughput while avoiding exceeding the allotted memory. Claim(s) 8-9, 12-14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kang (“US 20230386449 A1,” 2023-11-30) in view of Li (“IMPROVING RNN TRANSDUCER MODELING FOR END-TO-END SPEECH RECOGNITION” 2019-09-26) further in view of Nguyen (“MLPerf v1.0 Training Benchmarks: Insights into a Record-Setting NVIDIA Performance,” 2021-07-30), further in view of Chen (US 20210042620 A1, 2021-02-11). Regarding claim 8, Kang teaches; A method for training a neural transducer ([0027] a process of training the RNN-T model), the method comprising: … computing a loss for the neural transducer based on the training with the joint embeddings ([0054] In a possible embodiment, a first network loss value can be determined according to the second joint result … When calculating the first network loss value, the forward-backward algorithm can be used); … update the parameters of the neural transducer to minimize the loss ([0054] The adjustment of the network parameters can be stopped when the first network loss value is within a preset error range); and generating a modified neural transducer (the resulting trained neural transducer … [0027] training the RNN-T model). Kang fails to explicitly teach but Li teaches; backpropagating the loss through the See equation (12) below) [pg. 3] PNG media_image3.png 191 872 media_image3.png Greyscale OBVIOUSNESS: Using the same reasoning from claim 1. Kang and Li fail to teach but Nguyen teaches; during training, dividing a computation of [joiner network outputs] into slices; (joint network training outputs generated from individual subbatches / slices … [pg. 23] joint net operates on a portion of the batch and loops through those subbatches one by one … tensors generated by the joint net) OBVIOUSNESS: Nguyen is analogous art to the present disclosure as it discloses a method for decreasing memory usage by an RNN-T model. It would have been obvious to one of ordinary skill in the art, before the effective filing date, to configure the joint network of Kang as modified by Li to divide its joint embeddings computation into slices / subbatches as taught by Nguyen. Nguyen teaches that the RNN-T joint network operates on substantially larger tensors and therefore processes portions of the batch sequentially to reduce GPU memory requirements ([Nguyen, pg. 23] RNN-T … the joint net takes much larger tensors … [to not exceed] the GPU memory capacity by having a huge tensor in the joint net, we employed a technique called batch splitting). A person of ordinary skill seeking to reduce the memory required by the joint network computation of Kang as modified by Li therefore would have had a reason to divide the computation performed on the encoder and decoder outputs into smaller slices as taught by Nguyen, with the predictable result of reducing peak memory consumption while performing the same joint network computation. Accordingly, the combination teaches; during training, dividing a computation (Nguyen) of joint embeddings (Kang, Li) into slices (Nguyen); … backpropagating the loss (Li) through the slices of the neural transducer (Nguyen) Kang, Li, and Nguyen fail to teach but Chen teaches; recalculating the [outputs] (taught above) during backpropagation; ([0062] outputs generated by an internal layer … recomputed … for the backpropagation function for the internal layer) OBVIOUSNESS: Chen is analogous art to the present disclosure as it pertains to neural network training techniques. Kang teaches a joint network of an RNN-T computing joint embeddings, Li teaches backpropagating loss through an RNN-T, Nguyen teaches dividing the RNN-T joint-network computation into subbatches / slices to reduce memory usage, while Chen teaches reducing neural network training memory by recomputing forward activations when those activations are required during backpropagation ([0077] The recomputation technique described above reduces peak memory requirement). Therefore, it would have been obvious to apply Chen’s known recomputation technique to Nguyen’s subbatch-wise joint network processing such that, during backpropagation, the joint network computation for each subbatch is re-executed to regenerate the corresponding joint embeddings when needed for gradient calculation, thereby avoiding the need to retain the joint embeddings from the forward pass and predictably reducing peak training memory consumption. Regarding claim 9, Kang and Li fail to teach but Nguyen teaches; wherein the slices of the computation of joint embeddings are divided based on a predetermined size. (joint network outputs generated from individual subbatches that are divided based on a predetermined size … [pg. 23] joint net operates on a portion of the batch and loops through those subbatches one by one … a batch splitting factor of 2 is used. In this case, the batch sizes of the ... joint net are B and B/2, respectively) OBVIOUSNESS: Using the same reasoning from claim 8 Regarding claim 12, Kang fails to explicitly teach but Li teaches; the computation of joint embeddings (z, see equation (3) below) includes applying a non-linear function (non-linear function ψ, see equation (4) below) after every matrix multiplication (Uh, Vh, see equation (4) below). [pg. 2] PNG media_image4.png 367 999 media_image4.png Greyscale OBVIOUSNESS: Using the same reasoning from claim 1. Regarding claim 13, Kang fails to explicitly teach but Li teaches; the loss is computed based on a difference ([pg. 2] The loss function of RNN-T is … L = −lnP(y|x)) between a model output generated by applying the neural transducer to training data ([pg. 2] input acoustic feature x … the forward-backward algorithm) and a ground truth output ([pg. 2] output label sequence y). OBVIOUSNESS: Using the same reasoning from claim 1. Regarding claim 14, Kang teaches; joint embeddings (using the same reasoning from claim 8) Kang and Li fail to teach but Nguyen teaches; slices of the [joint network output] (using the same reasoning from claim 8) OBVIOUSNESS: Using the same reasoning from claim 8. Kang, Li, and Nguyen fail to teach but Chen teaches; the recalculating of the [outputs] during backpropagation ([0062] outputs generated by an internal layer … recomputed … for the backpropagation function for the internal layer) includes recomputing the [outputs] in a same configuration of slices of the [outputs] used during a forward pass of the training ([0095] The backpropagation of the neural network mirrors the forward pass: beginning at the next time-step following the end of the forward pass, the last device computes the composite backpropagation function for the last composite layer on the last micro-batch in the plurality of micro-batches). OBVIOUSNESS: Using the same reasoning from claim 8. Chen further teaches that the backward pass mirrors the forward pass and processes the same micro-batches. Accordingly, applying Chen’s recomputation technique to Nguyen’s sliced joint network processing would predictably recompute the joint embeddings during backpropagation using the same subbatch / slice configuration employed during the forward pass. Claim(s) 10-11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kang (“US 20230386449 A1,” 2023-11-30) in view of Li (“IMPROVING RNN TRANSDUCER MODELING FOR END-TO-END SPEECH RECOGNITION” 2019-09-26) further in view of Nguyen (“MLPerf v1.0 Training Benchmarks: Insights into a Record-Setting NVIDIA Performance,” 2021-07-30), further in view of Chen (US 20210042620 A1, 2021-02-11) as applied to claim 8 above, further in view of Venkataramani Swagath et al. (hereinafter Swagath) (US 20200311536 A1, 2020-10-01). Regarding claim 10, Kang teaches; computation of joint embeddings ([0031] the first encoding result and the first prediction result can be input to a joint module for a joint processing to obtain the first joint result, wherein the joint module can be a trivial joiner network) Kang and Li fail to teach but Nguyen teaches; slices of [joint network outputs] (joint network outputs generated from individual subbatches / slices … [pg. 23] joint net operates on a portion of the batch and loops through those subbatches one by one … tensors generated by the joint net) OBVIOUSNESS: Using the same reasoning from claim 8. Kang, Li, Nguyen, and Chen fail to teach but Swagath teaches; divided based on a memory storage capacity of a computing system performing the method. ([0028] The spatial minibatch size may be determined based … the on-chip memory capacity of a chip … [0029] a given minibatch is split into multiple spatial minibatches) OBVIOUSNESS: Swagath is analogous art to the present disclosure because it pertains to memory efficient neural network training. It would have been obvious to one of ordinary skill in the art, before the effective filing date, to determine the size of Nguyen’s RNN-T joint network subbatches according to the available memory as taught by Swagath. Nguyen teaches dividing the joint network computation into smaller subbatches to avoid exceeding GPU memory, while Swagath teaches determining a particular minibatch size based on on-chip memory capacity. Applying Swagath’s memory-based sizing to Nguyen’s subbatch wise joint network processing in the RNN-T training of Kang as modified by Li, Nguyen, and Chen would predictably divide the corresponding slices of joint embedding computation based on the memory storage capacity of the computing system, while reducing memory requirements and avoiding memory overflow. Regarding claim 11, Kang teaches; computation of joint embeddings ([0031] the first encoding result and the first prediction result can be input to a joint module for a joint processing to obtain the first joint result, wherein the joint module can be a trivial joiner network) Kang and Li fail to teach but Nguyen teaches; slices of [joint network outputs] (joint network outputs generated from individual subbatches / slices … [pg. 23] joint net operates on a portion of the batch and loops through those subbatches one by one … tensors generated by the joint net) OBVIOUSNESS: Using the same reasoning from claim 8. Kang, Li, Nguyen, and Chen fail to teach but Swagath teaches; divided based on a computational capacity of a computing system performing the method. ([0018] choose a minibatch size that is large enough to make efficient use of resources (e.g., make use of all the cores), but small enough that a layer's output fits on chip) OBVIOUSNESS: Swagath is analogous art to the present disclosure because it pertains to memory efficient neural network training. It would have been obvious to one of ordinary skill in the art, before the effective filing date, to determine the size of Nguyen’s RNN-T joint network subbatches according to the available computing resources as taught by Swagath. Nguyen teaches dividing the joint network computation into smaller subbatches to avoid exceeding GPU memory, while Swagath teaches determining a particular minibatch size based on the available computational resources, to make efficient use of resources. Applying Swagath’s core-based sizing to Nguyen’s subbatch wise joint network processing in the RNN-T training of Kang as modified by Li, Nguyen, and Chen would predictably divide the corresponding slices of joint embedding computation based on the memory storage capacity of the computing system, while ensuring that the slices are large enough to make efficient use of computational resources. Claim(s) 15-17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kang (“US 20230386449 A1,” 2023-11-30) in view of Li (“IMPROVING RNN TRANSDUCER MODELING FOR END-TO-END SPEECH RECOGNITION” 2019-09-26) further in view of Nguyen (“MLPerf v1.0 Training Benchmarks: Insights into a Record-Setting NVIDIA Performance,” 2021-07-30) further in view of Swagath (US 20200311536 A1, 2020-10-01). Regarding claim 15, Kang teaches; A method for training a neural transducer ([0027] training the RNN-T model), the method comprising: … update the parameters of the neural transducer to minimize the loss ([0054] The adjustment of the network parameters can be stopped when the first network loss value is within a preset error range); and generating a modified neural transducer ([0054] the network parameters of the encoder network, the prediction network and the joiner network can be adjusted according to the first network loss value). Kang fails to explicitly teach but Li teaches; transformed embeddings ([pg. 2] The joint network is a feed-forward network that combines the encoder network output … and the prediction network output … as zt,u … ht,u = Wyzt,u +by) … computing a loss for the neural transducer based on the transformed embeddings ([pg. 2] P(k|t,u) = softmax(hkt,u) … loss L with respect to P(k|t,u)); backpropagating the loss through the neural transducer (See equation (12) below) [pg. 3] PNG media_image3.png 191 872 media_image3.png Greyscale OBVIOUSNESS: Using the same reasoning from claim 1. Kang and Li fail to teach but Nguyen teaches; during training, storing slices of [joint network outputs] (joint network training outputs generated from individual subbatches / slices, and stored for use in backpropagation / training … [pg. 23] joint net operates on a portion of the batch and loops through those subbatches one by one … tensors generated by the joint net, … are no longer needed after the backpropagation is completed, they can be released) OBVIOUSNESS: Nguyen is analogous art to the present disclosure as it discloses a method for decreasing memory usage by an RNN-T model. It would have been obvious to one of ordinary skill in the art, before the effective filing date, to configure the joint network of Kang as modified by Li to divide its joint computation into slices / subbatches as taught by Nguyen. Nguyen teaches that the RNN-T joint network operates on substantially larger tensors and therefore processes portions of the batch sequentially to reduce GPU memory requirements ([Nguyen, pg. 23] RNN-T … the joint net takes much larger tensors … [to not exceed] the GPU memory capacity by having a huge tensor in the joint net, we employed a technique called batch splitting). A person of ordinary skill seeking to reduce the memory required by the joint network computation of Kang as modified by Li therefore would have had a reason to divide the computation performed on the encoder and decoder outputs into smaller slices as taught by Nguyen, with the predictable result of reducing peak memory consumption while performing the same joint network computation. Kang, Li, and Nguyen fail to teach but Swagath teaches; and dynamically modifying at least one of: (a) training data batch size ([0003] dynamically resizing a minibatch in a neural network execution, wherein a size of the minibatch is configured such that the minibatch fits within on-chip memory … [0014] Minibatch is usually a subset of a training data) and/or (b) a quantity or size of slices of transformed embeddings to be stored, the dynamic modification being based on a memory storage maximum capacity and/or a predetermined threshold of memory storage of a computing system performing the method; OBVIOUSNESS TO COMBINE SWAGATH: Swagath is analogous art to the present disclosure as it pertains to dynamically modifying training data batch size. It would have been obvious to one of ordinary skill in the art, before the effective filing date, to dynamically resize the training minibatches used in the RNN-T training of Kang as modified by Li and Nguyen according to the available memory technique of Swagath. Nguyen teaches that RNN-T joint network processing produces substantially larger tensors and that dividing the batch into smaller portions permits training without exceeding available memory. Swagath similarly addresses memory limitations during neural network training and teaches determining available memory and dynamically resizing a training minibatch based on the available memory. A person of ordinary skill seeking to reduce the memory requirements of the RNN-T training of Kang as modified by Li and Nguyen would have had reason to dynamically resize the training minibatch according to available memory as taught by Swagath, because Swagath teaches that reducing a minibatch to fit on-chip memory saves memory bandwidth and provides a performance benefit ([0017] A method in one embodiment, to fit an output on chip, breaks down the minibatch size into smaller sizes … This method can save memory bandwidth as input need not be transferred from off-chip, thereby providing a performance benefit). The modification would predictably permit the RNN-T training workload to be processed within the available memory while improving memory utilization and performance. Regarding claim 16, Kang fails to explicitly teach but Li teaches transformed embeddings ([pg. 2] The joint network is a feed-forward network that combines the encoder network output … and the prediction network output … as zt,u … ht,u = Wyzt,u +by) OBVIOUSNESS: Using the same reasoning from claim 1. Kang and Li fail to teach but Nguyen teaches; the slices of [joint network outputs] are divided based on a predetermined size (joint network outputs generated from individual subbatches that are divided based on a predetermined size … [pg. 23] joint net operates on a portion of the batch and loops through those subbatches one by one … a batch splitting factor of 2 is used. In this case, the batch sizes of the ... joint net are B and B/2, respectively) OBVIOUSNESS: Using the same reasoning from claim 15 Regarding claim 17, Kang fails to explicitly teach but Li teaches transformed embeddings ([pg. 2] The joint network is a feed-forward network that combines the encoder network output … and the prediction network output … as zt,u … ht,u = Wyzt,u +by) OBVIOUSNESS: Using the same reasoning from claim 1. Kang and Li fail to teach but Nguyen teaches; slices of [joint network outputs] (joint network outputs generated from individual subbatches / slices … [pg. 23] joint net operates on a portion of the batch and loops through those subbatches one by one … tensors generated by the joint net) OBVIOUSNESS: Using the same reasoning from claim 15 Kang, Li, and Nguyen fail to teach but Swagath teaches; divided based on the memory storage capacity of the computing system performing the method. ([0028] The spatial minibatch size may be determined based … the on-chip memory capacity of a chip … [0029] a given minibatch is split into multiple spatial minibatches) OBVIOUSNESS: Using the same reasoning from claim 15. Accordingly, when Swagath’s memory capacity based minibatch sizing is applied to Nguyen’s subbatch processing of the joint network of Kang as modified by Li, the selected subbatch size correspondingly determines the size / division of the transformed embedding output generated for each subbatch. Claim(s) 18-19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kang (“US 20230386449 A1,” 2023-11-30) in view of Li (“IMPROVING RNN TRANSDUCER MODELING FOR END-TO-END SPEECH RECOGNITION” 2019-09-26) further in view of Nguyen (“MLPerf v1.0 Training Benchmarks: Insights into a Record-Setting NVIDIA Performance,” 2021-07-30) further in view of Swagath (US 20200311536 A1, 2020-10-01) as applied to claim 15 above, further in view of Marisa Kirisame at al. (hereinafter Kirisame) (“DYNAMIC TENSOR REMATERIALIZATION,” 2021-03-18), further in view of Sorokin (“CUDA-Warp RNN-Transducer,” 2021-08-23). Regarding claim 18, Kang and Li fail to explicitly teach but Nguyen teaches; storing slices … slices processed during the training (using the same reasoning as claim 15) OBVIOUSNESS: Using the same reasoning from claim 15. Kang, Li, Nguyen, and Swagath fail to teach but Kirisame teaches; storing [data] until the memory storage maximum capacity is reached and then discarding [additional data processed during the training] ([pg. 2] DTR first checks if sufficient memory is available. If so, it generates a fresh tensor identifier … allocates the requested memory, and returns a new tensor. If not, DTR heuristically selects and evicts resident tensors until the requested allocation can be accommodated) OBVIOUSNESS TO COMBINE KIRISAME: Kirisame is analogous art to the present disclosure because it pertains to reducing memory consumption during neural network training by dynamically managing tensors stored in memory. It would have been obvious to one of ordinary skill in the art, before the effective filing date, to manage the slices of joint network output data stored during the RNN-T training of Kang as modified by Li, Nguyen, and Swagath according to the memory management technique of Kirisame. Nguyen expressly teaches that the RNN-T joint network generates larger tensors and employs batch splitting to avoid exceeding GPU memory capacity, while Kirisame similarly addresses limited memory during neural network training and teaches retaining tensors while sufficient memory is available and, when additional allocation cannot be accommodated, evicting resident tensor data until sufficient memory is made available. A person of ordinary skill seeking to further control the memory required by Nguyen’s stored joint network slices therefore would have had reason to apply Kirisame’s memory capacity dependent tensor eviction technique, such that data is retained while memory capacity permits, and discarded when additional storage exceeds available memory. The modification would have predictably allowed the RNN-T training computation to proceed within the available memory capacity while reducing peak memory consumption and avoiding an out of memory condition ([Kirisame pg. 4] DTR enables training under restricted memory budgets and closely matches the performance of an optimal baseline). Kang, Li, Nguyen, Swagath, and Kirisame fail to teach but Sorokin teaches; … discarding (discarding can simply mean not storing, as stated by the applicant specification: [0081] discarding or otherwise refraining from storing other prediction outputs) probability data … other than probability data for a next blank output and a next token for each lattice node ([pg. 3] memory-efficient version that expects log_probs with the shape (N, T, U, 2) only for blank and labels values … trainable joint network with an output (N, T, U, V)) OBVIOUSNESS: Sorokin is analogous art to the present disclosure as it pertains to an RNN-T training framework. It would have been obvious to one of ordinary skill in the art, before the effective filing date, to configure the memory constrained RNN-T training of Kang as modified by Li, Nguyen, Swagath, and Kirisame such that, when probability data associated with additional transformed embedding slices is not retained due to the memory capacity limitation taught by Kirisame, the probability data retained for each lattice node is limited to the blank and corresponding label values as taught by Sorokin. Sorokin expressly teaches a memory efficient version of the RNN-T loss that uses probability data having shape (N, T, U, 2), containing only blank and label values, instead of the full joint network output having shape (N, T, U, V), and further teaches that this arrangement provides excellent performance for a large vocabulary. A person of ordinary skill seeking to reduce the amount of RNN-T probability data remaining resident after the memory capacity is reached therefore would have had reason to retain only the blank and next label probabilities required by the memory efficient RNN-T loss processing of Sorokin, rather than retaining the remaining vocabulary probabilities. The modification would predictably reduce the amount of probability data that must be stored while preserving the blank and next-token probability information required for the RNN-T lattice loss computation. Accordingly, the combination teaches; storing slices (Nguyen) until the memory storage maximum capacity is reached and then discarding (Kirisame) probability data (Sorokin) for additional slices (Nguyen) other than probability data for a next blank output and a next token for each lattice node (Sorokin) in additional slices processed during the training (Nguyen). Regarding claim 19, Claim 19 is substantially similar to claim 18 and is rejected using the same reasoning. Claim(s) 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Stefan Braun et al. (hereinafter Braun) (“NEURAL TRANSDUCER TRAINING: REDUCED MEMORY CONSUMPTION WITH SAMPLE-WISE COMPUTATION,” 2023-03-13) in view of Kang (“US 20230386449 A1,” 2023-11-30). Regarding claim 20, Braun teaches; A method for training a neural transducer ([Abstract] neural transducer … We propose a memory-efficient training method), the method comprising: obtaining a training batch of utterances ([pg. 3] When generating training batches … samples of similar acoustic sequence length are batched); removing padding used for one or more utterances of the training batch ([pg. 3] sample-wise computation allows us to remove all padding); generating joint embeddings ([pg. 2] The joint network fJ then combines the acoustic encodings and label encodings into the joint encodings z) and corresponding transformed embeddings ([pg. 2] The output layer fO projects the joint encodings to output scores h) of the utterances (acoustic sample x, see algorithm 1 below); computing a loss for the neural transducer based on the transformed embeddings (see line 7 of algorithm 1 below); backpropagating the loss through the neural transducer (see lines 10-11 of algorithm 1 below) [pg. 3] PNG media_image5.png 708 708 media_image5.png Greyscale Bruan fails to teach but Kang teaches; update the parameters of the neural transducer to minimize the loss ([0054] The adjustment of the network parameters can be stopped when the first network loss value is within a preset error range); and generating a modified neural transducer (the resulting trained neural transducer … [0027] a process of training the RNN-T model) OBVIOUSNESS: Braun and Kang both pertain to training RNN-T models using a calculated loss. It would have been obvious to one of ordinary skill in the art, before the effective filing date, to use Kang’s loss-based parameter adjustment with Bruan’s backpropagated parameter gradients to update the neural transducer parameters and reduce the loss. The modification would predictably result in a trained RNN-T having modified parameters and reduced loss. CONCLUSION Any inquiry concerning this communication or earlier communications from the examiner should be directed to Matthew Alan Cady whose telephone number is (571) 272-7229. The examiner can normally be reached Monday - Friday, 7:30 am - 5:00 pm ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached on (571)272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MATTHEW ALAN CADY/ Examiner, Art Unit 2145 /CESAR B PAULA/ Supervisory Patent Examiner, Art Unit 2145
Read full office action

Prosecution Timeline

Jun 12, 2024
Application Filed
Sep 11, 2026
Non-Final Rejection mailed — §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
0%
Grant Probability
0%
With Interview (+0.0%)
3y 4m (~1y 0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 1 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month