Prosecution Insights
Last updated: October 01, 2026
Application No. 18/526,148

Modular Training for Flexible Attention Based End-to-End ASR

Final Rejection §102§103
Filed
Dec 01, 2023
Priority
Dec 02, 2022 — provisional 63/385,959
Examiner
JONES, CHARLES JEFFREY
Art Unit
2122
Tech Center
2100 — Computer Architecture & Software
Assignee
Google LLC
OA Round
2 (Final)
26%
Grant Probability
At Risk
3-4
OA Rounds
1y 2m
Est. Remaining
63%
With Interview

Examiner Intelligence

Grants only 26% of cases
26%
Career Allowance Rate
6 granted / 23 resolved
-28.9% vs TC avg
Strong +37% interview lift
Without
With
+36.7%
Interview Lift
resolved cases with interview
Typical timeline
4y 0m
Avg Prosecution
22 currently pending
Career history
49
Total Applications
across all art units

Statute-Specific Performance

§101
30.5%
-9.5% vs TC avg
§103
38.7%
-1.3% vs TC avg
§102
15.6%
-24.4% vs TC avg
§112
14.9%
-25.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 23 resolved cases

Office Action

§102 §103
DETAILED ACTION This action is responsive to the amendment filed on 06/16/2026 for application 18/526,148. Claims 1-24 are pending in the case. Claims 1 and 13 are independent claims. Claims 1-2 and 13-14 are amended. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Regarding claims 1, 3-5, 12-13, 15-17 and 24: Claim(s) 1, 3-5, 12-13, 15-17 and 24 is/are rejected under 35 U.S.C. 102(a) as being anticipated by Houlsby et al.(“Parameter-Efficient Transfer Learning for NLP”, henceforth known as Houlsby) in view of Woo et al.(“ CBAM: Convolutional Block Attention Module”, henceforth known as Woo) Regarding claim 1: Houlsby discloses a computer-implemented method for training a modular neural network model, the computer-implemented method executed on data processing hardware that causes the data processing hardware to perform operations (Houlsby, Page 3, Col. 2, Paragraph 4,“All runs are trained on 4 Google Cloud TPUs with a batch size of 32” where running the model on a TPU corresponds to using a computer-implemented method for training a modular neural network mode and being executed on data processing hardware that causes the data processing hardware to perform operations) Houlsby discloses during an initial training stage training only a backbone model…comprising existing residual connections to provide a first model configuration of the modular neural network model(Houlsby, Page 3, Col. 1, Figure 2 and Paragraph 2, “…A skip-connection is applied across each of the sub-layers” where the left having feed-forward layers correspond to a backbone model with skip-connections corresponding to residual connections), the first model configuration comprising only the trained backbone model(Houlsby, Page 2, Col. 1, Paragraph 2, “Adapters are new modules added between layers of a pre-trained network.” where pre-trained network corresponds to a first model configuration where the backbone is trained without an adapter module is not added to the model yet and training only a backbone model and the adaptor module corresponds to an intrinsic sub-model being added to a backbone model)…adding an intrinsic sub-model…to the trained backbone model (Houlsby, Page 3, Col. 2, Figure 2 and Paragraph 2, “A skip-connection is applied across each of the sub-layers. The output of each sub-layer is fed into layer normalization. We insert two serial adapters after each of these sub-layers. The adapter is always applied directly to the output of the sub-layer, after the projection back to the input size, but before adding the skip connection back” where the pre-trained model containing residual connections in each sublayer and after insertion is the same number of residual connections corresponds to adding the intrinsic sub-model to the trained backbone model without requiring additional residual connection other than the existing residual connections or requiring a residual adaptor) Houlsby does not teach, however Woo discloses : during an initial training stage training only a backbone model comprising a non-attentive neural network adding an intrinsic sub-model comprising an attention-based sub-model to the trained backbone model without requiring any residual adaptors or additional residual connection other than the existing residual connections Woo discloses during an initial training stage training only a backbone model comprising a non-attentive neural network(Woo, Page 13, Paragraph 2, “We adopt Faster-RCNN [41] as our detection method and ImageNet pre-trained ResNet50 and ResNet101 [5] as our baseline network” where ResBlock/ResNet correspond to a backbone model) Woo discloses adding an intrinsic sub-model(Woo, Page 2, Paragraph 3, “In the ImageNet-1K dataset, we obtain accuracy improvement from various baseline networks by plugging our tiny module” where “plugging” in the module of CBAM corresponds to adding an intrinsic model) comprising an attention-based sub-model to the trained backbone model(Woo, Page 3, Paragraph 3, “The separate attention generation process…can be used as a plug-and-play module for pre-existing base CNN architectures… In our CBAM, we exploit both spatial and channel-wise attention based on an efficient architecture and empirically verify that exploiting both is superior to using only the channel-wise attention” where CBAM being added to a CNN architecture corresponds to adding an intrinsic sub-model comprising an attention-based sub-model to a trained backbone) without requiring any residual adaptors(Woo, Page 4, Paragraph 2, “The overall attention process can be summarized as F`=Mc(F)⊗F, F`` =Ms(F)⊗F,” where ⊗ denotes element-wise multiplication and is performed without a residual adaptor) or additional residual connection other than the existing residual connections(Woo, Page 6, Figure 3, where the long curved line along the bottom is the ResBlock’s existing identity/skip connection and the addition of CBAM does not increase the number of additional residual connection) References Houlsby and Woo are analogous art because they are from the field of endeavor of modifying an existing deep neural network architecture with a relatively small module. Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Houlsby and Woo before him or her, to modify the module of Houlsby to include the CBAM of Woo to emphasize useful features and suppressing irrelevant features. The suggestion/motivation for doing so would have been Woo, Page 14, Page 3, “Our final module (CBAM) learns what and where to emphasize or suppress and refines intermediate features effectively” Regarding claim 3: The rejection of claim 1 with prior art Houlsby-Woo is incorporated and further: Houlsby further discloses wherein the operations further comprise, after fine-tuning parameters of the intrinsic sub-model: removing the intrinsic sub-model(Houlsby, Page 2, Col. 2, Paragraph 3, “The adapter modules may also be ignored if not required” where each adapter for each task may be ignored if not required corresponds to removing the intrinsic sub-model as it’s no longer being considered in the network (See also Houlsby, Page 6, Col. 2, Paragraph 2, “For this, we remove some trained adapters and re-evaluate the model (without re-training)” where the removal of the adapters after training can also be considered removing an intrinsic sub-model after fine-tuning parameters) Houlsby further discloses adding another intrinsic sub-model(Houlsby, Page 1, Col. 1, Abstract, “Adapter modules yield a compact and extensible model; they add only a few trainable parameters per task, and new tasks can be added without revisiting' previous ones” where adding adapter modules for new tasks corresponds with disclosing adding another intrinsic sub-model as for each new task there is a new adapter) to the trained backbone model(Houlsby, Page 2, Col. 1, Paragraph 2, “Adapters are new modules added between layers of a pre-trained network.”) Houlsby further discloses during another fine-tuning training stage: freezing parameters of the trained backbone model(Houlsby, Page 2, Col. 2, Paragraph 2, “The weights of the original network are untouched…the parameters of the original network are frozen and therefore may be shared by many tasks.”) and fine-tuning parameters of the other intrinsic sub-model added to the trained backbone model while the parameters of the trained backbone model are frozen to provide a third model configuration, the third model configuration comprising the backbone model initially trained during the initial training stage and the other intrinsic sub-model having the parameters fine-tuned during the other fine-tuning stage. (Houlsby, Page 2, Col. 1, Paragraph 2, “ Only the new, task specific, parameters, v, are then trained…During training, only v are tuned” where only training the adapters task specific parameters v corresponds creating a third model configuration where the fine-tuning parameters of the intrinsic sub model while the pre-trained model parameters are frozen(See also Houlsby, Page 2, Col. 1, Paragraph 3, “…Adapters differ in that the tasks do not interact and the shared parameters are frozen” that further shows that each task has a corresponding adapter) Regarding claim 4: The rejection of claim 3 with prior art Houlsby-Woo is incorporated and further: Houlsby discloses during the fine-tuning training stage, the parameters of the intrinsic sub-model are trained on a first domain and/or first application(Houlsby, Page 1, Col. 1, Abstract, “Adapter modules yield a compact and extensible model; they add only a few trainable parameters per task, and new tasks can be added without revisiting previous ones” where each adapter is tied to a separate task which is considered a domain and having a task trained on a task corresponds to the intrinsic-sub-model being trained on a first domain(See also, Houlsby, Page 2, Col. 1, Paragraph 2, “…Only the new, task specific, parameters, v, are then trained.”)) Houlsby discloses during the other fine-tuning training stage, the parameters of the other intrinsic sub-model are trained on a second domain different than the first domain and/or a second application different than the first application(Houlsby, Page 1, Col. 1, Abstract, “Adapter modules yield a compact and extensible model; they add only a few trainable parameters per task, and new tasks can be added without revisiting previous ones …To demonstrate adapter’s effectiveness, we transfer the recently proposed BERT Transformer model to 26 diverse text classification tasks” where each task using an adapter and 26 tasks being defined as diverse with Table 1 and Table 2 showing the tasks being different corresponds to having other intrinsic sub-model are trained on a second domain different than the first domain) Regarding claim 5: The rejection of claim 4 with prior art Houlsby-Woo is incorporated and further: Houlsby discloses wherein the trained backbone model is domain-independent(Houlsby, Page 3, Col. 2, Paragraph 3, “We use the public, pre-trained BERT Transformer network as our base model.” where the pre-trained network used with adapters being public pre-trained network corresponds to being domain-independent(See also Houlsby, Page 1, Col. 1, Paragraph 1, “BERT, a Transformer network trained on large text corpora with an unsupervised loss”)) Regarding claim 12: The rejection of claim 1 with prior art Houlsby-Woo is incorporated and further: Houlsby discloses wherein during inference, the trained modular neural network model is configured to operate in any one of: the second model configuration comprising the backbone model initially trained during the initial training stage and the intrinsic sub-model having the parameters fine- tuned during the fine-tuning stage(Houlsby, Page 2, Col. 1, Paragraph 2, “For adapter tuning, a new function, ψw,v(x), is defined, where parameters w are copied over from pre-training. …During training, only v are tuned” where ψw,v(x) represents the output of running input x through the model that includes the pretrained parameters(w) and the adapter(v) corresponds to inferring , using the being the backbone model initially trained and the intrinsic sub-model having parameters fine-tuned(See also Houlsby, Page 2, Col. 1, Paragraph 4, “We demonstrate on a large and diverse set of text classification tasks that adapters yield parameter-efficient tuning for NLP.”)) The following limitations were not mapped as the claim language requires only one or more of the first, second or third model configurations be shown in prior art and second model configuration is shown in prior art as cited above: the first model configuration comprising only the trained backbone model and having the intrinsic sub-model removed; or a third model configuration comprising only the intrinsic sub-model having the parameters fine-tuned during the fine-tuning stage and the trained backbone model removed Regarding claim 13: Houlsby discloses data processing hardware; and memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations(Houlsby, Page 3, Col. 2, Paragraph 4,“All runs are trained on 4 Google Cloud TPUs with a batch size of 32” where running the model on a TPU corresponds to using data processing hardware and memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations) Houlsby discloses during an initial training stage training only a backbone model…comprising existing residual connections to provide a first model configuration of the modular neural network model(Houlsby, Page 3, Col. 1, Figure 2 and Paragraph 2, “…A skip-connection is applied across each of the sub-layers” where the left having feed-forward layers correspond to a backbone model with skip-connections corresponding to residual connections), the first model configuration comprising only the trained backbone model(Houlsby, Page 2, Col. 1, Paragraph 2, “Adapters are new modules added between layers of a pre-trained network.” where pre-trained network corresponds to a first model configuration where the backbone is trained without an adapter module is not added to the model yet and training only a backbone model and the adaptor module corresponds to an intrinsic sub-model being added to a backbone model)…adding an intrinsic sub-model…to the trained backbone model (Houlsby, Page 3, Col. 2, Figure 2 and Paragraph 2, “A skip-connection is applied across each of the sub-layers. The output of each sub-layer is fed into layer normalization. We insert two serial adapters after each of these sub-layers. The adapter is always applied directly to the output of the sub-layer, after the projection back to the input size, but before adding the skip connection back” where the pre-trained model containing residual connections in each sublayer and after insertion is the same number of residual connections corresponds to adding the intrinsic sub-model to the trained backbone model without requiring additional residual connection other than the existing residual connections or requiring a residual adaptor) Houlsby does not teach, however Woo discloses : during an initial training stage training only a backbone model comprising a non-attentive neural network adding an intrinsic sub-model comprising an attention-based sub-model to the trained backbone model without requiring any residual adaptors or additional residual connection other than the existing residual connections Woo discloses during an initial training stage training only a backbone model comprising a non-attentive neural network(Woo, Page 13, Paragraph 2, “We adopt Faster-RCNN [41] as our detection method and ImageNet pre-trained ResNet50 and ResNet101 [5] as our baseline network” where ResBlock/ResNet correspond to a backbone model) Woo discloses adding an intrinsic sub-model(Woo, Page 2, Paragraph 3, “In the ImageNet-1K dataset, we obtain accuracy improvement from various baseline networks by plugging our tiny module” where “plugging” in the module of CBAM corresponds to adding an intrinsic model) comprising an attention-based sub-model to the trained backbone model(Woo, Page 3, Paragraph 3, “The separate attention generation process…can be used as a plug-and-play module for pre-existing base CNN architectures… In our CBAM, we exploit both spatial and channel-wise attention based on an efficient architecture and empirically verify that exploiting both is superior to using only the channel-wise attention” where CBAM being added to a CNN architecture corresponds to adding an intrinsic sub-model comprising an attention-based sub-model to a trained backbone) without requiring any residual adaptors(Woo, Page 4, Paragraph 2, “The overall attention process can be summarized as F`=Mc(F)⊗F, F`` =Ms(F)⊗F,” where ⊗ denotes element-wise multiplication and is performed without a residual adaptor) or additional residual connection other than the existing residual connections(Woo, Page 6, Figure 3, where the long curved line along the bottom is the ResBlock’s existing identity/skip connection and the addition of CBAM does not increase the number of additional residual connection) References Houlsby and Woo are analogous art because they are from the field of endeavor of modifying an existing deep neural network architecture with a relatively small module. Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Houlsby and Woo before him or her, to modify the module of Houlsby to include the CBAM of Woo to emphasize useful features and suppressing irrelevant features. The suggestion/motivation for doing so would have been Woo, Page 14, Page 3, “Our final module (CBAM) learns what and where to emphasize or suppress and refines intermediate features effectively” Regarding claim 15: The rejection of claim 13 incorporated in claim 15. Claim 15 is rejected under the same rationale as set forth in the rejection of claim 3. Regarding claim 16: The rejection of claim 15 incorporated in claim 16. Claim 16 is rejected under the same rationale as set forth in the rejection of claim 4. Regarding claim 17: The rejection of claim 16 incorporated in claim 17. Claim 17 is rejected under the same rationale as set forth in the rejection of claim 5. Regarding claim 24: The rejection of claim 13 incorporated in claim 24. Claim 24 is rejected under the same rationale as set forth in the rejection of claim 12. Regarding claims 2 and 14: Claim(s) 2 and 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Houlsby et al.(“Parameter-Efficient Transfer Learning for NLP”, henceforth known as Houlsby) in view of Woo et al.(“ CBAM: Convolutional Block Attention Module”, henceforth known as Woo) and further in view of Gulati et al.(“Conformer: Convolution-augmented Transformer for Speech Recognition”, henceforth known as Gulati). Regarding claim 2: The rejection of claim 1 with prior art Houlsby-Woo is incorporated and further Houlsby-Woo does not disclose, however Gulati discloses the attention-based sub-model comprises a stack of one or more multi-head self-attention layers disposed between layers of the backbone model(Gulati, Figure 1, where Figure 1 showing a Multi-Head Self Attention Module in between layers corresponds to an attention-based sub-model comprises a stack of one or more multi-head self-attention layers disposed between layers of the backbone model as the modules are considered layers and the Multi-Head Self Attention Module is between modules ) References Houlsby-Woo and Gulati are analogous art because they are from the same field of endeavor of using Transformers and neural network architecture for deep learning Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Houlsby-Woo and Gulati before him or her, to modify model of Houlsby-Woo with the multi-head self-atatention latyers of Gulati to enhance prediction results. The suggestion/motivation for doing so would have been Gulati, Page 4, Col. 2, Paragraph 1(“In self-attention, each attention head learns to focus on different parts of the input, making it possible to improve predictions beyond the simple weighted average. We perform experiments to study the effect of varying the number of attention heads from 4 to 32 in our large model, using the same number of heads in all layers. We find that increasing attention heads up to 16 improves the accuracy, especially over the devother datasets,”). Regarding claim 14: The rejection of claim 13 incorporated in claim 14. Claim 14 is rejected under the same rationale as set forth in the rejection of claim 2. Regarding claims 6 and 18: Claim(s) 6 and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Houlsby et al.(“Parameter-Efficient Transfer Learning for NLP”, henceforth known as Houlsby) in view of Woo et al.(“ CBAM: Convolutional Block Attention Module”, henceforth known as Woo) with Kannan et al.(“ Large-Scale Multilingual Speech Recognition with a Streaming End-to-End Model”, henceforth known as Kannan) Regarding claim 6: The rejection of claim 4 with prior art Houlsby-Woo is incorporated and further: Houlsby-Woo doesn’t teach wherein the first domain is associated with speech recognition in a first language and the second domain is associated with speech recognition in a second language different than the first language. Kannan discloses wherein the first domain is associated with speech recognition in a first language(Kannan, Page 2, Col. 2, Paragraphs 4-5,“ Here we extend it to multilingual speech recognition…Adapter modules are effectively domain-specific (language-specific in our case) adjustments to the activations coming out of each layer.”) and the second domain is associated with speech recognition in a second language different than the first language(Kannan, Page 3, Col.1 , Table 1 and Paragraph 1, “Importantly, each adapter module contains separate parameters for each language” where each adapter being for a separate language corresponds to a second domain with speech recognition in a second language different than the first language (See also Kannan, Page 2, Col. 2, Figure 1, “For a Tamil utterance, only the Tamil adapters are applied to each activation” which emphases each adapter is focused on the domain of a specific language)) References Houlsby-Woo and Kannan are analogous art because they are from the same field of endeavor of using deep learning with adapters as a techniques and modular architectures with the goal of parameter efficiency. Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Houlsby-Woo and Kannan before him or her, to modify the domain of the adapters of Houlsby-Woo to include the language-specific domains for adapters of Kannan because multilingual models tend to be biased toward languages with more data and the use of adapters with language specific adapters to capture the difference of language. The suggestion/motivation for doing so would have been Kannan, Page 2, Col. 1, Paragraph 5(“Imbalanced data typically leads to having a model perform better on languages with larger data”), Kannan, Page 2, Col. 2, Paragraph 3(“it is typical to have varying amounts of transcribed data available for different languages. As a result, a multilingual model will be more influenced by languages which are over-represented in the training set”), Kannan, Page 2, Col. 2, Paragraph 4(“A second architecture extension that we investigate to handle data imbalance is adapter modules.”) and Kannan having Houlsby cited as a resource/reference. Regarding claim 18: The rejection of claim 16 incorporated in claim 18. Claim 18 is rejected under the same rationale as set forth in the rejection of claim 6. Regarding claims 7, 9, 19 and 21: Claim(s) 7, 9, 19 and 21 is/are rejected under 35 U.S.C. 103 as being unpatentable over Houlsby et al.(“Parameter-Efficient Transfer Learning for NLP”, henceforth known as Houlsby) in view of Woo et al.(“ CBAM: Convolutional Block Attention Module”, henceforth known as Woo) with Eeckt et al.(“USING ADAPTERS TO OVERCOME CATASTROPHIC FORGETTING IN END-TO-END AUTOMATIC SPEECH RECOGNITION”, henceforth known as Eeckt) Regarding claim 7: The rejection of claim 1 with prior art Houlsby-Woo is incorporated and further: Houlsby-Woo doesn’t teach the modular neural network model comprises an end-to-end speech recognition model comprising an audio encoder and a decoder; training only the backbone model comprises updating parameters of the audio encoder or the decoder; and fine-tuning the parameters of the intrinsic sub-model comprises updating the parameters of the audio encoder or the decoder. Eeckt discloses the modular neural network model comprises an end-to-end speech recognition model(Eeckt, Page 1, Abstract “The same applies to End-to-End (E2E) Automatic Speech Recognition (ASR) models, even for monolingual tasks. In this paper, we aim to overcome CF for E2E ASR by inserting adapters” and Eeckt, Page 1, Col. 2, Paragraph 3, “X ∈ RF×f the input utterance consisting of F frames of dimension f, and y the ground truth transcription of w word pieces” where the model being an Automatic Speech Recognition model and handling speech corresponds to an end-to-end speech recognition model) comprising an audio encoder and a decoder(Eeckt, Page 1, Col. 2, Paragraph 3, “The model is a hybrid encoder-decoder E2E model, consisting of a Conformer or Transformer encoder and Transformer decoder.”) Eeckt discloses training only the backbone model comprises updating parameters of the audio encoder or the decoder(Eeckt, Page 1, Col. 2, Equation 1 and Paragraph 3, “…We denote L(X,y;θ) the loss of the model with parameters θ ∈ RN (N is the number of parameters))” where the model is defined as a encoder and decoder with parameter vector θ representing all model parameters and the loss is minimized with respect to θ meaning the gradients are computed over the full model corresponds to training a backbone model updating parameters of the audio encoder and decoder(See also Eeckt, Page 1, Col. 2, Paragraph 2, “…we use adapters…which we insert into the model and make task-specific, meaning that each task uses its own adapters.”)) Eeckt discloses fine-tuning the parameters of the intrinsic sub-model comprises updating the parameters of the audio encoder or the decoder(Eeckt, Page 2, Col. 1, Paragraph 5, “We, too, consider this method, called A/Freeze. In addition, we consider A/CFT (Cautious Fine-Tuning) which consists of two stages: 1) train the adapters of the new task while freezing the shared parameters (as in A/Freeze); 2) adapt the entire model with a ten times smaller learning rate” where updating the entire model using A/CFT after adapters are inserted corresponds fine-tuning the intrinsic sub-model comprising updating parameters of the audio encoder or decode) References Houlsby-Woo and Eeckt are analogous art because they are from the same field of endeavor of using transfer learning with adapters as a techniques and modular architectures with the goal of parameter efficiency. Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Houlsby-Woo and Eeckt before him or her, to modify the encoder-only transformer of Houlsby-Woo with the encoder-decoder transformer model using A/CFT of Eeckt to protect shared parameters while adapting the model for new tasks better with A/CFT and handle sequence generation problems. The suggestion/motivation for doing so would have been Eeckt, Page 2, Col. 1, Paragraph 5(“…thus updating the shared parameters cautiously might improve on the new task while preventing forgetting of the old tasks.”) and Eeckt having Houlsby cited as a resource/reference. Regarding claim 9: The rejection of claim 7 with prior art Houlsby-Woo-Eeckt is incorporated and further: Eeckt further discloses wherein the operations further comprise training another modular neural network, the other modular neural network comprising the other one of the audio encoder or the decoder of the end-to-end speech recognition model(Eeckt, Page 1, Col. 2, Equation 1 and Paragraph 3, “…We denote L(X,y;θ) the loss of the model with parameters θ ∈ RN (N is the number of parameters))” where the model is defined as a encoder and decoder with parameter vector θ representing all model parameters and the loss is minimized with respect to θ meaning the gradients are computed over the full model corresponds to a training a backbone model updating parameters of the audio encoder and the decoder) Regarding claim 19: The rejection of claim 13 incorporated in claim 19. Claim 19 is rejected under the same rationale as set forth in the rejection of claim 7. Regarding claim 21: The rejection of claim 19 incorporated in claim 21. Claim 21 is rejected under the same rationale as set forth in the rejection of claim 9. Regarding claims 8 and 20 Claim(s) 8 and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Houlsby et al.(“Parameter-Efficient Transfer Learning for NLP”, henceforth known as Houlsby) in view of Woo et al.(“ CBAM: Convolutional Block Attention Module”, henceforth known as Woo) with Eeckt et al.(“USING ADAPTERS TO OVERCOME CATASTROPHIC FORGETTING IN END-TO-END AUTOMATIC SPEECH RECOGNITION”, henceforth known as Eeckt) and Kannan et al.(“ Large-Scale Multilingual Speech Recognition with a Streaming End-to-End Model”, henceforth known as Kannan) Regarding claim 8: The rejection of claim 7 with prior art Houlsby-Woo-Eeckt is incorporated and further: Houlsby-Woo-Eeckt does not teach wherein the end-to-end speech recognition model comprises a recurrent neural network-transducer (RNN-T) architecture Kannan discloses wherein the end-to-end speech recognition model comprises a recurrent neural network-transducer (RNN-T) architecture(Kannan, Page 1, Col. 2, Paragraph 2, “First, we present a streaming E2E multilingual system using the Recurrent Neural Network Transducer (RNN-T)”) References Houlsby-Woo-Eeckt and Kannan are analogous art because they are from the same field of endeavor of using deep learning with adapters as a techniques and modular architectures with the goal of parameter efficiency. Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Houlsby-Woo-Eeckt and Kannan before him or her, to modify encoder-decoder transformer model of Houlsby-Eeckt with the encoder-prediction-joint transducer model of Kannan to provide a straightforward streaming implementation. The suggestion/motivation for doing so would have been Kannan, Page 1, Col. 2, Paragraph 2(“prior E2E multilingual work has been limited to…models that do not admit a straightforward streaming implementation”) and Kannan having Houlsby cited as a resource/reference. Regarding claim 20: The rejection of claim 19 incorporated in claim 20. Claim 20 is rejected under the same rationale as set forth in the rejection of claim 8. Regarding claims 10-11 and 22-23: Claim(s) 10-11 and 22-23 is/are rejected under 35 U.S.C. 103 as being unpatentable over Houlsby et al.(“Parameter-Efficient Transfer Learning for NLP”, henceforth known as Houlsby) in view of Woo et al.(“ CBAM: Convolutional Block Attention Module”, henceforth known as Woo) with Gulati et al.(“Conformer: Convolution-augmented Transformer for Speech Recognition”, henceforth known as Gulati) and Zhao et al.(“Tiny-Attention Adapter: Contexts Are More Important Than the Number of Parameters”, henceforth known as Zhao) Regarding claim 10: The rejection of claim 1 with prior art Houlsby-Woo is incorporated and further: Houlsby-Woo does not teach the backbone model comprises a first half feedforward layer; a convolution layer; a second half feedforward layer; and a layer norm layer or the intrinsic sub-model comprises a stack of one or more multi-head self-attention layers. Gulati discloses the backbone model comprises a first half feedforward layer; a convolution layer; a second half feedforward layer; and a layer norm layer(Gulati, Page 1, Figure 1, “Conformer comprises of two macaron-like feed-forward layers with half step residual connections sandwiching the multi-headed self attention and convolution modules. This is followed by a post layernorm” where the convolution-augmented transformer has a corresponding first half feedforward layer, a convolution layer, a second half feedforward layer and a layer norm layer) References Houlsby-Woo and Gulati are analogous art because they are from the same field of endeavor of using Transformers and neural network architecture for deep learning Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Houlsby-Woo and Gulati before him or her, to modify transformer model of Houlsby-Woo with the convolution-augmented transformer model of Gulati to capture local patterns for better accuracy and efficiency. The suggestion/motivation for doing so would have been Gulati, Page 1, Col. 2, Paragraph 2(“…we propose a novel combination of self-attention and convolution…self-attention learns the global interaction whilst the convolutions efficiently capture…local correlations.”) and Gulati, Page 1, Col. 1, Abstract(“In this work, we achieve the best of both worlds…in a parameter-efficient way… Conformer significantly outperforms the previous Transformer and CNN based models achieving state-of-the-art accuracies”). Neither Houlsby-Woo nor Gulati discloses the intrinsic sub-model comprises a stack of one or more multi-head self-attention layers. Zhao discloses the intrinsic sub-model comprises a stack of one or more multi-head self-attention layers(Zhao, Page 2, Col. 2, Paragraph 2-3, “Our adapter has an attentive structure: as shown in Figure2b, at each position t, it takes as input the intermediate embeddings…from not only the current position…but also all the other positions… As shown in Figure2c, the internal architecture of our tiny-attention adapter resembles an ordinary multi-head attention mechanism” where the adapter has each token attend to all tokens in the same sequence corresponds to having a multi-head self-attention layers) References Houlsby-Woo -Gulati and Zhao are analogous art because they are from the same field of endeavor of using deep learning with adapters as a techniques and modular architectures with the goal of parameter efficiency. Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Houlsby-Woo-Gulati and Zhao before him or her, to modify adapter of Houlsby-Woo-Gulati with the tiny-attention adapter of Zhao to provide better context modeling. The suggestion/motivation for doing so would have been Zhao, Page 1, Col. 1, Abstract(“Our tiny attention adapter learns to modify the hidden states at each position directly conditioned on the hidden states at all the other positions, which is missed by the previously proposed adapters.”), Zhao, Page 1, Col. 2, Paragraphs 1-2, “Almost all the previously proposed adapter architectures are feed-forward neural networks. Thus, we suspect that their embedding modifications are not as contextually rich as they should. Therefore, we propose to use the attentive structure that allows the embedding modifications of each token to capture more contextual information” and Zhao Kannan having Houlsby cited as a resource/reference. Regarding claim 11: The rejection of claim 10 with prior art Houlsby-Woo-Gulati-Zhao combination is incorporated and further: Gulati further discloses wherein the second model configuration comprises: the first half feedforward layer; the stack of one or more multi-head self-attention layers; the convolution layer; the second half feedforward layer; and the layer norm layer(Gulati, Page 1, Figure 1, “Conformer comprises of two macaron-like feed-forward layers with half step residual connections sandwiching the multi-headed self attention and convolution modules. This is followed by a post layernorm” where the convolution-augmented transformer has a corresponding first half feedforward layer, a stack of one or more multi-head self-attention layers, a convolution layer, a second half feedforward layer and a layer norm layer) Regarding claim 22: The rejection of claim 13 incorporated in claim 22. Claim 22 is rejected under the same rationale as set forth in the rejection of claim 10. Regarding claim 23: The rejection of claim 22 incorporated in claim 23. Claim 23 is rejected under the same rationale as set forth in the rejection of claim 11. Relevant Art While not used in the current rejection the following was found to be relevant prior art: US20220207730A1, Systems and Methods for Automated Image Analysis, Arnold et al, as the publication discusses a strong similarity to architecture claimed. US20220309315A1, Extension of existing neural networks without affecting existing outputs, Aladahalli et al, as the publication discusses adding an augmented layer without residual connections or residual adaptors. US20200104706A1, Parameter-Efficient Multi-Task and Transfer Learning, Sandler et al, as the publication discusses adding an augmented layer without residual connections or residual adaptors. Attention Augmented Convolutional Networks, Bello et al, as the publication discusses adding an augmented layer without residual connections or residual adaptors. Bottleneck Transformers for Visual Recognition, Srinivas et al, as the publication discusses a strong similarity to architecture claimed. Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned, Voita et al, as the publication discusses Multi-Head Self-Attention. On the Relationship between Self-Attention and Convolutional Layers, Cordonnier et al, as the publication discusses Multi-Head Self-Attention with Convolutional. TReC: Transferred ResNet and CBAM for Detecting Brain Diseases, Xiao et al, as the publication discusses adding an augmented layer without residual connections or residual adaptors. CONTEXTUAL ADAPTERS FOR PERSONALIZED SPEECH RECOGNITION IN NEURAL TRANSDUCERS, Sathyendra et al, as the publication discusses using adapters as interchangeable sub-models. Response to Arguments Applicant's arguments filed 1-24 have been fully considered but they are not persuasive. The scope of the claims have been changed and a breakdown of the arguments can be found below: 102/103: Applicant appears to argue on pages 9-13 that Houlsby does not teach the amended language or scope of claims. Specifically, Houlsby teaching the backbone model comprising a non-attentive neural network without adding or requiring addition residual connections and with Sathyendra to disclose attention based biasing. Applicant’s arguments with respect to amended claims have been considered but are moot because the new ground of rejection does not rely on the mapping reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. as the amended language in claim 1 by new prior art Woo with the amended claim 2 now being mapped to prior art Gulati. Additionally, Applicant's arguments with respect to the dependent claims rely upon features argued to the independent claim and are thus similarly unpersuasive as they are moot in view of the of the new mapping in the current rejection of record. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to CHARLES JEFFREY JONES JR whose telephone number is (703)756-1414. The examiner can normally be reached Monday - Friday 8:00 - 5:00 EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kakali Chaki can be reached at 571-272-3719. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /C.J.J./Examiner, Art Unit 2122 /KAKALI CHAKI/Supervisory Patent Examiner, Art Unit 2122
Read full office action

Prosecution Timeline

Dec 01, 2023
Application Filed
Jul 03, 2025
Response after Non-Final Action
Apr 06, 2026
Non-Final Rejection mailed — §102, §103
Jun 16, 2026
Response Filed
Sep 04, 2026
Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12725007
ADAPTIVELY COMPRESSING A DEEP LEARNING MODEL
4y 11m to grant Granted Sep 01, 2026
Patent 12718061
NEURAL NETWORK MODEL AND LEARNING METHOD OF THE SAME
4y 2m to grant Granted Aug 25, 2026
Patent 12718091
LARGE KERNEL CONVOLUTIONAL NEURAL NETWORK
3y 11m to grant Granted Aug 25, 2026
Patent 12645930
APPARATUS AND METHOD FOR TRAINING LOW BIT-PRECISION DEEP NEURAL NETWORK
5y 2m to grant Granted Jun 02, 2026
Patent 12582959
DATA GENERATION DEVICE AND METHOD, AND LEARNING DEVICE AND METHOD
4y 7m to grant Granted Mar 24, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
26%
Grant Probability
63%
With Interview (+36.7%)
4y 0m (~1y 2m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 23 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month