Prosecution Insights
Last updated: October 01, 2026
Application No. 19/049,425

Using Synthetic Data to Improve Word Error Rate of Differentially Private ASR Models

Non-Final OA §103§112
Filed
Feb 10, 2025
Priority
Feb 29, 2024 — provisional 63/559,289
Examiner
AGAHI, DARIOUSH
Art Unit
Tech Center
Assignee
Google LLC
OA Round
1 (Non-Final)
84%
Grant Probability
Favorable
1-2
OA Rounds
11m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 84% — above average
84%
Career Allowance Rate
154 granted / 184 resolved
+23.7% vs TC avg
Strong +30% interview lift
Without
With
+30.1%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
20 currently pending
Career history
209
Total Applications
across all art units

Statute-Specific Performance

§101
23.7%
-16.3% vs TC avg
§103
55.0%
+15.0% vs TC avg
§102
10.7%
-29.3% vs TC avg
§112
6.7%
-33.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 184 resolved cases

Office Action

§103 §112
DETAILED ACTION This office action is in response to Applicant’s submission filed on 2/10/2025. Claims 1-20 are pending in the application of which Claims 1, and 11 are independent and have been examined. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Applicant’s claim for the benefit of a prior-filed application under 35 U.S.C. 119(e) or under 35 U.S.C. 120, 121, or 365 is acknowledged. The prior-filed application (Provisional application No. 63/559289 Filed on 2/29/2024) is acknowledged. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 6, 10, 16, and 20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claims 6, and 16, recites “… wherein the stack of conformer layers comprises a stack of 24 layers having about 600 million parameters.”, which appears to be indefinite since the scope of quantity is vague and unspecified. Claims 10, and 20, recites “… defines a maximum acceptable amount of information about individual training samples of the plurality of private training samples that may be revealed or leaked by the ASR model.”, which appears to be indefinite since the claim boundaries are ambiguous. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1-2, 4-5, 10 - 12, 14-15, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (US2023/017892A1)(herein "Chen"), and in further view of Greene et al. (US9563704B1)(herein "Greene"), Yang et al. (US20200218970A1)(herein “Yang”), and Pin-Yu Chen (US20240256894A1)(herein “Chen2”). Regarding claims 1, and 11 Chen teaches [A computer-implemented method executed on data processing hardware that causes the data processing hardware to perform operations comprising:- claim 1], and [A system comprising: data processing hardware; and memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising: - claim 11] (Chen, Par. 0004:” One aspect of the disclosure provides a computer-implemented method that when executed on data processing hardware causes the data processing hardware to perform operations for pre-training and audio encoder to jointly learn shared representations of speech and text.”, and Par. 0010:”… provides a system that includes data processing hardware and memory hardware storing instructions that when executed on the data processing hardware causes the data processing hardware to perform operations.”, and Par. 0062:”The method 500 may execute on data processing hardware 610 (FIG. 6) using instructions stored on memory hardware 620 …”). obtaining a public training utterance set; (Chen, Par. 0036:” FIGS. 3A and 3B illustrate an example training process 300 for pre-training the audio encoder 210 of the ASR model 200 (FIG. 2). The training process 300 may pre-train the audio encoder 210 using available training data [public data] that includes a set of unspoken textual utterances (Xtext) 320 and a set of un-transcribed non-synthetic speech utterances (Xunsup) 306.”) pre-training an audio encoder on the public training utterance set; (Chen, Par. 0036:” FIGS. 3A and 3B illustrate an example training process 300 for pre-training the audio encoder 210 of the ASR model 200 (FIG. 2). The training process 300 may pre-train the audio encoder 210 using available training data that includes a set of unspoken textual utterances (Xtext) 320 and a set of un-transcribed non-synthetic speech utterances (Xunsup) 306. “, and Par. 0045:” … the pre-trained audio encoder 210 converges on the un-transcribed non-synthetic speech utterances 306, …”). obtaining a plurality of private training samples, each private training sample comprising a corresponding non-synthetic speech utterance paired with a corresponding transcription; (Chen, Par. 0051:” … After pre-training the audio encoder 210, the training process 300 may fine-tune the pre-trained audio encoder on transcribed speech utterances that may include supervised training samples of both synthetic (e.g., synthesized speech) and non-synthetic (e.g., human speech).”) for each corresponding transcript of the predetermined number of transcripts, processing, using a text-to-speech (TTS) system, the corresponding transcript to generate a corresponding synthetic speech utterance, (Chen, Fig. 3A and 3B both teach a TTS, and Par. 0046:”Referring to FIG. 3B, the supervised loss part 300b of the training process 300 is configured to inject lexical information into the audio encoder 210 during pre-training based on supervised loss terms 342 derived from the synthesized speech representations 332 generated by the TTS system 330 for the unspoken textual utterances 320.”) wherein the corresponding transcript and the corresponding synthetic speech utterance form a corresponding synthetic training sample; (Chen, Fig. 3A and 3B both teach a TTS, and Par. 0046:”Referring to FIG. 3B, the supervised loss part 300b of the training process 300 is configured to inject lexical information into the audio encoder 210 during pre-training based on supervised loss terms 342 derived from the synthesized speech representations 332 generated by the TTS system 330 for the unspoken textual utterances 320.”) during a first fine-tuning stage for fine-tuning an automatic speech recognition (ASR) model comprising the pre-trained audio encoder and a decoder, fine-tuning the ASR model on each of the synthetic training samples to teach the ASR model to learn how to predict the transcripts from the corresponding synthetic speech utterances; and (Chen, Par. 0009:” … after pre-training the audio encoder, fine-tuning the pre-trained audio encoder on transcribed speech utterances.”, and Par. 0027:”… implementations are directed toward pre-training an audio encoder of the ASR model on training data that includes un-transcribed non-synthetic speech utterances, unspoken textual utterances for generating corresponding synthetic speech representations, and the transcribed non-synthetic speech utterances to jointly learn speech and text representations, and then fine-tuning (e.g., warm-start training) the pre-trained ASR model using the available transcribed non-synthetic speech utterances.”, and Par. 0036:” FIGS. 3A and 3B illustrate an example training process 300 for pre-training the audio encoder 210 of the ASR model 200 (FIG. 2).”, and Par. 0046:” Referring to FIG. 3B, the supervised loss part 300b of the training process 300 is configured to inject lexical information into the audio encoder 210 during pre-training based on supervised loss terms 342 derived from the synthesized speech representations 332 generated by the TTS system 330 for the unspoken textual utterances 320. Notably, the supervised loss part 300b leverages one or more auxiliary decoders 390 for generating the supervised loss terms 342. The auxiliary decoders 390 may include Connectionist Temporal Classification (CTC) decoders, Listen Attend Spell (LAS) decoders, or RNN-T decoders. … In some examples, the training process 300 applies data augmentation to at least one of the sample utterances of synthetic speech representations …”, and Par. 0045:”… After the pre-trained audio encoder 210 converges on the un-transcribed non-synthetic speech utterances 306, the pre-training procedure is repeated for the synthesized speech representations 332. Thus, the contrastive loss 316 is optimized for both real/human (non-synthetic) and synthetic (TTS audio) features, with … the training process 300 pre-trains the audio encoder 210 on the derived contrastive loss 316 applied on the corresponding encoded features 211 associated with each synthesized speech representation 332 and each un-transcribed non-synthetic speech utterance 306 provided as input to the audio encoder 210. Pre-training the audio encoder 210 may include updating parameters of the audio encoder based on the contrastive losses.”) during a second fine-tuning stage for fine-tuning the ASR model, fine-tuning, [[using a differentially private parameter-efficient-fine-tuning (DP-PEFT) technique]] the ASR model on the plurality of private training samples (Chen, Par. 0051:” … the training process 300 may pre-train the audio encoder 210 using the unpaired data loss function, unpaired, by updating parameters of the audio encoder 210 to effectively teach the audio encoder 210 to learn shared representations between speech and text. After pre-training the audio encoder 210, the training process 300 may fine-tune the pre-trained audio encoder on transcribed speech utterances that may include supervised training samples of both synthetic (e.g., synthesized speech) and non-synthetic (e.g., human speech).”) Chen, does not teach, however, Greene teaches from a corpus of text utterances, sampling a predetermined number of most frequent words that appear in the corpus of text utterances; (Greene, Col. 9, line 49- Col. 10, line 2:”… post and the transcript by determining a degree to which text in the social network post overlaps with words in the transcript. As a more particular example, in some embodiments, process 400 can compute a correlation between a social network post and the transcript based on how frequently words in the text associated with the social network post appear in a set of the most frequent words of the transcript. …The set of the most frequent words of the transcript can be determined using any suitable technique or combination of techniques, for example, the techniques to determine the most frequently used word(s) in a transcript described above in connection with block 406. The set of the most frequent words in the transcript can include any suitable number of words (e.g., one, five, ten, twenty, and/or any other suitable number).”) Greene is considered to be analogous to the claimed invention because it is in the same field of endeavor. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Chen further in view of Greene to obtain from a corpus of text utterances, sampling a predetermined number of most frequent words that appear in the corpus of text utterances. Motivation to do so would provide a clear random standard to see if real text patterns are special or just random chance Chen, as modified above, does not teach, however, Yang teaches randomly generating a predetermined number of transcripts, each transcript comprising a same number of words randomly sampled from the predetermined number of most frequent words; (Yang, Par. 0006:” … A second set of texts [number of transcripts] is created [generated] by inserting or replacing a randomly selected item from the list of keywords [predetermined number of most frequent words] into each of the first set of texts at a randomly chosen location within each of the first set. A third set of texts is created by inserting or replacing a randomly selected item from the list of to-be-excluded into each of the first set of texts at a randomly chosen location within each of the first set.”) Yang is considered to be analogous to the claimed invention because it is in the same field of endeavor. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Chen, as modified above, further in view of Yang to randomly generating a predetermined number of transcripts, each transcript comprising a same number of words randomly sampled from the predetermined number of most frequent words. Motivation to do so would strip out real topic meaning, and provide fair math comparisons due to the same number of words randomly selected for each transcript. Chen, as modified above, does not teach, however, Chen2 teaches [[during a second fine-tuning stage for fine-tuning the ASR model, fine-tuning,]] using a differentially private parameter-efficient-fine-tuning (DP-PEFT) technique [[the ASR model on the plurality of private training samples]] (Chen2, Par. 0028:” … the deep learning neural network can be pre-trained in a non-federated manner, and differentially private federated learning can be subsequently performed by having each client fully fine-tune, via DP-SGD, its own local instance of the pre-trained deep learning neural network. In other words, each client can update all of the trainable internal parameters of its local instance of the deep learning neural network during each iteration of federated learning, and such trainable internal parameters can begin in the first iteration with pre-trained values rather than with randomly initialized values. Because the deep learning neural network can be pre-trained, such full fine-tuning can involve fewer global-updates than can training from scratch, which can be considered as fewer opportunities for the clients' private training data to be leaked or attacked. Thus, full fine-tuning can achieve a better privacy-accuracy tradeoff as compared to training from scratch.”) Note: since fine tuning implies fewer parameters to be updated, it reads on parameter-efficient-fine-tuning. wherein the DP-PEFT technique updates only a subset of newly added or existing parameters of the pre-trained audio encoder. (Chen2, Par. 0028:” … the deep learning neural network can be pre-trained in a non-federated manner, and differentially private federated learning can be subsequently performed by having each client fully fine-tune, via DP-SGD, its own local instance of the pre-trained deep learning neural network. In other words, each client can update all of the trainable internal parameters of its local instance of the deep learning neural network during each iteration of federated learning, and such trainable internal parameters can begin in the first iteration with pre-trained values rather than with randomly initialized values. Because the deep learning neural network can be pre-trained, such full fine-tuning can involve fewer global-updates than can training from scratch, which can be considered as fewer opportunities for the clients' private training data to be leaked or attacked. Thus, full fine-tuning can achieve a better privacy-accuracy tradeoff as compared to training from scratch.”) Note: since fine tuning implies updates only a subset of existing parameters of the pre-trained model. Chen2 is considered to be analogous to the claimed invention because it is in the same field of endeavor. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Chen, as modified above, further in view of Chen2 to use a differentially private parameter-efficient-fine-tuning (DP-PEFT) technique, wherein the DP-PEFT technique updates only a subset of newly added or existing parameters of the pre-trained audio encoder. Motivation to do so would improve upon the privacy-accuracy tradeoff of differentially private federated learning (Pin-Yu Chen, Par. 0030). Regarding claims 2, and 12, Chen as modified above, teaches the computer-implemented method, and the system of claims 1, and 11, respectively. Chen as modified above, further teaches wherein the public training utterance set includes publicly-available utterances comprising a plurality of un-transcribed non-synthetic speech utterances, each un-transcribed non-synthetic speech utterance not paired with a corresponding transcription. (Chen, Par. 0004:” … Each un-transcribed non-synthetic speech utterance is not paired with a corresponding transcription.”, and Par. 0036:” … The training process 300 may pre-train the audio encoder 210 using available training data [public] that includes a set of unspoken textual utterances (Xtext) 320 and a set of un-transcribed non-synthetic speech utterances (Xunsup) 306. Each unspoken training text utterance 320 includes text-only data (i.e., unpaired data) such that each unspoken training text utterance 320 is not paired any corresponding spoken audio representation (speech) of the utterance. Each un-transcribed non-synthetic speech utterance 306 (also referred to as simply “un-transcribed speech utterance 306”) includes audio-only data (i.e., unpaired data) such that the un-transcribed speech utterance 306 is not paired with any corresponding transcription.”) Regarding claims 4, and 14, Chen as modified above, teaches the computer-implemented method, and the system of claims 1, and 11, respectively. Chen, as modified above, further teaches wherein the audio encoder comprises a stack of multi-head attention layers each including a multi-headed self- attention mechanism. (Chen, Par. 0035:” In some examples, the encoder network (i.e., audio encoder) 210 of the RNN-T model 200 includes a stack of self-attention layers/blocks, such as conformer blocks. Here, each conformer block includes a series of multi-headed self attention, depth wise convolution and feed-forward layers.”) Regarding claims 5, and 15, Chen as modified above, teaches the computer-implemented method, and the system of claims 4, and 14, respectively. Chen, as modified above, further teaches wherein the stack of multi-head attention layers comprises a stack of conformer layers. (Chen, Par. 0035:” … the RNN-T model 200 includes a stack of self-attention layers/blocks, such as conformer blocks. Here, each conformer block includes a series of multi-headed self attention, depth wise convolution and feed-forward layers. the prediction network 220 may include a stack of transformer or conformer blocks, or a embedding look-up table in lieu of LSTM layers. Finally, the joint network 230 may also have 640 hidden units.”) Regarding claims 10, and 20, Chen as modified above, teaches the computer-implemented method, and the system of claims 1, and 11, respectively. Chen, as modified above, does not teach, however, Chen2 further teaches wherein fine-tuning, using the DP-PEFT technique, the ASR model on the plurality of private training samples fine-tunes the ASR model according to a differential privacy budget that defines a maximum acceptable amount of information about individual training samples of the plurality of private training samples that may be revealed or leaked by the ASR model. (Chen2, Par. 0048:” … As explained above, existing techniques for performing federated learning (e.g., differentially private federated learning, in particular) implement training from scratch or fine-tuning. ... However, such very many global updates expose federated clients' private training data to progressively higher risk. … Accordingly, fine-tuning can be considered as providing a better privacy-accuracy tradeoff as compared to training from scratch.”) Claims 3, and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Chen, Greene, Yang, and Chen2, and in further view of Chen et al. (US 20230351558 A1)(herein "Chen3"). Regarding claims 3, and 13, Chen as modified above, teaches the computer-implemented method, and the system of claims 2, and 12, respectively. Chen as modified above, further teaches wherein pre-training the audio encoder on the public training utterance set comprises: for each corresponding un-transcribed non-synthetic speech utterance in the public training utterance set: generating, at each of a plurality of output steps, using a random-projection quantizer, a target quantized vector token and a target token index for a corresponding audio feature in a sequence of audio features associated with the corresponding un-transcribed non-synthetic speech utterance, ( Chen, Par. 0043:” … a quantizer 217 receives the encoded features 211, as input, and generates quantized vectors (i.e., target context vectors) 219 as output.”) Note: a random-projection quantizer results in both a quantized vector and a target token index for a corresponding audio feature. after masking a subset of the audio features in the sequence of audio features associated with the corresponding un-transcribed non-synthetic speech utterance, generating, by the audio encoder, contrastive context vectors from corresponding masked audio features; and (Chen, Par. 0043:” … After masking is applied, the linear layer 214 and the Conformer blocks 216 of the context network receives the masked encoded features 211m and outputs corresponding contrastive context vectors 215 from masked encoded features 211m. Moreover, a quantizer 217 receives the encoded features 211, as input, and generates quantized vectors (i.e., target context vectors) 219 as output. Thereafter, a contrastive loss module 315 derives a contrastive loss ( [Image Omitted] w2v) 316 between the contrastive context vectors 215 at the masked positions and the target context vectors 219 …”). deriving a contrastive loss term between the contrastive context vectors at the masked positions and the target token index; and (Chen, Par. 0043:” … a contrastive loss module 315 derives a contrastive loss ( [Image Omitted] w2v) 316 between the contrastive context vectors 215 at the masked positions and the target context vectors 219 …”). pre-training the audio encoder based on the contrastive loss terms derived for each of the un-transcribed non-synthetic speech utterances in the public training utterance set. (Chen, Par. 0045:” … the training process 300 pre-trains the audio encoder 210 on the derived contrastive loss 316 applied on the corresponding encoded features 211 associated with each synthesized speech representation 332 and each un-transcribed non-synthetic speech utterance 306 provided as input to the audio encoder 210. Pre-training the audio encoder 210 may include updating parameters of the audio encoder based on the contrastive losses.”) Chen, as modified above, does not teach, however, Chen3 teaches wherein the target token index maps the corresponding audio feature to the target quantized vector token stored in one or more codebooks; (Chen3, Par. 0076:” … For each feature vector, it is compared to the latent vectors of either the masked or unmasked codebooks, based on the associated indicator mask value. The latent vector that is found to be closest to the feature vector based on the elementwise subtraction operation is the resulting latent vector 430. Further, let ê∈ PNG media_image1.png 30 116 media_image1.png Greyscale be the quantized latent vectors 430 and PNG media_image2.png 38 218 media_image2.png Greyscale be the tokens 422 for feature vectors {circumflex over (f)}, where I({circumflex over (f)}, e, e′, m.sup.↓) represents the function that gets tokens for the first argument (e.g., a function that obtains indices (tokens) of the quantized latent vectors in ê from in the dual codebook 427). Chen3 is considered to be analogous to the claimed invention because it is in the same field of endeavor. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Chen, as modified above, further in view of Chen3 to wherein the target token index maps the corresponding audio feature to the target quantized vector token stored in one or more codebooks. Motivation to do so would provide extreme data compression, seamless language model integration and noise reduction. Claims 6, and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Chen, Greene, Yang, and Chen2, and in further view of Gulati et al. (Conformer: Convolution-augmented Transformer for Speech Recognition, 2020)(herein “Gulati”). Regarding claims 6, and 16, Chen as modified above, teaches the computer-implemented method, and the system of claims 5, and 15, respectively. Chen as modified above, does not teach, however, Gulati teaches wherein the stack of conformer layers comprises a stack of 24 layers having about 600 million parameters. (Gulati, section 3.2:” We identify three models, small, medium and large, with 10M, 30M, and 118M params, respectively, by sweeping different combinations of network depth, model dimensions, number of attention heads and choosing the best performing one within model parameter size constraints. We use a single-LSTM-layer decoder in all our models. Table 1 describes their architecture hyper-parameters.”) Note: Gulati does not specifically teach 600M parameters, however, it would be obvious to one of ordinary skill in the art at the time of the invention to vary the complexity of the architecture as needed because it allows adjustment of the approach per model constraint and is a matter of design choice. (MPEP 2144.04(VI)). Gulati is considered to be analogous to the claimed invention because it is in the same field of endeavor. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Chen, as modified above, further in view of Gulati to wherein the stack of conformer layers comprises a stack of 24 layers having about 600 million parameters. Motivation to do so would exhibit better accuracy and achieves a state-of-the-art performance (Gulati, Conclusion). Claims 7, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Chen, Greene, Yang, and Chen2, and in further view of Liu et al. (US20250156684A1)(herein “Liu”). Regarding claims 7, and 17, Chen as modified above, teaches the computer-implemented method, and the system of claims 4, and 14, respectively. Chen as modified above, does not teach, however, Liu teaches wherein the DP-PEFT technique comprises modifying the audio encoder to incorporate adapters that each incorporate two low-rank projection matrices and one activation layer, (Liu, Par. 0005:” … parameter-efficient fine-tuning (PEFT) techniques may be used in scenarios where computational resources or labeled task-specific data are limited. The PEFT techniques may enable to strike a balance between the knowledge encoded in pre-trained models and adapting to the specifics of a target task/domain with a reduced number of trainable parameters. For example, PEFT techniques only fine-tune a small set of parameters, which may be a subset of the existing parameters of the pre-trained models or a set of newly added parameters, thereby greatly reducing the computational and memory costs. PEFT techniques also allow to store only a small number of model parameters for domain adaptation in addition to the pre-trained model. …”, and Par. 0014:”… an adapter tuning technique based on Low-Rank Adaptation (LoRA) demonstrates performance comparable to full fine-tuning, despite having significantly fewer trainable parameters. In an example, LoRA may use low-rank decomposition matrices to approximate a parameter update of a weight matrix of a dense layer of a pre-trained model. In an example, the adapter tuning technique may update query and value projection matrices of a transformer architecture of a pre-trained model.”, and Par. 0029:” … The operation includes one or a combination of: … a nonlinear activation operation, and variants thereof.”, and Par. 0072:” … Typically, the set of layers of a path may include linear transformation functions, activation functions, and other operations.”, and Par. 0125:” … the architecture of the base model 104 first converts an input or input data into an n-dimensional embedding, which is then fed to an encoder. The layers 410, 412, 414 and 416 may implement encoder and decoder stacked on each other several times. The layers 410, 412, 414 and 416 include mainly feed-forward and multi-head attention layers.”) wherein only the parameters of the adapters are updated during the fine-tuning. (Liu, Par 0011:” … employed for fine-tuning LLMs for domain adaptation. Parameter-efficient fine-tuning refers to a process of adapting a pre-trained neural network model, such as an LLM to a specific task/domain while efficiently managing and updating a limited set of parameters. The use of PEFT for fine-tuning approach is particularly relevant when computational resources or labelled task-specific data are constrained, and it aims to achieve effective fine-tuning with a reduced number of parameters.”, and Par. 0014:” … For example, PEFT techniques only fine tune a small set of parameters, which may be a subset of the existing parameters of the pre-trained models or a set of newly added parameters, thereby greatly reducing the computational and memory costs.”) Liu is considered to be analogous to the claimed invention because it is in the same field of endeavor. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Chen, as modified above, further in view of Liu to wherein the DP-PEFT technique comprises modifying the audio encoder to incorporate adapters that each incorporate two low-rank projection matrices and one activation layer, wherein only the parameters of the adapters are updated during the fine-tuning. Motivation to do so would improve computational efficiency at inference by replacing a pre-trained weight matrix by its low-rank, quantization, or sparse approximation(Liu, Par. 0019). Claims 9, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Chen, Greene, Yang, and Chen2, and in further view of Wichern et al. (US 20250220375 A1)(herein “Wichern”), and Liu. Regarding claims 9, and 19, Chen as modified above, teaches the computer-implemented method, and the system of claims 4, and 14, respectively. Chen, as modified above, does not teach, however, Wichern teaches wherein the DP-PEFT technique comprises Bias-Term Fine-Tuning (BitFit), (Wichern, Par. 0094:” … the subject specific parameters may be the bias vectors for a subset of the network's hidden layers, in which case updating the neural network parameters for given a specific subject involves updating only a subset of bias terms corresponding to that specific subject. Such an approach is typically referred to as BitFit.”) Wichern is considered to be analogous to the claimed invention because it is in the same field of endeavor. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Chen, as modified above, further in view of Wichern to wherein the DP-PEFT technique comprises Bias-Term Fine-Tuning (BitFit). Motivation to do so would drastically reduce the storage needed for optimizer states and gradients during training. Chen, as modified above, does not teach, however, Liu teaches wherein only bias parameters of the stack of multi-head attention layers are updated during the fine-tuning. (Liu, Par. 0032:” … and as mentioned above, fine-tuning (whether full or partial) can involve training a deep learning neural network to handle new training data, by altering at least some trainable internal parameters (e.g., weight matrices, bias values, convolutional kernels) of the deep learning neural network, where such at least some trainable internal parameters have already been trained.”, and Par. 0068:” … As another example, any of such initial layer, one or more hidden layers, or final layer can be dense layers, whose learnable or trainable parameters can be weight matrices or bias values.”) Liu is considered to be analogous to the claimed invention because it is in the same field of endeavor. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Chen, as modified above, further in view of Liu to wherein only bias parameters of the stack of multi-head attention layers are updated during the fine-tuning. Motivation to do so would improve computational efficiency at inference by replacing a pre-trained weight matrix by its low-rank, quantization, or sparse approximation(Liu, Par. 0019). Allowable Subject Matter Claims 8 and 18 are objected to as being dependent upon a rejected base claim, but would be allowable if written in independent form including all of the limitations of the base claim and any intervening claims. Claims 8 and 18 recites:” wherein the DP-PEFT technique comprises modifying the audio encoder to incorporate two low-rank projection matrices parallel to feed-forward layers of the audio encoder, wherein only parameters of the two low-rank projection matrices are updated during the fine-tuning.”, which is allowable over the prior art. The closest teachings to the indicated allowable subject matter are the references that are cited in the current office action. However, none of the prior art of record teach the limitation as stated above specifically the underlined as shown including all supporting limitations thereof. Therefore, claims 8 and 18, are allowable. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. Gulati et al. (Conformer: Convolution-augmented Transformer for Speech Recognition, 2020) teaches in ABS:” Recently Transformer and Convolution neural network (CNN) based models have shown promising results in Automatic Speech Recognition (ASR), outperforming Recurrent neural networks (RNNs). Transformer models are good at capturing content-based global interactions, while CNNs exploit local features effectively. In this work, we achieve the best of both worlds by studying how to combine convolution neural networks and transformers to model both local and global dependencies of an audio sequence in a parameter-efficient way. To this regard, we propose the convolution-augmented transformer for speech recognition, named Conformer. Conformer significantly outperforms the previous Transformer and CNN based models achieving state-of-the-art accuracies. On the widely used LibriSpeech benchmark, our model achieves WER of 2.1 %/4.3% without using a language model and 1.9%/3.9% with an external language model on test/testother. We also observe competitive performance of 2.7%/6.3% with a small model of only 10M parameters.” Examiner's Note: Examiner has cited particular columns and line numbers and/or paragraph numbers in the references applied to the claims above for the convenience of the applicant. Although the specified citations are representative of the teachings of the art and are applied to specific limitations within the individual claim, other passages and figures may apply as well. It is respectfully requested from the applicant in preparing responses, to fully consider the references in entirety as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior art or disclosed by the Examiner. In the case of amending the Claimed invention, Applicant is respectfully requested to indicate the portion(s) of the specification which dictate(s) the structure relied on for proper interpretation and also to verify and ascertain the metes and bounds of the claimed invention. Any inquiry concerning this communication or earlier communications from the examiner should be directed to DARIOUSH AGAHI whose telephone number is (408)918-7689. The examiner can normally be reached Monday - Thursday and alternate Fridays, 7:30-4:30 PT. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bhavesh Mehta can be reached on 571-272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. DARIOUSH AGAHI, P.E. Primary Examiner /DARIOUSH AGAHI/Primary Examiner, Art Unit 2656
Read full office action

Prosecution Timeline

Feb 10, 2025
Application Filed
Aug 26, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12718030
LARGE LANGUAGE MODELS PROVIDING EVIDENCE MAPPINGS FOR GENERATED OUTPUT
3y 1m to grant Granted Aug 25, 2026
Patent 12718017
SYSTEM AND METHOD FOR PROVIDING LARGE LANGUAGE MODEL FOR SANCTIONS ARTIFICIAL INTELLIGENCE ASSISTED AUTOMATION
2y 11m to grant Granted Aug 25, 2026
Patent 12718819
INTERRUPTION DETECTION AND HANDLING BY DIGITAL ASSISTANTS
2y 3m to grant Granted Aug 25, 2026
Patent 12710915
VOICE MODIFICATION FOR WEARABLE DEVICE
3y 9m to grant Granted Aug 18, 2026
Patent 12712972
System and method for generating and managing a workflow using webhook technology
2y 4m to grant Granted Aug 18, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
84%
Grant Probability
99%
With Interview (+30.1%)
2y 7m (~11m remaining)
Median Time to Grant
Low
PTA Risk
Based on 184 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month