Prosecution Insights
Last updated: July 31, 2026
Application No. 18/662,550

Abridged Multilingual Speech Models For Automatic Speech Recognition

Final Rejection §103
Filed
May 13, 2024
Examiner
DUGDA, MULUGETA TUJI
Art Unit
2653
Tech Center
2600 — Communications
Assignee
ORACLE INTERNATIONAL Corporation
OA Round
2 (Final)
82%
Grant Probability
Favorable
3-4
OA Rounds
8m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 82% — above average
82%
Career Allowance Rate
44 granted / 54 resolved
+19.5% vs TC avg
Strong +22% interview lift
Without
With
+21.7%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
16 currently pending
Career history
74
Total Applications
across all art units

Statute-Specific Performance

§101
4.9%
-35.1% vs TC avg
§103
91.5%
+51.5% vs TC avg
§102
3.7%
-36.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 54 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims 1-20 are pending, and claims 1, 8 and 15 are independent claims. Response to Arguments Applicant's arguments, see Arguments pages 7-10, filed on 04/13/2026, with respect to 35 USC § 103 claim rejections have been fully considered but they are not persuasive. The Applicant argues, after listing the various combination of references for rejecting the various claims, that the cited references fail to teach or suggest the claimed limitation of "removing, from the multilingual ASR model, a second portion of the embedding matrix corresponding to the second subset of language tokens associated with the second language" as recited in Claim 1. The Applicant further asserts that the Office Action relies (a) on Wu to generally teach a multilingual Automatic Speech Recognition (ASR) model and (b) on Mavandadi to allegedly teach generating a language-specific ASR model. The Applicant also argues that the portions of Mavandadi cited by the Office Action do not disclose, either expressly or inherently, modifying a multilingual embedding matrix by removing language-specific portions of the embedding matrix and thus, the rejection is based on an improper characterization of the reference (Arguments, page 8). The Examiner respectfully disagrees. In Mavandadi, the multilingual Automatic Speech Recognition (ASR) model uses separate non-causal encoders and decoders for each language. Either predicted or inferred LID is used to activate one encoder-decoder pair for each utterance. A multilingual decoder can be used to generate hypotheses directly based on the causal encoder outputs without needing LID. We use the trained parameters from E0 to initialize the first pass encoders of our remaining experimental architectures and keep those parameters frozen during training in order to preserve the models’ capability to recognize multilingual speech using their first pass. The Causal Encoder has all the Multilingual components of the different languages and it retains everything and after going through LID, each Non-Causal Encoder and associated Decoder has the language specific tokens (Mavandadi, page 840, left col, 1st para; Mavandadi, Figure 1 and Figure 2, please see Caption). The Applicant argues that first, Mavandadi does not disclose a multilingual embedding matrix. Instead, Mavandadi describes a system composed of multiple, distinct language-specific encoder-decoder pairs. Mavandadi explains that the prediction network uses a "separate identically shaped embedding table for each language". This disclosure is fundamentally inconsistent with the claim recitations, in which a single multilingual embedding matrix includes tokens associated with multiple languages. The Applicant further argues that Mavandadi lacks a multilingual embedding matrix, it necessarily cannot teach removing a portion of the multilingual embedding matrix corresponding to a subset of language tokens. Then, the Applicant further argues that the Office Action appears to equate Mavandadi's language identification (LID)- based routing system with the claimed "removing" step, but this equivalence is not supported by the reference. Specifically, Mavandadi describes routing an input audio signal to a selected language-specific encoder-decoder pair using a switch controlled by LID information. The Office Action equates the non-selection of the remaining encoder-decoder pairs as the claimed step of "removing ... a second portion of the embedding matrix corresponding to the second subset of language tokens...". The Applicant further argues that this type of routing is not equivalent to removing a portion of an embedding matrix. The claim recites modifying a multilingual model by removing entries within a multilingual embedding matrix that correspond to certain language tokens. In contrast, according to the Applicant, Mavandadi leaves non-selected embedding tables, from the array of distinct embedding tables, intact and unmodified, merely selecting one for use while leaving the others unused but still present. The Applicant argues that the claimed "removing" step operates on a multilingual embedding matrix in which tokens from multiple languages coexist within a single embedding matrix. This coexistence permits the selective removal of entries within the shared matrix. Furthermore, entries within a shared multilingual embedding matrix allow for direct vector-space comparison and retrieval across languages. Mavandadi contains no such shared representation and no disclosure of selectively modifying a single embedding matrix. Instead, Mavandadi's architecture isolates languages into separate embedding tables, eliminating any possibility of removing a subset of language tokens from a unified embedding matrix. As such, even under a broad interpretation, Mavandadi cannot reasonably be understood to teach or suggest the claimed "removing" step (Arguments, page 8-10). The Examiner respectfully disagrees. Mavandadi does teach a multilingual embedding matrix. On Mavandadi, page 840, Figure 2, Fig. 2 Figure caption, we look at the Block diagram for E3 multilingual model. Model uses separate non-causal encoders and decoders for each language. Either predicted or inferred LID is used to activate one encoder-decoder pair for each utterance. A multilingual decoder can be used to generate hypotheses directly based on the causal encoder outputs without needing LID. On Mavandadi, page 840, right col, 1st para, for E3, they convert the non-causal encoder to be language-dependent as well, resulting in the entire second pass being replicated per-language. Fig. 2 illustrates the E3 architecture.; [i.e., 1st Pass Hypothesis – Multilingual Decoder as “multilingual ASR model”; LID= language identification (LID) information; After 1st pass, switching with LID to individual specific languages like en-US, zh-TW …en-GB basically “REMOVES” all other languages except those specific language going through the Non-Casual Encoder and Decoder before the 2nd pass] (Mavandadi, page 840, Figure 2, Fig. 2 Figure caption; Mavandadi, page 840, right col, 1st para, for E3). The Applicant argues that the Office Action appears to equate Mavandadi's language identification (LID)- based routing system with the claimed "removing" step, but this equivalence is not supported by the reference. Specifically, Mavandadi describes routing an input audio signal to a selected language-specific encoder-decoder pair using a switch controlled by LID information. The Office Action equates the non-selection of the remaining encoder-decoder pairs as the claimed step of "removing ... a second portion of the embedding matrix corresponding to the second subset of language tokens...". However, this type of routing is not equivalent to removing a portion of an embedding matrix. The claim recites modifying a multilingual model by removing entries within a multilingual embedding matrix that correspond to certain language tokens. In contrast, Mavandadi leaves non-selected embedding tables, from the array of distinct embedding tables, intact and unmodified, merely selecting one for use while leaving the others unused but still present. The Office Action's assertion that non-selection of other languages "basically 'REMOVES' [sic] all other languages" is therefore inaccurate, as no actual removal or modification of a portion of an embedding matrix occurs in Mavandadi. Moreover, the Applicant argues that, the claimed "removing" step operates on a multilingual embedding matrix in which tokens from multiple languages coexist within a single embedding matrix and as such, even under a broad interpretation, Mavandadi cannot reasonably be understood to teach or suggest the claimed "removing" step (Arguments, page 9-10). The Examiner respectfully disagrees. Prior to selection the embedding matrix is shared since all language IDs are present and as Applicant states the reference teaches embedding tables albeit separate. However, the selection based on the language ID creates a specific table to be selected from the step of tables. There is nothing in the claims that require or show the benefit of a shared table, how that table is shared, and how the table is used other than a selection thereof. Therefore, the current reference’s selection of a table of the series of tables would anticipate the claims. Therefore, the current reference’s selection of a table of the series of tables would anticipate the claims as stated. The Applicant argues that, even if Wu is assumed (solely for the sake of argument) to disclose a multilingual ASR model, the Office Action has not identified any teaching, suggestion, or motivation in Mavandadi that would lead one of ordinary skill in the art to modify such a model by removing portions of an embedding matrix corresponding to specific language tokens. To the contrary, Mavandadi adopts a different design approach in which monolingual models are duplicated and selected via routing. This approach does not suggest, and in fact discourages, altering a shared multilingual embedding matrix. Accordingly, there is no proper basis for combining Wu and Mavandadi in the manner proposed by the Office Action (Arguments, page 10). The Examiner respectfully disagrees. In response to applicant’s argument that there is no teaching, suggestion, or motivation to combine the references of Wu and Mavandadi, the examiner recognizes that obviousness may be established by combining or modifying the teachings of the prior art to produce the claimed invention where there is some teaching, suggestion, or motivation to do so found either in the references themselves or in the knowledge generally available to one of ordinary skill in the art. See In re Fine, 837 F.2d 1071, 5 USPQ2d 1596 (Fed. Cir. 1988), In re Jones, 958 F.2d 347, 21 USPQ2d 1941 (Fed. Cir. 1992), and KSR International Co. v. Teleflex, Inc., 550 U.S. 398, 82 USPQ2d 1385 (2007). Furthermore, the Examiner would like to note that the reference of Mavandadi does teach what is being asserted as noted above. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-3, 5, 8-10, 12, 15-17 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Wu et al. Pat App No. US 20250022457 A1 (Wu) in view of Mavandadi et al., "A Truly Multilingual First Pass and Monolingual Second Pass Streaming on-Device ASR System," 2022 IEEE Spoken Language Technology Workshop (SLT), Doha, Qatar, 2023, pp. 838-845 (Year: 2022). Regarding Claim 1, Wu discloses one or more non-transitory computer-readable media storing instructions which, when executed by one or more hardware processors, cause performance of operations (Wu, para 0045, processing units performing any of methods 400 and 500 may be executing instructions stored on a non-transitory computer-readable storage media) comprising: accessing a multilingual automatic speech recognition (ASR) model comprising a token embedding matrix corresponding to a plurality of language tokens, wherein the plurality of language tokens comprises at least (a) a first subset of language tokens associated with a first language and (b) a second subset of language tokens associated with a second language (Wu, para 0010-0013, providing systems and techniques that use a cascaded diffusion model (e.g., a machine learning model that includes a diffusion process and receives, as input, output from another machine learning model) for multi-lingual semi-supervised ASR. The cascaded diffusion model may receive as training data a plurality of (speech, text) pairs… A vector representation of the speech data may be concatenated with a vector representation of the text data of a (speech, text) pair to form an input tensor (e.g., one or more vectors or tensors within an embedding space) for training the diffusion model… In some embodiments, the input speech data and the output text data are in the same language. In some embodiments, the input speech data is in a first language, and the output text data is in a second language (e.g., language translation). In some embodiments, the input speech data includes multiple languages, and the output text data similarly includes multiple languages (e.g., multi-lingual support); para 0040, Figure 3, The (speech, text) pair may include a speech portion 310 and a text portion 312. Speech portion 310 may be represented as a sequence of mel-spectrograms, where each frame of the mel-spectrogram includes a vector within an embedding space (e.g., an 80-dimension vector). Text portion 312 may be represented as a sequence of discrete tokens. Text portion 312 may be converted (e.g., embedding 320) from discrete tokens into text vectors 314;); Wu does not specifically disclose generating a language-specific ASR model for the first language, retaining, from the multilingual ASR model, a first portion of the embedding matrix corresponding to the first subset of language tokens associated with the first language, removing, from the multilingual ASR model, a second portion of the embedding matrix corresponding to the second subset of language tokens associated with the second language, and applying a digital audio input, comprising spoken language in the first language, to the language-specific ASR model, to obtain a transcript of the digital audio input in the first language. However, Mavandadi, in the same field of endeavor, discloses generating a language-specific ASR model for the first language (Mavandadi, page 839, Figure 1, baseline monolingual architecture), at least by: retaining, from the multilingual ASR model, a first portion of the embedding matrix corresponding to the first subset of language tokens associated with the first language (Mavandadi, Figure 2 see Caption: Model uses separate non-causal encoders and decoders for each language. Either predicted or inferred LID is used to activate one encoder-decoder pair for each utterance. A multilingual decoder can be used to generate hypotheses directly based on the causal encoder outputs without needing LID; Mavandadi, page 840, left col, 1st para, We use the trained parameters from E0 to initialize the first pass encoders of our remaining experimental architectures and keep those parameters frozen during training in order to preserve the models’ capability to recognize multilingual speech using their first pass ; [i.e., From Figure 1, the Causal Encoder has all the Multilingual components of the different languages and it retains everything and after going through LID, each Non-Causal Encoder and associated Decoder has the language specific tokens]); removing, from the multilingual ASR model, a second portion of the embedding matrix corresponding to the second subset of language tokens associated with the second language (Mavandadi, page 840, Figure 2, Fig. 2 Figure caption: Block diagram for E3 multilingual model. Model uses separate non-causal encoders and decoders for each language. Either predicted or inferred LID is used to activate one encoder-decoder pair for each utterance. A multilingual decoder can be used to generate hypotheses directly based on the causal encoder outputs without needing LID; Mavandadi, page 840, right col, 1st para, for E3, we convert the non-causal encoder to be language-dependent as well, resulting in the entire second pass being replicated per-language. Fig. 2 illustrates the E3 architecture.; [i.e., 1st Pass Hypothesis – Multilingual Decoder as “multilingual ASR model”; LID= language identification (LID) information; After 1st pass, switching with LID to individual specific languages like en-US, zh-TW …en-GB basically “REMOVES” all other languages except those specific language going through the Non-Casual Encoder and Decoder before the 2nd pass] … …); applying a digital audio input, comprising spoken language in the first language, to the language-specific ASR model, to obtain a transcript of the digital audio input in the first language (Mavandadi, page 843, (1) Decoders with Language Specific Non-Causal Encoders produce 2nd pass Hypo and (2) Multilingual Decoder produced 1st Pass Hypo. Experimental model WERs for different Architectures, including E3, E3+ and E31 are provided in Table 4 and discussed. In conclusion: "Our best system combines per-language decoders and second-pass encoders with a shared multilingual encoder to achieve a 5.5% reduction of WER relative to monolingual models”). Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Mavandadi in the method of Wu because this would enable a truly multilingual first-pass and monolingual second-pass streaming on-device ASR system based on the recently developed Cascaded Encoders model so that the ASR systems will be more accurate, with low latency, and effectively handle language switching in order to be useful for the 60% of the world population that speaks more than one language (Mavandadi, Abstract). Regarding Claim 2, Wu in view of Mavandadi discloses the one or more media of claim 1, wherein the first subset of language tokens associated with the first language comprises tokens associated with a particular language script (Wu, para 0020-0021, In some embodiments, the speech portion of a (speech, text) pair and the text portion of the pair are in the same language. In some embodiments, the speech portion of a (speech, text) pair may be in a first language while the text portion is in a second language (e.g., language translation). In some embodiments, the speech portion includes speech in multiple languages and the text portion includes text in those same languages (e.g., multi-lingual support). Cascaded ASR module 120 may include a diffusion machine learning model that is trained using the (speech, text) pairs in training dataset 130. The speech input and the text input of a given (speech, text) pair may be combined to create a single input tensor for the diffusion model. In some embodiments, the speech data is represented as a sequence of mel-spectrograms, where each frame of the mel-spectrogram includes a vector within an embedding space (e.g., an 80-dimension vector). For example, a neural network (e.g., Wav2Vec) may be used to convert the input speech data into vectors within an embedding space. The text data may be a represented as a sequence of discrete tokens (e.g., words, phonemes, IPA symbols, etc.)). Regarding Claim 3, The one or more media of claim 1. Mavandadi further teaches: wherein generating the language-specific ASR model further comprises: retaining, from the multilingual ASR model, a third portion of the embedding matrix corresponding to numerical characters (Mavandadi, page 840, see Figure 2; Causal Encoder is inherently converting text to vectors which are just numbers/numerical characters). Regarding Claim 5, The one or more media of claim 1, the operations further comprising: identifying the first subset of language tokens associated with the first language, at least by tokenizing a corpus of text written in the first language (Wu, para 0020-0023, In some embodiments, the speech portion of a (speech, text) pair may be in a first language while the text portion is in a second language (e.g., language translation). In some embodiments, the speech portion includes speech in multiple languages and the text portion includes text in those same languages (e.g., multi-lingual support). Cascaded ASR module 120 may include a diffusion machine learning model that is trained using the (speech, text) pairs in training dataset 130…The text data may be a represented as a sequence of discrete tokens (e.g., words, phonemes, IPA symbols, etc.). The text may be converted from discrete tokens into vectors within an embedding space (e.g., vectors each having 512 dimensions). In some embodiments, the embedding space of the text data may be different from the embedding space of the speech data. To convert the discrete tokens to vectors, an embedding mapping may be used, where each token is replaced by a vector within the embedding space. In some embodiments, the embedding mapping is performed using a lookup table. In some embodiments, the embedding mapping is performed using a neural network (e.g., Word2Vec). In some embodiments, the text data is converted from a first set of discrete tokens (e.g., words, phonemes) to a second set of discrete tokens (e.g., international phonetic alphabet (IPA) symbols). Then the second set of discrete tokens may be converted to the vector representation… The text portion may be represented as one or more vectors within an embedding space, so a linear model with a softmax layer may be used to translate (e.g., “round”) the vector representations back into discrete text tokens). Regarding Claim 8, Wu discloses a system comprising: one or more hardware processors (Wu, para 0045, processing units performing any of methods 400 and 500 may be executing instructions stored on a non-transitory computer-readable storage media. In at least one embodiment, any of methods 400 and 500 may be performed using multiple processor threads (e.g., CPU threads and/or GPU threads)); one or more non-transitory computer-readable media (Wu, para 0045, processing units performing any of methods 400 and 500 may be executing instructions stored on a non-transitory computer-readable storage media); and program instructions stored on the one or more non-transitory computer readable media which, when executed by the one or more hardware processors, cause the system to perform operations (Wu, para 0045, processing units performing any of methods 400 and 500 may be executing instructions stored on a non-transitory computer-readable storage media) comprising: accessing a multilingual automatic speech recognition (ASR) model comprising a token embedding matrix corresponding to a plurality of language tokens, wherein the plurality of language tokens comprises at least (a) a first subset of language tokens associated with a first language and (b) a second subset of language tokens associated with a second language (Wu, para 0010-0013, providing systems and techniques that use a cascaded diffusion model (e.g., a machine learning model that includes a diffusion process and receives, as input, output from another machine learning model) for multi-lingual semi-supervised ASR. The cascaded diffusion model may receive as training data a plurality of (speech, text) pairs… A vector representation of the speech data may be concatenated with a vector representation of the text data of a (speech, text) pair to form an input tensor (e.g., one or more vectors or tensors within an embedding space) for training the diffusion model… In some embodiments, the input speech data and the output text data are in the same language. In some embodiments, the input speech data is in a first language, and the output text data is in a second language (e.g., language translation). In some embodiments, the input speech data includes multiple languages, and the output text data similarly includes multiple languages (e.g., multi-lingual support); para 0040, Figure 3, The (speech, text) pair may include a speech portion 310 and a text portion 312. Speech portion 310 may be represented as a sequence of mel-spectrograms, where each frame of the mel-spectrogram includes a vector within an embedding space (e.g., an 80-dimension vector). Text portion 312 may be represented as a sequence of discrete tokens. Text portion 312 may be converted (e.g., embedding 320) from discrete tokens into text vectors 314); Wu does not specifically disclose generating a language-specific ASR model for the first language, at least by, retaining, from the multilingual ASR model, a first portion of the embedding matrix corresponding to the first subset of language tokens associated with the first language, removing, from the multilingual ASR model, a second portion of the embedding matrix corresponding to the second subset of language tokens associated with the second language, and applying a digital audio input, comprising spoken language in the first language, to the language-specific ASR model, to obtain a transcript of the digital audio input in the first language. However, Mavandadi, in the same field of endeavor, discloses generating a language-specific ASR model for the first language (Mavandadi, page 839, Figure 1, baseline monolingual architecture), at least by: retaining, from the multilingual ASR model, a first portion of the embedding matrix corresponding to the first subset of language tokens associated with the first language (Mavandadi, pages 839 and 840, Figure 2, Figure 2 Figure Caption: Model uses separate non-causal encoders and decoders for each language. Either predicted or inferred LID is used to activate one encoder-decoder pair for each utterance. A multilingual decoder can be used to generate hypotheses directly based on the causal encoder outputs without needing LID; Mavandadi, page 840, left col, 1st para, We use the trained parameters from E0 to initialize the first pass encoders of our remaining experimental architectures and keep those parameters frozen during training in order to preserve the models’ capability to recognize multilingual speech using their first pass); ; [i.e., From Figure 1, the Causal Encoder has all the Multilingual components of the different languages and it retains everything and after going through LID, each Non-Causal Encoder and associated Decoder has the language specific tokens]); removing, from the multilingual ASR model, a second portion of the embedding matrix corresponding to the second subset of language tokens associated with the second language (Mavandadi, page 840, Figure 2, , Fig. 2 Figure caption: Block diagram for E3 multilingual model. Model uses separate non-causal encoders and decoders for each language. Either predicted or inferred LID is used to activate one encoder-decoder pair for each utterance. A multilingual decoder can be used to generate hypotheses directly based on the causal encoder outputs without needing LID; Mavandadi, page 840, right col, 1st para, for E3, we convert the non-causal encoder to be language-dependent as well, resulting in the entire second pass being replicated per-language. Fig. 2 illustrates the E3 architecture.; [i.e., 1st Pass Hypothesis – Multilingual Decoder as “multilingual ASR model”; LID= language identification (LID) information; After 1st pass, switching with LID to individual specific languages like en-US, zh-TW …en-GB basically “REMOVES” all other languages except those specific language going through the Non-Casual Encoder and Decoder before the 2nd pass]) applying a digital audio input, comprising spoken language in the first language, to the language-specific ASR model, to obtain a transcript of the digital audio input in the first language (Mavandadi, page 843, (1) Decoders with Language Specific Non-Causal Encoders produce 2nd pass Hypo and (2) Multilingual Decoder produced 1st Pass Hypo. Experimental model WERs for different Architectures, including E3, E3+ and E31 are provided in Table 4 and discussed. In conclusion: "Our best system combines per-language decoders and second-pass encoders with a shared multilingual encoder to achieve a 5.5% reduction of WER relative to monolingual models”). Therefore, it would have been obvious for one having ordinary skill in the art before theeffective filing date of the claimed invention to incorporate the method of Mavandadi in the method of Wu because this would enable a truly multilingual first-pass and monolingual second-pass streaming on-device ASR system based on the recently developed Cascaded Encoders model so that the ASR systems will be more accurate, with low latency, and effectively handle language switching in order to be useful for the 60% of the world population that speaks more than one language (Mavandadi, Abstract). Regarding Claim 9, Wu in view of Mavandadi discloses the system of claim 8, wherein the first subset of language tokens associated with the first language comprises tokens associated with a particular language script ((Wu, para 0020-0021, In some embodiments, the speech portion of a (speech, text) pair and the text portion of the pair are in the same language. In some embodiments, the speech portion of a (speech, text) pair may be in a first language while the text portion is in a second language (e.g., language translation). In some embodiments, the speech portion includes speech in multiple languages and the text portion includes text in those same languages (e.g., multi-lingual support). Cascaded ASR module 120 may include a diffusion machine learning model that is trained using the (speech, text) pairs in training dataset 130. The speech input and the text input of a given (speech, text) pair may be combined to create a single input tensor for the diffusion model. In some embodiments, the speech data is represented as a sequence of mel-spectrograms, where each frame of the mel-spectrogram includes a vector within an embedding space (e.g., an 80-dimension vector). For example, a neural network (e.g., Wav2Vec) may be used to convert the input speech data into vectors within an embedding space. The text data may be a represented as a sequence of discrete tokens (e.g., words, phonemes, IPA symbols, etc.))). Regarding Claim 10, Wu in view of Mavandadi discloses the system of claim 8. Mavandadi further teaches: wherein generating the language-specific ASR model further comprises: retaining, from the multilingual ASR model, a third portion of the embedding matrix corresponding to numerical characters (Mavandadi, page 840, see Figure 2: Causal Encoder is inherently converting text to vectors which are numbers/numerical characters). Regarding Claim 12, Wu in view of Mavandadi discloses the system of claim 8, the operations further comprising: identifying the first subset of language tokens associated with the first language, at least by tokenizing a corpus of text written in the first language (Wu, para 0020-0023, In some embodiments, the speech portion of a (speech, text) pair may be in a first language while the text portion is in a second language (e.g., language translation). In some embodiments, the speech portion includes speech in multiple languages and the text portion includes text in those same languages (e.g., multi-lingual support). Cascaded ASR module 120 may include a diffusion machine learning model that is trained using the (speech, text) pairs in training dataset 130… The text data may be a represented as a sequence of discrete tokens (e.g., words, phonemes, IPA symbols, etc.). The text may be converted from discrete tokens into vectors within an embedding space (e.g., vectors each having 512 dimensions). In some embodiments, the embedding space of the text data may be different from the embedding space of the speech data. To convert the discrete tokens to vectors, an embedding mapping may be used, where each token is replaced by a vector within the embedding space. In some embodiments, the embedding mapping is performed using a lookup table. In some embodiments, the embedding mapping is performed using a neural network (e.g., Word2Vec). In some embodiments, the text data is converted from a first set of discrete tokens (e.g., words, phonemes) to a second set of discrete tokens (e.g., international phonetic alphabet (IPA) symbols). Then the second set of discrete tokens may be converted to the vector representation… The text portion may be represented as one or more vectors within an embedding space, so a linear model with a softmax layer may be used to translate (e.g., “round”) the vector representations back into discrete text tokens). Regarding Claim 15, Wu discloses a method comprising: accessing a multilingual automatic speech recognition (ASR) model comprising a token embedding matrix corresponding to a plurality of language tokens, wherein the plurality of language tokens comprises at least (a) a first subset of language tokens associated with a first language and (b) a second subset of language tokens associated with a second language (Wu, para 0010-0013, providing systems and techniques that use a cascaded diffusion model (e.g., a machine learning model that includes a diffusion process and receives, as input, output from another machine learning model) for multi-lingual semi-supervised ASR. The cascaded diffusion model may receive as training data a plurality of (speech, text) pairs… A vector representation of the speech data may be concatenated with a vector representation of the text data of a (speech, text) pair to form an input tensor (e.g., one or more vectors or tensors within an embedding space) for training the diffusion model… In some embodiments, the input speech data and the output text data are in the same language. In some embodiments, the input speech data is in a first language, and the output text data is in a second language (e.g., language translation). In some embodiments, the input speech data includes multiple languages, and the output text data similarly includes multiple languages (e.g., multi-lingual support); para 0040, Figure 3, The (speech, text) pair may include a speech portion 310 and a text portion 312. Speech portion 310 may be represented as a sequence of mel-spectrograms, where each frame of the mel-spectrogram includes a vector within an embedding space (e.g., an 80-dimension vector). Text portion 312 may be represented as a sequence of discrete tokens. Text portion 312 may be converted (e.g., embedding 320) from discrete tokens into text vectors 314) Wu does not specifically disclose generating a language-specific ASR model for the first language, at least by: retaining, from the multilingual ASR model, a first portion of the embedding matrix corresponding to the first subset of language tokens associated with the first language, removing, from the multilingual ASR model, a second portion of the embedding matrix corresponding to the second subset of language tokens associated with the second language, applying a digital audio input, comprising spoken language in the first language, to the language-specific ASR model, to obtain a transcript of the digital audio input in the first language, wherein the method is performed by at least one device including a hardware processor. However, Mavandadi, in the same field of endeavor, discloses generating a language-specific ASR model for the first language (Mavandadi, page 839, Figure 1, baseline monolingual architecture), at least by: retaining, from the multilingual ASR model, a first portion of the embedding matrix corresponding to the first subset of language tokens associated with the first language (Mavandadi, pages 839 and 840, , Figure 1 and Figure 2, Figure 2 Figure Caption, Model uses separate non-causal encoders and decoders for each language. Either predicted or inferred LID is used to activate one encoder-decoder pair for each utterance. A multilingual decoder can be used to generate hypotheses directly based on the causal encoder outputs without needing LID; Mavandadi, page 840, left col, 1st para, We use the trained parameters from E0 to initialize the first pass encoders of our remaining experimental architectures and keep those parameters frozen during training in order to preserve the models’ capability to recognize multilingual speech using their first pass; [i.e., From Figure 1, the Causal Encoder has all the Multilingual components of the different languages and it retains everything and after going through LID, each Non-Causal Encoder and associated Decoder has the language specific tokens]); removing, from the multilingual ASR model, a second portion of the embedding matrix corresponding to the second subset of language tokens associated with the second language (Mavandadi, page 840, Figure 2, Fig. 2 Figure caption: Block diagram for E3 multilingual model. Model uses separate non-causal encoders and decoders for each language. Either predicted or inferred LID is used to activate one encoder-decoder pair for each utterance. A multilingual decoder can be used to generate hypotheses directly based on the causal encoder outputs without needing LID; Mavandadi, page 840, right col, 1st para, for E3, we convert the non-causal encoder to be language-dependent as well, resulting in the entire second pass being replicated per-language. Fig. 2 illustrates the E3 architecture.; [i.e., 1st Pass Hypothesis – Multilingual Decoder as “multilingual ASR model”; LID= language identification (LID) information; After 1st pass, switching with LID to individual specific languages like en-US, zh-TW …en-GB basically “REMOVES” all other languages except those specific language going through the Non-Casual Encoder and Decoder before the 2nd pass]); applying a digital audio input, comprising spoken language in the first language, to the language-specific ASR model, to obtain a transcript of the digital audio input in the first language (Mavandadi, page 843, (1) Decoders with Language Specific Non-Causal Encoders produce 2nd pass Hypo and (2) Multilingual Decoder produced 1st Pass Hypo. Experimental model WERs for different Architectures, including E3, E3+ and E31 are provided in Table 4 and discussed. In conclusion: "Our best system combines per-language decoders and second-pass encoders with a shared multilingual encoder to achieve a 5.5% reduction of WER relative to monolingual models”), wherein the method is performed by at least one device including a hardware processor (Mavandadi, page 838, Advances in hardware [1, 2] and algorithms [3–5] have made ASR an accessible and reliable approach for a variety of hands-free tasks in daily routines. In addition, the growing prevalence of personal computing devices has made it possible to bring ASR to much of the world’s population if models and algorithms are efficient enough to operate on hardware with a diverse range of computation capabilities). Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Mavandadi in the method of Wu because this would enable a truly multilingual first-pass and monolingual second-pass streaming on-device ASR system based on the recently developed Cascaded Encoders model so that the ASR systems will be more accurate, with low latency, and effectively handle language switching in order to be useful for the 60% of the world population that speaks more than one language (Mavandadi, Abstract). Regarding Claim 16, Wu in view of Mavandadi discloses the method of claim 15, wherein the first subset of language tokens associated with the first language comprises tokens associated with a particular language script (Wu, para 0020-0021, In some embodiments, the speech portion of a (speech, text) pair and the text portion of the pair are in the same language. In some embodiments, the speech portion of a (speech, text) pair may be in a first language while the text portion is in a second language (e.g., language translation). In some embodiments, the speech portion includes speech in multiple languages and the text portion includes text in those same languages (e.g., multi-lingual support). Cascaded ASR module 120 may include a diffusion machine learning model that is trained using the (speech, text) pairs in training dataset 130. The speech input and the text input of a given (speech, text) pair may be combined to create a single input tensor for the diffusion model. In some embodiments, the speech data is represented as a sequence of mel-spectrograms, where each frame of the mel-spectrogram includes a vector within an embedding space (e.g., an 80-dimension vector). For example, a neural network (e.g., Wav2Vec) may be used to convert the input speech data into vectors within an embedding space. The text data may be a represented as a sequence of discrete tokens (e.g., words, phonemes, IPA symbols, etc.)). Regarding Claim 17, Wu in view of Mavandadi discloses the method of claim 15. Mavandadi further teaches: wherein generating the language-specific ASR model further comprises: retaining, from the multilingual ASR model, a third portion of the embedding matrix corresponding to numerical characters (Mavandadi, page 840, see Figure 2: Causal Encoder is inherently converting text to vectors which are numbers/numerical characters). Regarding Claim 19, Wu in view of Mavandadi discloses the method of claim 15, the method further comprising: identifying the first subset of language tokens associated with the first language, at least by tokenizing a corpus of text written in the first language (Wu, para 0020-0023, In some embodiments, the speech portion of a (speech, text) pair may be in a first language while the text portion is in a second language (e.g., language translation). In some embodiments, the speech portion includes speech in multiple languages and the text portion includes text in those same languages (e.g., multi-lingual support). Cascaded ASR module 120 may include a diffusion machine learning model that is trained using the (speech, text) pairs in training dataset 130… The text data may be a represented as a sequence of discrete tokens (e.g., words, phonemes, IPA symbols, etc.). The text may be converted from discrete tokens into vectors within an embedding space (e.g., vectors each having 512 dimensions). In some embodiments, the embedding space of the text data may be different from the embedding space of the speech data. To convert the discrete tokens to vectors, an embedding mapping may be used, where each token is replaced by a vector within the embedding space. In some embodiments, the embedding mapping is performed using a lookup table. In some embodiments, the embedding mapping is performed using a neural network (e.g., Word2Vec). In some embodiments, the text data is converted from a first set of discrete tokens (e.g., words, phonemes) to a second set of discrete tokens (e.g., international phonetic alphabet (IPA) symbols). Then the second set of discrete tokens may be converted to the vector representation… The text portion may be represented as one or more vectors within an embedding space, so a linear model with a softmax layer may be used to translate (e.g., “round”) the vector representations back into discrete text tokens). Claims 4, 11 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Wu in view of Mavandadi, and further in view of Bălan, Dragoș Alexandru, "Improving the State-of-the-Art Frisian ASR by fine-tuning Large-Scale Cross-Lingual Pre-Trained Models." PhD diss., 2023 (Balan). Regarding Claim 4, Wu in view of Mavandadi disclose the one or more media of claim 1. Wu in view of Mavandadi do not specifically disclose wherein generating the language-specific ASR model further comprises: retaining, from the multilingual ASR model, a third portion of the embedding matrix corresponding to special characters used to direct operation of the multilingual ASR model. However, Balan, in the same field of endeavor, discloses wherein generating the language-specific ASR model further comprises: retaining, from the multilingual ASR model, a third portion of the embedding matrix corresponding to special characters used to direct operation of the multilingual ASR model (Balan, page 30, 3rd-4th para, To fine-tune the framework for speech recognition, a linear projection is added on top which is trained using tokens, and a CTC loss which outputs sequences of tokens. Tokens in this case correspond to characters. Employing the CTC method, we make predictions of tokens for every frame in our audio features. These tokens are selected from a predefined vocabulary…The special tokens consist of the whitespace token, a blank token that helps with decoding words that contain two letters next to each other (such as ”namme”), or an unknown token for characters that are out of the vocabulary. After tokens have been predicted for each frame, the non-space tokens that appear consecutively are combined into a single character, resulting in a character sequence for each group of tokens separated by a space token. The CTC loss function’s role is to minimize the differences between the predicted and the target character sequences so that the model outputs accurate transcriptions). Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Balan in the method of Wu in view of Mavandadi because this would enhance Frisian ASR performance and address the challenges posed by its lowresource status by focusing on fine-tuning the XLS-R model, a large-scale cross-lingual pre-trained model, that has shown promising results in multilingual ASR tasks (Balan, Abstract). Regarding Claim 11, Wu in view of Mavandadi discloses the system of claim 8. Wu in view of Mavandadi do not specifically disclose wherein generating the language-specific ASR model further comprises: retaining, from the multilingual ASR model, a third portion of the embedding matrix corresponding to special characters used to direct operation of the multilingual ASR model. However, Balan, in the same field of endeavor, discloses wherein generating the language-specific ASR model further comprises: retaining, from the multilingual ASR model, a third portion of the embedding matrix corresponding to special characters used to direct operation of the multilingual ASR model (Balan, page 30, 3rd-4th para, To fine-tune the framework for speech recognition, a linear projection is added on top which is trained using tokens, and a CTC loss which outputs sequences of tokens. Tokens in this case correspond to characters. Employing the CTC method, we make predictions of tokens for every frame in our audio features. These tokens are selected from a predefined vocabulary…The special tokens consist of the whitespace token, a blank token that helps with decoding words that contain two letters next to each other (such as ”namme”), or an unknown token for characters that are out of the vocabulary. After tokens have been predicted for each frame, the non-space tokens that appear consecutively are combined into a single character, resulting in a character sequence for each group of tokens separated by a space token. The CTC loss function’s role is to minimize the differences between the predicted and the target character sequences so that the model outputs accurate transcriptions ). Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Balan in the method of Wu in view of Mavandadi because this would enhance Frisian ASR performance and address the challenges posed by its lowresource status by focusing on fine-tuning the XLS-R model, a large-scale cross-lingual pre-trained model, that has shown promising results in multilingual ASR tasks (Balan, Abstract). Regarding Claim 18, Wu in view of Mavandadi discloses the method of claim 15. Wu in view of Mavandadi do not specifically disclose wherein generating the language-specific ASR model further comprises: retaining, from the multilingual ASR model, a third portion of the embedding matrix corresponding to special characters used to direct operation of the multilingual ASR model. However, Balan, in the same field of endeavor, discloses wherein generating the language-specific ASR model further comprises: retaining, from the multilingual ASR model, a third portion of the embedding matrix corresponding to special characters used to direct operation of the multilingual ASR model (Balan, page 30, 3rd-4th para, To fine-tune the framework for speech recognition, a linear projection is added on top which is trained using tokens, and a CTC loss which outputs sequences of tokens. Tokens in this case correspond to characters. Employing the CTC method, we make predictions of tokens for every frame in our audio features. These tokens are selected from a predefined vocabulary…The special tokens consist of the whitespace token, a blank token that helps with decoding words that contain two letters next to each other (such as ”namme”), or an unknown token for characters that are out of the vocabulary. After tokens have been predicted for each frame, the non-space tokens that appear consecutively are combined into a single character, resulting in a character sequence for each group of tokens separated by a space token. The CTC loss function’s role is to minimize the differences between the predicted and the target character sequences so that the model outputs accurate transcriptions ). Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Balan in the method of Wu in view of Mavandadi because this would enhance Frisian ASR performance and address the challenges posed by its lowresource status by focusing on fine-tuning the XLS-R model, a large-scale cross-lingual pre-trained model, that has shown promising results in multilingual ASR tasks (Balan, Abstract). Claims 7 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Wu in view of Mavandadi, and further in view of MacWhinney, Brian, and Davida Fromm. "Language sample analysis with TalkBank: An update and review." Frontiers in communication 7 (2022): 865498 (MacWhinney)). Regarding Claim 7, The one or more media of claim 1. Wu in view of Mavandadi do not specifically disclose the operations further comprising: identifying the second subset of language tokens as tokens having a token reading direction different from a reading direction of the first language. However, MacWhinney, in the same field of endeavor, discloses the operations further comprising: identifying the second subset of language tokens as tokens having a token reading direction different from a reading direction of the first language (MacWhinney, page 5, right col, 2nd para, Entry of characters from languages that write from right to left is possible. However, combining right-to-left script with the left-to-right features of CHAT can be tricky. For that reason, we recommend the use of romanization for languages with right to left orthographies). Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of MacWhinney in the method of Wu in view of Mavandadi because this would enable utilization of open and free methods that have been used for language sample analysis (LSA) with several clinical populations examine automatic and implementation of TalkBank methods that use ASR (automatic speech recognition), NLP (natural language processing), database technology, statistics in R and Python, and ML (machine learning) (MacWhinney, Abstract). Regarding Claim 14, The system of claim 8. Wu in view of Mavandadi do not specifically disclose the operations further comprising: identifying the second subset of language tokens as tokens having a token reading direction different from a reading direction of the first language. However, MacWhinney, in the same field of endeavor, discloses the operations further comprising: identifying the second subset of language tokens as tokens having a token reading direction different from a reading direction of the first language (MacWhinney, page 5, right col, 2nd para, Entry of characters from languages that write from right to left is possible. However, combining right-to-left script with the left-to-right features of CHAT can be tricky. For that reason, we recommend the use of romanization for languages with right to left orthographies). Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of MacWhinney in the method of Wu in view of Mavandadi because this would enable utilization of open and free methods that have been used for language sample analysis (LSA) with several clinical populations examine automatic and implementation of TalkBank methods that use ASR (automatic speech recognition), NLP (natural language processing), database technology, statistics in R and Python, and ML (machine learning) (MacWhinney, Abstract). Allowable Subject Matter Claims 6, 13 and 20 are objected to as being dependent upon rejected base claims, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The reasons for allowance are that the prior art of record do not specifically teach the limitations as recited in the mentioned claims Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MULUGETA T. DUGDA whose telephone number is (703)756-1106. The examiner can normally be reached Mon - Fri, 4:30am - 7:00pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Paras D. Shah can be reached at 571-270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MULUGETA TUJI DUGDA/Examiner, Art Unit 2653 /Paras D Shah/Supervisory Patent Examiner, Art Unit 2653 07/13/2026
Read full office action

Prosecution Timeline

May 13, 2024
Application Filed
Feb 04, 2026
Non-Final Rejection mailed — §103
Apr 10, 2026
Examiner Interview Summary
Apr 10, 2026
Applicant Interview (Telephonic)
Apr 13, 2026
Response Filed
Jul 15, 2026
Final Rejection mailed — §103
Jul 28, 2026
Response after Non-Final Action

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12694338
TECHNIQUES FOR TRAINING AND DEPLOYING A NAMED ENTITY RECOGNITION MODEL
3y 2m to grant Granted Jul 28, 2026
Patent 12670918
VOICE MODIFICATION
2y 5m to grant Granted Jun 30, 2026
Patent 12620387
VOICE GENERATION METHOD AND APPARATUS, DEVICE, AND COMPUTER READABLE MEDIUM
3y 2m to grant Granted May 05, 2026
Patent 12597424
METHOD AND APPARATUS FOR DETERMINING SKILL FIELD OF DIALOGUE TEXT
3y 6m to grant Granted Apr 07, 2026
Patent 12592244
REDUCED-BANDWIDTH SPEECH ENHANCEMENT WITH BANDWIDTH EXTENSION
3y 6m to grant Granted Mar 31, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
82%
Grant Probability
99%
With Interview (+21.7%)
2y 11m (~8m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 54 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month