DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant's arguments filed on 06/01/2026 have been fully considered but they are not persuasive.
Regarding the 35 USC 102 rejection, Applicant argues:
“Anticipation requires that each element of the claim be disclosed in a single prior art reference. Applicant respectfully submits that the amendments to claim 1 overcome the rejection of claim 1 under 35 U.S.C. § 102(a)(1). Similar language is also included in independent claims 14 and 20. Accordingly, Applicant respectfully requests that the rejection of claims 1-3, 5, 7-11, 14-16, and 18-20 under 35 U.S.C. § 102(a)(1) be withdrawn.” (Remarks, page 10)
In response, Examiner respectfully disagrees. Prabhavalkar discloses an embodiment of on-the-fly rescoring (para 0053). Prabhavalkar appears to disclose in Equation 5 a first score in log P(y|x) and a second score of λ log Pc(y) and the word spoken is determined by summing up both scores with λ “controlling how much the contextual language model influences the overall model score during beam search” (para 0054). So, the result is correctly predicting “stop the alarm” as shown in example of para 0047). Therefore, the rejection is maintained.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-3, 5, 7-11, 14-16, and 18-20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Prabhavalkar et al. (US Pub 2020/0357387).
Regarding claim 1, Prabhavalkar discloses a method comprising:
applying an automatic speech recognition (ASR) model to audio data to generate an ASR output representative of output likelihoods that the audio data comprise one or more spoken speech units (SUs) (see abstract);
generating, using the ASR output, a first score characterizing a first likelihood that the audio data comprise a first word spoken at a time interval of the audio data, where the first word is a dictionary word (see equations 1-3 and 5; and para 0051-0054 – the term log P(y|x) characterizes the likelihood audio comprises the hypothesized word sequence; para 0053 – “ ‘speller’ FST, S, which transduces a sequence of graphemes/word-pieces into the corresponding word.” And model’s word-level output space constitutes a dictionary and the score accumulates over the beam search module decoding time steps corresponding to the time interval of the spoken word – see para 0034);
generating, using the ASR output, a second score characterizing a second likelihood that the audio data comprise a second word spoken at the time interval, the second word including a target word from a plurality of target words that are identified based at least on a context of the audio data (para 0054 - λ log Pc(y) is the contextual score from bias phrases compiled into WFST - matching units add weight; para 0038 – “from the context data 111 directly, the bias phrase selector 113 compiles a list of phrases 114 that are likely to be spoken. In the illustrated example, the bias phrase selector 113 determines that a voice command is likely, and so the bias phrase selector 113 provides a set of contextual bias phrases 114, 114 a-n that includes commands such as ‘turn on the lights,’ ‘stop the alarm,’”), wherein the second score is based on the output likelihoods that individual SUs spoken during the time interval match one or more SUs of the target word (see fig. 3C and para 0057 – “pushes weights to each subword unit of the word. To avoid artificially giving weight to prefixes which are boosted early on but do not match the entire phrase, a subtractive cost is included, as indicated by the negative weights shown in FIG. 3C. By pushing the wegiths to each subword unit of the word, the OTF rescoring technique of FIG. 3C aims to help keep the word on the beam”); and
predicting, using the first score and the second score, a word spoken at the time interval of the audio data (see para 0054 and equation 5 – the word spoken is determined by summing up both scores with λ “controlling how much the contextual language model influences the overall model score during beam search” (para 0054). So the result is correctly predicting “stop the alarm” as shown in example of para 0047).
Regarding claim 2, Prabhavalkar discloses wherein the ASR output comprises, for one or more time intervals of the audio data:
a plurality of SU likelihoods, wherein an individual SU likelihood of the plurality of SU likelihoods characterizes a probability that a respective SU of a plurality of SUs was spoken during a respective time interval (para 0014 – “the first encoder includes a stacked, recurrent neural network (RNN) and/or the decoder includes a stacked, unidirectional RNN configured to compute a probability of a sequence of output tokens” RNN produces per frame probability).
Regarding claim 3, Prabhavalkar discloses wherein the generating the second score comprises:
generating a plurality of hypotheses, wherein an individual hypothesis of the plurality of hypotheses:
associates a portion of the audio data with a respective hypothesized word of the plurality of target words (para 0015 – hypotheses for each bias phrase), and
assigns, based at least on the ASR output, a score to the respective hypothesized word, wherein the assigned score characterizes a likelihood that the portion of the audio data comprises the respective hypothesized word (para 0018 – “the bias encoder is configured to encode a corresponding bias context vector for each bias phrase in the set of bias phrases”); and
identifying, using the assigned scores, the second word as a most likely word represented in the portion of the audio data (para 0017 – “decoder configured to determine likelihoods of sequences of speech elements based on output of the first attention module and output of the bias attention module”).
Regarding claim 5, Prabhavalkar discloses wherein the generating the second score comprises performing a plurality of iterations, wherein an individual iteration of the plurality of iterations is associated with a respective time interval of a plurality of time intervals of the audio data (para 0015), and wherein the individual iteration comprises:
for a respective time interval of the plurality of time intervals, selecting, using the ASR output, from at least:
an SU of the second word, the SU being spoken during the respective time interval, or
a state of no SU being spoken during the respective time interval (para 0018 – “additional bias context vector represents an option to not bias the likelihoods of sequences of speech elements determined by the decoder toward any of the bias phrases”).
Regarding claim 7, Prabhavalkar discloses wherein the state of no SU being spoken is selected responsive to no SU of the second word having a likelihood of being spoken above a threshold likelihood (para 0018 – “additional bias context vector represents an option to not bias the likelihoods of sequences of speech elements determined by the decoder toward any of the bias phrases”).
Regarding claim 8, Prabhavalkar discloses wherein the selecting from the at least the SU of the second word or the state of no SU comprises:
identifying a candidate SU (para 0012-0013);
obtaining, using the ASR output, a likelihood of the candidate SU being spoken during the respective time interval (para 0015); and
responsive to determining that the candidate SU matches the SU of the second word, enhancing the obtained likelihood of the candidate SU (para 0017 – “a decoder configured to determine likelihoods of sequences of speech elements based on output of the first attention module and output of the bias attention module”).
Regarding claim 9, Prabhavalkar discloses wherein the generating the second score comprises:
performing a plurality of iterations identified by a context graph associated with the plurality of target words, wherein the context graph comprises one or more root nodes associated with a starting SU of one or more target words of the plurality of target words (para 0016 - “each bias prefix in the list of bias prefixes represents an initial portion of one or more of the bias phrases in the set of bias phrases” – prefix represents starting portion such as root node).
Regarding claim 10, Prabhavalkar discloses wherein the ASR model is further to generate an additional ASR output comprising a plurality of recognized words in the audio data (for example para 0076).
Regarding claim 11, Prabhavalkar discloses further comprising:
replacing one or more words of the plurality of recognized words with the predicted spoken word (para 0017 – bias output can replace the standard recognition).
Regarding claims 14 and 20, see rejection of claim 1.
Regarding claim 15, see rejection of claim 3.
Regarding claim 16, see rejection of claim 5.
Regarding claim 18, see rejection of claim 8.
Regarding claim 19, Prabhavalkar discloses wherein the system is comprised in one of:
an in-vehicle infotainment system for an autonomous or semi-autonomous machine;
a system for performing one or more simulation operations;
a system for performing one or more digital twin operations;
a system for performing light transport simulation;
a system for performing collaborative content creation for 3D assets;
a system for performing one or more deep learning operations (para 0014 - RNN);
a system implemented using an edge device (para 0034 – “the speech recognition model 200 resides on a user device 106 associated with a user 102”);
a system for generating or presenting at least one of virtual reality content, mixed reality content, or augmented reality content;
a system implemented using a robot;
a system for performing one or more conversational AI operations (para 0009);
a system implementing one or more large language models (LLMs);
a system implementing one or more language models;
a system for performing one or more generative AI operations;
a system for generating synthetic data;
a system incorporating one or more virtual machines (VMs);
a system implemented at least partially in a data center (para 0034); or
a system implemented at least partially using cloud computing resources (para 0034).
Allowable Subject Matter
Claims 4, 6, 12-13, 17 objected to as being dependent upon a rejected base claim but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NAFIZ E HOQUE whose telephone number is (571)270-1811. The examiner can normally be reached M-F 8-5.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ahmad Matar can be reached at (571)272-7488. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/NAFIZ E HOQUE/Primary Examiner, Art Unit 2693