Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
The terminal disclaimer filed on 06/06/2026 disclaiming the terminal portion of any patent granted on this application which would extend beyond the expiration date of U.S. Patent No. 12073824 has been reviewed and is accepted. The terminal disclaimer has been recorded.
Response to Arguments
Applicant's arguments with respect to claims 1, 10, and 19 have been considered but are moot in view of the new ground(s) of rejection. Applicant’s arguments are directed to the amended subject matter; new prior art citations and explanation using existing art of Phillips is provided below.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1, 10, and 19 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by US 20110054899 A1 Phillips; Michael S. et al. (hereinafter Phillips).
Re claim 1, Phillips teaches
1. A method implemented by one or more processors, the method comprising: (fig. 1 professor on hardware)
training a two-pass automatic speech recognition (ASR) model to generate a text representation of a spoken utterance, wherein training the ASR model comprises: (ASR operations using speech recognition models for training using past utterances, user history, context, and corrections thereof for instance 0071 fig. 1 and fig. 2… e.g. the first pass and second pass involving a user command with corrections or two separate interactions, both are used to train the model 0105 and fig. 7b-7c)
training a first-pass portion of the ASR model, wherein training the first-pass portion of the ASR model comprises updating one or more portions of a shared encoder portion of the ASR model (the first pass is a literal first interaction with an ASR model by a user such as a first command 0105 and fig. 7b-7c, an ASR model containing encoded data or an encoder per se as in 0177 and 0185, ASR operations using speech recognition models for training using past utterances, user history, context, and corrections thereof for instance 0071 fig. 1 and fig. 2) AND/OR updating one or more portions of a recurrent neural network transformer (RNN-T) decoder portion of the first-pass portion of the ASR model based on processing a plurality of training instances; and
training a second-pass portion of the ASR model, wherein training the second-pass portion of the ASR model comprises updating one or more portions of an additional encoder portion of the ASR model (the second pass can be a literal second interaction with an ASR model by a user such as the correction to a first input of a first command 0105 and fig. 7b-7c, an ASR model containing encoded data or an encoder per se as in 0177 and 0185, ASR operations using speech recognition models for training using past utterances, user history, context, and corrections thereof for instance 0071 fig. 1 and fig. 2) AND/OR updating one or more portions of a listen attend spell (LAS) decoder based on processing the plurality of training instances;
subsequent to training the ASR model, processing audio data capturing a spoken utterance using the ASR model to generate a text representation of the spoken utterance, (ASR inherently and explicitly in Phillips, converts speech to text, particularly after iterations of model training, e.g. the first pass and second pass involving a user command 0105 and fig. 7b-7c, an ASR model containing encoded data or an encoder per se as in 0177 and 0185, ASR operations using speech recognition models for training using past utterances, user history, context, and corrections thereof for instance 0071 fig. 1 and fig. 2)
, wherein the audio data capturing the spoken utterance comprises a sequence of segments and wherein processing the audio data capturing the spoken utterance using the ASR model comprises: (each word is a segment in a sequence utterance, as in 0104 with fig. 7b, the user can select the suggested options specifically by replacing or selecting words by speaking as in 0064-0066 using ASR models, such actions would be considered a second pass per se, the alternative interpretation is any pass that is in the future at a separate time)
using both the first-pass portion of the ASR model and the second pass portion of the ASR model in processing each of the segments; (utilizing the concept that each word is a segment in a sequence utterance, as in 0104 with fig. 7b, the user can select the suggested options specifically by replacing or selecting words by speaking as in 0064-0066 using ASR models, such actions would be considered a second pass per se, the alternative interpretation is any pass that is in the future at a separate time)
causing a client device to perform one or more actions based on the text representation of the spoken utterance. (in fig. 7c the action is opening and populating the app fields just by uttering a command… ASR inherently and explicitly in Phillips, converts speech to text, particularly after iterations of model training, e.g. the first pass and second pass involving a user command 0105 and fig. 7b-7c, an ASR model containing encoded data or an encoder per se as in 0177 and 0185, ASR operations using speech recognition models for training using past utterances, user history, context, and corrections thereof for instance 0071 fig. 1 and fig. 2)
Re claim 10, this claim has been rejected for teaching a broader, or narrower claim based on general inclusion of hardware alone (e.g. processor, memory, instructions), representation of claim 1 omitting/including hardware for instance, otherwise amounting to a virtually identical scope
For instance, see fig. 1 and 0194 which contains the necessary memory and processors.
Re claim 19, this claim has been rejected for teaching a broader, or narrower claim based on general inclusion of hardware alone (e.g. processor, memory, instructions), representation of claim 1 omitting/including hardware for instance, otherwise amounting to a virtually identical scope
For instance, see fig. 1 and 0194 which contains the necessary memory and processors.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 2 and 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over US 20110054899 A1 Phillips; Michael S. et al. (hereinafter Phillips) in view of US 20190318261 A1 Deng; Yue et al. (hereinafter Deng).
Re claims 2 and 11, Phillips teaches
2. The method of claim 1, wherein training the first-pass portion of the ASR model further comprises: (in fig. 7c the action is opening and populating the app fields just by uttering a command… ASR inherently and explicitly in Phillips, converts speech to text, particularly after iterations of model training, e.g. the first pass and second pass involving a user command 0105 and fig. 7b-7c)
for each of the plurality of training instances and until one or more conditions are satisfied: (confidence levels met 0083… in fig. 7c the action is opening and populating the app fields just by uttering a command… ASR inherently and explicitly in Phillips, converts speech to text, particularly after iterations of model training, e.g. the first pass and second pass involving a user command 0105 and fig. 7b-7c)
processing an instance of training audio data portion of the training instance using the shared encoder to generate shared encoder training output, wherein the training audio data captures a spoken training utterance; (an ASR model containing encoded data or an encoder per se as in 0177 and 0185, ASR operations using speech recognition models for training using past utterances, user history, context, and corrections thereof for instance 0071 fig. 1 and fig. 2)
However, while Phillips teaches encoder or encoding based ASR model learning, it fails to teach RNN concepts:
processing the shared encoder training output using the RNN-T decoder portion to generate predicted RNN-T training output; (Deng a sequence learning on real time data RNN is analogous to RNN-T under BRI, training by using sequenced learning on real tie data through loss in comparing prediction with ground truth and updating encoders per se 0055-0057 and 0065)
determining a loss based on comparing the predicted RNN-T training output and a ground truth text representation of the training utterance; (Deng a sequence learning on real time data RNN is analogous to RNN-T under BRI, training by using sequenced learning on real tie data through loss in comparing prediction with ground truth and updating encoders per se 0055-0057 and 0065)
updating the one or more portions of the shared encoder (Deng a sequence learning on real time data RNN is analogous to RNN-T under BRI, training by using sequenced learning on real tie data through loss in comparing prediction with ground truth and updating encoders per se 0055-0057 and 0065) and/or updating the one or more portions of the RNN-T decoder based on the loss.
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Phillips to incorporate the above claim limitations as taught by Deng to allow for simple substitution of one known element for another to obtain predictable results such as the trainable model with encoding and ASR of Phillips with an RNN equivalent for speech to allow for superior mapping of variable-length audio to text by leveraging sequential memory, improving both recognition accuracy and adaptability to new data, by having the encoder learn complex temporal dynamics and context from raw audio e.g. streamed for using the ground truth (transcriptions) to update these encoders, often in an end-to-end (E2E) fashion functionally analogous to RNN with sequenced data.
Claims 6-8 and 15-17 is/are rejected under 35 U.S.C. 103 as being unpatentable over US 20110054899 A1 Phillips; Michael S. et al. (hereinafter Phillips) in view of US 20190318261 A1 Deng; Yue et al. (hereinafter Deng) and further in view of US 20180166066 A1 Dimitriadis; Dimitrios B. et al. (hereinafter Dimitriadis).
Re claims 6 and 15, while the combination teaches encoder or encoding based ASR model learning, it fails to teach speaking start/stop and RNN concepts:
6. The method of claim 2, wherein the RNN-T training output includes an end of query token indicating a human speaker has finished speaking using the first-pass portion of the ASR model. (Dimitriadis user start and stop areas determined by applying RNN to sequenced frames of real time audio data 0041 with 0063-0064 fig. 4-6 and fig. 9, evidenced further in 0050-0054 for training or having the RNN learn per se, a label analogous to a token with nth number of passes possible, to prevent erroneous speech)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Phillips in view of Deng to incorporate the above claim limitations as taught by Dimitriadis to allow for simple substitution of one known element for another to obtain predictable results such as the trainable model with encoding and ASR of Phillips with an RNN and start/stop labeling for nth iterations of speech to train a model analogous to the model of Phllips with a weighting to enforce good learning avoid or prevent garbage or wrong recognitions, to allow for improved speaker diarization and accuracy of who is talking and when they start/stop, with learning to prevent erroneous speech being analogous to loss per se which RNNs utilize to handle arbitrary input sequences.
Re claims 7 and 16, while the combination teaches encoder or encoding based ASR model learning, it fails to teach speaking start/stop and RNN concepts:
7. The method of claim 6, wherein determining the human speaker has finished speaking the utterance comprises determining the human speaker has finished speaking the utterance in response to identifying the end of query token in the RNN-T training output. (Dimitriadis user start and stop areas determined by applying RNN to sequenced frames of real time audio data 0041 with 0063-0064 fig. 4-6 and fig. 9, evidenced further in 0050-0054 for training or having the RNN learn per se, a label analogous to a token with nth number of passes possible, to prevent erroneous speech)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Phillips in view of Deng to incorporate the above claim limitations as taught by Dimitriadis to allow for simple substitution of one known element for another to obtain predictable results such as the trainable model with encoding and ASR of Phillips with an RNN and start/stop labeling for nth iterations of speech to train a model analogous to the model of Phllips with a weighting to enforce good learning avoid or prevent garbage or wrong recognitions, to allow for improved speaker diarization and accuracy of who is talking and when they start/stop, with learning to prevent erroneous speech being analogous to loss per se which RNNs utilize to handle arbitrary input sequences.
Re claims 8 and 17, while the combination teaches encoder or encoding based ASR model learning, it fails to teach speaking start/stop and RNN concepts:
8. The method of claim 7, wherein training the ASR model comprises penalizing the RNN-T decoder portion for generating the end of query token too early or too late. (Dimitriadis RNNs use loss to predict outcomes and to prevent erroneous speech 0073, user start and stop areas determined by applying RNN to sequenced frames of real time audio data 0041 with 0063-0064 fig. 4-6 and fig. 9, evidenced further in 0050-0054 for training or having the RNN learn per se, a label analogous to a token with nth number of passes possible)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Phillips in view of Deng to incorporate the above claim limitations as taught by Dimitriadis to allow for simple substitution of one known element for another to obtain predictable results such as the trainable model with encoding and ASR of Phillips with an RNN and start/stop labeling, in which an RNN inherently uses loss i.e. penalization under BRI, for nth iterations of speech to train a model analogous to the model of Phllips with a weighting to enforce good learning avoid or prevent garbage or wrong recognitions, to allow for improved speaker diarization and accuracy of who is talking and when they start/stop, with learning to prevent erroneous speech being analogous to loss per se which RNNs utilize to handle arbitrary input sequences.
Allowable Subject Matter
Claims 3-5 and 12-14
Claims 9 and 18
Are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. After searching through patent and non-patent literature, there was no evidence that there exists a limitation in direct relation or an obvious variant to such limitations as a whole as precisely limited. When searching for a secondary prior art for the limitation as recited in the above claims, the most relevant topics pertained to material from the same Inventor and Assignee but did not teach or suggest the aforementioned complex limitations as a whole as precisely limited.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US 20180247643 A1 BATTENBERG; Eric et al.
LAS concepts
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHAEL COLUCCI whose telephone number is (571)270-1847. The examiner can normally be reached on M-F 9 AM - 7 PM.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Flanders can be reached at (571)272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MICHAEL COLUCCI/Primary Examiner, Art Unit 2655 (571)-270-1847
Examiner FAX: (571)-270-2847
Michael.Colucci@uspto.gov