DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 8/7/2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Status of Claims
Claims 1-24 are pending in this application.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1, 8, 10-11, 13, 20 and 22-23 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Ding et al. (U.S. Patent Application Publication 2023/0298569).
As per claims 1 and 13, Ding et al. discloses:
A system (Figure 6 and paragraphs [0046-0051]) comprising:
data processing hardware (Figure 6, item 610 and paragraphs [0046-0051]); and
memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations (Figure 6, item 620 and paragraphs [0046-0051]) comprising:
obtaining a plurality of training samples, each respective training sample of the plurality of training samples (Paragraph [0028] -The disclosed model trainer literally obtains training samples, each comprising a speech utterance and a corresponding transcription, as claimed.) comprising:
audio data characterizing a corresponding speech utterance; and a transcription of the corresponding speech utterance (Paragraph [0028] -The disclosed model trainer literally obtains training samples, each comprising a speech utterance and a corresponding transcription, as claimed.);
training an automatic speech recognition (ASR) model on the plurality of training samples, the ASR model comprising a recurrent neural network-transducer (RNN-T) architecture (Paragraphs [0028] & [0032] – The disclosed model is an RNN-T model and is effectively being trained);
quantizing the trained ASR model to an integer target fixed-bit width (Paragraph [0028] – The disclosed model is quantized to an integer fixed-bit); and
providing the quantized trained ASR model to a user device (Paragraph [0031] – the trained model is explicitly provided to the user device).
Claim 1 is directed to the method of using the system of claim 13, so is rejected for similar reasons.
As per claims 8 and 20, Ding et al. discloses all of the limitations of claims 1 and 13 above. Ding et al. further discloses:
an audio encoder and a decoder, the decoder comprising a prediction network and a joint network (Paragraph [0033]).
As per claims 10 and 22, Ding et al. discloses all of the limitations of claims 8 and 20 above. Ding et al. further discloses:
the audio encoder comprises a plurality of multi- headed self attention layers each comprising a multi-headed attention mechanism (Paragraph [0035]).
As per claims 11 and 23, Ding et al. discloses all of the limitations of claims 10 and 22 above. Ding et al. further discloses:
the plurality of multi-headed self attention layers comprises conformer layers or transformer layers (Paragraph [0035]).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 2-5, 9, 14-17 and 21 are rejected under 35 U.S.C. 103 as being unpatentable over Ding et al. (U.S. Patent Application Publication 2023/0298569) in view of Zhen et al. (Non-Patent Literature “Sub-8-Bit Quantization for On-Device Speech Recognition: A Regularization Free Approach”, listed in IDS dated 8/7/2025).
As per claims 2 and 14, Ding et al. discloses all of the limitations of claims 1 and 13 above. Ding et al. fails to disclose, but Zhen et al. in the same field of endeavor teaches:
quantizing the trained ASR model comprises quantizing, using per-channel asymmetrical quantization with scale backpropagation, the trained ASR model to the target fixed-bit width, the target fixed-bit width equal to two (Abstract and Sections 1-4).
It would be obvious for a person having ordinary skill in the art at the effective filing date of the invention of modify the method and system of Ding et al. with the quantization capabilities of Zhen et al. because it is a case of simple substitution with predicable effects.
As per claims 3 and 15, Ding et al. discloses all of the limitations of claims 1 and 13 above. Ding et al. fails to disclose, but Zhen et al. in the same field of endeavor teaches:
quantizing the trained ASR model comprises quantizing, using per-channel asymmetrical quantization with scale backpropagation, clipping, and sub-channel split, the trained ASR model to the target fixed-bit width, the target fixed-bit width equal to two (Abstract and Sections 1-4).
It would be obvious for a person having ordinary skill in the art at the effective filing date of the invention of modify the method and system of Ding et al. with the quantization capabilities of Zhen et al. because it is a case of simple substitution with predicable effects.
As per claims 4 and 16, Ding et al. discloses all of the limitations of claims 1 and 13 above. Ding et al. fails to disclose, but Zhen et al. in the same field of endeavor teaches:
quantizing the trained ASR model comprises quantizing, using absolute max binarization, the trained ASR model to the target fixed-bit width, the target fixed-bit width equal to one (Abstract and Sections 1-4).
It would be obvious for a person having ordinary skill in the art at the effective filing date of the invention of modify the method and system of Ding et al. with the quantization capabilities of Zhen et al. because it is a case of simple substitution with predicable effects.
As per claims 5 and 17, Ding et al. discloses all of the limitations of claims 1 and 13 above. Ding et al. fails to disclose, but Zhen et al. in the same field of endeavor teaches:
quantizing the trained ASR model comprises: subtracting a per channel mean value from a plurality of weights of the trained ASR model to provide a plurality of scaled weights; and binarizing the plurality scaled weights to the target fixed-bit width, the target fixed-bit width equal to one (Abstract and Sections 1-4).
It would be obvious for a person having ordinary skill in the art at the effective filing date of the invention of modify the method and system of Ding et al. with the quantization capabilities of Zhen et al. because it is a case of simple substitution with predicable effects.
As per claims 9 and 21, Ding et al. discloses all of the limitations of claims 8 and 20 above. Ding et al. fails to disclose, but Zhen et al. in the same field of endeavor teaches:
quantizing the ASR model comprises quantizing the audio encoder and not quantizing the decoder (Abstract and Sections 1-4).
It would be obvious for a person having ordinary skill in the art at the effective filing date of the invention of modify the method and system of Ding et al. with the quantization capabilities of Zhen et al. because it is a case of simple substitution with predicable effects.
Claims 7 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Ding et al. (U.S. Patent Application Publication 2023/0298569) in view of Zhao et al. (U.S. Patent Application Publication 2020/0357388).
As per claims 7 and 19, Ding et al. discloses all of the limitations of claims 1 and 13 above. Ding et al. fails to disclose, but Zhao et al. in the same field of endeavor teaches:
receiving training data comprising: a corpus of transcribed non-synthetic speech utterances, each transcribed non-synthetic speech utterance paired with a corresponding transcription; and a corpus of un-transcribed non-synthetic speech utterances, each un-transcribed non-synthetic speech utterance not paired with a corresponding transcription: training a teacher ASR model on the corpus of transcribed non-synthetic speech utterances to teach the teacher ASR model to learn how to predict the corresponding transcriptions from the non-synthetic speech utterances; and processing, using the trained teacher ASR model, the corpus of transcribed non- synthetic speech utterances and the corpus of un-transcribed non-synthetic speech utterances to predict corresponding pseudo ground-truth labels (Paragraphs [0055-0059], [0090-0091], [0097], [0108]).
It would be obvious for a person having ordinary skill in the art at the effective filing date of the invention of modify the method and system of Ding et al. with the Teacher model capabilities of Zhao et al. because it is a case of combining prior art elements according to known methods to yield predicable results.
Claims 12 and 24 are rejected under 35 U.S.C. 103 as being unpatentable over Ding et al. (U.S. Patent Application Publication 2023/0298569) in view of Zhang et al. (U.S. Patent Application Publication 2023/0306958).
As per claims 12 and 24, Ding et al. discloses all of the limitations of claims 1 and 13 above. Ding et al. fails to disclose, but Zhang et al. in the same field of endeavor teaches:
the speech utterances and transcriptions of the plurality of training samples span multiple different languages; and the trained ASR model comprises a multilingual ASR model (Abstract and Paragraph [0047]).
It would be obvious for a person having ordinary skill in the art at the effective filing date of the invention of modify the method and system of Ding et al. with the multilingual ASR capabilities of Zhang et al. because it is a case of combining prior art elements according to known methods to yield predicable results.
Allowable Subject Matter
Claims 6 and 18 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Examiner Notes
The Examiner cites particular columns and line numbers in the references as applied to the claims above for the convenience of the Applicant. Although the specified citations are representative of the teachings in the art and are applied to the specific limitations within the individual claim, other passages and figures may apply as well. It is respectfully requested that, in preparing responses, the Applicant fully considers the references in its entirety as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior art or as disclosed by the Examiner.
Communications via Internet e-mail are at the discretion of the applicant and require written authorization. Should the Applicant wish to communicate via e-mail, including the following paragraph in their response will allow the Examiner to do so:
“Recognizing that Internet communications are not secure, I hereby authorize the USPTO to communicate with me concerning any subject matter of this application by electronic mail. I understand that a copy of these communications will be made of record in the application file.”
Should e-mail communication be desired, the Examiner can be reached at Edwin.Leland@USPTO.gov
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to EDWIN S LELAND III whose telephone number is (571)270-5678. The examiner can normally be reached 8:00 - 5:00 M-F.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Hai Phan can be reached at 571-272-6338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/EDWIN S LELAND III/Primary Examiner, Art Unit 2654