Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Examiner’s Note
Providing supporting paragraph(s) for each limitation of amended/new claim(s) in Remarks is strongly requested for clear and definite claim interpretations by Examiner (e.g., to avoid rejections under 35 U.S.C § 112(a) “Lack of written description”)
Applicant can schedule an interview at any stage of the prosecution (e.g., Non-Final, Final, and After-Final) to discuss any issues related to, for example, rejections under 35 U.S.C § 101 and § 103, for moving toward allowance.
Priority
Acknowledgment is made of applicant's claim for the present application filed on 08/24/2021.
Response to Arguments
Applicant's arguments filed on 02/19/2026 have been fully considered but they are not persuasive.
In Remarks, pp. 9-12, Applicant contends:
The present embodiments resolve the technical issue regarding peaky probability distribution over outputs when fusing an external language model with an end-to-end speech recognition model by relaxing the sharpness of the probability distribution by distorting the probability distribution and lowering the amplitudes of higher probability during decoding. See at least para. [0017], [0050] of the Specification as filed.
Thus, similar to Desjardins, the present embodiments resolve a technical problem in the machine learning field by improving the functioning of speech recognition models.
Examiner’s response:
The examiner understands the applicant’s assertion.
However, it appears that each processing step is just applying the abstract idea to a general field of endeavor with additional elements. In addition, improvements to technology or technical field are not necessarily reflected in the claims. Thus, the claim does not integrate the judicial exception into a practical application, and the claim does not amount to significantly more than the judicial exception.
The examiner understands the applicant’s assertion “distorting the probability distribution and lowering the amplitudes of higher probability during decoding” and “improving the functioning of speech recognition models”.
Pars 2 and 16 states “the probability distribution over output symbols (alphabet) is quite peaky, making the effective fusion of the RNN-T output and the ExternalLM difficult” and “Recurrent Neural Network Transducer (RNN-T) architectures lack explicit language models”. However, it is not clear how lowering the amplitudes of higher probability during decoding resolves the problems of the conventional approaches. The applicant may need to clearly explain how the recited claims provide improvements compared to the conventional approaches.
Currently, the limitations do not clearly show e.g., improvements in computer technology and improvements to other technical fields. Rather, the improvements in Remarks are about just improving the abstract ideas of the independent claims. It doesn’t seem that the specification and/or the independent claims clearly show how the inventive concept of the claims enables improvements and how they are tied together. The applicant may need to show specific improvements from the specification and/or amend the claims to show how the claim languages and improvements are tied together.
To find a valid improvement to a technology, MPEP 2106.04(d)(1) says the specification must explain the improvement and that the claim must reflect the disclosed improvement. Furthermore, the improvement should not be merely a consequence of the abstract idea. See MPEP 2106.05(a). An improvement in the abstract idea itself is not an improvement to technology.
For at least these reasons, Applicant's arguments are not convincing.
Applicant’s arguments regarding 35 USC 103 with respect to the independent claims have been considered but are moot because the arguments are directed to amended limitation(s) that has/have not been previously examined.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim(s) 1-20 is/are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim(s) 1 recite(s) the limitation “the alignments” (last line of the “suppressing” step). There is insufficient antecedent basis for this limitation in the claim. It is not clear what it is referring to. It appears that it is supposed to indicate “alignments of outputs of the end-to-end speech recognition model”, but it may not because it appears that “alignments of outputs of the end-to-end speech recognition model” consider all outputs while “alignments having same output symbol sequences” are only based on “same output symbol sequences”. Thus, it appears that the two “alignments” indicate different things. It appears it may need to read “alignments”, or something else. For the purposes of examination, “alignments” is used. In addition, claim(s) 10, 19 is/are rejected for the same reason.
Claim(s) 1-20 each recite(s) limitations that raise issues of indefiniteness as set forth above, and their dependent claims are rejected at least based on their direct and/or indirect dependency from the claims listed above. Appropriate explanation and/or amendment is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding claim 1
The claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1:
The limitations of
“… fusing an end-to-end speech recognition model with an external language model (ExternalLM), the method comprising:
obtaining an output of the end-to-end speech recognition model, the output being a first probability distribution;
suppressing a sharpness of the first probability distribution including peaks as outputs of the end-to-end speech recognition model above a threshold during decoding … that increases accuracy of the end-to-end speech recognition model … by transforming, …, the first probability distribution into a transformed probability distribution with lowered amplitudes of the peaks of the first probability distribution based on a threshold by utilizing a non-linear function for symbol posterior probabilities of alignments of outputs of the end-to-end speech recognition model obtained with the alignments having same output symbol sequences;
fusing the transformed probability distribution and an output of the ExternalLM being a second probability distribution to obtain a fused model of the end-to-end speech recognition model and the ExternalLM for decoding speech; and
performing speech recognition tasks with the fused model to control computing systems that communicate with users”, as drafted, are a process that, under its broadest reasonable interpretation, covers performance of the limitation based on mathematical relationships and/or mathematical formulas or equations and/or mathematical calculations. That is, nothing in the claim element precludes the step from practically being performed based on mathematical relationships and/or mathematical formulas or equations and/or mathematical calculations.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation based on mathematical relationships and/or mathematical formulas or equations and/or mathematical calculations, but for the recitation of generic computer components, then it falls within the “Mathematical concepts” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong 2: This judicial exception is not integrated into a practical application.
The claim recites additional elements that are mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. See MPEP 2106.05(f). In particular, the claim recites an additional element(s) (“computer-implemented”, “by a hardware processor”, “with the ExtetnalLM”) – using a device and/or a model to process data. The device and/or the model in each step is/are recited at a high-level of generality (i.e., as a generic computer performing a generic computer function of processing data) such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
As discussed above, with respect to integration of the abstract idea into a practical application, the additional elements of using a generic computer component to perform each step amount to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The claim is not patent eligible. MPEP 2106.05(f).
Regarding claim 2
The claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1:
The limitations of
“wherein the fusing comprising searching for a best output sequence in a decoding by applying a max function to the transformed probability distribution”, as drafted, are a process that, under its broadest reasonable interpretation, covers performance of the limitation based on mathematical relationships and/or mathematical formulas or equations and/or mathematical calculations. That is, nothing in the claim element precludes the step from practically being performed based on mathematical relationships and/or mathematical formulas or equations and/or mathematical calculations.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation based on mathematical relationships and/or mathematical formulas or equations and/or mathematical calculations, but for the recitation of generic computer components, then it falls within the “Mathematical concepts” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. In particular, the claim does not recite additional elements. Thus, the claim is directed to an abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Thus, the claim is not patent eligible.
Regarding claim 3
The claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1:
The limitations of
“wherein the transforming is performed by applying a non-linear function to the first probability distribution”, as drafted, are a process that, under its broadest reasonable interpretation, covers performance of the limitation based on mathematical relationships and/or mathematical formulas or equations and/or mathematical calculations. That is, nothing in the claim element precludes the step from practically being performed based on mathematical relationships and/or mathematical formulas or equations and/or mathematical calculations.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation based on mathematical relationships and/or mathematical formulas or equations and/or mathematical calculations, but for the recitation of generic computer components, then it falls within the “Mathematical concepts” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. In particular, the claim does not recite additional elements. Thus, the claim is directed to an abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Thus, the claim is not patent eligible.
Regarding claim 4
The claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1:
The limitations of
“wherein the transforming is performed by applying a logarithmic function to the first probability distribution”, as drafted, are a process that, under its broadest reasonable interpretation, covers performance of the limitation based on mathematical relationships and/or mathematical formulas or equations and/or mathematical calculations. That is, nothing in the claim element precludes the step from practically being performed based on mathematical relationships and/or mathematical formulas or equations and/or mathematical calculations.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation based on mathematical relationships and/or mathematical formulas or equations and/or mathematical calculations, but for the recitation of generic computer components, then it falls within the “Mathematical concepts” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. In particular, the claim does not recite additional elements. Thus, the claim is directed to an abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Thus, the claim is not patent eligible.
Regarding claim 5
The claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1:
The limitations of
“wherein the transforming is performed by applying a power function to the first probability distribution”, as drafted, are a process that, under its broadest reasonable interpretation, covers performance of the limitation based on mathematical relationships and/or mathematical formulas or equations and/or mathematical calculations. That is, nothing in the claim element precludes the step from practically being performed based on mathematical relationships and/or mathematical formulas or equations and/or mathematical calculations.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation based on mathematical relationships and/or mathematical formulas or equations and/or mathematical calculations, but for the recitation of generic computer components, then it falls within the “Mathematical concepts” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. In particular, the claim does not recite additional elements. Thus, the claim is directed to an abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Thus, the claim is not patent eligible.
Regarding claim 6
The claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1: The claim recites the abstract idea identified above regarding claim 1.
Step 2A Prong 2: This judicial exception is not integrated into a practical application.
In particular, the claim recites an additional element (“wherein the transformed probability distribution comprises a probability distribution amplitude controlling hyper parameter determined by a grid search using held-out data”). This is a recitation of a particular type or source of data to be used in performing the abstract idea. Limiting the abstract idea to a particular type or source of data is an attempt to limit the abstract idea to a particular field of use or technological environment, which does not integrate the abstract idea into a practical application. See MPEP 2106.05(h)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
This is a recitation of a particular type or source of data to be used in performing the abstract idea. Limiting the abstract idea to a particular type or source of data is an attempt to limit the abstract idea to a particular field of use or technological environment, which does not amount to significantly more than the abstract idea. See MPEP 2106.05(h).
Regarding claim 7
The claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1: The claim recites the abstract idea identified above regarding claim 1.
Step 2A Prong 2: This judicial exception is not integrated into a practical application.
In particular, the claim recites an additional element (“wherein the transformed probability distribution comprises a probability distribution amplitude controlling hyper parameter determined by a statistic of a probability distribution of the ExternalLM”). This is a recitation of a particular type or source of data to be used in performing the abstract idea. Limiting the abstract idea to a particular type or source of data is an attempt to limit the abstract idea to a particular field of use or technological environment, which does not integrate the abstract idea into a practical application. See MPEP 2106.05(h)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
This is a recitation of a particular type or source of data to be used in performing the abstract idea. Limiting the abstract idea to a particular type or source of data is an attempt to limit the abstract idea to a particular field of use or technological environment, which does not amount to significantly more than the abstract idea. See MPEP 2106.05(h).
Regarding claim 8
The claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1:
The limitations of
“wherein the sharpness of the first probability distribution is relaxed by reducing one or more amplitudes of the first probability distribution which are greater than a threshold amount”, as drafted, are a process that, under its broadest reasonable interpretation, covers performance of the limitation based on mathematical relationships and/or mathematical formulas or equations and/or mathematical calculations. That is, nothing in the claim element precludes the step from practically being performed based on mathematical relationships and/or mathematical formulas or equations and/or mathematical calculations.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation based on mathematical relationships and/or mathematical formulas or equations and/or mathematical calculations, but for the recitation of generic computer components, then it falls within the “Mathematical concepts” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. In particular, the claim does not recite additional elements. Thus, the claim is directed to an abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Thus, the claim is not patent eligible.
Regarding claim 9
The claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: The claim recites a method; therefore, it falls into the statutory category of processes.
Step 2A Prong 1:
The limitations of
“wherein the sharpness of the first probability distribution is relaxed by reducing one or more amplitudes of the first probability distribution by a threshold amount”, as drafted, are a process that, under its broadest reasonable interpretation, covers performance of the limitation based on mathematical relationships and/or mathematical formulas or equations and/or mathematical calculations. That is, nothing in the claim element precludes the step from practically being performed based on mathematical relationships and/or mathematical formulas or equations and/or mathematical calculations.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation based on mathematical relationships and/or mathematical formulas or equations and/or mathematical calculations, but for the recitation of generic computer components, then it falls within the “Mathematical concepts” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. In particular, the claim does not recite additional elements. Thus, the claim is directed to an abstract idea.
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Thus, the claim is not patent eligible.
Regarding claim 10
The claim recites “A computer program product for fusing an end-to-end speech recognition model with an external language model (ExternalLM), the computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to perform a method comprising:” to perform precisely the method of Claim 1. As performance of an abstract idea on generic computer components (see MPEP 2106.05(f)) cannot integrate the abstract idea into a practical application nor provide significantly more than the abstract idea itself, the claim is rejected for reasons set forth in the rejection of Claim 1.
Regarding claim 11
The claim is rejected for the reasons set forth in the rejection of Claim 2 under 35 U.S.C. 101, mutatis mutandis, as reciting an abstract idea without integrating the judicial exception into a practical application nor providing significantly more than the judicial exception.
Regarding claim 12
The claim is rejected for the reasons set forth in the rejection of Claim 3 under 35 U.S.C. 101, mutatis mutandis, as reciting an abstract idea without integrating the judicial exception into a practical application nor providing significantly more than the judicial exception.
Regarding claim 13
The claim is rejected for the reasons set forth in the rejection of Claim 4 under 35 U.S.C. 101, mutatis mutandis, as reciting an abstract idea without integrating the judicial exception into a practical application nor providing significantly more than the judicial exception.
Regarding claim 14
The claim is rejected for the reasons set forth in the rejection of Claim 5 under 35 U.S.C. 101, mutatis mutandis, as reciting an abstract idea without integrating the judicial exception into a practical application nor providing significantly more than the judicial exception.
Regarding claim 15
The claim is rejected for the reasons set forth in the rejection of Claim 6 under 35 U.S.C. 101, mutatis mutandis, as reciting an abstract idea without integrating the judicial exception into a practical application nor providing significantly more than the judicial exception.
Regarding claim 16
The claim is rejected for the reasons set forth in the rejection of Claim 7 under 35 U.S.C. 101, mutatis mutandis, as reciting an abstract idea without integrating the judicial exception into a practical application nor providing significantly more than the judicial exception.
Regarding claim 17
The claim is rejected for the reasons set forth in the rejection of Claim 8 under 35 U.S.C. 101, mutatis mutandis, as reciting an abstract idea without integrating the judicial exception into a practical application nor providing significantly more than the judicial exception.
Regarding claim 18
The claim is rejected for the reasons set forth in the rejection of Claim 9 under 35 U.S.C. 101, mutatis mutandis, as reciting an abstract idea without integrating the judicial exception into a practical application nor providing significantly more than the judicial exception.
Regarding claim 19
The claim recites “A computer processing system for fusing an end-to-end speech recognition model with an external language model (ExternalLM), the computer processing system comprising: a memory device for storing program code; and a hardware processor operatively coupled to the memory device for running the program code” to perform precisely the method of Claim 1. As performance of an abstract idea on generic computer components (see MPEP 2106.05(f)) and “Storing and retrieving information in memory” (see MPEP 2106.05(g) on Insignificant Extra-Solution Activity, and MPEP 2106.05(d) on Well-Understood, Routine, Conventional Activity) cannot integrate the abstract idea into a practical application nor provide significantly more than the abstract idea itself, the claim is rejected for reasons set forth in the rejection of Claim 1.
Regarding claim 20
The claim is rejected for the reasons set forth in the rejection of Claim 2 under 35 U.S.C. 101, mutatis mutandis, as reciting an abstract idea without integrating the judicial exception into a practical application nor providing significantly more than the judicial exception.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-4, 10-13, 19-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhao et al. (Shallow-Fusion End-to-End Contextual Biasing) in view of Chan et al. (Listen, Attend and Spell)
Regarding claim 1
Zhao teaches
A computer-implemented method for fusing an end-to-end speech recognition model with an external language model (ExternalLM), the method comprising:
(Zhao [sec(s) Abs] “Contextual biasing to a specific domain, including a user’s song names, app names and contact names, is an important component of any production-level automatic speech recognition (ASR) system.” [sec(s) 1] “End-to-end (E2E) systems combine the AM, PM, and LM as a single neural network.” [sec(s) 2.1] “Given a set of acoustic observations x = (x1, . . . , xK), E2E models provide posterior probabilities for a set of subword units y = (y1, . . . , yL) given these observations, that is, P(y|x). Shallow fusion interpolates the score from the E2E model with an external contextual LM during beam-search decoding, given by (1).
PNG
media_image1.png
85
781
media_image1.png
Greyscale
(1) Here, PC(y) is the score from the contextual LM and λ is a tunable hyperparameter controlling how much the contextual LM influences the overall model score during beam search.”; e.g., eq (1) read(s) on “fusing”.)
obtaining an output of the end-to-end speech recognition model, the output being a first probability distribution;
(Zhao [sec(s) 1] “End-to-end (E2E) systems combine the AM, PM, and LM as a single neural network.” [sec(s) 2.1] “Given a set of acoustic observations x = (x1, . . . , xK), E2E models provide posterior probabilities for a set of subword units y = (y1, . . . , yL) given these observations, that is, P(y|x). Shallow fusion interpolates the score from the E2E model with an external contextual LM during beam-search decoding, given by (1).
PNG
media_image1.png
85
781
media_image1.png
Greyscale
(1) Here, PC(y) is the score from the contextual LM and λ is a tunable hyperparameter controlling how much the contextual LM influences the overall model score during beam search.”; e.g., “E2E models provide posterior probabilities for a set of subword units y = (y1, . . . , yL) given these observations, that is, P(y|x)” read(s) on “probability distribution”.)
(Note: Hereinafter, if a limitation has bold brackets (i.e. [·]) around claim languages, the bracketed claim languages indicate that they have not been taught yet by the current prior art reference but they will be taught by another prior art reference afterwards.)
[suppressing a sharpness of] the first probability distribution including peaks as outputs of the end-to-end speech recognition model above a threshold during decoding with the ExternalLM that increases accuracy of the end-to-end speech recognition model with the ExternalLM by transforming, by a hardware processor, the first probability distribution into a transformed probability distribution [with lowered amplitudes of the peaks of the first probability distribution based on a threshold] by utilizing a non-linear function for symbol posterior probabilities of alignments of outputs of the end-to-end speech recognition model obtained with the alignments having same output symbol sequences;
(Zhao [sec(s) 1] “End-to-end (E2E) systems combine the AM, PM, and LM as a single neural network. … We report results across four different contextual test sets. We find our proposed changes to the contextual FST construction lead to significant improvements in shallow-fusion based biasing compared to past work [2, 4].” [sec(s) 2.1] “We will use a similar technique to build a contextual FST, and then incorporate it into the E2E decoding framework. Given a set of acoustic observations x = (x1, . . . , xK), E2E models provide posterior probabilities for a set of subword units y = (y1, . . . , yL) given these observations, that is, P(y|x). Shallow fusion interpolates the score from the E2E model with an external contextual LM during beam-search decoding, given by (1).
PNG
media_image1.png
85
781
media_image1.png
Greyscale
(1) Here, PC(y) is the score from the contextual LM and λ is a tunable hyperparameter controlling how much the contextual LM influences the overall model score during beam search.” [sec(s) 3.2] “All RNN-T models are trained in Tensorflow [22] on 8 × 8 Tensor Processing Units (TPU) slices with a batch size of 4,096.” [sec(s) 2.4] “Then we combine the category specific prefixes and proper nouns to generate the utterance text and use a synthetic speech data generator to create training sets with roughly 1 million utterances for each category. … As an alternate, we created synthetic training datasets by generating sentences with a variety of proper nouns and then synthesizing this data, using a concatenative TTS approach with only 1 voice [15].” [sec(s) 3.1] “The supervised training set used for experiments consists of 35 million English utterances (∼ 27,500 hours). The training utterances are anonymized and hand-transcribed, and are representative of Google’s voice search traffic. This data set is created by artificially corrupting clean utterances using a room simulator, adding varying degrees of noise and reverberation such that the overall SNR is between 0dB and 30dB, with an average SNR of 12dB [16]. The noise sources are from YouTube and noisy environmental recordings.”; e.g., “E2E models provide posterior probabilities for a set of subword units y = (y1, . . . , yL) given these observations, that is, P(y|x)” read(s) on “probability distribution”. In addition, e.g., “log P(y|x)” read(s) on “transformed probability distribution” and “non-linear function”. Furthermore, e.g., a smallest output read(s) on “threshold”. Moreover, e.g., “datasets” along with “training” read(s) on “same output symbol sequences” since same output symbol sequences may be generated during training.
Examiner notes that paragraph 17 of the Instant Specification describes “As used herein, the term "relaxed sharpness" refers to reducing the peakiness of a probability output by distorting the probability distribution and lowering the amplitudes of higher probability. In embodiments of the present invention, the non-linear function used to relax the sharpness can be log() or pow(). In embodiments of the present invention, parameters can be derived from the probability distribution of the external language model (ExternalLM).”)
fusing the transformed probability distribution and an output of the ExternalLM being a second probability distribution to obtain a fused model of the end-to-end speech recognition model and the ExternalLM for decoding speech; and
(Zhao [sec(s) Abs] “Contextual biasing to a specific domain, including a user’s song names, app names and contact names, is an important component of any production-level automatic speech recognition (ASR) system.” [sec(s) 1] “End-to-end (E2E) systems combine the AM, PM, and LM as a single neural network.” [sec(s) 2.1] “We will use a similar technique to build a contextual FST, and then incorporate it into the E2E decoding framework. Given a set of acoustic observations x = (x1, . . . , xK), E2E models provide posterior probabilities for a set of subword units y = (y1, . . . , yL) given these observations, that is, P(y|x). Shallow fusion interpolates the score from the E2E model with an external contextual LM during beam-search decoding, given by (1).
PNG
media_image1.png
85
781
media_image1.png
Greyscale
(1) Here, PC(y) is the score from the contextual LM and λ is a tunable hyperparameter controlling how much the contextual LM influences the overall model score during beam search.”; e.g., eq (1) read(s) on “fusing”. In addition, e.g., “log P(y|x)” read(s) on “transformed probability distribution”. Furthermore, e.g., “PC(y)” read(s) on “output of the ExternalLM being a second probability distribution”.)
performing speech recognition tasks with the fused model to control computing systems that communicate with users.
(Zhao [sec(s) Abs] “Contextual biasing to a specific domain, including a user’s song names, app names and contact names, is an important component of any production-level automatic speech recognition (ASR) system.” [sec(s) 1] “End-to-end (E2E) systems combine the AM, PM, and LM as a single neural network.” [sec(s) 2.1] “We will use a similar technique to build a contextual FST, and then incorporate it into the E2E decoding framework. Given a set of acoustic observations x = (x1, . . . , xK), E2E models provide posterior probabilities for a set of subword units y = (y1, . . . , yL) given these observations, that is, P(y|x). Shallow fusion interpolates the score from the E2E model with an external contextual LM during beam-search decoding, given by (1).
PNG
media_image1.png
85
781
media_image1.png
Greyscale
(1) Here, PC(y) is the score from the contextual LM and λ is a tunable hyperparameter controlling how much the contextual LM influences the overall model score during beam search.” [sec(s) 3] “The Songs test set contains media requests (e.g. play rihanna music) with biasing phrases containing popular songs and artist names in US-English. The Cnt-TTS test set contains communication requests (e.g. call John mobile) with biasing phrases containing popular USEnglish names. Finally, the Apps test set contains requests to interact with an app (e.g. open trivia game) with biasing phrases containing popular app names. Noise is artificially added to the synthetic data, similar to [16]. Note that none of the synthetic test sets have phrases that appear in the TTS synthesized training data.”;)
However, Zhao does not appear to explicitly teach:
[suppressing a sharpness of] the first probability distribution including peaks as outputs of the end-to-end speech recognition model above a threshold during decoding with the ExternalLM that increases accuracy of the end-to-end speech recognition model with the ExternalLM by transforming, by a hardware processor, the first probability distribution into a transformed probability distribution [with lowered amplitudes of the peaks of the first probability distribution based on a threshold] by utilizing a non-linear function for symbol posterior probabilities of alignments of outputs of the end-to-end speech recognition model obtained with the alignments having same output symbol sequences;
(Note: Hereinafter, if a limitation has one or more bold underlines, the one or more underlined claim languages indicate that they are taught by the current prior art reference, while the one or more non-underlined claim languages indicate that they have been taught already by one or more previous art references.)
Chan teaches
suppressing a sharpness of the first probability distribution including peaks as outputs of the end-to-end speech recognition model above a threshold during decoding with the ExternalLM that increases accuracy of the end-to-end speech recognition model with the ExternalLM by transforming, by a hardware processor, the first probability distribution into a transformed probability distribution with lowered amplitudes of the peaks of the first probability distribution based on a threshold by utilizing a non-linear function for symbol posterior probabilities of alignments of outputs of the end-to-end speech recognition model obtained with the alignments having same output symbol sequences;
(Chan [sec(s) 3.4 Decoding and Rescoring] “Decoding is performed with a simple left-to-right beam search algorithm similar to [17]. … We have vast quantities of text data [36], compared to the amount of transcribed speech utterances. We can use language models trained on text corpora alone similar to conventional speech systems [37]. To do so we can rescore our beams with the language model. We find that our model has a small bias for shorter utterances so we normalize our probabilities by the number of characters |y|c in the hypothesis and combine it with a language model probability PLM(y):
PNG
media_image2.png
157
1032
media_image2.png
Greyscale
(16) where λ is our language model weight and can be determined by a held-out validation set.” [sec(s) 4] “We used the DistBelief framework [38] with 32 replicas, each with a minibatch of 32 utterances.” [sec(s) Abs] “The network produces character sequences without making any independence assumptions between the characters. This is the key improvement of LAS over previous end-to-end CTC models. On a subset of the Google voice search task, LAS achieves a word error rate (WER) of 14.1% without a dictionary or a language model, and 10.3% with language model rescoring over the top 32 beams.”; e.g., “We find that our model has a small bias for shorter utterances so we normalize our probabilities by the number of characters |y|c in the hypothesis” read(s) on “suppressing a sharpness of the first probability distribution”. In addition, e.g., “P(y|x)” read(s) on “probability distribution”, and e.g., “
PNG
media_image3.png
157
305
media_image3.png
Greyscale
” read(s) on “transformed probability distribution with lowered amplitudes of the peaks of the first probability distribution”. Furthermore, e.g., a peak before normalization read(s) on “threshold”. Moreover, e.g., “the DistBelief framework [38]” read(s) on “hardware processor” since the reference “[38]” indicates Dean et al. (“Large Scale Distributed Deep Networks”), which states “In this paper, we consider the problem of training a deep network with billions of parameters using tens of thousands of CPU cores. We have developed a software framework called DistBelief that can utilize computing clusters with thousands of machines to train large models” in [sec(s) Abs] along with fig 1. Furthermore, e.g., eq (16) along with “the DistBelief framework [38]” read(s) on “transforming, by a hardware processor, the first probability distribution into a transformed probability distribution”.
Examiner notes that paragraph 17 of the Instant Specification describes “As used herein, the term "relaxed sharpness" refers to reducing the peakiness of a probability output by distorting the probability distribution and lowering the amplitudes of higher probability.”)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Zhao with the relaxed sharpness of Chan.
One of ordinary skill in the art would have been motived to combine in order to be able to generate multiple spelling variants naturally by modeling characters as outputs to produce diverse transcripts for the same utterance, unlike a previous approach which has trouble producing such diverse transcripts for the same utterance because of conditional independence assumptions between frames.
(Chan [sec(s) 1] “Key to our approach is the fact that we use a pyramidal RNN model for the listener, which reduces the number of time steps that the attention model has to extract relevant information from. Rare and out-of-vocabulary (OOV) words are handled automatically, since the model outputs the character sequence, one character at a time. Another advantage of modeling characters as outputs is that the network is able to generate multiple spelling variants naturally. For example, for the phrase “triple a” the model produces both “triple a” and “aaa” in the top beams (see section 4.5). A model like CTC may have trouble producing such diverse transcripts for the same utterance because of conditional independence assumptions between frames.”)
Regarding claim 2
The combination of Zhao, Chan teaches claim 1.
Zhao further teaches
wherein the fusing comprising searching for a best output sequence in a decoding by applying a max function to the transformed probability distribution.
(Zhao [sec(s) Abs] “Contextual biasing to a specific domain, including a user’s song names, app names and contact names, is an important component of any production-level automatic speech recognition (ASR) system.” [sec(s) 1] “We will use a similar technique to build a contextual FST, and then incorporate it into the E2E decoding framework. End-to-end (E2E) systems combine the AM, PM, and LM as a single neural network.” [sec(s) 2.1] “Given a set of acoustic observations x = (x1, . . . , xK), E2E models provide posterior probabilities for a set of subword units y = (y1, . . . , yL) given these observations, that is, P(y|x). Shallow fusion interpolates the score from the E2E model with an external contextual LM during beam-search decoding, given by (1).
PNG
media_image1.png
85
781
media_image1.png
Greyscale
(1) Here, PC(y) is the score from the contextual LM and λ is a tunable hyperparameter controlling how much the contextual LM influences the overall model score during beam search.”; e.g., “log P(y|x)” read(s) on “transformed probability distribution”.)
Regarding claim 3
The combination of Zhao, Chan teaches claim 1.
Zhao further teaches
wherein the transforming is performed by applying a non-linear function to the first probability distribution.
(Zhao [sec(s) 1] “End-to-end (E2E) systems combine the AM, PM, and LM as a single neural network.” [sec(s) 2.1] “Given a set of acoustic observations x = (x1, . . . , xK), E2E models provide posterior probabilities for a set of subword units y = (y1, . . . , yL) given these observations, that is, P(y|x). Shallow fusion interpolates the score from the E2E model with an external contextual LM during beam-search decoding, given by (1).
PNG
media_image1.png
85
781
media_image1.png
Greyscale
(1) Here, PC(y) is the score from the contextual LM and λ is a tunable hyperparameter controlling how much the contextual LM influences the overall model score during beam search.” [sec(s) 3.2] “All RNN-T models are trained in Tensorflow [22] on 8 × 8 Tensor Processing Units (TPU) slices with a batch size of 4,096.”; e.g., “log P(y|x)” read(s) on “applying a non-linear function to the first probability distribution”.)
Regarding claim 4
The combination of Zhao, Chan teaches claim 1.
Zhao further teaches
wherein the transforming is performed by applying a logarithmic function to the first probability distribution.
(Zhao [sec(s) 1] “End-to-end (E2E) systems combine the AM, PM, and LM as a single neural network.” [sec(s) 2.1] “Given a set of acoustic observations x = (x1, . . . , xK), E2E models provide posterior probabilities for a set of subword units y = (y1, . . . , yL) given these observations, that is, P(y|x). Shallow fusion interpolates the score from the E2E model with an external contextual LM during beam-search decoding, given by (1).
PNG
media_image1.png
85
781
media_image1.png
Greyscale
(1) Here, PC(y) is the score from the contextual LM and λ is a tunable hyperparameter controlling how much the contextual LM influences the overall model score during beam search.” [sec(s) 3.2] “All RNN-T models are trained in Tensorflow [22] on 8 × 8 Tensor Processing Units (TPU) slices with a batch size of 4,096.”; e.g., “log P(y|x)” read(s) on “applying a logarithmic function to the first probability distribution”.)
Regarding claim 10
The claim is a computer program product claim corresponding to the method claim 1, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the method claim.
Note that Zhao teaches “computer readable storage medium”, “computer” and “hardware processor”.
(Zhao [sec(s) 3.2] “All RNN-T models are trained in Tensorflow [22] on 8 × 8 Tensor Processing Units (TPU) slices with a batch size of 4,096.”;)
Regarding claim 11
The claim is a computer program product claim corresponding to the method claim 2, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the method claim.
Regarding claim 12
The claim is a computer program product claim corresponding to the method claim 3, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the method claim.
Regarding claim 13
The claim is a computer program product claim corresponding to the method claim 4, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the method claim.
Regarding claim 19
The claim is a system claim corresponding to the method claim 1, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the method claim.
Note that Zhao teaches “the computer processing system comprising: a memory device for storing program code; and a hardware processor operatively coupled to the memory device for running the program code to”.
(Zhao [sec(s) 3.2] “All RNN-T models are trained in Tensorflow [22] on 8 × 8 Tensor Processing Units (TPU) slices with a batch size of 4,096.”;)
Regarding claim 20
The claim is a system claim corresponding to the method claim 2, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the method claim.
Claim(s) 5, 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhao et al. (Shallow-Fusion End-to-End Contextual Biasing) in view of Chan et al. (Listen, Attend and Spell) in view of AGARWAL et al. (US 20190080225 A1)
Regarding claim 5
The combination of Zhao, Chan teaches claim 1.
Zhao further teaches
wherein the transforming is performed by applying a [power] function to the first probability distribution.
(Zhao [sec(s) 1] “End-to-end (E2E) systems combine the AM, PM, and LM as a single neural network.” [sec(s) 2.1] “Given a set of acoustic observations x = (x1, . . . , xK), E2E models provide posterior probabilities for a set of subword units y = (y1, . . . , yL) given these observations, that is, P(y|x). Shallow fusion interpolates the score from the E2E model with an external contextual LM during beam-search decoding, given by (1).
PNG
media_image1.png
85
781
media_image1.png
Greyscale
(1) Here, PC(y) is the score from the contextual LM and λ is a tunable hyperparameter controlling how much the contextual LM influences the overall model score during beam search.” [sec(s) 3.2] “All RNN-T models are trained in Tensorflow [22] on 8 × 8 Tensor Processing Units (TPU) slices with a batch size of 4,096.”;)
However, the combination of Zhao, Chan does not appear to explicitly teach:
wherein the transforming is performed by applying a [power] function to the first probability distribution.
AGARWAL teaches
wherein the transforming is performed by applying a power function to the first probability distribution.
(AGARWAL [fig(s) 4] [par(s) 39] “With a view to force the network to learn better separation of the embeddings (query embeddings), the above loss may be increased slightly for all predictions, i.e., irrespective of whether the prediction is right or wrong. For this, a square-root of all the probabilities in the prediction distribution Pi and then re-normalize to obtain the new probability distribution Qi. Qi has higher entropy than Pi, as depicted in FIG. 4. More specifically, FIG. 4 is a graphical representation illustrating a predicted Probability Distribution (P), new probability distribution obtained after square-root and normalization of P, and T is the target distribution in accordance with an embodiment of the present disclosure. As can be seen from FIG. 4, probability of high likely classes reduces, and the probability of low likely classes increases slightly. Instead of using the standard categorical_crossentropy loss, KLD(Ti∥Qi) which in the case of a deep network, this is equivalent to scaling the activations input to the final softmax layer by half.”;)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Zhao, Chan with the power function of AGARWAL.
One of ordinary skill in the art would have been motived to combine in order to help achieve better accuracy on BiLSTM classification as well as when attached to Siamese network iteratively.
(AGARWAL [par(s) 39] “As it can be observed from the evaluation results presented in Tables 1, 2 and 3, this proposed approach helps achieve better accuracy on BiLSTM classification as well as when attached to Siamese network iteratively (explained later in this section). This suggests that such an artificial increase of loss helps with better separation of the query embeddings. A similar technique was used by a conventional approach wherein the conventional approach took square of the predicted distribution and assumed it as auxiliary target distribution for clustering in unsupervised setting, while embodiments of the present disclosure and the proposed approach take square-root of the predicted distribution and use it to increase the loss, in the context of classification.”)
Regarding claim 14
The claim is a computer program product claim corresponding to the method claim 5, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the method claim.
Claim(s) 6, 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhao et al. (Shallow-Fusion End-to-End Contextual Biasing) in view of Chan et al. (Listen, Attend and Spell) in view of Meng et al. (MIXSPEECH: DATA AUGMENTATION FOR LOW-RESOURCE AUTOMATIC SPEECH RECOGNITION) in view of Unni et al. (COUPLED TRAINING OF SEQUENCE-TO-SEQUENCE MODELS FOR ACCENTED SPEECH RECOGNITION)
Regarding claim 6
The combination of Zhao, Chan teaches claim 1.
However, the combination of Zhao, Chan does not appear to explicitly teach:
wherein the transformed probability distribution comprises a probability distribution amplitude controlling hyper parameter determined by a grid search using held-out data.
Meng teaches
wherein the transformed probability distribution comprises a probability distribution amplitude controlling hyper parameter determined by a grid search using [held-out] data.
(Meng [fig(s) 1] [table(s) 3] “
PNG
media_image4.png
116
541
media_image4.png
Greyscale
” [sec(s) 2-2.3] “The Cross-Entropy loss can be written as:
PNG
media_image5.png
147
961
media_image5.png
Greyscale
, (4) where yu is the u-th target token. In LAS [4], the encoder and decoder are the stacked BiLSTM and LSTM respectively. In Transformer based structure [20], they are stacked of multi-head attention and feed-forward network. During training, following [21], we leverage the multi-task learning by combining CTC and Cross-Entropy loss as follows:
PNG
media_image6.png
70
868
media_image6.png
Greyscale
, (5) where β ∈ [0, 1] is a tunable hyper-parameter. By combining with MixSpeech in Eq. 2, the final training objective can be written as:
PNG
media_image7.png
73
1459
media_image7.png
Greyscale
. (6)” [sec(s) 3.4] “Both LAS and Transformer based models leverage multi-task learning to boost the performance and β in Equation 5 is a hyper-parameter to adjust the weight of them. When β is set to 0 or 1, there is only Cross-Entropy loss or CTC loss. We vary β and get the results of the baseline as shown in Table 3. The results show that the model cannot perform well with only one target and requires alignment information guided by CTC loss to help the attention decoder. In Table 3, β = 0.3 gets the lowest PER score, and thus is set as default in the baselines and models enhanced with MixSpeech.” [sec(s) 1] “MixSpeech is much simple with only a single hyper-parameter (the combination weight λ), unlike the complicated hyper-parameters used in SpecAugment.”; e.g., Table 3 along with “We vary β and get the results of the baseline as shown in Table 3” read(s) on “grid search” since the search is done for a grid of a single parameter.
Examiner notes that paragraph 58 of the Instant Specification describes “The parameter Ɣ is a hyper parameter to control the magnitude of the amplitude of the distribution and can be determined by: (i) grid search using held-out data; and (ii) statistics of the probability distribution of the ExtemalLM (e.g., average mean, variance, kurtosis, and skewness). Held-out data is data used only for training (and not testing).”)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Zhao, Chan with the probability distribution amplitude controlling hyper parameter of Meng.
One of ordinary skill in the art would have been motived to combine in order to improve the recognition accuracy compared with the baseline model and previous data augmentation method SpecAgument, and achieve competitive WER performance on WSJ dataset.
(Meng [sec(s) 4] “In this paper, we proposed MixSpeech, a new data augmentation method to apply the mixup technique on ASR tasks for low-resource scenarios. Experimental results show that our method improves the recognition accuracy compared with the baseline model and previous data augmentation method SpecAgument, and achieves competitive WER performance on WSJ dataset.”)
However, the combination of Zhao, Chan, Meng does not appear to explicitly teach:
wherein the transformed probability distribution comprises a probability distribution amplitude controlling hyper parameter determined by a grid search using [held-out] data.
Unni teaches
wherein the transformed probability distribution comprises a probability distribution amplitude controlling hyper parameter determined by a grid search using held-out data.
(Unni [fig(s) 1] [sec(s) 3] “The training objective to be maximized in LAS, with Pr(yjx) defined as in Eq. (1), is:
PNG
media_image8.png
64
584
media_image8.png
Greyscale
where y*1:j-1 corresponds to the ground-truth of previous characters. … Since the character sequence y is identical, both utterances will produce the same number of decoding time-steps. We introduce an L2 loss, Lpair, across the context vectors:
PNG
media_image9.png
138
545
media_image9.png
Greyscale
We now optimize a linear combination of the LAS and coupled objectives:
PNG
media_image10.png
58
632
media_image10.png
Greyscale
where
PNG
media_image11.png
129
451
media_image11.png
Greyscale
is a tunable hyperparameter.” [sec(s) 4] “The scaling factor
PNG
media_image12.png
96
67
media_image12.png
Greyscale
for the coupled loss Lpair was tuned on a held-out dataset and set to 0.0001. We used 150 sub-word units and a scheduled sampling rate of 0.3.”;)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Zhao, Chan with the probability distribution amplitude controlling hyper parameter with held-out data of Unni.
One of ordinary skill in the art would have been motived to combine in order to show significant improvements in WERs across diverse accented samples and data settings.
(Unni [sec(s) 6] “In this work, we proposed a new coupled training paradigm which imposes an L2 regularization between the context vectors for two utterances with the same text. We showed significant improvements in WERs across diverse accented samples and data settings.”)
Regarding claim 15
The claim is a computer program product claim corresponding to the method claim 6, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the method claim.
Claim(s) 7, 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhao et al. (Shallow-Fusion End-to-End Contextual Biasing) in view of Chan et al. (Listen, Attend and Spell) in view of Bai et al. (Integrating Knowledge into End-to-End Speech Recognition from External Text-Only Data)
Regarding claim 7
The combination of Zhao, Chan teaches claim 1.
However, the combination of Zhao, Chan does not appear to explicitly teach:
wherein the transformed probability distribution comprises a probability distribution amplitude controlling hyper parameter determined by a statistic of a probability distribution of the ExternalLM.
Bai teaches
wherein the transformed probability distribution comprises a probability distribution amplitude controlling hyper parameter determined by a statistic of a probability distribution of the ExternalLM.
(Bai [fig(s) 1-2] [algorithm(s) 1] “Hyper parameters 𝜆 and 𝑇.” [sec(s) II] “The parameters of the AED model are trained with cross-entropy loss:
PNG
media_image13.png
192
1438
media_image13.png
Greyscale
. (4) where (𝑋(𝑛), 𝑌(𝑛)) is the 𝑛-th pair of data in the corpus.” [sec(s) III] “In this section, we introduce the proposed LST training method. As shown in Fig. 1, the knowledge in the large-scale external text-only data is represented as an LM. Then, the knowledge is transferred to the AED model. Specifically, the LM which is trained on external text-only data provides probability distribution corresponding to each token in the transcription. And the Kullback-Leibler divergence (KLD) between the decoder of the AED model and the LM is minimized. Here, the LM is a general model to estimate the probability of a token in the sentence, given the context. This makes the method more flexible. We denote the probability of 𝑗-th token in the transcription given by the LM as 𝑃𝐿𝑀(𝑦𝑗|𝐶), where 𝐶 represents the context. The KLD between the decoder of the AED model and the LM for 𝑗-th token is:
PNG
media_image14.png
185
1552
media_image14.png
Greyscale
, (5) where S is the vocabulary with size 𝑀. Because the parameters of the LM are not updated when we train the AED model, the KLD can be simplified to cross-entropy. The LST loss is defined as:
PNG
media_image15.png
314
1216
media_image15.png
Greyscale
(6) where 𝐻 (𝑛) 𝑗 is the cross-entropy for the token at 𝑗-th position in 𝑛-th sample of the corpus, and 𝐶 (𝑛) is a part of the transcription in 𝑛-th sample as the context for the LM. … To combine the knowledge of the groundtruth of the transcription and the knowledge from the LM, 𝐿𝐶𝐸 (𝜃) and 𝐿𝐿𝑆𝑇 (𝜃) are added together:
PNG
media_image16.png
64
974
media_image16.png
Greyscale
(7) where 𝜆 ∈ [0, 1] is a weight to balance 𝐿𝐶𝐸 and 𝐿𝐿𝑆𝑇.”; e.g., “Hyper parameters 𝜆” along with eq (7) read(s) on “probability distribution amplitude controlling hyper parameter”. In addition, e.g., “Kullback-Leibler divergence (KLD) between the decoder of the AED model and the LM” and “cross-entropy” along with eq (5) read(s) on “a statistic of a probability distribution of the ExternalLM” since “KLD can be simplified to cross-entropy” indicates a statistic of a similarity of probability distributions.)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Zhao, Chan with the probability distribution amplitude controlling hyper parameter of Bai.
One of ordinary skill in the art would have been motived to combine in order to demonstrate the effectiveness of leveraging external text-only data and the whole context in a sentence with our proposed method, compared with baseline hybrid systems and AED model based systems.
(Bai [sec(s) Abs] “The experimental results demonstrate the effectiveness of leveraging external text-only data and the whole context in a sentence with our proposed method, compared with baseline hybrid systems and AED model based systems.”)
Regarding claim 16
The claim is a computer program product claim corresponding to the method claim 7, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the method claim.
Claim(s) 8, 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhao et al. (Shallow-Fusion End-to-End Contextual Biasing) in view of Chan et al. (Listen, Attend and Spell) in view of Yu et al. (US 9628499 B1)
Regarding claim 8
The combination of Zhao, Chan teaches claim 1.
Chan further teaches
wherein the sharpness of the first probability distribution is relaxed by reducing one or more amplitudes of the first probability distribution [which are greater than a threshold amount].
(Chan [sec(s) 3.4] “We have vast quantities of text data [36], compared to the amount of transcribed speech utterances. We can use language models trained on text corpora alone similar to conventional speech systems [37]. To do so we can rescore our beams with the language model. We find that our model has a small bias for shorter utterances so we normalize our probabilities by the number of characters |y|c in the hypothesis and combine it with a language model probability PLM(y):
PNG
media_image2.png
157
1032
media_image2.png
Greyscale
(16) where λ is our language model weight and can be determined by a held-out validation set.”; e.g., “We find that our model has a small bias for shorter utterances so we normalize our probabilities by the number of characters |y|c in the hypothesis” read(s) on “sharpness of the first probability distribution is relaxed”. Moreover, e.g., “P(y|x)” read(s) on “first probability distribution”.
Examiner notes that paragraph 17 of the Instant Specification describes “As used herein, the term "relaxed sharpness" refers to reducing the peakiness of a probability output by distorting the probability distribution and lowering the amplitudes of higher probability.”)
The combination of Zhao, Chan is combinable with Chan for the same rationale as set forth above with respect to claim 1.
However, the combination of Zhao, Chan does not appear to explicitly teach:
wherein the sharpness of the first probability distribution is relaxed by reducing one or more amplitudes of the first probability distribution [which are greater than a threshold amount].
Yu teaches
wherein the sharpness of the first probability distribution is relaxed by reducing one or more amplitudes of the first probability distribution which are greater than a threshold amount.
(Yu [fig(s) 4] [col 6 ln 47– col 8 ln 51] “At step 204, the processor 106 estimates a historical probability distribution corresponding to previously received samples of the signal. In a residual signal, each sample may be modeled with the same random variable, such that the samples are identically distributed. The probability distribution of the samples may be estimated using probability distribution estimation methods. In particular, estimating a historical probability distribution as depicted at step 204 may include invoking a call to an estimate historical probability distribution function, as depicted in FIG. 4. As is described in more detail in relation to FIG. 4, outlier samples may be identified and removed from the signal, and a histogram may be generated based on the remaining samples. The histogram may then be extended (by using a parametric function to estimate the tails of the histogram) and normalized to result in an estimated probability distribution.” [col 10 ln 66– col 11 ln 30] “An outlier may correspond to a sample in the set of data whose value exceeds a number of estimated standard deviations from an estimated mean. In particular, a range of values may be identified based on the estimated mean and the estimated standard deviation, and any values falling outside of the identified range may be labeled as outliers.”;)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Zhao, Chan with the probability distribution greater than a threshold amount of Yu.
One of ordinary skill in the art would have been motived to combine in order to accurately identify anomalies in network traffic patterns with low false alarm rates in order to avoid or prevent attacks.
(Yu [col 1 ln 11-col 1 ln 43] “Thus, network administrators need to proactively identify anomalies in order to avoid or prevent attacks. It is therefore important for managers of successful networks to accurately identify anomalies in network traffic patterns with low false alarm rates. Systems and methods to accurately detect anomalies would therefore be of great benefit in data analysis”)
Regarding claim 17
The claim is a computer program product claim corresponding to the method claim 8, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the method claim.
Claim(s) 9, 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhao et al. (Shallow-Fusion End-to-End Contextual Biasing) in view of Chan et al. (Listen, Attend and Spell) in view of Senior et al. (US 20170011738 A1)
Regarding claim 9
The combination of Zhao, Chan teaches claim 1.
Chan further teaches
wherein the sharpness of the first probability distribution is relaxed by reducing one or more amplitudes of the first probability distribution [by a threshold amount].
(Chan [sec(s) 3.4] “We have vast quantities of text data [36], compared to the amount of transcribed speech utterances. We can use language models trained on text corpora alone similar to conventional speech systems [37]. To do so we can rescore our beams with the language model. We find that our model has a small bias for shorter utterances so we normalize our probabilities by the number of characters |y|c in the hypothesis and combine it with a language model probability PLM(y):
PNG
media_image2.png
157
1032
media_image2.png
Greyscale
(16) where λ is our language model weight and can be determined by a held-out validation set.” [sec(s) 4] “We used the DistBelief framework [38] with 32 replicas, each with a minibatch of 32 utterances.”; e.g., “We find that our model has a small bias for shorter utterances so we normalize our probabilities by the number of characters |y|c in the hypothesis” read(s) on “sharpness of the first probability distribution is relaxed”. Moreover, e.g., “P(y|x)” read(s) on “first probability distribution”.
Examiner notes that paragraph 17 of the Instant Specification describes “As used herein, the term "relaxed sharpness" refers to reducing the peakiness of a probability output by distorting the probability distribution and lowering the amplitudes of higher probability.”)
The combination of Zhao, Chan is combinable with Chan for the same rationale as set forth above with respect to claim 1.
However, the combination of Zhao, Chan does not appear to explicitly teach:
wherein the sharpness of the first probability distribution is relaxed by reducing one or more amplitudes of the first probability distribution [by a threshold amount].
Senior teaches
wherein the sharpness of the first probability distribution is relaxed by reducing one or more amplitudes of the first probability distribution by a threshold amount.
(Senior [par(s) 53] “For example, some output distributions may indicate an extremely high confidence of one output target, such as a 99% likelihood or higher for a particular output label. Such an unbalanced distribution may be undesirable for training, since it approaches a binary decision rather than a distribution that encodes non-zero likelihoods for multiple output labels. However, the probabilities can be exponentiated and normalized to generate a less-extreme distribution that allocates a greater share of the probability to other labels. For example, each of the probability values, e.g., 0.9999, 0.00001, and so on, can each be raised to an exponent for a base of ten, e.g., 10^0.9999, 10^0.00001, and so on, and the resulting quantities can be normalized into a probability distribution. This softened probability distribution can be used as the output target for the second neural network 130 rather than the direct output of the first neural network 120. This can bring out the patterns embedded in the original distribution, with values of, e.g., 0.75, 0.1, 0.05, and so on, rather than a sharp distribution that is closer to a binary value and zeros for the rest of the values.” [par(s) 7] “These techniques can be used for training a model using the connectionist temporal classification (CTC) algorithm. As discussed below, a CTC model can be trained to indicate the presence of various phonetic units or a blank label that does not correspond to any phonetic unit. The CTC model is required to indicate the presence of each phonetic unit of an utterance, in the proper sequence” [par(s) 30] “The first neural network 120, the second neural network 130, or both may be trained using a CTC algorithm.”; e.g., a largest reduction amount for reducing “extremely high confidence of one output target, such as a 99% likelihood or higher” read(s) on “a threshold amount” since it is a maximum threshold level which is used for reducing a probability distribution.)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Zhao, Chan with the probability distribution reduction based on a threshold amount of Senior.
One of ordinary skill in the art would have been motived to combine in order to significantly reduce the amount of time required to train a neural network acoustic model, allowing the trained neural network to quickly reach a high level of accuracy.
(Senior [par 23] “The techniques described herein can provide a number of advantages and improvements. For example, the amount of time required to train a neural network acoustic model can be significantly reduced. In particular, a neural network that is trained based on the output distributions of a CTC-trained neural network can obtain the performance of a CTC network without the time-consuming and resource intensive process of CTC training. A simpler cross-entropy loss training technique can transfer the knowledge learned by one neural network to another, allowing the trained neural network to quickly reach a high level of accuracy.”)
Regarding claim 18
The claim is a computer program product claim corresponding to the method claim 9, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the method claim.
Prior Art
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Zou et al. (A COMPARABLE STUDY OF MODELING UNITS FOR END-TO-END MANDARIN SPEECH RECOGNITION) teaches a power function to the probability distribution in
PNG
media_image17.png
56
1173
media_image17.png
Greyscale
.
Toshniwal et al. (MULTILINGUAL SPEECH RECOGNITION WITH A SINGLE END-TO-END MODEL) teaches a multilingual model based on
PNG
media_image18.png
80
732
media_image18.png
Greyscale
. In addition, it teaches a probability distribution reduction based on a threshold amount (e.g., “We restricted ourselves to these values because for a very large λ, the language ID prediction task would dominate the primary task of ASR, while for a very small λ the additional task would have no effect on the training loss” along with {0.1, 0.01} for an empirically determined weight λ.)
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SEHWAN KIM whose telephone number is (571)270-7409. The examiner can normally be reached Mon - Thu 7:00 AM - 5:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michael J Huntley can be reached on (303) 297-4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SEHWAN KIM/Examiner, Art Unit 2129 4/15/2026