DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
In response to the office action from 4/22/2026, the applicant has submitted an amendment, filed 6/15/2026, amending claims 1, 8, 11, and 18, adding claim 20, while arguing to traverse the prior art and 101 rejections. Applicant’s arguments have been fully considered but are moot with respect to new grounds of rejections further in view of Fedorov et al. (US 2023/0229921) mandated by the latest amendments, and for the reasons explained in the response to arguments.
Response to Arguments
In what follows, applicant’s arguments and comments are addressed in the order presented, but similar arguments are grouped together.
Following a broad overview of the latest amendments on page 5 item “1”, item “2” discusses the last office action objections.
Due to the latest amendments, the said objection is withdrawn.
Section “3” paragraphs 1-2 on page 5 primarily are devoted to trivia pertaining to some broad MPEP teachings pertaining to subject matter eligibility, except that in paragraph two LINES 2+ it is recited that: “The examiner has alleged that application of a “sparsity mask” to “weights” “can be done by a human receiving or gathering some data of audio transcript pairs and deciding which ones to choose” and that for quantization “the human can assign a series of bits.” “Applicant submits that this characterization fundamentally mischaracterizes the claimed operations”. Presumably the reason for this assertion is: “A human mind is not equipped to perform native integer matrix multiplications on a neural network model, jointly prune and quantize billions of weights, or map pruned weights to zero-point of symmetric quantization” (page 6 lines 5+).
Respectfully as an initial matter the way the claims are drafted, neither expressly and nor inherently one can deduce dealing with “billions of weights”. Even if one assumes this to be inherently implied, merely mentioning expressly there are “billions of weights”, does not necessarily make the claim eligible. What makes them eligible, is the solution presented by the claim in dealing with an immense amount of data that can make it patent eligible and its impact on the e.g. processors and/or hardware and software used in that management.
Secondly performance of “integer matrix multiplication” itself is also not patent eligible, as it is directed to “Mathematical concept” type of abstract ideas.
Thirdly respectfully a human can easily map a “weight” to “zero” (“zero-point” “quantization”) if he judges it to be insignificant, in which can the choice of “symmetric quantization” reduces to an unnecessary additional element.
On page 6 he 2nd paragraph last 3 lines argument directed at the latest amendment is presented: i.e., “joint pruning and quantization with zero-point mapping on a pre-trained ASR model --- go beyond abstract mathematical concepts and are not operations that can practically be performed in the human mind”.
As regards the “quantiz[ation]”, based on the amplitude levels and thresholds the human can assign a series of bits, i.e., by simply converting the magnitude of the amplitude from base 10 to base 2, where every entry of one in the resulting amplitude in base 2 is associated with a bit. Given weights, one can likewise quantize each weight, and pruning them amounts to filtering undesired ones to zero. As a result, since the “pruning” is based on the “quantiz[ation]”, doing the “pruning” is contingent upon the “quantiz[ation]” which implies they are done jointly. Furthermore, the choice of “symmetric quantization” amounts to “insignificant extra-solution activity does not amount to an inventive concept, particularly when the activity is well-understood or conventional. Parker v. Flook, 437 U.S. 584, 588-89, 198 USPQ 193, 196 (1978).” That is so since these limitations pertaining to quantization and pruning as disclosed above do not preclude a human from performing them without the “symmetric quantization” which is at the core of the claimed invention.
Section “4” (pages 6-7) provide arguments directed at the latest amendments.
Please visit the new office action further in view of Fedorov et al. (US 2023/0229921).
Section “5” (pages 7-8) argue that the dependent claims 8, 18, 9, 10, and 19 “are at least patentable for depending from a non-obvious base claim”.
Since applicants have not argued the merits of these dependent claims, but assert patentability solely through their dependence on the allegedly patentable parent claims, they stand or fall with said parent claims and hence no further response to applicant’s arguments is necessary.
Finally, item “7” on page 8 is devoted to the “Double Patenting”, and it is concluded: “A terminal disclaimer” “will be filed upon resolution of all other issues in this case, if at that time the referenced claim is still co-pending and this rejection has not been withdrawn”.
The rejection is thus maintained.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 stand rejected:
The independent claims 1 and 11 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claims concern simply how to “prune” “weights” associated with ASR “training samples” (to be used for speech recognition) comprised of “speech utterance” samples and their associated “respective textual” “representations”. It begins by “obtaining” certain “plurality of training samples”. It then applies a “sparsity mask” (e.g. a “binary mask”) to each “weight” in which the following criterion is used: e.g., for a “weight” whose absolute value is above a threshold it gets weighted by “1”, and if it is below the threshold, it gets weighted by “0”; i.e., see Eq. 10 ¶ 0060 and also Fig. 5. Thereafter each weight is “quantized” by a “fixed-bit width” (e.g. Par. 0009 S1: “fixed-bit width is four”). This is then “provide[ed] as a “fine-tuned ASR model to a user device”.
These limitations under their broadest reasonable interpretations, cover performance of the limitations in the mind but for the recitation of generic computer component, i.e., “data processing hardware” (claims 1 and 11), “memory hardware” (claim 11). Therefore, other than reciting by the “data processing hardware, nothing in the claim limitations precludes the limitations from practically being performed in the mind. For example, application of a “sparsity mask” to “weights” can be done by a human receiving or gathering some data of audio transcript pairs and deciding which ones to choose. As regards to the “quantiz[ation]” based on the amplitude levels and thresholds the human can assign a series of bits. This can be done if the weights are known. Furthermore, the claims are silent on what if anything is done with the “obtain[ed]” “training samples”. They only appear to do pruning on the weights for which the model is already known and then quantizing each weight into some fixed width. As a result, the operations of “quantiz[aton]” and “pruning” (i.e., by simply filtering which weights to choose following weight quantization) can be done by a human, and since the “pruning” is based on the “quantiz[ation]”, doing the “pruning” is contingent upon the “quantiz[ation]” which implies they are done jointly. Furthermore “pruning” implies basically filtering or zeroing a weight which maps to the claim’s “weights” “set to zero directly to a zero-point of symmetric quantization”, and the choice of “symmetric quantization” amounts to “insignificant extra-solution activity does not amount to an inventive concept, particularly when the activity is well-understood or conventional. Parker v. Flook, 437 U.S. 584, 588-89, 198 USPQ 193, 196 (1978).”
The judicial exception is not integrated into a practical application. In particular, the claims recite only one additional element of “data processing hardware” to be responsible for the “obtaining …”, “fine-tuning …”, “pruning …”, “quantizing …” and “providing …” step. As such the “data processing hardware” in all these steps is recited at a high-level of generality (i.e., as a generic computer component performing generic computer functions) such that it amounts no more than mere instructions to apply the exception using the generic computer component. See also specification ¶ 0021: “data processing hardware, such as graphics processing units (GPUs) and tensor processing units (TPUs)” which as examples well known brands. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claims are thus directed to an abstract idea.
The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of using the “data processing hardware” to perform the steps of “obtaining …”, “fine-tuning …”, “pruning …”, “quantizing …” and “providing …” amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The claims are therefore not patent eligible.
Regarding claims 2, and 12, and/or 3 and 13, application of a “sparsity mask” of “binary” type to “weights” can be done by a human receiving or gathering some data of audio transcript pairs and deciding which ones to choose.
Regarding claims 4 and 14, application of a “sparsity mask” to a “plurality of” “weights” can be done by a human receiving or gathering some data of audio transcript pairs and deciding which ones to choose based on non-zero values only.
Regarding claims 5 and 15, the “quantiz[ation]” based on the amplitude levels and thresholds the human can assign a series of bits and in particular choose 4 bits for each e.g. amplitude.
Regarding claims 6 and 16, “quantiz[ation]” based on the amplitude levels and thresholds the human can assign a series of bits and in particular choose 4 bits for each e.g. amplitude.
Regarding claims 7 and 17, the “quantiz[ation]” based on the amplitude levels and thresholds the human can assign a series of bits and in particular choose 2 bits for each e.g. smaller amplitudes.
Regarding claims 8 or 18, the choice of “asymmetric quantization” and “sub-channel quantization” amount to defining specific parameters of a model and what it should comprise, where the “quantization” itself is determined that can be done by a human who can assign a series of bits based on the amplitude levels and thresholds.
Regarding claims 9-10 and 19-20, the choice of “multi-head attention layers” amount to defining specific parameters of a model and what they should comprise, where the multi-head attention layers are well known additional elements.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-3, 5-7, 11-13, 15-17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lai et al. (US 2023/0360642), in view of Afzal (US Patent 11,520,561), and further in view of Fedorov et al. (US 2023/0229921).
Regarding claim 1, Lai et al. do teach a computer-implemented method when executed on data processing hardware (¶ 0066 lines 5+: “programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer”)
causes the data processing hardware to perform operations comprising:
obtaining a plurality of training samples (¶ 0078 S1+: “Consider the low-resource ASR problem, where there is only a small, transcribed training set (x,y) Ɛ Dl” (obtaining a plurality of training samples)) ,
each respective training sample of the plurality of training samples comprising:
a respective speech utterance (¶ 0078 S1+: “Here x represents input audio” (a respective speech utterance in the training samples)); and
a respective textual utterance representing a transcription of the respective speech utterance (¶ 0078 S1+: “and y represents output transcription” (a respective utterance representing a transcription of the respective speech utterance));
fine-tuning, using quantization and sparsity aware training with 0” (based on “x” (training samples) a plurality of weights are determined as a pre-trained ASR process) “f(x;θ0) is then finetuned” (which are subsequently fine tuned where the “SSL” (“Self-supervised learning (SSL) speech model”) also known as “Wave2vec2” (¶ 0032) according to ¶ 0113 “consists of three modules” e.g. “a quantization module” (using quantization) and furthermore the “SSL” model according to step “204” in Fig. 2 undergoes a “TARGET SPARSITY AND INITIAL PRUNING MASK” (uses sparsity)),
the fine- tuning comprising:
pruning one or more weights of the plurality of weights using a sparsity mask (¶ 0079 lines 3+: “a binary pruning mask” (using a sparsity mask for pruning) “for the pre-trained weights θ0” (one or more weights of the plurality of weights)) ; and
and
providing the fine-tuned ASR model to a user device (¶ 0028 lines 5+: “program 150 may implement” (which can “resid[e] on any” “computing device” (e.g., a user device (¶ 0028 page 3 lines 1+)) “finetune” (provide the finetuned ASR model) “the initial subnetwork”).
Lai et al. do not specifically disclose:
Native integer operations;
quantizing each weight of the plurality of weights based on an integer with a fixed-bit width.
Afzal does teach:
Native integer operations; quantizing each weight of the plurality of weights based on an integer with a fixed-bit width (Col. 16 lines 43+: “a symmetric linear quantization” (quantizing) “scheme is used to map 32-bit floating points to 8-bit integers” (with fixed bit width which results in integer output or help generating integer operations) “as follows: “FP32(T)=scale-factor(sf)*8-bit Tensor” “The scaling factor” “can correspond to” “a weight scale factor” (weight) and is applied to “speech recognition”).
It would have therefore been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the “symmetric quantization” techniques of Afzal into the “quantization module” of Lai et al. would enable the combined systems and their associated methods to perform in combination as they do separately and to further enable Lai et al. to have its “weights” “represented using 8-bit integers” which will result in “Reducing the number of bits used to represent weights” and thus help “perform[ing] faster” as disclosed in Afzal Col. 16 lines 37-42.
Lai et al. in view of Afzal do not specifically disclose:
Jointly pruning and quantizing the plurality of weights using a prune-and-quantize approach
Sets pruned weights to zero
Mapping the pruned weights that are set to zero directly to a zero-point of symmetric quantization.
Fedorov et al. do teach:
Jointly pruning and quantizing the plurality of weights using a prune-and-quantize approach (¶ 0095 S2: “The weights” (weights) “for the neural network layers that have been pruned” (are pruned) “and quantized” (and jointly quantized)),
Sets pruned weights to zero (¶ 0094 last lines: “the plurality of weights are pruned by setting each quantized offset weight” (jointly pruning and quantization each weight) “value with the range of weight values to be pruned to zero” (wherein pruned weights are set to zero)),
Mapping the pruned weights that are set to zero directly to a zero-point of symmetric quantization (¶ 0042 last S: “employ” “symmetric quantization given by Equation (3)”, where the “quantization” involves “quantizing the offset weight values” (Abstract line 6), which according to ¶ 0107 “quantized offset weight values are pruned” “to zero” (“quantized offset weight values” (weights )are mapped to zero (that are pruned are set to zero-point quantization), wherein the “quantization” used is “symmetric quantization” (is symmetric quantization according to ¶ 0042 quoted above)).
It would have therefore been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the “ANN” (a neural network) technique including e.g., its “uniform symmetric quantization” of Fedorov et al. into the “quantization module” of Lai et al. in Lai et al. in view of Afzal would enable the combined systems and their associated methods to perform in combination as they do separately and to further enable Lai et al. in view of Afzal to “achieve a particular level of accuracy” by “optimizing the connection weights” as disclosed in Fedorov et al. ¶ 0026 S1.
Regarding claim 2, Lai et al. do teach the method of claim 1, wherein the sparsity mask comprises a binary mask (¶ 0079 lines 3+: “a binary pruning mask” (using a binary mask for pruning) “for the pre-trained weights θ0”).
Regarding claim 3, Lai et al. do teach The method of claim 1, wherein pruning the one or more weights comprises: generating a binary mask; and applying the binary mask to the plurality of weights (¶ 0079 lines 3+: “a binary pruning mask” (generating a binary mask) “for the pre-trained weights θ0” (to apply to one or more weights of the plurality of weights)).
Regarding claim 5, Lai et al. do not specifically disclose the method of claim 1, wherein the fixed-bit width is four.
Afzal do teach the method of claim 1, wherein the fixed-bit width is four (Col. 16 lines 36-39: “research has demonstrated that weights” (for weights) “and activations can be represented using 8-bit integers (INT8) without incurring significant loss in accuracy. Even lower bit widths” (fixed bit width) “such as 4-bit” (being four) “2-bit and 1-bit”).
For obviousness to combine Lai et al. and Afzal see claim 1.
Regarding claim 6, Lai et al. do not specifically disclose the method of claim 5, wherein quantizing each weight of the plurality of weights comprises applying symmetric quantization.
Afzal does teach:
the method of claim 5, wherein quantizing each weight of the plurality of weights comprises applying symmetric quantization (Col. 16 lines 43+: “a symmetric linear quantization” (applying a symmetric quantization) “scheme is used to map 32-bit floating points to 8-bit integers” “as follows: “FP32(T)=scale-factor(sf)*8-bit Tensor” “The scaling factor” “can correspond to” “a weight scale factor” (to each weight)).
For obviousness to combine Lai et al. and Afzal see claim 1.
Regarding claim 7, Lai et al. do not specifically disclose the method of claim 1, wherein the fixed-bit width is two.
Afzal does teach the method of claim 1, wherein the fixed-bit width is two (Col. 16 lines 36-39: “research has demonstrated that weights” (for weights) “and activations can be represented using 8-bit integers (INT8) without incurring significant loss in accuracy. Even lower bit widths” (fixed bit width) “such as 4-bit” “2-bit” (being two) and 1-bit”).
For obviousness to combine Lai et al. and Afzal see claim 1.
Regarding claim 11, Lai et al. do teach a system comprising: data processing hardware; and memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware (¶ 0054 S1: “Program 150 may be stored in persistent storage 1605 and in memory 1602” (memory) “for execution by one or more of the respective computer processor(s) 1601” (data processing hardware) “via cache 1603”);
cause the data processing hardware to perform operations comprising:
obtaining a plurality of training samples (¶ 0078 S1+: “Consider the low-resource ASR problem, where there is only a small, transcribed training set (x,y) Ɛ Dl” (obtaining a plurality of training samples)) ,
each respective training sample of the plurality of training samples comprising:
a respective speech utterance (¶ 0078 S1+: “Here x represents input audio” (a respective speech utterance in the training samples)); and
a respective textual utterance representing a transcription of the respective speech utterance (¶ 0078 S1+: “and y represents output transcription” (a respective utterance representing a transcription of the respective speech utterance));
fine-tuning, using quantization and sparsity aware training with parameters” “and obtains the pre-trained weights θ0” (based on “x” (training samples) a plurality of weights are determined as a pre-trained ASR process) “f(x;θ0) is then finetuned” (which are subsequently fine tuned where the “SSL” (“Self-supervised learning (SSL) speech model”) also known as “Wave2vec2” (¶ 0032) according to ¶ 0113 “consists of three modules” e.g. “a quantization module” (using quantization) and furthermore the “SSL” model according to step “204” in Fig. 2 undergoes a “TARGET SPARSITY AND INITIAL PRUNING MASK” (uses sparsity)),
the fine- tuning comprising:
pruning one or more weights of the plurality of weights using a sparsity mask (¶ 0079 lines 3+: “a binary pruning mask” (using a sparsity mask for pruning) “for the pre-trained weights θ0” (one or more weights of the plurality of weights)) ; and
and
providing the fine-tuned ASR model to a user device (¶ 0028 lines 5+: “program 150 may implement” (which can “resid[e] on any” “computing device” (e.g., a user device (¶ 0028 page 3 lines 1+)) “finetune” (provide the finetuned ASR model) “the initial subnetwork”).
Lai et al. do not specifically disclose:
Native integer operations;
quantizing each weight of the plurality of weights based on an integer with a fixed-bit width.
Afzal does teach:
Native integer operations; quantizing each weight of the plurality of weights based on an integer with a fixed-bit width (Col. 16 lines 43+: “a symmetric linear quantization” (quantizing) “scheme is used to map 32-bit floating points to 8-bit integers” (with fixed bit width which results in integer output or help generating integer operations) “as follows: “FP32(T)=scale-factor(sf)*8-bit Tensor” “The scaling factor” “can correspond to” “a weight scale factor” (weight) and is applied to “speech recognition”).
It would have therefore been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the “symmetric quantization” techniques of Afzal into the “quantization module” of Lai et al. would enable the combined systems and their associated methods to perform in combination as they do separately and to further enable Lai et al. to have its “weights” “represented using 8-bit integers” which will result in “Reducing the number of bits used to represent weights” and thus help “perform[ing] faster” as disclosed in Afzal Col. 16 lines 37-42.
Lai et al. in view of Afzal do not specifically disclose:
Jointly pruning and quantizing the plurality of weights using a prune-and-quantize approach
Sets pruned weights to zero
Mapping the pruned weights that are set to zero directly to a zero-point of symmetric quantization.
Fedorov et al. do teach:
Jointly pruning and quantizing the plurality of weights using a prune-and-quantize approach (¶ 0095 S2: “The weights” (weights) “for the neural network layers that have been pruned” (are pruned) “and quantized” (and jointly quantized)),
Sets pruned weights to zero (¶ 0094 last lines: “the plurality of weights are pruned by setting each quantized offset weight” (jointly pruning and quantization each weight) “value with the range of weight values to be pruned to zero” (wherein pruned weights are set to zero)),
Mapping the pruned weights that are set to zero directly to a zero-point of symmetric quantization (¶ 0042 last S: “employ” “symmetric quantization given by Equation (3)”, where the “quantization” involves “quantizing the offset weight values” (Abstract line 6), which according to ¶ 0107 “quantized offset weight values are pruned” “to zero” (“quantized offset weight values” (weights )are mapped to zero (that are pruned are set to zero-point quantization), wherein the “quantization” used is “symmetric quantization” (is symmetric quantization according to ¶ 0042 quoted above)).
It would have therefore been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the “ANN” (a neural network) technique including e.g., its “uniform symmetric quantization” of Fedorov et al. into the “quantization module” of Lai et al. in Lai et al. in view of Afzal would enable the combined systems and their associated methods to perform in combination as they do separately and to further enable Lai et al. in view of Afzal to “achieve a particular level of accuracy” by “optimizing the connection weights” as disclosed in Fedorov et al. ¶ 0026 S1.
Regarding claim 12, Lai et al. do teach the system of claim 11, wherein the sparsity mask comprises a binary mask (¶ 0079 lines 3+: “a binary pruning mask” (using a binary mask for pruning) “for the pre-trained weights θ0”).
Regarding claim 13, Lai et al. do teach the system of claim 11, wherein pruning the one or more weights comprises: generating a binary mask; and applying the binary mask to the plurality of weights (¶ 0079 lines 3+: “a binary pruning mask” (generating a binary mask) “for the pre-trained weights θ0” (to apply to one or more weights of the plurality of weights)).
Regarding claim 15, Lai et al. do not specifically disclose the system of claim 11, wherein the fixed-bit width is four.
Afzal do teach the system of claim 11, wherein the fixed-bit width is four (Col. 16 lines 36-39: “research has demonstrated that weights” (for weights) “and activations can be represented using 8-bit integers (INT8) without incurring significant loss in accuracy. Even lower bit widths” (fixed bit width) “such as 4-bit” (being four) “2-bit and 1-bit”).
For obviousness to combine Lai et al. and Afzal see claim 11.
Regarding claim 16, Lai et al. do not specifically disclose the method of claim 15, wherein quantizing each weight of the plurality of weights comprises applying symmetric quantization.
Afzal does teach:
the system of claim 15, wherein quantizing each weight of the plurality of weights comprises applying symmetric quantization (Col. 16 lines 43+: “a symmetric linear quantization” (applying a symmetric quantization) “scheme is used to map 32-bit floating points to 8-bit integers” “as follows: “FP32(T)=scale-factor(sf)*8-bit Tensor” “The scaling factor” “can correspond to” “a weight scale factor” (to each weight)).
For obviousness to combine Lai et al. and Afzal see claim 11.
Regarding claim 17, Lai et al. do not specifically disclose the system of claim 11, wherein the fixed-bit width is two.
Afzal does teach the system of claim 11, wherein the fixed-bit width is two (Col. 16 lines 36-39: “research has demonstrated that weights” (for weights) “and activations can be represented using 8-bit integers (INT8) without incurring significant loss in accuracy. Even lower bit widths” (fixed bit width) “such as 4-bit” “2-bit” (being two) and 1-bit”).
For obviousness to combine Lai et al. and Afzal see claim 11.
Claim(s) 4, 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lai et al. in view of Afzal and Fedorov et al., and further in view of Xu et al. (US 2022/0318604).
Regarding claim 4, Lai et al. in view of Afzal and Fedorov et al. do not specifically disclose the method of claim 3, wherein the binary mask is based on an N:M sparsity pattern, wherein M represents a consecutive number of weights of the plurality of weights and N represents a maximum number of non-zero values.
Xu et al. do teach:
The method of claim 3, wherein the binary mask is based on an N:M sparsity pattern, wherein M represents a consecutive number of weights of the plurality of weights and N represents a maximum number of non-zero values (¶ 0025 lines 15+: “ the sparsity” (e.g., the binary mask is a sparsity pattern and) “criteria can be expressed” (is based on) “in terms of a maximum number of non-zero weight values”(maximum number of non-zero values) “regardless of the total number of weights” “e.g., to ensure the weight matrix is within a storage limit” “and the number of non-zero weight values” (and a consecutive number of weights) “in initial weight tensor 202 can be counted to determine if the sparsity criteria is met”).
It would have therefore been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the “sparsity” features of Xu et al. into the “SPARSITY” step of Lai et al. in Lai et al. in view of Afzal and Fedorov et al., would enable the combined systems and their associated methods to perform in combination as they do separately and to further enable Lai et al. in view of Afzal and Fedorov et al. to “adjust[]” “sparsity” based on “the training process” as disclosed in Xu et al. ¶ 0025 last sentence.
Regarding claim 14, Lai et al. in view of Afzal and Fedorov et al. do not specifically disclose the system of claim 13, wherein the binary mask is based on an N:M sparsity pattern, wherein M represents a consecutive number of weights of the plurality of weights and N represents a maximum number of non-zero values.
Xu et al. do teach:
The system of claim 13, wherein the binary mask is based on an N:M sparsity pattern, wherein M represents a consecutive number of weights of the plurality of weights and N represents a maximum number of non-zero values (¶ 0025 lines 15+: “ the sparsity” (e.g., the binary mask is a sparsity pattern and) “criteria can be expressed” (is based on) “in terms of a maximum number of non-zero weight values”(maximum number of non-zero values) “regardless of the total number of weights” “e.g., to ensure the weight matrix is within a storage limit” “and the number of non-zero weight values” (and a consecutive number of weights) “in initial weight tensor 202 can be counted to determine if the sparsity criteria is met”).
It would have therefore been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the “sparsity” features of Xu et al. into the “SPARSITY” step of Lai et al. in Lai et al. in view of Afzal and Fedorov et al., would enable the combined systems and their associated methods to perform in combination as they do separately and to further enable Lai et al. in view of Afzal and Fedorov et al. to “adjust[]” “sparsity” based on “the training process” as disclosed in Xu et al. ¶ 0025 last sentence.
Claim(s) 8, 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lai et al. in view of Afzal and Fedorov et al., and further in view of Dong XUJIONG (CN 116468082).
Regarding claim 8, Lai et al. in view of Afzal and Fedorov et al. do not specifically disclose the method of claim 7, wherein quantizing each weight of the plurality of weights comprises quantizing each weight of the plurality of weights using asymmetric quantization and sub-channel quantization.
Dong XUJIONG et al. do teach the method of claim 7, wherein quantizing each weight of the plurality of weights comprises quantizing each weight of the plurality of weights using asymmetric quantization and sub-channel quantization (¶ n0030 S1: “ asymmetric quantization” (asymmetric quantization) “the quantization of the weight” (for quantizing weights) “tensor comprises hierarchical quantization or sub-channel quantization” (also being sub-channel quantization)).
It would have therefore been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the “quantization” procedures of Dong XUJIONG et al. into the “quantization module” of Lai et al. in Lai et al. in view of Afzal would enable the combined systems and their associated methods to perform in combination as they do separately and to further enable Lai et al. in view of Afzal and Fedorov et al. to adjust “quantization” of “different channels” “according to different quantization parameters” as disclosed in Don XUJIONG et al. ¶ n0030 last sentence.
Regarding claim 18, Lai et al. in view of Afzal and Fedorov et al. do not specifically disclose the system of claim 17, wherein quantizing each weight of the plurality of weights comprises quantizing each weight of the plurality of weights using asymmetric quantization and sub-channel quantization.
Dong XUJIONG et al. do teach the system of claim 17, wherein quantizing each weight of the plurality of weights comprises quantizing each weight of the plurality of weights using asymmetric quantization and sub-channel quantization (¶ n0030 S1: “ asymmetric quantization” (asymmetric quantization) “the quantization of the weight” (for quantizing weights) “tensor comprises hierarchical quantization or sub-channel quantization” (also being sub-channel quantization)).
It would have therefore been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the “quantization” procedures of Dong XUJIONG et al. into the “quantization module” of Lai et al. in Lai et al. in view of Afzal and Fedorov et al. would enable the combined systems and their associated methods to perform in combination as they do separately and to further enable Lai et al. to adjust “quantization” of “different channels” “according to different quantization parameters” as disclosed in Don XUJIONG et al. ¶ n0030 last sentence.
Claim(s) 9-10, 19-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lai et al. in view of Afzal and Fedorov et al., and further in view of TU JIHUI et al. (CN 116721673).
Regarding claim 9, Lai et al. in view of Afzal and Fedorov et al. do not specifically disclose the method of claim 1, wherein the ASR model comprises one or more multi- head attention layers.
TU JIHUI et al. do teach the method of claim 1, wherein the ASR model comprises one or more multi- head attention layers (¶ n0039 P. 22 ¶ 3 lines 3-5: “global features are extracted from the feature spectrograms” (As part of an ASR operation) “through the multi-head self-attention mechanism” (using multi-head attention layers) “in the Transformer model”).
It would have therefore been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the “multi-head self-attention mechanism” of TU JIHUI et al. into the “self-attention heads” of Lai et al. ¶ 0113 in Lai et al. in view of Afzal and Fedorov et al. would enable the combined systems and their associated methods to perform in combination as they do separately and to further enable Lai et al. in view of Afzal and Fedorov et al. to reduce “the number of parameters in the trained model” as disclosed in TU JIHUI et al. ¶ n0057 ¶ before last.
Regarding claim 10, Lai et al. in view of Afzal and Fedorov et al. do not specifically disclose the method of claim 9, wherein the one or more multi-head attention layers comprise one or more conformer layers or one or more transformer layers.
TU JIHUI et al. do teach the method of claim 9, wherein the one or more multi-head attention layers comprise one or more conformer layers or one or more transformer layers (¶ n0039 P. 22 ¶ 3 lines 3-5: “global features are extracted from the feature spectrograms” (As part of an ASR operation) “through the multi-head self-attention mechanism” (using multi-head attention layers) “in the Transformer model” (comprised one or more transformer layers)).
For obviousness to combine Lai et al. in view of Afzal and Fedorov et al. and TU JIHUI et al. see claim 9.
Regarding claim 19, Lai et al. in view of Afzal and Fedorov et al. do not specifically disclose the system of claim 11, wherein the ASR model comprises one or more multi- head attention layers.
TU JIHUI et al. do teach the system of claim 11, wherein the ASR model comprises one or more multi- head attention layers (¶ n0039 P. 22 ¶ 3 lines 3-5: “global features are extracted from the feature spectrograms” (As part of an ASR operation) “through the multi-head self-attention mechanism” (using multi-head attention layers) “in the Transformer model”).
It would have therefore been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the “multi-head self-attention mechanism” of TU JIHUI et al. into the “self-attention heads” of Lai et al. ¶ 0113 in Lai et al. in view of Afzal and Fedorov et al. would enable the combined systems and their associated methods to perform in combination as they do separately and to further enable Lai et al. in view of Afzal and Fedorov et al. to reduce “the number of parameters in the trained model” as disclosed in TU JIHUI et al. ¶ n0057 ¶ before last.
Regarding claim 20, Lai et al. in view of Afzal and Fedorov et al. do not specifically disclose the system of claim 19, wherein the one or more multi-head attention layers comprise one or more conformer layers or one or more transformer layers.
TU JIHUI et al. do teach the system of claim 19, wherein the one or more multi-head attention layers comprise one or more conformer layers or one or more transformer layers (¶ n0039 P. 22 ¶ 3 lines 3-5: “global features are extracted from the feature spectrograms” (As part of an ASR operation) “through the multi-head self-attention mechanism” (using multi-head attention layers) “in the Transformer model” (comprised one or more transformer layers)).
For obviousness to combine Lai et al. in view of Afzal and Fedorov et al. and TU JIHUI et al. see claim 19.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claims 1, 5, 9, 10, 11, 15 and 19 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1, 2, 7, 8, 11, 12, 17 respectively of U.S. Patent No. 12374323 in view of Lai et al. because:
18/826,135
A computer-implemented method when executed on data processing hardware causes the data processing hardware to perform operations comprising:
obtaining a plurality of training samples, each respective training sample of the
plurality of training samples comprising:
a respective speech utterance; and
a respective textual utterance representing a transcription of the respective speech utterance;
fine-tuning, using quantization and sparsity aware training with native integer operations, a pre-trained automatic speech recognition (ASR) model on the plurality of training samples; the pre-trained ASR model comprising a plurality of weights, the fine-tuning comprising:
pruning one or more weights of the plurality of weights
quantizing each weight of the plurality of weights based on an integer with a fixed-bit width; and providing the quantized fine-tuned ASR model to a user device.
12374323
A computer-implemented method when executed on data processing hardware causes the data processing hardware to perform operations comprising:
obtaining a plurality of training samples, each respective training sample of the
plurality of training samples comprising:
a respective speech utterance; and
a respective textual utterance representing a transcription of the respective speech utterance;
training, using quantization aware training with native integer operations, an
automatic speech recognition (ASR) model on the plurality of training samples;
quantizing the trained ASR model to an integer target fixed-bit width, the quantized trained ASR model comprising a plurality of weights, each weight of the plurality of weights comprising an integer with the target fixed-bit width; and providing the quantized trained ASR model to a user device.
12374323 does not specifically disclose using a sparsity mask for pruning one or more weights.
Lai et al. do teach using a sparsity mask for pruning one or more weights (¶ 0079 lines 3+: “a binary pruning mask” (using a sparsity mask for pruning) “for the pre-trained weights θ0” (one or more weights of the plurality of weights))).
It would have therefore been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the “binary pruning mask” of Lai et al. into 18/826,135 would enable the combined systems and their associated methods to perform in combination as they do separately and to further enable 18/826,135 in “Reducing the number of bits used to represent weights” and thus help “perform[ing] faster” as disclosed in Afzal Col. 16 lines 37-42.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to FARZAD KAZEMINEZHAD whose telephone number is (571)270-5860. The examiner can normally be reached 10:30 am to 11:30 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Paras D. Shah can be reached at (571) 270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Farzad Kazeminezhad/
Art Unit 2653
August 28th 2026.