Prosecution Insights
Last updated: August 17, 2026
Application No. 18/639,937

SYSTEM AND METHOD FOR HARDWARE-AWARE PRUNING OF CONFORMER NETWORKS

Non-Final OA §101§103§112
Filed
Apr 18, 2024
Priority
Aug 07, 2023 — provisional 63/531,146
Examiner
SIPPEL, MOLLY CLARKE
Art Unit
Tech Center
Assignee
Samsung Electronics Co., Ltd.
OA Round
1 (Non-Final)
52%
Grant Probability
Moderate
1-2
OA Rounds
1y 5m
Est. Remaining
78%
With Interview

Examiner Intelligence

Grants 52% of resolved cases
52%
Career Allowance Rate
12 granted / 23 resolved
-7.8% vs TC avg
Strong +26% interview lift
Without
With
+26.1%
Interview Lift
resolved cases with interview
Typical timeline
3y 9m
Avg Prosecution
17 currently pending
Career history
41
Total Applications
across all art units

Statute-Specific Performance

§101
35.4%
-4.6% vs TC avg
§103
29.7%
-10.3% vs TC avg
§102
11.3%
-28.7% vs TC avg
§112
22.6%
-17.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 23 resolved cases

Office Action

§101 §103 §112
DETAILED ACTION This action is responsive to the application filed on 04/18/2024. Claims 1-20 are pending in the case. Claims 1, 15, and 20 are independent claims. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Acknowledgement is made of applicant’s claim for domestic priority based on a provisional application filed on 08/07/2023. Information Disclosure Statement The information disclosure statement (IDS) submitted on 04/18/2024 is being considered by the examiner. Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim 8 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 8 recites the limitation "the power of the floor of the ratio of the index of the training epoch" in lines 5-6. There is insufficient antecedent basis for this limitation in the claim. It is unclear if applicant is attempting to refer to previous claim elements or if applicant is attempting to recite new claim elements. For examination purposes, this limitation is being interpreted to mean “a power of a floor of a ratio of the index of the training epoch” reciting new claim elements. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Regarding claim 1: Step 1 Statutory Category: Claim 1 is directed to a method, which falls under one of the four statutory categories. Step 2A Prong 1 Judicial exception: Claim 1 recites, in part, “training a neural network, the training comprising: performing a first pruning operation, on the neural network, after a first training epoch, and performing a second pruning operation, on the neural network, after a second training epoch and after the first pruning operation, wherein each of the pruning operations results in a respective pruning fraction, the respective pruning fraction being a function of an index of a training epoch preceding the pruning operation”. This limitation, under the broadest reasonable interpretation, covers the recitation of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case an evaluation. See MPEP § 2106.04(a)(2)(III). Additionally, this limitation, under the broadest reasonable interpretation, and in light of applicant’s specification paragraphs 0068-0082, covers the recitation of mathematical concepts, including calculations and relationships, see MPEP §2106.04(a)(2)(I). Step 2A Prong 2 Integration into a practical application: This judicial exception is not integrated into a practical application. The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application. Step 2B Significantly more: The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible. Regarding claim 2, the rejection of claim 1 is incorporated, and further, the claim recites: “wherein an increase in a pruning fraction during the second pruning operation is less than an increase in the pruning fraction during the first pruning operation”. This limitation recites mathematical concepts in addition to those identified in the rejection of the parent claim. Thus, the claim recites a judicial exception. The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible. Regarding claim 3, the rejection of claim 2 is incorporated, and further, the claim recites: “further comprising training the neural network in a third training epoch, after the first training epoch, and before the second training epoch”. This limitation is an additional element that amounts to adding insignificant extra-solution activity to the judicial exception. See MPEP §2106.05(g). Further, the additional element is well‐understood, routine, and conventional as taught by activity is supported under Berkheimer Option 2, Boren et al., U.S. Patent Application Publication No. 20260161472, Paragraph 0043, 12-21, “As known, performance of an ML model may be defined and measured in terms of any applicable metrics, including accuracy, precision and recall, etc. Each of these models may typically include millions, or even tens and hundreds of millions, of trainable parameters, and its each iteration of training may require hundreds of millions or billions of computations, per each epoch. Every model training-run usually includes tens to several hundred epochs to get to a model having satisfying performance”. The claim is not patent eligible. Regarding claim 4, the rejection of claim 1 is incorporated, and further, the claim recites: “performing a third pruning operation on the neural network after the second pruning operation, wherein an increase in the respective pruning fraction during the third pruning operation is less than an increase in the respective pruning fraction during the second pruning operation”. This limitation, under the broadest reasonable interpretation, covers the recitation of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case an evaluation. See MPEP § 2106.04(a)(2)(III). Additionally, this limitation, under the broadest reasonable interpretation, and in light of applicant’s specification paragraphs 0068-0082, covers the recitation of mathematical concepts, including calculations and relationships, see MPEP §2106.04(a)(2)(I). The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible. Regarding claim 5, the rejection of claim 1 is incorporated, and further, the claim recites: “the training comprises performing a sequence of four or more pruning operations, each of the pruning operations of the sequence increases a respective pruning fraction by an amount less than a preceding pruning operation of the sequence, the sequence ends at a stopping epoch for pruning, and the stopping epoch for pruning is a training epoch after which a final pruning operation of the sequence is performed”. This limitation recites mental processes and mathematical concepts in addition to those identified in the rejection of the parent claim. Thus, the claim recites a judicial exception. The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible. Regarding claim 6, the rejection of claim 5 is incorporated, and further, the claim recites: “wherein each of the pruning operations results in a respective pruning fraction, the respective pruning fraction being a function of: an index of a training epoch, a number of training epochs per pruning operation, and an initial pruning fraction”. This limitation recites mathematical concepts in addition to those identified in the rejection of the parent claim. Thus, the claim recites a judicial exception. The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible. Regarding claim 7, the rejection of claim 6 is incorporated, and further, the claim recites: “wherein the function has: a value of zero when the index of the training epoch is less than the number of training epochs per pruning operation, and a value of the initial pruning fraction, after the first pruning operation”. This limitation recites mathematical concepts in addition to those identified in the rejection of the parent claim. Thus, the claim recites a judicial exception. The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible. Regarding claim 8, the rejection of claim 7 is incorporated, and further, the claim recites: “for each training epoch less than or equal to the stopping epoch for pruning: the function includes a second term subtracted from a first term, the first term is one, the second term is a first difference raised to the power of the floor of the ratio of the index of the training epoch and the number of training epochs per pruning operation, and the first difference is one less the initial pruning fraction; and for each training epoch greater than the stopping epoch, the function is equal to the respective pruning fraction after the last pruning operation”. This limitation recites mathematical concepts in addition to those identified in the rejection of the parent claim. Thus, the claim recites a judicial exception. The claim does not include any additional elements that amount to an integration of the judicial exception into a practical application, nor to significantly more than the judicial exception. The claim is not patent eligible. Regarding claim 9, the rejection of claim 1 is incorporated, and further, the claim recites: “the first pruning operation comprises removing a row or a column of the fully connected layer”. This limitation is a continuation of the “training a neural network, the training comprising: performing a first pruning operation, on the neural network, after a first training epoch, and performing a second pruning operation, on the neural network, after a second training epoch and after the first pruning operation, wherein each of the pruning operations results in a respective pruning fraction, the respective pruning fraction being a function of an index of a training epoch preceding the pruning operation” limitation identified as an abstract idea in the rejection of the parent claim. Thus, the claim recites a judicial exception. Further, the claim recites: “the neural network comprises a fully connected layer”. This limitation is an additional element that amounts to generally linking the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Elements that merely amount to generally linking the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. The claim is not patent eligible. Regarding claim 10, the rejection of claim 1 is incorporated, and further, the claim recites: “the first pruning operation comprises removing a row or a column of the multi-head self-attention block”. This limitation is a continuation of the “training a neural network, the training comprising: performing a first pruning operation, on the neural network, after a first training epoch, and performing a second pruning operation, on the neural network, after a second training epoch and after the first pruning operation, wherein each of the pruning operations results in a respective pruning fraction, the respective pruning fraction being a function of an index of a training epoch preceding the pruning operation” limitation identified as an abstract idea in the rejection of the parent claim. Thus, the claim recites a judicial exception. Further, the claim recites: “the neural network comprises a multi-head self-attention block”. This limitation is an additional element that amounts to generally linking the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Elements that merely amount to generally linking the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. The claim is not patent eligible. Regarding claim 11, the rejection of claim 1 is incorporated, and further, the claim recites: “the first pruning operation comprises removing a row or a column or a channel of a convolutional kernel of the convolutional layer”. This limitation is a continuation of the “training a neural network, the training comprising: performing a first pruning operation, on the neural network, after a first training epoch, and performing a second pruning operation, on the neural network, after a second training epoch and after the first pruning operation, wherein each of the pruning operations results in a respective pruning fraction, the respective pruning fraction being a function of an index of a training epoch preceding the pruning operation” limitation identified as an abstract idea in the rejection of the parent claim. Thus, the claim recites a judicial exception. Further, the claim recites: “the neural network comprises a convolutional layer”. This limitation is an additional element that amounts to generally linking the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Elements that merely amount to generally linking the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. The claim is not patent eligible. Regarding claim 12, the rejection of claim 1 is incorporated, and further, the claim recites: “the neural network comprises a first conformer layer and a second conformer layer, and the training comprises setting each of a plurality of weights of the second conformer layer equal to a respective weight of the first conformer layer”. This limitation is an additional element that amounts to generally linking the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Elements that merely amount to generally linking the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. The claim is not patent eligible. Regarding claim 13, the rejection of claim 1 is incorporated, and further, the claim recites: “wherein the training comprises knowledge distillation”. This limitation is an additional element that amounts to generally linking the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Elements that merely amount to generally linking the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. The claim is not patent eligible. Regarding claim 14, the rejection of claim 1 is incorporated, and further, the claim recites: “further comprising receiving, by the neural network, a raw signal”. This limitation is an additional element that amounts to mere data gathering. It is necessary to acquire the data in order to use the recited judicial exception. Therefore, this limitation is insignificant extra-solution activity to the judicial exception, see MPEP §2106.05(g). Further, this limitation is directed to receiving or transmitting data over a network which courts have recognized as well-understood, routine, and conventional when they are claimed in a generic manner, see MPEP §2106.05(d)(II). Further, the claim recites: “producing, by the neural network, an output, the output comprising an enhanced signal corresponding to the raw signal”. This limitation is an additional element that amounts to generally linking the use of the judicial exception to a particular technological environment or field of use. See MPEP §2106.05(h). Elements that merely amount to generally linking the use of the judicial exception to a particular technological environment or field of use cannot provide an inventive concept. The claim is not patent eligible. Regarding claim 15: Step 1 Statutory Category: Claim 15 is directed to a system, which falls under one of the four statutory categories. Step 2A Prong 1 Judicial exception: Claim 15 recites, in part, “training a neural network, the training comprising: performing a first pruning operation, on the neural network, after a first training epoch, and performing a second pruning operation, on the neural network, after a second training epoch and after the first pruning operation, wherein each of the pruning operations results in a respective pruning fraction, the respective pruning fraction being a function of an index of a training epoch preceding the pruning operation”. This limitation, under the broadest reasonable interpretation, covers the recitation of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case an evaluation. See MPEP § 2106.04(a)(2)(III). Additionally, this limitation, under the broadest reasonable interpretation, and in light of applicant’s specification paragraphs 0068-0082, covers the recitation of mathematical concepts, including calculations and relationships, see MPEP §2106.04(a)(2)(I). Step 2A Prong 2 Integration into a practical application: This judicial exception is not integrated into a practical application. In particular the claim recites: “A system comprising: one or more processors; and a memory storing instructions which, when executed by the one or more processors, cause performance”. This limitation is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f). Step 2B Significantly more: The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element: “A system comprising: one or more processors; and a memory storing instructions which, when executed by the one or more processors, cause performance” amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. Elements that merely amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process cannot provide an inventive concept. The claim is not patent eligible. Regarding claim 16, the rejection of claim 15 is incorporated, and further, claim 16 is substantially similar to claim 9 respectively, and is rejected in the same manner and reasoning applying. Regarding claim 17, the rejection of claim 15 is incorporated, and further, claim 17 is substantially similar to claim 10 respectively, and is rejected in the same manner and reasoning applying. Regarding claim 18, the rejection of claim 15 is incorporated, and further, claim 18 is substantially similar to claim 11 respectively, and is rejected in the same manner and reasoning applying. Regarding claim 19, the rejection of claim 15 is incorporated, and further, claim 19 is substantially similar to claim 12 respectively, and is rejected in the same manner and reasoning applying. Regarding claim 20: Step 1 Statutory Category: Claim 20 is directed to a system, which falls under one of the four statutory categories. Step 2A Prong 1 Judicial exception: Claim 20 recites, in part, “training a neural network, the training comprising: performing a first pruning operation, on the neural network, after a first training epoch, and performing a second pruning operation, on the neural network, after a second training epoch and after the first pruning operation, wherein each of the pruning operations results in a respective pruning fraction, the respective pruning fraction being a function of an index of a training epoch preceding the pruning operation”. This limitation, under the broadest reasonable interpretation, covers the recitation of a mental process that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper (including an observation, evaluation, judgment, opinion), in this case an evaluation. See MPEP § 2106.04(a)(2)(III). Additionally, this limitation, under the broadest reasonable interpretation, and in light of applicant’s specification paragraphs 0068-0082, covers the recitation of mathematical concepts, including calculations and relationships, see MPEP §2106.04(a)(2)(I). Step 2A Prong 2 Integration into a practical application: This judicial exception is not integrated into a practical application. In particular the claim recites: “A system comprising: means for processing; and a memory storing instructions which, when executed by the means for processing, cause performance”. This limitation is an additional element that amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. See MPEP §2106.05(f). Step 2B Significantly more: The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element: “A system comprising: means for processing; and a memory storing instructions which, when executed by the means for processing, cause performance” amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process. Elements that merely amount to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer in its ordinary capacity as a tool to perform an existing process cannot provide an inventive concept. The claim is not patent eligible. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-7, 13-15, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Wei et al., Model Compression by Iterative Pruning with Knowledge Distillation and Its Application to Speech Enhancement, 18-22 September 2022, https://www.isca-archive.org/interspeech_2022/wei22_interspeech.pdf, hereinafter referred to as “Wei” in view of Kim et al., U.S. Patent Application Publication No. 20220318633, hereinafter referred to as “Kim”. Regarding claim 1, Wei teaches A method, comprising: training a neural network (Wei, Page 941, Abstract, Lines 8-9, we propose a compression strategy based on iterative pruning and knowledge distillation”), the training comprising: performing a first pruning operation, on the neural network, after a first training epoch, and performing a second pruning operation, on the neural network, after a second training epoch and after the first pruning operation (Wei, Page 942, Section 2.2, Lines 9-12, “Iterative pruning is gradually pruning, which means the number of the parameters reduce slowly each time, and the model will be fine-tuned well in the end. Therefore, this paper use iterative pruning strategy to compress model”; see also Wei, Page 942, Figure 1). Wei does not explicitly teach wherein each of the pruning operations results in a respective pruning fraction, the respective pruning fraction being a function of an index of a training epoch preceding the pruning operation. Kim teaches wherein each of the pruning operations results in a respective pruning fraction, the respective pruning fraction being a function of an index of a training epoch preceding the pruning operation (Kim, Paragraph 0055, “Additionally, a pruning ratio may be adapted based on the iteration. The pruning ratio may serve as a mask and may be gradually increased as given by: p c = p t + p i - p t 1 - c - c 0 n 3 , where p i is the initial pruning ratio (e.g., p i = 0 ), p t is the target pruning ratio, n represents the training epoch and p c represents the current pruning ratio for c ∈ { c 0 , … , c 0 + n } , where c represents the current epoch” the “pruning ratio” is considered to be the “pruning fraction”). It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to have modified the pruning method of Wei to include pruning operations resulting in pruning fractions as taught by Kim. The motivation to do so would have been that pruning achieve reduced model size, increasing sparsity and reducing computations of the neural network (Kim, Paragraph 0022), and using the function discloses by Kim allows the pruning ratio to be gradually increased by adapting it based on the iteration (Kim, Paragraph 0055). Regarding claim 2, the rejection of claim 1 is incorporated, and further, the proposed combination teaches wherein an increase in a pruning fraction during the second pruning operation is less than an increase in the pruning fraction during the first pruning operation (Kim, Paragraph 0055, “Additionally, a pruning ratio may be adapted based on the iteration. The pruning ratio may serve as a mask and may be gradually increased as given by: p c = p t + p i - p t 1 - c - c 0 n 3 , where p i is the initial pruning ratio (e.g., p i = 0 ), p t is the target pruning ratio, n represents the training epoch and p c represents the current pruning ratio for c ∈ { c 0 , … , c 0 + n } , where c represents the current epoch”; A person of ordinary skill in the art would recognize that as the index of the epoch increases, the increase of the pruning fraction is smaller because the curve uses a cubed term which flattens out as c approaches n (the total number of epochs). Regarding claim 3, the rejection of claim 2 is incorporated, and further, the proposed combination teaches further comprising training the neural network in a third training epoch, after the first training epoch, and before the second training epoch (Wei, Page 942, Section 2.2, Lines 9-12, “Iterative pruning is gradually pruning, which means the number of the parameters reduce slowly each time, and the model will be fine-tuned well in the end. Therefore, this paper use iterative pruning strategy to compress model”; see also Wei, Page 942, Figure 1, Each “teach” arrow is considered to be a “training epoch”). Regarding claim 4, the rejection of claim 1 is incorporated, and further, the proposed combination teaches further comprising performing a third pruning operation on the neural network after the second pruning operation, wherein an increase in the respective pruning fraction during the third pruning operation is less than an increase in the respective pruning fraction during the second pruning operation (Kim, Paragraph 0055, “Additionally, a pruning ratio may be adapted based on the iteration. The pruning ratio may serve as a mask and may be gradually increased as given by: p c = p t + p i - p t 1 - c - c 0 n 3 , where p i is the initial pruning ratio (e.g., p i = 0 ), p t is the target pruning ratio, n represents the training epoch and p c represents the current pruning ratio for c ∈ { c 0 , … , c 0 + n } , where c represents the current epoch”; A person of ordinary skill in the art would recognize that as the index of the epoch increases, the increase of the pruning fraction is smaller because the curve uses a cubed term which flattens out as c approaches n (the total number of epochs). Regarding claim 5, the rejection of claim 1 is incorporated, and further, the proposed combination teaches wherein: the training comprises performing a sequence of four or more pruning operations (Wei, Page 942, Section 2.2, Lines 9-12, “Iterative pruning is gradually pruning, which means the number of the parameters reduce slowly each time, and the model will be fine-tuned well in the end. Therefore, this paper use iterative pruning strategy to compress model”; see also Wei, Page 942, Figure 1, The figure shows at least 4 “prune” arrows with an “iterative pruning n times” showing it could include more than 4 pruning operations), each of the pruning operations of the sequence increases a respective pruning fraction by an amount less than a preceding pruning operation of the sequence (Kim, Paragraph 0055, “Additionally, a pruning ratio may be adapted based on the iteration. The pruning ratio may serve as a mask and may be gradually increased as given by: p c = p t + p i - p t 1 - c - c 0 n 3 , where p i is the initial pruning ratio (e.g., p i = 0 ), p t is the target pruning ratio, n represents the training epoch and p c represents the current pruning ratio for c ∈ { c 0 , … , c 0 + n } , where c represents the current epoch”; A person of ordinary skill in the art would recognize that as the index of the epoch increases, the increase of the pruning fraction is smaller because the curve uses a cubed term which flattens out as c approaches n (the total number of epochs), the sequence ends at a stopping epoch for pruning, and the stopping epoch for pruning is a training epoch after which a final pruning operation of the sequence is performed (Kim, Paragraph 0055, “Additionally, a pruning ratio may be adapted based on the iteration. The pruning ratio may serve as a mask and may be gradually increased as given by: p c = p t + p i - p t 1 - c - c 0 n 3 , where p i is the initial pruning ratio (e.g., p i = 0 ), p t is the target pruning ratio, n represents the training epoch and p c represents the current pruning ratio for c ∈ { c 0 , … , c 0 + n } , where c represents the current epoch”; “ c 0 + n ” is considered to be the “stopping epoch”). Regarding claim 6, the rejection of claim 5 is incorporated, and further, the proposed combination teaches wherein each of the pruning operations results in a respective pruning fraction, the respective pruning fraction being a function of: an index of a training epoch, a number of training epochs per pruning operation, and an initial pruning fraction (Kim, Paragraph 0055, “Additionally, a pruning ratio may be adapted based on the iteration. The pruning ratio may serve as a mask and may be gradually increased as given by: p c = p t + p i - p t 1 - c - c 0 n 3 , where p i is the initial pruning ratio (e.g., p i = 0 ), p t is the target pruning ratio, n represents the training epoch and p c represents the current pruning ratio for c ∈ { c 0 , … , c 0 + n } , where c represents the current epoch”). Regarding claim 7, the rejection of claim 6 is incorporated, and further, the proposed combination teaches wherein the function has: a value of zero when the index of the training epoch is less than the number of training epochs per pruning operation, and a value of the initial pruning fraction, after the first pruning operation (Kim, Paragraph 0055, “Additionally, a pruning ratio may be adapted based on the iteration. The pruning ratio may serve as a mask and may be gradually increased as given by: p c = p t + p i - p t 1 - c - c 0 n 3 , where p i is the initial pruning ratio (e.g., p i = 0 ), p t is the target pruning ratio, n represents the training epoch and p c represents the current pruning ratio for c ∈ { c 0 , … , c 0 + n } , where c represents the current epoch”; “ p i ” is the “initial pruning fraction” and is also “0”). Regarding claim 13, the rejection of claim 1 is incorporated, and further, the proposed combination teaches wherein the training comprises knowledge distillation (Wei, Page 941, Abstract, Lines 8-9, we propose a compression strategy based on iterative pruning and knowledge distillation”). Regarding claim 14, the rejection of claim 1 is incorporated, and further, the proposed combination teaches further comprising receiving, by the neural network, a raw signal, and producing, by the neural network, an output, the output comprising an enhanced signal corresponding to the raw signal (Wei, Page 941, Introduction, Paragraph 4, “In order to evaluate the proposed method, we choose two representative models, LSTM and Gated CRN [22], for speech enhancement. Experimental results show that the model size is greatly reduced about 10x and 40x for LSTM and GCRN respectively without significant performance loss”; Wei, Page 943, Section 3.2, Lines 1-2, “For all experiments, we take a piece of speech through the sampling rate of 16kHz”). Regarding claim 15, Wei teaches A system comprising: one or more processors; and a memory storing instructions which, when executed by the one or more processors (Wei, Page 943, Section 3.1, “WSJ0 SI-84 [24], which includes 7138 utterances from 83 speakers (42 males and 41 females), is used to evaluate the proposed method. 77 speakers are used for training and remaining 6 are used for evaluation. For the training set, 10000 noises from a sound effect library (available at https://www.sound ideas.com) [25] are used. Specifically, we mix a randomly selected training utterance with a random cut from the 10000 noises. The SNR is randomly sampled between-5 and 0 dB. The training set includes 300,000 mixtures. Then, creating 30,000 mixtures as validate set by using the same method”; A person of ordinary skill in the art would recognize that these experiments require the use of a generic computer, providing evidence for “one or more processors” and “a memory storing instructions”), cause performance of: training a neural network (Wei, Page 941, Abstract, Lines 8-9, we propose a compression strategy based on iterative pruning and knowledge distillation”), the training comprising: performing a first pruning operation, on the neural network, after a first training epoch, and performing a second pruning operation, on the neural network, after a second training epoch and after the first pruning operation (Wei, Page 942, Section 2.2, Lines 9-12, “Iterative pruning is gradually pruning, which means the number of the parameters reduce slowly each time, and the model will be fine-tuned well in the end. Therefore, this paper use iterative pruning strategy to compress model”; see also Wei, Page 942, Figure 1). Wei does not explicitly teach wherein each of the pruning operations results in a respective pruning fraction, the respective pruning fraction being a function of an index of a training epoch preceding the pruning operation. Kim teaches wherein each of the pruning operations results in a respective pruning fraction, the respective pruning fraction being a function of an index of a training epoch preceding the pruning operation (Kim, Paragraph 0055, “Additionally, a pruning ratio may be adapted based on the iteration. The pruning ratio may serve as a mask and may be gradually increased as given by: p c = p t + p i - p t 1 - c - c 0 n 3 , where p i is the initial pruning ratio (e.g., p i = 0 ), p t is the target pruning ratio, n represents the training epoch and p c represents the current pruning ratio for c ∈ { c 0 , … , c 0 + n } , where c represents the current epoch” the “pruning ratio” is considered to be the “pruning fraction”). It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to have modified the pruning method of Wei to include pruning operations resulting in pruning fractions as taught by Kim. The motivation to do so would have been that pruning achieve reduced model size, increasing sparsity and reducing computations of the neural network (Kim, Paragraph 0022), and using the function discloses by Kim allows the pruning ratio to be gradually increased by adapting it based on the iteration (Kim, Paragraph 0055). Regarding claim 20, Wei teaches A system comprising: means for processing; and a memory storing instructions which, when executed by the means for processing (Wei, Page 943, Section 3.1, “WSJ0 SI-84 [24], which includes 7138 utterances from 83 speakers (42 males and 41 females), is used to evaluate the proposed method. 77 speakers are used for training and remaining 6 are used for evaluation. For the training set, 10000 noises from a sound effect library (available at https://www.sound ideas.com) [25] are used. Specifically, we mix a randomly selected training utterance with a random cut from the 10000 noises. The SNR is randomly sampled between-5 and 0 dB. The training set includes 300,000 mixtures. Then, creating 30,000 mixtures as validate set by using the same method”; A person of ordinary skill in the art would recognize that these experiments require the use of a generic computer, providing evidence for “a means for processing” (considered “processing circuitry”) and “a memory storing instructions”), cause performance of: training a neural network (Wei, Page 941, Abstract, Lines 8-9, we propose a compression strategy based on iterative pruning and knowledge distillation”), the training comprising: performing a first pruning operation, on the neural network, after a first training epoch, and performing a second pruning operation, on the neural network, after a second training epoch and after the first pruning operation (Wei, Page 942, Section 2.2, Lines 9-12, “Iterative pruning is gradually pruning, which means the number of the parameters reduce slowly each time, and the model will be fine-tuned well in the end. Therefore, this paper use iterative pruning strategy to compress model”; see also Wei, Page 942, Figure 1). Wei does not explicitly teach wherein each of the pruning operations results in a respective pruning fraction, the respective pruning fraction being a function of an index of a training epoch preceding the pruning operation. Kim teaches wherein each of the pruning operations results in a respective pruning fraction, the respective pruning fraction being a function of an index of a training epoch preceding the pruning operation (Kim, Paragraph 0055, “Additionally, a pruning ratio may be adapted based on the iteration. The pruning ratio may serve as a mask and may be gradually increased as given by: p c = p t + p i - p t 1 - c - c 0 n 3 , where p i is the initial pruning ratio (e.g., p i = 0 ), p t is the target pruning ratio, n represents the training epoch and p c represents the current pruning ratio for c ∈ { c 0 , … , c 0 + n } , where c represents the current epoch” the “pruning ratio” is considered to be the “pruning fraction”). It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to have modified the pruning method of Wei to include pruning operations resulting in pruning fractions as taught by Kim. The motivation to do so would have been that pruning achieve reduced model size, increasing sparsity and reducing computations of the neural network (Kim, Paragraph 0022), and using the function discloses by Kim allows the pruning ratio to be gradually increased by adapting it based on the iteration (Kim, Paragraph 0055). Claims 9-11 and 16-18 are rejected under 35 U.S.C. 103 as being unpatentable over Wei in view of Kim in further view of Jiang et al., Accurate and Structured Pruning for Efficient Automatic Speech Recognition, 05/31/2023, https://arxiv.org/pdf/2305.19549, hereinafter referred to as “Jiang”. Regarding claim 9, the rejection of claim 1 is incorporated. The proposed combination thus far does not explicitly teach the neural network comprises a fully connected layer; and the first pruning operation comprises removing a row or a column of the fully connected layer. Jiang teaches the neural network comprises a fully connected layer; and the first pruning operation comprises removing a row or a column of the fully connected layer (Jiang, Page 2, Paragraph 3, Final 3 lines and bullet 2, “To attain this goal, we propose four types of binary masks z ∈{0,1} to control the sparsity of different modules, as illustrated in Fig. 1a… FNN INTERMEDIATE MASK z F F N . We use masks z F F N j i to prune the i-th FNN j-th channel in intermediate dimensions”). It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to have modified the pruning method of the proposed combination to include removing a row or column of the fully connected layer as taught by Jiang. The motivation to do so would have been to structurally remove unimportant parameters enabling faster inference while preserving accuracy, and to achieve optimal results all four modules of each conformer encoder layer are accounted for (Jiang, Page 2, Paragraph 2, “our objective is to structurally remove unimportant parameters for ASR models, thereby enabling faster inference on standard hardware while preserving accuracy. While prior research has primarily focused on pruning transformer or CNN models, our approach is tailored to the Conformer architecture, which features a unique combination of convolution and transformer modules. To achieve optimal results, we must develop a specific pruning policy that accounts for all four modules that comprise each Conformer encoder layer: a feed-forward module (FFN1), a multi-head self-attention module (MHA), a convolution module, and a second feed-forward module (FFN2)”). Regarding claim 10, the rejection of claim 1 is incorporated. The proposed combination does not explicitly teach the neural network comprises a multi-head self-attention block; and the first pruning operation comprises removing a row or a column of the multi-head self-attention block. Jiang teaches the neural network comprises a multi-head self-attention block; and the first pruning operation comprises removing a row or a column of the multi-head self-attention block (Jiang, Page 2, Paragraph 3, Final 3 lines and bullet 1, “To attain this goal, we propose four types of binary masks z ∈{0,1} to control the sparsity of different modules, as illustrated in Fig. 1a…HEAD MASK z H E A D . We use z h e a d j i to determine whether j-th head of the i-th encoder layer should be kept or removed”). It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to have modified the pruning method of the proposed combination to include removing a row or column of the multi-head self-attention block as taught by Jiang. The motivation to do so would have been to structurally remove unimportant parameters enabling faster inference while preserving accuracy, and to achieve optimal results all four modules of each conformer encoder layer are accounted for (Jiang, Page 2, Paragraph 2, “our objective is to structurally remove unimportant parameters for ASR models, thereby enabling faster inference on standard hardware while preserving accuracy. While prior research has primarily focused on pruning transformer or CNN models, our approach is tailored to the Conformer architecture, which features a unique combination of convolution and transformer modules. To achieve optimal results, we must develop a specific pruning policy that accounts for all four modules that comprise each Conformer encoder layer: a feed-forward module (FFN1), a multi-head self-attention module (MHA), a convolution module, and a second feed-forward module (FFN2)”). Regarding claim 11, the rejection of claim 1 is incorporated. The proposed combination thus far does not explicitly teach the neural network comprises a convolutional layer; and the first pruning operation comprises removing a row or a column or a channel of a convolutional kernel of the convolutional layer. Jiang teaches the neural network comprises a convolutional layer; and the first pruning operation comprises removing a row or a column or a channel of a convolutional kernel of the convolutional layer (Jiang, Page 2, Paragraph 3, Final 3 lines and bullet 3, “To attain this goal, we propose four types of binary masks z ∈{0,1} to control the sparsity of different modules, as illustrated in Fig. 1a…CONV MODULE MASK z C O N V . Although the convolution module has comparatively few parameters, it contributes significantly to the total number of floating-point operations (FLOPs) in the model. To maximize inference efficiency, we introduce a gate mask, denoted as z c o n v i , which allows for the pruning of the 𝑖-th entire convolution module in each block”). It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to have modified the pruning method of the proposed combination to include removing a row or column or a channel of the convolutional kernel as taught by Jiang. The motivation to do so would have been to structurally remove unimportant parameters enabling faster inference while preserving accuracy, and to achieve optimal results all four modules of each conformer encoder layer are accounted for (Jiang, Page 2, Paragraph 2, “our objective is to structurally remove unimportant parameters for ASR models, thereby enabling faster inference on standard hardware while preserving accuracy. While prior research has primarily focused on pruning transformer or CNN models, our approach is tailored to the Conformer architecture, which features a unique combination of convolution and transformer modules. To achieve optimal results, we must develop a specific pruning policy that accounts for all four modules that comprise each Conformer encoder layer: a feed-forward module (FFN1), a multi-head self-attention module (MHA), a convolution module, and a second feed-forward module (FFN2)”). Regarding claim 16, the rejection of claim 15 is incorporated, and further, claim 16 is substantially similar to claim 9 respectively, and is rejected in the same manner and reasoning applying. Regarding claim 17, the rejection of claim 15 is incorporated, and further, claim 17 is substantially similar to claim 10 respectively, and is rejected in the same manner and reasoning applying. Regarding claim 18, the rejection of claim 15 is incorporated, and further, claim 18 is substantially similar to claim 11 respectively, and is rejected in the same manner and reasoning applying. Claims 12 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Wei in view of Kim in further view of Bai et al., Parameter-Efficient Conformers via Sharing Sparsely-Gated Experts for End-to-End Speech Recognition, 09/17/2022, https://arxiv.org/pdf/2209.08326, hereinafter referred to as “Bai”. Regarding claim 12, the rejection of claim 1 is incorporated. The proposed combination thus far does not explicitly teach the neural network comprises a first conformer layer and a second conformer layer, and the training comprises setting each of a plurality of weights of the second conformer layer equal to a respective weight of the first conformer layer. Bai teaches the neural network comprises a first conformer layer and a second conformer layer, and the training comprises setting each of a plurality of weights of the second conformer layer equal to a respective weight of the first conformer layer (Bai, Page 1, Abstract, Lines 7-11, “this paper proposes a parameter-efficient conformer via sharing sparsely-gated experts. Specifically, we use sparsely-gated mixture-of-experts (MoE) to extend the capacity of a conformer block without increasing computation. Then, the parameters of the grouped conformer blocks are shared”). It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention, to have modified the pruning method of the proposed combination to include parameter sharing as taught by Bai. The motivation to do so would have been to reduce the parameters of the model without losing capacity and thus maintaining model performance (Bai, Page 1, Abstract, Lines 1-8, “While transformers and their variant conformers show promising performance in speech recognition, the parameterized property leads to much memory cost during training and inference. Some works use cross-layer weight-sharing to reduce the parameters of the model. However, the inevitable loss of capacity harms the model performance. To address this issue, this paper proposes a parameter-efficient conformer via sharing sparsely-gated experts”) Regarding claim 19, the rejection of claim 15 is incorporated, and further, claim 19 is substantially similar to claim 12 respectively, and is rejected in the same manner and reasoning applying. Conclusion Claim 8 would be allowable if rewritten to overcome the rejection(s) under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), 2nd paragraph, and 35 U.S.C. 101 set forth in this Office action and to include all of the limitations of the base claim and any intervening claims. A complete prior art search was performed for this claim; however, no prior art was uncovered that disclose or fairly suggest the following claimed features: “for each training epoch less than or equal to the stopping epoch for pruning: the function includes a second term subtracted from a first term, the first term is one, the second term is a first difference raised to the power of the floor of the ratio of the index of the training epoch and the number of training epochs per pruning operation, and the first difference is one less the initial pruning fraction; and for each training epoch greater than the stopping epoch, the function is equal to the respective pruning fraction after the last pruning operation”. Pertinent art (Kim et al., U.S. Patent Application Publication No. 20220318633) discloses a pruning function with a second term subtracted from a first term, where the first term is one (Kim, Paragraph 0055, “Additionally, a pruning ratio may be adapted based on the iteration. The pruning ratio may serve as a mask and may be gradually increased as given by: p c = p t + p i - p t 1 - c - c 0 n 3 , where p i is the initial pruning ratio (e.g., p i = 0 ), p t is the target pruning ratio, n represents the training epoch and p c represents the current pruning ratio for c ∈ { c 0 , … , c 0 + n } , where c represents the current epoch”). However, the second term is not a first difference raised to the power of the floor of the ratio of the index of the training epoch and the number of training epochs per pruning operation as required by the claims. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Dekhovich et al., NEURAL NETWORK RELIEF: A PRUNING ALGORITHM BASED ON NEURAL ACTIVITY, 06/27/2022, https://arxiv.org/pdf/2109.10795v2 discloses an iterative pruning strategy introducing an importance score metric that deactivates unimportant connections to tackle overparameterization in DNNs. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MOLLY CLARKE SIPPEL whose telephone number is (571)272-3270. The examiner can normally be reached Monday - Friday, 7:30 a.m. - 4:30 p.m. ET.. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kakali Chaki can be reached at (571)272-3719. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /M.C.S./Examiner, Art Unit 2122 /KAKALI CHAKI/Supervisory Patent Examiner, Art Unit 2122
Read full office action

Prosecution Timeline

Apr 18, 2024
Application Filed
Aug 06, 2026
Non-Final Rejection mailed — §101, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12670387
SYSTEM, METHOD, AND COMPUTER-READABLE MEDIA FOR LEAKAGE CORRECTION IN GRAPH NEURAL NETWORK BASED RECOMMENDER SYSTEMS
4y 1m to grant Granted Jun 30, 2026
Patent 12664398
SYSTEM, METHOD AND NON-TRANSITORY COMPUTER READABLE MEDIUM
3y 9m to grant Granted Jun 23, 2026
Patent 12657427
Systems, Methods, and Computer Program Products for Determining Uncertainty from a Deep Learning Classification Model
4y 1m to grant Granted Jun 16, 2026
Patent 12632779
HYPERPARAMETER SELECTION USING BUDGET-AWARE BAYESIAN OPTIMIZATION
4y 5m to grant Granted May 19, 2026
Patent 12626098
METHOD AND SYSTEM FOR CREATING AN ENSEMBLE OF NEURAL NETWORK-BASED CLASSIFIERS THAT OPTIMIZES A DIVERSITY METRIC
3y 8m to grant Granted May 12, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
52%
Grant Probability
78%
With Interview (+26.1%)
3y 9m (~1y 5m remaining)
Median Time to Grant
Low
PTA Risk
Based on 23 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month