Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
This Office action has been issued in response to amendment filed on 07/22/2026, Claims (1-24) are pending. Applicants' arguments have been carefully and respectfully considered and addressed. Accordingly, this action has been made FINAL necessitated by amendment.
Claims (1-24) are presented for examination.
Response to Arguments
Applicants' arguments have been carefully and respectfully considered and addressed. The arguments presented are moot based on amendment.
Regarding claim 13, the Applicant’s argument does not provide evidence whether the claim invokes or doesn’t invoke 35 USC § 112 (f), therefore; the argument is not persuasive.
Regarding Applicant arguments that pertain to the existing claims’ limitations and the amendment, the arguments were fully considered and are moot in view of the new ground rejection wherein Yamamoto Kohei et al. Foreign Patent Application Publication WO 2020105341 A1 (hereinafter Yamamoto) in view of Ashok et al. N2N Learning: Network to Network Compression via Policy Gradient Reinforcement Learning, 2017 (hereinafter Ashok) and further in view of Sanh et al., DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter, 2020 (hereinafter Sanh) and further in view of Vaswani et al. Attention Is All You Need, 2017 (hereinafter Vaswani) and further in view of GE, Shi-ming et al. Foreign Application publication CN 112199717 A (hereinafter GE), wherein the combination of the new ground rejection teaches the claims’ limitations based on the amendment.
Claim Interpretation - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
Use of the word “means” (or “step for”) in a claim with functional language creates a rebuttable presumption that the claim element is to be treated in accordance with 35 U.S.C. 112(f) (pre-AIA 35 U.S.C. 112, sixth paragraph). The presumption that 35 U.S.C. 112(f) (pre-AIA 35 U.S.C. 112, sixth paragraph) is invoked is rebutted when the function is recited with sufficient structure, material, or acts within the claim itself to entirely perform the recited function.
Absence of the word “means” (or “step for”) in a claim creates a rebuttable presumption that the claim element is not to be treated in accordance with 35 U.S.C. 112(f) (pre-AIA 35 U.S.C. 112, sixth paragraph). The presumption that 35 U.S.C. 112(f) (pre-AIA 35 U.S.C. 112, sixth paragraph) is not invoked is rebutted when the claim element recites function but fails to recite sufficiently definite structure, material or acts to perform that function.
Claim elements in this application that use the word “means” (or “step for”) are presumed to invoke 35 U.S.C. 112(f) except as otherwise indicated in an Office action. Similarly, claim elements that do not use the word “means” (or “step for”) are presumed not to invoke 35 U.S.C. 112(f) except as otherwise indicated in an Office action.
Independent claim 13 describe claim limitations “means for generating an initial student model… means for removing a layer… means for providing an input… and means for applying a model loss…” which have been interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because they use a generic placeholder “unit” or “model” coupled with functional language “perform and/or configured” without reciting sufficient structure to achieve the function. Furthermore, the generic placeholder is not preceded by a structural modifier.
Since the claim limitations invokes 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, independent claims 13-18 have been interpreted to cover the corresponding structure described in the specification that achieves the claimed function, and equivalents thereof.
A review of the specification shows that the following appears to be the corresponding structure described in the specification for the 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph limitation: Paragraphs [0005-0007] of the specification describe the functional language as connected to physical hardware when each of the units are performing its duty.
By this interpretation, the examiner is satisfied with the description of the functional language and the use of sufficient structure as described in the specification of the application.
Thus, it appears that claims 13-18 are properly invoking 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
For more information, see MPEP § 2173 et seq. and Supplementary Examination Guidelines for Determining Compliance With 35 U.S.C. 112 and for Treatment of Related Issues in Patent Applications, 76 FR 7162, 7167 (Feb. 9, 2011).
If applicant does not intend to have the claim limitation(s) treated under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may amend the claim so that they will clearly not invoke 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, or present a sufficient showing that the claim recites/recite sufficient structure, material, or acts for performing the claimed function to preclude application of 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
For more information, see MPEP § 2173 et seq. and Supplementary Examination Guidelines for Determining Compliance With 35 U.S.C. 112 and for Treatment of Related Issues in Patent Applications, 76 FR 7162, 7167 (Feb. 9,2011).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-3, 7-9, 13-15 and 19-21 are rejected under AIA 35 U.S.C. 103(a) as being unpatentable over Yamamoto Kohei et al. Foreign Patent Application Publication WO 2020105341 A1 (hereinafter Yamamoto) in view of Ashok et al. N2N Learning: Network to Network Compression via Policy Gradient Reinforcement Learning, 2017 (hereinafter Ashok) and further in view of Sanh et al., DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter, 2020 (hereinafter Sanh) and further in view of Vaswani et al. Attention Is All You Need, 2017 (hereinafter Vaswani) and further in view of GE, Shi-ming et al. Foreign Application publication CN 112199717 A (hereinafter GE).
Regarding claim 1, Yamamoto teaches A method comprising: generating an initial student model from a teacher model (Page. 3, paragraphs 3-7, page. 5, paragraph 5 wherein Yamamoto generates student model based on the input of the teacher model).
Yamamoto teaches removing a layer from the initial student model to generate an intermediate student model (page. 4, paragraphs 2-4, page. 6, paragraphs 2-7, page. 8, paragraphs 3-4 wherein Yamamoto describes different methods to generate an intermediate student model, wherein the method includes deleting certain models from the student models or excluding the auxiliary layer from the network from processing layer of the student model to the middle layer) providing an input to the teacher model and the intermediate student model (page. 3, paragraphs 4-7, page. 4, paragraphs 4-7, page. 8, paragraph 3, page. 9, paragraph 3-4 wherein Yamamoto teaches providing input to the teacher model and the intermediate student model) and outputting a final student model based on the intermediate student model (page. 3, paragraphs 5-7 wherein Yamamoto output the student model based the intermediate student model).
Yamamoto does not teach removing from the initial student model, a layer having a same type as a successive layer of the initial student model to generate an intermediate student model;
Ashok teaches generating compressed student architecture through sequential decisions to retain or remove layers. Algorithm 1 initializes the architecture state from the teacher and applies layer-removal actions to produce reduced candidate architectures. Ashok therefore teaches the architecture-generation and layer-removal framework (§§3.1–3.2.1 and Algorithm 1).
Sanh teaches a student having the same general architecture as its teacher, with the number of layers reduced by a factor of two. Sanh further teaches initializing the student using alternating teacher layers and explains that reducing layer count improves computational efficiency. These disclosures provide a specific reason to omit layers when constructing a teacher-derived student. (See Reference 2, §3, p. 2).
Vaswani teaches an encoder comprising a stack of six identical layers, each containing the same types of constituent sublayers. Vaswani also teaches common output dimensionality throughout the stack. Accordingly, at the complete encoder-layer level, an interior layer and its immediately following layer have the same architectural type. (See Reference 3, §3.1, p. 2).
It would have been obvious to a person in the ordinary skill in the art before the effective filing date of the claimed invention to combine Yamamoto with Ashok, Sanh and Vaswani by incorporating the method of removing from the initial student model, a layer having a same type as a successive layer of the initial student model to generate an intermediate student model of Ashok, Sanh and Vaswani into the method generating an initial student model from a teacher model of Yamamoto. It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to apply Ashok’s layer-removal process to a teacher-derived candidate comprising repeated Transformer encoder layers, as supported by Sanh and Vaswani. In the proposed modification, the candidate architecture before deletion constitutes the initial student model. An interior encoder layer is removed while its immediately following encoder layer is retained. Because both layers have the same encoder-layer structure in the initial student model, the removed layer has the same type as a successive layer, thereby satisfying the recited relationship. A person of ordinary skill would have been motivated to make this modification to reduce model depth, computational requirements, and parameter storage while retaining teacher-guided training. Sanh’s successful reduction of Transformer depth supplies a reasonable expectation that a useful compressed student could be obtained, and Vaswani’s common layer dimensions support connecting the remaining encoder layers after removal. (See Reference 2, §§3–4, and Reference 3, §3.1.)
Accordingly, under the stated construction of “layer,” Ashok in view of Sanh and Vaswani supports an obviousness finding for the amended same-type successive-layer limitation. This finding rests on the identified modification and combined teachings; a rejection of the complete claim additionally requires mapping its remaining limitations.
Yamamoto does not teach applying a model loss function to adjust a set of parameters of the intermediate student model.
However in analogous art of kernel-guided architecture search and knowledge distillation, GE teaches applying a model loss function to adjust a set of parameters of the intermediate student model (page. 3, paragraphs 11-15, page. 5, paragraphs 6-8, page. 7, paragraph 5-6 wherein GE describes monitoring loss function and cross entroy loss function and adjusting super-parameter).
It would have been obvious to a person in the ordinary skill in the art before the effective filing date of the claimed invention to combine Yamamoto with GE by incorporating the method of applying a model loss function to adjust a set of parameters of the intermediate student model of GE into the method of providing an input to the teacher model and the intermediate student model of Yamamoto for the purpose of solving the problem that the precision of the privacy student model is not high (GE: Abstract).
Regarding claim 2, Yamamoto as modified by Ashok, Sanh and Vaswani and GE teach operating the final student model to generate an output based on the input (page. 8, paragraphs 3-4 wherein Yamamoto generates output based on the input).
Regarding claim 3, Yamamoto as modified by Ashok, Sanh and Vaswani and GE teach in which the layer removed from the initial student model has a same type as second a preceding layer of the initial student model (page. 4, paragraphs 2-4, page. 6, paragraphs 2-7, page. 8, paragraphs 3-4 wherein Yamamoto describes different method to generate an intermediate student model, wherein the method includes deleting certain models from the student models or excluding the auxiliary layer from the network from processing layer of the student model to the middle layer), Ashok teaches generating compressed student architecture through sequential decisions to retain or remove layers. Algorithm 1 initializes the architecture state from the teacher and applies layer-removal actions to produce reduced candidate architectures. Ashok therefore teaches the architecture-generation and layer-removal framework (§§3.1–3.2.1 and Algorithm 1). Sanh teaches a student having the same general architecture as its teacher, with the number of layers reduced by a factor of two. Sanh further teaches initializing the student using alternating teacher layers and explains that reducing layer count improves computational efficiency. These disclosures provide a specific reason to omit layers when constructing a teacher-derived student. (See Reference 2, §3, p. 2). Vaswani teaches an encoder comprising a stack of six identical layers, each containing the same types of constituent sublayers. Vaswani also teaches common output dimensionality throughout the stack. Accordingly, at the complete encoder-layer level, an interior layer and its immediately following layer have the same architectural type. (See Reference 3, §3.1, p. 2).
Claims 7, 13, 19 are similar in scope to claim 1 therefore the claims are rejected under similar rationale.
Claims 8, 14, 20 are similar in scope to claim 2 therefore the claims are rejected under similar rationale.
Claims 9, 15, 21 are similar in scope to claim 3 therefore the claims are rejected under similar rationale.
Claims 4-5, 10-11, 16-17 and 22-23 are rejected under AIA 35 U.S.C. 103(a) as being unpatentable over Yamamoto Kohei et al. Foreign Patent Application Publication WO 2020105341 A1 (hereinafter Yamamoto) in view of Ashok et al. N2N Learning: Network to Network Compression via Policy Gradient Reinforcement Learning, 2017 (hereinafter Ashok) and further in view of Sanh et al., DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter, 2020 (hereinafter Sanh) and further in view of Vaswani et al. Attention Is All You Need, 2017 (hereinafter Vaswani) and further in view of GE, Shi-ming et al. Foreign Application publication CN 112199717 A (hereinafter GE) and further in view of Jafari et al. US Patent Application Publication US 20210383238 A1 (hereinafter Jafari).
Regarding claim 4, Yamamoto, Ashok, Sanh, Vaswani and GE do not teach further comprising training the final student model based on a minimization of a cross entropy loss function between the teacher model and the intermediate student model.
However in analogous art of kernel-guided architecture search and knowledge distillation, Jafari teaches training the final student model based on a minimization of a cross entropy loss function between the teacher model and the intermediate student model ([0005], [0048-0049] wherein Jafari describes alleviating the output of the teacher model and enhancing the probability distribution of the output of the student DNN and enhances capturing this information, Jafari calculates the cross-entropy loss function hyper parameter for controlling tradeoff between two loss functions and perform minimization to train the student model ).
It would have been obvious to a person in the ordinary skill in the art before the effective filing date of the claimed invention to combine Jafari with Yamamoto, Ashok, Sanh, Vaswani and GE by incorporating the method of training the final student model based on a minimization of a cross entropy loss function between the teacher model and the intermediate student model of Jafari into the method of providing an input to the teacher model and the intermediate student model of Yamamoto and Yamamoto, Ashok, Sanh, Vaswani and GEfor the purpose of incorporating a minimization step as student NN model that is learning parameters to minimize a loss (Jafari: [0048]).
Regarding claim 5, Yamamoto as modified by Yamamoto, Ashok, Sanh, Vaswani, GE and Jafari teach applying one or more of quantization or pruning to the final student model ([0047] wherein Jafari describes a trained teacher NN model that may be a DNN model that comprises several hidden layers and a large set of learned parameters that configure the operations of such layers and wherein an untrained student NN model that may be a compressed DNN model relative to teacher NN model. For example, compared to teacher NN model, student NN model that may be compressed in one or more of the following ways: fewer number of layers; reduced number of weight parameters per layer; and use of quantized parameters and/or features to simplify computations).
Claims 10, 16, 22 are similar in scope to claim 5 therefore the claims are rejected under similar rationale.
Claims 11, 17, 23 are similar in scope to claim 5 therefore the claims are rejected under similar rationale.
Claims 6, 12, 18 and 24 are rejected under AIA 35 U.S.C. 103(a) as being unpatentable over Yamamoto Kohei et al. Foreign Patent Application Publication WO 2020105341 A1 (hereinafter Yamamoto) in view of Ashok et al. N2N Learning: Network to Network Compression via Policy Gradient Reinforcement Learning, 2017 (hereinafter Ashok) and further in view of Sanh et al., DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter, 2020 (hereinafter Sanh) and further in view of Vaswani et al. Attention Is All You Need, 2017 (hereinafter Vaswani) and further in view of GE, Shi-ming et al. Foreign Application publication CN 112199717 A (hereinafter GE) and further in view of Jafari et al. US Patent Application Publication US 20230222326 A1 (hereinafter Jafari2).
Regarding claim 6, Yamamoto, Ashok, Sanh, Vaswani and GE do not teach setting a hyper parameter of the final student model to control a tradeoff between a model accuracy and a memory consumption.
However in analogous art of kernel-guided architecture search and knowledge distillation, Jafari2 teaches setting a hyper parameter of the final student model to control a tradeoff between a model accuracy and a memory consumption ([0004], [0021] wherein Jafari describes sets the hyper parameter for controlling the trade-off between two losses that includes sophisticated neural network model with faster inference time and reduced computing resource and memory space cost that may with less effort on consumer computing devices).
It would have been obvious to a person in the ordinary skill in the art before the effective filing date of the claimed invention to combine Jafari2 with Yamamoto, Ashok, Sanh, Vaswani and GE by incorporating the method of setting a hyper parameter of the final student model to control a tradeoff between a model accuracy and a memory consumption of Jafari2 into the method of providing an input to the teacher model and the intermediate student model of Yamamoto, Ashok, Sanh, Vaswani and GE for the purpose of determining if the computed updated set of the SNN model parameters improves a performance of the SNN model relative to updated sets of SNN model parameters previously computed during the first training phase in respect of a development dataset that includes a set of development data samples and respective expected outputs, and when the computed updated set of the SNN model parameters does improve the performance, update the SNN model parameters to the computed updated set of the SNN model parameters prior to a next epoch (Jafari2: [0021]).
Claims 12, 18, 24 are similar in scope to claim 5 therefore the claims are rejected under similar rationale.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HASSAN MRABI whose telephone number is (571)272-8875. The examiner can normally be reached on Monday-Friday, 7:30am-5pm. Alt, Friday, EST.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached on 571-270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HASSAN MRABI/Examiner, Art Unit 2144