Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Regarding the previous rejection of claims under 35 U.S.C. 112, amendments to the claims have overcome the rejections, which are withdrawn.
Regarding the previous rejection of claims under 35 U.S.C. 102 or 103, Applicant’s arguments are directed towards amended claims which have not been previously examined, and for which new grounds of rejection under 35 U.S.C. 103 are given below.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1–4 rejected under 35 U.S.C. 103 over Leontiadis et al., “It’s always personal: Using Early Exits for Efficient On-Device CNN Personalisation,” 2021, The 22nd International Workshop on Mobile Computing Systems and Applications (HotMobile ’21), https://doi.org/10.1145/3446382.3448359 (hereafter Leontiadis) in view of Garcia et al., US Pre-Grant Publication No. 2019/0258928 (hereafter Garcia).
Regarding claim 1 and analogous claim 3:
Leontiadis teaches:
(bold only) “for a machine learning model that includes a plurality of preliminarily trained layers and a plurality of output layers”: Leontiadis, section 3.1.1, “Deriving a multi-exit model. Given a CNN model [a plurality of … layers], PersEPhonEE converts it into a multi-exit network by attaching a number of early exits [a plurality of output layers] along its path. For the exit architecture, we utilise a uniform design for all early exits, adopting the intermediate classifier structure of MSDNet [10].”
“the plurality of output layers being formed according to a downstream task and each having a same configuration”: Leontiadis, section 3.1.2, “The multi-exit global model is then trained offine by utilising a generic training set for the target task [being formed according to a downstream task]”; Leontiadis, section 3.1.1., “Deriving a multi-exit model. Given a CNN model, PersEPhonEE converts it into a multi-exit network by attaching a number of early exits along its path. For the exit architecture, we utilise a uniform design for all early exits [each having a same configuration], adopting the intermediate classifier structure of MSDNet [10].”
“the plurality of output layers including a first output layer and a plurality of second output layers, the first output layer being coupled to a final layer of the plurality of preliminarily trained layers, each of the plurality of second output layers being coupled to an output of a corresponding one of the plurality of preliminarily trained layers other than the final layer of the plurality of preliminarily trained layers”: , Fig. 1,
PNG
media_image1.png
284
500
media_image1.png
Greyscale
[showing, under Personalized Training, a series of Personalized Early Classifiers coupled to outputs of the Frozen Backbone, hence, the plurality of output layers including a first output layer and a plurality of second output layers, the first output layer being coupled to a final layer of the plurality of preliminarily trained layers, each of the plurality of second output layers being coupled to an output of a corresponding one of the plurality of preliminarily trained layers other than the final layer of the plurality of preliminarily trained layers, “final layer of the plurality of preliminarily trained layers” interpreted as the layer coupled to the last early classifier].
“training only the first output layer and the plurality of second output layers of the machine learning model using the downstream task”: Leontiadis, section 3.1.3, “To overcome these limitations, PersEPhonEE introduces personalised multi-exit networks and an efficient on-device training scheme. We adopt a frozen-backbone training approach that updates the parameters of the early exits only [training only the first output layer and the plurality of second output layers of the machine learning model using the downstream task] (Fig 2).”
“training the entire machine learning model that includes the plurality of preliminarily trained layers and the plurality of output layers using the downstream task”: Leontiadis, section 3.1.2, “If the supplied CNN is not trained, we jointly train the backbone and intermediate exits from scratch using the cost function introduced in [13 [training the entire machine learning model that includes the plurality of preliminarily trained layers and the plurality of output layers using the downstream task].”
Leontiadis does not explicitly teach:
“[a] non-transitory computer-readable recording medium storing a machine learning program for causing a computer to execute a process comprising.”
(bold only) “for a machine learning model that includes a plurality of preliminarily trained layers and a plurality of output layers.”
Garcia teaches:
“[a] non-transitory computer-readable recording medium storing a machine learning program for causing a computer to execute a process comprising”: Garcia, paragraph 0110, “The techniques may be implemented by computer software which, when executed by a computer, causes the computer to implement the method described above and/or to implement the resulting ANN. Such computer software may be stored by a non-transitory machine-readable medium such as a hard disk, optical disk, flash memory or the like, and implemented by data processing apparatus comprising one or more processing elements [non-transitory computer-readable recording medium storing a machine learning program for causing a computer to execute a process comprising]”; Garcia, paragraph 0006, “The present disclosure provides a computer-implemented method of generating a derived artificial neural network (ANN) from a base ANN, the method comprising.”
(bold only) “for a machine learning model that includes a plurality of preliminarily trained layers and a plurality of output layers”: Garcia, paragraph 0076, ““Embodiments of the present disclosure can provide techniques to use an approximation method to modify the structure of a previously trained neural network model (a base ANN) to a new structure ( of a derived ANN) to avoid training from scratch every time [a machine learning model that includes a plurality of preliminarily trained layers]. In the present examples, the previously trained network is the base ANN 800, the new structure is that of the derived ANN 830, and these processes can be performed by the module 832.”
Garcia and Leontiadis are analogous arts as they are both related to model design and training methods. It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to have combined the use of pre-trained models of Garcia with the teachings of Leontiadis to arrive at the present invention, in order to improve training speed, as stated in Garcia, paragraph 0076, ““Embodiments of the present disclosure can provide techniques to use an approximation method to modify the structure of a previously trained neural network model (a base ANN) to a new structure ( of a derived ANN) to avoid training from scratch every time.”
Regarding claim 2:
Leontiadis as modified by Garcia teaches “[t]he non-transitory computer-readable recording medium according to claim 1.”
Garcia further teaches:
“for a preliminarily trained model that includes the plurality of preliminarily trained layers”: Garcia, paragraph 0076, ““Embodiments of the present disclosure can provide techniques to use an approximation method to modify the structure of a previously trained neural network model (a base ANN) to a new structure ( of a derived ANN) to avoid training from scratch every time [a preliminarily trained model that includes the plurality of preliminarily trained layers]. In the present examples, the previously trained network is the base ANN 800, the new structure is that of the derived ANN 830, and these processes can be performed by the module 832.”
“the recording medium storing the machine learning program for causing the computer to execute the process further comprising”: Garcia, paragraph 0110, “The techniques may be implemented by computer software which, when executed by a computer, causes the computer to implement the method described above and/or to implement the resulting ANN. Such computer software may be stored by a non-transitory machine-readable medium such as a hard disk, optical disk, flash memory or the like, and implemented by data processing apparatus comprising one or more processing elements [the recording medium storing the machine learning program for causing the computer to execute the process further comprising].”
“replacing an output layer coupled to the final layer of the plurality of layers with the first output layer formed according to the downstream task”: Garcia, Figs. 12a-12d,
PNG
media_image2.png
898
425
media_image2.png
Greyscale
[showing layer 1240 acting as a final layer of the plurality of layers]; Garcia, paragraph 0106, “FIG. 12b: the layers 1220, 1230, are replaced by a replacement layer 1225 [replacing an output layer coupled to the final layer of the plurality of layers with the first output layer formed according to the downstream task, formed according to the downstream task interpreted as designed to be part of the same model]. Here, the step 1100 involves detecting data signals (when training data is applied) for a first position x1 , . . . , xN (such as the input to the layer 1220) and a second position y1, . . . , yN (such as the output of the layer 1230) in the ordered series of layers of neurons; the insertion layer is the layer 1225 and the step 1120 involves deriving an initial approximation of at least a set of weights (Winit and/or binit) for the insertion layer 1225 using a least squares approximation from the data signals detected for the first position and a second position. This provides an example of providing the insertion layer to replace one or more layers of the base ANN.”
“and coupling a respective one of the plurality of second output layers that have the same configuration as the first output layer to the respective outputs of the layers other than the final layer of the plurality of layers to generate the machine learning model”: Garcia, Figs. 12a-12d,
PNG
media_image2.png
898
425
media_image2.png
Greyscale
[showing replacement layer 1225 in Fig. 12b connected to previous layer 1210, and further layer 1240, hence all layers in the model remain connected, hence, coupling a respective one of the plurality of second output layers that have the same configuration as the first output layer to the respective outputs of the layers other than the final layer of the plurality of layers to generate the machine learning model];
Garcia and Leontiadis are analogous arts as they are both related to the training of machine learning models. It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to have combined the layer replacement of Garcia with the teachings of Leontiadis to arrive at the present invention, in order to provide an efficient mechanism for trying different model variants, as stated in Garcia, paragraph 0004, “Designing and training such a DNN is typically very time consuming. When a new DNN is developed for a given task, many so-called hyper-parameters (parameters related to the overall structure of the network) must be chosen empirically. For each possible combination of structural hyper-parameters, a new network is typically trained from scratch and evaluated. While progress has been made on hardware (such as Graphical Processing Units providing efficient single instruction multiple data (SIMD) execution) and software (such as a DNN library developed by NVIDIA called cuDNN) to speed-up the training time of a single structure of a DNN, the exploration of a large set of possible structures remains still potentially slow.”
Regarding claim 4:
Claim 4 is analogous to claim 2, except that lacks the use of the non-transitory computer-readable recording medium. Claim 4 is therefore rejected by the same reasoning.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Geng et al., “RomeBERT: Robust Training of Multi-Exit BERT,” 2021, arXiv:2101.09755v1, discloses a model with multiple exits in which the model as a whole is trained and separately, the exits are trained.
Kouris et al., “Multi-Exit Semantic Segmentation Networks,” July 2022, arXiv:2106.03527v3, discloses a model which is trained end-to-end with “vanilla” early-exit heads, in which one head per iteration is updated, then the backbone is trained and all exits are trained jointly (see section 4.3).
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to VINCENT SPRAUL whose telephone number is (703) 756-1511. The examiner can normally be reached M-F 9:00 am - 5:00 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, MICHAEL HUNTLEY can be reached at (303) 297-4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/VAS/Examiner, Art Unit 2129
/MICHAEL J HUNTLEY/Supervisory Patent Examiner, Art Unit 2129