DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s arguments regarding the 103 rejection have been fully considered but are respectfully not persuasive. The amended limitation part of an or clause. As the first part option has been mapped and cited, the limitation is met. Nonetheless even if it were to be mapped, the second data being processed first data is met by the Li reference since the input data would be fed into the neural network, processed (which one having ordinary skill in the art would recognize is a linear transformation) and the output would be the second data.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-9 are rejected under 35 U.S.C. 103 as being unpatentable over Li et al. US Patent Application Publication US-2016-0078339-A1 (as cited by the IDS on 5/31/2022) in view of Zhou, Peng, et al. "M2kd: Multi-model and multi-level knowledge distillation for incremental learning." arXiv preprint arXiv:1904.01769 (2019).
Regarding claim 1, Li teaches An information processing method to be executed by a computer that performs execution of deep learning on a edge device, the information processing method comprising: obtaining first data; (Paragraph [0020], "Turning now to FIG.1, a block diagram is provided showing aspects of one example of a system architecture suitable for implementing an embodiment of the invention.."[0033] " Training component 126 also receives un-labeled data for training the student DNN from accessing component 112" [0023] “In one embodiment, the client device is capable of receiving input data such as audio and image information usable by a DNN system…” In other words, the receiving data for training corresponds to obtaining first data.)
calculating a first prediction result by inputting the first data into a first prediction model; (Paragraph [0043], "In particular, for each iteration, a small piece of unlabeled (or un-transcribed) data 310 is provided to both student DNN 301 and teacher DNN 302." In other words, the providing (inputting) of data 310 to the teacher DNN 302 would correspond to inputting first data to obtain a first result.)
calculating a second prediction result by inputting the first data into a second prediction model; (Paragraph [0034] “provides the same un-labeled data to the teacher and student DNN models” and [0043], "In particular, for each iteration, a small piece of unlabeled (or un-transcribed) data 310 is provided to both student DNN 301 and teacher DNN 302." In other words, the providing (inputting) of data 310 to the student DNN 301 would correspond to inputting first data to obtain a first second result by the student DNN.).
calculating a degree of similarity between the first prediction result and the second prediction result; [Li, [0059] “At step 550, the student DNN output distribution is evaluated against the teacher DNN output distribution. The evaluation process of step 550 may be carried out by an evaluating component 128 as described in connection to evaluating component 128 of FIG. 1. In one embodiment of step 550, from the output distributions determined in step 540 (determined from the mini-batch or subset of training data used in step 540), a difference is first determined between the output distribution of the student DNN and the output distribution of the teacher DNN.” And [0043] “The error signal may be calculated by determining the KL divergence between distributions 351 and 352,… a vector”]
determining second data which is training data for machine learning based on the degree of similarity; [[0055] “one embodiment step 530 comprises receiving a large amount of un-labeled data for use in training the student DNN, in subsequent steps of method 500… , with each iteration of steps 540 through 560, a new portion (or subset) of un-labeled data may be received and used for determining the output distributions.”]
calculating a third prediction result by inputting the second data into the first prediction model; [abstract “In one embodiment, an iterative process is applied to train the student DNN by minimize the divergence of the output distributions from the teacher and student DNN models. For each iteration until convergence, the difference in the output distributions is used to update the student DNN model, and output distributions are determined again, using the unlabeled training data” the process is iterative, i.e. the data is input into the model multiple times resulting in third, fourth, prediction results etc. Paragraph [0043], "In particular, for each iteration, a small piece of unlabeled (or un-transcribed) data 310 is provided to both student DNN 301 and teacher DNN 302." In other words, the providing (inputting) of data 310 to the teacher DNN 302 would correspond to inputting first data to obtain a first result]
calculating a fourth prediction result by inputting the second data into the second prediction model; (Paragraph [0034] “provides the same un-labeled data to the teacher and student DNN models” and [0043], "In particular, for each iteration, a small piece of unlabeled (or un-transcribed) data 310 is provided to both student DNN 301 and teacher DNN 302." In other words, the providing (inputting) of data 310 to the student DNN 301 would correspond to inputting first data to obtain a first second result by the student DNN.). and
training the second prediction model by machine learning using the second data so that behavior of the second prediction model is brought closer to behavior of the first prediction model, wherein [[0056] “In steps 540 through 560, the student DNN is trained using an iterative process to optimize its output distribution to approximate the output distribution of the teacher DNN; for example, in one embodiment steps 540-560 are repeated until the student output distribution sufficiently converges with (or otherwise becomes close to) the output distribution of the teacher”]
either:
the degree of similarity is whether or not the first prediction result and the second prediction result match, and ; [[0043-44] “if the output has not converged, and in some embodiments the output still appears to be converging, then the student DNN 301 is trained based on the error. For example, as shown at 370, using back propagation the weights of student DNN 301 are updated using the error signal” (note: error signal indicate degree of similarity, if below threshold training continues)]
in the determining, when the first prediction result and the second prediction result do not match, data generated by processing the first data which has been inputted to the first prediction model and the second prediction model is determined as the second data [[0057] “At step 540 … some embodiments one or more full sweepings of the training data are used over successive iterations” (note: full sweep of data reused in successive training iteration discloses first data determined to be used as second data for continuous training before convergence)]
or the degree of similarity is a degree of similarity between a magnitude of a first prediction value in the first prediction result and a magnitude of a second prediction value in the second prediction result, and in the determining, when a difference between the first prediction value and the second prediction value greater than or equal to a threshold value, the data generated by processing the first data which has been inputted to the first prediction model and the second prediction model is determined as the second data, and the second data obtained by processing the first data is data obtained by applying geometric transformation to the first data, data obtained by imparting noise to a value of the first data, or data obtained by applying linear transformation to the value of the first data
While Li teaches KL divergence which functionally minimizes cross-entropy, Zhou more explicitly teaches by calculating training parameters so that a distance between the third prediction result and the fourth prediction result decreases, and updating the second prediction model using the training parameters, the distance being obtained by cross-entropy (pg. 3 §3.1,
PNG
media_image1.png
292
340
media_image1.png
Greyscale
and last ¶ into pg. 4 “Formally, we minimize the cross entropy for the logits between the current model and corresponding teacher models from previous incremental steps”)
It would have been obvious to one having ordinary skill in the art at the time that the invention was filed to combine the teachings of Li with that of Zhou since a combination of known methods would yield predictable results. Cross-entropy is a known statistical distance metric and would operate in a known and predictable manner with the systems above.
Regarding claim 2, Li teaches The information processing method according to claim 1, wherein the first prediction model has a configuration different from a configuration of the second prediction model. (Li [0042], "system 300 includes teacher DNN 302 and a smaller student DNN 301, which is depicted as having fewer nodes on each of its layers 341." See Figure 3. The broadest reasonable interpretation of a model with a different configuration is interpreted to include a model with a different number of nodes (i.e. network architecture). Therefore, the student DNN (second prediction model) have fewer nodes in its layers than the teacher DNN (first prediction model) would correspond to the models having different configurations.).
Regarding claim 3, Li teaches The information processing method according to claim 1, wherein the first prediction model has a processing accuracy different from a processing accuracy of the second prediction model.(Li [0005], "Embodiments of the invention are directed to systems and methods for providing a more accurate DNN model of reduced size for deployment on devices by "learning" the deployed DNN from a DNN with larger capacity (number of hidden nodes). To learn a DNN with a smaller number of hidden nodes, a larger size (more accurate) "teacher" DNN is used to train the smaller "student" DNN." In other words, the teacher DNN (first prediction model) is a "larger…(more accurate)" model than the student DNN (second prediction model). Therefore, it is understood that the models have difference accuracies.).
Regarding claim 4, Li teaches The information processing method according to claim 2, wherein the second prediction model is obtained by making the first prediction model lighter.(Li [0016], "To learn a DNN with a smaller number of hidden nodes, a larger size (more accurate) "teacher" DNN is used for training the smaller "student" DNN." The broadest reasonable interpretation of a lighter prediction model is interpreted to include "a DNN with a smaller number of hidden nodes" In other words, the smaller student DNN (second prediction model) would correspond to a lighter version of the teacher DNN (the first prediction model).).
Regarding claim 5, Li teaches The information processing method according to claim 3, wherein the second prediction model is obtained by making the first prediction model lighter.(Li [0016], "To learn a DNN with a smaller number of hidden nodes, a larger size (more accurate) "teacher" DNN is used for training the smaller "student" DNN." The broadest reasonable interpretation of a lighter prediction model is interpreted to include "a DNN with a smaller number of hidden nodes" In other words, the smaller student DNN (second prediction model) would correspond to a lighter version of the teacher DNN (the first prediction model).).
Regarding claim 6, Li teaches the information processing method according to claim 1, wherein in the training, the second prediction model is trained using more of the second data than other training data.. [0068] “As described previously an advantage of some embodiments of method 500 is that the student DNN may be trained using un-labeled (or un-transcribed data) because its supervised signal (P.sub.L(s|x), which is the output distribution of the teacher DNN) is obtained by passing the un-labeled training data through the teacher DNN model. Without the need for labeled or transcribed training data, much more data becomes available for training.” ]
Regarding claim 7, Li teaches The information processing method according to claim 1, wherein the first prediction model and the second prediction model are neural network models.(Li [0005], "Embodiments of the invention are directed to systems and methods for providing a more accurate DNN model of reduced size for deployment on devices by "learning" the deployed DNN from a DNN with larger capacity (number of hidden nodes). To learn a DNN with a smaller number of hidden nodes, a larger size (more accurate) "teacher" DNN is used to train the smaller "student" DNN." In other words, both the "teacher" (first prediction model) and the "student" (second prediction model) are deep neural networks, which are neural networks.).
Claim 8 recite a system that implements the methods of claims 1 with substantially the same limitations, respectively. Therefore the rejection applied to claims 1 also applies to claims 8. In addition, Claim 8 recites a An information processing system that implement the method of claim 1 with substantially the same limitations, therefore the rejection applied to claim 1 also apply to claim 8. Further , Li discloses processor, memory and non-transitory computer readable medium in [0080] “With reference to FIG.7, computing device 700 includes…memory 712, one or more processors 714, one or more presentation components 716, one or more input/output (I/O) ports 618, one or more 1/0 components 720, and an illustrative power supply 722.” In other words, the obtainer, calculators and trainer are interpreted as corresponding to functionality implemented on a generic computer perform (Instant specification, Page 9, “Obtainer 10,prediction result calculator 20,similarity calculator 30, determiner 40, and trainer 50 are realized by the processor or the like that executes the programs stored in the memory.”). Therefore, the computing environment, including a process and memory, is interpreted to correspond to the components comprising the information processing system.) to perform substantially the same limitations, respectively, as the method of claims 1.
Claim 9 recite a system that implements the methods of claims 1 with substantially the same limitations, respectively. Therefore the rejection applied to claims 1 also applies to claims 9. Li further discloses processor, memory and computer readable medium in [0080], “With reference to FIG.7, computing device 700 includes…memory 712, one or more processors 714, one or more presentation components 716, one or more input/output (I/O) ports 618, one or more 1/0 components 720, and an illustrative power supply 722.” In other words, the obtainer, controller and outputter are interpreted as corresponding to functionality implemented on a generic computer perform. Therefore, the computing environment in figure 7, including a process and memory, is interpreted to correspond to the components comprising the information processing system. In addition, the broadest reasonable interpretation of sensing data is interpreted to include data related to human senses. Therefore, visual or image data is interpreted to correspond sensing data.) to perform substantially the same limitations, respectively, as the method of claims 1.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KEVIN W FIGUEROA whose telephone number is (571)272-4623. The examiner can normally be reached Monday-Friday, 10AM-6PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, MIRANDA HUANG can be reached at (571)270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
KEVIN W FIGUEROA
Primary Examiner
Art Unit 2124
/Kevin W Figueroa/Primary Examiner, Art Unit 2124