Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 12/15/2022 was filed. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Response to Arguments
Applicant’s arguments filed 5/12/2026 have been fully considered, but are not fully persuasive.
Regarding the claim objection for claim 11, the amendment has overcome the objection; thus, it has been withdrawn.
Regarding the rejections under 35 USC § 101:
Applicant asserts that the amended claims include execution of a trained teacher model, computation of a model reliability based on similarity between datasets, and iterative machine learning training of a student model based on a derived output target, thus requiring operation of trained machine learning models which is not an abstract idea (p.9 last ¶). Examiner respectfully submits that execution of the trained teacher model merely performs the identified abstract ideas of deriving an output target; computing a reliability degree, which incorporates similar ideas to Claim 3, is merely an abstract idea. Claim 1's trained machine learning model is therefore generally linked to the abstract idea.
Applicant further asserts that the Memorandum clarifies that training a machine learning model is not an abstract idea and thus cannot be characterized as a mental step (p. 10 ¶2). Examiner respectfully submits that the training step is not identified as a mental step. Additionally, the Memorandum does not state that training a machine-learning model makes the identified abstract ideas as eligible, but rather that the Examiner should not make "overly broad interpretation" of the machine-learning model's role in the claimed invention. However, the claims do not recite an improvement towards a particular architecture of the machine learning model; thus, the claims remain directed towards an abstract idea.
Applicant further asserts that claim 1 improves the training method by dynamically adjusting the influence of the teacher model based on reliability and training of the student model (p.10 ¶3-4). However, Examiner respectfully submits that the improvement to training is via merely performing the identified abstract ideas (which do not currently require the machine learning architecture); thus, the claims are directed towards the abstract ideas.
Applicant further asserts that claim 8 applies a function for training the student model, which improves training of a machine learning model (p.11 ¶2). Similar to above, the abstract ideas (deriving a target output), not the training of the student model using the target output, are the general crux of the claimed inventions which do not describe a particular architecture of a machine learning model. Thus, the improvement is directed towards the abstract idea of deriving the target output.
Regarding the rejections under 35 USC § 102/103:
Applicant asserts that Ok derives the reliability degree during its own training process instead of the post-process in the present claims (p.13 ¶2). Examiner respectfully submits that the details of post training derivation are not recited in the claims. Thus, this distinction cannot be read into the relationship of the sample data and the training data.
Applicant further asserts that the combination of Ishii and Ok are through impermissible hindsight reconstruction, as it would require redefining Ishii’s class probability output and Ok’s verification neural network (p.13 last ¶). Examiner respectfully submits that, similar to above, the requirement that the reliability measure is for post-training is not readily apparent in the claimed invention; thus, the combination of Ishii and Ok regarding the reliability degree are obvious to the person having ordinary skill in the art. Additionally, the details regarding the reliability measure in the current claims does not preclude from being interpreted as a probability from Ishii.
Applicant further asserts that Claims 14 and 15 are allowable for the same rationale provided above (P.14 ¶1). Examiner respectfully submits that they remain rejected for the same reasons from the responses provided above.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-15 are rejected under 35 U.S.C. 101 because the claimed invention is directed to
an abstract idea without significantly more. The claims recite mental processes and mathematical concepts. This judicial exception is not integrated into a practical application because the claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception, as explained below.
Step 1 for all Claims:
Claims 1-13 are directed to an apparatus. Claims 14 and 15 are directed to methods. Therefore, Claims 1-15 are directed to one of the four statutory categories of invention, i.e., process, machine, manufacture, or composition of matter.
Regarding Claim 1:
Step 2A, Prong 1:
derive a reliability degree of a trained teacher model based on training data used for training the teacher model and predetermined sample data according to a similarity between the training data and the sample data; As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses deriving the reliability of a teacher model which is making an evaluation based upon training and sample data which can be feasibly performed in the human mind (see MPEP 2106.04(a)(2)(III)).
…and obtain output data…; As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses deriving the output target of a teacher model which is making an evaluation upon sample data and reliability degree which can be feasibly performed in the human mind (see MPEP 2106.04(a)(2)(III)).
derive an output target based on the output data outputted by the teacher model the reliability degree of the teacher model; As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses deriving the output target of a teacher model which is making an evaluation upon sample data and reliability degree which can be feasibly performed in the human mind (see MPEP 2106.04(a)(2)(III)).
Step 2A, Prong 2:
An information processing apparatus comprising at least one processor, wherein the at least one processor is configured to: This limitation is recited at a high level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. The elements are recited at a high-level of generality with no detail of the information processing such that it amounts to no more than mere instructions to apply the exception using a generic computer component (See MPEP 2106.05(f)).
input the sample data to the teacher model… outputted by the teacher model, This limitation is recited at a high level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. The training of a student model is recited at a high-level of generality with no detail of the training process such that it amounts to no more than mere instructions to apply the exception using a generic computer component (See MPEP 2106.05(f)).
and perform training on a student model such that output data obtained by inputting the sample data to the student model approaches the output target. This limitation is recited at a high level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. The training of a student model is recited at a high-level of generality with no detail of the training process such that it amounts to no more than mere instructions to apply the exception using a generic computer component (See MPEP 2106.05(f)).
Step 2B:
An information processing apparatus comprising at least one processor, wherein the at least one processor is configured to: This limitation is recited at a high level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. The information processing is recited at a high-level of generality with no detail of the information processing such that it amounts to no more than mere instructions to apply the exception using a generic computer component (See MPEP 2106.05(f)).
input the sample data to the teacher model… outputted by the teacher model, This limitation is recited at a high level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. The training of a student model is recited at a high-level of generality with no detail of the training process such that it amounts to no more than mere instructions to apply the exception using a generic computer component (See MPEP 2106.05(f)).
and train a student model such that output data obtained by inputting the sample data to the student model approaches the output target. This limitation is recited at a high level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. The training of a student model is recited at a high-level of generality with no detail of the training process such that it amounts to no more than mere instructions to apply the exception using a generic computer component (See MPEP 2106.05(f)).
Regarding Claim 2:
Step 2A, Prong 1:
The information processing apparatus according to claim 1, wherein the at least one processor is configured to: derive the reliability degree for each of a plurality of the teacher models different from each other; As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses deriving the reliability of a model which is making an evaluation based upon data which can be feasibly performed in the human mind (see MPEP 2106.04(a)(2)(III)).
and derive, as the output target, a weighted average according to the reliability degrees with respect to a plurality of pieces of output data obtained by inputting the sample data to each of the plurality of teacher models. This limitation is directed to a mathematical concept such as mathematical relationships, mathematical formulas or equations, or mathematical calculations (MPEP 2106.04(a)(2)(I)).
Regarding Claim 4:
Step 2A, Prong 1:
The information processing apparatus according to claim 1, wherein the at least one processor is configured to: in a case where there are a plurality of pieces of the training data, derive a similarity degree with the sample data for each piece of the training data; As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses deriving the similarity degree which is making an evaluation based upon data groups which can be feasibly performed in the human mind (see MPEP 2106.04(a)(2)(III)).
and derive the reliability degree based on an average of all the similarity degrees derived for each piece of the training data. As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses deriving the reliability degree which is making an evaluation based upon a similarity between data groups which can be feasibly performed in the human mind (see MPEP 2106.04(a)(2)(III)).
Regarding Claim 5:
Step 2A, Prong 1:
The information processing apparatus according to claim 1, wherein the at least one processor is configured to: in a case where there are a plurality of pieces of the training data, derive a similarity degree with the sample data for each piece of the training data; As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses deriving the similarity degree which is making an evaluation based upon data groups which can be feasibly performed in the human mind (see MPEP 2106.04(a)(2)(III)).
and derive the reliability degree based on an average of the similarity degrees selected by a predetermined number in descending order of the similarity degrees among all the similarity degrees derived for each piece of the training data. As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses deriving the reliability degree which is making an evaluation based upon similarities of data groups which can be feasibly performed in the human mind (see MPEP 2106.04(a)(2)(III)).
Regarding Claim 6:
Step 2A, Prong 1:
The information processing apparatus according to claim 1, wherein: and the at least one processor is configured to derive the reliability degree based on a loss value representing a magnitude of an error of output data obtained by inputting the training input data to the teacher model with respect to the training correct answer data. As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses deriving the reliability degree which is making an evaluation based upon loss values and the data they come from which can be feasibly performed in the human mind (see MPEP 2106.04(a)(2)(III)).
Step 2A, Prong 2:
the training data includes a combination of training input data and training correct answer data serving as output data in a case where the training input data is input to the teacher model… This limitation amounts to extra-solution activity of gathering data and outputting for use in the claimed process. As described in MPEP 2106.05(g), limitations that amount to merely adding insignificant extra-solution activity to a judicial exception do not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application.
Step 2B:
the training data includes a combination of training input data and training correct answer data serving as output data in a case where the training input data is input to the teacher model… As discussed above, the additional elements of data gathering which is recited at a high level of generality and amounts to extra-solution activity of transmitting data. The courts have found limitations directed to obtaining information electronically, recited at a high level of generality, to be well-understood, routine, and conventional (see MPEP 2106.05(d)(II), “receiving or transmitting data over a network”, "electronic record keeping," and "storing and retrieving information in memory").
Regarding Claim 7:
Step 2A, Prong 1:
The information processing apparatus according to claim 1 wherein the at least one processor is configured to derive the reliability degree based on an evaluation value representing a degree of matching of output data obtained by inputting evaluation input data included in evaluation data to the teacher model with evaluation correct answer data, the evaluation data including a combination of the evaluation input data and the evaluation correct answer data serving as output data in a case where the evaluation input data is input to the teacher model. As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses deriving the reliability degree which is making an evaluation based upon loss values and the data they come from which can be feasibly performed in the human mind (see MPEP 2106.04(a)(2)(III)).
Regarding Claim 8:
Step 2A, Prong 1:
The claim does not recite any judicial exception. However, it is still directed to the same abstract idea as identified in Claim 1.
Step 2A, Prong 2:
The information processing apparatus according to claim 1, wherein the at least one processor is configured to perform training on the student model such that a loss value representing a magnitude of an error of the output data obtained by inputting the sample data to the student model with respect to the output target is minimized. This limitation is recited at a high level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. The information processing and training is recited at a high-level of generality with no detail of the information processing nor the training such that it amounts to no more than mere instructions to apply the exception using a generic computer component (See MPEP 2106.05(f)).
Step 2B:
The information processing apparatus according to claim 1, wherein the at least one processor is configured to train the student model such that a loss value representing a magnitude of an error of the output data obtained by inputting the sample data to the student model with respect to the output target is minimized. This limitation is recited at a high level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. The information processing and training is recited at a high-level of generality with no detail of the information processing nor the training such that it amounts to no more than mere instructions to apply the exception using a generic computer component (See MPEP 2106.05(f)).
Regarding Claim 9:
Step 2A, Prong 1:
The information processing apparatus according to claim 6, wherein the at least one processor is configured to derive the loss value using at least one measure of cross entropy, Kullback-Leibler divergence, or mean squared error. This limitation is directed to a mathematical concept such as mathematical relationships, mathematical formulas or equations, or mathematical calculations (MPEP 2106.04(a)(2)(I)).
Regarding Claim 10:
Step 2A, Prong 1:
The claim does not recite any judicial exception. However, it is still directed to the same abstract idea as identified in Claim 1.
Step 2A, Prong 2:
The information processing apparatus according to claim 1, wherein: the teacher model and the student model are models in which an input is text data and an output is classification for each character included in the text data, the training data includes a combination of text data and classification for each character included in the text data, and the sample data includes text data. The limitation on the specific usage of text data as input data into the teacher and student models amount to no more than generally linking the use of a judicial exception to a particular technological environment or field of use. As explained by the Supreme Court, a claim directed to a judicial exception cannot be made eligible "simply by having the applicant acquiesce to limiting the reach of the patent for the formula to a particular technological use." Diamond v. Diehr, 450 U.S. 175, 192 n.14, 209 USPQ 1, 10 n. 14 (1981). Thus, limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception do not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application.
Step 2B:
The information processing apparatus according to claim 1, wherein: the teacher model and the student model are models in which an input is text data and an output is classification for each character included in the text data, the training data includes a combination of text data and classification for each character included in the text data, and the sample data includes text data. The limitation on the specific usage of text data as input data into the teacher and student models amount to no more than generally linking the use of a judicial exception to a particular technological environment or field of use. As explained by the Supreme Court, a claim directed to a judicial exception cannot be made eligible "simply by having the applicant acquiesce to limiting the reach of the patent for the formula to a particular technological use." Diamond v. Diehr, 450 U.S. 175, 192 n.14, 209 USPQ 1, 10 n. 14 (1981). Thus, limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception do not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application.
Regarding Claim 11:
Step 2A, Prong 1:
The information processing apparatus according to claim 10, wherein the at least one processor is configured to derive the reliability degree based on a similarity degree with respect to at least one of a meaning, a structure, or appearance of the words between the text data included in the training data and the text data included in the sample data. As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses deriving the reliability degree which is making an evaluation based upon similarities which can be feasibly performed in the human mind (see MPEP 2106.04(a)(2)(III)).
Regarding Claim 12:
Step 2A, Prong 1:
derive an output target based on the probability distribution of the NE label obtained by inputting the sample data to the teacher model and the reliability degree; As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses deriving the output target which is making an evaluation based upon a probability distribution which can be feasibly performed in the human mind (see MPEP 2106.04(a)(2)(III)).
Step 2A, Prong 2:
The information processing apparatus according to claim 10, wherein:
the teacher model and the student model are models in which an input is text data and an output is a probability distribution of an NE label indicating a type of named entity represented by the character, which is given for each character included in the text data, and the at least one processor is configured to: The limitation of text data and probability distributions amount to no more than generally linking the use of a judicial exception to a particular technological environment or field of use. As explained by the Supreme Court, a claim directed to a judicial exception cannot be made eligible "simply by having the applicant acquiesce to limiting the reach of the patent for the formula to a particular technological use." Diamond v. Diehr, 450 U.S. 175, 192 n.14, 209 USPQ 1, 10 n. 14 (1981). Thus, limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception do not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application.
and train the student model such that the probability distribution of the NE label obtained by inputting the sample data to the student model approaches the output target. This limitation is recited at a high level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. The training is recited at a high-level of generality with no detail of the training such that it amounts to no more than mere instructions to apply the exception using a generic computer component (See MPEP 2106.05(f)).
Step 2B:
The information processing apparatus according to claim 10, wherein:
the teacher model and the student model are models in which an input is text data and an output is a probability distribution of an NE label indicating a type of named entity represented by the character, which is given for each character included in the text data, and the at least one processor is configured to: The limitation of text data and probability distributions amount to no more than generally linking the use of a judicial exception to a particular technological environment or field of use. As explained by the Supreme Court, a claim directed to a judicial exception cannot be made eligible "simply by having the applicant acquiesce to limiting the reach of the patent for the formula to a particular technological use." Diamond v. Diehr, 450 U.S. 175, 192 n.14, 209 USPQ 1, 10 n. 14 (1981). Thus, limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception do not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application.
and train the student model such that the probability distribution of the NE label obtained by inputting the sample data to the student model approaches the output target. This limitation is recited at a high level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. The training is recited at a high-level of generality with no detail of the training such that it amounts to no more than mere instructions to apply the exception using a generic computer component (See MPEP 2106.05(f)).
Regarding Claim 13:
Step 2A, Prong 1:
derive an output target based on the probability distribution of the BIO label obtained by inputting the sample data to the teacher model and the reliability degree; As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses deriving the output target which is making an evaluation based upon the probability distribution which can be feasibly performed in the human mind (see MPEP 2106.04(a)(2)(III)).
Step 2A, Prong 2:
The information processing apparatus according to claim 10, wherein: the teacher model and the student model are models in which an input is text data and an output is a probability distribution of a BIO label indicating whether the character corresponds to any of a start position, an internal position, and an external position of the named entity, which is given for each character included in the text data, and the at least one processor is configured to: The limitation of text data and probability distributions amount to no more than generally linking the use of a judicial exception to a particular technological environment or field of use. As explained by the Supreme Court, a claim directed to a judicial exception cannot be made eligible "simply by having the applicant acquiesce to limiting the reach of the patent for the formula to a particular technological use." Diamond v. Diehr, 450 U.S. 175, 192 n.14, 209 USPQ 1, 10 n. 14 (1981). Thus, limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception do not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application.
and perform training on the student model such that the probability distribution of the BIO label obtained by inputting the sample data to the student model approaches the output target. (MPEP 2106.05(f) mere instructions to apply an abstract idea on a computer is not enough to integrate the claim into a practical application.)
Step 2B:
The information processing apparatus according to claim 10, wherein: the teacher model and the student model are models in which an input is text data and an output is a probability distribution of a BIO label indicating whether the character corresponds to any of a start position, an internal position, and an external position of the named entity, which is given for each character included in the text data, and the at least one processor is configured to: The limitation of text data and probability distributions amount to no more than generally linking the use of a judicial exception to a particular technological environment or field of use. As explained by the Supreme Court, a claim directed to a judicial exception cannot be made eligible "simply by having the applicant acquiesce to limiting the reach of the patent for the formula to a particular technological use." Diamond v. Diehr, 450 U.S. 175, 192 n.14, 209 USPQ 1, 10 n. 14 (1981). Thus, limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception do not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application.
and train the student model such that the probability distribution of the BIO label obtained by inputting the sample data to the student model approaches the output target. This limitation is recited at a high level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. The training is recited at a high-level of generality with no detail of the training such that it amounts to no more than mere instructions to apply the exception using a generic computer component (See MPEP 2106.05(f)).
Regarding Claim 14:
Claim 14 recites An information processing method comprising, thus a process, one of the four statutory categories of patentable subject matter. However, Claim 14 recites precisely the abstract ideas and additional elements of Claim 1. Therefore, Step 2A Prong 1, Step 2A Prong 2, and Step 2B analyses remain the same. Claim 14 is rejected as subject-matter ineligible for reasons set forth in the rejections of Claim 1.
Regarding Claim 15:
Claim 15 recites A non-transitory computer-readable medium, thus an article of manufacture, one of the four statutory categories of patentable subject matter. However, Claim 15 recites storing an information processing program for causing a computer to execute a process comprising precisely the abstract ideas and additional elements of Claim 1. Therefore, Step 2A Prong 1 analysis remains the same. As for Step 2A Prong 2 and Step 2B: performance on a computer cannot integrate an abstract idea into a practical application (Step 2A Prong 2) nor provide significantly more than the abstract idea itself (Step 2B) (MPEP 2106.05(f)), and thus Claim 15 is rejected as subject-matter ineligible for reasons set forth in the rejections of Claim 1.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 6-9, 14, and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Ishii (US 20220366678 A1) (hereinafter Ishii), in view of Ok et al. (US 20210406688 A1) (hereinafter Ok).
Regarding Claim 1:
Ishii teaches: An information processing apparatus comprising at least one processor, wherein the at least one processor is configured to: ([0084] “The processor 13 is a computer such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or the like, and controls the entire object detection apparatus 100 by executing a program prepared in advance. Specifically, the processor 13 trains an object detection model described later.”)
derive a reliability degree of a trained teacher model based on training data used for training the teacher model and predetermined sample data… ([0051] “The class classification unit 72 performs a class classification with respect to detection targets based on the feature map, and outputs a classification result. In the example in FIG. 1, the detection targets are four classes of a “bicycle”, a “horse”, a “dog”, and a “car”, and the class classification unit 72 outputs a degree of reliability (a probability) for each class as the classification result.”
[0060] “On the other hand, the teacher model 90 is a model that has been trained in advance using a large number of images, and includes a feature extraction unit 91, a class classification unit 92, and a rectangular position detection unit 93. The input image is also input to the teacher model 90. In the teacher model 90, the feature extraction unit 91 generates a feature map from the input image. The class classification unit 92 outputs a class classification result for detection targets based on the feature map.”
Examiner’s Note: The input image is read as sample data. The training process is conducted by inputting training data, which is read as a large number of images, into the teacher model. The output of the class classification unit is a class classification result, which includes an output of reliability degrees. This can be further confirmed by comparing Fig. 1 and Fig. 2. The bar graphs displaying the probabilities (or reliability degrees) are found at the end of the teacher model.)
Although Ishii teaches deriving a reliability degree, it does not explicitly teach that it is according to a similarity between the training data and the sample data.
However, Ok further teaches:
according to a similarity between the training data and the sample data.
(Ok [0017] “In one general aspect, one or more embodiments include a non-transitory computer readable medium including instructions, which when executed by a processor, configure the processor perform any one, any combination, or all operations and methods described herein.”
[0023] “The training of the verification neural network may further include either of a first reliability model that determines a reliability of a sample point corresponding to sample data with an attribute similar to the training data based on a distance between the sample point and a first central point corresponding to the training data…” Examiner’s Note: Reliability is read as reliability degree. Distance is read as similarity degree between the training data and sample data point.)
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the apparatus which derives reliability degrees of trained teacher models and trains student models taught by Ishii with the feature of deriving the reliability degree based on a similarity degree between the training and sample data taught by Ok in order to verify the reliability of the classification result and to determine future actions based on verification. ([0007-0008] “The method may further include verifying the classification result of the input data when the generated determination of the reliability of the classification result meets a predetermined verification threshold. The method may further include selectively controlling performance of operations of a computing device based on whether the classification result is verified.”)
input the sample data to the teacher model and obtain output data outputted by the teacher model; derive an output target based on output data outputted by the teacher model and the reliability degree by the teacher model; ([0051] “The class classification unit 72 performs a class classification with respect to detection targets based on the feature map, and outputs a classification result. In the example in FIG. 1, the detection targets are four classes of a “bicycle”, a “horse”, a “dog”, and a “car”, and the class classification unit 72 outputs a degree of reliability (a probability) for each class as the classification result.
[0060] “On the other hand, the teacher model 90 is a model that has been trained in advance using a large number of images, and includes a feature extraction unit 91, a class classification unit 92, and a rectangular position detection unit 93. The input image is also input to the teacher model 90. In the teacher model 90, the feature extraction unit 91 generates a feature map from the input image. The class classification unit 92 outputs a class classification result for detection targets based on the feature map. Moreover, the rectangular position detection unit 93 outputs coordinates of each rectangular position to be detected based on the feature map.”
[0061] “A difference between the class classification result output by the student model 80 and the class classification result output by the teacher model 90 is calculated as a classification loss Lcls, and a difference between the coordinates of each rectangular positions output by the student model 80 and the coordinates of each rectangular position output by the teacher model 90 is calculated as a regression loss Lreg.”
Examiner’s Note: The class classification result includes reliability degrees. The class classification result is derived from inputting sample data into the teacher model. The output of the teacher model is then used as an output target to compare with the results of the student model.)
and perform training on a student model such that output data obtained by inputting the sample data to the student model approaches the output target. ([0061] “A difference between the class classification result output by the student model 80 and the class classification result output by the teacher model 90 is calculated as a classification loss Lcls, and a difference between the coordinates of each rectangular positions output by the student model 80 and the coordinates of each rectangular position output by the teacher model 90 is calculated as a regression loss Lreg. As the regression loss Lreg, a difference between the coordinates for each rectangular positions output by the student model 80 and a true value may be used. Then, the student model 80 is trained so that the total loss L represented by the above equation (1) is minimized.”)
Regarding Claim 6:
Ishii/Ok teaches all the limitations of Claim 1. Ishii, via Ishii/Ok, fruther teaches: The information processing apparatus according to claim 1, wherein: the training data includes a combination of training input data and training correct answer data serving as output data in a case where the training input data is input to the teacher model… ([0052] “Correct answer data (also referred to as a true value (“ground truth”)) are prepared in advance for the input image. A difference Lcls in class classification (hereinafter, also referred to as a “classification loss”, and a “loss” is also referred to as a “loss”) is calculated based on a class classification result by the class classification unit 72 and the correct answer data of the class classification.”
[0123] “When training data and a true value for the training data are input, the teacher model 51 infers from the training data and outputs an inference result (step S11).” Examiner’s Note: True value is read as correct answer data, which is included along with the training data.)
and the at least one processor is configured to derive the reliability degree based on a loss value representing a magnitude of an error of output data obtained by inputting the training input data to the teacher model with respect to the training correct answer data. ([0051] “The class classification unit 72 performs a class classification with respect to detection targets based on the feature map, and outputs a classification result. In the example in FIG. 1, the detection targets are four classes of a “bicycle”, a “horse”, a “dog”, and a “car”, and the class classification unit 72 outputs a degree of reliability (a probability) for each class as the classification result.”
[0056] “First, a method called a “Focal Loss (hereinafter, also referred to as “FL”)” will be described. RetinaNet illustrated in FIG. 1 is a method for embedding an “anchor” having a spread for each pixel on the feature map extracted by the feature extraction unit 71, performing the class classification for each anchor, and detecting a rectangular position for each anchor. In particular, the focal loss focuses on a notable anchor among a plurality of anchors included in the feature map for learning. For instance, among a plurality of anchors set on the feature map, anchors that are expected to presence of detection targets are focused on, rather than anchors that correspond to a background. Specifically, a degree of attention is higher with respect to an anchor that is difficult to predict by DNN (Deep Neural Network), that is, an anchor with a greater difference between a correct answer and a prediction. The focal loss FL(p) is expressed by the following equation. Note that “a” is a constant determined based on a class balance of training data.
Examiner’s Note: Paragraph 56 goes into more detail of how it preforms class classifications for image pixels, resulting in reliability degrees, by using feature map extraction. The focal loss is read as the loss value and accounts for correct answers and predictions based on training data.)
Regarding Claim 7:
Ishii/Ok teaches all the limitations of Claim 1. Ishii, via Ishii/Ok, fruther teaches: The information processing apparatus according to claim 1, wherein the at least one processor is configured to derive the reliability degree based on an evaluation value representing a degree of matching of output data obtained by inputting evaluation input data included in evaluation data to the teacher model with evaluation correct answer data, the evaluation data including a combination of the evaluation input data and the evaluation correct answer data serving as output data in a case where the evaluation input data is input to the teacher model. ([0051] “The class classification unit 72 performs a class classification with respect to detection targets based on the feature map, and outputs a classification result. In the example in FIG. 1, the detection targets are four classes of a “bicycle”, a “horse”, a “dog”, and a “car”, and the class classification unit 72 outputs a degree of reliability (a probability) for each class as the classification result.”
[0052] “Correct answer data (also referred to as a true value (“ground truth”)) are prepared in advance for the input image. A difference Lcls in class classification (hereinafter, also referred to as a “classification loss”, and a “loss” is also referred to as a “loss”) is calculated based on a class classification result by the class classification unit 72 and the correct answer data of the class classification. Moreover, a difference (hereinafter, also referred to as a “regression loss”) Lreg between coordinates of the rectangular position detected by the rectangular position detection unit 73 and the correct answer data at the coordinates of the rectangular position is calculated... Accordingly, the learning model is trained so as to minimize a sum (also referred to as a “total loss”) L of the classification loss Lcls and the regression loss Lreg shown below.”
Examiner’s Note: The difference in class classification is read as evaluation value since both are the values pertaining to the difference between the output data and the correct answer data.)
Regarding Claim 8:
Ishii/Ok teaches all the limitations of Claim 1. Ishii, via Ishii/Ok, fruther teaches: The information processing apparatus according to claim 1, wherein the at least one processor is configured to perform training on the student model such that a loss value representing a magnitude of an error of the output data obtained by inputting the sample data to the student model with respect to the output target is minimized. [0061] “A difference between the class classification result output by the student model 80 and the class classification result output by the teacher model 90 is calculated as a classification loss Lcls, and a difference between the coordinates of each rectangular positions output by the student model 80 and the coordinates of each rectangular position output by the teacher model 90 is calculated as a regression loss Lreg. As the regression loss Lreg, a difference between the coordinates for each rectangular positions output by the student model 80 and a true value may be used. Then, the student model 80 is trained so that the total loss L represented by the above equation (1) is minimized.” Examiner’s Note: The output of the teacher model is read as the output target.)
Regarding Claim 9:
Ishii/Ok teaches all the limitations of Claims 1 and 6. Ishii, via Ishii/Ok, further teaches: The information processing apparatus according to claim 6, wherein the at least one processor is configured to derive the loss value using at least one measure of cross entropy, Kullback-Leibler divergence, or mean squared error. ([0063] “Next, an ADL (Adaptive Distillation knowledge Loss) will be described. The ADL is a learning method that applies an idea of the focal loss to the distillation, and a model is trained according to the following policies.”
[0066] From the above policies, ADL is expressed by the following equation. [Math 4] ADL=(1−exp[—KL(q∥p)−βT(q)]).sup.γKL(q∥p))
[0068] “The “KL” denotes a KL Divergence and is also likened to a “KL distance” or simply a “distance”. The KL(q∥p) denotes a function that measures a closeness between values of q and p, and takes a minimum value “0” when q=p.
[0069] The “T” denotes an entropy function and is given by T(q)=−q log[q].”
Examiner’s Note: ADL value is a calculated loss value which includes KL divergence and entropy terms. Using the ADL function is one of the three listed methods (focal loss, distillation, ADL) described by Ishii in the specification that are used to calculate the loss value depending on the case. Focal loss was described in Claim 6.)
Regarding Claim 14:
Independent Claim 14 recites An information processing method comprising precisely the methods of Claim 1. Thus, Claim 14 is rejected for reasons set forth in Claim 1.
Regarding Claim 15:
Independent Claim 15 recites A non-transitory computer-readable storage medium storing an information processing program for causing a computer to execute a process comprising: (0086] “The recording medium 15 is a non-volatile and non-transitory recording medium such as a disk-shaped recording medium or a semiconductor memory, and is formed to be removable from the object detection apparatus 100. The recording medium 15 records various programs executed by the processor 13.”) precisely the methods of Claim 1. Thus, Claim 15 is rejected for reasons set forth in Claim 1.
Claim 2 is rejected under 35 U.S.C. 103 as being unpatentable over Ishii/Ok further in view of Inoshita et al. (US 20220301293 A1) and Gillian et al. (US 20240338572 A1), hereinafter referred to as Inoshita and Gillian, respectively.
Regarding Claim 2:
Ishii/Ok teaches all the limitations of Claim 1, but fails to particularly teach: The information processing apparatus according to claim 1, wherein the at least one processor is configured to: derive the reliability degree for each of a plurality of the teacher models different from each other; and derive, as the output target, a weighted average according to the reliability degrees with respect to a plurality of pieces of output data obtained by inputting the sample data to each of the plurality of teacher models.
However, Inoshita teaches: derive the reliability degree for each of a plurality of the teacher models different from each other; ([0033] “Next, learned teacher models A to C are prepared using a large-scale network in advance. Each of the teacher models A to C recognizes an input image data. Here, since the target classes of the student models are the “person”, the “car”, and the “signal”, models that recognizes the “person”, the “car”, and the “signal” are prepared as the teacher models A to C, respectively. Specifically, the teacher model A recognizes whether the recognition target is the “person” and image data show the “person” or a “non-person” (hereinafter, it is shown using “Not”). Then, as a recognition result, the teacher model A outputs a degree of reliability indicating an accuracy of the recognition for each of the class “person” and the class “Not-person”. Similarly, the teacher model B recognizes whether the recognition target is the “car” and the image data show the “car” or a “Not-car”. Then, as a recognition result, the teacher model 13 outputs a degree of reliability indicating an accuracy of the recognition for each of the class “car” and the class “Not-car”.” Examiner’s Note: Each teacher model (A, B, C) outputs its own reliability degree.)
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the apparatus which derives reliability degrees of trained teacher models and trains student models taught by Ishii with finding the reliability degrees of multiple teacher models in order to add adaptability to the distillation process by using highly-specific and accurate teacher models. ([0038] “In this technique, when a model of the negative-type two class is prepared for various recognition targets as a teacher model, it becomes possible to adapt to any types of target classes of each student model… Therefore, it becomes possible to generate a new target model by combining high-accuracy teacher models in accordance with various needs.”)
However, Inoshita fails to teach: and derive, as the output target, a weighted average according to the reliability degrees with respect to a plurality of pieces of output data obtained by inputting the sample data to each of the plurality of teacher models.
However, Gillian teaches: and derive, as the output target, a weighted average according to the reliability degrees with respect to a plurality of pieces of output data obtained by inputting the sample data to each of the plurality of teacher models. ([0040] “In some examples, a second data set of training data can have predetermined labels associated with each data element.”
[0046] Thus, the output of the encoder decoder teacher model can include, for each data element, one or more labels. In some examples, the one or more labels can each have an associated confidence or weight representing the likelihood that the data element should have that label. For example, if the data element is an image and the labels represent animals that appear in the image, a particular image could have the labels lion and jaguar with lion having a confidence value of 70% and jaguar having a confidence value of 30%.
[0047] “In some examples, the provisional labels generated by each of the plurality of encoder decoder teacher models can be combined together into an aggregated provisional label. The aggregated provisional label can represent an average estimation of the correct label for a particular data element.”
Examiner’s Note: The provisional labels have confidence values associated with them, which are read as reliability degrees. The provisional labels are read as the output target.
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the apparatus which derives reliability degrees of trained teacher models and trains student models taught by Ishii with using an average of reliability degrees of teacher models taught by Gillian in order to train student models accurately with a fraction of the resources. ([0053] “Once the student models have been trained using the provisionally labeled data, the student models can classify new unlabeled data elements received with significantly fewer parameters than are required for the teacher models. The provisional labels provided by the plurality of encoder decoder teacher models enable the student models to be trained to accurately classify unlabeled data elements with a fraction of the resources used by the encoder decoder teacher models.”)
Claims 4-5 rejected under 35 U.S.C. 103 as being unpatentable over Ishii/Ok further in view of Kim et al. (US 20210264209 A1), hereinafter referred to as Kim.
Regarding Claim 4:
Ishii/Ok teach all the limitations of Claims 1, but fails to teach: The information processing apparatus according to claim 1, wherein the at least one processor is configured to: in a case where there are a plurality of pieces of the training data, derive a similarity degree with the sample data for each piece of the training data; and derive the reliability degree based on an average of all the similarity degrees derived for each piece of the training data.
However, Kim teaches: The information processing apparatus according to claim 1, wherein the at least one processor is configured to: in a case where there are a plurality of pieces of the training data, derive a similarity degree with the sample data for each piece of the training data; ([0158] “For example, the processor 110 may calculate an average value of the distance between each of the data and the classification reference… [0113] “As the similarity between the first probability distribution and the second probability distribution increases, the training data set generated by the pseudo anomalous data set may become similar to the first sample data set.”
Examiner’s Note: Distance between the data and classification reference is interpreted as the similarity degree.)
and derive the reliability degree based on an average of all the similarity degrees derived for each piece of the training data. ([0158] “For example, the processor 110 may calculate an average value of the distance between each of the data and the classification reference, and evaluate that the suitability of the training data set is lower as the average value of the distance decreases.”
Examiner’s Note: Data and classification reference are interpreted as training data and sample data, respectively. Suitability is interpreted as reliability degree.)
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the apparatus which derives reliability degrees of trained teacher models and trains student models taught by Ishii with the feature of deriving a similarity degree for each data point and reliability degree based on an average of all similarity degrees taught by Kim in order to help the training of the neural network since the closer the data is in similarity, the more efficient the learning will be. ([0164] “The closer the data is to the classification reference, the more helpful it is to train the classification reference of the neural network. Therefore, if the data included in an arbitrary data set are close to the classification reference on average, the data set may be suitable for learning the neural network. In addition, when comparing two data sets having an average distance of substantially the same range, a case where the data are dense may be a case suitable for learning the classification reference.”)
Regarding Claim 5:
Ishii and Ok teach all the limitations of Claims 1, but fails to teach: wherein the at least one processor is configured to: in a case where there are a plurality of pieces of the training data, derive a similarity degree with the sample data for each piece of the training data;
and derive the reliability degree based on an average of the similarity degrees selected by a predetermined number in descending order of the similarity degrees among all the similarity degrees derived for each piece of the training data.
However, Kim further teaches: The information processing apparatus according to claim 1, wherein the at least one processor is configured to: in a case where there are a plurality of pieces of the training data, derive a similarity degree with the sample data for each piece of the training data; ([0158] “For example, the processor 110 may calculate an average value of the distance between each of the data and the classification reference… [0113] “As the similarity between the first probability distribution and the second probability distribution increases, the training data set generated by the pseudo anomalous data set may become similar to the first sample data set.”
Examiner’s Note: Distance between the data and classification reference is interpreted as the similarity degree.)
and derive the reliability degree based on an average of the similarity degrees selected by a predetermined number in descending order of the similarity degrees among all the similarity degrees derived for each piece of the training data. ([0158] “For example, the processor 110 may calculate an average value of the distance between each of the data and the classification reference, and evaluate that the suitability of the training data set is lower as the average value of the distance decreases.”
[0056] The auto encoder structure may have a structure in which the number of nodes in the hidden layer included in the encoder decreases as a distance from the input layer increases. When the number of nodes in the bottleneck layer (a layer having a smallest number of nodes positioned between an encoder and a decoder) is too small, a sufficient amount of information may not be delivered, and as a result, the number of nodes in the bottleneck layer may be maintained to be a specific number or more (e.g., half of the input layers or more).
Examiner’s Note: Data and classification reference are interpreted as training data and sample data, respectively. Suitability is interpreted as reliability degree. The input layer nodes are read as the descending order of similarity degrees.)
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the apparatus which derives reliability degrees of trained teacher models and trains student models taught by Ishii with the feature of deriving a similarity degree for each data point and reliability degree based on an average of all similarity degrees taught by Kim in order to help the training of the neural network since the closer the data is in similarity, the more efficient the learning will be. ([0164] “The closer the data is to the classification reference, the more helpful it is to train the classification reference of the neural network. Therefore, if the data included in an arbitrary data set are close to the classification reference on average, the data set may be suitable for learning the neural network. In addition, when comparing two data sets having an average distance of substantially the same range, a case where the data are dense may be a case suitable for learning the classification reference.”)
Claims 10 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Ishii/Ok further in view of Park et al. (US 20220207431 A1), hereinafter referred to as Park.
Regarding Claim 10:
Ishii/Ok teaches all the limitations of Claims 1, but fails to teach: the teacher model and the student model are models in which an input is text data and an output is classification for each character included in the text data, the training data includes a combination of text data and classification for each character included in the text data, and the sample data includes text data.
However, Park teaches: The information processing apparatus according to claim 1, wherein: the teacher model and the student model are models in which an input is text data and an output is classification for each character included in the text data, the training data includes a combination of text data and classification for each character included in the text data, and the sample data includes text data. ([0028] “In some example embodiments, the teacher model 12 and the student model 13 may process the training data 11 to generate a score with respect to each class in N predefined classes, where N is an integer greater than 1. The training data 11 may be an arbitrary data that may be classified into classes, and may include, for example, images or text segments. According to an example embodiment, the classes may be predefined or predetermined. In addition, each of the teacher model 12 and the student model 13 may refer to a machine learning model having an arbitrary structure that may be trained based on the training data 11.”)
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the apparatus which derives reliability degrees of trained teacher models and trains student models taught by Ishii with the feature that input into the models are text data and classification data taught by Park in order to properly conduct knowledge distillation from the teacher models to the student model. ([0029] “The knowledge distillation may include two phases. In a first phase, the teacher model 12 may be trained based on the training data 11 and label data LA, and in a second phase, the student model 13 may be trained based on the training data 11, the teacher model 12, and the label data LA. The label data LA is an ideal result corresponding to the training data 11 and may represent classes to which inputs included in the training data 11 belong respectively.”)
Regarding Claim 11:
Ishii/Ok/Park teach all the limitations of Claims 1 and 10. Ok, via Ishii/Ok/Park, further teaches: The information processing apparatus according to claim 10, wherein the at least one processor is configured to derive the reliability degree based on a similarity degree with respect to at least one of a meaning, a structure, or appearance of the words between the text data included in the training data and the text data included in the sample data. ([0065] “As an example, when the classification neural network classifies information of an image or speech, as non-limiting examples of input data to the classification neural network, as being of a particular class, additional processes may be implemented to determine whether this classification result of the classification neural network is appropriate to be relied on with respect to further predetermined applications.” Editor’s Note: Information of speech is interpreted as at least one of a meaning, structure, or appearance of words. Refer to [0023] in Claim 3’s rejection deriving the reliability degree based on similarity degree between the training data and the sample data.)
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the apparatus which derives reliability degrees of trained teacher models and trains student models taught by Ishii with the feature of deriving the reliability degree based on a similarity degree between the training and sample data taught by Ok in order to verify the reliability of the classification result and to determine future actions based on verification. ([0003-0004] “A neural network may be used to classify predetermined objects or patterns represented in input data that may include information of such objects or patterns, such as in an input image information… However, input data to a neural network may be substantially different from training data that was used in the training of the neural network.
[0005] In one general aspect, a processor-implemented method includes implementing a classification neural network to generate a classification result of data input to the classification neural network by generating, with respect to the input data, intermediate hidden values of one or more hidden layers of the classification neural network, generating the classification result of the input data based on the generated intermediate hidden values, and generating a determination of a reliability of the classification result by implementing a verification neural network, input the intermediate hidden values, to generate the determination of the reliability.
[0007-0008] “The method may further include verifying the classification result of the input data when the generated determination of the reliability of the classification result meets a predetermined verification threshold. The method may further include selectively controlling performance of operations of a computing device based on whether the classification result is verified.”)
Claim 12 rejected is under 35 U.S.C. 103 as being unpatentable over Ishii/Ok further in view of Park in view of Sungchul Kim et al. (US 20230143721 A1), hereinafter referred to as Sungchul Kim.
Regarding Claim 12:
Ishii/Ok/Park teach all the limitations of Claims 1 and 10, but fail to teach: the teacher model and the student model are models in which an input is text data and an output is a probability distribution of an NE label indicating a type of named entity represented by the character, which is given for each character included in the text data, and the at least one processor is configured to:
derive an output target based on the probability distribution of the NE label obtained by inputting the sample data to the teacher model and the reliability degree;
and train the student model such that the probability distribution of the NE label obtained by inputting the sample data to the student model approaches the output target.
However, Sungchul Kim teaches: The information processing apparatus according to claim 10, wherein: the teacher model and the student model are models in which an input is text data and an output is a probability distribution of an NE label indicating a type of named entity represented by the character, which is given for each character included in the text data, and the at least one processor is configured to: ([0020] “By the framework, the technology described herein trains the NER model with annotations of the new classes, while distilling from the previous model with both the synthetic data and real text from the new training data.” [0023] “For instance, some functions are carried out by a processor executing instructions stored in memory.”
[0027] “In the NER model, the structured prediction is assigning a class or no class to each word or token in an unstructured text. The encoder 112 receives a text (or token) string 101 and outputs emission scores 103, which are a representation of the likelihood of the word being a certain class. The emission scores 103 are received by the CRF layer 114, which generates the label tag for each word. The label tag may be described as a transition score, which is the likelihood of a word being a certain tag considering the previous word was a certain tag.” Examiner’s Note: Distillation is the process of compressing teacher models into a student model. The previous model is read as the teacher model. Assigning a class is read as assigning a NE label.)
derive an output target based on the probability distribution of the NE label obtained by inputting the sample data to the teacher model and the reliability degree; ([0040] “The distillation from M.sup.t-1 to M.sup.t involves matching the output probability distributions between M.sup.t to M.sup.t-1. A probability distribution provides a probability that an input maps to each of a set of classes, rather than only outputting the most likely class.”)
and train the student model such that the probability distribution of the NE label obtained by inputting the sample data to the student model approaches the output target. ([0037] “The distillation feeds the synthetic embeddings 124 into the original model 110 and the updated model 130. The updated model 130 is trained by updating the parameters of the updated model until the loss measured between the result of the original model 110 and the updated model 130 is minimized. In other words, the updated model 130 is trained to produce a label sequence similar to that produced by the original model 110 in response to the synthetic embedding. The process is repeated with the natural language training data, which is input to both models. The updated model is trained to reduce differences between the output generated by the two models in response to the natural language data.” Examiner’s Note: The original model is read as teacher model, and the updated model is read as the student model.)
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the apparatus which derives reliability degrees of trained teacher models and trains student models taught by Ishii with the use of NE labels taught by Sungchul Kim in order to help re-train models with minimal amount of new training for new classes. ([0001] “It is a technical challenge to teach the classifier to learn new classes without degrading the accuracy when classifying the data of old classes (i.e., avoid catastrophic forgetting).”
[0002] “Embodiments of the technology described herein include a machine classifier that is able to learn new classes with a minimal amount of new training instances for the new classes.”)
Claim 13 rejected is under 35 U.S.C. 103 as being unpatentable over Ishii/Ok further in view of Park in view of Bui et al. (US 20220114476 A1), hereinafter referred to as Bui.
Regarding Claim 13:
Ishii/Ok/Park teach all the limitations of Claims 1 and 10, but fail to teach: the teacher model and the student model are models in which an input is text data and an output is a probability distribution of a BIO label indicating whether the character corresponds to any of a start position, an internal position, and an external position of the named entity, which is given for each character included in the text data, and the at least one processor is configured to:
derive an output target based on the probability distribution of the BIO label obtained by inputting the sample data to the teacher model and the reliability degree;
and train the student model such that the probability distribution of the BIO label obtained by inputting the sample data to the student model approaches the output target.
However, Bui teaches: The information processing apparatus according to claim 10, wherein: the teacher model and the student model are models in which an input is text data and an output is a probability distribution of a BIO label indicating whether the character corresponds to any of a start position, an internal position, and an external position of the named entity, which is given for each character included in the text data, and the at least one processor is configured to: ([0019] “For example, in various implementations, the text sequence labeling system trains a text sequence labeling teacher model (“teacher model”) to generate text sequence labels. The text sequence labeling system then creates pseudo-ground truth text sequence labels utilizing the teacher model and trains a text sequence labeling student model (“student model”) utilizing these pseudo-ground truth text sequence labels… Upon completing this training process, the text sequence labeling system utilizes the trained text sequence labeling model to generate text sequence labels (e.g., keyphrases) from input documents.”
([0103] “In one or more implementations, the text sequence labels (e.g., the predicted text sequence labels 612, the pseudo-text sequence label 712, or other extracted text sequence labels) include beginning, inside, and outside labels (e.g., BIO labels). For example, for each word provided to a text sequence labeling model from text data, the text sequence labeling model outputs a tag indicating whether the word corresponds to the beginning (or first word) in a keyphrase, the inside (or non-first word) of a keyphrase, or whether the word does not correspond to a keyphrase (e.g., the word is outside the keyphrase).”)
derive an output target based on the probability distribution of the BIO label obtained by inputting the sample data to the teacher model and the reliability degree; ([0053] “As mentioned above, in one or more implementations, the text sequence labeling system 106 trains the student model 110 using the trained teacher model 108 as a baseline model. Further, in various implementations, the text sequence labeling system 106 determines whether the trained student model 110 yields improved results over the teacher model 108. Indeed, in various implementations, the text sequence labeling system 106 compares validation outputs between the two models to determine whether the student model 110 has improved over the teacher model 108.”
[0103] “In one or more implementations, the text sequence labels (e.g., the predicted text sequence labels 612, the pseudo-text sequence label 712, or other extracted text sequence labels) include beginning, inside, and outside labels (e.g., BIO labels).”)
and perform training on the student model such that the probability distribution of the BIO label obtained by inputting the sample data to the student model approaches the output target. ([0053] “As mentioned above, in one or more implementations, the text sequence labeling system 106 trains the student model 110 using the trained teacher model 108 as a baseline model. Further, in various implementations, the text sequence labeling system 106 determines whether the trained student model 110 yields improved results over the teacher model 108. Indeed, in various implementations, the text sequence labeling system 106 compares validation outputs between the two models to determine whether the student model 110 has improved over the teacher model 108.”
[0103] “In one or more implementations, the text sequence labels (e.g., the predicted text sequence labels 612, the pseudo-text sequence label 712, or other extracted text sequence labels) include beginning, inside, and outside labels (e.g., BIO labels).”)
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the apparatus which derives reliability degrees of trained teacher models and trains student models taught by Ishii with the use of BIO labels taught by Bui in order to accurately and efficiently improve text sequence labeling machine-learning models. ([0002] Implementations of the present disclosure provide benefits and/or solve one or more of the foregoing or other problems in the art with systems, non-transitory computer-readable media, and methods that accurately and efficiently utilize a joint-learning self-distillation approach to improve text sequence labeling machine-learning models.)
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JOSEP HAN whose telephone number is (703)756-1346. The examiner can normally be reached Mon-Fri 9am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kakali Chaki can be reached on (571) 272-3719. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/J.H./Examiner, Art Unit 2122
/KAKALI CHAKI/Supervisory Patent Examiner, Art Unit 2122