DETAILED ACTION
1. This communication is in response to the Application No. 18/393,263 filed on December 21, 2023 in which Claims 1-20 are presented for examination.
Notice of Pre-AIA or AIA Status
2. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
3. The information disclosure statements submitted on 12/21/2023 and 05/08/2024 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statements are being considered by the examiner.
Claim Interpretation
4. The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
5. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
6. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier.
Such claim limitation(s) is/are:
“risk management component” in Claims 8-16
“training set characterization component” in Claims 8-16
“baseline characterization component” in Claims 8-16
“monitoring system” in Claims 8-16
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 101
7. 35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
8. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding Claim 1:
Step 1: Claim 1 is a method type claim. Therefore, Claims 1-7 are directed to either a process, machine, manufacture, or composition of matter.
2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation by mathematical calculation but for the recitation of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract ideas.
determining a distribution of feature values within a first dataset of samples used to train an expert system (mental process – determining a distribution of feature values within a first dataset of samples may be performed manually by a user observing/analyzing the first dataset of samples and accordingly using judgement/evaluation to determine a distribution of features based on said analysis of the samples (i.e., if the samples comprise a document, a user is capable of observing/analyzing the document and correspondingly using judgement/evaluation to determine features of the document including layout, format, inclusion of visuals, etc.))
determining, for each of a plurality of baseline datasets received at the expert system, a Kullback–Leibler divergence between the feature values in the baseline dataset and the feature values in the first dataset to provide set of Kullback–Leibler divergence values (mathematical process – determining, for each of a plurality of baseline datasets, a kullback-leibler divergence between the feature values in the baseline dataset and the feature values in the first dataset to provide a set of kullback-leibler divergence values may be performed by mathematical process utilizing an equation for calculating a kullback-leibler divergence – see, for example, Applicant’s specification Par. [0014-0015])
determining a measure of central tendency and a measure of statistical dispersion of the set of Kullback–Leibler divergence values (mathematical process – determining a measure of central tendency and a measure of statistical dispersion of the set of divergence values may be performed by mathematical process utilizing an equation for calculating a measure of central tendency (e.g., mean, median, etc. as supported by Applicant’s specification Par. [0012]) and a measure of statistical dispersion (e.g., variance, standard deviation, range, etc. as supported by Applicant’s specification Par. [0013]))
determining a threshold Kullback–Leibler divergence value, representing an unacceptable divergence of the distribution of feature values within a new dataset from the distribution of feature values within the first dataset of samples, from a desired confidence interval for new datasets, the measure of central tendency of the dataset of Kullback–Leibler divergence values, and the measure of statistical dispersion of the dataset of Kullback–Leibler divergence values (mental process – determining a threshold kullback-leibler divergence value which represents an unacceptable divergence of the distribution of feature values within a new dataset from the distribution of feature values within the first dataset of samples may be performed manually by a user observing/analyzing a desired confidence interval for the new datasets, the measure of central tendency, and the measure of statistical dispersion and accordingly using judgement/evaluation to determine a threshold kullback-leibler divergence value based on the preceding analysis)
determining, for a second dataset received at the expert system, a Kullback Leibler divergence value between the feature values of the second dataset and the feature values of the first dataset (mathematical process – determining, for a second dataset, a kullback-leibler divergence between the feature values of the second dataset and the feature values of the first dataset may be performed by mathematical process utilizing an equation for calculating a kullback-leibler divergence – see, for example, Applicant’s specification Par. [0014-0015])
determining that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples if the Kullback–Leibler divergence value between the feature values of the second dataset and the feature values of the first dataset exceeds the threshold value (mental process – determining that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples may be performed manually by a user observing/analyzing the kullback-leibler divergence between the feature values of the second dataset and the feature values of the first dataset and accordingly using judgement/evaluation to compare the divergence value to the previously determined threshold value to determine that the value exceeds the threshold and hence represents an unacceptable divergence)
2A Prong 2: This judicial exception is not integrated into a practical application.
Additional elements:
[…] train an expert system (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of training a machine learning model with previously determined data without significantly more)
[…] for each of a plurality of baseline datasets received at the expert system […] (Adding insignificant extra-solution activity to the judicial exception – see MPEP 2106.05(g))
[…] for a second dataset received at the expert system […] (Adding insignificant extra-solution activity to the judicial exception – see MPEP 2106.05(g))
2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
[…] train an expert system (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of training a machine learning model with previously determined data without significantly more. This cannot provide an inventive concept)
[…] for each of a plurality of baseline datasets received at the expert system […] (MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer)
[…] for a second dataset received at the expert system […] (MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer)
For the reasons above, Claim 1 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 2-7. The additional limitations of the dependent claims are addressed below.
Regarding Claim 2:
Step 2A Prong 1:
See the rejection of Claim 1 above, which Claim 2 depends on.
Step 2A Prong 2 & Step 2B:
further comprising retraining the expert system with a third dataset of samples in response to determining that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner's note: high level recitation of training a machine learning model with previously determined data without significantly more. This cannot provide an inventive concept)
Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1.
Regarding Claim 3:
Step 2A Prong 1:
See the rejection of Claim 1 above, which Claim 3 depends on.
Step 2A Prong 2 & Step 2B:
taking the expert system offline, such that no further samples are received, in response to determining that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner's note: high level recitation of taking a system offline without significantly more. This cannot provide an inventive concept)
Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1.
Regarding Claim 4:
Step 2A Prong 1:
See the rejection of Claim 1 above, which Claim 4 depends on.
Step 2A Prong 2 & Step 2B:
retraining the expert system with the first dataset of samples in response to determining that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner's note: high level recitation of training a machine learning model with previously determined data without significantly more. This cannot provide an inventive concept)
Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1.
Regarding Claim 5:
Step 2A Prong 1:
See the rejection of Claim 1 above, which Claim 5 depends on.
Step 2A Prong 2 & Step 2B:
wherein the measure of central tendency of the dataset of Kullback–Leibler divergence values is a mean of the dataset of Kullback–Leibler divergence values and the measure of statistical dispersion of the dataset of Kullback Leibler divergence values is a standard deviation of the dataset of Kullback Leibler divergence values (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the measure of central tendency is a mean and the measure of statistical dispersion is a standard deviation does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h))
Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1.
Regarding Claim 6:
Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 6 depends on.
determining the threshold Kullback–Leibler divergence value from the desired confidence interval for new datasets, the measure of central tendency of the dataset of Kullback–Leibler divergence values, and the measure of statistical dispersion of the dataset of Kullback–Leibler divergence values comprises determining a threshold Kullback–Leibler divergence value from the desired confidence interval for new datasets, the measure of central tendency of the dataset of Kullback Leibler divergence values, the measure of statistical dispersion of the dataset of Kullback–Leibler divergence values, and Chebyshev’s inequality (mental process – determining a threshold kullback-leibler divergence value may be performed manually by a user observing/analyzing the desired confidence interval for new datasets, the measure of central tendency, the measure of statistical dispersion, and Chebyshev’s inequality and accordingly using judgement/evaluation to determine a threshold kullback-leibler divergence value based on said analysis)
Step 2A Prong 2 & Step 2B:
Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1.
Regarding Claim 7:
Step 2A Prong 1:
See the rejection of Claim 1 above, which Claim 7 depends on.
Step 2A Prong 2 & Step 2B:
wherein the expert system is a document classification system, and each sample of the first dataset of samples represents a document (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the system is a document classification system and each sample represents a document does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h))
Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1.
Independent Claim 8 recites substantially the same limitations as Claim 1, in the form of a system, including generic computer components. The claim is also directed to performing mental processes/mathematical calculations without significantly more, therefore it is rejected under the same rationale.
For the reasons above, Claim 8 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 9-16. The additional limitations of the dependent claims are addressed below.
Regarding Claim 9:
Step 2A Prong 1:
See the rejection of Claim 8 above, which Claim 9 depends on.
Step 2A Prong 2 & Step 2B:
the expert system comprising a document classification system and the system further comprising a feature extractor that generates each of the first dataset of samples, the second dataset, and the plurality of baseline datasets from corresponding sets of documents (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the expert system comprises a document classification system and the system comprises a feature extractor does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h))
Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 8.
Regarding Claim 10:
Step 2A Prong 1: See the rejection of Claim 9 above, which Claim 10 depends on.
[…] generates each of the first dataset of samples, the second dataset, and the plurality of baseline datasets from corresponding sets of documents using a bag-of-words approach (mental process – generating each of the first dataset of samples, the second dataset, and the plurality of baseline datasets may be performed manually by a user observing/analyzing the sets of documents and accordingly applying a bag-of-words approach to generate the corresponding datasets)
Step 2A Prong 2 & Step 2B:
feature extractor […] (mere instructions to apply the exception using generic computer components cannot provide an inventive concept)
Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 8.
Regarding Claim 11:
Step 2A Prong 1:
See the rejection of Claim 8 above, which Claim 11 depends on.
Step 2A Prong 2 & Step 2B:
retraining component that retrains the expert system in response to a determination at the monitoring system that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner's note: high level recitation of training a machine learning model with previously determined data without significantly more. This cannot provide an inventive concept)
Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 8.
Claim 12 recites substantially the same limitations as Claim 2, in the form of a system, including generic computer components. The claim is also directed to performing mental processes/mathematical calculations without significantly more, therefore it is rejected under the same rationale.
Regarding Claim 13:
Step 2A Prong 1:
See the rejection of Claim 11 above, which Claim 13 depends on.
Step 2A Prong 2 & Step 2B:
retraining component restores at least a portion of the expert system from a backup in response to the determination that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner's note: high level recitation of restoring a portion of the expert system from a backup without significantly more. This cannot provide an inventive concept)
Accordingly, under Step 2A Prong 2 and Step 2B, these additional elements do not integrate the abstract idea into practical application because they do not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1.
Claim 14 recites substantially the same limitations as Claim 4, in the form of a system, including generic computer components. The claim is also directed to performing mental processes/mathematical calculations without significantly more, therefore it is rejected under the same rationale.
Claim 15 recites substantially the same limitations as Claim 5, in the form of a system, including generic computer components. The claim is also directed to performing mental processes/mathematical calculations without significantly more, therefore it is rejected under the same rationale.
Claim 16 recites substantially the same limitations as Claim 6, in the form of a system, including generic computer components. The claim is also directed to performing mental processes/mathematical calculations without significantly more, therefore it is rejected under the same rationale.
Regarding Claim 17:
Step 1: Claim 17 is a method type claim. Therefore, Claims 17-20 are directed to either a process, machine, manufacture, or composition of matter.
2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation by mathematical calculation but for the recitation of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract ideas.
determining a distribution of feature values within a first dataset of samples used to train an expert system, each of the samples in the first dataset of samples representing a document (mental process – determining a distribution of feature values within a first dataset of samples may be performed manually by a user observing/analyzing the first dataset of samples representing a document and accordingly using judgement/evaluation to determine a distribution of features based on said analysis of the samples (i.e., if the samples comprise a document, a user is capable of observing/analyzing the document and correspondingly using judgement/evaluation to determine features of the document including layout, format, inclusion of visuals, etc.))
determining, for each of a plurality of baseline datasets received at the expert system, a Kullback–Leibler divergence between the feature values in the baseline dataset and the feature values in the first dataset to provide set of Kullback–Leibler divergence values (mathematical process – determining, for each of a plurality of baseline datasets, a kullback-leibler divergence between the feature values in the baseline dataset and the feature values in the first dataset to provide a set of kullback-leibler divergence values may be performed by mathematical process utilizing an equation for calculating a kullback-leibler divergence – see, for example, Applicant’s specification Par. [0014-0015])
determining a mean and standard deviation of the set of Kullback–Leibler divergence values (mathematical process – determining a mean and standard deviation of the set of kullback-leibler divergence values may be performed by mathematical process utilizing an equation for calculating mean and standard deviation of the set of divergence values – see Applicant’s specification Par. [0012-0013])
determining a threshold Kullback–Leibler divergence value, representing an unacceptable divergence of the distribution of feature values within a new dataset from the distribution of feature values within the first dataset of samples, from a desired confidence interval for new datasets, the mean and standard deviation of the dataset of Kullback–Leibler divergence values, and Chebyshev’s inequality (mental process – determining a threshold kullback-leibler divergence value which represents an unacceptable divergence of the distribution of feature values within a new dataset from the distribution of feature values within the first dataset of samples may be performed manually by a user observing/analyzing a desired confidence interval for the new datasets, the mean, the standard deviation, and Chebyshev’s inequality and accordingly using judgement/evaluation to determine a threshold kullback-leibler divergence value based on the preceding analysis)
determining, for a second dataset received at the expert system, a Kullback-Leibler divergence value between the feature values of the second dataset and the feature values of the first dataset (mathematical process – determining, for a second dataset, a kullback-leibler divergence between the feature values of the second dataset and the feature values of the first dataset may be performed by mathematical process utilizing an equation for calculating a kullback-leibler divergence – see, for example, Applicant’s specification Par. [0014-0015])
determining that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples if the Kullback–Leibler divergence value between the feature values of the second dataset and the feature values of the first dataset exceed the threshold value (mental process – determining that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples may be performed manually by a user observing/analyzing the kullback-leibler divergence between the feature values of the second dataset and the feature values of the first dataset and accordingly using judgement/evaluation to compare the divergence value to the previously determined threshold value to determine that the value exceeds the threshold and hence represents an unacceptable divergence)
2A Prong 2: This judicial exception is not integrated into a practical application.
Additional elements:
[…] train an expert system […] (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of training a machine learning model with previously determined data without significantly more)
[…] for each of a plurality of baseline datasets received at the expert system […] (Adding insignificant extra-solution activity to the judicial exception – see MPEP 2106.05(g))
[…] for a second dataset received at the expert system […] (Adding insignificant extra-solution activity to the judicial exception – see MPEP 2106.05(g))
2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Additional elements:
[…] train an expert system […] (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of training a machine learning model with previously determined data without significantly more. This cannot provide an inventive concept)
[…] for each of a plurality of baseline datasets received at the expert system […] (MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer)
[…] for a second dataset received at the expert system […] (MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer)
For the reasons above, Claim 17 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 18-20. The additional limitations of the dependent claims are addressed below.
Claim 18 recites substantially the same limitations as Claim 2. The claim is also directed to performing mental processes/mathematical calculations without significantly more, therefore it is rejected under the same rationale.
Claim 19 recites substantially the same limitations as Claim 13. The claim is also directed to performing mental processes/mathematical calculations without significantly more, therefore it is rejected under the same rationale.
Claim 20 recites substantially the same limitations as Claim 4. The claim is also directed to performing mental processes/mathematical calculations without significantly more, therefore it is rejected under the same rationale.
Claim Rejections - 35 USC § 103
9. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
10. Claims 1-2, 4-5, and 7 are rejected under 35 U.S.C. 103 as being unpatentable over Basterrech et al. (hereinafter Basterrech) (“Tracking changes using Kullback-Leibler Divergence for the Continual Learning”), in view of Huynh et al. (hereinafter Huynh) (“Threshold Text Classification with Kullback-Leibler Divergence Approach”).
Regarding Claim 1, Basterrech teaches a method (Basterrech, Pg. 1, Abstract, “This article introduces a novel method for monitoring changes in the probabilistic distribution of multi-dimensional data streams. As a measure of the rapidity of changes, we analyze the popular Kullback-Leibler divergence.”, thus, a method is disclosed) comprising:
determining a distribution of feature values within a first dataset of samples (Basterrech, Pgs. 4-5, “This section presents a pipeline for a KL-divergence-based concept drift detector (KLD). Firstly, we show how to determine the empirical distribution from streaming data, calculate KL-divergence between two data chunks, and detect the drift using a simple threshold scheme.”, thus, as also shown by expression (3) on Pg. 5, a distribution of feature values is determined within a first dataset of samples (streaming data)) used to train an expert system (Basterrech, Pg. 8, “Experimental protocol and implementation. The implementation of the proposed method and the experimental environment are done using the Python v3.9 programming language and a few libraries: NumPy v1.19.5, tsmoothie v1.0.4 and stream-learn v0.8.16. We used the implementation of CART Decision Tree (DT) and Gaussian Naïve Bayes Classifier form sk-learn v1.0.2, and learning protocol was based on the Test-Than-Train [15] evaluation protocol.”, thus, the samples may be used to train an expert system (CART decision tree and Gaussian naïve bayes classifier));
determining, for each of a plurality of baseline datasets received at the expert system, a Kullback–Leibler divergence between the feature values in the baseline dataset and the feature values in the first dataset to provide set of Kullback–Leibler divergence values (Basterrech, Pg. 5, Section 3.2 Proposed KL-divergence-based similarity metric, “Hence, each bin has associated a probability mass function, then we may see changes through any two chunks Si and St computing […] Finally, similarity between any two chunks Si and St is defined […] Note that, in expression (5) each bin has the same relevance in the total sum. We present a slight variation that considers the probability of sampling points in each bin. Then, the proposed similarity measure for comparing two chunks is given by the weighted sum […]”, thus, as also shown by line 6 in Algorithm 1 and Figure 1 on Pg. 7, a kullback-leibler divergence is determined between the feature values in the baseline dataset and the feature values in the first dataset (for each chunk and corresponding bin of continuous data));
determining a measure of central tendency and a measure of statistical dispersion of the set of Kullback–Leibler divergence values (Basterrech, Pgs. 5-6, “We identify a critical point location (i.e., timestamps when a concept drift occurs) when a gradient point li /∈ [¯ l− ασ(l),¯ l+ ασ(l)], where ¯ l denotes the mean of the sequence, σ(l) is the standard deviation, and α is a real-value control parameter.”, therefore, a measure of central tendency (mean) and a measure of statistical dispersion (standard deviation) of the set of kullback-leibler divergence values are determined);
determining a threshold Kullback–Leibler divergence value, representing an unacceptable divergence of the distribution of feature values within a new dataset from the distribution of feature values within the first dataset of samples, from a desired confidence interval for new datasets, the measure of central tendency of the dataset of Kullback–Leibler divergence values, and the measure of statistical dispersion of the dataset of Kullback–Leibler divergence values (Basterrech, Pg. 6, Section 3.4 Main parameters, “Tuning of the decision rule. The α-threshold value presented in the decision rule may produce impact in the result performance (accuracy, confusion matrix, etc.). A lower value of α makes a larger interval, then it is possible to fall in the error of detecting false positive drifts. On the other, a large value of α means that the method may increase the false negative error. That decision rule is commonly used for outlier detection and artifact removal over signals [3, 2]. Note that the mean and standard deviation can be done over a segment of the sequence (for instance last n points), it does not need to be over the full sequence. In addition, in practice the parameter can be dynamically corrected according to the performance of the matching matrix (in cases when the information about changes on the distribution arrives at certain moment)”, therefore, a threshold kullback-leibler divergence value, representing an unacceptable divergence of the distribution of feature values, is determined from the confidence interval (error/accuracy), measure of central tendency and the measure of statistical dispersion);
determining, for a second dataset received at the expert system, a Kullback Leibler divergence value between the feature values of the second dataset and the feature values of the first dataset (Basterrech, Pg. 5, Section 3.2 Proposed KL-divergence-based similarity metric, “Hence, each bin has associated a probability mass function, then we may see changes through any two chunks Si and St computing […] Finally, similarity between any two chunks Si and St is defined […] Note that, in expression (5) each bin has the same relevance in the total sum. We present a slight variation that considers the probability of sampling points in each bin. Then, the proposed similarity measure for comparing two chunks is given by the weighted sum […]”, thus, as also shown by line 6 in Algorithm 1 and Figure 1 on Pg. 7, a kullback-leibler divergence may be determined between the feature values of the second dataset and the feature values of the first dataset (for each chunk and corresponding bin of continuous data). The data stream is also further described by Section 4.2 Experimental setup on Pgs. 7-8, which details how the stream has 10,000 chunks with 250 instances); and
determining that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples if the Kullback–Leibler divergence value between the feature values of the second dataset and the feature values of the first dataset exceeds the threshold value (See introduction of Huynh reference below).
While Basterrech discloses the use of a α-threshold value presented in the decision rule and considers this α-threshold value in Algorithm 1 depicting the KL-divergence-based concept drift detector disclosed on Pg. 7, Basterrech does not explicitly disclose determining that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples if the Kullback–Leibler divergence value between the feature values of the second dataset and the feature values of the first dataset exceeds the threshold value
However, Huynh teaches determining that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples if the Kullback–Leibler divergence value between the feature values of the second dataset and the feature values of the first dataset exceeds the threshold value (Huynh, Pg. 4, “Step 3: Clustering terms, we cluster a pair of terms whose Jensen–Shannon divergence is greater than the threshold. The formula of Jensen–Shannon divergence’s is as follows [3]: J(m1,m2) = log2+ 1 2 m′C h(Pm′ |m1 + Pm′ |m2 −h(Pm′m1 −h(Pm′m2)}”, therefore, Huynh determines that there is an unacceptable divergence from the distribution of features between two datasets if the kullback-leibler divergence (jensen-shannon divergence, which leverages a kullback-leibler divergence) exceeds, or is greater than, the threshold value).
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method, as disclosed by Basterrech to include determining that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples if the Kullback–Leibler divergence value between the feature values of the second dataset and the feature values of the first dataset exceeds the threshold value, as disclosed by Huynh. One of ordinary skill in the art would have been motivated to make this modification to enable the use of a threshold value, which may be utilized to manage probability distributions and control data scale, hence improving system efficiency and accuracy (Huynh, Pg. 2, “Threshold text classification using a vast of keywords has become popular, and this really helps save time for tasks of aggregating, searching, and managing data information [1, 2]. Text classification tasks were proposed in a vast of research which can be applied in numerous fields including searching, retrieving, and extracting information, automated news aggregator, and useful applications in electronic library. Documents are automatically labeled (class/topic) conducted from the extracted content or full-text content using a vast of thresholds [3].” & Pg. 9, “We presented a novel approach using keywords by divergence method with the thresholds, combining Kullback divergence. The method has leveraged a pair of terms whose mutual information for recommending on the trend using similarity measures, support, and confidence to do topic classification tasks. Our method reveals encouraging results using Kullback method where this method outperforms Jaccard method from 4 to 48%.”).
Regarding Claim 2, Basterrech in view of Huynh teaches the method of claim 1, further comprising retraining the expert system with a third dataset of samples in response to determining that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples (Basterrech, Pg. 7, Algorithm 1, which depicts the KL-divergence-based concept drift detector (KLD) and utilizes the α-threshold in expression (7) shown in line 12 to determine whether the divergence is acceptable or unacceptable (if the expression (7) is satisfied) – if the divergence is unacceptable (expression (7) is not satisfied), the while loop continues to ingest and process new chunks of data (third dataset of samples)).
Regarding Claim 4, Basterrech in view of Huynh teaches the method of claim 1, further comprising retraining the expert system with the first dataset of samples in response to determining that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples (Basterrech, Pg. 8, “We used the implementation of CART Decision Tree (DT) and Gaussian Naïve Bayes Classifier form sk-learn v1.0.2, and learning protocol was based on the Test-Than-Train [15] evaluation protocol.”, therefore, utilizing a test-then-train evaluation protocol involves first testing the system then training based on the testing results, hence the system may be tested based on the first dataset of samples and then trained on the same first dataset of samples based on evaluating the divergence).
Regarding Claim 5, Basterrech in view of Huynh teaches the method of claim 1, wherein the measure of central tendency of the dataset of Kullback–Leibler divergence values is a mean of the dataset of Kullback–Leibler divergence values and the measure of statistical dispersion of the dataset of Kullback-Leibler divergence values is a standard deviation of the dataset of Kullback-Leibler divergence values (Basterrech, Pgs. 5-6, “We identify a critical point location (i.e., timestamps when a concept drift occurs) when a gradient point li /∈ [¯ l− ασ(l),¯ l+ ασ(l)], where ¯ l denotes the mean of the sequence, σ(l) is the standard deviation, and α is a real-value control parameter.”, therefore, a measure of central tendency (mean) and a measure of statistical dispersion (standard deviation) of the set of kullback-leibler divergence values are determined).
Regarding Claim 7, Basterrech in view of Huynh teaches the method of claim 1, wherein the expert system is a document classification system, and each sample of the first dataset of samples represents a document (Huynh, Pg. 1, Abstract, “Text classification based on thresholds belongs to the supervised learning method which assigns text material to predefined classes or categories based on different thresholds with divergence approach. These categories are identified by a set of documents trained by an automated algorithm. This work presents an approach of text classification using an automatic keyword extraction algorithm based on the Kullback–Leibler divergence approach. The proposed method is evaluated on 2000 documents in Vietnamese, covering ten topics, collected from various e-journals and news portal Web sites including vietnamnet.vn, vnexpress.net, and so on to generate a completely new set of keywords.”, therefore, the system may comprise a document classification system and each sample may represent a document).
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of claim 1, as disclosed by Basterrech in view of Huynh to include wherein the expert system is a document classification system, and each sample of the first dataset of samples represents a document, as disclosed by Huynh. One of ordinary skill in the art would have been motivated to make this modification to enable the use of the kullback-leibler divergence for document classification, which may improve accuracy and reduce resource consumption for text classification tasks (Huynh, Pg. 2, “Threshold text classification using a vast of keywords has become popular, and this really helps save time for tasks of aggregating, searching, and managing data information [1, 2]. Text classification tasks were proposed in a vast of research which can be applied in numerous fields including searching, retrieving, and extracting information, automated news aggregator, and useful applications in electronic library. Documents are automatically labeled (class/topic) conducted from the extracted content or full-text content using a vast of thresholds [3]. For a variety of proposed methods, the authors usually consider documents’ full-text content to perform text classification tasks [5, 6, 11–14]. This leads to the classification tasks have to process with a huge number of features while the authors often face the limitations of computation resources. As seen from many real applications, we frequently use a set of keywords to look for materials, scientific article, and etc. We can see that such keywords enable us to find the crux of the document and core content[7].”).
11. Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over Basterrech et al. (hereinafter Basterrech) (“Tracking changes using Kullback-Leibler Divergence for the Continual Learning”), in view of Huynh et al. (hereinafter Huynh) (“Threshold Text Classification with Kullback-Leibler Divergence Approach”), further in view of Dossa et al. (hereinafter Dossa) (“An Empirical Investigation of Early Stopping Optimizations in Proximal Policy Optimization”).
Regarding Claim 3, Basterrech in view of Huynh teaches the method of claim 1.
Basterrech in view of Huynh do not explicitly disclose taking the expert system offline, such that no further samples are received, in response to determining that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples.
However, Dossa teaches taking the expert system offline, such that no further samples are received, in response to determining that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples (Dossa, Pg. 1, Abstract, “In this paper, we investigate the effect of one such optimization known as ‘‘early stopping’’ implemented for PPO in the popular openai/spinningup library but not in openai/baselines. This optimization technique, which we refer to as KLE-Stop, can stop the policy update within an epoch if the mean Kullback-Leibler (KL) Divergence between the target policy and current policy becomes too high. More specifically, we conduct experiments to examine the empirical importance of KLE-Stop and its conservative variant KLE-Rollback when they are used in conjunction with other common code-level optimizations.”, thus, in response to determining that there is an unacceptable divergence between datasets, the system may be taken offline as a part of an early stopping optimization).
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of claim 1, as disclosed by Basterrech in view of Huynh to include taking the expert system offline, such that no further samples are received, in response to determining that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples, as disclosed by Dossa. One of ordinary skill in the art would have been motivated to make this modification to enable taking the system offline/early stopping in response to an unacceptable divergence, hence improving system performance by reducing resource consumption and mitigating performance sensitivities (Dossa, Pg. 1, Abstract, “The main findings of our experiments are 1) the performance of PPO is sensitive to the number of update iterations per epoch (K), 2) Early stopping optimizations (KLE-Stop and KLE-Rollback) mitigate such sensitivity by dynamically adjusting the actual number of update iterations within an epoch, 3) Early stopping optimizations could serve as a convenient alternative to tuning on K.”).
12. Claims 6, 17-18, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Basterrech et al. (hereinafter Basterrech) (“Tracking changes using Kullback-Leibler Divergence for the Continual Learning”), in view of Huynh et al. (hereinafter Huynh) (“Threshold Text Classification with Kullback-Leibler Divergence Approach”), further in view of Mardia et al. (hereinafter Mardia) (“Concentration Inequalities for the Empirical Distribution of Discrete Distributions: Beyond the Method of Types”).
Regarding Claim 6, Basterrech in view of Huynh teaches the method of claim 1, wherein determining the threshold Kullback–Leibler divergence value from the desired confidence interval for new datasets, the measure of central tendency of the dataset of Kullback–Leibler divergence values, and the measure of statistical dispersion of the dataset of Kullback–Leibler divergence values comprises determining a threshold Kullback–Leibler divergence value from the desired confidence interval for new datasets, the measure of central tendency of the dataset of Kullback Leibler divergence values, the measure of statistical dispersion of the dataset of Kullback–Leibler divergence values (Basterrech, Pg. 6, Section 3.4 Main parameters, “Tuning of the decision rule. The α-threshold value presented in the decision rule may produce impact in the result performance (accuracy, confusion matrix, etc.). A lower value of α makes a larger interval, then it is possible to fall in the error of detecting false positive drifts. On the other, a large value of α means that the method may increase the false negative error. That decision rule is commonly used for outlier detection and artifact removal over signals [3, 2]. Note that the mean and standard deviation can be done over a segment of the sequence (for instance last n points), it does not need to be over the full sequence. In addition, in practice the parameter can be dynamically corrected according to the performance of the matching matrix (in cases when the information about changes on the distribution arrives at certain moment)”, therefore, a threshold kullback-leibler divergence value, representing an unacceptable divergence of the distribution of feature values, is determined from the confidence interval (error/accuracy), measure of central tendency and the measure of statistical dispersion), and Chebyshev’s inequality (See introduction of Mardia reference below).
Basterrech in view of Huynh does not explicitly disclose determining a threshold kullback-leibler divergence value from […] Chebyshev’s inequality.
However, Mardia teaches determining a threshold kullback-leibler divergence value from […] Chebyshev’s inequality (Mardia, Pg. 2, “This paper focuses on obtaining concentration inequalities of the Kullback–Leibler (KL) divergence between the empirical distribution and the true distribution for discrete distributions. The KL divergence is not an integral probability metric, which makes it difficult to apply the VC inequality and chaining to obtain tight bounds. However, there is fundamental importance in understanding the behavior of the KL divergence.” & Pg. 5, “In this context, it may be more insightful to provide bounds for the centered concentration, i.e. […] for which Theorem 1 provides a bound via Chebyshev’s inequality, but we suspect stronger (exponential) bounds are within reach.”, therefore, Chebyshev’s inequality may be used to determine a threshold kullback-leibler divergence value (bounds)).
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of claim 1, as disclosed by Basterrech in view of Huynh to include determining a threshold kullback-leibler divergence value from […] Chebyshev’s inequality, as disclosed by Mardia. One of ordinary skill in the art would have been motivated to make this modification to enable the use of Chebyshev’s inequality, which may be used to efficiently and accurately bound the kullback-leibler divergence value (Mardia, Pg. 5, “[…] for which Theorem 1 provides a bound via Chebyshev’s inequality, but we suspect stronger (exponential) bounds are within reach. The main technique we use in the variance bound for Theorem 1 is to break the quantity into smooth (pi > 1 n) and non-smooth (pi ≤ 1 n) regions.”).
Regarding Claim 17, Basterrech teaches a method (Basterrech, Pg. 1, Abstract, “This article introduces a novel method for monitoring changes in the probabilistic distribution of multi-dimensional data streams. As a measure of the rapidity of changes, we analyze the popular Kullback-Leibler divergence.”, thus, a method is disclosed) for risk management in a document classification system (See introduction of Huynh reference below), the method comprising:
determining a distribution of feature values within a first dataset of samples (Basterrech, Pgs. 4-5, “This section presents a pipeline for a KL-divergence-based concept drift detector (KLD). Firstly, we show how to determine the empirical distribution from streaming data, calculate KL-divergence between two data chunks, and detect the drift using a simple threshold scheme.”, thus, as also shown by expression (3) on Pg. 5, a distribution of feature values is determined within a first dataset of samples (streaming data)) used to train an expert system (Basterrech, Pg. 8, “Experimental protocol and implementation. The implementation of the proposed method and the experimental environment are done using the Python v3.9 programming language and a few libraries: NumPy v1.19.5, tsmoothie v1.0.4 and stream-learn v0.8.16. We used the implementation of CART Decision Tree (DT) and Gaussian Naïve Bayes Classifier form sk-learn v1.0.2, and learning protocol was based on the Test-Than-Train [15] evaluation protocol.”, thus, the samples may be used to train an expert system (CART decision tree and Gaussian naïve bayes classifier)), each of the samples in the first dataset of samples representing a document (See introduction of Huynh reference below);
determining, for each of a plurality of baseline datasets received at the expert system, a Kullback–Leibler divergence between the feature values in the baseline dataset and the feature values in the first dataset to provide set of Kullback–Leibler divergence values (Basterrech, Pg. 5, Section 3.2 Proposed KL-divergence-based similarity metric, “Hence, each bin has associated a probability mass function, then we may see changes through any two chunks Si and St computing […] Finally, similarity between any two chunks Si and St is defined […] Note that, in expression (5) each bin has the same relevance in the total sum. We present a slight variation that considers the probability of sampling points in each bin. Then, the proposed similarity measure for comparing two chunks is given by the weighted sum […]”, thus, as also shown by line 6 in Algorithm 1 and Figure 1 on Pg. 7, a kullback-leibler divergence is determined between the feature values in the baseline dataset and the feature values in the first dataset (for each chunk and corresponding bin of continuous data));
determining a mean and standard deviation of the set of Kullback–Leibler divergence values (Basterrech, Pgs. 5-6, “We identify a critical point location (i.e., timestamps when a concept drift occurs) when a gradient point li /∈ [¯ l− ασ(l),¯ l+ ασ(l)], where ¯ l denotes the mean of the sequence, σ(l) is the standard deviation, and α is a real-value control parameter.”, therefore, a mean and standard deviation of the set of kullback-leibler divergence values are determined);
determining a threshold Kullback–Leibler divergence value, representing an unacceptable divergence of the distribution of feature values within a new dataset from the distribution of feature values within the first dataset of samples, from a desired confidence interval for new datasets, the mean and standard deviation of the dataset of Kullback–Leibler divergence values (Basterrech, Pg. 6, Section 3.4 Main parameters, “Tuning of the decision rule. The α-threshold value presented in the decision rule may produce impact in the result performance (accuracy, confusion matrix, etc.). A lower value of α makes a larger interval, then it is possible to fall in the error of detecting false positive drifts. On the other, a large value of α means that the method may increase the false negative error. That decision rule is commonly used for outlier detection and artifact removal over signals [3, 2]. Note that the mean and standard deviation can be done over a segment of the sequence (for instance last n points), it does not need to be over the full sequence. In addition, in practice the parameter can be dynamically corrected according to the performance of the matching matrix (in cases when the information about changes on the distribution arrives at certain moment)”, therefore, a threshold kullback-leibler divergence value, representing an unacceptable divergence of the distribution of feature values, is determined from the confidence interval (error/accuracy), measure of central tendency and the measure of statistical dispersion), and Chebyshev’s inequality (See introduction of Mardia reference below);
determining, for a second dataset received at the expert system, a Kullback Leibler divergence value between the feature values of the second dataset and the feature values of the first dataset (Basterrech, Pg. 5, Section 3.2 Proposed KL-divergence-based similarity metric, “Hence, each bin has associated a probability mass function, then we may see changes through any two chunks Si and St computing […] Finally, similarity between any two chunks Si and St is defined […] Note that, in expression (5) each bin has the same relevance in the total sum. We present a slight variation that considers the probability of sampling points in each bin. Then, the proposed similarity measure for comparing two chunks is given by the weighted sum […]”, thus, as also shown by line 6 in Algorithm 1 and Figure 1 on Pg. 7, a kullback-leibler divergence may be determined between the feature values of the second dataset and the feature values of the first dataset (for each chunk and corresponding bin of continuous data). The data stream is also further described by Section 4.2 Experimental setup on Pgs. 7-8, which details how the stream has 10,000 chunks with 250 instances); and
determining that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples if the Kullback–Leibler divergence value between the feature values of the second dataset and the feature values of the first dataset exceed the threshold value (See introduction of Huynh reference below).
Basterrech does not explicitly disclose a method for risk management in a document classification system […] each of the samples in the first dataset of samples representing a document
However, Huynh teaches a method for risk management in a document classification system […] each of the samples in the first dataset of samples representing a document (Huynh, Pg. 1, Abstract, “Text classification based on thresholds belongs to the supervised learning method which assigns text material to predefined classes or categories based on different thresholds with divergence approach. These categories are identified by a set of documents trained by an automated algorithm. This work presents an approach of text classification using an automatic keyword extraction algorithm based on the Kullback–Leibler divergence approach. The proposed method is evaluated on 2000 documents in Vietnamese, covering ten topics, collected from various e-journals and news portal Web sites including vietnamnet.vn, vnexpress.net, and so on to generate a completely new set of keywords.”, therefore, the system may comprise a document classification system and each sample may represent a document).
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of claim 17, as disclosed by Basterrech to include wherein the method is a method for risk management in a document classification system and each of the samples in the first dataset of samples representing a document, as disclosed by Huynh. One of ordinary skill in the art would have been motivated to make this modification to enable the use of the kullback-leibler divergence for document classification, which may improve accuracy and reduce resource consumption for text classification tasks (Huynh, Pg. 2, “Threshold text classification using a vast of keywords has become popular, and this really helps save time for tasks of aggregating, searching, and managing data information [1, 2]. Text classification tasks were proposed in a vast of research which can be applied in numerous fields including searching, retrieving, and extracting information, automated news aggregator, and useful applications in electronic library. Documents are automatically labeled (class/topic) conducted from the extracted content or full-text content using a vast of thresholds [3]. For a variety of proposed methods, the authors usually consider documents’ full-text content to perform text classification tasks [5, 6, 11–14]. This leads to the classification tasks have to process with a huge number of features while the authors often face the limitations of computation resources. As seen from many real applications, we frequently use a set of keywords to look for materials, scientific article, and etc. We can see that such keywords enable us to find the crux of the document and core content[7].”).
Basterrech in view of Huynh does not explicitly disclose determining a threshold kullback-leibler divergence value from […] Chebyshev’s inequality.
However, Mardia teaches determining a threshold kullback-leibler divergence value from […] Chebyshev’s inequality (Mardia, Pg. 2, “This paper focuses on obtaining concentration inequalities of the Kullback–Leibler (KL) divergence between the empirical distribution and the true distribution for discrete distributions. The KL divergence is not an integral probability metric, which makes it difficult to apply the VC inequality and chaining to obtain tight bounds. However, there is fundamental importance in understanding the behavior of the KL divergence.” & Pg. 5, “In this context, it may be more insightful to provide bounds for the centered concentration, i.e. […] for which Theorem 1 provides a bound via Chebyshev’s inequality, but we suspect stronger (exponential) bounds are within reach.”, therefore, Chebyshev’s inequality may be used to determine a threshold kullback-leibler divergence value (bounds)).
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of claim 17, as disclosed by Basterrech in view of Huynh to include determining a threshold kullback-leibler divergence value from […] Chebyshev’s inequality, as disclosed by Mardia. One of ordinary skill in the art would have been motivated to make this modification to enable the use of Chebyshev’s inequality, which may be used to efficiently and accurately bound the kullback-leibler divergence value (Mardia, Pg. 5, “[…] for which Theorem 1 provides a bound via Chebyshev’s inequality, but we suspect stronger (exponential) bounds are within reach. The main technique we use in the variance bound for Theorem 1 is to break the quantity into smooth (pi > 1 n) and non-smooth (pi ≤ 1 n) regions.”).
Regarding Claim 18, Basterrech in view of Huynh in view of Mardia teaches the method of claim 17, further comprising retraining the expert system with a third dataset of samples in response to determining that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples (Basterrech, Pg. 7, Algorithm 1, which depicts the KL-divergence-based concept drift detector (KLD) and utilizes the α-threshold in expression (7) shown in line 12 to determine whether the divergence is acceptable or unacceptable (if the expression (7) is satisfied) – if the divergence is unacceptable (expression (7) is not satisfied), the while loop continues to ingest and process new chunks of data (third dataset of samples)).
Regarding Claim 20, Basterrech in view of Huynh in view of Mardia teaches the method of claim 17, further comprising retraining the expert system with the first dataset of samples in response to determining that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples (Basterrech, Pg. 8, “We used the implementation of CART Decision Tree (DT) and Gaussian Naïve Bayes Classifier form sk-learn v1.0.2, and learning protocol was based on the Test-Than-Train [15] evaluation protocol.”, therefore, utilizing a test-then-train evaluation protocol involves first testing the system then training based on the testing results, hence the system may be tested based on the first dataset of samples and then trained on the same first dataset of samples based on evaluating the divergence).
13. Claims 8-12 and 14-15 is rejected under 35 U.S.C. 103 as being unpatentable over Basterrech et al. (hereinafter Basterrech) (“Tracking changes using Kullback-Leibler Divergence for the Continual Learning”), in view of Kamulete (hereinafter Kamulete) (US PG-PUB 20200410403), further in view of Huynh et al. (hereinafter Huynh) (“Threshold Text Classification with Kullback-Leibler Divergence Approach”).
Regarding Claim 8, Basterrech teaches a system (Basterrech, Pg. 8, “Experimental protocol and implementation. The implementation of the proposed method and the experimental environment are done using the Python v3.9 programming language and a few libraries: NumPy v1.19.5, tsmoothie v1.0.4 and stream-learn v0.8.16. We used the implementation of CART Decision Tree (DT) and Gaussian Naïve Bayes Classifier form sk-learn v1.0.2, and learning protocol was based on the Test-Than-Train [15] evaluation protocol.”, thus, a system is disclosed) comprising: a processor; and a non-transitory computer readable medium storing computer readable instructions executable by the processor to provide: an expert system (Basterrech, Pg. 8, “Experimental protocol and implementation. The implementation of the proposed method and the experimental environment are done using the Python v3.9 programming language and a few libraries: NumPy v1.19.5, tsmoothie v1.0.4 and stream-learn v0.8.16. We used the implementation of CART Decision Tree (DT) and Gaussian Naïve Bayes Classifier form sk-learn v1.0.2, and learning protocol was based on the Test-Than-Train [15] evaluation protocol.”, thus, an expert system is disclosed); and a risk management component (See introduction of Kamulete reference below) comprising:
a training set characterization component that determines a distribution of feature values within a first dataset of samples (Basterrech, Pgs. 4-5, “This section presents a pipeline for a KL-divergence-based concept drift detector (KLD). Firstly, we show how to determine the empirical distribution from streaming data, calculate KL-divergence between two data chunks, and detect the drift using a simple threshold scheme.”, thus, as also shown by expression (3) on Pg. 5, a distribution of feature values is determined within a first dataset of samples (streaming data). Regarding the term “training set characterization component”, See 35 U.S.C. 112(f) claim interpretation above – using broadest reasonable interpretation, the functions of this component are analogous to that of the KL-divergence-based concept drift detector presented by Basterrech) used to train the expert system (Basterrech, Pg. 8, “Experimental protocol and implementation. The implementation of the proposed method and the experimental environment are done using the Python v3.9 programming language and a few libraries: NumPy v1.19.5, tsmoothie v1.0.4 and stream-learn v0.8.16. We used the implementation of CART Decision Tree (DT) and Gaussian Naïve Bayes Classifier form sk-learn v1.0.2, and learning protocol was based on the Test-Than-Train [15] evaluation protocol.”, thus, the samples may be used to train an expert system (CART decision tree and Gaussian naïve bayes classifier));
a baseline characterization component that determines, for each of a plurality of baseline datasets received at the expert system, a Kullback–Leibler divergence between the feature values in the baseline dataset and the feature values in the first dataset to provide a set of Kullback–Leibler divergence values (Basterrech, Pg. 5, Section 3.2 Proposed KL-divergence-based similarity metric, “Hence, each bin has associated a probability mass function, then we may see changes through any two chunks Si and St computing […] Finally, similarity between any two chunks Si and St is defined […] Note that, in expression (5) each bin has the same relevance in the total sum. We present a slight variation that considers the probability of sampling points in each bin. Then, the proposed similarity measure for comparing two chunks is given by the weighted sum […]”, thus, as also shown by line 6 in Algorithm 1 and Figure 1 on Pg. 7, a kullback-leibler divergence is determined between the feature values in the baseline dataset and the feature values in the first dataset (for each chunk and corresponding bin of continuous data). Regarding the term “baseline characterization component”, See 35 U.S.C. 112(f) claim interpretation above – using broadest reasonable interpretation, the functions of this component are analogous to that of the KL-divergence-based concept drift detector presented by Basterrech), determines a measure of central tendency and a measure of statistical dispersion of the set of Kullback–Leibler divergence values (Basterrech, Pgs. 5-6, “We identify a critical point location (i.e., timestamps when a concept drift occurs) when a gradient point li /∈ [¯ l− ασ(l),¯ l+ ασ(l)], where ¯ l denotes the mean of the sequence, σ(l) is the standard deviation, and α is a real-value control parameter.”, therefore, a measure of central tendency (mean) and a measure of statistical dispersion (standard deviation) of the set of kullback-leibler divergence values are determined), and determines a threshold Kullback–Leibler divergence value, representing an unacceptable divergence of the distribution of feature values within a new dataset from the distribution of feature values within the first dataset of samples, from a desired confidence interval for new datasets, the measure of central tendency of the dataset of Kullback–Leibler divergence values, and the measure of statistical dispersion of the dataset of Kullback–Leibler divergence values (Basterrech, Pg. 6, Section 3.4 Main parameters, “Tuning of the decision rule. The α-threshold value presented in the decision rule may produce impact in the result performance (accuracy, confusion matrix, etc.). A lower value of α makes a larger interval, then it is possible to fall in the error of detecting false positive drifts. On the other, a large value of α means that the method may increase the false negative error. That decision rule is commonly used for outlier detection and artifact removal over signals [3, 2]. Note that the mean and standard deviation can be done over a segment of the sequence (for instance last n points), it does not need to be over the full sequence. In addition, in practice the parameter can be dynamically corrected according to the performance of the matching matrix (in cases when the information about changes on the distribution arrives at certain moment)”, therefore, a threshold kullback-leibler divergence value, representing an unacceptable divergence of the distribution of feature values, is determined from the confidence interval (error/accuracy), measure of central tendency and the measure of statistical dispersion); and
a monitoring system that determines, for a second dataset received at the expert system, a Kullback–Leibler divergence value between the feature values of the second dataset and the feature values of the first dataset (Basterrech, Pg. 5, Section 3.2 Proposed KL-divergence-based similarity metric, “Hence, each bin has associated a probability mass function, then we may see changes through any two chunks Si and St computing […] Finally, similarity between any two chunks Si and St is defined […] Note that, in expression (5) each bin has the same relevance in the total sum. We present a slight variation that considers the probability of sampling points in each bin. Then, the proposed similarity measure for comparing two chunks is given by the weighted sum […]”, thus, as also shown by line 6 in Algorithm 1 and Figure 1 on Pg. 7, a kullback-leibler divergence may be determined between the feature values of the second dataset and the feature values of the first dataset (for each chunk and corresponding bin of continuous data). The data stream is also further described by Section 4.2 Experimental setup on Pgs. 7-8, which details how the stream has 10,000 chunks with 250 instances. Regarding the term “monitoring system”, See 35 U.S.C. 112(f) claim interpretation above – using broadest reasonable interpretation, the functions of this component are analogous to that of the concept drift monitoring presented by Basterrech – See Pg. 2) and determines that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples if the Kullback–Leibler divergence value between the feature values of the second dataset and the feature values of the first dataset exceed the threshold value (See introduction of Huynh reference below).
Basterrech does not explicitly disclose a system comprising: a processor; and a non-transitory computer readable medium storing computer readable instructions executable by the processor to provide: […] and a risk management component comprising: […]
However, Kamulete teaches a system comprising: a processor; and a non-transitory computer readable medium storing computer readable instructions executable by the processor (Kamulete, Claim 19, “A computer system comprising: a processor; a memory in communication with the processor, the memory storing instructions that, when executed by the processor cause the processor to perform the method of claim 1.”, thus, a system comprising a processor and a non-transitory computer readable medium storing instructions (see Kamulete claim 20) is disclosed) to provide: […] and a risk management component (Kamulete, Par. [0225], “Systems and methods for D-SOS, as disclosed herein, can be used, by way of example, to detect dataset shift for models in credit risk, marketing science (booking propensity) and anti-money laundering, such as in the context of a financial institution.”, thus, a risk management component is disclosed) comprising: […]
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of claim 8, as disclosed by Basterrech to include where the system comprises a processor; and a non-transitory computer readable medium storing computer readable instructions executable by the processor to provide: […] and a risk management component, as disclosed by Kamulete. One of ordinary skill in the art would have been motivated to make this modification to equip the system with components, such as a processor and memory, which are capable of efficiently and accurately performing distribution-based risk management (Kamulete, Par. [0025], “According to another aspect, there is provided a computer system comprising: a processor; a memory in communication with the processor, the memory storing instructions that, when executed by the processor cause the processor to perform a method as described herein.”).
Basterrech in view of Kamulete do not explicitly disclose determining that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples if the Kullback–Leibler divergence value between the feature values of the second dataset and the feature values of the first dataset exceed the threshold value
However, Huynh teaches determining that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples if the Kullback–Leibler divergence value between the feature values of the second dataset and the feature values of the first dataset exceed the threshold value (Huynh, Pg. 4, “Step 3: Clustering terms, we cluster a pair of terms whose Jensen–Shannon divergence is greater than the threshold. The formula of Jensen–Shannon divergence’s is as follows [3]: J(m1,m2) = log2+ 1 2 m′C h(Pm′ |m1 + Pm′ |m2 −h(Pm′m1 −h(Pm′m2)}”, therefore, Huynh determines that there is an unacceptable divergence from the distribution of features between two datasets if the kullback-leibler divergence (jensen-shannon divergence, which leverages a kullback-leibler divergence) exceeds, or is greater than, the threshold value).
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system, as disclosed by Basterrech in view of Kamulete to include determining that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples if the Kullback–Leibler divergence value between the feature values of the second dataset and the feature values of the first dataset exceeds the threshold value, as disclosed by Huynh. One of ordinary skill in the art would have been motivated to make this modification to enable the use of a threshold value, which may be utilized to manage probability distributions and control data scale, hence improving system efficiency and accuracy (Huynh, Pg. 2, “Threshold text classification using a vast of keywords has become popular, and this really helps save time for tasks of aggregating, searching, and managing data information [1, 2]. Text classification tasks were proposed in a vast of research which can be applied in numerous fields including searching, retrieving, and extracting information, automated news aggregator, and useful applications in electronic library. Documents are automatically labeled (class/topic) conducted from the extracted content or full-text content using a vast of thresholds [3].” & Pg. 9, “We presented a novel approach using keywords by divergence method with the thresholds, combining Kullback divergence. The method has leveraged a pair of terms whose mutual information for recommending on the trend using similarity measures, support, and confidence to do topic classification tasks. Our method reveals encouraging results using Kullback method where this method outperforms Jaccard method from 4 to 48%.”).
Regarding Claim 9, Basterrech in view of Kamulete in view of Huynh teaches the system of claim 8, the expert system comprising a document classification system (Huynh, Pg. 1, Abstract, “Text classification based on thresholds belongs to the supervised learning method which assigns text material to predefined classes or categories based on different thresholds with divergence approach. These categories are identified by a set of documents trained by an automated algorithm. This work presents an approach of text classification using an automatic keyword extraction algorithm based on the Kullback–Leibler divergence approach. The proposed method is evaluated on 2000 documents in Vietnamese, covering ten topics, collected from various e-journals and news portal Web sites including vietnamnet.vn, vnexpress.net, and so on to generate a completely new set of keywords.”, therefore, the system may comprise a document classification system and each sample may represent a document) and the system further comprising a feature extractor that generates each of the first dataset of samples, the second dataset, and the plurality of baseline datasets from corresponding sets of documents (Huynh, Pg. 4, “Step 1: Preprocessing, we extract words representing the content of the documents. We ignore words that include stop words as suggested in [19].”, therefore, the system may extract features by processing the corresponding sets of documents and generating datasets of words representing the content of the documents)
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of claim 8, as disclosed by Basterrech in view of Kamulete in view of Huynh to include wherein the expert system comprising a document classification system and the system further comprising a feature extractor that generates each of the first dataset of samples, the second dataset, and the plurality of baseline datasets from corresponding sets of documents, as disclosed by Huynh. One of ordinary skill in the art would have been motivated to make this modification to enable the use of the kullback-leibler divergence for document classification, which may improve accuracy and reduce resource consumption for text classification tasks (Huynh, Pg. 2, “Threshold text classification using a vast of keywords has become popular, and this really helps save time for tasks of aggregating, searching, and managing data information [1, 2]. Text classification tasks were proposed in a vast of research which can be applied in numerous fields including searching, retrieving, and extracting information, automated news aggregator, and useful applications in electronic library. Documents are automatically labeled (class/topic) conducted from the extracted content or full-text content using a vast of thresholds [3]. For a variety of proposed methods, the authors usually consider documents’ full-text content to perform text classification tasks [5, 6, 11–14]. This leads to the classification tasks have to process with a huge number of features while the authors often face the limitations of computation resources. As seen from many real applications, we frequently use a set of keywords to look for materials, scientific article, and etc. We can see that such keywords enable us to find the crux of the document and core content[7].”).
Regarding Claim 10, Basterrech in view of Kamulete in view of Huynh teaches the system of claim 9, wherein the feature extractor generates each of the first dataset of samples, the second dataset, and the plurality of baseline datasets from corresponding sets of documents using a bag-of-words approach (Huynh, Pg. 3, “A sentence can include some or many words which can be separated by some stop mark such as “.”, “?”, or “!”. Two words or terms in a sentence are determined to co-occur at the same time. Each sentence can be exhibited as a “basket.” We skip grammatical requirement except and the word order while we extract sequences of words [18]. Frequency of a word can be obtained by counting the appearance of the word. For instance, Table 1 reveals the top 10 most frequent terms (denoted as G) [3] and shows the probabilities of occurrence which were normalized to the sum of them to be 1”, therefore, the feature extractor may generate each of the samples using a bag-of-words approach (creating a vocabulary of unique words and frequencies)).
The reasons of obviousness have been noted in the rejection of Claim 8 above and applicable herein.
Regarding Claim 11, Basterrech in view of Kamulete in view of Huynh teaches the system of claim 8, further comprising a retraining component that retrains the expert system in response to a determination at the monitoring system that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples (Basterrech, Pg. 8, “We used the implementation of CART Decision Tree (DT) and Gaussian Naïve Bayes Classifier form sk-learn v1.0.2, and learning protocol was based on the Test-Than-Train [15] evaluation protocol.”, therefore, utilizing a test-then-train evaluation protocol involves first testing the system then training based on the testing results, hence the system may be tested based on the first dataset of samples and then trained/retrained on the same first dataset of samples based on evaluating the divergence).
Regarding Claim 12, Basterrech in view of Kamulete in view of Huynh teaches the system of claim 11, wherein the retraining component retrains the expert system with a third dataset of samples in response to the determination the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples (Basterrech, Pg. 7, Algorithm 1, which depicts the KL-divergence-based concept drift detector (KLD) and utilizes the α-threshold in expression (7) shown in line 12 to determine whether the divergence is acceptable or unacceptable (if the expression (7) is satisfied) – if the divergence is unacceptable (expression (7) is not satisfied), the while loop continues to ingest and process new chunks of data (third dataset of samples)).
Regarding Claim 14, Basterrech in view of Kamulete in view of Huynh teaches the system of claim 11, wherein the retraining component retrains the expert system with the first dataset of samples in response to the determination that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples (Basterrech, Pg. 8, “We used the implementation of CART Decision Tree (DT) and Gaussian Naïve Bayes Classifier form sk-learn v1.0.2, and learning protocol was based on the Test-Than-Train [15] evaluation protocol.”, therefore, utilizing a test-then-train evaluation protocol involves first testing the system then training based on the testing results, hence the system may be tested based on the first dataset of samples and then trained on the same first dataset of samples based on evaluating the divergence)..
Regarding Claim 15, Basterrech in view of Kamulete in view of Huynh teaches the system of claim 8, wherein the measure of central tendency of the dataset of Kullback–Leibler divergence values is a mean of the dataset of Kullback–Leibler divergence values and the measure of statistical dispersion of the dataset of Kullback Leibler divergence values is a standard dispersion of the dataset of Kullback Leibler divergence values (Basterrech, Pgs. 5-6, “We identify a critical point location (i.e., timestamps when a concept drift occurs) when a gradient point li /∈ [¯ l− ασ(l),¯ l+ ασ(l)], where ¯ l denotes the mean of the sequence, σ(l) is the standard deviation, and α is a real-value control parameter.”, therefore, a measure of central tendency (mean) and a measure of statistical dispersion (standard deviation) of the set of kullback-leibler divergence values are determined).
14. Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over Basterrech et al. (hereinafter Basterrech) (“Tracking changes using Kullback-Leibler Divergence for the Continual Learning”), in view of Kamulete (hereinafter Kamulete) (US PG-PUB 20200410403), further in view of Huynh et al. (hereinafter Huynh) (“Threshold Text Classification with Kullback-Leibler Divergence Approach”), further in view of Dossa et al. (hereinafter Dossa) (“An Empirical Investigation of Early Stopping Optimizations in Proximal Policy Optimization”).
Regarding Claim 13, Basterrech in view of Kamulete in view of Huynh teaches the system of claim 11.
Basterrech in view of Kamulete in view of Huynh does not explicitly disclose wherein the retraining component restores at least a portion of the expert system from a backup in response to the determination that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples.
However, Dossa teaches wherein the retraining component restores at least a portion of the expert system from a backup in response to the determination that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples (Dossa, Pg. 2, “Therefore, we also propose to study a more conservative variant KLE-Rollback, which rolls back to the policy just before the threshold was crossed. We conductexperiments on PPO that is augmented with all the code-level optimization implemented in openai/baselines to study if early stopping optimizations (KLE-Stop and KLE Rollback) could improve PPO’s performance.”, thus, the retraining component may restore at least a portion of the expert system from backup (rollback) in response to an unacceptable divergence).
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of claim 11, as disclosed by Basterrech in view of Kamulete in view of Huynh to include the retraining component restores at least a portion of the expert system from a backup in response to the determination that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples, as disclosed by Dossa. One of ordinary skill in the art would have been motivated to make this modification to enable restoring the system from backup in response to an unacceptable divergence, hence improving system performance by reducing resource consumption and mitigating performance sensitivities (Dossa, Pg. 1, Abstract, “The main findings of our experiments are 1) the performance of PPO is sensitive to the number of update iterations per epoch (K), 2) Early stopping optimizations (KLE-Stop and KLE-Rollback) mitigate such sensitivity by dynamically adjusting the actual number of update iterations within an epoch, 3) Early stopping optimizations could serve as a convenient alternative to tuning on K.”).
15. Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over Basterrech et al. (hereinafter Basterrech) (“Tracking changes using Kullback-Leibler Divergence for the Continual Learning”), in view of Kamulete (hereinafter Kamulete) (US PG-PUB 20200410403), further in view of Huynh et al. (hereinafter Huynh) (“Threshold Text Classification with Kullback-Leibler Divergence Approach”), further in view of Mardia et al. (hereinafter Mardia) (“Concentration Inequalities for the Empirical Distribution of Discrete Distributions: Beyond the Method of Types”).
Regarding Claim 16, Basterrech in view of Kamulete in view of Huynh teaches the system of claim 15, wherein determining the threshold Kullback–Leibler divergence value from the desired confidence interval for new datasets, the measure of central tendency of the dataset of Kullback–Leibler divergence values, and the measure of statistical dispersion of the dataset of Kullback–Leibler divergence values comprises determining a threshold Kullback–Leibler divergence value from the desired confidence interval for new datasets, the measure of central tendency of the dataset of Kullback Leibler divergence values, the measure of statistical dispersion of the dataset of Kullback–Leibler divergence values (Basterrech, Pg. 6, Section 3.4 Main parameters, “Tuning of the decision rule. The α-threshold value presented in the decision rule may produce impact in the result performance (accuracy, confusion matrix, etc.). A lower value of α makes a larger interval, then it is possible to fall in the error of detecting false positive drifts. On the other, a large value of α means that the method may increase the false negative error. That decision rule is commonly used for outlier detection and artifact removal over signals [3, 2]. Note that the mean and standard deviation can be done over a segment of the sequence (for instance last n points), it does not need to be over the full sequence. In addition, in practice the parameter can be dynamically corrected according to the performance of the matching matrix (in cases when the information about changes on the distribution arrives at certain moment)”, therefore, a threshold kullback-leibler divergence value, representing an unacceptable divergence of the distribution of feature values, is determined from the confidence interval (error/accuracy), measure of central tendency and the measure of statistical dispersion), and Chebyshev’s inequality (See introduction of Mardia reference below).
Basterrech in view of Kamulete in view of Huynh does not explicitly disclose determining a threshold kullback-leibler divergence value from […] Chebyshev’s inequality.
However, Mardia teaches determining a threshold kullback-leibler divergence value from […] Chebyshev’s inequality (Mardia, Pg. 2, “This paper focuses on obtaining concentration inequalities of the Kullback–Leibler (KL) divergence between the empirical distribution and the true distribution for discrete distributions. The KL divergence is not an integral probability metric, which makes it difficult to apply the VC inequality and chaining to obtain tight bounds. However, there is fundamental importance in understanding the behavior of the KL divergence.” & Pg. 5, “In this context, it may be more insightful to provide bounds for the centered concentration, i.e. […] for which Theorem 1 provides a bound via Chebyshev’s inequality, but we suspect stronger (exponential) bounds are within reach.”, therefore, Chebyshev’s inequality may be used to determine a threshold kullback-leibler divergence value (bounds)).
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of claim 15, as disclosed by Basterrech in view of Kamulete in view of Huynh to include determining a threshold kullback-leibler divergence value from […] Chebyshev’s inequality, as disclosed by Mardia. One of ordinary skill in the art would have been motivated to make this modification to enable the use of Chebyshev’s inequality, which may be used to efficiently and accurately bound the kullback-leibler divergence value (Mardia, Pg. 5, “[…] for which Theorem 1 provides a bound via Chebyshev’s inequality, but we suspect stronger (exponential) bounds are within reach. The main technique we use in the variance bound for Theorem 1 is to break the quantity into smooth (pi > 1 n) and non-smooth (pi ≤ 1 n) regions.”).
16. Claim 19 is rejected under 35 U.S.C. 103 as being unpatentable over Basterrech et al. (hereinafter Basterrech) (“Tracking changes using Kullback-Leibler Divergence for the Continual Learning”), in view of Huynh et al. (hereinafter Huynh) (“Threshold Text Classification with Kullback-Leibler Divergence Approach”), in view of Mardia et al. (hereinafter Mardia) (“Concentration Inequalities for the Empirical Distribution of Discrete Distributions: Beyond the Method of Types”), further in view of Dossa et al. (hereinafter Dossa) (“An Empirical Investigation of Early Stopping Optimizations in Proximal Policy Optimization”).
Regarding Claim 19, Basterrech in view of Huynh in view of Mardia teaches the method of claim 17.
Basterrech in view of Huynh in view of Mardia do not explicitly disclose restoring at least a portion of the expert system from a backup in response to determining that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples.
However, Dossa teaches restoring at least a portion of the expert system from a backup in response to determining that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples (Dossa, Pg. 2, “Therefore, we also propose to study a more conservative variant KLE-Rollback, which rolls back to the policy just before the threshold was crossed. We conduct experiments on PPO that is augmented with all the code-level optimization implemented in openai/baselines to study if early stopping optimizations (KLE-Stop and KLE Rollback) could improve PPO’s performance.”, thus, the system may restore at least a portion of the expert system from backup (rollback) in response to an unacceptable divergence).
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of claim 17, as disclosed by Basterrech in view of Huynh in view of Mardia to include restoring at least a portion of the expert system from a backup in response to determining that the second dataset represents an unacceptable divergence from the distribution of features within the first dataset of samples, as disclosed by Dossa. One of ordinary skill in the art would have been motivated to make this modification to enable restoring the system from backup in response to an unacceptable divergence, hence improving system performance by reducing resource consumption and mitigating performance sensitivities (Dossa, Pg. 1, Abstract, “The main findings of our experiments are 1) the performance of PPO is sensitive to the number of update iterations per epoch (K), 2) Early stopping optimizations (KLE-Stop and KLE-Rollback) mitigate such sensitivity by dynamically adjusting the actual number of update iterations within an epoch, 3) Early stopping optimizations could serve as a convenient alternative to tuning on K.”).
Conclusion
17. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Devika S Maharaj whose telephone number is (571)272-0829. The examiner can normally be reached Monday - Thursday 8:30am - 5:30pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571)270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DEVIKA S MAHARAJ/Examiner, Art Unit 2123