Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This action is in response to the application and claims filed 04/16/2024. Claims 1-18 are pending and have been examined. Claims 1-18 are rejected.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 04/16/2024 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Priority
Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The present application claims foreign priority based on KR10-2023-0050916 and KR10-2023-0141649 filed 04/18/2023 and 10/23/2023 respectively. The examiner notes that a certified copy (in Korean) of the above-noted application was retrieved on 5/20/2024. Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55.
Claim Objections
Claim 4 objected to because of the following informalities:
“a domain adaptation process” should read “the domain adaptation”. Since it is the same “domain adaptation” as recited in claim 1.
Appropriate correction is required.
Claim 12 objected to because of the following informalities:
“calculating a representative vale” should read “value”.
Appropriate correction is required.
Specification
The disclosure is objected to because of the following informalities:
Paragraph [0169], [0170], [0173] and [0193] state “features 193” but feature 193 does not exist in the drawings.
Paragraph [0103] state “first predicted label 84 and the second predicted label 85” should read “83 and 84”. (85 does not exist in the drawings.)
Paragraph [0108] states “evaluation dataset 93” should read “92”, (93 is the degree of dispersion)
Paragraph [0138] states “the model 114” should read “144”
Paragraph [0028] “representative vale” should read “representative value”
Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 16 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 16 recites the limitation "evaluating the performance of the first model by comparing the pseudo label and the predicted label." There is insufficient antecedent basis for this limitation in the claim. Claim 16 introduces “predicting a label…” but claim 1 also previously introduced “a predicted label…” in the last limitation of claim 1. Therefore, it is unclear if “the predicted label” recited in claim 16 refers to claim 1 or claim 16. Suggested fix: recite “the label predicted through the first model” or similar.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claim 1, 3, 15-18 provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over claim 1, 2, 4, 8, 12, 14 of copending Application No. 18/213,478 for claims filed 04/30/2026 (reference application) in view of Balaji et al., "Instance Adaptive Adversarial Training: Improved Accuracy Tradeoffs in Neural Nets," (hereinafter Balaji).
Although the claims at issue are not identical, they are not patentably distinct from each other because all of the limitations of the instant application's claims are contained in their counterpart claims of the reference application, with the exception that the instant application recites “source domain”, “target domain”, "adjusting an upper limit of a size constraint of adversarial noise based on a predefined factor" and deriving the adversarial noise "within a range that satisfies the size constraint according to the adjusted upper limit". In plain terms, the reference application derives adversarial noise for a data sample but says nothing about limiting its size, while the instant application caps the size of the noise and adjusts the cap based on a predefined factor. This is not considered to be a significant distinction as the claims of the instant application are directed to the same performance evaluation method comprising the same first model trained on a labeled dataset, second model built by domain adaptation of the first model, pseudo label generated from a noisy sample of an unlabeled evaluation dataset, and evaluation of the first model using the pseudo label. This is a provisional nonstatutory double patenting rejection because the patentably indistinct claims have not in fact been patented. A comparison chart of the claims follows, followed by an analysis.
Instant Application (18/636,548)
Reference Application (18/213,478) Claims filed 04/30/2026
1. A performance evaluation method performed by at least one computing device, the method comprising:
1. A method for evaluating performance, the method being performed by at least one computing device and comprising:
obtaining a first model trained using a labeled dataset of a source domain;
obtaining a first model trained using a labeled dataset;
obtaining a second model built by performing domain adaptation to a target domain on the first model;
obtaining a second model built by performing unsupervised domain adaptation on the first model;
generating a pseudo label for an evaluation dataset of the target domain using the second model; and
generating pseudo labels for an evaluation dataset using the second model, wherein the evaluation dataset is an unlabeled dataset; and
evaluating a performance of the first model using the pseudo label,
evaluating performance of the first model using the pseudo labels,
wherein the evaluation dataset is an unlabeled dataset, and the generating of the pseudo label comprises:
wherein the generating of the pseudo labels comprises:
adjusting an upper limit of a size constraint of adversarial noise based on a predefined factor;
deriving the adversarial noise for a data sample belonging to the evaluation dataset within a range that satisfies the size constraint according to the adjusted upper limit;
deriving adversarial noise for a data sample belonging to the evaluation dataset;
generating a noisy sample by reflecting the derived adversarial noise in the data sample; and
generating a noisy sample by reflecting the derived adversarial noise in the data sample; and
generating a pseudo label for the data sample based on a predicted label of the noisy sample obtained through the second model.
generating a pseudo label for the data sample based on a predicted label of the noisy sample obtained through the second model.
3. The method of claim 1, wherein the domain adaptation and the generating of the pseudo label are performed without using the labeled dataset.
2. The method of claim 1, wherein the unsupervised domain adaptation and the generating of the pseudo labels are performed without using the labeled dataset.
15. The method of claim 1, wherein the deriving of the adversarial noise comprises:
obtaining a first predicted label for the data sample through the second model;
generating a specific noisy sample by reflecting a value of a noise parameter in the data sample;
obtaining a second predicted label for the specific noisy sample through the second model;
updating the value of the noise parameter in a direction to increase a difference between the first predicted label and the second predicted label; and
calculating the adversarial noise for the data sample based on the updated value of the noise parameter.
4. The method of claim 1, wherein the deriving of the adversarial noise comprises:
obtaining a first predicted label for the data sample through the second model;
generating a noisy sample by reflecting a value of a noise parameter in the data sample;
obtaining a second predicted label for the noisy sample through the second model;
updating the value of the noise parameter in a direction to increase a difference between the first predicted label and the second predicted label; and
calculating adversarial noise for the data sample based on the updated value of the noise parameter.
16. The method of claim 1, wherein the evaluating of the performance of the first model comprises:
predicting a label of the evaluation dataset through the first model; and
evaluating the performance of the first model by comparing the pseudo label and the predicted label.
8. The method of claim 1, wherein the evaluating of the performance of the first model comprises:
predicting labels of the evaluation dataset through the first model; and
evaluating the performance of the first model by comparing the pseudo labels and the predicted labels.
17. A performance evaluation system comprising:
one or more processors; and
a memory configured to store a computer program which is to be executed by the one or more processors,
wherein the computer program comprises instructions for performing:
an operation of obtaining a first model trained using a labeled dataset of a source domain;
an operation of obtaining a second model built by performing domain adaptation to a target domain on the first model;
an operation of generating a pseudo label for an evaluation dataset of the target domain using the second model; and
an operation of evaluating a performance of the first model using the pseudo label,
wherein the evaluation dataset is an unlabeled dataset, and the operation of generating the pseudo label comprises:
an operation of adjusting an upper limit of a size constraint of adversarial noise based on a predefined factor;
an operation of deriving the adversarial noise for a data sample belonging to the evaluation dataset within a range that satisfies the size constraint according to the adjusted upper limit;
an operation of generating a noisy sample by reflecting the derived adversarial noise in the data sample; and
an operation of generating a pseudo label for the data sample based on a predicted label of the noisy sample obtained through the second model.
12. A system for evaluating performance, the system comprising:
a memory configured to store one or more instructions; and
one or more processors configured to execute the one or more stored instructions to perform:
obtaining a first model trained using a labeled dataset;
obtaining a second model built by performing unsupervised domain adaptation on the first model;
generating pseudo labels for an evaluation dataset using the second model, wherein the evaluation dataset is an unlabeled dataset; and
evaluating performance of the first model using the pseudo labels,
wherein the generating of the pseudo labels comprises:
deriving adversarial noise for a data sample belonging to the evaluation dataset;
generating a noisy sample by reflecting the derived adversarial noise in the data sample; and
generating a pseudo label for the data sample based on a predicted label of the noisy sample obtained through the second model.
18. A non-transitory computer-readable recording medium configured to store a computer program to be executed by one or more processors to perform:
obtaining a first model trained using a labeled dataset of a source domain;
obtaining a second model built by performing domain adaptation to a target domain on the first model;
generating a pseudo label for an evaluation dataset of the target domain using the second model;
evaluating a performance of the first model using the pseudo label,
wherein the evaluation dataset is an unlabeled dataset, and the generating of the pseudo label comprises:
adjusting an upper limit of a size constraint of adversarial noise based on a predefined factor;
deriving the adversarial noise for a data sample belonging to the evaluation dataset within a range that satisfies the size constraint according to the adjusted upper limit;
generating a noisy sample by reflecting the derived adversarial noise in the data sample; and
generating a pseudo label for the data sample based on a predicted label of the noisy sample obtained through the second model.
14. A non-transitory computer-readable recording medium storing computer program executable by at least one processor to perform:
obtaining a first model trained using a labeled dataset;
obtaining a second model built by performing unsupervised domain adaptation on the first model;
generating pseudo labels for an evaluation dataset using the second model, wherein the evaluation dataset is an unlabeled dataset; and
evaluating performance of the first model using the pseudo labels,
wherein the generating of the pseudo labels comprises:
deriving adversarial noise for a data sample belonging to the evaluation dataset;
generating a noisy sample by reflecting the derived adversarial noise in the data sample; and
generating a pseudo label for the data sample based on a predicted label of the noisy sample obtained through the second model.
Claims 3, 15 and 16 are identical to claims 2, 4 and 8 of the reference application, as shown in the comparison chart above. The differences are only in wording ("pseudo label" for "pseudo labels", "a label" for "labels", "specific noisy sample" for "noisy sample", and "domain adaptation" for "unsupervised domain adaptation"), and none of them changes what is claimed. The reference application's "unsupervised domain adaptation" is one kind of domain adaptation and therefore falls within the broader "domain adaptation" of the instant claims.
Claim 1 is essentially identical to claim 1 of the reference application. Both claims obtain a first model trained with a labeled dataset, build a second model by domain adaptation of the first model, generate a pseudo label for an unlabeled evaluation dataset by adding adversarial noise to a data sample and taking the second model's predicted label for the noisy sample, and evaluate the first model with that pseudo label. Claim 1 of the reference application does not use the words "source domain" and "target domain", but those words only name where the two datasets come from. Domain adaptation means adapting a model from the domain it was trained on to a different domain. In claim 1 of the reference application, the first model is trained on the labeled dataset (the source domain) and is adapted by unsupervised domain adaptation so that the second model can label the unlabeled evaluation dataset (the target domain). The instant application describes it the same way (instant specification, Para. [0064], "the evaluation system 10 may build a temporary model 22 by performing unsupervised domain adaptation to the target domain on the source model 11"). Naming the domains adds no step to the method. The only limitations of claim 1 that are not in claim 1 of the reference application are "adjusting an upper limit of a size constraint of adversarial noise based on a predefined factor" and deriving the adversarial noise "within a range that satisfies the size constraint according to the adjusted upper limit". In plain terms, claim 1 caps the size of the adversarial noise and adjusts the cap based on a predefined factor. Balaji teaches this, as set out below.
Balaji teaches:
adjusting an upper limit of a size constraint of adversarial noise based on a predefined factor (Section 3, p. 3, "Like vanilla adversarial training, we solve this by sampling mini-batches of images {xi}, crafting adversarial perturbations {δi} of size at most {ϵi}, and then updating the network model using the perturbed images." -- EN: the perturbation δi denotes the adversarial noise and ϵi, its size at most, denotes the upper limit of the size constraint; Section 3, p. 3, "The proposed algorithm is distinctive in that it uses a different ϵi for each image xi. ... After crafting a perturbation for xi, we check if the perturbed image was a successful adversarial example. If PGD succeeded in finding an image with a different class label, then ϵi is too big, so we replace ϵi ← ϵi − γ. If PGD failed, then we set ϵi ← ϵi + γ." -- EN: raising or lowering ϵi by γ denotes the adjusting of the upper limit; Algorithm 2, p. 4, "Require: β: Smoothing constant, γ: Discretization for ϵ search. ... 11: ϵi ← (1 − β)ϵmem[j − 1, i] + βϵi" -- EN: the smoothing constant β and the discretization γ, both fixed before the search, denote the predefined factor, the adjusted upper limit being computed from them);
-- EN: the instant specification's predefined factor denotes a scale factor for adjusting the size of the upper limit (instant specification, Para. [0093], "the predefined factor is a scale factor for adjusting the size of the upper limit and may include, for example, a factor related to the characteristics of an evaluation dataset (or a target domain), a factor related to the characteristics of an individual data sample, etc."). Balaji's smoothing constant β, which scales the stored radius and the newly selected radius in the update of Algorithm 2, line 11, and its discretization γ, by which the radius is raised or lowered, correspond to that factor.
deriving the adversarial noise for a data sample belonging to the evaluation dataset within a range that satisfies the size constraint according to the adjusted upper limit (Algorithm 1, p. 3, "Require: PGDk(x, y, ϵ): Function to generate PGD-k adversarial samples with ϵ norm-bound ... 6: Choose ϵi using Alg 2 ... 8: xiadv = PGD(xi, yi, ϵi)" -- EN: the perturbation of xi is generated with the norm-bound ϵi chosen by Algorithm 2, which corresponds to deriving the noise within the range that satisfies the constraint according to the adjusted upper limit, the images xi being in the combination the data samples of the evaluation dataset recited in the reference application's claims; Section 2, p. 3, "All these methods use the gradient of the loss function with respect to inputs to construct additive perturbations with a norm-constraint." -- EN: PGD derives the noise from the loss gradient under the norm-constraint; Algorithm 2, p. 4, "4: if fθ(PGDk(xi, yi, ϵ1)) predicts as yi then 5: Set ϵi = ϵ1" -- EN: the label yi against which the perturbation is crafted and checked is, in the combination, the label that the second model of the reference application's claims predicts for the data sample).
It would have been obvious to one of ordinary skill in the art, before the effective filing date, to have modified the elements in claims 1, 12 and 14 of the reference application to include adjusting an upper limit of a size constraint of the adversarial noise based on a predefined factor and deriving the adversarial noise within a range that satisfies the size constraint according to the adjusted upper limit, as taught by Balaji.
The motivation for doing so would be to avoid using one perturbation size that is too large for inputs near the decision boundary and too small for the others, since each input gets its own bound that shrinks as soon as the perturbed input changes class, as taught by Balaji where the benefit is disclosed (Section 1, p. 1 and Fig. 1, p. 2).
Similarly, claim 17 is rejected since it is a system claim that recites the same limitations as claim 12 of the reference application. Both claims recite a memory storing a program of instructions and one or more processors that execute those instructions to perform the steps of claim 1.
Similarly, claim 18 is rejected since it is a non-transitory computer-readable recording medium claim that recites the same limitations as claim 14 of the reference application.
Claims 17 and 18 differ from claims 12 and 14 of the reference application only in the same way that claim 1 differs from claim 1 of the reference application, and the same analysis and motivation apply.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Examiner’s Note: Some rejections will include an Examiner’s Note (labeled ‘EN’) to provide additional context or rationale explaining the basis for the rejection.
Claims 1-18 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Claim 1:
Step 1: The claim recites a method; therefore, it is directed to the statutory category of processes.
Step 2A prong 1: The claim recites the following abstract ideas:
generating a pseudo label for an evaluation dataset of the target domain (...); and (A person mentally or with a pen and paper can assign a label, i.e., a pseudo label, to each data sample of a dataset.)
evaluating a performance of the first model using the pseudo label, (A person mentally or with a pen and paper can compare predicted labels with pseudo labels to evaluate the performance of a model.)
adjusting an upper limit of a size constraint of adversarial noise based on a predefined factor; (A person mentally or with a pen and paper can adjust an upper limit of a size constraint based on a predefined factor.)
generating a pseudo label for the data sample based on a predicted label of the noisy sample (...). (A person mentally or with a pen and paper can designate the predicted label of the noisy sample as the pseudo label of the data sample.)
Step 2A prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
A performance evaluation method performed by at least one computing device, the method comprising: (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. -- EN: Generic, off the shelf computing device used as a tool to perform the abstract ideas.)
obtaining a first model trained using a labeled dataset of a source domain; (Insignificant extra-solution activity as the limitation amounts to receiving/obtaining data (MPEP 2106.05(g)). -- EN: Merely obtaining a previously trained model, i.e., mere data gathering.)
obtaining a second model built by performing domain adaptation to a target domain on the first model; (Insignificant extra-solution activity as the limitation amounts to receiving/obtaining data (MPEP 2106.05(g)). -- EN: Obtaining a model built by generic domain adaptation, i.e., generic training, is mere data gathering.)
...using the second model; and (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. -- EN: Generic model used as a tool to obtain a predicted label, no additional details.)
wherein the evaluation dataset is an unlabeled dataset, and the generating of the pseudo label comprises: (Adding insignificant extra-solution activity to the judicial exception (MPEP 2106.05(g)). -- EN: Merely specifies the type of data (unlabeled) on which the abstract idea is performed.)
deriving the adversarial noise for a data sample belonging to the evaluation dataset within a range that satisfies the size constraint according to the adjusted upper limit; (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. -- EN: Deriving adversarial noise is a generic, well-known machine learning technique recited without details.)
generating a noisy sample by reflecting the derived adversarial noise in the data sample; and (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. -- EN: Applying generic noise to a data sample using a generic computer as a tool.)
...obtained through the second model. (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. -- EN: Generic model used as a tool to obtain the predicted label of the noisy sample.)
Step 2B:
A performance evaluation method performed by at least one computing device, the method comprising: (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. -- EN: Generic, off the shelf computing device used as a tool to perform the abstract ideas.)
obtaining a first model trained using a labeled dataset of a source domain; (MPEP 2106.05(d)(II) indicates that merely gathering data is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
obtaining a second model built by performing domain adaptation to a target domain on the first model; (MPEP 2106.05(d)(II) indicates that merely gathering data is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
...using the second model; and (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. -- EN: Generic model used as a tool to obtain a predicted label, no additional details.)
wherein the evaluation dataset is an unlabeled dataset, and the generating of the pseudo label comprises: (The limitation merely selects the type of data to be manipulated, which is insignificant extra-solution activity (MPEP 2106.05(g)). Further, MPEP 2106.05(d)(II) indicates that merely gathering data is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
deriving the adversarial noise for a data sample belonging to the evaluation dataset within a range that satisfies the size constraint according to the adjusted upper limit; (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. -- EN: Deriving adversarial noise is a generic, well-known machine learning technique recited without details.)
generating a noisy sample by reflecting the derived adversarial noise in the data sample; and (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. -- EN: Applying generic noise to a data sample using a generic computer as a tool.)
...obtained through the second model. (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. -- EN: Generic model used as a tool to obtain the predicted label of the noisy sample.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 2:
Step 1: A process, as above.
Step 2A prong 1: See the rejection of Claim 1 above, which claim 2 depends on.
Step 2A prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
wherein the domain adaptation is performed using an unlabeled dataset of the target domain. (Adding insignificant extra-solution activity to the judicial exception (MPEP 2106.05(g)). -- EN: Merely specifies the type of data used by the generic domain adaptation.)
Step 2B:
wherein the domain adaptation is performed using an unlabeled dataset of the target domain. (MPEP 2106.05(d)(II) indicates that merely gathering data is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 3:
Step 1: A process, as above.
Step 2A prong 1: See the rejection of Claim 1 above, which claim 3 depends on.
Step 2A prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
wherein the domain adaptation and the generating of the pseudo label are performed without using the labeled dataset. (The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)). -- EN: Merely limits the abstract idea to an environment where the labeled dataset is not used.)
Step 2B:
wherein the domain adaptation and the generating of the pseudo label are performed without using the labeled dataset. (The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)). -- EN: Merely limits the abstract idea to an environment where the labeled dataset is not used.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 4:
Step 1: A process, as above.
Step 2A prong 1: See the rejection of Claim 1 above, which claim 4 depends on. Claim 4 further recites:
wherein the obtaining of the second model comprises: monitoring a training loss calculated during a domain adaptation process; (A person mentally or with a pen and paper can monitor a series of calculated loss values.)
determining a time when an amount of change in the training loss is equal to or less than a reference value as an early stop time; and (This is interpreted as a mental process and recitation of mathematical concepts. A person with a pen and paper can calculate a change in the loss, compare it to a reference value and note the time.)
Step 2A prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
obtaining the second model by stopping the domain adaptation at the determined early stop time. (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. -- EN: Merely stopping generic training at the time determined by the abstract idea.)
Step 2B:
obtaining the second model by stopping the domain adaptation at the determined early stop time. (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. -- EN: Merely stopping generic training at the time determined by the abstract idea.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 5:
Step 1: A process, as above.
Step 2A prong 1: See the rejection of Claim 1 above, which claim 5 depends on. Claim 5 further recites:
wherein the adjusting of the upper limit of the size constraint comprises: adjusting a first upper limit applied to a size constraint of a first data sample belonging to the evaluation dataset; and (A person mentally or with a pen and paper can adjust an upper limit for a first data sample.)
adjusting a second upper limit applied to a size constraint of a second data sample belonging to the evaluation dataset, (A person mentally or with a pen and paper can adjust an upper limit for a second data sample.)
wherein the adjusted first upper limit is different from the adjusted second upper limit. (A person mentally or with a pen and paper can adjust the two upper limits to different values.)
Step 2A prong 2: The claim does not recite additional elements therefore the judicial exception is not integrated into a practical application.
Step 2B: The claim does not recite additional elements that amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 6:
Step 1: A process, as above.
Step 2A prong 1: See the rejection of Claim 1 above, which claim 6 depends on. Claim 6 further recites:
wherein the predefined factor comprises a first factor related to characteristics of the evaluation dataset and a second factor related to characteristics of the data sample. (A person mentally or with a pen and paper can define a factor comprising a dataset-related factor and a sample-related factor.)
Step 2A prong 2: The claim does not recite additional elements therefore the judicial exception is not integrated into a practical application.
Step 2B: The claim does not recite additional elements that amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 7:
Step 1: A process, as above.
Step 2A prong 1: See the rejection of Claim 1 above, which claim 7 depends on. Claim 7 further recites:
wherein the adjusting of the upper limit of the size constraint comprises: measuring predictive uncertainty of the second model for the data sample; and (This is interpreted as a mental process and recitation of mathematical concepts. A person with a pen and paper can calculate an uncertainty value, e.g., a standard deviation, from a set of predicted confidence scores.)
adjusting the upper limit based on the predictive uncertainty. (A person mentally or with a pen and paper can adjust an upper limit based on an uncertainty value.)
Step 2A prong 2: The claim does not recite additional elements therefore the judicial exception is not integrated into a practical application.
EN: "The second model" merely refers back to the generic model of claim 1.
Step 2B: The claim does not recite additional elements that amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 8:
Step 1: A process, as above.
Step 2A prong 1: See the rejection of Claim 7 above, which claim 8 depends on. Claim 8 further recites:
wherein the measuring of the predictive uncertainty comprises: obtaining a plurality of predicted labels by (...) repeating prediction for the data sample; (A person mentally or with a pen and paper can repeat a prediction for a data sample and record the plurality of predicted labels.)
determining a class with a highest average confidence score among the plurality of predicted labels; and (This limitation falls within the mathematical concepts grouping because it involves averaging confidence scores and selecting the maximum.)
measuring the predictive uncertainty for the data sample based on confidence scores of the determined class included in the plurality of predicted labels, (This limitation falls within the mathematical concepts grouping because it involves calculating a statistical value, e.g., a standard deviation, of the confidence scores of the determined class.)
wherein values of the plurality of predicted labels are confidence scores for each class. (This limitation falls within the mathematical concepts grouping because it merely specifies that the predicted labels are numerical values (probabilities) used in the above calculations.)
Step 2A prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
...applying a drop-out technique to at least a portion of the second model and... (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. -- EN: Generic, well-known neural network technique applied to a generic model without any details.)
Step 2B:
...applying a drop-out technique to at least a portion of the second model and... (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. -- EN: Generic, well-known neural network technique applied to a generic model without any details.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 9:
Step 1: A process, as above.
Step 2A prong 1: See the rejection of Claim 1 above, which claim 9 depends on. Claim 9 further recites:
wherein the adjusting of the upper limit of the size constraint comprises: obtaining a first predicted label for the data sample (...); (A person mentally or with a pen and paper can obtain a predicted label for a data sample.)
obtaining a second predicted label for the data sample (...); and (A person mentally or with a pen and paper can obtain a second predicted label for the data sample.)
adjusting the upper limit based on a difference between the first predicted label and the second predicted label. (A person mentally or with a pen and paper can calculate a difference between two predicted labels and adjust an upper limit based on the difference.)
Step 2A prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
...through the first model; (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. -- EN: Generic model used as a tool to obtain a predicted label, no additional details.)
...through the second model; and (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. -- EN: Generic model used as a tool to obtain a predicted label, no additional details.)
Step 2B:
...through the first model; (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. -- EN: Generic model used as a tool to obtain a predicted label, no additional details.)
...through the second model; and (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. -- EN: Generic model used as a tool to obtain a predicted label, no additional details.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 10:
Step 1: A process, as above.
Step 2A prong 1: See the rejection of Claim 9 above, which claim 10 depends on. Claim 10 further recites:
wherein the difference between the first predicted label and the second predicted label is calculated based on Jensen-Shannon divergence (JSD). (This limitation falls within the mathematical concepts grouping because it recites a specific mathematical formula (Jensen-Shannon divergence) for calculating the difference between two probability distributions.)
Step 2A prong 2: The claim does not recite additional elements therefore the judicial exception is not integrated into a practical application.
Step 2B: The claim does not recite additional elements that amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 11:
Step 1: A process, as above.
Step 2A prong 1: See the rejection of Claim 1 above, which claim 11 depends on. Claim 11 further recites:
wherein the adjusting of the upper limit of the size constraint comprises adjusting the upper limit based on a degree of dispersion of data samples of the evaluation dataset. (This is interpreted as a mental process and recitation of mathematical concepts. A person with a pen and paper can calculate a degree of dispersion, e.g., a standard deviation, of data samples and adjust an upper limit based on it.)
Step 2A prong 2: The claim does not recite additional elements therefore the judicial exception is not integrated into a practical application.
Step 2B: The claim does not recite additional elements that amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 12:
Step 1: A process, as above.
Step 2A prong 1: See the rejection of Claim 11 above, which claim 12 depends on. Claim 12 further recites:
calculating a representative vale of each of the data samples based on a pixel value of each of the data samples; and (This limitation falls within the mathematical concepts grouping because it involves calculating a representative value, e.g., an average, of the pixel values of each image.)
measuring the degree of dispersion of the data samples based on a degree of dispersion of calculated representative values. (This limitation falls within the mathematical concepts grouping because it involves calculating a degree of dispersion, e.g., a standard deviation, of the calculated representative values.)
Step 2A prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
wherein the data samples are images, and the adjusting of the upper limit based on the degree of dispersion of the data samples of the evaluation dataset comprises: (Adding insignificant extra-solution activity to the judicial exception (MPEP 2106.05(g)). -- EN: Merely specifies the type of data (images) on which the calculations are performed.)
Step 2B:
wherein the data samples are images, and the adjusting of the upper limit based on the degree of dispersion of the data samples of the evaluation dataset comprises: (The limitation merely selects the type of data to be manipulated, which is insignificant extra-solution activity (MPEP 2106.05(g)). Further, MPEP 2106.05(d)(II) indicates that merely gathering data is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 13:
Step 1: A process, as above.
Step 2A prong 1: See the rejection of Claim 12 above, which claim 13 depends on. Claim 13 further recites:
wherein the degree of dispersion of the representative values is a first degree of dispersion, and the measuring of the degree of dispersion of the data samples based on the degree of dispersion of the calculated representative values comprises: obtaining a second degree of dispersion of data samples of the labeled dataset (...); and (This is interpreted as a mental process and recitation of mathematical concepts. A person with a pen and paper can obtain a degree of dispersion, e.g., a standard deviation, of a set of data samples.)
calculating a degree of dispersion of the data samples based on a ratio of the first degree of dispersion to the second degree of dispersion. (This limitation falls within the mathematical concepts grouping because it involves calculating a ratio of two dispersion values.)
Step 2A prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
...from training history data of the labeled dataset of the source domain without accessing the labeled dataset... (Insignificant extra-solution activity as the limitation amounts to obtaining data from a particular data source (MPEP 2106.05(g)). -- EN: Merely specifies the data source (training history data) from which the value is obtained.)
Step 2B:
...from training history data of the labeled dataset of the source domain without accessing the labeled dataset... (MPEP 2106.05(d)(II) indicates that merely gathering data is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 14:
Step 1: A process, as above.
Step 2A prong 1: See the rejection of Claim 1 above, which claim 14 depends on. Claim 14 further recites:
wherein the adjusting of the upper limit of the size constraint comprises adjusting the upper limit based on a number of classes of the evaluation dataset. (This is interpreted as a mental process and recitation of mathematical concepts. A person with a pen and paper can count the number of classes and adjust an upper limit based on the count.)
Step 2A prong 2: The claim does not recite additional elements therefore the judicial exception is not integrated into a practical application.
Step 2B: The claim does not recite additional elements that amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 15:
Step 1: A process, as above.
Step 2A prong 1: See the rejection of Claim 1 above, which claim 15 depends on. Claim 15 further recites:
wherein the deriving of the adversarial noise comprises: obtaining a first predicted label for the data sample (...); (A person mentally or with a pen and paper can obtain a predicted label for a data sample.)
obtaining a second predicted label for the specific noisy sample (...); (A person mentally or with a pen and paper can obtain a predicted label for the noisy sample.)
updating the value of the noise parameter in a direction to increase a difference between the first predicted label and the second predicted label; and (This is interpreted as a mental process and recitation of mathematical concepts. A person with a pen and paper can calculate a difference between two predicted labels and update a value in the direction that increases the difference.)
calculating the adversarial noise for the data sample based on the updated value of the noise parameter. (This limitation falls within the mathematical concepts grouping because it involves calculating a noise value from the updated parameter value.)
Step 2A prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
...through the second model; (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. -- EN: Generic model used as a tool to obtain a predicted label, no additional details.)
generating a specific noisy sample by reflecting a value of a noise parameter in the data sample; (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. -- EN: Applying generic noise to a data sample using a generic computer as a tool.)
...through the second model; (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. -- EN: Generic model used as a tool to obtain a predicted label, no additional details.)
Step 2B:
...through the second model; (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. -- EN: Generic model used as a tool to obtain a predicted label, no additional details.)
generating a specific noisy sample by reflecting a value of a noise parameter in the data sample; (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. -- EN: Applying generic noise to a data sample using a generic computer as a tool.)
...through the second model; (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. -- EN: Generic model used as a tool to obtain a predicted label, no additional details.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 16:
Step 1: A process, as above.
Step 2A prong 1: See the rejection of Claim 1 above, which claim 16 depends on. Claim 16 further recites:
wherein the evaluating of the performance of the first model comprises: predicting a label of the evaluation dataset (...); and (A person mentally or with a pen and paper can predict a label for each data sample of a dataset.)
evaluating the performance of the first model by comparing the pseudo label and the predicted label. (A person mentally or with a pen and paper can compare the pseudo label with the predicted label to evaluate the performance of a model.)
Step 2A prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
...through the first model; and (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. -- EN: Generic model used as a tool to obtain a predicted label, no additional details.)
Step 2B:
...through the first model; and (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. -- EN: Generic model used as a tool to obtain a predicted label, no additional details.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 17:
Step 1: The claim recites a system; therefore, it is directed to the statutory category of machine.
Step 2A prong 1: The claim recites the following abstract ideas:
an operation of generating a pseudo label for an evaluation dataset of the target domain (...); and (A person mentally or with a pen and paper can assign a label, i.e., a pseudo label, to each data sample of a dataset.)
an operation of evaluating a performance of the first model using the pseudo label, (A person mentally or with a pen and paper can compare predicted labels with pseudo labels to evaluate the performance of a model.)
an operation of adjusting an upper limit of a size constraint of adversarial noise based on a predefined factor; (A person mentally or with a pen and paper can adjust an upper limit of a size constraint based on a predefined factor.)
an operation of generating a pseudo label for the data sample based on a predicted label of the noisy sample (...). (A person mentally or with a pen and paper can designate the predicted label of the noisy sample as the pseudo label of the data sample.)
Step 2A prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
A performance evaluation system comprising: one or more processors; and a memory configured to store a computer program which is to be executed by the one or more processors, wherein the computer program comprises instructions for performing: (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. -- EN: Generic, off the shelf processors and memory used as tools to perform the abstract ideas.)
an operation of obtaining a first model trained using a labeled dataset of a source domain; (Insignificant extra-solution activity as the limitation amounts to receiving/obtaining data (MPEP 2106.05(g)). -- EN: Merely obtaining a previously trained model, i.e., mere data gathering.)
an operation of obtaining a second model built by performing domain adaptation to a target domain on the first model; (Insignificant extra-solution activity as the limitation amounts to receiving/obtaining data (MPEP 2106.05(g)). -- EN: Obtaining a model built by generic domain adaptation, i.e., generic training, is mere data gathering.)
...using the second model; and (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. -- EN: Generic model used as a tool to obtain a predicted label, no additional details.)
wherein the evaluation dataset is an unlabeled dataset, and the operation of generating the pseudo label comprises: (Adding insignificant extra-solution activity to the judicial exception (MPEP 2106.05(g)). -- EN: Merely specifies the type of data (unlabeled) on which the abstract idea is performed.)
an operation of deriving the adversarial noise for a data sample belonging to the evaluation dataset within a range that satisfies the size constraint according to the adjusted upper limit; (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. -- EN: Deriving adversarial noise is a generic, well-known machine learning technique recited without details.)
an operation of generating a noisy sample by reflecting the derived adversarial noise in the data sample; and (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. -- EN: Applying generic noise to a data sample using a generic computer as a tool.)
...obtained through the second model. (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. -- EN: Generic model used as a tool to obtain the predicted label of the noisy sample.)
Step 2B:
A performance evaluation system comprising: one or more processors; and a memory configured to store a computer program which is to be executed by the one or more processors, wherein the computer program comprises instructions for performing: (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. -- EN: Generic, off the shelf processors and memory used as tools to perform the abstract ideas.)
an operation of obtaining a first model trained using a labeled dataset of a source domain; (MPEP 2106.05(d)(II) indicates that merely gathering data is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
an operation of obtaining a second model built by performing domain adaptation to a target domain on the first model; (MPEP 2106.05(d)(II) indicates that merely gathering data is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
...using the second model; and (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. -- EN: Generic model used as a tool to obtain a predicted label, no additional details.)
wherein the evaluation dataset is an unlabeled dataset, and the operation of generating the pseudo label comprises: (The limitation merely selects the type of data to be manipulated, which is insignificant extra-solution activity (MPEP 2106.05(g)). Further, MPEP 2106.05(d)(II) indicates that merely gathering data is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
an operation of deriving the adversarial noise for a data sample belonging to the evaluation dataset within a range that satisfies the size constraint according to the adjusted upper limit; (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. -- EN: Deriving adversarial noise is a generic, well-known machine learning technique recited without details.)
an operation of generating a noisy sample by reflecting the derived adversarial noise in the data sample; and (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. -- EN: Applying generic noise to a data sample using a generic computer as a tool.)
...obtained through the second model. (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. -- EN: Generic model used as a tool to obtain the predicted label of the noisy sample.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 18:
Step 1: The claim recites a non-transitory computer-readable recording medium; therefore, it is directed to the statutory category of manufacture.
Step 2A prong 1: The claim recites the following abstract ideas:
generating a pseudo label for an evaluation dataset of the target domain (...); and (A person mentally or with a pen and paper can assign a label, i.e., a pseudo label, to each data sample of a dataset.)
evaluating a performance of the first model using the pseudo label, (A person mentally or with a pen and paper can compare predicted labels with pseudo labels to evaluate the performance of a model.)
adjusting an upper limit of a size constraint of adversarial noise based on a predefined factor; (A person mentally or with a pen and paper can adjust an upper limit of a size constraint based on a predefined factor.)
generating a pseudo label for the data sample based on a predicted label of the noisy sample (...). (A person mentally or with a pen and paper can designate the predicted label of the noisy sample as the pseudo label of the data sample.)
Step 2A prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
A non-transitory computer-readable recording medium configured to store a computer program to be executed by one or more processors to perform: (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. -- EN: Generic recording medium and processors used as tools to perform the abstract ideas.)
obtaining a first model trained using a labeled dataset of a source domain; (Insignificant extra-solution activity as the limitation amounts to receiving/obtaining data (MPEP 2106.05(g)). -- EN: Merely obtaining a previously trained model, i.e., mere data gathering.)
obtaining a second model built by performing domain adaptation to a target domain on the first model; (Insignificant extra-solution activity as the limitation amounts to receiving/obtaining data (MPEP 2106.05(g)). -- EN: Obtaining a model built by generic domain adaptation, i.e., generic training, is mere data gathering.)
...using the second model; and (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. -- EN: Generic model used as a tool to obtain a predicted label, no additional details.)
wherein the evaluation dataset is an unlabeled dataset, and the generating of the pseudo label comprises: (Adding insignificant extra-solution activity to the judicial exception (MPEP 2106.05(g)). -- EN: Merely specifies the type of data (unlabeled) on which the abstract idea is performed.)
deriving the adversarial noise for a data sample belonging to the evaluation dataset within a range that satisfies the size constraint according to the adjusted upper limit; (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. -- EN: Deriving adversarial noise is a generic, well-known machine learning technique recited without details.)
generating a noisy sample by reflecting the derived adversarial noise in the data sample; and (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. -- EN: Applying generic noise to a data sample using a generic computer as a tool.)
...obtained through the second model. (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. -- EN: Generic model used as a tool to obtain the predicted label of the noisy sample.)
Step 2B:
A non-transitory computer-readable recording medium configured to store a computer program to be executed by one or more processors to perform: (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. -- EN: Generic recording medium and processors used as tools to perform the abstract ideas.)
obtaining a first model trained using a labeled dataset of a source domain; (MPEP 2106.05(d)(II) indicates that merely gathering data is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
obtaining a second model built by performing domain adaptation to a target domain on the first model; (MPEP 2106.05(d)(II) indicates that merely gathering data is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
...using the second model; and (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. -- EN: Generic model used as a tool to obtain a predicted label, no additional details.)
wherein the evaluation dataset is an unlabeled dataset, and the generating of the pseudo label comprises: (The limitation merely selects the type of data to be manipulated, which is insignificant extra-solution activity (MPEP 2106.05(g)). Further, MPEP 2106.05(d)(II) indicates that merely gathering data is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
deriving the adversarial noise for a data sample belonging to the evaluation dataset within a range that satisfies the size constraint according to the adjusted upper limit; (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. -- EN: Deriving adversarial noise is a generic, well-known machine learning technique recited without details.)
generating a noisy sample by reflecting the derived adversarial noise in the data sample; and (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. -- EN: Applying generic noise to a data sample using a generic computer as a tool.)
...obtained through the second model. (This additional element recites a mere instruction to apply an exception with a recitation of the words "apply it" (or an equivalent) as identified in MPEP 2106.05(f), and does not provide integration into a practical application. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. -- EN: Generic model used as a tool to obtain the predicted label of the noisy sample.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Examiner’s Note: Some rejections will include an Examiner’s Note (labeled ‘EN’) to provide additional context or rationale explaining the basis for the rejection.
Claim(s) 1, 2, 3, 5, 7, 8, 9 and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al., "Detecting Errors and Estimating Accuracy on Unlabeled Data with Self-training Ensembles," (hereinafter Chen) in view of Liang et al., "Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain Adaptation," (hereinafter Liang) in view of Balaji et al., "Instance Adaptive Adversarial Training: Improved Accuracy Tradeoffs in Neural Nets," (hereinafter Balaji) in view of Alarab et al., "Uncertainty estimation based adversarial attack in multi-class classification," (hereinafter Alarab).
Regarding claim 1, Chen teaches:
A performance evaluation method performed by at least one computing device, the method comprising (Abstract, p. 1, "For safe deployment, it is essential to estimate the accuracy of the pre-trained model on the test data." C.1.1, p. 21, "We run all experiments with PyTorch and NVIDIA GeForce RTX 2080Ti GPUs."):
obtaining a first model trained using a labeled dataset of a source domain (Section 3, p. 3, "Let D be a set of labeled data from the training distribution PX,Y." -- EN: D denotes the labeled dataset; Section 3, p. 3, "Given a model f : X → Y trained on D, together with D and UX, the goal of unsupervised accuracy estimation is to get an estimate" -- EN: the pre-trained model f denotes the first model; Section 1, p. 2, "Data distribution in the real world may be wildly different from the training dataset for various reasons such as covariate shift due to domain divergence" -- EN: the training distribution from which D is drawn denotes the source domain and the test distribution denotes the target domain; Framework 1, p. 4, "Input: A training dataset D, an unlabeled test dataset UX, a pre-trained model f" -- EN: receiving the pre-trained model f as an input of the framework corresponds to the obtaining);
obtaining a second model built by performing domain adaptation to a target domain … (Algorithm 3, p. 7, "TRM: Ensemble via Representation Matching Input: D, UX, R, N, initial pre-trained model h0, parameters α, γ ... 1: Fine-tune h0 for N epochs using the objective: ... + α · d(pDφ, pUXφ) ... where d(pDφ, pUXφ) is the distance between the distribution of φ(x) on D and that on UX." -- EN: the check model h fine-tuned with the representation matching loss denotes the second model; Section 6, p. 7, "Algorithm 3 describes another method designed with the representation matching technique for domain adaptation, which can potentially improve the accuracy of the ensemble on some test inputs related to the training data" -- EN: the representation matching of the check model to the unlabeled test set UX denotes the domain adaptation, and the test distribution from which UX is drawn denotes the target domain; Section 6, p. 7, "For our experiments, we use the representation matching loss from the classic DANN [12]." -- EN: DANN is a domain adaptation method; Appendix A.1, p. 14, "Both proxy risk and one instance of our framework (Algorithm 3) use domain-invariant representations (DIR) to improve the accuracy of the check models on the target domain." -- EN: the check models are adapted to the target domain)
generating a pseudo label for an evaluation dataset of the target domain using the second model (Algorithm 1, p. 6, "4: Let ỹx be the majority vote of {hi}i=1N: ỹx := arg maxj∈Y (1/N) Σi=1N I[hi(x) = j]." -- EN: ỹx denotes the pseudo label and the check models hi whose predictions on x generate it denote the second model; Framework 1, p. 4, "Input: A training dataset D, an unlabeled test dataset UX" -- EN: the unlabeled test dataset UX, drawn from the test distribution, denotes the evaluation dataset of the target domain; Section 4, p. 4, footnote 2, "This includes one single h as a special case." -- EN: the ensemble of check models includes the case of a single check model, which denotes the second model); and
evaluating a performance of the first model using the pseudo label (Algorithm 1, p. 6, "5: Set R = {(x, ỹx) : x ∈ UX, ỹx ≠ f(x)}. ... Output: Estimated mis-classified points RX = {x : (x, y) ∈ R}, estimated accuracy |UX \ RX| / |UX|." -- EN: the estimated accuracy of f, computed from the points at which the pseudo label ỹx differs from f(x), denotes the performance of the first model evaluated using the pseudo label; Section 3, p. 3, "the goal of unsupervised accuracy estimation is to get an estimate" -- EN: acc(f), the accuracy of the pre-trained model f, is the performance estimated),
wherein the evaluation dataset is an unlabeled dataset (Abstract, p. 1, "(1) unsupervised accuracy estimation, which aims to estimate the accuracy of a pre-trained classifier on a set of unlabeled test inputs" -- EN: the unlabeled test inputs denote the unlabeled evaluation dataset; Section 3, p. 3, "Let UX = {x : (x, yx) ∈ U} be the set of input feature vectors in U." -- EN: UX holds the inputs without their labels yx),
Chen does not explicitly teach:
[obtaining a second model built by performing domain adaptation to a target domain] on the first model;
However, Liang teaches:
[obtaining a second model built by performing domain adaptation to a target domain] on the first model; (Abstract, p. 1, "This work tackles a practical setting where only a trained source model is available and investigates how we can effectively utilize such a model without source data to solve UDA problems." -- EN: the trained source model denotes the first model and UDA denotes the domain adaptation; Abstract, p. 1, "SHOT freezes the classifier module (hypothesis) of the source model and learns the target-specific feature extraction module by exploiting both information maximization and self-supervised pseudo-labeling to implicitly align representations from the target domains to the source hypothesis." -- EN: the source model with its classifier frozen and its feature extraction module learned on the target domain denotes the second model built by performing domain adaptation on the first model; Section 3.1, p. 3, "We consider to develop a deep neural network and learn the source model fs : Xs → Ys by minimizing the following standard cross-entropy loss" -- EN: fs is trained on the labeled source samples, Chen teaching the first model trained on the labeled dataset as mapped above; Algorithm 1, p. 9, "Input: source model fs = gs ∘ hs, target data {xti}i=1nt, maximum number of epochs Tm, trade-off parameter β. Initialization: Freeze the final classifier layer ht = hs, and copy the parameters from gs to gt as initialization." -- EN: the target model is initialized from the parameters of the source model itself and keeps the source classifier, which corresponds to the domain adaptation being performed on the first model; Fig. 2, p. 5, "SHOT keeps the hypothesis frozen and utilizes the feature learning module as initialization for target domain learning." -- EN: the target domain learning of the source model's own feature learning module denotes the domain adaptation to the target domain);
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Chen to include a check model that is a copy of the first model f, with the classifier of the copy kept fixed and its feature extractor re-trained on the unlabeled test data, as taught by Liang, in order to build the check model from the first model itself instead of training a separate model (Section 3.2, page 3-4). Chen itself suggests unsupervised domain adaptation of the check model to the test inputs (Section 4, p. 4).
The motivation for doing so would be to improve the accuracy of the check model on the target-domain test data, since a copy of the source model re-trained in this way classifies the target data far better than the source model itself, as taught by Liang where the benefit is disclosed (Section 4.3, p. 6 and Table 2, p. 6).
Chen in view of Liang do not explicitly teach:
and the generating of the pseudo label comprises:
adjusting an upper limit of a size constraint of adversarial noise based on a predefined factor;
deriving the adversarial noise for a data sample belonging to the evaluation dataset within a range that satisfies the size constraint according to the adjusted upper limit;
generating a noisy sample by reflecting the derived adversarial noise in the data sample; and
However, Balaji teaches:
and the generating of the pseudo label comprises:
adjusting an upper limit of a size constraint of adversarial noise based on a predefined factor (Section 3, p. 3, "Like vanilla adversarial training, we solve this by sampling mini-batches of images {xi}, crafting adversarial perturbations {δi} of size at most {ϵi}, and then updating the network model using the perturbed images." -- EN: the perturbation δi denotes the adversarial noise and ϵi, its size at most, denotes the upper limit of the size constraint; Section 3, p. 3, "The proposed algorithm is distinctive in that it uses a different ϵi for each image xi. ... After crafting a perturbation for xi, we check if the perturbed image was a successful adversarial example. If PGD succeeded in finding an image with a different class label, then ϵi is too big, so we replace ϵi ← ϵi − γ. If PGD failed, then we set ϵi ← ϵi + γ." -- EN: raising or lowering ϵi by γ denotes the adjusting of the upper limit; Algorithm 2, p. 4, "Require: β: Smoothing constant, γ: Discretization for ϵ search. ... 11: ϵi ← (1 − β)ϵmem[j − 1, i] + βϵi" -- EN: the smoothing constant β and the discretization γ, both fixed before the search, denote the predefined factor, the adjusted upper limit being computed from them);
-- EN: the instant specification's predefined factor denotes a scale factor for adjusting the size of the upper limit (instant specification, Para. [0093], "the predefined factor is a scale factor for adjusting the size of the upper limit and may include, for example, a factor related to the characteristics of an evaluation dataset (or a target domain), a factor related to the characteristics of an individual data sample, etc."). Balaji's smoothing constant β, which scales the stored radius and the newly selected radius in the update of Algorithm 2, line 11, and its discretization γ, by which the radius is raised or lowered, correspond to that factor.
deriving the adversarial noise for a data sample belonging to the evaluation dataset within a range that satisfies the size constraint according to the adjusted upper limit (Algorithm 1, p. 3, "Require: PGDk(x, y, ϵ): Function to generate PGD-k adversarial samples with ϵ norm-bound ... 6: Choose ϵi using Alg 2 ... 8: xiadv = PGD(xi, yi, ϵi)" -- EN: the perturbation of xi is generated with the norm-bound ϵi chosen by Algorithm 2, which corresponds to deriving the noise within the range that satisfies the constraint according to the adjusted upper limit, the images xi being in the combination the unlabeled test inputs of Chen; Section 2, p. 2, "All these methods use the gradient of the loss function with respect to inputs to construct additive perturbations with a norm-constraint." -- EN: PGD derives the noise from the loss gradient under the norm-constraint);
generating a noisy sample by reflecting the derived adversarial noise in the data sample; and (Algorithm 1, p. 3, "8: xiadv = PGD(xi, yi, ϵi)" -- EN: xiadv, the image xi with its perturbation δi added, denotes the noisy sample; Section 3, p. 3, "crafting adversarial perturbations {δi} of size at most {ϵi}, and then updating the network model using the perturbed images" -- EN: the perturbed images denote the noisy samples; Section 2, p. 2, "All these methods use the gradient of the loss function with respect to inputs to construct additive perturbations with a norm-constraint." -- EN: the perturbation is added to the input, which corresponds to reflecting the noise in the data sample)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Chen in view of Liang to include perturbing each unlabeled test input within a bound of its own, raised by a fixed step γ when the perturbed input keeps its class and lowered by γ when it does not, as taught by Balaji, in order to fit the bound of each input to its distance from the decision boundary (Section 1, p. 2).
The motivation for doing so would be to avoid using one perturbation size that is too large for inputs near the decision boundary and too small for the others, since each input gets its own bound that shrinks as soon as the perturbed input changes class, as taught by Balaji where the benefit is disclosed (Section 1, p. 1 and Fig. 1, p. 2).
Chen in view of Liang in view of Balaji do not explicitly teach:
generating a pseudo label for the data sample based on a predicted label of the noisy sample obtained through the second model.
However, Alarab teaches:
generating a pseudo label for the data sample based on a predicted label of the noisy sample obtained through the second model (Algorithm 1, p. 1523, "for all ϵi ∈ I do xϵi ← x + ϵi·sign(grad); ŷϵi ← M(xϵi); return ŷϵi end for" -- EN: xϵi, the input with a perturbation added, denotes the noisy sample, and ŷϵi, the output of the model M denotes the predicted label of the noisy sample obtained through the model, M being in the combination the check model of Chen as adapted by Liang. And the noisy sample being the input perturbed within its bound as set out above by Balaji; Section 3.1, p. 1523, "Referring to [2], the predictive mean can be computed as follows: ŷ = pMC-AA(y|x) ≈ (1/T) Σi=1T ŷϵi, (2)" -- EN: ŷ, the label assigned to the data sample x from the outputs ŷϵi for its noisy samples, denotes the pseudo label generated based on the predicted label of the noisy sample).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Chen in view of Liang in view of Balaji to include labeling each test input with the check model's prediction for the noisy version of that input, rather than for the original input, as taught by Alarab, in order to base each pseudo label on how the check model classifies the input after it is pushed toward the decision boundary (Section 3.1, p. 1523).
The motivation for doing so would be to reduce the number of test inputs that the check model labels wrongly with high confidence, since a wrong but confident prediction on the original input is exposed when the input is pushed toward the decision boundary, as taught by Alarab where the benefit is disclosed (Section 2, p. 1522 and Section 7, p. 1533).
Regarding claim 2, Chen in view of Liang in view of Balaji in view of Alarab teaches all the limitations of claim 1, Liang further teaches:
wherein the domain adaptation is performed using an unlabeled dataset of the target domain (Section 3, p. 3, "and also nt unlabeled samples {xti}i=1nt from the target domain Dt where xti ∈ Xt." -- EN: the unlabeled target samples denote the unlabeled dataset of the target domain; Section 3.2, p. 3, "Here SHOT aims to learn a target function ft : Xt → Yt and infer {yti}i=1nt, with only {xti}i=1nt and the source function fs : Xs → Ys available." -- EN: the adaptation is carried out with the unlabeled target inputs and the source model alone; Section 3.3, p. 4, "To summarize, given the source model fs = gs ∘ hs and pseudo labels generated as above, SHOT freezes the hypothesis from source ht = hs and learns the feature encoder gt with the full objective as" -- EN: the objective of Eq. (7) that follows is formed over the unlabeled target samples Xt alone, which corresponds to performing the domain adaptation using the unlabeled dataset; Algorithm 1, p. 9, "Sample a batch from target data and get the corresponding pseudo labels. Update the parameters in gt via L(gt) in Eq. (7)." -- EN: the feature encoder is updated from batches of the target data).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Chen to include a check model that is a copy of the first model f, with the classifier of the copy kept fixed and its feature extractor re-trained on the unlabeled test data, as taught by Liang, in order to build the check model from the first model itself instead of training a separate model (Section 3.2, page 3-4). Chen itself suggests unsupervised domain adaptation of the check model to the test inputs (Section 4, p. 4).
The motivation for doing so would be to improve the accuracy of the check model on the target-domain test data, since a copy of the source model re-trained in this way classifies the target data far better than the source model itself, as taught by Liang where the benefit is disclosed (Section 4.3, p. 6 and Table 2, p. 6).
Regarding claim 3, Chen in view of Liang in view of Balaji in view of Alarab teaches all the limitations of claim 1, Liang, Chen, Balaji, further teach:
wherein the domain adaptation (Liang Abstract, p. 1, "This work tackles a practical setting where only a trained source model is available and investigates how we can effectively utilize such a model without source data to solve UDA problems." -- EN: the source data, the labeled dataset on which the source model was trained, does not enter the adaptation; Section 1, p. 2, "It aims to learn a target-specific feature encoding module to generate target data representations that are well aligned with source data representations, without accessing source data and labels for target data." -- EN: neither the source data nor target labels enter the adaptation) and the generating of the pseudo label are performed without using the labeled dataset (Chen Algorithm 1, p. 6, "4: Let ỹx be the majority vote of {hi}i=1N: ỹx := arg maxj∈Y (1/N) Σi=1N I[hi(x) = j]." -- EN: the pseudo label is formed from the check models' predictions on the test input x alone, without the labeled dataset D; Balaji Algorithm 1, p. 3, "8: xiadv = PGD(xi, yi, ϵi)" -- EN: the label yi used to craft the perturbation is, in the combination, the check model's pseudo label for the unlabeled input as set out for claim 1, so no labeled data enters the perturbation)
The motivation to combine is the same as set forth above under claim 1.
Regarding claim 5, Chen in view of Liang in view of Balaji in view of Alarab teaches all the limitations of claim 1, Balaji further teaches:
wherein the adjusting of the upper limit of the size constraint comprises:
adjusting a first upper limit applied to a size constraint of a first data sample belonging to the evaluation dataset (Balaji Section 3, p. 3, "The proposed algorithm is distinctive in that it uses a different ϵi for each image xi." -- EN: the radius ϵi of a first image xi denotes the first upper limit applied to the size constraint of the first data sample, the images being in the combination the unlabeled test inputs of Chen; Algorithm 2, p. 4, "Require: i: Sample index, j: Epoch index ... 12: Update ϵmem[j, i] ← ϵi 13: Return ϵi" -- EN: the radius is adjusted and stored per sample index i); and
adjusting a second upper limit applied to a size constraint of a second data sample belonging to the evaluation dataset (Balaji Section 3, p. 3, "The proposed algorithm is distinctive in that it uses a different ϵi for each image xi." -- EN: the radius ϵi of a second image denotes the second upper limit; Section 1, p. 2, "The above observation naturally motivates adversarial training with instance adaptive perturbation radii that are customized to each training image." -- EN: each image receives its own adjusted radius),
wherein the adjusted first upper limit is different from the adjusted second upper limit (Balaji Fig. 2, p. 4, "ϵ = 0.20 ϵ = 0.83 ϵ = 1.71 ϵ = 1.25 ... ϵ = 28.07 ϵ = 28.13 ϵ = 28.23 ϵ = 28.57" -- EN: the radii displayed below the images differ from image to image; Fig. 2, p. 4, "The left panel shows samples that are assigned small ϵ (displayed below images) during adaptive training. These images are close to class boundaries, and change class when perturbed with ϵ ≥ 8. The right panel show images that are assigned large ϵ." -- EN: a sample assigned a small radius and a sample assigned a large radius have different adjusted upper limits; Section 5.2, p. 8, "Also, each sample has a different ϵ profile - for some, ϵ increases well beyond the commonly use radius of ϵ = 8, while for others, it converges below it." -- EN: the adjusted radii of different samples differ).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Chen in view of Liang to include perturbing each unlabeled test input within a bound of its own, raised by a fixed step γ when the perturbed input keeps its class and lowered by γ when it does not, as taught by Balaji, in order to fit the bound of each input to its distance from the decision boundary (Section 1, p. 2).
The motivation for doing so would be to avoid using one perturbation size that is too large for inputs near the decision boundary and too small for the others, since each input gets its own bound that shrinks as soon as the perturbed input changes class, as taught by Balaji where the benefit is disclosed (Section 1, p. 1 and Fig. 1, p. 2).
Regarding claim 7, Chen in view of Liang in view of Balaji in view of Alarab teaches all the limitations of claim 1, Alarab and Balaji further teach:
wherein the adjusting of the upper limit of the size constraint comprises:
measuring predictive uncertainty of the second model for the data sample (Alarab Section 3.2, p. 1524, "In this method, the dropout is activated during the testing phase which is applied after each weight layer in a neural network." -- EN: the dropout kept active in the layers of the network at test time lets the model be run repeatedly on the data sample itself, the model being in the combination the adapted check model; Section 3.2, p. 1524, "By drawing T samples from Bernoulli distribution, this produces {W₁ᵗ, …, W_Lᵗ}ₜ₌₁ᵀ which so allows to express the approximated predictive mean of a given input as: E_q(y|x)(y) ≈ (1/T) Σₜ₌₁ᵀ ŷ(x, W₁ᵗ, …, W_Lᵗ) = p_MC(y|x) (8)" -- EN: p_MC(y|x) is formed from T runs of the model on the unperturbed data sample x; Section 3.2, p. 1525, "Hence the predictive uncertainty using mutual information can be expresses as: Î(y|x, D_train) = Ĥ(y|x, D_train) + Σ_c (1/T) Σₜ₌₁ᵀ p(y = c|x, w) log p(y = c|x, w) (9)" -- EN: the mutual information Î(y|x, D_train), computed from the T outputs of the model for the data sample, denotes the measured predictive uncertainty, which is obtained from the data sample alone and so before any adversarial noise is derived for it); and
adjusting the upper limit based on the predictive uncertainty (Alarab teaches the predictive uncertainty measured for the data sample as mapped above. Balaji further teaches selecting the bound of each sample from a quantity computed for that sample alone, an ambiguous sample receiving a smaller bound and a sample whose class is unambiguous receiving a larger one (Balaji Algorithm 2, p. 4, "4: if f_θ(PGD_k(x_i, y_i, ϵ₁)) predicts as y_i then 5: Set ϵ_i = ϵ₁ 6: else if f_θ(PGD_k(x_i, y_i, ϵ₂)) predicts as y_i then 7: Set ϵ_i = ϵ₂ 8: else 9: Set ϵ_i = ϵ₃"; Section 1, p. 2, "samples with small radii are often ambiguous and have nearby images of another class, while images with large radii have unambiguous class labels that are difficult to manipulate."). Thus, in the proposed combination, the per-sample quantity from which Balaji selects the bound of a test input is the predictive uncertainty that Alarab measures for that input by running the adapted check model on it several times with dropout active, that uncertainty being measured before any adversarial noise is derived for the input, such that the upper limit is adjusted based on the predictive uncertainty).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Chen in view of Liang in view of Balaji to include an uncertainty measured for each test input by running the check model on that input several times with its dropout switched on, used to set that input's bound ϵ_i in place of Balaji's three-trial test, as taught by Alarab, in order to obtain a per-input measure of the check model's uncertainty without first perturbing the input (Section 3.2, page 1524-1525). Chen's check model for the Digits datasets already carries a dropout layer (Appendix C.1.4, p. 22).
The motivation for doing so would be to fit each input's bound to how sure the check model is about that input, since running the existing model a few times with dropout turned on approximates the Bayesian posterior at a fraction of its cost, as taught by Alarab where the benefit is disclosed (Section 1, p. 1520 and Section 2, p. 1521).
Regarding claim 8, Chen in view of Liang in view of Balaji in view of Alarab teaches all the limitations of claims 1 and 7, Alarab further teaches:
wherein the measuring of the predictive uncertainty comprises:
obtaining a plurality of predicted labels by applying a drop-out technique to at least a portion of the second model and repeating prediction for the data sample (Alarab Section 3.2, p. 1524, "In this method, the dropout is activated during the testing phase which is applied after each weight layer in a neural network." -- EN: the dropout kept active in the layers of the network at test time denotes the drop-out technique applied to at least a portion of the model, the model being in the combination the adapted check model; Section 3.2, p. 1524, "By drawing T samples from Bernoulli distribution, this produces {W1t, …, WLt}t=1T which so allows to express the approximated predictive mean of a given input as: Eq(y|x)(y) ≈ (1/T) Σt=1T ŷ(x, W1t, …, WLt) = pMC(y|x) (8)" -- EN: the T outputs ŷ(x, W1t, …, WLt), one for each of the T dropout realizations of the weights, denote the plurality of predicted labels obtained by repeating the prediction for the data sample x);
determining a class with a highest average confidence score among the plurality of predicted labels (Alarab Section 3.2, p. 1524, "Eq(y|x)(y) ≈ (1/T) Σt=1T ŷ(x, W1t, …, WLt) = pMC(y|x) (8)" -- EN: pMC(y|x), the mean of the T outputs, holds the average confidence score of each class; Section 4, p. 1526, "Predictive mean provides the correct or incorrect classification with respect to the actual labels." -- EN: classifying the data sample from the predictive mean denotes determining the class with the highest average confidence score; Algorithm 1, p. 1523, "c = argmax(M(x)) is the predicted class." -- EN: the class of the highest score is the predicted class); and
measuring the predictive uncertainty for the data sample based on confidence scores of the determined class included in the plurality of predicted labels (Alarab Section 3.2, p. 1525, "Hence the predictive uncertainty using mutual information can be expresses as: Î(y|x, Dtrain) = Ĥ(y|x, Dtrain) + Σc (1/T) Σt=1T p(y = c|x, w) log p(y = c|x, w) (9)" -- EN: the mutual information is computed from the confidence scores p(y = c|x, w) of the classes c across the T predictions, the determined class being among them, which corresponds to measuring the uncertainty based on the confidence scores of the determined class included in the plurality of predicted labels, the claim doesn’t exclude the confidence scores of the other classes from also entering the measure; Section 3.2, p. 1525, "Ĥ(y|x, Dtrain) = −Σc pMC(y = c|x, w) log pMC(y = c|x, w) (10)" -- EN: the entropy of the mean confidence scores enters the measure),
wherein values of the plurality of predicted labels are confidence scores for each class (Alarab Section 3.2, p. 1525, "p(y = c|x, w)" -- EN: each predicted label holds the probability of each class c, which denotes the confidence score for each class; Section 5.2, p. 1529, "The output layer is followed by softmax function to output the class prediction which is one of the handwritten digits from 0 to 9." -- EN: the softmax output of the network holds a score for each class).
The same motivation supplied under claim 7 also applies here.
Regarding claim 9, Chen in view of Liang in view of Balaji in view of Alarab teaches all the limitations of claim 1, Chen and Balaji further teach:
wherein the adjusting of the upper limit of the size constraint comprises:
obtaining a first predicted label for the data sample through the first model (Chen Section 3, p. 3, "Given a model f : X → Y trained on D" -- EN: f maps each input to a label; Framework 1, p. 4, "4: Identify those points on which the ensemble and f disagree as mis-classified points: RX := {x ∈ UX : Prh∼T{h(x) = f(x)} < τ}." -- EN: f(x), the label that the pre-trained model f assigns to the test input x, denotes the first predicted label obtained for the data sample through the first model);
obtaining a second predicted label for the data sample through the second model (Chen Algorithm 1, p. 6, "4: Let ỹx be the majority vote of {hi}i=1N: ỹx := arg maxj∈Y (1/N) Σi=1N I[hi(x) = j]." -- EN: hi(x), the label that the check model assigns to the same test input x, denotes the second predicted label obtained for the data sample through the second model, the check model being in the combination the adapted check model set out for claim 1; Section 4, p. 4, footnote 2, "This includes one single h as a special case." -- EN: a single check model is the second model); and
adjusting the upper limit based on a difference between the first predicted label and the second predicted label (Chen teaches forming, for each test input, the difference between the label of f and the label of the check model, marking the input as mis-classified when the two differ, and treating the inputs so they are marked differently from the rest ( Chen Framework 1, p. 4, "4: Identify those points on which the ensemble and f disagree as mis-classified points: RX := {x ∈ UX : Prh∼T{h(x) = f(x)} < τ}."; Algorithm 1, p. 6, "5: Set R = {(x, ỹx) : x ∈ UX, ỹx ≠ f(x)}."; Chen Section 4, p. 3, "The fundamental idea is to identify a point x as mis-classified if h disagrees with f on x."; Section 4, p. 4, "For each mis-classified data point x identified by the ensemble, we assign it a pseudo-label that is different from f(x)"). Balaji further teaches adjusting the bound of each input by a step γ up or down according to the outcome of a test run on that input alone, the bound of an ambiguous input being made smaller and the bound of an unambiguous input larger (Balaji Section 3, p. 3, "If PGD succeeded in finding an image with a different class label, then ϵi is too big, so we replace ϵi ← ϵi − γ. If PGD failed, then we set ϵi ← ϵi + γ."; Balaji Section 1, p. 2, "samples with small radii are often ambiguous and have nearby images of another class, while images with large radii have unambiguous class labels that are difficult to manipulate."). Thus, in the proposed combination, the bound of each test input is raised or lowered by the step γ according to whether the label of the check model for the input differs from the label f(x) of the pre-trained model, such that the upper limit is adjusted based on a difference between the first predicted label and the second predicted label).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Chen in view of Liang in view of Balaji in view of Alarab to include a bound for each test input that is raised or lowered by the step γ according to whether the label the check model assigns to the input differs from the label f(x) the pre-trained model assigns to it, as taught by Chen and Balaji, in order to perturb the inputs on which the two models disagree within a bound of their own (Chen Section 4, page 3-4; Balaji Section 1, p. 2).
The motivation for doing so would be to avoid one bound that is too large for the inputs the two models label differently and too small for the inputs they label alike, since Chen marks the inputs on which the check model departs from f as the likely errors of f and Balaji shrinks the bound of an ambiguous input and grows the bound of an unambiguous one, as taught by Chen and Balaji where the benefit is disclosed (Chen Section 4, page 3-4 and Framework 1, p. 4; Balaji Section 1, p. 2).
Regarding claim 16, Chen in view of Liang in view of Balaji in view of Alarab teaches all the limitations of claim 1, Chen further teaches:
wherein the evaluating of the performance of the first model comprises:
predicting a label of the evaluation dataset through the first model (Chen Algorithm 1, p. 6, "5: Set R = {(x, ỹx) : x ∈ UX, ỹx ≠ f(x)}." -- EN: f(x), the prediction of the pre-trained model f for each input x of the unlabeled test dataset UX, denotes the label predicted through the first model; Section 3, p. 3, "Given a model f : X → Y trained on D" -- EN: f maps each input to a label); and
evaluating the performance of the first model by comparing the pseudo label and the predicted label (Chen Algorithm 1, p. 6, "5: Set R = {(x, ỹx) : x ∈ UX, ỹx ≠ f(x)}. ... Output: Estimated mis-classified points RX = {x : (x, y) ∈ R}, estimated accuracy |UX \ RX| / |UX|." -- EN: testing whether the pseudo label ỹx equals f(x) for each input denotes the comparing, and the estimated accuracy formed from the points that pass the test denotes the performance evaluated by the comparing; Appendix B.2, p. 18, "Let arx(f, T) be the agreement rate between f and h's on a point x" -- EN: the agreement between the label of f and the label of the check model is the comparison; Appendix B.2, p. 18, "Recall that we are using ar(f, T) as an estimate of the accuracy of f on U." -- EN: the agreement rate over the test set is the estimated accuracy of the first model).
Claim(s) 4 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen in view of Liang in view of Balaji in view of Alarab in view of Keras Team, "EarlyStopping," class EarlyStopping in keras/callbacks.py, Keras v2.11.0 (released Nov. 14, 2022) (hereinafter Keras).
Regarding claim 4, Chen in view of Liang in view of Balaji in view of Alarab teaches all the limitations of claim 1 including “wherein the obtaining of the second model”. Liang further teaches:comprises:
… a domain adaptation process (Liang Algorithm 1, p. 9, "Input: source model fs = gs ∘ hs, target data {xti}i=1nt, maximum number of epochs Tm, trade-off parameter β." -- EN: the epochs over which the source model is adapted to the target data denote the domain adaptation process)
Chen in view of Liang in view of Balaji in view of Alarab does not explicitly teach: monitoring a training loss calculated during [a domain adaptation process];
determining a time when an amount of change in the training loss is equal to or less than a reference value as an early stop time; and
obtaining the second model by stopping the domain adaptation at the determined early stop time
However, Keras teaches:
monitoring a training loss calculated during [a domain adaptation process]; (Keras docstring, lines 1944-1950, "Stop training when a monitored metric has stopped improving. Assuming the goal of a training is to minimize the loss. With this, the metric to be monitored would be `'loss'`, and mode would be `'min'`. A `model.fit()` training loop will check at end of every epoch whether the loss is no longer decreasing, considering the `min_delta` and `patience` if applicable." -- EN: reading the training loss 'loss' at the end of every epoch of the training loop denotes the monitoring of the training loss, the training loop being in the combination of the adaptation of Liang);
determining a time when an amount of change in the training loss is equal to or less than a reference value as an early stop time (Keras Args, lines 1958-1963, "min_delta: Minimum change in the monitored quantity to qualify as an improvement, i.e. an absolute change of less than min_delta, will count as no improvement. patience: Number of epochs with no improvement after which training will be stopped." -- EN: min_delta denotes the reference value, and the change of the loss of an epoch from the best loss recorded so far denotes the amount of change in the training loss; Keras _is_improvement(), lines 2113-2114, "def _is_improvement(self, monitor_value, reference_value): return self.monitor_op(monitor_value - self.min_delta, reference_value)" -- EN: with mode 'min' the operator is np.less and min_delta is negated (lines 2033-2034 and 2049-2050), so an epoch counts as an improvement only when the loss has fallen by more than min_delta, and a fall equal to or less than min_delta counts as no improvement, which corresponds to the amount of change being equal to or less than the reference value; Keras on_epoch_end(), lines 2069-2085, "self.wait += 1 ... if self.wait >= self.patience and epoch > 0: self.stopped_epoch = epoch self.model.stop_training = True" -- EN: stopped_epoch, the epoch at which the count of epochs without improvement reaches patience, denotes the early stop time determined from the change in the loss); and
obtaining the second model by stopping the domain adaptation at the determined early stop time (Keras docstring, lines 1950-1951, "Once it's found no longer decreasing, `model.stop_training` is marked True and the training terminates." -- EN: the training that terminates is in the combination the adaptation of the check model, which corresponds to stopping the domain adaptation at the early stop time; Keras on_epoch_end(), lines 2083-2085, "if self.wait >= self.patience and epoch > 0: self.stopped_epoch = epoch self.model.stop_training = True" -- EN: the stop is marked at the stopped epoch; Keras Args, lines 1976-1979, "restore_best_weights: Whether to restore model weights from the epoch with the best value of the monitored quantity. If False, the model weights obtained at the last step of training are used." -- EN: the model with the weights obtained at the last step of the training, the step at which the training was stopped, denotes the second model obtained by stopping the domain adaptation, the model being in the combination the adapted check model of Liang).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Chen in view of Liang in view of Balaji in view of Alarab to include an early stopping callback on the adaptation of the check model, which reads the training loss at the end of every epoch, counts an epoch as no improvement when the loss has not fallen by more than min_delta, and marks the training to stop after patience such epochs, as taught by Keras, in order to end the adaptation once its loss is no longer decreasing (docstring, lines 1946-1951). Liang runs the adaptation for a fixed maximum number of epochs Tm (Algorithm 1, p. 9).
The motivation for doing so would be to spare the epochs that no longer lower the loss, since the callback ends the training loop after the patience epochs without improvement so that only the epochs up to the stop are run rather than all the epochs requested, as taught by Keras where the benefit is disclosed (Example, lines 1990-1999).
Claim(s) 6 and 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen in view of Liang in view of Balaji in view of Alarab in view of Vitorino et al., "Adaptative Perturbation Patterns: Realistic Adversarial Learning for Robust Intrusion Detection," (hereinafter Vitorino).
Regarding claim 6, Chen in view of Liang in view of Balaji in view of Alarab teaches all the limitations of claim 1 including “wherein the predefined factor comprises” … “and a second factor related to characteristics of the data sample” (Balaji Section 3, p. 3, "The proposed algorithm is distinctive in that it uses a different ϵi for each image xi. ... After crafting a perturbation for xi, we check if the perturbed image was a successful adversarial example. If PGD succeeded in finding an image with a different class label, then ϵi is too big, so we replace ϵi ← ϵi − γ. If PGD failed, then we set ϵi ← ϵi + γ." -- EN: the step γ applied to the radius of image xi up or down according to the outcome of that image's own perturbation denotes the second factor related to characteristics of the data sample, the predefined factor being construed as set out above for claim 1).
Vitorino teaches:
a first factor related to characteristics of the evaluation dataset (Section 3.1, p. 5, "The Interval pattern encapsulates a mechanism that records the valid intervals to create perturbations tailored to the characteristics of each feature (Figure 2)." -- EN: the recorded interval is a characteristic of the data; Section 3.1, p. 5, Eq. (1)-(2), "mi = mi−1 ∗ k + min(xi) ∗ (1 − k) (1) Mi = Mi−1 ∗ k + max(xi) ∗ (1 − k) (2) where min(xi) and max(xi) are the actual minimum and maximum values of the samples xi of batch i." -- EN: the minimum mi and the maximum Mi are taken over the samples of the data, which corresponds to a characteristic of the dataset rather than of one sample; Section 3.1, p. 5, Eq. (3), "The random number ε ∈ (0, 1] acts as a ratio to scale the interval. To restrict its possible values, it is generated within the standard range of [0.1, 0.3], although other ranges can be configured. For a given feature, a perturbation Pi of a batch i can be represented as: Pi = (Mi − mi) ∗ ε (3)" -- EN: the interval width Mi − mi, by which the ratio ε is multiplied to give the size of the perturbation, denotes the first factor, the largest perturbation 0.3 (Mi − mi) being the upper limit that it sets; Section 4.4, p. 10, "A2PM was applied to perform adversarial attacks against the fine-tuned models for a maximum of 50 iterations, by adapting to the data of the holdout evaluation sets." -- EN: the intervals are recorded from the samples of the evaluation set, whose characteristics they are, the evaluation set being in the combination the unlabeled test dataset of Chen).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Chen in view of Liang to include perturbing each unlabeled test input within a bound of its own, raised by a fixed step γ when the perturbed input keeps its class and lowered by γ when it does not, as taught by Balaji, in order to fit the bound of each input to its distance from the decision boundary (Section 1, p. 2).
The motivation for doing so would be to avoid using one perturbation size that is too large for inputs near the decision boundary and too small for the others, since each input gets its own bound that shrinks as soon as the perturbed input changes class, as taught by Balaji where the benefit is disclosed (Section 1, p. 1 and Fig. 1, p. 2).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Chen in view of Liang in view of Balaji in view of Alarab to include an interval from the minimum to the maximum value recorded from the test inputs, by which the size of each input's perturbation is scaled and at which the perturbed input is capped, as taught by Vitorino, in order to keep every perturbed input within the values that the data actually takes (Section 3.1, p. 5).
The motivation for doing so would be to keep the perturbed inputs realistic, since a perturbation scaled by the interval recorded from the data and capped at that interval cannot leave the valid minimum and maximum values of the data, as taught by Vitorino where the benefit is disclosed (Abstract, p. 1 and Section 3.1, p. 5).
Regarding claim 11, Chen in view of Liang in view of Balaji in view of Alarab teaches all the limitations of claim 1 including “wherein the adjusting of the upper limit of the size constraint comprises”. Vitorino teaches:
adjusting the upper limit based on a degree of dispersion of data samples of the evaluation dataset (Section 3.1, p. 5, Eq. (1)-(2), "mi = mi−1 ∗ k + min(xi) ∗ (1 − k) (1) Mi = Mi−1 ∗ k + max(xi) ∗ (1 − k) (2) where min(xi) and max(xi) are the actual minimum and maximum values of the samples xi of batch i." -- EN: the interval from the minimum to the maximum value of the samples, the range of the samples, denotes the degree of dispersion of the data samples; Section 3.1, p. 5, "Instead of a static interval, moving intervals can be utilized after the first batch to enable an incremental adaptation to new data, according to a configured momentum." -- EN: the interval, and with it the largest perturbation, is re-set from each batch of samples, which corresponds to the adjusting based on the dispersion of the samples; Section 3.1, p. 5, Eq. (3), "To restrict its possible values, it is generated within the standard range of [0.1, 0.3], although other ranges can be configured. For a given feature, a perturbation Pi of a batch i can be represented as: Pi = (Mi − mi) ∗ ε (3)" -- EN: the largest perturbation, 0.3 (Mi − mi), denotes the upper limit, which is larger for a wider interval and smaller for a narrower one; Section 3.1, p. 5, "The resulting value is capped at the current interval to ensure it remains within the valid minimum and maximum values of that feature." -- EN: the cap at the interval is a further limit set from the dispersion of the samples; Section 4.4, p. 10, "by adapting to the data of the holdout evaluation sets" -- EN: the samples from which the interval is recorded are those of the evaluation set, which in the combination is the unlabeled test dataset of Chen).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Chen in view of Liang in view of Balaji in view of Alarab to include an interval from the minimum to the maximum value recorded from the test inputs, by which the size of each input's perturbation is scaled and at which the perturbed input is capped, as taught by Vitorino, in order to keep every perturbed input within the values that the data actually takes (Section 3.1, p. 5).
The motivation for doing so would be to keep the perturbed inputs realistic, since a perturbation scaled by the interval recorded from the data and capped at that interval cannot leave the valid minimum and maximum values of the data, as taught by Vitorino where the benefit is disclosed (Abstract, p. 1 and Section 3.1, p. 5).
Claim(s) 10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen in view of Liang in view of Balaji in view of Alarab in view of Qiao et al., "Deep Co-Training for Semi-Supervised Image Recognition," (hereinafter Qiao).
Regarding claim 10, Chen in view of Liang in view of Balaji in view of Alarab teaches all the limitations of claims 1 and 9 including “wherein the difference between the first predicted label and the second predicted label is calculated”. Qiao teaches:
[difference between labels of a first and second label] based on Jensen-Shannon divergence (JSD) (Section 2.1, p. 4, "In other words, we want networks p1(x) = f1(v1(x)) and p2(x) = f2(v2(x)) to have close predictions on U. Therefore, we use a natural measure of similarity, the Jensen-Shannon divergence between p1(x) and p2(x), i.e., Lcot(x) = H(½(p1(x) + p2(x))) − ½(H(p1(x)) + H(p2(x))) (4) where x ∈ U and H(p) is the entropy of p." -- EN: p1(x) and p2(x), the predictions of two different networks for the same unlabeled sample x, denote the first predicted label and the second predicted label, which in the combination are the predictions of the pre-trained model f and of the check model for the test input as set out for claim 9, and the Jensen-Shannon divergence between them denotes the difference calculated based on Jensen-Shannon divergence).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Chen in view of Liang in view of Balaji in view of Alarab to include, as the difference between the label of the pre-trained model f and the label of the check model for a test input, the Jensen-Shannon divergence between the two models' predicted class distributions for that input, as taught by Qiao, in order to measure how close the two models' predictions for the input are (Section 2.1, p. 4).
The motivation for doing so would be to grade the difference between the two models' predictions for an input rather than only test whether their top classes match, since the Jensen-Shannon divergence compares the two networks' full predicted distributions p1(x) and p2(x) on the input and is a natural measure of their similarity, as taught by Qiao where the benefit is disclosed (Section 2.1, p. 4).
Claim(s) 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen in view of Liang in view of Balaji in view of Alarab in view of Vitorino in view of Li et al., "Uncertainty Modeling for Out-of-Distribution Generalization," (hereinafter Li).
Regarding claim 12, Chen in view of Liang in view of Balaji in view of Alarab in view of Vitorino teaches all the limitations of claims 1 and 11 including “wherein the data samples are images” (Chen Abstract, p. 1, "Experimental results on 59 tasks over five dataset categories including image classification and sentiment classification datasets" -- EN: the test inputs of the image classification datasets are images; Chen Section 7.1, p. 7, "Specifically, we use the following dataset categories: Digits (including MNIST [26], MNIST-M [12], SVHN [29], USPS [19]), Office-31 [33], CIFAR10-C [24], iWildCam [1] and Amazon Review [2]." -- EN: the Digits, Office-31, CIFAR10-C and iWildCam inputs are images) “, and the adjusting of the upper limit based on the degree of dispersion of the data samples of the evaluation dataset” (see mapping above for claim 11).
Li teaches:
comprises:
calculating a representative vale of each of the data samples based on a pixel value of each of the data samples (Section 3.1, p. 3, Eq. (1), "Given x ∈ RB×C×H×W to be the encoded features in the intermediate layers of the network, we denote μ ∈ RB×C and σ ∈ RB×C as the channel-wise feature mean and standard deviation of each instance in a mini-batch, respectively, which can be formulated as: μ(x) = (1/HW) Σh=1H Σw=1W xb,c,h,w (1)" -- EN: μ(x), the mean over the H × W positions of the values of instance b, denotes the representative value calculated for each data sample; Section 5, p. 7, "Here we name the positions of ResNet after first Conv, Max Pooling layer, 1,2,3,4-th ConvBlock as 0,1,2,3,4,5 respectively." -- EN: at position 0 the values xb,c,h,w are the outputs of the first convolution applied to the pixels of image b, so the mean is calculated based on the pixel values of the image; Section 5, p. 7, "Based on the analysis, we plug the module into positions 0-5 in all experiments." -- EN: the means are calculated at position 0 among others; Section 3.1, p. 3, "As the abstraction of features, feature statistics can capture informative characteristics of the corresponding domain (such as color, texture, and contrast)" -- EN: the mean summarizes the image); and
-- EN: the instant specification's representative value denotes one value summarizing the pixel values of an image, an average pixel value being the example given (instant specification, Para. [0107], "the evaluation system 10 may extract a representative value from each of data samples belonging to an evaluation dataset"; Para. [0108], "the evaluation system 10 may extract a representative value (e.g., an average pixel value) from each of image samples of the source dataset 91"). Li's channel-wise mean μ(x) of the values of an image at the first position of the network corresponds to that value.
measuring the degree of dispersion of the data samples based on a degree of dispersion of calculated representative values (Section 3.2.1, p. 4, Eq. (3), "we propose a simple yet effective non-parametric method for uncertainty estimation, utilizing the variance of the feature statistics to provide some instructions: Σμ2(x) = (1/B) Σb=1B (μ(x) − Eb[μ(x)])2 (3)" -- EN: Σμ2, the variance over the B images of their means μ(x), denotes the degree of dispersion of the data samples measured based on the dispersion of the calculated representative values; Section 3.2.1, p. 4, "Although the underlying distribution of the domain shifts is unpredictable, the uncertainty estimation captured from the mini-batch can provide an appropriate and meaningful variation range for each feature channel" -- EN: the variance of the per-image means sets the variation range of the perturbation, which in the combination is the dispersion by which the upper limit is adjusted as set out for claim 11; Section 3.2.2, p. 4, Eq. (5), "β(x) = μ(x) + εμ Σμ(x), εμ ∼ N(0, 1)" -- EN: the perturbation εμ Σμ(x) is scaled by the dispersion Σμ).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Chen in view of Liang in view of Balaji in view of Alarab in view of Vitorino to include, for each test image, a channel-wise mean of its values at the first position of the check model and, over the test images, the variance of those means as the dispersion of the evaluation data by which the perturbation of the images is scaled, as taught by Li, in order to obtain the variation range of the perturbation from the statistics of the images themselves (Section 3.2.1, p. 4).
The motivation for doing so would be to give the perturbation an appropriate and meaningful variation range, since the variance of the per-image statistics across the batch reveals how much each channel may change while a range set without that instruction harms the model, as taught by Li where the benefit is disclosed (Section 3.2.1, p. 4 and Section 5, page 7-8).
Claim(s) 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen in view of Liang in view of Balaji in view of Alarab in view of Vitorino in view of Li in view of Ishii et al., "Source-free Domain Adaptation via Distributional Alignment by Matching Batch Normalization Statistics," (hereinafter Ishii).
Regarding claim 13, Chen in view of Liang in view of Balaji in view of Alarab in view of Vitorino in view of Li teaches all the limitations of claims 1, 11, and 12 including “wherein the degree of dispersion of the representative values is a first degree of dispersion” (Li Section 3.2.1, p. 4, Eq. (3), "Σμ2(x) = (1/B) Σb=1B (μ(x) − Eb[μ(x)])2 (3)" -- EN: the variance Σμ2 of the per-image means, mapped for claim 12 as the degree of dispersion of the representative values, is the first degree of dispersion) “, and the measuring of the degree of dispersion of the data samples based on the degree of dispersion of the calculated representative values” (see mapping above for claim 12).
Ishii teaches:
comprises:
obtaining a second degree of dispersion of data samples of the labeled dataset from training history data of the labeled dataset of the source domain without accessing the labeled dataset (Abstract, p. 1, "In this setting, we cannot access source data during adaptation, while unlabeled target data and a model pretrained with source data are given." -- EN: the source data on which the model was pretrained denotes the labeled dataset of the source domain, which in the combination is the labeled dataset D of Chen on which the first model was trained, and the source data is not accessed; Abstract, p. 1, "we propose utilizing batch normalization statistics stored in the pretrained model to approximate the distribution of unobserved source data." -- EN: the stored statistics stand in for the source data; Section 3.1, p. 3, "the BN statistics stored in the first BN in the classifier can be seen as the statistics of the source features extracted by the pretrained encoder. We approximate the source-feature distribution by using these statistics. Specifically, we simply use a Gaussian distribution for each channel denoted by N(μ̂c, σ̂c2) where μ̂c and σ̂c2 are the mean and variance of the Gaussian distribution which are the stored BN statistics corresponding to the c-th channel." -- EN: σ̂c2, the stored variance of the source features, denotes the second degree of dispersion of the data samples of the labeled dataset, obtained from the stored statistics rather than from the data; Section 2.2, p. 2, "BN stores the exponentially weighted averages of the BN statistics in the training phase and uses them in the inference phase to compute z̃ in Eq. (2)" -- EN: the averages of the means and variances accumulated over the training phase of the model denote the training history data of the labeled dataset; Section 3.2, p. 4, "Note that this optimization can be conducted without the source data, which means that we do not need to access to the source data during adaptation." -- EN: the variance is obtained and used without accessing the source data); and
-- EN: the instant specification's training history data denotes data recorded from the training of the first model on the source dataset, including the results of normalization operations that used the average and standard deviation of the source samples (instant specification, Para. [0109], "since the training history data of the first model includes the results of preprocessing operations (e.g., normalization operations using average, standard deviation, etc.) on the image samples of the source dataset 91, the evaluation system 10 can obtain representative values for the image samples from this training history data without accessing the source dataset 91"). Ishii's batch normalization statistics, the averages of the means and variances of the source features accumulated in the training phase and stored in the model, correspond to that data.
calculating a degree of dispersion of the data samples based on a ratio of the first degree of dispersion to the second degree of dispersion (Section 3.1, p. 3, Eq. (3), "LBNM({xi}i=1B, θ) = (1/C) Σc=1C KL(N(μ̂c, σ̂c2) || N(μc, σc2)) = (1/2C) Σc=1C (log(σc2/σ̂c2) + (σ̂c2 + (μ̂c − μc)2)/σc2 − 1)," -- EN: the term log(σ_c²/σ̂_c²) is formed from the ratio of the target-side dispersion σ_c² to the stored source-side dispersion σ̂_c² mapped above, which corresponds to calculating a degree of dispersion based on a ratio of the first degree of dispersion to the second degree of dispersion).
-- EN: the instant specification's degree of dispersion calculated based on the ratio denotes the relative dispersion of the evaluation dataset with respect to the source dataset (instant specification, Para. [0107], "use a ratio (i.e., relative degree of dispersion) of the degree of dispersion of the evaluation dataset to the degree of dispersion of the source dataset (or another target dataset) as the value of the third factor"). Ishii's term log(σc2/σ̂c2), the logarithm of the ratio of the target variance to the source variance, corresponds to that relative dispersion.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Chen in view of Liang in view of Balaji in view of Alarab in view of Vitorino in view of Li to include the variance of the source data taken from the batch normalization statistics stored in the first model from its training, and a term formed from the ratio of the variance of the test images to that stored variance, as taught by Ishii, in order to compare the test data with the source data when the source data cannot be accessed (Abstract, p. 1). Ishii itself works in the source-free setting of Liang, which it cites (Section 1, p. 1).
The motivation for doing so would be to obtain the dispersion of the source data without the source data, since the batch normalization statistics accumulated while the model trained on the source data are stored in the model and represent the distribution of the unobserved source data, as taught by Ishii where the benefit is disclosed (Abstract, p. 1 and Section 3.1, p. 3).
Claim(s) 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen in view of Liang in view of Balaji in view of Alarab in view of Laves et al., "Well-calibrated Model Uncertainty with Temperature Scaling for Dropout Variational Inference," (hereinafter Laves).
Regarding claim 14, Chen in view of Liang in view of Balaji in view of Alarab teaches all the limitations of claim 1 including “wherein the adjusting of the upper limit of the size constraint comprises adjusting the upper limit…”. Laves teaches:
[adjusting the upper limit] based on a number of classes of the evaluation dataset (Laves Section 2, p. 2, Eq. (3), "Entropy of the softmax likelihood is used to describe uncertainty of prediction [4]. In contrast to confidence as a measure of goodness of prediction, uncertainty takes into account the likelihoods of all C classes. We introduce normalization to scale the values to a range between 0 and 1: H̃(p) := −(1/log C) Σc=1C p(c) log p(c), H̃ ∈ [0, 1]." – EN: this denotes the entropy of a Monte Carlo dropout model's averaged softmax output is normalized by dividing it by the logarithm of the number of classes.)
-- EN: the instant specification's adjusting based on a number of classes denotes scaling the upper limit by a factor computed from the number of classes, the logarithm of the number of classes being one example given (instant specification, Para. [0114], "the evaluation system 10 may calculate the class complexity by taking the log (e.g., natural log) or square root of the number of classes"; Para. [0115], "the evaluation system 10 may adjust the upper limit of the size constraint to be applied to the adversarial noise of a corresponding data sample based on the measured value"). Laves' factor 1/log C, applied to the uncertainty by which the bound is adjusted, corresponds to such a factor.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Chen in view of Liang in view of Balaji in view of Alarab to include, for each test input, an uncertainty equal to the entropy of the check model's averaged dropout output divided by log C, C being the number of classes, as taught by Laves, in order to scale the uncertainty values to a range between 0 and 1 (Section 2, p. 2).
The motivation for doing so would be to bound the uncertainty of every input between 0 and 1 whatever the number of classes, since dividing the entropy by log C scales the values to the range between 0 and 1 in which one threshold on the uncertainty can be applied to every input, as taught by Laves where the benefit is disclosed (Section 2, p. 2 and Section 3.3, p. 4).
Claim(s) 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen in view of Liang in view of Balaji in view of Alarab in view of Miyato et al., "Virtual Adversarial Training: A Regularization Method for Supervised and Semi-Supervised Learning," (hereinafter Miyato).
Regarding claim 15, Chen in view of Liang in view of Balaji in view of Alarab teaches all the limitations of claim 1 including “wherein the deriving of the adversarial noise”.
Miyato teaches:
comprises:
obtaining a first predicted label for the data sample through the second model (Section 3.2, p. 4, "Therefore, in this study, we use the current estimate p(y|x, θ̂) in place of q(y|x)." -- EN: p(y|x, θ̂), the output distribution of the model for the unperturbed input x, denotes the first predicted label obtained through the model, which in the combination is the adapted check model; Section 3.2, p. 4, "LDS(x∗, θ) := D[p(y|x∗, θ̂), p(y|x∗ + rvadv, θ)] (5)" -- EN: p(y|x∗, θ̂) is the model's prediction for the data sample x∗; Section 1, p. 2, "In other words, even in the absence of label information, virtual adversarial direction can be defined on an unlabeled data point as if there is a “virtual” label" -- EN: the model's own prediction stands in for a label of the unlabeled data sample);
generating a specific noisy sample by reflecting a value of a noise parameter in the data sample (Section 3.3, p. 5, "To summarize, we can approximate rvadv with the repeated application of the following update: d ← ∇rD(r, x, θ̂)|r=ξd. (13)" -- EN: r, the perturbation added to x, denotes the noise parameter, and r = ξd, the direction vector d scaled by ξ, denotes its value; Algorithm 1, p. 6, "2) Generate a random unit vector d(i) ∈ RI using an iid Gaussian distribution. 3) Calculate rvadv via taking the gradient of D with respect to r on r = ξd(i) on each input data point x(i)" -- EN: x(i) + r with r = ξd(i), the data point with the current value of the noise parameter added, denotes the specific noisy sample);
obtaining a second predicted label for the specific noisy sample through the second model (Section 3.3, p. 5, "where g = ∇rD[p(y|x, θ̂), p(y|x + r, θ̂)]|r=ξd. (15)" -- EN: p(y|x + r, θ̂), the output distribution of the model for the noisy sample x + r, denotes the second predicted label obtained through the model; Algorithm 1, p. 6, "g(i) ← ∇rD[p(y|x(i), θ̂), p(y|x(i) + r, θ̂)]|r=ξd(i)" -- EN: the model is evaluated on x(i) + r);
updating the value of the noise parameter in a direction to increase a difference between the first predicted label and the second predicted label (Section 3.2, p. 4, "rvadv := arg maxr;‖r‖2≤ϵ D[p(y|x∗, θ̂), p(y|x∗ + r)], (6) which defines our virtual adversarial perturbation." -- EN: D, the divergence between the model's output distributions for the data sample and for the noisy sample, denotes the difference between the first predicted label and the second predicted label, and the perturbation sought is the one that maximizes it; Section 3.3, p. 5, "To summarize, we can approximate rvadv with the repeated application of the following update: d ← ∇rD(r, x, θ̂)|r=ξd. (13)" -- EN: replacing d by the gradient of D with respect to r, the direction along which D grows, denotes updating the value of the noise parameter in a direction to increase the difference); and
calculating the adversarial noise for the data sample based on the updated value of the noise parameter (Section 3.3, p. 5, "rvadv ≈ ϵ g/‖g‖2 (14)" -- EN: rvadv, the virtual adversarial perturbation computed by normalizing the updated direction g and scaling it to ϵ, denotes the adversarial noise calculated from the updated value of the noise parameter).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Chen in view of Liang in view of Balaji in view of Alarab to include computing the perturbation of each test input as a virtual adversarial perturbation, that is, the direction that changes the check model's output for the input the most, found by a few gradient steps and scaled to the input's bound, in place of the PGD perturbation, as taught by Miyato, in order to perturb each input in the direction that changes the check model's output the most (Section 1, p. 2).
The motivation for doing so would be to compute the perturbation of each unlabeled test input without needing any label for it, since the virtual adversarial direction is defined from the model's output alone, as taught by Miyato where the benefit is disclosed (Abstract, p. 1 and Section 3.2, p. 4).
Claim(s) 17 and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen in view of Liang in view of Balaji in view of Alarab in view of Lee et al. (US 2021/0133585 A1) (hereinafter Lee).
Regarding claim 17, it is a performance evaluation system claim that recites substantially the same limitations as claim 1, and is therefore rejected under the same rationale as claim 1. Chen in view of Liang in view of Balaji in view of Alarab do not explicitly teach:
A performance evaluation system comprising:
one or more processors; and
a memory configured to store a computer program which is to be executed by the one or more processors,
wherein the computer program comprises instructions for performing:
However, Lee teaches:
A performance evaluation system comprising: ([0083], "An illustrated computing environment 10 includes a computing device 12. In one embodiment, the computing device 12 may be an apparatus 100 for unsupervised domain adaptation")
one or more processors; and ([0084], "The computing device 12 includes at least one processor 14, a computer-readable storage medium 16, and a communication bus 18.")
a memory configured to store a computer program which is to be executed by the one or more processors, ([0085], "A program 20 stored in the computer-readable storage medium 16 includes a set of instructions executable by the processor 14. … the computer-readable storage medium 16 may be a memory (volatile memory such as a random access memory, non-volatile memory, or any suitable combination thereof)")
wherein the computer program comprises instructions for performing: ([0084], "The one or more programs may include one or more computer-executable instructions, which, when executed by the processor 14, may be configured to cause the computing device 12 to perform operations according to the exemplary embodiment." ).
The remaining limitations of claim 17 are substantially the same as that of method claim 1. Therefore, claim 17 is rejected under the same rationale as claim 1.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Chen in view of Liang in view of Balaji in view of Alarab to include one or more processors and a memory holding the evaluation method as computer-executable instructions, as taught by Lee, in order to run the method on a programmed computing device ([0084]).
The motivation for doing so would be to run the evaluation method on general-purpose computing hardware, as taught by Lee where the benefit is disclosed ([0084] and [0085]).
Regarding claim 18, it is a non-transitory computer-readable recording medium claim that recites substantially the same limitations as claim 1, and is therefore rejected under the same rationale as claim 1. Chen in view of Liang in view of Balaji in view of Alarab do not explicitly teach:
A non-transitory computer-readable recording medium configured to store a computer program to be executed by one or more processors to perform:
However, Lee teaches:
A non-transitory computer-readable recording medium configured to store a computer program ([0085], "The computer-readable storage medium 16 is configured to store the computer-executable instruction or program code, program data, and/or other suitable forms of information. A program 20 stored in the computer-readable storage medium 16 includes a set of instructions executable by the processor 14."; [0085], "non-volatile memory, or any suitable combination thereof), one or more magnetic disk storage devices, optical disk storage devices, flash memory devices") to be executed by one or more processors to perform: ([0084], "The one or more programs may include one or more computer-executable instructions, which, when executed by the processor 14, may be configured to cause the computing device 12 to perform operations according to the exemplary embodiment.").
The remaining limitations of claim 18 are substantially the same as that of method claim 1. Therefore, claim 17 is rejected under the same rationale as claim 1.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Chen in view of Liang in view of Balaji in view of Alarab to include a computer-readable recording medium holding the evaluation method as computer-executable instructions, as taught by Lee, in order to run the method on a programmed computing device ([0085]).
The motivation for doing so would be to run the evaluation method on general-purpose computing hardware, as taught by Lee where the benefit is disclosed ([0084] and [0085]).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NAYMUR RAHMAN ALI whose telephone number is (571)272-0007. The examiner can normally be reached Mon-Fri. 9:30-6:30 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571)270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/NAYMUR RAHMAN ALI/ Examiner, Art Unit 2123
/ALEXEY SHMATOV/ Supervisory Patent Examiner, Art Unit 2123