DETAILED ACTION
Notices to Applicant
This communication is a non-final rejection. Claims 1-12, as filed 08/19/2024, are currently pending and have been considered below.
Priority is generally acknowledged as shown on the filing receipt with the earliest date being 03/02/2022.
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon and the rationale supporting the rejection would be the same under either status.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
Use of the word “means” (or “step for”) in a claim with functional language creates a rebuttable presumption that the claim element is to be treated in accordance with 35 U.S.C. 112(f) (pre-AIA 35 U.S.C. 112, sixth paragraph). The presumption that 35 U.S.C. 112(f) (pre-AIA 35 U.S.C. 112, sixth paragraph) is invoked is rebutted when the function is recited with sufficient structure, material, or acts within the claim itself to entirely perform the recited function.
Absence of the word “means” (or “step for”) in a claim creates a rebuttable presumption that the claim element is not to be treated in accordance with 35 U.S.C. 112(f) (pre-AIA 35 U.S.C. 112, sixth paragraph). The presumption that 35 U.S.C. 112(f) (pre-AIA 35 U.S.C. 112, sixth paragraph) is not invoked is rebutted when the claim element recites function but fails to recite sufficiently definite structure, material or acts to perform that function.
Claim elements in this application that use the word “means” (or “step for”) are presumed to invoke 35 U.S.C. 112(f) except as otherwise indicated in an Office action. Similarly, claim elements that do not use the word “means” (or “step for”) are presumed not to invoke 35 U.S.C. 112(f) except as otherwise indicated in an Office action.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
The claim limitations “first selector,” “calculator,” “second selector,” and “generator” have been interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because it uses/they use a generic placeholder “*” coupled with functional language “*” without reciting sufficient structure to achieve the function. Furthermore, the generic placeholder is not preceded by a structural modifier.
Since the claim limitation(s) invokes 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, claim 11 has been interpreted to cover the corresponding structure described in the specification that achieves the claimed function, and equivalents thereof.
A review of the specification provides structure for each limitation: low-frequency class selector 10, inter-class distance calculator 20, mixed class selector 30, and data mixer 70 with label mixer 80 ([0053] of the specification as filed). Each of these functional limitations is “normally realized by a program executor such as a processor reading out and executing software (a program) recorded in a recording medium such as a ROM” (id.). FIGS. 3, 4, 5, and 7 provide algorithms for these functions. Accordingly, the structure of these terms is a processor executing the algorithms of FIGS. 3, 4, 6, and 7.
If applicant wishes to provide further explanation or dispute the examiner’s interpretation of the corresponding structure, applicant must identify the corresponding structure with reference to the specification by page and line number, and to the drawing, if any, by reference characters in response to this Office action.
If applicant does not intend to have the claim limitation(s) treated under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112 , sixth paragraph, applicant may amend the claim(s) so that it/they will clearly not invoke 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, or present a sufficient showing that the claim recites/recite sufficient structure, material, or acts for performing the claimed function to preclude application of 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
For more information, see MPEP § 2173 et seq. and Supplementary Examination Guidelines for Determining Compliance With 35 U.S.C. 112 and for Treatment of Related Issues in Patent Applications, 76 FR 7162, 7167 (Feb. 9, 2011).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Wen (Yeming Wen et al., "Combining Ensembles and Data Augmentation can Harm your Calibration" arXiv:2010.09875v1) in view of Kono (Reference C1 on IDS dated 08/19/2024).
Regarding claims 1 and 11, Wen discloses: A training data generation method for generating training data for training a recognition model that is input with image data and outputs one of a plurality of classes as a class of an object present in the image data (“Ideally, we prefer our model to be confident when it is predicting over easy classes such as cars and ships. For harder classes like cats and dogs, the model is encouraged to be less confident to achieve better calibration,” page 5; “Applying the correction produces new state-of-the art in uncertainty calibration across CIFAR-10, CIFAR-100, and ImageNet.,” Abstract), the training data generation method comprising:
--selecting a first class from the plurality of classes based on a recognition accuracy of the recognition model (“instead of a uniform Mixup hyperparameter for all classes, we propose to adjust the Mixup hyperparameter of each class by the difference between its accuracy and confidence,” page 5; “under-confident” on page 6; “We compute the accuracy and confidence on a validation dataset after each training epoch,” page 6);
--generating the training data by mixing image data and labels of each of the first class selected and the second class selected (“Mixup: Mixup (Zhang et al., 2018) manipulates both the features and the labels in order to encourage linearly interpolating predictions. Given an example (xi; yi), Mixup applies [equation 3],” page 3).
Wen does not expressly disclose but Kono teaches:
--calculating an inter-class distance that is a distance between the first class and each of two or more other classes among the plurality of classes (“Interclass distance is defined as a distance between class centroids in latent space. In other words, if the center of gravity of class y is denoted as…” page 3);
--selecting a second class for generating the training data from the two or more other classes, based on the inter-class distance (“Instead of interpolation between all two classes among target classes, a method of selecting only a part of the classes is introduced. For this purpose, the interclass distance is introduced so that two classes with a distance greater than a threshold are not interpolated,” page 3).
One of ordinary skill in the art before the effective filing date would have been motivated to expand Wen’s per-class mixing to include the inter-class distance selection of Kono because choosing the class to be mixed with the first class based on the distance between them avoids generating composite images from class pairs that are too separate and for which augmentation provides little benefit (“The reason for this is that two classes with a large distance between the classes are less likely to be confused and there is little need for data augmentation, but because of the distance, images of classes that are not relevant are generated, or rather than that, there is a risk of adverse effects,” page 3).
Claims 2, 3, and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Wen in view of Kono, and Abad (US20200110969A1).
Regarding claim 2 and 12, Wen and Kono do not expressly disclose but Abad teaches: extracting, as at least one candidate class, at least one of a class for which the recognition accuracy is at most a first threshold, or a class for which the recognition accuracy is at least a second threshold higher than the first threshold, among the plurality of classes, wherein in the selecting of the first class, the first class is selected from the at least one candidate class (“an accuracy threshold value, e.g., a value that is lower than the desired or target accuracy, may be used to quickly improve the accuracy of all relevant classes up to the accuracy threshold value,” [0016]; “In block 708, it is determined whether the accuracy for each class meets or exceeds an accuracy threshold value. This value may be different from the desired accuracy,” [0055]; “In bock 710, the number of the samples associated with classes that fall below the accuracy threshold value may be increased in the training batch in order to generate an adjusted training batch,” [0056]; “the classification model 110 may be an image classifier and may include at least five different classes: class A, class B, class C, class D, and class E,” [0019]; “Embodiments may also be at least partly implemented as instructions contained in or on a non-transitory computer-readable medium, which may be read and executed by one or more processors to enable performance of the operations described herein,” [0062]).
One of ordinary skill in the art before the effective filing date would have been motivated to expand the training data generation of Wen and Kono to include the accuracy thresholds of Abad because comparing each class’s accuracy against a threshold ensures that all classes reach an acceptable level of accuracy even which the desired accuracy cannot be achieved (Abad [0016]).
Regarding claim 3, Wen, Kono, and Abad do not expressly disclose: wherein in the selecting of the first class, the first class is selected at random from the at least one candidate class (“In block 708, it is determined whether the accuracy for each class meets or exceeds an accuracy threshold value,” [0055]), but the references do not provide selection logic among the candidate classes. Choosing one class rather than another is a choice among a finite number of identified, predictable alternatives. Thus, selecting one at random would have been obvious to one of ordinary skill in the art before the effective filing date with a reasonable expectation of success.
Claims 4 and 7 are rejected under 35 U.S.C. 103 as being unpatentable over Wen in view of Kono, and Li (Bingcong Li et al., "Confusable Learning for Large-class Few-Shot Classification” arXiv:2011.03154).
Regarding claim 4, Wen and Kono do not expressly disclose but Li teaches: wherein in the calculating of the inter-class distance, the inter-class distance is calculated based on a likelihood of each of the plurality of classes, output by the recognition model (“Instead of counting the times of classifying the class I as class j, we utility the probability of classifying instances of class i as class j: [equation 3]” page 6; “soft confusion matrix” page 6; “the classes with higher confusion are more likely to be sampled,” page 7).
One of ordinary skill in the art before the effective filing date would have been motivated to expand the training data generation of Wen and Kono to include the calculation of the class-to-class measure from the class probabilities output by the model as taught by Li because using this probability exploits more detailed information about the error the model made for each instance (Li page 6: “We will show in the experiment section that by using such a definition Confusable Learning exploit more detailed information about the error the model made for each instance.”).
Additionally, each element is taught by either Wen, Kono, or Li. Li’s calculation of class-to-class measure from probabilities the model outputs does not affect the normal functioning of the elements taught by Wen and Kono, so the combination is no more than an assembly of old elements, each performing as it did separately, with predictable results. It would therefore have been obvious to a POSITA before the effective filing date to combine the teachings of Li, Wen, and Kono.
Regarding claim 7, Wen does not expressly disclose but Kono teaches: wherein the inter-class distance is a Mahalanobis distance, a Euclidean distance, a Manhattan distance, or a cosine similarity (“Euclidean distance” page 3). The motivation to combine is the same as in claims 1 and 4.
Claim Objections and Allowable Subject Matter
Claims 5, 6, and 8-10 are objected to because they depend from rejected claims. These claims would be allowable if re-written in independent form.
Nothing in the prior art teaches or suggests the combinations of limitations presented in claims 5, 6, 8, 9, and 10.
The closest prior art is Wen, Kono, Li, and Abad. Wen adjusts a mixing hyperparameter for each class according to the accuracy of the class and measured confidence on a validation dataset (“We compute the accuracy and confidence on a validation dataset after each training epoch” Wen page 6). Wen’s example is drawn without regard to which class it belongs to (page 3). Wen therefore has no measure between classes, no selection of a second class, and logic for selecting a second class.
Kono calculates a distance between class centroids in latent space and does not interpolate pairs that are too distant (“For this purpose, the interclass distance is introduced so that two classes with a distance greater than a threshold are not interpolated,” page 3). Kono’s distance is not calculated from a likelihood or its variance. Li calculates “the probability of classifying instances of class i as class j: [equation 3]” (page 6), but similarly lacks this distance calculation. Abad determines an accuracy for each class of an image classifier and increases the number of samples for low-accuracy samples ([0056]: “the number of the samples associated with classes that fall below the accuracy threshold value may be increased in the training batch in order to generate an adjusted training batch”), but Abad does not calculate any measure between classes and does not generate training data by mixing the two classes.
Subject Matter Eligibility – 35 U.S.C. 101
Claims 1-12 are not rejected under 101.
While the claims likely recite abstract ideas such as mathematical concepts and mental processes (Step 2A Prong One) such as calculating a distance between classes and comparing values to choose a class, the claim as a whole integrates any abstract idea into a practical application (Step 2A Prong Two) by reciting technical details that improve the operation of the recognition model itself rather than the result that the model produces. The claims do not merely use a trained model as a tool to obtain an outcome. Instead, they recite how the training data is selected and constructed, namely, identifying a class from the model’s own measured recognition accuracy, calculating an inter-class distance from that class to each of the two or more other classes, choosing the second class from that distance, and mixing the image data and the labels of the two chosen classes. This technology “makes it possible to improve the recognition performance of the recognition model by training the recognition model using such image data… if the recognition accuracy for the first class is low, training data can be generated that can effectively improve the recognition accuracy for classes having low recognition accuracy” ([0026] of the specification as published).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Sharma (US-10664722-B1) discloses: “recognition statistics for the lower-ranked classes can be improved by enlarging the training image sets for the lower-ranked classes,” (col. 19 lines 44-46), and “the product depictions in the 500 training images for each class were each segmented from their respective backgrounds, and augmented with different backgrounds—four different backgrounds for each segmented depiction, yielding 2000 additional training images per class,” (col. 19 lines 57-62).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JOSHUA BLANCHETTE whose telephone number is (571)272-2299. The examiner can normally be reached on Monday - Thursday 7:30AM - 6:00PM, EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Shahid Merchant, can be reached on (571) 270-1360. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JOSHUA B BLANCHETTE/Primary Examiner, Art Unit 3624