DETAILED ACTION
This action is responsive to the application filed on 6/10/2024. Claims 1-20 are pending in the case. Claims 1, 8, and 15 are independent claims.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
The application claims the benefit of U.S. Provisional application No. 63/521,581 filed on 6/16/2023. The date used for determinations of prior art is 6/16/2023.
Claim Objections
Claims 5, 12, and 19 are objected to because of the following informalities:
Claims 5, 12, and 19 recite the use of supplemental metrics like diversity and memorization. Given the specification and the examiners best understanding, memorization is synonymous to overfitting in terms of machine learning language. Thereby, to say that an ideal quality metric is not correlated to a degradation in memorization can be misconstrued to mean that a decrease in overfitting is a suboptimal characteristic of both models and evaluation metrics; the examiner comes to interpret this as not being the intended understanding of the limitation. The examiner notes that this is merely an objection, but suggests rewording of the limitation to more clearly indicate what degradation of diversity and memorization entails.
Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed towards an abstract idea without significantly more.
Step 1: Claims 1-7 are directed towards a machine, Claims 8-14 are directed towards a process, and Claims 15-20 are directed towards an article of manufacture. Therefore, Claims 1-20 are directed towards one of the 4 statutory categories: process, machine, article of manufacture, or composition of matter.
With respect to claim 1:
Step 2A Prong 1: The claim is directed towards a judicial exception.
identifying a set of generative models trained to generate data samples based on a reference data set (Mental Process: One could identify a set of generative models trained to generate data samples based on a reference data set, mentally or using pen and paper)
identifying, for each generative model in the set of generative models, a generated data set generated by the generative model (Mental process: One could identify for each generative model in the set of generative models, a generated data set generated by the generative model, mentally or using pen and paper)
determining a plurality of model rankings for a corresponding plurality of candidate model metrics, each comparative model performance ranking describing comparative performance of the set of generative models based on the generated data set of each generative model evaluated by the corresponding candidate model metric (Mental Process: one could determine a plurality of model rankings for a corresponding plurality of candidate model metrices, each ranking describing comparative performance of the set of generative models based on the generated data set of each generative model evaluated by the corresponding candidate model metric, mentally or using pen and paper)
identifying a manual model ranking describing comparative performance of the set of generative models based on the generated data set of each generative model evaluated by human evaluation; and… (Mental Process: One could identify a manual model ranking describing comparative performance of the set of generative models based on the generated data set of each generative model evaluated by human evaluation, mentally or using pen and paper)
selecting an automated quality metric for evaluating model quality from the plurality of candidate model metrics based on a similarity between the manual model ranking and the plurality of model rankings (Mental Process: One could select an automated quality metric for evaluating model quality from the plurality of candidate model metrics based on a similarity between the manual model ranking and the plurality of model rankings, mentally or using pen and paper)
Step 2A Prong 2: The judicial exceptions as a whole are not integrated into a practical application
Additional Elements:
A system comprising: one or more processing elements that executes instructions; and a non-transitory computer-readable medium comprising instructions executable by the processing elements for… (A system comprising a processing element and non-transitory computer-readable medium holding instructions is simply a generic computer with generic components. Adding generic computer components to perform the method is not sufficient. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f))
Step 2B: The claim does not include additional elements that amount to significantly more than the judicial exceptions.
Additional Elements:
A system comprising: one or more processing elements that executes instructions; and a non-transitory computer-readable medium comprising instructions executable by the processing elements for… (A system comprising a processing element and non-transitory computer-readable medium holding instructions is simply a generic computer with generic components. Adding generic computer components to perform the method is not sufficient. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f))
With respect to claim 2:
Step 2A Prong 1: The claim is directed towards a judicial exception, including those inherited from claim 1 via dependency.
applying the automated quality metric to evaluate a first generative model and a second generative model; and… (Mental Process: One could apply the automated quality metric to evaluate a first generative model and a second generative model, mentally or using pen and paper)
selecting a preferred model from the first model and second model based on the evaluation with the automated quality metric (Mental Process: One could select a preferred model from the first model and second model based on the evaluation with the automated quality metric, mentally or using pen and paper)
Step 2A Prong 2: The judicial exceptions as a whole are not integrated into a practical application
Additional Elements:
The system of claim 1, wherein the instructions are further executable for… (See above)
Step 2B: The claim does not include additional elements that amount to significantly more than the judicial exceptions.
Additional Elements:
The system of claim 1, wherein the instructions are further executable for… (See above)
With respect to claim 3:
Step 2A Prong 1: The claim is directed towards a judicial exception, including those inherited from claim 2 via dependency.
…wherein the first generative model and the second generative model are trained on a data set different from the reference data set (Mental Process: One could train the first and second generative model on a data set different from the reference data set by iteratively applying mathematical and mental processes, mentally or using pen and paper)
Step 2A Prong 2: The judicial exceptions as a whole are not integrated into a practical application
Additional Elements:
The system of claim 2… (See above)
Step 2B: The claim does not include additional elements that amount to significantly more than the judicial exceptions.
Additional Elements:
The system of claim 2… (See above)
With respect to claim 4:
Step 2A Prong 1: The claim is directed towards a judicial exception, including those inherited from claim 1 via dependency.
…wherein the similarity between the manual model ranking and the plurality of model rankings is measured by a statistical correlation (Mathematical Concept: Measuring similarity between two rankings using a statistical correlation is a very broadly claimed application of a mathematical concept)
Step 2A Prong 2: The judicial exceptions as a whole are not integrated into a practical application
Additional Elements:
The system of claim 1… (See above)
Step 2B: The claim does not include additional elements that amount to significantly more than the judicial exceptions.
Additional Elements:
The system of claim 1… (See above)
With respect to claim 5:
Step 2A Prong 1: The claim is directed towards a judicial exception, including those inherited from claim 1 via dependency.
evaluating the set of generative models according to a supplemental metric related to at least diversity or memorization (Mental Process: One could evaluate the set of generative models according to a supplemental metric related to at least diversity or memorization, mentally or using pen and paper. Evaluation is a mental process.)
determining whether each candidate model metric is correlated with degradation of the supplemental metric; and… (Mental Process: One could determine whether each candidate model metric is correlated with degradation of the supplemental metric, mentally or using pen and paper. Determining is a mental process.)
wherein selecting the automated quality metric comprises selecting a candidate model metric that is not correlated with degradation of the supplemental metric (Mental Process: One could select the automated quality metric by selecting a candidate model metric that is not correlated with degradation of the supplemental metric, mentally or using pen and paper)
Step 2A Prong 2: The judicial exceptions as a whole are not integrated into a practical application
Additional Elements:
The system of claim 1, wherein the instructions are further executable for… (See above)
Step 2B: The claim does not include additional elements that amount to significantly more than the judicial exceptions.
Additional Elements:
The system of claim 1, wherein the instructions are further executable for… (See above)
With respect to claim 6:
Step 2A Prong 1: The claim is directed towards a judicial exception, including those inherited from claim 1 via dependency.
wherein the human evaluation includes an experimental comparation of the generated data set and the reference data set (Mental Process: One could perform human evaluation including an experimental comparation of the generative data set and the reference data set, mentally or using pen and paper. The examiner notes that the use of a human mind to perform any task is by definition a mental process.)
Step 2A Prong 2: The judicial exceptions as a whole are not integrated into a practical application
Additional Elements:
The system of claim 1… (See above)
Step 2B: The claim does not include additional elements that amount to significantly more than the judicial exceptions.
Additional Elements:
The system of claim 1… (See above)
With respect to claim 7:
Step 2A Prong 1: The claim is directed towards a judicial exception, including those inherited from claim 1 via dependency.
Step 2A Prong 2: The judicial exceptions as a whole are not integrated into a practical application
Additional Elements:
The system of claim 1… (See above)
wherein at least one metric applies an encoding model to the generated samples and applies a scoring function to encoded data samples (Describes the information that the abstract idea operates on rather than an additional element to integrate the exception into a practical application, see MPEP § 2106.05(e), which discusses other meaningful limitations; This merely provides an example of a performance metric that could be used to evaluate the models, which is not a practical application.)
Step 2B: The claim does not include additional elements that amount to significantly more than the judicial exceptions.
Additional Elements:
The system of claim 1… (See above)
wherein at least one metric applies an encoding model to the generated samples and applies a scoring function to encoded data samples (Describes the information that the abstract idea operates on rather than an additional element to integrate the exception into a practical application, see MPEP § 2106.05(e), which discusses other meaningful limitations; This merely provides an example of a performance metric that could be used to evaluate the models, which is not a practical application.)
With respect to claims 8-14:
See the rejections for claims 1-7 above. Note that the primary difference between the two sets of claims is that claims 1-7 are directed towards a system performing the method whereas claims 8-14 are directed towards the method; this method is performed by one or more processors which are generic computer components as per MPEP § 2106.05(f), noted in the rejection for claim 1 above.
With respect to claims 15-20:
See the rejections for claims 1-6 above. Note that the primary difference between the two sets of claims is that claims 1-6 are directed towards a system performing the method whereas claims 15-20 are directed towards an article of manufacture holding instructions for the method to be executed by a processor; see the rejection for claim 1 above which notes how these elements are generic computer components as per MPEP § 2106.05(f).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1, 4, 6-8, 11, 13-15, 18, and 20 are rejected under 35 U.S.C 103 as being unpatentable over Kirstain et. al. (Kirstain, Yuval, et al. Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation, 2 May 2023, arxiv.org/abs/2305.01569v1) in view of Bigaj et al. (US 20210089942 A1).
Regarding claim 1, Kristain et al. teaches a method for:
identifying a set of generative models trained to generate data samples based on a reference data set (5. Model Evaluation, Other Evaluation Metrics, “This set of examples contains images generated by 45 different models - four different backbone models, each with different guidance scales.” The examiner comes to understand that a set of 45 different generative models (4 different backbone models with different guidance scales) is identified. This set of generative models was trained to generate images (data samples) based on a reference data set (prompts from the Pick-a-Pic test set, see Other Evaluation Metrics and Pick-a-Pic dataset sections). The examiner notes that the same process is performed at a smaller scale in the section from Model Evaluation titled “FID,” which evaluates the FID metric that assumes a set of ground truth images unlike other evaluation metrics; this smaller scale method also identifies a set of generative models.)
identifying, for each generative model in the set of generative models, a generated data set generated by the generative model (5. Model Evaluation, Other Evaluation Metrics, “This set of examples contains images generated by 45 different models.” The examiner comes to understand that set of example generated images (a generated data set) was generated (identified) for each of the 45 generative models (set of generative models). See Figure 7 which depicts examples of generated data sets identified for 4 of the 9 models evaluated using the smaller scale procedure in the “FID” section.)
determining a plurality of model rankings for a corresponding plurality of candidate model metrics, each comparative model performance ranking describing comparative performance of the set of generative models based on the generated data set of each generative model evaluated by the corresponding candidate model metric (Figure 8; 5. Model Evaluation, Other Evaluation Metrics, “We then use these real user preferences to calculate Elo ratings [4] for the different models. Afterward, we repeat the process [of ranking] while replacing the real user preferences with CLIP-H[7], and PickScore predictions… we also compare against… ImageReward [18] and HPS[17].” As seen in Figure 8, the examiner comes to understand that a plurality of model ratings (rankings) for each of a plurality of candidate model metrics (CLIP-H, ImageReward, HPS, and PickScore) is determined. Each set of rankings provides a comparable numerical score (Elo Rating) for each model of the set of generative models by evaluating the generated data set of each model using the corresponding candidate metric. Note, all of the metrics listed are known in the art (PickScore being described in this reference) to evaluate generative models using their generated outputs.)
identifying a manual model ranking describing comparative performance of the set of generative models based on the generated data set of each generative model evaluated by human evaluation (Figure 8; 5. Model Evaluation, Other Evaluation Metrics, “We use these real user preferences to calculate Elo ratings [4] for the different models.” Additionally, as seen in 2. Pick-a-Pic Dataset, Model Selection and Evaluation, “by examining user preferences between images generated by different backbone models, we can determine which model is preferred more by users.” The examiner comes to understand that (as in Figure 8) the user preference based Elo ratings assigned to the generative models (manual model rankings) are identified. As seen in the additional citation, these rankings describe comparative performance of the set of generative models based on their generated images (generated data set) that are preferred by users (human evaluation).); and
selecting an automated quality metric for evaluating model quality from the plurality of candidate model metrics based on a similarity between the manual model ranking and the plurality of model rankings (5. Model Evaluation, Other Evaluation Metrics, “Since the Elo rating system is iterative by nature, we repeat the process 50 times. Each time we randomly shuffle the order of the examples and calculate the correlation with human ratings. Finally, for each metric, we output the mean and standard deviation of its 50 corresponding correlations. Figure 8 displays the correlation between real users’ Elo ratings with different metrics’ ratings, showing that PickScore exhibits a stronger correlation (0.790 ± 0.054) with real users than all other automatic metrics.” 5. Model Evaluation, “[We] show that when evaluating state-of-the-art text-to-image models, PickScore is more aligned with human judgements than other automatic models.” The examiner comes to understand that PickScore is selected as the best automated quality metric for evaluating model quality from the plurality of candidate model metrics (CLIP-H, ImageReward, HPS, and PickScore) based on a correlation (similarity) between the human ratings (manual model rankings) and the plurality of model rankings (model Elo ratings from the metrics).)
Kirstain et al. does not distinctly disclose:
A system comprising: one or more processing elements that executes instructions; and a non-transitory computer-readable medium comprising instructions executable by the processing elements for…
However, Bigaj et al. teaches these limitations:
A system comprising: one or more processing elements that executes instructions; and a non-transitory computer-readable medium comprising instructions executable by the processing elements for… (Paragraph [0076], “Embodiments of the invention may be implemented together with virtually any type of computer.” Paragraph [0078], “The components of computer system/server 500 may include, but are not limited to, one or more processors or processing units 502, a system memory 504… Computer system/server 500 typically includes a variety of computer system readable media. Such media may be any available media that is accessible by computer system/server 500, and it includes both, volatile and non-volatile media, removable and non-removable media.” Paragraph [0084], “The present invention may be embodied as a system, a method, and/or a computer program product. The computer program product may include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present invention.” The examiner notes that these citations provide examples of the systems or articles of manufacture that can be used to implement any embodiment of the claimed invention. This reference relates to systems and methods for determining an evaluation metric of a machine learning model, but merely serves as an example of the many potential references that provide the structure for the additional embodiments of the claimed invention. It would take no inventive effort to use the cited systems and computer components to perform the computer reliant method claimed in Kirstain et al.)
Before the effective filing date of the claimed invention it would have been obvious to one of ordinary skill in the art to combine the model evaluation metric selection technique of Kirstain et al. (including identifying a set of generative models and their generated outputs, determining a set of model rankings for the models based on a plurality of candidate metrics, determining a set of manual model rankings for the models based on human evaluation, and determining an automated quality metric by comparing the similarity between the two sets of rankings) with the systems and computer components for determining an evaluation metric of a machine learning model in Bigaj et al. in order to perform the claimed method in an efficient and accurate environment, such as the environment provided by those computing devices.
Regarding claim 4, Kirstain et al. as modified by Bigaj et al. teaches all of the limitations of the system of claim 1 as cited above and Kirstain et al. further teaches:
wherein the similarity between the manual model ranking and the plurality of model rankings is measured by a statistical correlation (5. Model Evaluation, Other Evaluation Metrics, “Each time we randomly shuffle the order of the examples and calculate the correlation with human ratings. Finally, for each metric, we output the mean and standard deviation of its 50 corresponding correlations.” The examiner comes to understand that the correlation (similarity) determined between the human ratings (manual model rankings) and the per metric ratings (plurality of model rankings) is measured using mean and standard deviation of the correlations. These are understood to be measures of statistical correlation.)
Regarding claim 6, Kirstain et al. as modified by Bigaj et al. teaches all of the limitations of the system of claim 1 as cited above and Kirstain et al. further teaches:
wherein the human evaluation includes an experimental comparation of the generated data set and the reference data set (2. Pick-a-Pic Dataset, The Pick-a-Pic Web App, “At each turn, the user is presented with two generated images (conditioned on their prompt), and asked to select their preferred option or indicate a tie if they have no strong preference. Upon selection, the rejected (non-preferred) image is replaced with a newly generated image, and the process repeats. The user can also clear or edit the prompt at any time, and the app will generate new images appropriately.” The examiner comes to understand that the above provides for an example of how the human evaluation used is structured. This human evaluation is used to determine user preferences which are then used to calculate Elo ratings for ranking the models. Additionally, the examiner comes to understand that this human evaluation includes an experimental comparation in which the user compares sets of generated images (generated data sets) to their prompt and selects the generated image that they prefer. Note that these prompts and corresponding generated images are included in the reference data set and generated data set, respectively, used in the rejection for claim 1 above; 5. Model Evaluation, Other Evaluation Metrics, “we take all the 14,000 collected preferences that correspond to prompts from the Pick-a-Pic test set. This set of examples contains images generated by 45 different models…”)
Regarding claim 7, Kirstain et al. as modified by Bigaj et al. teaches all of the limitations of the system of claim 1 as cited above and Kirstain et al. further teaches:
wherein at least one metric applies an encoding model to the generated samples and applies a scoring function to encoded data samples (5. Model Evaluation, Other Evaluation Metrics, “Afterward, we repeat the process while replacing the real user preferences with CLIP-H [7], and PickScore predictions… we also compare against… ImageReward [18] and HPS [17].” The examiner comes to understand that all of the metrics listed in the plurality of candidate model metrics apply an encoding model to the generated samples and apply a scoring function to the encoded samples. Using PickScore as an example, see section 3, PickScore: Model, Objective, and Training. This section describes how the PickScore metric follows a CLIP architecture which includes an encoder (as described in paragraph [0025] of the present application) and a scoring function. The examiner comes to understand that the other metrics compared are known in the art to also comprise an encoder and some sort of scoring function.)
Regarding claim 8, see the rejection for claim 1 above. Note that the primary difference between claim 1 and claim 8 is that claim 1 is directed towards a machine performing the method, whereas claim 8 is directed towards the method. The citations used in the rejection of claim 1 cover the additional embodiment as a method.
Regarding claim 11, see the rejection for claim 4 above. Note that the primary difference between claim 4 and claim 11 is that claim 4 is directed towards a machine performing the method, whereas claim 11 is directed towards the method. The citations used in the rejection of claim 1 cover the additional embodiment as a method.
Regarding claim 13, see the rejection for claim 6 above. Note that the primary difference between claim 6 and claim 13 is that claim 6 is directed towards a machine performing the method, whereas claim 13 is directed towards the method. The citations used in the rejection of claim 1 cover the additional embodiment as a method.
Regarding claim 14, see the rejection for claim 7 above. Note that the primary difference between claim 7 and claim 14 is that claim 7 is directed towards a machine performing the method, whereas claim 14 is directed towards the method. The citations used in the rejection of claim 1 cover the additional embodiment as a method.
Regarding claim 15, see the rejection for claim 1 above. Note that the primary difference between claim 1 and claim 15 is that claim 1 is directed towards a machine performing the method, whereas claim 15 is directed an article of manufacture holding instructions for the method. The citations used in the rejection of claim 1 cover the additional embodiment as an article of manufacture holding instructions for a method.
Regarding claim 18, see the rejection for claim 4 above. Note that the primary difference between claim 4 and claim 18 is that claim 4 is directed towards a machine performing the method, whereas claim 18 is directed an article of manufacture holding instructions for the method. The citations used in the rejection of claim 1 cover the additional embodiment as an article of manufacture holding instructions for a method.
Regarding claim 20, see the rejection for claim 6 above. Note that the primary difference between claim 6 and claim 20 is that claim 6 is directed towards a machine performing the method, whereas claim 20 is directed an article of manufacture holding instructions for the method. The citations used in the rejection of claim 1 cover the additional embodiment as an article of manufacture holding instructions for a method.
Claims 2-3, 9-10, and 16-17 are rejected under 35 U.S.C 103 as being unpatentable over Kirstain et. al. (Kirstain, Yuval, et al. Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation, 2 May 2023, arxiv.org/abs/2305.01569v1) in view of Bigaj et al. (US 20210089942 A1), further in view of Mitra et al. (US 20230075453 A1).
Regarding claim 2, Kirstain et al. as modified by Bigaj et al. teaches all of the limitations of the system of claim 1 as cited above and Kirstain et al. further teaches:
applying the automated quality metric to evaluate a first generative model and a second generative model (5. Model Evaluation, FID, “we generate images from 9 different models based on the same set of prompts… Specifically, we use Stable Diffusion 1.5, Stable Diffusion 2.1, and Dreamlike Photoreal 2.0 combined with three different classifier-free guidance scales (3, 6, and 9)… repeat the labeling process (over the same images) using FID, and PickScore instead of humans to determine preferences.” The examiner comes to understand that in an experiment different from the one cited above, Kirstain et al. takes a set of 9 generative models and applies the PickScore metric (selected automated quality metric, determined above), as well as FID, to separately rank (evaluate) the models (See Figure 6). Note that the set of 9 different models is understood to include generative models, comprising a first and second generative model to which the PickScore metric is applied for evaluation.)
Kirstain et al. as modified by Bigaj et al. does not distinctly* disclose:
selecting a preferred model from the first model and second model based on the evaluation with the automated quality metric.
(* The examiner notes that Kirstain et al. still provides a visual representation of preference in the models as seen in Figures 6 and 7, but does not distinctly include the exact language for selecting a preferred model from the first model and second model based on the evaluation with the automated quality metric.)
However, Mitra et al. teaches this limitation:
selecting a preferred model from the first model and second model based on the evaluation with the automated quality metric (Paragraph [0005], “For each of the plurality of machine learning based models the system… determines the value of the model metric for the machine learning based model. The system selects a machine learning based model based on comparison of values of the model metric for the different machine learning based models.” The examiner comes to understand that for a plurality of machine learning models, a quality metric is applied to evaluate each of the machine learning models. Subsequently, the system selects a preferred machine learning model based on these evaluations using the quality metric. The examiner notes that it would take no inventive effort to apply this to the above citation in order to select a preferred model from a first and second model of the 9 different models. In this application, the model metric used for evaluation could be the automated quality metric, PickScore. See paragraph [0007] which exemplifies how the model metric used can be chosen depending on use case of the application, in this case use being to evaluate generative image models, i.e. selecting PickScore which can achieve that use.)
Before the effective filing date of the claimed invention it would have been obvious to one or ordinary skill in the art to combine the model evaluation metric selection technique of Kirstain et al. (including identifying a set of generative models and their generated outputs, determining a set of model rankings for the models based on a plurality of candidate metrics, determining a set of manual model rankings for the models based on human evaluation, and determining an automated quality metric by comparing the similarity between the two sets of rankings) as modified by the systems and structures for performing a method of Bigaj et al. with the technique for selecting models using a determined metric of Mitra et al. in order to improve the efficiency of selecting a model via evaluation based on one metric (Mitra, Paragraph [0068], “For example, a system that evaluates 5 different metrics for various models is likely to take five times the effort and resources compared to the system according to various embodiments as disclosed.”) and allow for selection of a single best performing model based on a determined quality metric. (Mitra, Paragraph [0069], “if the user specifies a specific model metric, for example, based on the use case, the system does not have to select across different models. For one model metric, the system determines a single one top machine learning model.”)
Regarding claim 3, Kirstain et al. as modified by Bigaj et al. and further modified by Mitra et al. teaches all of the limitations of the system of claim 2 as cited above and Kirstain et al. further teaches:
wherein the first generative model and the second generative model are trained on a data set different from the reference data set (5. Model Evaluation, FID, “To provide the most convenient settings for the FID metric we use MS-COCO captions, selecting 100 random captions from the dataset’s test split.” The examiner comes to understand that the 9 different generative models are trained using the MS-COCO data set which is different than the Pick-a-Pic data set (reference data set) used in the other experiment cited above for the generative model rankings. The examiner notes that the additional elements and modifications of this experiment to evaluate using FID in addition to PickScore do not change the fact that PickScore is applied to a first and second generative model (trained on a dataset different from the reference dataset) to select a preferred model for claims 2 or 3.)
Regarding claim 9, see the rejection for claim 2 above. Note that the primary difference between claim 2 and claim 9 is that claim 2 is directed towards a machine performing the method, whereas claim 9 is directed towards the method. The citations used in the rejection of claim 1 cover the additional embodiment as a method.
Regarding claim 10, see the rejection for claim 3 above. Note that the primary difference between claim 3 and claim 10 is that claim 3 is directed towards a machine performing the method, whereas claim 10 is directed towards the method. The citations used in the rejection of claim 1 cover the additional embodiment as a method.
Regarding claim 16, see the rejection for claim 2 above. Note that the primary difference between claim 2 and claim 16 is that claim 2 is directed towards a machine performing the method, whereas claim 16 is directed an article of manufacture holding instructions for the method. The citations used in the rejection of claim 1 cover the additional embodiment as an article of manufacture holding instructions for a method.
Regarding claim 17, see the rejection for claim 3 above. Note that the primary difference between claim 3 and claim 17 is that claim 3 is directed towards a machine performing the method, whereas claim 17 is directed an article of manufacture holding instructions for the method. The citations used in the rejection of claim 1 cover the additional embodiment as an article of manufacture holding instructions for a method.
Claims 5, 12, and 19 are rejected under 35 U.S.C 103 as being unpatentable over Kirstain et. al. (Kirstain, Yuval, et al. Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation, 2 May 2023, arxiv.org/abs/2305.01569v1) in view of Bigaj et al. (US 20210089942 A1), further in view of Jiralerspong et al. (Jiralerspong, Marco, et al. Feature Likelihood Score: Evaluating Generalization of Generative Models Using Samples, 29 May 2023, arxiv.org/abs/2302.04440v2.)
Regarding claim 5, Kirstain et al. as modified by Bigaj et al. teaches all of the limitations of the system of claim 1 as cited above and Kirstain et al. teaches:
…the set of generative models… and …the candidate model metric[s]… (See the citations in the rejection for claim 1 above which identify both of these elements.)
Kirstain et al. as modified by Bigaj et al. does not distinctly disclose:
evaluating… generative models according to a supplemental metric related to at least diversity or memorization;
determining whether each… metric is correlated with degradation of the supplemental metric; and
wherein selecting the automated quality metric comprises selecting a… metric that is not correlated with degradation of the supplemental metric
However, Jiralerspong et al. teaches these limitations:
evaluating… generative models according to a supplemental metric related to at least diversity or memorization (Table 1; 4. Experiments, 4.4 Comparison of Evaluation of State-of-the-art Models, “We perform a large-scale evaluation of various generative models in Tab. 1 using different metrics on CIFAR10… Interestingly, we observe that modern generative models art overfitting in a benign way…” The examiner notes that as seen in Table 1, the generative models are evaluated on their overfitting (R (-50%)) as described in Appendix A.1. The examiner notes that it would take no inventive effort to perform this method on the set of generative models identified in the rejection for claim 1.);
determining whether each… metric is correlated with degradation of the supplemental metric (4. Experiments, 4.3 Detecting Overfitting and Memorization with FLS, “We find that FLS generally decreases as more copies are added whereas FID improves for mild transformations and worsens for heavier transformations (due to the decrease in quality). Moreover, even for less drastic transformations, FLS detects almost every copy as the percentage of overfit Gaussians scales linearly with the number of copies. For these transformations, the decrease in FLS is driven by the increase in overfitting,” 4. Experiments, 4.6 Difference and Correlation with FID, “We now investigate the relationship between FLS and FID, performing a large-scale comparison of the same set of models and datasets… We find evidence of a strong correlation between FLS and FID for all datasets, which highlights that FLS captures sample quality/diversity in a similar vein to FID… We see a clear pattern amongst different models in that FLS requires significantly fewer samples to report a score that remains consistent across increasing test set sizes, while FID requires a large amounts of sample to report a stable score.” Additionally, see figure 10. The examiner comes to understand that FLS is determined to not be correlated with degradations in either supplement metric of diversity or memorization, i.e. the FLS metric does not increase as memorization/overfitting in a model occurs OR fluctuate with degradation in diversity. On the other hand, it is indicated that FID is correlated with degradation in either or both supplemental metrics. The examiner notes that it would take no inventive effort to apply this analysis to the set of candidate model metrics, as identified in the rejection for claim 1 above, to determine their correlation to the degradation of these supplemental metrics.); and
wherein selecting the automated quality metric comprises selecting a… metric that is not correlated with degradation of the supplemental metric (4. Experiments, 4.5 Specific Application of FLS, “We now apply FLS to a variety of application to demonstrate utility and flexibility.” Also, see 5. Conclusion, “…unlike previous approaches, FLS provides more explainable insights into the overfitting and memorization behavior of trained generative models. We empirically demonstrate both on synthetic and real-world datasets that FLS can diagnose important failure modes such as memorization/overfitting.” The examiner comes to understand that FLS is determined as the best quality metric for evaluating models following comparison of the potential evaluation metrics to the supplemental metrics (diversity and/or memorization/overfitting). The examiner notes that it would take no inventive effort to apply this analysis technique to the selection of the automated quality metric, as identified in the rejection for claim 1, by selecting a candidate model metric, also identified in the rejection for claim 1, that is not correlated with the degradation of the supplemental metric.)
Before the effective filing date of the claimed invention it would have been obvious to one or ordinary skill in the art to combine the model evaluation metric selection technique of Kirstain et al. (including identifying a set of generative models and their generated outputs, determining a set of model rankings for the models based on a plurality of candidate metrics, determining a set of manual model rankings for the models based on human evaluation, and determining an automated quality metric by comparing the similarity between the two sets of rankings) as modified by the systems and structures for performing a method of Bigaj et al. with the technique for determining a quality metric not correlated with degradations in diversity or memorization of Jiralerspong et al. in order to validate the connection of the selected quality metric to ideal model criteria (4. Experiments, “Through our experiments, we seek to validate the correlation between FLS and sample fidelity, diversity, and novelty.” 5. Conclusion, “FLS is easy to compute, broadly applicable to all generative models, and evaluates generation quality, diversity, and generalization.”) and determine the best metric for evaluating generative models. (Abstract, “We empirically demonstrate the ability of FLS to identify specific overfitting problem cases, where previously proposed metrics fail.”)
Regarding claim 12, see the rejection for claim 5 above. Note that the primary difference between claim 5 and claim 12 is that claim 5 is directed towards a machine performing the method, whereas claim 12 is directed towards the method. The citations used in the rejection of claim 1 cover the additional embodiment as a method.
Regarding claim 19, see the rejection for claim 5 above. Note that the primary difference between claim 5 and claim 19 is that claim 5 is directed towards a machine performing the method, whereas claim 19 is directed an article of manufacture holding instructions for the method. The citations used in the rejection of claim 1 cover the additional embodiment as an article of manufacture holding instructions for a method.
Citation of Pertinent Prior Art
The prior art made of record and not relied upon is considered pertinent to applicant's
disclosure. Xu et al. teaches an ImageReward metric for evaluating generative models and follows a very similar procedure to the one detailed in Kirstain et al. This procedure comprises evaluating models using a variety of candidate metrics and then ranking the outputs of the model to determine the best one. A process of human evaluation on the same models and outputs is also performed to demonstrate that ImageReward is the optimal metric for evaluating generative models. The candidate model metrics include those that make use of encoding models like CLIP. Alaa et al. teaches a method for ranking generative models using different evaluation metrics as well as using a supplemental metric of diversity or memorization to compare the evaluation metrics. Hashimoto et al. proposes an evaluation metric called HUSE that combines statistical evaluation and human evaluation of different machine learning models. They also conduct experiments using the metric to determine its performance in a variety of testing environments with variations in diversity of the test sets. Kocmi et al. discloses a machine translation method focused towards finding an automatic metric for evaluating models which is best suited for ranking systems and measuring how well we can rely on the metric chosen. Hodosh et al. teaches a study on associating images with natural language translations and includes a comparison of human and automatic evaluation metrics. Sajjadi et al. discloses experiments relating to the evaluation of generative models. These experiments teach that certain metrics are not sufficient to capture different failure cases of state-of-the-art generative models, like overfitting, diversity, and memorization. Some of the experiments additionally compare metrics to one another based on their correlations to degradation of supplemental metrics like diversity and overfitting. Naeem et al. discusses comparison of different metrics for evaluating generative models. The comparison of these metrics focuses on evaluating them in comparison to the supplemental metrics of fidelity and diversity. Friedman et al. teaches a generative model evaluation metric called Vendi Score. Experiments of the proposed metric include comparisons to other evaluation metrics as well as comparisons to model diversity. Kynkäänniemi et al. teaches new evaluation metrics for evaluating generative models and draws comparisons between their proposed metrics and existing metrics. These comparisons include experiments relating to diversity of the models being tested and the correlation of metrics to changes in diversity.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Zane A Rawlings whose telephone number is (571)270-3372. The examiner can normally be reached M-F, 8am to 5pm ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571) 270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Z.A.R./Examiner, Art Unit 2123
/ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123