DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Drawings
The drawings are objected to as failing to comply with 37 CFR 1.84(p)(5) because they include the following reference character(s) not mentioned in the description: 1418.
Corrected drawing sheets in compliance with 37 CFR 1.121(d), or amendment to the specification to add the reference character(s) in the description in compliance with 37 CFR 1.121(b) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
The drawings are objected to as failing to comply with 37 CFR 1.84(p)(5) because they do not include the following reference sign(s) mentioned in the description: 1416.
Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-2, 5-10, 13-16, and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Li (US 20230237772 A1) in view of Ye (US 20250054322 A1).
Regarding Claim 1, Li discloses “A method of filtering image-text data, comprising:” (Li, Paragraph [0012], discloses “…a method of pre-training a MED (Multi-modal encoder-decoder) model for downstream tasks and dataset bootstrapping…”; Li, Paragraph [0020], discloses “A filter (e.g., a pre-trained image-grounded text encoder) may be finetuned using the small set of human annotated image-text pairs based on image-text contrastive loss and image-text matching loss. Then the captioner is used to generate a caption for an image from the noisy training data, and the filter is used to filter original noisy captions and/or the generated captions from the noisy training data. The resulting filtered images and texts can then form a dataset for pre-training any new vision-language models. The captioner and the filter work together to achieve substantial performance improvement on various downstream tasks by bootstrapping the captions.”); (Li, Paragraph [0019], discloses “In view of the need for a unified VLP framework to learn from noisy image-text pairs, a multimodal mixture of encoder-decoder (MED) architecture is used for effective multi-task pre-training and flexible transfer learning. Specifically, the MED can operate either as a text-only encoder, or an image-grounded text encoder, or an image-grounded text decoder. Thus, the model is jointly pre-trained with three objectives: image-text contrastive learning, image-text matching, and language modeling using even very noisy image-text training data (e.g., from the web). In this way, the multiple training objectives help to enhance the model's ability to learn image-text matching.”; Li, Paragraph [0020], discloses “In another embodiment, a two-model mechanism is provided to improve the quality of noisy image-text training data. A captioner (e.g., a pre-trained image-grounded text decoder) may be finetuned using a small set of human annotated image-text pairs based on language modeling loss. A filter (e.g., a pre-trained image-grounded text encoder) may be finetuned using the small set of human annotated image-text pairs based on image-text contrastive loss and image-text matching loss. Then the captioner is used to generate a caption for an image from the noisy training data, and the filter is used to filter original noisy captions and/or the generated captions from the noisy training data. The resulting filtered images and texts can then form a dataset for pre-training any new vision-language models. The captioner and the filter work together to achieve substantial performance improvement on various downstream tasks by bootstrapping the captions.”); (Li, Paragraphs [0035]-[0037] and Figure 4, disclose the following:
PNG
media_image1.png
251
518
media_image1.png
Greyscale
PNG
media_image2.png
493
695
media_image2.png
Greyscale
It is important to note that for Paragraphs [0035]-[0037], that the filter removes the noisy image-text pairs to make room for the high quality image-text pairs. The image-text pairs in the initial data set are also removed based off of their irrelevance of what the context of the image is depicting for a specific object, making this similar to the image-text matching metric.). Li does not explicitly disclose “constructing instruction data on a plurality of image-text pair quality scoring tasks”, or “evaluating a quality of each image-text pair from a dataset by the fine-tuned machine learning model using a plurality of metrics, wherein the plurality of metrics comprise an Image-Text Matching (ITM) metric, an Object Detail Fulfillment (ODF) metric, and a Caption Text Quality (CTQ) metric”. However, in an analogous field of endeavor, Ye discloses that “In some implementations, the task can be an instruction following task. Machine-learned model(s) I can be configured to process input(s) 2 that represent instructions to perform a function and to generate output(s) 3 that advance a goal of satisfying the instruction function (e.g., at least a step of a multi-step procedure to perform the function). Output(s) 3 can represent data of the same or of a different modality as input(s) 2. For instance, input(s) 2 can represent textual data (e.g., natural language instructions for a task to be performed) and machine-learned model(s) I can process input(s) 2 to generate output(s) 3 that represent textual data responsive to the instructions (e.g., natural language responses, programming language responses, machine language responses, etc.). Input(s) 2 can represent image data (e.g., image-based instructions for a task to be performed, optionally accompanied by textual instructions) and machine-learned model(s) 1 can process input(s) 2 to generate output(s) 3 that represent textual data responsive to the instructions (e.g., natural language responses, programming language responses, machine language responses, etc.). One or more output(s) 3 can be iteratively or recursively generated to sequentially process and accomplish steps toward accomplishing the requested functionality. For instance, an initial output can be executed by an external system or be processed by machine-learned model(s) 1 to complete an initial step of performing a function. Multiple steps can be performed, with a final output being obtained that is responsive to the initial instructions.” (Ye, Paragraph [0307]). Here, we can see that the inputs (textual data, image data) are being fed into a machine learning model to be used to develop outputs (textual data, image data) based on a series of tasks that creates a training dataset to be used for different steps in filtering image-text data or finding high-quality image pairs. Ye also discloses “In some implementations, the above training loop can be implemented for fine-tuning a machine-learned model. Fine-tuning can include, for instance, smaller-scale training on higher-quality (e.g., labeled, curated, etc.) data. Fine-tuning can affect all or a portion of the parameters of a machine-learned model” (Ye, Paragraph [0068]). Ye discloses each of the plurality of metrics mentioned in Claim 1 in detail, which will be described in the following paragraph.
For the “Image-Text Matching (ITM) metric”, Ye discloses “The method can include obtaining, by a computing system including one or more processors, image data and text data. The image data can be descriptive of one or more objects. The text data can be descriptive of a particular object associated with the image data. The method can include processing, by the computing system, the text data with a language model to determine a plurality of candidate attributes. The plurality of candidate attributes can include attributes predicted to be candidate terms that describe attributes of the particular object. The method can include processing, by the computing system, the image data, text data, and candidate attribute with a pre-trained image-text model to determine a probability score for the candidate attribute for each of the plurality of candidate attributes. The probability score can be descriptive of a likelihood the candidate attribute is associated with the image data. The method can include determining, by the computing system, a particular attribute of the plurality of candidate attributes is associated with the particular object depicted in the image data based on the plurality of probability scores” (Li, Paragraph [0006]) and “the model is only trained to match image-text pairs” (Li, Paragraph [0174]). From the above paragraphs, it is evident that the text is being compared to distinct features that are present within the image, ultimately allowing each feature to be matched independently to generate a probability score (a scoring task and a metric) to determine whether the text and image data align with one another. For “an Object Detail Fulfillment (ODF) metric”, Ye discloses “The systems and methods can include determining a particular attribute of the plurality of candidate attributes is associated with a particular object depicted in the image data based on the plurality of probability scores. The particular attribute and the text data can be processed to generate a caption that can then be displayed and/or stored. The particular attribute may be a determined characteristic and/or property associated with the particular object” (Ye, Paragraph [0044]). Paragraph [0044] further goes in depth of how the probability scores are used to determine whether an attribute or property (from the associated text) is associated closely to the object in the image, with the metric being the “plurality of probability scores”, which may seem as though it is apart of the ITM metric, but each of these probability scores combined associate with the features for a specific object within the image. For “a Caption Text Quality metric”, Ye discloses that “Output sequence 7 can include or otherwise represent the same or different data types as input sequence 5. For instance, input sequence 5 can represent textual data, and output sequence 7 can represent textual data” (Ye, Paragraph [0244]) and “Output sequence 7 can be generated autoregressively. For instance, for some applications, an output of one or more prediction layer(s) 6 can be passed through one or more output layers (e.g., softmax layer) to obtain a probability distribution over an output vocabulary (e.g., a textual or symbolic vocabulary) conditioned on a set of input elements in a context window” (Ye, Paragraph [0246]. The two aforementioned paragraphs demonstrate how the quality of text can be evaluated based off of the output vocabulary, which is important for grammatical correctness and, fluency, length, and structure of the text, as all of it goes hand in hand to determine how well the caption will read, and if it needs to be either filtered or adjusted to be used with an image-text pair. As we can see in Paragraph [0246], the metric or scoring task seen is the “probability distribution” across the output vocabulary, which is taken from the input elements that were provided to the model.
Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the techniques of filtering image-text data, fine-tuning of machine models, and selecting high quality image-text pairs seen in Li with the techniques of constructing instruction data and evaluating image-text pairs using a series of image quality metrics seen in Ye to achieve a complete filtering method for filtering image-text data. By using the Li techniques for filtering image text data for high-quality image-text pairs and fine-tuning a machine learning model and combining it with the Ye image quality metrics and instruction data acquisition, one of ordinary skill in the art can effectively associate text captions for features or objects in an image accurately without having unpredictable results. The metrics (or scoring tasks) that Ye provide in the combination also allows for one of ordinary skill in the art to have an appropriate quality scale to determine whether an image-text pair fits the context of the image. Therefore, it would have been obvious for one of ordinary skill in the art to combine the Li and Ye references to achieve the same system described in Claim 1.
Regarding Claim 2, the combination of Li and Ye discloses “The method of claim 1, further comprising:” (Please refer to the above-described analysis for Claim 1); “training another machine learning model on the selected high-quality image-text pairs to improve a performance of the other machine learning model” (Li, Paragraph [0037], discloses the following:
PNG
media_image1.png
251
518
media_image1.png
Greyscale
). Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to use the technique of using the high-quality text pairs to improve the performance of another machine learning model as seen in the combination of Li and Ye to improve the image-filtering method in the same way. By using the high-quality image-text pairs on another the machine learning model (such as the captioner), one of ordinary skill in the art can increase the accuracy of the model by determining whether the image-text pair in future data fits what is going on within a particular image. Therefore, it would have been obvious for one of ordinary skill in the art to use the Li and Ye references to achieve the same method described in Claim 2.
Regarding Claim 5, the combination of Li and Ye discloses “The method of claim 1, further comprising:” (Please refer to the above-described analysis regarding Claim 1) “generating a balanced instruction dataset by sampling the instruction data” (Li, Paragraphs [0059]-[0061] discloses the following:
PNG
media_image3.png
689
517
media_image3.png
Greyscale
In the above paragraphs, we can see that the image-grounded text encoder (or captioner) generates a dataset using the second training datasets by calculating the loss between the positive and negative pairs of image-text. These positive and negative pairs offer a diversity of quality, which are used to for the image-text contrastive loss to then fine-tune the model further); “wherein the balanced instruction dataset comprises image-text pairs with diverse quality levels; generating a mixed instruction dataset by mixing the balanced instruction dataset with other instruction datasets corresponding to other vision-language tasks; and fine-tuning the machine learning model using the mixed instruction dataset” (Li, Paragraphs [0062]-[0066], discloses the following:
PNG
media_image4.png
549
514
media_image4.png
Greyscale
In this next set of paragraphs, the mixed instruction dataset is clearly described when the filtered image text-pairs from the first dataset is combined with the second dataset to create the third dataset. Although the fine tuning comes before the second dataset is added, it still allows for the dataset to be adjusted appropriately to train other machine learning models for vision-language related tasks). Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to use the technique of generating a balanced instruction dataset and mixed instruction dataset to fine-tune the machine learning model seen in the combination of Li and Ye to improve the image-text filtering method in the same way. By using the technique of generating the balanced and mixed datasets, one of ordinary skill in the art does not limit the model to focus just on the training data, but a mix of data to gain a holistic understanding of what is happening in multiple images. This way, when the model is used in different scenarios, it can develop stronger associations with the text that pertain to various features of the image to make informed decisions on which image-text pairs align appropriately with the image and captions provided. Therefore, it would have been obvious for one of ordinary skill in the art to use the Li and Ye references to achieve the same image-text filtering method seen in Claim 5.
Regarding Claim 6, the combination of Li and Ye discloses “The method of claim 1, wherein the ITM metric is configured to evaluate whether a text in an image-text pair accurately represents primary features of an image in the image-text pair, wherein the ODF metric is configured to evaluate whether the text depicts detailed properties of objects in the image, and wherein the CTQ is configured to evaluate a quality of the text based on a grammatical correctness, diversity of vocabulary, fluency, readability, length, and structure of the text” (Ye, Paragraphs [0006], [0174], [0044], [0244], [0246]; For the sake of brevity, the descriptions and explanations for each of these paragraphs have been omitted as they have been described before in Claim 1. For more information on how these paragraphs correlate with each of the three metrics described in this limitation, please refer to the above-described analysis for Claim 1). Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed inventions to use the ITM, ODF, and CTQ metrics seen in the combination of Li and Ye to improve the image-filtering method in the same way. Each of these metrics seen in the combination of Li and Ye offer a quantitative and quantitative way to assess the quality of the image-text pairs and how they relate to the context of the image. Therefore, it would have been obvious for one of ordinary skill in the art to use the Li and Ye references to achieve the same image-text filtering method described in Claim 6.
Regarding Claim 7, the combination of Li and Ye discloses “The method of claim 1, wherein the evaluating a quality of each image-text pair from a dataset by the fine-tuned machine learning model comprises:” (Please refer to the above-described analysis for Claim 1) “generating a first score indicative of an ITM quality level of each image-text pair” (Ye, Paragraph [0006] discloses: “The probability score can be descriptive of a likelihood the candidate attribute is associated with the image data.”; Here, the candidate attribute is referring to the feature defined in the text, meaning the score computes whether the text matches with the actual characteristics of the image); “generating a second score indicative of an ODF quality level of each image-text pair” (Ye, Paragraph [0006], discloses: “The method can include determining, by the computing system, a particular attribute of the plurality of candidate attributes is associated with the particular object depicted in the image data based on the plurality of probability scores.”; As described before in the analysis of Claim 1, the “plurality of probability scores” are used together to determine whether the image-text pair aligns with specific features of an object seen within the observed image(s)); and generating a third score indicative of a CTQ quality level of each image-text pair” (Ye, Paragraph [0044] discloses “The particular attribute and the text data can be processed to generate a caption that can then be displayed and/or stored. The particular attribute may be a determined characteristic and/or property associated with the particular object.”; Paragraph [0244] discloses: “Output sequence 7 can include or otherwise represent the same or different data types as input sequence 5. For instance, input sequence 5 can represent textual data, and output sequence 7 can represent textual data.”; Paragraph [0246] discloses: “Output sequence 7 can be generated autoregressively. For instance, for some applications, an output of one or more prediction layer(s) 6 can be passed through one or more output layers (e.g., softmax layer) to obtain a probability distribution over an output vocabulary (e.g., a textual or symbolic vocabulary) conditioned on a set of input elements in a context window.”; From the analysis of Claim 1, the “probability distribution” represents the scoring that represents the CTQ metric). Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to use the first, second, and third scores for the ITM, ODF, and CTQ metrics to improve the image-filtering method to have improved quality scoring for the instruction data used.
Regarding Claim 8, the combination of Li and Ye discloses “The method of claim 1, wherein the selecting high-quality image-text pairs from the dataset based on one or more of the plurality of metrics comprises:” (Please see the above-described analysis for Claim 1); “selecting the high-quality image-text pairs from the dataset based on at least one of the first score, a second score, or a third score of each image-text pair” (Ye, Paragraph [0185], discloses the following:
PNG
media_image5.png
312
517
media_image5.png
Greyscale
). Here, we can see that the ITM metric is clearly shown as they calculate the loss between the image and text, whereas the loss calculation is analogous to a first score as that will further be computed to determine whether the image-text pair is of high-quality or not in this contrastive prompting example. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to use the selection of high-quality pairs seen in the combination of Li and Ye to improve the image-filtering method in the same way.
Claim 9 recites a system with elements corresponding to the steps recited in Claim 1. Therefore, the recited elements of this claim are mapped to the proposed combination in the same manner as the corresponding steps in its corresponding method claim. Additionally, the rationale and motivation to combine the Li and Ye references, presented in rejection of 1, apply to this claim. Finally, the combination of Li and Ye references discloses a processor and a memory (for example, see Yang, Paragraphs [0040] and [0043]).
Claim 10 recites a system with elements corresponding to the steps recited in Claim 2. Therefore, the recited elements of this claim are mapped to the proposed combination in the same manner as the corresponding steps in its corresponding method claim. Additionally, the rationale and motivation to combine the Li and Ye references, presented in rejection of Claim 2, apply to this claim.
Claim 13 recites a system with elements corresponding to the steps recited in Claim 5. Therefore, the recited elements of this claim are mapped to the proposed combination in the same manner as the corresponding steps in its corresponding method claim. Additionally, the rationale and motivation to combine the Li and Ye references, presented in rejection of Claim 5, apply to this claim.
Claim 14 recites a system with elements corresponding to the steps recited in Claim 7. Therefore, the recited elements of this claim are mapped to the proposed combination in the same manner as the corresponding steps in its corresponding method claim. Additionally, the rationale and motivation to combine the Li and Ye references, presented in rejection of Claim 7, apply to this claim.
Claim 15 recites a computer-readable storage medium storing a program with instructions corresponding to the steps recited in Claim 1. Therefore, the recited programming instructions of this claim are mapped to the proposed combination in the same manner as the corresponding steps in its corresponding method claim. Additionally, the rationale and motivation to combine the Li and Ye references, presented in rejection of Claim 1, apply to this claim. Finally, the combination of Li and Ye references discloses a computer readable storage medium (for example, see Li, Paragraph [0043]).
Claim 16 recites a computer-readable storage medium storing a program with instructions corresponding to the steps recited in Claim 2. Therefore, the recited programming instructions of this claim are mapped to the proposed combination in the same manner as the corresponding steps in its corresponding method claim. Additionally, the rationale and motivation to combine the Li and Ye references, presented in rejection of Claim 2, apply to this claim.
Claim 19 recites a computer-readable storage medium storing a program with instructions corresponding to the steps recited in Claim 5. Therefore, the recited programming instructions of this claim are mapped to the proposed combination in the same manner as the corresponding steps in its corresponding method claim. Additionally, the rationale and motivation to combine the Li and Ye references, presented in rejection of Claim 5, apply to this claim.
Claim 20 recites a computer-readable storage medium storing a program with instructions corresponding to the steps recited in Claim 7. Therefore, the recited programming instructions of this claim are mapped to the proposed combination in the same manner as the corresponding steps in its corresponding method claim. Additionally, the rationale and motivation to combine the Li and Ye references, presented in rejection of Claim 7, apply to this claim.
Claims 3, 11, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Li in view of Ye and Li 2 (US 20220391755 A1).
Regarding Claim 3, the combination of Li and Ye discloses “The method of claim 1, further comprising:” (Please refer to the above-described analysis for Claim 1); (Ye, Paragraphs [0006], [0174], [0044], [0244], [0246]; For the sake of brevity, the descriptions and explanations for each of these paragraphs have been omitted as they have been described before in Claim 1. For more information on how these paragraphs correlate with each of the three scoring tasks (defined as metrics) described in this limitation, please refer to the above-described analysis for Claim 1). The combination of Li and Ye does not explicitly disclose “constructing the instruction data on the plurality of image-text pair quality scoring tasks using a teacher model”. However, in an analogous field of endeavor, Li 2 discloses “In one embodiment, in order to improve learning, such as in the presence of noisy input data for training the model, pseudo-targets are generated using momentum distillation (MoD) as an alternative of original noisy data for training the model. For all of the encoders (e.g., the image encoder 212, the text encoder 222, and the multimodal encoder 240), pseudo-targets are generated by a momentum model 260. The momentum model is a continuously-evolving teacher model which includes exponential-moving average versions of all of the encoders, including the unimodal and multimodal encoders.” (Li 2, Paragraph [0041] discloses “In one embodiment, in order to improve learning, such as in the presence of noisy input data for training the model, pseudo-targets are generated using momentum distillation (MoD) as an alternative of original noisy data for training the model. For all of the encoders (e.g., the image encoder 212, the text encoder 222, and the multimodal encoder 240), pseudo-targets are generated by a momentum model 260. The momentum model is a continuously-evolving teacher model which includes exponential-moving average versions of all of the encoders, including the unimodal and multimodal encoders.”). Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to combine the image-filtering method seen in the combination of Li and Ye with the technique of using a teacher model to construct instruction data seen in Li 2 to achieve a more complete image-filtering method. By combining the image filtering method seen in the combination of Li and Ye with the technique of using generating instruction data using a teacher model seen in Li 2, one of ordinary skill in the art can efficiently train the model to understand different contexts of text to produce consistent results for many image-text datasets to follow. Therefore, it would have been obvious for one of ordinary skill in the art to combine the Li, Ye, and Li 2 references to achieve the same image-filtering method described in Claim 3.
Claim 11 recites a system with elements corresponding to the steps recited in Claim 3. Therefore, the recited elements of this claim are mapped to the proposed combination in the same manner as the corresponding steps in its corresponding method claim. Additionally, the rationale and motivation to combine the Li, Ye, and Li 2 references, presented in rejection of Claim 3, apply to this claim.
Claim 17 recites a computer-readable storage medium storing a program with instructions corresponding to the steps recited in Claim 3. Therefore, the recited programming instructions of this claim are mapped to the proposed combination in the same manner as the corresponding steps in its corresponding method claim. Additionally, the rationale and motivation to combine the Li, Ye, and Li 2 references, presented in rejection of Claim 3, apply to this claim.
Claims 4, 12, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Li in view of Ye and Li 2, and further in view of Wu (CN 119537636 A).
Regarding Claim 4, the combination of Li, Ye, and Li 2 discloses “The method of claim 3, further comprising:” (Please refer to the above-described analysis for Claim 1); (Ye, Paragraphs [0006], [0174], [0044], [0244], [0246]; For the sake of brevity, the descriptions and explanations for each of these paragraphs have been omitted as they have been described before in Claim 1. For more information on how these paragraphs correlate with each of the three image-text quality scoring tasks (defined as metrics) described in this limitation, please refer to the above-described analysis for Claim 1); description, and then an expansion image-text pair for training the target image generation model is determined at least based on the first image description and at least two generated images corresponding to the prediction descriptions. And further, the extended image-text pairs can be utilized to continuously train and update the target image generation model, so that the image generation capacity of the target image generation model is better” (Wu, Paragraph [0111]). It is important to note that although Paragraph [0111] does not explicitly disclose a teacher model, the image generation model is being continuously trained by the image-text pairs found by using the text generation model to continuously improve it.
Wu also discloses “And the quality scores corresponding to the first description image pairs are calculated, and then the target description image pairs with better quality are screened out from all the first description image pairs based on the quality scores corresponding to the first description image pairs, and then the better-quality images and the first image descriptions in the target description image pairs are utilized to obtain better-quality extended image-text pairs so as to better optimize the image generation capacity of the target image generation model by utilizing the extended image-text pairs, so that the target image generation model can obtain better-quality images for the image descriptions which are not extended by the target text generation model, and better-quality image generation services can be provided for users, namely, the better-quality images are generated based on the image descriptions input by the users” (Wu, Paragraph [0112]). This paragraph covers the inputting of an image-text pair into the teacher model, as they are being used to improve the model and make it stronger to make better decisions on whether image-text pairs are high-quality or not. Additionally, we see how the quality scores are also generated for each image-text pair, with score explanations provided in the form of filtering out the image text pairs that have lower-quality scores compared to higher quality scores, thus prompting the teacher model to generate a score accordingly based off of the image-text pair given by the user.
Therefore, it would have been obvious for one of ordinary skill in the art to combine the image-text data filtering method seen in the combination of Li, Ye, and Li 2 with the technique of inputting an image-text pair and a description to the teacher model and prompting the teacher model to generate a score and scoring explanation for the image-text pair seen in Wang to achieve a complete image-text data filtering method. By using the teacher model to input image text-pair and a description and generate a score and scoring explanation for the image-text pair, one of ordinary skill in the art is able to effectively train the current model to understand the associations between specific images and text to be more efficient with the comparisons in different scenarios that may be fed into the model. Thus, it would have been obvious for one of ordinary skill in the art to use the Li, Ye, Li 2, and Wu references to achieve the same image-text data filtering method seen in Claim 4.
Claim 12 recites a system with elements corresponding to the steps recited in Claim 4. Therefore, the recited elements of this claim are mapped to the proposed combination in the same manner as the corresponding steps in its corresponding method claim. Additionally, the rationale and motivation to combine the Li, Ye, Li 2, and Wu references, presented in rejection of Claim 4, apply to this claim.
Claim 18 recites a computer-readable storage medium storing a program with instructions corresponding to the steps recited in Claim 4. Therefore, the recited programming instructions of this claim are mapped to the proposed combination in the same manner as the corresponding steps in its corresponding method claim. Additionally, the rationale and motivation to combine the Li, Ye, Li 2, and Wu references, presented in rejection of Claim 4, apply to this claim.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Li 3 (US 20240161520 A1) teaches systems and methods for a vision-language pretraining framework that bootstraps language-image pre-training with frozen image encoders and large language models.
Cao (CN 114090815 A) teaches a training method and training device of image description model.
Wang (CN 118734091 A) teaches a visual positioning and finger division method, system, device and storage medium based on mask finger modelling.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SORIE I KOROMA JR whose telephone number is (571)272-9259. The examiner can normally be reached Monday - Friday 8AM-6:00PM; Alternate Fridays Off.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amandeep Saini can be reached at 571-272-3382. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SORIE I KOROMA JR/Examiner, Art Unit 2662
/AMANDEEP SAINI/Supervisory Patent Examiner, Art Unit 2662