Prosecution Insights
Last updated: October 04, 2026
Application No. 19/095,344

IMAGE PROCESSING MODEL

Non-Final OA §101§103
Filed
Mar 31, 2025
Priority
Apr 08, 2024 — IL 312001
Examiner
TSWEI, YU-JANG
Art Unit
Tech Center
Assignee
B. G. Negev Technologies and Applications Ltd.
OA Round
1 (Non-Final)
84%
Grant Probability
Favorable
1-2
OA Rounds
8m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 84% — above average
84%
Career Allowance Rate
388 granted / 464 resolved
+23.6% vs TC avg
Strong +16% interview lift
Without
With
+16.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 3m
Avg Prosecution
46 currently pending
Career history
507
Total Applications
across all art units

Statute-Specific Performance

§101
5.9%
-34.1% vs TC avg
§103
72.8%
+32.8% vs TC avg
§102
6.0%
-34.0% vs TC avg
§112
7.4%
-32.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 464 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed in parent Application No. IL312001, filed on 2024/04/08. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claim 19 rejected under 35 U.S.C. 101 because the claimed inventions are directed to non-statutory subject matter. Claim 19 is directed to a computer program which does not fall within at least one of the four categories of patent eligible subject matter recited in 35 U.S.C. 101 (process, machine, manufacture, or composition of matter). Program claimed as computer to execute per se, i.e., the descriptions or expressions of the programs are not physical "things." They are neither computer components nor statutory processes, as they are not "acts" being performed. Such claimed computer programs do not define any structural and functional interrelationships between the computer program and other claimed elements of a computer which permit the computer program's functionality to be realized. In contrast, a claimed non-transitory computer-readable medium encoded with a computer program is a computer element which defines structural and functional interrelationships between the computer program and the rest of the computer which permit the computer program's functionality to be realized, and is thus statutory. See Lowry, 32 F.3d at 1583-84, 32 USPQ2d at 1035. Thus, claim 16 is rejected under 35 U.S.C. 101 because, giving the claims their broadest reasonable interpretation, the claimed "computer program" encompasses non-statutory subject matter. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-5, 12-17, 19, 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Fashandi et al. (US 20240290119 A1, hereinafter Fashandi) in view of Bai et al. (US 20230334834 A1, hereinafter Bai) Regarding Claim 1, Fashandi teaches a computer-implemented method comprising (Fashandi, Paragraph [0002], “The present disclosure relates to a device and method for harvesting data from unlabeled sources, in the field of artificial intelligence (AI).”): generating, using an image-to-text model, image descriptions of images in an original training set of images (Fashandi, Paragraph [0153], "the image-captioning model can be a pre-trained Meshed-Memory Transformer"), generating, using an image-to-text model, image descriptions (Fashandi, Paragraph [0154], "the pre-trained Meshed-Memory Transformer (MMT) can combine visual features extracted from the input image with textual embeddings to generate descriptive captions"), determining, using at least one large language model, LLM, and based on the image descriptions, at least one domain and/or class which is unrepresented or under-represented in the original training set (Fashandi, Paragraph [0140], "Similar to other real-life applications, scene graph generation (SGG) tasks often deal with imbalanced data. Moreover, the available data in this domain also suffers from missing links or noisy annotations. Part of the issue is that our visual world and how we describe it is biased. For example, rare cases found within the long-tail distribution are often overlooked or underrepresented"; [0143], "In this example base dataset, there are 50 predicates and 150 object classes. Eleven predicates out of the fifty are in the head (e.g., n>10,000 samples) and body categories (e.g., 10,000≥n≥5,000), and the remaining predicates are located in the tail categories (e.g., n<5,000 samples)"), [[generating, using a second LLM and based on the determination of the at least one domain and/or class, at least one instruction for a third LLM to generate at least one text prompt]], [[generating, using the third LLM and based on the at least one instruction, the at least one text prompt for a text-to-image model]], [[generating, using the text-to-image model and based on the at least one text prompt, at least one synthetic image]], generating an enhanced training set of images for use in training an image processing machine learning, ML, model, the enhanced training set of images comprising the original training set of images and the at least one synthetic image (Fashandi, Paragraph [0196], "The filtered training samples selected from among the harvested training samples can be converted to the same format as the base dataset and merged together with the base dataset, in order to produce an enhanced dataset that has an improved tail distribution"). But Fashandi does not explicitly disclose generating, using a second LLM and based on the determination of the at least one domain and/or class, at least one instruction for a third LLM to generate at least one text prompt and generating, using the third LLM and based on the at least one instruction, the at least one text prompt for a text-to-image model and generating, using the text-to-image model and based on the at least one text prompt, at least one synthetic image. However, Bai teaches generating, using a second LLM and based on the determination of the at least one domain and/or class, at least one instruction for a third LLM to generate at least one text prompt (Bai, Paragraph [0015], “a corresponding output may be generated for a given input after the training. The generation of the model may be based on a machine learning technique.”; [0033], "The image classification task may be configured to distinguish between a plurality of classes (e.g., k classes) with a plurality of class names. In this case, the associated training labels 216 or the text prompts 212 at least indicate the plurality of class names. Concretely, for a k-way classification, there may be k class names C={c1, . . . , ck}, the text prompts 212 may be generated from the k class names "), generating, using a second LLM and based on the determination, at least one instruction for a third LLM (Bai, Paragraph, "In some example embodiments, the text prompt generator 205 may generate one or more text prompts 212 by filling one or more class names into a text template 204"), generating, using the third LLM and based on the at least one instruction, the at least one text prompt for a text-to-image model (Bai, Paragraph [0035], "In some example embodiments, considering that only using the label names as inputs might limit the diversity of synthesized images and cause bottlenecks for validating the effectiveness of synthetic data, the text prompt generator 205 may alternatively or in addition, the text prompt generator 205 to increase the diversity of text prompts and the generated synthetic images, so as to achieve language enhancement (LE) and to better unleash the potential of synthesized data. The text prompt generator 205 may generate one or more text prompts 212 by providing one or more class names into a trained word-to-sentence model 206, to obtain at least one text prompt 212 generated by the word-to-sentence model 206"), the at least one text prompt for a text-to-image model (Bai, Paragraph [0035], "The word-to-sentence model 206 can generates diversified sentences containing the class names as language prompts for the text-to-image generation process"), generating, using the text-to-image model and based on the at least one text prompt, at least one synthetic image (Bai, Paragraph [0029], "The model training system 200 uses a text-to-image generation model 210 for synthetic data generation. The text-to-image generation model 210 is a generative model for image generation. The input to the text-to-image generation model 210 is a text prompt which describes the expected image to be generated. The output from the text-to-image generation model 210 is an image corresponding to the text prompt"; [0031], " The model training system 200 provides each of a plurality of text prompts 212-1, 212-2, . . . , 212-N (collectively or individually referred to as text prompts 212) into the text-to-image generation model 210, so as to generate a plurality of synthetic images 214-1, 214-2, . . . , 214-N (collectively or individually referred to as synthetic images 214) "), generating an enhanced training set of images comprising the original training set of images and the at least one synthetic image (Bai, Paragraph [0032], "The model training system 200 performs training of a target model 220, which is configured to perform an image classification task, based at least in part on the plurality of synthetic images 214 and the associated training labels"). Fashandi and Bai are analogous since both deal with improving training of image processing ML models via harvesting/generating additional training samples for rare tail categories. Fashandi provided a way of harvesting training data from unlabeled sources using image captioning model that generates descriptive captions and filtering tail categories to produce enhanced dataset with improved tail distribution. Bai provided a way of generating diversified text prompts containing class names via word-to-sentence model and generating synthetic images via text-to-image generation model. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to incorporate the text prompt generation and text-to-image synthesis taught by Bai into the modified invention of Fashandi such that after determining underrepresented predicates/classes via caption analysis, text prompt generator 205 generates instruction naming that class and word-to-sentence model 206 generates prompt for text-to-image model to synthesize images for enhanced training set. The motivation is to automatically produce more training data samples for rare cases found in the long-tail distribution discussed by in Paragraph and to reduce noise and enhance robustness and improve diversity of synthesized images discussed by Bai in Paragraph. [0033][0034][0035][0029][0031][0136][0044] Regarding Claim 2, the combination of Fashandi and Bai teaches the invention in Claim 1. The combination further teaches wherein determining the at least one domain and/or class which is unrepresented or under-represented in the original training set comprises: determining prevalent terms among the image descriptions (Fashandi, Paragraph[0176], "FIG. 9 shows an example word cloud for predicates output by the SGG model and a smaller word cloud that corresponds to the predicates used by the original dataset (e.g., the image captioning model has a larger vocabulary than the base dataset)"), determining prevalent terms among the image descriptions (Fashandi, Paragraph [0187], "Also, FIG. 10 shows an example word cloud for objects output by the captioning model and the textual SGG model, and a smaller word cloud that corresponds to the objects used by the original dataset (e.g., the image captioning model has a larger vocabulary for objects than the base dataset)"), inferring, using a first LLM and based on the prevalent terms, domains represented in the original training set and a number of images in the original training set representing each domain (Fashandi, Paragraph [0157], "The textual scene graph generator (SGG) model receives the captioned sentences from image-captioning model and generates a scene graph. The nodes of the scene graph can represent entities in the captioned sentences (e.g., objects and subjects), and the edges between the nodes can represent relations or predicates between those entities"), domains represented <read on predicates> (Fashandi, Paragraph [0142], "FIG. 5 shows the number of each predicate in the VG200 dataset as an example to help illustrate issues regarding the long-tail distribution"), wherein determining the at least one domain and/or class which is unrepresented or under-represented further comprises: (Fashandi, Paragraph [0176], "FIG. 9 shows an example word cloud for predicates output by the SGG model and a smaller word cloud that corresponds to the predicates used by the original dataset (e.g., the image captioning model has a larger vocabulary than the base dataset)"), determining prevalent terms among the image descriptions (, Paragraph, "Also, FIG. 10 shows an example word cloud for objects output by the captioning model and the textual SGG model, and a smaller word cloud that corresponds to the objects used by the original dataset (e.g., the image captioning model has a larger vocabulary for objects than the base dataset)"),: determining that a domain represented in the original training set is under-represented if the number or proportion of images in the original training set representing the domain is below a domain threshold (Fashandi, Paragraph [0227], "the filter can select only those training samples that have predicates that are located in the tail categories (e.g., n<5,000 samples), and the other samples that fall within the head and body can be discarded"), [[determining, using a fourth LLM, if at least one domain exists which is not represented by any of the images in the original training set, and if it is determined that at least one domain exists which is not represented by any of the images in the original training set, determining the at least one domain as at least one unrepresented domain]]. But Fashandi does not explicitly disclose determining using a fourth LLM if at least one domain does not represent. However Bai teaches determining prevalent terms among the image descriptions (Bai, Paragraph [0048], "The training data distribution of the text-to-image generation model GLIDE would exhibit bias and produce different domain gaps with different datasets"), inferring, using a first LLM and based on the prevalent terms, domains represented in the original training set and a number of images in the original training set representing each domain (Bai, Paragraph [0065], "As shown in Table 5, both the RF strategy and the RG strategy provide performance gains upon the B strategy which is the best strategy in the zero-shot setting. This demonstrates the importance of utilizing the domain knowledge from few-shot images for preparing the synthetic data"), determining that a domain represented in the original training set is under-represented if the number or proportion of images in the original training set representing the domain is below a domain threshold (Bai, Paragraph [0067], "As for synthetic data, the statistical difference between different domains can provide good attribution"). determining, using a fourth LLM, if at least one domain exists which is not represented by any of the images in the original training set, and if it is determined that at least one domain exists which is not represented by any of the images in the original training set, (Bai, Paragraph [0040], "Regarding few-shot setting (or few-shot learning), the same aims to recognize new classes when provided with one or a few labeled real images of these classes. Considering the scarce real images for zero-shot setting < read on domain exists which is not represented > or few-shot setting, in embodiments of the present disclosure, given a label space for the zero-shot or few-shot task (i.e., given the k class names), the text-to-image generation model 210 is utilized to generate synthetic images 214 given the class names"; [0041], "In some example embodiments, for the zero-shot setting < read on domain exists which is not represented >where no real training images of the target classes are available, a multi-modal model may be pre-trained with large-scale image-caption pairs, and the similarities between paired image features (from an image encoder g) and text features (from a text encoder h) are maximized during pre-training"), determining the at least one domain as at least one unrepresented domain (Bai, Paragraph [0042], "In this way, as compared with the traditional way of zero-shot learning, labeled training data of synthetic images corresponding to the new classes can become available for model fine-tuning"). Fashandi and Bai analogous since both deal with long-tail / domain gap via counting samples per predicate/domain below threshold n<5,000. Fashandi provides domain gap discussion, Bai provides prevalent term word cloud counting. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to incorporate infer domains and threshold taught by Bai into modified invention of Fashandi the motivation is to handle zero-shot unrepresented domains and long-tail underrepresented domains to improve recall discussed by Bai in Paragraph [0040] - [0042]. Regarding Claim 3, the combination of Fashandi and Bai teaches the invention in Claim 1. The combination further teaches determining the at least one domain and/or class which is unrepresented or under-represented in the original training set comprises determining, based on metadata and/or labels associated with the images, a number of images in the original training set associated with each class (Fashandi, Paragraph [0159], " Further in this example, with reference to the lower path in the of the caption-based pipeline shown in FIG. 7 , the object detection model (e.g., object detector) can receive unlabeled images as an input, and outputs annotated images that include information for bounding boxes and their corresponding class labels"), wherein determining the at least one domain and/or class which is unrepresented or under-represented further comprises: (Fashandi, Paragraph[0176], "FIG. 9 shows an example word cloud for predicates output by the SGG model and a smaller word cloud that corresponds to the predicates used by the original dataset (e.g., the image captioning model has a larger vocabulary than the base dataset)") determining that a class represented in the original training set is under-represented if the number or proportion of images in the original training set representing the class is below a class threshold (Fashandi, Paragraph [0143], "In this example base dataset, there are 50 predicates and 150 object classes. Eleven predicates out of the fifty are in the head (e.g., n>10,000 samples) and body categories (e.g., 10,000≥n≥5,000), and the remaining predicates are located in the tail categories (e.g., n<5,000 samples)"), class threshold <read on n<5,000> (Fashandi, Paragraph [0194], "the filter can select only those training samples that have predicates that are located in the tail categories (e.g., n<5,000 samples), and the other samples that fall within the head and body can be discarded"), [[determining, using a sixth LLM, if at least one class exists which is not represented by any of the images in the original training set, and if it is determined that at least one class exists which is not represented by any of the images in the original training set, determining the at least one class as at least one unrepresented class]]. However, Bai teaches determining, using a sixth LLM, if at least one class exists which is not represented by any of the images in the original training set, and if it is determined that at least one class exists which is not represented by any of the images in the original training set (Bai, Paragraph [0040], "Regarding few-shot setting (or few-shot learning), the same aims to recognize new classes when provided with one or a few labeled real images of these classes. Considering the scarce real images for zero-shot setting or few-shot setting, in embodiments of the present disclosure, given a label space for the zero-shot <read on class exists which is not represented > or few-shot task (i.e., given the k class names), the text-to-image generation model 210 is utilized to generate synthetic images 214 given the class names"), determining the at least one class as at least one unrepresented class (, Bai, Paragraph [0042], "In this way, as compared with the traditional way of zero-shot learning, labeled training data of synthetic images corresponding to the new classes can become available for model fine-tuning"). As explained in rejection of claim 1, the obviousness for combining of Bai into Fashandi is provided above. Regarding Claim 4, the combination of Fashandi and Bai teaches the invention in Claim 1. The combination further teaches wherein generating the at least one instruction comprises: when an under-represented or unrepresented class has been determined, generating an instruction naming the under-represented or unrepresented class (, Fashandi, Paragraph [0194], "the filter can select only those training samples that have predicates that are located in the tail categories (e.g., n<5,000 samples), and the other samples that fall within the head and body can be discarded"), when an under-represented or unrepresented class has been determined, generating an instruction naming the under-represented or unrepresented class (Fashandi, Paragraph [0228], "The filtered training samples selected from among the harvested training samples can be converted to the same format as the base dataset and merged together with the base dataset, in order to produce an enhanced dataset that has an improved tail distribution"), and when an under-represented or unrepresented domain has been determined, generating an instruction naming the under-represented or unrepresented domain (Fashandi, Paragraph [0136], "According to an embodiment, the AI device 100 can harvest training data from unlabeled sources for improving scene graph generation and related downstream tasks. Particularly, the AI device 100 can produce more training data samples for rare cases found in the long-tail distribution, which can help improve the recall rate, zero-shot recall rate, and mean recall rate among other improvements"). Bai further teaches when an under-represented or unrepresented class has been determined, generating an instruction naming the under-represented or unrepresented class (Bai, Paragraph [0033], "The image classification task may be configured to distinguish between a plurality of classes (e.g., k classes) with a plurality of class names. In this case, the associated training labels 216 or the text prompts 212 at least indicate the plurality of class names. Concretely, for a k-way classification, there may be k class names C={c1,..., ck}, the text prompts 212 may be generated from the k class names. The model training system 200 comprises a text prompt generator 205 to generate the text prompts 212 from the k class names"), generating an instruction naming the under-represented or unrepresented class (, Bai, Paragraph [0034], "In some example embodiments, the text prompt generator 205 may generate one or more text prompts 212 by filling one or more class names into a text template 204"), generating an instruction naming the under-represented or unrepresented domain (Bai, Paragraph [0065], "As shown in Table 5, both the RF strategy and the RG strategy provide performance gains upon the B strategy which is the best strategy in the zero-shot setting. This demonstrates the importance of utilizing the domain knowledge from few-shot images for preparing the synthetic data"). As explained in rejection of claim 1, the obviousness for combining of Bai into Fashandi is provided above. Regarding Claim 5, the combination of Fashandi and Bai teaches the invention in Claim 1. The combination further teaches wherein generating the at least one synthetic image comprises generating a synthetic set comprising a plurality of synthetic images (Fashandi, Paragraph [0193], "The output of the matching block is annotated/labeled image training samples that have been converted/matched to use the same words/vocabulary that is used in the base dataset. However, these harvested data samples are distributed across the head, body and tail for the predicates, and a filtering step can used to select only those training samples that fall within the tail categories, in order to provide more samples for these rare situations"), wherein the computer-implemented method further comprises performing a cleaning process comprising cleaning the synthetic set by removing any synthetic image determined to be an outlier to generate a cleaned synthetic set of synthetic images (Fashandi, Paragraph [0169], "In order to align the vocabulary in newly harvested training samples with the same vocabulary used by the base dataset, the matching block can perform a matching/converting operation for the objects and the predicates. According to an embodiment, the matching block can implement a rules based approach, which is explained in more detail below"), cleaning the synthetic set <read on matching/converting and discarding> (, Paragraph, "the filter can select only those training samples that have predicates that are located in the tail categories (e.g., n<5,000 samples), and the other samples that fall within the head and body can be discarded"), and wherein the enhanced training set comprises the original training set of images and the cleaned synthetic set of synthetic images (Fashandi, Paragraph [0196], "The filtered training samples selected from among the harvested training samples can be converted to the same format as the base dataset and merged together with the base dataset, in order to produce an enhanced dataset that has an improved tail distribution"). Bai further teaches generating a synthetic set comprising a plurality of synthetic images (Bai, Paragraph [0031], "The model training system 200 provides each of a plurality of text prompts 212-1, 212-2,..., 212-N (collectively or individually referred to as text prompts 212) into the text-to-image generation model 210, so as to generate a plurality of synthetic images 214-1, 214-2,..., 214-N (collectively or individually referred to as synthetic images 214)"), performing a cleaning process comprising cleaning the synthetic set by removing any synthetic image determined to be an outlier to generate a cleaned synthetic set of synthetic images (Bai, Paragraph [0044], "Further, in some example embodiments, to reduce noise and enhance robustness, the synthetic images 214 generated from the text-to-image model 210 may be filtered to remove low-quality samples"), cleaning the synthetic set by removing any synthetic image determined to be an outlier (Bai, Paragraph [0057], "In a first strategy, referred to as a real filtering (RF) strategy, given the synthetic images of one class ci, the features of few-shot non-synthetic images to filter out synthetic images whose features are very close to the features of real non-synthetic images that belong to other classes different from the class ci"), outlier <read on low-quality / low-reliable> (, Paragraph [0057], "If a second feature similarity is higher than the first feature similarity, which means that the first synthetic image is much closer to the second class than to the first class, then the first synthetic image is discarded. In this way, low-reliable synthetic images in a class can be filtered out"). Fashandi and Bai analogous since both deal with long-tail / domain gap via counting samples per predicate/domain below threshold n<5,000. Fashandi provides merging filtered tail samples into enhanced dataset, Bai provides synthetic set plurality and filtering low-quality/outlier via quality scores and feature similarity. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to incorporate infer domains and threshold taught by Bai into modified invention of Fashandi, the motivation to clean synthetic set before merging into enhanced set to reduce noise and enhance robustness discussed in Bai in Paragraph [0031][0057]. Regarding Claim 12, the combination of Fashandi and Bai teaches the invention in Claim 1. The combination further teaches comprising training the image processing ML model using the enhanced training set of images (Fashandi, Paragraph [0197], "FIG. 11A shows example statistics of the harvested data according to the caption-based (CB) approach, and FIG. 11B shows an example of those training samples merged with the base dataset which have been added to the tail categories to generate an enhanced dataset. In this way, the caption-based (CB) approach can automatically generate new triplet training samples from unlabeled image sources in a manner that saves time and reduces costs, and enhances the rare tail categories, which can improve recall rates and accuracy for downstream tasks"). Bai further teaches training the image processing ML model using the enhanced training set of images (Bai, Paragraph [0032], "The model training system 200 performs training of a target model 220, which is configured to perform an image classification task, based at least in part on the plurality of synthetic images 214 and the associated training labels"), training the image processing ML model using the enhanced training set (Bai, Paragraph [0088], "In some example embodiments, the target model comprises a pre-trained model and the image classification task is a downstream task for the pre-trained model with zero-shot setting for the plurality of classes. In some example embodiments, to performing the training of the target model, the model training system 220 may fine-tune the pre-trained model based on the plurality of synthetic images and the associated training labels"). As explained in rejection of claim 1, the obviousness for combining of Bai into Fashandi is provided above. Regarding Claim 13, the combination of Fashandi and Bai teaches the invention in Claim 12. The combination further teaches wherein the computer-implemented method comprises using the image processing ML model after training (Fashandi, Paragraph [0147], "This harvested data can be merged with the original base dataset in order to create an enhanced dataset, which can be used to train a model, such as a scene graph generation (SGG) model and provide improved performance for various downstream applications"), using the image processing ML model after training (Fashandi, Paragraph [0136], "According to an embodiment, the AI device 100 can harvest training data from unlabeled sources for improving scene graph generation and related downstream tasks"). Regarding Claim 14, the combination of Fashandi and Bai teaches the invention in Claim 1. The combination further teaches wherein the image processing ML model comprises an image retrieval model (Fashandi, Paragraph [0062], “Machine learning is defined as an algorithm that enhances the performance of a certain task through a steady experience with the certain task.”; [0139], “Examples of such reasonings on the scene graph structure are visual question answering, image-captioning, image editing and retrieval, and visual grounding”). Regarding Claim 15, the combination of Fashandi and Bai teaches the invention in Claim 14. The combination further teaches wherein the image retrieval model is for searching (Fashandi, Paragraph [0139], "Examples of such reasonings on the scene graph structure are visual question answering, image-captioning, image editing and retrieval, and visual grounding."), among video frames (Fashandi, Paragraph [0005], "A scene graph (SG) is a structured representation of the visual content of an image or video."; [0006], "For instance, SGs are often constructed based on a process of detailed annotation, where humans, e.g., subject matter experts (SMEs), manually identify and label the objects, relationships, and attributes in an image or video."), [[for at least one image similar to a query image]]. But Fashandi does not explicitly disclose for at least one image similar to a query image. However, Bai teaches wherein the image retrieval model is for searching among video frames for at least one image similar to a query image (Bai, Paragraph [0032], "The model training system 200 performs training of a target model 220, which is configured to perform an image classification task, based at least in part on the plurality of synthetic images 214 and the associated training labels."), for at least one image similar to a query image (Bai, Paragraph [0090], "In some example embodiments, to fine-tune the pre-trained model, the model training system 220 may filter the plurality of synthetic images by: for a first synthetic image associated with a first class, determining a first feature similarity between the first synthetic image and a first non-synthetic image in the first class, and a second feature similarity between the first synthetic image and a second non-synthetic image in a second class, and in accordance with a determination that the second feature similarity is higher than the first feature similarity, discarding the first synthetic image;"), for at least one image similar to a query image (Bai, Paragraph [0026], "synthetic images are generated by providing respective text prompts into a text-to-image generation model."). Bai and Fashandi are analogous since both deal with training ML models with enhanced/synthetic data for retrieval and improving recall for rare cases. Fashandi provided way of reasoning on scene graph structure for retrieval over image or video. Bai provided way of determining first feature similarity between synthetic image and non-synthetic image in same class to filter and keep similar images to query.Therefore, it would have been obvious to incorporate searching for similar image via feature similarity taught by Bai into modified invention of Fashandi such that image retrieval model is for searching among video frames for at least one image similar to query image. The motivation is to improve recall rate and mean recall rate for rare cases and reduce noise and enhance robustness discussed by Bai in Paragraph [0032][0090][0002]. Regarding Claim 16, the combination of Fashandi and Bai teaches the invention in Claim 15. The combination further teaches wherein the query image comprises an object and the video frames comprises video frames from a surveillance (Bai, Paragraph [0090], "In some example embodiments, to fine-tune the pre-trained model, the model training system 220 may filter the plurality of synthetic images by: for a first synthetic image associated with a first class, determining a first feature similarity between the first synthetic image and a first non-synthetic image in the first class, and a second feature similarity between the first synthetic image and a second non-synthetic image in a second class, and in accordance with a determination that the second feature similarity is higher than the first feature similarity, discarding the first synthetic image;"), for at least one image similar to a query image (Bai, Paragraph [0031], "The model training system 200 provides each of a plurality of text prompts 212-1, 212-2,..., 212-N into the text-to-image generation model 210, so as to generate a plurality of synthetic images 214-1, 214-2,..., 214-N"). Fashandi and Bai analogous since both deal with long-tail/domain gap via counting samples per predicate/domain below threshold. Fashandi provides merging filtered tail samples into enhanced dataset, Bai provides video frames from a surveillance. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to incorporate different video source taught by Bai into modified invention of Fashandi, the motivation is to improve recall for rare tail cases and enhance surveillance retrieval discussed by [Bai] in Paragraph [0090][0031]. Regarding Claim 17, the combination of Fashandi and Bai teaches the invention in Claim 15. The combination further teaches wherein the query image comprises a vehicle (Fashandi, Paragraph [0070], "Self-driving refers to a technique of driving for oneself, and a self-driving vehicle refers to a vehicle that travels without an operation of a user or with a minimum operation of a user."; [0072], "The vehicle can include a vehicle having only an internal combustion engine, a hybrid vehicle having an internal combustion engine and an electric motor together, and an electric vehicle having only an electric motor, and can include not only an automobile but also a train, a motorcycle, and the like."), and/or the video frames comprises video frames from a traffic camera video (Fashandi, Paragraph [0134], "Alternatively, the robot 100 a that interacts with the self-driving vehicle 100 b can provide information or assist the function to the self-driving vehicle 100 b outside the self-driving vehicle 100 b. For example, the robot 100 a can provide traffic information including signal information and the like, such as a smart signal, to the self-driving vehicle 100 b, and automatically connect an electric charger to a charging port by interacting with the self-driving vehicle 100 b like an automatic electric charger of an electric vehicle."; [0234], "According to an embodiment, the AI device 100 can used the enhanced dataset to train a scene graph generation model. The trained scene graph generation model can be used for various applications, such as computer vision applications (e.g., self-driving, surveillance, robot guidance, etc.), question and answering systems or recommendations systems."), Regarding Claim 19, it recites limitations similar in scope to the limitations of claim 1 and the combination of Fashandi and Bai teaches all the limitations as of Claim 1. And Fashandi discloses these features can be implemented on a computer readable storage medium (Fashandi, Paragraph [0238], “Various aspects of the embodiments described herein can be implemented in a computer-readable medium”; [0104], “The learning model can be implemented in hardware, software, or a combination of hardware and software. If all or part of the learning models are implemented in software, one or more instructions that constitute the learning model can be stored in the memory”; [0117], “The method can be the form of an executable application or program”). Regarding Claim 20, it recites limitations similar in scope to the limitations of claim 1, but in an apparatus. As shown in the rejection, the combination of Fashandi and Bai disclose the limitations of claims 1. Additionally, Fashandi discloses an apparatus that maps to Fig. 1, Paragraph [0076] (Fashandi, Fig. 1, Element 180, Processor, Element 170, Memory, Paragraph [0076], “Referring to FIG. 1, the AI device 100 can include … a memory 170, and a processor 180 ( e.g., a controller).). Thus, Claim 20 is met by Fashandi according to the mapping presented in the rejection of claims 1, given the method corresponds to the apparatus. Claim(s) 6, 7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Fashandi et al. (US 20240290119 A1, hereinafter Fashandi) in view of Bai et al. (US 20230334834 A1, hereinafter Bai) as applied to Claim 1 above and further in view of Matamoros et al. (US 20250209309 A1, hereinafter Matamoros). Regarding Claim 6, the combination of Fashandi and Bai teaches the invention in Claim 5. The combination further teaches wherein cleaning the synthetic set (Fashandi , Paragraph [0194], " the filter can select only those training samples that have predicates that are located in the tail categories (e.g., n<5,000 samples), and the other samples that fall within the head and body can be discarded ") to generate the cleaned synthetic set comprises: generating first embeddings of the images in the original training set (Fashandi, Paragraph [0233], "In line 5 of Algorithm 3, the fsemantic vector is calculated, which is the concatenation of two, 100-dimensinal GloVE embeddings of the subject label (e.g., cs) and the object label (e.g., co), which are the labels that were obtained by the object detector."; [0222], "With reference to line 4 of Algorithm 3, the equations in lines 1-3 are used to calculate the fspatial vector, which is a 22 dimensional vector."; [0224], "Further in this example, in line 6, the visual features (e.g., fvisual vector) are calculated, which is the 1024-dimensional features of the object detector's ConvNet backbone followed by an ROI-align”, [[generating second embeddings of the synthetic images in the synthetic set which are associated with a class which is represented in the original training set; computing an average embedding for each class of images in the original training set; for each class of images in the original training set, computing an average distance of distances of the first embeddings of the images of the class from the average embedding of the class; and for each second embedding, comparing the distance between the second embedding and the average embedding for the corresponding class with a class outlier threshold which is based on the average distance for the corresponding class and, if the distance is greater than the class outlier threshold, removing the synthetic image corresponding to the second embedding from the synthetic set]]. Fashandi does not explicitly disclose but Bai teaches generating second embeddings of the synthetic images in the synthetic set which are associated with a class which is represented in the original training set (Bai, Paragraph [0031], "The model training system 200 provides each of a plurality of text prompts 212-1, 212-2,..., 212-N (collectively or individually referred to as text prompts 212) into the text-to-image generation model 210, so as to generate a plurality of synthetic images 214-1, 214-2,..., 214-N (collectively or individually referred to as synthetic images 214)."), cleaning the synthetic set by removing the synthetic image corresponding to the second embedding from the synthetic set (Bai, Paragraph [0044], "Further, in some example embodiments, to reduce noise and enhance robustness, the synthetic images 214 generated from the text-to-image model 210 may be filtered to remove low-quality samples."), comparing the distance between the second embedding and the average embedding for the corresponding class with a class outlier threshold (Bai, Paragraph [0090], "In some example embodiments, to fine-tune the pre-trained model, the model training system 220 may filter the plurality of synthetic images by: for a first synthetic image associated with a first class, determining a first feature similarity between the first synthetic image and a first non-synthetic image in the first class, and a second feature similarity between the first synthetic image and a second non-synthetic image in a second class, and in accordance with a determination that the second feature similarity is higher than the first feature similarity, discarding the first synthetic image;"), if the distance is greater than the class outlier threshold, removing the synthetic image (Bai, Paragraph [0083], "filter the plurality of synthetic images based on the determined respective quality scores, to discard at least one of the plurality of synthetic images;"). As explained in rejection of claim 1, the obviousness for combining of Bai into Fashandi is provided above. But the combination does not explicitly disclose computing an average embedding for each class of images in the original training set; for each class of images in the original training set, computing an average distance of distances of the first embeddings of the images of the class from the average embedding of the class. However, Matamoros teaches generating first embeddings of the images in the original training set; generating second embeddings of the synthetic images in the synthetic set which are associated with a class which is represented in the original training set; computing an average embedding for each class of images in the original training set (Matamoros, Paragraph [0010], "applying an embedding transformation to each of the seed data objects in the first set of seed data objects to create a first modified set of seed data objects; retrieving a first plurality of candidates from a database of data objects based on similarity to the first modified set of seed data objects;"), comparing the distance between the second embedding and the average embedding (, Paragraph, "wherein the similarity measure is a distance measure."), , for each second embedding, comparing the distance between the second embedding and the average embedding for the corresponding class with a class outlier threshold which is based on the average distance for the corresponding class and, if the distance is greater than the class outlier threshold, removing the synthetic image corresponding to the second embedding from the synthetic set (Matamoros, Paragraph [0091], "In other examples, a similarity distance threshold (e.g., a Euclidean distance measured between a seed embedding and a candidate embedding in any direction within the embedding space) may be used to determine the number of candidate embeddings 445 to extract for each respective seed embedding 435."), Fashandi and Matamoros are analogous since both deal with beddings of the images in the original training set. Fashandi provides embedding vectors fspatial/fsemantic/fvisual and filtering tail, Matamoros provides embedding transformation to seed embedding and similarity distance threshold. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to incorporate similarity distance calculation taught by Matamoros into modified invention of Fashandi the motivation is to improve computational efficiency discussed by Matamoros in Paragraph [0010][0012][0091][0014][0008]. Regarding Claim 7, the combination of Fashandi, Bai and Matamoros teaches the invention in Claim 6. The combination further teaches wherein the class outlier threshold for a given class comprises the average distance for the class multiplied by a diversity factor (Bai, Paragraph [0035], "to increase the diversity of text prompts and the generated synthetic images, so as to achieve language enhancement (LE) and to better unleash the potential of synthesized data."; [0057], "If a second feature similarity is higher than the first feature similarity, which means that the first synthetic image is much closer to the second class than to the first class, then the first synthetic image is discarded."). As explained in rejection of claim 1, the obviousness for combining of Bai into Fashandi is provided above. Claim(s) 8-11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Fashandi et al. (US 20240290119 A1, hereinafter Fashandi) in view of Bai et al. (US 20230334834 A1, hereinafter Bai) as applied to Claim 1 above and further in view of Shreshtha et al. (US 20220222873 A1, hereinafter Shreshtha). Regarding Claim 8, the combination of Fashandi and Bai teaches the invention in Claim 5. The combination further teaches [[performing a checking process comprising checking the cleaned synthetic set to determine whether additional synthetic images are required and, if it is determined that additional synthetic images are required, ]] performing a cleaning compensation process comprising: generating, using the second LLM, at least one further instruction for the third LLM to generate at least one text prompt (Bai, Paragraph [0035], "The text prompt generator 205 may generate one or more text prompts 212 by providing one or more class names into a trained word-to-sentence model 206, to obtain at least one text prompt 212 generated by the word-to-sentence model 206."); and generating, using the text-to-image model and based on the at least one text prompt, at least one further synthetic image (Bai, Paragraph [0228], "The filtered training samples selected from among the harvested training samples can be converted to the same format as the base dataset and merged together with the base dataset, in order to produce an enhanced dataset that has an improved tail distribution"). But the combination does not explicitly disclose checking whether additional synthetic images required. However, Shreshtha teaches performing a checking process comprising checking the cleaned synthetic set to determine whether additional synthetic images are required (Shreshtha, Paragraph [0035], "Datasets are collections of data used to build an ML mathematical model, so as to make data-driven predictions or decisions. Three types of ML datasets (also designated as ML sets) are typically dedicated to three respective kinds of tasks: training, i.e. fitting the parameters, validation, i.e. tuning ML hyper-parameters (which are parameters used to control the learning process), and testing or evaluation i.e. checking independently of a training dataset exploited for building a mathematical model that the latter model provides satisfying results."). Shreshtha and Fashandi are analogous since both of them are dealing with building and training ML mathematical models using image datasets for data-driven predictions.Fashandi provided a way of training a scene graph generation model based on an enhanced dataset formed by merging base dataset with filtered harvested training samples to improve tail distribution. Shreshtha provided a way of using three types of ML datasets for training, validation and testing/evaluation, where validation is for tuning ML hyper-parameters and testing/evaluation is for checking independently of a training dataset that the model provides satisfying results. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to incorporate checking independently of a training dataset that the model provides satisfying results taught by Shreshtha into modified invention of Fashandi such that the cleaned synthetic set is checked via validation/testing to determine whether additional synthetic images are required before using the enhanced dataset for training. The motivation is to make data-driven predictions or decisions and to ensure the mathematical model provides satisfying results independently of the training dataset and to tune hyper-parameters which control the learning process discussed by in Paragraph.[Shreshtha][0035] Regarding Claim 9, the combination of Fashandi and Bai teaches the invention in Claim 5. The combination further teaches performing a training feedback process comprising evaluating performance of a trained image processing ML model to determine whether further additional synthetic images are required (Fashandi, Paragraph [0147], "This harvested data can be merged with the original base dataset in order to create an enhanced dataset, which can be used to train a model, such as a scene graph generation (SGG) model and provide improved performance for various downstream applications."), performing a training feedback process comprising evaluating performance of a trained image processing ML model (Fashandi, Paragraph [0136], "According to an embodiment, the AI device 100 can harvest training data from unlabeled sources for improving scene graph generation and related downstream tasks. Particularly, the AI device 100 can produce more training data samples for rare cases found in the long-tail distribution, which can help improve the recall rate, zero-shot recall rate, and mean recall rate among other improvements."), and, if it is determined that further additional synthetic images are required, [[ performing a weak class compensation process ]] comprising (Fashandi, Paragraph [0194], "For example, in FIG. 7, the filter can select only those training samples that have predicates that are located in the tail categories (e.g., n<5,000 samples): [[generating, using the second LLM, at least one further instruction for the third LLM to generate at least one text prompt]], [[generating, using the third LLM and based on the at least one further instruction, the at least one text prompt for the text-to-image model]], [[generating, using the text-to-image model and based on the at least one text prompt, at least one further additional synthetic image]]. But does not explicitly disclose generating, using the second LLM, at least one further instruction for the third LLM to generate at least one text prompt; generating, using the third LLM and based on the at least one further instruction, the at least one text prompt for the text-to-image model; and generating, using the text-to-image model and based on the at least one text prompt, at least one further additional synthetic image. However, Bai teaches performing a training feedback process comprising evaluating performance of a trained image processing ML model to determine whether further additional synthetic images are required (Bai, Paragraph [0046], "To validate the performances of using synthetic images for model fine-tuning with zero-shot setting, the inventors have selected a number of diverse datasets covering object-level (CIFAR-10 and CIFAR-100, Caltech101, Caltech256, ImageNet), scene-level (SUN397), fine-grained (Aircraft, Birdsnap, Cars, CUB, Flower, Food, Pets), textures (DTD), satelite images (EuroSAT) and robustness (ImageNetSketch, ImageNet-R) for zero-shot image classification."), generating, using the second LLM, at least one further instruction for the third LLM to generate at least one text prompt (Bai, Paragraph [0035], "the text prompt generator 205 may generate one or more text prompts 212 by providing one or more class names into a trained word-to-sentence model 206, to obtain at least one text prompt 212 generated by the word-to-sentence model 206."), generating, using the third LLM and based on the at least one further instruction, the at least one text prompt for the text-to-image model (Bai, Paragraph [0031], "The model training system 200 provides each of a plurality of text prompts 212-1, 212-2,..., 212-N (collectively or individually referred to as text prompts 212) into the text-to-image generation model 210, so as to generate a plurality of synthetic images 214-1, 214-2,..., 214-N (collectively or individually referred to as synthetic images 214)."), and generating, using the text-to-image model and based on the at least one text prompt, at least one further additional synthetic image (Bai, Paragraph [0031], "The model training system 200 provides each of a plurality of text prompts 212-1, 212-2,..., 212-N (collectively or individually referred to as text prompts 212) into the text-to-image generation model 210, so as to generate a plurality of synthetic images 214-1, 214-2,..., 214-N (collectively or individually referred to as synthetic images 214)."). As explained in rejection of claim 1, the obviousness for combining of Bai into Fashandi is provided above. But the combination does not explicitly disclose performing a weak class compensation process comprising generating using second LLM further instruction naming weak class as at least one weak class. However, Shreshtha teaches and, if it is determined that further additional synthetic images are required, performing a weak class compensation process comprising: generating, using the second LLM, at least one further instruction for the third LLM to generate at least one text prompt; generating, using the third LLM and based on the at least one further instruction, the at least one text prompt for the text-to-image model; and generating, using the text-to-image model and based on the at least one text prompt, at least one further additional synthetic image <read on weak class compensation> (Shreshtha, Paragraph [0016], "In an example of the example preceding method, the method may further include: obtaining a second set of seed data objects based on a second identified desired attribute; applying an embedding transformation to each of the seed data objects in the second set of seed data objects to create a second modified set of seed data objects; retrieving a second plurality of candidates from the database of data objects based on similarity to the second modified set of seed data objects; using the LLM, annotating the second plurality of candidates based on the list of defined labels; appending the training dataset to include the second plurality of annotated candidates; and training the machine learning model using the training dataset."). Shreshtha and Fashandi are analogous since both deal with generating synthetic training data for tail/weak classes to improve model performance. Fashandi provided way of harvesting tail category samples to enhance dataset for SGG training. Shreshtha provided way of generating text prompts via word-to-sentence model and feeding into text-to-image model to generate synthetic images and validating performance via diverse datasets. Therefore, it would have been obvious to incorporate generating further instruction via second LLM for third LLM and generating further synthetic images taught by Shreshtha and obtaining second set based on second desired attribute taught into modified invention of Fashandi such that when training feedback determines further synthetic required, weak class compensation generates further instruction via second LLM for third LLM to generate text prompt and further synthetic images for weak class.Motivation is to produce more training data samples for rare cases in long-tail distribution to improve recall rate and reduce noise and enhance robustness discussed by Shreshtha in Paragraph [0016]. Regarding Claim 10, the combination of Fashandi, Bai and Matamoros teaches the invention in Claim 9. The combination further teaches wherein the training feedback process comprises: training the image processing ML model using the enhanced training set of images to generate the trained image processing ML model (Fashandi, Paragraph [0014], "training, via the processor, a scene graph generation model based on the enhanced dataset to generate a trained scene graph generation model, in which the trained scene graph generation model includes at least one trained neural network that is trained based on the enhanced dataset."; [0147], "This harvested data can be merged with the original base dataset in order to create an enhanced dataset, which can be used to train a model, such as a scene graph generation (SGG) model and provide improved performance for various downstream applications."), and evaluating performance of the trained image processing ML model using test images (Fashandi, Paragraph [0136], "According to an embodiment, the AI device 100 can harvest training data from unlabeled sources for improving scene graph generation and related downstream tasks. Particularly, the AI device 100 can produce more training data samples for rare cases found in the long-tail distribution, which can help improve the recall rate, zero-shot recall rate, and mean recall rate among other improvements.") [[and when the performance of the trained image processing ML model is below a performance threshold in respect of any class of the test images, determining that further additional synthetic images are required]] , and determining the class as at least one weak class (Fashandi, Paragraph [0094], "For example, in FIG. 7, the filter can select only those training samples that have predicates that are located in the tail categories (e.g., n<5,000 samples), and the other samples that fall within the head and body can be discarded, but embodiments are not limited thereto."). But Fashandi does not explicitly disclose evaluating performance of the trained image processing ML model using test images and when the performance of the trained image processing ML model is below a performance threshold in respect of any class of the test images, determining that further additional synthetic images are required. However, Bai teaches training the image processing ML model using the enhanced training set of images to generate the trained image processing ML model (Bai, Paragraph [0032], "The model training system 200 performs training of a target model 220, which is configured to perform an image classification task, based at least in part on the plurality of synthetic images 214 and the associated training labels."), and evaluating performance of the trained image processing ML model using test images (Bai, Paragraph [0047], "All results are top-1 accuracy on the test dataset."), and evaluating performance of the trained image processing ML model using test images (Bai, Paragraph [0053], "A performance of 28.74% top-1 accuracy on CIFAR-100 test dataset is achieved, which is much lower than the performance of the pre-trained CLIP model"), and when the performance of the trained image processing ML model is below a performance threshold in respect of any class of the test images, determining that further additional synthetic images are required (Bai, Paragraph [0060], "As shown in FIG. 3A, with only few-shot real images for training in the dataset SUN397, according to the curve 310 of performance scores over non-synthetic image number for a CT with initial strategy (classifier weights initialized from CLIP text embeddings) in accordance with some embodiments of the present disclosure, the CT with initial strategy performs comparably"), and determining the class as at least one weak class (Bai, Paragraph [0090], "In some example embodiments, to fine-tune the pre-trained model, the model training system 220 may filter the plurality of synthetic images by: for a first synthetic image associated with a first class, determining a first feature similarity between the first synthetic image and a first non-synthetic image in the first class"). As explained in rejection of claim 1, the obviousness for combining of Bai into Fashandi is provided above. But the combination does not explicitly disclose performance threshold and determining class as at least one weak class based on threshold. However, Shreshtha teaches training the image processing ML model using the enhanced training set of images to generate the trained image processing ML model (Shreshtha, Paragraph [0010], "and training a machine learning model using the training dataset."), and evaluating performance of the trained image processing ML model using test images and when the performance of the trained image processing ML model is below a performance threshold in respect of any class of the test images, determining that further additional synthetic images are required (Shreshtha, Paragraph [0091], "In other examples, a similarity distance threshold (e.g., a Euclidean distance measured between a seed embedding and a candidate embedding in any direction within the embedding space) may be used to determine the number of candidate embeddings 445 to extract for each respective seed embedding 435."), and determining the class as at least one weak class (Shreshtha, Paragraph [0016], "In an example of the example preceding method, the method may further include: obtaining a second set of seed data objects based on a second identified desired attribute; applying an embedding transformation to each of the seed data objects in the second set of seed data objects to create a second modified set of seed data objects; retrieving a second plurality of candidates from the database of data objects based on similarity to the second modified set of seed data objects; using the LLM, annotating the second plurality of candidates based on the list of defined labels; appending the training dataset to include the second plurality of annotated candidates; and training the machine learning model using the training dataset."). Shreshtha and Fashandi are analogous since both deal with generating synthetic images for tail/weak classes and training models with enhanced datasets. Fashandi provided way of merging harvested samples for tail categories to generate enhanced dataset to train SGG model and improve recall rate. Shreshtha provided way of training target model based on synthetic images, evaluating top-1 accuracy on test dataset, and filtering synthetic images by feature similarity to determine weak samples.Therefore, it would have been obvious to incorporate evaluating performance using test images and determining performance below threshold and determining weak class taught by Shreshtha into modified invention of Fashandi such that training feedback evaluates performance using test images and when below threshold determines further synthetic required and weak class. The motivation is to improve recall rate and mean recall rate for rare tail categories and reduce noise and enhance robustness discussed by Shreshtha in Paragraph [0010][0091][0016]. Regarding Claim 11, the combination of Fashandi, Bai and Matamoros teaches the invention in Claim 10. The combination further teaches comprising successively the training feedback process and the weak class compensation process (Matamoros, Paragraph [0114], "In other embodiments, candidate embeddings 445 may be extracted iteratively, for example, the embedding retriever 440 may extract a defined/predetermined number of candidate embeddings 445 for providing to the annotator 460. In examples, the annotator 460 may evaluate whether the applied labels are representative of the desired attribute, and if they are, may instruct the embedding retriever 440 to proceed to extract another defined/predetermined number of candidates for annotation. This process may continue until the applied annotations are no longer representative of the desired attribute."), until it is determined in the training feedback process that no further additional synthetic images are required (Matamoros, Paragraph [0114], "This process may continue until the applied annotations are no longer representative of the desired attribute."), or until a training threshold number of iterations has been performed (Matamoros, Paragraph [0122], "Optionally, at operation 510, operations 502 to 508 may be repeated one or more times, for example, for a defined number of iterations, to add an additional plurality of annotated candidates to the training dataset."), or until a training threshold number of iterations has been performed (Matamoros, Paragraph [0051], "Training may be carried out iteratively until a convergence condition is met (e.g., a predefined maximum number of iterations has been performed, or the value outputted by the ML model is sufficiently converged with the desired target value), after which the ML model is considered to be sufficie"), comprising successively iterating (, Paragraph, "In this way, these steps may be repeated to produce a more performant trained ML model."). Matamoros and Fashandi are analogous since both of them are dealing with generating synthetic/enhanced training data and training ML models with iterative improvement for tail/weak classes. Fashandi provided way of merging filtered tail samples to generate enhanced dataset to train SGG model. Matamoros provided way of phase-wise training of target model with synthetic images and evaluating performance on diverse datasets. Therefore, it would have been obvious to incorporate successively iterating until no further required or until defined number of iterations taught by Matamoros into modified invention of Fashandisuch that training feedback and weak class compensation are successively iterated. The motivation is to produce more performant trained ML model and continue extraction until annotations no longer representative discussed by Matamoros in Paragraph [0114][0122][0051][0050]. Claim(s) 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Fashandi et al. (US 20240290119 A1, hereinafter Fashandi) in view of Bai et al. (US 20230334834 A1, hereinafter Bai) as applied to Claim 1 above and further in view of Zhang et al. (US 20240203106 A1, hereinafter Zhang). Regarding Claim 18, the combination of Fashandi and Bai teaches the invention in Claim 14. The combination further teaches wherein the image processing ML model comprises an image retrieval model (Fashandi, Paragraph, "Examples of such reasonings on the scene graph structure are visual question answering, image-captioning, image editing and retrieval, and visual grounding."), [[wherein the image retrieval model is for face recognition]]. But the combination does not explicitly disclose wherein the image retrieval model is for face recognition. However, Zhang teaches wherein the image processing ML model comprises an image retrieval model (Zhang, Paragraph, "With the development of computer technologies, the retrieval technology for retrieving specified resources from the Internet is no longer limited to text search, but further supports users in picture search. For example, a user can enter a query picture for retrieval, so that pictures similar to the query picture entered by the user can be found from a database."), wherein the image retrieval model is for face recognition (, Paragraph, "For face recognition tasks (including face recognition data sets MS1Mv3 and IJB-C), true acceptance rates (TARs) under different false acceptance rates (FARs) can be calculated"). Zhang and Fashandi analogous since both deal with image retrieval systems extracting features and searching similar pictures from database. Fashandi provided way of reasoning on scene graph for retrieval. Zhang provided way of query picture retrieval returning same or similar picture and evaluation on face recognition data sets. Therefore, it would have been obvious to incorporate face recognition retrieval taught by Zhang into modified invention of Fashandi such that image retrieval model is used for face recognition. The motivation is to support picture search for similar pictures and verify on large-scale retrieval sets including face recognition discussed by Zhang in Paragraph [0003], [0142], [0139], [0032] and [0121]. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. US 20250278816 A1 CUSTOM IMAGE AND CONCEPT COMBINER USING DIFFUSION MODELS US 20250265831 A1 BUILDING VISION-LANGUAGE MODELS USING MASKED DISTILLATION FROM FOUNDATION MODELS US 20250209309 A1 METHODS AND SYSTEMS FOR GENERATING LABELED TRAINING DATA US 20250139957 A1 IMAGE COMPRESSION USING OVER-FITTING AND TEXT-TO-IMAGE MODEL US 20240256866 A1 GENERATING AI DATASET USING 3D ENGINE US 20220391755 A1 SYSTEMS AND METHODS FOR VISION-AND-LANGUAGE REPRESENTATION LEARNING US 20220261984 A1 METHODS AND APPARATUS FOR GRADING IMAGES OF COLLECTABLES USING IMAGE SEGMENTATION AND IMAGE ANALYSIS US 10504004 B2 Systems and methods for deep model translation generation Any inquiry concerning this communication or earlier communications from the examiner should be directed to YUJANG TSWEI whose telephone number is (571)272-6669. The examiner can normally be reached 8:30am-5:30pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kent Chang can be reached on (571) 272-7667. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /YuJang Tswei/Primary Examiner, Art Unit 2614
Read full office action

Prosecution Timeline

Mar 31, 2025
Application Filed
Sep 10, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749275
DIRECT MANIPULATION OF IMPLICITLY DEFINED DIGITAL 3D SHAPES
2y 4m to grant Granted Sep 29, 2026
Patent 12743679
SYSTEMS AND METHODS FOR TEMPLATE IMAGE EDITS
2y 5m to grant Granted Sep 22, 2026
Patent 12743795
Determining Object Structure Using Camera Devices With Views Of Moving Objects
2y 4m to grant Granted Sep 22, 2026
Patent 12718420
INFORMATION PROCESSING DEVICE AND METHOD
2y 2m to grant Granted Aug 25, 2026
Patent 12675993
AUGMENTED, VIRTUAL AND MIXED-REALITY CONTENT SELECTION & DISPLAY FOR BANK NOTE
4y 4m to grant Granted Jul 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
84%
Grant Probability
99%
With Interview (+16.0%)
2y 3m (~8m remaining)
Median Time to Grant
Low
PTA Risk
Based on 464 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month