Prosecution Insights
Last updated: October 02, 2026
Application No. 19/088,259

METHOD AND APPARATUS WITH VISUAL MEDIUM GENERATION

Non-Final OA §103
Filed
Mar 24, 2025
Priority
May 14, 2024 — RE 10-2024-0063381 +1 more
Examiner
TSWEI, YU-JANG
Art Unit
Tech Center
Assignee
Samsung Electronics Co., Ltd.
OA Round
1 (Non-Final)
84%
Grant Probability
Favorable
1-2
OA Rounds
8m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 84% — above average
84%
Career Allowance Rate
388 granted / 464 resolved
+23.6% vs TC avg
Strong +16% interview lift
Without
With
+16.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 3m
Avg Prosecution
44 currently pending
Career history
507
Total Applications
across all art units

Statute-Specific Performance

§101
5.9%
-34.1% vs TC avg
§103
72.8%
+32.8% vs TC avg
§102
6.0%
-34.0% vs TC avg
§112
7.4%
-32.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 464 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed in parent Application No. KR10-2024-0063381, KR10-2024-0086393, filed on 05/14/2024 and 07/01/2024 respectively. Should applicant desire to obtain the benefit of foreign priority under 35 U.S.C. 119(a)-(d) prior to declaration of an interference, a certified English translation of the foreign application must be submitted in reply to this action. 37 CFR 41.154(b) and 41.202(e). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-11, 15-19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Mitchell et al. (US 20250022185 A1, hereinafter Mitchell), in view of Zhang et al. (US 20220253981 A1, hereinafter Zhang). Regarding Claim 1, Mitchell teaches A processor-implemented method comprising: (Mitchell, Paragraph [0013], “The presently disclosed systems and methods solve these and other problems associated with text-to-image generators”; [0058], " a computer system 500 in which at least some operations described herein can be implemented. As shown, the computer system 500 can include: one or more processors 502, main memory 506, non-volatile memory 510"). obtaining a plurality of prompts indicating image quality with different levels; (Mitchell, Paragraph [0028], "the image tuning system can accept multiple prompts for generation of a single output (or a single related collection of outputs). For example, the image tuning system can accept multiple text prompts arising from different users and generate one or more corresponding images through the image generation model."; [0035], "the image tuning system 114 can generate a modified prompt 218 as a seed artifact for further image generation. As an illustrative example, the image generation model can generate a third image 220 based on the modified prompt 218."; [0034], "the image tuning system 114 can generate a control token based on the determined accuracy or aesthetics of images generated by the image generation model. A control token can include a token (as described above) that aids in further tuning or controlling the properties of subsequent generated images."; it is noted the original prompt 202 and the modified prompt 218 (and further iterations) constitute a plurality of prompts each associated with (indicating) different targeted quality levels for the generated image; the control tokens explicitly encode quality-related properties. (Mitchell, Paragraph [0027], "Negative prompts can include indications of aesthetic-based or accuracy-based qualities that are undesired in the requested media. For example, a negative prompt includes aesthetic information relating to an unwanted color, mood, or style within the image."; it is noted aesthetic/accuracy quality descriptors are embedded in the prompts themselves. generating a plurality of visual media of the same content corresponding to the respective prompts by applying the obtained prompts to a visual medium generation model, (Mitchell, Paragraph [0045], "the image tuning system 114 generates a second image approximating the first image by applying an image generation model to the image generation prompt. As an illustrative example, the system can input the text string corresponding to the image generation prompt into an image generation model (e.g., a text-to-image generator) in order to generate a digital image that includes the elements requested by the user based on the image generation prompt."; [0035], "the image generation model can generate a third image 220 based on the modified prompt 218."; [0037], "the image generation model 303 can include algorithms that include variational autoencoders, flow-based models, generative adversarial networks, and/or diffusion-based models. As such, in some cases, the image generation model can accept a variety of input data types, such as text strings, other images, audio, video, or other seed artifacts. As output, the image generation model 303 can generate images based on the one or more inputs, such as a generated image 304."; it is noted the second image 304 and third image 220 approximate the same requested content but at different quality levels, read on "same content.") But Mitchell does not explicitly disclose the visual medium generation model is trained based on a loss function related to a level of image quality evaluated for an output visual medium and a level of image quality indicated by an input prompt. However, Zhang teaches wherein the visual medium generation model is trained based on a loss function related to a level of image quality evaluated for an output visual medium and a level of image quality indicated by an input prompt (Zhang, Paragraph [0071], "the image enhancement model may adopt a convolutional neural network or a generative adversarial network <read on visual medium generation model>."; "A high-quality image serving as a reference image may be used to calculate the value of a loss function. The value of the loss function is continuously updated until the difference between the enhanced image <read on output visual medium> and the reference image <read on level of image quality indicated by an input prompt> is minimum."; [0070], "utilizing the plurality of sets training data to train the image enhancement model until the difference feature value between the output of the image enhancement model and the data of the second image is minimum <read on loss function related to level of image quality evaluated for an output visual medium and a level of image quality indicated by an input prompt>."); [0061], "Parameters that measure image quality may include resolution, a signal-to-noise ratio, and hue difference. The image quality of the second image is higher than that of the first image."; it is noted CNN/GAN generator is trained with an explicit loss function whose value quantifies the difference between the output-image quality and the target quality level defined by the reference image (measured on resolution/SNR/hue). Zhang and Mitchell are analogous since both train image-producing machine-learning models against explicit image-quality objectives (resolution, contrast/hue, aesthetic/perceptual attributes). Mitchell provided a way of driving image generation from prompts that carry quality-related descriptors and iteratively evaluating output accuracy/aesthetic against quality targets. Zhang provided a way of training the generative model itself with an explicit loss (difference-feature-value minimization) that ties the output's quality level to a quality level specified through the training pair. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to incorporate Zhang's difference-feature-value loss training scheme into Mitchell's prompt-conditioned image generation model such that Mitchell's text-to-image generator is trained with a loss coupling the evaluated output-image quality to the quality level indicated by the input prompt's quality descriptors, rather than relying on inference-time iterative re-prompting alone. The motivation is to improve image quality intrinsically at the model level as Zhang discussed in Paragraph [0006]. Regarding Claim 2, the combination of Mitchell and Zhang teaches the invention in Claim 1. The combination further teaches the training of the visual medium generation model comprises fine-tuning a previously trained generative model [[based on the loss function]] (Mitchell, Paragraph [0037], "the image generation model 303 can include algorithms that include variational autoencoders, flow-based models, generative adversarial networks, and/or diffusion-based models <read on previously trained generative model>."); (Mitchell, Paragraph [0049], "the image tuning system 114 can fine-tune generated images such that they are pleasing to look at and of higher quality, thereby improving the performance and quality of image generation tasks."). But Mitchell does not explicitly disclose [[ fine-tuning a previously trained generative model ]] based on the loss function. However, Zhang teaches fine-tuning a previously trained generative model based on the loss function (Zhang, Paragraph [0071], "the image enhancement model may adopt a convolutional neural network or a generative adversarial network <read on previously trained generative model>. The image enhancement model contains thousands of parameters."; [0071], "A high-quality image serving as a reference image may be used to calculate the value of a loss function. The value of the loss function is continuously updated until the difference between the enhanced image and the reference image is minimum <read on fine-tuning ... based on the loss function>."). Zhang and Mitchell are analogous since both train image-generating neural networks against explicit image-quality objectives. Mitchell provided a way of iteratively steering pre-trained generative models (VAE/GAN/diffusion) toward higher output quality via prompt refinement at inference time. Zhang provided a way of updating a CNN/GAN generator's parameters through a loss-function value continuously minimized against a quality-reference image. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to fine-tune Mitchell's pre-trained VAE/GAN/diffusion-based image generation model with Zhang's loss-function-driven training so that the pre-trained generator learns to reduce the quality gap against the quality-indicated reference at training time. The motivation is to internalize quality adherence into the generator's weights as Zhang discussed in Paragraph [0071]. Regarding Claim 3, the combination of Mitchell and Zhang teaches the invention in Claim 1. The combination further teaches the loss function is determined using a determination network (Mitchell, Paragraph [0013], "the systems disclosed herein evaluate both an accuracy value and an aesthetic value of images generated from a text-to-image generator using, for example, an image classification convolutional neural network (CNN) <read on determination network>."); trained to determine a difference between the level of the image quality evaluated for the output visual medium and the level of the image quality indicated by the input prompt (Mitchell, Paragraph [0038], "the aesthetic evaluation model 308 can quantitatively evaluate the generated image 304 based on a variety of aesthetic-related factors... the aesthetic evaluation model 308 can include any machine learning models or algorithms, such as artificial neural networks"; [0039], "the accuracy evaluation model 306 can utilize an image recognition model to generate probability values that the generated image 304 includes one or more objects or elements from the text prompt 302 <read on level of image quality indicated by the input prompt>. The image recognition model and/or the accuracy evaluation model 306 can include supervised or unsupervised machine learning algorithms, such as deep convolutional networks"). Regarding Claim 4, the combination of Mitchell and Zhang teaches the invention in Claim 1. The combination further teaches the prompt comprises level information about one or more image quality elements (Mitchell, Paragraph [0025], "the image generation prompt 202 includes a descriptive token such as the word 'fun,' which describes an aesthetic, subjective, and/or emotional quality of the image requested by the user. An aesthetic quality can include any indication of a trait, characteristic, or attribute of an object (e.g., an image) that characterizes the pleasing nature or appearance of the image"; [0027], "a negative prompt includes aesthetic information relating to an unwanted color, mood, or style within the image."; [0050], "the aesthetic evaluation model is trained to output aesthetic metrics based on at least one of depth of field, lighting, position, focus, contrast, color, and brightness <read on image quality elements>."; it is noted the prompt carries descriptive/negative tokens that specify particular image-quality element levels (e.g., "fun" style, unwanted color/mood, aesthetic descriptors mapping onto contrast/lighting/color). Regarding Claim 5, the combination of Mitchell and Zhang teaches the invention in Claim 1. The combination further teaches a first prompt of the prompts comprises first level information about a first image quality element, and a second prompt of the prompts comprises second level information about the first image quality element (Mitchell, Paragraph [0023], " the image generation prompt 202 can include a request from a user for a generated image corresponding to a “person wearing a fun hat.””; [0034], " the system can generate the control token 214 to include the phrase or word “on their head”... Additionally or alternatively, the image tuning system 114 generates negative prompts to handle situations where the generated image includes elements that are undesired. For example, the image tuning system 114 generates a token that includes a negative prompt, such as the text string “arctic animals””; [0035], "the image tuning system 114 can generate a modified prompt 218 as a seed artifact for further image generation."; it is noted original prompt (e.g., 'fun hat') and modified prompt (with added control token specifying stronger/different aesthetic level for the same aesthetic quality element) constitute a first prompt and a second prompt each carrying different level information about the same image quality element (aesthetic). Regarding Claim 6, the combination of Mitchell and Zhang teaches the invention in Claim 1. The combination further teaches a first prompt of the prompts comprises level information about a first image quality element, and a second prompt of the prompts comprises level information about a second image quality element (Mitchell, Paragraph [0050], "the aesthetic evaluation model is trained to output aesthetic metrics based on at least one of depth of field, lighting, position, focus, contrast, color, and brightness <read on plurality of image quality elements>."; [0034], "the image tuning system can determine an additional token (e.g., the control token 214) that includes any missing properties for the second image... Additionally or alternatively, the image tuning system 114 generates negative prompts to handle situations where the generated image includes elements that are undesired."; [0027], "a negative prompt includes aesthetic information relating to an unwanted color, mood, or style within the image."). Regarding Claim 7, the combination of Mitchell and Zhang teaches the invention in Claim 1. The combination further teaches obtaining training data of a visual medium-based model based on one or more of the prompts and visual media (Mitchell, Paragraph [0030], "the image tuning system can generate training images and corresponding image descriptions through the users' prompts and any generated or provided images, thereby enabling further improvement or training of the image labeling model."; [0038], "the aesthetic evaluation model 308 can accept training data of images (e.g., original or altered) and corresponding aesthetic descriptors or tokens."). Mitchell does not explicitly disclose but Zhang teaches training the visual medium-based model based on the training data (Zhang, Paragraph [0070], "utilizing the plurality of sets training data to train the image enhancement model <read on visual medium-based model> until the difference feature value between the output of the image enhancement model and the data of the second image is minimum."; Paragraph [0126], "N sets of images are employed to train an image enhancement model."). Zhang and Mitchell are analogous since both feed curated (image, quality-descriptor) training data to a downstream image-processing neural network. Mitchell provided a way of producing (prompt, generated-image) pairs and quality descriptors suitable as training data. Zhang provided a way of consuming plural training-data sets to train an image-processing neural network via a difference-minimizing procedure. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to feed Mitchell's (prompt, generated-image) training data into Zhang's training procedure so that the downstream visual-medium-based model is actually trained end-to-end on the curated pairs. The motivation is to actually realize a functioning downstream model, per Zhang Paragraph [0126]. Regarding Claim 8, it recites limitations similar in scope to the limitations of claim 1 and the combination of Mitchell and Zhang teaches all the limitations as of Claim 1. And Mitchell discloses these features can be implemented on a computer readable storage medium (Mitchell, Paragraph [0061], "the machine-readable medium 526 can include any medium that is capable of storing, encoding, or carrying a set of instructions for execution by the computing system 500. The machine-readable medium 526 can be non-transitory or comprise a non-transitory device."); (Mitchell, Paragraph [0062], "Examples of machine-readable storage media, machine-readable media, or computer-readable media include recordable-type media such as volatile and non-volatile memory 510, removable flash memory, hard disk drives, optical disks"); (Mitchell, Paragraph [0063], "When read and executed by the processor 502, the instruction(s) cause the computing system 500 to perform operations to execute elements involving the various aspects of the disclosure."). Regarding Claim 9, Mitchell teaches a processor-implemented method comprising: (Mitchell, Paragraph [0013], “The presently disclosed systems and methods solve these and other problems associated with text-to-image generators”; [0058], " a computer system 500 in which at least some operations described herein can be implemented. As shown, the computer system 500 can include: one or more processors 502, main memory 506, non-volatile memory 510"). obtaining a plurality of prompts indicating image quality with different levels; (Mitchell, Paragraph [0028], "the image tuning system can accept multiple prompts for generation of a single output (or a single related collection of outputs). For example, the image tuning system can accept multiple text prompts arising from different users and generate one or more corresponding images through the image generation model."; [0035], "the image tuning system 114 can generate a modified prompt 218 as a seed artifact for further image generation. As an illustrative example, the image generation model can generate a third image 220 based on the modified prompt 218."; [0034], "the image tuning system 114 can generate a control token based on the determined accuracy or aesthetics of images generated by the image generation model. A control token can include a token (as described above) that aids in further tuning or controlling the properties of subsequent generated images."; it is noted the original prompt 202 and the modified prompt 218 (and further iterations) constitute a plurality of prompts each associated with (indicating) different targeted quality levels for the generated image; the control tokens explicitly encode quality-related properties. (Mitchell, Paragraph [0027], "Negative prompts can include indications of aesthetic-based or accuracy-based qualities that are undesired in the requested media. For example, a negative prompt includes aesthetic information relating to an unwanted color, mood, or style within the image."; it is noted aesthetic/accuracy quality descriptors are embedded in the prompts themselves) generating a plurality of visual media of the same content corresponding to the respective prompts by applying the obtained prompts to a visual medium generation model; (Mitchell, Paragraph [0045], "the image tuning system 114 generates a second image approximating the first image by applying an image generation model to the image generation prompt. As an illustrative example, the system can input the text string corresponding to the image generation prompt into an image generation model (e.g., a text-to-image generator) in order to generate a digital image that includes the elements requested by the user based on the image generation prompt."; [0035], "the image generation model can generate a third image 220 based on the modified prompt 218."; [0037], "the image generation model 303 can include algorithms that include variational autoencoders, flow-based models, generative adversarial networks, and/or diffusion-based models. As such, in some cases, the image generation model can accept a variety of input data types, such as text strings, other images, audio, video, or other seed artifacts. As output, the image generation model 303 can generate images based on the one or more inputs, such as a generated image 304."; it is noted the second image 304 and third image 220 approximate the same requested content but at different quality levels, read on "same content."). generating training data of a visual medium-based model based on one or more of the prompts and visual media; (Mitchell, Paragraph [0030], "The image labeling model can be trained with training images and corresponding image descriptions. For example, the image tuning system can generate training images and corresponding image descriptions through the users' prompts and any generated or provided images, thereby enabling further improvement or training of the image labeling model."); (Mitchell, Paragraph [0038], "the aesthetic evaluation model 308 can accept training data of images (e.g., original or altered) and corresponding aesthetic descriptors or tokens."). Mitchell explicitly uses (prompt, generated image) pairs together with quality descriptors as training data for a downstream evaluation/labeling model) [[training the visual medium-based model based on the training data]] to some extent (Mitchell, Paragraph [0030], "thereby enabling further improvement or training of the image labeling model"); (Mitchell, Paragraph [0050], "the aesthetic evaluation model is trained to output aesthetic metrics based on at least one of depth of field, lighting, position, focus, contrast, color, and brightness."). Mitchell describes training accuracy/aesthetic evaluation models from generated images and descriptors, but to the extent claim 9 is read to require a training procedure with a specific loss based on the generated pairs, Zhang further supports this. But Mitchell does not explicitly disclose the full training pipeline for the visual medium-based model with the loss based on the generated pairs. However, Zhang teaches wherein the visual medium generation model is trained based on a loss function related to a level of image quality evaluated for an output visual medium and a level of image quality indicated by an input prompt (Zhang, Paragraph [0071], "the image enhancement model may adopt a convolutional neural network or a generative adversarial network <read on visual medium generation model>."; "A high-quality image serving as a reference image may be used to calculate the value of a loss function. The value of the loss function is continuously updated until the difference between the enhanced image <read on output visual medium> and the reference image <read on level of image quality indicated by an input prompt> is minimum."; [0070], "utilizing the plurality of sets training data to train the image enhancement model until the difference feature value between the output of the image enhancement model and the data of the second image is minimum <read on loss function related to level of image quality evaluated for an output visual medium and a level of image quality indicated by an input prompt>."); [0061], "Parameters that measure image quality may include resolution, a signal-to-noise ratio, and hue difference. The image quality of the second image is higher than that of the first image."; it is noted CNN/GAN generator is trained with an explicit loss function whose value quantifies the difference between the output-image quality and the target quality level defined by the reference image (measured on resolution/SNR/hue). Zhang and Mitchell are analogous since both train image-producing machine-learning models against explicit image-quality objectives (resolution, contrast/hue, aesthetic/perceptual attributes). Mitchell provided a way of driving image generation from prompts that carry quality-related descriptors and iteratively evaluating output accuracy/aesthetic against quality targets. Zhang provided a way of training the generative model itself with an explicit loss (difference-feature-value minimization) that ties the output's quality level to a quality level specified through the training pair. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to incorporate Zhang's difference-feature-value loss training scheme into Mitchell's prompt-conditioned image generation model such that Mitchell's text-to-image generator is trained with a loss coupling the evaluated output-image quality (Mitchell's accuracy/aesthetic metric) to the quality level indicated by the input prompt's quality descriptors, rather than relying on inference-time iterative re-prompting alone. The motivation is to improve image quality intrinsically at the model level as Zhang discussed in Paragraph [0006]. Regarding Claim 10, the combination of Mitchell and Zhang teaches the invention in Claim 9. The combination further teaches applying the training data to the visual medium-based model and applying the prompts to a generative model (Mitchell, Paragraph [0037], "the image tuning system 114 can input text prompt 302 into an image generation model 303... the image generation model 303 can generate images based on the one or more inputs, such as a generated image 304."); (Mitchell, Paragraph [0038], "the aesthetic evaluation model 308 can accept training data of images (e.g., original or altered) and corresponding aesthetic descriptors or tokens."). Mitchell does not explicitly disclose but Zhang teaches training the visual medium-based model by on a loss determined based on a result of the applying of the training data to the visual medium-based model and a result of the applying of the prompts to the generative model (Zhang, Paragraph [0126], "A low-quality image, i.e., the first image is input into the image enhancement model <read on visual medium-based model> for generating an enhanced image, and a high-quality image, i.e., the second image is a real image for reference, and is used to calculate the value of a loss function <read on training ... by on a loss>."; [0126], "The value of the loss function is continuously updated until the difference between an enhanced image generated by the image enhancement model and the image for reference is minimum."). Zhang and Mitchell are analogous. Mitchell provided a way of running two parallel branches — a text-prompt-driven generative branch and an image-evaluation branch — that produce results usable for comparison. Zhang provided a way of computing a training loss from the two branches' results and updating the model until that loss is minimized. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine Mitchell's two-branch architecture with Zhang's loss-based training such that the visual medium-based model is trained on a loss computed from its own output and the corresponding prompt-driven generator's output. The motivation is to ground the downstream model's training in a quantitative supervision signal derived from both branches, as Zhang discussed in Paragraph [0126]. Regarding Claim 11, the combination of Mitchell and Zhang teaches the invention in Claim 9. The combination further teaches the visual medium-based model comprises a model that generates image quality evaluation data of an input visual medium (Mitchell, Paragraph [0038], "the aesthetic evaluation model 308 can quantitatively evaluate the generated image 304 based on a variety of aesthetic-related factors, such as depth of field, lighting, position, focus, contrast, color, and brightness of the image... the image tuning system 114 can generate a quantitative measure (e.g., the aesthetic metric 312 or 210, as described above) evaluating such factors."; [0051], "The aesthetic recognition model can be trained to output text describing aesthetic properties for input images."), and the generating of the training data comprises generating ground truth (GT) of image quality evaluation data of the visual medium generated in response to each of the prompts by applying each of the prompts to a generative model (Mitchell, Paragraph [0038], "the aesthetic evaluation model 308 can accept training data of images (e.g., original or altered) and corresponding aesthetic descriptors <read on GT of image quality evaluation data> or tokens."; [0030], "the image tuning system can generate training images and corresponding image descriptions through the users' prompts and any generated or provided images"). Regarding Claim 15, it recites limitations similar in scope to the limitations of claim 1, but in an apparatus. As shown in the rejection, the combination of Mitchell and Zhang disclose the limitations of claims 1. Additionally, Mitchell discloses an apparatus that maps to Paragraph [0063]-[0064] (Mitchell, Paragraph [0064], "the computer system 500 can include: one or more processors 502, main memory 506, non-volatile memory 510"); (Mitchell, Paragraph [0063], "When read and executed by the processor 502, the instruction(s) cause the computing system 500 to perform operations to execute elements involving the various aspects of the disclosure."). Thus, Claim 15 is met by the combination of Mitchell and Zhang according to the mapping presented in the rejection of claims 15, given the method corresponds to the apparatus. Regarding Claim 16, it recites limitations similar in scope to the limitations of Claim 2 and therefore is rejected under the same rationale. Regarding Claim 17, it recites limitations similar in scope to the limitations of Claim 3 and therefore is rejected under the same rationale. Regarding Claim 18, it recites limitations similar in scope to the limitations of Claim 7 and therefore is rejected under the same rationale. Regarding Claim 19, it recites limitations similar in scope to the limitations of Claim 11 and therefore is rejected under the same rationale. Claim(s) 12-14, 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Mitchell et al. (US 20250022185 A1, hereinafter Mitchell), in view of Zhang et al. (US 20220253981 A1, hereinafter Zhang) as applied to Claim 1, 15 above and further in view of Bakunov et al.(US 20240296535 A1, hereinafter Bakunov). Regarding Claim 12, the combination of Mitchell and Zhang teaches the invention in Claim 9. The combination further teaches the visual medium-based model comprises a model that generates image quality [[comparative]] evaluation data of an input visual medium (Mitchell, Paragraph [0038], "the aesthetic evaluation model 308 can quantitatively evaluate the generated image 304 based on a variety of aesthetic-related factors <read on image quality evaluation data of an input visual medium>."; [0039], "the accuracy evaluation model 306 can utilize an image recognition model to generate probability values that the generated image 304 includes one or more objects or elements from the text prompt 302."). [[the generating of the training data comprises generating GT of image quality comparative evaluation data of the visual media]] by applying the prompts to a generative model (Mitchell, Paragraph [0037], "the image tuning system 114 can input text prompt 302 into an image generation model 303 <read on applying the prompts to a generative model>… the image generation model 303 can generate images based on the one or more inputs."). But Mitchell does not explicitly disclose [[ the visual medium-based model comprises a model that generates image quality ]] comparative [[ evaluation data of an input visual medium, ]] and the generating of the training data comprises generating GT of image quality comparative evaluation data of the visual media by applying the prompts to a generative model. However, Bakunov teaches a model that generates image quality comparative evaluation data of an input visual medium (Bakunov, Paragraph [0135], "the first image (the image selected from the first set of images generated by the first automated image generator) and the second image (the image selected from the second set of images generated by the second automated image generator) are automatically compared by the image quality evaluation system 236 <read on model that generates image quality comparative evaluation data>. Comparisons and rankings <read on comparative quality data> (as described below) may be automatically carried out, e.g., by one or more machine learning models or by other computing components."; [0025], "A first ranking may be generated based on a comparison between the selected first image and the selected second image."). generating GT of image quality comparative evaluation data of the visual media by applying the prompts to a generative model (Bakunov, Paragraph [0023], "a first set of images generated by a first automated image generator and a second set of images generated by a second automated image generator are accessed. The first set of images and the second set of images are both automatically generated based on a first text prompt <read on applying the prompts to a generative model>. The automated image generators may be text-to-image machine learning models, such as diffusion models."; [0125], "A first machine learning model, trained to evaluate quality based on a first quality metric, is applied by the image quality evaluation system 236 to generate a first quality indicator for each of the images in the first set of images and each of the images in the second set of images <read on GT of image quality comparative evaluation data>"). Bakunov and Mitchell/Zhang are analogous since all three concern automated image-quality assessment/generation using ML models fed prompts and evaluating image quality. Mitchell/Zhang provided a way of producing (prompt, generated-image, quality-metric) training data and training a downstream image-processing model on that data. Bakunov provided a way of producing comparative quality data through rankings between images generated by different prompt for generator paths via automated ML-scoring and comparison. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to extend the Mitchell+Zhang pipeline with Bakunov's automated comparative-evaluation and ranking so that the downstream visual-medium-based model of Mitchell/Zhang is trained to output comparative evaluation data, using Bakunov-style rankings as ground truth. The motivation is to enable objective, human-perception-consistent model selection and inter-image comparison as Bakunov discussed in Paragraph [0022]. Regarding Claim 13, the combination of Mitchell and Zhang teaches the invention in Claim 9. The combination further teaches the visual medium-based model comprises a model that generates a visual medium with improved image quality of an input visual medium (Mitchell, Paragraph [0035], "the image generation model can generate a third image 220 based on the modified prompt 218. The image tuning system 114 can proceed to evaluate the third image 220 for its accuracy… In some implementations, the image tuning system 114 can repeat or iterate the process of evaluating the images and generating new prompts or seed artifacts for further image generation."); (Mitchell, Paragraph [0055], "the image tuning system can generate the seed artifact based on the second image… enabling the image generation model to iteratively improve the generated image"). the generating of the training data comprises generating an image quality improvement prompt for converting a visual medium generated in response to a first prompt into a visual medium corresponding to a second prompt (Mitchell, Paragraph [0034], "the image tuning system 114 can generate a control token based on the determined accuracy or aesthetics of images generated by the image generation model. A control token can include a token (as described above) that aids in further tuning or controlling the properties of subsequent generated images <read on image quality improvement prompt>."); (Mitchell, Paragraph [0035], "by generating control tokens based on the evaluation of the generated second image 204, the image tuning system 114 can generate a modified prompt 218 as a seed artifact for further image generation."). [[by applying the first prompt and the second prompt of the prompts to a generative model, and ]] training the visual medium-based model to output a visual medium generated in response to the second prompt based on a visual medium generated in response to the first prompt and the image quality improvement prompt (Zhang, Paragraph [0126], "A low-quality image, i.e., the first image is input into the image enhancement model <read on based on a visual medium generated in response to the first prompt> for generating an enhanced image, and a high-quality image, i.e., the second image is a real image for reference <read on visual medium generated in response to the second prompt>… continuously updated until the difference between an enhanced image generated by the image enhancement model and the image for reference is minimum <read on training the visual medium-based model to output a visual medium generated in response to the second prompt>."). As explained in rejection of claim 1, the obviousness for combining of Zhang into Mitchell is provided above. But Mitchell/Zhang do not explicitly disclose by applying the first prompt and the second prompt of the prompts to a generative model to obtain the image quality improvement prompt. However, Bakunov teaches by applying the first prompt and the second prompt of the prompts to a generative model (Bakunov, Paragraph [0023], "a first set of images generated by a first automated image generator and a second set of images generated by a second automated image generator are accessed. The first set of images and the second set of images are both automatically generated based on a first text prompt <read on applying prompts to a generative model>. The automated image generators may be text-to-image machine learning models, such as diffusion models."; [0145], "the second automated image generator (e.g., second generative model) generates a fourth set of images, also based on the second text prompt. Accordingly, each automated image generator automatically generates a further set of images, based on the second prompt <read on second prompt applied to a generative model Bakunov and Mitchell/Zhang are analogous. since all three concern automated image-quality assessment/generation using ML models fed prompts and evaluating image quality. Mitchell provided a way of generating an improvement (modified-prompt) instruction from evaluation feedback and iteratively converting a generated image toward improved quality. Zhang provided a way of training the conversion model with a paired loss (first-image → second-image reference). Bakunov provided a way of applying multiple text prompts to text-to-image generative models and comparing the resulting images to determine relative quality. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine Mitchell's modified-prompt (improvement-prompt) generation and Zhang's paired-loss training with Bakunov's dual-prompt / dual-generator comparison, so that the improvement prompt is generated by applying two prompts to a generative model and the conversion model is trained on the resulting pair to convert the first-prompt image into the second-prompt image. The motivation is to ground the improvement-prompt generation in an objective comparison across prompt-conditioned outputs rather than heuristic per-image thresholds, per Bakunov Paragraph [0022]. Regarding Claim 14, the combination of Mitchell and Zhang teaches the invention in Claim 9. The combination further teaches extracting two prompts among the prompts (Mitchell, Paragraph [0028], "the image tuning system can accept multiple prompts for generation of a single output (or a single related collection of outputs). For example, the image tuning system can accept multiple text prompts… the image tuning system can concatenate text strings corresponding to the various prompts… <read on extracting two prompts>"). generating the image quality improvement prompt for converting a visual medium generated in response to the first prompt indicating relatively lower image quality into a visual medium corresponding to the second prompt indicating relatively higher image quality (Mitchell, Paragraph [0033], "If either of these metrics is below the threshold metrics, the image tuning system 114 can determine to regenerate, improve, or tune the image to further improve its accuracy or aesthetics."); (Mitchell, Paragraph [0035], "by generating control tokens based on the evaluation of the generated second image 204, the image tuning system 114 can generate a modified prompt 218 as a seed artifact for further image generation <read on image quality improvement prompt for converting… into… higher image quality>."). But Mitchell/Zhang do not explicitly disclose based on a relative image quality superiority determination result of the two extracted prompts. However, Bakunov teaches based on a relative image quality superiority determination result of the two extracted prompts (Bakunov, Paragraph [0025], "A first ranking may be generated based on a comparison between the selected first image and the selected second image. The first ranking is a ranking of the first automated image generator and the second automated image generator <read on relative image quality superiority determination result of the two extracted prompts>."; [0134], "the image quality evaluation system 236 may automatically generate a combined score that takes the aesthetic, alignment and visual realism scores for each image into account, and the image with a highest combined score may be selected."; [0135], "the first image (the image selected from the first set of images generated by the first automated image generator) and the second image (the image selected from the second set of images generated by the second automated image generator) are automatically compared by the image quality evaluation system 236."; it is noted the ranking based on comparison of two prompt-generated images directly provides the relative-superiority signal identifying which of the two prompts corresponds to relatively higher / lower image quality). Bakunov and Mitchell/Zhang are analogous. since all three concern automated image-quality assessment/generation using ML models fed prompts and evaluating image quality Mitchell/Zhang provided a way of generating and training on modified-prompt improvement instructions once a quality gap is identified. Bakunov provided a way of automatically comparing two prompt-generated image sets and outputting a ranking that identifies the superior/inferior side. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to use Bakunov's ranking output as the trigger and directional signal for Mitchell's improvement-prompt generation, so that two extracted prompts of relatively lower and higher quality are identified via Bakunov's comparison, and Mitchell's modified-prompt generation is directed from lower toward higher. The motivation is to replace fixed per-image thresholds with an objective inter-prompt comparison, per Bakunov Paragraph [0022]. Regarding Claim 20, it recites limitations similar in scope to the limitations of Claim 12 and therefore is rejected under the same rationale. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. US 20240394835 A1 AUTOMATICALLY ENHANCING IMAGE QUALITY IN MACHINE LEARNING TRAINING DATASET BY USING DEEP GENERATIVE MODELS US 12167000 B2 Techniques for predicting video quality across different viewing parameters US 20240169537 A1 METHODS AND SYSTEMS FOR AUTOMATIC CT IMAGE QUALITY ASSESSMENT US 20240119575 A1 TECHNIQUES FOR GENERATING A PERCEPTUAL QUALITY MODEL FOR PREDICTING VIDEO QUALITY ACROSS DIFFERENT VIEWING PARAMETERS US 11763581 B1 Methods and apparatus for end-to-end document image quality assessment using machine learning without having ground truth for characters US 10916003 B2 Image quality scorer machine Any inquiry concerning this communication or earlier communications from the examiner should be directed to YUJANG TSWEI whose telephone number is (571)272-6669. The examiner can normally be reached 8:30am-5:30pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kent Chang can be reached on (571) 272-7667. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /YuJang Tswei/Primary Examiner, Art Unit 2614
Read full office action

Prosecution Timeline

Mar 24, 2025
Application Filed
Aug 26, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749275
DIRECT MANIPULATION OF IMPLICITLY DEFINED DIGITAL 3D SHAPES
2y 4m to grant Granted Sep 29, 2026
Patent 12743679
SYSTEMS AND METHODS FOR TEMPLATE IMAGE EDITS
2y 5m to grant Granted Sep 22, 2026
Patent 12743795
Determining Object Structure Using Camera Devices With Views Of Moving Objects
2y 4m to grant Granted Sep 22, 2026
Patent 12718420
INFORMATION PROCESSING DEVICE AND METHOD
2y 2m to grant Granted Aug 25, 2026
Patent 12675993
AUGMENTED, VIRTUAL AND MIXED-REALITY CONTENT SELECTION & DISPLAY FOR BANK NOTE
4y 4m to grant Granted Jul 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
84%
Grant Probability
99%
With Interview (+16.0%)
2y 3m (~8m remaining)
Median Time to Grant
Low
PTA Risk
Based on 464 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month