Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
Response to Amendment
This is in response to applicant’s amendment/response filed on 05/11/2026, which has been entered and made of record. Claims 1, 10, 16 have been amended. No claim has been cancelled. No claim has been added. Claims 1-20 are pending in the application.
Response to Arguments
Applicant’s arguments on 05/11/2026 have been fully considered but are moot because the arguments do not apply to any of the references being used in the current rejection.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-2, 4-6, 9-10, 12, 14-19 are rejected under 35 U.S.C. 103 as being unpatentable over Meng, Chenlin, et al. ("Sdedit: Guided image synthesis and editing with stochastic differential equations." arXiv preprint arXiv:2108.01073 (2021)) in view of Azarian Yazdi et al. (US Pub 2025/0166236 A1) and Batra et al. (US Pub 2023/0162413 A1).
As to claim 1, Meng discloses a method comprising: obtaining an input image and a fidelity parameter, wherein the input image depicts a semantic entity and the fidelity parameter indicates a level of fidelity to the input image (Meng, abstract, “The key challenge is balancing faithfulness to the user inputs (e.g., hand-drawn colored strokes) and realism of the synthesized images.” Fig. 1, Page 2, “Given an input image with user guidance input, such as a stroke painting or an image with stroke edits, we can add a suitable amount of noise to smooth out undesirable artifacts and distortions (e.g., unnatural details at stroke pixels), while still preserving the overall structure of the input user guide. We then initialize the SDE with this noisy input, and progressively remove the noise to obtain a denoised result that is both realistic and faithful to the user guidance input (see Fig. 2).” Faithfulness is corresponding to fidelity. Page 2, “given a user guide in a form of manipulating RGB pixels, SDEdit adds Gaussian noise to the guide and then run the reverse SDE to synthesize images. SDEdit naturally finds a trade-off between realism and faithfulness: when we add more Gaussian noise and run the SDE for longer, the synthesized images are more realistic but less faithful. We can use this observation to find the right balance between realism and faithfulness.” Page 2, “SDEdit achieves a better faithfulness score and outperforms the baselines by up to 83.73% on overall satisfaction score in user studies.” Page 8, “generates realistic images that preserve semantics of the input stroke painting.”);
adding noise to the input image based on the fidelity parameter to obtain an intermediate noise image (Meng, Page 2, “SDE-based generative models smoothly convert an initial Gaussian noise vector to a realistic image sample through iterative denoising” “We then initialize the SDE with this noisy input, and progressively remove the noise to obtain a denoised result that is both realistic and faithful to the user guidance input (see Fig. 2).”); and
generating, using an image generation model, a synthetic image based on the intermediate noise image, wherein the synthetic image includes a (Meng, Page 2, “SDEdit naturally finds a trade-off between realism and faithfulness: when we add more Gaussian noise and run the SDE for longer, the synthesized images are more realistic but less faithful. We can use this observation to find the right balance between realism and faithfulness.” Fig. 3, Page 4. Page 7, 5.1 STROKE-BASED IMAGE SYNTHESIS.).
Meng does not explicitly disclose “a level of semantic adherence to the input” and “the synthetic image includes a vectorizable depiction of the semantic entity according to a semantic similarity transferred from the input to the synthetic image based on the level of semantic adherence to the input indicated by the fidelity parameter.
Azarian Yazdi discloses “a level of semantic adherence to the input” (Azarian Yazdi, ¶0006, “the input indicating a semantic importance for each of at least one of the one or more words associated with the at least one of the one or more input elements” ¶0027, “semantic image generation using localized cross-attention offers solutions to limitations in current diffusion models. By tuning attention weights according to image regions, spatial alignment between visual features and corresponding text improves. This provides a technical advancement over inconsistent semantics due in part to dominant themes overpowering local details. Specifically, attenuating a highest semantic attention weight per patch reduces unrelated or unrealistic combinations within each region.” ¶0040, “Semantic map conditioning parameter 306 refers to a latent feature map that encodes semantic information about an image; such a conditioning parameter can provide additional guidance about content and style of a desired image beyond a text conditioning parameter 308.”);
“the synthetic image includes a vectorizable depiction of the semantic entity according to a semantic similarity transferred from the input to the synthetic image based on the level of semantic adherence to the input indicated by the fidelity parameter (¶0038, “a latent vector representation of the image (e.g., noisy image) becomes less noisy and more refined.” ¶0046, “This adjustment to the noise level refines the predicted noise amount, or a related conditioning signal, for each patch, thereby boosting its alignment with local semantics. By modifying or attenuating the dominant semantic concept per patch, the guidance module 330j-330j+n can reduce the interference from text tokens unrelated to the patch content.” ¶0048, “The guidance modules 402j-402j+n associate text semantics with spatial image regions through cross-attention, including generating cross-attention weights based on key vectors K derived from text tokens and query vectors Q derived from a patch.” ¶0049, “the scaled weight 424 tailors the semantic blending to emphasize fidelity to the localized inferred semantics within each patch 408 (e.g., image region).” ¶0050, “the scaled cross-attention weights 516 tailor the semantic blending to emphasize fidelity to the localized inferred semantics within each patch (e.g., image region) in accordance with a user designated or user selectable scaling.” ¶0060, “the produced image 608 illustrates may show improved or reduced precision in areas connected to the chosen element, such as depicting a more or less detailed dog depending on the semantic importance level selected.”).
Meng and Azarian Yazdi are considered to be analogous art because all pertain to image generation. It would have been obvious before the effective filing date of the claimed invention to have modified Meng with the features of “a level of semantic adherence to the input” and “the synthetic image includes a vectorizable depiction of the semantic entity according to a semantic similarity transferred from the input to the synthetic image based on the level of semantic adherence to the input indicated by the fidelity parameter.” as taught by Azarian Yazdi. The suggestion/motivation would have been in order to enhance each patch of the image to better resemble the respective object in the patch (Azarian Yazdi, ¶0028).
Meng does not explicitly disclose vectorizable.
However it is obvious to one of ordinary skill in the art because the stochastic
differential equation (SDE) is based on vector (Meng, Page 2, “SDE-based generative models smoothly convert an initial Gaussian noise vector to a realistic image sample through iterative denoising”).
Batra discloses vectorizable (Batra, abstract, “A stroke-guided vectorization system is described that generates, from an input sketch and guide image depicting an approximate vector representation of the sketch, an aligned guide image depicting an improved vector representation of the sketch.”).
Meng, Azarian Yazdi and Batra are considered to be analogous art because all pertain to image generation. It would have been obvious before the effective filing date of the claimed invention to have modified Meng with the features of “vectorizable” as taught by Batra. The suggestion/motivation would have been in order to generate an aligned vector representation of the input sketch to be output as the aligned guide image (Batra, ¶0003).
As to claim 2, claim 1 is incorporated and the combination of Meng, Azarian Yazdi and Batra discloses obtaining a detail parameter indicating a level of detail for the synthetic image (Meng, Page 2, “enabling people with or without artistic expertise to produce photo-realistic images from different levels of details.” Page 4, “The user provides a full resolution image x (g) in a form of manipulating RGB pixels, which we call a “guide”. The guide may contain different levels of details; a high-level guide contains only coarse colored strokes, a mid-level guide contains colored strokes on a real image, and a low-level guide contains image patches on a target image.”); and
generating style guidance based on the detail parameter, wherein the synthetic image is generated based on the style guidance and includes the level of detail indicated by the detail parameter (Fig. 1, Page 15, “Described in Appendix D.2, the human-stroke-simulation algorithm uses different numbers of colors to generate stroke guides with different levels of detail.” Table 4. Page 22. D.2 Page 24, “Fig. 35 presents the image generation results based on input stroke paintings with various levels of details.”).
As to claim 4, claim 2 is incorporated and the combination of Meng, Azarian Yazdi and Batra discloses weighting the style guidance based on the detail parameter (Meng, Page 3, “The overall training objective is a weighted sum over t of each individual learning objective Lt, and various weighting procedures have been discussed in Ho et al. (2020); Song et al. (2020; 2021).”).
As to claim 5, claim 2 is incorporated and the combination of Meng, Azarian Yazdi and Batra discloses providing the style guidance to the image generation model at a diffusion step selected based on the detail parameter (Meng, Abstract, “Stochastic Differential Editing (SDEdit), based on a diffusion model generative prior, which synthesizes realistic images by iteratively denoising through a stochastic differential equation (SDE).”. Page 2, “Similar to the closely related diffusion models (SohlDickstein et al., 2015; Ho et al., 2020), SDE-based generative models smoothly convert an initial Gaussian noise vector to a realistic image sample through iterative denoising”).
As to claim 6, claim 1 is incorporated and the combination of Meng, Azarian Yazdi and Batra discloses adding the noise comprises: selecting a noise level based on the fidelity parameter, wherein the noise level decreases as the fidelity parameter increases (Meng, Page 2, “SDE-based generative models smoothly convert an initial Gaussian noise vector to a realistic image sample through iterative denoising”. “Given an input image with user guidance input, such as a stroke painting or an image with stroke edits, we can add a suitable amount of noise to smooth out undesirable artifacts and distortions (e.g., unnatural details at stroke pixels), while still preserving the overall structure of the input user guide. We then initialize the SDE with this noisy input, and progressively remove the noise to obtain a denoised result that is both realistic and faithful to the user guidance input (see Fig. 2).” Page 2, “Given an input image with user guidance input, such as a stroke painting or an image with stroke edits, we can add a suitable amount of noise to smooth out undesirable artifacts and distortions (e.g., unnatural details at stroke pixels), while still preserving the overall structure of the input user guide. We then initialize the SDE with this noisy input, and
progressively remove the noise to obtain a denoised result that is both realistic and faithful to the user guidance input (see Fig. 2).”)
As to claim 9, claim 1 is incorporated and the combination of Meng, Azarian Yazdi and Batra discloses generating a vector image based on the synthetic image (Meng, Page 2, “SDE-based generative models smoothly convert an initial Gaussian noise vector to a realistic image sample through iterative denoising” Batra, ¶0024, “the guide image 112 is a vector image that includes an artist's representation of the input sketch 110.”).
As to claim 10, the combination of Meng, Azarian Yazdi and Batra discloses a non-transitory computer readable medium storing code, the code comprising instructions executable by a processor to: obtain an input image, a style text, a fidelity parameter, and a detail parameter; add noise to the input image based on the fidelity parameter to obtain an intermediate noise image; generate a style guidance based on the style text and the detail parameter; and generate a synthetic image based on the intermediate noise image and the style guidance, wherein the synthetic image has a level of fidelity to the input image indicated by the fidelity parameter and has a level of detail indicated by the detail parameter (See claim 1 for detailed analysis. See Azarian Yazdi for style text, ¶0040, “Representation conditioning parameter 310 refers to encoded representations of images that capture style, text, composition, etc., which allow a model to replicate elements from a reference image. Image conditioning parameter 312 refers to representations (e.g., in pixel space) of images that capture style, text, composition, etc., which allow a model to replicate elements from a reference image.”).
As to claim 12, claim 10 is incorporated and the combination of Meng, Azarian Yazdi and Batra discloses the code further comprising instructions executable by the processor to: weight the style guidance based on the detail parameter (See claim 4 for detailed analysis.).
As to claim 14, claim 10 is incorporated and the combination of Meng, Azarian Yazdi and Batra discloses the code further comprising instructions executable by the processor to: provide the style guidance to the image generation model at a diffusion step selected based on the detail parameter (See claim 5 for detailed analysis.).
As to claim 15, claim 10 is incorporated and the combination of Meng, Azarian Yazdi and Batra discloses the code further comprising instructions executable by the processor to: generate a vector image based on the synthetic image (See claim 9 for detailed analysis.).
As to claim 16, the combination of Meng, Azarian Yazdi and Batra discloses an apparatus comprising: at least one processor; at least one memory storing instructions executable by the at least one processor; and the apparatus further comprising an image generation model comprising parameters stored in the at least one memory and configured to obtain an input image and a fidelity parameter, wherein the input image depicts a semantic entity and the fidelity parameter indicates a level of semantic adherence to the input image, add noise to an input image based on the fidelity parameter to obtain an intermediate noise image, and generate a synthetic image based on the intermediate noise image, wherein the synthetic image includes a vectorizable depiction of the semantic entity according to a semantic similarity transferred from the input image to the synthetic image based on level of semantic adherence to the input image as indicated by a fidelity parameter and has a level of detail as indicated by a detail parameter (See claim 1 for detailed analysis.).
As to claim 17, claim 16 is incorporated and the combination of Meng, Azarian Yazdi and Batra discloses a style prior model configured to generate a style guidance based on the detail parameter (See claim 4 for detailed analysis.).
As to claim 18, claim 16 is incorporated and the combination of Meng, Azarian Yazdi and Batra discloses a vectorization component configured to generate a vector image based on the synthetic image (See claim 9 for detailed analysis.).
As to claim 19, claim 16 is incorporated and the combination of Meng, Azarian Yazdi and Batra discloses the image generation model comprises a latent diffusion model (See claim 5 for detailed analysis.).
Claims 3 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Meng, Chenlin, et al. ("Sdedit: Guided image synthesis and editing with stochastic differential equations." arXiv preprint arXiv:2108.01073 (2021)) in view of Azarian Yazdi et al. (US Pub 2025/0166236 A1), Batra et al. (US Pub 2023/0162413 A1) and Chen et al. (US Pub 2025/0322570 A1).
As to claim 3, claim 2 is incorporated and the combination of Meng and Batra does not disclose obtaining a text prompt; and augmenting the text prompt based on the detail parameter, wherein the style guidance is generated based on the augmented text prompt.
Chen teaches obtaining a text prompt; and augmenting the text prompt based on the detail parameter, wherein the style guidance is generated based on the augmented text prompt (Chen, ¶0002, “receiving, at a client device, a style prompt; constructing, via a prompt construction unit, a first prompt by appending a font mask of a reference character and the style prompt to a first instruction string, the first instruction string including instructions to a first text-to-image model to iteratively generate salient content based on the style prompt and concentrate the salient content within the font mask of the reference character as a first image of the reference character; providing as an input the first prompt to the first text-to-image model and receiving as an output the first image from the first text-to-image model” ¶0015, “A shape-adaptive attention scheme is applied during inference to improve prompt fidelity. The generation model can interpret a given shape of a character and strategically plans pixel distributions within an irregular canvas. To achieve this, the generation model curates a high-quality shape-adaptive image-text dataset and incorporates the segmentation mask as a visual condition to steer the image generation process within the irregular-shaped canvas.” ¶0082, “the style prompt is a text prompt (e.g., the text prompt 150 in FIG. 1B: croissant) or an image prompt (e.g., the Easter egg image prompt in FIG. 2B).”).
Meng, Azarian Yazdi, Batra and Chen are considered to be analogous art because all pertain to image generation. It would have been obvious before the effective filing date of the claimed invention to have modified Meng with the features of “obtaining a text prompt; and augmenting the text prompt based on the detail parameter, wherein the style guidance is generated based on the augmented text prompt” as taught by Chen. The suggestion/motivation would have been in order curates a high-quality shape-adaptive image-text dataset and incorporates the segmentation mask as a visual condition to steer the image generation process within the irregular-shaped canvas (Chen, ¶0015).
As to claim 11, claim 10 is incorporated and the combination of Meng, Azarian Yazdi, Batra and Mohamed Ghouse discloses the code further comprising instructions executable by the processor to: obtain a text prompt; and augment the text prompt based on the detail parameter, wherein the style guidance is generated based on the augmented text prompt (See claim 3 for detailed analysis.).
Claims 7-8, 13, 20 are rejected under 35 U.S.C. 103 as being unpatentable over Meng, Chenlin, et al. ("Sdedit: Guided image synthesis and editing with stochastic differential equations." arXiv preprint arXiv:2108.01073 (2021)) in view of Azarian Yazdi et al. (US Pub 2025/0166236 A1), Batra et al. (US Pub 2023/0162413 A1) and Mohamed Ghouse et al. (US Pub 2024/0121398 A1).
As to claim 7, claim 1 is incorporated and the combination of Meng and Batra does not disclose generating the synthetic image comprises: selecting a diffusion sampling schedule based on the fidelity parameter.
Mohamed Ghouse teaches selecting a diffusion sampling schedule based on the fidelity parameter (Mohamed, ¶0162, “the input can include user input received via a user interface that can be used to configure the machine learning system. In some cases, a user can provide input to a graphical user interface to modify a sampling schedule of the diffusion model 526 (providing a user-friendly knob to control the perception-distortion tradeoff), which can increase or decrease the number of sampling steps of the diffusion model 526 (e.g., from 250 to 100 sampling steps, from 50 to 100 sampling steps, etc.). In one illustrative example, the user input can indicate a specific number of steps, which corresponds to a specific perceptual quality-fidelity trade-off. In another illustrative example, the user input can indicate a desired perceptual quality, a desired fidelity, or a desired perceptual quality-fidelity trade-off, and based on the input, the system can determine the number of steps needed to satisfy the perceptual quality, fidelity, or perceptual quality-fidelity trade-off.”).
Meng, Azarian Yazdi, Batra and Mohamed Ghouse are considered to be analogous art because all pertain to image generation. It would have been obvious before the effective filing date of the claimed invention to have modified Meng with the features of “selecting a diffusion sampling schedule based on the fidelity parameter” as taught by Mohamed Ghouse. The suggestion/motivation would have been in order to indicate a specific number of steps, which corresponds to a specific perceptual quality-fidelity trade-off (Mohamed Ghouse, ¶0162).
As to claim 8, claim 7 is incorporated and the combination of Meng, Azarian Yazdi, Batra and Mohamed Ghouse discloses an initial diffusion step of the diffusion sampling schedule increases as the fidelity parameter increases (Mohamed, ¶0162, “the input can include user input received via a user interface that can be used to configure the machine learning system. In some cases, a user can provide input to a graphical user interface to modify a sampling schedule of the diffusion model 526 (providing a user-friendly knob to control the perception-distortion tradeoff), which can increase or decrease the number of sampling steps of the diffusion model 526 (e.g., from 250 to 100 sampling steps, from 50 to 100 sampling steps, etc.). In one illustrative example, the user input can indicate a specific number of steps, which corresponds to a specific perceptual quality-fidelity trade-off. In another illustrative example, the user input can indicate a desired perceptual quality, a desired fidelity, or a desired perceptual quality-fidelity trade-off, and based on the input, the system can determine the number of steps needed to satisfy the perceptual quality, fidelity, or perceptual quality-fidelity trade-off.”).
As to claim 13, claim 10 is incorporated and the combination of Meng, Azarian Yazdi, Batra and Mohamed Ghouse discloses the code further comprising instructions executable by the processor to: select a diffusion sampling schedule based on the fidelity parameter (See claim 7 for detailed analysis.).
As to claim 20, claim 16 is incorporated and the combination of Meng, Azarian Yazdi, Batra and Mohamed Ghouse discloses a user interface including a fidelity parameter element and a detail parameter element (Meng, abstract, “faithfulness to the user inputs” Page 1, “a user specifies a general guide” Page 5, “ask the user whether the sample should be more faithful or more realistic; from the responses, we can obtain a reasonable t0 via binary search.” Page 2, “enabling people with or without artistic expertise to produce photo-realistic images from different levels of details.” Mohamed Ghouse, ¶0162, “providing a user-friendly knob to control the perception-distortion tradeoff”” the user input can indicate a desired perceptual quality, a desired fidelity, or a desired perceptual quality-fidelity trade-off”).
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to YU CHEN whose telephone number is (571)270-7951. The examiner can normally be reached on M-F 8-5 PST Mid-day flex.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Xiao Wu can be reached on 571-272-7761. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/YU CHEN/
Primary Examiner, Art Unit 2613