Prosecution Insights
Last updated: September 17, 2026
Application No. 19/196,755

METHOD, COMPUTING DEVICE, AND NON-TRANSITORY COMPUTER-READABLE RECORDING MEDIUM FOR GENERATING CUSTOMIZED IMAGES

Non-Final OA §103
Filed
May 02, 2025
Priority
Jan 09, 2025 — TW 114100845
Examiner
TRUONG, KARL DUC
Art Unit
2614
Tech Center
2600 — Communications
Assignee
Seidman Int'L Trading Ltd.
OA Round
1 (Non-Final)
65%
Grant Probability
Moderate
1-2
OA Rounds
1y 4m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 65% of resolved cases
65%
Career Allowance Rate
33 granted / 51 resolved
+2.7% vs TC avg
Strong +34% interview lift
Without
With
+34.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 8m
Avg Prosecution
24 currently pending
Career history
81
Total Applications
across all art units

Statute-Specific Performance

§101
1.4%
-38.6% vs TC avg
§103
87.4%
+47.4% vs TC avg
§102
7.6%
-32.4% vs TC avg
§112
1.8%
-38.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 51 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed in parent Application No. TW114100845, filed on 9th January, 2025. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 4, 6-8, 11, 13-15, 18, and 20-21 are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (US 20250117973 A1), hereinafter referenced as Chen, in view of Kharbanda et al. (US 20240378237 A1), hereinafter referenced as Kharbanda. Regarding Claim 1, Chen discloses a method for generating customized images (Chen, [0100]: teaches a method for generating synthetic images <read on customized images>), the method being executed after a computer program product is loaded and executed by a computing device (Chen, [0125]: teaches the method of an image generation model, which is performed by "a system <read on computing device> including a processor executing a set of codes <read on computer program product> to control functional elements of an apparatus such as the image generation model"), the method comprising the following steps of:receiving customization demand information (Chen, [0070]: teaches a text encoder 505 and style encoder 510 receiving text prompt 520 <read on customization demand information> and style input 525 respectively; Note: the text prompt and style input are both being interpreted as "customization demand information"); inputting the customization demand information into a semantic analysis model (Chen, [0070]: teaches the text encoder 505 <read on semantic analysis model> and the style encoder 510 receiving text prompt 520 <read on customization demand information> and style input 525 respectively); performing a semantic analysis on the customization demand information through the semantic analysis model (Chen, [0070]: teaches the text encoder 505 <read on semantic analysis model> and the style encoder 510 receiving text prompt 520 <read on customization demand information> and style input 525 respectively; Note: it should be noted that text encoders in neural networks inherently use semantic analysis), and generating and outputting analyzed customization information based on the customization demand information (Chen, [0070]: teaches the text encoder 505 and the style encoder 510 generating text embedding 530 <read on analyzed customization information> and style embedding 535 respectively, which are based on the inputted text prompt 520 <read on customization demand information> and style input 525); [[searching for a plurality of semantic images associated with the analyzed customization information from an image classification database, and]] [[receiving the plurality of semantic images, and/or]] inputting at least one of the customization demand information and the analyzed customization information into an image generation model (Chen, [0070]: teaches an image generation model 515 generating synthetic image 540 based on text embedding 530 <read on analyzed customization information> and style embedding 535 as shown in FIG. 5, where both text embedding 530 and style embedding 535 <read on customization demand information> are based on text prompt 520 and style input 525 respectively), and PNG media_image1.png 622 441 media_image1.png Greyscale generating and outputting at least one initial image based on at least one of the customization demand information and the analyzed customization information through the image generation model (Chen, [0034]: teaches the image generation system generating a synthetic image <read on initial image> "based on the text embedding <read on analyzed customization information> and the style embedding using the image generation model by providing the text embedding as an input to the image generation model during initial iterations of an image generation process and providing both the text embedding and the style embedding as inputs to the image generation model during subsequent iterations of the image generation process," where the text embedding and the style embedding are both based on the text prompt <read on customization demand information> and the style input respectively); [[inputting the plurality of semantic images and/or the at least one initial image into an image analysis and screening model; and]] [[performing image analysis and screening on the plurality of semantic images and/or the at least one initial image through the image analysis and screening model, and]] [[screening the plurality of semantic images and/or the at least one initial image based on at least one of the customization demand information and the analyzed customization information to obtain at least one first retrieved image, wherein]] the semantic analysis model, the image generation model [[and the image analysis and screening model]] are respectively trained artificial intelligence engines (Chen, [0176]: teaches a machine learning model 1915 <read on semantic analysis model> that is trained <read on trained AI engine> to generate a text embedding based on a text prompt describing image content, and generate a style embedding based on a style input describing an image style; [0210]: teaches a trained <read on trained AI engine> diffusion model <read on image generation model>). However, Chen does not expressly disclose searching for a plurality of semantic images associated with the analyzed customization information from an image classification database, and receiving the plurality of semantic images, and/or inputting the plurality of semantic images and/or the at least one initial image into an image analysis and screening model; and performing image analysis and screening on the plurality of semantic images and/or the at least one initial image through the image analysis and screening model, and screening the plurality of semantic images and/or the at least one initial image based on at least one of the customization demand information and the analyzed customization information to obtain at least one first retrieved image, wherein the semantic analysis model, the image generation model and the image analysis and screening model are respectively trained artificial intelligence engines. Kharbanda discloses searching for a plurality of semantic images associated with the analyzed customization information from an image classification database (Kharbanda, [0047]: teaches an image evaluation module 210 including/accessing image search space 215, which includes "intermediate representations for a plurality of stored images <read on searched semantic images>"; [0047]: further teaches the image search space 215 being an embedding space that includes embeddings <read on analyzed customization information> generated for images stored in a data store (e.g., a database <read on image classification database>) that stores and indexes a large volume of images to facilitate visual search services), and receiving the plurality of semantic images (Kharbanda, [0047]: teaches the image evaluation module 210 accessing image search space 215, which includes a database of stored images <read on received semantic images>), and/or inputting the plurality of semantic images and/or the at least one initial image into an image analysis and screening model (Kharbanda, [0047]: teaches the image evaluation module 210 including a machine-learned visual search model 214 <read on image analysis and screening model> that is trained to identify images that are visually similar to the query image 206 <read on semantic images/initial image> from a corpus of stored image data); and performing image analysis and screening on the plurality of semantic images and/or the at least one initial image through the image analysis and screening model (Kharbanda, [0047]: teaches the image evaluation module 210 including the machine-learned visual search model 214 <read on image analysis and screening model> that is trained to identify <read on performing image analysis and screening> images that are visually similar to the query image 206 <read on semantic images/initial image> from a corpus of stored image data), and screening the plurality of semantic images and/or the at least one initial image based on at least one of the customization demand information and the analyzed customization information to obtain at least one first retrieved image (Kharbanda, [0047]: teaches the image evaluation module 210 selecting <read on screening> result images <read on obtained first retrieved image>; [0048]: teaches the result image 212 being one or more result images obtained due to a similarity between the result images and the query image 206), wherein the semantic analysis model, the image generation model and the image analysis and screening model are respectively trained artificial intelligence engines (Kharbanda, [0047]: teaches the machine-learned visual search model 214 <read on image analysis and screening model> being trained <read on read on trained AI engine> to "identify images that are visually similar to the query image 206 from a corpus of stored image data"). Kharbanda is analogous art with respect to Chen because they are from the same field of endeavor, namely processing input requests for multimodal system applications. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to incorporate a visual search module that can access a large database of images that closely matches the query input for a multimodal image diffusion model as taught by Kharbanda into the teaching of Chen. The suggestion for doing so would allow the system to present at least one image result to the user, which can then be used as input for the multimodal image diffusion model to further clarify the intent and request of the user. Therefore, it would have been obvious to combine Kharbanda with Chen. Regarding Claim 8, it recites the limitations that are similar in scope to Claim 1, but in a computing device. As shown in the rejection, the combination of Chen and Kharbanda discloses the limitations of Claim 1. Additionally, Chen discloses a computing device for generating customized images (Chen, [0050]: teaches an image generation apparatus <read on computing device> that generates images <read on customized images>), comprising: a storage module, configured to store a computer program product (Chen, [0050]: teaches the image generation apparatus including a memory subsystem <read on storage module>; [0048]: teaches a user device 110 including software <read on stored computer program product> that "displays a user interface (e.g., a graphical user interface) provided by image generation apparatus 115"); and a processing module, configured to be coupled to the storage module (Chen, [0166]: teaches memory subsystem 1810 including one or more memory devices, where "memory is used to store computer-readable, computer-executable software including instructions that, when executed, cause a processor <read on processing module> to perform various functions"), wherein after the processing module loads and executes the computer program product, the processing module performs a method for generating customized images (Chen, [0166]: teaches memory subsystem 1810 including one or more memory devices, where "memory is used to store computer-readable, computer-executable software including instructions that, when executed, cause a processor <read on processing module> to perform various functions"; [0100]: teaches a method for generating synthetic images <read on customized images>), the method for generating customized images comprising the following steps of (Chen, [0100]: teaches the method for generating synthetic images <read on customized images>):… Thus, Claim 8 is met by Chen according to the mapping presented in the rejection of Claim 1, given the method corresponds to a computing device. Regarding Claim 15, it recites the limitations that are similar in scope to Claim 1, but in a non-transitory computer-readable recording medium. As shown in the rejection, the combination of Chen and Kharbanda discloses the limitations of Claim 1. Additionally, Chen discloses a non-transitory computer-readable recording medium for generating customized images (Chen, [0050]: teaches an image generation apparatus, which includes a memory subsystem (i.e., computer-readable media), that generates images <read on customized images>; [0222]: teaches computer-readable media being a form of memory, such as non-transitory computer storage media, that is accessible by a computer, where it stores data or code), after a computing device loads a computer program product stored in the non-transitory computer-readable recording medium and executes the computer program product, the computing device performs a method for generating customized images (Chen, [0164]: teaches computing device 1800 including one or more processors that can execute instructions <read on computer program product> stored in memory subsystem 1810 to perform image generation; [0100]: teaches a method for generating synthetic images <read on customized images>), and the method for generating customized images comprising the following steps of (Chen, [0100]: teaches the method for generating synthetic images <read on customized images>):… Thus, Claim 15 is met by Chen according to the mapping presented in the rejection of Claim 1, given the method corresponds to a non-transitory computer-readable recording medium. Regarding Claims 4, 11, 18, the combination of Chen and Kharbanda discloses the method, the computing device, and the non-transitory computer-readable recording medium of Claims 1, 8, and 15 respectively. Additionally, Chen further discloses wherein the step of performing a semantic analysis on the customization demand information through the semantic analysis model, and generating and outputting the analyzed customization information based on the customization demand information comprises the following sub- steps: [[performing language translation on the customization demand information through the semantic analysis model, and]] [[generating translated customization information based on the customization demand information; and]] [[performing semantic disassembly on the translated customization information through the semantic analysis model, and]] generating and outputting the analyzed customization information based on the [[translated]] customization information (Chen, [0070]: teaches the text encoder 505 and the style encoder 510 generating text embedding 530 <read on analyzed customization information> and style embedding 535 respectively, which are based on the inputted text prompt 520 <read on customization demand information> and style input 525). However, Chen does not expressly disclose performing language translation on the customization demand information through the semantic analysis model, and generating translated customization information based on the customization demand information; and performing semantic disassembly on the translated customization information through the semantic analysis model, and generating and outputting the analyzed customization information based on the translated customization information. Kharbanda discloses performing language translation on the customization demand information through the semantic analysis model (Kharbanda, [0165]: teaches a machine-learned model <read on semantic analysis model> processing text <read on customization demand information>, such as a prompt, for translation <read on language translation>), and generating translated customization information based on the customization demand information (Kharbanda, [0165]: teaches a machine-learned model processing text <read on customization demand information> to generate a translation output <read on translated customization information>); and performing semantic disassembly on the translated customization information through the semantic analysis model (Kharbanda, [0076]: teaches performing semantic understanding <read on semantic disassembly> of prompt 208 <read on translated customization information>), and generating and outputting the analyzed customization information based on the translated customization information (Kharbanda, [0165]: teaches a machine-learned model processing text to generate a translation output <read on translated customization information>, where the translation output is interpreted to be a text prompt). Kharbanda is analogous art with respect to Chen because they are from the same field of endeavor, namely processing input requests for multimodal system applications. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to incorporate a plurality of machine-learned models, such as a language translation model, as an add-on neural network of a multimodal image diffusion system as taught by Kharbanda into the teaching of Chen. The suggestion for doing so would allow the system to perform image searches and image generations at any language, thereby allowing for widespread use of the tool in any language. Therefore, it would have been obvious to combine Kharbanda with Chen. Regarding Claims 6, 13, 20, the combination of Chen and Kharbanda discloses the method, the computing device, and the non-transitory computer-readable recording medium of Claims 1, 8, and 15 respectively. Additionally, Chen further discloses wherein the method further comprises the following steps of: receiving a customized merchandise image corresponding to merchandise information for customization (Chen, [0118]: teaches using an input image <read on customized merchandise image> and a style input <read on merchandise information for customization> for generating a synthetic image); inputting [[the at least one first retrieved image and]] the customized merchandise image into an image synthesis model (Chen, [0122]: teaches an image generation model <read on image synthesis model> using an input image <read on customized merchandise image> and style input; Note: it should be noted that the "image generation model" and the "image synthesis model" are being treated as identical terms); and performing image synthesis on [[the at least one first retrieved image and]] the customized merchandise image through the image synthesis model (Chen, [0122]: teaches the image generation model <read on image synthesis model> performing a synthetic image generation process <read on performing image synthesis> using the input image <read on customized merchandise image> and style input), and generating and outputting at least one synthesized image based on [[the at least one first retrieved image and]] the customized merchandise image (Chen, [0122]: teaches the image generation model generating a stylized synthetic image <read on synthesized image> based on the input image <read on customized merchandise image> and style input), wherein the image synthesis model is a trained artificial intelligence engine (Chen, [0210]: teaches a trained <read on trained AI engine> diffusion model <read on image synthesis model>). However, Chen does not expressly disclose inputting the at least one first retrieved image and the customized merchandise image into an image synthesis model; and performing image synthesis on the at least one first retrieved image and the customized merchandise image through the image synthesis model, and generating and outputting at least one synthesized image based on the at least one first retrieved image and the customized merchandise image. Kharbanda discloses inputting the at least one first retrieved image and the customized merchandise image into an image synthesis model (Kharbanda, [0047]: teaches the image evaluation module 210 selecting result images <read on first retrieved image>); and performing image synthesis on the at least one first retrieved image and the customized merchandise image through the image synthesis model (Kharbanda, [0047]: teaches the image evaluation module 210 selecting result images <read on first retrieved image>), and generating and outputting at least one synthesized image based on the at least one first retrieved image and the customized merchandise image (Kharbanda, [0047]: teaches the image evaluation module 210 selecting result images <read on first retrieved image>). Kharbanda is analogous art with respect to Chen because they are from the same field of endeavor, namely processing input requests for multimodal system applications. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to incorporate a visual search module add-on that can access a large database of images that closely matches the query input for a multimodal image diffusion model as taught by Kharbanda into the teaching of Chen. The suggestion for doing so would allow the image generation model to use found images as part of the input image data query, thereby offering the neural network more contextual data that can be used for more desirable image generations. Therefore, it would have been obvious to combine Kharbanda with Chen. Regarding Claims 7, 14, 21, the combination of Chen and Kharbanda discloses the method, the computing device, and the non-transitory computer-readable recording medium of Claims 6, 13, and 20 respectively. Additionally, Chen further discloses wherein the method further comprises the following steps of: selecting from the at least one synthesized image by a user to obtain a selected image (Chen, [0079]: teaches a user interface 800 displaying synthetic images <read on selected image> based on a selected style element 810 as shown in FIG. 8) and PNG media_image2.png 665 443 media_image2.png Greyscale outputting the selected image to an ordering system (Chen, FIG. 8 teaches the synthetic images being displayed on user interface 800 <read on ordering system>). Claims 2-3, 9-10, and 16-17 are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (US 20250117973 A1), hereinafter referenced as Chen, in view of Kharbanda et al. (US 20240378237 A1), hereinafter referenced as Kharbanda as applied to Claims 1, 8, and 15 above respectively, and further in view of Jain et al. (US 5893095 A), hereinafter referenced as Jain. Regarding Claims 2, 9, 16, the combination of Chen and Kharbanda discloses the method, the computing device, and the non-transitory computer-readable recording medium of Claims 1, 8, and 15 respectively. Additionally, Chen further discloses wherein the method further comprises the following steps of: [[determining whether a total number of the at least one first retrieved image is less than a predetermined value; and]] [[when the total number of the at least one first retrieved image is less than the predetermined value,]] inputting at least one of the customization demand information and the analyzed customization information into the image generation model (Chen, [0070]: teaches an image generation model 515 generating synthetic image 540 based on text embedding 530 <read on analyzed customization information> and style embedding 535 as shown in FIG. 5, where both text embedding 530 and style embedding 535 <read on customization demand information> are based on text prompt 520 and style input 525 respectively), and generating and outputting at least one generated image based on at least one of the customization demand information and the analyzed customization information through the image generation model (Chen, [0034]: teaches the image generation system generating a synthetic image <read on generated image> "based on the text embedding <read on analyzed customization information> and the style embedding using the image generation model by providing the text embedding as an input to the image generation model during initial iterations of an image generation process and providing both the text embedding and the style embedding as inputs to the image generation model during subsequent iterations of the image generation process," where the text embedding and the style embedding are both based on the text prompt <read on customization demand information> and the style input respectively), wherein [[a total number of the at least one generated image is greater than or equal to a difference between the total number of the at least one first retrieved image and the predetermined value.]] However, the combination of Chen and Kharbanda does not expressly disclose determining whether a total number of the at least one first retrieved image is less than a predetermined value; and when the total number of the at least one first retrieved image is less than the predetermined value, inputting at least one of the customization demand information and the analyzed customization information into the image generation model, and a total number of the at least one generated image is greater than or equal to a difference between the total number of the at least one first retrieved image and the predetermined value. Jain discloses determining whether a total number of the at least one first retrieved image is less than a predetermined value (Jain, [Cols. 25-26, Lines 57-67 and 1-4]: teaches count C <read on total number> starting at zero, which is the number of results found <read on first retrieved images>, then being incremented when a result is found (i.e., an image is found that satisfies a threshold) and added to a query results list when the number of results found so far C is less than a desired number of results N <read on predetermined value>); and when the total number of the at least one first retrieved image is less than the predetermined value, inputting at least one of the customization demand information and the analyzed customization information into the image generation model (Jain, [Cols. 25-26, Lines 57-67 and 1-4]: teaches count C <read on total number> starting at zero, which is the number of results found <read on first retrieved images>, then being incremented when a result is found (i.e., an image is found that satisfies a threshold) and added to a query results list when the number of results found so far C is less than a desired number of results N <read on predetermined value>), and a total number of the at least one generated image is greater than or equal to a difference between the total number of the at least one first retrieved image and the predetermined value (Jain, [Col. 26, Lines 28-29]: teaches when value C is not less than desired number of results N (i.e., C <read on total number> is greater than or equal to N <read on predetermined value>); Note: it should be noted that although a "difference between the total number of the at least one first retrieved image and the predetermined value" is not expressly stated, the total value C being at least equal to the threshold is being interpreted as a "difference between the total number of the at least one first retrieved image and the predetermined value"; for example, if the total number is 1 and the predetermined value is 5, then the difference is 4; as the count increments, it will eventually be greater than or equal to 4). Jain is analogous art with respect to Chen, in view of Kharbanda because they are from the same field of endeavor, namely searching for images from an image database/repository. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to implement a generalized image search add-on to a multimodal image generation system as taught by Jain into the teaching of Chen, in view of Kharbanda. The suggestion for doing so would allow the system to search for additional reference images that are related to the input query. Therefore, it would have been obvious to combine Jain with Chen, in view of Kharbanda. Regarding Claims 3, 10, 17, the combination of Chen, Kharbanda, and Jain discloses the method, the computing device, and the non-transitory computer-readable recording medium of Claims 2, 9, and 16 respectively. Additionally, Chen further discloses wherein the method further comprises the following steps of: [[inputting the at least one generated image into the image analysis and screening model;]] [[performing image analysis and screening on the at least one generated image through the image analysis and screening model, and]] [[screening the at least one generated image based on at least one of the customization demand information and the analyzed customization information to obtain at least one second retrieved image;]] [[determining whether a total number of the at least one second retrieved image is less than the difference; and]] [[when the total number of the at least one second retrieved image is less than the difference,]] generating and outputting the at least one generated image again based on at least one of the customization demand information and the analyzed customization information through the image generation model (Chen, [0034]: teaches the image generation system generating a synthetic image <read on generated image> "based on the text embedding <read on analyzed customization information> and the style embedding using the image generation model by providing the text embedding as an input to the image generation model during initial iterations of an image generation process and providing both the text embedding and the style embedding as inputs to the image generation model during subsequent iterations of the image generation process," where the text embedding and the style embedding are both based on the text prompt <read on customization demand information> and the style input respectively), and [[screening the at least one generated image again based on at least one of the customization demand information and the analyzed customization information through the image analysis and screening model to obtain the at least one second retrieved image, until the total number of the at least one the second retrieved image is greater than or equal to the difference.]] However, Chen does not expressly disclose inputting the at least one generated image into the image analysis and screening model; performing image analysis and screening on the at least one generated image through the image analysis and screening model, and screening the at least one generated image based on at least one of the customization demand information and the analyzed customization information to obtain at least one second retrieved image; determining whether a total number of the at least one second retrieved image is less than the difference; and when the total number of the at least one second retrieved image is less than the difference, generating and outputting the at least one generated image again based on at least one of the customization demand information and the analyzed customization information through the image generation model, and screening the at least one generated image again based on at least one of the customization demand information and the analyzed customization information through the image analysis and screening model to obtain the at least one second retrieved image, until the total number of the at least one the second retrieved image is greater than or equal to the difference. Kharbanda discloses inputting the at least one generated image into the image analysis and screening model (Kharbanda, [0047]: teaches the image evaluation module 210 including a machine-learned visual search model 214 <read on image analysis and screening model> that is trained to identify images that are visually similar to the query image 206 <read on generated image> from a corpus of stored image data); performing image analysis and screening on the at least one generated image through the image analysis and screening model (Kharbanda, [0047]: teaches the image evaluation module 210 including the machine-learned visual search model 214 <read on image analysis and screening model> that is trained to identify <read on performing image analysis and screening> images that are visually similar to the query image 206 <read on generated image> from a corpus of stored image data), and screening the at least one generated image based on at least one of the customization demand information and the analyzed customization information to obtain at least one second retrieved image (Kharbanda, [0047]: teaches the image evaluation module 210 selecting <read on screening> result images <read on second retrieved images>; [0048]: teaches the result image 212 being one or more result images <read on generated images> obtained due to a similarity between the result images and the query image 206); [[determining whether a total number of the at least one second retrieved image is less than the difference; and]] [[when the total number of the at least one second retrieved image is less than the difference, generating and outputting the at least one generated image again based on at least one of the customization demand information and the analyzed customization information through the image generation model, and]] screening the at least one generated image again based on at least one of the customization demand information and the analyzed customization information through the image analysis and screening model to obtain the at least one second retrieved image, [[until the total number of the at least one the second retrieved image is greater than or equal to the difference]] (Kharbanda, [0047]: teaches the image evaluation module 210 selecting <read on screening> result images <read on second retrieved images>; [0048]: teaches the result image 212 being one or more result images <read on generated images> obtained due to a similarity between the result images and the query image 206; [0082]: teaches the search aspect being iterative). Kharbanda is analogous art with respect to Chen because they are from the same field of endeavor, namely processing input requests for multimodal system applications. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to incorporate a visual search module that can access a large database of images that closely matches the query input for a multimodal image diffusion model as taught by Kharbanda into the teaching of Chen. The suggestion for doing so would allow the system to present at least one image result to the user, which can then be used as input for the multimodal image diffusion model to further clarify the intent and request of the user. Therefore, it would have been obvious to combine Kharbanda with Chen. However, the combination of Chen and Kharbanda does not expressly disclose determining whether a total number of the at least one second retrieved image is less than the difference; and when the total number of the at least one second retrieved image is less than the difference, generating and outputting the at least one generated image again based on at least one of the customization demand information and the analyzed customization information through the image generation model, and screening the at least one generated image again based on at least one of the customization demand information and the analyzed customization information through the image analysis and screening model to obtain the at least one second retrieved image, until the total number of the at least one the second retrieved image is greater than or equal to the difference. Jain discloses determining whether a total number of the at least one second retrieved image is less than the difference (Jain, [Col. 26, Lines 28-29]: teaches when value C <read on total number> is less than desired number of results N, which is interpreted to be less than the difference between the total number and the predetermined value); and when the total number of the at least one second retrieved image is less than the difference, generating and outputting the at least one generated image again based on at least one of the customization demand information and the analyzed customization information through the image generation model (Jain, [Col. 26, Lines 28-29]: teaches value C <read on total number> being less than the desired number of results N, which is interpreted to be less than the difference between the total number and the predetermined value), and screening the at least one generated image again based on at least one of the customization demand information and the analyzed customization information through the image analysis and screening model to obtain the at least one second retrieved image, until the total number of the at least one the second retrieved image is greater than or equal to the difference (Jain, [Col. 26, Lines 28-29]: teaches value C <read on total number> being greater than or equal to the desired number of results N, which is interpreted to be less than the difference between the total number and the predetermined value). Jain is analogous art with respect to Chen, in view of Kharbanda because they are from the same field of endeavor, namely searching for images from an image database/repository. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to implement a generalized image search add-on to a multimodal image generation system as taught by Jain into the teaching of Chen, in view of Kharbanda. The suggestion for doing so would allow the system to search for additional reference images that are related to the input query. Therefore, it would have been obvious to combine Jain with Chen, in view of Kharbanda. Claims 5, 12, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (US 20250117973 A1), hereinafter referenced as Chen, in view of Kharbanda et al. (US 20240378237 A1), hereinafter referenced as Kharbanda as applied to Claims 1, 8, and 15 above respectively, and further in view of Peng et al. (US 20220414959 A1), hereinafter referenced as Peng. Regarding Claims 5, 12, 19, the combination of Chen and Kharbanda discloses the method, the computing device, and the non-transitory computer-readable recording medium of Claims 1, 8, and 15 respectively. The combination of Chen and Kharbanda does not expressly disclose the limitations of Claims 5, 12, and 19; however, Peng discloses wherein the method further comprises the following steps of: inputting the at least one first retrieved image into an image editing model (Peng, [0048]: teaches obtaining a standard image sample set <read on first retrieved images> for an image editing model); and performing image editing on the at least one first retrieved image through the image editing model (Peng, [0048]: teaches the image editing model editing <read on image editing> each initial image from the standard image sample set <read on first retrieved images>), and generating and outputting at least one edited image based on the at least one first retrieved image (Peng, [0048]: teaches an edited image from the image editing model), wherein the image editing model is a trained artificial intelligence engine (Peng, [0145]: teaches the image editing model being pre-trained <read on trained AI engine>). Peng is analogous art with respect to Chen, in view of Kharbanda because they are from the same field of endeavor, namely multimodal image generation systems. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to implement an image editing model that edits input images as taught by Peng into the teaching of Chen, in view of Kharbanda. The suggestion for doing so would allow the user to refine the generated image based on their preferences, thereby leading to a more personalized experience. Therefore, it would have been obvious to combine Peng with Chen, in view of Kharbanda. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Gong et al. (US 20250298840 A1) discloses a text-to-image retrieval system that finds the right picture when a user types a query about an object; Kharbanda et al. (US 20240378236 A1) discloses a visual search system based on a similarity between a query image and a result image; Maschmeyer et al. (US 20240161258 A1) discloses generating a multimodal image diffusion system that uses interaction data of a customer; Saraee et al. (US 20240378856 A1) discloses a multimodal image generation system that searches for images for use in e-commerce; Wong et al. (US 20240160662 A1) discloses a graphics-specific search engine; Zeng et al. (US 20240169623 A1) discloses a multi-modal image generation system; and Zhang et al. (US 20230418861 A1) discloses a multimodal image generation system that provides customizable search tools. Any inquiry concerning this communication or earlier communications from the examiner should be directed to KARL TRUONG whose telephone number is (703)756-5915. The examiner can normally be reached 10:30 AM - 7:30 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kent Chang can be reached at (571) 272-7667. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /K.D.T./Examiner, Art Unit 2614 /KENT W CHANG/Supervisory Patent Examiner, Art Unit 2614
Read full office action

Prosecution Timeline

May 02, 2025
Application Filed
Aug 25, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737979
INFORMATION PROCESSING APPARATUS AND METHOD, AND STORAGE MEDIUM
2y 6m to grant Granted Sep 15, 2026
Patent 12725260
GENERATING PANOPTIC SEGMENTATION LABELS
3y 5m to grant Granted Sep 01, 2026
Patent 12706011
DYNAMIC ARBITRARY BORDER GAIN
2y 11m to grant Granted Aug 11, 2026
Patent 12700060
HIGH RESOLUTION SYNTHESIS USING SHADERS
3y 2m to grant Granted Aug 04, 2026
Patent 12694605
Bounding Volume Hierarchy with Bounding Volumes in Prior Space corresponding to Subset of Transform Sub-Tree Bounds
2y 6m to grant Granted Jul 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
65%
Grant Probability
99%
With Interview (+34.3%)
2y 8m (~1y 4m remaining)
Median Time to Grant
Low
PTA Risk
Based on 51 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month