Prosecution Insights
Last updated: October 04, 2026
Application No. 19/037,762

Method and System for Generating Images

Non-Final OA §103
Filed
Jan 27, 2025
Priority
Jan 31, 2024 — RE 10-2024-0015288
Examiner
AMIN, JWALANT B
Art Unit
Tech Center
Assignee
Gengenai Inc.
OA Round
1 (Non-Final)
79%
Grant Probability
Favorable
1-2
OA Rounds
11m
Est. Remaining
95%
With Interview

Examiner Intelligence

Grants 79% — above average
79%
Career Allowance Rate
510 granted / 643 resolved
+19.3% vs TC avg
Strong +16% interview lift
Without
With
+15.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
19 currently pending
Career history
652
Total Applications
across all art units

Statute-Specific Performance

§101
15.3%
-24.7% vs TC avg
§103
56.4%
+16.4% vs TC avg
§102
7.0%
-33.0% vs TC avg
§112
11.0%
-29.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 643 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Objections Claim 10 is objected to because of the following informalities: On line 5, in the phrase “synthesized image , and”, there is space between the words “image” and “,”. Appropriate correction is required. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 1, and 13-14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Gandelsman et al. (US 2024/0169604, hereinafter Gandelsman). Regarding claim 1, Gandelsman teaches a method performed by one or more processors ([0005]: A method, apparatus, and non-transitory computer readable medium for image generation are described; [0006]: One or more embodiments of the apparatus and method include one or more processors; one or more memories including instructions executable by the one or more processors to; obtaining user input that indicates a target color and a semantic label for a region of an image to be generated; generating a noise map including noise biased towards the target color in the region indicated by the user input; and generating the image based on the noise map and the semantic label for the region using a diffusion model, wherein the image includes an object in the region that is described by the semantic label and that has the target color), the method comprising: receiving a first dot (user drawn rough layout on the image canvas such as in the first selection region 915 as shown in fig. 9) associated with a first object (user draws color layout of one or more objects at a particular position on the image canvas; [0029]: the user input includes layout information indicating a plurality of regions of an image canvas, and wherein each of the plurality of regions is associated with a corresponding target color and a corresponding semantic label; [0087]: As used herein, the term “color layout” includes color information of an object and layout information (e.g., position, size, and orientation) of the object to be generated in an output image; [0092]: At operation 805, the system displays a user interface to a user, where the user interface includes a label input field, a color input field, and a selection tool for selecting the region … For example, a user can use the selection tool (e.g., a virtual brush) to draw a rough layout of one or more objects on an image canvas of the user interface; [0095]: User input 905 includes semantic label of an entity (e.g., class), color input (e.g., color picker), and size of a selection tool (e.g., size picker). User interface 900 includes editing image canvas 910 where users can draw or indicate color layout of one or more objects; [0096]: For example, the user inputs a text prompt and a rough color layout of entities (e.g., objects) in a 2D canvas (e.g., image canvas 910). User interface 900 receives user commands where the user commands include text prompt in an image canvas. The user selects a region on the image canvas using a virtual brush associated with the phrase “hedgehog”. The user also inputs an intended color of the entity “hedgehog”. Then, the user draws a rough layout of this entity); based on the first dot, generating, using an image generation model (diffusion model, [0094]), a first synthesized image (output image 930, fig. 9; [0094]: At operation 815, the system generates an image based on the user input using the diffusion model … The diffusion model generates the image based on the noise map and the semantic label for the region, where the image includes an object in the region that is described by the semantic label and that has the target color. The combination of using a color layout guidance and text guidance enables precise control of the layout of a set of objects in the generated image; [0098]: The user repeats this process for one or more additional entities as the user desires to include in output image 930. As an example, the user also desires to have a “calculator” and an “eye” to be generated in output image 930, thereby inputs second selection region 920 and third selection region 925 with corresponding user input 905. Output image 930 shows a hedgehog next to a calculator. The scene depicted in output image 930 is consistent with the color layout and object relations in the image canvas. The scene depicted in output image 930 is also consistent with the text prompt); and outputting the first synthesized image ([0089]: At operation 620, the system displays the output image to the user. In some cases, the operations of this step refer to, or may be performed by, an image generation apparatus as described with reference to FIGS. 1 and 2. In some examples, image generation apparatus displays the output image to the user via the user interface; [0098]: Output image 930 shows a hedgehog next to a calculator; [0106]: At operation 1015, the system generates the image based on the noise map and the semantic label for the region using a diffusion model, where the image includes an object in the region that is described by the semantic label and that has the target color. In some cases, the operations of this step refer to, or may be performed by, a diffusion model as described with reference to FIG. 2. The generated image is displayed to the user), wherein the first dot comprises first class information (semantic label (class) of the entity (object); fig. 9 shows class as a semantic label “hedgehog”) and information indicating a first position (layout information such as position of the object; first selection region 915 as shown in fig. 9) associated with the first object ([0029]: the user input includes layout information indicating a plurality of regions of an image canvas, and wherein each of the plurality of regions is associated with a corresponding target color and a corresponding semantic label; [0087]: As used herein, the term “color layout” includes color information of an object and layout information (e.g., position, size, and orientation) of the object to be generated in an output image; [0093]: The user enters a corresponding class label into the label input field), and wherein the first synthesized image (output image 930, fig. 9) is a synthesized image in which the first object (hedgehog as shown in output image 930) corresponding to the first class information (class “hedgehog”, fig. 9) is placed at the first position (hedgehog is placed at the first select region 915; [0097]: First selection region 915 corresponds to entity “hedgehog”). Although Gandelsman does not explicitly teach to receive a dot to select a region on the image canvas to place the object in a generated image, Gandelsman teaches to select a region by drawing a rough layout on the image canvas (fig. 9). However, it would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention for the region on the image canvas to be selected as a dot rather a drawing rough layout. Whether the region is selected as a dot or a rough drawn layout is solely a matter of aesthetic design choice, and would not be sufficient to distinguish over the prior art. See MPEP 2144.04. Claim 13 is similar in scope to claim 1, and therefore the examiner provides similar rationale to reject these claims. Moreover, Gandelsman teaches a non-transitory computer-readable medium ([0079]: In FIGS. 6-11, a method, apparatus, and non-transitory computer readable medium for image generation are described. One or more embodiments of the method, apparatus, and non-transitory computer readable medium include obtaining user input that indicates a target color and a semantic label for a region of an image to be generated; generating a noise map including noise biased towards the target color in the region indicated by the user input; and generating the image based on the noise map and the semantic label for the region using a diffusion model, wherein the image includes an object in the region that is described by the semantic label and that has the target color). Regarding claim 14, Gandelsman teaches an apparatus (image generation apparatus 200, fig. 2/computing device 1400, fig. 14) comprising: a communication interface (communication interface 1415, fig. 14); a memory (memory unit 210, fig. 2/a memory subsystem 1410, fig. 14); one or more processors (processor unit 205, fig. 2/processor(s) 1405, fig. 14) coupled to the memory and configured to execute one or more computer-readable programs stored in the memory ([0042]: Memory unit 210 comprises a memory including instructions executable by processor unit 205. Examples of memory unit 210 include random access memory (RAM), read-only memory (ROM), or a hard disk. Examples of memory unit 210 include solid-state memory and a hard disk drive. In some examples, memory unit 210 is used to store computer-readable, computer-executable software including instructions that, when executed, cause a processor to perform various functions described herein; [0091]: FIG. 8 shows an example of a method for operating a user interface according to aspects of the present disclosure. In some examples, these operations are performed by a system including a processor executing a set of codes to control functional elements of an apparatus), wherein the one or more computer-readable programs comprise instructions that, when executed by the one or more processors, are configured to cause the apparatus to ([0006]: One or more embodiments of the apparatus and method include one or more processors; one or more memories including instructions executable by the one or more processors to; obtaining user input that indicates a target color and a semantic label for a region of an image to be generated; generating a noise map including noise biased towards the target color in the region indicated by the user input; and generating the image based on the noise map and the semantic label for the region using a diffusion model, wherein the image includes an object in the region that is described by the semantic label and that has the target color; [0128]: In some embodiments, computing device 1400 includes one or more processors 1405 that can execute instructions stored in memory subsystem 1410 to obtain user input that indicates a target color and a semantic label for a region of an image to be generated; generate a noise map including noise biased towards the target color in the region indicated by the user input; and generating the image based on the noise map and the semantic label for the region using a diffusion model, wherein the image includes an object in the region that is described by the semantic label and that has the target color): receive a first dot (user drawn rough layout on the image canvas such as in the first selection region 915 as shown in fig. 9) associated with a first object (user draws color layout of one or more objects at a particular position on the image canvas; [0029]: the user input includes layout information indicating a plurality of regions of an image canvas, and wherein each of the plurality of regions is associated with a corresponding target color and a corresponding semantic label; [0087]: As used herein, the term “color layout” includes color information of an object and layout information (e.g., position, size, and orientation) of the object to be generated in an output image; [0092]: At operation 805, the system displays a user interface to a user, where the user interface includes a label input field, a color input field, and a selection tool for selecting the region … For example, a user can use the selection tool (e.g., a virtual brush) to draw a rough layout of one or more objects on an image canvas of the user interface; [0095]: User input 905 includes semantic label of an entity (e.g., class), color input (e.g., color picker), and size of a selection tool (e.g., size picker). User interface 900 includes editing image canvas 910 where users can draw or indicate color layout of one or more objects; [0096]: For example, the user inputs a text prompt and a rough color layout of entities (e.g., objects) in a 2D canvas (e.g., image canvas 910). User interface 900 receives user commands where the user commands include text prompt in an image canvas. The user selects a region on the image canvas using a virtual brush associated with the phrase “hedgehog”. The user also inputs an intended color of the entity “hedgehog”. Then, the user draws a rough layout of this entity); based on the first dot, generate, using an image generation model (diffusion model, [0094]), a first synthesized image (output image 930, fig. 9; [0094]: At operation 815, the system generates an image based on the user input using the diffusion model … The diffusion model generates the image based on the noise map and the semantic label for the region, where the image includes an object in the region that is described by the semantic label and that has the target color. The combination of using a color layout guidance and text guidance enables precise control of the layout of a set of objects in the generated image; [0098]: The user repeats this process for one or more additional entities as the user desires to include in output image 930. As an example, the user also desires to have a “calculator” and an “eye” to be generated in output image 930, thereby inputs second selection region 920 and third selection region 925 with corresponding user input 905. Output image 930 shows a hedgehog next to a calculator. The scene depicted in output image 930 is consistent with the color layout and object relations in the image canvas. The scene depicted in output image 930 is also consistent with the text prompt); and output the first synthesized image ([0089]: At operation 620, the system displays the output image to the user. In some cases, the operations of this step refer to, or may be performed by, an image generation apparatus as described with reference to FIGS. 1 and 2. In some examples, image generation apparatus displays the output image to the user via the user interface; [0098]: Output image 930 shows a hedgehog next to a calculator; [0106]: At operation 1015, the system generates the image based on the noise map and the semantic label for the region using a diffusion model, where the image includes an object in the region that is described by the semantic label and that has the target color. In some cases, the operations of this step refer to, or may be performed by, a diffusion model as described with reference to FIG. 2. The generated image is displayed to the user), wherein the first dot comprises first class information (semantic label (class) of the entity (object); fig. 9 shows class as a semantic label “hedgehog”) and information indicating a first position (layout information such as position of the object; first selection region 915 as shown in fig. 9) associated with the first object ([0029]: the user input includes layout information indicating a plurality of regions of an image canvas, and wherein each of the plurality of regions is associated with a corresponding target color and a corresponding semantic label; [0087]: As used herein, the term “color layout” includes color information of an object and layout information (e.g., position, size, and orientation) of the object to be generated in an output image; [0093]: The user enters a corresponding class label into the label input field), and wherein the first synthesized image (output image 930, fig. 9) is a synthesized image in which the first object (hedgehog as shown in output image 930) corresponding to the first class information (class “hedgehog”, fig. 9) is placed at the first position (hedgehog is placed at the first select region 915; [0097]: First selection region 915 corresponds to entity “hedgehog”). Although Gandelsman does not explicitly teach to receive a dot to select a region on the image canvas to place the object in a generated image, Gandelsman teaches to select a region by drawing a rough layout on the image canvas (fig. 9). However, it would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention for the region on the image canvas to be selected as a dot rather a drawing rough layout. Whether the region is selected as a dot or a rough drawn layout is solely a matter of aesthetic design choice, and would not be sufficient to distinguish over the prior art. See MPEP 2144.04. Regarding claim 4, Gandelsman teaches the method according to claim 1, further comprising receiving a second dot (user drawn rough layout on the image canvas such as in the second selection region 920 as shown in fig. 9) associated with a second object (user draws color layout of one or more objects at a particular position on the image canvas; [0029]: the user input includes layout information indicating a plurality of regions of an image canvas, and wherein each of the plurality of regions is associated with a corresponding target color and a corresponding semantic label; [0087]: As used herein, the term “color layout” includes color information of an object and layout information (e.g., position, size, and orientation) of the object to be generated in an output image; [0092]: At operation 805, the system displays a user interface to a user, where the user interface includes a label input field, a color input field, and a selection tool for selecting the region … For example, a user can use the selection tool (e.g., a virtual brush) to draw a rough layout of one or more objects on an image canvas of the user interface; [0095]: User input 905 includes semantic label of an entity (e.g., class), color input (e.g., color picker), and size of a selection tool (e.g., size picker). User interface 900 includes editing image canvas 910 where users can draw or indicate color layout of one or more objects; [0096]: For example, the user inputs a text prompt and a rough color layout of entities (e.g., objects) in a 2D canvas (e.g., image canvas 910). User interface 900 receives user commands where the user commands include text prompt in an image canvas), wherein the second dot comprises second class information (semantic label (class) of the entity (object) is calculator; [0097]: Second selection region 920 corresponds to entity “calculator”) and information indicating a second position ([0097]: Second selection region 920 corresponds to entity “calculator”) associated with the second object ([0097]: Second selection region 920 corresponds to entity “calculator”; [0098]: the user also desires to have a “calculator” and an “eye” to be generated in output image 930, thereby inputs second selection region 920 and third selection region 925 with corresponding user input 905), and the first synthesized image (output image 930, fig. 9) is a synthesized image in which the first object (hedgehog) corresponding to the first class information (class: hedgehog, fig. 9) is placed at the first position (selection region 915, fig. 9) and the second object (calculator) corresponding to the second class information (class: calculator) is placed at the second position (selection region 920, fig. 9). Regarding claim 9, Gandelsman teaches the method according to claim 1, wherein the first dot (region 915, fig. 9) is generated by a user selecting a specific position on a display (user drawn rough layout on the image canvas such as in the first selection region 915 as shown in fig. 9; [0029]: the user input includes layout information indicating a plurality of regions of an image canvas, and wherein each of the plurality of regions is associated with a corresponding target color and a corresponding semantic label; [0087]: As used herein, the term “color layout” includes color information of an object and layout information (e.g., position, size, and orientation) of the object to be generated in an output image; [0092]: At operation 805, the system displays a user interface to a user, where the user interface includes a label input field, a color input field, and a selection tool for selecting the region … For example, a user can use the selection tool (e.g., a virtual brush) to draw a rough layout of one or more objects on an image canvas of the user interface; [0095]: User input 905 includes semantic label of an entity (e.g., class), color input (e.g., color picker), and size of a selection tool (e.g., size picker). User interface 900 includes editing image canvas 910 where users can draw or indicate color layout of one or more objects; [0096]: For example, the user inputs a text prompt and a rough color layout of entities (e.g., objects) in a 2D canvas (e.g., image canvas 910). User interface 900 receives user commands where the user commands include text prompt in an image canvas. The user selects a region on the image canvas using a virtual brush associated with the phrase “hedgehog”. The user also inputs an intended color of the entity “hedgehog”. Then, the user draws a rough layout of this entity) and inputting class information associated with the specific position ([0029]: the user input includes layout information indicating a plurality of regions of an image canvas, and wherein each of the plurality of regions is associated with a corresponding target color and a corresponding semantic label; [0087]: As used herein, the term “color layout” includes color information of an object and layout information (e.g., position, size, and orientation) of the object to be generated in an output image; [0092]: At operation 805, the system displays a user interface to a user, where the user interface includes a label input field, a color input field, and a selection tool for selecting the region … For example, a user can use the selection tool (e.g., a virtual brush) to draw a rough layout of one or more objects on an image canvas of the user interface; [0093]: The user enters a corresponding class label into the label input field; [0095]: User input 905 includes semantic label of an entity (e.g., class), color input (e.g., color picker), and size of a selection tool (e.g., size picker). User interface 900 includes editing image canvas 910 where users can draw or indicate color layout of one or more objects; [0096]: For example, the user inputs a text prompt and a rough color layout of entities (e.g., objects) in a 2D canvas (e.g., image canvas 910). User interface 900 receives user commands where the user commands include text prompt in an image canvas. The user selects a region on the image canvas using a virtual brush associated with the phrase “hedgehog”. The user also inputs an intended color of the entity “hedgehog”. Then, the user draws a rough layout of this entity). Claim(s) 2, 10-11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Gandelsman, and further in view of Lamm (US 2021/0146265). Regarding claim 2, Gandelsman does not explicitly teach the method according to claim 1, wherein a background of the first synthesized image is determined based on the first class information and the information indicating the first position. Lamm teaches a background of the first synthesized image is determined based on the first class information and the information indicating the first position ([0135]: The output is provided at step 1735. The output can comprise a set of coordinates defining an identified object in the image. With the object location and type identified, the augmented reality module can then superimpose a desired background (e.g. a background selected from the augmented background library) over every pixel in the image that was not identified as the object, as shown at step 1740). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply Lamm’s knowledge of determining a background based on the object location and type as taught and modify the process of Gandelsman because such a process integrates real world objects with computer generated virtual or augmented reality to enhance the user’s experience ([0002]). Regarding claim 10, Gandelsman teaches the method according to claim 1, wherein the generating the first synthesized image comprises: receiving condition information (color layout comprising the color information and the layout information of an object) associated with the first object ([0021]: In some examples, different target colors and semantic labels corresponding to a set of objects are drawn on an image canvas via a custom user interface. A noise version of a color layout is provided as input to a text-guided diffusion model. The text-guided diffusion model includes a perception model that enforces intermediate image outputs to comply with the text prompt (e.g., semantic labels). The image generation apparatus has precise control over the layout of the set of objects in an output image to be generated; [0032]: An image canvas includes color layout of two objects (e.g., hedgehog and calculator) and their target colors; [0087]: the term “color layout” includes color information of an object and layout information (e.g., position, size, and orientation) of the object to be generated in an output image; [0093]: the user selects a first region using the selection tool (e.g., the virtual brush), indicates the first region as hedgehog, and designates a color as yellow. The user further selects a second region and a third region. The user enters a corresponding class label into the label input field and a corresponding color into the color input field. The machine learning model receives the user input as guides (text guide and color layout guide); [0095]: User interface 900 includes editing image canvas 910 where users can draw or indicate color layout of one or more objects; [0096]: user interface 900 enables a user to control the object layout in an image to be generated (e.g., output image 930). User interface 900 receives a command from a user. For example, the user inputs a text prompt and a rough color layout of entities (e.g., objects) in a 2D canvas (e.g., image canvas 910). User interface 900 receives user commands where the user commands include text prompt in an image canvas. The user selects a region on the image canvas using a virtual brush associated with the phrase “hedgehog”. The user also inputs an intended color of the entity “hedgehog”. Then, the user draws a rough layout of this entity); and based on the first dot and the condition information, generating, using the image generation model, the first synthesized image ([0033]: the output image from image generation apparatus 110 includes a scene of two objects that matches the text prompt (semantic labels). In some examples, image generation apparatus 110 determines whether the first object (e.g., the hedgehog) overlaps the second object (e.g., the calculator) based on the text input and color layout. Image generation apparatus 110 generates an output image showing a hedgehog sitting on a calculator based on the text input and the intended color layout; [0079]: One or more embodiments of the method, apparatus, and non-transitory computer readable medium include obtaining user input that indicates a target color and a semantic label for a region of an image to be generated; generating a noise map including noise biased towards the target color in the region indicated by the user input; and generating the image based on the noise map and the semantic label for the region using a diffusion model, wherein the image includes an object in the region that is described by the semantic label and that has the target color; [0094]: the diffusion model generates a noise map including noise biased towards the target color in the region indicated by the user input. The diffusion model generates the image based on the noise map and the semantic label for the region, where the image includes an object in the region that is described by the semantic label and that has the target color. The combination of using a color layout guidance and text guidance enables precise control of the layout of a set of objects in the generated image. According to an embodiment, the user interface includes a virtual brush that represents color and text pairing; [0098]: The user repeats this process for one or more additional entities as the user desires to include in output image 930. As an example, the user also desires to have a “calculator” and an “eye” to be generated in output image 930, thereby inputs second selection region 920 and third selection region 925 with corresponding user input 905. Output image 930 shows a hedgehog next to a calculator. The scene depicted in output image 930 is consistent with the color layout and object relations in the image canvas. The scene depicted in output image 930 is also consistent with the text prompt), and the condition information represents structural information objects in the first synthesized image ([0094]: the diffusion model generates a noise map including noise biased towards the target color in the region indicated by the user input. The diffusion model generates the image based on the noise map and the semantic label for the region, where the image includes an object in the region that is described by the semantic label and that has the target color. The combination of using a color layout guidance and text guidance enables precise control of the layout of a set of objects in the generated image. According to an embodiment, the user interface includes a virtual brush that represents color and text pairing; [0098]: The user repeats this process for one or more additional entities as the user desires to include in output image 930. As an example, the user also desires to have a “calculator” and an “eye” to be generated in output image 930, thereby inputs second selection region 920 and third selection region 925 with corresponding user input 905. Output image 930 shows a hedgehog next to a calculator. The scene depicted in output image 930 is consistent with the color layout and object relations in the image canvas. The scene depicted in output image 930 is also consistent with the text prompt). Gandelsman does not explicitly teach the condition information represents structural information of a background in the first synthesized image. Lamm teaches the condition information (position of the object) represents structural information of a background (a desired background is selected based on the position information of the object) in the first synthesized image ([0135]: The output is provided at step 1735. The output can comprise a set of coordinates defining an identified object in the image. With the object location and type identified, the augmented reality module can then superimpose a desired background (e.g. a background selected from the augmented background library) over every pixel in the image that was not identified as the object, as shown at step 1740). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply Lamm’s knowledge of determining a background based on the object location and type as taught and modify the process of Gandelsman because such a process integrates real world objects with computer generated virtual or augmented reality to enhance the user’s experience ([0002]). Regarding claim 11, Gandelsman teaches the method according to claim 10, wherein the condition information comprises at least one of text prompt information, collage information, image information, layout information (color layout comprising the color information and the layout information of an object; [0021]: In some examples, different target colors and semantic labels corresponding to a set of objects are drawn on an image canvas via a custom user interface. A noise version of a color layout is provided as input to a text-guided diffusion model. The text-guided diffusion model includes a perception model that enforces intermediate image outputs to comply with the text prompt (e.g., semantic labels). The image generation apparatus has precise control over the layout of the set of objects in an output image to be generated; [0032]: An image canvas includes color layout of two objects (e.g., hedgehog and calculator) and their target colors; [0087]: the term “color layout” includes color information of an object and layout information (e.g., position, size, and orientation) of the object to be generated in an output image; [0093]: the user selects a first region using the selection tool (e.g., the virtual brush), indicates the first region as hedgehog, and designates a color as yellow. The user further selects a second region and a third region. The user enters a corresponding class label into the label input field and a corresponding color into the color input field. The machine learning model receives the user input as guides (text guide and color layout guide); [0095]: User interface 900 includes editing image canvas 910 where users can draw or indicate color layout of one or more objects; [0096]: user interface 900 enables a user to control the object layout in an image to be generated (e.g., output image 930). User interface 900 receives a command from a user. For example, the user inputs a text prompt and a rough color layout of entities (e.g., objects) in a 2D canvas (e.g., image canvas 910). User interface 900 receives user commands where the user commands include text prompt in an image canvas. The user selects a region on the image canvas using a virtual brush associated with the phrase “hedgehog”. The user also inputs an intended color of the entity “hedgehog”. Then, the user draws a rough layout of this entity), bounding box information, or edge information. Claim(s) 8 is/are rejected under 35 U.S.C. 103 as being unpatentable over Gandelsman, and further in view of Suzuki et al. (US 2020/0160575, hereinafter Suzuki). Regarding claim 8, Gandelsman teaches the method according to claim 4, further comprising: the first dot (selection region 915, fig. 9) and the second dot (selection region 920, fig. 9), and a synthesized image (output image 930, fig. 9), wherein the synthesized image is a synthesized image in which the first object corresponding to the first class information (semantic label (class) of the entity (object); fig. 9 shows class as a semantic label “hedgehog”) is placed at the first position (selection region 915, fig. 9) and a third object corresponding to the third class information (semantic label (class) of the entity (object); fig. 9 shows class for selection region 920 is a calculator) is placed at the second position (selection region 920, fig. 9; [0089]: At operation 620, the system displays the output image to the user. In some cases, the operations of this step refer to, or may be performed by, an image generation apparatus as described with reference to FIGS. 1 and 2. In some examples, image generation apparatus displays the output image to the user via the user interface; [0097]: In some examples, image canvas 910 includes first selection region 915, second selection region 920, and third selection region 925. First selection region 915 corresponds to entity “hedgehog”. Second selection region 920 corresponds to entity “calculator”. Third selection region 925 corresponds to entity “eye” of the hedgehog, which may or may not be recited in the text prompt. Embodiments of the present disclosure are not limited to the three selection regions mentioned herein; [0098]: The user repeats this process for one or more additional entities as the user desires to include in output image 930. As an example, the user also desires to have a “calculator” and an “eye” to be generated in output image 930, thereby inputs second selection region 920 and third selection region 925 with corresponding user input 905. Output image 930 shows a hedgehog next to a calculator. The scene depicted in output image 930 is consistent with the color layout and object relations in the image canvas. The scene depicted in output image 930 is also consistent with the text prompt; [0106]: At operation 1015, the system generates the image based on the noise map and the semantic label for the region using a diffusion model, where the image includes an object in the region that is described by the semantic label and that has the target color. In some cases, the operations of this step refer to, or may be performed by, a diffusion model as described with reference to FIG. 2. The generated image is displayed to the user). Gandelsman does not explicitly teach receiving an input to change the second class information (set target area 1100 and target class, fig. 3) associated with the second dot (B1, fig. 3) to third class information (object class in fig. 3 is changed into the target class as illustrated); and based on the first dot (B2, fig. 3) and the second dot (B1, fig. 3), generating, using the image generation model, a third synthesized image (1300, fig. 3; [0022]: Accordingly, for an image including pictures of various objects, an object class for a portion of the image can be changed (for example, changing a dog into a cat, changing a dog into a lion and so on). Also, in addition to changing the class of a whole object, for example, only an image of a portion of an object, such as dog's ears, may be edited, or an editing degree of an image may be continuously changed (for example, changing an object into an intermediate class between a dog and a cat); [0042]: In the example as illustrated in FIG. 3, the user operates to set the target area 1101 for the image patch 1100 in the bounding box B1 by painting it with the color corresponding to the target class, for example; [0043]: When the target area and the target class are set, the changed class information C is generated from the class information of the to-be-edited object. It is assumed that the changed class information C is information on a data structure having the same resolution as the image patch of the object. For example, if the pre-changed class information C is also a class map, one or more ones corresponding to the target area in values forming the class map represented by the pre-changed class information C are changed into values corresponding to the target class in the changed class information C. However, the class information C may be appropriately resized at inputting the changed class information C to the generative model, and accordingly the changed class information C may not have the same resolution as the image patch of the object; [0045]: the user can operate the class information C for the image patch 1100 spatially in a free (namely, the class information C can be changed for any area of an object in an image) and continuous manner; [0047]: For example, if the object class is “Shiba dog” and the target class for the target area (such as ears) is “Shepherd dog”, the model selection unit 114 may select a generative model trained with dogs' images. Also, for example, if the object class is “dog” and the target class for the target area (such as a face) is “lion”, the model selection unit 114 may select a generative model trained with dogs' images and lions' images. The model selection unit 114 may select a generative model corresponding to only either the object class or the target class from the memory unit 120; [0048]: the model selection unit 114 may select a generative model suitable for editing the image patch of the object corresponding to at least one of the object class and the target class (namely, a trained generative model that can generate images at a high accuracy by changing the object class into the target class); [0049]: Also, when presenting the user with the multiple generative model candidates, the model selection unit 114 may provide some scores for these candidates (for example, a measure or the like to indicate an accuracy of image generation whose object class is changed into the target class)). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply Suzuki’s knowledge of changing an object class to generate an image as taught and modify the process of Gandelsman because such a process achieves flexible and extensive variation of image editing ([0022]). The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Mellado Bataller et al. (US 2021/0183008) describes establishing, by the computing system, secondary objects of interest based on the identified object class, and determining, by the computing system, that the background image contains at least one of the secondary objects of interest, where placing the transformed object in the background image in accordance with the target position value comprises placing the transformed object to be adjacent to at least one of the secondary objects of interest. Kamo (US 2023/0032860) describes to determine the size of the object based on the placement position of the object ([0064]: If the type of processing concerns placing of an object which can be transformed and placed on an image, the placement position determiner 150 may first process the object and determine the placement position so that the size of the object can be maximized in a region where the degree of importance of a background image is smaller than or equal to a specific value). Allowable Subject Matter Claims 3, 5-7 and 12 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Regarding claim 3, none of the cited prior art references of record, teach, either individually or in combination, “wherein a size of the first object is determined based on the background of the first synthesized image, the first class information, and the information indicating the first position”. Regarding claim 5, none of the cited prior art references of record, teach, either individually or in combination, “wherein a background of the first synthesized image is determined based on the first class information, the information indicating the first position, the second class information, and the information indicating the second position”. Regarding claim 6, none of the cited prior art references of record, teach, either individually or in combination, “a size of the first object placed in the first synthesized image is determined based on a background of the first synthesized image, the first class information, the information indicating the first position, the second class information, the information indicating the second position, and the size information of the second object”. Regarding claim 7, none of the cited prior art references of record, teach, either individually or in combination, “receiving an input to remove the second dot; and based on the first dot, generating, using the image generation model, a second synthesized image, wherein the second synthesized image is a synthesized image, from which the second object has been removed, that comprises the first object”. Regarding claim 12, none of the cited prior art references of record, teach, either individually or in combination, “receiving a training segmentation map and a training image corresponding to the training segmentation map; extracting, based on the training segmentation map, training dot data; using pairs of the training dot data and the training image as training data; and training, based on the training data, the image generation model”. Any inquiry concerning this communication or earlier communications from the examiner should be directed to JWALANT B AMIN whose telephone number is (571)272-2455. The examiner can normally be reached Monday-Friday 10am - 630pm CST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Said Broome can be reached at 571-272-2931. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JWALANT AMIN/Primary Examiner, Art Unit 2612
Read full office action

Prosecution Timeline

Jan 27, 2025
Application Filed
Sep 10, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737999
Synchronized Analysis of Mixed-Reality OLAP and Supply Chain Network Visualizations
1y 7m to grant Granted Sep 15, 2026
Patent 12731347
METHODS TO IMPROVE PASSTHROUGH EXPERIENCE IN LOW LIGHT CONDITIONS
2y 5m to grant Granted Sep 08, 2026
Patent 12704718
Head-Worn Wearable Device Providing Indications of Received and Monitored Sensor Data, and Methods and Systems of Use Thereof
3y 3m to grant Granted Aug 11, 2026
Patent 12688617
METHOD AND APPARATUS FOR PROCESSING THREE DIMENSIONAL GRAPHIC DATA, DEVICE, STORAGE MEDIUM AND PRODUCT
3y 5m to grant Granted Jul 21, 2026
Patent 12675157
PAUSING DEVICE OPERATION BASED ON FACIAL MOVEMENT
2y 5m to grant Granted Jul 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
79%
Grant Probability
95%
With Interview (+15.5%)
2y 7m (~11m remaining)
Median Time to Grant
Low
PTA Risk
Based on 643 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month