Prosecution Insights
Last updated: October 01, 2026
Application No. 18/956,508

CONTROLLABLE IMAGE SYNTHESIS USING EDITABLE IMAGE ELEMENTS

Non-Final OA §102§103
Filed
Nov 22, 2024
Examiner
CAI, PHUONG HAU
Art Unit
2673
Tech Center
2600 — Communications
Assignee
Adobe Inc.
OA Round
1 (Non-Final)
77%
Grant Probability
Favorable
1-2
OA Rounds
1y 0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 77% — above average
77%
Career Allowance Rate
90 granted / 117 resolved
+14.9% vs TC avg
Strong +26% interview lift
Without
With
+26.4%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
25 currently pending
Career history
150
Total Applications
across all art units

Statute-Specific Performance

§101
22.3%
-17.7% vs TC avg
§103
42.6%
+2.6% vs TC avg
§102
23.3%
-16.7% vs TC avg
§112
11.4%
-28.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 117 resolved cases

Office Action

§102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement(s) The Information disclosure statement (IDS) filed on November 22nd, 2024 has been acknowledged and considered by the examiner. Claim Objections Claim 1 is objected to because of the following informalities: The reference “withing”, in line 6, should be read as “within” to follow proper language. Appropriate correction is required. Claim 16 is objected to because of the following informalities: The reference “withing”, in line 9, should be read as “within” to follow proper language. Appropriate correction is required. Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitation(s) that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that use the word “means” or “step” but are nonetheless not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph because the claim limitation(s) recite(s) sufficient structure, materials, or acts to entirely perform the recited function. Claims 18-20 recite(s) limitation(s) that use words like “means” (or “step”) or similar terms with functional language and do invoke 35 U.S.C. 112(f): Claim 18; recites the limitation, “a text encoder configured to…,” [Line 1]. Claim 19; recites the limitation, “a segmentation component configured to segment…” [Line 1]. Claim 20; recites the limitation, “a user interface configured to obtain…” [Line 1]. Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. After a careful analysis, as disclosed above, and a careful review of the specification the following limitations in claims 18-20; (i) “a text encoder configured to…” which is supported by the instant specification, filed on November 22nd, 2024, Par. [0038] discloses ”a text encoder includes a tokenizer, a token-embedding lookup table, and a transformer-based artificial neural network ANN” thus have sufficient structure or material/act wherein is a ANN neural network with the use of a tokenizer, a token-embedding lookup table. (i) “a segmentation component configured to…” which is supported by the instant specification, filed on November 22nd, 2024, Par. [0041] discloses “segmentation component 210 include a Segment Anything Model” and Par. [0115] discloses “segmentation component including a model known as Semantic SAM” thus have sufficient structure or material/act wherein is a Segment Anything Model (SAM). (i) “a user interface configured to…” which is supported by the instant specification, filed on November 22nd, 2024, Par. [0034] discloses ”the user interface may include an audio device, such as an external speaker system, an external display device such as a display screen, or an input device (e.g., remote control device interfaced with the user interface directly or through an IO controller module. In some cases, a user interface may be a graphical user interface GUI)” thus have sufficient structure or material/act such as discussed including a display screen, or a graphical user interface GUI. If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 10 and 12 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Maciej Pesko et. al. (“US 2025/0245886 A1” hereinafter as “Pesko”). Regarding claim 10, Pesko explicitly teaches a non-transitory computer readable medium storing code for image processing, the code comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising (Par. [0094] discloses “this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, data processing apparatus”): segmenting an image to obtain a first region of the image (Par. [0069] discloses “the image features can be predetermined using a face segmentation model to generate a mask for different face segments” indicating that the face segment is segment of the image [segmenting an image] being analogous to the recited “first region”); encoding, using an encoder of an image generation model (Par. [0007] discloses “processing the image pair using a style space encoder model” moreover, the encoder is part of a Dataset Generation System as illustrated in FIG. 1, which is analogous to the recited “image generation model”), the first region of the image to obtain a first encoded image element (Par. [0007] discloses “processing the image pair using a style space encoder model to generate an embedding of the original face image and an embedding of the expressive face image in an embedding space” wherein the embedding is analogous to the recited “first encoded image elements” moreover, Par. [0069] discloses “the image features can be predetermined using a face segmentation model to generate a mask for different face segments” indicating that the face segment is segment of the image being analogous to the recited “first region”); editing the first encoded image element based on a user edit to obtain a transformed image element (Par. [0064] discloses “using the decoder to generate a generated expressive image of the initial overall editing vector” wherein, the generated expressive image is analogous to the recited “edited image”; Par. [0053] discloses “the decoder that maps the intermediate embeddings to style space, e.g., using affine transformations” indicating the decoder generate the expressive image based on the transformed image element on the embeddings of the face being transformed [the modified object]; Furthermore, Par. [0039] discloses “apply the overall editing vector to the embedding of the original image(s) in style space to generate corresponding generated expression embeddings and can decode the generated expression embedding(s) into corresponding generated expression image(s)” indicating editing the embedding [first encoded image element] to obtain the transformed element [the modified object]; moreover, Par. [0010] discloses “the optimized overall editing vector is generalizable…process user-inputted text specifying a change to facial features” indicating using a user input to generalize the editing vector); and generating, using a decoder of the image generation model, an edited image (Par. [0064] discloses “using the decoder to generate a generated expressive image of the initial overall editing vector” wherein, the generated expressive image is analogous to the recited “edited image”) based on the transformed image element (Par. [0053] discloses “the decoder that maps the intermediate embeddings to style space, e.g., using affine transformations” indicating the decoder generate the expressive image based on the transformed image element on the embeddings of the face being transformed [the modified object]). Regarding claim 12, Pesko explicitly teaches the non-transitory computer readable medium of claim 10, wherein encoding the first region comprises: individually encoding (Par. [0007] discloses “processing the image pair using a style space encoder model” which is computed based on embeddings of each image being encoded individually [since each image has a face segment to be encoded for its own], therefore, the model when determined its performance can be used to further perform the same process as have been mapped on images of future processing) each of a plurality of regions ([0069] discloses “output representation of image patches…face segments in the corresponding original image…to identify different face segments as image patches”; moreover, Par. [0089] discloses “the system can embed the original face image in the style space, e.g., using the style space encoder model…the set of losses can include one or more of a regularization loss, a Laplacian loss indicative of a measure of image sharpness, and a perceptual patch image loss as a measure of matching content between the intermediate expression image and the target expression image” indicating calculating losses to determine the accuracy of performance of the model, wherein the losses are being computed based on embeddings of each image being encoded individually [since each image has a face segment to be encoded for its own], therefore, the model when determined its performance can be used to further perform the same process as have been mapped on images of future processing) to obtain a patch embedding for each of the plurality of regions [0089] discloses “the system can embed the original face image in the style space, e.g., using the style space encoder model…the set of losses can include one or more of a regularization loss, a Laplacian loss indicative of a measure of image sharpness, and a perceptual patch image loss as a measure of matching content between the intermediate expression image and the target expression image” indicating calculating losses to determine the accuracy of performance of the model, wherein the losses are being computed based on embeddings of each image being encoded individually [since each image has a face segment to be encoded for its own], therefore, the model when determined its performance can be used to further perform the same process as have been mapped on images of future processing; and the face image as a patch with its embedding is analogous to the recited “a patch embedding”). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1, 3-4, 9, 16 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Maciej Pesko et. al. (“US 2025/0245886 A1” hereinafter as “Pesko”) in view of Nicolas LIVET et. al. (“US 2022/0301295 A1” hereinafter as “LIVET”). Regarding claim 1, Pesko explicitly teaches a method comprising (Abstract): obtaining an image (Par. [0007] discloses “a method of receiving a plurality of image pairs”); encoding, using an encoder of an image generation model (Par. [0007] discloses “processing the image pair using a style space encoder model” moreover, the encoder is part of a Dataset Generation System as illustrated in FIG. 1, which is analogous to the recited “image generation model”), a first region of the image to obtain a first encoded image element (Par. [0007] discloses “processing the image pair using a style space encoder model to generate an embedding of the original face image and an embedding of the expressive face image in an embedding space” wherein the embedding is analogous to the recited “first encoded image elements” moreover, Par. [0069] discloses “the image features can be predetermined using a face segmentation model to generate a mask for different face segments” indicating that the face segment is segment of the image being analogous to the recited “first region”); applying a transformation to the first encoded image element to obtain a transformed image element (Par. [0059] discloses “the style space engine can apply one or more kernel transformations…in order to map the embeddings into a kernel space in which the original embeddings are linearly separatable” the indicating a transformed embedding, which is analogous to the recited “transformed image element”), wherein the transformation modifies an object (Par. [0060] discloses “apply one or more kernel transformations, e.g., using a polynomial or radial basis function kernel transformation, in order to map the embeddings to a kernel space” indicating the transformation modified the embeddings; moreover, Par. [0007] discloses “processing the image pair using a style space encoder model to generate an embedding of the original face image and an embedding of the expressive face image in an embedding space” wherein the embedding is analogous to the recited “first encoded image elements”; furthermore, Par. [0069] discloses “the image features can be predetermined using a face segmentation model to generate a mask for different face segments” indicating a face, which is analogous to the recited “an object” hence, the transformation modifies the object/the face) located withing the first region of the image (Par. [0069] discloses “the image features can be predetermined using a face segmentation model to generate a mask for different face segments” indicating a face, which is analogous to the recited “an object” hence, located within the first region of the image, being a segment of the image); and generating, using a decoder of the image generation model, an edited image (Par. [0064] discloses “using the decoder to generate a generated expressive image of the initial overall editing vector” wherein, the generated expressive image is analogous to the recited “edited image”) depicting the modified object based on the transformed image element (Par. [0053] discloses “the decoder that maps the intermediate embeddings to style space, e.g., using affine transformations” indicating the decoder generate the expressive image based on the transformed image element on the embeddings of the face being transformed [the modified object]). However, Pesko does not explicitly teach the image depicting a scene with the object located within the scene. In the same field of facial feature and facial expression detection (Abstract and Par. [0092], LIVET), LIVET explicitly teaches the image depicting a scene with the object located within the scene (Par. [0089] discloses “in particular when the environment contains human bodies, human faces…require accurately-positioned features (e.g., face features, human body features), segmentation masks…detecting movements, face emotions” indicating that the image depicting a scene [environment] to detect human faces with particular facial features and facial emotions using segmentation, which is analogous to the segmentation of Pesko to obtain segments of the face within an image, moreover, LIVET further teaches the image is of an environment wherein the segmentation happens; Therefore, it would have been obvious to one or ordinary skill of the art at the time the invention was made to perform segmentation on an image to obtain segments of face, moreover, the image can be of an environment where the segmentation takes place to segment face regions. Thus in order to have perform a particular task with higher focus such as segmentation in an environment to reduce load of processing and inference outputs can be refined, see LIVET’s Abstract). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing data of the claimed invention was made to combine the teachings of Pesko of a method comprising obtaining an image; encoding, using an encoder of an image generation model, a first region of the image to obtain a first encoded image element; applying a transformation to the first encoded image element to obtain a transformed image element, wherein the transformation modifies an object located withing the first region of the image; and generating, using a decoder of the image generation model, an edited image depicting the modified object based on the transformed image element. Moreover, Pesko’s image depicting an image can be modified to be an image depicting a scene with the object located within the scene as taught in LIVET. Such a modification is the result of combing prior art elements. Pesko and LIVET share the same field of endeavor of face detection and facial expressions detection. The motivation for the proposed modification would have been to have a method comprising obtaining an image depicting a scene; encoding, using an encoder of an image generation model, a first region of the image to obtain a first encoded image element; applying a transformation to the first encoded image element to obtain a transformed image element, wherein the transformation modifies an object in the scene located withing the first region of the image; and generating, using a decoder of the image generation model, an edited image depicting the scene with the modified object based on the transformed image element. Thus in order to have perform a particular task with higher focus such as segmentation in an environment to reduce load of processing and inference outputs can be refined, see LIVET’s Abstract. Regarding claim 3, Pesko in view of LIVET, in combination, explicitly teaches the method of claim 1, wherein Pesko explicitly teaches the encoding of the first region comprises: individually encoding (Par. [0007] discloses “processing the image pair using a style space encoder model” which is computed based on embeddings of each image being encoded individually [since each image has a face segment to be encoded for its own], therefore, the model when determined its performance can be used to further perform the same process as have been mapped on images of future processing) each of a plurality of regions ([0069] discloses “output representation of image patches…face segments in the corresponding original image…to identify different face segments as image patches”; moreover, Par. [0089] discloses “the system can embed the original face image in the style space, e.g., using the style space encoder model…the set of losses can include one or more of a regularization loss, a Laplacian loss indicative of a measure of image sharpness, and a perceptual patch image loss as a measure of matching content between the intermediate expression image and the target expression image” indicating calculating losses to determine the accuracy of performance of the model, wherein the losses are being computed based on embeddings of each image being encoded individually [since each image has a face segment to be encoded for its own], therefore, the model when determined its performance can be used to further perform the same process as have been mapped on images of future processing) to obtain a patch embedding for each of the plurality of regions [0089] discloses “the system can embed the original face image in the style space, e.g., using the style space encoder model…the set of losses can include one or more of a regularization loss, a Laplacian loss indicative of a measure of image sharpness, and a perceptual patch image loss as a measure of matching content between the intermediate expression image and the target expression image” indicating calculating losses to determine the accuracy of performance of the model, wherein the losses are being computed based on embeddings of each image being encoded individually [since each image has a face segment to be encoded for its own], therefore, the model when determined its performance can be used to further perform the same process as have been mapped on images of future processing; and the face image as a patch with its embedding is analogous to the recited “a patch embedding”). Regarding claim 4, Pesko in view of LIVET, in combination, explicitly teaches the method of claim 1, wherein Pesko explicitly teaches the first encoded image element includes location information and size information of the first region (Par. [0044] discloses “a style space embedding that characterizes a magnitude and direction of change that can be applied to the embedding of an input image”; wherein, the embedding is analogous to the first encoded image element as discussed, moreover, a magnitude is analogous to a size since it indicates a overall size, scale, and the direction is analogous a location information as claimed, of the first region/the face segment). Regarding claim 9, Pesko in view of LIVET, in combination, explicitly teaches the method of claim 1, wherein Pesko explicitly teaches further comprising: obtaining an additional encoded image element from a different image including a different object (Par. [0007] discloses “processing the image pair using a style space encoder model to generate an embedding of the original face image and an embedding of the expressive face image in an embedding space” indicating two embeddings are generated by the encoder of two different image depicting two different image objects/faces, any one of the image pair that is analogous to a different image to the other which has a different object to the other, and its embedding is analogous to the recited “additional encoded element”), wherein the edited image is generated based on the additional encoded image element and with the different object (Par. [0064] discloses “using the decoder to generate a generated expressive image of the initial overall editing vector” indicating generating of an edited image using both of the embeddings, including the additional encoded image element). However, Pesko does not explicitly teach depicts the scene with the different object. In the same field of facial feature and facial expression detection (Abstract and Par. [0092], LIVET), LIVET explicitly teaches depicts the scene with the different object (Par. [0089] discloses “in particular when the environment contains human bodies, human faces…require accurately-positioned features (e.g., face features, human body features), segmentation masks…detecting movements, face emotions” indicating that the image depicting a scene [environment] to detect human faces with particular facial features and facial emotions using segmentation, which is analogous to the segmentation of Pesko to obtain segments of the face within an image, moreover, LIVET further teaches the image is of an environment wherein the segmentation happens and the environment can be obtained a plurality of images containing a plurality of faces for the processing, even of the same person having different facial emotions; Therefore, it would have been obvious to one or ordinary skill of the art at the time the invention was made to perform segmentation on an image to obtain segments of faces, even of the same person having different emotions between frames, moreover, the image can be of an environment where the segmentation takes place to segment face regions. Thus in order to have perform a particular task with higher focus such as segmentation in an environment to reduce load of processing and inference outputs can be refined, see LIVET’s Abstract). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing data of the claimed invention was made to combine the teachings of Pesko of a method of obtaining an additional encoded image element from a different image including a different object, wherein the edited image is generated based on the additional encoded image element and depicts the scene with the different object. Moreover, Pesko’s additional image can be modified to depicts the scene with the different object as taught in LIVET. Such a modification is the result of combing prior art elements. Pesko and LIVET share the same field of endeavor of face detection and facial expressions detection. The motivation for the proposed modification would have been to have a method of obtaining an additional encoded image element from a different image including a different object, wherein the edited image is generated based on the additional encoded image element and depicts the scene with the different object. Thus in order to have perform a particular task with higher focus such as segmentation in an environment to reduce load of processing and inference outputs can be refined, see LIVET’s Abstract. Regarding claim 16, Pesko explicitly teaches an apparatus comprising: at least one processor; at least one memory storing instructions that, when executed by the at least one processor, cause the processor to perform operations comprising (Abstract and Par. [0095] discloses “including by way of example a programmable processor, a computer, or multiple processors or computers” indicating the use of a computer including a processor to execute instructions stored in a memory to carry out the invention): obtaining an image (Par. [0007] discloses “a method of receiving a plurality of image pairs”); encoding, using an encoder of an image generation model (Par. [0007] discloses “processing the image pair using a style space encoder model” moreover, the encoder is part of a Dataset Generation System as illustrated in FIG. 1, which is analogous to the recited “image generation model”), a first region of the image to obtain a first encoded image element (Par. [0007] discloses “processing the image pair using a style space encoder model to generate an embedding of the original face image and an embedding of the expressive face image in an embedding space” wherein the embedding is analogous to the recited “first encoded image elements” moreover, Par. [0069] discloses “the image features can be predetermined using a face segmentation model to generate a mask for different face segments” indicating that the face segment is segment of the image being analogous to the recited “first region”); applying a transformation to the first encoded image element to obtain a transformed image element (Par. [0059] discloses “the style space engine can apply one or more kernel transformations…in order to map the embeddings into a kernel space in which the original embeddings are linearly separatable” the indicating a transformed embedding, which is analogous to the recited “transformed image element”), wherein the transformation modifies an object (Par. [0060] discloses “apply one or more kernel transformations, e.g., using a polynomial or radial basis function kernel transformation, in order to map the embeddings to a kernel space” indicating the transformation modified the embeddings; moreover, Par. [0007] discloses “processing the image pair using a style space encoder model to generate an embedding of the original face image and an embedding of the expressive face image in an embedding space” wherein the embedding is analogous to the recited “first encoded image elements”; furthermore, Par. [0069] discloses “the image features can be predetermined using a face segmentation model to generate a mask for different face segments” indicating a face, which is analogous to the recited “an object” hence, the transformation modifies the object/the face) located withing the first region of the image (Par. [0069] discloses “the image features can be predetermined using a face segmentation model to generate a mask for different face segments” indicating a face, which is analogous to the recited “an object” hence, located within the first region of the image, being a segment of the image); and generating, using a decoder of the image generation model, an edited image (Par. [0064] discloses “using the decoder to generate a generated expressive image of the initial overall editing vector” wherein, the generated expressive image is analogous to the recited “edited image”) depicting the modified object based on the transformed image element (Par. [0053] discloses “the decoder that maps the intermediate embeddings to style space, e.g., using affine transformations” indicating the decoder generate the expressive image based on the transformed image element on the embeddings of the face being transformed [the modified object]). However, Pesko does not explicitly teach the image depicting a scene with the object located within the scene. In the same field of facial feature and facial expression detection (Abstract and Par. [0092], LIVET), LIVET explicitly teaches the image depicting a scene with the object located within the scene (Par. [0089] discloses “in particular when the environment contains human bodies, human faces…require accurately-positioned features (e.g., face features, human body features), segmentation masks…detecting movements, face emotions” indicating that the image depicting a scene [environment] to detect human faces with particular facial features and facial emotions using segmentation, which is analogous to the segmentation of Pesko to obtain segments of the face within an image, moreover, LIVET further teaches the image is of an environment wherein the segmentation happens; Therefore, it would have been obvious to one or ordinary skill of the art at the time the invention was made to perform segmentation on an image to obtain segments of face, moreover, the image can be of an environment where the segmentation takes place to segment face regions. Thus in order to have perform a particular task with higher focus such as segmentation in an environment to reduce load of processing and inference outputs can be refined, see LIVET’s Abstract). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing data of the claimed invention was made to combine the teachings of Pesko of a apparatus comprising obtaining an image; encoding, using an encoder of an image generation model, a first region of the image to obtain a first encoded image element; applying a transformation to the first encoded image element to obtain a transformed image element, wherein the transformation modifies an object located withing the first region of the image; and generating, using a decoder of the image generation model, an edited image depicting the modified object based on the transformed image element. Moreover, Pesko’s image depicting an image can be modified to be an image depicting a scene with the object located within the scene as taught in LIVET. Such a modification is the result of combing prior art elements. Pesko and LIVET share the same field of endeavor of face detection and facial expressions detection. The motivation for the proposed modification would have been to have a apparatus comprising obtaining an image depicting a scene; encoding, using an encoder of an image generation model, a first region of the image to obtain a first encoded image element; applying a transformation to the first encoded image element to obtain a transformed image element, wherein the transformation modifies an object in the scene located withing the first region of the image; and generating, using a decoder of the image generation model, an edited image depicting the scene with the modified object based on the transformed image element. Thus in order to have perform a particular task with higher focus such as segmentation in an environment to reduce load of processing and inference outputs can be refined, see LIVET’s Abstract. (best understood based on the 112f interpretation section above) Regarding claim 20, Pesko in view of LIVET, in combination, explicitly teaches the apparatus of claim 16, wherein Pesko explicitly teaches further comprising: a user interface configured to obtain the transformed image element (Par. [0064] discloses “using the decoder to generate a generated expressive image of the initial overall editing vector” wherein, the generated expressive image is analogous to the recited “edited image”; Par. [0053] discloses “the decoder that maps the intermediate embeddings to style space, e.g., using affine transformations” indicating the decoder generate the expressive image based on the transformed image element on the embeddings of the face being transformed [the modified object]; Furthermore, Par. [0039] discloses “apply the overall editing vector to the embedding of the original image(s) in style space to generate corresponding generated expression embeddings and can decode the generated expression embedding(s) into corresponding generated expression image(s)” indicating editing the embedding [first encoded image element] to obtain the transformed element [the modified object]; moreover, Par. [0010] discloses “the optimized overall editing vector is generalizable…process user-inputted text specifying a change to facial features” indicating using a user input to generalize the editing vector; furthermore, Par. [0105] discloses “a user device…for purposes of displaying data to and receiving user input from” indicating the data being processed are further being displayed, the user device here is analogous to the recited user interface for displaying). Claim 2 is rejected under 35 U.S.C. 103 as being unpatentable over Maciej Pesko et. al. (“US 2025/0245886 A1” hereinafter as “Pesko”) in view of Nicolas LIVET et. al. (“US 2022/0301295 A1” hereinafter as “LIVET”) and Ji Chen (“US 2008/0037836 A1” hereinafter as “Chen”). Regarding claim 2, Pesko in view of LIVET, in combination, explicitly teaches the method of claim 1. However, Pesko in view of LIVET, in combination, does not explicitly teach segmenting the image to obtain an initial plurality of regions; and performing linear clustering on the initial plurality of regions to obtain the first region. In the same field of facial feature analysis (Title and Abstract, Chen), Chen explicitly teaches segmenting the image to obtain an initial plurality of regions (Par. [0030] discloses “positions of eyes, nose and mouth on the target image are detected…adopts a traditional algorithm for detecting the face…valid Harr wavelets” indicating obtaining different regions from a face based on key points traction according to Fig. 2, indicating segmenting the image into these key point regions corresponding to the different face features); and performing linear clustering on the initial plurality of regions to obtain the first region (Par. [0033] discloses “a fitting or regression is performed for the positions of eyes, nose and mouth on the target image and the average face, such as the curve of key points of eyes, nose and mouth on the target image fits a certain point” indicating using a regression [clustering] to process these feature positions and obtain a face information [first region as claimed]; moreover, Par. [0047] discloses “a fitting or regression is performed on the estimated position of the key points…of the average face to obtain a warp parameters (the fitting or regression in accordance with this embodiment is conducted by a linear transformation)” indicating a linear regression [linear clustering] indicating fitting of points into a linear regression line [clustering of the points/key points]; Therefore, it would have been obvious to one of ordinary skill in the art at the time the invention was made to perform detecting of facial features within a target image, wherein the target face region can be obtained by processing these facial features using a linear regression; thus in order to use such method of linear regression to it the key points to an average face to detect a face region more accurately with less error, see Chen’s Pars. [0046-0047]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing data of the claimed invention was made to combine the teachings of Pesko in view of LIVET of a method of obtaining a first region within an image as taught in Pesko. Moreover, Pesko’s obtaining of the first region can be modified to comprise segmenting the image to obtain an initial plurality of regions; and performing linear clustering on the initial plurality of regions to obtain the first region as taught in Chen. Such a modification is the result of combing prior art elements. Pesko and LIVET and Chen share the same field of endeavor of face detection and facial expressions detection. The motivation for the proposed modification would have been to have a method of obtaining a first region within an image, wherein the obtaining of the first region can be modified to comprise segmenting the image to obtain an initial plurality of regions; and performing linear clustering on the initial plurality of regions to obtain the first region. Thus in order to use such method of linear regression to it the key points to an average face to detect a face region more accurately with less error, see Chen’s Pars. [0046-0047]. Claims 5-6 are rejected under 35 U.S.C. 103 as being unpatentable over Maciej Pesko et. al. (“US 2025/0245886 A1” hereinafter as “Pesko”) in view of Nicolas LIVET et. al. (“US 2022/0301295 A1” hereinafter as “LIVET”) further in view of Stephan Marcel MANDT et. al. (“US 2019/0393903 A1” hereinafter as “MANDT”) and Zhang Liting et. al. (Foreign Patent Document “CN 121934855 A” hereinafter as “Liting”). Regarding claim 5, Pesko in view of LIVET, in combination, explicitly teaches the method of claim 1. However, Pesko in view of LIVET, in combination, does not explicitly teach wherein the image generation model is trained by training the encoder simultaneously with a training decoder and replacing the training decoder with a diffusion-based decoder. In the same field of encoder-decoder training model (Title and Abstract, MANDT), MANDT explicitly teaches wherein the image generation model is trained by training the encoder simultaneously with a training decoder (Par. [0050] discloses “to optimize the neural network parameters for the encoder model and decoder model….the encoder model and decoder model are trained and generated simultaneously using the same training input sequences” indicating the encoder and the decoder are trained simultaneously; therefore, it would have been obvious for a person of ordinary skill in the art at the time the invention was made to have a neural network with an encoder and the decoder being trained, moreover, the training of the encoder and the decoder can be performed simultaneously. Thus in order to optimize the neural network using the same training input sequence being used to train both encoder and decoder simultaneously so to produce lower reconstruction error, see MANDT’s Par. [0050]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing data of the claimed invention was made to combine the teachings of Pesko in view of LIVET of a method of training an encoder and decoder as taught in Pesko. Moreover, Pesko’s training of the encoder and decoder can be modified to have wherein the image generation model is trained by training the encoder simultaneously with a training decoder as taught in MANDT. Such a modification is the result of combing prior art elements. Pesko and MANDT share the same field of endeavor of training encoder-decoder neural network. The motivation for the proposed modification would have been to have a method of training an encoder and decoder, wherein the image generation model is trained by training the encoder simultaneously with a training decoder. Thus in order to optimize the neural network using the same training input sequence being used to train both encoder and decoder simultaneously so to produce lower reconstruction error, see MANDT’s Par. [0050]. However, Pesko in view of LIVET and MANDT, in combination, does not explicitly teach replacing the training decoder with a diffusion-based decoder. In the same field of encoder and decoder training (Abstract, Liting), Liting explicitly teaches replacing the training decoder with a diffusion-based decoder (Page 10, 6th Par., discloses “node replacement refers to replacing a performance bottleneck node with a functionally similar but more efficient one…a standard autoencoder decoder (VAE decoder) can be replaced with a node encapsulating a Dedicated Simplified Autoencoder for Stable Diffusion (TAESD decoder)” indicating replacing a standard decoder for a diffusion-based decoder; Therefore, it would be obvious for one person of ordinary skill in the art at the time the invention was made to train an encoder-decoder system, wherein the decoder can be replaced with a diffusion-based decoder. Thus in order to perform node replacement efficiently to use a more simplified version of the decoder to perform the decoding faster and more efficiently, see Liting’s Page 10, 6th Par.). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing data of the claimed invention was made to combine the teachings of Pesko in view of LIVET and MANDT of a method of training an encoder and decoder as taught in Pesko. Moreover, Pesko’s training of the encoder and decoder can be modified to have wherein the image generation model is trained by replacing the decoder with a diffusion-based decoder as taught in Liting. Such a modification is the result of combing prior art elements. Pesko and MANDT and Liting share the same field of endeavor of training encoder-decoder neural network. The motivation for the proposed modification would have been to have a method of training an encoder and decoder, and replacing the training decoder with a diffusion-based decoder. Thus in order to perform node replacement efficiently to use a more simplified version of the decoder to perform the decoding faster and more efficiently, see Liting’s Page 10, 6th Par. Regarding claim 6, Pesko in view of LIVET and further in view of MANDT and Liting, in combination, explicitly teaches the method of claim 5. However, Pesko in view of LIVET and MANDT, in combination, does not explicitly teach wherein the image generation model is trained by freezing parameters of the encoder while updating parameters of the diffusion-based decoder. In the same field of encoder and decoder training (Abstract, Liting), Liting explicitly teaches wherein the image generation model is trained by freezing parameters of the encoder while updating parameters of the diffusion-based decoder (Page 10, 6th Par., discloses “node parameter tuning refers to adjusting the parameters in the performance bottleneck node”; moreover, Page 10, 8th Par., discloses “specific process of performing node pruning on the updated workflow may include…selecting a target end point node from the end point nodes contained in the updated workflow…finally, pruning the nodes in the updated workflow other than the target endpoint node and dependent nodes to obtain the target workflow” indicating only some nodes can be pruned over the others, indicating some nodes such as some encoder can be not selected for the workflow including the parameter adjusting, some other are being performed in the workflow including the parameter adjustment; Therefore, it would be obvious for one person of ordinary skill in the art at the time the invention was made to train an encoder-decoder system, wherein the some nodes including encoders can be not selected for the workflow. Thus in order to perform node replacement efficiently to use a more simplified version of the decoder to perform the decoding faster and more efficiently, see Liting’s Page 10, 6th Par.). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing data of the claimed invention was made to combine the teachings of Pesko in view of LIVET and MANDT of a method of training an encoder and decoder as taught in Pesko. Moreover, Pesko’s training of the encoder and decoder can be modified to have an image generation model is trained by freezing parameters of the encoder while updating parameters of the diffusion-based decoder as taught in Liting. Such a modification is the result of combing prior art elements. Pesko and MANDT and Liting share the same field of endeavor of training encoder-decoder neural network. The motivation for the proposed modification would have been to have a method of training an encoder and decoder, and an image generation model is trained by freezing parameters of the encoder while updating parameters of the diffusion-based decoder. Thus in order to perform node replacement efficiently to use a more simplified version of the decoder to perform the decoding faster and more efficiently, see Liting’s Page 10, 6th Par. Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Maciej Pesko et. al. (“US 2025/0245886 A1” hereinafter as “Pesko”) in view of Nicolas LIVET et. al. (“US 2022/0301295 A1” hereinafter as “LIVET”) and Raviteja Vemulapalli et. al. (“US 2020/0151438 A1” hereinafter as “Vemulapalli”). Regarding claim 7, Pesko in view of LIVET, in combination, explicitly teaches the method of claim 1. However, Pesko in view of LIVET, in combination, does not explicitly teach the image generation model is trained by obtaining a plurality of encoded image elements and dropping one or more of the plurality of encoded image elements to obtain a reduced set of encoded image elements. In the same field of facial expression detection (Title and Abstract, Vemulapalli), Vemulapalli explicitly teaches the image generation model is trained by obtaining a plurality of encoded image elements (Par. [0045] discloses “there is no requirement for the image triplets to include two images that match and one which does not, the processing involved with selecting the image triplets may be reduced. For example, any three arbitrary images can be used to provide useful training information” indicating obtaining training information for training of a facial expression detection model based on training images, moreover, the training information are encoded information, Par. [0035] discloses “the facial expression information encoded by the facial expression embedding” which is analogous to the recited “encoded image elements”) and dropping one or more of the plurality of encoded image elements to obtain a reduced set of encoded image elements (Par. [0045] discloses “there is no requirement for the image triplets to include two images that match and one which does not, the processing involved with selecting the image triplets may be reduced. For example, any three arbitrary images can be used to provide useful training information” indicating a reduced selecting of training information, indicating a dropping of some of the encoded image elements to obtained a reduced training information [analogous to the recited “reduced set of encoded image elements”]; Therefore, it would be obvious for one person of ordinary skill in the art at the time of the invention was made to have a process of obtaining encoded information as training information for training of a facial expression detection model, wherein the training is carried out by reducing selecting of images for the training as training information; Thus in order to perform training with reduced approach of selecting training method such as discussed, so that expenditure of time and financial expense can be further reduced for performance of an effective model, see Vemulapalli’s Par. [0045]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing data of the claimed invention was made to combine the teachings of Pesko in view of LIVET of a method of training of a model as taught in Pesko. Moreover, Pesko’s training of the model can be modified to have wherein the image generation model is trained by obtaining a plurality of encoded image elements and dropping one or more of the plurality of encoded image elements to obtain a reduced set of encoded image elements as taught in Vemulapalli. Such a modification is the result of combing prior art elements. Pesko and MANDT and Liting share the same field of endeavor of training neural network. The motivation for the proposed modification would have been to have a method of training of a model, wherein the image generation model is trained by obtaining a plurality of encoded image elements and dropping one or more of the plurality of encoded image elements to obtain a reduced set of encoded image elements. Thus in order to perform training with reduced approach of selecting training method such as discussed, so that expenditure of time and financial expense can be further reduced for performance of an effective model, see Vemulapalli’s Par. [0045]. Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Maciej Pesko et. al. (“US 2025/0245886 A1” hereinafter as “Pesko”) in view of Nicolas LIVET et. al. (“US 2022/0301295 A1” hereinafter as “LIVET”) further in view of Stephan Marcel MANDT et. al. (“US 2019/0393903 A1” hereinafter as “MANDT”) and Zhang Liting et. al. (Foreign Patent Document “CN 121934855 A” hereinafter as “Liting”) and Bao Tran (“US 2023/0351102 A1” hereinafter as “Tran”). Regarding claim 8, Pesko in view of LIVET and further in view of MANDT and Liting, in combination, explicitly teaches the method of claim 6, wherein Pesko explicitly teaches generating the edited image comprises: conditioning the generation of the edited image with the text features (Par. [0010] discloses “a generative adversarial network that can process user-inputted text specifying a change to facial features”). However, Pesko in view of LIVET and further in view of MANDT and Liting, in combination, does not explicitly teach encoding a text prompt to obtain text features. In the same field of user-input text prompt for command execution (Title and Abstract, Tran), Tran explicitly teaches encoding a text prompt to obtain text features (Par. [0004] discloses “using a transformer with an encoder on the text prompt”; moreover, Par. [0156] discloses “encode abstracts/summaries into idea representations” and Par. [0157] discloses “a concept encoder is used in addition…to embed the respective inputs” indicating encode a text prompt to carry out a command/task, which is analogous to Pesko’s using of user-input text to carry out a task; Therefore, it would be obvious for one person of ordinary skill in the art at the time the invention was made to perform a task/command based on having a user entering a text prompt to execute such task, by encoding the text prompt to obtain embedded features. Thus in order to perform such method to obtain task performed more interactively and speedily, see Tran’s Par. [0010]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing data of the claimed invention was made to combine the teachings of Pesko in view of LIVET and further in view of MANDT and Liting of a method of generating an edited image comprises conditioning a generation of the edited image with text features. Moreover, Pesko’s generating the edited image can be modified to perform encoding a text prompt to obtain text features as taught in Tran. Such a modification is the result of combing prior art elements. Pesko and TRan share the same field of endeavor of using user-input text prompt to carry out a task. The motivation for the proposed modification would have been to have a method of generating an edited image comprises: encoding a text prompt to obtain text features; and conditioning the generation of the edited image with the text features. Thus in order to perform such method to obtain task performed more interactively and speedily, see Tran’s Par. [0010]. Claim 11 is rejected under 35 U.S.C. 103 as being unpatentable over Maciej Pesko et. al. (“US 2025/0245886 A1” hereinafter as “Pesko”) in view of Ji Chen (“US 2008/0037836 A1” hereinafter as “Chen”). Regarding claim 11, Pesko explicitly teaches the non-transitory computer readable medium of claim 10. However, Pesko does not explicitly teach wherein segmenting the image comprises: performing linear clustering on an initial plurality of regions to obtain the first region. In the same field of facial feature analysis (Title and Abstract, Chen), Chen explicitly teaches wherein segmenting the image comprises (Par. [0030] discloses “positions of eyes, nose and mouth on the target image are detected…adopts a traditional algorithm for detecting the face…valid Harr wavelets” indicating obtaining different regions from a face based on key points traction according to Fig. 2, indicating segmenting the image into these key point regions corresponding to the different face features): and performing linear clustering on the initial plurality of regions to obtain the first region (Par. [0033] discloses “a fitting or regression is performed for the positions of eyes, nose and mouth on the target image and the average face, such as the curve of key points of eyes, nose and mouth on the target image fits a certain point” indicating using a regression [clustering] to process these feature positions and obtain a face information [first region as claimed]; moreover, Par. [0047] discloses “a fitting or regression is performed on the estimated position of the key points…of the average face to obtain a warp parameters (the fitting or regression in accordance with this embodiment is conducted by a linear transformation)” indicating a linear regression [linear clustering] indicating fitting of points into a linear regression line [clustering of the points/key points]; Therefore, it would have been obvious to one of ordinary skill in the art at the time the invention was made to perform detecting of facial features within a target image, wherein the target face region can be obtained by processing these facial features using a linear regression; thus in order to use such method of linear regression to it the key points to an average face to detect a face region more accurately with less error, see Chen’s Pars. [0046-0047]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing data of the claimed invention was made to combine the teachings of Pesko of a non-transitory computer readable medium obtaining a first region within an image as taught in Pesko. Moreover, Pesko’s obtaining of the first region can be modified to comprise segmenting the image to obtain an initial plurality of regions; and performing linear clustering on the initial plurality of regions to obtain the first region as taught in Chen. Such a modification is the result of combing prior art elements. Pesko and Chen share the same field of endeavor of face detection and facial expressions detection. The motivation for the proposed modification would have been to have a non-transitory computer readable medium of obtaining a first region within an image, wherein the obtaining of the first region can be modified to comprise segmenting the image to obtain an initial plurality of regions; and performing linear clustering on the initial plurality of regions to obtain the first region. Thus in order to use such method of linear regression to it the key points to an average face to detect a face region more accurately with less error, see Chen’s Pars. [0046-0047]. Claims 13-14 are rejected under 35 U.S.C. 103 as being unpatentable over Maciej Pesko et. al. (“US 2025/0245886 A1” hereinafter as “Pesko”) in view of Stephan Marcel MANDT et. al. (“US 2019/0393903 A1” hereinafter as “MANDT”) and Zhang Liting et. al. (Foreign Patent Document “CN 121934855 A” hereinafter as “Liting”). Regarding claim 13, Pesko explicitly teaches the non-transitory computer readable medium of claim 12. However, Pesko does not explicitly teach wherein the image generation model is trained by training the encoder simultaneously with a training decoder and replacing the training decoder with a diffusion-based decoder. In the same field of encoder-decoder training model (Title and Abstract, MANDT), MANDT explicitly teaches wherein the image generation model is trained by training the encoder simultaneously with a training decoder (Par. [0050] discloses “to optimize the neural network parameters for the encoder model and decoder model….the encoder model and decoder model are trained and generated simultaneously using the same training input sequences” indicating the encoder and the decoder are trained simultaneously; therefore, it would have been obvious for a person of ordinary skill in the art at the time the invention was made to have a neural network with an encoder and the decoder being trained, moreover, the training of the encoder and the decoder can be performed simultaneously. Thus in order to optimize the neural network using the same training input sequence being used to train both encoder and decoder simultaneously so to produce lower reconstruction error, see MANDT’s Par. [0050]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing data of the claimed invention was made to combine the teachings of Pesko of a non-transitory computer readable medium of training an encoder and decoder as taught in Pesko. Moreover, Pesko’s training of the encoder and decoder can be modified to have wherein the image generation model is trained by training the encoder simultaneously with a training decoder as taught in MANDT. Such a modification is the result of combing prior art elements. Pesko and MANDT share the same field of endeavor of training encoder-decoder neural network. The motivation for the proposed modification would have been to have non-transitory computer readable medium of training an encoder and decoder, wherein the image generation model is trained by training the encoder simultaneously with a training decoder. Thus in order to optimize the neural network using the same training input sequence being used to train both encoder and decoder simultaneously so to produce lower reconstruction error, see MANDT’s Par. [0050]. However, Pesko in view of MANDT, in combination, does not explicitly teach replacing the training decoder with a diffusion-based decoder. In the same field of encoder and decoder training (Abstract, Liting), Liting explicitly teaches replacing the training decoder with a diffusion-based decoder (Page 10, 6th Par., discloses “node replacement refers to replacing a performance bottleneck node with a functionally similar but more efficient one…a standard autoencoder decoder (VAE decoder) can be replaced with a node encapsulating a Dedicated Simplified Autoencoder for Stable Diffusion (TAESD decoder)” indicating replacing a standard decoder for a diffusion-based decoder; Therefore, it would be obvious for one person of ordinary skill in the art at the time the invention was made to train an encoder-decoder system, wherein the decoder can be replaced with a diffusion-based decoder. Thus in order to perform node replacement efficiently to use a more simplified version of the decoder to perform the decoding faster and more efficiently, see Liting’s Page 10, 6th Par.). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing data of the claimed invention was made to combine the teachings of Pesko in view of MANDT of a non-transitory computer readable medium of training an encoder and decoder as taught in Pesko. Moreover, Pesko’s training of the encoder and decoder can be modified to have wherein the image generation model is trained by replacing the decoder with a diffusion-based decoder as taught in Liting. Such a modification is the result of combing prior art elements. Pesko and MANDT and Liting share the same field of endeavor of training encoder-decoder neural network. The motivation for the proposed modification would have been to have a non-transitory computer readable medium of a method of training an encoder and decoder, and replacing the training decoder with a diffusion-based decoder. Thus in order to perform node replacement efficiently to use a more simplified version of the decoder to perform the decoding faster and more efficiently, see Liting’s Page 10, 6th Par. Regarding claim 14, Pesko in view of MANDT and Liting, in combination, explicitly teaches the non-transitory computer readable medium of claim 13. However, Pesko in view of MANDT and Liting, in combination, does not explicitly teach wherein the image generation model is trained by freezing parameters of the encoder while updating parameters of the diffusion-based decoder. In the same field of encoder and decoder training (Abstract, Liting), Liting explicitly teaches wherein the image generation model is trained by freezing parameters of the encoder while updating parameters of the diffusion-based decoder (Page 10, 6th Par., discloses “node parameter tuning refers to adjusting the parameters in the performance bottleneck node”; moreover, Page 10, 8th Par., discloses “specific process of performing node pruning on the updated workflow may include…selecting a target end point node from the end point nodes contained in the updated workflow…finally, pruning the nodes in the updated workflow other than the target endpoint node and dependent nodes to obtain the target workflow” indicating only some nodes can be pruned over the others, indicating some nodes such as some encoder can be not selected for the workflow including the parameter adjusting, some other are being performed in the workflow including the parameter adjustment; Therefore, it would be obvious for one person of ordinary skill in the art at the time the invention was made to train an encoder-decoder system, wherein the some nodes including encoders can be not selected for the workflow. Thus in order to perform node replacement efficiently to use a more simplified version of the decoder to perform the decoding faster and more efficiently, see Liting’s Page 10, 6th Par.). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing data of the claimed invention was made to combine the teachings of Pesko in view of MANDT of a non-transitory computer readable medium of training an encoder and decoder as taught in Pesko. Moreover, Pesko’s training of the encoder and decoder can be modified to have an image generation model is trained by freezing parameters of the encoder while updating parameters of the diffusion-based decoder as taught in Liting. Such a modification is the result of combing prior art elements. Pesko and MANDT and Liting share the same field of endeavor of training encoder-decoder neural network. The motivation for the proposed modification would have been to have a non-transitory computer readable medium of training an encoder and decoder, and an image generation model is trained by freezing parameters of the encoder while updating parameters of the diffusion-based decoder. Thus in order to perform node replacement efficiently to use a more simplified version of the decoder to perform the decoding faster and more efficiently, see Liting’s Page 10, 6th Par. Claim 15 is rejected under 35 U.S.C. 103 as being unpatentable over Maciej Pesko et. al. (“US 2025/0245886 A1” hereinafter as “Pesko”) in view of Raviteja Vemulapalli et. al. (“US 2020/0151438 A1” hereinafter as “Vemulapalli”). Regarding claim 15, Pesko explicitly teaches the non-transitory computer readable medium of claim 10. However, Pesko does not explicitly teach wherein the image generation model is trained by obtaining a plurality of encoded image elements and dropping one or more of the plurality of encoded image elements to obtain a reduced set of encoded image elements. In the same field of facial expression detection (Title and Abstract, Vemulapalli), Vemulapalli explicitly teaches wherein the image generation model is trained by obtaining a plurality of encoded image elements (Par. [0045] discloses “there is no requirement for the image triplets to include two images that match and one which does not, the processing involved with selecting the image triplets may be reduced. For example, any three arbitrary images can be used to provide useful training information” indicating obtaining training information for training of a facial expression detection model based on training images, moreover, the training information are encoded information, Par. [0035] discloses “the facial expression information encoded by the facial expression embedding” which is analogous to the recited “encoded image elements”) and dropping one or more of the plurality of encoded image elements to obtain a reduced set of encoded image elements (Par. [0045] discloses “there is no requirement for the image triplets to include two images that match and one which does not, the processing involved with selecting the image triplets may be reduced. For example, any three arbitrary images can be used to provide useful training information” indicating a reduced selecting of training information, indicating a dropping of some of the encoded image elements to obtained a reduced training information [analogous to the recited “reduced set of encoded image elements”]; Therefore, it would be obvious for one person of ordinary skill in the art at the time of the invention was made to have a process of obtaining encoded information as training information for training of a facial expression detection model, wherein the training is carried out by reducing selecting of images for the training as training information; Thus in order to perform training with reduced approach of selecting training method such as discussed, so that expenditure of time and financial expense can be further reduced for performance of an effective model, see Vemulapalli’s Par. [0045]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing data of the claimed invention was made to combine the teachings of Pesko of a non-transitory computer readable mediumof training of a model as taught in Pesko. Moreover, Pesko’s training of the model can be modified to have wherein the image generation model is trained by obtaining a plurality of encoded image elements and dropping one or more of the plurality of encoded image elements to obtain a reduced set of encoded image elements as taught in Vemulapalli. Such a modification is the result of combing prior art elements. Pesko and MANDT and Liting share the same field of endeavor of training neural network. The motivation for the proposed modification would have been to have a non-transitory computer readable mediumof training of a model, wherein the image generation model is trained by obtaining a plurality of encoded image elements and dropping one or more of the plurality of encoded image elements to obtain a reduced set of encoded image elements. Thus in order to perform training with reduced approach of selecting training method such as discussed, so that expenditure of time and financial expense can be further reduced for performance of an effective model, see Vemulapalli’s Par. [0045]. Claim 17 is rejected under 35 U.S.C. 103 as being unpatentable over Maciej Pesko et. al. (“US 2025/0245886 A1” hereinafter as “Pesko”) in view of Nicolas LIVET et. al. (“US 2022/0301295 A1” hereinafter as “LIVET”) and Zhang Liting et. al. (Foreign Patent Document “CN 121934855 A” hereinafter as “Liting”). Regarding claim 17, Pesko in view of LIVET, in combination, explicitly teaches the apparatus of claim 16. However, Pesko in view of LIVET, in combination, does not explicitly teach wherein: the image generation model comprises a diffusion model. In the same field of encoder and decoder training (Abstract, Liting), Liting explicitly teaches wherein: the image generation model comprises a diffusion model (Page 10, 6th Par., discloses “node replacement refers to replacing a performance bottleneck node with a functionally similar but more efficient one…a standard autoencoder decoder (VAE decoder) can be replaced with a node encapsulating a Dedicated Simplified Autoencoder for Stable Diffusion (TAESD decoder)” indicating replacing a standard decoder for a diffusion-based decoder; Therefore, it would be obvious for one person of ordinary skill in the art at the time the invention was made to train an encoder-decoder system, wherein the decoder can be replaced with a diffusion-based decoder. Thus in order to perform node replacement efficiently to use a more simplified version of the decoder to perform the decoding faster and more efficiently, see Liting’s Page 10, 6th Par.). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing data of the claimed invention was made to combine the teachings of Pesko in view of LIVET of aa apparatus of training an encoder and decoder as taught in Pesko. Moreover, Pesko’s training of the encoder and decoder can be modified to have wherein the image generation model comprises a diffusion model as taught in Liting. Such a modification is the result of combing prior art elements. Pesko and MANDT and Liting share the same field of endeavor of training encoder-decoder neural network. The motivation for the proposed modification would have been to have an apparatus of training an encoder and decoder wherein the image generation model comprises a diffusion model. Thus in order to perform node replacement efficiently to use a more simplified version of the decoder to perform the decoding faster and more efficiently, see Liting’s Page 10, 6th Par. Claim 18 is rejected under 35 U.S.C. 103 as being unpatentable over Maciej Pesko et. al. (“US 2025/0245886 A1” hereinafter as “Pesko”) in view of Nicolas LIVET et. al. (“US 2022/0301295 A1” hereinafter as “LIVET”) further in view of Bao Tran (“US 2023/0351102 A1” hereinafter as “Tran”) and Seoung Wug Oh et. al. (“US 2025/0119624 A1” hereinafter as “Oh”). (best understood based on the 112f interpretation above) Regarding claim 18, Pesko in view of LIVET and Liting, in combination, explicitly teaches apparatus of claim 16. However, Pesko in view of LIVET, in combination, does not explicitly teach further comprising: a text encoder configured to encode a text prompt. In the same field of user-input text prompt for command execution (Title and Abstract, Tran), Tran explicitly teaches further comprising: a text encoder configured to encode a text prompt (Par. [0004] discloses “using a transformer with an encoder on the text prompt”; moreover, Par. [0156] discloses “encode abstracts/summaries into idea representations” and Par. [0157] discloses “a concept encoder is used in addition…to embed the respective inputs” indicating encode a text prompt to carry out a command/task, which is analogous to Pesko’s using of user-input text to carry out a task; Therefore, it would be obvious for one person of ordinary skill in the art at the time the invention was made to perform a task/command based on having a user entering a text prompt to execute such task, by encoding the text prompt to obtain embedded features. Thus in order to perform such method to obtain task performed more interactively and speedily, see Tran’s Par. [0010]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing data of the claimed invention was made to combine the teachings of Pesko in view of LIVET of an apparatus of generating an edited image comprises conditioning a generation of the edited image with text features. Moreover, Pesko’s generating the edited image can be modified to further comprise a text encoder configured to encode a text prompt. as taught in Tran. Such a modification is the result of combing prior art elements. Pesko and Tran share the same field of endeavor of using user-input text prompt to carry out a task. The motivation for the proposed modification would have been to have an apparatus of generating an edited image further comprising a text encoder configured to encode a text prompt. Thus in order to perform such method to obtain task performed more interactively and speedily, see Tran’s Par. [0010]. However, Pesko in view of LIVET and Tran, in combination, does not explicitly teach wherein the text encoder is a ANN neural network with the use of a tokenizer, a token-embedding lookup table. In the same field of processing image using a text encoder (Abstract and Par. [0025], Oh), Oh explicitly teaches wherein the text encoder is a ANN neural network with the use of a tokenizer, a token-embedding lookup table (Par. [0025] discloses “the text-to-image diffusion models incorporate a pre-trained text encoder”; moreover, Par. [0038] discloses the processing includes the use of artificial neural network ANNs; moreover, Par. [0052] discloses “text encoder includes tokenizer and embedding lookup table”; Therefore, it would be obvious for one person of ordinary skill in the art at the time the invention was made to have an image generation that include text encoder, wherein the text encoder being a ANN neural network with a tokenizer and an embedding lookup table. Thus in order to use such method to break down text into smaller components for processing more efficiently and perform frame-wise token embedding and image processing more accurately, see Oh’s Par. [0052] and Abstract). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing data of the claimed invention was made to combine the teachings of Pesko in view of LIVET and Tran of an apparatus of generating an edited image comprises conditioning a generation of the edited image with text features. Moreover, Pesko’s generating the edited image can be modified to further comprise a text encoder configured to encode a text prompt as taught in Tran and Pesko’s text encoder can be modified to be a ANN neural network with the use of a tokenizer, a token-embedding lookup table as taught in Oh. Such a modification is the result of combing prior art elements. Pesko and Tran and Oh share the same field of endeavor of using user-input text prompt to carry out a task. The motivation for the proposed modification would have been to have an apparatus of generating an edited image further comprising a text encoder configured to encode a text prompt, wherein the encoder is a ANN neural network with the use of a tokenizer, a token-embedding lookup table. Thus in order to use such method to break down text into smaller components for processing more efficiently and perform frame-wise token embedding and image processing more accurately, see Oh’s Par. [0052] and Abstract. Claim 19 is rejected under 35 U.S.C. 103 as being unpatentable over Maciej Pesko et. al. (“US 2025/0245886 A1” hereinafter as “Pesko”) in view of Nicolas LIVET et. al. (“US 2022/0301295 A1” hereinafter as “LIVET”) and Chao Yang et. al. (“US 2026/0044941” hereinafter as “Yang”). (best understood based on the 112f interpretation section above) Regarding claim 19, Pesko in view of LIVET, in combination, explicitly teaches the apparatus of claim 16, wherein Pesko explicitly teaches further comprising: a segmentation component configured to segment the image to obtain a plurality of regions (Par. [0069] discloses “the image features can be predetermined using a face segmentation model to generate a mask for different face segments” indicating that the face segment is segment of the image [segmenting an image] being analogous to the recited “first region”; the processor performing this segmentation is analogous to the recited segmentation component). However, Pesko in view of LIVET, in combination, does not explicitly teach wherein the segmentation component being a SAM (Segment Anything Model). In the same field of image segmentation (Title and Abstract, Yang), Yang explicitly teaches wherein the segmentation component being a SAM (Segment Anything Model) (Par. [0006] discloses “the segment anything model (SAM)…is an advanced foundational model that shows outstanding performance in image segmentation”; Therefore, it would be obvious for one person of ordinary skill in the art at the time the invention was made to have an image generation model that perform image segmentation, wherein the image segmentation is a segment anything model [SAM]. Thus in order to perform more effective and robust image segmentation without any human intervention, see Yang’s Par. [0006]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing data of the claimed invention was made to combine the teachings of Pesko in view of LIVET of an apparatus of generating an edited image comprises conditioning using image segmentation. Moreover, Pesko’s image segmentation being a segment anything model (SAM) as taught in Yang. Such a modification is the result of combing prior art elements. Pesko and LIVET and Yang share the same field of endeavor of image segmentation. The motivation for the proposed modification would have been to have an apparatus of generating an edited image comprises conditioning using image segmentation, wherein the image segmentation being a segment anything model (SAM). Thus in order to perform more effective and robust image segmentation without any human intervention, see Yang’s Par. [0006]. Pertinent Prior Art(s) The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Guerrero, Paul et. al., “US 2026/0080644 A1”, discloses a computing system receives an image of a three-dimensional (“3D”) object, a text prompt, and a 3D transformation. The computing system generates an initial state for a diffusion model. The computing system generates, using the diffusion model, a second image of the 3D object and intermediate representations of the second image, based on the initial state, a depth map, and the text prompt. The computing system generates 3D representations of the second image based on the intermediate representations, transformed 3D representations of the second image by applying the 3D transformation, and edited intermediate representations. The computing system generates, using the diffusion model and the edited intermediate representations, an edited image of the 3D object, based on the initial state, a second depth map, and the text prompt and outputs the edited image. Mehr, Eloi et. al., “US 11468268 B2”, discloses A computer-implemented method for learning an autoencoder notably is provided. The method includes obtaining a dataset of images. Each image includes a respective object representation. The method also includes learning the autoencoder based on the dataset. The learning includes minimization of a reconstruction loss. The reconstruction loss includes a term that penalizes a distance for each respective image. The penalized distance is between the result of applying the autoencoder to the respective image and the set of results of applying at least part of a group of transformations to the object representation of the respective image. Such a method provides an improved solution to learn an autoencoder. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to PHUONG HAU CAI whose telephone number is (571)272-9424. The examiner can normally be reached M-F 8:30 am - 5:00pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chineyere Wills-Burns can be reached at (571) 272-9752. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /PHUONG HAU CAI/Examiner, Art Unit 2673 /CHINEYERE WILLS-BURNS/Supervisory Patent Examiner, Art Unit 2673
Read full office action

Prosecution Timeline

Nov 22, 2024
Application Filed
Jul 01, 2026
Non-Final Rejection mailed — §102, §103
Sep 14, 2026
Interview Requested
Sep 25, 2026
Applicant Interview (Telephonic)
Sep 26, 2026
Examiner Interview Summary

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12700217
SYSTEM AND METHOD FOR PROCESSING TRAINING DATASET ASSOCIATED WITH SYNTHETIC IMAGE
3y 8m to grant Granted Aug 04, 2026
Patent 12688683
Method and System for Optimization of a Human-Machine Team for Geographic Region Digitization
2y 2m to grant Granted Jul 21, 2026
Patent 12682605
SYSTEMS, METHODS, AND APPARATUS FOR IMAGE CLASSIFICATION WITH DOMAIN INVARIANT REGULARIZATION
3y 10m to grant Granted Jul 14, 2026
Patent 12639955
AUTOMATED VEHICLE IDENTIFICATION BASED ON CAR-FOLLOWING DATA WITH MACHINE LEARNING
3y 9m to grant Granted May 26, 2026
Patent 12632931
INSPECTION SYSTEM, IMAGE PROCESSING METHOD, AND DEFECT INSPECTION DEVICE
3y 7m to grant Granted May 19, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
77%
Grant Probability
99%
With Interview (+26.4%)
2y 11m (~1y 0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 117 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month