DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
Use of the word “means” (or “step for”) in a claim with functional language creates a rebuttable presumption that the claim element is to be treated in accordance with 35 U.S.C. 112(f) (pre-AIA 35 U.S.C. 112, sixth paragraph). The presumption that 35 U.S.C. 112(f) (pre-AIA 35 U.S.C. 112, sixth paragraph) is invoked is rebutted when the function is recited with sufficient structure, material, or acts within the claim itself to entirely perform the recited function.
Absence of the word “means” (or “step for”) in a claim creates a rebuttable presumption that the claim element is not to be treated in accordance with 35 U.S.C. 112(f) (pre-AIA 35 U.S.C. 112, sixth paragraph). The presumption that 35 U.S.C. 112(f) (pre-AIA 35 U.S.C. 112, sixth paragraph) is not invoked is rebutted when the claim element recites function but fails to recite sufficiently definite structure, material or acts to perform that function.
Claim elements in this application that use the word “means” (or “step for”) are presumed to invoke 35 U.S.C. 112(f) except as otherwise indicated in an Office action. Similarly, claim elements that do not use the word “means” (or “step for”) are presumed not to invoke 35 U.S.C. 112(f) except as otherwise indicated in an Office action.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier.
Such claim limitation(s) is/are:
Tokenization component in claim 15
**The component disclosed above has been interpreted as tied to the structure of a processor as disclosed in the originally filed specification at least in paragraphs [0075], [0094]-[0095], and [0098]
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-3, 14, 15, and 20 are rejected under 35 U.S.C. 102(a)(1)/(a)(2) as being anticipated by Daha (US 2024/0296595 A1, hereinafter referenced “Daha”).
In regards to claim 1. Daha discloses a method (Daha, Abstract) comprising:
-obtaining a text prompt describing an image element (Daha, para [0034]; Reference discloses wanting the customized image described, the user may input the textual prompt 106 of “Username holding his dog” or “Me holding my dog.”);
-generating, using a transformer prior model, an image embedding based on the text prompt, wherein the image embedding represents visual features of the image element (Daha, para [0024] and [0025]; Reference at para [0024] discloses one current example of an image-generating AI engine that can be used as the AI engine in the techniques described herein is DALL-E…DALL-E is a transformer language model similar to GPT but trained to output images in response to user input (i.e. transformer prior model). Para [0025] discloses DALL-E and other similar image generating AI engines process input by generating tokens. These tokens are created by encoding the textual description into a numerical representation, which is then fed into the model to generate an image. The tokens in DALL-E represent specific elements or attributes of the desired image, such as shape, color, and style (i.e. generate image embedding based on the text prompt, wherein the image embedding represents visual features of the image element));
-and generating, using an image generation model, a synthetic image depicting the image element based on the image embedding (Daha, para [0038] and [0040]; Reference at para [0038] discloses in this example, the user will provide as input to a personalized image generation system 110, an image or images 102 of the user and an image or images 104 of the user's dog. The user will also give a textual description 106 of the image to be generated, i.e., “Username holding his dog” or “Me holding my dog.” Para [0040] discloses In executing the command, the image generating AI engine 112 can use the tokens 108 and NLP layer 118 of the fine-tuning mechanism to personalize or customize the output image 116 (i.e. synthetic image generated from image generating AI or model based corresponding to the tokens (i.e. image embedding)).
In regards to claim 2. Daha discloses the method of claim 1.
Daha further discloses
-wherein generating the image embedding comprises: tokenizing the text prompt to obtain a plurality of text tokens, wherein the image embedding is generated based on the plurality of text tokens (Daha, para [0024] and [0025]; Reference at para [0024] discloses one current example of an image-generating AI engine that can be used as the AI engine in the techniques described herein is DALL-E…DALL-E is a transformer language model similar to GPT but trained to output images in response to user input. Para [0025] discloses DALL-E and other similar image generating AI engines process input by generating tokens. These tokens are created by encoding the textual description into a numerical representation, which is then fed into the model to generate an image (i.e. tokenizing the text prompt to obtain a plurality of text tokens, wherein the image embedding is generated based on the plurality of text tokens . The tokens in DALL-E represent specific elements or attributes of the desired image, such as shape, color, and style (i.e. generate image embedding based on the text prompt, wherein the image embedding represents visual features of the image element)).
In regards to claim 3. Daha discloses the method of claim 2.
Daha further discloses
-wherein generating the image embedding comprises: generating a plurality of text token embeddings based on the plurality of text tokens, respectively, wherein each of the plurality of text token embeddings represent text features, and wherein the image embedding is generated based on the plurality of text token embeddings (Daha, para [0037] and [0040]; Reference at para [0037] discloses In addition to the set of tokens 108, the fine-tuning mechanism may also include an additional Natural Language Processing (NLP) layer 118 to be added to the NLP system of the image-generating AI engine 112. The additional NLP layer 118 associates the relevant wording of the textual input to the corresponding tokens in the fine-tuning mechanism (i.e. generating a plurality of text token embeddings based on the plurality of text tokens, respectively, wherein each of the plurality of text token embeddings represent text features). Para [0040] discloses The fine-tuning mechanism 114 is thus transmitted to the image generating AI engine 112 and implemented, as described above. The command 106 is then executed by the image generating AI engine 112. In executing the command, the image generating AI engine 112 can use the tokens 108 and NLP layer 118 of the fine-tuning mechanism to personalize or customize the output image 116 (i.e. The fine-tuning mechanism 114 is thus transmitted to the image generating AI engine 112 and implemented, as described above. The command 106 is then executed by the image generating AI engine 112. In executing the command, the image generating AI engine 112 can use the tokens 108 and NLP layer 118 of the fine-tuning mechanism to personalize or customize the output image 116)).
In regards to claim 14. Daha discloses an apparatus comprising:
-a memory component; a processing device coupled to the memory component (Daha, para [0074]; Reference discloses the machine 600 may include processors 610, memory 630, and I/O components 650, which may be communicatively coupled via, for example, a bus 602);
-a transformer prior model comprising parameters stored in the memory component and trained to generate an image embedding based on a text prompt, wherein the image embedding represents visual features of the image element (Daha, para [0024] and [0025]; Reference at para [0024] discloses one current example of an image-generating AI engine that can be used as the AI engine in the techniques described herein is DALL-E…DALL-E is a transformer language model similar to GPT but trained to output images in response to user input (i.e. transformer prior model comprising parameters stored in the memory component). Para [0025] discloses DALL-E and other similar image generating AI engines process input by generating tokens. These tokens are created by encoding the textual description into a numerical representation, which is then fed into the model to generate an image. The tokens in DALL-E represent specific elements or attributes of the desired image, such as shape, color, and style (i.e. generate image embedding based on the text prompt, wherein the image embedding represents visual features of the image element));
-and an image generation model comprising parameters stored in the memory component and trained to generate a synthetic image depicting the image element based on the image embedding (Daha, para [0038] and [0040]; Reference at para [0038] discloses in this example, the user will provide as input to a personalized image generation system 110, an image or images 102 of the user and an image or images 104 of the user's dog. The user will also give a textual description 106 of the image to be generated, i.e., “Username holding his dog” or “Me holding my dog.” Para [0040] discloses In executing the command, the image generating AI engine 112 can use the tokens 108 and NLP layer 118 of the fine-tuning mechanism to personalize or customize the output image 116 (i.e. synthetic image generated from image generating AI or model comprising parameters stored in the memory component based on corresponding to the tokens (i.e. image embedding)).
In regards to claim 15. Daha discloses the apparatus of claim 14.
Daha further discloses
-wherein: the system comprises a tokenization component configured to tokenize the text prompt to obtain a plurality of text tokens, wherein the image embedding is generated based on the plurality of text tokens (Daha, para [0024] and [0025]; Reference at para [0024] discloses one current example of an image-generating AI engine that can be used as the AI engine in the techniques described herein is DALL-E…DALL-E is a transformer language model similar to GPT but trained to output images in response to user input. Para [0025] discloses DALL-E and other similar image generating AI engines process input by generating tokens (i.e. tokenization component). These tokens are created by encoding the textual description into a numerical representation, which is then fed into the model to generate an image (i.e. tokenizing the text prompt to obtain a plurality of text tokens, wherein the image embedding is generated based on the plurality of text tokens . The tokens in DALL-E represent specific elements or attributes of the desired image, such as shape, color, and style (i.e. generate image embedding based on the text prompt, wherein the image embedding represents visual features of the image element)).
In regards to claim 20. Daha discloses the apparatus of claim 14.
Daha further discloses
-further comprising: a user interface configured to display the synthetic image (Daha, para [0048]; Reference discloses the user device 210 will have a client application 214 with a user interface for receiving user input and commands. The client application 214 may be a purpose-specific application for generating customized images using the image generating AI engine 112).
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention.
Claims 8-13 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Shi (US 2024/0355022 A1, hereinafter referenced “Shi”).
In regards to claim 8. Shi discloses a method of training a machine learning model (Shi, Abstract and para [0005]), the method comprising:
-obtaining a training set including a text prompt and a ground-truth image, wherein the text prompt describes an image element (Shi, para [0045] and [0134]; Reference discloses at [0045] At operation 230, the image generation apparatus 120 can obtain the set 112 of images 115 and the text description, where the images 115 can be obtained from the user device 110 or database 140, and the text description can be obtained from the user 105 (i.e. wherein the text prompt describes an image element). Para [0134] discloses In various embodiments, the original image set, Xt, without cropping out the object region or masking out the background, is regarded as the ground-truth. During training, heavy augmentation, Figure US20240355022A1-20241024-P00005, can be used to obtain variations of masked images, Xs. The training data set may not include paired images of the same subject, so one (1) image may be used to train the model per subject, where N=1 in the training data set. The model can be separately trained for each category (i.e. obtaining a training set including a text prompt and a ground-truth image));
-and training, using the training set, a transformer prior model to generate an image embedding based on the text prompt, wherein the image embedding represents visual features of the image element (Shi, para [0083]-[0084]; Reference at para [0083] discloses CLIP is a multi-modal vision and language model, that can be used for image-text similarity and for zero-shot image classification. CLIP uses a ViT like transformer to get visual features and a causal language model to get the text features. Para [0084] discloses the image generation model 400 can utilize Stable Diffusion or a text-to-image transformer model, as the pretrained text-to-image model, where the model can be trained on text-image pairs).
In regards to claim 9. Shi discloses the method of claim 8.
-further comprising: generating a ground-truth image embedding based on the ground-truth image (Shi, para [0134] and [0136]; Reference at para [0134] discloses In various embodiments, the original image set, Xt, without cropping out the object region or masking out the background, is regarded as the ground-truth. Para [0136] discloses In various embodiments, the subject can be encoded into the model, where the input images can be converted into a textual subject embedding);
-generating a predicted image embedding based on the text prompt (Shi, para [0151]; Reference discloses an image generation model can be trained by obtaining a training data set including a plurality of training images and a text description, generating a new image based on an input test image and the text description using the image generation component, comparing the predicted noise and the ground truth noise, and updating parameters of the image generating component based on the comparison);
-computing a loss based on the ground-truth image embedding and the predicted image embedding (Shi, para [0149]; Reference discloses at operation 965, a loss value can be calculated for the comparison of the ground truth and the predicted noise);
-and updating parameters of the transformer prior model based on the loss (Shi, para [0149] and [0151]; Reference at para [0149] discloses at operation 965, a loss value can be calculated for the comparison of the ground truth and the predicted noise. Para [0151] discloses an image generation model can be trained by obtaining a training data set including a plurality of training images and a text description, generating a new image based on an input test image and the text description using the image generation component, comparing the predicted noise and the ground truth noise, and updating parameters of the image generating component based on the comparison).
In regards to claim 10. Shi discloses the method of claim 9.
Shi further discloses
-wherein: the loss comprises a mean squared error (MSE) loss (Shi, para [0137]; Reference illustrates the loss function formulation).
In regards to claim 11. Shi discloses the method of claim 8.
Shi further discloses
-further comprising: generating, using an image generation model, a synthetic image based on the image embedding (Shi, para [0065]; Reference discloses according to some aspects, image generation model 360 (e.g., diffusion model) generates an output image 125 including the original, identified subject from the original image(s) 115 and the modified content, where the new image 125 (i.e. synthetic image) can be generated using a diffusion model that takes a vector generated from a description by a text encoder 430 (e.g., a transformer) as input).
In regards to claim 12. Shi discloses the method of claim 11.
Shi further discloses
-wherein: the transformer prior model is trained independent of the image generation model (Shi, para [0031]; Reference discloses the original weights of the pre-trained model can be frozen (i.e., fixed) and the model can be extended with new trainable adapter layers. In various embodiments, an original transformer block contains a self-attention layer followed by a cross-attention layer that takes both the visual feature tokens and the textual embeddings as inputs for cross-attention learning. In each transformer block, a new learnable adapter layer can be added between the two frozen layers adding new layers for training on top of the previous pre-trained model that’s been frozen interpreted as the independent training of the models).
In regards to claim 13. Shi discloses the method of claim 8.
Shi further discloses
-wherein: the transformer prior model comprises parameters stored in a non-transitory computer readable medium that are optimized during the training (Shi, para [0005] and [0151]; Reference at [0005] discloses that one or more aspects of the method, apparatus, and non-transitory computer readable medium include obtaining a training data set including a training image….training the image generation model including a subject encoder and a diffusion model based on the training set, wherein the subject encoder is trained to encode an input image depicting a subject to obtain a subject embedding, and wherein the diffusion model is trained to generate an output image depicting the subject based on the subject embedding. Para [0151] discloses an image generation model can be trained by obtaining a training data set including a plurality of training images and a text description, generating a new image based on an input test image and the text description using the image generation component, comparing the predicted noise and the ground truth noise, and updating parameters of the image generating component based on the comparison (i.e. optimized parameters from training)).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 4-7 and 16-19 are rejected under 35 U.S.C. 103 as being unpatentable over Daha (US 2024/0296595 A1) in view of Shi (US 2024/0355022 A1).
In regards to claim 4. Daha discloses the method of claim 2.
Daha does not explicitly disclose but Shi teaches
-wherein generating the image embedding comprises: generating a plurality of partial image embeddings corresponding to the plurality of text tokens, respectively, wherein each of the plurality of partial image embeddings represents partial visual features; and combining the plurality of partial image embeddings to obtain the image embedding (Shi, para [0104]; Reference discloses In some cases, U-Net 700 takes additional input features to produce conditionally generated output. For example, the additional input features could include a vector representation of an input prompt (e.g., text description). The additional input features can be combined with the intermediate features 715 within the neural network at one or more layers. For example, a cross-attention module can be used to combine the additional input features and the intermediate features 715).
Daha and Shi are combinable because they are in the same field of endeavor regarding intelligent system image generation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention for the customized image generation system of Daha to include the personalized text-to-image generation features of Shi in order to provide the user with a system that allows for requesting a customized image from an image-generating artificial intelligence engine as taught by Daha, while incorporating the personalized text-to-image generation features of Shi to allow for use of techniques that provide a guidance embedding by combining a subject embedding and a text embedding to generate an output image based on the guidance embedding using a diffusion model of an image generation model to provide efficient, real-time personalized images that can retain fine-grained details, applicable to improving customized image generation systems such as those taught in Daha.
In regards to claim 5. Daha discloses the method of claim 2.
Daha does not explicitly disclose but Shi teaches
-wherein generating the image embedding comprises: obtaining a position embedding for each of the plurality of text tokens, wherein the image embedding is generated based on the position embedding (Shi, para [0107]; Reference discloses a transformer or transformer network is a type of neural network model used for natural language processing tasks. A transformer network transforms one sequence into another sequence using an encoder and a decoder. Encoder and decoder include modules that can be stacked on top of each other multiple times. The modules comprise multi-head attention and feed forward layers. The inputs and outputs (target sentences) are first embedded into an n-dimensional space. Positional encoding of the different words (i.e., give every word/part in a sequence a relative position since the sequence depends on the order of its elements) are added to the embedded representation (n-dimensional vector) of each word).
In regards to claim 6. Daha discloses the method of claim 1.
Daha does not explicitly disclose but Shi teaches
-wherein generating the synthetic image comprises: obtaining a noise input; and denoising the noise input based on the image embedding to generate the synthetic image (Shi, para [0094]; Reference discloses a diffusion model can include both a forward diffusion process 605 for adding noise to an image (or features in a latent space) and a reverse diffusion process 610 for denoising the images (or features) to obtain a denoised image. The forward diffusion process 605 can be represented as p(xt-1|xt), and the reverse diffusion process 610 can be represented as q(xt|xt-1). In some cases, the forward diffusion process 605 is used during training to generate images with successively greater noise, and a neural network is trained to perform the reverse diffusion process 610 (i.e., to successively remove the noise).).
Daha and Shi are combinable because they are in the same field of endeavor regarding intelligent system image generation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention for the customized image generation system of Daha to include the personalized text-to-image generation features of Shi in order to provide the user with a system that allows for requesting a customized image from an image-generating artificial intelligence engine as taught by Daha, while incorporating the personalized text-to-image generation features of Shi to allow for use of techniques that provide a guidance embedding by combining a subject embedding and a text embedding to generate an output image based on the guidance embedding using a diffusion model of an image generation model to provide efficient, real-time personalized images that can retain fine-grained details, applicable to improving customized image generation systems such as those taught in Daha.
In regards to claim 7. Daha discloses the method of claim 1.
Daha does not explicitly disclose but Shi teaches
-wherein: the transformer prior model is trained to generate image embeddings using a training set comprising a training text prompt and a ground-truth image embedding (Shi, para [0060] and [0142]; Reference at para [0060] discloses the text encoder can be, for example, a transformer, that receives a text prompt as input and generates a text embedding as output. Para [0142] discloses At operation 910, a plurality of images can be obtained for a data set, where each training image contains a subject, and a ground truth text description is associated with the image. Each image in the training data set can have a different subject).
Daha and Shi are combinable because they are in the same field of endeavor regarding intelligent system image generation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention for the customized image generation system of Daha to include the personalized text-to-image generation features of Shi in order to provide the user with a system that allows for requesting a customized image from an image-generating artificial intelligence engine as taught by Daha, while incorporating the personalized text-to-image generation features of Shi to allow for use of techniques that provide a guidance embedding by combining a subject embedding and a text embedding to generate an output image based on the guidance embedding using a diffusion model of an image generation model to provide efficient, real-time personalized images that can retain fine-grained details, applicable to improving customized image generation systems such as those taught in Daha.
In regards to claim 16. Daha discloses the apparatus of claim 15.
Daha does not explicitly disclose but Shi teaches
-wherein: the transformer prior model comprises a first transformer layer trained to generate a plurality of text token embeddings corresponding to the plurality of text tokens, respectively (Shi, para [0031]; Reference discloses in various embodiments, an original transformer block contains a self-attention layer followed by a cross-attention layer that takes both the visual feature tokens and the textual embeddings as inputs for cross-attention learning. In each transformer block, a new learnable adapter layer can be added between the two frozen layers).
Daha and Shi are combinable because they are in the same field of endeavor regarding intelligent system image generation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention for the customized image generation system of Daha to include the personalized text-to-image generation features of Shi in order to provide the user with a system that allows for requesting a customized image from an image-generating artificial intelligence engine as taught by Daha, while incorporating the personalized text-to-image generation features of Shi to allow for use of techniques that provide a guidance embedding by combining a subject embedding and a text embedding to generate an output image based on the guidance embedding using a diffusion model of an image generation model to provide efficient, real-time personalized images that can retain fine-grained details, applicable to improving customized image generation systems such as those taught in Daha.
In regards to claim 17. Daha in view of Shi teach the apparatus of claim 16.
Daha does not explicitly disclose but Shi teaches
-wherein: the transformer prior model comprises a second transformer layer trained to generate a plurality of intermediate embeddings corresponding to the plurality of text token embeddings, respectively (Shi, para [0031]; Reference discloses in various embodiments, an original transformer block contains a self-attention layer followed by a cross-attention layer that takes both the visual feature tokens and the textual embeddings as inputs for cross-attention learning. In each transformer block, a new learnable adapter layer can be added between the two frozen layers (i.e. adding new adaptable layers between frozen layers for the embeddings interpreted as the second transformer layer generating intermediate embeddings)).
Daha and Shi are combinable because they are in the same field of endeavor regarding intelligent system image generation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention for the customized image generation system of Daha to include the personalized text-to-image generation features of Shi in order to provide the user with a system that allows for requesting a customized image from an image-generating artificial intelligence engine as taught by Daha, while incorporating the personalized text-to-image generation features of Shi to allow for use of techniques that provide a guidance embedding by combining a subject embedding and a text embedding to generate an output image based on the guidance embedding using a diffusion model of an image generation model to provide efficient, real-time personalized images that can retain fine-grained details, applicable to improving customized image generation systems such as those taught in Daha.
In regards to claim 18. Daha in view of Shi teach the apparatus of claim 17.
Daha does not explicitly disclose but Shi teaches
-wherein: the transformer prior model comprises a third transformer layer trained to generate a plurality of partial image embeddings corresponding to the plurality of intermediate embeddings, respectively (Shi, para [0031]; Reference discloses in various embodiments, an original transformer block contains a self-attention layer followed by a cross-attention layer that takes both the visual feature tokens and the textual embeddings as inputs for cross-attention learning. In each transformer block, a new learnable adapter layer can be added between the two frozen layers (i.e. adding new adaptable layers between frozen layers for the embeddings interpreted as the third transformer layer generating partial image embeddings corresponding to intermediate embeddings)).
Daha and Shi are combinable because they are in the same field of endeavor regarding intelligent system image generation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention for the customized image generation system of Daha to include the personalized text-to-image generation features of Shi in order to provide the user with a system that allows for requesting a customized image from an image-generating artificial intelligence engine as taught by Daha, while incorporating the personalized text-to-image generation features of Shi to allow for use of techniques that provide a guidance embedding by combining a subject embedding and a text embedding to generate an output image based on the guidance embedding using a diffusion model of an image generation model to provide efficient, real-time personalized images that can retain fine-grained details, applicable to improving customized image generation systems such as those taught in Daha.
In regards to claim 19. Daha discloses the apparatus of claim 14.
Daha does not explicitly disclose but Shi teaches
-wherein: the image generation model includes a diffusion model (Shi, para [0037]; Reference discloses an image generation apparatus 120 can include a computer implemented network comprising a user interface, and a machine learning model, which can include a diffusion model).
Daha and Shi are combinable because they are in the same field of endeavor regarding intelligent system image generation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention for the customized image generation system of Daha to include the personalized text-to-image generation features of Shi in order to provide the user with a system that allows for requesting a customized image from an image-generating artificial intelligence engine as taught by Daha, while incorporating the personalized text-to-image generation features of Shi to allow for use of techniques that provide a guidance embedding by combining a subject embedding and a text embedding to generate an output image based on the guidance embedding using a diffusion model of an image generation model to provide efficient, real-time personalized images that can retain fine-grained details, applicable to improving customized image generation systems such as those taught in Daha.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: See the Notice of References Cited (PTO-892)
Any inquiry concerning this communication or earlier communications from the examiner should be directed to TERRELL M ROBINSON whose telephone number is (571)270-3526. The examiner can normally be reached 8am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, KENT CHANG can be reached at 571-272-7667. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/TERRELL M ROBINSON/Primary Examiner, Art Unit 2614