Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention
was made.
Claim(s) 1-2 and 11-12 are rejected under 35 U.S.C. 103 as being unpatentable over Povalyaev (US-20250111655-A1) in view of Palomera ( US-20260134227-A1) .
Regarding claim 1, Povalyaev teaches A system, comprising: a processor programmed to execute a text2image model, the processor to: access input text that describes an image to be generated (Para.79 and 84: teaches receiving natural language text prompts/ descriptions corresponding to images to be generated); generate a plurality of text tokens based on the input text (Para.79: the tokenizer generates multiple text tokens from the input prompt text, corresponding to the claimed plurality of text tokens); perform cross-attention between the transformed plurality of text tokens and the transformed plurality of image tokens (Para 79: teaches cross attention between text-derived embeddings/conditional embeddings and image generation network components); and generate an output image based on the cross-attention (Para. 78-80: teaches generating the final output image using the conditioned transformer/diffusion architecture employing the cross attention conditioning).
Povalyaev fails to teach transform, in a first transformer layer of the text2image model, the plurality of text tokens based on self-attention in which an importance of a text token is weighted relative to other ones of the plurality of text tokens; transform, in a second transformer layer of the text2image model, a plurality of image tokens based on self-attention in which an importance of an image token is weighted relative to other ones of the plurality of image tokens.
Palomera teaches transform, in a first transformer layer of the text2image model, the plurality of text tokens based on self-attention in which an importance of a text token is weighted relative to other ones of the plurality of text tokens (Para 35 and 39: teaches transformer layers using self-attention mechanisms to process token embeddings, where token relationships and importance are weighted relative to other tokens) ; transform, in a second transformer layer of the text2image model, a plurality of image tokens based on self-attention in which an importance of an image token is weighted relative to other ones of the plurality of image tokens (Para 35: teaches multimodal transformer models operating on image content using transformer attention mechanisms, including self-attention over token embeddings corresponding to image information. It would have been obvious to incorporate the transformer self-attention mechanisms of Palomera into the text to image generation framework of Povalyaev to improve contextual token weighting and multimodal attention processing, thereby improving generation accuracy.).
Regarding claim 2, Povalyaev in view of Palomera teaches The system of claim 1, wherein the text2image model comprises a diffusion model that (Povalyaev, Para 78: teaches a diffusion based text to image generation model using diffusion layers, when executed, is to initialize an image with random noise (Povalyaev, Para 74 and 78: teaches beginning from noisy/random latent representations corresponding to initialized random noise images), perform a plurality of diffusion iterations to denoise the image (Povalyaev, Para 78, and 80-81: teaches iterative denoising operations performed across multiple diffusion steps to reconstruct/generate the final image.), and wherein each diffusion iteration executes a plurality of stacks of transformation layers, each transformation stack comprising the first transformer layer, the second transformer layer, and the cross-attention ( Palomera, Para 35 036, 39 and 41: teaches transformer architectures composed of multiple transformer processing arrangements/ layers operating in series or parallel.).
Regarding claim 11, Povalyaev teaches A method, comprising: accessing input text that describes an image to be generated (Para.79 and 84: teaches receiving natural language text prompts/ descriptions corresponding to images to be generated); generate a plurality of text tokens based on the input text (Para.79: the tokenizer generates multiple text tokens from the input prompt text, corresponding to the claimed plurality of text tokens); perform cross-attention between the transformed plurality of text tokens and the transformed plurality of image tokens (Para 79: teaches cross attention between text-derived embeddings/conditional embeddings and image generation network components); and generate an output image based on the cross-attention (Para. 78-80: teaches generating the final output image using the conditioned transformer/diffusion architecture employing the cross attention conditioning).
Povalyaev fails to teach transform, in a first transformer layer of the text2image model, the plurality of text tokens based on self-attention in which an importance of a text token is weighted relative to other ones of the plurality of text tokens; transform, in a second transformer layer of the text2image model, a plurality of image tokens based on self-attention in which an importance of an image token is weighted relative to other ones of the plurality of image tokens.
Palomera teaches transform, in a first transformer layer of the text2image model, the plurality of text tokens based on self-attention in which an importance of a text token is weighted relative to other ones of the plurality of text tokens (Para 35 and 39: teaches transformer layers using self-attention mechanisms to process token embeddings, where token relationships and importance are weighted relative to other tokens) ; transform, in a second transformer layer of the text2image model, a plurality of image tokens based on self-attention in which an importance of an image token is weighted relative to other ones of the plurality of image tokens (Para 35: teaches multimodal transformer models operating on image content using transformer attention mechanisms, including self-attention over token embeddings corresponding to image information. It would have been obvious to incorporate the transformer self-attention mechanisms of Palomera into the text to image generation framework of Povalyaev to improve contextual token weighting and multimodal attention processing, thereby improving generation accuracy.).
Regarding claim 12, it falls under the same rejection as claim 2 it is similar in scope and dependent upon same references.
Claim(s) 3-6 and 13-16 are rejected under 35 U.S.C. 103 as being unpatentable over Povalyaev (US-20250111655-A1) in view of Palomera ( US-20260134227-A1) in further view of Kirazci (US-20140201229-A1).
Regarding claim 3, Povalyaev in view of Palomera teaches The system of claim 1, wherein to generate the plurality of text tokens (Povalyaev , Para. 79: teaches generating a plurality of text tokens from an input prompt using tokenizer), the processor is further programmed to: but fails to teach identify at least a first portion of the input and a second portion of the input; tokenize, based on an adaptive character-level tokenizer, the first portion via a first tokenization mechanism; and tokenize, based on an adaptive character-level tokenizer, the second portion via a second tokenization mechanism different than the first tokenization mechanism.
Kirazci teaches identify at least a first portion of the input and a second portion of the input; tokenize, based on an adaptive character-level tokenizer, the first portion via a first tokenization mechanism; and tokenize, based on an adaptive character-level tokenizer, the second portion via a second tokenization mechanism different than the first tokenization mechanism (Para 21: identifying multiple different portions/substrings of an input term corresponding to identifying first and second portions of the input. It would have been obvious to incorporate the adaptive character level tokenization technique of Kirazci into the text to image system of Povalyaev in view of Palomera to improve token generation and processing text portions prior to transformer based image generation) .
Regarding claim 4, Povalyaev in view of Palomera and in further view of Kirazci teaches The system of claim 3, wherein to tokenize the first portion, the processor is further programmed to: split the first portion into one or more words and tokenize each word from among the one or more words (Kirazci, Para 21: identifying multiple different portions/substrings of an input term corresponding to identifying first and second portions of the input.).
Regarding claim 5, Povalyaev in view of Palomera and in further view of Kirazci teaches The system of claim 3, wherein to tokenize the second portion, the processor is further programmed to: split the second portion into one or more characters and tokenize each character from among the one or more characters (Kirazci, Para 21: identifying multiple different portions/substrings of an input term corresponding to identifying first and second portions of the input.).
Regarding claim 6, Povalyaev in view of Palomera and in further view of Kirazci teaches The system of claim 3, wherein the processor is further programmed to: use the first portion as a description of the image to be generated and the second portion as literal characters to be included in the image (Povalyaev, Para84-86 and 102-103: teaches using portions of input text/ prompts as image descriptions and visual characteristic descriptions for generating images with the text to image model. Palomera: para 35, 48-39 and 51: teaches the embeddings and tokens from text prompts and literal text content. Kirazci: teaches character level tokens separately from boarder descriptive text. ) .
Regarding claim 13, it falls under the same rejection as claim 3 it is similar in scope and dependent upon same references.
Regarding claim 14, it falls under the same rejection as claim 4 it is similar in scope and dependent upon same references.
Regarding claim 15, it falls under the same rejection as claim 5 it is similar in scope and dependent upon same references.
Regarding claim 16, it falls under the same rejection as claim 6 it is similar in scope and dependent upon same references.
Claim(s) 7 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Povalyaev (US-20250111655-A1) in view of Palomera ( US-20260134227-A1) in further view of Yu (US-20250328752-A1).
Regarding claim 7, Povalyaev in view of Palomera teaches The system of claim 1, wherein a transformer block comprises a transformer block architecture but fails to teach that modulates the magnitude of the residual blocks and normalizes it back to have a predefined magnitude and instead of a normal summation between backbone and residual block activations.
Yu teaches that modulates the magnitude of the residual blocks and normalizes it back to have a predefined magnitude and instead of a normal summation between backbone and residual block activations (Para.101: teaches residual values being processed through normalization blocks within transformer layers. Yu further teaches normalization operations performed on outputs associated with residual values, which would normalize the activations to a normalized predefined magnitude range across dimensions. It would have been obvious to modify the transformer architecture of Povalyaev in view of Palomera with the normalized residual processing techniques of Yu in order to improve stability and consistency of transformer layer activations during image generation).
Regarding claim 17, it falls under the same rejection as claim 7 it is similar in scope and dependent upon same references.
Claim(s) 8-10 and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Povalyaev (US-20250111655-A1) in view of Palomera ( US-20260134227-A1) in further view of Yu (US-20250328752-A1) and SAKHINANA (US-20230045690-A1).
Regarding claim 8, Povalyaev in view of Palomera and Yu teaches The system of claim 7, wherein the processor is further programmed to: during a training phase of the text2image model, perform a gated summation between back and residual block activations to control a size of the activation norm of the backbone.
SAKHINANA teaches perform a gated summation between back and residual block activations ( Para.55: teaches performing a gated summation involving residual information flows/activations. It would have been obvious to modify the system of Povalyaev in view of Palomera and Yu to include the gated residual summation techniques of SAKHINANA to improve stability and control activation magnitudes during training of the text to image transformer network.)
Regarding claim 9, Povalyaev in view of Palomera, Yu, and SAKHINANA teaches The system of claim 8, wherein the gated summation comprises a weighted summation controlled by a gate value (SAKHINANA. Para 55: teaches a weighted summation operation controlled by a gate value).
Regarding claim 10, Povalyaev in view of Palomera, Yu, and SAKHINANA teaches The system of claim 9, wherein the gate value is a function of diffusion time ( Povalyaev, Para. 74,78,80 and 99-101: teaches a diffusion model are conditioned on diffusion time/timestep information. SAKHINANA. Para 55: teaches the gate summation. It would have been obvious to configure the gate value of the gated summation to vary as a function of the timestep/diffusion time taught by so that residual contribution and activation flow may be adaptively controlled at different stages of the diffusion process).
Regarding claim 18, it falls under the same rejection as claim 8 it is similar in scope and dependent upon same references.
Regarding claim 19, it falls under the same rejection as claim 9 it is similar in scope and dependent upon same references.
Regarding claim 20, it falls under the same rejection as claim 10 it is similar in scope and dependent upon same reference.
Claim 22 is rejected under 35 U.S.C. 103 as being unpatentable over Povalyaev (US-20250111655-A1) in view of Reynolds ( US-20260064721-A1) .
Regarding claim 22, Povalyaev teaches A system, comprising: a processor programmed to: access input text to encode for a text to image model (Para.49, 69 and 102: teaches accessing natural language prompts and encoding the prompts into tokens and embeddings for use by a text to image model) ; identify at least a first portion of the input text and at least a second portion of the input text (Para 69 and 79: teaches separating prompt text into multiple tokens and token attributes corresponding to different portion of the input text); tokenize the first portion based on word-level encoding or N-gram encoding (Para.69: teaches word level tokenization).
Povalyaev fails to teach determine that different portions of the input text should be encoded differently and tokenize the second portion based on character-level encoding.
Reynolds teaches to determine that different portions of the input text should be encoded differently and tokenize the second portion based on character-level encoding (para 11 and 40: teaches determining different tokenization granularities for different text portions, including words and individual characters depending on the required granularity. It would have been obvious to a person of skill in the art to incorporate the tokenization granularity techniques of Reynolds into the system of Povalyaev in order to improve flexibility and precision in processing different portions of textual prompts.).
Allowable Subject Matter
Claim 21 is allowed.
The following is a statement of reasons for the indication of allowable subject matter: None of the prior art teaches such limitations “set the gate value based on modulation of the modulation value and add a weighting value to increase the gate value; applying an activation function to the gate value; adjust the gate value based on an epsilon hyperparameter to bring the gate value within a range defined by the epsilon hyperparameter.”.
Response to Arguments
Applicant's arguments filed 09/03/2026, with respect to the rejection(s) of 1-2 and 11-12 are rejected under 35 U.S.C. 103 have been fully considered and are not persuasive. Therefore, the rejection has not been withdrawn. Applicant argues that Palomera fails to teach or suggest (i) transforming the plurality of text tokens in a first transformer layer using self-attention and (ii) transforming the plurality of image tokens in a second transformer layer using self-attention. Applicant further argues that Palomera concerns text-generation models and does not disclose the claimed separate processing of text tokens and image tokens before cross-attention. However, Applicant’s arguments are not persuasive. The rejection is based on the combined teachings of Povalyaev and Palomera, rather than on Palomera alone. Povalyaev teaches a text-to-image generation architecture in which an input prompt is tokenized, text-derived conditional embeddings are supplied through cross-attention to image-generation components, and an output image is generated using the conditioned diffusion architecture (Povalyaev, para 78–80). Palomera expressly teaches that its language models may be multimodal models operating on text and other content, including images (Palomera, para 35), and that its transformer-based architectures tokenize input, determine token embeddings, and incorporate self-attention and/or cross-attention mechanisms (Palomera, para 36). Palomera further teaches different configurations of self-attention and cross-attention followed by one or more neural-network layers (Palomera, para 37). Accordingly, a person of ordinary skill in the art would have found it obvious to incorporate Palomera’s self-attention processing into Povalyaev’s text-to-image architecture so that the respective text-token and image-token representations are contextually weighted before Povalyaev’s cross-attention operation, thereby improving contextual representation and the accuracy of the resulting text-conditioned image generation. Such an arrangement would have amounted to the predictable use of Palomera’s known transformer-attention techniques in Povalyaev’s compatible transformer/diffusion architecture.
Applicant's arguments filed 09/03/2026, with respect to the rejection(s) of claims 22 is rejected under 35 U.S.C. 103 have been fully considered and is not persuasive. Therefore, the rejection has not been withdrawn. Applicant argues that Reynolds merely identifies alternative token types and does not teach determining that different portions of the same input text should be encoded differently. However, the argument is not persuasive because the rejection is based on the combined teachings of Povalyaev and Reynolds. Povalyaev teaches parsing prompt text, identifying information associated with target visual characteristics, mapping that information to token attributes or custom tokens, and converting prompt words into tokens (para 79 and 102). Reynolds teaches selecting between word-level and individual-character token granularities according to the required granularity and separately treating special characters and other textual elements (para 11–12 and 40–41). Thus, the combined teachings support applying different tokenization granularities to respective portions of Povalyaev’s prompt.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to LATRELL ANTHONY CREARY whose telephone number is (703)756-1219. The examiner can normally be reached Mon - Fri 7:30am - 4:30pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Xiao WU can be reached at (571) 272-7761. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/LATRELL ANTHONY CREARY/Examiner, Art Unit 2613
/XIAO M WU/Supervisory Patent Examiner, Art Unit 2613