Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-2, 5, and 14-15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Balaji et al. (U.S. PGPUB 20240161250) in view of Wang et al. (U.S. PGPUB 20250095250).
With respect to claim 1, Balaji et al. disclose a method comprising: obtaining an input prompt indicating a visual text (paragraph 47, the image generating application 146 receives an input text 302 and, optionally, an input image 304); and
performing, using an image generation model, a first cross-attention operation based on the input prompt (paragraph 59, The attention map 520 cross attends between the text and image, and the attention map 520 is a matrix computed from queries 510 that are flattened image features and keys 512 and values 514 that are flattened text features). However, Balaji et al. do not expressly disclose performing, using the image generation model, a second cross-attention operation based on the visual text; and generating, using the image generation model, a synthetic image depicting the visual text based on the first cross-attention operation and the second cross-attention operation.
Wang et al., who also deal with generating an image, disclose a method for performing, using the image generation model, a second cross-attention operation based on the visual text (paragraph 109, In step S240, a second cross-attention feature of a second image feature of the reference image and the text feature is obtained); and
generating, using the image generation model, a synthetic image depicting the visual text based on the first cross-attention operation and the second cross-attention operation (paragraph 113, In step S250, the first cross-attention feature is edited based on the second cross-attention feature to obtain a third cross-attention feature. It may be understood that the third cross-attention feature is an edited first cross-attention feature, paragraph 127, In step S260, a result image feature of the time step is generated based on the third cross-attention feature and the text feature).
Balaji et al. and Wang et al. are in the same field of endeavor, namely computer graphics.
Before the effective filing date of the claimed invention, it would have been obvious to apply the method of performing, using the image generation model, a second cross-attention operation based on the visual text; and generating, using the image generation model, a synthetic image depicting the visual text based on the first cross-attention operation and the second cross-attention operation, as taught by Wang et al., to the Balaji et al. system, because information in the reference image can be continuously introduced into the image generation process of the diffusion model. Therefore, the information in the reference image can be effectively used to guide image generation of the diffusion model, thereby ensuring that a generated target image can be consistent with the reference image in terms of content and has a specified style (paragraph 153 of Wang et al.).
With respect to claim 2, Balaji et al. as modified by Wang et al. disclose the method of claim 1, further comprising: generating, using a word encoder, prompt features representing the input prompt (Balaji et al.: paragraph 49, the input text can be represented by a text embedding, extracted from a pretrained model such as CLIP or T5 text encoders); and generating intermediate image features for the synthetic image based on the prompt features, wherein the first cross-attention operation is based on the prompt features and the intermediate image features (Balaji et al.: paragraph 59, The attention map 520 cross attends between the text and image, and the attention map 520 is a matrix computed from queries 510 that are flattened image features and keys 512 and values 514 that are flattened text features. The vectors 509 are combined into a matrix 522 that is added to the attention map 520 to generate an updated attention map 524. The image generating application 146 then computes a softmax 526 of the updated attention map 524 and combines the result with a text embedding 514 to generate an embedding that is input into a next layer of an expert denoiser, such as one of the expert denoisers 150).
With respect to claim 5, Balaji et al. as modified by Wang et al. disclose the method of claim 1, wherein generating the synthetic image comprises: combining a result of the first cross-attention operation and a result of the second cross-attention operation to obtain combined image features, wherein the synthetic image is generated based on the combined image features (Wang et al.: paragraph 114, Accordingly, the first cross-attention feature, the second cross-attention feature, and the third cross-attention feature may each be divided into two sub-features, where one sub-feature corresponds to the content description text, and the other sub-feature corresponds to the style description text. Specifically, the first cross-attention feature includes a first content sub-feature corresponding to the content description text and a first style sub-feature corresponding to the style description text. The second cross-attention feature includes a second content sub-feature corresponding to the content description text and a second style sub-feature corresponding to the style description text. The third cross-attention feature includes a third content sub-feature corresponding to the content description text and a third style sub-feature corresponding to the style description text).
With respect to claim 14, Balaji et al. as modified by Wang et al. disclose a system (Balaji et al.: paragraph 38, FIG. 2 is a more detailed illustration of the computing device 140 of FIG. 1) comprising: a memory component; a processing device coupled to the memory component (Balaji et al.: paragraph 39, the computing device 140 includes, without limitation, the processor 142 and the memory 144 coupled to a parallel processing subsystem 212 via a memory bridge 205 and a communication path 213), the processing device configured to perform operations of claim 1; see rationale for rejection of claim 1.
With respect to claim 15, Balaji et al. as modified by Wang et al. disclose the system of claim 14, the system further comprising: a word encoder configured to encode the input prompt to obtain prompt features (Balaji et al.: paragraph 49, the input text can be represented by a text embedding, extracted from a pretrained model such as CLIP or T5 text encoders).
Claim(s) 3 and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Balaji et al. (U.S. PGPUB 20240161250) in view of Wang et al. (U.S. PGPUB 20250095250) and further in view of Gong et al. (U.S. PGPUB 20260112185).
With respect to claim 3, Balaji et al. as modified by Wang et al. disclose the method of claim 2. However, Balaji et al. as modified by Wang et al. do not expressly disclose performing optical character recognition (OCR) on the visual text, wherein the prompt features are based on the OCR.
Gong et al., who also deal with generating an image, disclose a method for performing optical character recognition (OCR) on the visual text, wherein the prompt features are based on the OCR (paragraph 56, The image based correction 310 is based on optical character recognition text (e.g., the visual text 308) generated from optical character recognition performed by the OCR model 306 based on the digital image 114).
Balaji et al., Wang et al., and Gong et al. are in the same field of endeavor, namely computer graphics.
Before the effective filing date of the claimed invention, it would have been obvious to apply the method of performing optical character recognition (OCR) on the visual text, wherein the prompt features are based on the OCR, as taught by Gong et al., to the Balaji et al. as modified by Wang et al. system, because this would improve the zero-shot performance when out-of-domain attribute values are reported from the text decoder 202 (paragraph 56 of Gong et al.).
With respect to claim 17, Balaji et al. as modified by Wang et al. and Gong et al. disclose the system of claim 14, the system further comprising: an OCR encoder configured to encode the input prompt to obtain OCR embeddings (Gong et al.: (paragraph 56, The image based correction 310 is based on optical character recognition text (e.g., the visual text 308) generated from optical character recognition performed by the OCR model 306 based on the digital image 114); see rationale for rejection of claim 3.
Claim(s) 4, 8, 10-11, and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Balaji et al. (U.S. PGPUB 20240161250) in view of Wang et al. (U.S. PGPUB 20250095250) and further in view of Xue et al. (U.S. PGPUB 20240119743).
With respect to claim 4, Balaji et al. as modified by Wang et al. disclose the method of claim 1. However, Balaji et al. as modified by Wang et al. do not expressly disclose generating, using a glyph encoder, text features representing the visual text; and generating intermediate image features for the synthetic image based on the text features, wherein the second cross-attention operation is based on the text features and the intermediate image features.
Xue et al., who also deal with generating an image, disclose a method for generating, using a glyph encoder, text features representing the visual text (paragraph 50, As a result, the character-aware text encoder 205 may extract the text instance embeddings te={t∈.sub.0, te.sub.1, . . . , te.sub.n-1} from the annotated text instances t={t.sub.0, t.sub.1, . . . , t.sub.n-1}. The character-aware text encoder 205 may encode the instance level textual information and neglect the relations between each pair of text instances. It may help to learn better visual text representations); and generating intermediate image features for the synthetic image based on the text features, wherein the second cross-attention operation is based on the text features and the intermediate image features (paragraph 52, Given sample images 205 and 305 in the first column, and column 2 show the attention maps 310 and 315 (from the attention layer 230 in the image encoder 115) that may be obtained from models with the character-aware text encoder 205).
Balaji et al., Wang et al., and Xue et al. are in the same field of endeavor, namely computer graphics.
Before the effective filing date of the claimed invention, it would have been obvious to apply the method of generating, using a glyph encoder, text features representing the visual text; and generating intermediate image features for the synthetic image based on the text features, wherein the second cross-attention operation is based on the text features and the intermediate image features, as taught by Xue et al., to the Balaji et al. as modified by Wang et al. system, because the character-aware text encoder 205 may attend better to text regions, leading to better learning of the scene text visual representations of the network backbone 225 (paragraph 52 of Xue et al.).
With respect to claim 8, Balaji et al. disclose a non-transitory computer readable medium storing code for image processing (paragraph 37, an image generating application 146 is stored in a memory 144, and executes on a processor 142, of the computing device 140, paragraph 112, any combination of one or more computer readable medium(s)), the code comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising: obtaining an input prompt indicating a visual text (paragraph 47, the image generating application 146 receives an input text 302 and, optionally, an input image 304);
encoding, using a word encoder of an image generation model, the input prompt to obtain prompt features (paragraph 49, the input text can be represented by a text embedding, extracted from a pretrained model such as CLIP or T5 text encoders). However, Balaji et al. do not expressly disclose encoding, using a glyph encoder of the image generation model, the input prompt to obtain text features; and generating, using the image generation model, a synthetic image depicting the visual text based on the prompt features and the text features.
Xue et al., who also deal with generating an image, disclose a method for encoding, using a glyph encoder of the image generation model, the input prompt to obtain text features (paragraph 50, As a result, the character-aware text encoder 205 may extract the text instance embeddings te={t∈.sub.0, te.sub.1, . . . , te.sub.n-1} from the annotated text instances t={t.sub.0, t.sub.1, . . . , t.sub.n-1}. The character-aware text encoder 205 may encode the instance level textual information and neglect the relations between each pair of text instances. It may help to learn better visual text representations).
Balaji et al. and Xue et al. are in the same field of endeavor, namely computer graphics.
Before the effective filing date of the claimed invention, it would have been obvious to apply the method of encoding, using a glyph encoder of the image generation model, the input prompt to obtain text features, as taught by Xue et al., to the Balaji et al. system, because the character-aware text encoder 205 may attend better to text regions, leading to better learning of the scene text visual representations of the network backbone 225 (paragraph 52 of Xue et al.).
Wang et al., who also deal with generating an image, disclose a method for generating, using the image generation model, a synthetic image depicting the visual text based on the prompt features and the text features (paragraph 89, the first cross-attention feature Me of the first image feature I.sub.t and the text feature Text is calculated, paragraph 110, The second image feature F of the reference image may be extracted by using the image encoder (for example, a CLIP image encoder), paragraph 113, In step S250, the first cross-attention feature is edited based on the second cross-attention feature to obtain a third cross-attention feature. It may be understood that the third cross-attention feature is an edited first cross-attention feature, paragraph 127, In step S260, a result image feature of the time step is generated based on the third cross-attention feature and the text feature).
Balaji et al., Xue et al., and Wang et al. are in the same field of endeavor, namely computer graphics.
Before the effective filing date of the claimed invention, it would have been obvious to apply the method of generating, using the image generation model, a synthetic image depicting the visual text based on the prompt features and the text features, as taught by Wang et al., to the Balaji et al. as modified by Xue et al. system, because information in the reference image can be continuously introduced into the image generation process of the diffusion model. Therefore, the information in the reference image can be effectively used to guide image generation of the diffusion model, thereby ensuring that a generated target image can be consistent with the reference image in terms of content and has a specified style (paragraph 153 of Wang et al.).
With respect to claim 10, Balaji et al. as modified by Xue et al. and Wang et al. disclose the non-transitory computer readable medium of claim 8, the code further comprising instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising: generating intermediate image features for the synthetic image based on the prompt features; and performing a cross-attention operation is based on the prompt features and the intermediate image features to obtain the synthetic image (Balaji et al.: paragraph 59, The attention map 520 cross attends between the text and image, and the attention map 520 is a matrix computed from queries 510 that are flattened image features and keys 512 and values 514 that are flattened text features. The vectors 509 are combined into a matrix 522 that is added to the attention map 520 to generate an updated attention map 524. The image generating application 146 then computes a softmax 526 of the updated attention map 524 and combines the result with a text embedding 514 to generate an embedding that is input into a next layer of an expert denoiser, such as one of the expert denoisers 150).
With respect to claim 11, Balaji et al. as modified by Xue et al. and Wang et al. disclose the non-transitory computer readable medium of claim 8, the code further comprising instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising: generating intermediate image features for the synthetic image based on the text features (Xue et al.: paragraph 50, As a result, the character-aware text encoder 205 may extract the text instance embeddings te={t∈.sub.0, te.sub.1, . . . , te.sub.n-1} from the annotated text instances t={t.sub.0, t.sub.1, . . . , t.sub.n-1}. The character-aware text encoder 205 may encode the instance level textual information and neglect the relations between each pair of text instances. It may help to learn better visual text representations); and performing a cross-attention operation is based on the text features and the intermediate image features to obtain the synthetic image (Xue et al.: paragraph 52, Given sample images 205 and 305 in the first column, and column 2 show the attention maps 310 and 315 (from the attention layer 230 in the image encoder 115) that may be obtained from models with the character-aware text encoder 205).
With respect to claim 16, Balaji et al. as modified by Wang et al. and Xue et al. disclose the system of claim 14, the system further comprising: a glyph encoder configured to encode the input prompt to obtain text features (Xue et al.: paragraph 50, As a result, the character-aware text encoder 205 may extract the text instance embeddings te={t∈.sub.0, te.sub.1, . . . , te.sub.n-1} from the annotated text instances t={t.sub.0, t.sub.1, . . . , t.sub.n-1}. The character-aware text encoder 205 may encode the instance level textual information and neglect the relations between each pair of text instances. It may help to learn better visual text representations); see rationale for rejection of claim 4.
Claim(s) 6 and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Balaji et al. (U.S. PGPUB 20240161250) in view of Wang et al. (U.S. PGPUB 20250095250) and further in view of Kusumi et al. (U.S. PGPUB 20260004404).
With respect to claim 6, Balaji et al. as modified by Wang et al. disclose the method of claim 1. However, Balaji et al. as modified by Wang et al. do not expressly disclose upscaling, using an additional image generation model, the synthetic image to obtain a high resolution image.
Kusumi et al., who also deal with generating an image, disclose a method for upscaling, using an additional image generation model, the synthetic image to obtain a high resolution image (paragraph 34, a low-resolution JPEG image is upscaled using a machine learning model, to thereby perform image processing for generating a high-resolution image with higher precision).
Balaji et al., Wang et al., and Kusumi et al. are in the same field of endeavor, namely computer graphics.
Before the effective filing date of the claimed invention, it would have been obvious to apply the method of upscaling, using an additional image generation model, the synthetic image to obtain a high resolution image, as taught by Kusumi et al., to the Balaji et al. as modified by Wang et al. system, because this leverage the power of artificial intelligence to improve image quality.
With respect to claim 18, Balaji et al. as modified by Wang et al. and Kusumi et al. disclose the system of claim 14, the system further comprising: an additional image generation model configured to upscale the synthetic image to obtain a high resolution synthetic image (Kusumi et al.: paragraph 34, a low-resolution JPEG image is upscaled using a machine learning model, to thereby perform image processing for generating a high-resolution image with higher precision); see rationale for rejection of claim 6.
Claim(s) 7 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Balaji et al. (U.S. PGPUB 20240161250) in view of Wang et al. (U.S. PGPUB 20250095250) and further in view of Sun (CN 120673428).
With respect to claim 7, Balaji et al. as modified by Wang et al. disclose the method of claim 1. However, Balaji et al. as modified by Wang et al. do not expressly disclose the image generation model is trained using an OCR loss.
Sun, who also deals with generating an image, discloses a method wherein the image generation model is trained using an OCR loss (paragraph 151, Based on the calculated loss value, the OCR model can be trained).
Balaji et al., Wang et al., and Sun are in the same field of endeavor, namely computer graphics.
Before the effective filing date of the claimed invention, it would have been obvious to apply the method wherein the image generation model is trained using an OCR loss, as taught by Sun, to the Balaji et al. as modified by Wang et al. system, because the recognition performance of the OCR model can be optimized by continuously adjusting the parameters of the OCR model, so as to improve the accuracy and recall of the OCR model (paragraph 151 of Sun).
With respect to claim 19, Balaji et al. as modified by Wang et al. and Sun disclose the system of claim 14, wherein: the image generation model is trained using an OCR loss (Sun: paragraph 151, Based on the calculated loss value, the OCR model can be trained); see rationale for rejection of claim 7.
Claim(s) 9 is/are rejected under 35 U.S.C. 103 as being unpatentable over Balaji et al. (U.S. PGPUB 20240161250) in view of Wang et al. (U.S. PGPUB 20250095250), Xue et al. (U.S. PGPUB 20240119743), and further in view of Gong et al. (U.S. PGPUB 20260112185).
With respect to claim 9, Balaji et al. as modified by Xue et al., Wang et al., and Gong et al. disclose the non-transitory computer readable medium of claim 8, the code further comprising instructions that, when executed by the at least one processor, cause the at least one processor to perform operations of claim 3; see rationale for rejection of claim 3.
Claim(s) 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Balaji et al. (U.S. PGPUB 20240161250) in view of Wang et al. (U.S. PGPUB 20250095250), Xue et al. (U.S. PGPUB 20240119743), and further in view of Kusumi et al. (U.S. PGPUB 20260004404).
With respect to claim 12, Balaji et al. as modified by Xue et al. and Wang et al. disclose the non-transitory computer readable medium of claim 8, the code further comprising instructions that, when executed by the at least one processor, cause the at least one processor to perform operations of claim 6; see rationale for rejection of claim 6.
Claim(s) 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Balaji et al. (U.S. PGPUB 20240161250) in view of Wang et al. (U.S. PGPUB 20250095250), Xue et al. (U.S. PGPUB 20240119743), and further in view of Sun (CN 120673428).
With respect to claim 13, Balaji et al. as modified by Xue et al., Wang et al., and Sun disclose the non-transitory computer readable medium of claim 8, wherein: the image generation model is trained using an OCR loss (Sun: paragraph 151, Based on the calculated loss value, the OCR model can be trained); see rationale for rejection of claim 7.
Claim(s) 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Balaji et al. (U.S. PGPUB 20240161250) in view of Wang et al. (U.S. PGPUB 20250095250) and further in view of Welsh et al. (U.S. PGPUB 20250022256).
With respect to claim 20, Balaji et al. as modified by Wang et al. disclose the system of claim 14. However, Balaji et al. as modified by Wang et al. do not expressly disclose the image generation model comprises a guided latent diffusion model.
Welsh et al., who also deal with generating an image, disclose a method wherein the image generation model comprises a guided latent diffusion model (paragraph 25, A guided latent diffusion model may be trained for use in a synthetic data generation system).
Balaji et al., Wang et al., and Welsh et al. are in the same field of endeavor, namely computer graphics.
Before the effective filing date of the claimed invention, it would have been obvious to apply the method wherein the image generation model comprises a guided latent diffusion model, as taught by Welsh et al., to the Balaji et al. as modified by Wang et al. system, because this would generate high-quality synthetic images based on a set of inputs while preserving the semantics of the original image (paragraph 25 of Welsh et al.).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ANDREW GUS YANG whose telephone number is (571)272-5514. The examiner can normally be reached M-F 9 AM - 5:30 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kent Chang can be reached at (571)272-7667. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ANDREW G YANG/Primary Examiner, Art Unit 2614
9/14/26