DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of Claims
Claims 1, 3, 12 and 16 are amended.
Claims 2, 4-11, 13-15 and 17-20 have been previously presented.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-2, 5-8 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Ma et al. (hereinafter Ma, “Text Style Transfer With Decorative Elements”, 2021 IEEE MIPR, pg. 330-336) in view of Xie et al (hereinafter Xie, “SmartBrush: Text and Shape Guided Object Inpainting with Diffusion Model”, pg. 1-10), and in further view of Guy et al.(hereinafter “Guy”, 2022/0413688).
Regarding claim 1, Ma teaches a method (pg. 331 sec. III 1st para. lines 1-4) comprising: obtaining, via a user interface, an input text including a plurality of characters (pg. 335 sec. V 1st para. lines 5-8); obtaining, via a user interface, a text effect prompt that describes a text effect for the input text (pg. 335 sec. V 1st para. lines 5-8); and generating, by an image generation (Figs. 2 & 7), an output image depicting the input text with the text effect described by the text effect prompt (pg. 331 ‘Text stye transfer’ 1st para. lines 1-23 and Fig. 7), wherein the output image depicts each of the plurality of characters from the input text (Fig. 7). However, Ma fails to teach generating, by an image generation model, wherein the text effect prompt comprises text different from the input text. Xie teaches generating, by an image generation model (pg. 5, sec. 4.4 2nd para. lines 1-6). Therefore, it would be obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the text input styles of Ma with the image generation models of Xie to enable a more controlled application of the text effect described by the text effect prompt in the generated output image. However, Ma and Xie fail to teach wherein the text effect prompt comprises text different from the input text,. Guy teaches wherein the text effect prompt (0137 lines 4-10) comprises text different from the input text (0101 lines 5-20, in which text and styles different from the input text is provided). Therefore, it would be obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the text input styles of Ma and image generation models of Xie with the different text styles of Guy because this modification would improve the visual appearance of user input text strings through providing different text styles to the user.
Regarding claim 2, Ma teaches identifying a font for the input text, wherein the output image is generated based on the font (pg. 330 3rd para. lines 7-12).
Regarding claim 5, Ma teaches identifying a text color, wherein the output image is generated based on the text color (Fig. 7).
Regarding claim 6, Ma teaches generating a mask for each character of the input text (pg. 334 sec. D 1st para. lines 1-3 and Fig. 10); and generating a character image for each character of the input text based on the mask (pg. 334 sec. D 1st para. lines 1-3 and Fig. 10), wherein the output image includes the character image for each character of the input text (Fig. 7).
Regarding claim 7, Ma fails to teach encoding the text effect prompt to obtain a text effect embedding, wherein the output image is generated based on the text effect embedding. Xie teaches encoding the text effect prompt to obtain a text effect embedding, wherein the output image is generated based on the text effect embedding (pg. 2 left col. 2nd para. lines 3-10). Therefore, it would be obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the text input styles of Ma and image generation models of Xie with the different text styles of Guy because this modification would improve the visual appearance of user input text strings through providing different text styles to the user.
Regarding claim 8, Ma teaches obtaining, via a styling interface, one or more styling parameters, wherein the output image is generated based on the one or more styling parameters (abst. lines 1-6 and pg. 330 3rd para. lines 7-12).
Regarding claim 16, Ma teaches obtaining, via a user interface, an input text including a plurality of characters (pg. 335 sec. V 1st para. lines 5-8); obtaining, via a user interface, a text effect prompt that describes a text effect for the input text (pg. 335 sec. V 1st para. lines 5-8); and generating, by an image generation (Figs. 2 & 7), an output image depicting the input text with the text effect described by the text effect prompt (pg. 331 ‘Text stye transfer’ 1st para. lines 1-23 and Fig. 7), wherein the output image depicts each of the plurality of characters from the input text (Fig. 7). However, Ma fails to teach an apparatus comprising: at least one processor; and at least one memory including instructions executable by the at least one processor to perform operations comprising generating, by an image generation model, wherein the text effect prompt comprises text different from the input text. Xie teaches generating, by an image generation model (pg. 5, sec. 4.4 2nd para. lines 1-6). Therefore, it would be obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the text input styles of Ma with the image generation models of Xie to enable a more controlled application of the text effect described by the text effect prompt in the generated output image. However, Ma and Xie fail to teach an apparatus comprising: at least one processor; and at least one memory including instructions executable by the at least one processor to perform operations comprising wherein the text effect prompt comprises text different from the input text. Guy teaches an apparatus comprising: at least one processor; and at least one memory including instructions executable by the at least one processor to perform operations (0004 lines 1-10) comprising wherein the text effect prompt (0137 lines 4-10) comprises text different from the input text (0101 lines 5-20, in which text and styles different from the input text is provided). Therefore, it would be obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the text input styles of Ma and image generation models of Xie with the different text styles of Guy because this modification would improve the visual appearance of user input text strings through providing different text styles to the user.
Claims 4, 11 and 17-20 are rejected under 35 U.S.C. 103 as being unpatentable over Ma, in view of Xie and further in view of Guy, and in further view of Smetanin et al, U.S Patent No 12,205,207 (hereinafter “Smetanin”).
Regarding claim 4, Ma, Xie and Guy fail to teach identifying a background color, wherein the output image is based on the background color. Smetanin teaches identifying a background color, wherein the output image is based on the background color (col. 18, lines 59-63, Fig. 10, set background 1006). Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ma, Xie and Guy to incorporate the teachings of Smetanin for identifying a background color, wherein the output image is generated based on the background color. Doing so would provide the user with an additional parameter for customizing the output image.
Regarding claim 11, Ma, Xie and Guy fail to teach identifying at least a portion of the text effect prompt as a negative text; and encoding the negative text to obtain a negative text effect embedding, wherein the output image is generated based on the negative text effect embedding. However, Smetanin teaches identifying at least a portion of the text effect prompt as a negative text (Smetanin, Col. 13, Lines 56-65, content moderation engine 408); and encoding the negative text to obtain a negative text effect embedding, wherein the output image is generated based on the negative text effect embedding (col. 15, lines 41-50, modified prompt). Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ma, Xie and Guy to incorporate the teachings of Smetanin for identifying a background color, wherein the output image is generated based on the background color. Doing so would provide the user with an additional parameter for customizing the output image.
Regarding Claim 17, Ma fails to teach wherein the image generation model comprises a diffusion model. Xie teaches wherein the image generation model comprises a diffusion model (Fig. 2). Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ma, Xie and Guy to incorporate the teachings of Smetanin for identifying a background color, wherein the output image is generated based on the background color. Doing so would provide the user with an additional parameter for customizing the output image.
Regarding claim 18, Ma and Xie fail to teach training a text effect encoder to encode at least a portion of the training text effect prompt to obtain a text effect embedding, wherein the output image is generated based on the text effect embedding. Smetanin teaches training a text effect encoder to encode at least a portion of the training text effect prompt to obtain a style embedding, wherein (col. 13, lines 31-39) wherein the output image is generated based on the text effect embedding (col. 15, lines 41-50). Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ma, Xie and Guy to incorporate the teachings of Smetanin for identifying a background color, wherein the output image is generated based on the background color. Doing so would provide the user with an additional parameter for customizing the output image.
Regarding claim 19, Ma, Xie and Guy fail to teach the text effect encoder comprises an aesthetic encoder configured to generate an aesthetic embedding and a style encoder configures to generate a style embedding, wherein the output image is generated based on the aesthetic embedding and style embedding. Smetanin teaches the text effect encoder comprises an aesthetic encoder configured to generate an aesthetic embedding and a style encoder configures to generate a style embedding, (col. 13, lines 31-39) wherein the output image is generated based on the aesthetic embedding and style embedding (col. 15, lines 41-50). Therefore, it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ma, Xie and Guy to incorporate the teachings of Smetanin for identifying a background color, wherein the output image is generated based on the background color. Doing so would provide the user with an additional parameter for customizing the output image.
Regarding claim 20, Ma teaches training a mask network to generate the mask for each character of the training input text (pg 5, Fig. 7, [a] style, [b] content, [i] ours), wherein the output image is generated based on the mask (Fig. 7).
Claims 9 and 10 are rejected under 35 U.S.C. 103 as being unpatentable over Ma in view of Xie, in view of Guy, in further view of Smetanin and further in view of Dehouche et al, "What is in a Text-to-Image Prompt: The Potential of Stable Diffusion in Visual Arts Education", arXiv, pages 1-11 (hereinafter "Dehouche").
Regarding Claim 9, Ma, Xie, Guy and Smetanin fail to teach generating a style embedding and an aesthetic embedding based on the text effect prompt wherein the output image is generated based on the style embedding and the aesthetic embedding. Dehouche teaches generating a style embedding and an aesthetic embedding based on the text effect prompt (Dehouche, Page 5-7, 4 Data and Methods; 5.1 Formalizing Stable Diffusion Prompts, Table 1, Table 2) wherein the output image is generated based on the style embedding and the aesthetic embedding (Dehouche, Page 8, Fig. 4). Therefore it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ma, Xie, Guy, and Smetanin to incorporate the teachings of Dehouche further comprising generating a style embedding and an aesthetic embedding based on the text effect wherein the output image is generated based on the style embedding and the aesthetic embedding. Doing so would provide the user with a more intuitive method for prompt generation using text input.
Regarding Claim 10, Ma, Xie, Guy and Smetanin fail to teach the text effect prompt comprises a style tag and the style embedding is based on the style tag. Dehouche teaches wherein the text effect prompt comprises a style tag (Page 5, Figure 2) and the style embedding is based on the style tag (Page 5-7, 4 Data and Methods; 5.1 Formalizing Stable Diffusion Prompts, Table 1, Table 2). Therefore it would be obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ma, Xie, Guy and Smetanin to incorporate the teachings of Dehouche further comprising generating a style embedding and an aesthetic embedding based on the text effect wherein the output image is generated based on the style embedding and the aesthetic embedding. Doing so would provide the user with a more intuitive method for prompt generation using text input.
Allowable Subject Matter
Regarding claim 3, the prior art fails to teach the limitations recited in claim 3. Therefore, claim 3 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
The following is a statement of reasons for the indication of allowable subject matter: Regarding claim 12, though prior art Smetanin teaches initializing an image generation model (col. 12 lines 43-59) the prior art fails to teach receiving training data including a training input text, a training image depicting the training input text, and a training text effect prompt that describes a text effect for the training input text; creating a training set by generating a mask for each character of the training input text, wherein the training set includes the training data and the mask generated for each character of the training input text; and training, using the training set including the mask, the image generation model to generate an output image based on the training input text and the training text effect prompt, wherein the output image comprises the text effect based on the training text effect prompt. Therefore claims 12-15 are allowable.
Response to Arguments
Applicant’s arguments with respect to claims 1-20 have been considered but are moot because claim 3 is object to as being dependent on a rejected claim, claims 12-15 are indicated allowable, and the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Said Broome whose telephone number is (571)272-2931. The examiner can normally be reached Monday - Friday 8:30am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Said Broome/Supervisory Patent Examiner, Art Unit 2612