Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s arguments with respect to claim(s) 1-12 and 17-20 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Applicant's arguments filed 13-16 have been fully considered but they are persuasive. Therefore, the previous office action are withdrawn. The claims 13-16 are allowed..
Claim Rejections -35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-6, 9-12 and 17-20 are rejected under 35 U.S.C 103 as being unpatentable over Mikhailiuk et al. (U.S. Pub. 2025/0150414 A1). in view of Kim et al. (U.S. pat. 12,456,243 B2)
With respect to claims 1, and 17, Mikhaliuk et al. discloses a data processing system comprising:
a processor, and
a machine-readable storage medium storing executable instructions which, when executed by the processor, cause the processor alone or in combination with other processors to perform the following operations:
receiving, via a user interface of a client device, a first prompt requesting an output visual content item to be generated(i.e.,” The system receives an image post from the user and generates a description of the image using an image-to-text model. User intent is determined based on the image and description. .If responding with an image is appropriate based on the user intent, the system generates a prompt using the image description and passes it to a text generation model to create an image description and caption. The image description and caption are used to synthesize a new image”(abstract)), the first prompt including a style visual content item and a topic content item (i.e., “a chatbot system responds to user posts comprising images with posts comprising images on an interaction system. The chatbot system leverages a large language model to support conversations on various topics and extend its capabilities to properly reply to image post.’(0021) and topic and properly is topic and style as claimed invention);
constructing a second prompt (fig. 4A shows step 408 is second prompt as claimed invention) by a prompt construction unit as an input to a first generative model, by appending the style visual content item and the topic content item to a first instruction string(i.e., “a chatbot system generates the prompt by appending style instructions to the post description. The style instructions include generating the image in an interactive platform post style.”(0027) description is topic of content as claimed invention and “In operation 410, in response to determining to respond with a chatbot interaction system post 444, the chatbot system 232 generates a prompt (e.g., prompt 502 of FIG. 5A) used o prompt a generative component 426 to generate
an image description 440 (e.g., image description 506 of FIG. 5B) and image caption 450 (e.g., image caption 508 of FIG. 5B) using the user intent 436.”(0083)), the first instruction string comprising instructions to the first generative model to generate a textual description combining text describing the topic with text describing the style as a third prompt i.e.,” The chatbot system generates a description of the post using an image-to-text model and determines a user intent based on the post and description. The chatbot system decides to respond with an image post based on the user intent. The chatbot system generates a prompt using the post description, and generates an image description and caption using the prompt. The chatbot system creates the image post using the image description and caption, and provides the image post to the user's client device..”(0026) and “the prompt 438 comprises a detailed text description of the image content to be generated such as, but not limited to, objects, scenes, people, colors, textures, styles, and the like. In some examples, the prompt 438 specifies the size, resolution and level of realism needed for an image.”(0083));
inputting the third prompt into a second generative model to generate the output visual content item visually depicting the topic by replacing one or more visual elements of the style visual content item with one or more different visual element based on the topic while preserving the style depicted in the style visual content item (i.e., “the prompt 438 comprises a detailed text description of the image content to be generated such as, but not limited to, objects, scenes, people, colors, textures, styles, and the like. In some examples, the prompt 438 specifies the size, resolution and level of realism needed for an image.”(0083) and “The SDK stored on the interaction server system 110 effectively provides the bridge between an external resource (e.g., applications 106 or applets) and the interaction client 104. This gives the user a seamless experience of communicating with other users on the interaction client 104 while also preserving the look and feel of the interaction client 104… the prompt 438 gives guidance on the graphic design, framing, lighting and composition. In some examples, the prompt 438 provides example images or artistic styles to emulate. In some examples, the prompt 438 may include keywords to make the image feel more humanmade rather than computer generated. By receiving such details in the prompt 438, the generative component 426 can produce a customized image description 440 and image caption 450 as requested by the chatbot system 232.”(0061));
providing the output visual content item to the client device (i.e., “The prompt 438 primes the generative component 426 to output an image description 440 and image
caption 450 tailored to the user intent, conversation history, and expected tone. By providing a customized prompt t 438, the chatbot system 232 can steer the generative component 426
to generate an image description 440 and image caption 450 that will lead to an image and caption that aligns with the user intent 436 and interaction goals.”(0084)). ; and
causing the user interface to present the output visual content item (fig. 3 shows user interface present the output visual content item or “he graphical user interface, presenting the modification performed by the transform system, may supply the user with additional interaction options. Such options may be based on the interface used to initiate the content capture and selection of a particular computer animation model (e.g., initiation from a content creator user interface). I”(0178)). But Mikhailiuk et al. does not disclose topic content item textually describing or visually depicting a topic requested by the user, the style specifying at least one of shadows, material or atmosphere. However, Kim discloses topic content item textually describing or visually depicting a topic requested by the user, the style specifying at least one of shadows, material or atmosphere (i.e., “upon determining that an object has been selected for modification, the scene-based image editing system provides one or more object attributes of the object for display via the graphical user interface displaying the object. For instance, in some cases, the scene-based image editing system retrieves a set of object attributes for the object (e.g., size, shape, or color) from the corresponding semantic scene graph and presents the set of object attributes for display in association with the object.”(col. 8, lines 1-10), “the scene-based image editing system enables user interactions that change the text of the displayed set of object attributes or select from a provided set of object attribute alternatives. Based on the user interactions, the scene-based image editing system modifies the digital image by modifying the one or more object attributes in accordance with the user interactions.”(col. 8, lines 11-20), “based on a generated semantic scene graph for an image generated utilizing various neural networks, the scene-based image editing system 106 determines objects, their attributes (position, depth, material, color, weight, size, label, etc.). The scene-based image editing system 106 utilizes the information of the semantic scene graph to edit an image intelligently as if the image were a real-world scene.”(col. 54, lines 50-57) and Examiner asserts that topic is object in reference and style is attribute of reference that including the material or atmosphere such as size, shape or color or fig. 60 user apply (prompt) shadow intensity value 6008, further, “the scene-based image editing system 106 provides a prompt for entry of textual user input. In some cases, upon detecting the user interaction with the object attribute indicator 1912c, the scene-based image editing system 106 maintains the corresponding object attribute for display, allowing user interactions to remove the object attribute in confirming that the object attribute has been targeted for modification.”(col. 68, lines 1-10) and textual (topic) and object attribute (style)).
It would have been obvious for a person of ordinary skill in the art, before the effective filing date of the claimed invention, to include Kim et al.’s feature in order to have flexibility, editing of digital images while efficiently reducing the user interaction typically required to make such edits for the stated purpose has been well known in the art as evidenced by teaching of Kim et al.(col. 1, lines 40-45). Father, both reference teach the same field such edit, transform and using Neural network to generate the digital image.
With respect to claims 2, and 18, Mikhaliuk et al. discloses wherein the machine-readable storage medium further includes instructions configured to cause the processor alone or in combination with other processors to perform operations of: receiving at least one user feedback on the output visual content item via the user interface (i.e.,. “interaction systems (e.g., social platforms, social media platforms, interactive platforms, extended reality platforms, messaging platforms, systems with which a user interacts, and the like) may provide ways for users to perform various functions, access information, and access entertainment.”(0019) and user interacts is feedback as claimed invention).
With respect to claims 3, and 19, Mikhaliuk et al. discloses the data processing system of claim 2, wherein the instructions to the first generative model further comprise instructions to construct a fourth prompt as an input to the first generative model, by appending the feedback and the output visual content item to another instruction string (i.e., “In some examples, a chatbot system generates the prompt by Appending style instructions to the post description. The style instructions include generating the image in an interactive platform post style.”(0027) and Examiner asserts that post description is instruction string as claimed invention, further indicate that the platform is interform for user interacts so can be fourth prompt or fifth prompt etc., ), the other instruction string comprising instructions to the first generative model to generate another textual description combining the feedback and the output visual content item as a fifth prompt ((i.e.,. “interaction systems (e.g., social platforms, social media platforms, interactive platforms, extended reality platforms, messaging platforms, systems with which a user interacts, and the like) may provide ways for users to perform various functions, access information, and access entertainment.”(0019) and user interacts is feedback as claimed invention and Fig. 4B shows the interaction or prompt with textual description with model 454, 456, 458, 460 and 462., and “ The image-to-text model applies this learning to generate a textual image description 430
summarizing the contents of the image 446 from the user interaction system post 428. The image description 430 is a textual summary reflecting the objects, people, actions, and scene captured in the image 446. The final textual image description 430 generated by the image processing component 202 provides a concise description of what the input image 446 depicts”(0072)) , and to input the fifth prompt into the second generative model to generate a subsequent output visual content item by replacing one or more visual elements of the output visual content item based on the feedback while preserving the topic and the style of the style visual content item (fig. 4B shows the subsequent output visual content item from 428 to 444 with user interacts feedback (0019) and “This gives the user a seamless experience of communicating with other users on the interaction client 104 while also preserving the look and feel of the interaction client 104”(0061)).
With respect to claims 4, and 20, Mikhaliuk et al. discloses the data processing system of claim 3, wherein the machine-readable storage medium further includes instructions configured to cause the processor alone or in combination with other processors to perform operations of: providing the subsequent output visual content item to the client device (i.e., “The prompt 438 primes the generative component 426 to output an image description
440 and image caption 450 tailored to the user intent, conversation history, and expected tone. By providing a customized prompt t 438, the chatbot system 232 can steer the generative component 426
to generate an image description 440 and image caption 450 that will lead to an image and caption that aligns with the user intent 436 and interaction goals.”(0084)).; and causing the user interface to present the subsequent output visual content item (fig. 3 shows user interface present the output visual content item or “he graphical user interface, presenting the modification performed by the transform system, may supply the user with additional interaction options. Such options may be based on the interface used to initiate the content capture and selection of a particular computer animation model (e.g., initiation from a content creator user interface). I”(0178))
With respect to claims 5 , Mikhaliuk et al. discloses the data processing system of claim 2, wherein the user feedback is collected via a user selection of at least one of a thumbs-up tab, a thumbs-down tab, a neutral tab, or a generating-more-image tab, a textual input, or a combination thereof (i.e., “ the image-to-text model is trained on large datasets of images labeled with textual captions and descriptions. This allows the image-to-text model to learn associations between visual patterns/features in images and corresponding textual descriptions. The image-to-text model applies this learning to generate a textual image description 430 summarizing the contents of the image 446 from the user interaction system post 428. The image description 430 is a textual summary reflecting the objects, people, actions, and scene captured in the image 446. The final textual image description 430 generated by the image processing component 202 provides a concise description of what the input image 446 depicts”(0072)).
With respect to claim 6, Mikhaliuk et al. discloses wherein the instructions to the first generative model further comprise instructions to check whether the third prompt contains the topic and the style, and to input the third prompt into the second generative model when the third prompt contains the topic and the style (fig. 7 B shows feature such as attributes, content, concept are model that training by machine-learning and “If responding with an image is appropriate based on the user intent, the system generates a prompt using the image description and passes it to a text generation model to create an image description and caption. ”(abstract)).
With respect to claim 9, Mikhaliuk et al. discloses wherein the output visual content item is a photo (fig. 3), a diagram, a chart, an image, an infographic, a video, an animation, a screenshot, a meme, a slide deck, a pictogram, an ideogram, or a software application background.
With respect to claim 10, Mikhaliuk et al. discloses wherein the topic content item comprises at least one of a visual content item or a textual content item (i.e., “ interaction with an interaction system may be enhanced by interacting with a chatbot system designed to simulate human conversation through voice commands or text chats”(0020) or “The image description 430 is a textual summary reflecting the objects, people, actions, and scene captured in the image 446”(0072)).
With respect to claim 11, Mikhaliuk et al. discloses wherein the first generative model is a language model or a multi-modal model (i.e., “A chatbot system may employ Natural Language Processing (NLP) and Machine Learning (ML)/artificial intelligence methodologies to understand and interpret a user's input and generate a response.”(0020)).
With respect to claim 12, Mikhaliuk et al. discloses wherein the second generative model is a text-to-image model or a vision model (i.e., “the collection management system 222 employs machine vision (or image recognition technology) and content rules to curate a content collection automatically.”(0056)).
Allowable Subject Matter
Claims 7-8 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims, since the prior art of record and considered pertinent to the applicant’s disclosure does not teach or suggest the claimed wherein the instructions to the first generative model further comprise instructions to construct a sixth prompt as an input to the first generative model when the third prompt misses at least one of the topic or the style, by appending the missed at least one of the topic or the style and the third prompt to another instruction string, the other instruction string comprising instructions to the first generative model to generate another textual description combining the missed at least one of the topic or the style and the third prompt as a seventh prompt, and to input the seventh prompt into the second generative model to generate a subsequent output visual content item by replacing one or more visual elements of the style visual content item based on the topic while preserving the style; wherein the machine-readable storage medium further includes instructions configured to cause the processor alone or in combination with other processors to perform operations of: providing the subsequent output visual content item to the client device; and causing the user interface to present the subsequent output visual content item.
Reasons for Allowance
Claims 13-15 are allowed.
The following is an examiner’s statement of reason for allowance:
With respect to claims 13-15, Applicant amended the claim including the allowed subject matter in the claim 7, therefore, the claims 13-15 are allowed.
Citation of Pertinent References
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
The patent to Gormley disclose Building travel itineraries using generative intelligent engine, U.S. Pub. No. 2025/0182222 A1.
The patent to Smith et al. disclose Building travel itineraries using generative intelligent engine, U.S. Pat. No. 12,299,858 B2.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HUNG T VY whose telephone number is (571)272-1954. The examiner can normally be reached M-F 8-5.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tony Mahmoudi can be reached at (571)272-4078. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HUNG T VY/Primary Examiner, Art Unit 2163 September 19, 2026