DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 10 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 10 recites “the stored prompt”. There is a lack of antecedent basis for the phrase. Therefore the scope of the phrase is indefinite.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 1, 3, 7-10,13-16 and 18-19 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by CHENG et al. (US Pat. Pub. No. 20250225430 “Cheng”).
Regarding claim 1 Cheng teaches An information processing apparatus (Fig. 6) for causing a generative AI to generate a content (“[0051] The second generative model 126b can be any visual generative model trained to generate visual content (e.g., image, video, and the like) blending topic(s) and style(s) seamlessly in response to natural language prompts”), the information processing apparatus comprising:
at least one memory that stores instructions; and at least one processor that executes the instructions (“[0119] The machine 600 may include processors 610, memory 630, and I/O components 650, which may be communicatively coupled via, for example, a bus 602”) to:
obtain a base image serving as a base of a content desired to be generated with the generative AI and a style image representing a style of the content (“[0038] In FIG. 1B, the request processing unit 122 receives a user style request 150 (including a style image/prompt 150a). [0039] The request processing unit 122 also receives a user content request 152 (including a topic such as a birthday cake). [0040] When the topic is a topic visual content 152b (e.g., a birthday cake image), the request processing unit 122 forwards the topic visual content 152b”);
extract, from the obtained style image, attribute information indicating the style represented by the style image (“[0038]….. To understand the image style, the system prompts the first generative model 126a endpoint to interpret the image, describe it, and then convert the image description into an LVM prompt. In one embodiment, the system provides detailed guidance with a meta prompt as shown in Table 1. The meta prompt in Table 1 provides context and guidance to the first generative model 126a, and helps the model capture the style in the style image/prompt 150a”; Here captured style is claimed attribute information); and
set, based on the extracted attribute information, a prompt for the generative AI to generate the content based on the base image (“[0040]…… The prompt construction unit 124 or the first generative model 126a can then integrate the topic textual prompt 152c with the style text prompt 150b into an integrated text prompt. The second generative model 126b (e.g., DALLE-3) can process the integrated text prompt to generate an image output 154, e.g., a modern living room styled birthday cake image 154a”).
Claim 18 is directed to a method claim and its steps are similar in scope and functions performed by the elements of the apparatus claim 1 and therefore claim 18 is also rejected with the same rationale as specified in the rejection of claim 1.
Claim 19 is directed to “A non-transitory computer readable storage medium” (“[0118] FIG. 6 is a block diagram illustrating components of an example machine 600 configured to read instructions from a machine-readable medium (for example, a machine-readable storage medium)”) claim and its elements are similar in scope and functions performed by the elements of the apparatus claim 1 and therefore claim 19 is also rejected with the same rationale as specified in the rejection of claim 1.
Regarding claim 3 Cheng teaches wherein in a case where a plurality of style images representing the style of the content are obtained, a plurality of pieces of the attribute information are extracted for the plurality of style images, and the prompt is set based on the extracted plurality of pieces of attribute information (“[0045] In some implementations, instead of image style transfer based on one style image and one content topic as the above-discussed example, the system can generate one image output based on a plurality of style images. [0070]……. The chat pane 225 further shows a field 225e with an instruction of “Explore other styles” and a field 225f with an instruction of “Explore other topics.””).
Regarding claim 7 Cheng teaches a storage that stores the prompt, wherein in a case where the storage has stored the prompt for the style image, the prompt according to the style image stored in the storage is set (“[0061] All the above-discussed visual content library 142 (storing e.g., topics, styles, elements, or the like), request, prompts and responses 144, extracted/inferred user data 146 (e.g., user activities, preferences, or the like), and other asset data 148 can be stored in the enterprise data storage 140”).
Regarding claim 8 Cheng teaches wherein the at least one processor executes the instructions further to correct the prompt stored in the storage based on a user operation, wherein in the setting of the prompt, the corrected prompt is set (“[0043]….. In one embodiment, the meta prompt 156 can include instructions that guides the agent on how to improve its own instructions based on user positive, neutral, or negative feedback on the image output 154, such as a user selection of a thumbs-up tab, a thumbs-down tab, a neutral tab, or a generating-more-image tab, a textual input, or the like. The system can then create another image output based on the refined integrated text prompt, and serve the refined image output to the user”).
Regarding claim 9 Cheng teaches wherein the prompt is corrected with the user operation performed on a UI screen displayed on a display (“[0043]….. In one embodiment, the meta prompt 156 can include instructions that guides the agent on how to improve its own instructions based on user positive, neutral, or negative feedback on the image output 154, such as a user selection of a thumbs-up tab, a thumbs-down tab, a neutral tab, or a generating-more-image tab, a textual input, or the like. The system can then create another image output based on the refined integrated text prompt, and serve the refined image output to the user”).
Regarding claim 10 Cheng teaches a storage that stores the attribute information, wherein in a case where the storage has stored the attribute information corresponding to the style image, the stored prompt is set (“[0061] All the above-discussed visual content library 142 (storing e.g., topics, styles, elements, or the like), request, prompts and responses 144, extracted/inferred user data 146 (e.g., user activities, preferences, or the like), and other asset data 148 can be stored in the enterprise data storage 140”).
Regarding claim 13 Cheng teaches a user interface that accepts whether to execute conversion by the generative AI with the set prompt for an input new content (“[0037] The first generative model 126a (e.g., an LLM, LMM, or the like) can interpret and convert style elements of a user style image by generating a style text prompt. Either prompt construction unit 124, or the first or second generative model can integrate the style text prompt with a topic extracted from a user content request into an integrated text prompt. This user content request may be presented in either text or image format. The second generative model 126b (e.g., a LVM, text-to-image model, or the like) can process the integrated text prompt to generate visual content item(s) that embodies both the user's style and topic preferences. [0070] FIG. 2B continues from FIG. 2A upon a selection of the mini application tile 225a. In this example, the chat pane 225 shows a prompt enter box 225c with an instruction of “Select or drop a style image and a content image/text” and several style images for the user to select”. Here generative model does conversion for style for selected new content).
Regarding claim 14 Cheng teaches wherein the user interface accepts whether to execute the conversion each time the new content is input (“[0070]….. After the user drops the modern living room image 150a and the birthday cake image 152b into the prompt enter box 225c, the chat pane 225 shows a prompt enter box 225d with an instruction of “Select one of the following images or generate other images of the same style” and several image outputs generated by the generative model(s) 126 for the user in FIG. 2C. The chat pane 225 further shows a field 225e with an instruction of “Explore other styles” and a field 225f with an instruction of “Explore other topics.”).
Regarding claim 15 Cheng teaches wherein the at least one processor executes the conversion in a case of accepting a user instruction with the user interface (“[0071] FIG. 2D continues from FIG. 2C upon a selection of another style image (e.g., a breakfast shop with a milk cardboard box design) in the field 225e while maintaining the same birthday cake topic. In this example, the chat pane 225 shows the prompt enter box 225d with an instruction of “Display style image and output image side-by-side” and the respective style image and output image underneath”).
Regarding claim 16 Cheng teaches wherein the attribute information is extracted as at least one of information indicating a season, information indicating an event, information indicating an image style, information indicating an atmosphere, information indicating an emotion, or information indicating an expression (“[0090]….. a style visual content item (e.g., the modern living room image 150a in FIG. 2B)”).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 2 is rejected under 35 U.S.C. 103 as being unpatentable over Cheng in view of CHENG_JING (US Pat. Pub. No. 20240169476 “Cheng_Jing”).
Regarding claim 2 Cheng is silent about wherein the attribute information is extracted by performing image recognition using an image recognition model.
Cheng_Jing teaches attribute information is extracted by performing image recognition using an image recognition model (“[0069] In some embodiments, before the original sub-image and the mask image are input into the target image generation network, the method further includes performing image style recognition on the original image to obtain a target image style and querying preset correspondences according to the target image style to obtain the target image generation network. ….For example, the image style recognition performed on the original image may be completed by using an image style classification model. The image style classification model may be, for example, a convolutional neural network (CNN) image classification model”);
Cheng and Cheng_Jing are analogous art as both of them are related to image processing.
Therefore it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Cheng by having attribute information that is extracted by performing image recognition using an image recognition model as taught by Cheng_Jing.
The motivation for the above is to use an automated method of style extraction from an image.
Claim(s) 4 is rejected under 35 U.S.C. 103 as being unpatentable over Cheng in view of Chavez et al. (US Pat. No. 10685057 “Chavez”) and Peddinti (US Pat. Pub. No. 20250124476 “Peddinti”).
Regarding claim 4 Cheng is silent about classify the extracted plurality of pieces of attribute information by type; and set the prompt based on a most frequent piece of attribute information in each of the types used for the classification.
Chavez teaches classify plurality of pieces of attribute information by type (Col 9 lines 11-25 “The memory 232 includes style classification data 244. The style classification data 244 may include information about style classifications available for the image search. The information may be metadata and/or labels identifying parameters for each of the style classifications. The style classification data 244 may identify a number of style classes such as lightning, blizzard, alpenglow, sunny, cloudy, snowfall, etc. for weather-related patterns, or midday, nighttime, twilight, dusk, dawn, etc. for time-of-day-related patterns, or autumn, summer, wintertime, etc. for seasonal patterns, or a certain corporate color style, or an arbitrary type of style class depending on implementation. The parameters may indicate a range of vector values that correspond to a particular style class such that the image search engine 244 may correlate the extracted image vector to vector values for a given style class”) and
Peddinti teaches set prompt based on a most frequent piece of attribute information in each of types (“[0050]….. In this example, each output from the LLM received by the review summarization module 219 may include a summarized review for a corresponding item. The LLM may generate a summarized review for an item based on a set of guidelines included in a prompt. In the above examples, based on a set of guidelines also included in each prompt, each summarized review for an item may have a length that does not exceed a character or word limit, include and highlight words or brands that are likely to be relevant to the customer, prioritize words that appear most frequently and that are also unique to the item, etc.”);
Cheng and Chavez and Peddinti are analogous art as both of them are related to image processing.
Therefore it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Cheng by classifying the extracted plurality of pieces of attribute information by type; and set the prompt based on a most frequent piece of attribute information in each of the types used for the classification based on the teaching of classify plurality of pieces of attribute information by type as taught by Chavez and set prompt based on a most frequent piece of attribute information in each of types as taught by Peddinti.
The motivation for the above is generate image based on most common style, thereby increasing the applicability of generated image.
Claim(s) 5 is rejected under 35 U.S.C. 103 as being unpatentable over Cheng in view of SREEDHAR et al. (US Pat. Pub. No. 20250211550 “Sreedhar”).
Regarding claim 5 Cheng is silent about wherein the attribute information which includes preset setting information is extracted;
Sreedhar teaches attribute information which includes preset setting information is extracted and prompt is set based on the attribute information with the extracted setting information (“[0093]…….If this dialogue flow 234 includes using the language model, a prompt may be generated for the language model using (for example) a predefined prompt style associated with this dialogue flow 234”);
Cheng and Sreedhar are analogous art as both of them are related to data processing.
Therefore it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Cheng by having attribute information which includes preset setting information is extracted and prompt is set based on the attribute information with the extracted setting information as taught by Sreedhar.
The motivation for the above is to provide controllability in generating a prompt.
Claim(s) 6 is rejected under 35 U.S.C. 103 as being unpatentable over Cheng in view of Azmandian et al. (US Pat. Pub. No. 20250128165 “Azmandian”).
Regarding claim 6 Cheng teaches the plurality of pieces of attribute information as shown above but is silent about generate a plurality of prompts and set a prompt selected by the user operation accepted via the user interface among the generated plurality of prompts;
Azmandian teaches a user interface that accepts a user operation, wherein, in the setting of the prompt, generate a plurality of prompts and set a prompt selected by the user operation accepted via the user interface among the generated plurality of prompts (“Claim 1 …….. the detection causing automatic generation of a plurality of additional prompts for editing one or more attributes of the specific portion, the plurality of additional prompts presented on the user interface for user selection; and receiving the user selection of an additional prompt from the plurality of additional prompts rendered at the user interface, the additional prompt selected by the user used to edit a current version of the specific portion of the storyline”);
Cheng and Azmandian are analogous art as both of them are related to data processing.
Therefore it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Cheng by having
a user interface that accepts a user operation, wherein, in the setting of the prompt, the at least one processor executes the instructions further to: generate a plurality of prompts from the plurality of pieces of attribute information; and set a prompt selected by the user operation accepted via the user interface among the generated plurality of prompts similar to having a user interface that accepts a user operation, wherein, in the setting of the prompt, generate a plurality of prompts and set a prompt selected by the user operation accepted via the user interface among the generated plurality of prompts as taught by Azmandian.
The motivation for the above is to provide user an option to generate image interactively.
Claim(s) 11-12 are rejected under 35 U.S.C. 103 as being unpatentable over Cheng in view of KIM (US Pat. Pub. No. 20250030920 “Kim”).
Regarding claim 11 Cheng is silent about correct the attribute information stored in the storage based on a user operation;
Kim teaches correct attribute information stored in the storage based on a user operation (“[0075] Furthermore, the style settings screen 410 may include a user settings item 440. The user may select the user settings item 440 and input a style image. This is described in detail with reference to FIG. 5.
[0077] In addition, according to some embodiments, the display device 100 may download or update styles from an external device or server periodically or based on a user input.
[0088] The style storage 610 may store a plurality of style features received from an external device. In addition, the style storage 610 may store style features extracted by the style feature extractor 620”);
Cheng and Kim are analogous art as both of them are related to image processing.
Therefore it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Cheng by correcting attribute information stored in the storage based on a user operation as taught by Kim.
The motivation for the above is to provide user the capability to change style image according to his choice.
Cheng modified by Kim teaches wherein the prompt is set based on the corrected attribute information (Cheng has option to explore other styles and based on updated style prompts get set.)
Regarding claim 12 Cheng modified by Kim teaches wherein the attribute information is corrected with the user operation performed on a UI screen displayed on a display (Kim “[0077] In addition, according to some embodiments, the display device 100 may download or update styles from an external device or server periodically or based on a user input”);
Claim(s) 17 is rejected under 35 U.S.C. 103 as being unpatentable over Cheng in view of Jiang et al. (US Pat. No. 10496698 “Jiang”).
Regarding claim 17 Cheng teaches wherein the at least one processor executes the instructions further to generate a product incorporating a content generated by the generative AI (“[0078]…… The system can instruct the generative model(s) 126 to generate a single-shot prompt (i.e., including a single example or instruction to guide the language model's response) or a multi-shot prompt (i.e., including multiple examples or instructions to give the model more context and improve its understanding of the task) for generating the virtual content output”) but is silent about wherein the style image is obtained from a template prepared for the product.
Jiang teaches style image is obtained from a template prepared for a product (Col 10 lines 65-67 “ Based on one or more images 502 and one or more phrases/sentences 503, style generation module 504 generates a list of style candidates based on style templates 165. Each content style candidate contains at least one of images 502 and one or more phrases/sentences 503”);
Cheng and Jiang are analogous art as both of them are related to image processing.
Therefore it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Cheng by obtaining style image from a template prepared for a product as taught by Jiang.
The motivation for the above is to provide structural consistency in providing style image.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SAPTARSHI MAZUMDER whose telephone number is (571)270-3454. The examiner can normally be reached 8 am-4 pm PST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Said Broome can be reached at (571)272-2931. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SAPTARSHI MAZUMDER/ Primary Examiner, Art Unit 2612