DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claim 5 objected to because of the following informalities:
The element “the generated image” lacks antecedent basis. (NOTE: Claim 1 recites “acquired generated image”.) Appropriate correction is required.
Claim 8 is objected to under 37 CFR 1.75 as being a substantial duplicate of claim 7. When two claims in an application are duplicates or else are so close in content that they both cover the same thing, despite a slight difference in wording, it is proper after allowing one claim to object to the other as being a substantial duplicate of the allowed claim. See MPEP § 608.01(m).
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-3, and 7-9 are rejected under 35 U.S.C. 102(a1) as being anticipated by Yin et al. (Yin, H., Zhang, Z., & Liu, Y. (2023). The Exploration of Integrating the Midjourney Artificial Intelligence Generated Content Tool into Design Systems to Direct Designers towards Future-Oriented Innovation. Systems, 11(12), 566. https://doi.org/10.3390/systems11120566, hereinafter “Yin”).
Regarding claim 7,
Yin teaches:
An image generation supporting apparatus comprising a hardware processor that: (Yin: Page 1 Section 1, “. . . Midjourney, rooted in the Stable Diffusion AI painting paradigm, is a text-driven image generation tool. . .; Page 9, Fig. 10.; NOTE: Figure 10 illustrates users and computers as the supporting apparatus that accesses the Midjourney tool. Computers inherently comprise hardware processors.)
acquires text data generated based on an image (Yin: Page 7, Section 2.3, “. . . After analyzing the ‘impenetrable exoskeleton’ styling sensibility of the Cybertruck we converted its design elements—such as morphological style, materials, and color scheme—into textual descriptions to prepare for composing the prompt. AIGC Midjourney and our design team collaboratively worked on the exterior. To enhance the accuracy of the prompt, we used Midjourney’s “/describe” command to input an image of the Cybertruck, allowing the AI to extract its stylistic elements.; NOTE: Yin uses the /describe command to extract text data (textual descriptions) based on the image input, which is the image of the Cybertruck);
displays the acquired text data on a display part (Yin: Page 7, Section 2.3, “. . . The prompt serves as the primary means of interaction between designers and AIGC, and selecting the correct prompt words is crucial to maximizing AIGC’s efficiency. . .”; Page 8, Section 2.4, “. . .further optimized the Prompt cue words for the iteration. This interdisciplinary communication model for “idea visualization + language expression” . . .”; NOTE: The display part is the laptop screen as shown in Fig. 10. As discussed above, the acquired text are the textual descriptions extracted from the Cybertruck image. The textual descriptions are then displayed on the laptop screens of the users from the /describe feature for interaction.);
edits the acquired text data after displaying the text data (Yin: Page 7, Section 2.3, “. . . Additionally, we used AMP-Cards to input the prompt alongside the design concept: Flying car design, industrial design, black and white split, cold stamping, modernism, futurism, Tesla, perspective, white background, OC rendering, studio lighting, 4K. . .”; Page 8, Section 2.4, “. . . group members discussed the scheme and further optimized the Prompt cue words for the iteration. . .”; NOTE: After extracting textual description from the Cybertruck image, users can then edit the textual description by inputting AMP-Cards, which are style elements alongside the extracted words to generate an image, see Fig. 8 in page 8. Cybertruck Image – Extract textual description using /describe – Editing by inputting AMP-Cards alongside the textual description.);
acquires an image generated based on the edited text data (NOTE: See Fig. 8 in page 8 where an image of a “flying Cybertruck” is acquired based on the textual description in addition to AMP Cards discussed above.);
and displays the acquired generated image on the display part (NOTE: See Fig. 10 page 9. The display part constituting the laptop screens displays design images.).
Regarding claim 8,
System claim 8 is drawn to the apparatus corresponding to the configuration of using same as claimed in the apparatus of claim 7. Therefore, system claim 8 corresponds to the configuration of the apparatus of claim 7, and is rejected for the same reasons of anticipation as used above.
Regarding claim 9,
Method claim 9 is drawn to the method corresponding to the configuration of using same as claimed in the apparatus of claim 7. Therefore, method claim 9 corresponds to the configuration of the apparatus of claim 7, and is rejected for the same reasons of anticipation as used above.
Regarding claim 1,
Yin teaches:
A non-transitory computer-readable recording medium storing a program executable by a computer (NOTE: Laptop computers illustrated in Fig. 10 inherently comprise non-transitory CRM such as RAM/ROM which stores programs to access tools such as Midjourney.),
CRM claim 1 is drawn to the apparatus corresponding to the configuration of using same as claimed in the apparatus of claim 7. Therefore, CRM claim 1 corresponds to the configuration of the apparatus of claim 7, and is rejected for the same reasons of anticipation as used above.
Regarding claim 2, depending on 1,
Yin teaches:
The recording medium according to claim 1,
Yin further teaches:
wherein the editing includes receiving the edited text data that is changed according to an input after displaying the text data (Yin: Page 7, Section 2.3, “. . . Additionally, we used AMP-Cards to input the prompt alongside the design concept: Flying car design, industrial design, black and white split, cold stamping, modernism, futurism, Tesla, perspective, white background, OC rendering, studio lighting, 4K. . .”; Page 8, Section 2.4, “. . . group members discussed the scheme and further optimized the Prompt cue words for the iteration. . .”; NOTE: The received edited text data that is changed is the textual description extracted from the Cybertruck image + the inputted AMP cards to generate ).
Regarding claim 3, depending on 1,
Yin teaches:
The recording medium according to claim 1,
Yin further teaches:
wherein the editing includes, after displaying the text data, changing the displayed text data to the edited text data that is associated with the displayed text data. ((Yin: Page 7, Section 2.3, “. . . Additionally, we used AMP-Cards to input the prompt alongside the design concept: Flying car design, industrial design, black and white split, cold stamping, modernism, futurism, Tesla, perspective, white background, OC rendering, studio lighting, 4K. . .”; Page 8, Section 2.4, “. . . group members discussed the scheme and further optimized the Prompt cue words for the iteration. . .”; NOTE: By adding AMP Cards to the features extracted from the /describe command, the displayed text data will change accordingly. Figure 9 illustrates multiple iterations where users optimizing the prompt changing the displayed text data to the edited text data for every iteration, the edited text data is associated with the displayed text data because they are used to optimize the prompt. Every optimization of the prompt includes changes to the displayed text data, including the addition of the AMP-Card elements.)
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over Yin in view of Wang et al. (Wang, Z., Huang, Y., Song, D., Ma, L., & Zhang, T. (2024, May). Promptcharm: Text-to-image generation through multi-modal prompting and refinement. In Proceedings of the 2024 CHI conference on human factors in computing systems (pp. 1-21, hereinafter “Wang”).
Regarding claim 4, depending on 1,
Yin teaches:
The recording medium according to claim 1,
However, Yin fails to teach: wherein the text data includes a plurality of pieces of text data, and the editing includes, after displaying the plurality of pieces of text data, editing a priority order of the displayed plurality of pieces of text data.
The analogous art Wang teaches:
wherein the text data includes a plurality of pieces of text data (NOTE: Fig. 3 in page 7 Step A, illustrates text data in the prompt editor box including plurality of pieces of text data (“A” “painting” . . . . “moon”).),
and the editing includes, after displaying the plurality of pieces of text data (NOTE: : Fig. 3 in page 7 Step D, illustrates the displayed prompt where a user selects to edit the token “Thomas Kinkade” Attention attribute),
editing a priority order of the displayed plurality of pieces of text data (Wang: Page 9-10 Section 5, “. . . Nevertheless, she observes that the model fails to include the “human child” in the image and, intriguingly, includes multiple instances of the “wolf” object. Therefore, Alice resolves to refine her prompt with PromptCharm’s assistance to rectify her image. Alice discerns that the word “wolf” has a very high model attention value (Fig. 4 d). When she hovers over the word “wolf”, a large portion of the image is highlighted. She interprets the model may probably over-attend to the word “wolf”. Therefore, she clicks on the word “wolf” and select the Attention option. She proceeds to reduce the model’s attention to this word by a factor of 0.5 before regenerating the image (Fig. 4 d). As a result, the image now accurately features both the “human child” and the “wolf” objects (Fig. 6c).; NOTE: The user alice changes the attention attribute for the text data “wolf” effectively editing its relative priority in generating desired image. In the example, fig. 6A and 6B prioritizes wolf images to be generated, by changing the attention of the wolf, it effectively changed the “wolf” priority for generation, and the result is fig 6C with the wolf and the child in the generated image. Fig. 3 shows editing the priority order using a slider. The higher the attention for a token, the higher the priority it gets when generating an image.).
It would have been obvious to a person having ordinary skill in the art (PHOSITA) before the effective filing date of the claimed invention to combine Yin, and Wang and include: wherein the text data includes a plurality of pieces of text data, and the editing includes, after displaying the plurality of pieces of text data, editing a priority order of the displayed plurality of pieces of text data.
The reason for doing so is “helping users iteratively improve one generated image through multi-modal prompting by adjusting the model’s attention to keywords in the prompt. By adjusting the model’s attention, the user does not need to rewrite their prompts to align the model’s interpretation with their creative intention. Therefore, they can avoid risking completely changing the image content when revising the prompt” (Wang Page 3 Section 2.2).
Claim(s) 5-6 are rejected under 35 U.S.C. 103 as being unpatentable over Yin in view of Hogeg et al. (US 20140173424 A1, hereinafter “Hogeg”).
Regarding claim 5, depending on 1,
Yin teaches:
The recording medium according to claim 1,
However, Yin fails to teach: wherein the program causes the computer to further execute: reflecting a predetermined image on the generated image; and displaying an image obtained by reflecting the predetermined image on the generated image on the display part.
The analogous art Hogeg teaches:
wherein the program causes the computer to further execute (Hogeg” ¶30, “. . . software instructions being executed by a computer. . .”):
reflecting a predetermined image on the generated image (Hogeg: ¶40-42, “. . . editing comprises processing one or more images . . ., adding visual content to an image or one or more frames of a video file, for example graphic elements, adding audible content to an image and/or a video file, overlying or embedding text boxes, overlying or embedding graphical elements, . . . visual content added to the edited visual content is referred to as a visual overlay or overlay. . . This communication allows the system 100 to receive requests for visual content editing functions and to respond with a list of selected visual content editing functions,”; NOTE: An overlay is a reflected image on a generated image. It is predetermined because the image is provided by a server illustrated in Fig. 3 reference 301.);
and displaying an image obtained by reflecting the predetermined image on the generated image on the display part (Hogeg: ¶79, “As shown at 208, the adjusted visual content may now be outputted, for example uploaded to a server, posted an/or shared with other subscribers of the system 100, for example as a visual twit, such as a Mobli.TM. message, Twitter.TM. message and/or an Instagram.TM. message, uploaded to a social network webpage and/or forwarded to one or more friends, for example as an electronic message such as an multimedia messaging service (MMS) message. “; NOTE: Fig. 3 step 208 outputs the image obtained constituting the adjusted visual content. This includes the actual image, and the reflected image (overlay). The image is outputted to the display part of the user’s/other subscriber electronic device to view visual twit or Instagram posts. ).
It would have been obvious to a person having ordinary skill in the art (PHOSITA) before the effective filing date of the claimed invention to combine Yin, and Hogeg and include: wherein the program causes the computer to further execute: reflecting a predetermined image on the generated image; and displaying an image obtained by reflecting the predetermined image on the generated image on the display part.
The reason for doing so is to provide a visual content that include “descriptive features that allow classifying depicted scenes and/or characters and/or for identifying certain objects”, “For example, if a car is identified in the visual content, a visual content editing function with an overlay that say "I just pimped my ride" may be selected. If detected a plate or a food portion is identified, a visual content editing function with an overlay and/or a sound overlay that includes promotional content of a food company is presented. In another example, of smiling faces are identified, promotional content to a toothbrush maybe presented” (Hogeg: ¶73)
Regarding claim 6, depending on 5,
The combination of Yin, and Hogeg teaches:
The recording medium according to claim 5,
Hogeg further teaches:
wherein the predetermined image includes a catchphrase of the generated image (Hogeg: ¶73, “. . . a visual content editing function with an overlay that say "I just pimped my ride" may be selected. . .”; ¶61, “. . . In another example, an overlay with a specific slogan is selected. . .”).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PATRICK GALERA whose telephone number is (571)272-5070. The examiner can normally be reached Mon-Fri 0800-1700 ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, King Poon can be reached at 571-270-0728. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PATRICK P GALERA/Examiner, Art Unit 2617 /KING Y POON/Supervisory Patent Examiner, Art Unit 2617