DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
The present application has claimed priority under 35 U.S.C. 119 from Japanese Patent Application No. JP2023-176690 filed 10/12/2023. Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55
Information Disclosure Statement
The information disclosure statement (IDS) was filed on 10/9/2024. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Specification
The lengthy specification has not been checked to the extent necessary to determine the presence of all possible minor errors. Applicant’s cooperation is requested in correcting any errors of which applicant may become aware in the specification.
Applicant is reminded of the proper language and format for an abstract of the disclosure.
The abstract should be in narrative form and generally limited to a single paragraph on a separate sheet within the range of 50 to 150 words in length. The abstract should describe the disclosure sufficiently to assist readers in deciding whether there is a need for consulting the full patent text for details.
The language should be clear and concise and should not repeat information given in the title. It should avoid using phrases which can be implied, such as, “The disclosure concerns,” “The disclosure defined by this invention,” “The disclosure describes,” etc. In addition, the form and legal phraseology often used in patent claims, such as “means” and “said,” should be avoided.
The abstract of the disclosure is objected to because it uses the term “disclosure”. A corrected abstract of the disclosure is required and must be presented on a separate sheet, apart from any other text. See MPEP § 608.01(b).
Drawings
The drawings filed 10/9/2024 were accepted.
Claim Objections
Claim 9 is objected to because of the following informalities: “fist image” should be “first image”. Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-17 are rejected under 35 U.S.C. 101.
Claim 1 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claims do not fall within at least one of the four categories of patent eligible subject matter because they are directed to an abstract idea without significantly more. The claims recite the abstract idea of specifying an image and setting prompt information.
Step 2A, Prong 1
The limitations that describe the specifying an image and setting prompt information are processes that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. The claims also include elements of a memory, a processor, a generative model which generates an image, and presenting the image on a screen, however nothing in the claims precludes the steps from practically being performed in the mind.
Step 2A, Prong 2
The judicial exception is not integrated into a practical application because the additional elements regarding a memory, a processor, a generative model which generates an image, and presenting the image on a screen are considered insignificant extra-solution activity. These limitations are not considered improvements to the functioning of a technology or technical field. The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the extrasolutionary elements are not considered significantly more than just applying the steps of specifying an image and setting prompt information.
Step 2B
In addition to the abstract idea, the claims have the memory, processor, generative model which generates an image, and presenting the image on a screen, but they represent only well-understood, routine, conventional activity that can be performed on generic computers. The memory and processor are considered as generic computer parts used to apply the exception. Mere instructions to apply an exception using generic computer components cannot provide an inventive concept. Brade et al (Brade, S., Wang, B., Sousa, M., Sageev Oore, & Grossman, T. (April 2023). Promptify: Text-to-Image Generation through Interactive Prompt Exploration with Large Language Models. https://doi.org/10.1145/3586183.3606725) discloses well-understood, routine, and conventional the generating and displaying of images is: abstract: “Text-to-image generative models have demonstrated remarkable capabilities in generating high-quality images based on textual prompts.” The claims are not patent eligible.
As per claim 2, this claim has similar image presentation and is rejected similarly to claim 1. Claim 2 also recites an additional abstract idea of a user selecting an image element. The selecting of an image is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. There are no other additional elements.
As per claim 3, this claim has similar selecting and presenting steps and is rejected similarly to claim 2.
Claim 4 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. As per claim 4, this claim recites an additional abstract idea of editing the prompt. The editing is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. There are no other additional elements.
As per claim 5, this claim has similar setting steps and is rejected similarly to claim 1. It also has similar selecting steps of claim 2 and is rejected similarly to claim 2.
As per claim 6, this claim has similar setting steps and is rejected similarly to claim 1.
Claim 7 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. As per claim 7, this claim recites an additional abstract idea of analyzing an element selected by a user. The analyzing an element selected by a user is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. This claim also has similar setting steps and is rejected similarly to claim 1.
As per claim 8, this claim has similar setting and analyzing steps and is rejected similarly to claim 7.
As per claim 9, this claim has similar setting and is rejected similarly to claim 1.
As per claim 10, this claim has similar image generation and is rejected similarly to claim 1.
Claim 11 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. As per claim 11, this claim recites an additional element of applying post-processing on the image.
(Step 2A, prong 2) The judicial exception is not integrated into a practical application because the additional elements regarding applying post-processing on the image are considered insignificant extra-solution activity. These limitations are not considered improvements to the functioning of a technology or technical field.
(Step 2B) The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the applying post-processing on the image are not considered significantly more than the judicial exception. The additional elements represent only well-understood, routine, conventional activity that can be performed on generic computer systems. Hertzfeld et al (US20070260979A1; filed 7/10/2006) discloses how well-understood, routine, and conventional processing an image is: Hertzfeld et al, paragraph 46: “As shown, the user interface 500 includes a pull down menu 505 that includes a number of effects that may be applied to the image 110. In the example shown, a user may apply… a grayscale effect 530… to the image 110.” The claims are not patent eligible.
Claim 12 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. As per claim 12, this claim recites an additional abstract idea of determining a flag. The determining of a flag is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. This claim also has similar generating images using the generative model and is rejected similarly to claim 1.
Claim 13 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. As per claim 13, this claim recites an additional element of a server.
(Step 2A, prong 2) The judicial exception is not integrated into a practical application because the additional elements regarding the server are considered insignificant extra-solution activity. These limitations are not considered improvements to the functioning of a technology or technical field.
(Step 2B) The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the server is not considered significantly more than the judicial exception. The additional elements represent only well-understood, routine, conventional activity that can be performed on generic computer systems. Rogers et al (US20210117859A1; filed 9/9/2020) discloses how well-understood, routine, and conventional processing an image is: Rogers et al, paragraph 20: “Technologies such as machine learning are increasingly being relied upon for a variety of different tasks. For many applications, it may be desirable to have the machine learning hosted on a remote server or other such system, rather than a client device.” The claims are not patent eligible.
As per claim 14, this claim has similar apparatus (the memory and processor) and is rejected similarly to claim 1.
As per claim 15, this claim has similar document (included in the presenting step) and is rejected similarly to claim 1.
Claim 16 recites substantially similar limitations to claim 1 and is thus rejected along the same rationale.
Claim 17 recites substantially similar limitations to claim 1 and is thus rejected along the same rationale.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-10, 12, and 14-17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Brade et al (Brade, S., Wang, B., Sousa, M., Sageev Oore, & Grossman, T. (April 2023). Promptify: Text-to-Image Generation through Interactive Prompt Exploration with Large Language Models. https://doi.org/10.1145/3586183.3606725) in view of Lin et al (Lin, Jinpeng, et al. AutoPoster: A Highly Automatic and Content-Aware Design System for Advertising Poster Generation. August, 2023, https://arxiv.org/abs/2308.01095.)
With regards to claim 1, Brade et al discloses An apparatus for editing… on an editing screen, the apparatus comprising: at least one memory that stores instructions; at least one processor that executes the instructions to (Brade et al, page 8, section 7: “Both Promptify and Automatic1111 were installed on a desktop computer that had an NVIDIA Titan RTX GPU”):
specify a generation-source image to be used for generating a new image (Brade et al, page 1: “Image Layout and Clustering allows users to view, organize, and cluster their generated images. Users can interact with images and clusters by dragging, dropping, and zooming. They can also access prompt refinement suggestions based on individual images or clusters.” The source images are interpreted as the images being used for prompt refinement);
set prompt information in a prompt input area (Brade et al, fig. 1: Prompt controls A);
cause a generative model to generate a new image using the specified generation-source image and the set prompt information (Brade et al, p. 4, section 5: “Promptify improves the workflow of text-to-image generation by offering three main features, including 1) automatic prompt extension and suggestion, 2) image layout and clustering by similarity, and 3) automatic prompt refinement suggestions. By integrating these features, Promptify establishes a feedback loop that enables users to generate high-quality images from initial prompts, examine how the model interprets their prompt, and receive suggestions to edit the initial prompts for enhancing output.”); and
present, on the editing screen, the generated new image as an element (Brade et al, Fig. 2: the figure shows the iterative process of refining the prompt and generating and review images).
However, Brade et al does not disclose editing a layout of a document… an element to be used for editing the layout of the document.
Lin et al teaches editing a layout of a document… an element to be used for editing the layout of the document (Lin et al, abstract: “With only product images and titles as inputs, AutoPoster can automatically produce posters of varying sizes”).
It would have been obvious to a person of ordinary skill in the art before the effective filing date to have combined Brade et al and Lin et al such that the generated images are used to edit a document layout. This would have enabled the invention to automatically generate documents such as posters (Lin et al, abstract: “Creating a poster involves multiple steps and necessitates design experience and creativity. This paper introduces AutoPoster, a highly automatic and content-aware system for generating advertising posters”).
With regards to claim 2, which depends on claim 1, Brade et al discloses wherein the specified generation-source image is a first image element selected by the user from a set of elements arranged on a screen (Brade et al, page 1: “Image Layout and Clustering allows users to view, organize, and cluster their generated images. Users can interact with images and clusters by dragging, dropping, and zooming. They can also access prompt refinement suggestions based on individual images or clusters.” The source images are interpreted as the images being used for prompt refinement).
However, Brade et al does not disclose for editing the layout of the document.
Lin et al teaches for editing the layout of the document (Lin et al, abstract: “With only product images and titles as inputs, AutoPoster can automatically produce posters of varying sizes”).
It would have been obvious to a person of ordinary skill in the art before the effective filing date to have combined Brade et al and Lin et al such that the generated images are used to edit a document layout. This would have enabled the invention to automatically generate documents such as posters (Lin et al, abstract: “Creating a poster involves multiple steps and necessitates design experience and creativity. This paper introduces AutoPoster, a highly automatic and content-aware system for generating advertising posters”).
With regards to claim 3, which depends on claim 1, Brade et al discloses wherein the specified generation-source image is a second image specified to replace a first image element selected by the user from a set of elements arranged on a screen (Brade et al, page 6, section 5.2: “Moreover, users can adjust the distance between images using the scale slider to avoid any disruptive overlap. They can compare and contrast different aspects of the images by dragging and rearranging them.” The dragging of elements can be used to swap the positions of elements, thus “replacing” image elements at their positions).
However, Brade et al does not disclose for editing the layout of the document.
Lin et al teaches for editing the layout of the document (Lin et al, abstract: “With only product images and titles as inputs, AutoPoster can automatically produce posters of varying sizes”).
It would have been obvious to a person of ordinary skill in the art before the effective filing date to have combined Brade et al and Lin et al such that the generated images are used to edit a document layout. This would have enabled the invention to automatically generate documents such as posters (Lin et al, abstract: “Creating a poster involves multiple steps and necessitates design experience and creativity. This paper introduces AutoPoster, a highly automatic and content-aware system for generating advertising posters”).
With regards to claim 4, which depends on claim 1, Brade et al discloses wherein the set prompt information can be edited based on an instruction from a user (Brade et al, p. 5, Fig. 2: “Users can skip Promptify’s suggestion features at any time and manually write parts of the prompt if they choose to”).
With regards to claim 5, which depends on claim 1, Brade et al discloses wherein the set prompt information is initially set with prompt information stored in association with an element selected by a user from a set of elements arranged on a screen (Brade et al, Fig. 2: The prompt is stored (if not in a database, then it is stored locally in a text box) for the iterative refinement of the group of images, which is “in association with” the first image; the term “initially” could be interpreted as being at the beginning of one of the iterations).
However, Brade et al does not disclose for editing the layout of the document.
Lin et al teaches for editing the layout of the document (Lin et al, abstract: “With only product images and titles as inputs, AutoPoster can automatically produce posters of varying sizes”).
It would have been obvious to a person of ordinary skill in the art before the effective filing date to have combined Brade et al and Lin et al such that the generated images are used to edit a document layout. This would have enabled the invention to automatically generate documents such as posters (Lin et al, abstract: “Creating a poster involves multiple steps and necessitates design experience and creativity. This paper introduces AutoPoster, a highly automatic and content-aware system for generating advertising posters”).
With regards to claim 6, which depends on claim 3, Brade et al discloses wherein the set prompt information is initially set with prompt information stored in association with the first image element (Brade et al, Fig. 2: The prompt is stored (if not in a database, then it is stored locally in a text box) for the iterative refinement of the group of images, which is “in association with” the first image; the term “initially” could be interpreted as being at the beginning of one of the iterations).
With regards to claim 7, which depends on claim 1, Brade et al discloses wherein the set prompt information is initially set with information extracted by analyzing an element selected by a user from a set of elements arranged on a screen (Brade et al, Fig. 2: the prompt is refined based on the images; the term “initially” could be interpreted as being at the beginning of one of the iterations).
However, Brade et al does not disclose for editing the layout of the document.
Lin et al teaches for editing the layout of the document (Lin et al, abstract: “With only product images and titles as inputs, AutoPoster can automatically produce posters of varying sizes”).
It would have been obvious to a person of ordinary skill in the art before the effective filing date to have combined Brade et al and Lin et al such that the generated images are used to edit a document layout. This would have enabled the invention to automatically generate documents such as posters (Lin et al, abstract: “Creating a poster involves multiple steps and necessitates design experience and creativity. This paper introduces AutoPoster, a highly automatic and content-aware system for generating advertising posters”).
With regards to claim 8, which depends on claim 3, Brade et al discloses wherein the set prompt information is initially set with information extracted by analyzing the first image element (Brade et al, Fig. 2: the prompt is refined based on the images; the term “initially” could be interpreted as being at the beginning of one of the iterations, or at the first iteration).
With regards to claim 9, which depends on claim 6, Brade et al discloses in a case where the second image does not satisfy an application condition prevent the prompt information stored in association with the fist [sic] image element from being initially set as the set prompt information (Brade et al, p. 2, section 1: “Promptify also assisted users in managing and comparing a large number of images to identify and reinforce desired features while ignoring unwanted ones in future iterations;” p. 10, section 7.5.3: “The simplest use case that users highlighted was the ability to ignore entire clusters that did not appeal to them.” The ignoring of a set cluster so that it no longer effects the prompt refinement, and therefore no longer effects the iterative image generation, is interpreted as preventing prompt information stored in association with the images in that cluster from being used for the prompt information).
With regards to claim 10, which depends on claim 1, Brade et al discloses wherein the new image is generated by the specified generation-source image, the set prompt information, and information indicating an application area for the prompt information being input to the generative model (Brade et al, p. 4, section 5: “three main features, including 1) automatic prompt extension and suggestion, 2) image layout and clustering by similarity, and 3) automatic prompt refinement suggestions. By integrating these features, Promptify establishes a feedback loop that enables users to generate high-quality images from initial prompts, examine how the model interprets their prompt, and receive suggestions to edit the initial prompts for enhancing output;” Fig. 2: Brade uses the generated images (image source) and their clustering positions (application area) to refine the prompt (prompt information) for generating new images).
With regards to claim 12, which depends on claim 1, Brade et al discloses determine whether a flag associated with the specified generation-source image indicates permission allowing generating the new image by using the specified generation-source image; execute control to cause the generative model to generate the new image in a case where it is determined that the flag indicates the permission; and execute control to prevent the generative model from generating the new image in a case where it is determined that the flag does not indicate the permission (Brade et al, p. 2, section 1: “Promptify also assisted users in managing and comparing a large number of images to identify and reinforce desired features while ignoring unwanted ones in future iterations;” p. 10, section 7.5.3: “The simplest use case that users highlighted was the ability to ignore entire clusters that did not appeal to them.” The ignoring of a set cluster so that it no longer effects the prompt refinement, and therefore no longer effects the iterative image generation, is interpreted as the flag that prevents those images in the cluster from being used to generate additional images).
With regards to claim 14, which depends on claim 1, Brade et al discloses wherein the generative model is included in the apparatus (Brade et al, page 8, section 7: “Both Promptify and Automatic1111 were installed on a desktop computer that had an NVIDIA Titan RTX GPU”).
With regards to claim 15, which depends on claim 1, Brade et al does not disclose wherein the document is at least any one of a poster and a flyer.
However, Lin et al teaches wherein the document is at least any one of a poster and a flyer (Lin et al, abstract: “With only product images and titles as inputs, AutoPoster can automatically produce posters of varying sizes”).
It would have been obvious to a person of ordinary skill in the art before the effective filing date to have combined Brade et al and Lin et al such that the generated images are used to edit a document layout. This would have enabled the invention to automatically generate documents such as posters (Lin et al, abstract: “Creating a poster involves multiple steps and necessitates design experience and creativity. This paper introduces AutoPoster, a highly automatic and content-aware system for generating advertising posters”).
Claim 16 recites substantially similar limitations to claim 1 and is thus rejected along the same rationale.
Claim 17 recites substantially similar limitations to claim 1 and is thus rejected along the same rationale.
Claim(s) 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Brade et al in view of Lin et al, and further in view of Hertzfeld et al (US20070260979A1; filed 7/10/2006).
With regards to claim 11, which depends on claim 1, Brade et al and Lin et al do not disclose apply post-processing for executing additional processing on the new image generated by the generative model, and wherein the additional processing includes at least one of background transparency processing and monochrome processing.
Hertzfeld et al teaches apply post-processing for executing additional processing on the new image generated by the generative model, and wherein the additional processing includes at least one of background transparency processing and monochrome processing (Hertzfeld et al, paragraph 46: “As shown, the user interface 500 includes a pull down menu 505 that includes a number of effects that may be applied to the image 110. In the example shown, a user may apply… a grayscale effect 530… to the image 110.”).
It would have been obvious to a person of ordinary skill in the art before the effective filing date to have combined Brade et al, Lin et al, and Hertzfeld et al such that the user can edit an image to grayscale/monochrome. This would have enabled the user to mute the colors of an image (Hertzfeld et al, paragraph 46: “The user can mute the colors of the image 110 by selecting the grayscale effect 530 from the menu 505”).
Claim(s) 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Brade et al in view of Lin et al, and further in view of Rogers et al (US20210117859A1; filed 9/9/2020).
With regards to claim 13, which depends on claim 1, Brade et al and Lin et al do not disclose wherein the generative model is included in a server.
Rogers et al teaches wherein the generative model is included in a server (Rogers et al, paragraph 20: “Technologies such as machine learning are increasingly being relied upon for a variety of different tasks. For many applications, it may be desirable to have the machine learning hosted on a remote server or other such system, rather than a client device”).
It would have been obvious to a person of ordinary skill in the art before the effective filing date to have combined Brade et al, Lin et al, and Rogers et al such that the model is executed remotely. This would have enabled the local device to require fewer resources to run the application (Rogers et al, paragraph 20: “For many applications, it may be desirable to have the machine learning hosted on a remote server or other such system, rather than a client device, due in part to the resources needed to execute the machine learning”).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Hertz et al (Hertz, A., Mokady, R., Tenenbaum, J., Aberman, K., Pritch, Y., & Cohen-Or, D. (2022). Prompt-to-Prompt Image Editing with Cross Attention Control. arXiv (Cornell University). https://doi.org/10.48550/arxiv.2208.01626): Teaches iteratively editing an image using prompts.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to BRODERICK C ANDERSON whose telephone number is (313)446-6566. The examiner can normally be reached Monday-Tuesday, Thursday-Saturday 9-5 PST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Stephen Hong can be reached at 5712724124. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/B.C.A/Examiner, Art Unit 2178
/STEPHEN S HONG/Supervisory Patent Examiner, Art Unit 2178