Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
Allowable Subject Matter
Claims 6-9, 17 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1-3, 5, 12-14, 18-19 is/are rejected under 35 U.S.C. 102(a)(1) as being clearly anticipated by Davidson U.S. Patent/PG Publication 20240264723.
Regarding claim 1 (independent):
A device, comprising: a processor system and storage accessible to the processor system and comprising instructions executable by the processor system to: (Davidson [0022] Additional embodiments include a computer, computing device, or system or combination thereof capable of carrying out the method and its implementations. The computer, computing device, system or combination thereof can include one or more processors capable of executing the computer-readable code, computer-readable instructions, computer-executable instructions, or “software”, one or more interface capable of providing input or output, one or more databases and a set of instructions (e.g., software) stored in a memory of the computer, computing device, or system or combination thereof for carrying out the method and its implementations. The computer, computing device, or system or combination thereof can include one or more stand-alone computer, such as a desktop computer, a portable computer, such as a tablet, laptop, PDA, or smartphone, or a set of computers or devices connected through a network including a client-server configuration and one or more database servers.)
receive a prompt to a generative image model (Davidson [0025] A user logs into an application of a generative AI model hosted on the Internet. After logging in, the application displays the graphical user interface shown in FIG. 1. The application provides a text box or prompt 110 for the user to enter initial image characteristic information.).
augment the prompt with data related to one or more user preferences indicated via user input (Davidson [0026] The application then provides static options for the user to choose from, selectable by check boxes, to add more information such as: [0027] Style 115: anime, abstract, classic, comic, Rockwell, Renoir, etc. In the example shown in FIG. 1, the user has selected “classic”. A user may be able to provide examples of a particular artist (for example by scanning and/and or uploading representative images) to the application to provide the AI with the style that the user wants its image mimic. [0028] Color palate 120: vibrant, subdued, grey scale, black and white, colorable line drawing (e.g., like a coloring book). In the example shown in FIG. 1, the user has selected “vibrant”. This causes a vibrancy scale 125 to appear below, which is selectable with a slidable indicator 130.).
provide the augmented prompt as input to the generative image model and based on providing the augmented prompt as input to the generative image model, receive an output from the generative image model, the output indicating a generative image in conformance with the augmented prompt (Davidson [0029] The user inputs “A girl rides a bike down a street” either typed into the text box 110 or spoken. The application than uses generative AI to assimilate the keywords and additional information into an image for the user.).
Regarding claim 2:
The device of claim 1, has all of its limitations taught by Davidson. Davidson further teaches wherein the instructions are executable to:
augment the prompt by using the data to alter the prompt to indicate the one or more user preferences (Davidson [0039] The second option provides a word table or icon table of various items (houses, dogs, cats, trees, etc.) on a graphical user interface.)(Davidson [0021] The remote computer(s) can receive image feature, context, or characteristic information inputted by the user through the user interface (either provided as the application on the user's computer or device, or accessed through the web server) and can have a memory capable of housing storages of the inputted information. The memory of the remote computer(s) can house a trained generative AI model or models such as text-to-image diffusion models programmed as computer-readable code used to generate images from text prompts.)(Davidson [0017] The generated text includes a set of options or questions that refine the initial subject matter entered in the text prompt.).
Regarding claim 3:
The device of claim 1, has all of its limitations taught by Davidson. Davidson further teaches wherein the instructions are executable to:
augment the prompt by appending the data to the prompt as an addition to the prompt (Davidson [0030] The application then provides additional selectable options on the graphical user interface shown in FIG. 2, based on the initial input to obtain additional information to refine the initial image such as: [0031] Age 210: 3-6, 7-10, 11-13, 14-17 [0032] Type of area 215: rural, suburban, urban [0033] Type of street 220: windy, straight, hilly, tree-lined, cul-de-sac [0034] Time of day 225: morning, afternoon, twilight, night [0035] Weather 230: sunny, cloudy, raining, snowing, windy.)(Davidson Claim 1 inputting or receiving information on one or more image characteristics from a user by way of a first user prompt displayed on a graphical user interface; outputting one or more questions or options for additional details of the one or more image characteristics on a second user prompt displayed on the graphical user interface by way of a generative artificial intelligence language model performed on one or more processor, wherein the questions or options of the second user prompt are generated based on the information obtained from the first user prompt; inputting or receiving the additional details from the user by way of the second user prompt displayed on the graphical user interface) since the additional information comes after the original prompt.
Regarding claim 5:
The device of claim 4, has all of its limitations taught by Davidson. Davidson further teaches wherein the audible, verbal input relates to an object in a geographic area (Davidson [0016] The present disclosure relates to an application, website, or program that allows a person to speak or type one or more image characteristics. The image characteristics can include subject matter, features, or context.)(Davidson [0032] Type of area 215: rural, suburban, urban [0033] Type of street 220: windy, straight, hilly, tree-lined, cul-de-sac [0034] Time of day 225: morning, afternoon, twilight, night [0035] Weather 230: sunny, cloudy, raining, snowing, windy.).
Regarding claim 12 (independent):
The claim is a/an parallel version of claim 1. As such it is rejected under the same teachings.
Regarding claim 13:
The claim is a/an parallel version of claim 1. As such it is rejected under the same teachings.
Regarding claim 14:
The method of claim 12, has all of its limitations taught by Davidson. Davidson further teaches comprising:
training the generative image model to augment received prompts with user preferences to produce generative outputs that incorporate the one or more user preferences (Davidson [0017] According to some implementations, the images are generated from text by way of a trained generative AI model or models. The generative AI models can be trained using a training set comprising millions or billions of images paired with text. In one implementation, the training set is an open-source data set. […] A trained text-to-image diffusion model can be hosted as an application on a web server that receives text input of image characteristics from users logged in to the application remotely, or can be hosted as an application on the user's computer or computing device directly. The trained text-to-image diffusion model can be combined with a trained language-based AI model such as a Generative Pre-trained Transformer model (e.g., GPT-3, an autoregressive language model developed by OpenAI), to generate additional options or questions for the user to refine the image characteristics based on input of initial image characteristics obtained from a text prompt. The generative AI language model uses deep learning to produce human-like text from the text prompt of initial image characteristics entered by the user. The generated text includes a set of options or questions that refine the initial subject matter entered in the text prompt. Instead of being trained on images, the language-based AI model is trained on volumes of text data such as books and articles, and generates text based on patterns found in the text data using neural networks as transformers for processing the text data.)(Davidson [0021] The memory of the remote computer(s) can house a trained generative AI model or models such as text-to-image diffusion models programmed as computer-readable code used to generate images from text prompts. The memory can also house trained generative AI language models programmed as computer-readable code used to generate text prompts based on initial input of text of desired image characteristics from a user. The generated text prompts refine the image characteristics by eliciting additional details for the desired image. A training set comprising a repository of regular or compressed images paired with text can also be stored in stored in memory on a server or servers that communicate with the user's computer or computing device(s) or other remote computer(s). The training set can be used to further train the generative AI model(s). The remote computer(s) can include the set of computer-executable instructions stored in memory that implement the generative AI models or produce the text prompts.).
Regarding claim 18 (independent):
The claim is a/an parallel version of claim 1. As such it is rejected under the same teachings.
Regarding claim 19:
The at least one CRSM of claim 18, has all of its limitations taught by Davidson. Davidson further teaches wherein the model comprises a generative image model, and wherein the output comprises a generative image (Davidson [0002] In general, in a first aspect, the disclosure features a method. The method includes inputting or receiving information on one or more image characteristics from a graphical user interface, outputting one or more questions or options for additional details on the one or more image characteristics on a graphical user interface by way of a generative artificial intelligence language model performed on one or more processor, inputting or receiving the additional details from the graphical user interface, and outputting one or more images by way of a generative artificial intelligence text-to-image model performed on one or more processor based on the one or more image characteristics and the additional details.).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 4, 10-11, 15-16, 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Davidson U.S. Patent/PG Publication 20240264723 in view of Sadr U.S. Patent/PG Publication 20240330381.
Regarding claim 4:
The device of claim 1, has all of its limitations taught by Davidson. Davidson further teaches wherein the instructions are executable to:
identify the one or more user preferences based on audible, verbal input from a user as received (Davidson [0016] The present disclosure relates to an application, website, or program that allows a person to speak or type one or more image characteristics. The image characteristics can include subject matter, features, or context.).
Davidson does not teach preferences before submitting the prompt. In a related field of endeavor, Sadr teaches:
identify the one or more user preferences based on (Sadr [0080] For example, the model-generated user-specific terms 212 can be generated by machine-learned models (e.g., LLM) based user personalization data, search engine data, and merchant data. The user personalization data can include explicit personalization that is user generated and/or user controlled. Additionally, the user personalization data can include interest graph data. The search engine data can include fashion knowledge data, recent trends data, and implicit personalization data. The implicit personalization can be based on a plurality of characteristics derived from the user (e.g., time, third-party data, location, intent). The merchant data can include merchant assets data. In some instances, the user personalization data, search engine data, and merchant data can be inputted into a machine-learned model (e.g., large language model (LLM)) to generate user-specific terms for a user. The LLM is a type of machine learning model that can perform a variety of natural language processing NLP tasks, such as generating and classifying terms. The user personalization data can include a prompt input, historical data (e.g., data descriptive of user search history, user purchase history, user browsing history, etc.), profile data, and/or preference data. The prompt input can include a freeform prompt input and/or a generated prompt input generated based on one or more tile selections of a user interface. The prompt input can be descriptive of one or more attributes a user is requesting to be rendered in a generated image. The prompt input can include a subject of the image (e.g., an environment and/or one or more objects) and one or more details for the subject (e.g., a color, a style, a material, etc.).).
Therefore, it would have been obvious before the effective filing date of the claimed invention to use prior data as taught by Sadr. The motivation for doing so would have been to provide the user with greater personalization and customization, with tailored results. Therefore it would have been obvious to combine Sadr with Davidson to obtain the invention.
Regarding claim 10:
The device of claim 1, has all of its limitations taught by Davidson. Davidson further teaches wherein the instructions are executable to:
identify the one or more user preferences based on a user’s Internet browser (Davidson [0021] The graphical user interfaces can be downloaded together as an application from cloud storage services providing retail application downloading to the user's computer or computing device, or can be hosted on a remote computer which acts as a web server and accessed through the Internet as webpages through an internet browser on the user's computer or computing device.).
Davidson does not teach browser history. In a related field of endeavor, Sadr teaches:
identify the one or more user preferences based on a user’s Internet browser history (Sadr [0080] For example, the model-generated user-specific terms 212 can be generated by machine-learned models (e.g., LLM) based user personalization data, search engine data, and merchant data. The user personalization data can include explicit personalization that is user generated and/or user controlled. Additionally, the user personalization data can include interest graph data. The search engine data can include fashion knowledge data, recent trends data, and implicit personalization data. The implicit personalization can be based on a plurality of characteristics derived from the user (e.g., time, third-party data, location, intent). The merchant data can include merchant assets data. In some instances, the user personalization data, search engine data, and merchant data can be inputted into a machine-learned model (e.g., large language model (LLM)) to generate user-specific terms for a user. The LLM is a type of machine learning model that can perform a variety of natural language processing NLP tasks, such as generating and classifying terms. The user personalization data can include a prompt input, historical data (e.g., data descriptive of user search history, user purchase history, user browsing history, etc.), profile data, and/or preference data. The prompt input can include a freeform prompt input and/or a generated prompt input generated based on one or more tile selections of a user interface. The prompt input can be descriptive of one or more attributes a user is requesting to be rendered in a generated image. The prompt input can include a subject of the image (e.g., an environment and/or one or more objects) and one or more details for the subject (e.g., a color, a style, a material, etc.).).
Therefore, it would have been obvious before the effective filing date of the claimed invention to use a browser history as taught by Sadr. The motivation for doing so would have been to provide the user with greater personalization and customization, with tailored results. Therefore it would have been obvious to combine Sadr with Davidson to obtain the invention.
Regarding claim 11:
The device of claim 1, has all of its limitations taught by Davidson. Davidsonfurther teaches wherein the instructions are executable to:
identify the one or more user preferences based on (Davidson [0026] The application then provides static options for the user to choose from, selectable by check boxes, to add more information such as: [0027] Style 115: anime, abstract, classic, comic, Rockwell, Renoir, etc. In the example shown in FIG. 1, the user has selected “classic”. A user may be able to provide examples of a particular artist (for example by scanning and/and or uploading representative images) to the application to provide the AI with the style that the user wants its image mimic. [0028] Color palate 120: vibrant, subdued, grey scale, black and white, colorable line drawing (e.g., like a coloring book). In the example shown in FIG. 1, the user has selected “vibrant”. This causes a vibrancy scale 125 to appear below, which is selectable with a slidable indicator 130.).
Davidson does not teach social media. In a related field of endeavor, Sadr teaches:
identify the one or more user preferences based on a user’s social media data (Sadr [0095] For example, stored and/or learned information 402 (e.g., fashion knowledge, personalization (e.g., based on stored data associated with a user), and/or trends (e.g., purchase trends, social media trends, and/or search trends)) can be obtained and utilized to generate a prompt and/or to suggest prompt inputs for selection via selectable user interface elements.).
Therefore, it would have been obvious before the effective filing date of the claimed invention to use a browser history as taught by Sadr. The motivation for doing so would have been to provide the user with greater personalization and customization, with tailored results. Therefore it would have been obvious to combine Sadr with Davidson to obtain the invention.
Regarding claim 15:
The method of claim 12, has all of its limitations taught by Davidson. Davidson further teaches comprising:
identifying the one or more user preferences based on user input as received (Davidson [0016] The present disclosure relates to an application, website, or program that allows a person to speak or type one or more image characteristics. The image characteristics can include subject matter, features, or context.).
Davidson does not teach preferences before submitting the prompt. In a related field of endeavor, Sadr teaches:
identifying the one or more user preferences based on user input as received prior to receipt of the prompt (Sadr [0080] For example, the model-generated user-specific terms 212 can be generated by machine-learned models (e.g., LLM) based user personalization data, search engine data, and merchant data. The user personalization data can include explicit personalization that is user generated and/or user controlled. Additionally, the user personalization data can include interest graph data. The search engine data can include fashion knowledge data, recent trends data, and implicit personalization data. The implicit personalization can be based on a plurality of characteristics derived from the user (e.g., time, third-party data, location, intent). The merchant data can include merchant assets data. In some instances, the user personalization data, search engine data, and merchant data can be inputted into a machine-learned model (e.g., large language model (LLM)) to generate user-specific terms for a user. The LLM is a type of machine learning model that can perform a variety of natural language processing NLP tasks, such as generating and classifying terms. The user personalization data can include a prompt input, historical data (e.g., data descriptive of user search history, user purchase history, user browsing history, etc.), profile data, and/or preference data. The prompt input can include a freeform prompt input and/or a generated prompt input generated based on one or more tile selections of a user interface. The prompt input can be descriptive of one or more attributes a user is requesting to be rendered in a generated image. The prompt input can include a subject of the image (e.g., an environment and/or one or more objects) and one or more details for the subject (e.g., a color, a style, a material, etc.).).
Therefore, it would have been obvious before the effective filing date of the claimed invention to use a prior data as taught by Sadr. The motivation for doing so would have been to provide the user with greater personalization and customization, with tailored results. Therefore it would have been obvious to combine Sadr with Davidson to obtain the invention.
Regarding claim 16:
The method of claim 15, has all of its limitations taught by Davidson in view of Sadr. Davidson further teaches wherein the user input relates to an aspect of a geographic area (Davidson [0016] The present disclosure relates to an application, website, or program that allows a person to speak or type one or more image characteristics. The image characteristics can include subject matter, features, or context.)(Davidson [0032] Type of area 215: rural, suburban, urban [0033] Type of street 220: windy, straight, hilly, tree-lined, cul-de-sac [0034] Time of day 225: morning, afternoon, twilight, night [0035] Weather 230: sunny, cloudy, raining, snowing, windy.).
Regarding claim 20:
The at least one CRSM of claim 18, has all of its limitations taught by Davidson. Davidson further teaches wherein the model comprises a large language model, and wherein the output comprises generative text (Davidson [0016] The application, program, or website can then use artificial intelligence (AI) to prompt a user to input additional characteristics of the image based on the initial text description of the image, such as by outputting a series of questions or options on details of the image characteristics.).
Davidson discloses outputting text describe above. However, for the purposes of compact prosecution and for further clarity, in a related field of endeavor, Sadr teaches:
wherein the model comprises a large language model, and wherein the output comprises generative text (Sadr [0042] For example, the system can process the user personalization data and the merchant assets data with a text generation model to generate model-generated terms that improve the search process.)
Therefore, it would have been obvious before the effective filing date of the claimed invention to output text as taught by Sadr. The rationale for doing so would have been that it combines prior art elements according to known methods to yield predictable results, where Davidson take a prompt and preferences to generate text, where the text is then used for image generation and Sadr takes preferences and generates text, and the text can be modified by the user, and optionally an image is generated, where the end result of both is image generation with generative text outputs before the image generation. Therefore it would have been obvious to combine Sadr with Davidson to obtain the invention.
Conclusion
For the prior art referenced and the prior art considered pertinent to Applicant’s disclosure but not relied upon, see PTO-892 “Notice of References Cited”.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JASON PRINGLE-PARKER whose telephone number is (571) 272-5690 and e-mail is jason.pringle-parker@uspto.gov. The examiner can normally be reached on 8:30am-5:00pm est Monday-Friday. If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, King Poon can be reached on (571) 270-0728. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, seehttp://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JASON A PRINGLE-PARKER/
Primary Examiner, Art Unit 2617