DETAILED ACTION
1. This action is responsive to Application no.19/003,766 filed 12/27/2024. All claims have been examined and are currently pending.
Notice of Pre-AIA or AIA Status
2. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
3. The information disclosure statement (IDS) submitted is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 102
4. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
5. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
6. Claims 1, 6-7, 10-11, 16-20 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Trzyna (2024/0338860).
Regarding claim 1 Trzyna (2024/0338860) teaches A method for generating a multimodal text (0003 systems and methods for providing live image generation; summarization; image), comprising:
generating, by a large language model, a text information corresponding to a prompt information based on the prompt information, in response to a multimodal text generation request comprising the prompt information being received (0003: a segment of the live text transcript is extracted and included in a first-language model (LM) prompt; the first LM prompt includes a request for summarization of the transcript segment; the first LM prompt is provided to a large language model (LLM), and a summarization is received in response);
generating, by the large language model, an image information corresponding to the text information based on the text information (0003 a second LM prompt is generated including the summarization and a request for an image of the summarization); and
calling, by the large language model, a multimodal text rendering tool based on the text information and the image information to render the multimodal text comprising the text information and the image information (0003 summarization; the image is displayed on a display screen; 12 displaying images while presenting information ; 28 live text transcript;
figures 2, 3C-3G, 4;
0017-0018;
0022: a single multi-modal generative AI model that is capable of generating and/or processing multiple forms of inputs and outputs, such as text, images, and or audio. In such examples where a multimodal generative AI model is used, the first LM prompt and the second LM prompt may be effectively combined. For instance, a multi-modal prompt may be generated that includes a set of instructions to generate an image based on a summarization of the segment of the text transcript or even a segment of the audio itself.).
Regarding claim 6 Trzyna teaches The method according to claim 1, wherein the calling, by the large language model, a multimodal text rendering tool based on the text information and the image information to render the multimodal text comprising the text information and the image information comprises:
performing, by the large language model, a layout generation task based on the text information and the image information to generate a layout information for the multimodal text (figures 3C-3G; 0031: the image-generation priming instructions may include instructions such as, “generate a beautiful picture of:”, “generate a sad picture of:”, “generate a black-and-white line drawing of:”, “generate a cartoon drawing of:”, etc., followed by the summarization 230 generated by the LLM 108.); and
calling, by the large language model, the multimodal text rendering tool based on the layout information to render the multimodal text (figures 3C-3G; 0018: image generator;
0033: For instance, if the summarization 230 includes language about a “cute white fluffy poodle” and the image-generation priming instructions include instructions to generate a picture in a sad tone, the image-generation AI model 118 may generate an image 222 of a poodle with a sad expression or with a gloomy background.).
Regarding claim 7 Trzyna teaches. The method according to claim 6, further comprising:
determining a background image information for the multimodal text based on the image information (0033),
wherein the calling, by the large language model, the multimodal text rendering tool based on the layout information to render the multimodal text comprises:
calling, by the large language model, the multimodal text rendering tool based on the layout information and the background image information to render the multimodal text (0033: For instance, if the summarization 230 includes language about a “cute white fluffy poodle” and the image-generation priming instructions include instructions to generate a picture in a sad tone, the image-generation AI model 118 may generate an image 222 of a poodle with a sad expression or with a gloomy background.).
Regarding claim 10 Trzyna teaches A method for acquiring a multimodal text, comprising:
transmitting, in response to a prompt information being received, a multimodal text generation request comprising the prompt information (0003: a segment of the live text transcript is extracted and included in a first-language model (LM) prompt; the first LM prompt includes a request for summarization of the transcript segment; the first LM prompt is provided to a large language model (LLM), and a summarization is received in response; 0035); and
presenting the multimodal text, in response to acquiring the multimodal text generated in response to the multimodal text generation request (0003 summarization; image),
wherein the multimodal text is generated by using the method according to claim 1.
rejected for similar rationale and reasoning as claim 1
Regarding claim 11 Trzyna teaches An electronic device (fig 5; 0014), comprising:
at least one processor (fig 5); and
a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions, when executed by the at least one processor, are configured to cause the at least one processor (fig 5; 0014) to at least:
generate, by a large language model, a text information corresponding to a prompt information based on the prompt information, in response to a multimodal text generation request comprising the prompt information being received;
generate, by the large language model, an image information corresponding to the text information based on the text information; and
call, by the large language model, a multimodal text rendering tool based on the text information and the image information to render the multimodal text comprising the text information and the image information.
Claim recites limitations similar to claim 1 and is rejected for similar rationale and reasoning
Claims 16-17 recite limitations similar to claims 6-7 and are rejected for similar rationale and reasoning
Regarding claim 18 Trzyna teaches An electronic device (fig 5; 14), comprising:
at least one processor (fig 5); and
a memory communicatively connected to the at least one processor (fig 5), wherein the memory stores instructions executable by the at least one processor, and the instructions, when executed by the at least one processor, are configured to cause the at least one processor to implement the method of claim 10.
rejected for similar rationale and reasoning as claim 1/10
Regarding claim 19 Trzyna teaches A non-transitory computer-readable storage medium having computer instructions stored therein, wherein the computer instructions are configured to cause a computer to at least:
generate, by a large language model, a text information corresponding to a prompt information based on the prompt information, in response to a multimodal text generation request comprising the prompt information being received;
generate, by the large language model, an image information corresponding to the text information based on the text information; and
call, by the large language model, a multimodal text rendering tool based on the text information and the image information to render the multimodal text comprising the text information and the image information.
Claim recites limitations similar to claim 1/11 and is rejected for similar rationale and reasoning
Regarding claim 20 Trzyna teaches A non-transitory computer-readable storage medium having computer instructions stored therein, wherein the computer instructions are configured to cause a computer to implement the method of claim 10.
Claim recites limitations similar to claim 1/10/18 and is rejected for similar rationale and reasoning
Claim Rejections - 35 USC § 103
7. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
8. Claims 2-5, 12-15 are rejected under 35 U.S.C. 103 as being unpatentable over Trzyna in view of Deutsch et al (12,051,205).
Regarding claim 2 Trzyna does not specifically teach where Deutsch et al (12,051,205) teaches The method according to claim 1, wherein the generating, by a large language model, a text information corresponding to a prompt information based on the prompt information comprises:
processing, by the large language model, the prompt information to generate a first decision information, wherein the first decision information comprises a first indication information indicating whether to search a first database, and in a case that the first indication information indicates to search the first database, the first decision information further comprises a search statement (Col 2 l. 20-40; l. 31-37: The method may further include generating an output at least in part by applying the input data to the multimodal machine learning model, the multimodal machine learning model configured using prompt engineering to identify a location in the image conditioned on the image and the textual prompt, wherein the output includes a first location indication.;
col 8 l. 27-30: Prompt engineering refers to an AI engineering technique implemented in order to optimize machine learning models such as multimodal LLMs for particular tasks and outputs.;
col 8 l. 37-49: Prompt engineering may involve different techniques for text-to-text, text-to-image, and non-text prompts. Text-to-text techniques may include Chain-of-thought, Generated knowledge prompting, Least-to-most prompting, Self-consistency decoding, Complexity-based prompting, Self-refine, Tree-of-thought, Maieutic prompting, and Directional-stimulus prompting. Text-to-text techniques may also include automated generation such as Retrieval-augmented generation. Text-to-image techniques may include Prompt formats, Artist styles, and Negative prompts. And non-text prompts may include Textual inversion and embeddings, Image prompting, and Using gradient descent to search for prompts.); and
performing, by the large language model, a retrieval-augmented generation task based on the search statement to generate the text information, in response to the first indication information indicating to search the first database (col 8 l 45: retrieval-augmented generation).
It would have been obvious to one of ordinary skill in the art before the effective filing date to incorporate Deutsch to allow for the best prompt engineering technique to be used for improved and more efficient processing and results. Trzyna already teaches prompts for executing LLM text and image generation, and one could look to Deutsch to perform specific LLM prompt engineering technique to allow the model to output the desired response (Deutsch col 8 l. 32), while still allowing to enhance the impact of spoken content and provide visual indicators and feedback to the speakers (Trzyna 0012).
Regarding claim 3 Trzyna teaches The method according to claim 2, wherein the generating, by a large language model, a text information corresponding to a prompt information based on the prompt information further comprises:
performing, by the large language model, a text generation task based on the prompt information to generate the text information, in response to the first indication information indicating not to search the first database (0003; 0011; 0017;
0028: the prompt generator 204 generates summarization instructions for the first LM prompt 215 corresponding to summarizing the segment of the live text transcript 210. For instance, the summarization instructions may include directives to the LLM 108, such as “summarize the following:”. In some examples, the summarization instructions include text length instructions (e.g., “limit the summarization to N words or less”, where N is a predetermined number). In some examples, the summarization instructions include element-of-interest instructions corresponding to detecting a particular context or theme of the live text transcript segment and including text corresponding to the context or theme in the summarization 230. For instance, the instructions may cause the first LM to prioritize certain words, attributes, themes, or another feature of the live text transcript segment. In one example, the summarization instructions include instructions to focus on content in the live text transcript segment that can be represented visually through an image;
0030: the LLM 108 analyzes the first LM prompt 215 to generate a relevant response including a summarization 230 of the live text transcript segment. In examples, the LLM 108 uses information included in the summarization instructions to understand the intent and context of the live text transcript segment. According to examples, the term “context” is used to describe information that can influence an interpretation and execution of the request to generate a summarization 230. For instance, if the live text transcript segment includes language about a specific topic, the LLM 108 can generate a summarization 230 that includes the topic.).
Regarding claim 4 Trzyna does not specifically teach where Deutsch teaches The method according to claim 1, wherein the generating, by the large language model, an image information corresponding to the text information based on the text information comprises:
processing, by the large language model, an input information to generate a second decision information, wherein the input information is obtained based on the prompt information and the text information, the second decision information comprises a second indication information indicating whether to search a second database (Deutsch Col 2 l. 20-40; l. 31-37; col 8 l. 27-30; col 8 l. 37-49); and
in a case that the second indication information indicates to search the second database, the second decision information further comprises a search parameter (Deutsch Col 2 l. 20-40; l. 31-37; col 8 l. 27-30; col 8 l. 37-49); and
performing, by the large language model, an image search task based on the search parameter to obtain the image information, in response to the second indication information indicating to search the second database (Deutsch
Col 2 l. 20-40; l. 31-37:;
col 8 l. 27-30: Prompt engineering refers to an AI engineering technique implemented in order to optimize machine learning models such as multimodal LLMs for particular tasks and outputs.;
col 8 l. 37-49:).
Rejected for similar rationale and reasoning as claim 2, where it would be obvious to also incorporate the prompt engineering techniques for image generation for improved and more efficient processing and results while still presenting a reasonable expectation of success.
Regarding claim 5 Trzyna teaches The method according to claim 4, wherein in a case that the second indication information indicates not to search the second database, the second decision information further comprises an image description statement (0031: image generation priming instructions), and
wherein the generating, by the large language model, an image information corresponding to the text information based on the text information further comprises:
performing, by a text-to-image model, an image generation task based on the image description statement to generate the image information, in response to the second indication information indicating not to search the second database
(0003: A second LM prompt is generated including the summarization and a request for an image of the summarization. The second LM prompt is provided to a text-to-image model, and an image is received in response; 0011;
0018: the second LM prompt includes image-generation priming instructions and the summarization generated by the LLM 108. In an example implementation, the second LM is a text-to-image model, herein referred to as an image-generation AI model 118. For example, the image-generation AI model 118 may be an LM based on a transformer architecture that is trained to generate images based on textual descriptions, such as the DALL-E model from OpenAI. According to an example, the image-generation AI model 118 uses a combination of natural language processing and computer vision to generate images from textual descriptions. For instance, the image-generation AI model 118 is trained on a large dataset of image-caption pairs and can generate a wide range of images based on textual input. Images, for example can include objects, scenes, anthropomorphic creatures, etc.;
0031: where the image-generation priming instructions include one or more priming words that describe a specific context, theme, mood, desired image type (e.g., photo, drawing, painting, a stained glass image), etc., that causes the image-generation AI model 118 to generate an image 222 representative of the summarization 230 generated by the first LM and further representative of the one or more priming words.).
Claims 12-15 recite limitations similar to claims 2-5 and are rejected for similar rationale and reasoning
9. Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Trzyna in view of Zhou et al (2024/03692093)
Regarding claim 8 Trzyna does not specifically teach where Zhou et al (2024/03692093) teaches The method according to claim 1, wherein an information generated by the large language model comprises a generation information corresponding to a task performed by the large language model and a confidence level of the generation information, and the generation information comprises at least one of the text information, the image information, or the multimodal text; and
wherein the method further comprises:
re-performing, by the large language model, a task of generating the generation information, in response to the confidence level of the generation information being less than a confidence level threshold ([0032] The response confidence engine 140 may determine a confidence score for an NL based response generated by the LLM response generation engine 136. If the confidence score is above a threshold score, the NL based response is provided to the client device for output to the user. If the confidence score is below the threshold score, the NL based response system may cause the client device to request further information and/or clarifications.).
It would have been obvious to one of ordinary skill in the art before the effective filing date to incorporate Zhou to incorporate LLM confidence for improved optimization. Trzyna already teaches performing LM processing to generate text and images, and one could look to Zhou to further incorporate confidence scores to ensure the generated components are of the highest quality.
10. Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Trzyna in view of Flockhart et al (6,463,346).
Regarding claim 9 Trzyna teaches The method according to claim 1, further comprising:
generating an information flow for the large language model in a process of generating the multimodal text, the information flow indicating an input information of the large language model and an output information of the large language model (fig 2; 4; 0003; 0017-18);
but does not specifically teach where Flockhart teaches
presenting the information flow (abstract: workflow; col 1 l. 55); and
determining the information flow as a target information flow, in response to a selection operation on the information flow (abstract: The flow of work items (40) through a workflow process (50) is optimized by repeatedly reordering (FIG. 3) work items enqueued in inbox queues (21) of workflow process tasks (500) to maximize results according to a given business strategy expressed through target times; col 1l. 55-58), {wherein the large language model is fine-tuned with the target information flow}.
It would have been obvious to one of ordinary skill in the art before the effective filing date to incorporate Flockhart for improved processing. Trzyna already teaches LLM for image and text generation with a particular workflow, and one could look to Flockhart to further determine a target workflow/information flow to optimize the processing and ultimately the text and image generation;
The incorporation Allowing to teach
wherein the large language model is fine-tuned with the target information flow.
Conclusion
11. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: See PTO-892.
Jain et al (12,106,205)
Col 19 l 30-46: Based on providing the modified inputs (e.g., prompts) to a suitable LLM, the data generation platform 102, through the generative model engine 120, can generate an output. For example, the data generation platform 102 generates a response to a query posed within the prompt of the output generation request. To illustrate, the output can include generated natural language (e.g., in the form of alphanumeric strings of characters), code (e.g., portions of code, such as code samples), or other generated outputs. The output can include one or more images, videos, audio, and/or combinations thereof. For example, a model can output a combination of an image, text, and/or a video (e.g., multi-modal outputs). In some implementations, the LLM generates audio data (e.g., corresponding to speech), videos, or images based on the input prompt. As such, the data generation platform 102 can include flexible, modular generative machine learning models for a variety of applications.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHAUN A ROBERTS whose telephone number is (571)270-7541. The examiner can normally be reached Monday-Friday 9-5 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Flanders can be reached on 571-272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov.
For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SHAUN ROBERTS/Primary Examiner, Art Unit 2655