DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s arguments with respect to claim(s) 1-20 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
Claim(s) 1,4,8,10,11, 17 and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. (US 2022/0108417 A1), in view of Socher et al. (US 2024/0020538 A1).
Regarding claim 1, Liu discloses one or more processors comprising: one or more circuits to (Liu, fig. 5, [0076] Content Server 520, includes multiple processors such as central processing units (CPUs) and graphics processing units (GPUs):
identify text corresponding to an image in an electronic document (Liu, figs 3-5, [0050-0051;0060], text data can be provided as input to one or more feature extractors 206. The extractor can include one or more neural network or transformers that analysis the received text data and identify features in the text data and generate corresponding images based on the identified features in the received text data input. The network can receive text input such as “create an image of a lake in a forest with mountains in back,” and generate an associated image in a presentation/document);
store a representation of the text (image generated based on received text input) in association with an identifier (label of the image (Liu [0053;57;0060-0061;0076] when text data is provided to an extractor 206 as an input, the extractor, extracts relevant features from the identified text to generate corresponding images of the text data. once an acceptable image is generated, a user can cause that image to be saved in a content database 534, exported, or otherwise utilized for its intended purpose. Each image/object or regions of the generated images will be associated with a label or another type of identifier. The system may present a user interface which provides label options that enable a user to select or specify a label for a specific region. A user can select such a label before creating a new region or choose a label after selecting a created region, among other such options);
receive an input prompt for a machine-learning model (Liu [0067] a text-to-image model receives text as an input and the model is trained to generate a corresponding image/object of the identified text data); and
generate a response to the input prompt using the machine-learning model (Liu [0067] a text-to-image model receives text as an input and the model is trained to generate a corresponding image/object of the identified text data).
While Liu discloses the capability of generating by machine-learning model an image representative of a text data provided as an input to the machine-learning model.
Liu did not explicitly disclose the response comprising (i) text data generated using the machine-learning model and (ii) the image responsive to identifying the representation of the text using a searching function and text data generated using the machine-learning model.
Socher discloses the response (output from a chatbot in response to a user request) comprising (i) text data generated using the machine-learning model (Large Language Model (LLM) based chat agent) (Socher, figs. 10A-10E [0136;0148], a generative AI system provides a text generation tool implemented as a search assistance tool, such as an artificial intelligence (AI) chatbot that conduct a conversation with a user during a search and provide a summary of search results in response to user interested search topics. A user may provide input to Large Language Model (LLM) based chat agent asking/requesting for a picture of an animal, such as a cat, and the search assistance tool may produce the image directly in the output alongside text. In [0039-0040;0043-0044] provide example operations of a generation server. During a first operation, a music clip that a user recorded at Times Square, New York may be uploaded (provided as an input) to the generation server, which may in turn identify the music clip belongs to Broadway musical Phantom of the Opera, and may generate an output (with both text and/or images relating to the musical) containing a short description of the musical. Using a vison-language model and/or other multi-modal model employed at a text generation server 110, the text generation server 110 may generate a text captioning of an image retrieved from a webpage following a search result link); and
(ii) the image (image/picture of a cat) responsive to identifying the representation of the text (responsive to a text input requesting for a picture of a cat) using a searching function and text data generated using the machine-learning model (Socher, figs. 10A-10E [0136;0148], a user may provide input/text request to a Large Language Model (LLM) based chat agent asking/requesting for a picture of an animal, such as a cat, and the search assistance tool may produce the image of a cat directly in the output alongside text. [0043-0044] a text generation server 110 may receive input 122 that contains a request “please write a paragraph about the history of direct current vs. alternate current,” in addition to generate a text summary based on various search results, the NLP models 115 may further insert web images of Thomas Edison and Nikola Tesla into the output 125 for illustration).
One of ordinary skill in the art would have been motivated to combine Liu and Socher because these teachings are from the same field of endeavor with respect to disclosing techniques for the use of Large Language Model (LLM) based chat agent in responding to user request.
Therefore, before the effective filing date of the invention, it would have been obvious to a person of ordinary skill in the art to incorporate the strategies by Socher into the invention of Liu. The motivation would have been to train a machine learning model such that the model correctly utilize the input information from both a user and search results to provide tailored output that is relevant and timely for the user, thus, enhancing the user experience, Socher, [Abstract; 0148].
Regarding claim 4, Liu modified by Socher disclose the one or more processors of claim 1, wherein the one or more circuits are to: generate the representation of the text by providing the text as input to an embeddings model (Text-to-image model) (Liu [0067] discloses text-to-image models that can receive text data as an input to the model and generates an image output).
The motivation to combine is similar to that of claim 1.
Regarding claim 8, Liu modified by Socher disclose the one or more processors of claim 1, wherein the one or more circuits are to: present the output of the machine-learning model with the image via a graphical user interface responsive to the input prompt (Socher [0043-0044] a text generation server 110 may receive input 122 that contains a request “please write a paragraph about the history of direct current vs. alternate current,” in addition to generate a text summary based on various search results, the NLP models 115 may further insert web images of Thomas Edison and Nikola Tesla into the output 125 for illustration. [0080] Outputs 125/result 430 of a user’s request is transmitted to the user device for displaying via a graphical user interface or some other type of user output device).
The motivation to combine is similar to that of claim 1.
Regarding claim 10, Liu modified by Socher disclose the one or more processors of claim 1, wherein the one or more processors are comprised in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing conversational AI operations; a system for performing generative AI operations using a multi-modal language model; a system for performing generative AI operations using a large language model (LLM); a system for performing generative AI operations using a video language model (VLM); a system for generating synthetic data; a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources (Liu, fig. 21, [0251] deep learning application processor 2100 uses instructions that, if executed by deep learning application processor 2100, cause deep learning application processor 2100 to perform some or all of processes and techniques for generating images using one or more neural networks. In figs. 6A and/or 6B, deep learning application processor 2100 is used to train a machine learning model, such as a neural network, to predict or infer information provided to deep learning application processor 2100. Deep learning application processor 2100 is used to infer or predict information based on a trained machine learning model (e.g., neural network) that has been trained by another processor or system or by deep learning application processor 2100).
The motivation to combine is similar to that of claim 1.
Regarding claim 11, Liu discloses a system, comprising: one or more processors to: (Liu, , figs. 5 & 8, [0076;0100] Content Server 520, includes multiple processors such as central processing units (CPUs) and graphics processing units (GPUs). A computer system, which may be a system with interconnected devices and components, the computer system includes a processor capable of executing instructions to):
receive an input prompt for a machine-learning model (Liu [0067] a text-to-image model receives text as an input and the model is trained to generate a corresponding image/object of the identified text data);
generate a response message using the input prompt and the machine-learning model (Liu [0067] a text-to-image model receives text as an input and the model is trained to generate a response message in the form of a corresponding image/object of the identified text data);
Liu did not explicitly disclose the response message comprising text data; identify encoded text data using a searching function and the text data of the response message, the encoded text data stored in association with an identifier of an image; and provide the response message and the image for display in response to the input prompt.
Socher discloses the response message comprising text data (Socher, figs. 10A-10E [0039-0040;0043-0044; 0136;0148], a generative AI system provides a text generation tool implemented as a search assistance tool, such as an artificial intelligence (AI) chatbot that conduct a conversation with a user during a search and provide a summary of search results in response to user interested search topics. A user may provide input to Large Language Model (LLM) based chat agent asking/requesting for a picture of an animal, such as a cat, and the search assistance tool may produce the image directly in the output alongside text. Using a vison-language model and/or other multi-modal model employed at a text generation server 110, the text generation server 110 may generate a text captioning of an image retrieved from a webpage following a search result link);
identify encoded text data using a searching function and the text data of the response message (Socher [0039-0040;0043-0044] provide example operations of a generation server. During a first operation, a music clip that a user recorded at Times Square, New York may be uploaded (provided as an input) to the generation server, which may in turn identify the music clip belongs to Broadway musical Phantom of the Opera, and may generate an output (with both text and/or images relating to the musical) containing a short description of the musical, and/or available schedule and tickets for the musical and a link to purchase. Using a vison-language model and/or other multi-modal model employed at a text generation server 110, the text generation server 110 may generate a text captioning of an image retrieved from a webpage following a search result link);
the encoded text data stored in association with an identifier (Topic) of an image (Socher,[0021;0093;0136;0145] user device 310 may include database 318 may store user profile relating to the user 340, predictions previously viewed or saved by the user 340, historical data received from the server 330, and/or the like. Historical data saved in the database 318 may be saved based of different topics. An artificial intelligence (AI) chatbot may conduct a conversation with a user during a search and provide a summary of search results in response to user interested search topics. The search assistance tool may determine to initiate a completely new search when the user input switches to a new topic); and
provide the response message and the image for display (fig. 5, 500) in response to the input prompt (Socher [0043-0044] a text generation server 110 may receive input 122 that contains a request “please write a paragraph about the history of direct current vs. alternate current,” in addition to generate a text summary based on various search results, the NLP models 115 may further insert web images of Thomas Edison and Nikola Tesla into the output 125 for illustration. [0080] Outputs 125/result 430 of a user’s request is transmitted to the user device for displaying via a graphical user interface or some other type of user output device).
The motivation to combine is similar to that of claim 1.
Regarding claim 16, the claim is rejected with rational similar to that of claim 10.
Regarding claim(s) 17and 20 the claim(s) are rejected with rational similar to that of claim(s) 1and 4.
Claim(s) 2,3,5-7,9,12-15 and 18-19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. (US 2022/0108417 A1), in view of Socher et al. (US 2024/0020538 A1)., further in view of Ramesh et al. (US 2018/0314715 A1).
Regarding claim 2, Liu modified by Socher disclose the one or more processors of claim 1, but did not explicitly disclose wherein the one or more circuits are to: identify the text corresponding to the image by extracting the text proximate to the image in the electronic document.
Ramesh discloses wherein the one or more circuits are to: identify the text corresponding to the image by extracting the text proximate to the image in the electronic document. (Ramesh [0032] A Text Extraction Logic 130 may be configured to identify text within a website that specifically refers to the image, and/or text disposed proximate to the image or proximate to text that refers to the image. Also, Text Extraction Logic 130 is configured to identify text that refers to the image and then extract an entire paragraph including that text, or 1-5 sentences adjacent to the reference).
One of ordinary skill in the art would have been motivated to combine Liu, Socher and Ramesh because these teachings are from the same field of endeavor with respect to disclosing techniques for the use of images and associated text to train a neural network, or other machine learning systems.
Therefore, before the effective filing date of the invention, it would have been obvious to a person of ordinary skill in the art to incorporate the strategies by Ramesh into the invention of Liu and Socher. The motivation would have been to train a model to evolve a neural network to generate attribute vectors and/or feature vectors that better match those of an associated image, Ramesh, [0033].
Regarding claim 3, Liu, Socher and Ramesh disclose the one or more processors of claim 2, wherein the one or more circuits are to: identify the text corresponding to the image by extracting a predetermined portion of the text proximate to the image in the electronic document (Ramesh [0032] an image management system includes a Text Extraction Logic 130 configured to extract text from multimedia content found to include images identified and/or tracked using Tracking Logic 125. If an image is found on a specific blog or website, the location of the is predetermined and Text Extraction Logic 130 130 may be configured to identify text within a website that specifically refers to the image, and/or text disposed proximate to the image or proximate to text that refers to the image).
The motivation to combine is similar to that of claim 2.
Regarding claim 5, Liu modified by Socher disclose the one or more processors of claim 1, but did not explicitly disclose wherein the one or more circuits are to: store the representation of the text in a vector database.
Ramesh wherein the one or more circuits are to: store the representation of the text in a vector database (Ramesh: storing keywords in association with images in image library 110); and store the image in an image database (fig. 1 - Image Library 110), wherein the image is identified in the image database by the identifier (Image tags) of the image (Ramesh [0024;0035;0039] discloses an image library 110 that optionally stores in association with attribute vectors, image feature vectors, keywords, and/or the like. An Image Management System 100 optionally includes an Image Tagging System 140 configured to associate image tags with images stored within the image library 110. These image tags can include keywords, attributed vectors and/or feature vectors, and are optionally used in the search for images within Image Library 110).
The motivation to combine is similar to that of claim 2.
Regarding claim 6, Liu modified by Socher disclose the one or more processors of claim 1, but did not explicitly disclose wherein the one or more circuits are to: identify a plurality of images using the searching function and the text data generated using the machine- learning model; and select at least one of the plurality of images for inclusion in the response based at least on an image selection parameter.
Ramesh discloses wherein the one or more circuits are to: identify a plurality of images using the searching function and the text data generated using the machine- learning model; and select at least one of the plurality of images for inclusion in the response based at least on an image selection parameter (Ramesh [0038] FIG. 2 discloses an Image Selection System 200, configured for selecting an image from a library of images, such as Image Library 110. The selection is based on received text used to generate an output of a neural network. Optionally, the selection is further based received keywords. For example, keywords may be used to first select an initial set of images from Image Library 110 and then a subset of this initial set may be selected using a greater amount of text and the neural network. The neural network is optionally trained using Image Management System 100).
The motivation to combine is similar to that of claim 2.
Regarding claim 7, Liu, Socher and Ramesh disclose the one or more processors of claim 6, wherein the one or more circuits are to: receive the image selection parameter with the input prompt for the machine-learning model (Ramesh [0033-0035] Image Management System 100 optionally includes an Image Tagging System 140 configured to associate image tags with images within the image library. These image tags can include keywords, attributed vectors and/or feature vectors, and are optionally used in the search for images within Image Library 110 as described elsewhere herein. The tags are provided to a machine learning model as input to the model to train the model so it may evolve the neural network to generate attribute vectors and/or feature vectors that better match those of an associated image).
The motivation to combine is similar to that of claim 2.
Regarding claim 9, Liu modified by Socher disclose the one or more processors of claim 1, but did not explicitly disclose wherein the searching function comprises a vector similarity searching function.
Ramesh discloses (Interface Logic 210) wherein the searching function comprises a vector similarity searching function (Ramesh [0041-0042] an Interface Logic 210 may have a text field to receive a full paragraph and text fields to receive 1-5 keywords, such as “Fog,” “Harbor” and “Night.” The keywords “Fog,” “Harbor” and “Night” may be used to select an initial set of images being associated with similar image tags, the full paragraph may then be searched using a similarity function to select images from this initial set using a neural network trained using Image Management System 100).
The motivation to combine is similar to that of claim 2.
Regarding claim 12, Liu modified by Socher disclose the system of claim 11, but did not explicitly disclose wherein the encoded text data comprises embeddings data, and wherein the searching function is a vector search function.
Ramesh discloses wherein the encoded text data comprises embeddings data, and wherein the searching function is a vector search function (Ramesh [0012] discloses a content based image management and selection system configured to observe how images are used by third parties and to train a machine learning system to better search for and select images based on these observations. Once the machine learning system is trained, a sample of embedded text from multimedia content can be used to search for images likely to be used with that text. This search is optionally also based on one or more keyword vectors. The search for images can be based on significant sections of text, e.g., entire sentences, paragraphs or more. This often produces search results that better match a subject matter of the text, relative to results based on a simple keyword search).
The motivation to combine is similar to that of claim 2.
Regarding claim 13, Liu, Socher and Ramesh disclose the system of claim 12, wherein the one or more processors are to: identify a set of search results including the encoded text data (text data/words/attributes resulting from a search text input); and select the encoded text data (selecting text similar to the input text and corresponding image) based at least on a similarity between the encoded text data and the text data response message (Ramesh [0013;0025-0026] automated image selection system is configured to analyze text input and output a set of text data and corresponding images. The system selects one or more images for publication in mixed media content that includes both the text and at least one of the selected images. The selection is based on processing of the similarity between the input text, the output text and the corresponding image. The automated image selection system optionally includes an image tagging system).
The motivation to combine is similar to that of claim 12.
Regarding claim 14, Liu, Socher and Ramesh disclose the system of claim 11, wherein the one or more processors are to: extract document text data from an electronic document (website), the document text data proximate to the image (Ramesh [0032] discloses a Text Extraction Logic 130 configured to identify text within a website that specifically refers to the image, and/or text disposed proximate to the image or proximate to text that refers to the image. Text Extraction Logic 130 is configured to identify text that refers to the image and then extract an entire paragraph including that text, or 1-5 sentences adjacent to the reference);
encode the document text data to generate the encoded text data; and store the identifier (image tag) of the image in association with the encoded text data in a database (Image library 110) (Ramesh [0035] an Image Management System 100 optionally includes an Image Tagging System 140 configured to associate image tags with images generated from text received as input to the management system and stored within the image library. These image tags can include keywords, attributed vectors and/or feature vectors, and are optionally used in the search for images within Image Library 110).
The motivation to combine is similar to that of claim 12.
Regarding claim 15, Liu, Socher and Ramesh disclose the system of claim 14, wherein the one or more processors are to: encode the document text data using an embeddings model corresponding to the machine-learning model (Liu [0067] Using text-to-image models, text is inputted into the model and using an embedded model, the input data is text and is converted to a corresponding image).
The motivation to combine is similar to that of claim 12.
Regarding claim(s) 18-19 the claim(s) are rejected with rational similar to that of claim(s) 2-3.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. The following publications show the state of the art related to displaying images in chatbot responses.
Yuen et al. (US 2017/0103072 A1)
Hunter et al. (US 2021/0089570 A1)
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DIXON F DABIPI whose telephone number is (571)270-3673. The examiner can normally be reached on Monday - Friday from 9:00 am – 5:00 pm.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Christopher L Parry, can be reached at telephone number 571-272-8328. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from Patent Center. Status information for published applications may be obtained from Patent Center. Status information for unpublished applications is available through Patent Center to authorized users only. Should you have questions about access to the USPTO patent electronic filing system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free).
Examiner interviews are available via a variety of formats. See MPEP § 713.01. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) Form at https://www.uspto.gov/InterviewPractice.
/D.F.D/Examiner, Art Unit 2451
/Chris Parry/Supervisory Patent Examiner, Art Unit 2451