Prosecution Insights
Last updated: August 17, 2026
Application No. 19/037,333

Artificial Intelligence Agent Access Multi-Modal Input In A Database System

Non-Final OA §102
Filed
Jan 27, 2025
Priority
Sep 13, 2024 — provisional 63/694,675
Examiner
VY, HUNG T
Art Unit
2163
Tech Center
2100 — Computer Architecture & Software
Assignee
Salesforce Inc.
OA Round
1 (Non-Final)
86%
Grant Probability
Favorable
1-2
OA Rounds
1y 1m
Est. Remaining
88%
With Interview

Examiner Intelligence

Grants 86% — above average
86%
Career Allowance Rate
795 granted / 922 resolved
+31.2% vs TC avg
Minimal +2% lift
Without
With
+2.1%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
15 currently pending
Career history
941
Total Applications
across all art units

Statute-Specific Performance

§101
11.4%
-28.6% vs TC avg
§103
39.7%
-0.3% vs TC avg
§102
25.4%
-14.6% vs TC avg
§112
6.1%
-33.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 922 resolved cases

Office Action

§102
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claim(s) 1-20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Thompson, III, (U.S. Pub. 2025/0371318 A1 ). With respect to claims 1, 15 and 20, Thompson, III discloses a method comprising: receiving input from a client machine at a computing services environment via a communication channel (i.e., “A round of dialog as used herein may refer to a user input and an associated system-generated response, e.g., a reply to the user input that is generated at least in part via a generative artificial intelligence model. Any dialog or thread can include one or more different types of digital content, including natural language text, audio, video, digital imagery, hyperlinks, and/or multimodal content such as web pages.”(0064)), the input including natural language input text and non-text media having an associated media type (i.e., “A round of dialog as used herein may refer to a user input and an associated system-generated response, e.g., a reply to the user input that is generated at least in part via a generative artificial intelligence model. Any dialog or thread can include one or more different types of digital content, including natural language text, audio, video, digital imagery, hyperlinks, and/or multimodal content such as web pages.”(0064), 0059),; autonomously selecting, by an autonomous agent instantiated by an agent service within the computing services environment (i.e., “ responsive to receiving input via one or more components of the environment 101, the automated agent 102 can be dynamically configured or reconfigured to perform a task or a series of tasks, via one or more components of the distributed multi-agent system 105.”(0066)), a non-text media summarization model accessible via the computing services environment, the non-text media summarization model having a model type corresponding to the associated media type (i.e., “A generative artificial intelligence (GAI) model or generative model uses artificial intelligence technology, e.g., machine learning models, e.g., neural networks, to machine-generate digital content based on model inputs and the previously existing data with which the model has been trained. A generative language model is a particular type of GAI model that is capable of generating content in response to model input. ”(0037). 0525 and fig. 7A at step 710); determining a non-text media textual summary describing the non-text media in natural language by applying the non-text media summarization model to the non-text media (i.e., “Portions of the task description can be in the form of natural language text, such as a question or a statement. Alternatively or in addition, a task description or prompt can include non-text forms of content, such as digital imagery and/or digital audio.”(0037) and “Language models, including large language models and other generative models, can be implemented using transformer models. A generative model can be constructed using a neural network-based machine learning model architecture…. the one or more embedding are generated based on the task description by a pre-processor, the embeddings are input to the generative language model, and the generative language model outputs digital content, e.g., natural language text or a combination of natural language text and non-text output, based on the embeddings.’(0581)); retrieving one or more database records from a database system within the computing services environment based on the natural language input text and the non-text media textual summary (i.e., “embedding-based retrieval can be used to match vector representations of digital content stored in a vector database with a vector representation of a query, question, or request. With in-context learning, the retrieved content is used as input to an LM or LLM, which generates a response to the input including the RAG content. ”(0056) or “Actions (e.g., tools or skills) can be performed, for example, by a planner agent explicitly or through an LLM selection. For instance, a tool can be configured to, given some context (e.g., task to be performed), select and invoke the most appropriate action, including argument population and self-repairing retries.”(0165) or “ memory is queried across all layers and then the most relevant portions of the retrieved information are selected and extracted from the query results and mapped to their respective arguments in the action call. An LLM or other type of model or tool can be used to map the relevant portions of the retrieved information to their respective arguments. For example, the information retrieved by querying can be included in an LLM prompt (e.g., a micro-prompt) that has been specially configured to cause an LLM to perform the specific task of function calling, and then the prompt is passed to the LLM which maps the information to arguments and generates the function call.”(0283), “ if input from the environment 201 satisfies a condition or threshold criteria, which can be stored in and retrieved from the multi-layer memory 218, the start automated agent process 202 obtains or creates an agent identifier (e.g., the agent ID can be or include the session ID). The process 202 passes the agent ID to a create instance of automated agent process 208 via flow 205 and also passes the agent ID to messaging service 210 via flow 206.” (0117), 0096); determining a response message including novel natural language text by completing an input prompt via a generative language model (i.e., “portions of the tool or AI service response are communicated back to orchestrator 216 via flow 244, messaging service 210, and flow 234.”(0123) and “The service endpoints 468 can include APIs 470, which can be application programming interfaces that can enable the interaction or integration of the tools 422, the action planners 408, the agent planners 404, or the pluggable messaging service 462 with the automated agent system 200 or other external or third-party systems or services.”(0152)), the input prompt being determined based on the natural language input text, the non-text media textual summary, and the one or more data records (i.e., “A generative language model is a particular type of GAI model that is capable of generating content in response to model input. A generative language model is a particular type of GAI model that is capable of generating content in response to model input. The model input includes a task description, also referred to as a prompt. The task description can include instructions (e.g., natural language or multimodal instructions such as “please generate a summary of these search results” or a video recording of a demonstration of how to perform a task) and/or examples of digital content, such as text or multimodal content (e.g., examples of digital images, videos, articles, audio, or other content produced using a particular language, format, writing style, or tone).”(0037)); and transmitting the response message from the autonomous agent to the client machine via the communication channel (i.e., “This is because complex tasks especially require the autonomous agents to perform consistently and generate output in a user-expected and reliable manner, but the inherent nature of LLMs is that the output of the LLMs can be unpredictable.”(0039) and “ responsive to receiving input via one or more components of the environment 101, the automated agent 102 can be dynamically configured or reconfigured to perform a task or a series of tasks, via one or more components of the distributed multi-agent system 105.”(0068)). With respect to claims 2 and 16, Thompson, III discloses the method further comprising: instantiating the autonomous agent via the agent service based on an autonomous agent definition stored in the database system in accordance with a unified metadata framework defining a plurality of actions corresponding to operations capable of being performed via the computing services environment (i.e., “Any sub-agent 106A, 106B, . . . , 106N can include or be defined by a combination of computer code, data, memory, AI services 114, data resources 116, and/or tools 118, which are arranged or configured to perform a specific task or action.”(0073) or “Task or action as used herein can refer to an atomic action or operation that an agent or sub-agent is configured to perform, either alone or in combination with other tasks. Workflow as used herein can refer to an arrangement, sequence, or series of possible tasks that can be used by an agent to respond to a request or complete a goal or objective, from which an agent can select one or more specific tasks to complete the request, goal, or objective”(0079)), wherein instantiating the autonomous agent includes determining an initial runtime context for the autonomous agent, the initial runtime context including a plurality of data values accessible to the autonomous agent (i.e., “The service endpoints 468 can include runtime environments 472, which can be components or modules that can provide the necessary infrastructure or resources for executing the plans 410 or actions. ”(0153)). With respect to claims 3, 17, Thompson, III discloses the method recited in claim 2, the method further comprising: determining an initial execution plan for the autonomous agent by selecting a first subset of the plurality of actions based on the autonomous agent definition (i.e., “A context model can include various types of information, such as preferences, policies, profile data, historical user activity data, sensor data, network data, model parameters, and/or any other data that can be used to configure an agent, workflow, plan, or task. For example, examples may create, initialize and/or update a context model for a user associated with an automated agent 102 by extracting information from the user's online profile (e.g., a profile page and/or the user's online activity history on a social network)”(0085)) and the initial runtime context via the generative language model, wherein the non-text media summarization model is triggered by a first action of the first subset of the plurality of actions (i.e., “One or more signals generated by the environment 201 can trigger a start automated agent process 202”(0117) and “The event can be any trigger or input that initiates or updates a task, objective or goal for the automated agent.”(0127)). With respect to claim 4, Thompson, III discloses the method recited in claim 3, the method further comprising: determining an updated runtime context based on the non-text media textual summary, wherein the one or more database records are retrieved based on the updated run-time context (i.e., “Dynamic can also or alternatively be used herein to indicate that one or more system components, data structures or data stores, e.g., agents, workflows, databases, vector stores, memory layers, etc., are updated, reconfigured, or refreshed within a time interval that is less than the time interval between two different inputs to a computer system.”(0046)). With respect to claim 5, Thompson, III discloses the method recited in claim 4, the method further comprising: determining an updated execution plan for the autonomous agent by selecting a second subset of the plurality of actions based on the autonomous agent definition (i.e., “the method can continue to block 320, where the adaptive machine learning-based orchestrator 216 can update the portion of the context model used to perform the action (e.g., to adjust the weights applied to different portions of the context model) and then return to block 304 to obtain the updated context model and re-determine the objective (e.g., determine whether the objective remains the same or needs revision). If the objective is modified based on the updated context model, the orchestrator 216 can regenerate or update the plan at block 308 and begin executing actions of the regenerated or updated plan at block 310.”(0137)) and the updated runtime context via the generative language model, wherein the non-text media summarization model is triggered by a second action of the first subset of the plurality of actions (i.e., “The architecture 400 can include various components or modules that can be used by one or more agent planners 404 to generate or update workflows 406 and by one or more action planners to generate or update plans 410.”(0141) and “The service endpoints 468 can include runtime environments 472, which can be components or modules that can provide the necessary infrastructure or resources for executing the plans 410 or actions. The runtime environments 472 can include hardware, software, or network components or modules that can support the operation of the tools 422, the action planners 408, the agent planners 402, the pluggable modules 424, or the asynchronous distributed coordination 460.”(0153) and “A workflow within the architecture 400 can use a structured approach,… active planning and decision-making functions, and culminating with communication outputs and memory updates. Each or any phase of this process can be informed by the agent's profile, memory, planning, and action mechanisms, allowing the agent to manage complex tasks effectively. The communication outputs may comprise an instruction or trigger sent to another entity in the environment to trigger an action at the other entity.”(0204)). With respect to claims 6 and 18, Thompson, III discloses the method further comprising: determining natural language clarification query text based on the non-text media textual summary (i.e., “The AI services 230 can include various types of machine learning models and/or algorithms that can enhance the capabilities and performance of the automated agent 102, such as but not limited to one or more large language models (LLMs), Bayesian inference learning (BIL) models, machine learning (ML) models, or any other AI service that can provide, for instance, natural language understanding, natural language generation, dialog management, task execution, intent classification, entity extraction, information extraction, code generation, embedding generation models, similarity prediction,”(0114)); transmitting the natural language clarification query text from the autonomous agent to the client machine via the communication channel (i.e.,“ responsive to receiving input via one or more components of the environment 101, the automated agent 102 can be dynamically configured or reconfigured to perform a task or a series of tasks, via one or more components of the distributed multi-agent system 105.”(0068)); and receiving clarification input text from the client machine via the communication channel, wherein the one or more database records are retrieved based at least in part on the clarification input text (i.e., “An example of a zero-shot prompt is “classify the user input [input1] into action_a, action_b, or action_c,” where [input1] is a placeholder for the user input and/or associated context data and action_a, action_b, and action_c are possible intents into which the large language model may classify input1. A few-shot prompt includes examples along with an instruction to cause the large language model to follow the examples provided when processing an input. An example of a few-shot prompt is “‘software engineering’.fwdarw.job_search; ‘fill’.fwdarw.job_candidate_search; what is the intent of [input1] ?” where ‘software engineering’.fwdarw.job_search and ‘fill’.fwdarw.job_candidate_search are examples of how to classify inputs into search categories, and input1 includes user input and/or context data.”(0100)). With respect to claim 7, Thompson, III discloses the method recited in claim 6, wherein the natural language clarification query text identifies two or more database records, and wherein the clarification input text identifies a first database record of the two or more database records (i.e., “An example of a few-shot prompt is “‘software engineering’ .fwdarw.job_search; ‘fill’.fwdarw.job_candidate_search; what is the intent of [input1] ?” where ‘software engineering’.fwdarw.job_search and ‘fill’.fwdarw.job_candidate_search are examples of how to classify inputs into search categories, and input1 includes user input and/or context data.”(0100) and “ a dialog or conversation can have an associated user identifier, session identifier, conversation identifier, or dialog identifier, and an associated timestamp”(0064)). With respect to claim 8, Thompson, III discloses wherein the non-text media includes an image or video (i.e., “aspects of the disclosed technologies can be used to receive input and/or generate output that includes non-text forms of content, such as digital imagery, videos, multimedia, audio, hyperlinks, and/or platform-independent file formats.”(0059)), and wherein the non-text media summarization model includes an object recognition model, and wherein the non-text media textual summary includes a description of an object represented in the non-text media (i.e., “ The model input includes a task description, also referred to as a prompt. The task description can include instructions (e.g., natural language or multimodal instructions such as “please generate a summary of these search results” or a video recording of a demonstration of how to perform a task) and/or examples of digital content, such as text or multimodal content (e.g., examples of digital images, videos, articles, audio, or other content produced using a particular language, format, writing style, or tone). ” (0037)). With respect to claim 9, Thompson, III discloses wherein the method recited in claim 8, wherein the one or more database records include a knowledge article corresponding to the object, and wherein the non-text media textual summary describes of an object represented in the non-text media summarization model (i.e., “ RAG or RAFT can be used to perform domain-specific fine tuning of a pre-trained machine learning model using, e.g., samples of digital content that represent the desired domain-specific knowledge. Using RAG, digital content can be stored in and retrieved from a data store, e.g., a database such as a vector database, using queries that are configured to measure the similarity between the digital content in the vector database and the query, question, or request being asked” (0056) and “ the embeddings are input to the generative language model, and the generative language model outputs digital content, e.g., natural language text or a combination of natural language text and non-text output, based on the embeddings.”(0581)). With respect to claim 10, Thompson, III discloses wherein the non-text media includes an image or video, and wherein the non-text media summarization model includes a text recognition model, and wherein the non-text media textual summary includes a text portion represented in the non- text media (i.e., “The model input includes a task description, also referred to as a prompt. The task description can include instructions (e.g., natural language or multimodal instructions such as “please generate a summary of these search results” or a video recording of a demonstration of how to perform a task) and/or examples of digital content, such aa text or multimodal content (e.g., examples of digital images, videos, articles, audio, or other content produced using a particular language, format, writing style, or tone). Portions of the task description can be in the form of natural language text, such as a question or a statement. Alternatively or in addition, a task description or prompt can include non-test forms of content, such as digital imagery and/or digital audio.”(0037) and “ A language model or large language model can be configured to perform one or more natural language processing (NLP) tasks, such as generating content, classifying content, answering questions in a conversational manner, and translating content from one language to another.”(0038)). With respect to claim 11, Thompson, III discloses the method recited in claim 10, wherein a database record of the one or more database records correspond to the text portion (i.e., “RAG or RAFT can be used to perform domain-specific fine tuning of a pre-trained machine learning model using, e.g., samples of digital content that represent the desired domain-specific knowledge. Using RAG, digital content can be stored in and retrieved from a data store, e.g., a database such as a vector database, using queries that are configured to measure the similarity between the digital content in the vector database and the query, question, or request being asked” (0056) and “document databases, document-oriented databases, column-oriented data stores, or document stores, such as NOSQL document stores, can be used to implement portions of the multi-layered memory structures.”(0078)). With respect to claim 12, Thompson, III discloses wherein the autonomous agent is configured as a conversational chat assistant, and wherein the input includes user input received via a chat session conducted via the communication channel (i.e.,. “At dialog element 1202, the automated agent opens the dialog, e.g., “chat session” with the user in a personalized way, with an introduction 1204 that mentions the user's name” (0343)). With respect to claim 13, Thompson, III discloses wherein the non-text media includes an audio segment or a video segment, and wherein the non-text media summarization model includes a speech recognition model (i.e., “The model input includes a task description, also referred to as a prompt. The task description can include instructions (e.g., natural language or multimodal instructions such as “please generate a summary of these search results” or a video recording of a demonstration of how to perform a task) and/or examples of digital content, such aa text or multimodal content (e.g., examples of digital images, videos, articles, audio, or other content produced using a particular language, format, writing style, or tone). Portions of the task description can be in the form of natural language text, such as a question or a statement. Alternatively or in addition, a task description or prompt can include non-test forms of content, such as digital imagery and/or digital audio.”(0037)). With respect to claim 14, Thompson, III discloses wherein the non-text media summarization model receives as input all or a portion of the natural language input text in addition to the non-text media (i.e., “The model input includes a task description, also referred to as a prompt. The task description can include instructions (e.g., natural language or multimodal instructions such as “please generate a summary of these search results” or a video recording of a demonstration of how to perform a task) and/or examples of digital content, such aa text or multimodal content (e.g., examples of digital images, videos, articles, audio, or other content produced using a particular language, format, writing style, or tone). Portions of the task description can be in the form of natural language text, such as a question or a statement. Alternatively or in addition, a task description or prompt can include non-test forms of content, such as digital imagery and/or digital audio.”(0037)). With respect to claim 19, Thompson, III discloses all limitation of claims of claims 9-11 above (see rejection above). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to HUNG T VY whose telephone number is (571)272-1954. The examiner can normally be reached M-F 8-5. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tony Mahmoudi can be reached at (571)272-4078. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /HUNG T VY/ Primary Examiner, Art Unit 2163 July 11, 2026
Read full office action

Prosecution Timeline

Jan 27, 2025
Application Filed
Jul 15, 2026
Non-Final Rejection mailed — §102 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705220
SYSTEM AND METHOD FOR TRACEABLE PUSH-BASED ETL PIPELINE AND MONITORING OF DATA QUALITY IN ETL ENVIRONMENTS
1y 8m to grant Granted Aug 11, 2026
Patent 12693996
SCALABLE, SECURE, EFFICIENT, AND ADAPTABLE DISTRIBUTED DIGITAL LEDGER TRANSACTION NETWORK
1y 4m to grant Granted Jul 28, 2026
Patent 12675467
HORIZONTAL PROCESSING OF SEQUENTIAL DATA (HPSD), AND DETACHED FIELD BATCH PROCESS (DFBP)
1y 10m to grant Granted Jul 07, 2026
Patent 12657519
CLUSTER BASED TRAINING HOST SELECTION IN ASYNCHRONOUS FEDERATED LEARNING MODEL COLLECTION
3y 0m to grant Granted Jun 16, 2026
Patent 12651180
STATE PREDICTION RELIABILITY MODELING
3y 9m to grant Granted Jun 09, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
86%
Grant Probability
88%
With Interview (+2.1%)
2y 7m (~1y 1m remaining)
Median Time to Grant
Low
PTA Risk
Based on 922 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month