Prosecution Insights
Last updated: August 17, 2026
Application No. 18/427,304

METADATA DRIVEN PROMPT GROUNDING FOR GENERATIVE ARTIFICIAL INTELLIGENCE APPLICATIONS

Non-Final OA §103
Filed
Jan 30, 2024
Priority
Sep 11, 2023 — provisional 63/581,729
Examiner
MORALES, PEDRO JESUS
Art Unit
Tech Center
Assignee
Salesforce Inc.
OA Round
1 (Non-Final)
62%
Grant Probability
Moderate
1-2
OA Rounds
1y 1m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 62% of resolved cases
62%
Career Allowance Rate
8 granted / 13 resolved
+1.5% vs TC avg
Strong +56% interview lift
Without
With
+55.6%
Interview Lift
resolved cases with interview
Typical timeline
3y 8m
Avg Prosecution
20 currently pending
Career history
34
Total Applications
across all art units

Statute-Specific Performance

§101
24.4%
-15.6% vs TC avg
§103
47.5%
+7.5% vs TC avg
§102
11.9%
-28.1% vs TC avg
§112
13.1%
-26.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 13 resolved cases

Office Action

§103
DETAILED ACTION This non-final office action is responsive to application 18/427,304 as submitted on January 30th 2024. Claim status is currently pending and under examination for claims 1-20 of which independent claims are 1, 9 and 17. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The following are the references relied upon in the rejections below: Zhou (US 20240362093 A1) Mansour (US 20250005263 A1) Claims 1-3, 5-11 and 13-19 are rejected under 35 U.S.C. 103 as being unpatentable over Zhou in view of Mansour. Regarding Claim 1, Zhou teaches: A method for grounding a generative artificial intelligence (AI) prompt for a tenant of a [single-tenant] … system, comprising (The Examiner interprets “grounding” according to its broadest reasonable interpretation (BRI) in view of the Applicant’s specification as encompassing conditioning a large language model when generating a response to a user query. This interpretation is consistent with the illustrative descriptions in the Applicant’s specification at [0019], (see excerpt below). Applicant’s written description at [0019]: “grounding refers to the processing of making a prompt (e.g., a request or instruction) more specific, clear, and unambiguous so the generative AI application or LLM can generate more accurate and contextually relevant responses. Additionally or alternatively, the process of grounding may include providing the LLM with a corpus of data from which it is to generate the content of a response … The process of grounding involves providing additional context or details to inform the LLM's understanding of the requested task/query and/or the universe of information from which the LLM may draw content.” [0005] “provide a client device with limited computational capabilities (e.g., limited processing power, memory or battery power) the ability to tailor LLM responses on user documents stored on the client device (i.e. a custom corpus of documents associated with a user), without having to run or store the LLM locally on the client device” [0025] “the client device 110, the one or more additional client devices, and/or any other computing devices of a user can form an ecosystem of devices that can employ techniques described herein. These additional client devices and/or computing devices may be in communication with the client device 110 (e.g., over the network(s) 199). As another example, a given client device can be utilized by multiple users in a shared setting (e.g., a group of users of a household, a workplace, or the like).” [0003] “utilizing a custom corpus of documents to condition an LLM when generating a response to a user query (e.g., a user submitted query or an automatically generated query) for rendering (e.g., audibly and/or graphically) by a user device. The LLM processes a received user query to generate one or more API queries for one or more external applications that each has access to a respective custom corpus of documents. The one or more external applications select one or more relevant documents from their respective document corpus based on the API query and return data representative of the selected one or more documents (e.g., a copy of the document, a snippet of the document and/or an embedding representing the document or a snippet thereof) to the LLM. The LLM uses data representing at least one of the selected documents as contextual data when generating a response to the user query.” Documents from a custom corpus can be used as contextual data for conditioning an LLM when generating a response to a user query (‘generative AI prompt’). A custom corpus contains documents related to a specific subject/topic (see [0007]), therefore when retrieving relevant documents from the custom corpus, the retrieved documents provide specialized, relevant context for the LLM to generate a contextual and relevant response for the user query (and therefore grounding a generative AI prompt). A client device (‘tenant’) is comprised of multiple users of a shared setting (workplace) and is connected to a network, therefore a client device storing a custom corpus is a tenant of a single-tenant system.): receiving user input indicating a configuration that identifies one or more large language models (LLMs) configurable for the tenant and an indication of a subset of documents of a set of documents stored at the [single]-tenant … system ([0028] “The LLM selection engine 132 can, in response to receiving a query, determine which, if any, of multiple generative model(s) (LLM(s) and/or other generative model(s)) to utilize in generating response(s) to render responsive to the query. For example, the LLM selection engine 132 can select none, one, or multiple generative model(s) to utilize in generating response(s) to render responsive to a query.” [0042] “the user query is accompanied by access control data, such as a token or key, for allowing an LLM to access external application(s) and/or a custom corpus of documents. This can allow the system to access one or more access restricted custom corpora of documents. In some versions of those implementations, the access control data may have a limit to the number of times it can be used by the system, e.g., only allow access to the custom corpus once” A user’s query (‘received user input’) is used to determine which LLMs to use to generate responses to the user’s query (therefore a user query indicates a configuration that identifies one or more LLMs configurable for a tenant (client device)). A user query is accompanied by access control data which indicates which custom corpora of documents (‘subset of documents of a set of documents’ stored at a client device (‘tenant’), see [0005]) can be accessed by an LLM.), the subset of documents indicated in the configuration as being available to the tenant for processing with the one or more LLMs (A user query is accompanied by access control data which indicates which custom corpora of documents can be accessible (or restricted) to an LLM (therefore access control data (‘configuration’) indicates which corpora (‘subset of documents’) is available to a user of a client device (‘tenant’) for LLM processing).); generating one or more respective vectorizations of content of each document of the subset of documents ([0075] “At block 454, the system compares the context vector to a set of precomputed embeddings, each embedding representing a respective document in the custom corpus of documents. Each precomputed embedding, or each precomputed embedding in a subset of the embeddings, is compared with the context vector” [0034] “utilize document embeddings 170 of documents in the custom corpus 168 of documents accessible to the external application(s)”); receiving, from a user associated with the tenant, a request to generate a generative response with the one or more LLMs ([0029] “The LLM input engine 134 can, in response to receiving a query, generate LLM input that is to be processed using an LLM in generating an NL based response to the query.” See [0025] describing multiple users are connected to a client device (‘tenant’)); generating the generative AI prompt using the content of one or more documents of the subset of documents to ground the generative AI prompt ([0004] “the LLM can generate query responses that are tailored to particular set of documents (i.e., the documents included in the custom corpus) without any fine-tuning or additional training of the LLM. … This enables the query responses to be tailored to the custom corpus without any having to perform any of the aforementioned fine-tuning or additional training of the LLM.” [0056] “At block 260 the system generates a response to the user query using the LLM. The LLM is conditioned on the data representative of one or more of the documents in the custom corpus of documents received from the external application, e.g., data representing one or more of the received documents may be input into the LLM as contextual data for generating the response.” A set of documents (‘subset of documents’) is used as input to an LLM to generate a response to a user query (‘generative AI prompt’). The inputted set of documents and the user query are used to generate a response, therefore the inputted set of documents and user query are a ‘generative AI prompt’ that is generated using the content of a subset of documents. When the set of documents are input into the LLM to generate a response, the set of documents (contextual data) are used to provide more specific and contextually relevant content to the LLM, therefore making a user query more specific (and therefore grounding the user query (‘generative AI prompt’) by generating a generative AI prompt (inputted set of documents and user query) that makes the user query more specific by appending contextual data (set of documents) to the query).), wherein the subset of documents is identified based at least in part on a comparison between a vectorization of the request and at least one of the one or more respective vectorizations of the subset of documents and based at least in part on a determination that the user associated with the tenant is permitted to access the subset of documents ([0048] “the LLM generates a current context vector based on the user query. The current context vector may be an embedded representation of the user query” [0073-0074] “At block 452, the system receives an API query comprising a context vector representing a user query. … the API query may include access control data, such as a token or key. The system may compare the access control data in the API query to access control data associated with a custom corpus to determine if the entity sending the query has permission to access the custom corpus. If the access control data in the API query matches the access control data associated with a custom corpus, then the system proceeds to block 454” [0075-0076] “At block 454, the system compares the context vector to a set of precomputed embeddings, each embedding representing a respective document in the custom corpus of documents. Each precomputed embedding, or each precomputed embedding in a subset of the embeddings, is compared with the context vector. Each precomputed embedding comprises a vector in an embedding space that is representative of the contents of a corresponding document in the custom corpus.” [0079] “At block 456, the system selects one or more documents from the custom corpus based on the comparison of the context vector to the set of precomputed embeddings” A user query is embedded as a context vector (‘vectorization of the request’) by an LLM. An API query uses the context vector to determine if a user sending the user query has permission to access a custom corpus (therefore determining that a user associated with a tenant is permitted to access a subset of documents). If the user has permission, then the context vector is compared to precomputed embeddings representing each document in a corpus of documents (‘one or more respective vectorizations of the subset of documents’). The comparison of the context vector and precomputed embeddings are used to select one or more documents from the custom corpus (therefore identifying a subset of documents).); and presenting a response to the generative AI prompt, the response generated by the one or more LLMs using the generative AI prompt ([0054] “At block 258, the system receives a response to the API query from the external application. The response includes data representative of one or more documents in the custom corpus of documents accessible by the external application that are selected by the external application based on their relevance to the user query as described in further detail herein (e.g., with reference to FIG. 4).” [0056] “At block 260 the system generates a response to the user query using the LLM. The LLM is conditioned on the data representative of one or more of the documents in the custom corpus of documents received from the external application, e.g., data representing one or more of the received documents may be input into the LLM as contextual data for generating the response.” [0058] “At block 262, the system causes the response to the user query to be rendered at the user device. For example, the system can cause the response to be rendered graphically in an interface of an application of a client device via which the query was submitted.”). However, Zhou does not teach grounding a generative AI prompt for a tenant of a multi-tenant database system which is taught by Mansour: A method for grounding a generative artificial intelligence (AI) prompt for a tenant of a multi-tenant database system (The Examiner interprets “grounding” according to its BRI in view of the Applicant’s specification as encompassing adding contextual information to a prompt and defining a format for a prompt. This interpretation is consistent with the illustrative descriptions in the Applicant’s specification at [0019], (see excerpt below). Applicant’s written description at [0019]: “Examples of grounding include adding contextual information to the prompt, defining the expected input/output format, instructing the model to avoid certain terminology, etc.” [0134] “a multiplatform service provider can host a suite of collaboration tools. For example, a multiplatform service provider may host, for its clients, a multitenant issue tracking system, a multitenant code repository service, and a multitenant documentation service. In this example, an organization that is a customer/client of the service provider may be a tenant of each of the issue tracking system, the code repository service, and the documentation service.” [0095] “a prompt provided as input to a generative output engine can be engineered from user input. For example, in some cases, a user input can be inserted into an engineered template prompt that itself is stored in a database. For example, an engineered prompt template can include one or more fields into which user input portions thereof can be inserted. In some cases, an engineered prompt template can include contextual information that narrows the scope of the prompt, increasing the specificity thereof.” [0096] “For example, some engineered prompt templates can include example input/output format cues or requests that define for a generative output engine, as described herein, how an input format is structured and/or how output should be provided by the generative output engine.” [0207] “the system can include a prompt context analysis instance or other service that monitors user input and/or generative output for compliance with a set of policies or content guidelines associated with the tenant or organization. For instance, the service may monitor the content of a user input and block potential ethical violations including hate speech, derogatory language, or other content that may violate a set of policies or content guidelines. The service may also monitor output of the generative engine to ensure the generative content or response is also in compliance with policies or guidelines.” A user prompt (‘generative AI prompt’) can be inserted into an engineered prompt template to add contextual information that narrows the scope of the prompt and increases the specificity of the prompt and defines an input format for the prompt, therefore ‘grounding’ a generative AI prompt.), receiving user input indicating a configuration that identifies one or more large language models (LLMs) configurable for the tenant and an indication of a subset of documents of a set of documents stored at the multi-tenant database system ([0184] “the prompt management service 114 is configured to structure an API request to the generative engine service 116. The API request can include the modified prompt as an attribute of a structured data object that serves as a body of the API request. Other attributes of the body of the API request can include, but are not limited to: an identifier of a particular LLM or generative engine to receive and continue the modified prompt; a user authentication token; a tenant authentication token; an API authorization token; a priority level at which the generative engine service 116 should process the request” [0206] “the system 100 can include a prompt context analysis instance configured to determine whether a user issuing a request has permission to access the resources required to service that request. For example, a prompt from a user may be “Generate a text summary in Document123 of all changes to Kanban board 456 that do not have a corresponding issue tagged in the issue tracking system.” In respect of this example, the prompt context analysis instance may determine whether the requesting user has permission to access Document123, whether the requesting user has written permission to modify Document123, whether the requesting user has read access to Kanban board 456, and whether the requesting user has read access to referenced issue tracking system.” An API request (‘received user input’) includes a modified user prompt and includes an identifier (‘configuration’) that describes which LLM to receive (identifies an LLM that is configurable for a user of a tenant). A prompt also includes a request (‘indication’) to access resources (‘subset of documents’) of a Kanban board (see [0134] describing a multitenant issue tracking system (‘multi-tenant database system’)).), the subset of documents indicated in the configuration as being available to the tenant for processing with the one or more LLMs ([0133] “The organization can create and/or purchase user accounts for its employees so that each employee has access to both messaging and project management functionality. In some cases, the organization may limit seats in each tenancy of each platform so that only certain users have access to messaging functionality and only certain users have access to project management functionality;” See [0206] describing that it is first determined if a user has permission to access requested resources (‘subset of documents’). A user can access resources it has permission for, therefore indicating a ‘subset of documents’ as being available to a tenant for processing with an LLM.); Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to combine the method of Zhou with the multi-tenant system disclosed by Mansour to perform prompt engineering for tenants of a multi-tenant system. By performing prompt engineering for tenants of a multi-tenant system, a shared infrastructure can be used to modify the prompts of tenants into specific and contextual prompts according to each tenant’s preferences and policies, thereby saving computational resources. Regarding Claims 2, 10 and 18, the combined method of Zhou/Mansour teaches: The method of claim 1, further comprising: storing configuration metadata comprising the indication of the one or more LLMs and the indication of the subset of documents ([0088] “a user can specify, as part of a query or via interface element(s) in conjunction with a query (e.g., selectable interface element(s) provided near a query input field), desired formatting option(s) for an NL based response. For example, a desired formatting option could be “list format”, “graph format”, “top 5”, “in the style of”, etc. … the specified format can be used to select an LLM for the selected format. For example, if “list format” is specified, an LLM that is trained a list format prompt can be selected as the LLM to utilize in generating a NL based response” [0047] “Where the user query provided at block 252 includes access control data, the LLM may incorporate the access control data into the API query, e.g., incorporate a token or key received in the user query into the API query. Alternatively, the LLM may itself be assigned a set of access control data by an organization. Such access control data may be incorporated into the API query.” A user query includes a desired formatting option and access control data. The desired formatting option (‘configuration metadata’) is used to select a trained LLM (therefore the user query stores configuration metadata since the query includes a desired formatting option (‘indication’) that determines which LLM to use). Access control data (‘configuration metadata’) is used to determine which custom corpora of documents a user can access (see [0042]), therefore a user query stores configuration metadata since the query includes access control data used to determine (indicate) which documents to use.); wherein generating the one or more respective vectorizations is based at least in part on the configuration metadata ([0074] “the API query may include access control data, such as a token or key. The system may compare the access control data in the API query to access control data associated with a custom corpus to determine if the entity sending the query has permission to access the custom corpus.” [0075-0076] “At block 454, the system compares the context vector to a set of precomputed embeddings, each embedding representing a respective document in the custom corpus of documents. Each precomputed embedding, or each precomputed embedding in a subset of the embeddings, is compared with the context vector. Each precomputed embedding comprises a vector in an embedding space that is representative of the contents of a corresponding document in the custom corpus. The embedding vectors may, for example, be outputs of an encoder model, such as the encoder portion of a variational autoencoder (VAE) model trained on text documents”); and wherein generating the generative AI prompt is based at least in part on the configuration metadata ([0056] “At block 260 the system generates a response to the user query using the LLM. The LLM is conditioned on the data representative of one or more of the documents in the custom corpus of documents received from the external application, e.g., data representing one or more of the received documents may be input into the LLM as contextual data for generating the response.” A set of documents from a custom corpus is used as input to an LLM to generate a response to a user query. The inputted set of documents and the user query are used to generate a response, therefore the inputted set of documents and user query are a ‘generative AI prompt’ that is generated using the content of a set of documents. A user query includes access control data (‘configuration metadata’) that is used to determine which documents a user has access to, and the most relevant and accessible documents are used as input to an LLM (therefore generating the generative AI prompt based on configuration metadata.). Regarding Claims 3, 11 and 19, the combined method of Zhou/Mansour teaches: The method of claim 2, wherein the configuration metadata further comprises an indication of one or more fields of the subset of documents, an indication of a quantity of the subset of documents whose content is to be vectorized, an indication of a vectorization model, an indication of one or more specifications of the subset of documents, or any combination thereof ([0042] “the user query is accompanied by access control data, such as a token or key, for allowing an LLM to access external application(s) and/or a custom corpus of documents. This can allow the system to access one or more access restricted custom corpora of documents. In some versions of those implementations, the access control data may have a limit to the number of times it can be used by the system, e.g., only allow access to the custom corpus once. In additional or alternative versions of those implementations, the access control data may be time-limited, i.e., provide access to the custom corpus for a predetermined time after the user query is submitted.” Access control data (‘configuration metadata’) specifies how many times a system can access a custom corpus of documents (‘subset of documents’) and a predetermined time for accessing the corpus, therefore the access control data is used to determine (indicate) specifications of a subset of documents.). Regarding Claims 5 and 13, the combined method of Zhou/Mansour teaches: The method of claim 1, wherein generating the one or more respective vectorizations of content further comprises: transmitting, to the one or more LLMs, a first vectorization request to generate the one or more respective vectorizations of the content of each document of the subset of documents ([0077] “The precomputed embeddings may be determined continuously, e.g. whenever a new document is added to the custom corpus or an existing document updated, a corresponding embedding is determined. Additionally, or alternatively, the custom corpus may be checked periodically to determine if any new documents have been added or any existing documents have been updated, and corresponding embeddings determined if so.” [0037] “The document embeddings 170 are latent representations of the documents (e.g., vectors in an embedding space or lower-dimensional latent space) that each represent a respective document or set of documents in the custom corpus. The embeddings may be generated from an intermediate output of an LLM, such as output of an intermediate layer of the LLM that has processed the document.” When a new document is added or when an existing document of a corpus is updated, an LLM generates a document embedding for the document. For the LLM to generate an embedding whenever a document is added or updated, the LLM must receive an instruction to generate the embedding in the form of tokens (vectorization request), therefore a first vectorization request to generate vectorizations of the content of each document is implied.); and receiving, from the one or more LLMs, the one or more respective vectorizations of the content ([0037] “The embeddings may be generated from an intermediate output of an LLM”). Regarding Claims 6 and 14, the combined method of Zhou/Mansour teaches: The method of claim 1, wherein generating the generative AI prompt further comprises: transmitting, to the one or more LLMs, a second vectorization request to generate the vectorization of the request ([0029] “The LLM input engine 134 can, in response to receiving a query, generate LLM input that is to be processed using an LLM in generating an NL based response to the query.” [0048] “the LLM generates a current context vector based on the user query. The current context vector may be an embedded representation of the user query, e.g., an intermediate output of the LLM/output of an intermediate layer of the LLM, and may further be based on recent user queries provided to the system, user data (such as a user identity or profile), a client device state, and/or other contextual information.” A user provides a query (‘request’) for an LLM to generate a response. The user query is processed into a context vector (‘a vectorization of the request’) by an LLM. To generate the embedding of the query (context vector), the LLM must receive an instruction to generate the embedding in the form of tokens (‘second vectorization request’), therefore transmitting a second vectorization request to an LLM to generate the vectorization of the request (context vector) is implied.); and receiving, from the one or more LLMs, the vectorization of the request ([0048] “the LLM generates a current context vector based on the user query.”). Regarding Claims 7 and 15, the combined method of Zhou/Mansour teaches: The method of claim 1, further comprising: storing access metadata comprising an indication of the subset of documents, an indication of one or more fields of the subset of documents, or any combination thereof ([0047] “Where the user query provided at block 252 includes access control data, the LLM may incorporate the access control data into the API query, e.g., incorporate a token or key received in the user query into the API query. Alternatively, the LLM may itself be assigned a set of access control data by an organization. Such access control data may be incorporated into the API query.” A user query includes access control data. Access control data (‘access metadata’) is used to determine which custom corpora of documents a user can access (see [0042]), therefore a user query stores access metadata since the query includes access control data used to determine (indicate) which subset of documents to use.); and querying a central authentication service to produce an authentication result indicating that the tenant is permitted to access the one or more of the subset of documents, the one or more fields, or any combination thereof ([0074] “the API query may include access control data, such as a token or key. The system may compare the access control data in the API query to access control data associated with a custom corpus to determine if the entity sending the query has permission to access the custom corpus. If the access control data in the API query matches the access control data associated with a custom corpus, then the system proceeds to block 454. If the access control data does not match the access control data associated with a custom corpus, then the system prevents access to the custom corpus.” [0053] “the external application may utilize the access control data provided in the API query to determine whether the user of the client device and/or the LLM has permission to access the custom corpus of documents, as described in further detail herein (e.g., with reference to FIG. 4)” [0008] “the entity/user that manages the external application and/or custom corpus of documents may be a third-party entity/user that is a distinct entity from a first-party entity/user that trains or manages the LLM. … the third-party entity/user can provide the LLM access to the third-party external application and the third-party custom corpus of documents via the third-party external application, or the third-party entity/user can provide the LLM access directly to the third-party custom corpus of documents. As a result, the LLM can not only generate responses based on data that was utilized in training the LLM, but also based on data that is present in the third-party custom corpus of documents.” An API query includes access control data and an external application uses the access control data provided in the API query to determine if a user of a client device (‘tenant’) has permission to access a corpus of documents (‘subset of documents’). The external data application (‘central authentication service’) compares access control data to determine access, therefore querying a central authentication service to produce an authentication result.); wherein presenting the response to the generative AI prompt is based at least in part on the authentication result ([0074] “If the access control data in the API query matches the access control data associated with a custom corpus, then the system proceeds to block 454.” [0075] “At block 454, the system compares the context vector to a set of precomputed embeddings, each embedding representing a respective document in the custom corpus of documents” [0056] “data representing one or more of the received documents may be input into the LLM as contextual data for generating the response.” If the access control data in an API query matches access control data associated with a corpus (the user is authenticated), the documents from the corpus are compared to a context vector (embedding of user query). Then based on the comparison, a relevant set of documents are selected from the corpus (see [0079]) to be used as input to an LLM model for generating a response. The response is then rendered at a user device (see [0058]).). Regarding Claims 8 and 16, the combined method of Zhou/Mansour teaches: The method of claim 1, wherein generating the generative AI prompt is based at least in part on a runtime context associated with the request to generate the generative response ([0062] “the system determines if the query is directed to one or more custom corpus accessible by the system. The system may process the query to determine a user intent, for example using the LLM to identify entities or subjects referred to explicitly or implicitly in the query. Further, the system may determine whether the query is directed to one or more custom corpus accessible by the system based on the determined intent. Additional contextual information (e.g. a user profile or identity, a recent query/interaction history of the user, a location etc.) may be used to make the determination of whether the query is directed to one or more custom corpus accessible by the system.” [0063] “For example, if a user employed by a particular entity uses the term “our” in a query, the system may infer that the user is referring to the entity, and consequently determine that the query is directed towards a custom corpus managed by the entity. As another example, if a user uses the term “my” in a query, the system may infer that the user is referring to personal documents, and consequently determine that the query is directed towards a personal custom corpus of the user.” A user query (‘request’) is processed by an LLM to identify a user’s intent (‘runtime context’). Depending on the user’s intent, a specific custom corpus of documents is selected for the user. When a custom corpus of documents is selected, document embeddings representing the documents in the corpus are compared to a context vector (see [0075]). Based on the comparison, the most relevant documents are selected from the corpus (see [0079-0080]). The relevant documents are inputted into an LLM as contextual data to generate a response to the user query (see [0056]). The inputted documents and the user query are used to generate a response, therefore the inputted documents and user query are a ‘generative AI prompt’ that is generated based on user intent (‘runtime context’).). Regarding Claim 9, the rejection of claim 1 is incorporated. The difference in scope being: An apparatus … comprising: one or more memories storing processor-executable code ([0113] “non-transitory computer readable storage media storing computer instructions); and one or more processors coupled with the one or more memories and individually or collectively operable to execute the code to cause the apparatus to ([0113] “non-transitory computer readable storage media storing computer instructions executable by one or more processors”).). Regarding Claim 17, the rejection of claim 1 is incorporated. The difference in scope being: A non-transitory computer-readable medium storing code, the code comprising instructions executable by one or more processors to ([0113] “non-transitory computer readable storage media storing computer instructions executable by one or more processors”). The following are the references relied upon in the rejections below: Smith (US 20240419709 A1) Claims 4, 12 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Zhou in view of Mansour, further in view of Smith. Regarding Claims 4, 12 and 20, the combined method of Zhou/Mansour teaches the method of claim 1, however the combination does not teach a request indicates generation metadata indicating one or more fields of one or more documents of a subset of documents, which is taught by Smith: the request indicates generation metadata indicating one or more fields of the one or more documents of the subset of documents, or any combination thereof ([0028] “the query may include a query tag that points to an LLM tag that may be located in the input data 108.” [0024] “input data 108 may include any type of data, including text documents” [0078] “the LLM may identify the LLM tag in the text or body of the input document.” [0051] “the LLM tag 110 may include a breakpoint tag in the input data 108” [0058] “the breakpoint tag may identify any type of additional information that the LLM 104 may utilize. For example, the breakpoint tag may identify a request for authorization to authorize a particular portion of the input data 108. The authorization may be based on access level, security level, need-to-know, any other authorization, and combinations thereof, of the requesting entity.” A query (‘request’) includes a query tag (‘generation metadata’) that points to an LLM tag. The LLM tag can include a breakpoint tag that identifies a request for authorizing a portion (‘one or more fields’) of an input document.); and generating the generative AI prompt is based at least in part on the generation metadata ([0081] “Requesting external input may help to improve the relevance of the response of the LLM to the query. … receiving external input authorizing access to a particular input document or set of input documents may allow the LLM to prepare the response to the query based on the authorized input document or the authorized input data.” A query tag (‘generation metadata’) is used to authorize a portion of an input document. The authorized portion of the input document is used by an LLM to prepare a response to a user query. By using the authorized input document to generate a response to query, the LLM uses both the document and query as input to generate the response. Therefore, the authorized document and query are a ‘generative AI prompt’ that is generated to create a response.). Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined method of Zhou/Mansour with the query tags disclosed by Smith to use query tags to retrieve authorized portions of documents. By using query tags to retrieve authorized portions of documents, users are restricted to accessing only portions of a document they have permission for, thereby limiting an LLM to only use data a user is allowed to see in generated responses. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Austin et al. (US 20240420012 A1) teaches creating a modified query by combining a user query with a curated LLM prompt that is generated based on a context of the user query. Srinivas Aluri (“Context-Aware IDE Systems Using Large Language Models and Semantic Memory Architectures”) teaches implementing a retrieval-augmented generation pipeline to fetch relevant documents before generating LLM responses, thereby reducing hallucination. Jain et al. (US 20240256582 A1) teaches retrieving a set of search results for a user query and using relevant search results as part of an input prompt to guide an LLM model. Any inquiry concerning this communication or earlier communications from the examiner should be directed to PEDRO J MORALES whose telephone number is (571)272-6106. The examiner can normally be reached 8:30 AM - 6:00 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, MIRANDA M HUANG can be reached at (571)270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /PEDRO J MORALES/Examiner, Art Unit 2124 /MIRANDA M HUANG/Supervisory Patent Examiner, Art Unit 2124
Read full office action

Prosecution Timeline

Jan 30, 2024
Application Filed
Jul 21, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12664036
ANOMALY DETECTION IN COMPUTER SYSTEMS
4y 0m to grant Granted Jun 23, 2026
Patent 12651179
REINFORCEMENT LEARNING MODEL FOR BALANCED UNIT RECOMMENDATION
4y 5m to grant Granted Jun 09, 2026
Patent 12639625
BIAS ADJUSTMENT DEVICE, INFORMATION PROCESSING DEVICE, INFORMATION PROCESSING METHOD, AND INFORMATION PROCESSING PROGRAM
4y 1m to grant Granted May 26, 2026
Patent 12591803
SYSTEMS AND METHODS FOR APPLYING MACHINE LEARNING BASED ANOMALY DETECTION IN A CONSTRAINED NETWORK
3y 11m to grant Granted Mar 31, 2026
Patent 12530412
SEARCH-QUERY SUGGESTIONS USING REINFORCEMENT LEARNING
4y 2m to grant Granted Jan 20, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
62%
Grant Probability
99%
With Interview (+55.6%)
3y 8m (~1y 1m remaining)
Median Time to Grant
Low
PTA Risk
Based on 13 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month