DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1-20 are presented for examination.
Examiner Remarks:
Paragraph [0018] of the specification defines “computer-readable storage medium” as not to be construed as transitory signals per se.
Paragraph [0038] of the specification defines “number of” as “one or more items”. Accordingly, Examiner is interpreting any claim limitation reciting a “number of” as “one or more”.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on April 16, 2024, is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Specification
The abstract of the disclosure is objected to because of the following informalities:
The first sentence is incomplete
“managing prompt” should read “managing prompts”
“training dataset comprises” should read “training dataset comprising”
A corrected abstract of the disclosure is required and must be presented on a separate sheet, apart from any other text. See MPEP § 608.01(b).
The disclosure is objected to because of the following informalities:
Title: “Learing” should read “Learning”
[0002]: "using huge amount of data" should read "using huge amounts of data"
[0005]: "training dataset comprises" should read "training dataset comprising"
[0037]: "training dataset comprises" should read "training dataset comprising"
[0038]: "used in herein" should read "used herein"
[0053]: "identify pattern" should read "identify a pattern"; "Pattern for the tabular data" should read "A pattern for the tabular data"; "the value corresponds to each centroid" should read "the value corresponding to each centroid"
[0060]: "parameters for machine learning model 256 continuously refined" should read "parameters for machine learning model 256 are continuously refined"
[0067]: "updated version" should read "updated versions"
[0072]: "dataset 246 that contain" should read "dataset 246 that contains"
[0073]: "used as reference or benchmark" should read "used as references or benchmarks"
[0075]: "learning sample validation data" should read "learning sample in validation data"
[0077]: "backpropagate feedback" should read "backpropagates feedback"
[0080]: "and combined vector." should read "and combined vector 284."
[0081]: "generates training dataset" should read "generates a training dataset"
Appropriate correction is required.
Claim Objections
Claims 2, 4-7, 9, and 11-20 are objected to because of the following informalities:
Claim 2: "1 further comprising" should read "1, further comprising"; "the foundation models" should read "the foundation model"
Claim 4: "comprises the number" should read "comprising the number"; “prompts tokens” should read “prompt tokens”
Claim 5: "4 further comprising" should read "4, further comprising"; "comprises numerical representation" should read "comprises a numerical representation"
Claim 6: "5 further comprising" should read "5, further comprising"
Claim 7: “6 further comprising” should read “6, further comprising”
Claim 9: "the foundation models are pre-trained general purpose models" should read "the foundation model is a pre-trained general purpose model"
Claim 11: "comprises the number" should read "comprising the number"; “prompts tokens” should read “prompt tokens”
Claim 12: "comprises numerical representation" should read "comprises a numerical representation"
Claim 15: "media, cause" should read "media, which cause"
Claim 16: "the foundation models are pre-trained general purpose models" should read "the foundation model is a pre-trained general purpose model"
Claim 17: “the program instructions, collectively stored in the set of one or more storage media, the operation performed by the processor set comprises” should read “the program instructions, collectively stored in the set of one or more storage media, further cause the processor set to:”
Claim 18: "comprises the number" should read "comprising the number"; “the program instructions, the operation performed by the processor set comprises” should read “the program instructions, collectively stored in the set of one or more storage media, further cause the processor set to”; “prompts tokens” should read “prompt tokens”
Claim 19: “comprises numerical representation" should read "comprises a numerical representation"
Claims 13, 14, and 20 are objected to due to dependency on an objected-to base claim.
Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Claim 1
Step 1: The claim is directed to a computer implemented method and thus is directed to the statutory category of processes.
Step 2A Prong One: The claim recites, inter alia:
“determining… patterns of data in a sample dataset to identify representative data from the sample dataset”: This limitation encompasses mentally determining patterns of data in a sample dataset to mentally identify representative data from the sample dataset.
“combining… the representative data with words for a number of tasks to generate a number of simple prompts, wherein each simple prompt in the number of simple prompts comprises a portion of the representative data and words for a task from the number of tasks”: This limitation encompasses mentally combining the representative data with words for a number of tasks to mentally generate a number of simple prompts.
Step 2A Prong Two: This judicial exception is not integrated into a practical application. The claim further recites that the method is “computer implemented” and that the determining and combining steps are performed “by the processor set”; however, these limitations amount to mere instructions to apply a judicial exception on a generic computer with a generic class of computer algorithms (MPEP 2106.05(f)). The claim additionally recites “training, by the processor set, the machine learning model using a training dataset comprising the number of simple prompts, wherein the machine learning model is trained to identify priorities of words in the number of simple prompts”; however this limitation amounts to generally linking the use of the judicial exception to the field of use/technological environment of model training (MPEP 2106.05(h)).
Step 2B: The claim does not contain significantly more than the judicial exception. The analysis at this step mirrors that of Step 2A Prong Two above. As an ordered whole, the claim is directed to a mentally performable process of determining patterns of data in a sample dataset to identify representative data from the sample dataset and combining the representative data with words for a number of tasks to generate a number of simple prompts. Nothing in the claim provides significantly more than this. As such, the claim is not patent eligible.
Claim 2
Step 1: A process, as above.
Step 2A Prong One: The claim recites, inter alia:
“modifying… an input prompt for a foundation model based on the priorities of words to perform the number of tasks for additional data, wherein the foundation models is a pre-trained general purpose model that can be used for performing tasks for data”: This limitation encompasses mentally modifying an input prompt for a foundation model based on the priorities of words, such as by mentally removing words from the prompt with low priorities.
Step 2A Prong Two: This judicial exception is not integrated into a practical application. The claim further recites that the modifying is performed “by the processor set using the machine learning model”; however, this limitation amounts to mere instructions to apply a judicial exception on a generic computer with a generic class of computer algorithms (MPEP 2106.05(f)).
Step 2B: The claim does not contain significantly more than the judicial exception. The analysis at this step mirrors that of Step 2A Prong Two above.
Claim 3
Step 1: A process, as above.
Step 2A Prong One: The claim recites, inter alia:
“selecting… a portion of tabular data from the sample dataset, wherein the portion of tabular data comprises no duplicated data”: This limitation encompasses mentally selecting a portion of tabular data from the sample dataset.
“clustering… data in columns for the portion of tabular data to generate a first set of clusters”: This limitation encompasses mentally clustering data in columns for the portion of tabular data to generate a first set of clusters, such as by mentally grouping similar data into clusters.
“identifying… patterns of data in the sample dataset based on the first set of clusters”: This limitation encompasses mentally identifying patterns of data in the sample dataset based on the first set of clusters.
“identifying…representative data from the sample dataset based on the patterns of data in the sample dataset”: This limitation encompasses mentally identifying representative data from the sample dataset based on the patterns of data in the sample dataset.
Step 2A Prong Two: This judicial exception is not integrated into a practical application. The claim further recites that the selecting, clustering, and identifying steps are performed “by the processor set”; however, this limitation amounts to mere instructions to apply a judicial exception on a generic computer with a generic class of computer algorithms (MPEP 2106.05(f)).
Step 2B: The claim does not contain significantly more than the judicial exception. The analysis at this step mirrors that of Step 2A Prong Two above.
Claim 4
Step 1: A process, as above.
Step 2A Prong One: The claim recites, inter alia:
“splitting…each simple prompt in the number of simple prompts into a number of prompt tokens, wherein each prompt token in the number of prompt tokens comprises a word or part of a word from a simple prompt from the number of simple prompts”: This limitation encompasses mentally splitting each single prompt in the number of single prompts into a number of prompt tokens.
“converting…each prompt token from the number of prompt tokens into a numerical vector”: This limitation encompasses mentally converting each prompt token from the number of prompt tokens into a numerical vector.
“determining…correlations between words in the number of prompt tokens using the numerical vectors for the number of prompts tokens”: This limitation encompasses mentally determining correlations between words in the number of prompt tokens using the numerical vectors for the number of prompt tokens.
Step 2A Prong Two: This judicial exception is not integrated into a practical application. The claim further recites that the splitting, converting, and determining steps are performed “by the processor set” and that the determining step is performed “using the machine learning model”; however, this limitation amounts to mere instructions to apply an exception on a generic computer using a generic class of computer algorithms (MPEP 2106.05(f)). The claim further recites “training, by the processor set, the machine learning model to identify priorities of words in each prompt token based on the correlations between words in the number of prompt tokens”; however, this limitation amounts to generally linking the judicial exception to the field of use/technological environment of model training (MPEP 2106.05(h)).
Step 2B: The claim does not contain significantly more than the judicial exception. The analysis at this step mirrors that of Step 2A Prong Two above.
Claim 5
Step 1: A process, as above.
Step 2A Prong One: The claim recites, inter alia:
“generating… a number of new prompt tokens for the number of simple prompts based on the correlations between words in the number of prompt tokens…”: This limitation encompasses mentally generating a number of new prompt tokens for the number of simple prompts based on the correlations between words in the number of prompt tokens.
“generating… a context vector for the number of simple prompts by combining the number of new prompt tokens, wherein the context vector comprises numerical representation for priority of words and pattern of words for all simple prompts in the number of simple prompts”: This limitation encompasses mentally generating a context vector for the number of simple prompts by mentally combining the number of new prompt tokens.
Step 2A Prong Two: This judicial exception is not integrated into a practical application. The claim further recites that generating a number of new prompt tokens is performed “by the processor set” and “using the machine learning model” and that generating a context vector is performed “by the processor set”; however, these limitations amount to mere instructions to apply an exception on a generic computer using a generic class of computer algorithms (MPEP 2106.05(f)).
Step 2B: The claim does not contain significantly more than the judicial exception. The analysis at this step mirrors that of Step 2A Prong Two above.
Claim 6
Step 1: A process, as above.
Step 2A Prong One: The claim recites, inter alia:
“generating… a number of validation prompts based on validation data from a validation dataset, wherein each validation prompt from the number of validation prompts comprises a portion of the validation data and context for a task from the number of tasks”: This limitation encompasses mentally generating a number of validation prompts based on validation data from a validation dataset.
“converting… each validation prompt from the number of validation prompts into a numerical vector”: This limitation encompasses mentally converting each validation prompt from the number of validation prompts into a numerical vector.
“combining…the context vector with the numerical vector for each validation prompt to form a number of combined vectors”: This limitation encompasses mentally combining the context vector with the numerical vector for each validation prompt.
“generating…an output…”: This limitation encompasses mentally generating an output.
Step 2A Prong Two: This judicial exception is not integrated into a practical application. The claim further recites that the generating, converting, and combining steps are performed “by the processor set” and that generating an output is performed “by inputting a combined vector from the number of combined vectors to a foundation model”; however, these limitations amount to mere instructions to apply an exception on a generic computer using a generic class of computer algorithms (MPEP 2106.05(f)).
Step 2B: The claim does not contain significantly more than the judicial exception. The analysis at this step mirrors that of Step 2A Prong Two above.
Claim 7
Step 1: A process, as above.
Step 2A Prong One: The claim recites, inter alia:
“generating… a feedback based on accuracy for the output by comparing the output to a ground truth for the validation prompt for the combined vector”: This limitation encompasses mentally generating a feedback based on accuracy for the output by mentally comparing the output to a ground truth for the validation prompt for the combined vector.
“adjusting… parameters for the machine learning model based on the feedback”: This limitation encompasses mentally adjusting parameters for the machine learning model based on the feedback.
Step 2A Prong Two: This judicial exception is not integrated into a practical application. The claim further recites that the generating and adjusting steps are performed “by the processor set”; however, these limitations amount to mere instructions to apply an exception on a generic computer using a generic class of computer algorithms (MPEP 2106.05(f)).
Step 2B: The claim does not contain significantly more than the judicial exception. The analysis at this step mirrors that of Step 2A Prong Two above.
Claims 8-14
Step 1: The claims are directed to a computer system and thus are directed to the statutory category of machines.
Step 2A Prong One: Claims 8-14 recite the same judicial exception as claims 1-7, respectively.
Step 2A Prong Two: This judicial exception is not integrated into a practical application. The analysis at this step mirrors that of claims 1-7, respectively, except insofar as claims 8-14 further recite “A computer system comprising: a processor set; a set of one or more computer-readable storage media; and program instructions, collectively stored in the set of one or more storage media, for causing the processor set to perform the following computer operations”; however, this limitation amounts to mere instructions to apply an exception on a generic computer using a generic class of computer algorithms (MPEP 2106.05(f)).
Step 2B: The claim does not contain significantly more than the judicial exception. The analysis at this step mirrors that of Step 2A Prong Two above.
Claims 15-20
Step 1: The claims are directed to a computer system and thus are directed to the statutory category of machines.
Step 2A Prong One: Claims 15-20 recite the same judicial exception as claims 1-6, respectively.
Step 2A Prong Two: This judicial exception is not integrated into a practical application. The analysis at this step mirrors that of claims 1-6, respectively, except insofar as claims 15-20 further recite “A computer system comprising: a processor set; a set of one or more computer-readable storage media; and program instructions, collectively stored in the set of one or more storage media, for causing the processor set to perform the following computer operations”; however, this limitation amounts to mere instructions to apply an exception on a generic computer using a generic class of computer algorithms (MPEP 2106.05(f)).
Step 2B: The claim does not contain significantly more than the judicial exception. The analysis at this step mirrors that of Step 2A Prong Two above.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1, 8, and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Barke et al. (Solving Data-centric Tasks using Large Language Models) (hereinafter “Barke”) in view of Jones et al. (US20240394549) (hereinafter “Jones”).
Regarding claim 1, Barke discloses “A… method for training a machine learning model for managing prompts, the… method comprising:
determining… patterns of data in a sample dataset to identify representative data from the sample dataset (Barke, 3.3 Cluster-then-select prompting technique: “To solve tasks on large tables, we propose a cluster-then-select technique which prompts the model with a representative sample of the input data. In order to capture the syntactic variation in the input data, we rely on an existing tool (Padhi et al., 2018), which takes as input a set of strings and synthesizes a small set of regular expressions (regexes), such that each input string matches one of the regexes… These regexes are then used to cluster the input strings, and we select some number of rows from each cluster. In Figure 1, we pick one row from each of the four distinct clusters (depicted with different colors)”; Examiner notes that clustering the input data corresponds to “determining patterns of data” and the selected rows from the clusters correspond to “representative data”);
combining… the representative data with words for a number of tasks to generate a number of simple prompts, wherein each simple prompt in the number of simple prompts comprises a portion of the representative data and words for a task from the number of tasks (Barke, C Our Prototype Tool: “For our running example, k is 1, Q is “create a new column in lowercase that concatenates the first initial and the last name.”, and T is Data({"Names":["John Smith", "Jack Will Anders", ...]}). At a high-level, the algorithm first clusters the data in T based on automatically synthesized regular expressions and stores them in a map M (line 2). It then extracts representative rows of the table using SELECT (line 3); combines the query Q and the rows R to create a prompt P using PROMPT (line 4)”; Examiner notes that query Q corresponds to “words for a number of tasks,” rows R corresponds to “the representative data,” and prompt P corresponds to “a number of simple prompts”).
Barke does not appear to explicitly disclose the further limitations of the claim.
However, Jones discloses “a computer-implemented method” and “a processor set” (Jones, [0040]: “The machine 400 may include processors 404 (including processors 408 and 412), memory/storage 406, and I/O components 418, which may be configured to communicate with each other such as via a bus 402. The memory/storage 406 may include a memory 414, such as a main memory, or other memory storage, and a storage unit 416, both accessible to the processors 404 such as via the bus 402. The storage unit 416 and memory 414 store the instructions 410 embodying any one or more of the methodologies or functions described herein”); and
“training, by the processor set… [a] machine learning model using a training dataset comprising… [a] number of simple prompts, wherein the machine learning model is trained to identify priorities of words in the number of simple prompts” (Jones, [0062] : “The tuning service may train a tuned LM by providing the training prompts in the training data 712 to the pre-trained LM” and [0065]: “The trained parameters of the self-attention hidden layers 812 enables the LM to weight the importance of each token in input text 802 sequence while taking into account the relationships between all tokens and token dependencies”).
Jones and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified Barke with the teachings of Jones such that the method is “computer-implemented,” the steps are performed “by the processor set,” and the method further includes “training, by the processor set, the machine learning model using a training dataset comprising the number of simple prompts, wherein the machine learning model is trained to identify priorities of words in the number of simple prompts,” and one would have been motivated to do so. Doing so would allow the model to learn the context and meaning of words included in the simple prompts, improving the performance of the model on downstream tasks (see Jones, [0068] and [0072]).
Regarding claim 8, Barke discloses
“determine patterns of data in a sample dataset to identify representative data from the sample dataset (Barke, 3.3 Cluster-then-select prompting technique: “To solve tasks on large tables, we propose a cluster-then-select technique which prompts the model with a representative sample of the input data. In order to capture the syntactic variation in the input data, we rely on an existing tool (Padhi et al., 2018), which takes as input a set of strings and synthesizes a small set of regular expressions (regexes), such that each input string matches one of the regexes… These regexes are then used to cluster the input strings, and we select some number of rows from each cluster. In Figure 1, we pick one row from each of the four distinct clusters (depicted with different colors)”; Examiner notes that clustering the input data corresponds to “determining patterns of data” and the selected rows from the clusters correspond to “representative data”);
combine the representative data with words for a number of tasks to generate a number of simple prompts, wherein each simple prompt in the number of simple prompts comprises a portion of the representative data and words for a task from the number of tasks (Barke, C Our Prototype Tool: “For our running example, k is 1, Q is “create a new column in lowercase that concatenates the first initial and the last name.”, and T is Data({"Names":["John Smith", "Jack Will Anders", ...]}). At a high-level, the algorithm first clusters the data in T based on automatically synthesized regular expressions and stores them in a map M (line 2). It then extracts representative rows of the table using SELECT (line 3); combines the query Q and the rows R to create a prompt P using PROMPT (line 4)”; Examiner notes that query Q corresponds to “words for a number of tasks,” rows R corresponds to “the representative data,” and prompt P corresponds to “a number of simple prompts”).
Barke does not appear to explicitly disclose the further limitations of the claim.
However, Jones discloses “A computer system comprising: a processor set; a set of one or more computer-readable storage media; and program instructions, collectively stored in the set of one or more storage media, for causing the processor set to perform the following computer operations: (Jones, [0040]: “The machine 400 may include processors 404 (including processors 408 and 412), memory/storage 406, and I/O components 418, which may be configured to communicate with each other such as via a bus 402. The memory/storage 406 may include a memory 414, such as a main memory, or other memory storage, and a storage unit 416, both accessible to the processors 404 such as via the bus 402. The storage unit 416 and memory 414 store the instructions 410 embodying any one or more of the methodologies or functions described herein”) …
train a machine learning model using a training dataset comprising… [a] number of simple prompts, wherein the machine learning model is trained to identify priorities of words in the number of simple prompts” (Jones, [0062] : “The tuning service may train a tuned LM by providing the training prompts in the training data 712 to the pre-trained LM” and [0065]: “The trained parameters of the self-attention hidden layers 812 enables the LM to weight the importance of each token in input text 802 sequence while taking into account the relationships between all tokens and token dependencies”).
Jones and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified Barke with the teachings of Jones such that the method is performed by “A computer system comprising: a processor set; a set of one or more computer-readable storage media; and program instructions, collectively stored in the set of one or more storage media, for causing the processor set to perform the following computer operations” and the method further includes “train the machine learning model using a training dataset comprising the number of simple prompts, wherein the machine learning model is trained to identify priorities of words in the number of simple prompts,” and one would have been motivated to do so. Doing so would allow the model to learn the context and meaning of words included in the simple prompts, improving the performance of the model on downstream tasks (see Jones, [0068] and [0072]).
Regarding claim 15, Barke discloses
“determine patterns of data in a sample dataset to identify representative data from the sample dataset (Barke, 3.3 Cluster-then-select prompting technique: “To solve tasks on large tables, we propose a cluster-then-select technique which prompts the model with a representative sample of the input data. In order to capture the syntactic variation in the input data, we rely on an existing tool (Padhi et al., 2018), which takes as input a set of strings and synthesizes a small set of regular expressions (regexes), such that each input string matches one of the regexes… These regexes are then used to cluster the input strings, and we select some number of rows from each cluster. In Figure 1, we pick one row from each of the four distinct clusters (depicted with different colors)”; Examiner notes that clustering the input data corresponds to “determining patterns of data” and the selected rows from the clusters correspond to “representative data”);
combine the representative data with words for a number of tasks to generate a number of simple prompts, wherein each simple prompt in the number of simple prompts comprises a portion of the representative data and words for a task from the number of tasks (Barke, C Our Prototype Tool: “For our running example, k is 1, Q is “create a new column in lowercase that concatenates the first initial and the last name.”, and T is Data({"Names":["John Smith", "Jack Will Anders", ...]}). At a high-level, the algorithm first clusters the data in T based on automatically synthesized regular expressions and stores them in a map M (line 2). It then extracts representative rows of the table using SELECT (line 3); combines the query Q and the rows R to create a prompt P using PROMPT (line 4)”; Examiner notes that query Q corresponds to “words for a number of tasks,” rows R corresponds to “the representative data,” and prompt P corresponds to “a number of simple prompts”).
Barke does not appear to explicitly disclose the further limitations of the claim.
However, Jones discloses “A computer program product for training a machine learning model for managing prompts, the computer program product comprising: a set of one or more computer-readable storage media; and program instructions, collectively stored in the set of one or more storage media, cause a processor set to perform the following computer operations: (Jones, [0040]: “The machine 400 may include processors 404 (including processors 408 and 412), memory/storage 406, and I/O components 418, which may be configured to communicate with each other such as via a bus 402. The memory/storage 406 may include a memory 414, such as a main memory, or other memory storage, and a storage unit 416, both accessible to the processors 404 such as via the bus 402. The storage unit 416 and memory 414 store the instructions 410 embodying any one or more of the methodologies or functions described herein”) …
train the machine learning model using a training dataset comprising… [a] number of simple prompts, wherein the machine learning model is trained to identify priorities of words in the number of simple prompts” (Jones, [0062] : “The tuning service may train a tuned LM by providing the training prompts in the training data 712 to the pre-trained LM” and [0065]: “The trained parameters of the self-attention hidden layers 812 enables the LM to weight the importance of each token in input text 802 sequence while taking into account the relationships between all tokens and token dependencies”).
Jones and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified Barke with the teachings of Jones such that the method is performed by “A computer program product for training a machine learning model for managing prompts, the computer program product comprising: a set of one or more computer-readable storage media; and program instructions, collectively stored in the set of one or more storage media, cause a processor set to perform the following computer operations:” and the method further includes “train the machine learning model using a training dataset comprising the number of simple prompts, wherein the machine learning model is trained to identify priorities of words in the number of simple prompts,” and one would have been motivated to do so. Doing so would allow the model to learn the context and meaning of words included in the simple prompts, improving the performance of the model on downstream tasks (see Jones, [0068] and [0072]).
Claims 2, 9, and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Barke in view of Jones, and further in view of Kong et al. (Prewrite: Prompt Rewriting with Reinforcement Learning) (hereinafter “Kong”) and Shim et al. (US20250131023) (hereinafter “Shim”).
Regarding claim 2, the rejection of claim 1 is incorporated. Neither Barke nor Jones appear to explicitly disclose the further limitations of the claim.
However, Kong discloses “modifying… using… [a] machine learning model, an input prompt for a foundation model… to perform… [a] number of tasks for additional data, wherein the foundation models is a pre-trained general purpose model that can be used for performing tasks for data” (2.2 Overview: “First, the prompt rewriter takes in an initial prompt p and rewrites it to another prompt p†. The initial prompt is usually crafted manually and can be sub-optimal. Observing the remarkable capability of LLMs, we instruct a LLM (e.g., PaLM 2-S) with a meta prompt m for rewriting as follows: p† = LLMR("{m}\nInstruction: {p}"). We call LLMR, rewriter LLM, which is to be differentiated from the task LLM, used for the end task. We list our meta prompts in Appendix B. Second, the rewritten prompt p† is then used by the task LLM to generate the task output. The task LLM is assumed to be a blackbox accessed via API and can be larger than the rewriter LLM”; Examiner notes that rewriter LLM corresponds to “a machine learning model”, task LLM corresponds to “a foundation model,” and the task performed by the task LLM corresponds to “a number of tasks for additional data” (see examples in Table 9)).
Kong and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Barke/Jones with the teachings of Kong such that the method further includes “modifying, by the processor set using the machine learning model, an input prompt for a foundation model…wherein the foundation models is a pre-trained general purpose model that can be used for performing tasks for data,” and one would have been motivated to do so. Doing so would improve the effectiveness of the initial input prompt (see Kong, 3.2 Results).
Kong does not appear to explicitly disclose that the modifying is “based on the priorities of words”.
However, Shim discloses “modifying… an input prompt… based on… priorities of words” ([0048]: “The term “word-level attention scores” is used herein to describe quantifiable values that may indicate the degree of importance or relevance assigned to individual tokens or words within a text sequence during the processing stages of AI models, such as LXMs” and [0076]: “In some embodiments, the components may be configured to use word-level attention scores to streamline text. In such embodiments, the components may be configured to reduce the length of the text prompt by prioritizing or focusing on the terms that are associated with higher attention scores”).
Shim and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Barke/Jones/Kong with the teachings of Shim such that the modifying is “based on the priorities of words,” and one would have been motivated to do so. Doing so would improve performance and power consumption of the computing device while retaining key elements of the prompt (see Shim, [0054]).
Regarding claim 9, the rejection of claim 8 is incorporated. Claim 9 is a computer system claim corresponding to method claim 2, thus the rejection of claim 9 follows the same rationale as the rejection of claim 2 above.
Regarding claim 16, the rejection of claim 15 is incorporated. Claim 16 is a computer program product claim corresponding to method claim 2, thus the rejection of claim 16 follows the same rationale as the rejection of claim 2 above.
Claims 4, 11, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Barke in view of Jones, and further in view of Omidi et al. (US20250232162) (hereinafter “Omidi”).
Regarding claim 4, the rejection of claim 1 is incorporated. Barke as modified by Jones further discloses “wherein training, by the processor set, a machine learning model using a training dataset comprises the number of simple prompts comprises:
splitting, by the processor set, each simple prompt in the number of simple prompts into a number of prompt tokens, wherein each prompt token in the number of prompt tokens comprises a word or part of a word from a simple prompt from the number of simple prompts (Jones, [0068]: “To train the pre-trained LM 661, each training prompt included in the training data may be tokenized and the sequence of tokens determined for each prompt may be fed into the pre-trained LM” and [0063]: “In various embodiments the encoder nodes 822 may tokenize the input text 802 into a sequence of tokens (e.g., individual words of sub-words)”);
converting, by the processor set, each prompt token from the number of prompt tokens into a numerical vector (Jones, [0068]: “The encoder layers of the pre-trained LM may determine input embeddings and positional encodings for each token in a sequence”);
…
training, by the processor set, the machine learning model to identify priorities of words in each prompt token based on the correlations between words in the number of prompt tokens” (Jones, [0069]: “In various embodiments, for each sample in the training dataset, the pre-trained LM 661 may be trained over multiple iterations” and Jones, [0065]: “The trained parameters of the self-attention hidden layers 812 enables the LM to weight the importance of each token in input text 802 sequence while taking into account the relationships between all tokens and token dependencies”).
Jones and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified Barke with the teachings of Jones such that training, by the processor set, a machine learning model using a training dataset comprises the number of simple prompts comprises: “splitting, by the processor set, each simple prompt in the number of simple prompts into a number of prompt tokens, wherein each prompt token in the number of prompt tokens comprises a word or part of a word from a simple prompt from the number of simple prompts; converting, by the processor set, each prompt token from the number of prompt tokens into a numerical vector… and training, by the processor set, the machine learning model to identify priorities of words in each prompt token based on the correlations between words in the number of prompt tokens” and one would have been motivated to do so. Doing so would allow the model to learn the context and meaning of words included in the simple prompts, improving the performance of the model on downstream tasks (see Jones, [0068] and [0072]).
Neither Barke nor Jones appear to explicitly disclose the further limitations of the claim.
However, Omidi discloses “determining… using… [a] machine learning model, correlations between words in… [a] number of prompt tokens using… numerical vectors for the number of prompts tokens” (Omidi, [0113]: “In some more recent approaches to attention, each token is converted to a Key vector, a Query vector, and a Value vector, by different learned weight matrices. The dot product between a Query vector for a given token and a Key vector for another token produces a value which is proportional to the importance of the relationship between the given token and the other token”).
Omidi and the instant application both relate to machine learning and are analogous. It would have been obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to have modified the combination of Barke/Jones with the teachings of Omidi to include “determining, by the processor set using the machine learning model, correlations between words in the number of prompt tokens using the numerical vectors for the number of prompts tokens”, and one would have been motivated to do so. Doing so would allow the model to capture token dependencies within a sequence of tokens, improving its ability to generate a relevant output sequence (see Omidi, [0111]).
Regarding claim 11, the rejection of claim 8 is incorporated. Claim 11 is a computer system claim corresponding to method claim 4, thus the rejection of claim 11 follows the same rationale as the rejection of claim 4 above.
Regarding claim 18, the rejection of claim 15 is incorporated. Claim 18 is a computer program product claim corresponding to method claim 4, thus the rejection of claim 18 follows the same rationale as the rejection of claim 4 above.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Pool-Search-Demonstrate: Improving Data-Wrangling LLMs via Better In-Context Examples (Huh et al.): Huh et al. discloses generating prompts for prompting foundation models to perform data wrangling tasks by combining a task description with task demonstrations.
US20210027503 (Kumaresan et al.): Kumaresan et al. discloses selecting representative samples from tabular data.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to GWYNEVERE A DETERDING whose telephone number is (571)272-7657. The examiner can normally be reached Mon-Fri. 9am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kamran Afshar can be reached at (571) 272-7796. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/G.A.D./Examiner, Art Unit 2125
/KAMRAN AFSHAR/Supervisory Patent Examiner, Art Unit 2125