Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This action is in response to the application and claims filed 01/25/2024. Claims 1-16 are pending and have been examined. Claims 1-16 are rejected.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 01/25/2024 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Objections
Claims 3 and 11 objected to because of the following informalities: "the ground truth completeness graphs associated with instructions as labels" should read "the ground truth completeness graphs associated with the instructions as labels." Appropriate correction is required.
Claims 5 and 13 objected to because of the following informalities: in claim 5, "comprises determining a Kullback-Leibler (KL) divergence" should read "comprises determining the divergence as a Kullback-Leibler (KL) divergence" for consistency with "a divergence" recited in claim 4; in claim 13, "by being configured to determine a Kullback-Leibler (KL) divergence" should read "by being configured to determine the divergence as a Kullback-Leibler (KL) divergence" for consistency with "a divergence" recited in claim 12. Appropriate correction is required.
Claims 8 and 16 objected to because of the following informalities: "the validity of one or more of a file type, node types, and edges of the generated completeness graph" should read "the validity of one or more of a file type, node types, or edges of the generated completeness graph." Appropriate correction is required.
Claims 9-16 objected to because of the following informalities: the recited operations are set forth in the imperative rather than the -ing form, e.g., in claim 9, "obtain a dataset…," "produce a generated completeness graph…," "evaluate the generated completeness graph…," and "re-train the active large language model…" should read "obtaining…," "producing…," "evaluating…," and "re-training…." Claims 10-16 recite the same informality and require similar correction. Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-16 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract
idea without significantly more.
Claim 1
Step 1: The claim recites a method; therefore, it is directed to the statutory category of process.
Step 2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the "Mental Processes" grouping of abstract ideas. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation by mathematical calculation but for the recitation of generic computer components, then it falls within the "Mathematical Concepts" grouping of abstract ideas. The claim recites the following abstract ideas:
• “generate completeness graphs” (This limitation is a mental process. A person mentally or with a pen and paper can produce, i.e., draw, a completeness graph)
• "produce a generated completeness graph (...) in response to a query based on instructions from the dataset;" (This limitation is a mental process. A person mentally or with a pen and paper can produce, i.e., draw, a completeness graph in response to a query based on instructions. The specification acknowledges that completeness graphs are conventionally generated manually by domain experts based on instructions, e.g., forms, rules, and regulations.)
• "evaluate the generated completeness graph (...) to produce a reward based on validity of the generated completeness graph and semantic similarity of the generated completeness graph and the associated ground truth completeness graph; and" (This limitation is a mental process. A person mentally or with a pen and paper can compare a generated completeness graph to the associated ground truth completeness graph, judge whether the generated completeness graph is valid and how similar the two graphs are, and assign a score, i.e., a reward, based on the evaluation.)
Step 2A Prong 2: This judicial exception is not integrated into a practical application because the additional elements are as follows:
• "A method of training a large language model to (...), comprising:" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim recites a generic, off-the-shelf large language model as a tool with which the recited abstract ideas are performed.)
• "obtaining a dataset comprising instructions and ground truth completeness graphs associated with the instructions;" (Data Gathering - Mere data gathering recited at a high level of generality, and thus is insignificant extra-solution activity (MPEP 2106.05(g)).)
• "...with an active large language model..." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim only recites a generic, off-the-shelf large language model, recited at a high level of generality, that is used as a tool to perform the abstract idea of producing a completeness graph. Therefore, it amounts to no more than mere instructions to apply the exception using a generic computer component.)
• "...with a reward model..." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim only recites a generic reward model, recited at a high level of generality, that is used as a tool to perform the abstract idea of evaluating the generated completeness graph to produce the reward.)
• "re-training the active large language model based at least partially on the reward." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim denotes generic re-training and a generic large language model with no additional details or limitations beyond a generic, off-the-shelf large language model.)
Step 2B:
• "A method of training a large language model to generate completeness graphs, comprising:" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim recites a generic, off-the-shelf large language model as a tool with which the recited abstract ideas are performed.)
• "obtaining a dataset comprising instructions and ground truth completeness graphs associated with the instructions;" (MPEP 2106.05(d)(II) indicates that merely gathering data is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
• "...with an active large language model..." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim only recites a generic, off-the-shelf large language model, recited at a high level of generality, that is used as a tool to perform the abstract idea of producing a completeness graph. Therefore, it amounts to no more than mere instructions to apply the exception using a generic computer component.)
• "...with a reward model..." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim only recites a generic reward model, recited at a high level of generality, that is used as a tool to perform the abstract idea of evaluating the generated completeness graph to produce the reward.)
• "re-training the active large language model based at least partially on the reward." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim denotes generic re-training and a generic large language model with no additional details or limitations beyond a generic, off-the-shelf large language model.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 2
Step 1: A process, as above.
Step 2A Prong 1: See the rejection of Claim 1 above, which claim 2 depends on.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
• "training a reference large language model to receive instructions from the dataset and in response to produce completeness graphs," (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim denotes generic training and a generic large language model (the reference large language model) with no additional details or limitations beyond a generic, off-the-shelf large language model that receives data and produces an output.)
• "wherein the active large language model is initialized based on the reference large language model." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): This limitation amounts to initializing (copying) one generic model from another generic model, which is merely an instruction to apply the exception using generic computer components.)
Step 2B:
• "training a reference large language model to receive instructions from the dataset and in response to produce completeness graphs," (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim denotes generic training and a generic large language model (the reference large language model) with no additional details or limitations beyond a generic, off-the-shelf large language model that receives data and produces an output.)
• "wherein the active large language model is initialized based on the reference large language model." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): This limitation amounts to initializing (copying) one generic model from another generic model, which is merely an instruction to apply the exception using generic computer components.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 3
Step 1: A process, as above.
Step 2A Prong 1: See the rejection of Claim 2 above, which claim 3 depends on.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
• "wherein training the reference large language model comprises supervised training using instructions from the dataset as prompts and the ground truth completeness graphs associated with instructions as labels." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim denotes generic supervised training of a generic, off-the-shelf large language model using the gathered data as prompts and labels, with no additional details or limitations beyond a generic, off-the-shelf training technique.)
Step 2B:
• "wherein training the reference large language model comprises supervised training using instructions from the dataset as prompts and the ground truth completeness graphs associated with instructions as labels." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim denotes generic supervised training of a generic, off-the-shelf large language model using the gathered data as prompts and labels, with no additional details or limitations beyond a generic, off-the-shelf training technique.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 4
Step 1: A process, as above.
Step 2A Prong 1: See the rejection of Claim 2 above, which claim 4 depends on. Claim 4 further recites:
• "producing a second generated completeness graph (...) in response to the query based on the instructions from the dataset; and" (This limitation is a mental process. A person mentally or with a pen and paper can produce, i.e., draw, a second completeness graph in response to the query based on the instructions.)
• "comparing the generated completeness graph (...) to the second generated completeness graph (...) to determine a divergence;" (This limitation is a mental process. A person mentally or with a pen and paper can compare two completeness graphs to determine a divergence, i.e., how much the two graphs differ from one another.)
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
• "...with the reference large language model...", "...produced with the active large language model...", and "...produced with the reference large language model..." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim only recites generic, off-the-shelf large language models, recited at a high level of generality, that are used as tools to perform the recited abstract ideas.)
• "wherein the re-training of the active large language model is further based at least partially on the divergence." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim denotes generic re-training of a generic, off-the-shelf large language model based on the result (the divergence) of the recited abstract idea, with no additional details or limitations beyond a generic, off-the-shelf large language model.)
Step 2B:
• "...with the reference large language model...", "...produced with the active large language model...", and "...produced with the reference large language model..." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim only recites generic, off-the-shelf large language models, recited at a high level of generality, that are used as tools to perform the recited abstract ideas.)
• "wherein the re-training of the active large language model is further based at least partially on the divergence." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim denotes generic re-training of a generic, off-the-shelf large language model based on the result (the divergence) of the recited abstract idea, with no additional details or limitations beyond a generic, off-the-shelf large language model.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 5
Step 1: A process, as above.
Step 2A Prong 1: See the rejection of Claim 4 above, which claim 5 depends on. Claim 5 further recites:
• "wherein comparing the generated completeness graph (...) to the second generated completeness graph (...) comprises determining a Kullback-Leibler (KL) divergence." (This limitation falls within the mathematical concepts grouping because determining a Kullback-Leibler (KL) divergence involves calculating a divergence value by evaluating a mathematical formula (see e.g., the KL-divergence term in paragraph [0076] of the specification).)
Step 2A Prong 2: The claim does not recite additional elements that integrate the judicial exception into a practical application. -- Examiner's Note (EN): The recited "active large language model" and "reference large language model" are addressed in the rejections of claims 1 and 4 above.
Step 2B: The claim does not recite additional elements that amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 6
Step 1: A process, as above.
Step 2A Prong 1: See the rejection of Claim 4 above, which claim 6 depends on. Claim 6 further recites:
• "determining a loss based on the reward and the divergence," (This limitation falls within the mathematical concepts grouping because it involves calculating a loss value based on the reward value and the divergence value using a mathematical loss function (see e.g., Equation 1 in paragraph [0077] of the specification).)
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
• "wherein the re-training of the active large language model is based on the loss." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim denotes generic re-training and a generic large language model with no additional details or limitations beyond a generic, off-the-shelf large language model.)
Step 2B:
• "wherein the re-training of the active large language model is based on the loss." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim denotes generic re-training and a generic large language model with no additional details or limitations beyond a generic, off-the-shelf large language model.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 7
Step 1: A process, as above.
Step 2A Prong 1: See the rejection of Claim 6 above, which claim 7 depends on. Claim 7 further recites:
• "...performing a Proximal Policy Optimization (PPO) algorithm." (This limitation falls within the mathematical concepts grouping because the Proximal Policy Optimization (PPO) algorithm is a mathematical optimization algorithm that optimizes a loss function through mathematical calculations (see e.g., paragraph [0077] of the specification, describing optimizing the loss function of Equation 1 using the PPO algorithm).)
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
• "wherein the re-training of the active large language model based on the loss comprises..." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The limitation amounts to using the recited mathematical algorithm to re-train a generic, off-the-shelf large language model, which is merely an instruction to apply the abstract idea using a generic computer component.)
Step 2B:
• "wherein the re-training of the active large language model based on the loss comprises..." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The limitation amounts to using the recited mathematical algorithm to re-train a generic, off-the-shelf large language model, which is merely an instruction to apply the abstract idea using a generic computer component.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 8
Step 1: A process, as above.
Step 2A Prong 1: See the rejection of Claim 1 above, which claim 8 depends on. Claim 8 further recites:
• "determining the validity of the generated completeness graph based on the validity of one or more of a file type, node types, and edges of the generated completeness graph;" (This limitation is a mental process. A person mentally or with a pen and paper can inspect a completeness graph and determine whether the graph is valid based on its file type, its node types, and its edges, e.g., by checking that each node represents a decision point and that each node is properly connected by edges (see e.g., paragraph [0062] of the specification).)
• "determining the semantic similarity of the generated completeness graph and the associated ground truth completeness graph based on a graph edit distance between the generated completeness graph and the associated ground truth completeness graph; and" (This limitation falls within the mathematical concepts grouping because it involves calculating a graph edit distance, which is a mathematical calculation of the minimum cost of an edit path (sequence of node and edge operations) to transform one graph into a graph isomorphic to the other graph (see e.g., paragraph [0061] of the specification).)
• "converting the validity and the semantic similarity to a scalar value as the reward." (This limitation falls within the mathematical concepts grouping because it involves mathematically combining values into a single scalar value, e.g., an average, a sum, or a weighted sum (see e.g., paragraph [0062] of the specification).)
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
• "wherein the evaluating the generated completeness graph with the reward model to produce the reward comprises:" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The "reward model" is a generic model, recited at a high level of generality, that is used as a tool to perform the recited abstract ideas.)
Step 2B:
• "wherein the evaluating the generated completeness graph with the reward model to produce the reward comprises:" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The "reward model" is a generic model, recited at a high level of generality, that is used as a tool to perform the recited abstract ideas.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 9
Step 1: The claim recites a system; therefore, it is directed to the statutory category of machine.
Step 2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the "Mental Processes" grouping of abstract ideas. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation by mathematical calculation but for the recitation of generic computer components, then it falls within the "Mathematical Concepts" grouping of abstract ideas. The claim recites the following abstract ideas:
• “generate completeness graphs” (This limitation is a mental process. A person mentally or with a pen and paper can produce, i.e., draw, a completeness graph)
• "produce a generated completeness graph (...) in response to a query based on instructions from the dataset;" (This limitation is a mental process. A person mentally or with a pen and paper can produce, i.e., draw, a completeness graph in response to a query based on instructions. The specification acknowledges that completeness graphs are conventionally generated manually by domain experts based on instructions, e.g., forms, rules, and regulations.)
• "evaluate the generated completeness graph (...) to produce a reward based on validity of the generated completeness graph and semantic similarity of the generated completeness graph and the associated ground truth completeness graph; and" (This limitation is a mental process. A person mentally or with a pen and paper can compare a generated completeness graph to the associated ground truth completeness graph, judge whether the generated completeness graph is valid and how similar the two graphs are, and assign a score, i.e., a reward, based on the evaluation.)
Step 2A Prong 2: This judicial exception is not integrated into a practical application because the additional elements are as follows:
• "A system of training a large language model to generate completeness graphs, comprising: one or more processors; and a memory coupled to the one or more processors and storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim recites a generic, off-the-shelf computer (one or more processors and a memory storing instructions) and a generic large language model as tools to perform the recited abstract ideas.)
• "obtain a dataset comprising instructions and ground truth completeness graphs associated with the instructions;" (Data Gathering - Mere data gathering recited at a high level of generality, and thus is insignificant extra-solution activity (MPEP 2106.05(g)).)
• "...with an active large language model..." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim only recites a generic, off-the-shelf large language model, recited at a high level of generality, that is used as a tool to perform the abstract idea of producing a completeness graph. Therefore, it amounts to no more than mere instructions to apply the exception using a generic computer component.)
• "...with a reward model..." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim only recites a generic reward model, recited at a high level of generality, that is used as a tool to perform the abstract idea of evaluating the generated completeness graph to produce the reward.)
• "re-train the active large language model based at least partially on the reward." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim denotes generic re-training and a generic large language model with no additional details or limitations beyond a generic, off-the-shelf large language model.)
Step 2B:
• "A system of training a large language model to generate completeness graphs, comprising: one or more processors; and a memory coupled to the one or more processors and storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim recites a generic, off-the-shelf computer (one or more processors and a memory storing instructions) and a generic large language model as tools to perform the recited abstract ideas.)
• "obtain a dataset comprising instructions and ground truth completeness graphs associated with the instructions;" (MPEP 2106.05(d)(II) indicates that merely gathering data is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer.)
• "...with an active large language model..." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim only recites a generic, off-the-shelf large language model, recited at a high level of generality, that is used as a tool to perform the abstract idea of producing a completeness graph. Therefore, it amounts to no more than mere instructions to apply the exception using a generic computer component.)
• "...with a reward model..." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim only recites a generic reward model, recited at a high level of generality, that is used as a tool to perform the abstract idea of evaluating the generated completeness graph to produce the reward.)
• "re-train the active large language model based at least partially on the reward." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim denotes generic re-training and a generic large language model with no additional details or limitations beyond a generic, off-the-shelf large language model.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 10
Step 1: A machine, as above.
Step 2A Prong 1: See the rejection of Claim 9 above, which claim 10 depends on.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
• "wherein the one or more processors are further configured to perform operations comprising..." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim recites generic, off-the-shelf processors as tools to perform the recited abstract ideas.)
• "train a reference large language model to receive instructions from the dataset and in response to produce completeness graphs," (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim denotes generic training and a generic large language model (the reference large language model) with no additional details or limitations beyond a generic, off-the-shelf large language model that receives data and produces an output.)
• "wherein the active large language model is initialized based on the reference large language model." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): This limitation amounts to initializing (copying) one generic model from another generic model, which is merely an instruction to apply the exception using generic computer components.)
Step 2B:
• "wherein the one or more processors are further configured to perform operations comprising..." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim recites generic, off-the-shelf processors as tools to perform the recited abstract ideas.)
• "train a reference large language model to receive instructions from the dataset and in response to produce completeness graphs," (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim denotes generic training and a generic large language model (the reference large language model) with no additional details or limitations beyond a generic, off-the-shelf large language model that receives data and produces an output.)
• "wherein the active large language model is initialized based on the reference large language model." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): This limitation amounts to initializing (copying) one generic model from another generic model, which is merely an instruction to apply the exception using generic computer components.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 11
Step 1: A machine, as above.
Step 2A Prong 1: See the rejection of Claim 10 above, which claim 11 depends on.
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
• "wherein the one or more processors are configured to perform the operation of train the reference large language model by being configured to supervised train using instructions from the dataset as prompts and the ground truth completeness graphs associated with instructions as labels." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim recites generic, off-the-shelf processors as tools to perform generic supervised training of a generic, off-the-shelf large language model using the gathered data as prompts and labels, with no additional details or limitations beyond a generic, off-the-shelf training technique.)
Step 2B:
• "wherein the one or more processors are configured to perform the operation of train the reference large language model by being configured to supervised train using instructions from the dataset as prompts and the ground truth completeness graphs associated with instructions as labels." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim recites generic, off-the-shelf processors as tools to perform generic supervised training of a generic, off-the-shelf large language model using the gathered data as prompts and labels, with no additional details or limitations beyond a generic, off-the-shelf training technique.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 12
Step 1: A machine, as above.
Step 2A Prong 1: See the rejection of Claim 10 above, which claim 12 depends on. Claim 12 further recites:
• "produce a second generated completeness graph (...) in response to the query based on the instructions from the dataset; and" (This limitation is a mental process. A person mentally or with a pen and paper can produce, i.e., draw, a second completeness graph in response to the query based on the instructions.)
• "compare the generated completeness graph (...) to the second generated completeness graph (...) to determine a divergence;" (This limitation is a mental process. A person mentally or with a pen and paper can compare two completeness graphs to determine a divergence, i.e., how much the two graphs differ from one another.)
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
• "wherein the one or more processors are further configured to perform operations comprising:" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim recites generic, off-the-shelf processors as tools to perform the recited abstract ideas.)
• "...with the reference large language model...", "...produced with the active large language model...", and "...produced with the reference large language model..." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim only recites generic, off-the-shelf large language models, recited at a high level of generality, that are used as tools to perform the recited abstract ideas.)
• "wherein the re-training of the active large language model is further based at least partially on the divergence." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim denotes generic re-training of a generic, off-the-shelf large language model based on the result (the divergence) of the recited abstract idea, with no additional details or limitations beyond a generic, off-the-shelf large language model.)
Step 2B:
• "wherein the one or more processors are further configured to perform operations comprising:" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim recites generic, off-the-shelf processors as tools to perform the recited abstract ideas.)
• "...with the reference large language model...", "...produced with the active large language model...", and "...produced with the reference large language model..." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim only recites generic, off-the-shelf large language models, recited at a high level of generality, that are used as tools to perform the recited abstract ideas.)
• "wherein the re-training of the active large language model is further based at least partially on the divergence." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim denotes generic re-training of a generic, off-the-shelf large language model based on the result (the divergence) of the recited abstract idea, with no additional details or limitations beyond a generic, off-the-shelf large language model.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 13
Step 1: A machine, as above.
Step 2A Prong 1: See the rejection of Claim 12 above, which claim 13 depends on. Claim 13 further recites:
• "...determine a Kullback-Leibler (KL) divergence." (This limitation falls within the mathematical concepts grouping because determining a Kullback-Leibler (KL) divergence involves calculating a divergence value by evaluating a mathematical formula (see e.g., the KL-divergence term in paragraph [0076] of the specification).)
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
• "wherein the one or more processors are configured to perform the operation of compare the generated completeness graph (...) to the second generated completeness graph (...) by being configured to..." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim recites generic, off-the-shelf processors and generic, off-the-shelf large language models as tools to perform the recited abstract idea.)
Step 2B:
• "wherein the one or more processors are configured to perform the operation of compare the generated completeness graph (...) to the second generated completeness graph (...) by being configured to..." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim recites generic, off-the-shelf processors and generic, off-the-shelf large language models as tools to perform the recited abstract idea.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 14
Step 1: A machine, as above.
Step 2A Prong 1: See the rejection of Claim 12 above, which claim 14 depends on. Claim 14 further recites:
• "determine a loss based on the reward and the divergence," (This limitation falls within the mathematical concepts grouping because it involves calculating a loss value based on the reward value and the divergence value using a mathematical loss function (see e.g., Equation 1 in paragraph [0077] of the specification).)
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
• "wherein the one or more processors are further configured to perform operations comprising..." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim recites generic, off-the-shelf processors as tools to perform the recited abstract ideas.)
• "wherein the re-training of the active large language model is based on the loss." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim denotes generic re-training and a generic large language model with no additional details or limitations beyond a generic, off-the-shelf large language model.)
Step 2B:
• "wherein the one or more processors are further configured to perform operations comprising..." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim recites generic, off-the-shelf processors as tools to perform the recited abstract ideas.)
• "wherein the re-training of the active large language model is based on the loss." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim denotes generic re-training and a generic large language model with no additional details or limitations beyond a generic, off-the-shelf large language model.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 15
Step 1: A machine, as above.
Step 2A Prong 1: See the rejection of Claim 14 above, which claim 15 depends on. Claim 15 further recites:
• "...perform a Proximal Policy Optimization (PPO) algorithm." (This limitation falls within the mathematical concepts grouping because the Proximal Policy Optimization (PPO) algorithm is a mathematical optimization algorithm that optimizes a loss function through mathematical calculations (see e.g., paragraph [0077] of the specification, describing optimizing the loss function of Equation 1 using the PPO algorithm).)
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
• "wherein the one or more processors are configured to perform the operation of re-train of the active large language model based on the loss by being configured to..." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The limitation amounts to using generic, off-the-shelf processors and the recited mathematical algorithm to re-train a generic, off-the-shelf large language model, which is merely an instruction to apply the abstract idea using generic computer components.)
Step 2B:
• "wherein the one or more processors are configured to perform the operation of re-train of the active large language model based on the loss by being configured to..." (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The limitation amounts to using generic, off-the-shelf processors and the recited mathematical algorithm to re-train a generic, off-the-shelf large language model, which is merely an instruction to apply the abstract idea using generic computer components.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim 16
Step 1: A machine, as above.
Step 2A Prong 1: See the rejection of Claim 9 above, which claim 16 depends on. Claim 16 further recites:
• "determine the validity of the generated completeness graph based on the validity of one or more of a file type, node types, and edges of the generated completeness graph;" (This limitation is a mental process. A person mentally or with a pen and paper can inspect a completeness graph and determine whether the graph is valid based on its file type, its node types, and its edges, e.g., by checking that each node represents a decision point and that each node is properly connected by edges (see e.g., paragraph [0062] of the specification).)
• "determine the semantic similarity of the generated completeness graph and the associated ground truth completeness graph based on a graph edit distance between the generated completeness graph and the associated ground truth completeness graph; and" (This limitation falls within the mathematical concepts grouping because it involves calculating a graph edit distance, which is a mathematical calculation of the minimum cost of an edit path (sequence of node and edge operations) to transform one graph into a graph isomorphic to the other graph (see e.g., paragraph [0061] of the specification).)
• "convert the validity and the semantic similarity to a scalar value as the reward." (This limitation falls within the mathematical concepts grouping because it involves mathematically combining values into a single scalar value, e.g., an average, a sum, or a weighted sum (see e.g., paragraph [0062] of the specification).)
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim further recites:
• "wherein the one or more processors are configured to perform the operation of evaluate the generated completeness graph with the reward model to produce the reward by being configured to perform:" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim recites generic, off-the-shelf processors and a generic reward model as tools to perform the recited abstract ideas.)
Step 2B:
• "wherein the one or more processors are configured to perform the operation of evaluate the generated completeness graph with the reward model to produce the reward by being configured to perform:" (Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)). -- Examiner's Note (EN): The claim recites generic, off-the-shelf processors and a generic reward model as tools to perform the recited abstract ideas.)
The additional elements considered individually or in combination do not amount to significantly more than the judicial exception. Therefore, the claim is not patent eligible.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Examiner’s Note: Some rejections will include an Examiner’s Note (labeled ‘EN’) to provide additional context or rationale explaining the basis for the rejection.
Claims 1-7 and 9-15 are rejected under 35 U.S.C. 103 as being unpatentable over Ouyang et al., "Training language models to follow instructions with human feedback," hereinafter "Ouyang" in view of You et al., "Graph Convolutional Policy Network for Goal-Directed Molecular Graph Generation," , hereinafter "You".
Claim 1
Ouyang teaches,
A method of training a large language model to generate (Abstract, "we show an avenue for aligning language models with user intent on a wide range of tasks by fine-tuning with human feedback... we collect a dataset of labeler demonstrations of the desired model behavior, which we use to fine-tune GPT-3 using supervised learning. We then collect a dataset of rankings of model outputs, which we use to further fine-tune this supervised model using reinforcement learning from human feedback." Sec. 3.5, "We start with the GPT-3 pretrained language models from Brown et al. (2020)." - EN: this denotes a method of training a large language model, where the model is trained to produce a desired output in response to an input.)
obtaining a dataset comprising instructions and ground truth (Sec. 3.2, "Our prompt dataset consists primarily of text prompts submitted to the OpenAI API... From these prompts, we produce three different datasets used in our fine-tuning procedure: (1) our SFT dataset, with labeler demonstrations used to train our SFT models..." Sec. 3.1, "Our labelers provide demonstrations of the desired behavior on the input prompt distribution..." Sec. 3.3, "For each natural language prompt, the task is most often specified directly through a natural language instruction (e.g. 'Write a story about a wise frog')..." Table 6, "Dataset sizes, in terms of number of prompts." - EN: this denotes obtaining a dataset in which each prompt is an instruction and each associated labeler demonstration is the known-correct desired output for that instruction, i.e., the ground truth associated with the instruction.)
producing a generated (Sec. 3.5, "The environment is a bandit environment which presents a random customer prompt and expects a response to the prompt." Sec. 3.1, Step 3, "Optimize a policy against the reward model using PPO." Sec. 3.5, "we fine-tuned the SFT model on our environment using PPO... We call these models 'PPO.'" - EN: this denotes that the PPO policy, which is the language model actively undergoing reinforcement learning, produces an output in response to a prompt drawn from the dataset, i.e., the active large language model producing an output in response to a query based on instructions from the dataset.)
evaluating the generated (Sec. 3.5, "Starting from the SFT model with the final unembedding layer removed, we trained a model to take in a prompt and response, and output a scalar reward." Sec. 3.5, "Given the prompt and response, it produces a reward determined by the reward model and ends the episode." Sec. 3.1, Step 2, "We then train a reward model to predict the human-preferred output." - EN: this denotes evaluating the generated output with a separately trained reward model to produce a scalar reward.)
re-training the active large language model based at least partially on the reward. (Sec. 3.1, Step 3, "We use the output of the RM as a scalar reward. We fine-tune the supervised policy to optimize this reward using the PPO algorithm (Schulman et al., 2017)." Eq. 2, objective(φ) = E(x,y)∼DπRL[rθ(x,y) − β log(πRLφ(y|x) / πSFT(y|x))] + γEx∼Dpretrain[log(πRLφ(x))]. - EN: this denotes re-training the active language model on the reward produced by the reward model, the reward being one component of the combined optimization objective, i.e., re-training based at least partially on the reward.)
Ouyang does not explicitly teach:
that the generated output and the associated ground truth output are completeness graphs (as struck through in the preamble and in the obtaining, producing, and evaluating steps above); and
the reward being based on validity of the generated completeness graph and semantic similarity of the generated completeness graph and the associated ground truth completeness graph (as struck through in the evaluating step above).
However, You teaches:
that the generated output and the associated ground truth output are completeness graphs (as struck through in the preamble and in the obtaining, producing, and evaluating steps above); (Abstract, "we propose Graph Convolutional Policy Network (GCPN), a general graph convolutional network based model for goal-directed graph generation through reinforcement learning." Sec. 3.1, "We represent a graph G as (A, E, F), where A ∈ {0, 1}n×n is the adjacency matrix, and F ∈ Rn×d is the node feature matrix assuming each node has d features." Sec. 3.5, "any ground truth molecule could be viewed as an expert trajectory for pretraining GCPN... given a molecule dataset, we randomly sample a molecular graph G..." Sec. 4.1, "we utilize the ZINC250k molecule dataset [14] that contains 250,000 drug like commercially available molecules whose maximum atom number is 38." - EN: this denotes an RL policy that produces a graph as its generated output, trained against a dataset of known-correct example graphs that serve as the ground truth for the generation process. Under the broadest reasonable interpretation, a "completeness graph" is a graph of nodes and edges representing decision points and their dependencies The claim recites no structural limitation confining the graph to question-flow content; "completeness" characterizes what the depicted graph represents in a particular field of use rather than imposing any further structure on the claimed training method, which operates identically regardless of what the nodes denote. The molecular graph G = (A, E, F) of You, which is a set of nodes and edges built up through a sequence of dependent decision steps subject to structural constraints, therefore falls within the scope of the claimed completeness graph. You itself teaches that its approach is not limited to molecular graphs: "Furthermore, the application of GCPN can extend well beyond molecule generation. The algorithm can be applied to generate graphs in many contexts, such as electric circuits, social networks, and explore graphs that can optimize certain domain specific properties." (You, Sec. 5).)
the reward being based on validity of the generated completeness graph (Sec. 3.3, "The intermediate rewards include step-wise validity rewards and adversarial rewards. A small positive reward is assigned if the action does not violate valency rules, otherwise a small negative reward is assigned." Sec. 3.3, "Domain-specific rewards also include penalization of unrealistic molecules according to various criteria, such as excessive steric strain and the presence of functional groups that violate ZINC functional group filters [14]." Sec. 7 (Appendix), "We define a molecule as valid if it is able to pass the sanitization checks in RDKit." Sec. 3.3, "Infeasible actions proposed by the policy network are rejected and the state remains unchanged." - EN: this denotes that a component of the reward is determined by whether the generated graph satisfies structural rules governing its node and edge types, i.e., a reward based on validity of the generated graph, consistent with paragraph 62 of the instant application, which determines validity from "the resulting format of the generated completeness graph, the node types, and edge restrictions of the completeness graph.")
and semantic similarity of the generated completeness graph and the associated ground truth completeness graph; (Sec. 3.3, "To ensure that the generated molecules resemble a given set of molecules, we employ the Generative Adversarial Network (GAN) framework [10] to define the adversarial rewards V(πθ, Dφ)... we use −V(πθ, Dφ) as an additional reward together with other rewards." Sec. 3.1, "We provide the model with a set of example graphs G ∼ pdata(G), and would like to incorporate such prior knowledge by regularizing the property optimization objective with EG,G′[J(G, G′)] under distance metric J(·,·)." Sec. 4.2, "the molecule similarity sim(G, G′) between the original and modified molecules is above a threshold δ." Sec. 4.2, "The diversity of a set of molecules is defined as the average pairwise Tanimoto distance between the Morgan fingerprints [33] of the molecules." - EN: this denotes a reward component computed as a distance metric J(G, G′) between the generated graph and the ground truth example graphs, which under the broadest reasonable interpretation is a semantic similarity because it measures the degree to which the generated graph corresponds in content to the ground truth graph rather than matching it literally; further, Sec. 3.3 teaches combining this similarity component with the validity component into the total reward optimized by policy gradient.)
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the reward-model-based reinforcement learning fine-tuning of a language model of Ouyang with the graph generation rewarded for validity and ground truth similarity of You. The motivation for doing so would be to put the format rules of the generated graph, and its closeness to a known-correct graph, directly into the reward signal, so that a malformed or off-target graph is penalized during training instead of being caught later by a human reviewer. As You elaborates regarding the benefit of this reward-based methodology in Section 1, "desired molecular properties such as drug-likeness [1, 29] and molecule constraints such as valency are complex and non-differentiable, thus they cannot be directly incorporated into the objective function of graph generative models. In contrast, reinforcement learning is capable of directly representing hard constraints and desired properties through the design of environment dynamics and reward function."
Examiner's Note: The term correspondences established in the rejection of claim 1 are carried forward and used consistently below: the "active large language model" is Ouyang's PPO policy (πRL); the "reward model" is Ouyang's reward model (RM); the "reference large language model" is Ouyang's supervised fine-tuned model (SFT model, πSFT); and, per You, the generated and ground truth outputs are "completeness graphs" (graphs of nodes and edges). Consistent with the rejection of claim 1, "completeness graph(s)" is struck through where the recitation is taught by Ouyang except for the graph nature of the output, which is supplied by You, and the rationale for combining Ouyang and You set forth for claim 1 applies to each such recitation.
Claim 2
Ouyang in view of You teaches the method of claim 1. Ouyang further teaches:
The method of claim 1, further comprising training a reference large language model to receive instructions from the dataset and in response to produce (Sec. 3.1, Step 1, "Collect demonstration data, and train a supervised policy. Our labelers provide demonstrations of the desired behavior on the input prompt distribution... We then fine-tune a pretrained GPT-3 model on this data using supervised learning." Sec. 3.5, "Supervised fine-tuning (SFT). We fine-tune GPT-3 on our labeler demonstrations using supervised learning." Sec. 3.5, Reinforcement learning (RL), "we fine-tuned the SFT model on our environment using PPO... We call these models 'PPO.'" - EN: this denotes training a second, reference language model (the SFT model) that receives prompts (instructions) from the dataset and in response produces outputs, and that the active large language model (the PPO policy mapped in claim 1) is initialized as a fine-tuned copy of that SFT (reference) model. Ouyang expressly initializes the PPO policy from the SFT model by fine-tuning the SFT model with PPO, i.e., the active model is initialized based on the reference model. As set forth in the rejection of claim 1, You teaches that the produced outputs are completeness graphs (as struck through above), and the same rationale for combining Ouyang and You stated for claim 1 applies.)
Claim 3
Ouyang in view of You teaches the method of claim 2. Ouyang further teaches:
The method of claim 2, wherein training the reference large language model comprises supervised training using instructions from the dataset as prompts and (Sec. 3.1, Step 1, "Our labelers provide demonstrations of the desired behavior on the input prompt distribution... We then fine-tune a pretrained GPT-3 model on this data using supervised learning." Sec. 3.2, "From these prompts, we produce three different datasets used in our fine-tuning procedure: (1) our SFT dataset, with labeler demonstrations used to train our SFT models..." Sec. 3.3, "For each natural language prompt, the task is most often specified directly through a natural language instruction..." Sec. 3.5, "Supervised fine-tuning (SFT). We fine-tune GPT-3 on our labeler demonstrations using supervised learning." - EN: this denotes supervised training of the reference SFT large language model using natural-language instructions from the dataset as prompts and the associated labeler demonstrations as the desired target outputs, i.e., labels, for the supervised training.)
Ouyang does not explicitly teach:
the ground truth completeness graphs associated with instructions as labels (as struck through above).
However, You teaches:
the ground truth completeness graphs associated with instructions as labels (as struck through above); (Sec. 3.5, "any ground truth molecule could be viewed as an expert trajectory for pretraining GCPN... where (st, at) pairs are obtained from ground truth molecules. Specifically, given a molecule dataset, we randomly sample a molecular graph G... and use the pair (st, at) to supervise the expert imitation objective." - EN: this denotes using known-correct, i.e., ground-truth, graph examples from a dataset as supervisory target information for training the graph-generating policy. As discussed with respect to claim 1, under the applied broadest reasonable interpretation, You's ground-truth molecular graphs correspond to the claimed ground truth completeness graphs. Thus, when You's graph-generation teachings are applied to Ouyang's supervised instruction-following language-model training, the instructions serve as prompts and the associated ground-truth graphs serve as the target labels for supervised training.)
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the supervised fine-tuning of an instruction-following large language model on labeler demonstrations of Ouyang with the use of ground truth graphs as supervised training targets of You during the supervised training step. The motivation for doing so would be to improve training stability and performance when training the model to generate graph-structured outputs, by starting the model from known-correct examples of the graphs it must produce. As You elaborates regarding the benefit of this expert pretraining methodology in Section 3.5, "It is known that pretraining a policy network with expert policies if they are available leads to better training stability and performance [24]."
Claim 4
Ouyang in view of You teaches the method of claim 2. Ouyang further teaches:
The method of claim 2, further comprising:
producing a second generated (Sec. 3.5, RL, "In addition, we add a per-token KL penalty from the SFT model at each token to mitigate overoptimization of the reward model." Eq. 2, in relevant part, objective(φ) = E(x,y)∼DπRL[rθ(x,y) − β log(πRLφ(y|x) / πSFT(y|x))] + γEx∼Dpretrain[log(πRLφ(x))]. - EN: this denotes that the reference (SFT) model is applied to the same query and produces its own response to that query, the reference model's response being represented by its per-token probabilities πSFT(y|x). You supplies the graph nature of the output (as struck through above), per claim 1. The same rationale for the combination applies.)
and comparing the generated (Sec. 3.5, RL, "we add a per-token KL penalty from the SFT model at each token to mitigate overoptimization of the reward model." Eq. 2, the term β log(πRLφ(y|x) / πSFT(y|x)). - EN: this denotes computing a divergence between the active (PPO) model's response and the reference (SFT) model's response to the same query, namely the KL penalty term of Eq. 2 that compares the active-model probabilities πRL(y|x) against the reference-model probabilities πSFT(y|x). You supplies the graph nature of the output (as struck through above), per claim 1. The same rationale for the combination applies.)
wherein the re-training of the active large language model is further based at least partially on the divergence. (Eq. 2, in relevant part, objective(φ) = E(x,y)∼DπRL[rθ(x,y) − β log(πRLφ(y|x) / πSFT(y|x))] + γEx∼Dpretrain[log(πRLφ(x))]. Sec. 3.5, "The KL reward coefficient, β, and the pretraining loss coefficient, γ, control the strength of the KL penalty and pretraining gradients respectively." - EN: this denotes that the divergence (the KL penalty) is a component of the objective that the PPO update maximizes when re-training the active model, i.e., the re-training of the active model is further based at least partially on the divergence.)
Claim 5
Ouyang in view of You teaches the method of claim 4. Ouyang further teaches:
The method of claim 4, wherein comparing the generated (Sec. 3.5, RL, "In addition, we add a per-token KL penalty from the SFT model at each token to mitigate overoptimization of the reward model." Eq. 2, the term β log(πRLφ(y|x) / πSFT(y|x)). - EN: this denotes that the divergence between the active model's response and the reference model's response is a Kullback-Leibler (KL) divergence, expressly the per-token KL penalty of Ouyang between the RL policy πRL and the reference SFT policy πSFT. You supplies the graph nature of the output (as struck through above), per claim 1. The same rationale for the combination applies.)
Claim 6
Ouyang in view of You teaches the method of claim 4. Ouyang further teaches:
The method of claim 4, further comprising determining a loss based on the reward and the divergence, wherein the re-training of the active large language model is based on the loss. (Eq. 2, in relevant part, objective(φ) = E(x,y)∼DπRL[rθ(x,y) − β log(πRLφ(y|x) / πSFT(y|x))] + γEx∼Dpretrain[log(πRLφ(x))]. Sec. 3.5, "we fine-tuned the SFT model on our environment using PPO... In addition, we add a per-token KL penalty from the SFT model at each token to mitigate overoptimization of the reward model." - EN: this denotes forming a single combined optimization objective that combines the reward rθ(x,y) produced by the reward model with the KL divergence penalty, and re-training (optimizing) the active model on that combined objective. Ouyang's combined objective corresponds to the recited "loss": Under the broadest reasonable interpretation, forming and optimizing Ouyang's reward-plus-KL objective is determining a loss based on the reward and the divergence and re-training the active model based on the loss.)
Claim 7
Ouyang in view of You teaches the method of claim 6. Ouyang further teaches:
The method of claim 6, wherein the re-training of the active large language model based on the loss comprises performing a Proximal Policy Optimization (PPO) algorithm. (Sec. 3.1, Step 3, "Optimize a policy against the reward model using PPO... We fine-tune the supervised policy to optimize this reward using the PPO algorithm (Schulman et al., 2017)." Sec. 3.5, "we fine-tuned the SFT model on our environment using PPO... We call these models 'PPO.'" - EN: this denotes that re-training the active model on the combined objective (loss) is performed using the Proximal Policy Optimization (PPO) algorithm, as expressly disclosed by Ouyang.)
Claim 9
Claim 9 recites a system that performs operations corresponding to the method of claim 1. Ouyang in view of You teaches each recited operation as set forth in the rejection of claim 1, and Ouyang in view of You further teaches the recited hardware.
A system of training a large language model to generate processors to perform operations comprising: (You, Sec. 4.1, "we set up the same objective functions for all methods, and run all the experiments on the same computing facilities using 32 CPU cores." Ouyang, Sec. C (Additional model details), "All models use fp16 weights and activations, with fp32 master copies of weights... All models are trained with the Adam optimizer, with β1 = 0.9 and β2 = 0.95." - EN: this denotes that the training method is carried out by a computer system having one or more processors (You's 32 CPU cores) and a memory coupled to the processors that stores both the executable program instructions and the model parameters operated upon during training (e.g., Ouyang's fp16/fp32 weight copies). A processor executing stored instructions from a coupled memory is the inherent and well-understood implementation of each reference's computer-implemented training method; You supplies the graph nature of the output, per claim 1.)
Examiner's Note: Each of Ouyang and You executes its training method on computer hardware (You, Sec. 4.1; Ouyang, Sec. C). It would have been obvious to a person having ordinary skill in the art before the effective filing date to implement the combined training method of Ouyang and You as instructions stored in a memory coupled to one or more processors, because executing machine learning training as stored instructions on one or more processors is the standard, well-known implementation in the art of machine learning for performing such training.
The remaining limitations of claim 9 are substantially the same as method claim 1, therefore claim 9 is rejected under the same rationale as claim 1.
Claims 10-15 recite substantially the same limitations as method claims 2-7 respectively. Therefore, claims 10-15 are rejected under the same rationale as claims 2-7.
Claims 8 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Ouyang in view of You, and further in view of Bai et al., "SimGNN: A Neural Network Approach to Fast Graph Similarity Computation," hereinafter "Bai".
Claim 8
Ouyang in view of You teaches the method of claim 1, as set forth in the rejection of claim 1 above, which is incorporated herein together with the rationale for combining Ouyang and You. Ouyang further teaches:
The method of claim 1, wherein the evaluating the generated (see the evaluating step in the rejection of claim 1; the struck-through "completeness graph" is the graph nature of the output supplied by You, as set forth for claim 1.)
converting (Sec. 3.5, "trained a model to take in a prompt and response, and output a scalar reward." - EN: this denotes outputting the evaluation as a single scalar reward value. The validity and the semantic similarity (struck through above) are the evaluation components taught by You as set forth below, and You likewise combines these components into a single reward: "The intermediate rewards include step-wise validity rewards and adversarial rewards." (You, Sec. 3.3); "We define the final rewards as a sum over domain-specific rewards and adversarial rewards." (You, Sec. 3.3). In the combination, the validity and the semantic similarity taught by You are converted to the scalar reward value output as taught by Ouyang.)
Ouyang does not explicitly teach:
determining the validity of the generated completeness graph based on the validity of one or more of a file type, node types, and edges of the generated completeness graph;
determining the semantic similarity of the generated completeness graph and the associated ground truth completeness graph based on a graph edit distance between the generated completeness graph and the associated ground truth completeness graph; and
However, You teaches:
determining the validity of the generated completeness graph based on the validity of one or more of a file type, node types, and edges of the generated completeness graph; (Sec. 3.3, "A small positive reward is assigned if the action does not violate valency rules, otherwise a small negative reward is assigned." Sec. 3.3, "Infeasible actions proposed by the policy network are rejected and the state remains unchanged." - EN: valency rules govern how nodes connect via edges. This denotes determining validity based on node types and edge restrictions of the generated completeness graph. Thereby satisfying the "node types, and edges" alternatives of the recited "one or more of" list.)
determining the semantic similarity of the generated completeness graph and the associated ground truth completeness graph based on (Sec. 3.1, "...regularizing the property optimization objective with EG,G′[J(G, G′)] under distance metric J(·,·)." Sec. 4.2, "the molecule similarity sim(G, G′) between the original and modified molecules is above a threshold δ." - EN: this denotes calculating the semantic similarity based on a distance metric between the generated completeness graph and the associated ground truth completeness graph, for the reasons set forth in the rejection of claim 1 above.)
Ouyang in view of You does not explicitly teach:
a graph edit distance (as struck through in the determining the semantic similarity limitation above).
However, Bai teaches:
a graph edit distance (as struck through in the determining the semantic similarity limitation above); (Sec. 2.1, "we choose one of the most popular graph similarity/distance metric, Graph Edit Distance (GED)..." Sec. 2.1, "...the edit distance between G1 and G2, denoted by GED(G1, G2)..." Sec. 2.1, "...an insertion or deletion of a vertex/edge or relabelling of a vertex..." Sec. 2.1, "...we transform it to a similarity score ranging between 0 and 1." Sec. 4.2, "...there is a one-to-one mapping between the GED and the similarity score." - EN: this denotes the graph edit distance, a distance metric defined between two graphs as the number of node and edge insertions, deletions, and relabelings in the optimal alignment that transforms one graph into the other, which Bai converts into a similarity score for the pair of graphs. In the combination, the graph edit distance of Bai is the distance metric J(·,·) of You, computed between the generated completeness graph and the associated ground truth completeness graph, so the semantic similarity is determined based on a graph edit distance between the two graphs as claimed.)
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the graph generation training of Ouyang and You, in which the reward is based on the similarity of the generated completeness graph and the associated ground truth completeness graph, with the graph edit distance of Bai as the distance metric between the two graphs. The motivation for doing so would be to score the similarity of the generated completeness graph and the associated ground truth completeness graph in the same node and edge terms that the validity component of the reward already inspects, so that the reward counts the node and edge operations that separate the two graphs. Bai further teaches that the graph edit distance maps one-to-one to a similarity score between 0 and 1 (Bai, Sec. 2.1 and Sec. 4.2), which gives the reward model a bounded scalar to combine with the validity component. As Bai elaborates regarding the benefit of this graph edit distance methodology in Section 2.1, "GED has been widely used in many applications, such as graph similarity search...".
Claim 16 recites substantially the same limitations as method claim 8; therefore, claim 16 is rejected under the same rationale as claim 8.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NAYMUR RAHMAN ALI whose telephone number is (571)272-0007. The examiner can normally be reached Mon-Fri. 9:30-6:30 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571)270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/NAYMUR RAHMAN ALI/Examiner, Art Unit 2123
/ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123