DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Interpretation
For clarity of record, the limitation in independent claim 1 of optimizing a performance of an Al model by selecting one or more internal settings for the Al model using reinforcement learning based on a task-specific reward function that measures the performance of the Al model on a specified task under broadest reasonable interpretation can be interpreted as either 1) selecting settings for the model and the model is using reinforcement learning, or 2) selecting settings, using reinforcement learning, for the model. Claims 9 and 25 contain similar language and allow similar interpretations.
Claim Rejections - 35 USC § 112(b)
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 31, 33, and 35 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Exemplary claim 31 recites wherein the off-policy algorithms for reinforcement learning comprise Soft Action-Critic (SAC) algorithms. However, Examiner believes the term of art is “Soft Actor-Critic (SAC).” The specification in paragraph 65 states, “The platform can employ reinforcement learning algorithms, such as Proximal Policy Optimization (PPO), Asynchronous Advantage Actor-Critic (A3C), or Soft Action-Critic (SAC) to optimize model settings like temperature and token length based on task-specific reward functions” (emphasis added), which supports the claim. However, the specification also states in paragraph 66, “Soft Actor-Critic is an off-policy reinforcement learning algorithm that combines the benefits of both policy-based and value-based methods. It is designed to maximize both the expected reward and the entropy of the policy, which encourages exploration and helps prevent the agent from getting stuck in suboptimal solutions. SAC aims to maximize the entropy of the policy alongside the expected return. By adding an entropy term to the objective function, the agent is encouraged to explore more diverse actions, leading to better exploration and improved sample efficiency. SAC is an off-policy algorithm, meaning it can learn from data collected by any policy, not just the current policy being optimized” (emphasis added), which contradicts the claim and previous specification paragraph. Due to this contradiction, the scope of the term Soft Action-Critic (SAC) algorithms is unclear. For examination purposes, the term is interpreted as Soft Actor-Critic. For this reason, the above listed claims are rejected for containing this language or being dependent on a claim that contains this language.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-2, 4, 6, 8-10, 12, 14, 16, 25-26, 28, 30-36 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: Claim 1 is a system claim. Claim 9 is a method claim. Claim 25 is a CRM claim. Therefore, claims 1, 9, and 25 are directed to either a process, machine, manufacture or composition of matter.
With respect to Claim 1:
Step 2A Prong 1:
optimizing performance of an Al model by selecting one or more hyperparameters for the Al model using reinforcement learning based on a task-specific reward function comprising a metric for comparing generated outputs from the Al model to reference outputs to measure the performance of the Al model on a specified task, wherein the reinforcement learning comprises on-policy optimization algorithms or off-policy optimization algorithms to sample a group of the generated outputs from the Al model based on an input query, compute one or more rewards based on the task-specific reward function, and update the one or more hyperparameters based on the computed rewards, wherein the one or more hyperparameters comprise at least one of temperature, token length, learning rate, batch size, attention head count, embedding size, dropout rate, or activation function (mental process – user can manually optimize a performance of an Al model by selecting one or more hyperparameters for the Al model that uses reinforcement learning based on a task-specific reward function comprising a metric for comparing generated outputs from the Al model to reference outputs to measure the performance of the Al model on a specified task, wherein the reinforcement learning comprises on-policy optimization algorithms or off-policy optimization algorithms to sample a group of the generated outputs from the Al model based on an input query, compute one or more rewards based on the task-specific reward function, and update the one or more hyperparameters based on the computed rewards, wherein the one or more hyperparameters comprise at least one of temperature, token length, learning rate, batch size, attention head count, embedding size, dropout rate, or activation function)
validating the performance of the Al model against one or more authoritative datastores, wherein the one or more authoritative datastores comprises a relational datastore, a NoSQL datastore, a graph datastore, a knowledge graph, a vector datastore, a document datastore, or a hybrid vectorized knowledge graph, or against one or more rule sets (mental process – user can manually validate the performance of the Al model against one or more authoritative datastores, wherein the one or more authoritative datastores comprises a relational datastore, a NoSQL datastore, a graph datastore, a knowledge graph, a vector datastore, a document datastore, or a hybrid vectorized knowledge graph, or against one or more rule sets)
optimizing a stability of the Al model using input perturbation (mental process – user can manually optimize a stability of the Al model using input perturbation)
optimizing a reliability of the Al model for the specified task using one or more techniques measured against a fitness function, wherein the techniques include model type search, attention mechanism search, model blending with weighted consensus, expert synthesis, retrieval augmented generation (RAG), knowledge graph verification, or composite vectorized knowledge graphs (mental process – user can manually optimize a reliability of the AI model for the specified task using one or more techniques measured against a fitness function, wherein the techniques include model type search, attention mechanism search, model blending with weighted consensus, expert synthesis, retrieval augmented generation (RAG}, knowledge graph verification, or composite vectorized knowledge graphs)
wherein a distributed computational graph automatically parallelizes optimizing the performance of the AI model, validating the performance of the AI model, optimizing the robustness of the AI model, optimizing the stability of the AI model, and optimizing the reliability of the AI model across heterogeneous computing resources (mental process – using a distributed computational graph that automatically parallelizes optimizing the performance of the AI model, validating the performance of the AI model, optimizing the robustness of the AI model, optimizing the stability of the AI model, and optimizing the reliability of the AI model across heterogeneous computing resources)
Step 2A Prong 2: This judicial exception is not integrated into a practical application. Additional elements:
one or more hardware processors (mere instructions to apply the exception using a generic computer component)
optimizing a robustness of the Al model using adversarial training (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. Additional elements:
one or more hardware processors (mere instructions to apply the exception using a generic computer component)
optimizing a robustness of the Al model using adversarial training (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
Conclusion: The claim is not patent eligible.
Claims 9 and 25 are rejected on the same grounds as claims 1. Claims 9 and 25 each include additional generic computer components which do not integrate the abstract idea into practical application or provide significantly more than the abstract idea.
Regarding Claims 2, 4, 8, 10, 12, 16, 26, 28, 31-36: These limitations, as drafted, are a process that, under its broadest reasonable interpretation, covers performance of the limitations in the mind. That is, nothing in the claim limitation precludes the step from practically being performed in the mind.
For claims 2, 10, 26: the limitation encompasses the user manually use wherein on-policy algorithms for reinforcement learning comprise Proximal Policy Optimization (PPO) or Asynchronous Advantage Actor-Critic (A3C) algorithms.
For claims 4, 12, 28: the limitation encompasses the user manually comparing responses of the AI model to responses provided by the one or more authorities.
For claims 8, 16: the limitation encompasses the user manually selecting from or blending outputs from multiple models or authoritative knowledge bases, each trained on, or obtained from, defined resource collections or with different retrieval strategies.
For claims 31, 33, 35: the limitation encompasses the user manually use wherein the off-policy algorithms for reinforcement learning comprise Soft Action-Critic (SAC) algorithms
For claims 32, 34, 36: the limitation encompasses the user manually use wherein the task-specific reward function comprises recall-oriented understudy for gisting evaluation (ROUGE) or bilingual evaluation understudy (BLEU) scores algorithms
These judicial exceptions are not integrated into a practical application. In particular, the claims do not recite any additional elements. Accordingly, this does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea.
The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, no additional elements are cited. Accordingly, the claim is not patent eligible.
Regarding Claims 6, 14, 30: The limitation, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. That is, other than the additional elements, nothing in the claim limitation precludes the step from practically being performed in the mind.
For claims 6, 14, 30: the limitation includes the additional element of incorporates malicious examples into training data to make the AI model more resilient against manipulated predictions.
These judicial exceptions are not integrated into a practical application. In particular, the additional elements are merely adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f)). Accordingly, this does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element are merely adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f). Accordingly, the claim is not patent eligible.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 4, 6, 8-9, 12, 14, 16, 25, 28, 30-31, 33, 35 is/are rejected under 35 U.S.C. 103 as being unpatentable over Xiao et al. (hereinafter Xiao), U.S. Patent Application Publication 2022/0385545, in view of Lucas et al. (hereinafter Lucas) Adversarial Training for Raw-Binary Malware Classifiers, further in view of Chunduru et al. (hereinafter Chunduru), U.S. Patent Application Publication 2025/0209737, further in view of Wang et al. (hereinafter Wang), A match made in consistency heaven: when large language models meet evolutionary algorithms, further in view of Vasiljevic et al. (hereinafter Vasiljevic), U.S. Patent 11,960,885.
Regarding Claim 1, Xiao discloses a computing system for optimizing generative Al models, the computing system comprising: one or more hardware processors configured for:
optimizing performance of an Al model by selecting one or more hyperparameters for the Al model using reinforcement learning based on a task-specific reward function [“using a Reinforcement Learning (RL) algorithm to refine the at least one hyperparameter of the autoencoder, wherein a reward function of the RL algorithm is calculated on the basis of the generated evaluation” ¶55] comprising a metric for comparing generated outputs from the Al model to reference outputs to measure the performance of the Al model on a specified task [“using an autoencoder to concentrate information in the data stream” ¶55; “detecting an event from the concentrated information, and, in step 130, generating an evaluation of the detected event on the basis of logical compatibility between the detected event and a knowledge base” ¶55; Fig. 1], wherein the reinforcement learning comprises on-policy optimization algorithms or off-policy optimization algorithms function [“agents can learn and optimize a policy for controlling a system or environment, such as the autoencoder of the method 100, based on observed states of the system and a reward system that is tailored towards achieving a particular goal.” ¶60] to sample a group of the generated outputs from the Al model based on an input query, compute one or more rewards based on the task-specific reward, and update the one or more hyperparameters based on the computed rewards [“agent selects Actions on the basis of system States with the aim of maximizing the expected future Reward. A Reward function may be defined such that a greater Reward is received for Actions that result in the system entering a state that approaches a target end state for the system, consistent with an overall goal of an entity managing the system. In the case of the method 100, the target end stat of the autoencoder may be a state in which the hyperparameters are such that event detection in the concentrated data stream has reached a desired accuracy threshold,” ¶60], wherein the one or more hyperparameters comprise at least one of temperature, token length, learning rate, batch size, attention head count, embedding size, dropout rate, or activation function [“dividing the accumulated data stream into a plurality of consecutive windows, each window corresponding to a different time interval,” ¶62; “The at least one hyper parameter may comprise a time interval associated with the time window, a scaling factor, and/or a layer number decreasing rate.” ¶63; Examiner Note: dividing the accumulated data stream into a plurality of consecutive windows is interpreted as a batch size, and a time interval associated with the time window is interpreted as an embedding size];
validating the performance of the Al model against one or more authoritative datastores, wherein the one or more authoritative datastores comprises a relational datastore, a NoSQL datastore, a graph datastore, a knowledge graph, a vector datastore, a document datastore, or a hybrid vectorized knowledge graph, or against one or more rule sets [“Examples of the present disclosure apply logical verification of results based on a knowledge base that may be populated without the need for specific domain knowledge. Such a knowledge base may be built from data including environmental, physical and business data, and may thus be considered as a "common sense" check that results are consistent with what is known about a monitored system and/or environment and about business requirements for a particular deployment.” ¶53; “A repository for the knowledge base.” ¶81].
However, Xiao fails to explicitly disclose optimizing a robustness of the AI model using adversarial training.
Lucas discloses optimizing a robustness of the AI model using adversarial training [“the effectiveness of using adversarial training methods to create malware-classification models that are more robust to some state-of-the-art attacks” Abstract];
It would have been obvious to one having ordinary skill in the art, having the teachings of Xiao and Lucas before him before the effective filing date of the claimed invention, to modify the system of Xiao to incorporate the adversarial training of Lucas.
Given the advantage of increased robustness of the model, one having ordinary skill in the art would have been motivated to make this obvious modification.
However, Xiao fails to explicitly disclose optimizing a stability of the AI model using input perturbation.
Chunduru discloses optimizing a stability of the AI model using input perturbation [“The key idea behind SmoothGrad is to introduce perturbations to the input data. Instead of attributing the prediction solely to the gradients calculated with respect to the original input, the gradients are averaged over multiple perturbed versions of the input. By averaging the gradients over multiple perturbed samples, SmoothGrad helps reduce the impact of noise or irrelevant features in the attribution maps. This is particularly beneficial when dealing with complex or noisy datasets. Perturbation techniques include adding Gaussian noise, random rotations, or random translations to the input data. These perturbations create variations in the input while preserving the essential features, leading to more stable and reliable attribution maps.” ¶68].
It would have been obvious to one having ordinary skill in the art, having the teachings of Xiao, Lucas, and Chunduru before him before the effective filing date of the claimed invention, to modify the combination to incorporate input perturbations of Chunduru.
Given the advantage of reducing the impact of noise or irrelevant features, one having ordinary skill in the art would have been motivated to make this obvious modification.
However, Xiao fails to explicitly disclose optimizing a reliability of the AI model for the specified task using one or more techniques measured against a fitness function, wherein the techniques include model type search, attention mechanism search, model blending with weighted consensus, expert synthesis, retrieval augmented generation (RAG), knowledge graph verification, or composite vectorized knowledge graphs.
Wang discloses optimizing a reliability of the AI model for the specified task using one or more techniques measured against a fitness function [“Individuals are sorted in descending order of fitness” §1.2 ¶2; Table I], wherein the techniques include model type search, attention mechanism search, model blending with weighted consensus, expert synthesis, retrieval augmented generation (RAG), knowledge graph verification, or composite vectorized knowledge graphs [“attention mechanism directly performs feature transformation” §1.3 ¶1; Table I];
It would have been obvious to one having ordinary skill in the art, having the teachings of Xiao, Lucas, Chunduru, and Wang before him before the effective filing date of the claimed invention, to modify the combination to incorporate both a fitness function and an attention mechanism of Wang.
Given the advantage of improving accuracy, one having ordinary skill in the art would have been motivated to make this obvious modification.
However, Xiao fails to explicitly disclose wherein a distributed computational graph automatically parallelizes optimizing the performance of the AI model, validating the performance of the AI model, optimizing the robustness of the AI model, optimizing the stability of the AI model, and optimizing the reliability of the AI model across heterogeneous computing resources ["execute an application data flow graph on a set of computational nodes" col. 5, lines 7-8; "parallel computing using heterogeneous networks of computational nodes" col. 4, lines 16-17].
Vasiljevic discloses wherein a distributed computational graph automatically parallelizes optimizing the performance of the AI model, validating the performance of the AI model, optimizing the robustness of the AI model, optimizing the stability of the AI model, and optimizing the reliability of the AI model across heterogeneous computing resources ["execute an application data flow graph on a set of computational nodes" col. 5, lines 7-8; "parallel computing using heterogeneous networks of computational nodes" col. 4, lines 16-17].
It would have been obvious to one having ordinary skill in the art, having the teachings of Xiao, Lucas, Chunduru, Wang, and Vasiljevic before him before the effective filing date of the claimed invention, to modify the combination to incorporate distributed processing on heterogeneous resources of Vasiljevic.
Given the advantage of faster processing and more efficient use of resources, one having ordinary skill in the art would have been motivated to make this obvious modification.
Regarding Claim 4, Xiao, Lucas, Chunduru, Wang, and Vasiljevic disclose the computing system of claim 1. Xiao further discloses wherein validating the performance of the AI model comprises comparing responses of the AI model to responses provided by the one or more authorities [“Examples of the present disclosure apply logical verification of results based on a knowledge base that may be populated without the need for specific domain knowledge. Such a knowledge base may be built from data including environmental, physical and business data, and may thus be considered as a "common sense" check that results are consistent with what is known about a monitored system and/or environment and about business requirements for a particular deployment.” ¶53; “A repository for the knowledge base.” ¶81].
Regarding Claim 6, Xiao, Lucas, Chunduru, Wang, and Vasiljevic disclose the computing system of claim 1.
However, Xiao fails to explicitly disclose wherein the adversarial training incorporates malicious examples into training data to make the AI model more resilient against manipulated predictions.
Lucas discloses wherein the adversarial training incorporates malicious examples into training data to make the AI model more resilient against manipulated predictions [“creating variants of malicious binaries, referred to as adversarial examples, that are transformed in a functionality-preserving way to evade detection” Abstract].
It would have been obvious to one having ordinary skill in the art, having the teachings of Xiao, Lucas, Chunduru, Wang, and Vasiljevic before him before the effective filing date of the claimed invention, to modify the combination to incorporate the adversarial training that incorporates malicious examples of Lucas.
Given the advantage of making the model more resilient, one having ordinary skill in the art would have been motivated to make this obvious modification.
Claims 9, 12, 14 are rejected on the same grounds as claims 1, 4, 6, respectively.
Claims 25, 28, 30 are rejected on the same grounds as claims 1, 4, 6, respectively.
Claim(s) 2, 10 and 26 is/are rejected under 35 U.S.C. 103 as being unpatentable over Xiao, Lucas, Chunduru, Wang, and Vasiljevic, further in view of Mermoud et al. (hereinafter Mermoud) U.S. Patent Application Publication 2025/0158895.
Regarding Claim 2, Xiao, Lucas, Chunduru, Wang, and Vasiljevic discloses the computing system of claim 1.
However, Xiao fails to explicitly disclose wherein on-policy algorithms for reinforcement learning comprise Proximal Policy Optimization (PPO) or Asynchronous Advantage Actor-Critic (A3C) algorithms.
Mermoud discloses wherein on-policy algorithms for reinforcement learning comprise Proximal Policy Optimization (PPO) or Asynchronous Advantage Actor-Critic (A3C) algorithms [“update the policy by using an appropriate algorithm, such as any of the following: Policy-based (e.g., Proximal Policy Optimization (PPO))” ¶¶138-139].
It would have been obvious to one having ordinary skill in the art, having the teachings of Xiao, Lucas, Chunduru, Wang, Vasiljevic, and Mermoud before him before the effective filing date of the claimed invention, to modify the combination to incorporate an on-policy algorithm for reinforcement learning such as Proximal Policy Optimization (PPO) of Mermoud.
Given the advantage of using an efficient policy for stable training, one having ordinary skill in the art would have been motivated to make this obvious modification.
Claims 10 and 26 are rejected on the same grounds as claims 2.
Claim(s) 8 and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Xiao, Lucas, Chunduru, Wang, and Vasiljevic, further in view of Brownlee, Blending Ensemble Machine Learning With Python.
Regarding Claim 8, Xiao, Lucas, Chunduru, Wang, and Vasiljevic disclose the computing system of claim 1.
However, Mermoud fails to explicitly disclose wherein model blending comprises selecting from or blending outputs from multiple models or authoritative knowledge bases, each trained on, or obtained from, defined resource collections or with different retrieval strategies or hyperparameters.
Brownlee discloses wherein model blending comprises selecting from or blending outputs from multiple models or authoritative knowledge bases, each trained on, or obtained from, defined resource collections or with different retrieval strategies or hyperparameters ["Blending may suggest developing a stacking ensemble where the base-models are machine learning models of any type, and the meta-model is a linear model that "blends" the predictions of the base models." pg. 3; Examiner Note: Model blending is recited for a technique in the alternative in claim 1. However, the interpretation and rejection do not use model blending. Therefore, under the broadest reasonable interpretation of the claim, prior art would not be required to be applied to this limitation to fully reject the claim. However, in the interest of compact prosecution, the current art rejection has been applied.].
It would have been obvious to one having ordinary skill in the art, having the teachings of Xiao, Lucas, Chunduru, Wang, Vasiljevic, and Brownlee before him before the effective filing date of the claimed invention, to modify the combination to incorporate the blending of Brownlee.
Given the advantage of boosted prediction accuracy, one having ordinary skill in the art would have been motivated to make this obvious modification.
Claim 16 is rejected on the same grounds as claim 8.
Claim(s) 31, 33 and 35 is/are rejected under 35 U.S.C. 103 as being unpatentable over Xiao, Lucas, Chunduru, Wang, and Vasiljevic, further in view of Haarnoja et al. (hereinafter Haarnoja), Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor.
Regarding Claim 31, Xiao, Lucas, Chunduru, Wang, and Vasiljevic disclose the computing system of claim 1.
However, Xiao fails to explicitly disclose wherein the off-policy algorithms for reinforcement learning comprise Soft Action-Critic (SAC) algorithms.
Haarnoja discloses wherein the off-policy algorithms for reinforcement learning comprise Soft Action-Critic (SAC) algorithms [“we propose soft actor-critic, an off policy actor-critic deep RL algorithm based on the maximum entropy reinforcement learning frame work.” Abstract; Examiner Note: Off-policy algorithms are recited for a technique in the alternative in claim 1. However, the interpretation and rejection do not use off-policy algorithms. Therefore, under the broadest reasonable interpretation of the claim, prior art would not be required to be applied to this limitation to fully reject the claim. However, in the interest of compact prosecution, the current art rejection has been applied].
It would have been obvious to one having ordinary skill in the art, having the teachings of Xiao, Lucas, Chunduru, Wang, Vasiljevic, and Haarnoja before him before the effective filing date of the claimed invention, to modify the combination to incorporate soft actor-critic, an off policy actor-critic deep RL algorithm of Haarnoja .
Given the advantage of achieving state-of-the-art performance on a range of continuous control bench mark tasks, outperforming prior on-policy and off-policy methods, being very stable, and achieving very similar performance across different random seeds, one having ordinary skill in the art would have been motivated to make this obvious modification.
Claims 33 and 35 are rejected on the same grounds as claim 31.
Claim(s) 32, 34 and 36 is/are rejected under 35 U.S.C. 103 as being unpatentable over Xiao, Lucas, Chunduru, Wang, and Vasiljevic, further in view of Shrivastava et al. (hereinafter Shrivastava), U.S. Patent Application Publication 2021/0374338.
Regarding Claim 32, Xiao, Lucas, Chunduru, Wang, and Vasiljevic discloses the computing system of claim 1.
However, Xiao fails to explicitly disclose wherein the task-specific reward function comprises recall-oriented understudy for gisting evaluation (ROUGE) or bilingual evaluation understudy (BLEU) scores algorithms.
Shrivastava discloses wherein the task-specific reward function comprises recall-oriented understudy for gisting evaluation (ROUGE) or bilingual evaluation understudy (BLEU) scores algorithms [“the r(.) function is ROUGE (Recall Oriented Understudy for Gisting Evaluation) function indicating a comparison between generated text summaries and reference text summaries. In other words, the ROUGE is used as a reward” ¶83].
It would have been obvious to one having ordinary skill in the art, having the teachings of Xiao, Lucas, Chunduru, Wang, Vasiljevic, and Shrivastava before him before the effective filing date of the claimed invention, to modify the combination to incorporate the ROUGE reward function of Shrivastava.
Given the advantage of quick and efficient measuring for a reward, one having ordinary skill in the art would have been motivated to make this obvious modification.
Claims 34 and 36 are rejected on the same grounds as claims 32.
Examiner’s Note
The Examiner respectfully requests of the Applicant in preparing responses, to fully consider the entirety of the reference(s) as potentially teaching all or part of the claimed invention. It is noted, REFERENCES ARE RELEVANT AS PRIOR ART FOR ALL THEY CONTAIN. “The use of patents as references is not limited to what the patentees describe as their own inventions or to the problems with which they are concerned. They are part of the literature of the art, relevant for all they contain.” In re Heck, 699 F.2d 1331, 1332-33, 216 USPQ 1038, 1039 (Fed. Cir. 1983) (quoting In re Lemelson, 397 F.2d 1006, 1009, 158 USPQ 275, 277 (CCPA 1968)). A reference may be relied upon for all that it would have reasonably suggested to one having ordinary skill in the art, including non-preferred embodiments (see MPEP 2123). The Examiner has cited particular locations in the reference(s) as applied to the claim(s) above for the convenience of the Applicant. Although the specified citations are representative of the teachings of the art and are applied to the specific limitations within the individual claim(s), typically other passages and figures will apply as well.
Additionally, any claim amendments for any reason should include remarks indicating clear support in the originally filed specification.
Response to Arguments
Regarding the 101 rejections, Applicant's arguments have been fully considered but have been found unpersuasive. Applicant argues that claim 1 has technical improvement as a result of each of the following claimed elements 1) using on-policy or off-policy reinforcement learning, and 2) the final limitation of wherein a distributed computational graph automatically parallelizes optimizing the performance of the AI model, validating the performance of the AI model, optimizing the robustness of the AI model, optimizing the stability of the AI model, and optimizing the reliability of the AI model across heterogeneous computing resources. Examiner disagrees for at least the following reasons.
First, hyperparameters are configuration settings of the model which a person having ordinary skill in the art understands are traditionally designed or selected by the person who is making the model. Therefore, selecting a hyperparameter can reasonable be done by a human mind. However, Applicant argues that reinforcement learning for hyperparameter selection cannot be performed by the human mind. However, that argument is narrower than the claim language requires. As explained in both the previous Office action dated 2/19/2026 as well as in the claim interpretation section of this Office action, that limitation has more than one reasonable interpretation. Under broadest reasonable interpretation, that limitation can be interpreted as either 1) selecting settings for the model and the model is using reinforcement learning, or 2) selecting settings, using reinforcement learning, for the model. Applicant is arguing the second interpretation. However, since the first interpretation is reasonable and can be performed by a human mind, the scope of the limitation is part of the judicial exception.
Second, the final limitation in the previous claim set was an additional element; however, the current amendments to that limitation have removed active language using the computing resources. Accordingly, the current version is part of the abstract idea. Under broadest reasonable interpretation, a distributed computation graph is a graph that shows distributed computation, wherein the graph conveys the parallelization of the several processes among heterogenous computing resources.
For at least these reasons, the rejections are maintained.
Regarding the prior art rejections, Applicant's arguments with respect to the claims have been considered but are moot because the arguments do not apply to the references being used in the current rejection of the limitations.
Conclusion
Any prior art made of record and not relied upon is considered pertinent to Applicant's disclosure. Applicant is reminded that in amending in response to a rejection of claims, the patentable novelty must be clearly shown in view of the state of the art disclosed by the references cited and the objections made. Applicant must also show how the amendments avoid such references and objections. See 37 CFR §1.111(c). Additionally when amending, in their remarks Applicant should particularly cite to the supporting paragraphs in the original disclosure for the amendments.
The following references were found during the examination of this patent application and were found to be relevant to patentability. Applicant is advised to review these references prior to responding to this Office action.
Jarrett et al. (U.S. Patent Application Publication 2025/0068919) discloses fine tuning a model using rewards determined from one or more metrics of the task, such as a BLEU (BiLingual Evaluation Understudy) metric.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ROBERT H BEJCEK II whose telephone number is (571)270-3610. The examiner can normally be reached Monday - Friday: 9:00am - 5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michelle T. Bechtold can be reached at (571) 431-0762. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/R.B./ Examiner, Art Unit 2148
/MICHELLE T BECHTOLD/ Supervisory Patent Examiner, Art Unit 2148