Prosecution Insights
Last updated: October 02, 2026
Application No. 18/168,774

REINFORCEMENT LEARNING FOR OPTIMIZING CROSS-CHANNEL COMMUNICATIONS

Final Rejection §101§103
Filed
Feb 14, 2023
Priority
Aug 16, 2022 — provisional 63/371,552
Examiner
SMITH, KEVIN LEE
Art Unit
2122
Tech Center
2100 — Computer Architecture & Software
Assignee
Optum Inc.
OA Round
2 (Final)
37%
Grant Probability
At Risk
3-4
OA Rounds
12m
Est. Remaining
57%
With Interview

Examiner Intelligence

Grants only 37% of cases
37%
Career Allowance Rate
52 granted / 141 resolved
-18.1% vs TC avg
Strong +20% interview lift
Without
With
+20.0%
Interview Lift
resolved cases with interview
Typical timeline
4y 7m
Avg Prosecution
31 currently pending
Career history
184
Total Applications
across all art units

Statute-Specific Performance

§101
31.2%
-8.8% vs TC avg
§103
40.3%
+0.3% vs TC avg
§102
10.7%
-29.3% vs TC avg
§112
13.3%
-26.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 141 resolved cases

Office Action

§101 §103
DETAILED ACTION 1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . 2. Applicant’s submission filed 08 May 2026 [hereinafter Response] has been entered, where: Claims 1-19 have been amended. Claims 1-20 are pending. Claims 1-20 are rejected. Claim Rejections - 35 U.S.C. § 101 3. 35 U.S.C. § 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. 4. Claims 1-20 are rejected under 35 U.S.C. § 101 because the claimed invention is directed to an abstract idea without significantly more. Claim 1 recites a computer-implemented method, which is a process, and thus one of the statutory categories of patentable subject matter. (35 U.S.C. § 101). However, under Step 2A Prong One, the claim recites the limitations of “providing, by the one or more processors, the set of one or more sequence embeddings to a predictive machine learning model to receive a prediction output comprising an agent action of a set of one or more agent actions that correspond to a predictive software agent.” The activity of “providing . . . the set of one or more sequence embeddings to a predictive machine learning model to receive a prediction output” includes limitations that can practically be performed in the human mind (prediction output), including, for example, observations, evaluations, judgments, and opinions, and accordingly, are mental processes, (MPEP § 2106.04(a)(2) sub III), which is one of the groupings of abstract ideas. (MPEP § 2106.04(a)(2)). The claim also recites more details or specifics to the abstract idea of “providing . . . the set of one or more sequence embeddings to a predictive machine learning model to receive a prediction output” “wherein: “(i) the agent action comprises a highest score of a set of one or more scores corresponding to the set of one or more agent actions,” “(ii) the highest score is determined from an action distribution based on an action reward value that corresponds to the agent action,” and “(iv) a set of one or more model parameters of the predictive machine learning model are modified,” and accordingly, are merely more specific to the abstract idea. Moreover, the activities of “determined” and “model parameters modified,” include limitations that can practically be performed in the human mind (prediction output), including, for example, observations, evaluations, judgments, and opinions, and accordingly, are mental processes, (MPEP § 2106.04(a)(2) sub III), which is one of the groupings of abstract ideas. (MPEP § 2106.04(a)(2)). Also, the claim recites more details or specifics to the abstract idea of “(iv) a set of one or model parameters of the predictive machine learning are modified” by “(b) generating a training action reward value for a training agent interaction of the set of one or more training agent interactions by applying a reward function to the training agent interaction,” “(c) combining a subset of the set of one or more training agent interactions, including the training agent interaction into a historical episode,” “(e) generating set of one or more historical episode combinations, . . .” and “(f) modifying the set of one or more model parameters based on the historical episode combination,” and accordingly, are merely more specific to the abstract idea. Moreover, the activities of “generating . . . a reward value,” “combining a subset,” “generating set,” and “modifying the set of . . . model parameters,” include limitations that can practically be performed in the human mind, including, for example, observations, evaluations, judgments, and opinions, and accordingly, are mental processes, (MPEP § 2106.04(a)(2) sub III), which is one of the groupings of abstract ideas. (MPEP § 2106.04(a)(2)). Further, the claim recites more details or specifics of the abstract idea of “(e) generating set of one or more historical episode combinations,” wherein: “(1) a historical episode combination of the set of one or more historical episode combinations comprises a subset of the set of one or more historic episodes,” and “(2) the subset of the set of one or more historic episodes comprise a most recent historical episode,” and accordingly, are merely more specific to the abstract idea. Accordingly, claim 1 recites an abstract idea. Under Step 2A Prong Two, the claim as a whole is not integrated into a practical application, because the additional elements recited in the claim beyond the identified judicial exception include “one or more processors,” “one or more communication channels,” and a “historical database,” which are recited at a high-level of generality, and accordingly, are generic computer components used to implement the abstract idea, (MPEP § 2106.05(f)), that does not serve to integrate the abstract idea into a practical application. Also, the claim recites “a state encoder machine learning model,” and a ”predictive software agent machine learning model,” which are recited at a high-level of generality, and accordingly are generic computer components used to implement the abstract idea, (MPEP § 2106.05(f)), that does not serve to integrate the abstract idea into a practical application. The claim also recites “providing, by the one or more processors, a historical event sequence data to a state encoder machine learning model to receive a set of one or more sequence embeddings,” “(iv) a set of one or model parameters of the predictive machine learning are modified” by “(a) retrieving training agent interaction data comprising a set of one or more training agent interactions that are associated with a set of one or more training agent actions, wherein the set of one or more training agent interactions corresponds to the set of one or more communication channels,” and “(d) storing the historical episode to a historical database that comprises a set of one or more historical episodes.” The activities of “retrieving” and “storing” are pre-processing and post-processing insignificant extra-solution activities of receiving and storing data, (MPEP § 2106.05(g)), that does not serve to integrate the abstract idea into a practical application. The claim also recites “initiating , by the one or more processors, the performance of the one or more optimal agent actions.” The activity of “initiating” is a post-processing insignificant extra-solution activities data output, respectively, (MPEP § 2106.05(g)), that do not serve to integrate the abstract idea into a practical application. Further, the plain meaning of “initiating, by the one or more processors, the performance of the one or more optimal agent actions” includes outputting data, which is not inconsistent with the Applicant’s disclosure, (MPEP § 2111; see, e.g., Specification ¶ 0081), and accordingly, under a broadest reasonable interpretation, is directed to the post-processing insignificant extra-solution activity of data output, (MPEP § 2106.05(g)), that does not serve to integrate the abstract idea into a practical application. Also, the claim recites more details or specifics to the additional element of “providing . . . a historical event sequence data” “wherein the historical event sequence data comprises usage of a communication channel of a set of one or more communication channels,” and accordingly, is merely more specific to the additional element. Therefore, claim 1 is directed to the abstract idea. Finally, under Step 2B, the additional elements, taken alone or in combination, do not represent significantly more than the abstract idea itself. The additional elements include “one or more processors,” “one or more communication channels,” and a “historical database,” which are recited at a high-level of generality, and accordingly, are generic computer components used to implement the abstract idea, (MPEP § 2106.05(f)), that does not amount to significantly more than the abstract idea. Also, the claim recites “a state encoder machine learning model,” and a ”predictive software agent machine learning model,” which are recited at a high-level of generality, and accordingly are generic computer components used to implement the abstract idea, (MPEP § 2106.05(f)), that does not amount to significantly more than the abstract idea. The claim also recites “providing, by the one or more processors, a historical event sequence data to a state encoder machine learning model to receive a set of one or more sequence embeddings,” “(iv) a set of one or model parameters of the predictive machine learning are modified” by “(a) retrieving training agent interaction data comprising a set of one or more training agent interactions that are associated with a set of one or more training agent actions, wherein the set of one or more training agent interactions corresponds to the set of one or more communication channels,” and “(d) storing the historical episode to a historical database that comprises a set of one or more historical episodes.” The activities of “retrieving” and “storing” are well-understood, routine, and conventional activities of storing and retrieving information in memory, (MPEP § 2106.05(d) sub II.iv), that does not amount to significantly more than the abstract idea. The claim also recites “initiating , by the one or more processors, the performance of the one or more optimal agent actions.” The activity of “initiating” is a well-understood, routine, and conventional activity of a result output (MPEP § 2016.05(d) sub II.i), that does not amount to significantly more than the abstract idea. Further, the plain meaning of “initiating , by the one or more processors, the performance of the one or more optimal agent actions” includes outputting data, which is not inconsistent with the Applicant’s disclosure, (MPEP § 2111; see, e.g., Specification ¶ 0081), and accordingly, under a broadest reasonable interpretation, is directed to the well-understood, routine, and conventional activity of data output, (MPEP § 2106.05(d) sub II.i) , that does not amount to significantly more than the abstract idea. Also, the claim recites more details or specifics to the additional element of “providing . . . a historical event sequence data” “wherein the historical event sequence data comprises usage of a communication channel of a set of one or more communication channels,” and accordingly, is merely more specific to the additional element. Therefore, claim 1 is subject-matter ineligible. Claim 9 recites a system, which is a product, and thus one of the statutory categories of patentable subject matter. (35 U.S.C. § 101). However, under Step 2A Prong One, the claim recites the limitations of “providing, by the one or more processors, the set of one or more sequence embeddings to a predictive machine learning model to receive a prediction output comprising an agent action of a set of one or more agent actions that correspond to a predictive software agent.” The activity of “providing . . . the set of one or more sequence embeddings to a predictive machine learning model to receive a prediction output” includes limitations that can practically be performed in the human mind (prediction output), including, for example, observations, evaluations, judgments, and opinions, and accordingly, are mental processes, (MPEP § 2106.04(a)(2) sub III), which is one of the groupings of abstract ideas. (MPEP § 2106.04(a)(2)). The claim also recites more details or specifics to the abstract idea of “providing . . . the set of one or more sequence embeddings to a predictive machine learning model to receive a prediction output” “wherein: “(i) the agent action comprises a highest score of a set of one or more scores corresponding to the set of one or more agent actions,” “(ii) the highest score is determined from an action distribution based on an action reward value that corresponds to the agent action,” and “(iv) a set of one or more model parameters of the predictive machine learning model are modified,” and accordingly, are merely more specific to the abstract idea. Moreover, the activities of “determined” and “model parameters modified,” include limitations that can practically be performed in the human mind (prediction output), including, for example, observations, evaluations, judgments, and opinions, and accordingly, are mental processes, (MPEP § 2106.04(a)(2) sub III), which is one of the groupings of abstract ideas. (MPEP § 2106.04(a)(2)). Also, the claim recites more details or specifics to the abstract idea of “(iv) a set of one or model parameters of the predictive machine learning are modified” by “(b) generating a training action reward value for a training agent interaction of the set of one or more training agent interactions by applying a reward function to the training agent interaction,” “(c) combining a subset of the set of one or more training agent interactions, including the training agent interaction into a historical episode,” “(e) generating set of one or more historical episode combinations, . . .” and “(f) modifying the set of one or more model parameters based on the historical episode combination,” and accordingly, are merely more specific to the abstract idea. Moreover, the activities of “generating . . . a reward value,” “combining a subset,” “generating set,” and “modifying the set of . . . model parameters,” include limitations that can practically be performed in the human mind, including, for example, observations, evaluations, judgments, and opinions, and accordingly, are mental processes, (MPEP § 2106.04(a)(2) sub III), which is one of the groupings of abstract ideas. (MPEP § 2106.04(a)(2)). Further, the claim recites more details or specifics of the abstract idea of “(e) generating set of one or more historical episode combinations,” wherein: “(1) a historical episode combination of the set of one or more historical episode combinations comprises a subset of the set of one or more historic episodes,” and “(2) the subset of the set of one or more historic episodes comprise a most recent historical episode,” and accordingly, are merely more specific to the abstract idea. Accordingly, claim 9 recites an abstract idea. Under Step 2A Prong Two, the claim as a whole is not integrated into a practical application, because the additional elements recited in the claim beyond the identified judicial exception include “one or more processors,” “one or more non-transitory computer readable media storing processor-executable instructions,” “one or more communication channels,” and a “historical database,” which are recited at a high-level of generality, and accordingly, are generic computer components used to implement the abstract idea, (MPEP § 2106.05(f)), that does not serve to integrate the abstract idea into a practical application. Also, the claim recites “a state encoder machine learning model,” and a ”predictive software agent machine learning model,” which are recited at a high-level of generality, and accordingly are generic computer components used to implement the abstract idea, (MPEP § 2106.05(f)), that does not serve to integrate the abstract idea into a practical application. The claim also recites “providing, by the one or more processors, a historical event sequence data to a state encoder machine learning model to receive a set of one or more sequence embeddings,” “(iv) a set of one or model parameters of the predictive machine learning are modified” by “(a) retrieving training agent interaction data comprising a set of one or more training agent interactions that are associated with a set of one or more training agent actions, wherein the set of one or more training agent interactions corresponds to the set of one or more communication channels,” and “(d) storing the historical episode to a historical database that comprises a set of one or more historical episodes.” The activities of “retrieving” and “storing” are pre-processing and post-processing insignificant extra-solution activities of receiving and storing data, (MPEP § 2106.05(g)), that does not serve to integrate the abstract idea into a practical application. The claim also recites “initiating , by the one or more processors, the performance of the one or more optimal agent actions.” The activity of “initiating” is a post-processing insignificant extra-solution activities data output, respectively, (MPEP § 2106.05(g)), that do not serve to integrate the abstract idea into a practical application. Further, the plain meaning of “initiating , by the one or more processors, the performance of the one or more optimal agent actions” includes outputting data, which is not inconsistent with the Applicant’s disclosure, (MPEP § 2111; see, e.g., Specification ¶ 0081), and accordingly, under a broadest reasonable interpretation, is directed to the post-processing insignificant extra-solution activity of data output, (MPEP § 2106.05(g)), that does not serve to integrate the abstract idea into a practical application. Also, the claim recites more details or specifics to the additional element of “providing . . . a historical event sequence data” “wherein the historical event sequence data comprises usage of a communication channel of a set of one or more communication channels,” and accordingly, is merely more specific to the additional element. Therefore, claim 9 is directed to the abstract idea. Finally, under Step 2B, the additional elements, taken alone or in combination, do not represent significantly more than the abstract idea itself. The additional elements include “one or more processors,” “one or more non-transitory computer readable media storing processor-executable instructions,” “one or more communication channels,” and a “historical database,” which are recited at a high-level of generality, and accordingly, are generic computer components used to implement the abstract idea, (MPEP § 2106.05(f)), that does not amount to significantly more than the abstract idea. Also, the claim recites “a state encoder machine learning model,” and a ”predictive software agent machine learning model,” which are recited at a high-level of generality, and accordingly are generic computer components used to implement the abstract idea, (MPEP § 2106.05(f)), that does not amount to significantly more than the abstract idea. The claim also recites “providing, by the one or more processors, a historical event sequence data to a state encoder machine learning model to receive a set of one or more sequence embeddings,” “(iv) a set of one or model parameters of the predictive machine learning are modified” by “(a) retrieving training agent interaction data comprising a set of one or more training agent interactions that are associated with a set of one or more training agent actions, wherein the set of one or more training agent interactions corresponds to the set of one or more communication channels,” and “(d) storing the historical episode to a historical database that comprises a set of one or more historical episodes.” The activities of “retrieving” and “storing” are well-understood, routine, and conventional activities of storing and retrieving information in memory, (MPEP § 2106.05(d) sub II.iv), that does not amount to significantly more than the abstract idea. The claim also recites “initiating , by the one or more processors, the performance of the one or more optimal agent actions.” The activity of “initiating” is a well-understood, routine, and conventional activity of a result output (MPEP § 2016.05(d) sub II.i), that does not amount to significantly more than the abstract idea. Further, the plain meaning of “initiating , by the one or more processors, the performance of the one or more optimal agent actions” includes outputting data, which is not inconsistent with the Applicant’s disclosure, (MPEP § 2111; see, e.g., Specification ¶ 0081), and accordingly, under a broadest reasonable interpretation, is directed to the well-understood, routine, and conventional activity of data output, (MPEP § 2106.05(d) sub II.i) , that does not amount to significantly more than the abstract idea. Also, the claim recites more details or specifics to the additional element of “providing . . . a historical event sequence data” “wherein the historical event sequence data comprises usage of a communication channel of a set of one or more communication channels,” and accordingly, is merely more specific to the additional element. Therefore, claim 9 is subject-matter ineligible. Claim 16 recites a one or more non-transitory computer-readable storage media, which is a product, and thus one of the statutory categories of patentable subject matter. (35 U.S.C. § 101). However, under Step 2A Prong One, the claim recites the limitations of “providing, by the one or more processors, the set of one or more sequence embeddings to a predictive machine learning model to receive a prediction output comprising an agent action of a set of one or more agent actions that correspond to a predictive software agent.” The activity of “providing . . . the set of one or more sequence embeddings to a predictive machine learning model to receive a prediction output” includes limitations that can practically be performed in the human mind (prediction output), including, for example, observations, evaluations, judgments, and opinions, and accordingly, are mental processes, (MPEP § 2106.04(a)(2) sub III), which is one of the groupings of abstract ideas. (MPEP § 2106.04(a)(2)). The claim also recites more details or specifics to the abstract idea of “providing . . . the set of one or more sequence embeddings to a predictive machine learning model to receive a prediction output” “wherein: “(i) the agent action comprises a highest score of a set of one or more scores corresponding to the set of one or more agent actions,” “(ii) the highest score is determined from an action distribution based on an action reward value that corresponds to the agent action,” and “(iv) a set of one or more model parameters of the predictive machine learning model are modified,” and accordingly, are merely more specific to the abstract idea. Moreover, the activities of “determined” and “model parameters modified,” include limitations that can practically be performed in the human mind (prediction output), including, for example, observations, evaluations, judgments, and opinions, and accordingly, are mental processes, (MPEP § 2106.04(a)(2) sub III), which is one of the groupings of abstract ideas. (MPEP § 2106.04(a)(2)). Also, the claim recites more details or specifics to the abstract idea of “(iv) a set of one or model parameters of the predictive machine learning are modified” by “(b) generating a training action reward value for a training agent interaction of the set of one or more training agent interactions by applying a reward function to the training agent interaction,” “(c) combining a subset of the set of one or more training agent interactions, including the training agent interaction into a historical episode,” “(e) generating set of one or more historical episode combinations, . . .” and “(f) modifying the set of one or more model parameters based on the historical episode combination,” and accordingly, are merely more specific to the abstract idea. Moreover, the activities of “generating . . . a reward value,” “combining a subset,” “generating set,” and “modifying the set of . . . model parameters,” include limitations that can practically be performed in the human mind, including, for example, observations, evaluations, judgments, and opinions, and accordingly, are mental processes, (MPEP § 2106.04(a)(2) sub III), which is one of the groupings of abstract ideas. (MPEP § 2106.04(a)(2)). Further, the claim recites more details or specifics of the abstract idea of “(e) generating set of one or more historical episode combinations,” wherein: “(1) a historical episode combination of the set of one or more historical episode combinations comprises a subset of the set of one or more historic episodes,” and “(2) the subset of the set of one or more historic episodes comprise a most recent historical episode,” and accordingly, are merely more specific to the abstract idea. Accordingly, claim 16 recites an abstract idea. Under Step 2A Prong Two, the claim as a whole is not integrated into a practical application, because the additional elements recited in the claim beyond the identified judicial exception include “[o]ne or more non-transitory computer-readable storage media including storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations,” which are recited at a high-level of generality, and accordingly, are generic computer components used to implement the abstract idea, (MPEP § 2106.05(f)), that does not serve to integrate the abstract idea into a practical application. Also, the claim recites “a state encoder machine learning model,” and a ”predictive software agent machine learning model,” which are recited at a high-level of generality, and accordingly are generic computer components used to implement the abstract idea, (MPEP § 2106.05(f)), that does not serve to integrate the abstract idea into a practical application. The claim also recites “providing, by the one or more processors, a historical event sequence data to a state encoder machine learning model to receive a set of one or more sequence embeddings,” “(iv) a set of one or model parameters of the predictive machine learning are modified” by “(a) retrieving training agent interaction data comprising a set of one or more training agent interactions that are associated with a set of one or more training agent actions, wherein the set of one or more training agent interactions corresponds to the set of one or more communication channels,” and “(d) storing the historical episode to a historical database that comprises a set of one or more historical episodes.” The activities of “retrieving” and “storing” are pre-processing and post-processing insignificant extra-solution activities of receiving and storing data, (MPEP § 2106.05(g)), that does not serve to integrate the abstract idea into a practical application. The claim also recites “initiating , by the one or more processors, the performance of the one or more optimal agent actions.” The activity of “initiating” is a post-processing insignificant extra-solution activities data output, respectively, (MPEP § 2106.05(g)), that do not serve to integrate the abstract idea into a practical application. Further, the plain meaning of “initiating , by the one or more processors, the performance of the one or more optimal agent actions” includes outputting data, which is not inconsistent with the Applicant’s disclosure, (MPEP § 2111; see, e.g., Specification ¶ 0081), and accordingly, under a broadest reasonable interpretation, is directed to the post-processing insignificant extra-solution activity of data output, (MPEP § 2106.05(g)), that does not serve to integrate the abstract idea into a practical application. Also, the claim recites more details or specifics to the additional element of “providing . . . a historical event sequence data” “wherein the historical event sequence data comprises usage of a communication channel of a set of one or more communication channels,” and accordingly, is merely more specific to the additional element. Therefore, claim 16 is directed to the abstract idea. Finally, under Step 2B, the additional elements, taken alone or in combination, do not represent significantly more than the abstract idea itself. The additional elements include “one or more processors,” “one or more non-transitory computer readable media storing processor-executable instructions,” “one or more communication channels,” and a “historical database,” which are recited at a high-level of generality, and accordingly, are generic computer components used to implement the abstract idea, (MPEP § 2106.05(f)), that does not amount to significantly more than the abstract idea. Also, the claim recites “a state encoder machine learning model,” and a ”predictive software agent machine learning model,” which are recited at a high-level of generality, and accordingly are generic computer components used to implement the abstract idea, (MPEP § 2106.05(f)), that does not amount to significantly more than the abstract idea. The claim also recites “providing, by the one or more processors, a historical event sequence data to a state encoder machine learning model to receive a set of one or more sequence embeddings,” “(iv) a set of one or model parameters of the predictive machine learning are modified” by “(a) retrieving training agent interaction data comprising a set of one or more training agent interactions that are associated with a set of one or more training agent actions, wherein the set of one or more training agent interactions corresponds to the set of one or more communication channels,” and “(d) storing the historical episode to a historical database that comprises a set of one or more historical episodes.” The activities of “retrieving” and “storing” are well-understood, routine, and conventional activities of storing and retrieving information in memory, (MPEP § 2106.05(d) sub II.iv), that does not amount to significantly more than the abstract idea. The claim also recites “initiating , by the one or more processors, the performance of the one or more optimal agent actions.” The activity of “initiating” is a well-understood, routine, and conventional activity of a result output (MPEP § 2016.05(d) sub II.i), that does not amount to significantly more than the abstract idea. Further, the plain meaning of “initiating , by the one or more processors, the performance of the one or more optimal agent actions” includes outputting data, which is not inconsistent with the Applicant’s disclosure, (MPEP § 2111; see, e.g., Specification ¶ 0081), and accordingly, under a broadest reasonable interpretation, is directed to the well-understood, routine, and conventional activity of data output, (MPEP § 2106.05(d) sub II.i) , that does not amount to significantly more than the abstract idea. Also, the claim recites more details or specifics to the additional element of “providing . . . a historical event sequence data” “wherein the historical event sequence data comprises usage of a communication channel of a set of one or more communication channels,” and accordingly, is merely more specific to the additional element. Therefore, claim 16 is subject-matter ineligible. Claims 2, 7, and 8 depend directly or indirectly from claim 1. Claims 10, 14, and 15 depend directly or indirectly from claim 9. Claims 19 and 20 depend directly or indirectly from claim 16. The claims recite more details or specifics of the additional element of the “predictive machine learning model,” (claims 2 and 10: wherein the predictive software agent machine learning model comprises a reinforcement learning machine learning model”; claims 7, 14, and 19: wherein the predictive machine learning model comprises a deep Q network”; claims 8, 15, and 20: “wherein the deep Q network comprises an exploration phase and an exploitation phase”), and accordingly, are merely more specific to the additional element. The abstract idea of these claims are not integrated into a practical application, (see MPEP § 2106.04(d)), nor do they amount to significantly more than the abstract idea, (MPEP § 2106.05 sub I; see also MPEP § 2106.05(a) – (h)), because the claims recite no more than the abstract idea. Therefore, claims 2, 7, 8, 10, 14, 15, 19 and 20 are subject-matter ineligible. Claim 3 depends from claim 1. Claim 11 depends from claim 9. Claim 17 depends from claim 16. The claims recite further limitations of “determining one or more monitored agent interactions associated with the performance of the agent actions,” and “generating the training agent interaction data based on the one or more agent interactions.” These activities of “monitoring” and “generating the training agent interaction data” are limitations that can practically be performed in the human mind, including, for example, observations, evaluations, judgments, and opinions, and accordingly, are mental processes, (MPEP § 2106.04(a)(2) sub III), which is one of the groupings of abstract ideas. (MPEP § 2106.04(a)(2)). Also, because the broadest reasonable interpretation of “generating” covers the selecting and/or collecting of information from the “monitoring,” which is not inconsistent with the Applicant’s disclosure, (MPEP § 2111), the activity of “generating” is a limitation that can practically be performed in the human mind. The abstract idea of these claims are not integrated into a practical application, (see MPEP § 2106.04(d)), nor do they amount to significantly more than the abstract idea, (MPEP § 2106.05 sub I; see also MPEP § 2106.05(a) – (h)), because the claims recite no more than the abstract idea. Therefore, claims 3, 11, and 17 are subject-matter ineligible. Claim 4 depends from claim 1. The claim recites further limitation of “filtering the action distribution based on one or more Boolean flags, wherein the one or more Boolean flags correspond to one or more selections of the set of one or more agent actions.” The activity of “filtering agent actions” is a limitation that can practically be performed in the human mind, including, for example, observations, evaluations, judgments, and opinions, and accordingly, a mental process, (MPEP § 2106.04(a)(2) sub III), which is one of the groupings of abstract ideas. (MPEP § 2106.04(a)(2)). The abstract idea of these claims are not integrated into a practical application, (see MPEP § 2106.04(d)), nor do they amount to significantly more than the abstract idea, (MPEP § 2106.05 sub I; see also MPEP § 2106.05(a) – (h)), because the claims recite no more than the abstract idea. Therefore, claim 4 is subject-matter ineligible. Claim 5 depends from claim 1. Claim 12 depends from claim 9. Claim 18 depends from claim 16. The claims recite more details or specifics to the additional element of “ “historical event sequence data”, “wherein the historical event sequence data comprises state representation data, and the computer-implemented method further comprises transforming the state representation data into the set of one or more sequence embeddings by: generating tokenized historical event sequence data based on the historical event sequence data; and normalizing the tokenized historical event sequence data into fixed-length vectors,” and accordingly, are merely more specific to the additional element. . Also, the activities of “tokenizing,” and “normalizing the tokenized historical event sequence data into the fixed-length vectors,” are converting information from one format to another, and thus are limitations that can practically be performed in the human mind, including, for example, observations, evaluations, judgments, and opinions, and accordingly, a mental process, (MPEP § 2106.04(a)(2) sub III), which is one of the groupings of abstract ideas. (MPEP § 2106.04(a)(2)). Therefore, claims 5, 12, and 18 are subject-matter ineligible. Claim 6 depends from claim 1. Claim 13 depends from claim 9. The claims recite more details or specifics to the abstract idea of “providing . . . the set of one or more sequence embeddings to a predictive machine learning model to receive a prediction output,” “ (ii) wherein the action distribution comprises the set of one or more training agent actions and a set of one or more action reward values that corresponds to the set of one or more training agent actions,” and accordingly, are merely more specific to the abstract idea. The abstract idea of these claims are not integrated into a practical application, (see MPEP § 2106.04(d)), nor do they amount to significantly more than the abstract idea, (MPEP § 2106.05 sub I; see also MPEP § 2106.05(a) – (h)), because the claims recite no more than the abstract idea. Therefore, claims 6 and 13 are subject-matter ineligible. Claim Rejections – 35 U.S.C. § 103 5. The following is a quotation of 35 U.S.C. § 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 6. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. § 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. 7. This application currently names joint inventors. In considering patentability of the claims the Examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the Examiner to consider the applicability of 35 U.S.C. § 102(b)(2)(C) for any potential 35 U.S.C. § 102(a)(2) prior art against the later invention. 8. Claims 1-4, 6-11, 13-17, 19 and 20 are rejected under 35 U.S.C. § 103 as being unpatentable over US Published Application 20210150417 to Fadel et al. [hereinafter Fadel] in view US Published Application 20220215244 to Le [hereinafter Le]. Regarding claims 1, 9, and 16, Fadel teaches [a] computer-implemented method (Fadel, Abstract, teaches “method for reinforcement machine learning uses a reinforcement learning system that has an environment and an agent”) of claim 1, [a] system comprising: one or more processors and one or more non-transitory computer readable media storing processor-executable instructions that , when executed by the one or more processors, cause the one or more processors to perform operations (Fadel ¶ 0179 teaches “[p]rocessors 602 can perform operations embodying the function, method, or operation by, for example, executing code (e.g., interpreting scripts) stored on memory 604; Fadel ¶ 0182 teaches “memory 604 include a non-transitory computer-readable media . . . that can be used to store program code in the form of instructions or data structures, and the like”) of claim 9, and [o]ne or more non-transitory computer-readable storage media storing instructions (Fadel ¶ 0182 teaches “[e]xamples of memory 604 include a non-transitory computer-readable media) of claim 16, comprising: providing, by one or more processors, a historical event sequence data to a state encoder machine learning model to receive a set of one or more sequence embeddings (Fadel, Fig. 4, teaches a system that possesses external knowledge and interacts with the agent during training [Examiner annotations in dashed-line text boxes]: PNG media_image1.png 598 784 media_image1.png Greyscale Fadel ¶ 0168 teaches “guide functions, implemented via the tutor 410, [(that is, a state encoder machine learning model)] are programmable functions that express domain heuristics that the agent 420 can use to guide its decisions (especially in moments of high uncertainty, e.g. start of the learning process). Each guide function takes the current state st and reward rt as inputs, and then outputs a vector to represent the weight of each preferred action according to the encoded domain heuristics [(that is, “vector to represent the weight of each preferred action according to encoded domain heuristics” is a set of one or more sequence embeddings)]”; [Examiner notes that the plain meaning of “sequence embeddings” are dense vector representations of ordered sets of events (e.g., transactions, social media posts, medical procedures, or news items) from “historical event sequence data” that capture both their temporal order and semantic meaning. They are derived by encoding the sequence of events into a fixed‑dimensional vector space; accordingly, the broadest reasonable interpretation of “sequence embeddings” covers the teachings of Fadel pertaining to state and action mappings and also the preferred action vectors thereof, which is not inconsistent with the Applicant’s disclosure (MPEP § 2111)]), . . . ; providing, by the one or more processors, the set of one or more sequence embeddings to a predictive machine learning model (Fadel ¶ 0160 teaches “a diagram of Tutor4RL, which is a particular implementation of an RL system [(that is, “Tutor4RL” is a predictive machine learning model)]”) to receive a prediction output comprising an agent action of a set of one or more agent actions that correspond to a predictive software agent (Fadel ¶ 0154 teaches “[a]fter observing the environment, the model applies those observations (Operation 304). Applying the observations can include the agent using its policy to determine its output based on the observed state (with or without the observed reward). As described herein, the output of the agent's policy can include a selection of an action predicted [(that is, a prediction output)] to obtain the highest reward [(that is, with an action obtaining a highest reward is a prediction output comprising an agent action of a set of one more agent actions that correspond to a predictive software agent)]”; Fadel ¶ 0008 teaches that, in “the RL framework, the agent can be the system's manager [(that is, a predictive software agent)] and the system and the environment can be the execution context. The agent can find the best configuration by modifying the configuration parameters (seen as actions)”), wherein: (i) the agent action comprises a highest score of a set of one or more scores corresponding to the set of one or more agent actions (Fadel ¶ 0095 teaches “[t]he values that a guide function outputs can be interpreted as the expected reward [(that is, a set of one or more scores corresponding to the set of one or more agent actions)] after performing each action, (e.g., similar to the values that the agent's policy outputs). In the end, the best action is the action with the highest value [(that is, the highest score)]”), (ii) the highest score is determined from an action distribution based on an action reward value that corresponds to the agent action (Fadel ¶ 0095 teaches “[g]uide functions are also programmable knowledge functions and express guidelines for the agent's behavior. These functions take, as input, the current RL state and reward, and output a vector of size L [(that is, the “vector” is an action distribution based on an action reward value)] , assigning a value that represents how ‘good’ is each action. . . . In the end, the best action is the action with the highest value [(that is, the highest score is determined from an action distribution based on an action reward value that corresponds to the agent action)]”), and initiating, by the one or more processors, performance of the agent action by the predictive software agent (Fadel ¶ 0121 teaches “the action with the highest value is chosen and performed by the agent 240. After the action is chosen and performed, the environment 250 reacts to the action, including possibly changing its state and giving a reward to the agent 240 [(that is, initiating, by the one or more processors, the performance of the one or more optimal agent actions)]”). Though Fadel teaches that a tutor guides an agent to make informed decisions during training of a reinforcement learning model, Fadel, however, does not explicitly teach- * * * [providing . . . a historical event sequence data], wherein the historical event sequence data comprises usage of a communication channel of a set of one or more communication channels; [providing, by the one or more processors, the set of one or more sequence embeddings to a predictive machine learning model], wherein: * * * (iii) the agent action comprises the predictive software agent using a communication channel of the set of one or more communication channels, and (iv) a set of one or more model parameters of the predictive machine learning model are modified by: (a) retrieving training agent interaction data comprising a set of one or more training agent interactions that are associated with a set of one or more training agent actions, wherein the set of one or more training agent interactions corresponds to the set of one or more communication channels, (b) generating a training action reward value for a training agent interaction of the set of one or more training agent interactions by applying a reward function to the training agent interaction, (c) combining a subset of the set of one or more training agent interactions, including the training agent interaction into a historical episode, (d) storing the historical episode to a historical database that comprises a set of one or more historical episodes, (e) generating set of one or more historical episode combinations, wherein: (1) a historical episode combination of the set of one or more historical episode combinations comprises a subset of the set of one or more historic episodes, and (2) the subset of the set of one or more historic episodes comprise a most recent historical episode, and (f) modifying the set of one or more model parameters based on the historical episode combination; and * * * But Le teaches – * * * [providing . . . a historical event sequence data] wherein the historical event sequence data comprises usage of a communication channel of a set of one or more communication channels (Le, Fig. 3, teaches a communication channel of a set of one or more communication channels [Examiner annotations in dashed-line text boxes]: PNG media_image2.png 878 792 media_image2.png Greyscale Le ¶ 0052 teaches that a trained reinforcement learning agent model 302 “may be trained on one or more datasets of information [utilizing] historical offline, synthetic online, and experimental online data sets. Historical offline data may be initial data from historical sessions with combined user response data (e.g., session clicks) having different determined and/or presumed goals [(that is, wherein the historical event sequence data comprises usage of a communication channel of a set of one or more communication channels)]”; Le ¶ 0045 & Fig. 3 teaches that “Communication paths 328, 330, and 332 may include the Internet, a mobile phone network, a mobile voice or data network (e.g., a 5G or LTE network), a cable network, a public switched telephone network, or other types of communications networks or combinations of communications networks [(that is, a set of one or more communication channels)]”); [providing, by the one or more processors, the set of one or more sequence embeddings to a predictive machine learning model], wherein: * * * (iii) the agent action comprises the predictive software agent using a communication channel of the set of one or more communication channels (Le, Fig. 3, teaches generating alternative content via a trained reinforcement learning agent [Examiner annotations in dashed-line text boxes]: PNG media_image3.png 871 843 media_image3.png Greyscale Le ¶ 0043 teaches “the devices [322, 324] may have neither user input interface nor displays and may instead receive and display content using another device . . . . Additionally, the devices [332, 324] in system 300 may run an application (or another suitable program). The application may cause the processors and/or control circuitry to perform operations related to generating alternative content”; Le ¶ 0045 teaches “[c]ommunication paths 328, 330, and 332 may include the Internet, a mobile phone network, a mobile voice or data network (e.g., a 5G or LTE network), a cable network, a public switched telephone network, or other types of communications networks or combinations of communications networks [(that is, usage of a communication channel of a set of one or more communication channels)]”; Le ¶ 0053 teaches “model 302 [(that is, predictive software agent)] may predict alternative content. For example, the system may determine that particular user data and/or previous user actions are more likely to be indicative of a desired or intent for a particular piece of alternative content. In some embodiments, the model (e.g., model 302) may automatically perform actions based on output 306”), and (iv) a set of one or more model parameters of the predictive machine learning model are modified (Le ¶ 0049 teaches “, model 302 may update [(that is, “update” is modified)] its configurations (e.g., weights, biases, or other parameters) [(that is, a set of one or more model parameters of the predictive machine learning model are modified)] based on the assessment of its prediction (e.g., outputs 306) and reference feedback information (e.g., user indication of accuracy, reference labels, or other information)”) by: (a) retrieving training agent interaction data comprising a set of one or more training agent interactions that are associated with a set of one or more training agent actions (Le ¶ 0012 teaches “actions and returned values may comprise streaming data [combining] the outputs of multiple machine learning models. This combination allows the cross-channel interaction learning, even though the multiple machine learning model may have different algorithms, architecture, training data, and/or goals [(that is, “cross-channel interaction learning” is retrieving training agent interaction data comprising a set of one or more training agent actions that are associated with a set of one or more training agent actions)]”; Le ¶ 0034 teaches “agent 204 may be a reinforcement learning agent which receives observations and a returned value from the environment. Using its policy, the agent selects an action based on the observations and returned value and sends the action to the environment. During training, the agent may continuously update the policy parameters based on the action, observations, and returned value”; also, Le ¶ 0046 & Fig. 3 teaches that “[c]loud components 310 may be a database . . . [that] may include user data that the system has collected about the user through prior interactions”), wherein the set of one or more training agent interactions corresponds to the set of one or more communication channels (Le ¶ 0062 teaches that “’actions’ may be the agent's prediction of the best task (e.g., content) to present. For each action, the agent may collect an algorithmically determined returned value via returned value function [(that is, the set of one or more training agent interactions)]. The actions and returned values may comprise streaming data [combining] the outputs of multiple machine learning models. This combination allows the cross-channel interaction learning [(that is, “cross-channel interaction learning” is corresponds to the set of one or more communication channels)]”), (b) generating a training action reward value for a training agent interaction of the set of one or more training agent interactions by applying a reward function to the training agent interaction (Le, Fig. 5, teaches dynamically selecting alternative content based on real-time events during device sessions through the use of a cross-channel, time-bound deep reinforcement machine learning with a DDPG architecture [Examiner annotations in dashed-line text boxes]: PNG media_image4.png 620 1008 media_image4.png Greyscale Le ¶ 0012 teaches “[f]or each action, the agent may collect an algorithmically determined returned value [(reward r)] via returned value function [(that is, generating a training action reward value for a training agent interaction of the set of one or more training agent interactions by applying a reward function to the training agent interaction)]”), (c) combining a subset of the set of one or more training agent interactions, including the training agent interaction into a historical episode (Le ¶ 0052 teaches “[h]istorical offline data may be initial data from historical sessions with combined user response data (e.g., session clicks) [(that is, combining a subset of the set of one or more training agent interactions, including the training agent interaction into a historical episode)] having different determined and/or presumed goals”), (d) storing the historical episode to a historical database that comprises a set of one or more historical episodes (Le ¶ 0067 teaches the “policy parameters may comprise parameter values, which at any time, may represent accumulated [(that is, combining a subset)] learning experience of agent 502 up to that point [(that is, the historical episode)]. The policy parameters may be stored in a behavior signature database 508. In some embodiments, the behavior signature database may allow system 500 to statistically reenact a user's response [(that is, “statistically reenact” is storing the historical episode to a historical database that comprise a set of one or more historical episodes)]”), (e) generating set of one or more historical episode combinations (Le ¶ 0052 teaches “[h]istorical offline data may be initial data from historical sessions with combined user response data (e.g., session clicks) having different determined and/or presumed goals [(that is, generating set of one . . . historical episode combinations)]. This data may be labeled with a determined goal and/or have associated features and timestamps”), wherein: (1) a historical episode combination of the set of one or more historical episode combinations comprises a subset of the set of one or more historic episodes (Le ¶ 0052 teaches “[m]odel 302 may be trained on one or more datasets of information. For example, the system may utilize historical offline, synthetic online, and experimental online data sets. Historical offline data may be initial data from historical sessions with combined user response data (e.g., session clicks) having different determined and/or presumed goals [(that is, a historical episode combination of the set of one or more historical episode combinations)]. This data may be labeled with a determined goal and/or have associated features and timestamps. Historical data may be updated in real-time [(that is, “labeled data” is a subset of the set of one or more historic episodes)]”), and (2) the subset of the set of one or more historic episodes comprise a most recent historical episode (Le ¶ 0052 teaches “[h]istorical data may be updated in real-time [(that is, “updated historical data” necessarily entails the subset of the set of one or more historic episodes comprise a most recent historical episode)]”), and (f) modifying the set of one or more model parameters based on the historical episode combination (Le ¶ 0049 teaches “model 302 [(a trained reinforcement learning agent)] may update its configurations [(that is, modifying the set of one or more model parameters )] (e.g., weights, biases, or other parameters) based on the assessment of its prediction (e.g., outputs 306) and reference feedback information (e.g., user indication of accuracy, reference labels, or other information) [(that is, “prediction assessment” and “reference feedback information” is based on the historical episode combination)]”); and * * * Fadel and Le are from the same or similar field of endeavor. Fadel teaches reinforcement machine learning uses a reinforcement learning system that has an environment and an agent. Le teaches dynamically selecting alternative content based on real-time events during device sessions using a cross-channel, time-bound deep reinforcement machine learning. Thus, it would have been obvious to a person having ordinary skill in the art as of the effective filing date of the Applicant’s claimed invention to modify Fadel pertaining to reinforcement machine learning with the cross-channel reinforcement learning of Le. The motivation to do so is because “overcome these technical problems with respect to conventional machine learning models, methods and systems are described herein for dynamically selecting alternative content based on real-time events during device sessions through the use of a cross-channel, time-bound deep reinforcement machine learning. The use of this architecture allows for alternative content to be selected in a time-bound and continuous manner that provides predictions in a dynamic environment (e.g., an environment in which user data is continuously changing and new events are continuously occurring) and with an increased success rate (e.g., new data and events are factored into each prediction).” (Le ¶ 0010). Regarding claims 2 and 10, the combination of Fadel and Le teaches all of the limitations of claims 1 and 9, respectively, as described above in detail. Fadel teaches - wherein the predictive machine learning model comprises a reinforcement learning machine learning model (Fadel, Abstract, teaches a “method for reinforcement machine learning uses a reinforcement learning system that has an environment and an agent [(that is, the predictive machine learning model comprises a reinforcement learning machine learning model)]”). Regarding claims 3, 11, and 17, the combination of Fadel and Le teaches all of the limitations of claims 1, 9, and 16, respectively, as described above in detail. Le teaches - further comprising: determining one or more agent interactions associated with the performance of the agent action (Le ¶ 0062 teaches “’actions’ may be the agent's prediction of the best task (e.g., content) to present [(that is, one or more prediction-based actions)]. For each action, the agent may collect an algorithmically determined returned value via returned value function [(that is, “returned value” is determining one or more agent interactions associated with the performance of the agent actions)]”); and generating the training agent interaction data based on the one or more monitored agent interactions (Le ¶ 0012 teaches “actions and returned values [(that is, “returned values” is monitored)] may comprise streaming data [combining] the outputs of multiple machine learning models [(that is, “combining” is generating the training agent interaction data)]. This combination allows the cross-channel interaction learning [(that is, generating the training agent interaction data based on the one or more monitored agent interactions)], even though the multiple machine learning model may have different algorithms, architecture, training data, and/or goals [(that is, “cross-channel interaction learning” is retrieving training agent interaction data comprising a set of one or more training agent actions that are associated with a set of one or more training agent actions)]”). Regarding claim 4, the combination of Fadel and Le teaches all of the limitations of claim 1 as described above in detail. Fadel teaches - filtering the action distribution (Fadel ¶ 0039 teaches a “constraining function being configured to take as its input the current state and to return an action mask indicating which of the actions [(that is, action distribution)] are enabled or disabled [(that is, the “action mask” is filtering the action distribution)]”) based on one or more Boolean flags, wherein the one or more Boolean flags correspond to one or more selections of the set of one or more agent actions (Fadel ¶ 0110 teaches “[e]ach constraining function ƒi(⋅) takes as its input the state of the environment 250 at time step t, and outputs an action mask . . . in the form of a vector with an entry for each action in the action space, where the value of each entry is either a 1 or 0 [(that is, an “action mask” is based on one or more Boolean flags, wherein the one or more Boolean flags correspond to one or more selections of the set of one or more agent actions)]”). Regarding claims 6 and 13, the combination of Fadel and Le teaches all of the limitations of claims 1 and 9, respectively, as described above in detail. Fadel teaches - wherein the action distribution comprises the set of one or more training agent actions and a set of one or more action reward values that corresponds to the set of one or more training agent actions (Fadel ¶ 0004 teaches that “[t]o control the environment, the agent can perform a set of actions [(that is, the action distribution)] that may alter the state of this environment. For each action performed, the agent observes the change in the environment's state and a numerical signal, usually called a reward, that indicates if the action performed moved the agent closer or further to the completion of its goal [(that is, the set of one or more training agent actions and a set of one or more action reward values that corresponds to the set of one or more training agent actions)]”). Regarding claims 7, 14, and 19, the combination of Fadel and Le teaches all of the limitations of claims 1, 9, and 16, respectively, as described above in detail. Fadel teaches - wherein the predictive machine learning model comprises a deep Q network (in relation to Fig. 4, Fadel ¶ 0171 teaches “[a]n embodiment ofTutor4RL was implemented by modifying the Deep Q-Networks (DQN) agent (see Mnih), using the library Keras-RL (see Plappert et al, keras-rl (2016) available at github.com, the entire contents of which is hereby incorporated by reference herein) along with Tensorflow [(that is, the predictive machine learning model comprises a deep Q network)]”). Regarding claims 8, 15, and 20, the combination of Fadel and Le teaches all of the limitations of claims 7, 14, and 19, respectively, as described above in detail. Fadel teaches - wherein the deep Q network comprises an exploration phase (Fadel ¶ 0090 teaches “Embodiments also relate to reinforcement learning exploration [(that is, the deep Q network comprises an exploration phase)]”) and an exploitation phase (Fadel ¶ 0091 teaches “a [weakly supervised reinforcement learning (WSLR)] model is deployed by modifying and existing RL model such that it can incorporate and exploit knowledge functions [(that is, the deep-Q network comprises . . . an exploitation phase)]”). 9 Claims 5, 12, and 18 are rejected under 35 U.S.C. § 103 as being unpatentable over US Published Application 20210150417 to Fadel et al. [hereinafter Fadel] in view of US Published Application 20220215244 to Le [hereinafter Le] and US Published Application 20220058345 to Guo et al. [hereinafter Guo]. Regarding claims 5, 12, and 18, the combination of Fadel and Le teaches all of the limitations of claims 1, 9, and 16, respectively, as described above in detail. Though Fadel and Le teach the features of a tutor guides an agent to make informed decisions during training of a reinforcement learning model in a cross-channel context, the combination of Fadel and Le, however, do not explicitly teach – wherein the historical event sequence data comprises state representation data, and the computer-implemented method further comprises transforming the state representation data into the set of one or more sequence embeddings by: generating tokenized historical event sequence data based on the historical event sequence data; and normalizing the tokenized historical event sequence data into fixed-length vectors. But Guo teaches - wherein the historical event sequence data comprises state representation data, and the computer-implemented method further comprises transforming the state representation data into the set of one or more sequence embeddings by: [(c.1)] generating tokenized historical event sequence data based on the historical event sequence data (Guo, Fig. 2, teaches a model architecture of a reading comprehension-based action prediction model [Examiner annotations in dashed-line text boxes]: PNG media_image5.png 750 943 media_image5.png Greyscale Guo ¶ 0045 teaches “tokenize the observation 234 and the verb phrase 236 into words shown at 202, 204, then embed these words into word vector representation, for example, using embeddings 206, 208 such as pre-trained GloVe embeddings. GloVe (Global Vectors for Word Representation) is an algorithm that generates word embeddings by aggregating global word-word co-occurrence matrix from a corpus”; Guo ¶¶ 0052-53 teaches “Past Observation Retrieval [(that is, historical)] . . . . [O]bservations from different time steps . . . are separated by a special token [(that is, generating tokenized historical event sequence data based on the historical event sequence data)]”); and normalizing the tokenized historical event sequence data into fixed-length vectors (Guo ¶¶ 0044-45 teaches “Observation and Verb Representations . . . . A shared encoder block shown at 210, 212 that includes layer normalization (e.g., Layer-Norm) 214 and a neural network (e.g., Bidirectional gated recurrent unit (GRU)) 216, processes the observation and verb word embeddings to obtain the separate observation and verb representation. . . . Briefly, layer normalization normalizes internal layer features of a neural network, and it stabilizes the neural network training and substantially reduces the training time [(that is, normalizing the tokenized historical event sequence data into fixed-length vectors.)]”; [Examiner notes that a vector inherently has a fixed length, and accordingly, the broadest reasonable interpretation of the term “fixed length vectors” covers the teachings of Fadel (see above), which is not inconsistent with the Applicant’s disclosure. (MPEP § 2111). Further, an inherent “internal layer feature of a neural network” includes a data size, such as a “fixed vector length”]). Fadel, Le, and Guo are from the same or similar field of endeavor. Fadel teaches reinforcement machine learning uses a reinforcement learning system that can be integrated with reinforcement learning (RL) to promote efficient RL agent learning. Le teaches dynamically selecting alternative content based on real-time events during device sessions using a cross-channel, time-bound deep reinforcement machine learning. Guo teaches a neural network that tokenizes inputs using natural language processing Thus, it would have been obvious to a person having ordinary skill in the art as of the effective filing date of the Applicant’s claimed invention to modify the combination of Fadel and Le pertaining to reinforcement machine learning with the cross-channel reinforcement learning with the reinforcement learning input tokenization of Guo. The motivation to do so is to “provide for understanding of natural language and generating accurate responses or actions in an efficient manner. The system and/or method in an embodiment may implement Multi-Passage Reading Comprehension (MPRC) and harness MPRC techniques to solve the huge action space and partial observability challenges.” (Guo ¶ 0023). Response to Arguments 10. Examiner has fully considered Applicant’s arguments, and responds below accordingly. Section 101 11. Under Step 2A Prong One, Applicant submits that “Claim 1 describes a reinforcement learning technique that is achieved by training a predictive machine learning model using historical episode combinations derived from training agent interactions, where the model parameters are modified based on action reward values generated by applying a reward function to the training agent interactions, enabling the predictive software agent to select optimal agent actions across various communication channels, something a human mind is not equipped to do. For example, claim 1 recites, inter alia: providing . . . a historical event sequence data to a state encoder machine learning model to receive a set of one or more sequence embeddings . . . ; providing . . . the set of one or more sequence embeddings to a predictive machine learning model to receive a prediction output comprising an agent action of a set of one or more agent actions that correspond to a predictive software agent, * * * (b) generating a training action reward value for a training agent interaction of the set of one or more training agent interactions by applying a reward function to the training agent interaction, (c) combining a subset of the set of one or more training agent interactions, including the training agent interaction into a historical episode, * * * (e) generating a set of one or more historical episode combinations . . . , and (f) modifying the set of one or more model parameters based on the historical episode combination; and (emphasis added). A human mind is not equipped to (i) provide a historical event sequence data to a state encoder machine learning model, (ii) provide a set of one or more sequence embeddings to a predictive machine learning model, and (iii) perform a reinforcement learning process of a machine learning model comprising generating a training action reward value, combining a subset of a set of one or more training agent interactions, generating a set of one or more historical episode combinations, and modifying a set of one or more model parameters. Accordingly, Applicant respectfully requests withdrawal of the rejection under 35 U.S.C. § 101 at least because the claimed invention is not directed to a judicial exception under prong one of Step 2A.” (Response at pp. 17-18). Examiner’s Response: Examiner submits that under Step 2A Prong One, the rejection hereinabove identifies the judicial exception (that is, abstract idea) by referring to what is recited in the claim and explain why it is considered an exception. For example, as the claim is directed to an abstract idea, the rejection identifies the abstract idea as it is recited in the claim and explains why it is an abstract idea. (MPEP § 2106.07(a)). For example, the limitation of providing a set of one or more sequence embeddings to a predictive machine learning model to receive a prediction output, (see, e.g., claim 1, lines 12-17), which includes limitations that can practically be performed in the human mind (prediction output), including, for example, observations, evaluations, judgments, and opinions, and accordingly, are mental processes, (MPEP § 2106.04(a)(2) sub III). Moreover, further details or specifics of arriving at a prediction output include activities of “determining” a highest score from an action distribution, and also, by way of example, of “modifying” model parameters of the model by “generating” reward values for training agent interactions, as set out above in detail. (see, claim 1, lines 21-23; claim 1, lines 30-31). Accordingly, the claims are directed to an abstract idea. 12. Under Step 2A Prong Two, Applicant submits that “[a]s amended, claim 1 recites a machine learning model training technique that improves machine learning technology and system performance. See Specification, as filed, paragraphs [0060]-[0062]. The elements of claim 1 constitute an improvement in a technical field (e.g., machine learning) and system performance such that the claim, as a whole, integrates any alleged abstract idea into a practical application.” (Response at p. 19). Applicant submits “The second paragraph of MPEP § 2106.05(a), subsection I, has been revised, following Ex Parte Desjardins, to add the following example xiv to the list of examples that show improvement in computer functionality: xiv. Improvements to computer component or system performance based upon adjustments to parameters of a machine learning model associated with tasks or workstreams; Ex Parte Desjardins, Appeal No. 2024-000567 (PTAB September 26, 2025, Appeals Review Panel Decision) (precedential). Desjardins Memorandum, page 4. As explained above, the claims recite an improvement to system performance (e.g., reduce the number of ineffective agent actions that lead to higher accuracy of performing predictive operations by a predictive software agent system, see Specification, as filed, paragraph [0062]) based upon adjustment to parameters (e.g., optimal model parameters are determined for training of a predictive machine learning model, Id. at paragraphs [0087]-[0088]) of a machine learning model (e.g., a predictive machine learning model, Id.) associated with tasks or workstreams (e.g., directing a predictive software agent system to facilitate contact with a client computing entity by using a given one of a plurality of communication channels, such as web, mobile, telephone, IVR, email, SMS, or in-app messaging, Id. at paragraph [0081]).” (Response at pp. 19-20). Further, Applicant submits “Claim 1 recites an improvement to machine learning technology through a reinforcement learning technique. Specifically, claim 1 describes generating a training action reward value for a training agent interaction, combining a subset of training agent interactions into a historical episode, generating a historical episode combination from a set of one or more historic episodes, and modifying the set of one or more model parameters based on the historical episode combination. By doing so, machine learning technology is improved because temporal sequences of past interactions and their outcomes are systematically incorporated into a model parameter optimization process. See Specification, as filed, paragraphs [0060] and [0063]. This model parameter optimization process enhances predictive accuracy and training efficiency of a machine learning model. Id. In particular, a feedback loop created by the model parameter optimization process that maximizes action reward values of agent actions based on historical event sequence data is directed to an improvement in machine learning model technology.” (Response at p. 21 (quoting Specification ¶ 0061 (“fragmented ecosystems”))). Referring to “fragmented ecosystems,” Applicant submits that “Claim 1 solves this problem by reciting, providing . . . the set of one or more sequence embeddings to a predictive machine learning model to receive a prediction output comprising an agent action of a set of one or more agent actions that correspond to a predictive software agent, wherein: (i) the agent action comprises a highest score of a set of one or more scores corresponding to the set of one or more agent actions, [(claim 1, lines 12-19 (emphasis added by Applicant))].” (Response at p. 22 (quoting Specification ¶ 0074)). In sum, Applicant submits “independent claim 1 recites a combination of additional elements that improves computer functionality and a technical field (e.g., machine learning) such that the claim, as a whole, integrates the alleged abstract idea into a practical application. Accordingly, Applicant respectfully requests withdrawal of the rejection under 35 U.S.C. § 101 because independent claim 1 integrates an alleged abstract idea into a practical application under Prong Two of Step 2A.” (Response at p. 23). Examiner Response: Under Step 2A Prong Two, the rejection identifies any additional elements recited in the claim beyond the identified judicial exception (i.e., abstract idea); and evaluate the integration of the judicial exception into a practical application by explaining that the claim as a whole, looking at the additional elements individually and in combination, does not integrate the judicial exception into a practical application using the considerations set forth in MPEP §§ 2106.04(d), 2106.05(a)-(c) and (e)-(h). “Integration” may be based on the improvements in the functioning of a computer or an improvement to any other technology or technical field. (MPEP § 2106.04(d)(1)). The evaluation requires, [i]n sum, that (1) the specification should be evaluated to determine if the disclosure provides sufficient details such that one of ordinary skill in the art would recognize the claimed invention as providing an improvement. Next, (2) if the specification sets forth such an improvement, the claim must be evaluated to ensure that the claim itself reflects the disclosed improvement. By way of example to Desjardins, the MPEP provides under Step 2A Prong Two that “the [Desjardins] specification identified improvements as to how the machine learning model itself operates, including training a machine learning model to learn new tasks while protecting knowledge about previous tasks to overcome the problem of ‘catastrophic forgetting’ encountered in continual learning systems. Importantly, the [appeals review panel (ARP)] evaluated the claims as a whole in discerning at least the limitation ‘adjust the first values of the plurality of parameters to optimize performance of the machine learning model on the second machine learning task while protecting performance of the machine learning model on the first machine learning task’ reflected the improvement disclosed in the specification. Accordingly, the claims as a whole integrated what would otherwise be a judicial exception instead into a practical application at Step 2A Prong Two, and therefore the claims were deemed to be outside any specific, enumerated judicial exception (Step 2A: NO).” (MPEP § 2106.04(d) sub III; see “Advance Notice of Change to the MPEP in light of Ex Parte Desjardins” (05 December 2025) at p. 2)). With regard to “fragmented ecosystems,” the disclosure recites that fragmented ecosystems lead to excessive communications that lead to alert fatigue and disengagement with other future, important communications. Furthermore, these types of communications are created in siloes inside each individual application, leading to a fragmented experience by the client computing entities 102. (Specification ¶ 0061). The disclosure submits that the present invention provides for predicting “optimal agent actions,” that may reduce the number of ineffective agent actions and maximize benefit to recipients (e.g., client computing entities 102) of the agent actions. This technique will lead to higher accuracy of performing predictive operations by the predictive software agent system 101 as needed. In doing so, the techniques described herein improve efficiency and speed of training predictive machine learning models, thus reducing the number of computational operations needed and/or the amount of training data entries needed to train predictive machine learning models. Accordingly, the techniques described herein improve the computational efficiency, storage-wise efficiency, and/or speed of training predictive machine learning models. (Specification ¶ 0062). With regard to “one or more communication channels,” or rather “a communication channel,” the disclosure recites “performing at least one of the one or more optimal agent actions, e.g., using one or more communication channels (e.g., web, mobile, telephone, IVR, email, SMS, in-app messaging) to facilitate contact with, for example, claimant computing entities 102.” (Specification ¶ 0023). Further, the Specification recites that in relation to “one or more communication channels,” “state representation data” may comprise device type, device specification, device location, and communication channel capability (e.g., web, mobile application, phone, email, or SMS). In another example, state representation data may comprise characteristics of a user associated with a client computing entity 102, such as demographics, social determinants, communication channel usage or preference, and medical history of a user. (Specification ¶ 0048 (emphasis added by Examiner)). In sum, the disclosure recites an agent action is performed “at a given time and communication channel in a manner that prevents duplicative and extraneous agent action.” (Specification ¶ 0063). However, the claims do not appear to reflect an improvement as set out by the disclosure under the second leg of MPEP § 2106.05(d)(1). For instance, the claims do not reflect the occurrence of an agent action performed at a given time and communication channel in a manner that prevents duplicative and extraneous agent actions, in kind with the guidance of Desjardins. That is, the claims do not integrate a judicial exception into a practical application of the exception (abstract idea) in a manner that imposes a meaningful limit on the judicial exception, such that the claim is more than a drafting effort designed to monopolize or preempt the judicial exception. (see 2024 Guidance, 89 Fed. Reg. 137 at p. 58136 (17 July 2024)). Moreover, to the extent the disclosure improves the efficiency of the abstract idea, an improved abstract idea remains an abstract idea. Therefore, as set out above in detail, the claims are subject-matter ineligible. 13. Applicant submits that “Step 2B of the Alice/Mayo test focuses on whether the additional limitations present in the claim and their combination is unconventional and provides an inventive concept. MPEP § 2106.05.II. Applicant respectfully submits that the added limitations cannot be considered to be well-understood, routine, or known within the industry at least because they do not appear to be taught by the prior art of record. Accordingly, Applicant respectfully submits that the Office Action improperly rejects independent claim 1 (and the claims depending therefrom) as being directed to patent ineligible subject matter and requests withdrawal of the rejection.” (Response at p. 23). Examiner Response: Examiner submits that for Step 2B, the rejection hereinabove explains why the additional elements, taken individually and in combination, do not result in the claim, as a whole, amounting to significantly more than the identified judicial exception. For example, the additional elements include “one or more processors,” “one or more communication channels,” and a “historical database,” which are recited at a high-level of generality, and accordingly, are generic computer components used to implement the abstract idea, (MPEP § 2106.05(f)), that does not amount to significantly more than the abstract idea. Also, for example, activities of “providing,” “retrieving” and “storing” are well-understood, routine, and conventional activities of storing and retrieving information in memory, (MPEP § 2106.05(d) sub II.iv), that does not amount to significantly more than the abstract idea. In sum, uses components in the ordinary and customary manner for a two-stage machine learning pipeline: a state encoder converts historical interaction sequences into fixed-length embeddings, and a predictive software agent model uses those embeddings to choose the best next action. Accordingly, the claims are subject-matter ineligible, as set out above in detail. Section 103 14. Applicant submits that “Independent claim 1 is amended herein to recite, inter alia, the following elements that are not taught or suggested by any of the references cited in the Office Action: providing . . . a historical event sequence data to a state encoder machine learning model to receive a set of one or more sequence embeddings, wherein the historical event sequence data comprises usage of a communication channel of a set of one or more communication channels; providing . . . the set of one or more sequence embeddings to a predictive machine learning model to receive a prediction output comprising an agent action of a set of one or more agent actions that correspond to a predictive software agent, wherein: * * * (iv) a set of one or more model parameters of the predictive machine learning model are modified by: (a) retrieving training agent interaction data comprising a set of one or more training agent interactions that are associated with a set of one or more training agent actions, wherein the set of one or more training agent interactions corresponds to the set of one or more communication channels, * * * (emphasis added). Claims 9 and 16 are similarly amended. None of the cited references, either alone or in combination teach or suggest such features.” (Response at p. 24). Examiner Response: Examiner agrees the cited reference of Fadel and/or Guo do not teach or disclose the subject matter of the instant claims. The cited reference of Le is relied upon as teaching the features of a real-time events during device sessions using a cross-channel, time-bound deep reinforcement machine learning, as set out above in detail. Conclusion 15. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. 16. The prior art made of record and not relied upon is considered pertinent to Applicant's disclosure: (Narang et al., “Mobile Marketing 2.0: State of the Art and Research Agenda,” Emerald Publishing (2019)) teaches that mobile marketing, the two- or multi-way communication and promotion of an offer between a firm and its customers using a mobile medium, device, platform, or technology, has made rapid strides in the past several years. Mobile marketing has entered its second phase or Mobile Marketing 2.0. The surpassing of desktop by mobile devices in digital media consumption, diffusion of wearable devices among customers, and an overall integration and interconnectedness of devices characterize this phase. Against this backdrop, we present a synthesis of the most recent literature in mobile marketing. We discuss three key advances in mobile marketing research relating to mobile targeting, personalization, and mobile-led cross-channel effects. (US Published Application 20210319478 to Dejardins) teaches optimization and personalization of marketing actions using cloud , hybrid , and quantum - based computing techniques . An example method commences with iteratively selecting , from a pool of prospective clients , at least one subgroup of the prospective clients based on predetermined criteria . The method further includes performing at least one marketing action on the at least one subgroup of the prospective clients . The method then continues with receiving a feedback from a prospective client belonging to the at least one subgroup of the prospective clients in response to the at least one marketing action . The method further includes scoring , by a machine learning technique , the feedback received from the prospective client . The method further includes modifying the at least one marketing action until the at least one marketing action is optimized for the prospective client based on the scoring of the feedback. (US Published Application 20190235936 to Murdock et al.) teaches providing personalized notification management . Notifications can be communicated to a user upon receipt or queued for subsequent handling based on a probability that the user will interact with the notification within a threshold elapsed time from presentation , if it is presented . The probability is determined based on a user's past interactions with similar notifications. (US Published Application 20230401416 to Mistor) teaches determining a priority service context specific (SCS) channel among multiple SCS channels, according to a priority status metric (PSM), and sends an advisory message to the priority channel. The system further determines a priority SCS channel having a PSM higher than at least some of the other SCS channels, generates an advisory message for the priority SCS channel, and sends, the advisory message to the respective system device of the priority SCS channel. 17. Any inquiry concerning this communication or earlier communications from the Examiner should be directed to KEVIN L. SMITH whose telephone number is (571) 272-5964. Normally, the Examiner is available on Monday-Thursday 0730-1730. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, Applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the Examiner by telephone are unsuccessful, the Examiner’s supervisor, KAKALI CHAKI can be reached on 571-272-3719. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /K.L.S./ Examiner, Art Unit 2122 /KAKALI CHAKI/Supervisory Patent Examiner, Art Unit 2122
Read full office action

Prosecution Timeline

Show 2 earlier events
Mar 11, 2026
Interview Requested
Mar 25, 2026
Applicant Interview (Telephonic)
Mar 26, 2026
Examiner Interview Summary
May 08, 2026
Response Filed
Jul 27, 2026
Final Rejection mailed — §101, §103
Jul 31, 2026
Interview Requested
Aug 12, 2026
Examiner Interview Summary
Aug 12, 2026
Applicant Interview (Telephonic)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12731054
REINFORCEMENT LEARNING USING A RELATIONAL NETWORK FOR GENERATING DATA ENCODING RELATIONSHIPS BETWEEN ENTITIES IN AN ENVIRONMENT
3y 6m to grant Granted Sep 08, 2026
Patent 12664451
SYSTEM AND METHOD FOR GENERATING A PREDICTIVE MODEL
6y 3m to grant Granted Jun 23, 2026
Patent 12657425
DYNAMIC CACHE MANAGEMENT IN BEAM SEARCH
5y 3m to grant Granted Jun 16, 2026
Patent 12591815
METHOD AND SYSTEM FOR UPDATING MACHINE LEARNING BASED CLASSIFIERS FOR RECONFIGURABLE SENSORS
4y 10m to grant Granted Mar 31, 2026
Patent 12585917
REINFORCEMENT LEARNING USING ADVANTAGE ESTIMATES
4y 0m to grant Granted Mar 24, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
37%
Grant Probability
57%
With Interview (+20.0%)
4y 7m (~12m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 141 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month