Prosecution Insights
Last updated: October 02, 2026
Application No. 18/626,081

GENERATIVE COUNTERFACTUAL EXPLANATIONS FROM HUMAN PREFERENCES

Non-Final OA §101§103
Filed
Apr 03, 2024
Examiner
GOLAN, MATTHEW BRYCE
Art Unit
Tech Center
Assignee
Dell Products L.P.
OA Round
1 (Non-Final)
0%
Grant Probability
At Risk
1-2
OA Rounds
1y 3m
Est. Remaining
0%
With Interview

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 8 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 9m
Avg Prosecution
25 currently pending
Career history
43
Total Applications
across all art units

Statute-Specific Performance

§101
25.5%
-14.5% vs TC avg
§103
46.6%
+6.6% vs TC avg
§102
6.0%
-34.0% vs TC avg
§112
21.1%
-18.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 8 resolved cases

Office Action

§101 §103
DETAILED ACTION This communication is in response to Application No. 18/626,081 filed on April 3rd, 2024 in which claims 1-20 are presented for examination. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement submitted on 04/03/2024 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement was considered by the examiner. Specification The contents of the specification are sufficient for examination purposes. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to abstract ideas without significantly more. Regarding Claim 1: Step 1: Claim 1 is a process claim. Therefore, claims 1-10 are directed to a statutory category of eligible subject matter. Step 2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the "Mental Processes" grouping of abstract ideas. Here, elements of the claimed subject matter are mental processes. Specifically, the claim recites “A method, comprising . . . recognize instances of time series data” (mental process – amounts to observing information, which may be aided by pen and paper); “generate counterfactual explanations (CEs) for anomalies detected in time series data . . . generate CEs for different types of anomalous time series instances” (mental process – amounts to exercising judgment to form an opinion on counterfactual explanations, based on known or observed detected anomalies in time series data, which may be aided by pen and paper); and “evaluate CEs . . . and to designate respective scores (assigned by human subject matter experts) to the CEs based on the evaluation of the CEs” (mental process – amounts to exercising judgment to form an opinion on respective scores, based on subject matter expertise and known or observed counterfactuals, which may be aided by pen and paper). Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim recites the additional elements: “an MLU that is able to . . . an MLS that is able to . . . a reward large language model (LLM) to . . . generated by the MLS . . . the RLMLS is able to” (amounts to mere instructions to apply the judicial exception on generic and unspecialized computer components, which do not impose any meaningful limits on practicing the abstract idea) and “performing unsupervised training of a multi-modal large language model (MLLM) so as to define an MLU . . . performing supervised training of the MLU so as to define an MLS . . . training a reward large language model (LLM) . . . creating a reinforcement learning MLS (RLMLS) model from the MLS, and performing a fine-tuning process using the RLMLS model and the MLS so that, after fine-tuning” (amounts to merely generally linking the use of the judicial exception to a particular technological environment or field of use, which do not impose any meaningful limits on practicing the abstract idea). Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. The claim recites the additional elements: “an MLU that is able to . . . an MLS that is able to . . . a reward large language model (LLM) to . . . generated by the MLS . . . the RLMLS is able to” (mere instructions to apply the exception using generic computer components does not provide an inventive concept) and “performing unsupervised training of a multi-modal large language model (MLLM) so as to define an MLU . . . performing supervised training of the MLU so as to define an MLS . . . training a reward large language model (LLM) . . . creating a reinforcement learning MLS (RLMLS) model from the MLS, and performing a fine-tuning process using the RLMLS model and the MLS so that, after fine-tuning” (merely generally linking the use of the judicial exception to a particular technological environment or field of use does not provide an inventive concept). For the reasons above, Claim 1 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 2-10. The additional limitations of the dependent claims are addressed below. Regarding Claim 2: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 2 depends on. Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim recites the additional elements: “wherein, prior to the unsupervised training, the MLLM was trained with multi-modal data” (amounts to merely generally linking the use of the judicial exception to a particular technological environment or field of use, which do not impose any meaningful limits on practicing the abstract idea). Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. The claim recites the additional elements: “wherein, prior to the unsupervised training, the MLLM was trained with multi-modal data” (merely generally linking the use of the judicial exception to a particular technological environment or field of use does not provide an inventive concept). Accordingly, Claim 2 is rejected as being directed to an abstract idea without significantly more. Regarding Claim 3: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 3 depends on. Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim recites the additional elements: “wherein the unsupervised training is performed using multi-modal time-series data comprising text and images” (amounts to merely generally linking the use of the judicial exception to a particular technological environment or field of use, which do not impose any meaningful limits on practicing the abstract idea). Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. The claim recites the additional elements: “wherein the unsupervised training is performed using multi-modal time-series data comprising text and images” (merely generally linking the use of the judicial exception to a particular technological environment or field of use does not provide an inventive concept). Accordingly, Claim 3 is rejected as being directed to an abstract idea without significantly more. Regarding Claim 4: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 4 depends on. Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim recites the additional elements: “wherein the supervised training is performed using a dataset that comprises multiple elements, each of which has a form {anomaly instance, description in counterfactual form}” (amounts to merely generally linking the use of the judicial exception to a particular technological environment or field of use, which do not impose any meaningful limits on practicing the abstract idea). Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. The claim recites the additional elements: “wherein the supervised training is performed using a dataset that comprises multiple elements, each of which has a form {anomaly instance, description in counterfactual form}” (merely generally linking the use of the judicial exception to a particular technological environment or field of use does not provide an inventive concept). Accordingly, Claim 4 is rejected as being directed to an abstract idea without significantly more. Regarding Claim 5: Step 2A Prong 1: See the rejection of Claim 4 above, which Claim 5 depends on. Here, the claim recites additional limitations that are mental processes. Specifically, the claim recites “wherein the description in counterfactual form is generated by a human” (mental process – amounts to exercising judgment to form an opinion on a description, based on a known or observed counterfactual, which may be aided by pen and paper). Step 2A Prong 2 & Step 2B: There are no elements left for consideration of implementation within a practical application or for consideration of significantly more. Accordingly, Claim 5 is rejected as being directed to an abstract idea without significantly more. Regarding Claim 6: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 6 depends on. Here, the claim recites additional limitations that are mental processes. Specifically, the claim recites “wherein the scores, together with one or more formulas, enable a human subject matter expert (SME) to rank the CEs” (mental process – amounts to exercising judgment to form an opinion on a ranking, based on known or observed counterfactuals, scores, and formulas, which may be aided by pen and paper). Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim recites the additional elements: “generated by the MLS” (amounts to mere instructions to apply the judicial exception on generic and unspecialized computer components, which do not impose any meaningful limits on practicing the abstract idea). Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. The claim recites the additional elements: “generated by the MLS” (mere instructions to apply the exception using generic computer components does not provide an inventive concept). Accordingly, Claim 6 is rejected as being directed to an abstract idea without significantly more. Regarding Claim 7: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 7 depends on. Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim recites the additional elements: “wherein the reward LLM has fewer parameters than the MLS” (amounts to merely generally linking the use of the judicial exception to a particular technological environment or field of use, which do not impose any meaningful limits on practicing the abstract idea). Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. The claim recites the additional elements: “wherein the reward LLM has fewer parameters than the MLS” (merely generally linking the use of the judicial exception to a particular technological environment or field of use does not provide an inventive concept). Accordingly, Claim 7 is rejected as being directed to an abstract idea without significantly more. Regarding Claim 8: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 8 depends on. Here, the claim recites additional limitations that are mental processes. Specifically, the claim recites “performed using numerical rankings of the CEs” (mental process – amounts to exercising judgment to evaluate processes, with reference to calculated rankings, which may be aided by pen and paper). Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim recites the additional elements: “wherein training the reward LLM is” (amounts to merely generally linking the use of the judicial exception to a particular technological environment or field of use, which do not impose any meaningful limits on practicing the abstract idea) and “that were generated by the MLS” (amounts to mere instructions to apply the judicial exception on generic and unspecialized computer components, which do not impose any meaningful limits on practicing the abstract idea). Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. The claim recites the additional elements: “wherein training the reward LLM is” (merely generally linking the use of the judicial exception to a particular technological environment or field of use does not provide an inventive concept) and “that were generated by the MLS” (mere instructions to apply the exception using generic computer components does not provide an inventive concept). Accordingly, Claim 8 is rejected as being directed to an abstract idea without significantly more. Regarding Claim 9: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 9 depends on. Here, the claim recites additional limitations that are mental processes. Specifically, the claim recites “comparing respective CE outputs . . . to identify divergences between the CE outputs . . . and the CE outputs” (mental process – amounts to exercising judgment to evaluate counterfactuals to form opinions on divergencies, which may be aided by pen and paper); “merging the divergences with scores of the CE outputs to form a merged output” (mental process – amounts to exercising judgment to form an opinion on merged values, which may be aided by pen and paper); and “providing the merged output to a proximal policy optimization (PPO) process; and with the PPO process, using the merged output” (mental process – amounts to exercising judgment to perform an optimization process, with reference to known or observed merged outputs, which may be aided by pen and paper). Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim recites the additional elements: “wherein the fine-tuning comprises: inputting a common group of time series anomalies to both the MLS and the RLMLS . . . to fine tune the RLMLS” (amounts to merely generally linking the use of the judicial exception to a particular technological environment or field of use, which do not impose any meaningful limits on practicing the abstract idea) and “of the MLS and the RLMLS . . . of the MLS . . . of the RLMLS” (amounts to mere instructions to apply the judicial exception on generic and unspecialized computer components, which do not impose any meaningful limits on practicing the abstract idea). Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. The claim recites the additional elements: “wherein the fine-tuning comprises: inputting a common group of time series anomalies to both the MLS and the RLMLS . . . to fine tune the RLMLS” (merely generally linking the use of the judicial exception to a particular technological environment or field of use does not provide an inventive concept) and “of the MLS and the RLMLS . . . of the MLS . . . of the RLMLS” (mere instructions to apply the exception using generic computer components does not provide an inventive concept). Accordingly, Claim 9 is rejected as being directed to an abstract idea without significantly more. Regarding Claim 10: Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 10 depends on. Here, the claim recites additional limitations that are mental processes. Specifically, the claim recites “generating . . . a respective CE for one or more of the anomalies in the set of time-series data, and the CEs are comprehensible by a human” (mental process – amounts to exercising judgment to form an opinion on counterfactual explanations, such that the opinions are comprehensible to humans, based on known or observed detected anomalies in time series data, which may be aided by pen and paper). Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim recites the additional elements: “receiving . . . a set of time-series data” (receiving data amounts to extra-solution activity because transmission of data is incidental to the claimed subject matter); “that comprises one or more anomalies” (amounts to merely generally linking the use of the judicial exception to a particular technological environment or field of use, which do not impose any meaningful limits on practicing the abstract idea); and “by the RLMLS . . . by the RLMLS” (amounts to mere instructions to apply the judicial exception on generic and unspecialized computer components, which do not impose any meaningful limits on practicing the abstract idea). Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. The claim recites the additional elements: “receiving . . . a set of time-series data” (transmitting data is well‐understood, routine, and conventional, see Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362; see also buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014); therefore the limitation, which is recited with a high level of generality, remains insignificant extra-solution activity even upon reconsideration); “that comprises one or more anomalies” (merely generally linking the use of the judicial exception to a particular technological environment or field of use does not provide an inventive concept); and “by the RLMLS . . . by the RLMLS” (mere instructions to apply the exception using generic computer components does not provide an inventive concept). Accordingly, Claim 10 is rejected as being directed to an abstract idea without significantly more. Regarding Claim 11: Step 1: Claim 11 is a machine claim. Therefore, claims 11-20 are directed to a statutory category of eligible subject matter. Step 2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the "Mental Processes" grouping of abstract ideas. Here, the claim recites limitations that are substantially the same as the limitations of Claim 1. As a result, and as elaborated above, these limitations are abstract ideas because they are mental processes. Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim recites the additional elements: “A non-transitory storage medium having stored therein instructions that are executable by one or more hardware processors to perform operations comprising . . . an MLU that is able to . . . an MLS that is able to . . . a reward large language model (LLM) to . . . generated by the MLS . . . the RLMLS is able to” (amounts to mere instructions to apply the judicial exception on generic and unspecialized computer components, which do not impose any meaningful limits on practicing the abstract idea) and “performing unsupervised training of a multi-modal large language model (MLLM) so as to define an MLU . . . performing supervised training of the MLU so as to define an MLS . . . training a reward large language model (LLM) . . . creating a reinforcement learning MLS (RLMLS) model from the MLS, and performing a fine-tuning process using the RLMLS model and the MLS so that, after fine-tuning” (amounts to merely generally linking the use of the judicial exception to a particular technological environment or field of use, which do not impose any meaningful limits on practicing the abstract idea). Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. The claim recites the additional elements: “A non-transitory storage medium having stored therein instructions that are executable by one or more hardware processors to perform operations comprising . . . an MLU that is able to . . . an MLS that is able to . . . a reward large language model (LLM) to . . . generated by the MLS . . . the RLMLS is able to” (mere instructions to apply the exception using generic computer components does not provide an inventive concept) and “performing unsupervised training of a multi-modal large language model (MLLM) so as to define an MLU . . . performing supervised training of the MLU so as to define an MLS . . . training a reward large language model (LLM) . . . creating a reinforcement learning MLS (RLMLS) model from the MLS, and performing a fine-tuning process using the RLMLS model and the MLS so that, after fine-tuning” (merely generally linking the use of the judicial exception to a particular technological environment or field of use does not provide an inventive concept). For the reasons above, Claim 11 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 12-20. The additional limitations of the dependent claims are addressed below. Regarding Claim 12, the claim recites limitations that are all substantially the same as limitations of Claim 2, in the form of a non-transitory computer-readable medium. The claim is also directed to performing mental processes without integration into a practical component or significantly more. Accordingly, Claim 12 is rejected under the same rationale. Regarding Claim 13, the claim recites limitations that are all substantially the same as limitations of Claim 3, in the form of a non-transitory computer-readable medium. The claim is also directed to performing mental processes without integration into a practical component or significantly more. Accordingly, Claim 13 is rejected under the same rationale. Regarding Claim 14, the claim recites limitations that are all substantially the same as limitations of Claim 4, in the form of a non-transitory computer-readable medium. The claim is also directed to performing mental processes without integration into a practical component or significantly more. Accordingly, Claim 14 is rejected under the same rationale. Regarding Claim 15, the claim recites limitations that are all substantially the same as limitations of Claim 5, in the form of a non-transitory computer-readable medium. The claim is also directed to performing mental processes without integration into a practical component or significantly more. Accordingly, Claim 15 is rejected under the same rationale. Regarding Claim 16, the claim recites limitations that are all substantially the same as limitations of Claim 6, in the form of a non-transitory computer-readable medium. The claim is also directed to performing mental processes without integration into a practical component or significantly more. Accordingly, Claim 16 is rejected under the same rationale Regarding Claim 17, the claim recites limitations that are all substantially the same as limitations of Claim 7, in the form of a non-transitory computer-readable medium. The claim is also directed to performing mental processes without integration into a practical component or significantly more. Accordingly, Claim 17 is rejected under the same rationale. Regarding Claim 18, the claim recites limitations that are all substantially the same as limitations of Claim 8, in the form of a non-transitory computer-readable medium. The claim is also directed to performing mental processes without integration into a practical component or significantly more. Accordingly, Claim 18 is rejected under the same rationale. Regarding Claim 19, the claim recites limitations that are all substantially the same as limitations of Claim 9, in the form of a non-transitory computer-readable medium. The claim is also directed to performing mental processes without integration into a practical component or significantly more. Accordingly, Claim 19 is rejected under the same rationale. Regarding Claim 20, the claim recites limitations that are all substantially the same as limitations of Claim 10, in the form of a non-transitory computer-readable medium. The claim is also directed to performing mental processes without integration into a practical component or significantly more. Accordingly, Claim 20 is rejected under the same rationale. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 3-11, and 13-20 are rejected under 35 U.S.C. 103 as being unpatentable over Ouyang et al. (hereinafter Ouyang) (“Training language models to follow instructions with human feedback”) in view of Sulem et al. (hereinafter Sulem) (“Diverse Counterfactual Explanations for Anomaly Detection in Time Series”) and Wu et al. (hereinafter Wu) (“Multimodal Large Language Models: A Survey”). Regarding Claim 1, Ouyang teaches a method, comprising (Pg. 3, Fig. 2, “A diagram illustrating the three steps of our method: . . . See Section 3 for more details on our method”; see also Pg. 1, Abstract, “In this paper, we show an avenue for aligning language models with user intent on a wide range of tasks by fine-tuning with human feedback. Starting with a set of labeler-written prompts and prompts submitted through the OpenAI API, we collect a dataset of labeler demonstrations of the desired model behavior, which we use to fine-tune GPT-3 using supervised learning. We then collect a dataset of rankings of model outputs, which we use to further fine-tune this supervised model using reinforcement learning from human feedback. We call the resulting models InstructGPT”): performing unsupervised training of a . . . large language model (MLLM) so as to define an MLU that is able to recognize instances of . . . [input] data (Pg. 6, Para. 1, “We start with a pretrained language model (Radford et al., 2019; Brown et al., 2020; Fedus et al., 2021; Rae et al., 2021; Thoppilan et al., 2022)” and Ouyang, Pg. 24, Cit. 9, “Radford . . . Language models are unsupervised multitask learners”, where the “pretrained language model” is a “Language model” generated by performing “unsupervised” training so as to define an MLU that is able to recognize instances of data, “a pretrained language model”; see also Pg. 8, Para. 5, “We start with the GPT-3 pretrained language models from Brown et al. (2020)”, where a person of ordinary skill in the art would understand that “the GPT-3 pretrained language models from Brown et al. (2020)” are an LLM generated using unsupervised training performed using data); performing supervised training of the MLU so as to define an MLS that is able to generate . . . [outputs] (Pg. 6, Para. 2, “Step 1: Collect demonstration data, and train a supervised policy. Our labelers provide demonstrations of the desired behavior on the input prompt distribution (see Section 3.2 for details on this distribution). We then fine-tune a pretrained GPT-3 model on this data using supervised learning” and Pg. 8, Para. 6, “Supervised fine-tuning (SFT). We fine-tune GPT-3 on our labeler demonstrations using supervised learning . . . we find that our SFT models overfit on validation loss after 1 epoch”, where supervised training of the MLU, “We then fine-tune a pretrained GPT-3 model on this data using supervised learning”, defines a MLS, “SFT models”, that is able to generate outputs in response to an “input prompt”; see also Pg. 3, Fig. 2, “A diagram illustrating the three steps of our method: (1) supervised fine-tuning (SFT)”); training a reward large language model (LLM) to evaluate . . . [the outputs] generated by the MLS (Pg. 6, Para. 3, “Step 2: Collect comparison data, and train a reward model. We collect a dataset of comparisons between model outputs, where labelers indicate which output they prefer for a given input. We then train a reward model to predict the human-preferred output”; Pg. 8, Para. 7, “Reward modeling (RM). Starting from the SFT model with the final unembedding layer removed, we trained a model to take in a prompt and response, and output a scalar reward”, and Pg. 3, Fig. 2, “A diagram illustrating the three steps of our method: . . . (2) reward model (RM) training”, where a reward large language model (LLM), is trained, “train a reward model”, see Pg. 3, Para. 1, “all of our models use the GPT-3 architecture”, to evaluate the outputs generated by the MLS, “We collect a dataset of comparisons between model outputs . . . then train a reward model to predict the human-preferred output”), and to designate respective scores (assigned by human subject matter experts) to the . . . [outputs] based on the evaluation of the . . . [outputs] (Pg. 7, Para. 4, “Our training tasks are from two sources: (1) a dataset of prompts written by our labelers and (2) a dataset of prompts submitted to early InstructGPT models on our API” and Pg. 7, Para. 6, “Our aim was to select a group of labelers who were sensitive to the preferences of different demographic groups, and who were good at identifying outputs that were potentially harmful. Thus, we conducted a screening test designed to measure labeler performance on these axes. We selected labelers who performed well on this test”, where the “labelers” are subject matter experts because they are the creators of the subject matter, “a dataset of prompts written by our labelers”, and display sufficient expertise to be “sensitive to the preferences of different demographic groups, and who were good at identifying outputs that were potentially harmful”, as demonstrated by “a screening test”; see also Pg. 3, Fig. 2, “A diagram illustrating the three steps of our method: . . . (2) reward model (RM) training . . . In Step 2, boxes A-D are samples from our models that get ranked by labelers”, and Pg. 39, Fig. 12, “Screenshots of our labeling interface. (a) For each output, labelers give a Likert score for overall quality on a 1-7 scale”, where the “labelers give . . . score[s] . . . on a 1-7 scale”, which the reward model is trained to designated to the outputs, “reward model (RM) training . . . In Step 2, boxes A-D are samples from our models that get ranked by labelers”, see Pg. 2, Para. 3, “This technique uses human preferences as a reward signal to fine-tune our models . . . we collect a dataset of human-labeled comparisons between outputs from our models on a larger set of API prompts. We then train a reward model (RM) on this dataset to predict which model output our labelers would prefer”, where the “reward model” designates “human” scores to outputs to “predict which model output our labelers would prefer”, and, as depicted in Fig. 2, the scores, together with one or more formulas, “D>C>A=B”, enable the “labeler” to “rank” the outputs generated by the MLS, “fine-tun[ed] GPT-3 with supervised learning” “model outputs are sampled”); and creating a reinforcement learning MLS (RLMLS) model from the MLS, and performing a fine-tuning process using the RLMLS model and the MLS so that, after fine-tuning, the RLMLS is able to generate . . . [outputs] for different types of . . . [input] instances (Pg. 6, Para. 4, “Step 3: Optimize a policy against the reward model using PPO. We use the output of the RM as a scalar reward. We fine-tune the supervised policy to optimize this reward using the PPO algorithm” and Pg. 9, Para. 2, “Reinforcement learning (RL). Once again following Stiennon et al. (2020), we fine-tuned the SFT model on our environment using PPO (Schulman et al., 2017). The environment is a bandit environment which presents a random customer prompt and expects a response to the prompt. Given the prompt and response, it produces a reward determined by the reward model and ends the episode”; and Pg. 3, Fig. 2, “A diagram illustrating the three steps of our method: . . . (3) reinforcement learning via proximal policy optimization (PPO) on this reward model”, where a reinforcement learning MLS (RLMLS) model is created from the MLS on an “episode[ic]” basis, such that fine-tuning is performed using the RLMLS model and the MLS, “Optimize a policy against the reward model using PPO” where “we fine-tuned the SFT model on our environment using PPO”, so that, after fine-tuning, the RLMLS is able to generate outputs for different types of input instances, “We fine-tune the supervised policy to optimize this reward using the PPO algorithm”; see also Pg. 14, Para. 5, “We can minimize performance regressions on public NLP datasets by modifying our RLHF fine-tuning procedure”; Pg. 7, Para. 4, “These prompts are very diverse and include generation, question answering, dialog, summarization, extractions, and other natural language tasks”; and Pg. 12, Para. 2, “all of our InstructGPT models still greatly outperform the GPT-3 baselines. Thus, our InstructGPT models aren’t simply overfitting to the preferences of our training labelers”; see generally Pg. 42, Section “C.4 Details of RLHF training”). Ouyang does not explicitly disclose . . . multi-modal . . . time series . . . (where the pretrained model is not specifically described as multi-modal and the input data is not specifically described as time-series data) . . . counterfactual explanations (CEs) for anomalies detected in time series data. . . (where the outputs of the MLS and RLMLS are not specifically described as counterfactual explanations (CEs) for anomalies detected in time series input data; redundant recitations of CEs and anomalies in time series data omitted). However, Sulem teaches . . . [inputting time series data into a model to generate] counterfactual explanations (CEs) for anomalies detected in time series data . . . (Pg. 4, Para. 3-5, “We assume that we are given an anomaly detection model which we can use to predict – or rather in this context, detect – anomalies on a time series of any given length . . . Examples of anomaly detection models . . . where . . . the context is implicit and the whole training set is considered as normal data and thus the context of anomalies detected in a test time series”; Pg. 2, Para. 5, “We introduce a model-agnostic method that generates counterfactual ensemble explanations for anomalies detected by an algorithm in a univariate or multivariate time series”; and Pg. 5, Fig. 1, “The original input (1a) is a univariate time series window containing an anomalous subsequence (highlighted in red)”). Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the MLU that is able to recognize instances of input data; performing supervised training of the MLU so as to define an MLS that is able to generate outputs; training a reward large language model (LLM) to evaluate the outputs generated by the MLS, and to designate respective scores (assigned by human subject matter experts) to the outputs based on the evaluation of the outputs; and creating a reinforcement learning MLS (RLMLS) model from the MLS, and performing a fine-tuning process using the RLMLS model and the MLS so that, after fine-tuning, the RLMLS is able to generate outputs for different types of input instances of Ouyang with the inputting time series data into a model to generate counterfactual explanations (CEs) for anomalies detected in time series data of Sulem in order to utilize the machine learning performance improvements and truthfulness benefits of the method disclosed by Ouyang (Ouyang, Pg. 1, Abstract, “In human evaluations on our prompt distribution, outputs from the 1.3B parameter InstructGPT model are preferred to outputs from the 175B GPT-3, despite having 100x fewer parameters. Moreover, InstructGPT models show improvements in truthfulness and reductions in toxic output generation while having minimal performance regressions on public NLP datasets”) for a model-agnostic method in an area with unmet demand for machine learning solutions (Sulem, Pg. 1, Abstract, “Data-driven methods that detect anomalies in times series data are ubiquitous in practice, but they are in general unable to provide helpful explanations for the predictions they make. In this work we propose a model-agnostic algorithm that generates counterfactual ensemble explanations for time series anomaly detection models”) and to generate useful explanations for model-detected anomalies in time series, which can improve time series interpretability and reveal actionable insights (Sulem, Pg. 2, Para. 4, “Counterfactual explanations are particularly useful for explaining anomalies in time series detected by a given model, in which case a counterfactual is another time series (or subsequence) that does not contain observations detected as anomalous and is similar to the anomalous subsequence. It thus corresponds to the closest normal or expected behaviour according to the model. For example, if the anomalous time series is a temporal record of the blood glucose level of a patient at risk, a counterfactual could be an alternative series of values contained in a non-critical interval. Hence, counterfactual explanations can reveal the boundaries of the normal time series distribution according to the model and, consequently, its sensitivity to possible issues”; see also Sulem, Pg. 12, Para. 3, “This work proposed a novel method for generating explanations for time series anomalies and detection models. Our real-world experiments show that the counterfactual framework, augmented with an ensemble approach, improves the interpretability of time series anomaly detection models, and can help their users identify the possible actions to take in consequence”). Additionally, Wu teaches . . . [performing unsupervised training of a] multi-modal [large language model (MLLM) so as to define an MLU that is able to recognize instances of] . . . time series [data] . . . (Pg. 2248, Col. 1, Fig. 1, “The definition of multimodal”, where the “Multimodal” “Training” uses “Input” data comprising “Text” and images, “Visual”; see also Pg. 2247, Col. 1, Para. 2, “A multimodal model combines multiple data types, including images, text, audio, and more” and Pg. 2248, Col. 1, Para. 3, “From a data perspective, multimodal data can be seen as a combination of different data types, such as images, numerical data, text, symbols, audio, time series”, where “time series” data comprising “text” and “images” is suggested for use in unsupervised training, see Pg. 2250-2251, “Generative pre-training is an important method and training objective in self-supervised learning, where the model learns how to generate data without relying on labels or manual . . . enabling large-scale self-supervised pre-training” and Pg. 2247, Col. 1, Para. 2, “To address this issue, multimodal LLMs integrate multiple data types, overcoming the limitations of pure text models and opening up possibilities for handling diverse data types. GPT-4 [6] serves as an excellent example of a multimodal LLM”, where a person of ordinary skill in the art would understand “GPT-4” to be trained using unsupervised learning; see generally Pg. 3, Col. 1, Para. 2, “Large-scale multimodal (2020-?) . . . CLIP liberates the burden of assembling massive datasets with predetermined class counts. Instead, CLIP empowers the collection of image-text pairs and leverages unsupervised techniques to either predict their similarity or generate them”). Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the performing unsupervised training of a large language model (MLLM) so as to define an MLU that is able to recognize instances of time series data of Ouyang in view of Sulem with the performing unsupervised training of a multi-modal large language model (MLLM) so as to define an MLU that is able to recognize instances of time series data of Wu in order to allow the models to handle diverse text and image data types (Wu, Pg. 2247, Col. 1, Para. 2, “Pure text LLMs, such as GPT-3 [3], BERT [4], and RoBERTa [5], excel in tasks like text generation and encoding, but they lack a comprehensive understanding and processing of other data types. To address this issue, multimodal LLMs integrate multiple data types, overcoming the limitations of pure text models and opening up possibilities for handling diverse data types”), which will increase the model’s performance and utility in high-value domains (Wu, Pg. 2247, Col. 1-2, Para. 1-1, “the application of multimodal inputs greatly expands the potential of language models in high-value domains, such as multimodal robotics, document intelligence, and robot technology . . . multimodal LLMs have shown superior performance in common-sense reasoning compared to single-modality models”). Regarding Claim 3, Ouyang in view of Sulem and Wu teach the method as recited in claim 1, wherein the unsupervised training is performed using multi-modal time-series data comprising text and images (Ouyang, Pg. 6, Para. 1, “We start with a pretrained language model (Radford et al., 2019; Brown et al., 2020; Fedus et al., 2021; Rae et al., 2021; Thoppilan et al., 2022)” and Ouyang, Pg. 24, Cit. 9, “Radford . . . Language models are unsupervised multitask learners”, where the “pretrained language model” is a “are unsupervised multitask learners”; see also Ouyang, Pg. 8, Para. 5, “We start with the GPT-3 pretrained language models from Brown et al. (2020)”, where a person of ordinary skill in the art would understand that “the GPT-3 pretrained language models from Brown et al. (2020)” are generated using unsupervised training performed using data, which, in view of Sulem and Wu, is time-series data comprising text and images, see Sulem, Pg. 4, Para. 3-5, “We assume that we are given an anomaly detection model which we can use to predict – or rather in this context, detect – anomalies on a time series of any given length . . . Examples of anomaly detection models . . . where . . . the context is implicit and the whole training set is considered as normal data and thus the context of anomalies detected in a test time series” and Wu, Pg. 2248, Col. 1, Fig. 1, “The definition of multimodal”, where the “Multimodal” “Training” uses “Input” data comprising “Text” and images, “Visual”; see also Wu, Pg. 2247, Col. 1, Para. 2, “A multimodal model combines multiple data types, including images, text, audio, and more” and Wu, Pg. 2248, Col. 1, Para. 3, “From a data perspective, multimodal data can be seen as a combination of different data types, such as images, numerical data, text, symbols, audio, time series”, where “time series” data comprising “text” and “images” is suggested for use in unsupervised training, see Wu, Pg. 2250-2251, “Generative pre-training is an important method and training objective in self-supervised learning, where the model learns how to generate data without relying on labels or manual . . . enabling large-scale self-supervised pre-training” and Wu, Pg. 2247, Col. 1, Para. 2, “To address this issue, multimodal LLMs integrate multiple data types, overcoming the limitations of pure text models and opening up possibilities for handling diverse data types. GPT-4 [6] serves as an excellent example of a multimodal LLM”, where a person of ordinary skill in the art would understand “GPT-4” to be trained using unsupervised learning; see generally Wu, Pg. 3, Col. 1, Para. 2, “Large-scale multimodal (2020-?) . . . CLIP liberates the burden of assembling massive datasets with predetermined class counts. Instead, CLIP empowers the collection of image-text pairs and leverages unsupervised techniques to either predict their similarity or generate them”). The reasons for obviousness were discussed in regard to the rejection of claim 1 above and remain applicable here. Regarding Claim 4, Ouyang in view of Sulem and Wu teach the method as recited in claim 1, wherein the supervised training is performed using a dataset that comprises multiple elements, each of which has a form {anomaly instance, description in counterfactual form} (Ouyang, Pg. 39, Fig. 12, “Screenshots of our labeling interface. (a) For each output, labelers give a Likert score for overall quality on a 1-7 scale, and also provide various metadata labels”; Ouyang, Pg. 6, Para. 2, “Collect demonstration data, and train a supervised policy. Our labelers provide demonstrations of the desired behavior on the input prompt distribution (see Section 3.2 for details on this distribution). We then fine-tune a pretrained GPT-3 model on this data using supervised learning”; and Ouyang, Pg. 3, Fig. 2, wherein the “supervised learning” is performed using a dataset that comprises multiple elements, see Wu, Pg. 2248, Col. 1, Para. 3, “From a data perspective, multimodal data can be seen as a combination of different data types, such as images, numerical data, text, symbols, audio, time series”, each of which, when incorporated into a supervised training dataset would require an association that is within the broadest reasonable interpretation of {data instance, description label}, see generally, Ouyang, Pg. 48, Fig. 23, “In the few-shot setting, there are 15 additional text / label pairs”, which, in view of Sulem, is {anomaly instance, description in counterfactual form}, see Sulem, Pg. 2, Para. 5, “We introduce a model-agnostic method that generates counterfactual ensemble explanations for anomalies detected by an algorithm in a univariate or multivariate time series” and Sulem, Pg. 5, Fig. 1, “The original input (1a) is a univariate time series window containing an anomalous subsequence (highlighted in red)”). The reasons for obviousness were discussed in regard to the rejection of claim 1 above and remain applicable here. Regarding Claim 5, Ouyang in view of Sulem and Wu teach the method as recited in claim 4, wherein the description in counterfactual form is generated by a human (Ouyang, Pg. 7, Para. 4, “Our training tasks are from two sources: (1) a dataset of prompts written by our labelers and (2) a dataset of prompts submitted to early InstructGPT models on our API” and Ouyang, Pg. 3, Fig. 2, “A diagram illustrating the three steps of our method: (1) supervised fine-tuning (SFT)”, where, as depicted in Fig. 2, humans, “labelers”, generate “the desired output behavior” in description form, such as “Some people went to the moon . . .”, which “is used to fine-tune GPT-3 with supervised learning”, and where, in view of Sulem, the generating is a description in counterfactual form, see Sulem, Pg. 1, Abstract, “we propose a model-agnostic algorithm that generates counterfactual ensemble explanations for time series anomaly detection models”; see also Sulem, Pg. 10, Para. 4, “On Figure 2, we present a visualization of our method on anomalies from the KPI dataset, detected by NCAD and USAD. We observe that the counterfactual ensemble explanation from DPE (in red color scale), ICE (in green), and FS (in purple) are quite dissimilar” and Sulem, Pg. 11, Fig. 2, “Figure 2: Time series windows contains an anomaly and our counterfactual ensemble explanations”). The reasons for obviousness were discussed in regard to the rejection of claim 1 above and remain applicable here. Regarding Claim 6, Ouyang in view of Sulem and Wu teach the method as recited in claim 1, wherein the scores, together with one or more formulas, enable a human subject matter expert (SME) to rank the CEs generated by the MLS (Ouyang, Pg. 7, Para. 4, “Our training tasks are from two sources: (1) a dataset of prompts written by our labelers and (2) a dataset of prompts submitted to early InstructGPT models on our API” and Ouyang, Pg. 7, Para. 6, “Our aim was to select a group of labelers who were sensitive to the preferences of different demographic groups, and who were good at identifying outputs that were potentially harmful. Thus, we conducted a screening test designed to measure labeler performance on these axes. We selected labelers who performed well on this test”, where, as discussed above, the “labelers” are subject matter experts because they are the creators of the subject matter, “a dataset of prompts written by our labelers”, and display sufficient expertise to be “sensitive to the preferences of different demographic groups, and who were good at identifying outputs that were potentially harmful”, as demonstrated by “a screening test”; see also Ouyang, Pg. 3, Fig. 2, “A diagram illustrating the three steps of our method: . . . (2) reward model (RM) training . . . In Step 2, boxes A-D are samples from our models that get ranked by labelers”, and Ouyang, Pg. 39, Fig. 12, “Screenshots of our labeling interface. (a) For each output, labelers give a Likert score for overall quality on a 1-7 scale”, where the “labelers give . . . score[s] . . . on a 1-7 scale” and, as depicted in Fig. 2, the scores, together with one or more formulas, “D>C>A=B”, enable the “labeler” to “rank” the outputs generated by the MLS, “fine-tun[ed] GPT-3 with supervised learning” “model outputs are sampled”, which, in view of Sulem, the outputs are CEs, see Sulem, Pg. 1, Abstract, “we propose a model-agnostic algorithm that generates counterfactual ensemble explanations for time series anomaly detection models”). The reasons for obviousness were discussed in regard to the rejection of claim 1 above and remain applicable here. Regarding Claim 7, Ouyang in view of Sulem and Wu teach the method as recited in claim 1, wherein the reward LLM has fewer parameters than the MLS (Ouyang, Pg. 2-3, Para. 4-1, “We train three model sizes (1.3B, 6B, and 175B parameters), and all of our models use the GPT-3 architecture”; but see Ouyang, Pg. 8, Para. 7, “Reward modeling (RM) . . . In this paper we only use 6B RMs, as this saves a lot of compute, and we found that 175B RM training could be unstable and thus was less suitable to be used as the value function during RL”, where the reward LLM, “Reward modeling (RM)”, has fewer parameters, “we only use 6B RMs”, than the “175B SFT”, which is the MLS, “Supervised fine-tuning (SFT)”, see Ouyang, Pg. 17, Para. 4, “training our 175B SFT model requires 4.9 petaflops/s-days” and Ouyang, Pg. 8, Para. 6, “Supervised fine-tuning (SFT). We fine-tune GPT-3 on our labeler demonstrations using supervised learning”). Regarding Claim 8, Ouyang in view of Sulem and Wu teach the method as recited in claim 1, wherein training the reward LLM is performed using numerical rankings of the CEs that were generated by the MLS (Ouyang, Pg. 6, Para. 3, “Step 2: Collect comparison data, and train a reward model. We collect a dataset of comparisons between model outputs, where labelers indicate which output they prefer for a given input. We then train a reward model to predict the human-preferred output”; Ouyang, Pg. 8, Para. 7, “Reward modeling (RM). Starting from the SFT model with the final unembedding layer removed, we trained a model to take in a prompt and response, and output a scalar reward”; and Ouyang, Pg. 3, Fig. 2, “A diagram illustrating the three steps of our method: . . . (2) reward model (RM) training . . . In Step 2, boxes A-D are samples from our models that get ranked by labelers”, wherein training the reward LLM, “We then train a reward model” that is the LLM “SFT model with the final unembedding layer removed”, is performed using numerical rankings, “In Step 2, boxes A-D are samples from our models that get ranked by labelers”, of the outputs that were generated, “We collect a dataset of comparisons between model outputs, where labelers indicate which output they prefer for a given input. We then train a reward model to predict the human-preferred output”, by the MLS, “the SFT model”, which, in view of Sulem, the outputs are CEs, see Sulem, Pg. 1, Abstract, “we propose a model-agnostic algorithm that generates counterfactual ensemble explanations for time series anomaly detection models”). The reasons for obviousness were discussed in regard to the rejection of claim 1 above and remain applicable here. Regarding Claim 9, Ouyang in view of Sulem and Wu teach the method as recited in claim 1, wherein the fine-tuning comprises: inputting a common group of time series anomalies to both the MLS and the RLMLS (Ouyang, Pg. 9, Para. 2-3, “Reinforcement learning (RL). Once again following Stiennon et al. (2020), we fine-tuned the SFT model on our environment using PPO . . . we add a per-token KL penalty from the SFT model at each token to mitigate overoptimization of the reward model. . . We maximize the following combined objective function in RL training: PNG media_image1.png 80 612 media_image1.png Greyscale ”, where a common group of data, “x”, is input to both the MLS, “SFT”, and the RLMLS, “RL”, which, in view of Sulem, the input is time series anomalies, see Sulem, Pg. 2, Para. 5, “We introduce a model-agnostic method that generates counterfactual ensemble explanations for anomalies detected by an algorithm in a univariate or multivariate time series” and Sulem, Pg. 5, Fig. 1, “The original input (1a) is a univariate time series window containing an anomalous subsequence (highlighted in red)”); comparing respective CE outputs of the MLS and the RLMLS to identify divergences between the CE outputs of the MLS and the CE outputs of the RLMLS (Ouyang, Pg. 9, Para. 2-3, “Reinforcement learning (RL). Once again following Stiennon et al. (2020), we fine-tuned the SFT model on our environment using PPO . . . we add a per-token KL penalty from the SFT model at each token to mitigate overoptimization of the reward model. . . We maximize the following combined objective function in RL training: PNG media_image1.png 80 612 media_image1.png Greyscale ”, where, through the “KL penalty”, the “objective function” compares the outputs, “y”, of the MLS, “SFT”, and RLMLS, “RL”, in order to identify the “KL”, which a person of ordinary skill in the art would know stands for Kullback–Leibler divergence, between the outputs, “y”, of the MLS, “SFT”, and RLMLS, “RL”, and where, in view of Sulem, the outputs are CE outputs, see Sulem, Pg. 2, Para. 5, “We introduce a model-agnostic method that generates counterfactual ensemble explanations for anomalies detected by an algorithm in a univariate or multivariate time series”); merging the divergences with scores of the CE outputs to form a merged output (Ouyang, Pg. 9, Para. 2-3, “Reinforcement learning (RL). Once again following Stiennon et al. (2020), we fine-tuned the SFT model on our environment using PPO . . . we add a per-token KL penalty from the SFT model at each token to mitigate overoptimization of the reward model. . . We maximize the following combined objective function in RL training: PNG media_image1.png 80 612 media_image1.png Greyscale ”, where the divergences, “βlog( . . .)”, are merged with scores of the outputs, “rθ(x,y)”, to form a merged output, depicted in Fig. 2 as “The reward . . . rk”, see Ouyang, Pg. 3, Fig. 2, and where, in view of Sulem, the outputs are CE outputs, see Sulem, Pg. 2, Para. 5, “We introduce a model-agnostic method that generates counterfactual ensemble explanations for anomalies detected by an algorithm in a univariate or multivariate time series”); providing the merged output to a proximal policy optimization (PPO) process (Ouyang, Pg. 3, Fig. 2, “The reward is used to update the policy using PPO”, where the merged output, “The reward”, must be provided to the “PPO” for it to be used to “update the policy”); and with the PPO process, using the merged output to fine tune the RLMLS (Ouyang, Pg. 3, Fig. 2, “The reward is used to update the policy using PPO” and Ouyang, Pg. 2, Para. 3, “Finally, we use this RM as a reward function and fine-tune our supervised learning baseline to maximize this reward using the PPO algorithm”, where, with the PPO process, “using PPO”, using the merged output, “reward”, to “fine-tune” the RLMLS, “the policy”). The remaining limitations are substantially the same as limitations of Claim 1, therefore it is rejected under the same rationale. Regarding Claim 10, Ouyang in view of Sulem and Wu teach the method as recited in claim 1, further comprising: receiving, by the RLMLS, a set of time-series data that comprises one or more anomalies (Ouyang, Pg. 6, Para. 4, “Step 3: Optimize a policy against the reward model using PPO . . . We fine-tune the supervised policy” and Ouyang, Pg. 3, Fig. 2, “(3) reinforcement learning via proximal policy optimization (PPO)”, where, as depicted in Fig. 2, the RLMLS, the “fine-tune[d] . . . supervised policy” with “reinforcement learning”, receives an input, “A new prompt is sampled from the dataset”, which, in view of Sulem, the input data is time-series data that comprises one or more anomalies, see Sulem, Pg. 2, Para. 5, “We introduce a model-agnostic method that generates counterfactual ensemble explanations for anomalies detected by an algorithm in a univariate or multivariate time series” and Sulem, Pg. 5, Fig. 1, “The original input (1a) is a univariate time series window containing an anomalous subsequence (highlighted in red)”); and generating, by the RLMLS, a respective CE for one or more of the anomalies in the set of time-series data (Ouyang, Pg. 6, Para. 4, “Step 3: Optimize a policy against the reward model using PPO . . . We fine-tune the supervised policy” and Ouyang, Pg. 3, Fig. 2, “(3) reinforcement learning via proximal policy optimization (PPO)”, where, as depicted in Fig. 2, the RLMLS, the “fine-tune[d] . . . supervised policy” with “reinforcement learning”, generates output data, “The policy generates an output”, which, in view of Sulem, the output data is a respective CE for one or more of the anomalies in the set of time-series data, see Sulem, Pg. 2, Para. 5, “We introduce a model-agnostic method that generates counterfactual ensemble explanations for anomalies detected by an algorithm in a univariate or multivariate time series” and Sulem, Pg. 5, Fig. 1, “The original input (1a) is a univariate time series window containing an anomalous subsequence (highlighted in red)”), and the CEs are comprehensible by a human (Sulem, Pg. 2, Para. 5, “We propose an interpretable visualization of the counterfactual ensemble explanation, in particular by associating the counterfactual examples and their prediction scores under the model”, where the CEs, “counterfactual examples” are comprehensible by a human, “We propose an interpretable visualization”; see also Sulem, Pg. 10, Para. 4, “On Figure 2, we present a visualization of our method on anomalies from the KPI dataset, detected by NCAD and USAD. We observe that the counterfactual ensemble explanation from DPE (in red color scale), ICE (in green), and FS (in purple) are quite dissimilar” and Sulem, Pg. 11, Fig. 2, “Figure 2: Time series windows contains an anomaly and our counterfactual ensemble explanations”). The reasons for obviousness were discussed in regard to the rejection of claim 1 above and remain applicable here. Regarding Claim 11, Ouyang teaches a non-transitory storage medium having stored therein instructions that are executable by one or more hardware processors to perform operations comprising: . . . (Pg. 3, Fig. 2, “A diagram illustrating the three steps of our method: (1) supervised fine-tuning (SFT), (2) reward model (RM) training, and (3) reinforcement learning via proximal policy optimization (PPO)on this reward model. Blue arrows indicate that this data is used to train one of our models . . . See Section 3 for more details on our method”, where a person of ordinary skill in the art would understand that performance of the “method”, which includes “(1) supervised fine-tuning (SFT), (2) reward model (RM) training, and (3) reinforcement learning via proximal policy optimization (PPO)”, requires a non-transitory storage medium storing instruction that are executed by hardware processors to perform operations; see also Pg. 10, Para. 8, “We are releasing samples from our models on all of the sampling-based NLP tasks.7 . . . 7Accessible here: https://github.com/openai/following-instructions-human-feedback”, where code “release[ed]” on “github” is instructions stored on a non-transitory storage medium that must be executed by processors to perform “sampling-based NLP tasks”; see generally Pg. 17, Para. 4, “The cost of collecting our data and the compute for training runs, including experimental runs is a fraction of what was spent to train GPT-3: training our 175B SFT model requires 4.9 petaflops/s-days and training our 175B PPO-ptx model requires 60 petaflops/s-days, compared to 3,640 petaflops/s-days for GPT-3 (Brown et al., 2020). At the same time, our results show that RLHF is very effective at making language models more helpful to users, more so than a 100x model size increase”). The remaining limitations are substantially the same as limitations of Claim 1, therefore it is rejected under the same rationale. Regarding Claim 13, the additional elements of the dependent claim are substantially the same as limitations of Claim 3, therefore it is rejected under the same rationale. Regarding Claim 14, the additional elements of the dependent claim are substantially the same as limitations of Claim 4, therefore it is rejected under the same rationale. Regarding Claim 15, the additional elements of the dependent claim are substantially the same as the limitations of Claim 5, therefore it is rejected under the same rationale. Regarding Claim 16, the additional elements of the dependent claim are substantially the same as the limitations of Claim 6, therefore it is rejected under the same rationale. Regarding Claim 17, the additional elements of the dependent claim are substantially the same as the limitations of Claim 7, therefore it is rejected under the same rationale. Regarding Claim 18, the additional elements of the dependent claim are substantially the same as limitations of Claim 8, therefore it is rejected under the same rationale. Regarding Claim 19, the additional elements of the dependent claim are substantially the same as the limitations of Claim 9, therefore it is rejected under the same rationale. Regarding Claim 20, the additional elements of the dependent claim are substantially the same as the limitations of Claim 10, therefore it is rejected under the same rationale. Claims 2 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Ouyang in view of Sulem, Wu, and Aharoni et al. (hereinafter Aharoni) (“Unsupervised Domain Clusters in Pretrained Language Models”). Regarding Claim 2, Ouyang in view of Sulem and Wu teach the method as recited in claim 1, wherein, . . . [during] the unsupervised training, the MLLM was trained with multi-modal data (Ouyang, Pg. 6, Para. 1, “We start with a pretrained language model (Radford et al., 2019; Brown et al., 2020; Fedus et al., 2021; Rae et al., 2021; Thoppilan et al., 2022)” and Ouyang, Pg. 24, Cit. 9, “Radford . . . Language models are unsupervised multitask learners”, where the “pretrained language model” is a “are unsupervised multitask learners”; see also Ouyang, Pg. 8, Para. 5, “We start with the GPT-3 pretrained language models from Brown et al. (2020)”, where a person of ordinary skill in the art would understand that “the GPT-3 pretrained language models from Brown et al. (2020)” are generated using unsupervised training performed using data, which in view of Wu, is a multi-modal model trained with multi-modal data, see Wu, Pg. 2248, Col. 1, Fig. 1, “The definition of multimodal”, where the “Multimodal” “Training” uses “Input” data comprising “Text” and images, “Visual”; see also Wu, Pg. 2247, Col. 1, Para. 2, “A multimodal model combines multiple data types, including images, text, audio, and more” and Wu, Pg. 2248, Col. 1, Para. 3, “From a data perspective, multimodal data can be seen as a combination of different data types, such as images, numerical data, text, symbols, audio, time series”, where “time series” data comprising “text” and “images” is suggested for use in unsupervised training, see Wu, Pg. 3, Col. 1, Para. 2, “Large-scale multimodal (2020-?) . . . CLIP liberates the burden of assembling massive datasets with predetermined class counts. Instead, CLIP empowers the collection of image-text pairs and leverages unsupervised techniques to either predict their similarity or generate them” and Wu, Pg. 2250-2251, “Generative pre-training is an important method and training objective in self-supervised learning, where the model learns how to generate data without relying on labels or manual . . . enabling large-scale self-supervised pre-training). The remaining limitations are substantially the same as limitations of Claim 1, therefore it is rejected under the same rationale. Ouyang in view of Sulem and Wu do not explicitly disclose . . . prior to . . . . However, Aharoni teaches . . . [pretraining of a model] prior to [the unsupervised training] . . . (Pg. 7748, Col. 1, Para. 3, “we propose . . . positive-unlabeled fine-tuning of pretrained language models” and Pg. 7755, Col. 2, Para. 3, “We showed that massive pre-trained language models are highly effective in mapping data to domains in a fully-unsupervised manner”, where “pre-trained language models” are “fine-tun[ed]” “in a fully-unsupervised manner”). Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the MLLM trained with multi-modal data during unsupervised training of Ouyang in view of Sulem and Wu with the pretraining of a model prior to the unsupervised training of Aharoni in order to achieve the benefits of supervised fine-tuning, while reducing its resource expenditure and lessening its complex data requirements by first performing unsupervised fine-tuning on the pretrained multi-modal model (compare Ouyang, Pg. 6, Para. 2, “Collect demonstration data, and train a supervised policy. Our labelers provide demonstrations of the desired behavior on the input prompt distribution (see Section 3.2 for details on this distribution). We then fine-tune a pretrained GPT-3 model on this data using supervised learning”, where “supervised learning” allows for training on the “desired behavior on the input”, but requires complex “data” generated by resource constrained “labelers” with Aharoni, Pg. 7747, Col. 1, Abstract, “We show that massive pre trained language models implicitly learn sentence representations that cluster by domains without supervision–suggesting a simple data driven definition of domains in textual data”, where “simple data” “without supervision” can be used for “domains in textual data”; see also Aharoni, Pg. 7747, Col. 1, “domain labels are usually unavailable–e.g. in large-scale web-crawled data”) using an effective training method than performs similarly or better than other established methods (Aharoni, Pg. 7755, Col. 2, Para. 3, “We proposed new methods to harness this property for domain data selection using distance-based ranking in vector space and pretrained LM fine tuning, requiring only a small set of in-domain data. We demonstrated the effectiveness of our methods on a new, improved data split we created for a previously studied multi-domain machine translation benchmark. Our methods perform similarly or better than an established data selection method and oracle in-domain training across all five domains in the benchmark” but see Wu, Pg. 2247, Col. 1, Para. 2, “A multimodal model combines multiple data types, including images, text, audio, and more”, where downstream supervised learning may still be needed for other data modalities and complex behaviors). Regarding Claim 12, the additional elements of the dependent claim are substantially the same as limitations of Claim 2, therefore it is rejected under the same rationale. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Brown et al. (“Language Models are Few-Shot Learners”) discloses background information on the pretrained model used by primary reference Ouyang. Sun et al. (“Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervision”) discloses a large language model pipeline that incorporates annotations from subject matter experts to supervise model processes. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MATTHEW BRYCE GOLAN whose telephone number is (571)272-5159. The examiner can normally be reached Monday through Friday, 8:00 AM to 5:00 PM ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571) 270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MATTHEW BRYCE GOLAN/Examiner, Art Unit 2123 /ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123
Read full office action

Prosecution Timeline

Apr 03, 2024
Application Filed
Aug 12, 2026
Non-Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
0%
Grant Probability
0%
With Interview (+0.0%)
3y 9m (~1y 3m remaining)
Median Time to Grant
Low
PTA Risk
Based on 8 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month