Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. § 101 because the claimed invention is directed to an abstract idea without significantly more.
Independent Claims
Step 2A Prong One: Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes, the independent claims recite an abstract idea in the form of mental processes, as identified below. A mental process is a process that “can be performed in the human mind, or by a human using a pen and paper” (MPEP § 2106.04(a)(2)(III), paragraph 1). Examples of mental processes include “observations, evaluations, judgments, and opinions” (MPEP § 2106.04(a)(2)(III), paragraph 2).
Claim 1:
“identifying at least one emotion of a person making the selection” [This limitation is a mental process that can be performed by observation, evaluation, judgment, and opinion because it merely recites “identifying” at a high level of generality, without any specific details that distinguish from being a mental process.]
“responsive to the emotion satisfying an inconsistency threshold with respect to the selection, discounting the selection” [This limitation is a mental process that can be performed by observation, evaluation, judgment, and opinion because it merely recites “discounting” at a high level of generality, without any specific details that distinguish from being a mental process.]
Claim 10:
“associate at least one emotion signal indicating an emotion of a person with at least one evaluation input indicating an evaluation of the person of at least one audio and/or video output of at least one machine learning (ML) model trained to generate text and/or images” [This limitation is a mental process that can be performed by observation, evaluation, judgment, and opinion because it merely recites “associating” at a high level of generality, without any specific details that distinguish from being a mental process. Furthermore, the limitation of the “output of the at least one machine learning (ML) model” only defines the subject matter of the “associate” step and does not require a step of performing the outputting. Thus, the output” limitation merely defines the subject matter of the mental process and is not an additional element besides the abstract idea.]
Claim 18:
“correlating the emotion signal with the output” [This limitation is a mental process that can be performed by observation, evaluation, judgment, and opinion because it merely recites “associating” at a high level of generality, without any specific details that distinguish from being a mental process.]
Therefore, the independent claims recite a judicial exception.
Step 2A Prong Two: Does the claim recite additional elements that integrate the judicial exception into a practical application?
No. The judicial exception recited in the above discussed claims is not integrated into a practical application.
Independent claims 1, 10, and 18 recite the following additional elements, but these additional elements are not sufficient to integrate the judicial exception into a practical application:
“presenting at least a first output from at least one machine learning (ML) model on at least one display and/or at least one speaker; receiving a selection with respect to the output” (claim 1) and “presenting output from a machine learning (ML) model; receiving at least one emotion signal from at least one person” (claim 18) [This element constitutes “adding insignificant extra-solution activity to the judicial exception” (MPEP § 2106.05(g)) since it merely amounts to necessary data gathering or outputting, which identifies is identified in MPEP § 2106.05(g) as a form of extra-solution activity.]
“in training the ML model” (claim 1), “execute reinforcement learning from human feedback (RLHF) on the ML model according to the association of the emotion signal with the evaluation input” (claim 10) and “executing reinforcement learning on the ML model according to the correlating of the emotion signal with the output.” (claim 18) [These elements constitute no more than mere instructions to apply the judicial exception using generic computer functions (MPEP § 2106.04(d)(I)), namely the generic computer function of machine learning (as in claim 1) or reinforcement learning (as in claims 10 and 18), which is regarded as a generic machine learning function when stated only generically, as a tool to apply an abstract idea. Alternatively, or additionally, these elements do no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)), namely the technological environment of training or training reinforcement, since the term “in” or “according to” does not specify any specific implementation details of the training process, but instead only generally links the mental process to the training process.]
“a processor system” (claim 10) and “A device comprising: a computer memory that is not a transitory signal and that comprises instructions executable by at least one processor system for” (claim 18) [These elements constitute no more than mere instructions to apply the judicial exception using generic computer components (MPEP § 2106.04(d)(I)). These additional elements merely invoke the use of generic computer components, namely a processor system or a device comprising memory, as tools to perform the abstract idea, and do not place any limitations on the abstract idea other than the use of such generic computer components. Therefore, these additional elements do not integrate the judicial exception into a practical application.]
Therefore, under MPEP 2106.04(d), the additional elements of the claims do not integrate the judicial exception into a practical application.
Step 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception?
No. The claims do not include additional elements that are sufficient for the claims to amount to significantly more than the judicial exception.
Additional elements that are mere instructions to apply an exception or merely generally linking or generally linking the use of a judicial exception to a particular technological environment or field of use do not constitute significantly more than a judicial exception under MPEP § 2106.05(I)(A).
Additional elements that are considered to be extra-solution activity do not amount to significantly more if, upon their reevaluation in Step 2B, they are also merely appending “well-understood, routine, conventional activities previously known to the industry, specified at a high level of generality, to the judicial exception” (MPEP § 2106.05(I)(A)). Here, the additional elements that were previously identified as extra-solution activity are reevaluated as follows:
“presenting at least a first output from at least one machine learning (ML) model on at least one display and/or at least one speaker; receiving a selection with respect to the output” (claim 1) and “presenting output from a machine learning (ML) model; receiving at least one emotion signal from at least one person” (claim 18) [These elements are well-understood, routine, conventional activity because presenting information on a device or speaker, and receiving inputs from a user is conventional. See, e.g., US 2026/0236303 A1, [0039]: “User interface 108 can include conventional hardware components such as a display, keyboard, keypad, touchscreen, speakers, microphone, and/or any other component that can be operated to receive input from and/or present output to a user.” US 2023/0308505 A1, [0056]: “Generally, conventional computing devices receive input from users by way of a plurality of input devices. Conventional input devices may include a mouse, camera, joystick, keyboard, trackpad, or microphone.” US 2025/0284921 A1, [0079]: “The input devices enable the user to communicate information and select commands to the electronic system. The input devices 830 include alphanumeric keyboards, pointing devices (also called “cursor control devices”), audio input devices (such as microphones), etc. The output devices 835 display images and information generated by the electronic system 800. The output devices 835 include display devices, such as liquid crystal displays (LCD) and organic light emitting diode (OLED) displays, as well as other conventional output devices, such as printers and audio speakers. Some embodiments include devices such as a touchscreen that functions as both input and output devices.” Furthermore, the act of “receiving” is regarded as a limitation of “receiving or transmitting data over a network” or “storing and retrieving information in memory,” which MPEP § 2106.05(d)(II) identifies as an example of well‐understood, routine, and conventional computer functions when recited in a generic manner.]
Dependent Claims
The remaining dependent claims being rejected do not recite additional elements, whether considered individually or in combination, that are sufficient to integrate the judicial exception into a practical application or amount to significantly more than the judicial exception.
Claim 2: “wherein discounting the selection comprises discarding the selection” [This limitation is a mental process that can be performed by observations, evaluations, judgments, and opinions because the act of discarding is recited at a high degree of generality.] “from training the ML model” [This element is treated as an additional limitation besides the abstract idea. However, since the effect of the discarding on the model training is not specifically stated beyond the preparation of data, and neither is it stated or implied how the selection and the discarding of it affects the model, this elements does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)), namely the technological environment of training or training reinforcement.]
Claim 3: “wherein discounting the selection comprises associating a first weight to the selection in training the ML model, the first weight being less than a second weight of a non-discounted selection.” [This limitation is a mental process that can be performed by observation, evaluation, judgment, and opinion because the act of “associating” covers data analysis that can be performed by a human.]
Claim 4: “receiving the selection via a point-and-click device.” [This limitation merely defines additional characteristics the insignificant extra-solution activity identified in the parent independent claim. Thus, it is an additional element that is insignificant extra-solution activity for the same reasons discussed for the corresponding limitation in the parent claim. As discussed in the rejection of the parent independent claim, the use of a point and click device such as a mouse for user input is well-understood, routine, conventional activity, as discussed in the analysis for the parent independent claim.]
Claim 5: “receiving the selection via a camera.” [This limitation merely defines additional characteristics the insignificant extra-solution activity identified in the parent independent claim. Thus, it is an additional element that is insignificant extra-solution activity for the same reasons discussed for the corresponding limitation in the parent claim. Furthermore, the receipt of a selection with a camera is well-understood, routine, conventional activities, as discussed in the analysis for the parent independent claim.]
Claim 6: “receiving the selection via a microphone.” [This limitation merely defines additional characteristics the insignificant extra-solution activity identified in the parent independent claim. Thus, it is an additional element that is insignificant extra-solution activity for the same reasons discussed for the corresponding limitation in the parent claim. As discussed in the rejection of the parent independent claim, the use of a microphone for input is well-understood, routine, conventional activities, as discussed in the analysis for the parent independent claim.]
Claim 7: “wherein the output comprises at least one image.” [This element is an additional element besides the abstract idea but does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)), namely the technological environment of a model that outputs an image.]
Claim 8: “wherein the output comprises text.” [This element is an additional element besides the abstract idea but does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)), namely the technological environment of a model that outputs text.]
Claim 9: “determining whether the emotion satisfies the inconsistency threshold by comparing the emotion to the selection.” [This limitation is a mental process that can be performed by observation, evaluation, judgment, and opinion because a comparison can be performed by a human.]
Claim 11: “wherein the output of the ML model comprises text.” [This element merely defines the subject matter of the output in the mental process that is recited in the parent independent claim. As discussed for the parent independent claim, the output” limitation merely defines the subject matter of the mental process and is not an additional element in the form of an active process.]
Claim 12: “wherein the output of the ML model comprises images.” [This element merely defines the subject matter of the output in the mental process that is recited in the parent independent claim. As discussed for the parent independent claim, the output” limitation merely defines the subject matter of the mental process and is not an additional element in the form of an active process.]
Claim 13: “wherein the output comprises audio.” [This element merely defines the subject matter of the output in the mental process that is recited in the parent independent claim. As discussed for the parent independent claim, the output” limitation merely defines the subject matter of the mental process and is not an additional element in the form of an active process.]
Claim 14: “wherein the output comprises at least one image.” [This element merely defines the subject matter of the output in the mental process that is recited in the parent independent claim. As discussed for the parent independent claim, the output” limitation merely defines the subject matter of the mental process and is not an additional element in the form of an active process.]
Claim 15: “wherein the processor system is configured to execute RLHF on the ML model” [This limitation is an additional element besides the abstract idea but constitutes mere instructions to apply for the same reasons given for the corresponding limitation of “execute reinforcement learning from human feedback…” in the parent independent claim. Although the remainder of this claim recites “using,” the claim does not recite how the evaluation input affects the reinforcement learning process.] “using the evaluation input at a first weight responsive to the association being a first association and execute RLHF on the ML model using the evaluation input at a second weight less than the first weight responsive to the association being a second association.” [This limitation is a mental process that can be performed by observations, evaluations, judgments, and opinions because the association of weights with responses is a process that can be performed by a human.]
Claim 16: “wherein the second weight is zero such that the evaluation input does not affect RLHF on the ML model.” [This element merely defines the subject matter of the weight in the mental process that is recited in the parent independent claim and does not constitute an additional element besides the abstract idea.]
Claim 17: “wherein the processor system is configured to execute RLHF on the ML model using the emotion signal.” [This limitation is an additional element besides the abstract idea but constitutes mere instructions to apply for the same reasons given for the corresponding limitation of “execute reinforcement learning from human feedback…” in the parent independent claim. Although the remainder of this claim recites “using,” the claim does not recite the implementation details of how the emotion signal affects the reinforcement learning process.]
Claim 19: “wherein the instructions are executable for: executing reinforcement learning on the ML model according to evaluation input originated by the person.” [This limitation is an additional element besides the abstract idea but constitutes mere instructions to apply for the same reasons given for the corresponding limitation of “execute reinforcement learning from human feedback…” in the parent independent claim. Although this claim also “according to” the claim does not recite the implementation details of how the evaluation input affects the reinforcement learning process.]
Claim 20: “wherein the instructions are executable for: executing reinforcement learning on the ML model according to a relationship between evaluation input originated by the person and the emotion signal.” [This limitation is an additional element besides the abstract idea but constitutes mere instructions to apply for the same reasons given for the corresponding limitation of “execute reinforcement learning from human feedback…” in the parent independent claim. Although this claim also “according to” the claim does not recite the implementation details of how the relationship affects the reinforcement learning process.]
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 18-20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Li et al., “Facial feedback for reinforcement learning: a case study and offline analysis using the TAMER framework,” Autonomous Agents and Multi-Agent Systems (2020) 34:22 (“Li”).
As to claim 18, Li teaches a device comprising:
a computer memory that is not a transitory signal and that comprises instructions executable by at least one processor system for: [Since Li teaches “interactive reinforcement learning” (abstract) involving “computational models” (§ 2.2, paragraph 1) that learns in real time (see § 3.2, paragraph 1: “An agent implemented according to TAMER learns from real-time evaluations of its behavior, provided by a human teacher who observes the agent. These evaluations are taken as human reward signals.”), the instant limitations of generic components of a computer are implied.]
presenting output from a machine learning (ML) model; [§ 3.2, paragraph 3: “The TAMER agent will select actions with the value function model to get the most accumulated discounted human reward. Figure 2 shows the diagram of agent learning in the TAMER framework.” The selection of the action is from a machine learning model, namely a reinforcement learning model. See § 3.1, paragraph 3: “As in traditional reinforcement learning (RL), an interactive RL agent learns to make sequential decisions in a task, represented by a policy deciding the action to be taken by the agent in an environmental state.”]
receiving at least one emotion signal from at least one person; [§ 4, paragraph 2: “We recorded the training data including state observation, action, human reward, the time of human reward being given, score, the ending time for each time step, and video data of facial expressions and keypresses on the keyboard for each trainer during training.” § 6.2.2, paragraph 3: “we model the spatio-temporal dynamics of the expressions for distinguishing between negative and positive feedback classes.” In this case, “positive” and “negative” facial expressions are emotions of the person making the selection. See also § 1, paragraph 4: “We recorded the facial expressions of all trainers during training and, in some conditions, told participants that their facial expressions would be used as encouraging explicit feedback, e.g., happy and sad expressions would map to positive and negative reward respectively, in addition to keypresses, to train the agent.”]
correlating the emotion signal with the output; [This limitation is taught by Li because as noted above, the emotional signal is collected in response to the output. See § 3.2, paragraph 1: “An agent implemented according to TAMER learns from real-time evaluations of its behavior, provided by a human teacher who observes the agent. These evaluations are taken as human reward signals. Therefore, TAMER is a typical interactive reinforcement learning method.”] and
executing reinforcement learning on the ML model [Li, § 3.2, paragraph 3: “The TAMER agent will select actions with the value function model to get the most accumulated discounted human reward. Figure 2 shows the diagram of agent learning in the TAMER framework.” The selection of the action is from a machine learning model, namely a reinforcement learning model. See Li, § 3.1, paragraph 3: “As in traditional reinforcement learning (RL), an interactive RL agent learns to make sequential decisions in a task, represented by a policy deciding the action to be taken by the agent in an environmental state.” Li, § 4, paragraph 2: “We recorded the training data including state observation, action, human reward, the time of human reward being given, score, the ending time for each time step, and video data of facial expressions and keypresses on the keyboard for each trainer during training.”] according to the correlating of the emotion signal with the output. [See FIG. 9 and § 6.3: “We compare the average learning performance of the four conditions in terms of learning from keypress feedback, learning from binary keypress feedback, learning from random feedback and learning from predicted binary feedback. The differences between these four kinds of feedback are illustrated in Fig. 9.” Note that in the case of “learning from predicted feedback,” the feedback that is utilized is the emotion signal, as described in § 6.3, paragraph 1 (“learn from the human trainer’s facial expressions with the trained CNN–RNN model in Sect. 6.2.2 to predict human reward based on human trainer’s facial expressions… The closest approximation is to get predicted human reward based on facial expressions at the time when keypress feedback was given for the complete training trajectory of each trainer. Thus, we use the predicted feedback…to train the agent for the complete trajectory.”), Furthermore, since the instant claim merely says “according to the correlating,” the other training methods, including the keypress feedback would also read on the instant claim, because both keypress and emotional signals were recorded together, both correlating to the action output of the model, and “according to the correlating” is met by using part of the data that that obtained in the act of correlating the model actions with the feedback.]
As to claim 19, Li teaches the device of claim 18, wherein the instructions are executable for:
executing reinforcement learning on the ML model according to evaluation input originated by the person. [§ 3.1, paragraph 3: “As in traditional reinforcement learning (RL), an interactive RL agent learns to make sequential decisions in a task, represented by a policy deciding the action to be taken by the agent in an environmental state.” See also § 6.3: “We compare the average learning performance of the four conditions in terms of learning from keypress feedback, learning from binary keypress feedback, learning from random feedback and learning from predicted binary feedback. The differences between these four kinds of feedback are illustrated in Fig. 9.” As noted in the rejection of the parent claim, the “predicted human reward” is output from the CNN-RNN model, and this output corresponds to an “evaluation input originated from the person,” noting that the instant claim does not precisely define “evaluation input” (the output of the CNN-RNN) and the “facial expression” (the emotion signal) as being, for example, parallel inputs that are compared with one another for consistency.]
As to claim 20, Li teaches the device of claim 18, wherein the instructions are executable for:
executing reinforcement learning on the ML model according to a relationship between evaluation input originated by the person and the emotion signal. [As shown in FIG. 8, the output that is predicted by the CNN-RNN model has a relationship with the “facial expression” (the emotion signal) by being the predicted feedback predicted by the CNN-RNN. The Examiner notes that the instant claim does not precisely define the “relationship.” Therefore, an input-output relationship qualifies as a relationship within the scope of this claim.]
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
1. Claims 1-2 and 9 are rejected under 35 U.S.C. 103 as being unpatentable over Li et al., “Facial feedback for reinforcement learning: a case study and offline analysis using the TAMER framework,” Autonomous Agents and Multi-Agent Systems (2020) 34:22 (“Li”) in view of Cheng et al., “RIME:Robust Preference-based Reinforcement Learning with Noisy Preferences,” arXiv:2402.17257v2 [cs.LG] 12 Mar 2024 (“Cheng”).
As to claim 1, Li teaches a method comprising:
presenting at least a first output from at least one machine learning (ML) model [§ 3.2, paragraph 3: “The TAMER agent will select actions with the value function model to get the most accumulated discounted human reward. Figure 2 shows the diagram of agent learning in the TAMER framework.” The selection of the action is from a machine learning model, namely a reinforcement learning model. See § 3.1, paragraph 3: “As in traditional reinforcement learning (RL), an interactive RL agent learns to make sequential decisions in a task, represented by a policy deciding the action to be taken by the agent in an environmental state.”] on at least one display and/or at least one speaker; [As shown in FIG. 2, the action is conveyed to a human using a “sensory display.” Furthermore, in the example of the Mario domain described in this reference, the action is displayed. See § 3.3, paragraph 2: “In Infinite Mario, the Mario avatar must move towards the right of the screen as fast as possible and at the same time collect as many points as possible”; § 4, paragraph 1: “In each group, each participant sat at her own table, facing away from the other members and their screens. There was also a camera on the screen in front of the trainer’s face.” Here, “screen” refers to the screen that displays the progress of the Mario task.]
receiving a selection with respect to the output; [§ 3.2, last paragraph: “The TAMER agent learns by repeatedly taking an action, sensing reward, and updating the predictive model ̂R H and corresponding value function model. The trainer observes and evaluates the agent’s behavior. In our experiment, she can give reward by pressing two buttons on the keyboard, which are assigned to the agent’s most recent actions. Each press of the two buttons is mapped to a numeric reward of − 1 or + 1 respectively.”]
identifying at least one emotion of a person making the selection; [§ 4, paragraph 2: “We recorded the training data including state observation, action, human reward, the time of human reward being given, score, the ending time for each time step, and video data of facial expressions and keypresses on the keyboard for each trainer during training.” § 6.2.2, paragraph 3: “we model the spatio-temporal dynamics of the expressions for distinguishing between negative and positive feedback classes.” In this case, “positive” and “negative” facial expressions are emotions of the person making the selection. See also § 1, paragraph 4: “We recorded the facial expressions of all trainers during training and, in some conditions, told participants that their facial expressions would be used as encouraging explicit feedback, e.g., happy and sad expressions would map to positive and negative reward respectively, in addition to keypresses, to train the agent.”] and
Li does not explicitly teach “responsive to the emotion satisfying an inconsistency threshold with respect to the selection, discounting the selection in training the ML model.”
Cheng teaches “responsive to the emotion satisfying an inconsistency threshold with respect to the selection, discounting the selection in training the ML model.” [§ 4.1, paragraph 1: “we theoretically establish a lower bound on the KL divergence between the predicted preference Pψ and the annotated preference ˜y for corrupted samples, in order to filter out large-loss corrupted samples.” This lower bound is described in more detail in expressions (4) and (5). See text following equation (5): “At each training step for the reward model, we apply the threshold in Equation (5) to identify trustworthy sample dataset Dt, as described below.” See also algorithm 1 (on page 11 of this reference), which teaches in line 24, “Filter trustworthy samples Dt using lower bound τlower as in Equation (6).” This filtering out (i.e., discounting the selection) is based on the threshold (lower bound), which is an inconsistency threshold because it measures the KL divergence (inconsistency) between the predicted and annotated preference.]
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Li with the teachings of Cheng by implementing the training such that “responsive to the emotion satisfying an inconsistency threshold with respect to the selection, discounting the selection in training the ML model.” The motivation for doing so would have been to implement a system that achieves “robustness to noisy preference labels” (Cheng, § 1 paragraph 3: “In this work, we aim to improve the robustness of preference based RL methods on noisy and quantitatively limited preferences.”), noting that the concept of the annotated preference and the predicted preference are analogous to the emotion and the selection of the instant claim, because it is already known from Li that “humans may get tired of giving explicit feedback (e.g., button presses to indicate positive or negative reward) as training time progresses” (see Li, § 1, paragraph 3).
As to claim 2, the combination of Li and Cheng teaches the method of claim 1, as set forth above.
Cheng further teaches “wherein discounting the selection comprises discarding the selection from training the ML model.” [§ 4.1, paragraph 1: “Inspired by this, a theoretical lower bound on the KL divergence between the predicted preference Pψ and the annotated preference ˜y for corrupted samples could be established to filter out large-loss corrupted samples.”]
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Li and Cheng to have also arrived at the claimed invention of the instant dependent claim. The motivation for doing so is covered by the motivation given for the teachings of Cheng in the rejection of the parent claim.
As to claim 9, the combination of Li and Cheng teaches the method of claim 1, as set forth above.
Cheng further teaches “comprising determining whether the emotion satisfies the inconsistency threshold by comparing the emotion to the selection” [§ 4.1, paragraph 1: “Inspired by this, a theoretical lower bound on the KL divergence between the predicted preference Pψ and the annotated preference ˜y for corrupted samples could be established to filter out large-loss corrupted samples.” This lower bound is described in more detail in expressions (4) and (5). See text following equation (5): “At each training step for the reward model, we apply the threshold in Equation (5) to identify trustworthy sample dataset Dt, as described below.” See also algorithm 1, which teaches in line 20, “Filter trustworthy samples Dt using lower bound τlower as in Equation (6).” This filtering out (i.e., discounting the selection) is based on the threshold (lower bound), which is inconsistency threshold because it measures the KL divergence (inconsistency) between the predicted and annotated preference.]
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Li and Cheng to have also arrived at the claimed invention of the instant dependent claim. The motivation for doing so is covered by the motivation given for the teachings of Cheng in the rejection of the parent claim.
2. Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over Li in view of Cheng and further in view of Ren et al., “Learning to Reweight Examples for Robust Deep Learning,” arXiv:1803.09050v3 [cs.LG] 5 May 2019 (“Ren”).
As to claim 3, the combination of Li and Cheng teaches the method of claim 1, as set forth above, but does not explicitly teach the further limitations of the instant dependent claim.
Ren teaches “wherein discounting the selection comprises associating a first weight to the selection in training the ML model, the first weight being less than a second weight of a non-discounted selection.” [Abstract: “To determine the example weights, our method performs a meta gradient descent step on the current mini-batch example weights (which are initialized from zero) to minimize the loss on a clean unbiased validation set.” See § 3.1 for a description of setting weights of examples that are noisy, wherein high noise corresponds to low weights. See also FIG. 3 for examples of weight distribution and § 4.3: “As shown in the left figure of Figure 3, our model correctly pushes most noisy images to zero weights.” Note that “noise” in this context is analogous to the concept of inconsistency in the existing combination of references.]
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of the references combined thus far with the teachings of Ren by implementing the use of weights so as to arrive at the claimed invention. The motivation for doing so would have been to implement weighting to enable robustness against training set biases (see Ren, § 2, paragraph 3: “One crucial advantage of reweighting examples is robustness against training set bias.”).
3. Claims 4 and 6 are rejected under 35 U.S.C. 103 as being unpatentable over Li in view of Cheng and further in view of Wang et al., “Incorporating Voice Instructions in Model-Based Reinforcement Learning for Self-Driving Cars,” arXiv:2206.10249v1 [cs.HC] 21 Jun 2022 (“Wang”).
As to claim 4, the combination of Li and Cheng teaches the method of claim 1, as set forth above, but does not teach the further limitations of the instant dependent claim.
Wang teaches “comprising receiving the selection via a point-and-click device” [§ 2, paragraph 5: “Various types of human interventions have been explored in HILL. For example, the simplest one is the hardware delivered feedback, mainly including mouse clicks, keyboard keys, and other sensors [16, 21].”]
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of the references combined thus far with the teachings of Wang by using a mouse (a point and click device) for user feedback, so as to arrive at the claimed invention of the instant dependent claim. Doing so would have been obvious as a simple combination of prior art elements according to known methods to yield predictable results. Since Wang teaches that use of a mouse to provide feedback for reinforcement learning is known in the art as a simple method for a user to provide feedback (see parts cited above), one of ordinary skill in the art could have combined the elements as claimed by known methods, and that in combination, each element merely performs the same function as it does separately; and one of ordinary skill in the art would have recognized that the results of the combination were predictable, namely the result of receiving feedback through a particular type of device.
As to claim 6, the combination of Li and Cheng teaches the method of claim 1, as set forth above, but does not teach the further limitations of the instant dependent claim.
Wang teaches “comprising receiving the selection via a microphone” [Page 6 (§ 4), second-to-last paragraph: “First, we use Google’s speech-to-text (STT) tool [10], which can convert audio into text with real-time speech recognition. Each time the microphone receives voice instructions, it sends it to the server of STT and returns the corresponding text quickly. This service enables us to process a human coach’s live instructions like in a behind-the-wheel driving class.”]
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention the teachings of the references combined thus far with the teachings of Wang by implementing the method to use a microphone to receive the selection, so as to arrive at the claimed invention of the instant dependent claim. The motivation would have been to utilize a manner of communicating feedback that is natural for a human (see Wang, § 2, paragraph 5: “we can use mores natural ways to provide feedback such as facial feedback [2, 5], natural language feedback [17, 22] and gesture feedback [7, 32]”).
4. Claims 5 are rejected under 35 U.S.C. 103 as being unpatentable over Li in view of Cheng and further in view of Trick et al., “Interactive Reinforcement Learning With Bayesian Fusion of Multimodal Advice,” in IEEE Robotics and Automation Letters, vol. 7, no. 3, pp. 7558-7565, July 2022 (“Trick”).
As to claim 5, the combination of Li and Cheng teaches the method of claim 1, as set forth above, but does not teach the method comprising the further limitations of the instant dependent claim.
Trick teaches “receiving the selection via a camera.” [§ III.B, paragraph 3: “Besides speech commands, humans also use nonverbal cues to communicate intentions, in particular when they refer to objects [30]. Therefore we also chose arm gestures as an advice modality. The gestures are predefined, namely pointing gestures for objects and a 2-arm symbolic gesture for the action pour. Using an RGB-D camera (Intel Realsense D435), the human skeleton is tracked based on Openpose [31]. Missing skeleton frames are interpolated using univariate splines. The tracked joint positions of arms and shoulders are aligned with the neck joint and scaled to uniform length in order to become invariant to the human-camera distance.”]
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention combined the teachings of the references combined thus far with the teachings of Trick by implementing the method to further comprise receiving the selection via a camera. The motivation for doing so would have been implement a manner of natural interaction for a human to train a reinforcement learning system, as suggested by Trick (abstract: “Interactive Reinforcement Learning (IRL) has shown promising results in decreasing the learning times of Reinforcement Learning algorithms by incorporating human feedback and advice. In particular, the integration of multimodal feedback channels such as speech and gestures into IRL systems can enable more versatile and natural interaction of everyday users.”).
5. Claims 7-8 are rejected under 35 U.S.C. 103 as being unpatentable over Li in view of Cheng and further in view of Chung et al (US 2024/0256965 A1) (“Chung”).
As to claim 7, the combination of Li and Cheng teaches the method of claim 1, as set forth above, but does not teach the further limitations of the instant dependent claim.
Chung teaches “wherein the output comprises at least one image.” [[0050]: “Furthermore, it is to be understood that inputs and/or outputs can be unimodal or multimodal. For example, inputs or outputs can include data from multiple different data modalities (e.g., text, image, audio, video, etc.).” Note that the context is a model that can be subject to alignment (see [0205]-[0206]), including RLHF (see [0167]: “An example fine-tuning approach includes reinforcement learning. Reinforcement learning can be based on user feedback on model performance during use.”). See also [0174]-[0175].].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of the references combined thus far with the teachings of Chung by implementing the output to comprise at least one image. The motivation for doing so would have been to use a model that is adapted for an image processing task, such as image scaling or generation (see Chung, [0243]: “For example, various different input(s) 2 and output(s) 3 can be used for various different tasks….As another example, machine-learned model(s) 1 can process the image data to generate an upscaled image data output.” [0255]: “In some implementations, the task can be an image generation task.”).
As to claim 8, the combination of Li and Cheng teaches the method of claim 1, as set forth above, but does not teach the further limitations of the instant dependent claim.
Chung teaches “wherein the output comprises text.” [[0050]: “Furthermore, it is to be understood that inputs and/or outputs can be unimodal or multimodal. For example, inputs or outputs can include data from multiple different data modalities (e.g., text, image, audio, video, etc.).” Note that the context is a model that can be subject to alignment (see [0205]-[0206]), including RLHF (see [0167]: “An example fine-tuning approach includes reinforcement learning. Reinforcement learning can be based on user feedback on model performance during use.”). See also [0174]-[0175].].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of the references combined thus far with the teachings of Chung by implementing the output to comprise at least one image. The motivation for doing so would have been to use a model that is adapted for tasks such as question answering (see Chung, [0254]: “In some implementations, the task can be a question answering task….machine-learned model(s) 1 can process input(s) 2 to generate output(s) 3 that represent textual data responsive to the question”).
6. Claims 10-14 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Li in view of Chung.
As to claim 10, Li teaches a processor system [Since Li teaches “interactive reinforcement learning” (abstract) involving “computational models” (§ 2.2, paragraph 1), the instant limitation of a processor system is implied.] configured to:
associate at least one emotion signal indicating an emotion of a person § 4, paragraph 2: “We recorded the training data including state observation, action, human reward, the time of human reward being given, score, the ending time for each time step, and video data of facial expressions and keypresses on the keyboard for each trainer during training.” § 6.2.2, paragraph 3: “we model the spatio-temporal dynamics of the expressions for distinguishing between negative and positive feedback classes.” In this case, “positive” and “negative” facial expressions are emotions of the person making the selection. See also § 1, paragraph 4: “We recorded the facial expressions of all trainers during training and, in some conditions, told participants that their facial expressions would be used as encouraging explicit feedback, e.g., happy and sad expressions would map to positive and negative reward respectively, in addition to keypresses, to train the agent.”] with at least one evaluation input indicating an evaluation of the person [See FIG. 9 and § 6.3: “We compare the average learning performance of the four conditions in terms of learning from keypress feedback, learning from binary keypress feedback, learning from random feedback and learning from predicted binary feedback. The differences between these four kinds of feedback are illustrated in Fig. 9.” Note that in the case of “learning from predicted feedback,” the feedback that is utilized is the emotion signal, as described in § 6.3, paragraph 1 (“learn from the human trainer’s facial expressions with the trained CNN–RNN model in Sect. 6.2.2 to predict human reward based on human trainer’s facial expressions… The closest approximation is to get predicted human reward based on facial expressions at the time when keypress feedback was given for the complete training trajectory of each trainer. Thus, we use the predicted feedback…to train the agent for the complete trajectory.” That is, the “predicted human reward” is output from the CNN-RNN model, and this output corresponds to an “evaluation input originated from the person” associated with the facial expression (emotion signal), noting that the instant claim does not precisely define “evaluation input” (the output of the CNN-RNN) and the “facial expression” (the emotion signal) as being, for example, parallel inputs that are compared with one another for consistency.] of at least one audio and/or video output [As shown in FIG. 2, the action is conveyed to a human using a “sensory display.” Furthermore, in the example of the Mario domain described in this reference, the action is displayed. See § 3.3, paragraph 2: “In Infinite Mario, the Mario avatar must move towards the right of the screen as fast as possible and at the same time collect as many points as possible”; § 4, paragraph 1: “In each group, each participant sat at her own table, facing away from the other members and their screens. There was also a camera on the screen in front of the trainer’s face.” Here, “screen” refers to the screen that displays the progress of the Mario task.] of at least one machine learning (ML) model [§ 3.2, paragraph 3: “The TAMER agent will select actions with the value function model to get the most accumulated discounted human reward. Figure 2 shows the diagram of agent learning in the TAMER framework.” The selection of the action is from a machine learning model, namely a reinforcement learning model. See § 3.1, paragraph 3: “As in traditional reinforcement learning (RL), an interactive RL agent learns to make sequential decisions in a task, represented by a policy deciding the action to be taken by the agent in an environmental state.”] […]
execute reinforcement learning from human feedback (RLHF) on the ML model according to the association of the emotion signal with the evaluation input. [See § 3.1, paragraph 3: “As in traditional reinforcement learning (RL), an interactive RL agent learns to make sequential decisions in a task, represented by a policy deciding the action to be taken by the agent in an environmental state.” As shown in FIG. 8, the output that is predicted by the CNN-RNN model has a relationship with the “facial expression” (the emotion signal) by being the predicted feedback predicted by the CNN-RNN. The Examiner notes that the instant claim does not precisely define the “relationship.” Therefore, an input-output relationship qualifies as a relationship within the scope of this claim.]
Li does not explicitly teach the limitation that the model is “trained to generate text and/or images.”
Chung teaches “trained to generate text and/or images” [[0050]: “Furthermore, it is to be understood that inputs and/or outputs can be unimodal or multimodal. For example, inputs or outputs can include data from multiple different data modalities (e.g., text, image, audio, video, etc.).” Note that the context is a model that can be subject to alignment (see [0205]-[0206]), including RLHF (see [0167]: “An example fine-tuning approach includes reinforcement learning. Reinforcement learning can be based on user feedback on model performance during use.”). See also [0174]-[0175] and [0243]-[0254] for various examples of text and image outputs, including “image generation” ([0255]) and “question answering” ([0254]).].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of the references combined thus far with the teachings of Chung by implementing the model so that it is trained to generate text and/or images. The motivation for doing so would have been to use a model that is adapted for an image processing task, such as image scaling or generation (see Chung, [0243]: “For example, various different input(s) 2 and output(s) 3 can be used for various different tasks….As another example, machine-learned model(s) 1 can process the image data to generate an upscaled image data output.” [0255]: “In some implementations, the task can be an image generation task.”), and/or for tasks such as question answering (see Chung, [0254]: “In some implementations, the task can be a question answering task….machine-learned model(s) 1 can process input(s) 2 to generate output(s) 3 that represent textual data responsive to the question”).
As to claim 11, the combination of Li and Chung teaches the processor system of claim 10, as set forth above.
Chung further teaches “wherein the output of the ML model comprises text.” [[0050]: “Furthermore, it is to be understood that inputs and/or outputs can be unimodal or multimodal. For example, inputs or outputs can include data from multiple different data modalities (e.g., text, image, audio, video, etc.).” Note that the context is a model that can be subject to alignment (see [0205]-[0206]), including RLHF (see [0167]: “An example fine-tuning approach includes reinforcement learning. Reinforcement learning can be based on user feedback on model performance during use.”). See also [0174]-[0175] and [0243]-[0254] for various examples of text outputs, including “question answering” ([0254]).].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of the references combined thus far to have also arrived at the claimed invention of the instant dependent claim. The motivation for doing so is covered by the one given for Chung in the rejection of the parent claim.
As to claim 12, the combination of Li and Chung teaches the processor system of claim 10, as set forth above.
Chung further teaches “wherein the output of the ML model comprises images.” [[0050]: “Furthermore, it is to be understood that inputs and/or outputs can be unimodal or multimodal. For example, inputs or outputs can include data from multiple different data modalities (e.g., text, image, audio, video, etc.).” Note that the context is a model that can be subject to alignment (see [0205]-[0206]), including RLHF (see [0167]: “An example fine-tuning approach includes reinforcement learning. Reinforcement learning can be based on user feedback on model performance during use.”). See also [0174]-[0175] and [0243]-[0254] for various examples of image outputs, including “image generation” ([0255]).].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of the references combined thus far to have also arrived at the claimed invention of the instant dependent claim. The motivation for doing so is covered by the one given for Chung in the rejection of the parent claim.
As to claim 13, the combination of Li and Chung teaches the processor system of claim 10, as set forth above.
Chung further teaches “wherein the output comprises audio.” [[0050]: “Furthermore, it is to be understood that inputs and/or outputs can be unimodal or multimodal. For example, inputs or outputs can include data from multiple different data modalities (e.g., text, image, audio, video, etc.).” Note that the context is a model that can be subject to alignment (see [0205]-[0206]), including RLHF (see [0167]: “An example fine-tuning approach includes reinforcement learning. Reinforcement learning can be based on user feedback on model performance during use.”). See also [0174]-[0175] and [0243]-[0254] for various examples of image tasks, e.g., [0250]: “For example, the task may be an audio compression task. The input may include audio data and the output may include compressed audio data.”; [0256]: “In some implementations, the task can be an audio generation task.”].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of the references combined thus far, including the above further teachings of Chung, by implementing the model so that its output comprises audio. The motivation for doing so would have been to use a model that is adapted for an audio processing task, such as audio compression and audio generation (Chung, [0250] and [0256] as quoted above.)
As to claim 14, the combination of Li and Chung teaches the processor system of claim 10, as set forth above.
Chung further teaches “wherein the output comprises at least one image.” [[0050]: “Furthermore, it is to be understood that inputs and/or outputs can be unimodal or multimodal. For example, inputs or outputs can include data from multiple different data modalities (e.g., text, image, audio, video, etc.).” Note that the context is a model that can be subject to alignment (see [0205]-[0206]), including RLHF (see [0167]: “An example fine-tuning approach includes reinforcement learning. Reinforcement learning can be based on user feedback on model performance during use.”). See also [0174]-[0175] and [0243]-[0254] for various examples of image outputs, including “image generation” ([0255]).].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of the references combined thus far to have also arrived at the claimed invention of the instant dependent claim. The motivation for doing so is covered by the motivation given for Chung in the rejection of the parent claim.
As to claim 17, the combination of Li and Chung teaches the processor system of claim 10, wherein the processor system is configured to execute RLHF on the ML model using the emotion signal. [Li, § 3.2, paragraph 3: “The TAMER agent will select actions with the value function model to get the most accumulated discounted human reward. Figure 2 shows the diagram of agent learning in the TAMER framework.” The selection of the action is from a machine learning model, namely a reinforcement learning model. See Li, § 3.1, paragraph 3: “As in traditional reinforcement learning (RL), an interactive RL agent learns to make sequential decisions in a task, represented by a policy deciding the action to be taken by the agent in an environmental state.” Li, § 4, paragraph 2: “We recorded the training data including state observation, action, human reward, the time of human reward being given, score, the ending time for each time step, and video data of facial expressions and keypresses on the keyboard for each trainer during training.”]
7. Claims 15-16 are rejected under 35 U.S.C. 103 as being unpatentable over Li in view of Chung, and further in view of Ren.
As to claim 15, the combination of Li and Chung teaches the processor system of claim 10, but does not teach the further limitations of the instant dependent claim
Ren teaches “wherein the processor system is configured to execute RLHF on the ML model using the evaluation input at a first weight responsive to the association being a first association and execute RLHF on the ML model using the evaluation input at a second weight less than the first weight responsive to the association being a second association.” [Abstract: “To determine the example weights, our method performs a meta gradient descent step on the current mini-batch example weights (which are initialized from zero) to minimize the loss on a clean unbiased validation set.” See § 3.1 for a description of setting weights of examples that are noisy, wherein high noise corresponds to low weights. See also FIG. 3 for examples of weight distribution and § 4.3: “As shown in the left figure of Figure 3, our model correctly pushes most noisy images to zero weights.” Note that the limitations of “first association” and “second association” are met by training samples, and in the context of Li such examples are associations because they are associations of the input and output of the CNN-RNN, as discussed above. Furthermore, since Ren teaches that weights are assigned to different samples (analogous to associations), different associations would have its own respective value of its weight. The limitation of “weight less than” is met because there are different values as shown in FIG. 3, including zero and non-zero values, with zero being less than non-zero.]
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of the references combined thus far with the teachings of Ren by implementing the use of weights, so as to arrive at the claimed invention. The motivation for doing so would have been to implement weighting to enable robustness against training set biases (see Ren, § 2, paragraph 3: “One crucial advantage of reweighting examples is robustness against training set bias.”).
As to claim 16, the combination of Li, Chung, and Ren teaches the processor system of claim 15, as set forth above.
Ren further teaches “wherein the second weight is zero such that the evaluation input does not affect RLHF on the ML model.” [FIG. 3 and § 4.3: “As shown in the left figure of Figure 3, our model correctly pushes most noisy images to zero weights.”]
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of the references combined thus far to have also arrived at the claimed invention of the instant dependent claim. The motivation for doing so is covered by the motivation given for Ren in the rejection of the parent dependent claim.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. The following document depicts the state of the art.
Cruz et al., “Multi-modal Feedback for Affordance-driven Interactive Reinforcement Learning,” 2018 International Joint Conference on Neural Networks (IJCNN), Rio de Janeiro, Brazil, 2018, pp. 1-8, doi: 10.1109/IJCNN.2018.8489237 teaches the use of multi-modal feedback and comparing between different feedback types in reinforcement learning, similar to the techniques of the instant application.
Lawrence et al. (US 2023/0142625 A1) teaches the use of emotional state recognize in a machine learning system wherein the emotional state of a user is used to reweigh the feedback information given by said user.
Asada et al. (US 2026/0237386 A1) teaches a system in which a feedback score may indicate a degree to which the observed emotional state of the user 104, such as the user emotional state 306, corresponds to and/or matches the expected emotional state of the user 104. (see [0102]).
Gomez et al. (US 2023/0173683 A1) teaches the use of multiple agents, including an emotional agent in conjunction with a keyboard agent. See [0341]: “In Loop Maze, learning performance of two agents (the keyboard agent and the emotional agent) is evaluated with two indexes including the number of time steps executed by the agent and the total number of feedbacks provided by the human trainer.”
Any inquiry concerning this communication or earlier communications from the examiner should be directed to YAO DAVID HUANG whose telephone number is (571)270-1764. The examiner can normally be reached Monday - Friday 9:00 am - 5:30 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Miranda Huang can be reached at (571) 270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Y.D.H./Examiner, Art Unit 2124 /VINCENT GONZALES/Primary Examiner, Art Unit 2124