Notice of Pre-AIA or AIA Status
This Non-Final communication is in response to application no. 18/802,731 filed on 8/13/2024, which claims priority to PCT/JP2022/012936 filed on 3/22/2022. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-7, and 13-17 are rejected under 35 U.S.C. 103 as being unpatentable by Nakada (US 2019/0244133 A1) in view of Chung (NPL: “Battlesnake Challenge: A multi-agent Reinforcement Learning Playground with Human-in-the-loop” (accessed from applicant IDS)) and Cruz (NPL: “Interactive Explanations: Diagnosis and Repair of Reinforcement Learning Based Agent Behaviors” (Accessed from applicant IDS)).
Regarding claim 1, Nakada teaches:
A human collaborative agent device to perform multi-agent learning, comprising processing circuitry, wherein the processing circuitry performs processes of, ([0185] “A learning apparatus including: … a correcting section configured to correct the reinforcement learning model on a basis of user input to the reinforcement learning model information.”).
acquiring environmental information from environment including the human collaborative agent device; ([0037] “Specifically, in a case where the agent exists in a virtual world such as a simulation, the environment setting section 11 of the PC 10 builds a surrounding environment of the agent in the virtual world on the basis of an operation environment file and the like of the agent. Then, the environment setting section 11 generates an environment map (environment information). The environment map is a GUI (Graphical User Interface) image depicting the surrounding environment.”)
presenting information of a behavior inferred by the human collaborative agent device, information of a reason for the behavior inferred by the human collaborative agent device, or the environmental information acquired from the environment to a user who operates the human collaborative agent device, on a basis of the environmental information acquired from the environment; and ([0039] “On the basis of an initial value of a value function or a movement policy supplied from the receiving section 16, the initialization section 12 initializes a reinforcement learning model that learns the movement policy of the agent. At this time, an initial value of a reward function used for the reinforcement learning model is also set. Here, although a reward function model is assumed to be a linear basis function model that performs a weighted addition on a predetermined reward basis function group selected from a reward basis function group registered in advance, the reward function model is not limited thereto. The initialization section 12 supplies the initialized reinforcement learning model to the learning section 13.”)
acquiring information of the behavior corrected by the user, or information of the reason for the behavior corrected by the user. ([0042] “The receiving section 16 receives input from the user. For example, the receiving section 16 receives the initial value of the value function or the movement policy input from the user, and supplies the initial value of the value function or the movement policy to the initialization section 12. Further, the receiving section 16 receives, from the user who has seen the policy information and the like displayed on the display section 15, input of a movement path as indirect teaching of the movement policy with respect to the policy information, and supplies the movement path to the correcting section 17.”)
Nakada does not teach:
multi-agent learning
presenting information of a reason for the behavior inferred by the human collaborative agent device
acquiring information of the reason for the behavior correct by the user
Chung teaches:
multi-agent learning (Page 1 right column: “To fill in this gap, we introduce the Battlesnake Challenge, an accessible and standardised framework that allows re searchers to effectively train and evaluate their multi-agent RL algorithms with various HILL methods.”)
Nakada and Chung are considered analogous art to the claimed invention because they are in the same field of endeavor being reinforcement learning. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the human in the loop fixes of Nakada with multi-agent reinforcement learning of Chung. One would be motivated to do this as having human-in-the-loop learning can improve policy training.
Cruz teaches:
presenting information of a reason for the behavior inferred by the human collaborative agent device (Page 4 right column “For the explanations, we use the MDP data to create explanations that characterize (E1) the most relevant variables in the current state to make a decision, (E2) the environment’s dynamics, (E3) the short-term goal that the agent is trying to achieve, and (E4) contrasting outcomes between different actions.”)
acquiring information of the reason for the behavior correct by the user (Page 4 left column: “We designed Algorithm 1 to bias the exploration process using the action a_fix and goal g_fix that the user suggests to the bot.” Also see table II on page 7.)
Nakada, Chung and Cruz are considered analogous art to the claimed invention because they are in the same field of endeavor being reinforcement learning. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the human in the loop fixes of Nakada with multi-agent reinforcement learning of Chung with the behavior reasoning and reasoning of the fixes of Cruz. One would be motivated to do this as having human-in-the-loop learning can improve policy training.
Regarding claim 2, Nakada in view of Chung and Cruz teaches Claim 1 as outlined above. Nakada further teaches:
the presenting includes presenting to the user using visual information. ([0040] “The learning section 13 supplies the optimized reinforcement learning model to the correcting section 17 and supplies the learned movement policy to the display control section 14.”)
Regarding claim 3, Nakada in view of Chung and Cruz teaches Claim 1 as outlined above. Cruz further teaches:
the presenting includes presenting to the user using a sentence. (Page 2 right column: “When we use our interactive explanations as input, users aid the agent with natural language templates. [8] propose an approach, similar to ours, to train RL agents through reward shaping by specifying the goal-states with natural language templates. Similarly, [24] map natural language to a set of rules that increase or decrease the probability of selecting specific actions during training in an RL setting. On the other hand, our interactive explanations approach provides users with a natural language template that lets them specify more elements besides goals or preferred action. Additionally, using our interactive explanation to tailor the elements of the underlying RL algorithm allows us to create fixing patches for the main policy in a fast manner, which is vital to have a good user experience.” Also see Table II on page 7 which shows explanations, fixes, and reasoning behind explanations and fixes.)
Regarding claim 4, Nakada in view of Chung and Cruz teaches Claim 3 as outlined above. Cruz further teaches:
the presenting includes presenting the information of the reason for the behavior inferred by the human collaborative agent device, using the sentence. (Page 2 right column: “When we use our interactive explanations as input, users aid the agent with natural language templates. [8] propose an approach, similar to ours, to train RL agents through reward shaping by specifying the goal-states with natural language templates. Similarly, [24] map natural language to a set of rules that increase or decrease the probability of selecting specific actions during training in an RL setting. On the other hand, our interactive explanations approach provides users with a natural language template that lets them specify more elements besides goals or preferred action. Additionally, using our interactive explanation to tailor the elements of the underlying RL algorithm allows us to create fixing patches for the main policy in a fast manner, which is vital to have a good user experience.” Also see Table II on page 7 which shows explanations, fixes, and reasoning behind explanations and fixes.)
Regarding claim 5, Nakada in view of Chung and Cruz teaches Claim 1 as outlined above. Chung further teaches:
the processing circuitry performs a process of interacting with a different agent device from the human collaborative agent device among an agent group with each other in such a way as to reflect the information of the behavior corrected by the user, or the information of the reason for the behavior corrected by the user on the agent group including the human collaborative agent device. (Page 3 left column: “We use the Battlesnake arena to evaluate different combinations of in-training and ad-hoc human guidance by allowing the agents to compete against each other”)
Regarding claim 6, Nakada in view of Chung and Cruz teaches Claim 1 as outlined above. Nakada further teaches:
the processing circuitry performs a process of storing, in a memory, at least one of: the information of the behavior inferred by the human collaborative agent device, the information of the reason for the behavior inferred by the human collaborative agent device, the environmental information, the information of the behavior corrected by the user, and the information of the reason for the behavior corrected by the user. ([0170] describes the memory that the computer uses which stores the various information. “The storage section 408 includes a hard disk, a non-volatile memory, and the like. The communication section 409 includes a network interface and the like. The drive 410 drives a removable medium 411 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.”).
Regarding claim 7, Nakada in view of Chung and Cruz teaches Claim 6 as outlined above. Nakada further teaches:
the processing circuitry performs a process of presenting at least one of: the information of the behavior inferred by the human collaborative agent device, the information of the reason for the behavior inferred by the human collaborative agent device, the environmental information, the information of the behavior corrected by the user, and the information of the reason for the behavior corrected by the user stored in the memory, and information related to interpretation of inference by the human collaborative agent device. ([0039] “On the basis of an initial value of a value function or a movement policy supplied from the receiving section 16, the initialization section 12 initializes a reinforcement learning model that learns the movement policy of the agent. At this time, an initial value of a reward function used for the reinforcement learning model is also set. Here, although a reward function model is assumed to be a linear basis function model that performs a weighted addition on a predetermined reward basis function group selected from a reward basis function group registered in advance, the reward function model is not limited thereto. The initialization section 12 supplies the initialized reinforcement learning model to the learning section 13.” Also, Cruz teaches presenting reasons for behavior or corrected behavior as demonstrated above)
Regarding claim 13, Nakada in view of Chung and Cruz teaches Claim 1 as outlined above. Chung further teaches:
the multi-agent learning is multi-agent reinforcement learning. ([0084] “To fill in this gap, we introduce the Battlesnake Challenge, an accessible and standardised framework that allows re searchers to effectively train and evaluate their multi-agent RL algorithms with various HILL methods..”)
Regarding claim 14, Nakada in view of Chung and Cruz teaches Claim 1 as outlined above. Chung further teaches:
the processing circuitry performs a process of learning and inferring in collaboration with an agent device which is different from the human collaborative agent device among an agent group including the human collaborative agent device. (Page 4 right column “While human rules are instilled with a goal to accelerate the learning procedure, they can also be biased and limiting the snakes’ performance. For instance, rule 3 could lead to snakes focusing too much on food, whereas rule 4 could result in over-aggressive snakes. To this end, we design our platform such that the impact of the heuristics can be controlled and even removed once an agent acquires some basic skills. Such heuristic impact break-down can also be phrased as a curriculum learning method where a logical ordering or hierarchy of simple skills is learnt during training”)
Regarding claim 15, Nakada in view of Hu and Cruz teaches Claim 1 as outlined above.
A system comprising: ([0006] “A learning apparatus according to one aspect of the present disclosure includes: a display control section configured to cause a display section to display reinforcement learning model information regarding a reinforcement learning model; and a correcting section configured to correct the reinforcement learning model on a basis of user input to the reinforcement learning model information.”).
the human collaborative agent device according to claim 1; and (See above).
an agent group to exchange information with the human collaborative agent device. ([0182] “In addition, the present disclosure can also be applied to a learning apparatus that performs reinforcement learning of policies of a plurality of agents (multiple agents) at a time.”)
Regarding claim 16, Nakada teaches:
A multi-agent learning method of a human collaborative agent device to perform multi-agent learning comprising: ([0202] “A learning method including: … a correcting step of the learning apparatus correct the reinforcement learning model on a basis of user input to the reinforcement learning model information.”).
acquiring environmental information from environment including the human collaborative agent device; ([0037] “Specifically, in a case where the agent exists in a virtual world such as a simulation, the environment setting section 11 of the PC 10 builds a surrounding environment of the agent in the virtual world on the basis of an operation environment file and the like of the agent. Then, the environment setting section 11 generates an environment map (environment information). The environment map is a GUI (Graphical User Interface) image depicting the surrounding environment.”)
presenting information of a behavior inferred by the human collaborative agent device, information of a reason for the behavior inferred by the human collaborative agent device, or the environmental information acquired from the environment to a user who operates the human collaborative agent device, on a basis of the environmental information acquired from the environment; and ([0039] “On the basis of an initial value of a value function or a movement policy supplied from the receiving section 16, the initialization section 12 initializes a reinforcement learning model that learns the movement policy of the agent. At this time, an initial value of a reward function used for the reinforcement learning model is also set. Here, although a reward function model is assumed to be a linear basis function model that performs a weighted addition on a predetermined reward basis function group selected from a reward basis function group registered in advance, the reward function model is not limited thereto. The initialization section 12 supplies the initialized reinforcement learning model to the learning section 13.”)
acquiring information of the behavior corrected by the user, or information of the reason for the behavior corrected by the user. ([0042] “The receiving section 16 receives input from the user. For example, the receiving section 16 receives the initial value of the value function or the movement policy input from the user, and supplies the initial value of the value function or the movement policy to the initialization section 12. Further, the receiving section 16 receives, from the user who has seen the policy information and the like displayed on the display section 15, input of a movement path as indirect teaching of the movement policy with respect to the policy information, and supplies the movement path to the correcting section 17.”)
Nakada does not teach:
multi-agent learning
presenting information of a reason for the behavior inferred by the human collaborative agent device
acquiring information of the reason for the behavior correct by the user
Chung teaches:
multi-agent learning (Page 1 right column: “To fill in this gap, we introduce the Battlesnake Challenge, an accessible and standardised framework that allows re searchers to effectively train and evaluate their multi-agent RL algorithms with various HILL methods.”)
Nakada and Chung are considered analogous art to the claimed invention because they are in the same field of endeavor being reinforcement learning. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the human in the loop fixes of Nakada with multi-agent reinforcement learning of Chung. One would be motivated to do this as having human-in-the-loop learning can improve policy training.
Cruz teaches:
presenting information of a reason for the behavior inferred by the human collaborative agent device (Page 4 right column “For the explanations, we use the MDP data to create explanations that characterize (E1) the most relevant variables in the current state to make a decision, (E2) the environment’s dynamics, (E3) the short-term goal that the agent is trying to achieve, and (E4) contrasting outcomes between different actions.”)
acquiring information of the reason for the behavior correct by the user (Page 4 left column: “We designed Algorithm 1 to bias the exploration process using the action a_fix and goal g_fix that the user suggests to the bot.” Also see table II on page 7.)
Nakada, Chung and Cruz are considered analogous art to the claimed invention because they are in the same field of endeavor being reinforcement learning. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the human in the loop fixes of Nakada with multi-agent reinforcement learning of Chung with the behavior reasoning and reasoning of the fixes of Cruz. One would be motivated to do this as having human-in-the-loop learning can improve policy training.
Regarding claim 17, Nakada teaches:
A non-transitory computer readable storage medium storing a program to be executed by processing circuitry of a human collaborative agent device to perform multi-agent learning, the program to enable the processing circuitry to perform processes of, ([0173] “In the computer 400, the program can be installed in the storage section 408 via the input/output interface 405 by attaching the removable medium 411 to the drive 410. Further, the program can be received by the communication section 409 via a wired or wireless transmission medium and installed in the storage section 408. Additionally, the program can be installed in the ROM 402 or the storage section 408 in advance..”).
acquiring environmental information from environment including the human collaborative agent device; ([0037] “Specifically, in a case where the agent exists in a virtual world such as a simulation, the environment setting section 11 of the PC 10 builds a surrounding environment of the agent in the virtual world on the basis of an operation environment file and the like of the agent. Then, the environment setting section 11 generates an environment map (environment information). The environment map is a GUI (Graphical User Interface) image depicting the surrounding environment.”)
presenting information of a behavior inferred by the human collaborative agent device, information of a reason for the behavior inferred by the human collaborative agent device, or the environmental information acquired from the environment to a user who operates the human collaborative agent device, on a basis of the environmental information acquired from the environment; and ([0039] “On the basis of an initial value of a value function or a movement policy supplied from the receiving section 16, the initialization section 12 initializes a reinforcement learning model that learns the movement policy of the agent. At this time, an initial value of a reward function used for the reinforcement learning model is also set. Here, although a reward function model is assumed to be a linear basis function model that performs a weighted addition on a predetermined reward basis function group selected from a reward basis function group registered in advance, the reward function model is not limited thereto. The initialization section 12 supplies the initialized reinforcement learning model to the learning section 13.”)
acquiring information of the behavior corrected by the user, or information of the reason for the behavior corrected by the user. ([0042] “The receiving section 16 receives input from the user. For example, the receiving section 16 receives the initial value of the value function or the movement policy input from the user, and supplies the initial value of the value function or the movement policy to the initialization section 12. Further, the receiving section 16 receives, from the user who has seen the policy information and the like displayed on the display section 15, input of a movement path as indirect teaching of the movement policy with respect to the policy information, and supplies the movement path to the correcting section 17.”)
Nakada does not teach:
multi-agent learning
presenting information of a reason for the behavior inferred by the human collaborative agent device
acquiring information of the reason for the behavior correct by the user
Chung teaches:
multi-agent learning (Page 1 right column: “To fill in this gap, we introduce the Battlesnake Challenge, an accessible and standardised framework that allows re searchers to effectively train and evaluate their multi-agent RL algorithms with various HILL methods.”)
Nakada and Chung are considered analogous art to the claimed invention because they are in the same field of endeavor being reinforcement learning. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the human in the loop fixes of Nakada with multi-agent reinforcement learning of Chung. One would be motivated to do this as having human-in-the-loop learning can improve policy training.
Cruz teaches:
presenting information of a reason for the behavior inferred by the human collaborative agent device (Page 4 right column “For the explanations, we use the MDP data to create explanations that characterize (E1) the most relevant variables in the current state to make a decision, (E2) the environment’s dynamics, (E3) the short-term goal that the agent is trying to achieve, and (E4) contrasting outcomes between different actions.”)
acquiring information of the reason for the behavior correct by the user (Page 4 left column: “We designed Algorithm 1 to bias the exploration process using the action a_fix and goal g_fix that the user suggests to the bot.” Also see table II on page 7.)
Nakada, Chung and Cruz are considered analogous art to the claimed invention because they are in the same field of endeavor being reinforcement learning. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the human in the loop fixes of Nakada with multi-agent reinforcement learning of Chung with the behavior reasoning and reasoning of the fixes of Cruz. One would be motivated to do this as having human-in-the-loop learning can improve policy training.
Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable by Nakada in view of Chung, Cruz and Howell (US 2022/0230540 A1).
Regarding claim 8, Nakada in view of Chung and Cruz teaches Claim 6 as outlined above. None of them teaches the limitations of claim 8, However, Howell does:
the memory is shared by the agent group including the human collaborative agent device. ([0078] “Some embodiments may use aspects of collaborative reinforcement learning, in which the agents to some extent share memories.”)
Nakada, Chung, Cruz and Howell are considered analogous art to the claimed invention because they are in the same field of endeavor being reinforcement learning. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the human in the loop fixes of Nakada with multi-agent reinforcement learning of Chung with the behavior reasoning and reasoning of the fixes of Cruz with the agent shared memory of Howell. One would be motivated to do this so the agents can easily access each other’s knowledge.
Claims 9-12 are rejected under 35 U.S.C. 103 as being unpatentable by Nakada in view of Chung, Cruz and Lyu (US 2024/0160945 A1).
Regarding claim 9, Nakada in view of Chung and Cruz teaches Claim 1 as outlined above. None of them teaches the limitations of claim 9. However, Lyu does:
the processing circuitry performs a process of learning autonomously without an operation by the user on a basis of the environmental information acquired from the environment. ([0050] “The method comprises obtaining parameters of a trained deep reinforcement learning model trained for autonomous control of the machine. To apply the model, the method comprises: receiving state information indicative of an environment of the machine; determining, by the trained deep reinforcement learning model in response to input of the state information, an agent action indicative of a control signal; and transmitting the control signal to the machine.”)
Nakada, Chung, Cruz and Lyu are considered analogous art to the claimed invention because they are in the same field of endeavor being reinforcement learning. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the human in the loop fixes of Nakada with multi-agent reinforcement learning of Chung with the behavior reasoning and reasoning of the fixes of Cruz with the human/autonomous switching learning capability’s of Lyu. One would be motivated to do this so the agents are not always reliant on the humans for training.
Regarding claim 10, Nakada in view of Chung, Cruz and Lyu teaches Claim 9 as outlined above. Lyu further teaches:
the processing circuitry performs a process of switching from autonomous learning to human collaborative learning on a basis of an operation by the user. ([0059] “To realize the human-in-the-loop framework within the reinforcement learning algorithm, the present disclosure combines LfD and LfI into a uniform architecture where humans can decide when to intervene and override the original policy action and provide their real-time actions as demonstrations. Thus, an online switch mechanism between agent exploration and human control is designed.”)
Regarding claim 11, Nakada in view of Chung and Cruz teaches Claim 1 as outlined above. None of them teaches the limitations of claim 11. However, Lyu does:
the human collaborative agent device has a different role from that of an agent device which is different from the human collaborative agent device among an agent group including the human collaborative agent device. ([0073] “Although the proposed objective function of the policy network 202 looks similar to the control authority transfer mechanism of real-time human guidance shown in Eq. (2), the principles of these two stages, namely, real-time human intervention and off-policy learning, are different in the proposed method.”)
Regarding claim 12, Nakada in view of Chung and Cruz teaches Claim 1 as outlined above. None of them teaches the limitations of claim 12. However, Lyu does:
the processing circuitry performs a different process of learning from that of an agent device which is different from the human collaborative agent device among an agent group including the human collaborative agent device. ([0086] “FIGS. 5a-5d illustrate the improved training performance of the proposed Hug-DRL method. In particular, FIGS. 5a-5d show the results of the initial training performance under four different methods”)
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DANIEL P GRUSZKA whose telephone number is (571)272-5259. The examiner can normally be reached M-F 9:00 AM - 6:00 PM ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Li Zhen can be reached at (571) 272-3768. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DANIEL GRUSZKA/Examiner, Art Unit 2121
/Li B. Zhen/Supervisory Patent Examiner, Art Unit 2121