Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: “estimation unit”, “self region determination unit”, and “action generation unit” in claims 1-12.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 112
Claims 1-12 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
The claims cite terms such as “self-region” . This term is not particularly pointed out or distinctly claimed. It’s not clear from the claims what a self region is. All that is known about it is that it is based on data that may indicate what a partner or control system expects. Language the clarifies what the self region is is necessary in order to have the claim distinctly claimed. In light of this, the examiner is treating the self region as the expectation or desire itself. The scope of the claim is not clear.
The term “action information for approaching the observation that the control target expects” is not particularly pointed out or distinctly claimed. It is not clear if this is a goal to align with the expectation, or for gathering more observational data. The scope of the claim is not clear.
The claims cite an “Action Form” . It is not particularly pointed out and distinctly claimed what this is. As the examiner understands it, it could be any action or observation from a user, but the scope is not clear. Thus the claim is not particularly pointed out or distinctly claimed.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim 12 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claim does not fall within at least one of the four categories of patent eligible subject matter because it refers to a software program. The claim is directed to software per se. Hence, the claim is patent ineligible.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 3-7, 9, and 11-12 are rejected under 35 U.S.C. 103 as being unpatentable over Sawada et al (Human-Robot Kinaesthetic Interactions Based on the Free-Energy Principle, please refer to attached NPL) in light of Dani et al (US Pub 2018/0032868 A1), hereafter known as Dani.
For Claim 1, Sawada teaches A method comprising:
configured to estimate a self-region of a partner based on preference information indicating an observation that the partner is estimated to expect for a control target and the partner, the preference information being derived based on a predetermined principle and observation sensor data acquired from the partner and the control target, the partner being a person or an autonomous system including a robot, and the control target being an autonomous system including a robot; (Fig. 2
Page 1, Paragraph 2
In this regard, the current study investigated human-robot kinaesthetic interaction by applying system neuroscience theory, the free energy principle, proposed by Friston [8] which is consonant with enactivism 19, 10]. Let us consider a situation in which 3 robet and 3 human dance, holding each other with both hands, executing memorized dance patterns. If the robot initiates a particular pattern from memory with strong intention, the human counterpart might follow il without resisting because of the strong counter force. On the other hand. if the robot generates $ pattern without strong
intention. the human counterpart might be able to shift to a different pattern without experiencing strong counter force. Let us consider another situation. If the human counterpart attempts to induce a movement pattern that is familiar to the robot, such guidance should proceed easily without strong counter force, since the robot can infer the intended pattern immediately and can move as anticipated. On the other hand, if the human counterpari attempts to induce a movement pattern unfamiliar to the robot, such guidance should experience strong counter force, since the movement is neither inferential nor predictable for the robot.
The robot can infer the intention of the human and adjust its controls immediately. It should be noted that the specification allows for a self region to include intentions, desires, expectations, or physical locations.
Page 3
2.1 Free Energy Principle
As mentioned in the Introduction. the current study used PV-RNN, based on the free energy principle, as the basic model. The free energy principle assumes that brains learn as well as inferred for given observations. by minimizing free energy. defined in Eq. 1.
I:
(1)
po(X|z) is the likelihood of the sensory observation X, given the probabilistic latent variables 2 which is parameterized by A. is the inference model parameterised by Q. As shown in Eq. i, free energy consists of two terms. The first term indicates the complexity, which is the divergence between the approximate posterior probability distribution and prior probability distribution for the latent variable 2. and the second term is the accuracy in predicting the sensory observation.
2.2 PV-RNN implementation
PV-RNN is operated in two distinct phases. it training phase and an interaction phase. In the training phase, we prepared a dataset (joint angle sequences of the robot) and trained the PV-RNN model using it. In the interaction phase, the robot and the human experimenter interact physically, such that the PV-RNN attempts to drive the robot's arms by predicting next-time-step target joint angles while it infers latent variables using actual joint angle readings. Each operation phase is explained in detail below.
configured to determine a self-region of the control target based on preference information indicating an observation that the control target expects for the partner and the control target, which is derived based on the predetermined principle, using the self- region of the partner and an intention of the control target as a goal to be achieved by cooperation between the partner and the control target; and (Fig. 2
Fig. 2
Page 1, Paragraph 2
In this regard, the current study investigated human-robot kinaesthetic interaction by applying system neuroscience theory, the free energy principle, proposed by Friston [8] which is consonant with enactivism 19, 10]. Let us consider a situation in which 3 robet and 3 human dance, holding each other with both hands, executing memorized dance patterns. If the robot initiates a particular pattern from memory with strong intention, the human counterpart might follow il without resisting because of the strong counter force. On the other hand. if the robot generates $ pattern without strong
intention. the human counterpart might be able to shift to a different pattern without experiencing strong counter force. Let us consider another situation. If the human counterpart attempts to induce a movement pattern that is familiar to the robot, such guidance should proceed easily without strong counter force, since the robot can infer the intended pattern immediately and can move as anticipated. On the other hand, if the human counterpari attempts to induce a movement pattern unfamiliar to the robot, such guidance should experience strong counter force, since the movement is neither inferential nor predictable for the robot.
The robot can infer the intention of the human and adjust its controls immediately. It should be noted that the specification allows for a self region to include intentions, desires, expectations, or physical locations.
Page 3
2.1 Free Energy Principle
As mentioned in the Introduction. the current study used PV-RNN, based on the free energy principle, as the basic model. The free energy principle assumes that brains learn as well as inferred for given observations. by minimizing free energy. defined in Eq. 1.
I:
(1)
po(X|z) is the likelihood of the sensory observation X, given the probabilistic latent variables 2 which is parameterized by A. is the inference model parameterised by Q. As shown in Eq. i, free energy consists of two terms. The first term indicates the complexity, which is the divergence between the approximate posterior probability distribution and prior probability distribution for the latent variable 2. and the second term is the accuracy in predicting the sensory observation.
2.2 PV-RNN implementation
PV-RNN is operated in two distinct phases. it training phase and an interaction phase. In the training phase, we prepared a dataset (joint angle sequences of the robot) and trained the PV-RNN model using it. In the interaction phase, the robot and the human experimenter interact physically, such that the PV-RNN attempts to drive the robot's arms by predicting next-time-step target joint angles while it infers latent variables using actual joint angle readings. Each operation phase is explained in detail below.
configured to generate action information for controlling an action of the control target based on the self-region of the control target, and to control a driving device for the control target using the action information, thereby controlling an operation of the control target.
(Fig. 2
Page 7
2.3 The employed robot controller
The current study uses a humanoid robot, Torobo ] to conduct human-robot kinesthetic interaction experiments. Torobo is equipped with a built-in force-feedback controller that enables humans to back-drive joint angles with subtle force. In the Torobo control system, when a human exerts certain torque 90 joints by pushing or pulling the limbs of Torobo, the exerted torque can be estimated by subtracting the torque inferred as necessary to account for the current static state as well as a dynamic state of the robot from the actual torque measured in the joints. By computing the next time-step joint target positions by adding the current positions with the estimated exerted torque multiplied by a constant gain and feeding them in the PID controller, Torobo's limbs move by following the force exerted on them by the human. This force-feedback controller is integrated with the PV-RNN, which generates the next time-step target joint angles (Fig.2).
Page 10,
4) Human–Robot Interaction: After the training phase, we conducted two types of human–robot interaction experiments using trained PV-RNNs. In these experiments, while Torobo was generating movement pattern transitions successively based on the training, the human experimenter attempted to induce various movement pattern transitions by grasping both arms of Torobo and exerting force on them. These movement pattern transitions included trained transitions (AB, AC, BD, CD, and DA) and untrained transitions (AD, DB, DC, BA, and CA) in which AB, for example, dictates that A pattern is forced to transit to the B pattern. Each transition from one pattern to another requires some guiding force by a human experimenter, even for trained transitions, since ongoing pat terns tend to repeat another cycle with high probability, more than 85% for all patterns. This means that there exist some conflicts between movement trajectories intended by the robot and the human experimentor.)
Sawada does not explicitly teach a control device
a self-region estimation unit
a self-region determination unit
an action generation unit
Dani, however, does teach units that perform the tasks. [0018] The learning system 100 may include a training unit 150 and an online learning unit 160, both of which may be hardware, software, or a combination of hardware and software. Generally, the training unit 150 may train a neural network (NN) 170 based on training data derived from one or more user demonstrations of reaching tasks; and the online learning unit 160 may predict intentions for new reaching tasks and may also update the NN 170 based on test data derived from these new reaching tasks.
[0019] In some embodiments, the learning system 100 may be implemented on a computer system and may execute a training algorithm and an online learning algorithm, executed respectively by the training unit 150 and the online learning unit 160, as described below. Further, the learning system 100 may be in communication with a camera 180, such as Microsoft Kinect® for Windows® or some other camera capable of capturing data in three dimensions, which may record demonstrations of reaching tasks as well as new reaching tasks after initial training of the NN 170.
Therefore, it would be obvious to one of ordinary skill in the art prior to the effective filing date to modify Sawada in light of Dani such that there is a self-region estimation unit
a self-region determination unit
an action generation unit
It would be obvious to modify Sawada in light of Dani in this way because units, processors, and CPUs are expected to be successful at computational tasks, and separating tasks amongst a number of different units allows the processors and memories to be specialized for specific tasks, which would be expected to increase efficiency and output.
Dani, however, does teach a control device. ([0071] In some embodiments, as shown in FIG. 5, the computer system 500 includes a processor 505, memory 510 coupled to a memory controller 515, and one or more input devices 545 and/or output devices 540, such as peripherals, that are communicatively coupled via a local I/O controller 535. These devices 540 and 545 may include, for example, a printer, a scanner, a microphone, and the like. Input devices such as a conventional keyboard 550 and mouse 555 may be coupled to the I/O controller 535. The I/O controller 535 may be, for example, one or more buses or other wired or wireless connections, as are known in the art. The I/O controller 535 may have additional elements, which are omitted for simplicity, such as controllers, buffers (caches), drivers, repeaters, and receivers, to enable communications.)
Therefore, it would be obvious to one of ordinary skill in the art prior to the effective filing date to modify Sawada in light of Dani such that there is a control device that carries out the steps because robotics controls are carried out by controllers, processors, and memory devices. It would be necessary that there be some sort of control device to carry out Sawada’s method.
For Claim 3, Sawada teaches The control device according to claim 1, wherein the action generation unit is configured to generate the action information for approaching the observation that the control target expects for the partner, based on the self-region of the control target, thereby controlling an operation of the control target based on the action information. (Fig. 1
Page 7
The overall control diagram is shown in Fig. 2, which includes a human experimenter interacting with Torobo, the PV-RNN target joint angle generator, the inverse model, and the PID joint controller. The inverse model and PID controller were developed by the manufacturer of Torobo.
Fig. 2
Page 1, Paragraph 2
In this regard, the current study investigated human-robot kinaesthetic interaction by applying system neuroscience theory, the free energy principle, proposed by Friston [8] which is consonant with enactivism 19, 10]. Let us consider a situation in which 3 robet and 3 human dance, holding each other with both hands, executing memorized dance patterns. If the robot initiates a particular pattern from memory with strong intention, the human counterpart might follow il without resisting because of the strong counter force. On the other hand. if the robot generates $ pattern without strong
intention. the human counterpart might be able to shift to a different pattern without experiencing strong counter force. Let us consider another situation. If the human counterpart attempts to induce a movement pattern that is familiar to the robot, such guidance should proceed easily without strong counter force, since the robot can infer the intended pattern immediately and can move as anticipated. On the other hand, if the human counterpari attempts to induce a movement pattern unfamiliar to the robot, such guidance should experience strong counter force, since the movement is neither inferential nor predictable for the robot.
The robot can infer the intention of the human and adjust its controls immediately. It should be noted that the specification allows for a self region to include intentions, desires, expectations, or physical locations.
Page 3
2.1 Free Energy Principle
As mentioned in the Introduction. the current study used PV-RNN, based on the free energy principle, as the basic model. The free energy principle assumes that brains learn as well as inferred for given observations. by minimizing free energy. defined in Eq. 1.
I:
(1)
po(X|z) is the likelihood of the sensory observation X, given the probabilistic latent variables 2 which is parameterized by A. is the inference model parameterised by Q. As shown in Eq. i, free energy consists of two terms. The first term indicates the complexity, which is the divergence between the approximate posterior probability distribution and prior probability distribution for the latent variable 2. and the second term is the accuracy in predicting the sensory observation.
2.2 PV-RNN implementation
PV-RNN is operated in two distinct phases. it training phase and an interaction phase. In the training phase, we prepared a dataset (joint angle sequences of the robot) and trained the PV-RNN model using it. In the interaction phase, the robot and the human experimenter interact physically, such that the PV-RNN attempts to drive the robot's arms by predicting next-time-step target joint angles while it infers latent variables using actual joint angle readings. Each operation phase is explained in detail below.
For Claim 4, Sawada teaches The control device according to claim 1, wherein the action generation unit is configured to generate the action information for confirming whether or not the estimated self- region of the partner matches a preference of an observation actually considered by the partner, thereby controlling an operation of the control target based on the action information. (Page 10 Section 3.2.1.
Experiment-1: First, we examined time-development of essential values during movement pattern transitions induced by the experimenter for each case with a different meta prior setting. Fig. 5 shows an example snapshot of future prediction and past reflection, which shifted every 7 time-steps of the current time during the transition AB performed under different settings of the meta-prior, wi = 0.01,0.05,0.1. Each snapshot shows one of the observed joint angles θ (dotted blue) and its prediction ¯θ (blue) in the top row, the KL divergences between the approximate posterior and the prior in the layer 1 (orange) and in the layer 2 (dark orange) in the second row, the prediction error (green) in the third row, and the excess torque (black) in the bottom row. The gray area represents the past window where the approximate posterior in terms of adaptive variables Aμ t and Aσ t is updated. We provide two supplementary videos for experiment-1 showing the interaction between the experimenter and Torobo, as well as network dynamics in the case with the meta-prior wi set to 0.01 (video-link1) and 0.1 (video-link2). Sequences of time-shifted snapshots in Fig. 4, show that excess torque appears first followed by rises in the prediction error (negative log-likelihood) and the KL-divergence. Later, the predicted joint angle pattern shifts from pattern A to B while the gap between the observed joint angle and the reconstructed joint angle remains in the past window. This is the same for all the three cases with different wi settings.)
For Claim 5, Sawada teaches The control device according to claim 1, wherein the predetermined principle is a free energy principle. ((Fig. 2
Fig. 2
Page 1, Paragraph 2
In this regard, the current study investigated human-robot kinaesthetic interaction by applying system neuroscience theory, the free energy principle, proposed by Friston [8] which is consonant with enactivism 19, 10]. Let us consider a situation in which 3 robet and 3 human dance, holding each other with both hands, executing memorized dance patterns. If the robot initiates a particular pattern from memory with strong intention, the human counterpart might follow il without resisting because of the strong counter force. On the other hand. if the robot generates $ pattern without strong
intention. the human counterpart might be able to shift to a different pattern without experiencing strong counter force. Let us consider another situation. If the human counterpart attempts to induce a movement pattern that is familiar to the robot, such guidance should proceed easily without strong counter force, since the robot can infer the intended pattern immediately and can move as anticipated. On the other hand, if the human counterpari attempts to induce a movement pattern unfamiliar to the robot, such guidance should experience strong counter force, since the movement is neither inferential nor predictable for the robot.
The robot can infer the intention of the human and adjust its controls immediately. It should be noted that the specification allows for a self region to include intentions, desires, expectations, or physical locations.
Page 3
2.1 Free Energy Principle
As mentioned in the Introduction. the current study used PV-RNN, based on the free energy principle, as the basic model. The free energy principle assumes that brains learn as well as inferred for given observations. by minimizing free energy. defined in Eq. 1.
I:
(1)
po(X|z) is the likelihood of the sensory observation X, given the probabilistic latent variables 2 which is parameterized by A. is the inference model parameterised by Q. As shown in Eq. i, free energy consists of two terms. The first term indicates the complexity, which is the divergence between the approximate posterior probability distribution and prior probability distribution for the latent variable 2. and the second term is the accuracy in predicting the sensory observation.
2.2 PV-RNN implementation
PV-RNN is operated in two distinct phases. it training phase and an interaction phase. In the training phase, we prepared a dataset (joint angle sequences of the robot) and trained the PV-RNN model using it. In the interaction phase, the robot and the human experimenter interact physically, such that the PV-RNN attempts to drive the robot's arms by predicting next-time-step target joint angles while it infers latent variables using actual joint angle readings. Each operation phase is explained in detail below.
For Claim 6, Sawada teaches The control device according to claim 5, wherein
the self-region determination unit is configured to determine the self-region of the control target using a generative model of the free energy principle. ((Fig. 2
Fig. 2
Page 1, Paragraph 2
In this regard, the current study investigated human-robot kinaesthetic interaction by applying system neuroscience theory, the free energy principle, proposed by Friston [8] which is consonant with enactivism 19, 10]. Let us consider a situation in which 3 robet and 3 human dance, holding each other with both hands, executing memorized dance patterns. If the robot initiates a particular pattern from memory with strong intention, the human counterpart might follow il without resisting because of the strong counter force. On the other hand. if the robot generates $ pattern without strong
intention. the human counterpart might be able to shift to a different pattern without experiencing strong counter force. Let us consider another situation. If the human counterpart attempts to induce a movement pattern that is familiar to the robot, such guidance should proceed easily without strong counter force, since the robot can infer the intended pattern immediately and can move as anticipated. On the other hand, if the human counterpari attempts to induce a movement pattern unfamiliar to the robot, such guidance should experience strong counter force, since the movement is neither inferential nor predictable for the robot.
The robot can infer the intention of the human and adjust its controls immediately. It should be noted that the specification allows for a self region to include intentions, desires, expectations, or physical locations.
Page 3
2.1 Free Energy Principle
As mentioned in the Introduction. the current study used PV-RNN, based on the free energy principle, as the basic model. The free energy principle assumes that brains learn as well as inferred for given observations. by minimizing free energy. defined in Eq. 1.
I:
(1)
po(X|z) is the likelihood of the sensory observation X, given the probabilistic latent variables 2 which is parameterized by A. is the inference model parameterised by Q. As shown in Eq. i, free energy consists of two terms. The first term indicates the complexity, which is the divergence between the approximate posterior probability distribution and prior probability distribution for the latent variable 2. and the second term is the accuracy in predicting the sensory observation.
2.2 PV-RNN implementation
PV-RNN is operated in two distinct phases. it training phase and an interaction phase. In the training phase, we prepared a dataset (joint angle sequences of the robot) and trained the PV-RNN model using it. In the interaction phase, the robot and the human experimenter interact physically, such that the PV-RNN attempts to drive the robot's arms by predicting next-time-step target joint angles while it infers latent variables using actual joint angle readings. Each operation phase is explained in detail below.
Sawada does not teach the self-region estimation unit is configured to estimate the self-region of the partner using a generative model of the free energy principle, and
Dani, however, does teach self-region estimation unit is configured to estimate the self-region of the partner using a neural network, ([0040] In some embodiments, the learning system 100 may use an approximate EM algorithm with modifications for handling state transition models trained using the NN 170. Using the fact that E.sub.X.sub.T{log [p(Z.sub.T|g)]|X.sub.Tĝ.sub.t}=log p(Z.sub.T|g), the log-likelihood defined in Formula 4 may be decomposed as log p(Z.sub.T|g)=E.sub.X.sub.T{log [p(Z.sub.T, X.sub.T|g)]|Z.sub.Tĝ.sub.t}−E.sub.X.sub.T{log [p(X.sub.T|Z.sub.T, g)]|Z.sub.Tĝ.sub.t} log p(Z.sub.T|g)=Q(g, ĝ.sub.t)−H(g, ĝ.sub.t) (hereinafter “Formula 6” and “Formula 7,” respectively).
[0041] In the above Formulas 6 and 7, where E.sub.X.sub.T is the expectation operator; ĝ.sub.t is an estimate of the intention g at time t; Q(g, ĝ.sub.t)=E.sub.X.sub.T{log [p(Z.sub.T, X.sub.T|g)]|Z.sub.Tĝ.sub.t} is an expected value of the complete data log-likelihood given the various measurements and intentions of the training data; and H(g, ĝ.sub.t)=E.sub.X.sub.T{log [p(X.sub.T|Z.sub.T, g)]|Z.sub.Tĝ.sub.t}. It can be shown using Jensen's inequality that H(g, ĝ.sub.t)≦H(ĝ.sub.t, ĝ.sub.t. Thus, to iteratively increase the log-likelihood, it may be required to choose g such that Q(g, ĝ.sub.t)≧Q(ĝ.sub.t, ĝ.sub.t).
[0042] As described in more detail below, the learning system 100 may compute the auxiliary function Q(g, ĝ.sub.t) given the observations Z.sub.T and a current estimate of the intention ĝ.sub.t. The learning system 100 may also compute the next intention estimate ĝ.sub.t+1, after the current intention estimate, by finding a value of g that maximizes Q(g, ĝ.sub.t).)
Therefore, it would be obvious to one of ordinary skill in the art prior to the effective filing date to modify Sawada in light of Dani such that the self-region estimation unit is configured to estimate the self-region of the partner using a generative model of the free energy principle, and
It would be obvious to one or ordinary skill in the art prior to the effective filing date to modify Sawada in light of Dani in this way because the free energy principle would be expected to be useful for determining the expectations for one agent. It would therefore also be expected to be useful for the other agent as well. Dani’s teaching of using particular models to determine partner’s intentions would make it obvious to use the free energy model of Sawada for the partner as well.
For Claim 7, Sawada teaches The control device according to claim 6, wherein the self-region estimation unit is configured to update the generative model based on a result of an analysis using the observation sensor data. (
Page 10, Section 3.2.1.
Experiment-1: First, we examined time-development of essential values during movement pattern transitions induced by the experimenter for each case with a different meta prior setting. Fig. 5 shows an example snapshot of future prediction and past reflection, which shifted every 7 time-steps of the current time during the transition AB performed under different settings of the meta-prior, wi = 0.01,0.05,0.1. Each snapshot shows one of the observed joint angles θ (dotted blue) and its prediction ¯θ (blue) in the top row, the KL divergences between the approximate posterior and the prior in the layer 1 (orange) and in the layer 2 (dark orange) in the second row, the prediction error (green) in the third row, and the excess torque (black) in the bottom row. The gray area represents the past window where the approximate posterior in terms of adaptive variables Aμ t and Aσ t is updated. We provide two supplementary videos for experiment-1 showing the interaction between the experimenter and Torobo, as well as network dynamics in the case with the meta-prior wi set to 0.01 (video-link1) and 0.1 (video-link2). Sequences of time-shifted snapshots in Fig. 4, show that excess torque appears first followed by rises in the prediction error (negative log-likelihood) and the KL-divergence. Later, the predicted joint angle pattern shifts from pattern A to B while the gap between the observed joint angle and the reconstructed joint angle remains in the past window. This is the same for all the three cases with different wi settings.
.)
For Claim 9, Sawada teaches The control device according to claim 1
wherein
the action generation unit is configured to
identify an action form of the partner based on the observation sensor data, and (Fig. 2
Page 7
2.3 The employed robot controller
The current study uses a humanoid robot, Torobo ] to conduct human-robot kinesthetic interaction experiments. Torobo is equipped with a built-in force-feedback controller that enables humans to back-drive joint angles with subtle force. In the Torobo control system, when a human exerts certain torque 90 joints by pushing or pulling the limbs of Torobo, the exerted torque can be estimated by subtracting the torque inferred as necessary to account for the current static state as well as a dynamic state of the robot from the actual torque measured in the joints. By computing the next time-step joint target positions by adding the current positions with the estimated exerted torque multiplied by a constant gain and feeding them in the PID controller, Torobo's limbs move by following the force exerted on them by the human. This force-feedback controller is integrated with the PV-RNN, which generates the next time-step target joint angles (Fig.2).
Page 10,
4) Human–Robot Interaction: After the training phase, we conducted two types of human–robot interaction experiments using trained PV-RNNs. In these experiments, while Torobo was generating movement pattern transitions successively based on the training, the human experimenter attempted to induce various movement pattern transitions by grasping both arms of Torobo and exerting force on them. These movement pattern transitions included trained transitions (AB, AC, BD, CD, and DA) and untrained transitions (AD, DB, DC, BA, and CA) in which AB, for example, dictates that A pattern is forced to transit to the B pattern. Each transition from one pattern to another requires some guiding force by a human experimenter, even for trained transitions, since ongoing pat terns tend to repeat another cycle with high probability, more than 85% for all patterns. This means that there exist some conflicts between movement trajectories intended by the robot and the human experimentor.)
to generate action information based on the self-region of the control target, such that the action information corresponds to the identified action form. (Fig. 2
Page 7
2.3 The employed robot controller
The current study uses a humanoid robot, Torobo ] to conduct human-robot kinesthetic interaction experiments. Torobo is equipped with a built-in force-feedback controller that enables humans to back-drive joint angles with subtle force. In the Torobo control system, when a human exerts certain torque 90 joints by pushing or pulling the limbs of Torobo, the exerted torque can be estimated by subtracting the torque inferred as necessary to account for the current static state as well as a dynamic state of the robot from the actual torque measured in the joints. By computing the next time-step joint target positions by adding the current positions with the estimated exerted torque multiplied by a constant gain and feeding them in the PID controller, Torobo's limbs move by following the force exerted on them by the human. This force-feedback controller is integrated with the PV-RNN, which generates the next time-step target joint angles (Fig.2).
Page 10,
4) Human–Robot Interaction: After the training phase, we conducted two types of human–robot interaction experiments using trained PV-RNNs. In these experiments, while Torobo was generating movement pattern transitions successively based on the training, the human experimenter attempted to induce various movement pattern transitions by grasping both arms of Torobo and exerting force on them. These movement pattern transitions included trained transitions (AB, AC, BD, CD, and DA) and untrained transitions (AD, DB, DC, BA, and CA) in which AB, for example, dictates that A pattern is forced to transit to the B pattern. Each transition from one pattern to another requires some guiding force by a human experimenter, even for trained transitions, since ongoing pat terns tend to repeat another cycle with high probability, more than 85% for all patterns. This means that there exist some conflicts between movement trajectories intended by the robot and the human experimentor.)
For Claim 11, Sawada teaches A control method performed
a self-region estimation step of estimating a self- region of a partner based on preference information indicating an observation that the partner is estimated to expect for a control target and the partner, the preference information being derived based on a predetermined principle and observation sensor data acquired from the partner and the control target, the partner being a person or an autonomous system including a robot, and the control target being an autonomous system including a robot; ((Fig. 1
Fig. 2
Page 1, Paragraph 2
In this regard, the current study investigated human-robot kinaesthetic interaction by applying system neuroscience theory, the free energy principle, proposed by Friston [8] which is consonant with enactivism 19, 10]. Let us consider a situation in which 3 robet and 3 human dance, holding each other with both hands, executing memorized dance patterns. If the robot initiates a particular pattern from memory with strong intention, the human counterpart might follow il without resisting because of the strong counter force. On the other hand. if the robot generates $ pattern without strong
intention. the human counterpart might be able to shift to a different pattern without experiencing strong counter force. Let us consider another situation. If the human counterpart attempts to induce a movement pattern that is familiar to the robot, such guidance should proceed easily without strong counter force, since the robot can infer the intended pattern immediately and can move as anticipated. On the other hand, if the human counterpari attempts to induce a movement pattern unfamiliar to the robot, such guidance should experience strong counter force, since the movement is neither inferential nor predictable for the robot.
The robot can infer the intention of the human and adjust its controls immediately. It should be noted that the specification allows for a self region to include intentions, desires, expectations, or physical locations.
Page 3
2.1 Free Energy Principle
As mentioned in the Introduction. the current study used PV-RNN, based on the free energy principle, as the basic model. The free energy principle assumes that brains learn as well as inferred for given observations. by minimizing free energy. defined in Eq. 1.
I:
(1)
po(X|z) is the likelihood of the sensory observation X, given the probabilistic latent variables 2 which is parameterized by A. is the inference model parameterised by Q. As shown in Eq. i, free energy consists of two terms. The first term indicates the complexity, which is the divergence between the approximate posterior probability distribution and prior probability distribution for the latent variable 2. and the second term is the accuracy in predicting the sensory observation.
2.2 PV-RNN implementation
PV-RNN is operated in two distinct phases. it training phase and an interaction phase. In the training phase, we prepared a dataset (joint angle sequences of the robot) and trained the PV-RNN model using it. In the interaction phase, the robot and the human experimenter interact physically, such that the PV-RNN attempts to drive the robot's arms by predicting next-time-step target joint angles while it infers latent variables using actual joint angle readings. Each operation phase is explained in detail below.
a self-region determination step of determining a self- region of the control target based on preference information indicating an observation that the control target expects for the partner and the control target, which is derived based on the predetermined principle, using the self-region of the partner and an intention of the control target as a goal to be achieved by cooperation between the partner and the control target; and Fig. 2
Page 1, Paragraph 2
In this regard, the current study investigated human-robot kinaesthetic interaction by applying system neuroscience theory, the free energy principle, proposed by Friston [8] which is consonant with enactivism 19, 10]. Let us consider a situation in which 3 robet and 3 human dance, holding each other with both hands, executing memorized dance patterns. If the robot initiates a particular pattern from memory with strong intention, the human counterpart might follow il without resisting because of the strong counter force. On the other hand. if the robot generates $ pattern without strong
intention. the human counterpart might be able to shift to a different pattern without experiencing strong counter force. Let us consider another situation. If the human counterpart attempts to induce a movement pattern that is familiar to the robot, such guidance should proceed easily without strong counter force, since the robot can infer the intended pattern immediately and can move as anticipated. On the other hand, if the human counterpari attempts to induce a movement pattern unfamiliar to the robot, such guidance should experience strong counter force, since the movement is neither inferential nor predictable for the robot.
The robot can infer the intention of the human and adjust its controls immediately. It should be noted that the specification allows for a self region to include intentions, desires, expectations, or physical locations.
Page 3
2.1 Free Energy Principle
As mentioned in the Introduction. the current study used PV-RNN, based on the free energy principle, as the basic model. The free energy principle assumes that brains learn as well as inferred for given observations. by minimizing free energy. defined in Eq. 1.
I:
(1)
po(X|z) is the likelihood of the sensory observation X, given the probabilistic latent variables 2 which is parameterized by A. is the inference model parameterised by Q. As shown in Eq. i, free energy consists of two terms. The first term indicates the complexity, which is the divergence between the approximate posterior probability distribution and prior probability distribution for the latent variable 2. and the second term is the accuracy in predicting the sensory observation.
2.2 PV-RNN implementation
PV-RNN is operated in two distinct phases. it training phase and an interaction phase. In the training phase, we prepared a dataset (joint angle sequences of the robot) and trained the PV-RNN model using it. In the interaction phase, the robot and the human experimenter interact physically, such that the PV-RNN attempts to drive the robot's arms by predicting next-time-step target joint angles while it infers latent variables using actual joint angle readings. Each operation phase is explained in detail below.
an action generation step of generating action information for controlling an action of the control target based on the self-region of the control target, and controlling a driving device for the control target using the action information, thereby controlling an operation of the control target. (Fig. 2
Page 7
2.3 The employed robot controller
The current study uses a humanoid robot, Torobo ] to conduct human-robot kinesthetic interaction experiments. Torobo is equipped with a built-in force-feedback controller that enables humans to back-drive joint angles with subtle force. In the Torobo control system, when a human exerts certain torque 90 joints by pushing or pulling the limbs of Torobo, the exerted torque can be estimated by subtracting the torque inferred as necessary to account for the current static state as well as a dynamic state of the robot from the actual torque measured in the joints. By computing the next time-step joint target positions by adding the current positions with the estimated exerted torque multiplied by a constant gain and feeding them in the PID controller, Torobo's limbs move by following the force exerted on them by the human. This force-feedback controller is integrated with the PV-RNN, which generates the next time-step target joint angles (Fig.2).
Page 10,
4) Human–Robot Interaction: After the training phase, we conducted two types of human–robot interaction experiments using trained PV-RNNs. In these experiments, while Torobo was generating movement pattern transitions successively based on the training, the human experimenter attempted to induce various movement pattern transitions by grasping both arms of Torobo and exerting force on them. These movement pattern transitions included trained transitions (AB, AC, BD, CD, and DA) and untrained transitions (AD, DB, DC, BA, and CA) in which AB, for example, dictates that A pattern is forced to transit to the B pattern. Each transition from one pattern to another requires some guiding force by a human experimenter, even for trained transitions, since ongoing pat terns tend to repeat another cycle with high probability, more than 85% for all patterns. This means that there exist some conflicts between movement trajectories intended by the robot and the human experimentor.)
Sawada does not explicitly state a control device.
Dani, however, does teach a control device. ([0071] In some embodiments, as shown in FIG. 5, the computer system 500 includes a processor 505, memory 510 coupled to a memory controller 515, and one or more input devices 545 and/or output devices 540, such as peripherals, that are communicatively coupled via a local I/O controller 535. These devices 540 and 545 may include, for example, a printer, a scanner, a microphone, and the like. Input devices such as a conventional keyboard 550 and mouse 555 may be coupled to the I/O controller 535. The I/O controller 535 may be, for example, one or more buses or other wired or wireless connections, as are known in the art. The I/O controller 535 may have additional elements, which are omitted for simplicity, such as controllers, buffers (caches), drivers, repeaters, and receivers, to enable communications.)
Therefore, it would be obvious to one of ordinary skill in the art prior to the effective filing date to modify Sawada in light of Dani such that there is a control device that carries out the steps because robotics controls are carried out by controllers, processors, and memory devices. It would be necessary that there be some sort of control device to carry out Sawada’s method.
For Claim 12, Sawada teaches A method to
estimate a self-region of a partner based on preference information indicating an observation that the partner is estimated to expect for a control target and the partner, the preference information being derived based on a predetermined principle and observation sensor data acquired from the partner and the control target, the partner being a person or an autonomous system including a robot, and the control target being an autonomous system including a robot; ((Fig. 2
Page 1, Paragraph 2
In this regard, the current study investigated human-robot kinaesthetic interaction by applying system neuroscience theory, the free energy principle, proposed by Friston [8] which is consonant with enactivism 19, 10]. Let us consider a situation in which 3 robet and 3 human dance, holding each other with both hands, executing memorized dance patterns. If the robot initiates a particular pattern from memory with strong intention, the human counterpart might follow il without resisting because of the strong counter force. On the other hand. if the robot generates $ pattern without strong
intention. the human counterpart might be able to shift to a different pattern without experiencing strong counter force. Let us consider another situation. If the human counterpart attempts to induce a movement pattern that is familiar to the robot, such guidance should proceed easily without strong counter force, since the robot can infer the intended pattern immediately and can move as anticipated. On the other hand, if the human counterpari attempts to induce a movement pattern unfamiliar to the robot, such guidance should experience strong counter force, since the movement is neither inferential nor predictable for the robot.
The robot can infer the intention of the human and adjust its controls immediately. It should be noted that the specification allows for a self region to include intentions, desires, expectations, or physical locations.
Page 3
2.1 Free Energy Principle
As mentioned in the Introduction. the current study used PV-RNN, based on the free energy principle, as the basic model. The free energy principle assumes that brains learn as well as inferred for given observations. by minimizing free energy. defined in Eq. 1.
I:
(1)
po(X|z) is the likelihood of the sensory observation X, given the probabilistic latent variables 2 which is parameterized by A. is the inference model parameterised by Q. As shown in Eq. i, free energy consists of two terms. The first term indicates the complexity, which is the divergence between the approximate posterior probability distribution and prior probability distribution for the latent variable 2. and the second term is the accuracy in predicting the sensory observation.
2.2 PV-RNN implementation
PV-RNN is operated in two distinct phases. it training phase and an interaction phase. In the training phase, we prepared a dataset (joint angle sequences of the robot) and trained the PV-RNN model using it. In the interaction phase, the robot and the human experimenter interact physically, such that the PV-RNN attempts to drive the robot's arms by predicting next-time-step target joint angles while it infers latent variables using actual joint angle readings. Each operation phase is explained in detail below.
determine a self-region of the control target based on preference information indicating an observation that the control target expects for the partner and the control target, which is derived based on the predetermined principle, using the self- region of the partner and an intention of the control target as a goal to be achieved by cooperation between the partner and the control target; and (Fig. 2
Page 1, Paragraph 2
In this regard, the current study investigated human-robot kinaesthetic interaction by applying system neuroscience theory, the free energy principle, proposed by Friston [8] which is consonant with enactivism 19, 10]. Let us consider a situation in which 3 robet and 3 human dance, holding each other with both hands, executing memorized dance patterns. If the robot initiates a particular pattern from memory with strong intention, the human counterpart might follow il without resisting because of the strong counter force. On the other hand. if the robot generates $ pattern without strong
intention. the human counterpart might be able to shift to a different pattern without experiencing strong counter force. Let us consider another situation. If the human counterpart attempts to induce a movement pattern that is familiar to the robot, such guidance should proceed easily without strong counter force, since the robot can infer the intended pattern immediately and can move as anticipated. On the other hand, if the human counterpari attempts to induce a movement pattern unfamiliar to the robot, such guidance should experience strong counter force, since the movement is neither inferential nor predictable for the robot.
The robot can infer the intention of the human and adjust its controls immediately. It should be noted that the specification allows for a self region to include intentions, desires, expectations, or physical locations.
Page 3
2.1 Free Energy Principle
As mentioned in the Introduction. the current study used PV-RNN, based on the free energy principle, as the basic model. The free energy principle assumes that brains learn as well as inferred for given observations. by minimizing free energy. defined in Eq. 1.
I:
(1)
po(X|z) is the likelihood of the sensory observation X, given the probabilistic latent variables 2 which is parameterized by A. is the inference model parameterised by Q. As shown in Eq. i, free energy consists of two terms. The first term indicates the complexity, which is the divergence between the approximate posterior probability distribution and prior probability distribution for the latent variable 2. and the second term is the accuracy in predicting the sensory observation.
2.2 PV-RNN implementation
PV-RNN is operated in two distinct phases. it training phase and an interaction phase. In the training phase, we prepared a dataset (joint angle sequences of the robot) and trained the PV-RNN model using it. In the interaction phase, the robot and the human experimenter interact physically, such that the PV-RNN attempts to drive the robot's arms by predicting next-time-step target joint angles while it infers latent variables using actual joint angle readings. Each operation phase is explained in detail below.
generate action information for controlling an action of the control target based on the self-region of the control target, and to control a driving device for the control target using the action information, thereby controlling an operation of the control target. (Fig. 2
Page 7
2.3 The employed robot controller
The current study uses a humanoid robot, Torobo ] to conduct human-robot kinesthetic interaction experiments. Torobo is equipped with a built-in force-feedback controller that enables humans to back-drive joint angles with subtle force. In the Torobo control system, when a human exerts certain torque 90 joints by pushing or pulling the limbs of Torobo, the exerted torque can be estimated by subtracting the torque inferred as necessary to account for the current static state as well as a dynamic state of the robot from the actual torque measured in the joints. By computing the next time-step joint target positions by adding the current positions with the estimated exerted torque multiplied by a constant gain and feeding them in the PID controller, Torobo's limbs move by following the force exerted on them by the human. This force-feedback controller is integrated with the PV-RNN, which generates the next time-step target joint angles (Fig.2).
Page 10,
4) Human–Robot Interaction: After the training phase, we conducted two types of human–robot interaction experiments using trained PV-RNNs. In these experiments, while Torobo was generating movement pattern transitions successively based on the training, the human experimenter attempted to induce various movement pattern transitions by grasping both arms of Torobo and exerting force on them. These movement pattern transitions included trained transitions (AB, AC, BD, CD, and DA) and untrained transitions (AD, DB, DC, BA, and CA) in which AB, for example, dictates that A pattern is forced to transit to the B pattern. Each transition from one pattern to another requires some guiding force by a human experimenter, even for trained transitions, since ongoing pat terns tend to repeat another cycle with high probability, more than 85% for all patterns. This means that there exist some conflicts between movement trajectories intended by the robot and the human experimentor.)
Sawada does not explicitly teach A program for causing a computer to function as a control device, the program being configured to cause the computer to function as
a self-region estimation unit
a self-region determination unit
an action generation unit
Dani, however, does teach units that perform the tasks. [0018] The learning system 100 may include a training unit 150 and an online learning unit 160, both of which may be hardware, software, or a combination of hardware and software. Generally, the training unit 150 may train a neural network (NN) 170 based on training data derived from one or more user demonstrations of reaching tasks; and the online learning unit 160 may predict intentions for new reaching tasks and may also update the NN 170 based on test data derived from these new reaching tasks.
[0019] In some embodiments, the learning system 100 may be implemented on a computer system and may execute a training algorithm and an online learning algorithm, executed respectively by the training unit 150 and the online learning unit 160, as described below. Further, the learning system 100 may be in communication with a camera 180, such as Microsoft Kinect® for Windows® or some other camera capable of capturing data in three dimensions, which may record demonstrations of reaching tasks as well as new reaching tasks after initial training of the NN 170.
Therefore, it would be obvious to one of ordinary skill in the art prior to the effective filing date to modify Sawada in light of Dani such that there is a self-region estimation unit
a self-region determination unit
an action generation unit
It would be obvious to modify Sawada in light of Dani in this way because units, processors, and CPUs are expected to be successful at computational tasks, and separating tasks amongst a number of different units allows the processors and memories to be specialized for specific tasks, which would be expected to increase efficiency and output.
Dani, however, does teach a A program for causing a computer to function as a control device, the program being configured to cause the computer to function as. ([0006] In yet another embodiment, a computer program product for inferring an intention includes a computer readable storage medium having program instructions embodied therewith. The program instructions are executable by a processor to cause the processor to perform a method. The method includes recording, with a three-dimensional camera, one or more demonstrations of a user performing one or more reaching tasks. Further according to the method, training data is computed to describe the one or more demonstrations. One or more weights of a neural network are learned based on the training data, where the neural network is configured to estimate a goal location of the one or more reaching tasks. A partial trajectory of a new reaching task is recorded. An estimated goal location is computed, by applying the neural network to the partial trajectory of the new reaching task.
[0071] In some embodiments, as shown in FIG. 5, the computer system 500 includes a processor 505, memory 510 coupled to a memory controller 515, and one or more input devices 545 and/or output devices 540, such as peripherals, that are communicatively coupled via a local I/O controller 535. These devices 540 and 545 may include, for example, a printer, a scanner, a microphone, and the like. Input devices such as a conventional keyboard 550 and mouse 555 may be coupled to the I/O controller 535. The I/O controller 535 may be, for example, one or more buses or other wired or wireless connections, as are known in the art. The I/O controller 535 may have additional elements, which are omitted for simplicity, such as controllers, buffers (caches), drivers, repeaters, and receivers, to enable communications.)
Therefore, it would be obvious to one of ordinary skill in the art prior to the effective filing date to modify Sawada in light of Dani such that there is a program that causes a computer to act as control device that carries out the steps because robotics controls are carried out by controllers, processors, and memory devices. It would be necessary that there be some sort of control device to carry out Sawada’s method.
Claim 2 is rejected under 35 U.S.C. 103 as being unpatentable over Sawada in light of Dani in light of Bank et al (US Pub 2022/0203540 A1), hereafter known as Bank.
For Claim 2, Sawada teaches The control device according to claim 1,
Sawada does not teach wherein the action generation unit is configured to display at least one of the preference information indicating the self-region of the partner and the preference information indicating the self-region of the control target, on a display device visually recognizable by the partner.
Bank, however, does teach wherein the action generation unit is configured to display at least one of the preference information indicating the self-region of the partner and the preference information indicating the self-region of the control target, on a display device visually recognizable by the partner. ([0044] In some embodiments, a human participant may utilize a wearable device. The wearable device is configured to provide a notification to the wearer. For example, the wearable device may include an audio output or a vibration mechanism for alerting the wearer to an incoming notification. The incoming notification may include a programmable output of a programmable output detector. The programmable output detector generates and transmits a message as a notification to the wearable device. The wearable device receives the message and presents it to the wearer as a notification. In this way, the wearer receives the message received from the programmable output detector. The message may be based on an interaction image or a foreshadowing image relating to the co-operation of an autonomous machine with the wearer of the wearable device. The notification is not limited to notifications that are indicative of operational warnings. In addition, the notifications may include cooperative intent of the machine relating to a human interaction. For example, the machine may indicate an intention to receive an intermediate product from the human, such as a robot receiving the part for transport to another processing location. The robot may produce an interaction image indicating a precise location the robot expects the human to place the intermediate product. The interaction image may be detected by the programmable output detector and an appropriate message generated. The message is transmitted to the user through a wearable device or other communication device to inform the human of the robot's intention.
[0033] The goal of enhancing safe HRI by incorporating robot feedback to a human participant is to enable the bidirectional communication via visual (light projection techniques, augmented reality), sound and/or vibration (voice), touch (tactile, heat), or haptics. In one embodiment shown in FIG. 2, light projection is used. An interaction image 215i is projected to a space that is visually perceptible to the human participant. The projected interaction image 215i indicates to the human participant the immediate action of the robot. That is, projected interaction image 215i may indicate to the human participant, a restrictive zone that is representative of the freedom of motion of the robot so that the human may avoid those areas. A projected foreshadowing image 215f is projected alongside of the projected interaction image 215i to represent a future intention of the robot to indicate an area the robot intends to occupy in the near future. As the projected interaction image 215i and the projected foreshadowing image 215f are displayed simultaneously alongside one another, the projected interaction image 215i possesses characteristics that differentiate the projected interaction image 215i from the projected foreshadowing image 215f. Any type of differentiating characteristic may be used. For example, the projected interaction image 215i may be a higher intensity light than the projected foreshadowing image 215f. In another embodiment, the projected interaction image 215i may be a different color light than the projected foreshadowing image 215f. Other differentiating characteristics will be apparent to those of skill and may include flashing, pulsing, alternating colors and the like. As shown in FIG. 2, there may be more than one foreshadowing image 215f where a first foreshadowing image indicates an action of the robot that will occur before a second action that is represented by a second foreshadowing image 215f. As discussed above, the first projected foreshadowing image possesses characteristics so that the first projected foreshadowing image is distinguishable from the second projected foreshadowing image. According to some embodiments, the factory may include multiple robots working alongside with a human companion where a light projection creates a heat map around the vicinity of one of the robots involved in an interaction with the human companion. In some embodiments, the remaining schedule and list of tasks are projected and displayed to the human via a mixed-reality goggle worn by the human.)
Therefore, it would be obvious to one of ordinary skill in the art prior to the effective filing date to modify Sawada in light of Bank such that the preference information for a self region is shown on a display because it would allow a partner to observe what the robotic system is thinking. If it believes that the partner wants to hand over an item at a specific location, for example, then it may be useful to the partner to know that that is the expectation. It would allow the partner to lean into the robot’s thinking, and smooth the cooperation process out.
Claims 8 and 10 are rejected under 35 U.S.C. 103 as being unpatentable over Sawada in light of Dani in light of Takahashi et al (US Pub 2016/0101785 A1), hereafter known as Takahashi.
For Claim 8, Sawada teaches The control device according to claim 1, wherein
the self-region estimation unit is configured to store, information about the self-region of the partner with which the control target has performed a cooperative work before in association with the observation sensor data used for the estimation of the self-region, and
Page 3
2.1 Free Energy Principle
As mentioned in the Introduction. the current study used PV-RNN, based on the free energy principle, as the basic model. The free energy principle assumes that brains learn as well as inferred for given observations. by minimizing free energy. defined in Eq. 1.
I:
(1)
po(X|z) is the likelihood of the sensory observation X, given the probabilistic latent variables 2 which is parameterized by A. is the inference model parameterised by Q. As shown in Eq. i, free energy consists of two terms. The first term indicates the complexity, which is the divergence between the approximate posterior probability distribution and prior probability distribution for the latent variable 2. and the second term is the accuracy in predicting the sensory observation.
2.2 PV-RNN implementation
PV-RNN is operated in two distinct phases. it training phase and an interaction phase. In the training phase, we prepared a dataset (joint angle sequences of the robot) and trained the PV-RNN model using it. In the interaction phase, the robot and the human experimenter interact physically, such that the PV-RNN attempts to drive the robot's arms by predicting next-time-step target joint angles while it infers latent variables using actual joint angle readings. Each operation phase is explained in detail below.
the self-region estimation unit is configured to estimate a self-region of the partner, using the self-region corresponding to the partner, which is extracted from the database, when a cooperative work is performed with the partner. Page 3
2.1 Free Energy Principle
As mentioned in the Introduction. the current study used PV-RNN, based on the free energy principle, as the basic model. The free energy principle assumes that brains learn as well as inferred for given observations. by minimizing free energy. defined in Eq. 1.
I:
(1)
po(X|z) is the likelihood of the sensory observation X, given the probabilistic latent variables 2 which is parameterized by A. is the inference model parameterised by Q. As shown in Eq. i, free energy consists of two terms. The first term indicates the complexity, which is the divergence between the approximate posterior probability distribution and prior probability distribution for the latent variable 2. and the second term is the accuracy in predicting the sensory observation.
2.2 PV-RNN implementation
PV-RNN is operated in two distinct phases. it training phase and an interaction phase. In the training phase, we prepared a dataset (joint angle sequences of the robot) and trained the PV-RNN model using it. In the interaction phase, the robot and the human experimenter interact physically, such that the PV-RNN attempts to drive the robot's arms by predicting next-time-step target joint angles while it infers latent variables using actual joint angle readings. Each operation phase is explained in detail below.
Sawada does not explicitly teach that the information is stored in a predetermined database.
Takahashi, however, does teach that for systems that study users, data is stored and retrieved in a predetermined database. ([0060] The user information 231 is information related to the user who is using the in-vehicle device 200. User information 231 is recorded in each of the plurality of in-vehicle devices 200 that are connected to the telematics center 100 via the network 300, each of these items of user information 231 having different details.
[0061] The probe information 232 is information that has been acquired by the in-vehicle device 200 related to the operational state of the subject vehicle. The probe information 232 is read out from the storage device 230 at a predetermined timing by the probe information output unit 223, and is transmitted to the telematics center 100 by the communication unit 240 via the network 300.
[0062] The video information 233 is information including video that has been captured by the in-vehicle device 200. The video information 233 is read out from the storage device 230 at a predetermined timing by the video output unit 225, and is transmitted to the telematics center 100 by the communication unit 240 via the network 300.
[0063] The diagnosis result information 234 is information specifying result of driving characteristics diagnosis for the user who is using the in-vehicle device 200. In response to a request by the user, this diagnosis result information 234 is acquired from the telematics center 100, and is accumulated in the storage device 230.)
Therefore, it would be obvious to one of ordinary skill in the art prior to the effective filing date to modify Sawada in light of Takahashi such that the data is stored in a predetermined database because the data that Sawada is using needs to be stored and retrieved between teaching and experiments. Databases are expected to be successful at storing data, and using one that is predetermined would ensure that the system knows which particular database to access to get necessary information to control the robot.
For Claim 10, Sawada teaches The control device according to claim 1, wherein
the self-region determination unit is configured to identify an action target agent to which the self-region determined based on the preference information expected by the control target is to be transmitted, out of the persons or the autonomous systems included in the partner, and (Fig. 2
Page 7
2.3 The employed robot controller
The current study uses a humanoid robot, Torobo ] to conduct human-robot kinesthetic interaction experiments. Torobo is equipped with a built-in force-feedback controller that enables humans to back-drive joint angles with subtle force. In the Torobo control system, when a human exerts certain torque 90 joints by pushing or pulling the limbs of Torobo, the exerted torque can be estimated by subtracting the torque inferred as necessary to account for the current static state as well as a dynamic state of the robot from the actual torque measured in the joints. By computing the next time-step joint target positions by adding the current positions with the estimated exerted torque multiplied by a constant gain and feeding them in the PID controller, Torobo's limbs move by following the force exerted on them by the human. This force-feedback controller is integrated with the PV-RNN, which generates the next time-step target joint angles (Fig.2).
Page 10,
4) Human–Robot Interaction: After the training phase, we conducted two types of human–robot interaction experiments using trained PV-RNNs. In these experiments, while Torobo was generating movement pattern transitions successively based on the training, the human experimenter attempted to induce various movement pattern transitions by grasping both arms of Torobo and exerting force on them. These movement pattern transitions included trained transitions (AB, AC, BD, CD, and DA) and untrained transitions (AD, DB, DC, BA, and CA) in which AB, for example, dictates that A pattern is forced to transit to the B pattern. Each transition from one pattern to another requires some guiding force by a human experimenter, even for trained transitions, since ongoing pat terns tend to repeat another cycle with high probability, more than 85% for all patterns. This means that there exist some conflicts between movement trajectories intended by the robot and the human experimentor.)
(Fig. 2
Page 7
2.3 The employed robot controller
The current study uses a humanoid robot, Torobo ] to conduct human-robot kinesthetic interaction experiments. Torobo is equipped with a built-in force-feedback controller that enables humans to back-drive joint angles with subtle force. In the Torobo control system, when a human exerts certain torque 90 joints by pushing or pulling the limbs of Torobo, the exerted torque can be estimated by subtracting the torque inferred as necessary to account for the current static state as well as a dynamic state of the robot from the actual torque measured in the joints. By computing the next time-step joint target positions by adding the current positions with the estimated exerted torque multiplied by a constant gain and feeding them in the PID controller, Torobo's limbs move by following the force exerted on them by the human. This force-feedback controller is integrated with the PV-RNN, which generates the next time-step target joint angles (Fig.2).
Page 10,
4) Human–Robot Interaction: After the training phase, we conducted two types of human–robot interaction experiments using trained PV-RNNs. In these experiments, while Torobo was generating movement pattern transitions successively based on the training, the human experimenter attempted to induce various movement pattern transitions by grasping both arms of Torobo and exerting force on them. These movement pattern transitions included trained transitions (AB, AC, BD, CD, and DA) and untrained transitions (AD, DB, DC, BA, and CA) in which AB, for example, dictates that A pattern is forced to transit to the B pattern. Each transition from one pattern to another requires some guiding force by a human experimenter, even for trained transitions, since ongoing pat terns tend to repeat another cycle with high probability, more than 85% for all patterns. This means that there exist some conflicts between movement trajectories intended by the robot and the human experimentor.)the action generation unit is configured to identify an action form of the action target agent based on the observation sensor data, and to generate action information based on the self-region of the control target such that the action information corresponds to the identified action form.
(Fig. 2
Page 7
2.3 The employed robot controller
The current study uses a humanoid robot, Torobo ] to conduct human-robot kinesthetic interaction experiments. Torobo is equipped with a built-in force-feedback controller that enables humans to back-drive joint angles with subtle force. In the Torobo control system, when a human exerts certain torque 90 joints by pushing or pulling the limbs of Torobo, the exerted torque can be estimated by subtracting the torque inferred as necessary to account for the current static state as well as a dynamic state of the robot from the actual torque measured in the joints. By computing the next time-step joint target positions by adding the current positions with the estimated exerted torque multiplied by a constant gain and feeding them in the PID controller, Torobo's limbs move by following the force exerted on them by the human. This force-feedback controller is integrated with the PV-RNN, which generates the next time-step target joint angles (Fig.2).
Page 10,
4) Human–Robot Interaction: After the training phase, we conducted two types of human–robot interaction experiments using trained PV-RNNs. In these experiments, while Torobo was generating movement pattern transitions successively based on the training, the human experimenter attempted to induce various movement pattern transitions by grasping both arms of Torobo and exerting force on them. These movement pattern transitions included trained transitions (AB, AC, BD, CD, and DA) and untrained transitions (AD, DB, DC, BA, and CA) in which AB, for example, dictates that A pattern is forced to transit to the B pattern. Each transition from one pattern to another requires some guiding force by a human experimenter, even for trained transitions, since ongoing pat terns tend to repeat another cycle with high probability, more than 85% for all patterns. This means that there exist some conflicts between movement trajectories intended by the robot and the human experimentor.)
Sawada does not teach the partner includes a plurality of the persons and/or a plurality of the autonomous systems,
Takahashi does teach that for tools that diagnose and identify patterns the partner includes a plurality of the persons and/or a plurality of the autonomous systems, ([0043] The user information 131 is information for individually identifying various users who employ the driving characteristics diagnosis system of FIG. 1. The user information 131, for example, may include information such as the ID numbers of users or user's names or the like. It should be understood that while, for the simplicity of explanation, only one in-vehicle device 200 is shown in FIG. 1, actually a plurality of in-vehicle devices 200 may be connected to the telematics center 100 via the network 300, and these terminal devices may support a plurality of users who utilize the driving characteristics diagnosis system.)
Therefore, it would be obvious to one of ordinary skill in the art prior to the effective filing date to modify Sawada in light of Takahashi such that there are multiple persons or autonomous systems that interact. It would be obvious to do this because it would allow the system to be compatible with multiple people working with it, which would allow it to be used longer and for more applications. Storing data specifically for the users would be useful because different people or systems may have different expectations or preferences, and allowing the system to adapt to the different users would allow it to customize it’s actions for the different users, which would be expected to increase efficiency.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
O’Sullivan et al (US Pub 2017/0190051 A1) relates to robotic systems that predict human intent.
Ueno et al (US Pub 2020/0198148 A1) relates to determining interference between a robot and a human user.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to TRISTAN J GREINER whose telephone number is (571)272-1382. The examiner can normally be reached Mon - Fri 7:30-4:30.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tran Khoi can be reached at Monday-Thursday. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/T.J.G./Examiner, Art Unit 3656 /KHOI H TRAN/Supervisory Patent Examiner, Art Unit 3656