DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Remarks
This final office action is a response to the reply received on 05/26/2026. Claims 1 and 3-11 are pending. Claim 2 has have been canceled. Claims 1, 3-5, 7-8, and 11 have been amended.
Response to Arguments
Applicant’s arguments with respect to the claims have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 4, and 9-11 are rejected under 35 U.S.C. 103 as being unpatentable over Mizokawa et. al. (US 6901390 B2).
Regarding Claim 1, Mizokawa discloses:
A robot control device comprising one or more processors configured to: (See at least Col. 3 Lines 6-11 via "The device to be controlled in the present invention is not limited and includes any objects which interact with a user and which are allowed to output not only compulsory (commanded) behavior but also autonomous behavior for the benefit of the user. For example, home appliances, an industrial robot…" and Col. 5 Lines. 52-54 via "FIGS. 12-21 show an embodiment wherein the present invention is adapted to a pet robot which a user can keep as a pet" as well as Col. 26 Lines. 10-28 via the processing unit(s))
cause a robot to make motions in accordance with respective motion patterns each of which is made up of a combination of motion elements, (See at least Col. 14 Lines 22-23 via "…the pet robot is capable of conducting various actions" and Col. 18 Lines 52-60 via "commanded behavior generation unit F for generating behavior in accordance with the user's command based on the outcome of recognition and at least the preceding behavior previously outputted by this system itself; a new behavior pattern generation unit G for generating new behavior patterns which are usable in the emotion expression decision unit D and the autonomous behavior generation unit E based on the pseudo-personality information" as well as Claim 4 via "the patterns of autonomous behavior include movements or changes of the foregoing" and Col. 21 Lines 20-24 via "emotion expression can be indicated by movement of a tail or the pet robot itself. In the above, the emotion expression pattern memory 6 memorizes in advance patterns of each of parts of "face", and patterns of movement of the tail or the pet robot itself". Furthermore, see Col 2. Lines 35-42 via "autonomous behavior-forming unit memorizing a predetermined relationship between patterns of autonomous behavior…each pattern constituted by plural predetermined elements; (h) a new behavior pattern generation unit which generates new patterns of autonomous behavior by selecting and combining patterns' elements")
derive respective evaluation values of the motion patterns based on detection results by a sensor after the motions are made in accordance with the respective motion patterns; (See at least Col. 19 Lines 38-46 via "the user's evaluation recognition unit A1 establishes the user's evaluation as to behavior outputted by the pet robot and sets the result, for example, as information on the user's evaluation, based on the manner in which the user responds to the pet robot. The manner may be sensed by a tactile sensor, e.g., how the user touches or strokes the pet robot. If the manners of touching or stroking are correlated with evaluation values in advance, the evaluation values can be determined.")
and generate one or more new motion patterns by recombining some of the motion elements of the motion patterns within a series of the motions (See at least Col. 13 Lines 55-67 via "The new behavior pattern generation unit G calculates intervals of generating new behavior patterns, and then judges whether generation of new behavior patterns is activated based on the intervals and random elements. If generation is judged to be activated, the new behavior pattern generation unit G combines behavior elements stored in an emotion expression element database 8 and an autonomous behavior element database 9, which memorize in advance a number of emotion expression patterns and autonomous behavior patterns, respectively, thereby setting new emotion expression patterns and new autonomous behavior patterns. The above combining processes can be conducted randomly or by using genetic algorithms" as well as Claim 1 via " a new behavior pattern generation unit which generates new patterns of autonomous behavior (NBP) by selecting and combining patterns' elements (PEs) randomly or under predesignated rules and which overrides existing patterns in the autonomous behavior-establishing unit with the new patterns under predesignated conditions, said new behavior pattern generation unit having a database storing a pool of elements of patterns of autonomous behavior, wherein NBP is comprised of the sum of PEs").
However, Mizokawa does not explicitly disclose the recombination of motion elements being on the basis of the evaluation values. Nevertheless, Mizokawa discloses evaluating a user response to revise elements of a generated pattern: "After generating new emotion expression patterns and new autonomous behavior patterns…new patterns are checked and revised with a common sense database and the user information to eliminate extreme or obviously distaste patterns…if when expression of "anger" is associated with "growling", the user expresses uncomfortable reaction, the sound associated with "anger" is changed from "loud growling" to "quiet growling", for example. Checking based on the user's response or evaluation is conducted using the user information." [Mizokawa Col. 22 Lines 53-67]. Additionally, Mizokawa describes a priority of behavior elements, that is influenced by user reaction and pseudo-personality: Col 15 Lines 32-38 via "sets degrees of priority at each item of behavior information based on the pseudo-personality information, the pseudo-emotion information, and the user information; generates new priority behavior information; combines each item of behavior information based on the new priority behavior information having priority; and decides final output behavior".
Although Mizokawa does not explicitly state that the evaluation values/user reactions are directly input to produce a new combination/recombination of behavior elements, Mizokawa discloses using the user information (which is derived from evaluation(s) of the user's reaction(s)) to determine a priority of behaviors to combine (See Col 15 Lines 32-38), and further discloses modifying a newly generated pattern with respect to the user's reactions to specific behavior elements (See Col. 22 Lines 53-67). Thus, it would have been obvious to one of ordinary skill in the art prior to the effective filing date of the given invention to modify Mizokawa's robotic control to directly utilize the evaluation information when generating the new pattern because Mizokawa already discloses generating new patterns by combining behavior elements, prioritizing certain elements based on user response, and using evaluations of the user's response to modify elements. Thus, directly incorporating the evaluation information into the process of generating the new pattern would have predictably enabled the newly generated patterns to prioritize the behavior elements that yield(ed) a more positive user reaction and avoided behavior patterns that yield discomfort from the user.
Regarding Claim 4, Modified Mizokawa discloses the robot control device according to Claim 1.
Furthermore, Mizokawa discloses: wherein the one or more processors are configured to: select at least two motion patterns from among the motion patterns based on a descending order of the evaluation values; and (See at least Col 15 Lines 32-38 via "sets degrees of priority at each item of behavior information based on the pseudo-personality information, the pseudo-emotion information, and the user information; generates new priority behavior information; combines each item of behavior information based on the new priority behavior information having priority; and decides final output behavior" and additionally Col. 15 Lines 49-64 via "…if there is no required absolute behavior, a degree of priority (1 to 3: in this case, priority is 1, i.e., high priority) is determined for each item of behavior information based on a table of degrees of priority, and is set as new priority behavior information…The table of degrees of priority is a table defining degrees of priority as described in FIG. 10, for example: If the pseudo-emotion is "joy", the pseudo-personality is "obedience", and the state of the user is "anger", the commanded behavior is set at priority 1, the emotion expression behavior is set at priority 2, and the autonomous behavior is set at priority 3. As above, the table is a group of data wherein a degree of priority of each item of behavior information is set for a combination of all pseudo-emotions, all pseudo-personalities, and all user's states.". Also see Col. 8 Lines 39-43 via "the current state of the user includes a time period while the user is located within a range where the system can locate the user; the level of the user's emotions; and the user's evaluation as to the behavior outputted by the system" **Wherein the priority ranking is used to select multiple behavior patterns, priority rank "1" is the highest priority, and "3" is the lowest priority, which corresponds to a descending order under BRI)
generate the one or more new motion patterns by recombining the motion elements of the at least two motion patterns (See at least Claim 1 which states " paterns' " (plural, indicating at least two) via "a new behavior pattern generation unit which generates new patterns of autonomous behavior (NBP) by selecting and combining patterns' elements (PEs)…said new behavior pattern generation unit having a database storing a pool of elements of patterns of autonomous behavior, wherein NBP is comprised of the sum of PEs").
Regarding Claim 9, Modified Mizokawa discloses the robot control device according to Claim 1.
Furthermore, Mizokawa discloses: wherein the motion elements of the motion patterns include at least one of presence or absence of movement of a predetermined part, a speed of the movement of the predetermined part, a pitch of a sound to be output, or a length of the sound to be output (See at least Col. 21 Lines 17-25 via "The emotion expression includes, in practice, changes in shape of parts such as eyes and a mouth if a "face" is created by a visual output device. Further, emotion expression can be indicated by movement of a tail or the pet robot itself In the above, the emotion expression pattern memory 6 memorizes in advance patterns of each of parts of "face", and patterns of movement of the tail or the pet robot itself in relation to each level of each preliminary emotion." as well as Col. 22 Lines 61-65 via "Further, if when expression of "anger" is associated with "growling", the user expresses uncomfortable reaction, the sound associated with "anger" is changed from "loud growling" to "quiet growling", for example").
Regarding Claim 10, Modified Mizokawa discloses:
A robot comprising: the robot control device according to Claim 1; and the sensor (See at least Claim 1 rejection in addition to Col. 18 Lines. 4-5 via "a pet robot which a user can keep as a pet" and Col. 18 Lines. 24-26 via "In practice, in order to recognize the state of the user or the environments, this pet robot is equipped with a visual detection sensor, a touch detection sensor, a hearing-detection sensor").
Regarding Claim 11, Mizokawa discloses:
A robot control method that controls a robot including a sensor that detects an action from outside, comprising: (See at least Col. 6 Lines 8-16 via " In practice, in order to recognize the state of the user or the environments, this assist system comprises a visual detection sensor, a touch detection sensor, an auditory detection sensor, and a means for accessing external databases" and also Col. 3 Lines 19-21 via " In the present invention, another important aspect is to provide a method for adapting behavior of a device to user's characteristics" and Col. 1 Lines 16-18 via "various controlling methods have been available for controlling an object in accordance with a user's demands")
(Regarding the method steps, see Claim 1 rejection which illustrates the processing unit performing the steps).
Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over Mizokawa et. al. (US 6901390 B2) in view of Sabe et. al. (US 20030045203 A1, IDS dated 06/20/2025).
Regarding Claim 3, Modified Mizokawa discloses the robot control device according to Claim 1.
Furthermore, Mizokawa discloses: wherein the one or more processors are configured to: select, from among the motion patterns, a motion pattern having a higher evaluation value among the motion patterns with (See at least Col. 26 Lines. 10-28 via the processing unit(s), and see Col. 19 Lines 38-46 via "the user's evaluation recognition unit A1 establishes the user's evaluation as to behavior outputted by the pet robot and sets the result, for example, as information on the user's evaluation, based on the manner in which the user responds to the pet robot. The manner may be sensed by a tactile sensor, e.g., how the user touches or strokes the pet robot. If the manners of touching or stroking are correlated with evaluation values in advance, the evaluation values can be determined." as well as Col 15 Lines 32-38 via "sets degrees of priority at each item of behavior information based on the pseudo-personality information, the pseudo-emotion information, and the user information; generates new priority behavior information; combines each item of behavior information based on the new priority behavior information having priority; and decides final output behavior" and additionally Col. 15 Lines 49-64 via "…if there is no required absolute behavior, a degree of priority (1 to 3: in this case, priority is 1, i.e., high priority) is determined for each item of behavior information based on a table of degrees of priority, and is set as new priority behavior information…The table of degrees of priority is a table defining degrees of priority as described in FIG. 10, for example: If the pseudo-emotion is "joy", the pseudo-personality is "obedience", and the state of the user is "anger", the commanded behavior is set at priority 1, the emotion expression behavior is set at priority 2, and the autonomous behavior is set at priority 3. As above, the table is a group of data wherein a degree of priority of each item of behavior information is set for a combination of all pseudo-emotions, all pseudo-personalities, and all user's states." **Wherein Mizokawa discloses the evaluation based prioritization of behaviors. Additionally, see Claim 1 which states " patterns' " (plural, indicating at least two) via "a new behavior pattern generation unit which generates new patterns of autonomous behavior (NBP) by selecting and combining patterns' elements (PEs)" ).
However, although Mizokawa discloses behavior(s) having varying associated priorities, and also discloses a likelihood of pseudo emotions [See Col. 10 Lines 39-60]; Mizokawa does not explicitly disclose the higher evaluation value corresponding to a higher probability.
Nevertheless, Sabe--who is directed towards a robot apparatus and control method for judging character--discloses: a higher evaluation value among the motion patterns with a higher probability to select (See at least ¶0108 via "On the basis of this recognition result and a notice from the action switching module 71, the learning module 72 changes transition probability, to which the behavioral model 70.sub.1 to 70.sub.n corresponding thereto in the behavioral model library 70 corresponds in such a manner as to lower, when it "was patted (scolded)," the revelation probability of the action and to raise, when it "was petted (praised)," the revelation probability of the action." **Which illustrates the raising of the probability when there is a positive action (associated with a higher evaluation value)**. )
Therefore, it would have been obvious to one of ordinary skill in the art prior to the effective filing date of the given invention to modify Modified Mizokawa's prioritization of behavior elements and patterns in view of Sabe's higher probability of selecting the patterns with higher evaluation values in order to enable the robot's behavior to smoothly and naturally grow/adapt to the user's preferential behavior: "this robot apparatus is capable of reducing discontinuity in action output before and after change in state space to be used for action generation because the state space to be used for action generation continuously changes. Thereby, output actions can be changed smoothly and naturally, thus making it possible to realize a robot apparatus which improves the entertainment characteristics." [Sabe ¶0017].
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Mizokawa et. al. (US 6901390 B2) in view of Wei et. al. (US 11360441 B1).
Regarding Claim 5, Modified Mizokawa discloses the robot control device according to Claim 1.
Furthermore, Mizokawa discloses: wherein each new motion pattern is generated by selecting (See at least Claim 1 via "a new behavior pattern generation unit which generates new patterns of autonomous behavior (NBP) by selecting and combining patterns' elements (PEs)…said new behavior pattern generation unit having a database storing a pool of elements of patterns of autonomous behavior, wherein NBP is comprised of the sum of PEs" as well as Col 15 Lines 32-38 and Col. 15 Lines. 49-64 which disclose the priority for behavior(s) being determined partially based on the user information derived from evaluation of user reactions. Additionally see at least Col. 15 Lines 49-64 via "…if there is no required absolute behavior, a degree of priority (1 to 3: in this case, priority is 1, i.e., high priority) is determined for each item of behavior information based on a table of degrees of priority, and is set as new priority behavior information…The table of degrees of priority is a table defining degrees of priority as described in FIG. 10…" **Wherein there are multiple behavior patterns, and certain behaviors are prioritized based on user information which is derived from the user reaction(s), and the patterns are derived from the priority).
However, Mizokawa does not explicitly disclose "two" motion patterns, but instead discloses a plural " patterns' ". Thus, it would have been obvious to one of ordinary skill in the art prior to the effective filing date of the given invention to modify Modified Mizokawa to consider two patterns because two is encompassed in "patterns" (plural).
However, Modified Mizokawa does not explicitly disclose selecting elements with a predetermined probability from the motion patterns.
Nevertheless, Wei--who is directed towards self-learning control for a robot--discloses: each selected (See at least Col. 9 Lines 61-63 via " one expansion scheme at block B3 is that it depends on a crossover probability 0<p.sub.c<1 whether a control gain element needs to crossed" as well as Col. 11 Lines 20-26 via "a selection strategy combining roulette with elite retention are adopted to select the control gain element in the expanded control gain set into a new control gain set for next iteration. The roulette strategy enables a control gain element with greater fitness to be selected into a new control gain set with a higher probability…" **Wherein Wei discloses the combination of control gain elements of the sets/populations for robotic control with respect to corresponding crossover probabilities which corresponds to the concept of selecting motion elements from two motion patterns).
Therefore, it would have been obvious to one of ordinary skill in the art prior to the effective filing date of the given invention to modify Modified Mizokawa in view of the element crossover probabilities concept of Wei in order to improve the control and adaptability of the robot: "performing preferential iteration on control gain elements in the first control gain set to obtain a target control gain set" [Wei Col. 2 Lines 13-14] and "enable a control gain element with greater fitness to be selected into a new control gain set with a higher probability" [Wei Col. 11 Lines 20-26].
Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Mizokawa et. al. (US 6901390 B2) in view of Du et. al. (US 20120062399 A1).
Regarding Claim 6, Modified Mizokawa discloses the robot control device according to Claim 1.
Furthermore, Mizokawa discloses the motion elements: (See at least Claim 1 via "a new behavior pattern generation unit which generates new patterns of autonomous behavior (NBP) by selecting and combining patterns' elements (PEs)…said new behavior pattern generation unit having a database storing a pool of elements of patterns of autonomous behavior, wherein NBP is comprised of the sum of PEs")
However, Modified Mizokawa does not explicitly disclose the elements being represented by Boolean value.
Nevertheless, Du--who discloses generating improved solutions based on evaluation and bit mutation--discloses: wherein each of the motion elements is represented by a Boolean value (See at least ¶0035 via "the method can comprise randomly selecting a parent sequence of the parent binary sequences and applying two-point mutation to the parent sequence including generating a predefined number of offspring binary sequences from the parent sequence" and ¶0056 via "… the GA can be coded in a binary string…").
Therefore, it would have been obvious to one of ordinary skill in the art prior to the effective filing date of the given invention to modify Modified Mizokawa to represent the motion elements as Boolean/binary bit values such as in Du, in order to improve adaptability based on the fitness evaluation: "evolving the biphase sequences of bits with an evolutionary algorithm including bit climbing, wherein the bit climbing includes flipping the bits one by one and evaluating a pre-selected fitness function in connection with determining whether to retain a flipped or non-flipped value" [Du ¶0039]; thus, it would have been obvious to encode modified Mizokawa's motion elements in a similar manner to achieve the same goal of growth/evolution.
Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Mizokawa et. al. (US 6901390 B2) and Du et. al. (US 20120062399 A1) in view of Saito et. al. (US 20020016128 A1, IDS dated 06/20/2025).
Regarding Claim 7, Modified Mizokawa discloses the robot control device according to Claim 6.
Furthermore, Mizokawa discloses: wherein the one or more new motion patterns comprise a plurality of new motion patterns, and the one or more processors are configured to: (See at least Col. 2 Lines 40-42 via "a new behavior pattern generation unit which generates new patterns of autonomous behavior by selecting and combining patterns' elements").
However, Modified Mizokawa does not explicitly disclose the sum of the evaluation values being equal or less than a reference value or the Boolean values.
Nevertheless, Du discloses: invert the Boolean value of each of the motion elements of at least one of the plurality of new motion patterns (See at least ¶0035 where Du discloses representing elements as binary bits via "the method can comprise randomly selecting a parent sequence of the parent binary sequences and applying two-point mutation to the parent sequence including generating a predefined number of offspring binary sequences from the parent sequence" and ¶0056 via "… the GA can be coded in a binary string…"; additionally see at least ¶0074 via "the bits of the string are flipped by flipping the bits of parent sequences one by one. Their corresponding value in the fitness function f.sub.3 is then evaluated. If the newly flipped sequence has a higher fitness value, the newly flipped sequence replace the previous one in recognition of evolutionary superiority. For a single parent sequence, the fitness function value of L one-bit flipped versions of the parent sequence are computed").
Therefore, it would have been obvious to one of ordinary skill in the art prior to the effective filing date of the given invention to modify Modified Mizokawa to represent and mutate the motion elements as Boolean/binary bit values such as in Du, in order to improve adaptability based on the fitness evaluation: "evolving the biphase sequences of bits with an evolutionary algorithm including bit climbing, wherein the bit climbing includes flipping the bits one by one and evaluating a pre-selected fitness function in connection with determining whether to retain a flipped or non-flipped value" [Du ¶0039]. Furthermore, the concept of inverting/flipping the bit values as a part of the evolution in Du is an adaptation that supports modified Mizokawa's adaptation of an object/robot's behavior in order to create behavior highly responsive to the user [Mizokawa Col. 1 Lines 11-15], and additionally a shared goal of personality growth of the robot: "Since the pseudo-personality is expressed as the pseudo-emotions, the user can enjoy fostering a unique assist system during a growth period, similar to raising a child. Further, approximately after completion of the growth period, the user can enjoy establishing "friendship" with the assist system" [Mizokawa Col. 17 Lines 40-45].
However, Modified Mizokawa does not explicitly disclose the sum of evaluation values or the comparison to a reference value.
Nevertheless, Saito--who is directed towards an interactive toy and reaction behavior generating device--discloses: in response to a sum of the (See at least ¶0053 via "The point counting unit 15 counts a generated action point caused by the reaction behavior of the dog type robot 1. The action point is counted (added/subtracted) to the total value of the action points, and the latest total value is stored in the RAM. Here, an "action point" means a generated score caused by the reaction behavior (output) of the dog type robot 1. The total value of the action points corresponds to the level of communication between the dog type robot 1 and a user. It also becomes a base parameter related to the update of the character parameter XY, which determines the character state of the dog type robot 1.". Additionally, Saito compares the action point total to a reference value in order to determine when to shift to a next stage in the growth/evolution process: See at least ¶0066 for example: " The aggregate total value VTA corresponds to the amount of communication between a user and the dog type robot 1, and becomes a value for a determination when shifting from the first stage to the second stage." as well as Figures 12-14).
Therefore, it would have been obvious to one of ordinary skill in the art prior to the effective filing date of the given invention to modify Modified Mizokawa in view of the concept of Saito's sum and comparison to a reference value in order to have utilize a parameter to determine when to evolve/grow/change the behavior of the object/robot, which Mizokawa also does: "Since the pseudo-personality is expressed as the pseudo-emotions, the user can enjoy fostering a unique assist system during a growth period, similar to raising a child" [Mizokawa Col. 17 Lines 40-43] and "The pet robot itself has a pseudo-personality, which grows by interacting with the user and through its own experience, and pseudo-emotions, which change in accordance with interaction with the user and the surrounding environments." [Mizokawa Col. 18 Lines 15-19], which ultimately advances the shared goal of increasing adaptability to the user.
Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Mizokawa et. al. (US 6901390 B2) in view of Jonathan Hui (NPL: RL — Introduction to Deep Reinforcement Learning, previously attached in non-final dated 02/26/2026).
Regarding Claim 8, Modified Mizokawa discloses the robot control device according to Claim 1.
Furthermore, Mizokawa discloses: wherein the one or more processors are configured to: determine whether a user has touched the robot based on the detection results by the sensor; and derive the evaluation values (See at least Col. 18 Lines 24-26 via "In practice, in order to recognize the state of the user or the environments, this pet robot is equipped with a visual detection sensor, a touch detection sensor" as well as Col. 19 Lines 38-46 via " As shown in this figure, the user's evaluation recognition unit A1 establishes the user's evaluation as to behavior outputted by the pet robot and sets the result, for example, as information on the user's evaluation, based on the manner in which the user responds to the pet robot. The manner may be sensed by a tactile sensor, e.g., how the user touches or strokes the pet robot. If the manners of touching or stroking are correlated with evaluation values in advance, the evaluation values can be determined.")
Furthermore, Mizokawa discloses the robot making a motion, detecting a touch, and the evaluation value (See at least Col. 21 Lines 20-22 via "Further, emotion expression can be indicated by movement of a tail or the pet robot itself In the above", Col. 19 Lines 43-44 via "how the user touches or strokes the pet robot. If the manners of touching or stroking are correlated with evaluation values in advance, the evaluation values can be determined.")
However, Modified Mizokawa does not explicitly disclose th evaluation being based on the time from when the robot made a motion until when the sensor detected a touch.
Nevertheless, Hui--who is directed towards reinforcement learning--discloses: such that as a time from when the (See at least Page 3 via "The discount factor γ reduces the weight of future rewards, reflecting the principle that a delayed reward often holds less value than an immediate one.).
Therefore it would have been obvious to one of ordinary skill in the art prior to the effective filing date of the given invention to modify Modified Mizokawa in view of Hui's reinforcement learning principles where a delayed reward holds less value than an immediate reward (an immediate reward corresponding to a higher evaluation value) in order to improve the reinforcement learning for a robot to be able to perform more desired actions/behavior control (actions with higher evaluations or higher rewards): "The term “action” is equivalent to “control” in many contexts. Objectives can be framed in two equivalent ways: maximizing rewards or minimizing costs, where costs are simply the negative of rewards." [Hui Page 4] as well as to improve the robot learning stability: "helps some algorithms achieve stable convergence." [Hui Page 3].
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Cohen et. al. (US10898999B1): directed towards selective human-robot interaction
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KAYLA RENEE DOROS whose telephone number is (703)756-1415. The examiner can normally be reached Generally: M-F (8-5) EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Abby Lin can be reached on (571) 270-3976. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/K.R.D./Examiner, Art Unit 3657
/ABBY LIN/Supervisory Patent Examiner, Art Unit 3657