Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
This action is responsive to the following communication: Non-Provisional Application filed Jul. 23, 2024.
Claims 1-10 are pending in the case. Claims 1, 9 and 10 are independent claims.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-3, 5-7 are rejected under 35 U.S.C. 103 as being unpatentable over Faust et al. (hereinafter Faust) U.S. Patent Publication No. 2021/0334320 in view of Mguni et al. (hereinafter Mguni) U.S. Patent Publication No. 2024/0046154.
With respect to independent claim 1, Faust teaches a reinforcement learning method, the reinforcement learning method being performed by a reinforcement learning apparatus (see e.g., Fig. 1 Abstract Para [28]-[30]- “various approaches are presented for training various deep Q network (DQN) agents to perform various tasks associated with reinforcement learning, including hierarchical reinforcement learning, in challenging web navigation environments with sparse rewards and large state and action spaces. ”), the reinforcement learning method comprising:
obtaining information about an agent which is trained by reinforcement learning (see e.g., Para [4][58]-[60]- “These agents include a web navigation that can use learned value function(s) to automatically navigate through interactive web documents, as well as a training agent, referred to herein as a “meta-trainer,” that can be trained to generate synthetic training examples.”); and
performing reinforcement learning of the agent based on first reward (see e.g., Para [84]), and, after a predetermined point (see e.g., Para [84]- “At the beginning of training, the probability p may be initialized with a relatively large probability (e.g., greater than 0.5, such as 0.85) and may be gradually decayed towards 0.0 over some predefined number of steps. After this limit, the initial state of the environment will revert to the original state of the original DOM tree with a full natural language instruction.”), performing reinforcement learning of the agent (see e.g., Para [79]-[82]- “potential-based rewards may be employed for augmenting the environment reward function (which as described previously may be sparse). The environment reward is computed by evaluating if the final state is exactly equal to the goal state. Accordingly, a potential function (Potential(s, g)) may be defined that counts the number of matching DOM elements between a given state (s) and the goal state (g).”).
Faust does not expressly show transitioning reward to second reward having a density different from that of the first reward. However, Faust expressly indicated that the reward function can be augmented because rewards may be sparse (see e.g., Para [80]). Furthermore, Mguni teaches the transitioning feature (see e.g., Fig. 6 and Para [16][29][104]- “a first determining step comprising determining by means of the second agent function whether to use a second reward; (iii) if that determination has a negative outcome, refining the first agent function in dependence on the first reward; and if that determination has a positive outcome, computing the second reward according to a predetermined reward function and refining the first agent function in dependence on the first reward and the second reward; ““ At step 603, a first determining step comprises determining by means of the second agent function whether to use a second reward. At step 604, if that determination has a negative outcome, the first agent function is refined in dependence on the first reward; and if that determination has a positive outcome, the second reward is computed according to a predetermined reward function and the first agent function is refined in dependence on the first reward and the second reward.”). Both Faust and Mguni are directed to. Accordingly, it would have been obvious to the skilled artisan before the effective filing date of the claimed invention having Faust and Mguni in front of them to modify the system of Faust to include the above feature. The motivation to combine Faust and Mguni comes from Mguni. Mguni discloses the motivation to use a second reward so that training can be refined (see e.g., Abstract and Para [16]). This motivation for combination also applies to the remaining claims which depend on this combination.
With respect to dependent claim 2, the modified Faust teaches one of the first reward and the second reward is sparse reward provided depending on whether the agent has reached the goal (see e.g., Para [3][80]- “Reinforcement learning with sparse rewards results in the majority of the episodes generating no signal at all. “ “The environment reward is computed by evaluating if the final state is exactly equal to the goal state. Accordingly, a potential function (Potential(s, g)) may be defined that counts the number of matching DOM elements between a given state (s) and the goal state (g).”), and remaining reward is dense reward provided depending on whether the agent has reached the goal and proximity to the goal (see e.g., Para [80]-[83]- “Accordingly, a potential function (Potential(s, g)) may be defined that counts the number of matching DOM elements between a given state (s) and the goal state (g). This number may be normalized by the number of DOM elements in the goal state. Potential based reward may then be computed as the scaled difference between two potentials for the next state and current state”).
With respect to dependent claim 3, the modified Faust teaches performing the reinforcement learning of the agent comprises performing reinforcement learning using the sparse reward as the first reward (see e.g., Para [1][2]- “Reinforcement learning (“RL”) is challenging in environments having large state and action spaces, and especially when only sparse rewards are available. ”), and, at a predetermined point, transitioning from the dense reward to the second reward and then performing reinforcement learning (see e.g., Para [80] and Mguni Fig. 6 and Para [16][29][104]—See discussion above with respect to claim 1).
With respect to dependent claim 5, the modified Faust teaches a reinforcement learning apparatus, comprising: memory configured to store programs required for generation of an agent and reinforcement learning (see e.g., Para [13]- “a method such as one or more of the methods described above. Yet another implementation may include a system including memory and one or more processors operable to execute instructions, stored in the memory, to implement one or more modules or engines that, alone or collectively, perform a method such as one or more of the methods described above.”); and a controller configured to obtain information about the agent which is trained by reinforcement learning, and to perform reinforcement learning of the agent based on first reward (see e.g., Para [80]-[84]- “At the beginning of training, the probability p may be initialized with a relatively large probability (e.g., greater than 0.5, such as 0.85) and may be gradually decayed towards 0.0 over some predefined number of steps. After this limit, the initial state of the environment will revert to the original state of the original DOM tree with a full natural language instruction.”), and, after a predetermined point, transitioning reward to second reward having a density different from that of the first reward and then performing reinforcement learning of the agent (see e.g., Mguni Fig. 6 and Para [16][29][104]- “a first determining step comprising determining by means of the second agent function whether to use a second reward; (iii) if that determination has a negative outcome, refining the first agent function in dependence on the first reward; and if that determination has a positive outcome, computing the second reward according to a predetermined reward function and refining the first agent function in dependence on the first reward and the second reward; ““ At step 603, a first determining step comprises determining by means of the second agent function whether to use a second reward. At step 604, if that determination has a negative outcome, the first agent function is refined in dependence on the first reward; and if that determination has a positive outcome, the second reward is computed according to a predetermined reward function and the first agent function is refined in dependence on the first reward and the second reward.” The motivation to combine is discussed above with respect to claim 1).
With respect to dependent claim 6, the modified Faust teaches the controller determines any one of sparse reward provided depending on whether the agent has reached a goal and dense reward provided depending on whether the agent has reached the goal and proximity to the goal to be the first reward (see e.g., Mguni Fig. 6 and Para [16][29][104] – the motivation to combine is discussed above) and then performs reinforcement learning, and, after a predetermined point, determines remaining reward to be the second reward and then performs reinforcement learning (see e.g., Para [80]-[84]- “Accordingly, a potential function (Potential(s, g)) may be defined that counts the number of matching DOM elements between a given state (s) and the goal state (g). This number may be normalized by the number of DOM elements in the goal state. Potential based reward may then be computed as the scaled difference between two potentials for the next state and current state”).
With respect to dependent claim 7, the modified Faust teaches the controller performs reinforcement learning of the agent using the sparse reward as the first reward, and, at a predetermined point, transitions reward from the dense reward to the second reward and then performs reinforcement learning (see e.g., Para [80] and Mguni Fig. 6 and Para [16][29][104]—See discussion above with respect to claim 1).
Claim 9 is rejected for the similar reasons discussed above with respect to claim 1.
Claim 10 is rejected for the similar reasons discussed above with respect to claim 1.
Claims 4 and 8 are rejected under 35 U.S.C. 103 as being unpatentable over Faust in view of Mguni and further in view of Kimura et al. (hereinafter Kimura) U.S. Patent Publication No. 2019/0272465.
With respect to dependent claim 4, Faust does not expressly show performing the reinforcement learning of the agent comprises determining the dense reward using a density reward function calculated based on an L2 distance between a current state of the agent and the goal. However, Kimura teaches similar feature (see e.g., Fig. 6 and Para [102][103]- “The dense reward is a distance between the end-effector 208 and the point target 210. The sparse reward is based on a bonus for reaching. The dense reward function (6) and the sparse reward function (7) were employed as comparative examples (Experiments 1, 2).”). Both Faust and Kimura are directed to reinforced learning. Accordingly, it would have been obvious to the skilled artisan before the effective filing date of the claimed invention having Faust and Kimura in front of them to further modify the modified system of Faust to include the above feature. The motivation to combine Faust and Kimura comes from Kimura. Kimura discloses the motivation to use a density reward function calculated based on an L2 distance between a current state of the agent and the goal so that density can be measured (see e.g., Para [100]-[106]).
With respect to dependent claim 8, the modified Faust teaches the controller calculates the dense reward using a density reward function calculated based on an L2 distance between a current state of the agent and the goal (see e.g., Kimura Fig. 6 and Para [102][103]- “The dense reward is a distance between the end-effector 208 and the point target 210. The sparse reward is based on a bonus for reaching. The dense reward function (6) and the sparse reward function (7) were employed as comparative examples (Experiments 1, 2).” The motivation to combine is discussed above).
It is noted that any citation to specific pages, columns, lines, or figures in the prior art references and any interpretation of the references should not be considered to be limiting in any way. “The use of patents as references is not limited to what the patentees describe as their own inventions or to the problems with which they are concerned. They are part of the literature of the art, relevant for all they contain.” In re Heck, 699 F.2d 1331, 1332-33, 216 USPQ 1038, 1039 (Fed. Cir. 1983) (quoting In re Lemelson, 397 F.2d 1006, 1009, 158 USPQ 275, 277 (CCPA 1968)). Further, a reference may be relied upon for all that it would have reasonably suggested to one having ordinary skill the art, including nonpreferred embodiments. Merck & Co. v. Biocraft Laboratories, 874 F.2d 804, 10 USPQ2d 1843 (Fed. Cir.), cert. denied, 493 U.S. 975 (1989). See also Upsher-Smith Labs. v. Pamlab, LLC, 412 F.3d 1319, 1323, 75 USPQ2d 1213, 1215 (Fed. Cir. 2005); Celeritas Technologies Ltd. v. Rockwell International Corp., 150 F.3d 1354, 1361, 47 USPQ2d 1516, 1522-23 (Fed. Cir. 1998).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PEIYONG WENG whose telephone number is (571)270-1660. The examiner can normally be reached on Mon.-Fri. 8 am to 5 pm.
If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Matthew Ell, can be reached on (571) 270-3264. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://portal.uspto.gov/external/portal. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free).
/PEI YONG WENG/Primary Examiner, Art Unit 2141