Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: “a prior data collection unit configured to collect prior data…”, “a prior data processing unit configured to process the collected prior data…”, and “a policy learning unit configured to learn a policy of an ego agent…” in claim 1 and “the policy learning unit learns policy by…” in claim 6.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 4-8, and 11-13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Nair et al. (NPL from IDS: Overcoming Exploration in Reinforcement Learning with Demonstrations, published 2018, hereinafter “Nair”) in view of Schmidhuber et al. (US Pub. No. 2021/0089966, published March 2021, hereinafter “Schmidhuber”).
Regarding claim 1, Nair teaches a deep reinforcement learning based decision-making apparatus through prior data and selective imitation learning comprising:
a prior data collection unit configured to collect prior data from one or more other agents (Nair, Section IV Subsection A Paragraph 1 – “First, we maintain a second replay buffer RD where we store our demonstration data in the same format as R. In each minibatch, we draw an extra ND examples from RD to use as off-policy replay data for the update step.” – teaches a prior data collection unit configured to collect prior date from one or more other agents (collects demonstration data, which is data from one or more other agents, and stores demonstration data in a second replay buffer RD, thus collecting prior data from one or more other agents));
a prior data processing unit configured to process the collected prior data into data including state, action, next state, and reward (Nair, Section III Subsection B Paragraph 2 – “Concretely, DDPG maintains an actor function π(s) with parameters θπ, a critic function Q(s,a) with parameters θQ, and a replay buffer R as a set of tuples (st,at,rt,st+1) for each transition experienced.” and in Section IV Subsection A Paragraph 1 – “First, we maintain a second replay buffer RD where we store our demonstration data in the same format as R. In each minibatch, we draw an extra ND examples from RD to use as off-policy replay data for the update step.” – teaches a prior data processing unit configured to process the collected prior data into data including state, action, next state, and reward (replay buffer RD stores data in same format as replay buffer R, replay buffer R stores data in format of state st, action at, reward rt, and next state st+1. Thus, teaches processing collected prior data, or demonstration data, into data including state, action, next state, and reward)); and
a policy learning unit configured to learn policy of an ego agent using the processed prior data and interaction data including state, action, next state, and reward (Nair, Section III Subsection B Paragraph 2 – “Concretely, DDPG maintains an actor function π(s) with parameters θπ, a critic function Q(s,a) with parameters θQ, and a replay buffer R as a set of tuples (st,at,rt,st+1) for each transition experienced. DDPG alternates between running the policy to collect experience and updating the parameters.” and in Section IV Subsection A Paragraph 1 – “First, we maintain a second replay buffer RD where we store our demonstration data in the same format as R. In each minibatch, we draw an extra ND examples from RD to use as off-policy replay data for the update step.” – teaches a policy learning unit (actor function π(s)) configured to learn policy of an ego agent (actor) using the processed prior data (demonstration data stored in replay buffer RD, RD samples used as off-policy replay data in policy update step) and interaction data including state, action, next state, and reward (replay buffer R stores interaction data of actor in environment for each transition experiences, replay buffer R includes state st, action at, next state st+1, and reward rt)).
Nair fails to explicitly teach wherein the interaction data is obtained through real-time interaction with environment.
However, analogous to the field of the claimed invention, Schmidhuber teaches:
a policy learning unit configured to learn policy of an ego agent using the processed prior data and interaction data including state, action, and reward obtained through real-time interaction with environment (Schmidhuber, [0032], [0038] – “The computer system 100 has one or more input/output (I/O) devices 108 (e.g., to interact with and receive feedback from the external environment) and a replay buffer 110. The replay buffer 110 is a computer-based memory buffer that is configured to hold packets of data that is relevant to the external environment with which the computer system 100 is interacting (e.g., controlling/influencing and/or receiving feedback from). In a typical implementation, the replay buffer 110 stores data regarding previous command/control signals (to the external environment), observed results (rewards and associated time horizons), as well as other observed feedback from the external environment… This data trains, or at least helps train, the RNN” and in [0198] – “In what follows, s, a and r denote state, action, and reward respectively… A policy π:S.fwdarw.A is a function that selects an action in a given state. A stochastic policy maps a state to a probability distribution over actions. Each episode consists of an agent's interaction with its environment starting in an initial state and ending in a terminal state while following any policy.” – teaches learning a policy of an ego agent (policy π:S.fwdarw.A) using processed prior data (replay buffer 110 stores data that trains or at least helps train the RNN) and interaction data including state, action, and reward obtained through real-time interaction with environment (episode, or interaction data, includes state s, action a, and reward r that are obtained by an agent’s interaction with its environment, [0032] further teaches wherein the data is obtained from real-time interaction with an environment outside the computer system)).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the interaction data obtained through real-time interaction with the environment and system of Schmidhuber to the prior data, interaction data, and policy learning of Nair. Doing so would utilize data from interactions with an environment and data regarding previous command/control signals as well as other observed feedback to train, or at least help train, AI agents in an environment (Schmidhuber, [0032] & [0038]).
Claims 7-8 incorporate substantively all the limitations of claim 1 in an apparatus and method, and are rejected on similar grounds as above. Schmidhuber teaches the processors and memory of these claims at [0038] – “The computer system 100 of FIG. 1 has a computer-based processor 102, a computer-based storage device 104, and a computer-based memory 106. The computer-based memory 106 hosts an operating system and software that, when executed by the processor 102, causes the processor 102 to perform…”.
Regarding claim 4, the combination of Nair and Schmidhuber teaches the decision-making apparatus of claim 1, wherein
the processed prior data and the interaction data are stored in a single replay buffer (Nair, Section III Subsection D Paragraph 2 – “For every episode the agent experiences, we store it in the replay buffer twice: once with the original goal pursued in the episode and once with the goal corresponding to the final state achieved in the episode, as if the agent intended on reaching this state from the very beginning” – teaches wherein the prior data and interaction data are stored in a single replay buffer (every episode agent experiences is stored in replay buffer)).
Claim 11 is similar to claim 4, hence similarly rejected.
Regarding claim 5, the combination of Nair and Schmidhuber teaches the decision-making apparatus of claim 1, wherein
the processed prior data and the interaction data are stored in different replay buffers (Nair, Section IV Subsection A Paragraph 1 – “First, we maintain a second replay buffer RD where we store our demonstration data in the same format as R.” – teaches wherein the prior data and interaction data are stored in different replay buffers (teaches a second replay buffer RD different from first replay buffer R, thus data are stored in different replay buffers)).
Claim 12 is similar to claim 5, hence similarly rejected.
Regarding claim 6, the combination of Nair and Schmidhuber teaches the decision-making apparatus of claim 4, wherein
the policy learning unit learns policy by sampling the processed prior data and the interaction data from the single replay buffer or the different replay buffers by a preset number of samples (Nair, Section III Subsection B Paragraph 3 – “During each training step, DDPG samples a minibatch consisting of N tuples from R to update the actor and critic networks.” and in Section IV Subsection A Paragraph 1 – “In each minibatch, we draw an extra ND examples from RD to use as off-policy replay data for the update step. These examples are included in both the actor and critic update.” – teaches wherein the policy unit learns policy (actor-critic update) by sampling prior data and interaction data from the single replay buffer or the different replay buffers by a present number of samples (draws N or ND samples, thus a preset number of samples, from first replay buffer R and/or second replay buffer RD, thus sampling prior data and interaction data from the single replay buffer or the different replay buffers)).
Claim 13 is similar to claim 6, hence similarly rejected.
Claim(s) 2-3 and 9-10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Nair and Schmidhuber as applied to claims 1 and 7-8 above, and further in view of Oh et al. (NPL from IDS: Self-Imitation Learning, published June 2018, hereinafter “Oh”).
Regarding claim 2, the combination of Nair and Schmidhuber teaches the decision-making apparatus of claim 1.
The combination of Nair and Schmidhuber fails to explicitly teach wherein an objective function of policy network of the ego agent has a selective imitation learning term with a selective imitation learning weight that determines degree of imitation according to the magnitude of the reward of data sampled from the processed prior data and the interaction data.
However, analogous to the field of the claimed invention, Oh teaches: wherein
an objective function of policy network of the ego agent has a selective imitation learning term with a selective imitation learning weight that determines degree of imitation according to the magnitude of the reward of data sampled from the processed prior data and the interaction data (Oh, Algorithm 1 and Section 3 Subsection “Prioritized Replay” Paragraph 1 – “However, only good state-action pairs that satisfy R > Vθ can contribute to the gradient during self-imitation learning (Eq. 1)… This naturally increases the proportion of valid samples that satisfy the constraint (R − Vθ(s))+ in SIL objective and thus contribute to the gradient.” – teaches an objective function of policy network (SIL objective) of the ego agent has a selective imitation learning term (constraint of SIL objective) with a selective limitation learning weight (parameter θ) that determines degree of imitation according to the magnitude of the reward of data (R) sampled from the processed prior data and interaction data (in Algorithm 1, episode buffer is interaction data and replay buffer is prior data)).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the selective imitation learning term of Oh to the policy learning, agents, prior data, and interaction data of Nair and Schmidhuber. Doing so would provide algorithms that learn to imitate agent’s previous good decisions which lead to deeper exploration and learning to imitate state-action pairs only when the return in the past episode is greater than the agent’s value estimate (Oh, Introduction).
Claim 9 is similar to claim 2, hence similarly rejected.
Regarding claim 3, the combination of Nair, Schmidhuber, and Oh teaches the decision-making apparatus of claim 2, wherein
the selective imitation learning term in the objective function of the policy network is added when the reward of the sampled data is greater than a preset threshold (Oh, Section 3 Subsection “Prioritized Replay” Paragraph 1 – “However, only good state-action pairs that satisfy R > Vθ can contribute to the gradient during self-imitation learning (Eq. 1) … This naturally increases the proportion of valid samples that satisfy the constraint (R − Vθ(s))+ in SIL objective and thus contribute to the gradient.” – teaches wherein the selective imitation learning term in the objective function (constraint in SIL objective) of the policy network is added when the reward of the sampled data is greater than a preset threshold (selects only good state-action pairs that contribute to gradient during self-imitation learning, selection based on reward R greater than threshold Vθ)).
Therefore, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the selective imitation learning term and threshold of Oh to further modify the imitation learning, agents, prior data, and interaction data of Nair, Schmidhuber, and Oh. Doing so would provide algorithms that learn to imitate state-action pairs only when the return in the past episode is greater than the agent’s value estimate (Oh, Introduction).
Claim 10 is similar to claim 3, hence similarly rejected.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Shalev-Shwartz et al. (US Patent No. 10,514,705, published Dec. 2019) teaches a navigation system for a host vehicle that determines navigational constraints. Teaches autonomous vehicle navigation using reinforcement learning techniques. Teaches learning an initial policy using behavior cloning using real world data sets.
Moskovitz et al. (US Pub. No. 2024/0265263, filed Jan. 2024) teaches a method for iteratively training a policy model to control an agent interacting with an environment to perform a task subject to constraints. Teaches wherein each constraint may limit, to a corresponding threshold, the expected value of a corresponding constraint reward function. Teaches a history database for collecting accumulated trajectories and training performed parallel with action selection in online training. Teaches wherein the prior data collected is processed as a tuple storing a state, reward, action, and next state.
Wang et al. (NPL: SCRIMP: Scalable Communication for Reinforcement- and Imitation-Learning-Based Multi-Agent Pathfinding, published Oct. 2023) teaches a system for multi-agent pathfinding using reinforcement and imitation learning. Teaches a scalable and differentiable communication mechanism that allows agents to gain knowledge of other agents’ past observations, thereby mitigating risks posed by partial observations.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to LOUIS C NYE whose telephone number is 571-272-0636. The examiner can normally be reached Monday - Friday 9:00AM - 5:00PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, MATT ELL can be reached at 571-270-3264. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/LOUIS CHRISTOPHER NYE/Examiner, Art Unit 2141
/TAN H TRAN/Primary Examiner, Art Unit 2141