Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 21-22 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant regards), as the invention.
Claim 21 recites "a second state element subset comprises one or more but less than of the state elements" which appears to be missing a word. Compare the parallel first-subset clause, which recites "one or more but less than all of the state elements." It is unclear what "less than... of the state elements" means as written. For purposes of examination, this clause has been interpreted under BRI as "one or more but less than all of the state elements," consistent with the parallel first-subset clause.
Claim 21 further recites "determining a subset of potential next sub-actions for the first state element subset" but never establishes any corresponding subset of potential next sub-actions for the second state element subset. The later clause, "selecting a second sub-action from potential next sub-actions for the second state element subset", accordingly lacks antecedent basis, no earlier claim language establishes that any "potential next sub-actions for the second state element subset" exist. For purposes of examination, the "determining" clause has been interpreted under BRI as "...for the first state element subset and a subset of potential next sub-actions for the second state element subset."
Claim 22 is rejected as depending from rejected claim 21 and failing to cure the antecedent basis deficiencies identified above.Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-2, 7-14, 17-22, 27-29, and 33 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: Is the claim to a process, machine, manufacture or composition of matter?
Claims 1-2, 7-14, 17-22, and 27-29 are directed to a method (i.e., a process); and claim 33 is directed to an apparatus (i.e., a machine/apparatus); therefore, all pending claims are directed to one of the four categories of invention.
Independent Claims
Step 2A Prong One: Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes, independent claim 1 recites an abstract idea in the form of mental processes. A mental process is a process that “can be performed in the human mind, or by a human using a pen and paper” (MPEP§ 2106.04(a)(2)(III), paragraph 1). Examples of mental processes include “observations, evaluations, judgments, and opinions” (MPEP § 2106.04(a)(2)(III), paragraph 2).
The following limitations of claim 1 are mental processes:
evaluating a consequence of a previous action; [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for evaluating a consequence is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.]
based on the evaluated consequence of the previous action, determining a subset of potential next actions; [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for determining a subset is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.]
selecting an action from the determined subset of potential next actions; and [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for selecting an action is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.]
Therefore, the independent claims recite a judicial exception.
Step 2A Prong Two: Does the claim recite additional elements that integrate the judicial exception into a practical application?
No. The judicial exception recited in the above discussed claims is not integrated into a
practical application.
A method for reinforcement learning, the method comprising: [A method for reinforcement learning are components recited at a high level are construed as generic computer components used to implement the abstract idea. See MPEP 2106.05(f)(2). As such, the limitations do not integrate the abstract idea into a practical application. Nor to do they amount to significantly more.]
performing the selected action. [Performing an action are mere instructions to apply the abstract idea. Mere recitation that a judicial exception is to be performed using generic class of computer algorithms in their ordinary capacity, cannot meaningfully integrate the judicial exception into a practical application. See MPEP 2106.05(f).]
Therefore, under MPEP 2106.04(d), the additional elements of the claims do not integrate
the judicial exception into a practical application.
Step 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception?
No. The claims do not include additional elements that are sufficient for the claims to
amount to significantly more than the judicial exception.
Additional elements that are mere instructions to apply an exception or merely generally
linking or generally linking the use of a judicial exception to a particular technological
environment or field of use do not constitute significantly more than a judicial exception under
MPEP§2106.05(I)(A). Since the additional elements in the independent claims are all are mere
instructions to apply an exception or are merely generally linking or generally linking the use of
a judicial exception to a particular technological environment or field of use, they do not
constitute significantly more than a judicial exception.
Therefore, the additional elements identified in the Step 2A Prong Two analysis do not
constitute significantly more than a judicial exception.
Independent claim 33 recites the same relevant limitations and a similar analysis applies. Claim 33 recites the additional elements of “A reinforcement Learning (RL) agent comprising: processing circuitry; and a memory, the memory containing instructions executable by the processing circuitry, wherein the RL agent is configured to perform a process comprising:” [A memory containing instructions executable by the processing circuitry are components recited at a high level are construed as generic computer components used to implement the abstract idea. See MPEP 2106.05(f)(2). As such, the limitations do not integrate the abstract idea into a practical application. Nor to do they amount to significantly more.]
Therefore, the independent claims are not patent eligible.
Dependent Claims
The remaining dependent claims being rejected do not recite additional elements, whether considered individually or in combination, that are sufficient to integrate the judicial exception into a practical application or amount to significantly more than the judicial exception.
Claim 2
wherein evaluating the consequence of the previous action comprises performing a comparison of a set of one or more current monitored parameters to a set of one or more previous monitored parameters. [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for performing a comparison is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.]
Claim 7
wherein determining the subset of potential next actions comprises determining, for each potential next action, whether a dot product of a vector for the previous action and a vector for the potential next action is greater than a threshold. [This is a mathematical concept that describes mathematical relationships, mathematical formulas or equations, or mathematical calculations. The claim recites mathematical concept of a dot product.]
Claim 8
wherein, if the evaluated consequence of the previous action is a positive consequence, the determined subset of potential next actions comprises the potential next actions for which the dot product of the vector for the previous action and the vector for the potential next action is greater than the threshold. [This is a mathematical concept that describes mathematical relationships, mathematical formulas or equations, or mathematical calculations. The claim recites mathematical concept of a dot product.]
Claim 9
wherein, if the evaluated consequence of the previous action is a negative consequence, the determined subset of potential next actions comprises the potential next actions for which the dot product of the vector for the previous action and the vector for the potential next action is not greater than the threshold. [This is a mathematical concept that describes mathematical relationships, mathematical formulas or equations, or mathematical calculations. The claim recites mathematical concept of a dot product.]
Claim 10
wherein the threshold is 0. [This is a mathematical concept that describes mathematical relationships, mathematical formulas or equations, or mathematical calculations. The claim recites mathematical concept of a dot product.]
Claim 11
wherein determining the subset of potential next actions comprises determining, for each potential next action, whether an angle between a vector for the previous action and a vector for the potential next action is less than a threshold. [This is a mathematical concept that describes mathematical relationships, mathematical formulas or equations, or mathematical calculations. The claim recites mathematical concept of computing an angle.]
Claim 12
wherein, if the evaluated consequence of the previous action is a positive consequence, the determined subset of potential next actions comprises the potential next actions for which the angle between the vector for the previous action and the vector for the potential next action is less than the threshold. [This is a mathematical concept that describes mathematical relationships, mathematical formulas or equations, or mathematical calculations. The claim recites mathematical concept of computing an angle.]
Claim 13
wherein, if the evaluated consequence of the previous action is a negative consequence, the determined subset of potential next actions comprises the potential next actions for which the angle between the vector for the previous action and the vector for the potential next action is not less than the threshold. [This is a mathematical concept that describes mathematical relationships, mathematical formulas or equations, or mathematical calculations. The claim recites mathematical concept of computing an angle.]
Claim 14
wherein the threshold is π/2. [This is a mathematical concept that describes mathematical relationships, mathematical formulas or equations, or mathematical calculations. The claim recites mathematical concept of computing an angle.]
Claim 17
wherein the previous action and the potential next actions comprise state elements, and the vectors for the previous action and the potential next actions are based on all of the state elements. [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely indicates a field of use or technological environment in which a judicial exception is applied, namely the actions comprise state elements, and vectors are based on state elements.].
Claim 18
wherein the previous action and the potential next actions comprise state elements, and the vectors for the previous action and the potential next actions are based on a subset of the state elements. [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely indicates a field of use or technological environment in which a judicial exception is applied, namely the actions comprise state elements, and vectors are based on state elements.].
Claim 19
wherein the subset of the state elements comprise state elements that have inherent characteristics and/or a big impact on one or more performance metrics. [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely indicates a field of use or technological environment in which a judicial exception is applied, namely the subset of state elements have inherent characteristics or impact on performance metrics.].
Claim 20
wherein the state elements comprise x, y, and z-axis locations of a mobile base station (BS) and an antenna tilt value of the mobile BS, and the subset of the state elements comprises the x, y, and z-axis locations. [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely indicates a field of use or technological environment in which a judicial exception is applied, namely the state elements comprise x, y, and z-axis locations.].
Claim 21
wherein the previous action and the potential next actions comprise state elements, a first state element subset comprises one or more but less than all of the state elements, a second state element subset comprises one or more but less than of the state elements, the first and second state element subsets are different, determining the subset of potential next actions comprises determining a subset of potential next sub-actions for the first state element subset, and selecting an action from the determined subset of potential next actions comprises: [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely indicates a field of use or technological environment in which a judicial exception is applied, namely the actions comprise state elements].
selecting a first sub-action from the subset of potential next sub-actions for the first state element subset; [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for selecting a sub-action consequence is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.]
selecting a second sub-action from potential next sub-actions for the second state element subset; and [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for selecting a sub-action consequence is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.]
combining at least the first and second sub-actions. [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for combining a sub-action consequence is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.]
Claim 22
wherein the state elements comprise x, y, and z-axis locations of a mobile base station (BS) and an antenna tilt value of the mobile BS, the first state element subset comprises the x, y, and z-axis locations, and the second state element subset comprises the antenna tilt value. [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely indicates a field of use or technological environment in which a judicial exception is applied, namely the state elements comprise x, y, and z-axis locations.].
Claim 27
sending a message to one or more external nodes to request information reporting; and [sending a message is sending data, which is insignificant, extra-solution activity. See MPEP 2106.05(g). Transmitting data is well-understood, routine, and conventional. See MPEP 2106.05(d)(II)(i).]
receiving the requested information; [receiving information is sending data, which is insignificant, extra-solution activity. See MPEP 2106.05(g). Transmitting data is well-understood, routine, and conventional. See MPEP 2106.05(d)(II)(i).]
wherein determining the subset of potential next actions comprises using the requested information to reduce the number of potential next actions in the determined subset of potential next actions. [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely indicates a field of use or technological environment in which a judicial exception is applied, namely determining the subset comprises the requested information.].
Claim 28
determining whether to trigger sending the message to the one or more external nodes based on a current immediate reward, an accumulated reward in a current time window, an average reward in a current time window, and/or a value of one or more current key performance parameters. [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely indicates a field of use or technological environment in which a judicial exception is applied, namely the determination to sending a message is based on a current immediate reward, an accumulated reward, an average reward, or a value of a key performance parameter.].
Claim 29
evaluating a consequence of the selected action; and [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for evaluating a consequence is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.]
based on the evaluated consequence of the selected action, determining another subset of potential next actions. [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for determining a subset is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.]
The prior art used for rejections are provided below:
1. Knowledge-based Exploration for Reinforcement Learning in Self-Organizing Neural Networks (December 4, 2012) to Teng et al. (hereinafter Teng).
2. Autonomous Navigation and Configuration of Integrated Access Backhauling for UAV Base Station Using Reinforcement Learning (December 14, 2021) to Zhang et al. (hereinafter Zhang).
3. Learn What Not to Learn: Action Elimination with Deep Reinforcement Learning (September 6, 2018) to Zahavy et al. (hereinafter Zahavy).
4. Grid-based angle-constrained path planning (August 25, 2015) to Yakovlev et al. (hereinafter Yakovlev).
5. Teaching a Machine to Read Maps with Deep Reinforcement Learning (November 20, 2017) to Brunner et al. (hereinafter Brunner).
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1, 2, and 29 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Teng.
Per claim 1, Teng discloses: A method for reinforcement learning, the method comprising: [Teng, pg. 332 "This paper proposes a novel exploration strategy, known as Knowledge-based Exploration, for guiding the exploration of a family of self-organizing neural networks in reinforcement learning." (note: Teng discloses a method for reinforcement learning, as it is directed to an exploration strategy operating within a reinforcement learning process.)]
evaluating a consequence of a previous action; [Teng, pg. 332, Abstract "using the learned knowledge of the agent to identify prior action choices leading to low Q-values in similar situations." (note: Teng discloses evaluating the consequence (Q-value outcome) of a previous action choice, using the agent's learned knowledge of that action choice's outcome in similar past situations.);
pg. 332 "the learned knowledge is dichotomized into positive and negative chunks. While positive chunks refer to the knowledge of action choices that lead to desirable outcomes, negative chunks refer to the knowledge of action choices known to produce undesirable outcomes." (note: Teng discloses evaluating a consequence of a previous action by classifying it into a positive (desirable outcome) or negative (undesirable outcome) chunk based on that action's outcome.)]
based on the evaluated consequence of the previous action, determining a subset of potential next actions; [Teng, pg. 332, Abstract "exploration is directed towards unexplored and favorable action choices while steering away from those negative action choices that are likely to fail." (note: Teng discloses that based on the evaluated (negative) consequence, the possible next action choices considered for exploration is steered away from unfavorable choices, this is a subset of potential next actions which is determined based on the evaluated consequence.);
pg. 332 "the Knowledge-Based Exploration strategy directs the exploration of new action choices away from those action choices encoded by negative chunks for similar situations." (note: this discloses determining the subset of potential next actions based on the evaluated consequence, by excluding from that subset the action choices encoded as negative chunks for the situation.);
pg. 336, Algorithm 4 "Create a reduced action space Aʳ ≡ A⁺ ∪ Aᵘ" (note: Teng discloses the step of determining the subset of potential next actions (the reduced action space Aʳ) from the positive action set A⁺ and the unexplored action set Aᵘ, based on the prior evaluation of consequences.)]
selecting an action from the determined subset of potential next actions; and [Teng, pg. 336, Algorithm 4 "Randomly select an action choice a from Aʳ for exploration" (note: this discloses selecting an action from the determined subset of potential next actions, the reduced action space Aʳ)]
performing the selected action. [Teng, pg. 334, Algorithm 2, Line 9 "Use action choice a on state s for state s’" (note: this discloses performing the selected action, the agent applies th choses action choice a to a state s, transitioning to s’);
Algorithm 2, Line 5 “Use Exploration Strategy to select an action choice from action space” (note: this step, immediately preceding execution at line 9, is where Algorithm 2 invokes the Exploration Strategy, which is the procedure Teng sets out separately in Algorithm 4, which is relied upon above for the “determining a subset” and “selecting” claim 1 limitations. Algorithm 2 line 5 is the outer-loop call, and Algorithm 4 is the invoked procedure it calls.)]
Per claim 2, Teng discloses claim 1, further disclosing: wherein evaluating the consequence of the previous action comprises performing a comparison of a set of one or more current monitored parameters to a set of one or more previous monitored parameters. [Teng, pg. 334 "Upon receiving a feedback from the environment after performing the action, a TD formula is used to estimate the Q-value for performing the chosen action in the previous state." (note: discloses comparing current feedback received from the environment to the prior Q-value estimate using the Temporal Difference (TD) formula, which is a comparison of a current monitored parameter (feedback) to a previous monitored parameter (the prior Q-value estimate));
pg. 334 "TDₑᵣᵣ = r + γ maxₐ′ Q(s′, a′) − Q(s, a)" (note: the temporal error term itself is computed as a comparison between the current reward value and the existing (previous) Q-value estimate Q(s,a), which is a comparison of a current monitored parameter to a previous monitored parameter.)]
Per claim 29, Teng discloses claim 1, further disclosing: evaluating a consequence of the selected action; and
based on the evaluated consequence of the selected action, determining another subset of potential next actions. [Teng, pg. 334, Algorithm 2, line 10 "Evaluate effect of action choice a to derive a reward r from the environment";
pg. 334, Algorithm 2, line 14 "Repeat from Step 2 until s is a terminal state" (note: Algorithm 2, line 10 evaluates the consequence of the performed/selected action, and this evaluated reward goes into Algorithm 2’s own loop on line 14, which returns to line 5’s invocation of the Exploration Strategy (Algorithm 4), now operating on the updated knowledge produced by this evaluation, disclosing the determination of another subset of potential next actions based on the evaluated consequence)]
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 7, 8, 9, 10, 14, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Teng in view of Yakovlev and Brunner.
Per claim 7, Teng discloses claim 1.
Teng does not expressly disclose, but Teng combined with Yakovlev does teach:
wherein determining the subset of potential next actions comprises determining, for each potential next action, whether a…vector for the previous action and a vector for the potential next action is…a threshold [Yakovlev, pg. 7 “Then the potential successors of [a] are generated SUCC([a])=SUCC…After the set of potential successors is constructed it’s pruning is done…Third, the nodes that correspond to the cells that violate maximum angle of alteration constraints are discarded, e.g. the nodes [succi] that correspond to such cells succi: |α(⟨bp(a), a⟩, ⟨a, succi⟩)|>αm” (note: Yakovlev assembles, for the node the agent is expanding, a set of potential successor nodes — the moves available to it from that node — and then tests every candidate in that set individually against a threshold on the angle the candidate move makes with the move already performed into that node, discarding each candidate that fails the test so that the survivors are the reduced set the agent goes on to choose from, which constitutes determining the subset of potential next actions by determining, for each potential next action, whether a measure taken between a vector for the previous action and a vector for the potential next action stands in the prescribed relation to a threshold under BRI);
pg. 4 “Given two adjacent sections e1=⟨aij, alk⟩, e2=⟨alk, avw⟩ an angle of alteration is the angle between the vectors…which coordinates are (l - i, k - j) and (v - l, w - k) respectively” (note: the two quantities so compared are the vector of the section already traversed and the vector of the candidate section, each built from the coordinates of the cells the section runs between);
pg. 5 “while Theta*-LA validates also angle constraint, and if an angle between the sections defined by the trio: grandparent-parent-cell is greater than the predefined threshold αm, than parent cell is kept in the sequence” (note: Yakovlev names the value the comparison is made against a threshold)]
Teng and Yakovlev are analogous art because they are from the same field of endeavor, specifically automated agents that first construct a reduced set of candidate next moves and then take the move they perform from that reduced set. They are further reasonably pertinent to the same problem of sparing an agent the moves that the record of what it has already done marks as unprofitable.
Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to form the reduced action space of Teng by testing each candidate action, one by one, against a threshold on the direction it takes relative to the action already performed, in the manner Yakovlev describes, as claim 7 recites.
The suggestion/motivation for doing so would have been provided by Yakovlev itself, which teaches that testing every generated candidate against a threshold on the angle it makes with the move already performed, and discarding those that fail, is what yields the candidate set the agent then searches, [Yakovlev, pg. 7 “After the set of potential successors is constructed it’s pruning is done…Third, the nodes that correspond to the cells that violate maximum angle of alteration constraints are discarded”]. Teng already sorts its candidates against what its agent has learned about the action choices it has performed and carries only the survivors forward into the set it draws from, [Teng, pg. 336, Algorithm 4, line 12 “Create a reduced action space Ar ≡ A+ ∪ Au”, but leaves that sorting to the stored Q-value bounds of each candidate taken by itself, so a PHOSITA seeking a sorting criterion that relates each candidate directly to the action just performed would have looked to Yakovlev’s threshold test to supply one.
Teng combined with Yakovlev does not expressly disclose, but Teng combined with Brunner does teach:
…a dot product of…greater than… [Brunner, pg. 5, 3.5 Reactive Agent and Intrinsic Reward “For this we calculate an approximate two dimensional egomotion vector et from the egomotion probability distribution estimation st. Similarly we calculate a STTD vector dt−1 from the STTD distribution over {North, East, South, West} of the previous timestep. We calculate the exploitation intrinsic reward
I
t
e
x
p
l
o
i
t
as dot product between the two vectors” (note: Brunner forms one vector for the motion its reinforcement learning agent is making and a second vector for the direction that stood at the timestep before, takes the dot product of the two, and reads whether that dot product is greater than zero as the determination of whether the motion runs the same way as the earlier direction — expressly identifying that dot-product determination with the determination made on the angle between the very same two vectors — which constitutes determining whether a dot product of a vector for the previous action and a vector for the potential next action is greater than a threshold under BRI);
pg. 5, 3.5 Reactive Agent and Intrinsic Reward “Note that this reward is positive if and only if the angle difference between the two vectors is no bigger than 90 degrees, i.e., if the estimated egomotion was in the same direction as suggested by the STTD in the timestep before” (note: the sign of that dot product is expressly equated with the angle test on the two vectors)].
To the extent it is argued that Brunner's two vectors are not themselves a vector for a previous action and a vector for a potential next action, Brunner is not relied upon for those operands — Yakovlev supplies them, taking its angle between the vector of the move already performed and the vector of each candidate move [Yakovlev, pg. 4]. Brunner is relied upon only for the arithmetic by which that same directional determination is made, and for the threshold at which it is made.
Teng, Yakovlev and Brunner are analogous art because all three references reside within the same field of endeavor, specifically agents that decide which of the moves available to them to perform next by measuring a candidate move against the move already made. Brunner is further reasonably pertinent to the same problem of quantifying whether the motion an agent is about to make runs with or against the direction that stood before it.
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to carry out the directional threshold test of Teng combined with Yakovlev as a dot product taken between the vector for the previous action and the vector for the potential next action, in the manner Brunner describes, as claim 7 recites.
The suggestion/motivation for doing so would have been provided by Brunner itself, which teaches that the dot product of two such vectors is the measure of whether they run the same way and that its sign carries exactly the information the angle test carries, [Brunner, pg. 5, 3.5 Reactive Agent and Intrinsic Reward “Note that this reward is positive if and only if the angle difference between the two vectors is no bigger than 90 degrees, i.e., if the estimated egomotion was in the same direction as suggested by the STTD in the timestep before”]. Yakovlev states its own test on the angle taken between the two section vectors, [Yakovlev, pg. 4 “an angle of alteration is the angle between the vectors…which coordinates are (l - i, k - j) and (v - l, w - k) respectively”], and Brunner supplies the dot product as the arithmetic that returns that same determination directly from the two vectors without recovering the angle itself, so a PHOSITA implementing Yakovlev’s test on an agent of the kind Teng and Brunner both describe would have had reason to compute it as Brunner does.
Per claim 8, Teng-Yakovlev-Brunner disclose claim 7.
Teng combined with Brunner further teaches: wherein, if the evaluated consequence of the previous action is a positive consequence, [Teng, pg. 332, I. Introduction, "the learned knowledge is dichotomized into positive and negative chunks. While positive chunks refer to the knowledge of action choices that lead to desirable outcomes, negative chunks refer to the knowledge of action choices known to produce undesirable outcomes." (note: Teng discloses a positive/negative chunk dichotomy, actions associated with a positive (desirable) evaluated consequence are included in the subset, corresponding to the "greater than threshold" branch of the dot-product)] the determined subset of potential next actions comprises the potential next actions for which the dot product of the vector for the previous action and the vector for the potential next action is greater than the threshold [Brunner, pg. 5, 3.5 Reactive Agent and Intrinsic Reward “For this we calculate an approximate two dimensional egomotion vector et from the egomotion probability distribution estimation st. Similarly we calculate a STTD vector dt−1 from the STTD distribution over {North, East, South, West} of the previous timestep. We calculate the exploitation intrinsic reward
I
t
e
x
p
l
o
i
t
as dot product between the two vectors” (note: Brunner forms one vector for the motion its reinforcement learning agent is making and a second vector for the direction that stood at the timestep before, takes the dot product of the two, and reads whether that dot product is greater than zero as the determination of whether the motion runs the same way as the earlier direction — expressly identifying that dot-product determination with the determination made on the angle between the very same two vectors — which constitutes determining whether a dot product of a vector for the previous action and a vector for the potential next action is greater than a threshold under BRI);
pg. 5, 3.5 Reactive Agent and Intrinsic Reward “Note that this reward is positive if and only if the angle difference between the two vectors is no bigger than 90 degrees, i.e., if the estimated egomotion was in the same direction as suggested by the STTD in the timestep before” (note: the sign of that dot product is expressly equated with the angle test on the two vectors)].
The rationale to combine Teng with Brunner is the same as per claim 7.
Per claim 9, Teng-Yakovlev-Brunner disclose claim 7.
Teng further teaches: wherein, if the evaluated consequence of the previous action is a negative consequence, the determined subset of potential next actions comprises the potential next actions for which the dot product of the vector for the previous action and the vector for the potential next action is not greater than the threshold [Teng, pg. 332, I. Introduction, "the learned knowledge is dichotomized into positive and negative chunks. While positive chunks refer to the knowledge of action choices that lead to desirable outcomes, negative chunks refer to the knowledge of action choices known to produce undesirable outcomes." (note: Teng's negative-chunk branch shows actions associated with a negative (undesirable) evaluated consequence are excluded, corresponding to the "not greater than threshold" branch of the dot-product.)]
Per claim 10, Teng-Yakovlev-Brunner disclose claim 7.
Teng does not expressly disclose, but Teng combined with Brunner does teach: wherein the threshold is 0 [Brunner, pg. 5, 3.5 Reactive Agent and Intrinsic Reward, "Note that this reward is positive if and only if the angle difference between the two vectors is no bigger than 90 degrees, i.e., if the estimated egomotion was in the same direction as suggested by the STTD in the timestep before.." (note: Brunner discloses zero as the sign-boundary of a dot-product test, the reward is positive exactly when the dot product exceeds zero)]
The rationale to combine Teng with Brunner is the same as per claim 7.
Per claim 14, Teng-Yakovlev disclose claim 11.
Teng does not expressly disclose, but Teng combined with Brunner does teach: wherein the threshold is π/2 [Brunner, pg. 5, 3.5 Reactive Agent and Intrinsic Reward, "Note that this reward is positive if and only if the angle difference between the two vectors is no bigger than 90 degrees, i.e., if the estimated egomotion was in the same direction as suggested by the STTD in the timestep before.." (note: Brunner discloses 90 degrees (π/2 radians) as the angular boundary of a system, the point at which the underlying dot-product based reward changes sign)]
The rationale to combine Teng with Brunner is the same as per claim 7.
Per claim 17, Teng-Yakovlev-Brunner disclose claim 7.
Teng does not expressly disclose, but Teng combined with Yakovlev does teach: wherein the previous action and the potential next actions comprise state elements, and the vectors for the previous action and the potential next actions are based on all of the state elements [Yakovlev, pg. 4 “Given two adjacent sections e1=⟨aij, alk⟩, e2=⟨alk, avw⟩ an angle of alteration is the angle between the vectors…which coordinates are (l - i, k - j) and (v - l, w - k) respectively” (note: Yakovlev discloses that both the previous action vector and each potential next action vector are built from full cell-coordinate differences, both positional state elements are used in every instance, with no element ever dropped from the vector)]
The rationale to combine Teng with Yakovlev is the same as per claim 7.
Claims 11-13 are rejected under 35 U.S.C. 103 as being unpatentable over Teng in view of Yakovlev.
Per claim 11, Teng discloses claim 1.
Teng does not expressly disclose, but Teng combined with Yakovlev does teach: wherein determining the subset of potential next actions comprises determining, for each potential next action, whether an angle between a vector for the previous action and a vector for the potential next action is less than a threshold [Yakovlev, pg. 7 “Then the potential successors of [a] are generated SUCC([a])=SUCC… Third, the nodes that correspond to the cells that violate maximum angle of alteration constraints are discarded, e.g. the nodes [succi] that correspond to such cells succi: |α(⟨bp(a), a⟩, ⟨a, succi⟩)|>αm” (note: Yakovlev discloses testing, for each of multiple candidate next actions, whether the angle between the vector of the previous action ⟨bp(a), a⟩ and the vector of the candidate ⟨a, succi⟩ exceeds a threshold, the complement of this claims “less than a threshold” test, disclosing the same comparison structure);
pg. 4 “Given two adjacent sections e1=⟨aij, alk⟩, e2=⟨alk, avw⟩ an angle of alteration is the angle between the vectors…which coordinates are (l - i, k - j) and (v - l, w - k) respectively” (note: this discloses the two vectors, for the previous action and for the potential next action, between which the angle is measured)]
The rationale to combine Teng with Yakovlev is the same as per claim 7.
Per claim 12, Teng-Yakovlev disclose claim 11.
Teng combined with Yakovlev further teaches: wherein, if the evaluated consequence of the previous action is a positive consequence, [Teng, pg. 332, I. Introduction, "the learned knowledge is dichotomized into positive and negative chunks. While positive chunks refer to the knowledge of action choices that lead to desirable outcomes, negative chunks refer to the knowledge of action choices known to produce undesirable outcomes." (note Teng's positive chunk branch is the positive consequence branch, corresponding to the "less than threshold" branch of the angle test.)] the determined subset of potential next actions comprises the potential next actions for which the angle between the vector for the previous action and the vector for the potential next action is less than the threshold [Yakovlev, pg. 7 “Then the potential successors of [a] are generated SUCC([a])=SUCC…After the set of potential successors is constructed it’s pruning is done…Third, the nodes that correspond to the cells that violate maximum angle of alteration constraints are discarded, e.g. the nodes [succi] that correspond to such cells succi: |α(⟨bp(a), a⟩, ⟨a, succi⟩)|>αm” (note: Yakovlev assembles, for the node the agent is expanding, a set of potential successor nodes — the moves available to it from that node — and then tests every candidate in that set individually against a threshold on the angle the candidate move makes with the move already performed into that node, discarding each candidate that fails the test so that the survivors are the reduced set the agent goes on to choose from, which constitutes determining the subset of potential next actions by determining, for each potential next action, whether a measure taken between a vector for the previous action and a vector for the potential next action stands in the prescribed relation to a threshold under BRI);
pg. 4 “Given two adjacent sections e1=⟨aij, alk⟩, e2=⟨alk, avw⟩ an angle of alteration is the angle between the vectors…which coordinates are (l - i, k - j) and (v - l, w - k) respectively” (note: the two quantities so compared are the vector of the section already traversed and the vector of the candidate section, each built from the coordinates of the cells the section runs between);
The rationale to combine Teng with Yakovlev is the same as per claim 7.
Per claim 13, Teng-Yakovlev disclose claim 11.
Teng further teaches: wherein, if the evaluated consequence of the previous action is a negative consequence, the determined subset of potential next actions comprises the potential next actions for which the angle between the vector for the previous action and the vector for the potential next action is not less than the threshold [Teng, pg. 332, I. Introduction "the learned knowledge is dichotomized into positive and negative chunks. While positive chunks refer to the knowledge of action choices that lead to desirable outcomes, negative chunks refer to the knowledge of action choices known to produce undesirable outcomes." (note: Teng's negative chunk branch is the negative consequence branch, corresponding to the "not less than threshold" branch of the angle test.)]
Claims 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Teng in view of Yakovlev, Brunner, and Zhang.
Per claim 18, Teng-Yakovlev-Brunner disclose claim 7.
Teng does not expressly disclose, but Teng combined with Zhang does teach: wherein the previous action and the potential next actions comprise state elements, and the vectors for the previous action and the potential next actions are based on a subset of the state elements [Zhang, pg. 3, IV. Reinforcement Learning Algorithm Design "σₜ represents the electrical tilt value of the access and backhaul antenna, and {xₜ,yₜ,zₜ} denotes the 3-D location of the UAV-BS" (note: discloses a four element state vector (tilt plus x, y, z location) and, combined with the impact analysis quotation below, the basis for using only a subset (for example, x, y, z alone) of that state vector rather than all four elements.);
pg. 5, V. Performance Evaluation "both antenna tilt and UAV-BS height significantly impact drop rate performance" (note: this discloses that different state dimensions have unequal impact on performance metrics, which supports the use of only a subset of the full state-element set rather than all elements.)]
Teng and Zhang are analogous art because they are from the same field of endeavor of reinforcement learning based automated decision and control systems. They are further reasonably pertinent to the same problem of using external information to inform an RL agent’s evaluation of its state and future action selection.
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to incorporate Zhang’s data request and report signaling procedure into Teng’s reinforcement learning framework.
The suggestion/motivation for doing so would have been to increase optimization efficiency by incorporating Zhang’s signaling mechanism. [Zhang, pg. 3 “to assist the UAV-BS in optimizing its configuration and deployment”]
Per claim 19, Teng-Yakovlev-Brunner-Zhang disclose claim 18.
Teng does not expressly disclose, but Teng combined with Zhang does teach: wherein the subset of the state elements comprise state elements that have inherent characteristics and/or a big impact on one or more performance metrics [Zhang, pg. 5, V. Performance Evaluation "both antenna tilt and UAV-BS height significantly impact drop rate performance" (note: Zhang discloses state elements (tilt, height) significantly impact a performance metric (drop rate))]
The rationale to combine Teng with Zhang is the same as per claim 18.
Per claim 20, Teng-Yakovlev-Brunner-Zhang disclose claim 18.
Teng does not expressly disclose, but Teng combined with Zhang does teach: wherein the state elements comprise x, y, and z-axis locations of a mobile base station (BS) and an antenna tilt value of the mobile BS, and the subset of the state elements comprises the x, y, and z-axis locations [Zhang, pg. 3, IV. Reinforcement Learning Algorithm Design "σₜ represents the electrical tilt value of the access and backhaul antenna, and {xₜ,yₜ,zₜ} denotes the 3-D location of the UAV-BS" (note: Zhang discloses the identical x, y, z-axis locations and antenna tilt state element, with the x, y, z locations forming the recited subset.)]
The rationale to combine Teng with Zhang is the same as per claim 18.
Claims 27 and 28 are rejected under 35 U.S.C. 103 as being unpatentable over Teng in view of Zhang and Zahavy.
Per claim 27, Teng discloses claim 1.
Teng does not expressly disclose, but Teng combined with Zhang does teach:
sending a message to one or more external nodes to request information reporting; and [Zhang, pg. 3 " the donor-BS triggers data collection... by sending a data request message to the relevant users and BSs " (note: Zhang names the donor-BS, not the UAV-BS, as the sender. Under broadest reasonable interpretation, system-level reading, claim 27 recites a method comprising "sending... and receiving..." without requiring the same physical node to perform both acts)]
receiving the requested information [Zhang, pg. 3 "MC users, the donor-BS, and related on-ground BSs will send the requested data to the UAV-BS" (note: discloses receiving the requested information — verbatim, under the entity mapping used throughout this combination, where the UAV-BS corresponds to the claimed RL agent that receives.)]
Teng and Zhang are analogous art because they are from the same field of endeavor of reinforcement learning based automated decision and control systems. They are further reasonably pertinent to the same problem of using external information to inform an RL agent’s evaluation of its state and future action selection.
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to incorporate Zhang’s data request and report signaling procedure into Teng’s reinforcement learning framework.
The suggestion/motivation for doing so would have been to increase optimization efficiency by incorporating Zhang’s signaling mechanism. [Zhang, pg. 3 “to assist the UAV-BS in optimizing its configuration and deployment”]
Teng combined with Zhang does not expressly disclose, but Teng combined with Zahavy does teach:
wherein determining the subset of potential next actions comprises using the requested information to reduce the number of potential next actions in the determined subset of potential next actions [Zahavy, pg. 2 "restricting the available actions in each state to a subset of the most likely ones" (note: discloses using information to restrict (reduce) the number of potential next actions in the determined subset.);
pg. 2, Figure 1 caption "admissible actions set A′ " (note: this discloses the reduced potential next action structure, the admissible actions set A′);
pg. 7, Algorithm 1 ”A′ ← {a : E(s)ₐ − √(β·φ(s)ᵀV⁻¹φ(s)) < ℓ}” (note: computed from a per-step information signal and used to constrain the potential action set.)]
Teng and Zahavy are analogous art because they are from the same field of endeavor of RL frameworks that restrict candidate action spaces. They are further reasonably pertinent to the same problem of improving the efficiency of action selection in large discrete action spaces.
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to implement Teng’s subset determination step using Zahavy’s admissible actions set feature.
The suggestion/motivation for doing so would have been that Zahavy explicitly identifies improvement in efficiency. [Zahavy, pg. 5 “elimination algorithms can eliminate these actions substantially faster, and can, therefore, speed up the learning process approximately by A=A’ ”]
Per claim 28, Teng-Zhang-Zahavy disclose claim 27.
Teng further teaches: determining whether to trigger sending the message to the one or more external nodes based on a current immediate reward, an accumulated reward in a current time window, an average reward in a current time window, and/or a value of one or more current key performance parameters [Teng, pg. 334, Section III.B "Upon receiving a feedback from the environment after performing the action, a TD formula is used to estimate the Q-value for performing the chosen action in the previous state." (note: this discloses ‘a current immediate reward’. Teng’s quantity ‘r’ is received as feedback upon performing a single action and is used directly in the TD error computation stated in claim 2. Teng’s ‘r’ is applied as the reward based signal determination.)]
Claim 33 is rejected under 35 U.S.C. 103 as being unpatentable over Teng in view of Zahavy.
Per claim 33, Teng discloses: evaluating a consequence of a previous action; [Teng, pg. 332, Abstract "using the learned knowledge of the agent to identify prior action choices leading to low Q-values in similar situations." (note: Teng discloses evaluating the consequence (Q-value outcome) of a previous action choice, using the agent's learned knowledge of that action choice's outcome in similar past situations.);
pg. 332 "the learned knowledge is dichotomized into positive and negative chunks. While positive chunks refer to the knowledge of action choices that lead to desirable outcomes, negative chunks refer to the knowledge of action choices known to produce undesirable outcomes." (note: Teng discloses evaluating a consequence of a previous action by classifying it into a positive (desirable outcome) or negative (undesirable outcome) chunk based on that action's outcome.)]
based on the evaluated consequence of the previous action, determining a subset of potential next actions; [Teng, pg. 332, Abstract "exploration is directed towards unexplored and favorable action choices while steering away from those negative action choices that are likely to fail." (note: Teng discloses that based on the evaluated (negative) consequence, the possible next action choices considered for exploration is steered away from unfavorable choices, this is a subset of potential next actions which is determined based on the evaluated consequence.);
pg. 332 "the Knowledge-Based Exploration strategy directs the exploration of new action choices away from those action choices encoded by negative chunks for similar situations." (note: this discloses determining the subset of potential next actions based on the evaluated consequence, by excluding from that subset the action choices encoded as negative chunks for the situation.);
pg. 336, Algorithm 4 "Create a reduced action space Aʳ ≡ A⁺ ∪ Aᵘ" (note: Teng discloses the step of determining the subset of potential next actions (the reduced action space Aʳ) from the positive action set A⁺ and the unexplored action set Aᵘ, based on the prior evaluation of consequences.)]
selecting an action from the determined subset of potential next actions; and [Teng, pg. 336, Algorithm 4 "Randomly select an action choice a from Aʳ for exploration" (note: this discloses selecting an action from the determined subset of potential next actions, the reduced action space Aʳ)]
performing the selected action. [Teng, pg. 336, Algorithm 4 " return action choice a," (note: this discloses performing the selected action choice)]
Teng does not expressly disclose, but Teng combined with Zahavy does teach:
A reinforcement Learning (RL) agent comprising: [Teng, p. 332 “This paper proposes a novel exploration strategy, known as Knowledge-based Exploration, for guiding the exploration of a family of self-organizing neural networks in reinforcement learning.” (note: Teng discloses a reinforcement learning agent, the self-organizing neural network that carries out Teng’s exploration strategy is the agent operating within a reinforcement learning process)]
processing circuitry; and [(note: this is inherently part of the memory containing executable instructions and the stated process arrangment)]
a memory, [Teng, pg. 333 “The knowledge layer has the category field
F
2
c
for storing the committed and uncommitted cognitive nodes.” (note: Teng discloses a data-storage structure (memory))] the memory containing instructions executable by the processing circuitry, wherein the RL agent is configured to perform a process comprising: [Zahavy, pg. 7 “Initialize Replay Memory D to capacity N… The agent uses an Experience Replay (Lin, 1992) to store information about states, transitions, actions, and rewards” (note: this discloses the memory (Replay Memory D) holds content that is operated on by the agent’s defined process. )]
Teng and Zahavy are analogous art because they are from the same field of endeavor of RL frameworks that restrict candidate action spaces. They are further reasonably pertinent to the same problem of improving the efficiency of action selection in large discrete action spaces.
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to implement Teng’s subset determination step using Zahavy’s admissible actions set feature.
The suggestion/motivation for doing so would have been that Zahavy explicitly identifies improvement in efficiency. [Zahavy, pg. 5 “elimination algorithms can eliminate these actions substantially faster, and can, therefore, speed up the learning process approximately by A=A’ ”]
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Sayed M Shah whose telephone number is (571)272-9406. The examiner can normally be reached Monday-Friday 9:00 am - 5:00 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Miranda Huang can be reached at (571) 270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SAYED MUNEER SHAH/Examiner, Art Unit 2124
/ALAN CHEN/Primary Examiner, Art Unit 2125