Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Preliminary Amendments
This action is in response to preliminary amendments filed May 24th, 2024, in which Claims 1, 3, 4, 8, and 12 are amended. Claims 10, 11, 13, and 14 have been cancelled. Claims 15-23 are added. The amendments have been entered, and Claims 1-9, 12, and 15-23 are currently pending.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 5-7, 9, 18-20, and 22 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claims 6 and 19 recite a function
f
a
without providing any explanation about what the function might be, rendering the claim indefinite. Further, Claims 6, 7, 19, and 20 recite the non-standard operator
*
without providing any explanation as to what operation
*
might refer, and one of ordinary skill in the art would not understand whether
*
means scalar multiplication, vector multiplication, dot product, convolution, or any of several other possible operations, thus rendering the claim indefinite.
Further, the expression for
γ
i
in each of Claims 5-7 and 18-20, i.e. a discount value for the ith agent, does not depend on i in the expressions given, which makes the claim indefinite as to whether the discount factor for each different agent is required to be unique for each agent, as implied in Claims 1 and 3, or can be the same for each agent, as is implied in Claims 5-7 and 18-20.
Claims 9 and 22 recite the limitation the discount value. There is a lack of proper antecedent basis in the claims for this limitation – only a discount factor has previously been recited. For the purpose of examination, the claims will be interpreted as if the discount value refers to the discount factor of the independent claims.
Dependent claims are rejected for inheriting the indefiniteness of a parent claim.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-9, 12, and 15-23 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Claim 1 recites a method, thus a process, one of the four statutory categories of patentable subject matter. However, the claim further recites a step of determining a discount factor for each agent included in [a] group based on local observation data and global state data which, as a judgement, falls into the mental process grouping of abstract ideas. Therefore, the claim recites an abstract idea of determining a discount factor for each of a group of agents.
The claim does not recite any additional elements which could integrate the abstract idea into a practical application, because the additional elements consist of:
obtaining local observation data of a group of one or more agents, which is insignificant extra-solution activity required for all uses of the abstract idea (MPEP 2106.05(g))
obtaining global state data indicating collective performance of the group, which is insignificant extra-solution activity required for all uses of the abstract idea (MPEP 2106.05(g))
wherein the discount factor is a weight value of a future expected reward for each agent included in the group, which merely specifies the data which is to be determined in the execution of the abstract idea, i.e. the particular technological environment or field of use (MPEP 2106.05(h))
neither of which can integrate the abstract idea into a practical application. Therefore, the claim is directed to the abstract idea of determining a discount factor for each of a group of agents.
Further, the additional elements, taken alone and in combination, cannot provide significantly more than the abstract idea itself because the obtaining data steps are well-understood, routine, and conventional (by MPEP 2106.05(d), “transmitting and receiving data over a network), because specifying a field of use cannot do so (by MPEP 2106.05(h)), and because there is no nexus between the additional elements which could provide significantly more in combination. Thus, the claim is subject-matter ineligible.
Claims 2, 3, and 5-7, dependent upon Claim 1, recite only particulars about the data which is to be determined in the execution of the abstract idea, or used in the determination, i.e. the particular technological environment or field of use, which by MPEP 2106.05(h) can neither integrate the abstract idea into a practical application nor provide significantly more than the abstract idea itself.
Claim 4, dependent upon Claim 1, recites new additional elements of obtaining a plurality of weights¸ which is insignificant extra-solution activity of data gathering which can neither integrate the abstract idea into a practical application nor provide significantly more than the abstract idea itself (MPEP 2106.05(g) & MPEP 2106.05(d), “transmitting or receiving data over a network”), and of applying the obtained plurality of weights to the local observation data via the prediction neural network, which is merely using a computer or other machinery as a tool to perform the abstract idea, which by MPEP 2106.05(f)(2) neither integrates the abstract idea into a practical application nor provides significantly more than the abstract idea itself.
Claim 8, dependent upon Claim 4, further recites an additional mental process step (determining the plurality of weights) performed by a computer or other machinery as a tool (using a hypernetwork) which by MPEP 2106.05(f)(2) neither integrates the abstract idea into a practical application nor provides significantly more than the abstract idea itself.
Claim 9, dependent upon Claim 1, merely recites an additional mental process step (determining a prediction cumulative reward function for each agent in the group), but no additional elements, thus no additional elements which could integrate the abstract idea into a practical application nor provide significantly more than the abstract idea.
Claims 12 and 15-22 recite an apparatus comprising a memory and processing circuitry to perform precisely the methods recited in Claims 1-9, respectively. As performance of an abstract idea using generic computer components cannot integrate the abstract idea into a practical application nor provide significantly more than the abstract idea itself (MPEP 2106.05(f)(2)), Claims 12 and 15-22 are rejected for reasons set forth in the rejections of Claims 1-9, respectively. Similarly, 23 recites a non-transitory computer readable medium storing instructions to execute the method of Claim 1, and is thus similarly rejected based on MPEP 2106.05(f)(2) and the reasons set forth in the rejection of Claim 1.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
1-9, 12, and 15-23 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Chu et al., “Multi-agent reinforcement learning for networked system control.”
Regarding Claim 1, Chu teaches a method comprising: obtaining local observation data of a group of one or more agents, wherein the local observation data indicates performance of each agent in the group (Chu, title, “Multi-agent reinforcement learning” & pg. 3, 1st paragraph, “
S
i
and
A
i
are the local state space and action space of agent i … policy to choose its own action [as a function of current, thus observed and indicating performance of that agent, state]
s
i
,
t
at time t” & 2nd paragraph, “we enforce practical constraints only allow local observations and neighborhood communication”); obtaining global state data indicating collective performance of the group (Chu, pg. 3, 3rd paragraph, “All local rewards are shared globally [and] each agent i observes [message data from its neighborhood, i.e. global state data since it is not from the local node]”); based on the obtained local observation data and the obtained global state data, determining a discount factor for each agent in the group, wherein the discount factor is a weight value of a future expect reward for each agent included in the group (Chu, Abstract, “a spatial discount factor” & pg. 3, Eq. (2), “
α
d
i
j
” is a weight value of a future expected reward, see 1st paragraph, “to maximize [expected reward]” and is dependent upon the distance between the local agent and every other node, i.e. local observation data and global state data).
Regarding Claim 2, Chu teaches the method of Claim 1 (and thus the rejection of Claim 1 is incorporated). Chu further teaches wherein the local observation data indicates the performance of each agent in a current state of an environment, the global state indicates the collective performance of the group in the current state of the environment (the distance is determined by where the agents are, e.g. performance, see pg. 5, “Environment Setup … cooperative adaptive cruise control”); the discount factor is for a next sequential state of the environment that is after the current state of the environment (Chu, pg. 3, Eq. (2) the spatial discount factor is applied for all future states).
Regarding Claim 3, Chu teaches the method of Claim 1 (and thus the rejection of Claim 1 is incorporated). Chu further teaches wherein the group of agents includes a first agent and a second agent (Chu, title, “Multi-agent reinforcement learning,”), the local observation data indicates that performance of the first agent deviates from performance of the agents included in the group by a first degree, the local observation data indicates that the performance of the second agent deviates from the performance of the agents included in the group by a second degree, the first degree is greater than the second degree, and the discount factor of the first agent is less than the discount factor of the second agent (Chu, pg. 3, Eq. (2), “
α
d
i
j
” where
0
≤
α
≤
1
indicates that when distance, i.e. deviation from the performance of the other agents, is larger, then the discount factor
α
d
i
j
will be smaller).
Regarding Claim 4, Chu teaches the method of Claim 1 (and thus the rejection of Claim 1 is incorporated). Chu further teaches wherein determining the discount factor for each agent comprises: obtaining a plurality of weights of a prediction network; and applying the obtained plurality of weights to the local observation data via the prediction neural network, thereby determining the discount factors for the agents included in the group (Chu, pg. 4, 2nd paragraph, “Now we assume each agent is A2C, with parametric models for fitting the optimal policy and value function” whose weights
θ
i
are learned via Eq.(3); and actions are obtained by applying local observation state data to the policy network, see pg. 3, 1st paragraph, “each agent i follows a decentralized policy
π
i
” as a function of the state/observation data, which determines actions and thus position and thus distance to other agents, thereby determining the discount factors).
Regarding Claim 5, Chu teaches the method of Claim 4 (and thus the rejection of Claim 4 is incorporated). Chu further teaches wherein a discount value for each agent is a function of the weights and observations at a given time, which has already been shown in the rejection of Claim 4.
Regarding Claim 6, Chu teaches the method of Claim 5 (and thus the rejection of Claim 5 is incorporated). Chu further teaches wherein a discount value for each agent is a function of the product of observation data and weights at a given time (Chu, pg. 5, Fig. 1(a), at each agent, the local observation data
s
i
,
t
is encoded via a DNN, which applies weights to the input data and then performs further functions, to obtain actions, and thus the discount factor).
Regarding Claim 7, Chu teaches the method of Claim 6 (and thus the rejection of Claim 6 is incorporated). Chu further teaches wherein a discount value for each agent is a function of the product of observation data and weights at a given time (Chu, pg. 5, Fig. 1(a) at each agent, the local observation data
s
i
,
t
is encoded via a DNN, which applies weights to the input data and sums them to obtain a discount value for each agent – note that the claim does not require the discount factor to be calculated in this manner, but only a discount value).
Regarding Claim 8, Chu teaches the method of Claim 4 (and thus the rejection of Claim 4 is incorporated). Chu further teaches wherein obtaining the plurality of weights of the prediction network comprises determining the plurality of weights using a hypernetwork and global state data (Chu, pg. 4, last paragraph, “using a single meta-DNN” where the state data is input, see pg. 5, Fig. 1, where “meta-DNN” denotes a hypernetwork that is learned, i.e. where the weights of the prediction network are obtained).
Regarding Claim 9, Chu teaches the method of Claim 1 (and thus the rejection of Claim 1 is incorporated). Chu further teaches determining a prediction cumulative reward function for each agent included in the group using a current reward value of each agent, future reward values of each agent, and the discount value associated with each agent (Chu, pg. 3, Eq. (2) depends on all these values and is used to learn/determine the value function).
Claims 12 and 15-22 recite an apparatus comprising a memory and processing circuitry to perform precisely the methods recited in Claims 1-9, respectively. As Chu performs their method on a computer (Chu, pg. 2, footnote 1), Claims 12 and 15-22 are rejected for reasons set forth in the rejections of Claims 1-9, respectively. Similarly, 23 recites a non-transitory computer readable medium storing instructions to execute the method of Claim 1, and is thus similarly rejected based on Chu’s use of a computer and the reasons set forth in the rejection of Claim 1.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure: Liu et al., “Emergent coordination through competition,” discloses learning discount factors in a multi-agent environment based on observation data.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to BRIAN M SMITH whose telephone number is (469)295-9104. The examiner can normally be reached Monday - Friday, 8:00am - 4pm Pacific.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kakali Chaki can be reached at (571) 272-3719. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/BRIAN M SMITH/Primary Examiner, Art Unit 2122