DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This action is in response to an application filed on March 12th, 2024. Claims 1-20 are pending in the current application The IDS has been considered.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim(s) 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding claim 1, Under Step 1 of the Subject Matter Eligibility Test of Products and Processes, the claim is directed towards a process which is one of the four statutory categories.
Next, under a Step 2A Prong 1 Analysis, the claim recites the following limitations which are interpreted to be, under the broadest reasonable interpretation, abstract ideas.
generating the learned control policy for the environment… wherein the learned control policy provides a Q table comprising values specifying actions for an agent to take based on a state of the agent in order to complete a task. (mental process done in a computing environment)
detecting a change to the environment at a first location in the environment
defining a local region surrounding the first location
modifying the local Q table… based on the detected change to the environment
generating a diffusion model for propagating changes made in the local Q table across the Q table
and propagating… the changes made in the local Q table globally across the Q table to modify the learned control policy.
Therefore, we have to examine the claim under Step 2A prong 2, which considers the additional elements within the claim. The claim’s additional elements are:
using a reinforcement learning process
wherein the local region corresponds with a local Q table that is part of the Q table
using the reinforcement learning process
using the diffusion model
The limitations, “using a reinforcement learning process”, “using a reinforcement learning process”, and “using the diffusion model” are considered to be mere instructions to apply a judicial exception, as it instructs to use a reinforcement learning process and diffusion model to perform the abstract ideas. (See MPEP 2106.05(f)) The limitation, “the local region corresponds with a local Q table that is part of the Q table” merely indicates the field of use and technological environment, and “generally links” a local region corresponding with a local Q table that is part of the Q table, to the abstract idea. (See MPEP 2106.05(h)) Therefore, these additional elements do not integrate the abstract idea into a practical application. The claim is directed towards an abstract idea.
Under a Step 2B analysis, the claim’s additional elements do not amount to significantly
more than the judicial exception as explained above in Step 2A prong 2. Therefore, the claim is ineligible.
Regarding claim 10, Under Step 1 of the Subject Matter Eligibility Test of Products and Processes, the claim is directed towards a machine which is one of the four statutory categories.
Next, under a Step 2A Prong 1 Analysis, the claim recites the following limitations which are interpreted to be, under the broadest reasonable interpretation, abstract ideas.
detect a change to an environment at a first location in the environment wherein a learned control policy for the environment includes a Q table that comprises values specifying actions for an UAV to take based on a state for the UAV to complete a task
defining a local region surrounding the first location
modify the local Q table using a Q learning process based on the change to the environment
and propagate any changes made in the local Q table globally across the Q table to modify the learned control policy.
Therefore, we have to examine the claim under Step 2A prong 2, which considers the additional elements within the claim. The claim’s additional elements are:
at least one unmanned aerial vehicle (UAV)
and a computer in electrical communication with the at least one UAV,
wherein the local region corresponds with a local Q table that is part of the Q table
The limitations, “at least one unmanned aerial vehicle (UAV)”, and “a computer in electrical communication with the at least one UAV,” are considered to be mere instructions to apply a judicial exception, as it instructs to use a reinforcement learning process and diffusion model to perform the abstract ideas. (See MPEP 2106.05(f)) The limitation, “the local region corresponds with a local Q table that is part of the Q table” merely indicates the field of use and technological environment, and “generally links” a local region corresponding with a local Q table that is part of the Q table, to the abstract idea. (See MPEP 2106.05(h)) Therefore, these additional elements do not integrate the abstract idea into a practical application. The claim is directed towards an abstract idea.
Under a Step 2B analysis, the claim’s additional elements do not amount to significantly
more than the judicial exception as explained above in Step 2A prong 2. Therefore, the claim is ineligible.
Regarding claim 15, Under Step 1 of the Subject Matter Eligibility Test of Products and Processes, the claim is directed towards a manufacture, which is one of the four statutory categories.
Next, under a Step 2A Prong 1 Analysis, the claim recites the following limitations which are interpreted to be, under the broadest reasonable interpretation, abstract ideas.
generating a learned control policy for an environment using a reinforcement learning process, wherein the learned control policy provides a Q table comprising values specifying actions for an agent to take based on a state of the agent in order to complete a task. (mental process done in a computing environment)
detecting a change to the environment at a first location in the environment
defining a local region surrounding the first location
modifying the local Q table… based on the detected change to the environment
generating a diffusion model for propagating changes made in the local Q table across the Q table
and propagating… the changes made in the local Q table globally across the Q table to modify the learned control policy.
Therefore, we have to examine the claim under Step 2A prong 2, which considers the additional elements within the claim. The claim’s additional elements are:
using a reinforcement learning process
wherein the local region corresponds with a local Q table that is part of the Q table
using the reinforcement learning process
using the diffusion model
The limitations, “using a reinforcement learning process”, “using a reinforcement learning process”, and “using the diffusion model” are considered to be mere instructions to apply a judicial exception, as it instructs to use a reinforcement learning process and diffusion model to perform the abstract ideas. (See MPEP 2106.05(f)) The limitation, “the local region corresponds with a local Q table that is part of the Q table” merely indicates the field of use and technological environment, and “generally links” a local region corresponding with a local Q table that is part of the Q table, to the abstract idea. (See MPEP 2106.05(h)) Therefore, these additional elements do not integrate the abstract idea into a practical application. The claim is directed towards an abstract idea.
Under a Step 2B analysis, the claim’s additional elements do not amount to significantly
more than the judicial exception as explained above in Step 2A prong 2. Therefore, the claim is ineligible.
Regarding claims 2 and 16, the claims recite the limitation: “wherein the agent comprises an unmanned aerial vehicle (UAV).” The limitation, as drafted, merely indicates the field of use and technological environment, and “generally links” a UAV to the abstract idea. (See MPEP 2106.05(h)) Therefore, the claims are rejected on the basis as claims 1 and 15.
Regarding claims 3 and 17, the claims recite “the task comprises the UAV moving from an initial location to a target location in the environment.” The limitation, as drafted, is considered to be merely indicating the field of use and technological environment, and “generally links” the specific UAV tasks to the abstract idea. (See MPEP 2106.05(h)) Therefore, the claims are rejected on the basis as claims 2 and 16.
Regarding claims 4 and 18, the claims recite “identifying the state of the UAV based on a present location of the UAV in the environment; and determining an action for the UAV based on mapping the state of the UAV to the Q table.” The limitations, as drafted, are considered to be, under the broadest reasonable interpretation, “mental processes”, performed with the aid of a computer used as tool to perform the abstract idea. Therefore, the claims are rejected on the basis as claims 3 and 17.
Regarding claim 5, the claim recites “a reward is issued upon completion of the task, and wherein an amount of the reward is determined based on a length of a path traveled by the UAV in completion of the task.” The limitations, as drafted, are considered to be, under the broadest reasonable interpretation, “mental processes”, performed with the aid of a computer used as tool to perform the abstract idea. Therefore, the claim is rejected on the basis as claim 3.
Regarding claim 6, the claim recites “the reinforcement learning process comprises a Q learning process.” The limitation, as drafted, merely indicates the field of use and technological environment, and “generally links” a Q learning process to the abstract idea. (See MPEP 2106.05(h)) Therefore, the claims are rejected on the basis as claim 1.
Regarding claims 7 and 19, the claims recite “the change to the environment comprises a new object being identified at the first location in the environment.” The limitation, as drafted, is considered to be, under the broadest reasonable interpretation, a “mental process”, which is a grouping of abstract idea. Therefore, the claims are rejected on the basis as claims 1 and 15.
Regarding claim 8, the claim recites “receiving, from an image sensor of the agent, an image of the environment; and processing the image to identify the new object at the first location in the environment.”
Regarding claims 9 and 20, the claims recite “each location in the environment corresponds to a cell in the Q table.” The limitation, as drafted, merely indicates the field of use, and technological environment, and “generally links” location in the environment to cells in a Q table. (See MPEP 2106.05(h)) Therefore, the claims are rejected on the basis as claims 1 and 15.
Regarding claim 11, the claim recites to “generate a diffusion model for propagating any changes made in the local Q table across the Q table.” The limitation, as drafted, is considered to be, under the broadest reasonable interpretation, a “mental process”, with the aid of a computer used as tool to perform the abstract idea. Therefore, the claim is rejected on the same basis as claim 10.
Regarding claim 12, the claim recites “the UAV delivering a payload from an initial location to a target location in the environment, and wherein the state of the UAV is based on a current location of the UAV in the environment.” The limitation, as drafted, merely indicates the field of use and technological environment, and “generally links” delivering a payload and locations relative to a UAV and found within the environment to the abstract idea. (See MPEP 2106.05(h)) Therefore, the claim is rejected on the same basis as claim 10.
Regarding claim 13, the claim recites to “identify the state of the UAV based on the current location of the UAV in the environment; and determine an action for the UAV based on mapping the state of the UAV to the Q table.” The limitations, as drafted, are considered to be, under the broadest reasonable interpretation, “mental processes”, with the aid of a computer used as tool to perform the abstract idea. Therefore, the claim is rejected on the same basis as claim 12.
Regarding claim 14, the claim recites “a reward is issued upon completion of the task, and wherein an amount of the reward is determined based on a length of a path traveled by the UAV in completion of the task.” The limitations, as drafted, are considered to be, under the broadest reasonable interpretation, “mental processes”, which are groupings of abstract idea. Therefore, the claim is rejected on the basis as claim 10.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-6, 10, 11, and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Wei Yue et al. (Herein referred to as Yue) (Reinforcement Learning based Approach for Multi UAV Cooperative Searching in Unknown Environments) in view of Hsuan-Fu Wang et al. (Herein referred to as Wang) (RIS-assisted UAV Networks: Deployment Optimization with Reinforcement-Learning-Based Federated Learning)
Regarding claim 1, Yue teaches a method for updating a learned control policy in response to identifying a change to an environment, the method comprising: generating the learned control policy for the environment using a reinforcement learning process, wherein the learned control policy provides a Q table comprising values specifying actions for an agent to take based on a state of the agent in order to complete a task; (“In this section, Q-Learning is used to design the yaw angle decision u(k) of UAVs. When Vi at the state si = [xi(k), yi(k), (k)]T, the corresponding row of the state in the Q-table is found out and set as si(k) row, in which each value represents the effect of taking a decision.”, pg. 3, right column, bottom paragraph) (The decisions of Yue indicate choices of actions for a UAV to take.) detecting a change to the environment at a first location in the environment; (Delta pmn is the variable of probability, that is, when (m n) is not accessed by UAVs, due to other grids are accessed, the probability at the grid (m,n) changes…”, pg. 2, left column, third paragraph) and generating a diffusion model for propagating changes made in the local Q table across the Q table; (See the algorithm on pg. 4, right column. The state-decision model corresponds to a diffusion model which propagates changes across the Q-table. While the local Q-table is not taught with this reference, the model is capable of performing the limitation, as evidence by the muti-UAV application, which implicitly has local search areas and Q-tables associated with the local UAVs.)
However, Yue does not teach defining a local region surrounding the first location, wherein the local region corresponds with a local Q table that is part of the Q table; nor modifying the local Q table using the reinforcement learning process based on the detected change to the environment; nor propagating, using the diffusion model, the changes made in the local Q table globally across the Q table to modify the learned control policy.
Wang teaches defining a local region surrounding the first location, wherein the local region corresponds with a local Q table that is part of the Q table; (“Let denote the global Q-table and denote the local Q-table of the i-th UAV.”, pg. 4, right column, second paragraph; See also Fig. 5 on pg. 5) (The individual UAVs are assigned to a local region and have a local Q-table as part of a global Q-table.) modifying the local Q table using the reinforcement learning process based on the detected change to the environment; (“When a UAV takes an action, “Reward” is defined as the sum of changes in transmission rates of all users connecting to the UAV. Suppose the current state of the i-th UAV is the sum of transmission rates of all connected users is S, the new state after adopting the action is S’, and the new sum of transmission rates of all connected users is Ts’(i), then we have following equations… Once the i-th UAV adopts action A, the Q-table of the i-th UAV would be also updated according to (17)…”, pg. 4, left column; See Equations 14-17 on pg. 4) (The changing of states and transmission rates corresponds to a change in the environment, with the updating of the i-th UAV’s Q-table corresponding to modifying the local Q-table.) and propagating, using the diffusion model, the changes made in the local Q table globally across the Q table to modify the learned control policy. (“UAVs send their local Q-tables to BS, and then BS generates a global Q-table based on what it has received; at last, the global Q-table is returned to UAVs and UAVs keep on reinforcement learning with this global Q-table.”, pg. 3, right column, third paragraph) (Propagation is done from the BS which then returns Q-table to the UAV’s to continue to learn from, the Q-table used for reinforcement learning corresponding to a learned control policy.)
Therefore, it would have been considered obvious to someone of ordinary skill in the art, prior to the filing date of the current application, to combine the Multiple UAV system and Q-learning of Yue with the local and global Q-tables of Wang. Someone of ordinary skill in the art would have been motivated to combine the teachings, prior to the application’s filing date, as this allows for federated learning, speeding up the training process and convergence, as described in Wang. (“If there are UAVs in total, then the main process of federated learning can be written as (18). The reason for adopting federated learning is to help speed up the training process and enhance convergence [12].”, pg. 4, right column, second paragraph)
Regarding claim 10, Yue teaches a system comprising: at least one unmanned aerial vehicle (UAV); (“In this section, Q-Learning is used to design the yaw angle decision u(k) of UAVs”, pg. 3, right column, bottom paragraph) and a computer in electrical communication with the at least one UAV, (While not explicitly taught in the disclosure, one would implicitly need a computer in communication with UAVs to perform the method of Yue.) where the computer is operative to: detect a change to an environment at a first location in the environment, wherein a learned control policy for the environment includes a Q table that comprises values specifying actions for an UAV to take based on a state for the UAV to complete a task; (“When Vi is in a state si(k)=[xi(k),yi(k),
ϕ
i(k)]T, the policy ui(k) is selected according to the largest Q-value in the si(k) row in the Q table. After Vi executed it, Vi arrival state si(k+1), and the immediate reward or penalty value is used to update Q(si(k),ui(k)) in the Q table which is generated according to the efficiency from the state si(k) to the state si(k+1).”, pg. 4, left column, under “B. Q-value Update Process”)
However, Yue does not teach to define a local region surrounding the first location, wherein the local region corresponds with a local Q table that is part of the Q table; nor to modify the local Q table using a Q learning process based on the change to the environment; nor to propagate any changes made in the local Q table globally across the Q table to modify the learned control policy.
Wang teaches to define a local region surrounding the first location, wherein the local region corresponds with a local Q table that is part of the Q table; (“Let denote the global Q-table and denote the local Q-table of the i-th UAV.”, pg. 4, right column, second paragraph; See also Fig. 5 on pg. 5) (The individual UAVs are assigned to a local region and have a local Q-table as part of a global Q-table.) modify the local Q table using a Q learning process based on the change to the environment; (“When a UAV takes an action, “Reward” is defined as the sum of changes in transmission rates of all users connecting to the UAV. Suppose the current state of the i-th UAV is the sum of transmission rates of all connected users is S, the new state after adopting the action is S’, and the new sum of transmission rates of all connected users is Ts’(i), then we have following equations… Once the i-th UAV adopts action A, the Q-table of the i-th UAV would be also updated according to (17)…”, pg. 4, left column; See Equations 14-17 on pg. 4) (The changing of states and transmission rates corresponds to a change in the environment, with the updating of the i-th UAV’s Q-table corresponding to modifying the local Q-table.) and propagate any changes made in the local Q table globally across the Q table to modify the learned control policy. (“UAVs send their local Q-tables to BS, and then BS generates a global Q-table based on what it has received; at last, the global Q-table is returned to UAVs and UAVs keep on reinforcement learning with this global Q-table.”, pg. 3, right column, third paragraph) (Propagation is done from the BS which then returns Q-table to the UAV’s to continue to learn from, the Q-table used for reinforcement learning corresponding to a learned control policy.)
Therefore, it would have been considered obvious to someone of ordinary skill in the art, prior to the filing date of the current application, to combine the Multiple UAV system and Q-learning of Yue with the local and global Q-tables of Wang. Someone of ordinary skill in the art would have been motivated to combine the teachings, prior to the application’s filing date, as this allows for federated learning, speeding up the training process and convergence, as described in Wang. (“If there are UAVs in total, then the main process of federated learning can be written as (18). The reason for adopting federated learning is to help speed up the training process and enhance convergence [12].”, pg. 4, right column, second paragraph)
Regarding claim 2, Yue, as modified by Wang, teaches the method of claim 1, wherein the agent comprises an unmanned aerial vehicle (UAV). (Algorithm on pg. 4, right column of Yue. The multi-UAVs correspond to agents)
Regarding claim 3, Yue, as modified by Wang teaches the method of claim 2, wherein the task comprises the UAV moving from an initial location to a target location in the environment. (“The information in the search area is completely unknown, and the purpose of the search is to identify all targets and possible trends in the sea area… Each warship carries out its mission independently and four UAVs search the warships. The initial positions of UAVs are located at the four corners of the sea area, and the speed is 30 (m/s)”, pg. 4, right column, bottom paragraph; pg. 5, left column, under “A. Independent Random Distribution of Targets” (Yue)) (The searching for targets corresponds to a UAV moving from an initial location, (the start) to a target location in the environment (the targets).)
Regarding claim 4, Yue, as modified by Wang, teaches the method of claim 3, further comprising: identifying the state of the UAV based on a present location of the UAV in the environment; (“There are n yaw angles at each grid, so the number of rows in the Q table is Lx
×
Ly
×
n, which is the number of states of UAV. There are m optional control inputs for each UAV, so the number of columns in the Q table is m which is the number of decisions contained in the decision set A. Define variable Q(si(k),ui(k)) as the value that Vi selects the decision ui(k) in the state si(k)”, pg. 4, left column, top paragraph (Yue)) and determining an action for the UAV based on mapping the state of the UAV to the Q table. (“There are m optional control inputs for each UAV, so the number of columns in the Q table is m which is the number of decisions contained in the decision set A. Define variable Q(si(k),ui(k)) as the value that Vi selects the decision ui(k) in the state si(k)”, pg. 4, left column, top paragraph (Yue))
Regarding claim 5, Yue, as modified by Wang, teaches the method of claim 3, wherein a reward is issued upon completion of the task, and wherein an amount of the reward is determined based on a length of a path traveled by the UAV in completion of the task. (“When Vi is in a state si(k)=[xi(k),yi(k),
ϕ
i(k)]T, the policy ui(k) is selected according to the largest Q-value in the si(k) row in the Q table. After Vi executed it, Vi arrival state si(k+1), and the immediate reward or penalty value is used to update Q(si(k),ui(k)) in the Q table which is generated according to the efficiency from the state si(k) to the state si(k+1).”, pg. 4, left column, under “B. Q-value Update Process” (Yue)) (The “efficiency” metric of Yue is interpreted to be the most optimal path from one state to another, which is typically the shortest path.)
Regarding claim 6, Yue, as modified by Wang, teaches the method of claim 1, wherein the reinforcement learning process comprises a Q learning process. (“In this paper, the Boltzmann distribution mechanism is used to select the decision in the Q-learning process, that is, the probability that the policy set u(k) is selected in the state s(k) is as follows…”, pg. 3, left column, under “A. Establishment of Q-value Table” (Yue))
Regarding claim 11, Yue, as modified by Wang, teaches the system of claim 10, wherein the computer is further operative to: generate a diffusion model for propagating any changes made in the local Q table across the Q table. (“UAVs send their local Q-tables to BS, and then BS generates a global Q-table based on what it has received; at last, the global Q-table is returned to UAVs and UAVs keep on reinforcement learning with this global Q-table.”, pg. 3, right column, third paragraph (Wang)) (Propagation is done from the BS which then returns Q-table to the UAV’s to continue to learn from.)
Regarding claim 14, Yue, as modified by Wang, teaches the system of claim 10, wherein a reward is issued upon completion of the task, and wherein an amount of the reward is determined based on a length of a path traveled by the UAV in completion of the task. (“When Vi is in a state si(k)=[xi(k),yi(k),
ϕ
i(k)]T, the policy ui(k) is selected according to the largest Q-value in the si(k) row in the Q table. After Vi executed it, Vi arrival state si(k+1), and the immediate reward or penalty value is used to update Q(si(k),ui(k)) in the Q table which is generated according to the efficiency from the state si(k) to the state si(k+1).”, pg. 4, left column, under “B. Q-value Update Process” (Yue)) (The “efficiency” metric of Yue is interpreted to be the most optimal path from one state to another, which is typically the shortest path.)
Claim(s) 7, 8, and 15-19 are rejected under 35 U.S.C. 103 as being unpatentable over Yue in view of Wang and in further view of Frederik S. Keira et al. (Herein referred to as Keira) (Object detection, recognition, and tracking from UAVs using a thermal camera)
Regarding claim 7, Yue as modified by Wang, teaches the method of claim 1, but does not explicitly teach the change to the environment comprises a new object being identified at the first location in the environment.
Leira teaches the change to the environment comprises a new object being identified at the first location in the environment. (“These two modules could also be implemented to include a search component in the path planning, meaning that the path planner could choose to search for new undetected objects instead of only focusing on the objects currently being tracked. The user could also change the UAV's tracking priorities online and in real time to accommodate needs that arise due to, for example, a change in the environment.”, pg. 6, before “3 | OBJECT DETECTION AND RECOGNITION”)
Therefore, it would have been considered obvious to someone of ordinary skill in the art, prior to the filing date of the current application, to combine the system of Yue, as modified by Wang, with the modules of Leira. Someone of ordinary skill in the art would have been motivated to combine the teachings, prior to the application’s filing date, as this allows for the user to change object tracking priorities. (“The user could also change the UAV's tracking priorities online and in real time to accommodate needs that arise due to, for example, a change in the environment.”, pg. 6, before “3 | OBJECT DETECTION AND RECOGNITION”)
Regarding claim 8, Yue, as modified by Wang and Leira, teaches the method of claim 7, further comprising: receiving, from an image sensor of the agent, an image of the environment; (“These two modules could also be implemented to include a search component in the path planning, meaning that the path planner could choose to search for new undetected objects instead of only focusing on the objects currently being tracked. The user could also change the UAV's tracking priorities online and in real time to accommodate needs that arise due to, for example, a change in the environment.”, pg. 6, before “3 | OBJECT DETECTION AND RECOGNITION” (Leira)) and processing the image to identify the new object at the first location in the environment. (“when a new object is detected a Kalman filter is initialized with the object's measured position as the filter's position states
x
^
k
obj,
y
^
k
obj.”, pg. 9, right column, fifth paragraph (Leira))
Regarding claim 15, Yue teaches a computer-readable storage medium containing program instructions for a method being executed by an application, the application comprising code for one or more components that are called by the application during runtime, (See the Algorithm on pg. 4 of Yue for application code. While the medium is not explicitly taught, one would need said component to distribute the method of Yue.) wherein execution of the program instructions by one or more processors of a computer system causes the one or more processors to perform steps comprising: generating a learned control policy for an environment using a reinforcement learning process, wherein the learned control policy provides a Q table comprising values specifying actions for an agent to take based on a state of the agent in order to complete a task; (“In this section, Q-Learning is used to design the yaw angle decision u(k) of UAVs. When Vi at the state si = [xi(k), yi(k), (k)]T, the corresponding row of the state in the Q-table is found out and set as si(k) row, in which each value represents the effect of taking a decision.”, pg. 3, right column, bottom paragraph) (The decisions of Yue indicate choices of actions for a UAV to take.) generating a diffusion model for propagating changes made in the local Q table across the Q table; (See the algorithm on pg. 4, right column. The state-decision model corresponds to a diffusion model which propagates changes across the Q-table. While the local Q-table is not taught with this reference, the model is capable of performing the limitation, as evidenced by the muti-UAV application, which implicitly has local search areas and Q-tables associated with the local UAVs.)
However Yue does not teach detecting a change to the environment at a first location in the environment by: receiving, from an image sensor of the agent, an image of the environment; and processing the image to identify a new object at the first location in the environment; nor defining a local region surrounding the first location, wherein the local region corresponds with a local Q table that is part of the Q table; nor modifying the local Q table using the reinforcement learning process based on the detected change to the environment; nor propagating, using the diffusion model, the changes made in the local Q table globally across the Q table to modify the learned control policy.
Wang teaches defining a local region surrounding the first location, wherein the local region corresponds with a local Q table that is part of the Q table; (“Let denote the global Q-table and denote the local Q-table of the i-th UAV.”, pg. 4, right column, second paragraph; See also Fig. 5 on pg. 5) (The individual UAVs are assigned to a local region and have a local Q-table as part of a global Q-table.) modifying the local Q table using the reinforcement learning process based on the detected change to the environment; (“When a UAV takes an action, “Reward” is defined as the sum of changes in transmission rates of all users connecting to the UAV. Suppose the current state of the i-th UAV is the sum of transmission rates of all connected users is S, the new state after adopting the action is S’, and the new sum of transmission rates of all connected users is Ts’(i), then we have following equations… Once the i-th UAV adopts action A, the Q-table of the i-th UAV would be also updated according to (17)…”, pg. 4, left column; See Equations 14-17 on pg. 4) (The changing of states and transmission rates corresponds to a change in the environment, with the updating of the i-th UAV’s Q-table corresponding to modifying the local Q-table.) and propagating, using the diffusion model, the changes made in the local Q table globally across the Q table to modify the learned control policy. (“UAVs send their local Q-tables to BS, and then BS generates a global Q-table based on what it has received; at last, the global Q-table is returned to UAVs and UAVs keep on reinforcement learning with this global Q-table.”, pg. 3, right column, third paragraph) (Propagation is done from the BS which then returns Q-table to the UAV’s to continue to learn from, the Q-table used for reinforcement learning corresponding to a learned control policy.)
Therefore, it would have been considered obvious to someone of ordinary skill in the art, prior to the filing date of the current application, to combine the Multiple UAV system and Q-learning of Yue with the local and global Q-tables of Wang. Someone of ordinary skill in the art would have been motivated to combine the teachings, prior to the application’s filing date, as this allows for federated learning, speeding up the training process and convergence, as described in Wang. (“If there are UAVs in total, then the main process of federated learning can be written as (18). The reason for adopting federated learning is to help speed up the training process and enhance convergence [12].”, pg. 4, right column, second paragraph)
However, the combination does not teach detecting a change to the environment at a first location in the environment by: receiving, from an image sensor of the agent, an image of the environment; and processing the image to identify a new object at the first location in the environment.
Leira teaches detecting a change to the environment at a first location in the environment by: receiving, from an image sensor of the agent, an image of the environment; (“These two modules could also be implemented to include a search component in the path planning, meaning that the path planner could choose to search for new undetected objects instead of only focusing on the objects currently being tracked. The user could also change the UAV's tracking priorities online and in real time to accommodate needs that arise due to, for example, a change in the environment.”, pg. 6, before “3 | OBJECT DETECTION AND RECOGNITION”) and processing the image to identify a new object at the first location in the environment; (“when a new object is detected a Kalman filter is initialized with the object's measured position as the filter's position states
x
^
k
obj,
y
^
k
obj.”, pg. 9, right column, fifth paragraph)
Therefore, it would have been considered obvious to someone of ordinary skill in the art, prior to the filing date of the current application, to combine the system of Yue, as modified by Wang, with the image sensors for UAVs of Leira. Someone of ordinary skill in the art would have been motivated to combine the teachings, prior to the application’s filing date, as this allows for detection, recognition, and tracking of potential (new) objects of interest with unknown positions, as detailed in Leira. (“The tracking process does not only consider already known object positions, but also focuses on detecting, recognizing, and tracking potential objects of interest with unknown positions. This can be achieved by using an onboard sensor, typically a camera, mounted in a pan‐tilt gimbal, and a path controller.”, pg. 5, right column, second paragraph)
Regarding claim 16, Yue, as modified by Wang and Leira teaches the computer-readable storage medium of claim 15, wherein the agent comprises an unmanned aerial vehicle (UAV). (Algorithm on pg. 4, right column of Yue. The multi-UAVs correspond to agents)
Regarding claim 17, Yue, as modified by Wang and Leira teaches the computer-readable storage medium of claim 16, wherein the task comprises the UAV moving from an initial location to a target location in the environment. (“The information in the search area is completely unknown, and the purpose of the search is to identify all targets and possible trends in the sea area… Each warship carries out its mission independently and four UAVs search the warships. The initial positions of UAVs are located at the four corners of the sea area, and the speed is 30 (m/s)”, pg. 4, right column, bottom paragraph; pg. 5, left column, under “A. Independent Random Distribution of Targets” (Yue)) (The searching for targets corresponds to a UAV moving from an initial location, (the start) to a target location in the environment (the targets).)
Regarding claim 18, Yue, as modified by Wang and Leira, teaches the computer-readable storage medium of claim 17, further comprising: identifying the state of the UAV based on a present location of the UAV in the environment; (“There are n yaw angles at each grid, so the number of rows in the Q table is Lx
×
Ly
×
n, which is the number of states of UAV. There are m optional control inputs for each UAV, so the number of columns in the Q table is m which is the number of decisions contained in the decision set A. Define variable Q(si(k),ui(k)) as the value that Vi selects the decision ui(k) in the state si(k)”, pg. 4, left column, top paragraph (Yue)) and determining an action for the UAV based on mapping the state of the UAV to the Q table. (“There are m optional control inputs for each UAV, so the number of columns in the Q table is m which is the number of decisions contained in the decision set A. Define variable Q(si(k),ui(k)) as the value that Vi selects the decision ui(k) in the state si(k)”, pg. 4, left column, top paragraph (Yue))
Regarding claim 19, Yue, as modified by Wang and Leira, teaches the computer-readable storage medium of claim 15, wherein the change to the environment comprises a new object being identified at the first location in the environment. (“These two modules could also be implemented to include a search component in the path planning, meaning that the path planner could choose to search for new undetected objects instead of only focusing on the objects currently being tracked. The user could also change the UAV's tracking priorities online and in real time to accommodate needs that arise due to, for example, a change in the environment.”, pg. 6, before “3 | OBJECT DETECTION AND RECOGNITION” (Leira))
Claims 9, 12, and 13 is rejected under 35 U.S.C. 103 as being unpatentable over Yue in view of Wang and in further view of Anna Guerra et al. (Herein referred to as Guerra) (Multi-Agent Q-Learning in UAV Networks for Target Detection and Indoor Mapping)
Regarding claim 9, Yue, as modified by Wang, teaches the method of claim 1, but does not explicitly teach each location in the environment corresponds to a cell in the Q table.
Guerra teaches each location in the environment corresponds to a cell in the Q table. (“we define si,k ∈ S as the vector containing the states of the ith UAV at time instant k, that is, the ith UAV position, the map of the environment and a detection vector, i.e... where pi,k = [xi,k, yi,k]T ∈ R2 is the true UAV position, mk ∈ BNcell is the true map at time k described as a vector of Ncell cells that represent the map, and tk ∈ BN is the target vector (equal to one if the target is present and zero otherwise) with N, being the number of targets.”, pg. 2, left and right columns) (As this segment is for a Q-learning Algorithm for UAV navigation, it’s implied that the cells correspond to Q-table cells.)
Therefore, it would have been considered obvious to someone of ordinary skill in the art, prior to the filing date of the current application, to combine the system of Yue, as modified by Wang, with the UAV navigation and use of cells, as disclosed in Guerra. Someone of ordinary skill in the art would have been motivated to combine the teachings, prior to the application’s filing date, as this allows for optimized target detection and mapping accuracy. (“we consider a trajectory that is chosen by the UAVs to maximize the target detection and mapping accuracy subject to the mission time TM and collision avoidance. This optimization problem can be properly formulated by a Markov decision process (MDP), which is defined by a tuple containing the state space S, the action space A, the reward space R and the probability of transitioning from one state sk, at time instant k, to the state sk+1 at time instant k +1.”, pg. 2, left column, under “B. A Q-learning Algorithm for UAV Navigation”)
Regarding claim 12, Yue, as modified by Wang, teaches the system of claim 10, as well as the state of the UAV is based on a current location of the UAV in the environment. (“There are n yaw angles at each grid, so the number of rows in the Q table is Lx
×
Ly
×
n, which is the number of states of UAV. There are m optional control inputs for each UAV, so the number of columns in the Q table is m which is the number of decisions contained in the decision set A. Define variable Q(si(k),ui(k)) as the value that Vi selects the decision ui(k) in the state si(k)”, pg. 4, left column, top paragraph (Yue))
However, the combination does not teach the UAV delivering a payload from an initial location to a target location in the environment
Guerra teaches the UAV delivering a payload from an initial location to a target location in the environment. (“UAVs have played a central role in emergency situations in hazardous environments, for post natural disasters, or for search-and-rescue operations. In such events, UAVs have been used as a temporary network infrastructure for localization, communications, and for delivering items… More specifically, such agents cooperate in order to achieve two goals: (i) detection of targets that can be, for example, cooperative users that need to be rescued or hidden malicious targets whose unwanted communication is sniffed within a certain radiofrequency band;…”, pg. 1, left column, second paragraph; pg. 1, right column, bottom paragraph) (The UAV reaching a target destination to rescue targets or disrupt communication implicitly teaches the delivery of a payload at a target location in the environment.)
Therefore, it would have been considered obvious to someone of ordinary skill in the art, prior to the filing date of the current application, to combine the system of Yue, as modified by Wang, with the intended use of UAVs disclosed in Guerra. Someone of ordinary skill in the art would have been motivated to combine the teachings, prior to the application’s filing date, as this allows for the UAVs to be used in search and rescues scenarios that are too risky for human operators. (“UAVs are autonomous flying agents capable of performing multiple tasks, and they are usually deployed to carry out missions that are too risky for human operators. For example, UAVs have played a central role in emergency situations in hazardous environments, for post natural disasters, or for search-and-rescue operations.”, pg. 1, left column, under “I. INTRODUCTION”)
Regarding claim 13, Yue, as modified by Wang and Guerra, teaches the system of claim 12, wherein the computer is further operative to: identify the state of the UAV based on the current location of the UAV in the environment; (“There are n yaw angles at each grid, so the number of rows in the Q table is Lx
×
Ly
×
n, which is the number of states of UAV. There are m optional control inputs for each UAV, so the number of columns in the Q table is m which is the number of decisions contained in the decision set A. Define variable Q(si(k),ui(k)) as the value that Vi selects the decision ui(k) in the state si(k)”, pg. 4, left column, top paragraph (Yue)) and determine an action for the UAV based on mapping the state of the UAV to the Q table. (“There are m optional control inputs for each UAV, so the number of columns in the Q table is m which is the number of decisions contained in the decision set A. Define variable Q(si(k),ui(k)) as the value that Vi selects the decision ui(k) in the state si(k)”, pg. 4, left column, top paragraph (Yue))
Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Yue in view of Wang, in further view of Leira, and in further view of Guerra.
Regarding claim 20, Yue, as modified by Wang and Leira, teaches the computer-readable storage medium of claim 15, but does not explicitly teach each location in the environment corresponds to a cell in the Q table.
Guerra teaches each location in the environment corresponds to a cell in the Q table. (“we define si,k ∈ S as the vector containing the states of the ith UAV at time instant k, that is, the ith UAV position, the map of the environment and a detection vector, i.e... where pi,k = [xi,k, yi,k]T ∈ R2 is the true UAV position, mk ∈ BNcell is the true map at time k described as a vector of Ncell cells that represent the map, and tk ∈ BN is the target vector (equal to one if the target is present and zero otherwise) with N, being the number of targets.”, pg. 2, left and right columns) (As this segment is for a Q-learning Algorithm for UAV navigation, it’s implied that the cells correspond to Q-table cells.)
Therefore, it would have been considered obvious to someone of ordinary skill in the art, prior to the filing date of the current application, to combine the system of Yue, as modified by Wang and Leira, with the UAV navigation and use of cells, as disclosed in Guerra. Someone of ordinary skill in the art would have been motivated to combine the teachings, prior to the application’s filing date, as this allows for optimized target detection and mapping accuracy. (“we consider a trajectory that is chosen by the UAVs to maximize the target detection and mapping accuracy subject to the mission time TM and collision avoidance. This optimization problem can be properly formulated by a Markov decision process (MDP), which is defined by a tuple containing the state space S, the action space A, the reward space R and the probability of transitioning from one state sk, at time instant k, to the state sk+1 at time instant k +1.”, pg. 2, left column, under “B. A Q-learning Algorithm for UAV Navigation”)
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Tyler E Iles whose telephone number is (571)272-5442. The examiner can normally be reached 9:00am - 5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kakali Chaki can be reached at (571) 272-3719. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/T.E.I./ Patent Examiner, Art Unit 2122
/MICHAEL H HOANG/ PRIMARY EXAMINER, Art Unit 2122