DETAILED ACTION
Status of Claims
This Office action is in response to the amendment filed 06/24/2026. Claims 1-10 were previously canceled. With the filed amendment, claims 14-15 and 25 are also canceled. Claims 11-13, 16-24, and 26-30 are currently pending and are presented for examination.
Information Disclosure Statement
The information disclosure statement submitted on 05/26/2026 is in compliance with 37 C.F.R. 1.97 and is being considered by the examiner.
Response to Amendment/Arguments
The amendment filed 06/24/2026 has been entered and applicants’ arguments filed 06/24/2026 have been fully considered.
Regarding objections:
The applicants have argued that the objections to the drawings, specification, and claims are overcome by the filed amendment. The examiner agrees and has withdrawn these objections accordingly.
Regarding claim interpretation under 35 U.S.C. § 112(f):
The applicants have argued that the claim interpretation under 35 U.S.C. § 112(f) should no longer apply in light of the filed amendment. All of the nonce terms except for the “decision network” have been removed from the claims, and so the interpretation has been withdrawn for all the terms except the “decision network.”
Regarding claim rejections under 35 U.S.C. § 112(b):
The applicants have argued that the rejections under 35 U.S.C. § 112(b) are overcome by the filed amendment. The examiner agrees and has withdrawn the rejections accordingly.
Regarding claim rejections under 35 U.S.C. § 101:
The applicants have argued that the claims as amended should not be rejected under 35 U.S.C. § 101 since “they require specific operations that require manipulation of data structures stored in processor memory and cannot be practically performed in the human mind.” Citing to different portions of the instant specification, the applicants have argued that the amended step of “calculating, using the decision network and based on state information of the UAV server and the task request, decision information of the UAV server, the decision information comprising an action decision of the UAV server, available computing resources of the UAV server, available bandwidth of the UAV server, and an estimated execution time for a task” cannot be performed mentally since it requires “modifying numerical values across data structures stored in hardware memory as data propagates through the network” as described in the instant specification. The applicants have also argued that the step of “adjusting network parameters of an initial decision network based on a plurality of reward and punishment constraints” cannot be performed in the human mind since “This is a hardware-level operation that modified data structures in processor memory and is thus not something performable with pen and paper.”
The examiner respectfully disagrees; while the argued steps could be performed with the use of computer hardware as described in the specification, they could also be performed in the human mind. The decision network merely automates the steps as a means of performing the abstract idea which could otherwise be manually executed by a human with the help of pen and paper. Specifically, a human being could use pen and paper to perform mathematical calculations as described in the instant specification to calculate the decision information based on the state information and the task request. Further, a human could use pen and paper to adjust network parameters based on reward and punishment constraints as described in the instant specification. Employing generic and conventional computing components to perform an abstract idea which could otherwise be performed by a human does not preclude a claim from being directed to a judicial exception (see MPEP 2106.05(f)).
The applicants have also argued that the alleged abstract idea is integrated into a practical application because the claim limitations “define a particular computational pipeline with specific inputs, specific outputs tied to physical UAV-MEC operating parameters, and a specific training regime constrained by those parameters” rather than reciting the “decision network” at a high level of generality. The applicants have argued that the claimed invention resolves the problems of overlapping coverage areas and waste of resources as described in the instant specification, and that the steps of calculating decision information and adjusting network parameters output dimensions that “correspond to a physical resource of the UAV server that is allocated to service the task request.”
The examiner respectfully disagrees, because the described benefits are not required by the claims as they are currently drafted. While the claimed process could potentially be used to realize some of these benefits, the claims recite no limitations which actually achieve the benefits and integrate the abstract idea into a practical application. Even if the claimed generation and transmission of a service decision instruction is performed, nothing in the claim language requires that the service decision instruction is used in any way. Under the current broadest reasonable interpretation of the claims, the service decision instruction could be transmitted to the terminal and then ignored. For these reasons, the additional elements of the claims do not integrate the abstract idea into a practical application.
The applicants have further argued that the instant claims are analogous to the claims of Ex parte Desjardins and asserted that in the instant claims, “the claimed decision network and its training against UAV-MEC specific reward and punishment constraints improve the functioning of UAV-based mobile edge computing systems by reducing resource waste in overlapping coverage areas and are similarly patent eligible.”
The examiner respectfully disagrees, because the instant claims are not analogous to the claims of Ex parte Desjardins. There is no training of the network that occurs in the independent claims; rather, this training step only occurs in some of the dependent claims. The independent claims merely describe the decision network performing a calculation rather than a process of training the decision network in a way that would improve the functioning of computing systems as argued. For these reasons and the reasons explained above, the rejections under 35 U.S.C. § 101 have been maintained.
Regarding claim rejections under 35 U.S.C. § 103:
The applicants have argued that the previously cited prior art fails to teach the “decision information comprising an action decision of the UAV server, available computing resources of the UAV server, available bandwidth of the UAV server, and an estimated execution time for a task” because “Qin does not involve machine learning at all” and because “Chai’s binary action space does not teach or suggest this decision output in which the network determines both whether to serve a task and the specific resource quantities to allocate for that task.”
While the examiner agrees that the combination of Qin and Chai does not teach the decision information comprising “available bandwidth of the UAV server,” the examiner notes that Chai teaches the “decision information comprising an action decision of the UAV server, available computing resources of the UAV server… and an estimated execution time for a task.” Specifically, Chai ¶ 15 discloses to “Solve the task offloading model for minimizing latency and energy consumption in the UAV-mobile edge computing system using deep reinforcement learning,” which teaches the decision information comprising “available computing resources of the UAV server” as claimed. Further, Chai ¶ 19 discloses that “In deep reinforcement learning, the agent obtains the current Q value through Q-network training, selects action a in the current state s from action space A, and executes the action to obtain vector value reward r and the next state s´. The action space A includes the following two actions: executing tasks on the terminal device and unloading tasks to the UAV-mobile edge computing system.” This disclosure teaches the decision information comprising “an action decision of the UAV server” as claimed. Also, Chai ¶¶ 25-27 disclose a constraint in which “The UAV can only fly within a specified rectangular area, and the horizontal range of time slot t and the maximum distance to fly within time slot t are also specified,” which teaches the decision information comprising “an estimated execution time for a task” as claimed.
The applicants have further argued that the previously cited prior art fails to teach that “the decision network is obtained by adjusting network parameters of an initial decision network based on a plurality of reward and punishment constraints” including “an exclusivity constraint that penalizes a decision to service the terminal by zero UAV servers or by more than one UAV server in the overlapping service area.”
The examiner notes that the claims as currently drafted do not require consideration of the exclusivity constraint. Instead, the claims merely require adjusting network parameters based on “one or more of: available computing resources of the UAV server, available bandwidth of the UAV server, an exclusivity constraint that penalizes a decision to service the terminal by zero UAV servers or by more than one UAV server in the overlapping service area, an estimated execution time for the task request by the UAV server, and an execution time of a current training epoch.” Chai ¶¶ 25-27 teach the constraints being based on “an estimated execution time for the task request by the UAV server” and accordingly teaches the entire claim limitation since only “one or more” of the listed options is required by the claim.
The applicants have also argued that Chai does not teach the set of constraint conditions because the reward of Chai is a single composite signal that the agent optimizes, while the instant claims require that the decision network is obtained by adjusting parameters based on a plurality of reward and punishment constraints tied to specific UAV-MEC operating parameters, including available computing resources, available bandwidth, estimated execution time, training epoch time, and the exclusivity constraint.
The examiner respectfully disagrees, because the claims as currently drafted only require adjusting parameters based on “one or more” of the listed constraints. Accordingly, Chai teaches the limitation because Chai ¶¶ 25-27 disclose to “Determine whether the current iteration has reached the maximum number of iterations. If so, output the optimal unloading decision, where the optimal unloading decision means that the vector value reward obtained by the agent after performing action a is maximized. Otherwise, go to step 5. Furthermore, the task dependency constraints include: Constraint 1: The UAV can only fly within a specified rectangular area, and the horizontal range of time slot t and the maximum distance to fly within time slot t are also specified.” This at least teaches to use a constraint based on “an estimated execution time for the task request by the UAV server” as claimed.
The applicants’ remaining arguments are moot in view of the new grounds of rejection which are necessitated by the filed amendment.
Claim Objections
Claim 18 is objected to because of the following informality: In claim 18, it appears that the first instance of “the evaluation network” should be changed to “[[the]]an evaluation network.” Appropriate correction is required.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitations use a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitations are:
“generating, using a decision network, a service decision instruction based on the task request and the one or more other UAV servers” (claim 11)
“wherein each of the plurality of service decision instructions is generated using a decision network based on the task request” (claim 27)
“receive a plurality of service decision instructions from the plurality of UAV servers generated based on the task request using a decision network” (claim 30)
Because these claim limitations are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, they are being interpreted to cover the corresponding structure described in the specification (¶ 50: “The target decision network is a pre-trained neural network.”) as performing the claimed function, and equivalents thereof.
If applicant does not intend to have these limitations interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitations to avoid them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitations recite sufficient structure to perform the claimed function so as to avoid them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 11 and 26-30 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding claims 11, 27, and 29-30:
Step 1: Claims 11 and 27 are each directed to a computer-implemented method. Claims 29-30 are each directed to a service decision device. Claims 11, 27, and 29-30 are each directed to at least one of the four statutory categories.
Step 2A, prong 1: Claim 11 recites the abstract idea of generating a service decision instruction. This abstract idea is described at least in claim 11 by the mental process steps of determining that the terminal lies within an overlapping service area between a UAV server and one or more other UAV servers based on the position information; and generating a service decision instruction based on the task request and the one or more other UAV servers, wherein the generating comprises: calculating, based on state information of the UAV server and the task request, decision information of the UAV server, the decision information comprising an action decision of the UAV server, available computing resources of the UAV server, available bandwidth of the UAV server, and an estimated execution time for a task; and generating the service decision instruction based on the decision information, where the service decision instruction comprises an indication of whether the UAV server should service the task request. These steps fall into the mental processes grouping of abstract ideas as they include a human mentally identifying whether the terminal lies within the overlapping service area and using pen and paper to write out an instruction indicating whether the UAV server should service the task request based on calculating the decision information of the UAV server based on state information of the UAV server and the task request. The limitations as drafted are processes that, under their broadest reasonable interpretation, cover their performance in the mind if not for the recitation of generic computing components.
Claim 29 recites the abstract concept of generating a service decision instruction. This abstract idea is described at least in claim 29 by the mental process step of generating a service decision instruction based on the task request, wherein the generating comprises: calculating, based on state information of the UAV server and the task request, decision information of the UAV server, the decision information comprising an action decision of the UAV server, available computing resources of the UAV server, available bandwidth of the UAV server, and an estimated execution time for a task; and generating the service decision instruction based on the decision information, wherein the service decision instruction comprises an indication of whether the UAV server should service the task request. This step falls into the mental processes grouping of abstract ideas because it includes a human using pen and paper to write out an instruction indicating whether the UAV server should service the task request based on calculating the decision information of the UAV server based on state information of the UAV server and the task request. The limitations as drafted are processes that, under their broadest reasonable interpretation, cover their performance in the mind if not for the recitation of generic computing components.
Claims 27 and 30 recite the abstract concept of selecting a UAV server. This abstract idea is described at least in claims 27 and 30 by the mental process step of selecting a UAV server from the plurality of UAV servers to service the task request based on the plurality of service instructions. This step falls into the mental processes grouping of abstract ideas as it includes a human mentally choosing which UAV server should service the task request based on the service decision instructions. The limitations as drafted are processes that, under their broadest reasonable interpretation, cover their performance in the mind if not for the recitation of generic computing components.
With respect to claims 11, 27, and 29-30, other than reciting that the processes are “computer-implemented” and reciting “one or more processors,” “a decision network,” and a “service decision device,” nothing in the claims precludes the idea from practically being performed in the human mind. If not for the “computer,” “processor,” “decision network,” and “service decision device” language, the claims encompass a human operator performing each of the mental process steps in the mind with the help of pen and paper.
Step 2A, prong 2: The claims recite elements additional to the abstract concepts. However, these additional elements fail to integrate the abstract idea into a practical application.
Claim 11 recites that the method is “computer-implemented,” which amounts to specifying the use of generic computer components (as supported by ¶ 173 of the instant specification) that are simply employed as tools for performing the abstract idea. The use of such generic computing components for executing the abstract idea does not integrate the abstract idea into a practical application (see MPEP 2106.05(f)). Claim 11 also recites a step of receiving, by one or more processors, a task request from a terminal, wherein the task request comprises an identifier of the terminal, position information of the terminal, and/or task information of the terminal. This step is insignificant extra-solution activity, as it simply gathers data necessary to perform the abstract idea (i.e., all uses of the abstract idea require such data gathering). Additionally, the step of transmitting the service decision instruction to the terminal using the identifier of the terminal is insignificant extra-solution activity, because it is a data output step that does not impose meaningful limits on the claim such that it is not nominally or tangentially related to the invention. The recitation of such insignificant extra-solution activity does not integrate the abstract idea into a practical application (see MPEP 2106.05(g)). Claim 11 additionally specifies that the service decision instruction is generated using a decision network, wherein the decision network is obtained by adjusting network parameters of an initial decision network based on a plurality of reward and punishment constraints based on one or more of: available computing resources of the UAV server, available bandwidth of the UAV server, an exclusivity constraint that penalizes a decision to service the terminal by zero UAV servers of by more than one UAV server in the overlapping service area, an estimated time for the task request by the UAV server, and an execution time of a current training epoch, wherein the adjusting is performed cooperatively across the UAV server and the one or more other UAV servers. These limitations amount to general linking of the abstract idea to the particular technological field of machine learning that is performed across multiple UAV servers. Using such a decision network as described to generate a service decision instruction merely serves to employ machine learning technology to execute the abstract idea. The recitation of such general linking limitations does not integrate the abstract idea into a practical application (see MPEP 2106.05(h)).
Claim 27 recites that the method is “computer-implemented,” which amounts to specifying the use of generic computer components (as supported by ¶ 173 of the instant specification) that are simply employed as tools for performing the abstract idea. The use of such generic computing components for executing the abstract idea does not integrate the abstract idea into a practical application (see MPEP 2106.05(f)). Claim 27 also recites steps of transmitting a task request to each of a plurality of UAV servers within an overlapping coverage area accessible by a terminal, wherein the task request comprises an identifier of the terminal, position information of the terminal, and/or task information of the terminal; and receiving a plurality of service decision instructions from the plurality of UAV servers, wherein each of the plurality of service decision instructions comprises an indication of whether the corresponding UAV server should service the task request. These steps amount to insignificant extra-solution activity, because they are data gathering and data output steps that are necessary for performing the abstract idea (i.e., all uses of the abstract idea require such data gathering). Also, the step of transmitting the selection and the task request to the UAV server to service the task request is insignificant extra-solution activity, because it is a data output step that does not impose meaningful limits on the claim such that it is not nominally or tangentially related to the invention. Note that while the claim recites that the transmission is used for “the UAV server to service the task request,” this is merely the intended use of the claimed method and does not actually have to occur as part of the method. The recitation of such insignificant extra-solution activity does not integrate the abstract idea into a practical application (see MPEP 2106.05(g)). Claim 27 additionally specifies that each of the plurality of service decision instructions is generated using a decision network based on the task request, wherein the decision network is obtained by adjusting network parameters of an initial decision network based on a plurality of reward and punishment constraints based on one or more of: available computing resources of the UAV server, available bandwidth of the UAV server, an exclusivity constraint that penalizes a decision to service the terminal by zero UAV servers of by more than one UAV server in the overlapping coverage area, an estimated execution time for the task request by the corresponding UAV server, and an execution time of a current training epoch, wherein the adjusting is performed cooperatively across the plurality of UAV servers. This amounts to general linking of the abstract idea to the particular technological field of machine learning that is performed across multiple UAV servers. Using a decision network as described to generate each of the service decision instructions merely serves to employ machine learning technology to implement the abstract idea. Such general linking limitations do not integrate the abstract idea into a practical application (see MPEP 2106.05(h)).
Claim 29 recites a service decision device comprising one or more processors, which are generic computer components (as supported by ¶ 173 of the instant specification) that are simply employed as tools for performing the abstract idea. The use of such generic computing components for executing the abstract idea does not integrate the abstract idea into a practical application (see MPEP 2106.05(f)). Claim 29 also recites the step of receiving, via a communication interface, a task request sent by a terminal in an overlapping service area between a UAV server and one or more other UAV servers, wherein the task request comprises an identifier of the terminal, position information of the terminal and/or task information of the terminal. This step is insignificant extra-solution activity, as it simply gathers data that is necessary for performing the abstract idea (i.e., all uses of the abstract idea require such data gathering). Further, the step of transmitting, via the communication interface, the service decision instruction to the terminal using the identifier of the terminal is insignificant extra-solution activity, as it is a data output step that does not impose meaningful limits on the claim such that it is not nominally or tangentially related to the invention. Note that while the claim recites that “the service decision instruction is used for the terminal to select among the UAV server and the one or more UAV servers to service the task request based on the service decision instruction and service decision instructions transmitted by the one or more other UAV servers,” this is merely the intended use of the claimed service decision device and does not actually have to occur as part of the method. The recitation of such insignificant extra-solution activity does not integrate the abstract idea into a practical application (see MPEP 2106.05(g)). Claim 29 further specifies that the service decision instruction is generated using a decision network based on the task request, wherein the decision network is obtained by adjusting network parameters of an initial decision network based on a plurality of reward and punishment constraints based on one or more of: available computing resources of the UAV server, available bandwidth of the UAV server, an exclusivity constraint that penalizes a decision to service the terminal by zero UAV servers of by more than one UAV server in the overlapping service area, an estimated execution time for the task request by the corresponding UAV server, and an execution time of a current training epoch, where the adjusting is performed cooperatively across the UAV server and the one or more other UAV servers. This limitation amounts to general linking of the abstract idea to the particular technological field of machine learning that is performed across multiple UAV servers. Using the decision network as described to generate the service decision instruction merely serves to employ machine learning technology to implement the abstract idea. The recitation of such general linking limitations does not integrate the abstract idea into a practical application (see MPEP 2106.05(h)).
Claim 30 recites a service decision device comprising one or more processors, which are generic computer components (as supported by ¶ 173 of the instant specification) that are simply employed as tools to perform the abstract idea. The use of such generic computing components for executing the abstract idea does not integrate the abstract idea into a practical application (see MPEP 2106.05(f)). Claim 30 further recites the steps of transmitting, via a communication interface, a task request to each of a plurality of UAV servers within an overlapping coverage area accessible by a terminal, wherein the task request comprises an identifier of the terminal, position information of the terminal, and/or task information of the terminal; and receiving a plurality of service decision instructions from the plurality of UAV servers generated based on the task request, wherein each of the plurality of service decision instructions comprises an indication of whether the corresponding UAV server should service the task request. These steps are insignificant extra-solution activity, because they are data gathering and data output steps that are necessary for performing the abstract idea (i.e., all uses of the abstract idea require such data gathering and data output). The recitation of such insignificant extra-solution activity does not integrate the abstract idea into a practical application (see MPEP 2106.05(g)). Claim 30 further specifies that the plurality of service decision instructions are generated using a decision network, wherein the decision network is obtained by adjusting network parameters of an initial decision network based on a plurality of reward and punishment constraints based on one or more of: available computing resources of the UAV server, available bandwidth of the UAV server, an exclusivity constraint that penalizes a decision to service the terminal by zero UAV servers of by more than one UAV server in the overlapping coverage area, an estimated execution time for the task request by the corresponding UAV server, and an execution time of a current training epoch, wherein the adjusting is performed cooperatively across the plurality of UAV servers. This limitation amounts to general linking of the abstract idea to the particular technological field of machine learning that is performed across multiple UAV servers. Using the decision network to generate the plurality of service decision instructions merely serves to employ machine learning technology to implement the abstract idea. The recitation of such general linking limitations do not integrate the abstract idea into a practical application (see MPEP 2106.05(h)).
Step 2B: The additional elements are re-evaluated in Step 2B to determine if they are more than what is well-understood, routine, conventional activity in the field. The specification does not provide any indication that the recited computer components are anything other than generic computing components used within a conventional computer. The use of such generic and conventional computer components for executing the abstract idea does not amount to significantly more than the abstract idea itself (see MPEP 2106.05(f)).
MPEP 2106.05(d)(II), and the cases cited therein, including Intellectual Ventures I, LLC v. Symantec Corp., 838 F.3d 1307, 1321 (Fed. Cir. 2016), TLI Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610 (Fed. Cir. 2016), and OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363 (Fed. Cir. 2015), indicate that mere receipt or transmission of data over a network is a well‐understood, routine, and conventional function when it is claimed in a merely generic manner (as it is here). Accordingly, the steps of receiving the task request, transmitting the service decision instruction, transmitting the task request, receiving the plurality of service decision instructions, and transmitting the selection and task request merely amount to insignificant extra-solution activity that does not amount to significantly more than the abstract idea itself (see MPEP 2106.05(g)).
The limitations specifying that recited steps are performed by a decision network that can be adjusted cooperatively across a plurality of UAV servers are considered general linking limitations that do not impose meaningful limits on the claims. These limitations merely serve to link the abstract idea to the technological field of machine learning that is performed across multiple UAV servers, and specify that certain steps of the abstract idea are executed using machine learning with no additional details of how this occurs. Therefore, these general linking limitations do not amount to significantly more than the abstract idea itself (see MPEP 2106.05(h)).
For the above reasons, the additional elements do not amount to significantly more than the abstract idea itself, whether considered individually or in combination. Therefore, when considering the combination of elements and the claimed invention as a whole, claims 11, 27, and 29-30 are not patent-eligible.
Regarding claims 26 and 28:
Dependent claims 26 and 28 recite limitations that further define the claimed mental process. Specifically, claim 26 specifies that the task information comprises a data size, a computation strength, and/or a maximum allowable time delay for the task request, and claim 28 specifies that the plurality of service decision instructions indicate a UAV server that should service the task request. These limitations do not preclude the mental process from being performed in the human mind with the help of pen and paper, and are accordingly being analyzed as additional mental process steps.
Since dependent claims 26 and 28 merely recite limitations further defining the mental process, these claims contain no additional elements which integrate the abstract idea into a practical application or amount to significantly more than the abstract idea itself. Therefore, when considering the combination of elements and the claimed invention as a whole, claims 26 and 28 are not patent-eligible.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 11-13, 16-20, 23-24, and 26-30 are rejected under 35 U.S.C. 103 as being unpatentable over Qin et al. (CN 110177379 A), hereinafter referred to as Qin, in view of Chai et al. (CN 115827108 A), hereinafter referred to as Chai, and further in view of Arksey et al. (US 2022/0399936 A1), hereinafter referred to as Arksey.
Regarding claim 11:
Qin discloses the following limitations:
“A computer-implemented method, comprising: receiving, by one or more processors, a task request from a terminal.” (Qin ¶ 10: “When a mobile terminal is detected to be in a multi-base station coverage area, the mobile terminal sends a connection initialization request to the multiple base stations covering the area.” Also, Qin ¶ 122: “Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware.”)
“wherein the task request comprises an identifier of the terminal, position information of the terminal, and/or task information of the terminal.” (Qin ¶ 14: “in a first possible implementation of the first aspect, the connection initialization request carries the location information of the mobile terminal, the amount of data to be processed by the mobile terminal, and the required QoS service level.” This at least teaches the task request comprising “position information of the terminal” and “task information of the terminal” as claimed. Further, Qin ¶ 12 discloses that each of the base stations send response messages to the mobile terminal, which implies that an identifier of the terminal is used for sending the response messages to the correct mobile terminal.)
“determining that the terminal lies within an overlapping service area between a… server and one or more other … servers based on the position information” and “generating… a service decision instruction based on the task request and the one or more other … servers.” (Qin ¶ 58: “the multi-base station coverage area refers to the overlapping area between the coverage areas of multiple different base stations. The connection initialization request sent by the mobile terminal is used to send information about the mobile terminal itself and the services it needs to the base station. The base station calculates and feeds back the corresponding indicator parameters so that the mobile terminal can select the most suitable base station from multiple base stations to establish a connection based on these indicator parameters.” Also, Qin ¶ 16 discloses that the indicator parameters can include an indication of “whether the QoS service level can be met.”)
“wherein the service decision instruction comprises an indication of whether the … server should service the task request.” (Qin ¶ 58: “The base station calculates and feeds back the corresponding indicator parameters so that the mobile terminal can select the most suitable base station from multiple base stations to establish a connection based on these indicator parameters.” Also, Qin ¶ 16 discloses that the indicator parameters can include an indication of “whether the QoS service level can be met.”)
“and transmitting the service decision instruction to the terminal using the identifier of the terminal.” (Qin ¶¶ 11-12 disclose that in response to the connection request, “The mobile terminal receives various indicator parameters sent by each of the base stations.” Additionally, Qin ¶ 58: “The base station calculates and feeds back the corresponding indicator parameters so that the mobile terminal can select the most suitable base station from multiple base stations to establish a connection based on these indicator parameters.” Note that sending the parameters to the mobile terminal in response to the connection request message implies that an identifier of the mobile terminal is used for sending the response messages to the correct mobile terminal.)
The following limitations are not specifically disclosed by Qin, but are taught by Chai:
The servers are “UAV servers.” (Chai ¶ 12: “This system consists of F terminal devices and M UAVs. Each UAV carries an MEC server to offload tasks within a fixed area.”)
The service decision instruction is generated “using a decision network.” (Chai ¶ 15: “Solve the task offloading model for minimizing latency and energy consumption in the UAV-mobile edge computing system using deep reinforcement learning. The solution method is as follows: construct a task offloading model for each offloading task solved by deep reinforcement learning through a multi-objective Markov decision process.”)
“wherein the generating comprises: calculating, using the decision network and based on state information of the UAV server and the task request, decision information of the UAV server” and “generating the service decision instruction based on the decision information.” (Chai ¶¶ 18-19: “Step 5: The agent in deep reinforcement learning begins to interact with the MEC environment. On the one hand, the agent obtains the current state from the MEC environment… Step 6: In deep reinforcement learning, the agent obtains the current Q value through Q-network training, selects action a in the current state s from action space A, and executes the action to obtain vector value reward r and the next state s´. The action space A includes the following two actions: executing tasks on the terminal device and unloading tasks to the UAV-mobile edge computing system.”)
“the decision information comprising an action decision of the UAV server, available computing resources of the UAV server… and an estimated execution time for a task.” (Chai ¶ 15 discloses to “Solve the task offloading model for minimizing latency and energy consumption in the UAV-mobile edge computing system using deep reinforcement learning,” which teaches the decision information including “available computing resources of the UAV server” as claimed. Further, Chai ¶ 19: “Step 6: In deep reinforcement learning, the agent obtains the current Q value through Q-network training, selects action a in the current state s from action space A, and executes the action to obtain vector value reward r and the next state s´. The action space A includes the following two actions: executing tasks on the terminal device and unloading tasks to the UAV-mobile edge computing system.” This at least teaches the decision information comprising “an action decision of the UAV server” as claimed. Also, Chai ¶¶ 25-27 disclose a constraint in which “The UAV can only fly within a specified rectangular area, and the horizontal range of time slot t and the maximum distance to fly within time slot t are also specified,” which teaches the decision information comprising “an estimated execution time for a task” as claimed.)
“wherein the decision network is obtained by adjusting network parameters of an initial decision network based on a plurality of reward and punishment constraints based on one or more of: available computing resources of the UAV server, available bandwidth of the UAV server, an exclusivity constraint that penalizes a decision to service the terminal by zero UAV servers or by more than one UAV server in the overlapping service area, an estimated execution time for the task request by the UAV server, and an execution time of a current training epoch.” (Chai ¶¶ 25-27: “Determine whether the current iteration has reached the maximum number of iterations. If so, output the optimal unloading decision, where the optimal unloading decision means that the vector value reward obtained by the agent after performing action a is maximized. Otherwise, go to step 5. Furthermore, the task dependency constraints include: Constraint 1: The UAV can only fly within a specified rectangular area, and the horizontal range of time slot t and the maximum distance to fly within time slot t are also specified.” This at least teaches the constraints being based on “an estimated execution time for the task request by the UAV server” as claimed.)
Note that under the broadest reasonable interpretation (BRI) of claim 11, consistent with the instant specification, the limitation “wherein the decision network is obtained by adjusting network parameters of an initial decision network based on a plurality of reward and punishment constraints based on one or more of: available computing resources of the UAV server, available bandwidth of the UAV server, an exclusivity constraint that penalizes a decision to service the terminal by zero UAV servers or by more than one UAV server in the overlapping service area, an estimated execution time for the task request by the UAV server, and an execution time of a current training epoch” is treated as an alternative limitation. Applicant has elected to use the phrase “one or more” in the claim language, and therefore, the BRI covers the scenario in which only one of the limitations applies. Accordingly, while only the “estimated execution time for the task request by the UAV server” has been addressed here, the claim is still rejected in its entirety.
Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method of Qin by using UAV servers as taught by Chai with a reasonable expectation of success. A person having ordinary skill in the art could have been motivated to do this because Chai ¶ 3 teaches that “The edge servers mounted on UAVs can expand their communication coverage and reduce geographical constraints, thereby improving deployment efficiency and user service quality. UAV-MEC has advantages such as high flexibility, wide coverage, faster response, and low cost.” Additionally, it would have been obvious to use a decision network to generate the service decision as taught by Chai, because Chai ¶ 3 teaches that “Machine learning-based methods can dynamically adjust offloading strategies in UAV-MEC environments to adapt to rapid environmental changes.” Further, it would have been obvious to use reward and punishment constraints for the task request as taught by Chai, because Chai ¶ 38 teaches that “This invention incorporates task dependency constraints in the modeling of UAV-MEC systems, thereby improving the utilization rate of computing resources.”
The following limitations are not specifically taught by the combination of Qin and Chai, but are taught by Arksey:
“the decision information comprising … available bandwidth of the UAV server.” (Arksey ¶ 59: “The base stations may constantly map the radio frequencies and bandwidth available to them and can reposition base stations to better locations to ensure high bandwidth to the cloud systems.”)
“wherein the adjusting is performed cooperatively across the UAV server and the one or more other UAV servers.” (Arksey ¶ 115: “The workflow may be built as an integrated whole rather that has a single vision-built representation for the 3D world used in planning, piloting, and processing. … Processing can be done in real-time enabling real-time collaboration. The system can build 3D models with sensor data and can allow virtual cloud collaboration and dynamic changes in drone operation depending on what is found, all in real-time.”)
Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method disclosed by the combination of Qin and Chai by considering available bandwidth and by allowing the servers to work cooperatively as taught by Arksey with a reasonable expectation of success. A person having ordinary skill in the art could have been motivated to do this because Arksey ¶¶ 103 and 116 teach that this can help to “reposition hives to better locations to provide high bandwidth to the cloud systems” and can help with “facilitating real-time consultation improving decision making cycle times.”
Regarding claim 12:
The combination of Qin, Chai, and Arksey teaches “The computer-implemented method of claim 11,” and Qin also teaches the method “further comprising: servicing the task request in response to a terminal selection of the … server based on the service decision instruction and service decision instructions transmitted to the terminal by the one or more other … servers.” (Qin ¶ 58: “The base station calculates and feeds back the corresponding indicator parameters so that the mobile terminal can select the most suitable base station from multiple base stations to establish a connection based on these indicator parameters.” Also, Qin ¶ 62: “enables the selection of the target base station with the best data processing effect among multiple base stations to establish a communication connection, which can make fuller use of network resources and improve the quality of user service.”)
As explained regarding claim 11 above, Chai teaches the servers being “UAV servers.”
Regarding claim 13:
The combination of Qin, Chai, and Arksey teaches “The computer-implemented method of claim 12,” and Qin further teaches “wherein the service decision instruction and the service decision instructions … only indicate one … server that should service the task request.” (Qin ¶ 45: “This allows the system to select the target base station with the best data processing performance from among multiple base stations to establish a communication connection, thereby making fuller use of network resources and improving the quality of user service.”)
Qin does not specifically disclose the servers being “UAV servers,” or that the instructions indicating one UAV server that should service the task request are “transmitted to the terminal by the one or more other UAV servers.” However, Chai does teach these limitations. (Chai ¶ 12: “This system consists of F terminal devices and M UAVs. Each UAV carries an MEC server to offload tasks within a fixed area.” Also, Chai ¶ 19: “The action space A includes the following two actions: executing tasks on the terminal device and unloading tasks to the UAV-mobile edge computing system.” Further, Chai ¶ 110 discloses to “Determine whether the training has ended, and then decide whether to output the unloading decision.” It is implied that the final decision would need to be output to the terminal device to facilitate completion of the task.)
Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method of Qin by using UAV servers as taught by Chai with a reasonable expectation of success. A person having ordinary skill in the art could have been motivated to do this because Chai ¶ 3 teaches that “The edge servers mounted on UAVs can expand their communication coverage and reduce geographical constraints, thereby improving deployment efficiency and user service quality. UAV-MEC has advantages such as high flexibility, wide coverage, faster response, and low cost.” Additionally, it would have been obvious to transmit the service decision instruction to the terminal as taught by Chai, because this is a simple substitution of one known element (i.e., allowing a remote system to make the decision and transmit the result to the terminal as taught by Chai) for another (i.e., allowing the terminal to make the decision and establish the connection itself as disclosed by Qin) to obtain predictable results (see MPEP 2143(I)(B)). A person having ordinary skill in the art could have substituted the transmitting a final result to the terminal with the step of the terminal making the final decision to achieve the predictable result of offloading the decision process to a remote system capable of making an optimized decision.
Regarding claim 16:
The combination of Qin, Chai, and Arksey teaches “The computer-implemented method of claim 11,” and Chai additionally teaches the method “further comprising: training the decision network by iterating over one or more training epochs using sample environment data comprising a plurality of task requests each corresponding to each of a plurality of state information sets of the UAV server.” (Chai ¶ 18 discloses that Step 5 ends with a step to “update the preference experience pool W using the current iteration number.” Additionally, Chai ¶ 21: “Step 8: Train the experience samples: First, randomly select a portion of the experience samples from the experience buffer pool Φ; then, select the experience preferences from the preference experience pool W using a non-dominated ranking method. Train both the Q-network and the target Q-network simultaneously to maximize the vector value reward and obtain the optimal unloading decision. During training, let the input of the Q-network be the current state s, the experience preference, and the current preference, and output the Q-value. Let the input of the target Q-network be the next state s´, the experience preference, and the current preference, and output the target Q-value.”)
Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method disclosed by the combination of Qin and Arksey by training the model using iterations of experience samples as is taught by Chai with a reasonable expectation of success. A person having ordinary skill in the art could have been motivated to do this because Chai ¶ 38 teaches that “This invention adopts a dynamic weight adjustment strategy and uses a Q-network to train and optimize the current user's preferences and the previous user preferences simultaneously. The previous user preferences are obtained from the preference experience pool through a non-dominated ranking method, which can better maintain the previously learned strategy.”
Regarding claim 17:
The combination of Qin, Chai, and Arksey teaches “The computer-implemented method of claim 16,” and Chai also teaches “wherein each state information set comprises a location of the UAV server, available computing resources of the UAV server, available bandwidth of the UAV server, and/or a number of users within a coverage area of the UAV server.” (Chai ¶ 56: “The COP (Computation Offloading Problem) is modeled as a multi-objective optimization problem with added task dependency constraints, aiming to simultaneously minimize the latency and energy consumption of the UAV-MEC system.” Also, Chai ¶¶ 69-76 teach the consideration of “channel bandwidth” when determining the “time for transmitting the task to the drone” and the “latency and energy consumption of a task.” This at least teaches that the state information comprises available bandwidth of the UAV server as claimed.)
Note that under the broadest reasonable interpretation (BRI) of claim 17, consistent with the instant specification, the limitation “wherein each state information set comprises a location of the UAV server, available computing resources of the UAV server, available bandwidth of the UAV server, and/or a number of users within a coverage area of the UAV server” is treated as an alternative limitation. Applicant has elected to use the term “and/or” in the claim language, and therefore, the BRI covers the scenario in which only one of the limitations applies. Accordingly, while only the “available bandwidth of the UAV server” has been addressed here, the claim is still rejected in its entirety.
Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method disclosed by the combination of Qin and Arksey by configuring the decision network to consider the channel bandwidth as taught by Chai with a reasonable expectation of success. A person having ordinary skill in the art could have been motivated to do this based on recognizing that a UAV server should not be selected to service the task request if the UAV server is not capable of completing the task due to limitations such as not having enough bandwidth; a person having ordinary skill in the art would have recognized that selecting a UAV server that is incapable of completing the task would be inefficient and would require a reselection of another UAV server for servicing the task request, which would increase the time and resources required for completing the process.
Regarding claim 18:
The combination of Qin, Chai, and Arksey teaches “The computer-implemented method of claim 16,” and Chai also teaches the following limitations:
“wherein the iterating over one or more training epochs comprises: collecting experiences inside an experience pool by interacting with the sample environment data.” (Chai ¶ 21: “Step 8: Train the experience samples: First, randomly select a portion of the experience samples from the experience buffer pool Φ.”)
“updating internal weights of the evaluation network based on evaluation values obtained by the evaluation network; and updating internal weights of the decision network based on each of the collected experiences.” (Chai ¶ 38 discloses that “This invention adopts a dynamic weight adjustment strategy and uses a Q-network to train and optimize the current user's preferences and the previous user preferences simultaneously. The previous user preferences are obtained from the preference experience pool through a non-dominated ranking method, which can better maintain the previously learned strategy.” Also, Chai ¶ 116: “the present invention can quickly adjust the target weight to cope with changes in user preferences, thereby meeting user needs.”)
Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method disclosed by the combination of Qin and Arksey by adjusting internal weights of the decision network based on collected samples from an experience pool as taught by Chai with a reasonable expectation of success. A person having ordinary skill in the art could have been motivated to do this because Chai ¶ 38 teaches that “This invention adopts a dynamic weight adjustment strategy and uses a Q-network to train and optimize the current user's preferences and the previous user preferences simultaneously. The previous user preferences are obtained from the preference experience pool through a non-dominated ranking method, which can better maintain the previously learned strategy.”
Regarding claim 19:
The combination of Qin, Chai, and Arksey teaches “The computer-implemented method of claim 18,” and Chai further teaches “wherein the updating the internal weights of the decision network occurs after completing a full training epoch.” (Chai ¶¶ 21-23: “During training, let the input of the Q-network be the current state s, the experience preference, and the current preference, and output the Q-value. Let the input of the target Q-network be the next state s´, the experience preference, and the current preference, and output the target Q-value. Calculate the loss function L … Finally, the Q-network is updated using the loss function value, and the Q-network parameters are synchronized to the target Q-network every 300 generations.” Further, Chai ¶ 38: “Multiple objectives are optimized simultaneously, and the weights are dynamically adjusted to meet different user preferences.”)
Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method disclosed by the combination of Qin and Arksey by adjusting internal weights of the decision network after a training epoch as taught by Chai with a reasonable expectation of success. A person having ordinary skill in the art could have been motivated to do this because Chai ¶ 38 teaches that “This invention adopts a dynamic weight adjustment strategy and uses a Q-network to train and optimize the current user's preferences and the previous user preferences simultaneously. The previous user preferences are obtained from the preference experience pool through a non-dominated ranking method, which can better maintain the previously learned strategy.”
Regarding claim 20:
The combination of Qin, Chai, and Arksey teaches “The computer-implemented method of claim 18,” and Chai also teaches the following limitations:
“wherein the collecting experiences comprises: inputting a first sample environment data into the decision network to obtain decision information of the UAV server; evaluating the decision information using the evaluation network to obtain a final reward value.” (Chai ¶¶ 29-32: “Select action a using the Double DQN method, and determine action a using two action value functions: one for estimating the action, and the other for estimating the value of the action. … Executing action a in the current state s yields the next state s' and a vector value reward r.”)
“determining a second sample environment data by applying the decision information to the first sample environment; and storing the first sample environment data, the second sample environment data, the decision information, and the reward value inside the experience pool, wherein the experience pool comprises experiences collected from the UAV server and the one or more other UAV servers.” (Chai ¶ 18: “On the one hand, the agent obtains the current state from the MEC environment. On the other hand, the MEC environment returns the current reward vector value and the next state based on the action selected by the agent. The agent obtains the current state from the MEC environment and updates the preference experience pool. The method for updating the preference experience pool is as follows: select the current preference from the preference space Ψ and determine whether the current preference is in the encountered preference experience pool W. If it does not exist, add the current preference to the preference experience pool W. Otherwise, update the preference experience pool W using the current iteration number.”)
Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method disclosed by the combination of Qin and Arksey by assessing a final reward value and performing multiple iterations to provide an updated experience pool as taught by Chai with a reasonable expectation of success. A person having ordinary skill in the art could have been motivated to do this because Chai ¶ 38 teaches that “This invention adopts a dynamic weight adjustment strategy and uses a Q-network to train and optimize the current user's preferences and the previous user preferences simultaneously. The previous user preferences are obtained from the preference experience pool through a non-dominated ranking method, which can better maintain the previously learned strategy.”
Regarding claim 23:
The combination of Qin, Chai, and Arksey teaches “The computer-implemented method of claim 20,” and Chai also teaches the following limitations:
“wherein the evaluating comprises: calculating one or more reward and punishment values from the decision information and the first sample environment based on a plurality of constraints.” (Chai ¶ 84: “Total Energy Consumption (MUE) includes the energy consumption of TU and UAV during missions and the energy consumption of UAV during flight. In addition, during the task unloading process, we also need to follow the following task dependency constraints…” Further, Chai ¶¶ 103-105: “This invention aims to minimize latency and energy consumption, but to maximize the reward value, the opposite of latency and energy consumption is taken. … Therefore, maximizing is equivalent to minimizing total latency and total energy consumption.”)
“and aggregating the one or more calculated reward and punishment values to obtain a final reward value corresponding to the decision information.” (Chai ¶¶ 32-34 disclose the calculation of “a vector value reward r” using the equation below based on “the reward value for latency and the reward value for energy consumption.” The equation shows that a summation (i.e., aggregation) is used to calculate the final reward value.)
PNG
media_image1.png
84
425
media_image1.png
Greyscale
Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method disclosed by the combination of Qin and Arksey by aggregating reward values to obtain a final reward value based on specified constraints as taught by Chai with a reasonable expectation of success. A person having ordinary skill in the art could have been motivated to do this since Chai ¶ 38 teaches that “This invention incorporates task dependency constraints in the modeling of UAV-MEC systems, thereby improving the utilization rate of computing resources,” and that “the multi-objective Markov decision process expands the reward value into a vector value reward, where each element corresponds to an objective. Multiple objectives are optimized simultaneously, and the weights are dynamically adjusted to meet different user preferences.”
Regarding claim 24:
The combination of Qin, Chai, and Arksey teaches “The computer-implemented method of claim 23,” and Chai also teaches the method “further comprising: weighting the one or more calculated reward and punishment values using a reward factor.” (Chai ¶¶ 21-23 teach the calculation of a loss function L using the equations below, where “γ represents the reward discount factor.”)
PNG
media_image2.png
212
621
media_image2.png
Greyscale
Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method disclosed by the combination of Qin and Arksey by incorporating a reward factor for weighting the rewards as taught by Chai with a reasonable expectation of success. A person having ordinary skill in the art could have been motivated to do this because Chai ¶ 38 teaches that “the multi-objective Markov decision process expands the reward value into a vector value reward, where each element corresponds to an objective. Multiple objectives are optimized simultaneously, and the weights are dynamically adjusted to meet different user preferences.”
Regarding claim 26:
The combination of Qin, Chai, and Arksey teaches “The computer-implemented method of claim 11,” and Qin further teaches “wherein the task information comprises a data size, a computation strength, and/or a maximum allowable time delay for the task request.” (Qin ¶ 14: “the connection initialization request carries the location information of the mobile terminal, the amount of data to be processed by the mobile terminal, and the required QoS service level.” This at least teaches the task information comprising a data size as claimed.)
Note that under the broadest reasonable interpretation (BRI) of claim 26, consistent with the specification, the limitation that “the task information comprises a data size, a computation strength, and/or a maximum allowable time delay for the task request” is treated as an alternative limitation. Applicant has elected to use the term “and/or” in the claim language, and therefore, the BRI covers the scenario in which only one of the limitations applies. Accordingly, while only the “data size” has been addressed here, the claim is still rejected in its entirety.
Regarding claim 27:
Qin discloses the following limitations:
“A computer-implemented method comprising: transmitting a task request to each of a plurality of … servers within an overlapping coverage area accessible by a terminal.” (Qin ¶ 10: “When a mobile terminal is detected to be in a multi-base station coverage area, the mobile terminal sends a connection initialization request to the multiple base stations covering the area.” Also, Qin ¶ 122: “Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware.”)
“wherein the task request comprises an identifier of the terminal, position information of the terminal, and/or task information of the terminal.” (Qin ¶ 14: “in a first possible implementation of the first aspect, the connection initialization request carries the location information of the mobile terminal, the amount of data to be processed by the mobile terminal, and the required QoS service level.” This at least teaches the task request comprising “position information of the terminal” and “task information of the terminal” as claimed. Further, Qin ¶ 12 discloses that each of the base stations send response messages to the mobile terminal, which implies that an identifier of the terminal is used for sending the response messages to the correct mobile terminal.)
“receiving a plurality of service decision instructions from the plurality of … servers, wherein each of the plurality of service decision instructions comprises an indication of whether the corresponding … server should service the task request and wherein each of the plurality of service decision instructions is generated … based on the task request.” (Qin ¶ 58: “the multi-base station coverage area refers to the overlapping area between the coverage areas of multiple different base stations. The connection initialization request sent by the mobile terminal is used to send information about the mobile terminal itself and the services it needs to the base station. The base station calculates and feeds back the corresponding indicator parameters so that the mobile terminal can select the most suitable base station from multiple base stations to establish a connection based on these indicator parameters.” Also, Qin ¶ 16 discloses that the indicator parameters can include an indication of “whether the QoS service level can be met.”)
“selecting a … server from the plurality of … servers to service the task request based on the plurality of service instructions.” (Qin ¶ 58: “The base station calculates and feeds back the corresponding indicator parameters so that the mobile terminal can select the most suitable base station from multiple base stations to establish a connection based on these indicator parameters.”)
“and transmitting the selection and the task request to the … server to service the task request.” (Qin ¶ 58: “the mobile terminal can select the most suitable base station from multiple base stations to establish a connection based on these indicator parameters.” Additionally, Qin ¶ 62: “enables the selection of the target base station with the best data processing effect among multiple base stations to establish a communication connection, which can make fuller use of network resources and improve the quality of user service.”)
The following limitations are not specifically disclosed by Qin, but are taught by Chai:
The servers are “UAV servers.” (Chai ¶ 12: “This system consists of F terminal devices and M UAVs. Each UAV carries an MEC server to offload tasks within a fixed area.”)
The service decision instructions are generated “using a decision network.” (Chai ¶ 15: “Solve the task offloading model for minimizing latency and energy consumption in the UAV-mobile edge computing system using deep reinforcement learning. The solution method is as follows: construct a task offloading model for each offloading task solved by deep reinforcement learning through a multi-objective Markov decision process.”)
“wherein each of the plurality of service decision instructions is generated based on decision information comprising an action decision of the corresponding UAV server, available computing resources of the corresponding UAV server… and an estimated execution time for a task.” (Chai ¶ 15 discloses to “Solve the task offloading model for minimizing latency and energy consumption in the UAV-mobile edge computing system using deep reinforcement learning,” which teaches the decision information including “available computing resources of the UAV server” as claimed. Further, Chai ¶ 19: “Step 6: In deep reinforcement learning, the agent obtains the current Q value through Q-network training, selects action a in the current state s from action space A, and executes the action to obtain vector value reward r and the next state s´. The action space A includes the following two actions: executing tasks on the terminal device and unloading tasks to the UAV-mobile edge computing system.” This teaches the decision information comprising “an action decision of the UAV server” as claimed. Also, Chai ¶¶ 25-27 disclose a constraint in which “The UAV can only fly within a specified rectangular area, and the horizontal range of time slot t and the maximum distance to fly within time slot t are also specified,” which teaches the decision information comprising “an estimated execution time for a task” as claimed.)
“and wherein the decision network is obtained by adjusting network parameters of an initial decision network based on a plurality of reward and punishment constraints based on one or more of: available computing resources of the corresponding UAV server, available bandwidth of the corresponding UAV server, an exclusivity constraint that penalizes a decision to service the terminal by zero UAV servers or by more than one UAV server in the overlapping coverage area, an estimated execution time for the task request by the corresponding UAV server, and an execution time of a current training epoch.” (Chai ¶¶ 25-27: “Determine whether the current iteration has reached the maximum number of iterations. If so, output the optimal unloading decision, where the optimal unloading decision means that the vector value reward obtained by the agent after performing action a is maximized. Otherwise, go to step 5. Furthermore, the task dependency constraints include: Constraint 1: The UAV can only fly within a specified rectangular area, and the horizontal range of time slot t and the maximum distance to fly within time slot t are also specified.” This at least teaches the constraints being based on “an estimated execution time for the task request by the UAV server” as claimed.)
Note that under the broadest reasonable interpretation (BRI) of claim 27, consistent with the instant specification, the limitation “wherein the decision network is obtained by adjusting network parameters of an initial decision network based on a plurality of reward and punishment constraints based on one or more of: available computing resources of the corresponding UAV server, available bandwidth of the corresponding UAV server, an exclusivity constraint that penalizes a decision to service the terminal by zero UAV servers or by more than one UAV server in the overlapping coverage area, an estimated execution time for the task request by the corresponding UAV server, and an execution time of a current training epoch” is being treated as an alternative limitation. Applicant has elected to use the phrase “one or more” in the claim language, and therefore, the BRI covers the scenario in which only one of the limitations applies. Accordingly, while only the “estimated execution time for the task request by the corresponding UAV server” has been addressed here, the claim is still rejected in its entirety.
Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method of Qin by using UAV servers as taught by Chai with a reasonable expectation of success. A person having ordinary skill in the art could have been motivated to do this because Chai ¶ 3 teaches that “The edge servers mounted on UAVs can expand their communication coverage and reduce geographical constraints, thereby improving deployment efficiency and user service quality. UAV-MEC has advantages such as high flexibility, wide coverage, faster response, and low cost.” Additionally, it would have been obvious to use a decision network to generate the service decision as taught by Chai, because Chai ¶ 3 teaches that “Machine learning-based methods can dynamically adjust offloading strategies in UAV-MEC environments to adapt to rapid environmental changes.” Further, it would have been obvious to use reward and punishment constraints for the task request as taught by Chai, because Chai ¶ 38 teaches that “This invention incorporates task dependency constraints in the modeling of UAV-MEC systems, thereby improving the utilization rate of computing resources.”
The following limitations are not specifically taught by the combination of Qin and Chai, but are taught by Arksey:
the “decision information comprising … available bandwidth of the corresponding UAV server.” (Arksey ¶ 59: “The base stations may constantly map the radio frequencies and bandwidth available to them and can reposition base stations to better locations to ensure high bandwidth to the cloud systems.”)
“wherein the adjusting is performed cooperatively across the plurality of UAV servers.” (Arksey ¶ 115: “The workflow may be built as an integrated whole rather that has a single vision-built representation for the 3D world used in planning, piloting, and processing. … Processing can be done in real-time enabling real-time collaboration. The system can build 3D models with sensor data and can allow virtual cloud collaboration and dynamic changes in drone operation depending on what is found, all in real-time.”)
Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method disclosed by the combination of Qin and Chai by considering available bandwidth and by allowing the servers to work cooperatively as taught by Arksey with a reasonable expectation of success. A person having ordinary skill in the art could have been motivated to do this because Arksey ¶¶ 103 and 116 teach that this can help to “reposition hives to better locations to provide high bandwidth to the cloud systems” and can help with “facilitating real-time consultation improving decision making cycle times.”
Regarding claim 28:
The combination of Qin, Chai, and Arksey teaches “The computer-implemented method of claim 27,” and Qin also teaches “wherein the plurality of service decision instructions indicate a … server that should service the task request.” (Qin ¶ 45: “This allows the system to select the target base station with the best data processing performance from among multiple base stations to establish a communication connection, thereby making fuller use of network resources and improving the quality of user service.”)
As explained regarding claim 27 above, Chai teaches the servers being “UAV servers.”
Regarding claim 29:
Qin discloses the following limitations:
“A service decision device comprising: one or more processors configured to: receive, via a communication interface, a task request sent by a terminal in an overlapping service area between a… server and one or more other … servers.” (Qin ¶ 122: “Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware.” Further, Qin ¶ 10: “When a mobile terminal is detected to be in a multi-base station coverage area, the mobile terminal sends a connection initialization request to the multiple base stations covering the area.” It is implied that the base stations each include an interface for receiving the connection initialization request.)
“wherein the task request comprises an identifier of the terminal, position information of the terminal and/or task information of the terminal.” (Qin ¶ 14: “in a first possible implementation of the first aspect, the connection initialization request carries the location information of the mobile terminal, the amount of data to be processed by the mobile terminal, and the required QoS service level.” This at least teaches the task request comprising “position information of the terminal” and “task information of the terminal” as claimed. Further, Qin ¶ 12 discloses that each of the base stations send response messages to the mobile terminal, which implies that an identifier of the terminal is used for sending the response messages to the correct mobile terminal.)
“and generate… a service decision instruction based on the task request” and “wherein the service decision instruction comprises an indication of whether the … server should service the task request.” (Qin ¶ 58: “The base station calculates and feeds back the corresponding indicator parameters so that the mobile terminal can select the most suitable base station from multiple base stations to establish a connection based on these indicator parameters.” Further, Qin ¶ 16 discloses that the indicator parameters can include an indication of “whether the QoS service level can be met.”)
“and to transmit, via the communication interface, the service decision instruction to the terminal using the identifier of the terminal.” (Qin ¶¶ 11-12 disclose that in response to the connection request, “The mobile terminal receives various indicator parameters sent by each of the base stations.” Additionally, Qin ¶ 58: “The base station calculates and feeds back the corresponding indicator parameters so that the mobile terminal can select the most suitable base station from multiple base stations to establish a connection based on these indicator parameters.” Note that sending the parameters to the mobile terminal in response to the connection request message implies that an identifier of the mobile terminal is used for sending the response messages to the correct mobile terminal.)
“and wherein the service decision instruction is used for the terminal to select among the … server and the one or more other … servers to service the task request based on the service decision instruction and service decision instructions transmitted by the one or more other … servers.” (Qin ¶ 58: “The base station calculates and feeds back the corresponding indicator parameters so that the mobile terminal can select the most suitable base station from multiple base stations to establish a connection based on these indicator parameters.”)
The following limitations are not specifically disclosed by Qin, but are taught by Chai:
The servers are “UAV servers.” (Chai ¶ 12: “This system consists of F terminal devices and M UAVs. Each UAV carries an MEC server to offload tasks within a fixed area.”)
The service decision instruction is generated “using a decision network.” (Chai ¶ 15: “Solve the task offloading model for minimizing latency and energy consumption in the UAV-mobile edge computing system using deep reinforcement learning. The solution method is as follows: construct a task offloading model for each offloading task solved by deep reinforcement learning through a multi-objective Markov decision process.”)
“wherein the generating comprises: calculating, using the decision network and based on state information of the UAV server and the task request, decision information of the UAV server” and “generating the service decision instruction based on the decision information.” (Chai ¶¶ 18-19: “Step 5: The agent in deep reinforcement learning begins to interact with the MEC environment. On the one hand, the agent obtains the current state from the MEC environment… Step 6: In deep reinforcement learning, the agent obtains the current Q value through Q-network training, selects action a in the current state s from action space A, and executes the action to obtain vector value reward r and the next state s´. The action space A includes the following two actions: executing tasks on the terminal device and unloading tasks to the UAV-mobile edge computing system.”)
“the decision information comprising an action decision of the UAV server, available computing resources of the UAV server… and an estimated execution time for a task.” (Chai ¶ 15 discloses to “Solve the task offloading model for minimizing latency and energy consumption in the UAV-mobile edge computing system using deep reinforcement learning,” which teaches the decision information including “available computing resources of the UAV server” as claimed. Further, Chai ¶ 19: “Step 6: In deep reinforcement learning, the agent obtains the current Q value through Q-network training, selects action a in the current state s from action space A, and executes the action to obtain vector value reward r and the next state s´. The action space A includes the following two actions: executing tasks on the terminal device and unloading tasks to the UAV-mobile edge computing system.” This teaches the decision information comprising “an action decision of the UAV server” as claimed. Also, Chai ¶¶ 25-27 disclose a constraint in which “The UAV can only fly within a specified rectangular area, and the horizontal range of time slot t and the maximum distance to fly within time slot t are also specified,” which teaches the decision information comprising “an estimated execution time for a task” as claimed.)
“wherein the decision network is obtained by adjusting network parameters of an initial decision network based on a plurality of reward and punishment constraints based on one or more of: available computing resources of the UAV server, available bandwidth of the UAV server, an exclusivity constraint that penalizes a decision to service the terminal by zero UAV servers or by more than one UAV server in the overlapping service area, an estimated execution time for the task request by the UAV server, and an execution time of a current training epoch.” (Chai ¶¶ 25-27: “Determine whether the current iteration has reached the maximum number of iterations. If so, output the optimal unloading decision, where the optimal unloading decision means that the vector value reward obtained by the agent after performing action a is maximized. Otherwise, go to step 5. Furthermore, the task dependency constraints include: Constraint 1: The UAV can only fly within a specified rectangular area, and the horizontal range of time slot t and the maximum distance to fly within time slot t are also specified.” This at least teaches the constraints being based on “an estimated execution time for the task request by the UAV server” as claimed.)
Note that under the broadest reasonable interpretation (BRI) of claim 29, consistent with the instant specification, the limitation “wherein the decision network is obtained by adjusting network parameters of an initial decision network based on a plurality of reward and punishment constraints based on one or more of: available computing resources of the UAV server, available bandwidth of the UAV server, an exclusivity constraint that penalizes a decision to service the terminal by zero UAV servers or by more than one UAV server in the overlapping service area, an estimated execution time for the task request by the UAV server, and an execution time of a current training epoch” is treated as an alternative limitation. Applicant has elected to use the phrase “one or more” in the claim language, and therefore, the BRI covers the scenario in which only one of the limitations applies. Accordingly, while only the “estimated execution time for the task request by the UAV server” has been addressed here, the claim is still rejected in its entirety.
Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the system of Qin by using UAV servers as taught by Chai with a reasonable expectation of success. A person having ordinary skill in the art could have been motivated to do this because Chai ¶ 3 teaches that “The edge servers mounted on UAVs can expand their communication coverage and reduce geographical constraints, thereby improving deployment efficiency and user service quality. UAV-MEC has advantages such as high flexibility, wide coverage, faster response, and low cost.” Additionally, it would have been obvious to use a decision network to generate the service decision as taught by Chai, because Chai ¶ 3 teaches that “Machine learning-based methods can dynamically adjust offloading strategies in UAV-MEC environments to adapt to rapid environmental changes.” Further, it would have been obvious to use reward and punishment constraints for the task request as taught by Chai, because Chai ¶ 38 teaches that “This invention incorporates task dependency constraints in the modeling of UAV-MEC systems, thereby improving the utilization rate of computing resources.”
The following limitations are not specifically taught by the combination of Qin and Chai, but are taught by Arksey:
“the decision information comprising … available bandwidth of the UAV server.” (Arksey ¶ 59: “The base stations may constantly map the radio frequencies and bandwidth available to them and can reposition base stations to better locations to ensure high bandwidth to the cloud systems.”)
“wherein the adjusting is performed cooperatively across the UAV server and the one or more other UAV servers.” (Arksey ¶ 115: “The workflow may be built as an integrated whole rather that has a single vision-built representation for the 3D world used in planning, piloting, and processing. … Processing can be done in real-time enabling real-time collaboration. The system can build 3D models with sensor data and can allow virtual cloud collaboration and dynamic changes in drone operation depending on what is found, all in real-time.”)
Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the system disclosed by the combination of Qin and Chai by considering available bandwidth and by allowing the servers to work cooperatively as taught by Arksey with a reasonable expectation of success. A person having ordinary skill in the art could have been motivated to do this because Arksey ¶¶ 103 and 116 teach that this can help to “reposition hives to better locations to provide high bandwidth to the cloud systems” and can help with “facilitating real-time consultation improving decision making cycle times.”
Regarding claim 30:
Qin discloses “A service decision device comprising: one or more processors configured to” perform a process and a “communication interface” configured to transmit information. (Qin ¶¶ 10-12: “the mobile terminal sends a connection initialization request to the multiple base stations covering the area. In response to the base station receiving the connection initialization request sent by the mobile terminal, the base station calculates various indicator parameters based on the connection initialization request. The mobile terminal receives various indicator parameters sent by each of the base stations.” This implies that the mobile terminal includes an interface for sending the connection initialization request and receiving the indicator parameters. Also, Qin ¶ 122: “Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware.”)
The remaining limitations of claim 30 are taught by the combination of Qin, Chai, and Arksey using the same rationale applied to claim 27 above, mutatis mutandis.
Claims 21-22 are rejected under 35 U.S.C. 103 as being unpatentable over Qin in view of Chai and Arksey as applied to claim 20 above, and further in view of Zhan et al. (the non-patent article “Twin Delayed Multi-Agent Deep Deterministic Policy Gradient”), hereinafter referred to as Zhan.
Regarding claim 21:
The combination of Qin, Chai, and Arksey teaches “The computer-implemented method of claim 20,” but does not explicitly teach the limitations listed below. However, Zhan does teach these limitations:
“wherein: the evaluation network comprises a first evaluation model and a second evaluation model.” (Zhan p. 50 § C: “The model consists of multiple DDPG networks, each of which learns policy π (actor) and action value Q (critic), and has a target network for off-policy learning of Q-learning.”)
“and the updating the internal weights of the evaluation network comprises: comparing a first evaluation value by the first evaluation model and a second evaluation value by the second evaluation value to obtain a minimum evaluation value, calculating an error between the minimum evaluation value and a target evaluation value.” (Zhan p. 50 -§ E and FIG. 2 reproduced below: “With reference to the TD3 algorithm, the basic idea of this article is to use two sets of networks to represent different Q values, and by selecting the smallest one as our updated target (target Q value), the continuous overestimation is suppressed. The target network Q1(a’) and Q2(a’) in the above the Figure 2. take the minimum value min(Q1,Q2), instead of MADDPG’s Q’(a’) to calculate the update target.”)
PNG
media_image3.png
344
689
media_image3.png
Greyscale
“and updating the internal weights of the first evaluation model and the second evaluation model based on the calculated error using differential learning.” (Zhan p. 51 last paragraph: “Actor adjusts the strategy (actor neural network parameters) according to the critic’s score and tries to do better next time. The critic adjusts the score strategy (critic neural network parameters) based on the compensation provided by the system and the scores of other judges.”)
Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method disclosed by the combination of Qin, Chai, and Arksey by adjusting the decision network parameters based on a comparison of two evaluation models as taught by Zhan with a reasonable expectation of success. A person having ordinary skill in the art could have been motivated to do this because Zhan Abstract teaches that this “improves the actor network and critic network, optimizes the overestimation of Q value and adopts the update delayed method to make the actor training more stable.”
Regarding claim 22:
The combination of Qin, Chai, and Arksey teaches “The computer-implemented method of claim 20,” but does not explicitly teach “wherein the evaluation network is implemented using a multi-agent twin delayed deep deterministic policy gradient algorithm.” However, Zhan does teach this limitation. (Zhan proposes a “Twin Delayed Multi-Agent Deep Deterministic Policy Gradient,” where Zhan Abstract states that “Based on the traditional multi-agent reinforcement learning algorithm, this paper improves the actor network and critic network, optimizes the overestimation of Q value and adopts the update delayed method to make the actor training more stable.”)
Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the method disclosed by the combination of Qin, Chai, and Arksey by using a twin delayed multi-agent deep deterministic policy gradient as taught by Zhan with a reasonable expectation of success. A person having ordinary skill in the art could have been motivated to do this because Zhan Abstract teaches that this “improves the actor network and critic network, optimizes the overestimation of Q value and adopts the update delayed method to make the actor training more stable.”
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Madison R Inserra whose telephone number is (571)272-7205. The examiner can normally be reached Monday - Friday: 9:30 AM - 6:30 PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Aniss Chad can be reached at 571-270-3832. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Madison R. Inserra/Primary Examiner, Art Unit 3662