DETAILED ACTION
This action is responsive to the application filed on 02/12/2026. Claims 1-16 and 22, and 24 are pending and have been examined. This action is Final.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Applicant’s claim for the benefit of a prior-filed application under 35 U.S.C. 119(e) or under 35 U.S.C. 120, 121, 365(c), or 386(c) is acknowledged.
Response to Argument
Argument 1: Applicant argues on pages 16-21 that the claims are patent eligible because they are directed to a specific technical improvement in computer and RAN network functionality rather than an abstract idea. Applicant asserts that claim 1 recites a concrete architecture in which learning information is generated outside the learning learner only upon request, provided to the learning learner, and then deleted, which allegedly reduces storage burden, prevents data redundancy, and improves resource efficiency in a RAN control environment. Applicant also argues that the claim cannot practically be performed mentally because it involves execution of multiple operation determination models, reward-value calculations for reinforcement learning, separation of operation determination models from learning learners, and lifecycle management of learning information across system components. For Step 2A and Step 2B, applicant argues that the claim integrates any alleged abstract idea into a practical application and includes significantly more because the temporary external generation and deletion of learning information is allegedly non-conventional, solves a technical network-resource problem, and improves efficiency and scalability of RAN control systems.
Examiner Response to Argument 1: The examiner has considered applicant’s arguments, but they are not persuasive. As mapped, claim 1 still recites identifying parameter values and operation information in response to a request, producing learning information based on parameter values, operation information, and a reward value, providing that information to a learning learner, and then deleting the information after provision. These limitations remain directed to data identification, data evaluation, reward/value determination, information generation, transmission, and deletion, rather than a specific improvement to computer functionality or RAN operation. Although applicant characterizes the claim as a concrete RAN architecture, the claim does not recite a particular model architecture, training algorithm, memory-management structure, deletion protocol, cache-control mechanism, RAN-control protocol, or technical rule that changes how the computer or RAN operates. The recited plurality of operation determination models, separation from learning learners, request-based production, temporary external generation, and deletion after provision merely define where and when generic data handling occurs. Under Step 2A Prong 1, the identifying and reward-based production limitations remain mental processes and/or mathematical concepts because they involve selecting data and evaluating a reward value from parameter information. Under Step 2A Prong 2, the additional RAN/electronic-device limitations do not integrate the abstract idea into a practical application because the claim recites desired results, such as reducing storage burden or avoiding redundancy, without claiming a specific technical mechanism that achieves those results. Under Step 2B, the added limitations also do not provide significantly more because receiving requests, storing data, generating responsive information, transmitting information, and deleting information after use are well-understood, routine, and conventional computer data-management functions when claimed at this level of generality. Accordingly, the amended limitations have been considered, but they do not transform the abstract data-selection/reward-evaluation process into a patent-eligible technical improvement, and the rejection under 35 U.S.C. 101 is maintained.
Argument 2: Applicant argues on pages 12-16 that the cited combination of Katti, Devitt, and Shi fails to teach or suggest the amended claim limitations requiring learning information to be temporarily produced external to the learning learner, provided to the learning learner, and then deleted after provision. Applicant contends that Shi’s cited teachings about temporarily unavailable resources and state/resource allocation only relate to network-resource blocking or internal Q-learning state transitions, not to an external component generating learning information on request and deleting it after use. Applicant further argues that the Office Action improperly equates the Q-learning agent, operation determination model, and learning learner, whereas claim 1 allegedly requires a structural distinction between an operation determination model that outputs RAN operations and a corresponding learning learner that trains the model. Applicant also argues that new claim 24 is patentably distinct because it requires identifying a first parameter based on identification information of the learning learner included in the request, which allegedly enables selecting a parameter set optimized for a specific learning learner in a multi-learner environment.
Examiner Response to Argument 2: The examiner has considered applicant’s arguments, but they are not persuasive because the present rejection no longer relies on Shi as the basis for the amended limitations requiring the information for learning to be temporarily produced, external to the learning learner, and deleted after provision. As set forth in the updated mapping, Katti is relied upon for the RAN/RIC framework, including stored RAN state in the R-NIB, A1 and E2 information exchange, separated RIC non-RT learning and model-production functions, RIC near-RT runtime execution functions, and providing data feeds to train AI models. Thus, the rejection does not improperly equate the Q-learning agent, operation determination model, and learning learner, but instead maps Katti’s separated non-RT and near-RT architecture to the claimed distinction between learning learners and operation determination models. Devitt is relied upon for identifying relevant parameter or event values and producing performance or reward-type information based on parameter values and decision combinations. Most importantly, Bendre is relied upon for the amended lifecycle limitations because Bendre teaches temporary data storage for training data while an ML trainer process serves a corresponding ML training request, teaches that the ML trainer process refers to training data stored in temporary data storage to generate an ML model, and teaches deleting the training data from temporary data storage once the corresponding ML training request is complete. Therefore, applicant’s argument that Shi only teaches temporary resource blocking does not overcome the rejection because the rejection uses Bendre, not Shi, to teach temporary production or storage of learning information external to the learner and deletion after the learning or training request has been served. With respect to claim 24, Bendre further teaches request-based identification of training data and target-variable information for the requested ML model and assignment of the ML training request to a particular trainer process, which the examiner has reasonably mapped to identifying the relevant parameter information based on identification information associated with the learning learner or request. Accordingly, the combination of Katti, Devitt, and Bendre teaches or suggests the amended limitations, and the 35 U.S.C. 103 rejection is maintained.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-16 and 22, and 24 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: The claim directed to a method of operating an electronic device, which is one of the four statutory categories. Therefore, the claim satisfies Step 1.
Step 2A Prong 1:
(a) “in response to the request, identifying, among the at least one value, a value corresponding to at least one first parameter corresponding to the operation determination model and information associated with an operation corresponding to the at least one first parameter” -- This limitation is directed to identifying and selecting particular data based on a received request and a relationship between a parameter, a model, and operation information. Identifying a value and corresponding information based on observed data and a request can be performed in the human mind using observation, evaluation, and judgment, and therefore is directed to a mental process.
Step 2A Prong 2 and Step 2B:
(a) “A non-transitory computer-readable storage medium for storing instructions which, when executed individually and/or collectively by at least one processor of an electronic device, control the electronic device to perform:” -- The limitation recites a non-transitory CRM for storing instructions that will be recited executions to apply onto a computer/device and how the device will be controlled. The limitation amounts to no more than mere instructions to apply onto a computer, and it does not integrate to a practical application, nor does it provide significantly more than the judicial exception (see MPEP 2106.05(f)).
(b) “storing at least one value corresponding to each of a plurality of parameters associated with a radio access network (RAN), and information associated with an operation performed by the RAN” -- This limitation is directed to storing data related to RAN parameters and RAN operations. Storing data is mere data gathering or electronic recordkeeping, which is insignificant extra-solution activity and does not integrate the exception into a practical application (see MPEP 2106.05(g)). Further, storing and retrieving data in memory is well-understood, routine, and conventional activity and does not provide significantly more than the judicial exception (see MPEP 2106.05(d)(II)).
(c) “wherein a value corresponding to one or more parameters among the plurality of parameters is used based on a plurality of operation determination models executable by the electronic device to determine at least a part of the information associated with the operation, wherein each of the plurality of operation determination models is configured to output a respective operation to the RAN for the RAN to operate based on the respective operation, and each of the plurality of operation determination models is distinct from a respective learning learner and corresponds to the respective learning learner among a plurality of learning learners;” -- This limitation recites that parameter values are used by executable models and that the models are distinct from corresponding learners. However, the claim does not recite a specific model architecture, training algorithm, data structure, memory-management structure, or RAN-control protocol that improves the functioning of the computer or RAN itself. Rather, the limitation merely defines the environment and arrangement in which the abstract data identification, reward evaluation, and information production occur, which does not integrate to a practical application, nor provide significantly more than the judicial exception (see MPEP 2106.05(f)).
(d) “receiving, from a learning learner among the plurality of learning learners for learning and operation determination model, a request for information for learning the operation determination model selected from among the plurality of operation determination models…providing the information for learning the operation determination model to the learning learner” -- The limitation recites receiving a request for data and transmitting responsive data to another component. Receiving, sending, and outputting data are insignificant extra-solution activities and are well-understood, routine, and conventional computer functions, which does not integrate to a practical application (see MPEP 2106.05(g)). Furthermore, under Step 2B, sending/receiving data over a network is a well-understood, and routine activity (WURC) and does not provide significantly more than the judicial exception (see MPEP 2106.05(d)(II)).
(e) “after the information for learning the operation determination model is provided to the learning learner, deleting the information for learning the operation determination model” -- These limitations recite temporary generation and deletion of information outside the learning learner. However, the claim does not recite a particular technical mechanism for temporary storage, deletion, memory reclamation, cache control, data lifecycle enforcement, or learner-specific data management. Instead, the claim recites the result of producing information temporarily and deleting it after transmission. The limitation amounts to no more than mere further limiting to a field of use/environment, and it does not integrate to a practical application, nor provides significantly more than the judicial exception (see MPEP 2106.05(h)).
(f) “temporarily producing, external to the learning learner, the information for learning the operation determination model based on the value corresponding to the at least one first parameter, the information associated with the operation corresponding to the at least one first parameter, and a reward value identified based on at least a part of the value corresponding to the at least one first parameter” -- This limitation is directed to producing learning information based on parameter values, operation information, and a reward value. The limitation is directed to producing information for a model based on gathered collected data and associated parameters, which amounts to no more than mere further limiting to a field of use/environment, and thus it does not integrate to a practical application (see MPEP 2106.05(g)). Furthermore, under Step 2B, the claimed limitation also is a well-understood, routine and convention activity (WURC) that cannot provide significantly more than the judicial exception (see MPEP 21067.05(d)(II)).
Therefore, claim 1 is non-patent eligible under 35 USC 101. Claims 14 and 22 is analogous to claim 1, aside from claim type, and thus faces the same rejection as above. For claim 22, the main claim difference is the claim type and the below limitation in the preamble:
“An electronic device, comprising: a storage device storing instructions; and at least one processor, comprising processing circuitry, operatively connected to the storage device, wherein the instructions, when executed individually and/or collectively by the at least one processor, cause the electronic device to:” -- The limitation recites instructions to apply onto a computer a device, now comprising circuitry, to be executed, and it does not integrate to a practical application, nor provide significantly more than the judicial exception.
Regarding claim 2,
Step 1: The claim is directed to method, which is considered to be a process. The claim satisfies step 1.
Step 2A Prong 1:
“classifying” – The limitation is directed to classifying, in the context and broadest reasonable interpretation of the claim as a whole, values that will correspond to a parameter and information associated with the operation performed by the computer. The act of classifying is directed to a process that can be performed using evaluation, judgement, and observation in the human mind with aid of pen and paper, and thus is directed to a mental process.
Step 2A Prong 2 and Step 2B:
“The method of claim 1, wherein the storing of the at least one value corresponding to each of the plurality of parameters associated with the RAN, and the information associated with the operation performed by the RAN comprises: and storing, for each of a plurality of points in time, the at least one value corresponding to each of the at least one parameter and the information associated with the operation performed by the RAN.” – The limitation recites that instructions of storing the value that that will correspond to a parameter of the (neural network/computer) will comprise storing values for multiple points in time, that will correspond with a parameter and information associated by an operation of the computer (network), which is considered to be an insignificant, extra-solution activity that cannot be integrated to a practical application (see MPEP 2106.05(g)). Furthermore, under step 2B, the act of storing and/or retrieving data from memory (RAN) and electronic recordkeeping is a well-understood, routine, and conventional activity, and cannot provide significantly more than the judicial exception (see MPEP 2106.05(d)(II)).
Thus, claim 2 is non-patent eligible.
Regarding claim 3,
Step 1: The claim is directed to method, which is considered to be a process. The claim satisfies step 1.
There are no elements to be evaluated under Step 2A Prong 1.
Step 2A Prong 2 and Step 2B:
“The method of claim 1, further comprising: obtaining, from the learning learner, the operation determination model updated based on the provided information; obtaining a new value corresponding to the at least one first parameter from the RAN; obtaining information associated with a new operation which is a result obtained by applying the new value corresponding to the at least one first parameter to the updated operation determination model;” – The limitation recites obtaining and updating an operation determination model by a first learning learner and based on the provided information. The limitation involves updating a model (similar to updating an activity log) and obtaining that model based on gathered information. The limitation goes on to recite obtaining a new value that corresponds to a parameter on the network (RAN), and lastly obtains information associated with an operation to be manipulated and placed onto corresponding data (the parameter). All the limitations recited above is considered to be an insignificant, extra-solution activity and it cannot be integrated to a practical application (see MPEP 2106.05(g)). Furthermore, under step 2B, the act of updating a model and mere data gathering is considered to be directed to electronic recordkeeping, which is a well-understood, routine, and conventional activities that cannot provide significantly more than the judicial exception (see MPEP 2106.05(d)(II)).
“providing the information associated with the new operation to the RAN.” –The limitation recites providing information that is associated to a new operation of the network. Sending/receiving data over a network and merely outputting data is considered to be an insignificant, extra solution activity that cannot be integrated to a practical application (see MPEP 2106.05(g)). Furthermore, under step 2B, the act of transmitting data over a network is a well-understood, routine and conventional (WURC) activity, which cannot provide significantly more than the judicial exception (see MPEP 2106.05(d)(II)).
Thus, claim 3 is non-patent eligible.
Regarding claim 4,
Step 1: The claim is directed to method, which is considered to be a process. The claim satisfies step 1.
Step 2A Prong 1:
“identifying a parameter used by the updated operation determination model as at least one second parameter which is at least partially different from the at least one first parameter” – The limitation is directed to identifying a parameter that is to be a parameter that is different from the first parameter. Identifying a parameter is a process that can be performed in the human mind using evaluation, observation, and judgment, thus the limitation is considered to be a mental process.
Step 2A Prong 2 and Step 2B:
The majority of the limitations in this claim is analogous to claim 3. The following in (a) is considered analogous and thus will face the same reject as recited in claim 3 (insignificant, extra-solution activity under 2106.05(g) and WURC under 2106.05(d)(II):
“obtaining, from the learning learner, the operation determination model updated based on the provided information; obtaining a value corresponding to the at least one second parameter from the RAN; obtaining information associated with a new operation which is a result obtained by applying the value corresponding to the at least one second parameter to the updated operation determination model; and providing the information associated with the new operation to the RAN.”
Thus, claim 4 is non-patent eligible.
Regarding claim 5,
Step 1: The claim is directed to method, which is considered to be a process. The claim satisfies step 1.
Step 2A Prong 1:
“The method of claim 4, further comprising: identifying a new request for information for learning the operation determination model; in response to the new request, identifying a value corresponding to the at least one second parameter and information associated with an operation corresponding to the at least one second parameter;” – The limitation is directed to identifying requests for information and in response of the request, to identify a value that corresponds to a parameter and information. This limitation recites all processes that can be performed in the human mind using evaluation, observation, and judgment, as well as aid of pen and paper to perform the task, thus it is directed to a mental process.
Step 2A Prong 2 and Step 2B:
“providing, as new information, the value corresponding to the at least one second parameter, the information associated with the operation corresponding to the at least one second parameter, and a reward value identified based on at least a part of the value corresponding to the at least one second parameter.” – The limitation recites types of information (aka data) and the value that corresponds to a parameter. Providing information and values that correspond to parameters in the RAN is directed to selecting particular data to be manipulated and will also involve mere data gathering, which is an insignificant, extra-solution activity, and it cannot be integrated to a practical application (see MPEP 2106.05(g)). Furthermore, under step 2B, the act of transmitting data over a network is a well-understood, routine, and conventional activity (WURC), and it cannot provide significantly more than the judicial exception (see MPEP 2106.05(d)(II)).
Thus, claim 5 is non-patent eligible.
Regarding claim 6,
Step 1: The claim is directed to method, which is considered to be a process. The claim satisfies step 1.
Step 2A Prong 1:
“The method of claim 1, wherein, in response to the request, the identifying of the value corresponding to the at least one first parameter corresponding to the operation determination model among the at least one value and the information associated with the operation corresponding to the at least one first parameter comprises:” –The limitation is analogous to claim 1’s limitation: “in response to the request, identifying, among the at least one value, a value corresponding to at least one first parameter corresponding to the operation determination model and information associated with an operation corresponding to the at least one first parameter;” which was directed to a mental process, and thus the limitation of claim 6 is also directed to a mental process.
“identifying whether at least one value corresponding to each of the plurality of parameters supports the at least one first parameter, and based on the at least one first parameter being supported, identifying the value corresponding to the at least one first parameter corresponding to the operation determination model and the information associated with the operation corresponding to the at least one first parameter.” – The limitation is directed to identifying whether a value corresponds to a parameter and based on the supported parameter, identifying a value that corresponds to an operation of the model and information associated with an operation that corresponds to a parameter. Identifying a value and determining if that value should correspond with other data or values is directed to a process that can be performed in the human mind, and thus is directed to a mental process.
There are no elements to be evaluated under Step 2A Prong 2 and Step 2B.
Thus, claim 6 is non-patent eligible. Claim 15 is analogous to claim 6, aside from the added limitation below. Majority of claim 6’s mapping applies to claim 15. For claim 15, Step 2A Prong 2 and Step 2B, the limitation added “instructions, when executed by the at least one processor, cause the electronic device to:” is directed to mere instructions to apply onto computer, and thus does not integrate to a practical application, nor does it provide significantly more than the judicial exception (see MPEP 2106.05(f)).
Thus, claim 6 is non-patent eligible.
Regarding claim 7:
Step 1: The claim is directed to method, which is considered to be a process. The claim satisfies step 1.
Step 2A Prong 1:
(a) “producing the information for learning the operation determination model comprises: producing the information for learning the operation determination model based on the value corresponding to the at least one first parameter, the information associated with the operation corresponding to the at least one first parameter, and an entirety of a reward value identified based on the entirety of the value corresponding to the at least one first parameter” -- This limitation is directed to producing information for learning based on parameter values, operation information, and a reward value determined from the entirety of a parameter value. Producing information based on selected data and a reward value involves evaluating data and determining a value based on data, which is a mathematical concept and/or a mental process capable of being performed using observation, evaluation, and judgment.
Step 2A Prong 2 and Step 2B:
(a) “The method of claim 1, wherein the producing of the information for learning the operation determination model comprises:” -- This limitation further limits the data used to produce the learning information but does not recite a specific model architecture, training algorithm, memory-management mechanism, RAN-control protocol, or technical improvement to the electronic device. The limitation therefore does not integrate the abstract idea into a practical application nor does not provide significantly more than the judicial exception (see MPEP 2106.05(f)).
(b) “based on the value corresponding to the at least one first parameter, the information associated with the operation corresponding to the at least one first parameter, and an entirety of a reward value identified based on an entirety of the value corresponding to the at least one first parameter” -- This limitation merely specifies the data inputs used to produce learning information, including use of the entire parameter value and entire reward value. Selecting, using, and processing particular data are data-manipulation steps and, without more, amount to insignificant extra-solution activity and/or generic computer implementation, and thus does not integrate to a practical application (see MPEP 2106.05(g)). Furthermore, under Step 2B, reciting data relationships based on gathered data is a well-understood, routine, and conventional activity (WURC) that cannot be integrated to a practical application, nor provide significantly more than the judicial exception (see MPEP 2106.05(d)(II)).
Thus, claim 7 is non-patent eligible.
Regarding claim 8,
Step 1: The claim is directed to method, which is considered to be a process. The claim satisfies step 1.
There are no elements to be evaluated under Step 2A Prong 1.
Step 2A Prong 2 and Step 2B:
(a) “The method of claim 1, wherein the producing of the information for learning the operation determination model comprises…” -- This limitation narrows the production of learning information to use a selected part of a parameter value, but it does not recite a specific technical implementation for the selection, a specific data structure, a particular machine-learning training algorithm, or a technical improvement to RAN operation. The limitation merely applies the abstract idea using generic data processing and does not integrate into a practical application nor provide significantly more than the judicial exception (see MPEP 2106.05(f)).
(b) “selecting a part of the value corresponding to the at least one first parameter” -- Selecting a subset or portion of data is a data-selection operation. Data selection and preparation are insignificant extra-solution activities when recited at a high level of generality and used only as input to further abstract data processing, and thus the limitation does not integrate to a practical application (see MPEP 2106.05(g)). Furthermore, under Step 2B, the act of selecting data (like part of a value) corresponding to a parameter is a well-understood, routine, and conventional activity (WURC), and does not provide significantly more than the judicial exception (see MPEP 2106.05(d)(II)).
(c) “producing the information for learning the operation determination model based on the selected part, the information associated with the operation corresponding to the at least one first parameter, and a reward value identified based on the selected part” -- This limitation recites using selected data and a reward value to produce learning information. Producing information from selected input data and a reward value is generic data manipulation and mathematical evaluation. The claim does not recite a specific improvement to machine-learning training, such as a particular objective function, loss function, model update rule, memory-saving data structure, or RAN-control protocol, and is considered an insignificant, extra-solution activity that cannot be integrated to a practical application (see MPEP 2106.05(g)). Furthermore, under Step 2B, the limitation does not provide significantly more than the judicial exception (see MPEP 2106.05(d)(II)).
Thus, claim 8 is non-patent eligible. Thus, claim 8 is non-patent eligible. Claim 16 is analogous to claim 8, aside from the added limitation below.
Majority of claim 8’s mapping applies to claim 16.
For claim 16, Step 2A Prong 2 and Step 2B, the limitation added “instructions, when executed by the at least one processor, cause the electronic device to:” is directed to mere instructions to apply onto computer, and thus does not integrate to a practical application, nor does it provide significantly more than the judicial exception (see MPEP 2106.05(f)).
Regarding claim 9,
Step 1: The claim is directed to method, which is considered to be a process. The claim satisfies step 1.
Step 2A Prong 1:
“The method of claim 8, wherein the selecting of the part of the value corresponding to the at least one first parameter comprises: selecting the part of the value corresponding to the at least one first parameter based on at least one operation among: selecting the part based on priority of each of the at least one first parameter, selecting the part based on a point in time at which each value corresponding to the at least one first parameter is obtained, or selecting the part in a random manner.” -- The limitation is directed to selecting parts of a value corresponding to a parameters that and other parts that comprises also selecting based on at least an operation of at least a first parameter, a point in time, or in a random manner. The limitation is directed to a process that can be performed in the human mind using evaluation, observation, and judgement, with aid of pen and paper, and thus the limitation is directed to a mental process.
There are no elements to be evaluated under Step 2A Prong 2 and Step 2B.
Thus, claim 9 is non-patent eligible.
Regarding claim 10,
Step 1: The claim is directed to method, which is considered to be a process. The claim satisfies step 1.
Step 2A Prong 1:
“The method of claim 1, further comprising: identifying the reward value based on a reward determination scheme and at least a part of the value corresponding to the at least one first parameter,” – The limitation is directed to identifying a reward value based on a scheme and a part of a value that will correspond to a parameter. Identifying values based on a scheme and a value based on data can be performed in the human mind with aid of pen and paper, thus the limitation is directed to a mental process.
Step 2A Prong 2 and Step 2B:
“wherein the reward determination scheme is stored in advance in the electronic device or is received by the electronic device.” – The limitation recites that the determination scheme will either be stored in a electronic device (computer/network) or received by the device. This limitation is directed to mere data gathering, which is an insignificant, extra-solution activity that cannot be integrated to a practical application (see MPEP 2106.05(g)). Furthermore, under step 2B, the act of storing and/or receiving data over a network are both considered well-understood, routine and conventional activities (WURC), and it cannot provide significantly more than the judicial exception (see MPEP 2106.05(d)(II)).
Thus, claim 10 is non-patent eligible.
Regarding claim 11,
Step 1: The claim is directed to method, which is considered to be a process. The claim satisfies step 1.
Step 2A Prong 1:
“The method of claim 1, wherein, in response to the request, the identifying of the value corresponding to the at least one first parameter corresponding to the first operation determination model among the at least one value, and the information associated with the operation corresponding to the at least one first parameter comprises; identifying the at least one first parameter declared by the first operation determination model.” – This limitation is directed to identifying a value that corresponds to a parameter that will also correspond to an operation of the determination model, in response of a request, as well as identifying a parameter that was declared by an operation of the determination model. Identifying values based on data and initiating once getting the request and evaluating and observing it is directed to a mental process.
There are no elements to be evaluated under Step 2A Prong 2 and Step 2B.
Thus, claim 11 is non-patent eligible.
Regarding claim 12,
Step 1: The claim is directed to method, which is considered to be a process. The claim satisfies step 1.
Step 2A Prong 1:
“The method of claim 1, wherein, in response to the request, the identifying of the value corresponding to the at least one first parameter corresponding to the first operation determination model among the at least one value, and the information associated with the operation corresponding to the at least one first parameter comprises; identifying the at least one first parameter based on an external input.” – This limitation is directed to identifying a value that corresponds to a parameter that will also correspond to an operation of the determination model, in response of a request, as well as identifying a parameter based on an input. Identifying values based on data and initiating once getting the request and evaluating and observing it is directed to a mental process.
There are no elements to be evaluated under Step 2A Prong 2 and Step 2B.
Thus, claim 12 is non-patent eligible.
Regarding claim 13,
Step 1: The claim is directed to method, which is considered to be a process. The claim satisfies step 1.
There are no elements to be evaluated under Step 2A Prong 1.
Step 2A Prong 2 and Step 2B:
“The method of claim 1, wherein the experience information comprises: a value corresponding to the at least one first parameter at a first point in time, a first operation performed by the RAN at the first point in time, a value corresponding to the at least one first parameter at a second point in time after the first point in time according to a result of performing the first operation, and a reward value at the first point in time.” – The limitation recites a value, a first operation performed by the RAN and certain point in times (first and second) that a value corresponds to a parameter will exist once the first operation performs. All these limitations are merely further limiting the field of use/particular environment of the claim and it’s not integrating to a practical application nor providing significantly more than the judicial exception (see MPEP 2106.05(h)).
Thus, claim 13 is non-patent eligible.
Regarding claim 24,
Step 1: The claim is directed to method, which is considered to be a process. The claim satisfies step 1.
Step 2A Prong 1:
“The method of claim 1, further comprising: identifying, among the plurality of parameters, the at least one first parameter corresponding to the operation determination model based on identification information of the learning learner included in the request.” -- The limitation is directed to identifying, among parameters, the first parameter that corresponds to the operation determination model based on identifying information of the learning learner that was included in the request. The limitation is directed to a process that can be performed in the human mind using evaluation, observation, and judgement, with pen and paper, and thus the limitation is directed to a mental process.
There are no elements to be evaluated under Step 2A Prong 2 and Step 2B.
Thus, claim 24 is non-patent eligible.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this
Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not
identically disclosed as set forth in section 102, if the differences between the claimed invention and the
prior art are such that the claimed invention as a whole would have been obvious before the effective filing
date of the claimed invention to a person having ordinary skill in the art to which the claimed invention
pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1,14, 22, and 24 is/are rejected under 35 U.S.C. 103 as being unpatentable over NPL reference “O-RAN: Towards an Open and Smart RAN”, by Katti et. al. (referred herein as Katti) in view of US 8,255,524 B2, by Devitt et. al (in view of Devitt) further in view of US10380504B2, by Bendre et. al. (referred herein as Bendre).
Regarding claim 1, Katti teaches:
A method of operating an electronic device, the method comprising: storing at least one value corresponding to each of a plurality of parameters associated with a radio access network (RAN), and information associated with an operation performed by the RAN, ([Katti, page 11] “The RIC near-RT functions leverage a database called the Radio-Network Information Base (R-NIB) which captures the near real-time state of the underlying network via E2 and commands from RIC non-RT via A1.” and [Katti, page 13] “Network & UE-level information/context exposure from eNB/gNB to RIC non-RT to support various requirements such as network management, online learning and offline training of AI/ML models and driving non-RT optimization into the network.”, wherein the examiner interprets the Radio-Network Information Base capturing the near real-time state of the network and the exposure of network and UE-level information/context to be the same as storing at least one value corresponding to each of a plurality of parameters associated with a RAN and information associated with an operation performed by the RAN, because they are both directed to storing and maintaining RAN-related state, context, measurements, and operational information used for RAN management.)
wherein a value corresponding to one or more parameters among the plurality of parameters is used based on a plurality of operation determination models executable by the electronic device to determine at least a part of the information associated with the operation, wherein each of the plurality of operation determination models is configured to output a respective operation to the RAN for the RAN to operate based on the respective operation, and each of the plurality of operation determination models is distinct from a respective learning learner and corresponds to the respective learning learner among a plurality of learning learners; ([Katti, page 11] “Trained models and real-time control functions produced in the RIC non-RT are distributed to the RIC near-RT for runtime execution…Messages generated from AI-enabled policies and ML based training models in RIC non-RT are conveyed to RIC near-RT. The core algorithm of RIC non-RT is developed and owned by operators. It provides the capability to modify the RAN behaviors by deployment of different models optimized to individual operator policies and optimization objectives…While the E2 interface feeds data, including various RAN measurements, to the RIC near-RT to facilitate radio resource management, it is also the interface through which the RIC near-RT may initiate configuration commands directly to CU/DU.” and [Katti, page 12] “With the amount of L1/L2/L3 data collected from eNB/gNB (including CU/DU), useful data features and models can be learned to empower the intelligent management and control in RAN.”, wherein the examiner interprets the trained models and real-time control functions distributed to RIC near-RT for runtime execution, different models optimized to individual operator policies and optimization objectives, and configuration commands to CU/DU to be the same as a plurality of operation determination models executable by the electronic device and configured to output respective operations to the RAN, because they are both directed to multiple executable trained models that use RAN measurements and features to produce operational control outputs for RAN behavior. The examiner further interprets the RIC non-RT producing and training models and the RIC near-RT executing the trained models at runtime to be the same as operation determination models being distinct from respective learning learners, because they are both directed to a separated learning/training component and runtime execution/model-control component.)
receiving, from a learning learner among the plurality of learning learners for learning the operation determination model, a request for information for learning the operation determination model selected from among the plurality of operation determination models; ([Katti, page 11] “Messages generated from AI-enabled policies and ML based training models in RIC non-RT are conveyed to RIC near-RT” and [Katti, page 13] “The A1 interface supports communication & information exchange between Orchestration/NMS layer containing RIC non-RT and eNB/gNB containing RIC near-RT. Key functions that the A1 interface is expected to provide include: Network & UE-level information/context exposure from eNB/gNB to RIC non-RT to support various requirements such as network management, online learning and offline training of AI/ML models and driving non-RT optimization into the network. Support for policy-based guidance of RIC near-RT functions/use-cases, deploying/updating AI/ML models into RIC near-RT, and feedback mechanisms from RIC near-RT to ensure SLAs.”, wherein the examiner interprets the A1 interface supporting communication and information exchange between the RIC non-RT and RIC near-RT, including network and UE-level information exposure, feedback mechanisms, and deploying/updating AI/ML models, to be the same as receiving, from a learning learner, a request for information for learning an operation determination model selected from among a plurality of operation determination models, because they are both directed to message-based exchange between a learning/training entity and an executing RAN-control entity for obtaining information used in online or offline training and updating of AI/ML models.)
providing the information for learning the operation determination model to the learning learner; ([Katti, page 11] “RIC non-RT can distribute well-trained user mobility and traffic prediction models to the RIC near-RT so that near-real-time predictions and decisions related to user mobility and traffic load are efficiently executed…In a similar fashion, E2 interface can be leveraged to fetch data feeds from the radio nodes and provide those to the RIC non-RT to train AI models.”, wherein the examiner interprets providing data feeds from radio nodes to the RIC non-RT to train AI models to be the same as providing the information for learning the operation determination model to the learning learner, because they are both directed to transmitting RAN-related training information from an executing/network-side component to a learning/training component for AI model training.)
Katti does not teach in response to the request, identifying, among the at least one value, a value corresponding to at least one first parameter corresponding to the operation determination model, and information associated with an operation corresponding to the at least one first parameter; …based on the value corresponding to the at least one first parameter, the information associated with the operation corresponding to the at least one first parameter, and a reward value identified based on at least a part of the value corresponding to the at least one first parameter;…external to the learning learner;
Devitt teaches:
in response to the request, identifying, among the at least one value, a value corresponding to at least one first parameter corresponding to the operation determination model, and information associated with an operation corresponding to the at least one first parameter; ([Devitt, page 6] “A sensitivity analysis can be used to determine which events...have the strongest influence on the KPI...to perform a root cause analysis of predicted or actual KPI violations”, wherein the examiner interprets determining which events have the strongest influence on the KPI to be the same as identifying a value corresponding to at least one first parameter corresponding to the operation determination model, because they are both directed to locating the particular network parameter or event values most relevant to a performance model or decision model. [Katti, page 11] “E2 interface feeds data, including various RAN measurements, to the RIC near-RT to facilitate radio resource management, it is also the interface through which the RIC near-RT may initiate configuration commands directly to CU/DU.”, wherein the examiner interprets RAN measurements and configuration commands directly to CU/DU to be the same as information associated with an operation corresponding to the at least one first parameter, because they are both directed to network measurement values and corresponding RAN control operations used for adaptive RAN operation.)
based on the value corresponding to the at least one first parameter, the information associated with the operation corresponding to the at least one first parameter, and a reward value identified based on at least a part of the value corresponding to the at least one first parameter; ([Devitt, col. 6, lines 7-9] “The utility node is particularly adapted to assign a value to each quality evaluation based on parameter (variable) value and decision combinations.”, wherein the examiner interprets assigning a value to each quality evaluation based on parameter value and decision combinations to be the same as producing information for learning the operation determination model based on the value corresponding to the at least one first parameter, the information associated with the operation corresponding to the at least one first parameter, and a reward value identified based on at least a part of the value corresponding to the at least one first parameter, because they are both directed to generating performance-evaluation information from parameter values and decision or operation combinations, where the assigned quality value serves as a reward-type evaluation value for learning.)
Katti and Devitt do not teach temporarily producing, external to the learning learner, the information for learning the operation determination model; and after the information for learning the operation determination model is provided to the learning learner, deleting the information for learning the operation determination model.
temporarily producing…the information for learning the operation determination model; ([Bendre, page 21, col 21, lines 62-67] “The temporary data storage 628 may be configured to temporarily store training data. In particular, once the trainer device 606 obtains training data from the customer instance 610, the trainer device 606 may store that training data in the temporary data storage 628 while the ML trainer 626 process is serving a corresponding ML training request.”, wherein the examiner interprets storing training data in temporary data storage while an ML trainer process serves a corresponding ML training request to be the same as temporarily producing…the information for learning the operation determination model, because they are both directed to temporarily generating or maintaining training information outside the model-training process while a learning request is being served.)
external to the learning learner; ([Bendre, page 21, col 22, lines 1-4] “the ML trainer 626 process could refer to the training data stored in the temporary data storage 628, so as to learn from that training data for the purpose of generating an ML model.”, wherein the examiner interprets the ML trainer process referring to training data stored in temporary data storage to be the same as producing the information for learning external to the learning learner, because they are both directed to training information being stored or produced outside the learning process and accessed by that learning process for model learning.)
and after the information for learning the operation determination model is provided to the learning learner, deleting the information for learning the operation determination model. ([Bendre, page 21, col 22, lines 4-13] “However, once the trainer device 606 (e.g., the training controller 624) determines that the ML trainer 626 process completed the serving of the corresponding ML training request, the trainer device 606 may delete the training data from the temporary data storage 628. As such, the trainer device 606 could store training data for each ML training request being served at the trainer device 606 and, once service of a given ML training request is complete, the trainer device 606 may delete the training data stored in association with that given ML training request.”, wherein the examiner interprets deleting the training data from temporary data storage once the corresponding ML training request is complete to be the same as after the information for learning the operation determination model is provided to the learning learner, deleting the information for learning the operation determination model, because they are both directed to removing temporary training information after the requested learning/training operation has been served.)
Katti, Devitt, Bendre, and the instant application are analogous art because they are all directed to machine-learning-based model training and model execution in networked computer systems in which operational parameter data is used to generate training or learning information for models that control or predict system behavior.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of operating a RIC-based RAN control framework disclosed by Katti to include the “assign a value to each quality evaluation based on parameter value and decision combinations” disclosed by Devitt. One would be motivated to do so to effectively generate KPI-oriented performance information for training and updating RAN operation models, as suggested by Devitt ([Devitt, col. 6, lines 7-9] “assign a value to each quality evaluation based on parameter value and decision combinations.”).
It would have also been obvious to a person of ordinary skill in the art before the effective filing date of the invention to further modify the method of operating a RIC-based RAN control framework disclosed by Katti and Devitt to include the “temporary data storage” and deletion of training data disclosed by Bendre. One would be motivated to do so to securely and efficiently manage training information during distributed machine-learning requests, as suggested by Bendre ([Bendre, page 21, col 22, lines 14-16] “In this manner, due to the temporary storage of training data, the disclosed ML arrangement helps secure an enterprise's data against unauthorized access.”). Claims 14 and 22 is analogous to claim 1, aside from claim type, and thus faces the same rejection as above.
Regarding claim 24, Katti, Devitt, and Bendre teach The method of claim 1, (see rejection of claim 1).
Bendre further teaches wherein the identifying of the value corresponding to the at least one first parameter comprises: identifying, among the plurality of parameters, the at least one first parameter corresponding to the operation determination model based on identification information of the learning learner included in the request. ([Bendre, page 11, col 2, lines 4-10] “as training data that should be used as basis for generating an ML model . Additionally , the provided information could indicate a target variable to be predicted using the ML model . For example , the client device could request the network system to predict categories for any uncategorized information within certain fields of a data table”, [Bendre, page 27, col 33, lines 15-19] “In these embodiments, the scheduler device may assign the ML training request to the particular ML trainer process based at least on the determination that the particular ML trainer process is available to serve the ML training request.”, and [Bendre, page 24, col 27, lines 20-22] “Yet further, once the trainer device 606 generates the ML model 640, the trainer device 606 may send the generated ML model 640 to the customer instance 610.”, wherein the examiner interprets the ML training request indicating the training data to be used as the basis for the requested model and indicating the target variable to be predicted by the requested model to be the same as identifying, among the plurality of parameters, the at least one first parameter corresponding to the operation determination model, because they are both directed to using request-included model/training information to select the particular data or parameter information used for learning the selected model. The examiner further interprets assigning the ML training request to a particular ML trainer process on a particular trainer device and requesting training data from the customer instance to be the same as identifying the at least one first parameter based on identification information of the learning learner included in the request, because they are both directed to using information associated with the requesting learning/training entity and requested learning task to determine the corresponding training data or parameter set.)
Katti, Devitt, Bendre, and the instant application are analogous art because they are all directed to machine-learning-based network operation in which learning requests are used to identify model-related training data, parameter information, and information for learning or updating operation determination models.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method claim 1 disclosed by Katti, Devitt, and Bendre to include the request-based training-data identification framework disclosed by Bendre. One would be motivated to do so to efficiently identify the proper model-related parameter data for a requesting learning entity in a distributed learning environment, as suggested by Bendre ([Bendre, page 27, col 33, lines 15-19] “In these embodiments, the scheduler device may assign the ML training request to the particular ML trainer process based at least on the determination that the particular ML trainer process is available to serve the ML training request.”, and [Bendre, page 24, col 27, lines 20-22] “Yet further, once the trainer device 606 generates the ML model 640, the trainer device 606 may send the generated ML model 640 to the customer instance 610.”)
Claim(s) 2-5, 8, 10-12, and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Katti in view of Devitt in view of Bendre further in view of US11201784B2, by Peng et. al. (referred herein as Peng)
Regarding claim 2, Katti, Devitt, and Bendre teaches The method of claim 1 (see rejection of claim 1).
Katti, Devitt, and Bendre do not teach: wherein the storing of the at least one value corresponding to each of the at least one parameter associated with the RAN, and the information associated with the operation performed by the RAN comprises: classifying and storing, for each of a plurality of points in time, the at least one value corresponding to each of the at least one parameter and the information associated with the operation performed by the RAN.
Peng teaches:
wherein the storing of the at least one value corresponding to each of the at least one parameter associated with the RAN, and the information associated with the operation performed by the RAN comprises: classifying and storing, for each of a plurality of points in time, the at least one value corresponding to each of the at least one parameter and the information associated with the operation performed by the RAN. [Peng, col 7, lines 22-35] “As shown in FIG. 1, FIG. 1 is a flow chart of an artificial intelligence-based networking method for F-RANs, which 110: may include the following steps: Step a central computing logic module receives reported data which may include: measurement report data from user terminals, wireless transmission data from base stations, and operation and maintenance data from a radio access network. The measurement report data relates to user behavior information, the wireless transmission data relates to the performance indicators of the radio access network, and operation and maintenance data relates to service attributes.”, [Peng, FIG. 1], “Based on the reported data obtained during a cycle T1 and a proper machine learning algorithm, the central computing logic module configures an operating mode of the radio access network that matches the user behavior information, the service attributes, and the radio access network performance indicators ... The edge computing logic module receives the operating mode information from the central computing logic module. According to the operating mode, the edge computing logic module, during the cycle T2, determines whether the current configuration of the edge communication entity meets the networking aim.”, AND [Peng, col 13, 30-35], “The step 141 may further include, but is not limited to: using a deep reinforcement learning algorithm to learn the data related to the user terminals, and obtaining a strategy to perform configuration optimization of the edge communication entity.” wherein the examiner interprets “receives reported data which may include: measurement report data from user terminals, wireless transmission data from base stations, and operation and maintenance data from a radio access network” to be the same as “storing of the at least one value corresponding to each of the at least one parameter associated with the RAN, and the information associated with the operation performed by the RAN”, as both describe gathering and storing various types of network-related data. The examiner further interprets “based on the reported data obtained during a cycle T1” to be the same as “for each of a plurality of points in time”, as both describe organizing and analyzing data over distinct time intervals. Additionally, the examiner interprets “using a deep reinforcement learning algorithm to learn the data related to the user terminals, and obtaining a strategy to perform configuration optimization of the edge communication entity” to be the same as “the information associated with the operation performed by the RAN”, as both describe leveraging stored data to optimize the operation of network components.)
Katti, Devitt, Bendre, Peng, and the instant application are analogous art because they are all directed to machine-learning-based management of radio access network operations using collected network parameter data, operation information, and time-based network performance information.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method claim 1 disclosed by Katti, Devitt, and Bendre to include the “reported data obtained during a cycle T1” disclosed by Peng. One would be motivated to do so to efficiently classify and store RAN-related parameter and operation information over time so that the system can configure and optimize RAN operation based on time-indexed network conditions, as suggested by Peng ([Peng, col. 7, lines 22-35] “receives reported data which may include: measurement report data from user terminals, wireless transmission data from base stations, and operation and maintenance data from a radio access network.”).
Regarding claim 3, Katti, Devitt, and Bendre teaches The method of claim 1 (see rejection of claim 1).
Katti, Devitt, and Bendre do not teach further comprising: obtaining, from the first learning learner, a first operation determination model updated based on the provided experience information; obtaining a new value corresponding to the at least one first parameter from the RAN; obtaining information associated with a new operation which is a result obtained by applying the new value corresponding to the at least one first parameter to the updated first operation determination model; and providing the information associated with the new operation to the RAN.
Peng teaches:
further comprising: obtaining, from the first learning learner, a first operation determination model updated based on the provided experience information; [Peng, col. 11, lines 18-22], “The central computing logic module receives updated reported data which includes the measurement report data from the user terminals, the wireless transmission data from the base stations, and the operation and maintenance data from the radio access network.” wherein the examiner interprets “the central computing logic module receives updated reported data” to be the same as “obtaining, from the first learning learner, a first operation determination model updated based on the provided experience information”, as both describe receiving updated input that informs a subsequent decision-making or operational process.)
obtaining a new value corresponding to the at least one first parameter from the RAN; [Peng, col 18-19, 65-67, 1-5] “based on the reported data obtained during the cycle T1 and the proper machine learning algorithm, the central computing logic module configures more operating modes of radio access network that match the user behavior, the service attributes, and the performance indicators of the radio access network.” wherein the examiner interprets “reported data obtained during the cycle T1 and proper machine learning algorithm” to be the same as “obtaining a new value corresponding to the at least one first parameter from the RAN”, as both describe acquiring updated network data, including performance indicators, for further processing.)
obtaining information associated with a new operation which is a result obtained by applying the new value corresponding to the at least one first parameter to the updated first operation determination model; [Peng, col. 11, lines 27-33] “The edge computing logic module receives the operating mode information from the central computing logic module. According to the operating mode, the edge computing logic module, during a cycle T2, determines whether the current configuration of the edge communication entity meets the networking aim.” wherein the examiner interprets “the edge computing logic module receives the operating mode information from the central computing logic module” to be the same as “obtaining information associated with a new operation which is a result obtained by applying the new value corresponding to the at least one first parameter to the updated first operation determination model”, as both describe deriving operational configurations based on newly processed data.)
and providing the information associated with the new operation to the RAN. [Peng, col. 11, lines 65-67, col. 12, lines 1-3], “Step 140: If the current configuration meets the aim, the edge computing logic module allocates resources to the user terminals connected to the edge communication entity. The edge resources communication entities and user terminals that are allocated with proper resources are networked as an F-RAN. The resources may include radio resources, computing resources, and caching resources.” wherein the examiner interprets “allocates resources to the user terminals connected to the edge communication entity” to be the same as “providing the information associated with the new operation to the RAN”, as both describe transmitting the determined operational configuration to the network for implementation.)
Katti, Devitt, Bendre, Peng, and the instant application are analogous art because they are all directed to machine-learning-based RAN operation in which updated network parameter data and operation information are used to update models and determine new RAN operations or configurations.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method claim 1 disclosed by Katti, Devitt, and Bendre to include the updated reported data, operating mode information, and resource allocation disclosed by Peng. One would be motivated to do so to effectively update the RAN learning/control framework using current network conditions and implement newly determined RAN operations in the network, as suggested by Peng ([Peng, col. 11, lines 18-22] “receives updated reported data which includes the measurement report data from the user terminals, the wireless transmission data from the base stations, and the operation and maintenance data from the radio access network.”).
Regarding claim 4, Katti, Devitt, and Bendre teaches The method of claim 1 (see rejection of claim 1).
Katti, Devitt, and Bendre do not teach further comprising: obtaining, from the first learning learner, the first operation determination model updated based on the provided experience information; identifying a parameter used by the updated first operation determination model as at least one second parameter which is at least partially different from the at least one first parameter; obtaining a value corresponding to the at least one second parameter from the RAN; obtaining information associated with a new operation which is a result obtained by applying the value corresponding to the at least one second parameter to the updated first operation determination model;.
Peng teaches:
further comprising: obtaining, from the first learning learner, the first operation determination model updated based on the provided experience information; [Peng, col 13, lines 32-35] “using a deep reinforcement learning algorithm to learn the data related to the user terminals, and obtaining a strategy to perform configuration optimization of the edge communication entity.”, wherein the examiner interprets “using a deep reinforcement learning algorithm to learn the data related to the user terminals” to be the same as “obtaining, from the first learning learner, the first operation determination model updated based on the provided experience information”, as both describe acquiring an updated learning model that incorporates previously gathered experience data to refine operational decisions.)
identifying a parameter used by the updated first operation determination model as at least one second parameter which is at least partially different from the at least one first parameter; [Peng, col 3, lines 10-19] “during the cycle T2, monitoring, by the edge computing logic module, performance of the edge communication entity and checking whether a variation of a target performance indicator exceeds a preset threshold; if exceeds, determining that the current configuration of the edge communication entity does not meet the networking aim. Then there is a need for the edge computing logic module to optimize the current configuration of the edge communication entity.”, wherein the examiner interprets “monitoring, by the edge computing logic module, performance of the edge communication entity and checking whether a variation of a target performance indicator exceeds a preset threshold” to be the same as “identifying a parameter used by the updated first operation determination model as at least one second parameter which is at least partially different from the at least one first parameter”, as both describe evaluating a parameter that has changed and requires adjustment for the updated model.)
obtaining a value corresponding to the at least one second parameter from the RAN; [Peng, col 8, lines 48-54] “Step 1. the central computing logic module monitors the measurement report data from all the user terminals in the radio access network, and checks whether the obtained quality of service and the number of active user terminals exceed respective preset thresholds.”, wherein the examiner interprets “the central computing logic module monitors the measurement report data from all the user terminals in the radio access network, and checks whether the obtained quality of service and the number of active user terminals exceed respective preset thresholds” to be the same as “obtaining a value corresponding to the at least one second parameter from the RAN”, as both describe retrieving updated parameter values from network conditions for use in further processing.)
obtaining information associated with a new operation which is a result obtained by applying the value corresponding to the at least one second parameter to the updated first operation determination model; [Peng, col 13, lines 32-35] “using a deep reinforcement learning algorithm to learn the data related to the user terminals, and obtaining a strategy to perform configuration optimization of the edge communication entity.”, wherein the examiner interprets “using a deep reinforcement learning algorithm to learn the data related to the user terminals, and obtaining a strategy to perform configuration optimization of the edge communication entity” to be the same as “obtaining information associated with a new operation which is a result obtained by applying the value corresponding to the at least one second parameter to the updated first operation determination model”, as both describe processing updated data within a learning model to determine a new strategy or action for system optimization.)
and providing the information associated with the new operation to the RAN. [Peng, col 3, lines 10-19] “Then there is a need for the edge computing logic module to optimize the current configuration of the edge communication entity.”, wherein the examiner interprets “optimizing the current configuration of the edge communication entity” to be the same as “providing the information associated with the new operation to the RAN”, as both describe updating the network configuration based on newly obtained information.)
Katti, Devitt, Bendre, Peng, and the instant application are analogous art because they are all directed to machine-learning-based RAN operation in which updated network data is used to update or apply an operation determination model and determine new RAN operations or configurations.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method claim 1 disclosed by Katti, Devitt, and Bendre to include the updated model-based configuration optimization disclosed by Peng. One would be motivated to do so to effectively adapt RAN configuration using updated parameter values and learned strategies based on current network conditions, as suggested by Peng ([Peng, col. 13, lines 32-35] “using a deep reinforcement learning algorithm to learn the data related to the user terminals, and obtaining a strategy to perform configuration optimization of the edge communication entity.”).
Regarding claim 5, Katti, Devitt, Bendre, and Peng teaches The method of claim 4, (see rejection of claim 4).
Peng further teaches:
further comprising: identifying a new request for experience information for learning the first operation determination model; [Peng, col 17, lines 15-21], “In the step 2, the edge computing logic module enters the trigger state to optimize the resource allocation of the edge communication entity. Here, taking the deep reinforcement learning as an example, referring to FIG. 7, the edge computing logic module performs actions according to the rewards brought by different actions in the current state”, wherein the examiner interprets “the edge computing logic module enters the trigger state to optimize the resource allocation of the edge communication entity … taking deep reinforcement learning as an example” to be the same as “identifying a new request for experience information for learning the first operation determination model”, as both describe a process where the system initiates an update based on a machine learning technique for optimization.
in response to the new request, identifying a value corresponding to the at least one second parameter and information associated with an operation corresponding to the at least one second parameter; [Peng, col 17, lines 19-21], “the edge computing logic module performs actions according to the rewards brought by different actions in the current state” wherein the examiner interprets “the edge computing logic module performs actions according to the rewards brought by different actions in the current state” to be the same as “identifying a value corresponding to the at least one second parameter and information associated with an operation corresponding to the at least one second parameter”, as both describe evaluating parameters and corresponding operational actions in response to system conditions.
and providing, as new experience information, at least some among the value corresponding to the at least one second parameter, the information associated with the operation corresponding to the at least one second parameter, and a reward value identified based on at least a part of the value corresponding to the at least one second parameter. [Peng, col 17, lines 21-33] “The choice according to the resource allocation strategy obtained by deep reinforcement learning maximizes the benefits in a continuous time. In the step 3, after the resource adjustment is completed, the edge computing logic module monitors the performance and checks whether the networking aim is met. If it is not met, the edge computing logic module directly jumps to the cycle T2 and triggers the configuration optimization of edge communication entity; If it is met, then the edge computing logic module continues monitoring until the time reaches an integral multiple of the cycle T3, and performs the next round of resource allocation.”, wherein the examiner interprets “the choice according to the resource allocation strategy obtained by deep reinforcement learning maximizes the benefits in a continuous time” to be the same as “providing, as new experience information, at least some among the value corresponding to the at least one second parameter, the information associated with the operation corresponding to the at least one second parameter, and a reward value identified based on at least a part of the value corresponding to the at least one second parameter”, as both describe leveraging learned strategies to refine system performance based on received experience data, also using a learning approach.)
Katti, Devitt, Bendre, Peng, and the instant application are analogous art because they are all directed to machine-learning-based RAN operation in which updated network parameter values, operation information, and reward information are used to generate new experience information for learning and optimizing RAN operations. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method claim 4 disclosed by Katti, Devitt, Bendre, and Peng to include the new trigger-based resource optimization and reward-based experience information disclosed by Peng. One would be motivated to do so to effectively initiate further learning and generate updated experience information when network conditions require additional RAN optimization, as suggested by Peng ([Peng, col. 17, lines 21-33] “the choice according to the resource allocation strategy obtained by deep reinforcement learning maximizes the benefits in a continuous time.”).
Regarding claim 8, Katti, Devitt, and Bendre teach The method of claim 1, see rejection of claim 1.
Devitt further teaches selecting a part among the value corresponding to the at least one first parameter, ([Devitt, page 6] “A sensitivity analysis can be used to determine which events... have the strongest influence on the KPI... to perform a root cause analysis of predicted or actual KPI violations”, wherein the examiner interprets determining which events have the strongest influence on the KPI to be the same as selecting a part among the value corresponding to the at least one first parameter, because they are both directed to selecting or identifying the portion of network parameter information most relevant to a performance model or decision model.)
Katti, Devitt, and Bendre do not teach wherein the producing of the information for learning the operation determination model comprises…and producing the information for learning the operation determination model based on the selected part, the information associated with the operation corresponding to the at least one first parameter, and a reward value identified based on the selected part.
Peng teaches producing the information for learning the operation determination model based on the selected part, the information associated with the operation corresponding to the at least one first parameter, and a reward value identified based on the selected part. ([Peng, col. 15, lines 45-54] “The edge computing logic module selects an action according to certain rewards so that the resource allocation based on Deep Q Network (DQN) can maximize the benefit in a continuous period. The state is jointly defined by interference distribution, link status, buffer status, available computing resources, etc. The reward function is selected from one or more of the following: rate, energy efficiency, delay, etc.”, wherein the examiner interprets the state jointly defined by interference distribution, link status, buffer status, and available computing resources to be the same as the selected part of the value corresponding to the at least one first parameter, because they are both directed to using selected network-state parameter information for learning-based RAN operation. The examiner further interprets selecting an action according to certain rewards and the reward function selected from rate, energy efficiency, and delay to be the same as producing the information for learning based on the selected part, the information associated with the operation corresponding to the at least one first parameter, and a reward value identified based on the selected part, because they are both directed to using selected network-state information, a corresponding resource-allocation operation, and reward feedback to generate learning information for optimizing RAN control.)
Katti, Devitt, Bendre, Peng, and the instant application are analogous art because they are all directed to machine-learning-based RAN operation in which selected network-state parameter values, operation information, and reward information are used to generate learning information for optimizing RAN control decisions.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method claim 1 disclosed by Katti, Devitt, and Bendre to include the selected-state and reward-based DQN resource allocation disclosed by Peng. One would be motivated to do so to effectively generate reward-based learning information from selected RAN state and operation data for improved learning-based resource allocation, as suggested by Peng ([Peng, col. 15, lines 45-54] “The edge computing logic module selects an action according to certain rewards so that the resource allocation based on Deep Q Network (DQN) can maximize the benefit in a continuous period.”). Claim 16 is analogous to claim 8, and thus the rejection can apply to both.
Regarding claim 10, Katti, Devitt, and Bendre teaches The method of claim 1 (see rejection of claim 1).
Katti, Devitt, and Bendre do not teach further comprising: identifying the reward value based on a reward determination scheme and at least a part of the value corresponding to the at least one first parameter, wherein the reward determination scheme is stored in advance in the electronic device or is received by the electronic device.
Peng teaches further comprising: identifying the reward value based on a reward determination scheme and at least a part of the value corresponding to the at least one first parameter, wherein the reward determination scheme is stored in advance in the electronic device or is received by the electronic device. [Peng, col 14, lines 1-3, lines 17-19, lines 45-46, “The reward function of the DRL1 may refer to the number of user terminals whose outrage rate is greater than a preset threshold....The reward function of the DRL2 may refer to the average throughput of all access points....the reward function is a weighted sum of success and failure of data transmission to a node of the next hop.” wherein the examiner interprets “the reward function of the DRL1 may refer to the number of user terminals whose outrage rate is greater than a preset threshold,” and “The reward function of the DRL2 may refer to the average throughput of all access points,” and “the reward function is a weighted sum of success and failure of data transmission to a node of the next hop” to be the same as “identifying the reward value based on a reward determination scheme and at least a part of the value corresponding to the at least one first parameter”, as both describe determining a reward value using predefined criteria that evaluate network performance. The reward function(s) DRL1, DRL2, etc. are predefined or provided externally and hence the quote provided by Peng is also the same as “the reward determination scheme is stored in advance in the electronic device or is received by the electronic device”.)
Katti, Devitt, Bendre, Peng, and the instant application are analogous art because they are all directed to machine-learning-based RAN operation in which network parameter values and reward determination rules are used to generate reward information for learning and optimizing RAN control decisions.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method claim 1 disclosed by Katti, Devitt, and Bendre to include the reward functions disclosed by Peng. One would be motivated to do so to effectively determine reward values using predefined network-performance criteria for reinforcement-learning-based RAN optimization, as suggested by Peng ([Peng, col. 14, lines 1-3, 17-19, and 45-46] “the reward function is a weighted sum of success and failure of data transmission to a node of the next hop.”).
Regarding claim 11, Katti, Devitt, and Bendre teach The method of claim 1, (see rejection of claim 1.)
Devitt further teaches identifying the at least one first parameter declared by the first operation determination model. ([Devitt, col. 6, lines 43-55] “A Bayesian Network comprises as referred to above, a DAG structure with nodes representing statistical variables such as performance counters and the arcs represent the influential relationships between these nodes. In addition thereto there is an associated conditional probability distribution over said statistical variables, for example performance counters. The conditional probability distribution encodes the probability that the variables assume their different values given the values of other variables in the BN. According to different embodiments the probability distribution is assigned by an expert, learnt off-line from historical data or learnt on-line incrementally from a live feed of data. Most preferably the probabilities are learnt on-line on the network devices.”, wherein the examiner interprets nodes representing statistical variables such as performance counters in a Bayesian Network model to be the same as identifying the at least one first parameter declared by the first operation determination model, because they are both directed to parameters defined in or represented by the model and used by the model to determine network behavior or performance outcomes.)
Katti, Devitt, and Bendre do not teach wherein, in response to the request, the identifying of the value corresponding to the at least one first parameter corresponding to the first operation determination model among the at least one value, and the information associated with the operation corresponding to the at least one first parameter comprises:
Peng teaches wherein, in response to the request, the identifying of the value corresponding to the at least one first parameter corresponding to the first operation determination model among the at least one value, and the information associated with the operation corresponding to the at least one first parameter comprises: ([Peng, col. 1, lines 50-65] “The present disclosure provides an artificial intelligence-based networking method for fog radio access networks, including: receiving, by a central computing logic module, reported data which includes: measurement report data from user terminals, wireless transmission data from base stations, and operation and maintenance data from a radio access network. The measurement report data relates to user behavior, the wireless transmission data relates to performance indicators of the radio access network, and the operation and maintenance data relates to service attributes. Based on the reported data obtained during a cycle T1 and a proper machine learning algorithm, the central computing logic module configures an operating mode of the radio access network that matches the user behavior, the service attributes, and the performance indicators of the radio access network.”, wherein the examiner interprets receiving reported data including measurement report data, wireless transmission data, and operation and maintenance data from a radio access network to be the same as identifying the value corresponding to the at least one first parameter and information associated with the operation corresponding to the at least one first parameter, because they are both directed to retrieving and using RAN-related parameter values and operation information for model-based operational determination.)
Katti, Devitt, Bendre, Peng, and the instant application are analogous art because they are all directed to machine-learning-based RAN operation in which network parameter data, model-defined parameters, and operation information are used to generate learning information for optimizing RAN behavior.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method claim 1 disclosed by Katti, Devitt, and Bendre to include the model-defined parameter structure disclosed by Devitt and the reported network data processing disclosed by Peng. One would be motivated to do so to effectively identify parameters declared by or represented within the operation determination model and to efficiently retrieve and use RAN-related parameter values and operation information for model-based operational determination, as suggested by Devitt ([Devitt, col. 6, lines 43-55] “nodes representing statistical variables such as performance counters”), and Peng ([Peng, col. 1, lines 50-65] “receiving, by a central computing logic module, reported data which includes: measurement report data from user terminals, wireless transmission data from base stations, and operation and maintenance data from a radio access network.”).
Regarding claim 12, Katti, Devitt, and Bendre teaches The method of claim 1 (see rejection of claim 1).Devitt further teaches:
identifying the at least one first parameter based on an external input. ([Devitt, col 5, lines 40-55], “The performance parameters particularly comprise performance counters. The additional or correlation parameters ‘may comprise one or more of alarm, configuration action, KPI definition or external performance counters. In a preferred implementation a conditional probability distribution over the performance parameters (or variables) is provided which is adapted to encode the probability that the performance parameters (variables) assume different values when specific values are given for other performance variables or parameters. In one particular embodiment the arrangement is. adapted to receive the probability distribution on-line although there are also other ways to provide it, for example it may be provided by an expert, or learnt off-line e.g. from historical data.” wherein the examiner interprets “The additional or correlation parameters may comprise one or more of alarm, configuration action, KPI definition or external performance counters” to be the same as “identifying the at least one first parameter based on an external input”, as both describe determining at least one first parameter using externally provided information. )
Katti, Devitt, and Bendre do not teach wherein, in response to the request, the identifying of the value corresponding to the at least one first parameter corresponding to the first operation determination model among the at least one value, and the information associated with the operation corresponding to the at least one first parameter comprises:.
Peng further teaches wherein, in response to the request, the identifying of the value corresponding to the at least one first parameter corresponding to the first operation determination model among the at least one value, and the information associated with the operation corresponding to the at least one first parameter comprises: [Peng, col 1, lines 50- 65], “The present disclosure provides an artificial intelligence-based networking method for fog radio access networks, including: receiving, by a central computing logic module, reported data which includes: measurement report data from user terminals, wireless transmission data from base stations, and operation and maintenance data from a radio access network. The measurement report data relates to user behavior, the wireless transmission data relates to performance indicators of the radio access network, and the operation and maintenance data relates to service attributes. Based on the reported data obtained during a cycle T1 and a proper machine learning algorithm, the central computing logic module configures an operating mode of the radio access network that matches the user behavior, the service attributes, and the performance indicators of the radio access network.” wherein the examiner interprets “receiving, by a central computing logic module, reported data which includes: measurement report data from user terminals, wireless transmission data from base stations, and operation and maintenance data from a radio access network” to be the same as “in response to the request, the identifying of the value corresponding to the at least one first parameter corresponding to the first operation determination model among the at least one value, and the information associated with the operation corresponding to the at least one first parameter comprises;”, as both describe the process of retrieving and utilizing stored network-related data to support an operational determination.)
Katti, Devitt, Bendre, Peng, and the instant application are analogous art because they are all directed to machine-learning-based RAN operation in which externally provided network data, performance parameters, and operation information are used to identify relevant parameters and generate learning information for optimizing RAN behavior.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method claim 1 disclosed by Katti, Devitt, and Bendre to include the external parameter inputs disclosed by Devitt and the reported network data processing disclosed by Peng. One would be motivated to do so to effectively identify relevant network parameters using externally provided performance or correlation information and to efficiently retrieve and use RAN-related parameter values and operation information in response to a learning or optimization need, as suggested by Devitt ([Devitt, col. 5, lines 40-55] “The additional or correlation parameters may comprise one or more of alarm, configuration action, KPI definition or external performance counters.”), and Peng ([Peng, col. 1, lines 50-65] “receiving, by a central computing logic module, reported data which includes: measurement report data from user terminals, wireless transmission data from base stations, and operation and maintenance data from a radio access network.”).
Claims 6-7, 9, and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Katti in view of Devitt in view of Bendre in view of Peng further in view of Shi et. al, “Reinforcement Learning for Dynamic Resource Optimization in 5G Radio Access Network Slicing” (referred herein as Shi).
Regarding claim 6, Katti, Devitt, and Bendre teaches The method of claim 1, (see rejection of claim 1).
Katti, Devitt, and Bendre do not teach wherein, in response to the request, the identifying of the value corresponding to the at least one first parameter corresponding to the first operation determination model among the at least one value and the information associated with the operation corresponding to the at least one first parameter comprises: identifying whether at least one value corresponding to each of the at least one parameter supports the at least one first parameter, and based on the at least one first parameter being supported, identifying the value corresponding to the at least one first parameter corresponding to the first operation determination model and the information associated with the operation corresponding to the at least one first parameter.
Peng teaches wherein, in response to the request, the identifying of the value corresponding to the at least one first parameter corresponding to the first operation determination model among the at least one value and the information associated with the operation corresponding to the at least one first parameter comprises: [Peng, col 1, lines 50- 65], “The present disclosure provides an artificial intelligence-based networking method for fog radio access networks, including: receiving, by a central computing logic module, reported data which includes: measurement report data from user terminals, wireless transmission data from base stations, and operation and maintenance data from a radio access network. The measurement report data relates to user behavior, the wireless transmission data relates to performance indicators of the radio access network, and the operation and maintenance data relates to service attributes. Based on the reported data obtained during a cycle T1 and a proper machine learning algorithm, the central computing logic module configures an operating mode of the radio access network that matches the user behavior, the service attributes, and the performance indicators of the radio access network.” wherein the examiner interprets “receiving, by a central computing logic module, reported data which includes: measurement report data from user terminals, wireless transmission data from base stations, and operation and maintenance data from a radio access network” to be the same as “identifying of the value corresponding to the at least one first parameter corresponding to the first operation determination model among the at least one value and the information associated with the operation corresponding to the at least one first parameter comprises:” as both describe the process of retrieving and utilizing stored network-related data to support an operational determination.)
Katti, Devitt, Bendre, Peng does not teach identifying whether at least one value corresponding to each of the at least one parameter supports the at least one first parameter, and based on the at least one first parameter being supported, identifying the value corresponding to the at least one first parameter corresponding to the first operation determination model and the information associated with the operation corresponding to the at least one first parameter.
Shi teaches: identifying whether at least one value corresponding to each of the at least one parameter supports the at least one first parameter, [Shi, Sec 1.B], “In our Q-learning solution, the states correspond to the available resources that transition over time depending on how they are occupied (for granted network slicing requests) or released (for completed requests).”, wherein the examiner interprets “the states correspond to the available resources that transition over time depending on how they are occupied (for granted network slicing requests) or released (for completed requests)” to be the same as “identifying whether at least one value corresponding to each of the at least one parameter supports the at least one first parameter”, as both describe evaluating whether existing resource states (values corresponding to parameters) align with the requirements of network slicing requests (first parameter), determining whether they can be used in a given operational scenario.)
and based on the at least one first parameter being supported, identifying the value corresponding to the at least one first parameter corresponding to the first operation determination model and the information associated with the operation corresponding to the at least one first parameter. [Shi, Sec 1.B] “We show that Q-learning successfully allocates resources over a time horizon and provides major gains in network utility compared to myopic, random and first come first served (FCFS) resource allocation algorithms. As the number of UEs increases or priorities of network slicing requests change over time, we show that Q-learning successfully adapts to dynamic user demands.”, wherein the examiner interprets “Q-learning successfully allocates resources over a time horizon and provides major gains in network utility” to be the same as “identifying the value corresponding to the at least one first parameter corresponding to the first operation determination model and the information associated with the operation corresponding to the at least one first parameter”, as both describe determining and applying resource allocation decisions based on dynamic conditions and an optimization model.
Katti, Devitt, Bendre, Peng, Shi, and the instant application are analogous art because they are all directed to dynamically optimizing resource allocation in networked systems using machine learning techniques.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method of claim 1 disclosed by Katti, Devitt, Bendre, and Peng to include the approach in which “Q-learning successfully allocates resources over a time horizon and provides major gains in network utility” as disclosed by Shi. One would be motivated to do so to efficiently enhance resource allocation and adaptability in response to changing network conditions, as suggested by Shi (Shi, [Sec 1.B] “we show that Q-learning successfully adapts to dynamic user demands.”). Regarding claim 15, Majority of claim 15 is analogous to claim 6, aside from below:
Katti, Devitt, and Bendre teaches The electronic device of claim 14, (see rejection of claim 1 which is analogous to claims 14 and 22).Devitt further teaches:
wherein, in response to the request, as at least a part of the identifying of the value corresponding to the at least one first parameter corresponding to the first operation determination model among the at least one value and the information associated with the operation corresponding to the at least one first parameter, ([Devitt, col 8, lines 34-42] “For monitoring the KPI for a network device, the Decision Graph model subscribes to performance parameters and other events of interest, i.e. all those encoded in the Decision Graph, on the network device (or entire network). This means that any change to a performance parameter automatically is updated in the Decision Graph model. The basic functionality of a BN ensures that each individual change propagates through the Decision Graph changing the probabilities of related variables in the graph.”, wherein the examiner interprets “the Decision Graph model subscribes to performance parameters and other events of interest” to be the same as “in response to the request”, as both describe a system reacting to new input data. The examiner further interprets “any change to a performance parameter automatically is updated in the Decision Graph model” to be the same as “identifying of the value corresponding to the at least one first parameter corresponding to the first operation determination model among the at least one value”, as both describe detecting changes in network parameters and integrating them into a decision-making framework. Finally, the examiner interprets “this incremental learning process means that over time the Decision Graph will be able to make predictions about future behavior on the basis of past experience” to be the same as “the information associated with the operation corresponding to the at least one first parameter”, as both describe how historical performance data informs predictive decision-making for network optimization.)
Regarding claim 7, Katti, Devitt, and Bendre teach The method of claim 1, (see rejection of claim 1).
Katti, Devitt, and Bendre do not teach wherein the producing of the information for learning the operation determination model comprises: producing the information for learning the operation determination model based on the value corresponding to the at least one first parameter, the information associated with the operation corresponding to the at least one first parameter, and an entirety of a reward value identified based on an entirety of the value corresponding to the at least one first parameter.
Peng teaches an entirety of a reward value identified based on an entirety of the value corresponding to the at least one first parameter. ([Peng, Figure 1] “Based on the reported data obtained during a cycle T1 and a proper machine learning algorithm, the central computing logic module configures an operating mode of the radio access network that matches the user behavior information, the service attributes, and the radio access network performance indicators.” and [Peng, col. 15, lines 45-54] “The edge computing logic module selects an action according to certain rewards so that the resource allocation based on Deep Q Network (DQN) can maximize the benefit in a continuous period. The state is jointly defined by interference distribution, link status, buffer status, available computing resources, etc. The reward function is selected from one or more of the following: rate, energy efficiency, delay, etc.”, wherein the examiner interprets the reported data obtained during a cycle T1 and the state jointly defined by interference distribution, link status, buffer status, and available computing resources to be the same as an entirety of the value corresponding to the at least one first parameter, because they are both directed to using the complete set of relevant measured network-state values for model-based RAN operation. The examiner further interprets selecting an action according to certain rewards and the reward function to be the same as an entirety of a reward value, because they are both directed to using the full reward output for learning-based resource allocation.)
Katti, Devitt, Bendre, Peng, and the instant application are analogous art because they are all directed to machine-learning-based RAN operation in which network-state parameter values and reward information are used to generate learning information for optimizing RAN control decisions.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method claim 1 disclosed by Katti, Devitt, and Bendre to include the “state is jointly defined by interference distribution, link status, buffer status, available computing resources” and “reward function” disclosed by Peng. One would be motivated to do so to effectively generate complete reward-based learning information using the full relevant RAN state for improved learning-based resource allocation, as suggested by Peng ([Peng, col. 15, lines 45-54] “The edge computing logic module selects an action according to certain rewards so that the resource allocation based on Deep Q Network (DQN) can maximize the benefit in a continuous period.”).
Katti, Devitt, Bendre, and Peng do not teach wherein the producing of the information for learning the operation determination model comprises: producing the information for learning the operation determination model based on the value corresponding to the at least one first parameter, the information associated with the operation corresponding to the at least one first parameter, and an entirety of a reward value identified based on an entirety of the value corresponding to the at least one first parameter.
Shi teaches wherein the producing of the information for learning the operation determination model comprises: producing the information for learning the operation determination model based on the value corresponding to the at least one first parameter, the information associated with the operation corresponding to the at least one first parameter, and an entirety of a reward value identified based on an entirety of the value corresponding to the at least one first parameter. ([Shi, page 4] “The gNodeB applies Q-learning to compute the function Q : S × A → R to evaluate the quality of action A producing reward R at state S. Note that the gNodeB maintains Q as the Q-table. At each time t, the gNodeB selects an action at, observes a reward rt, and transitions from the current state st to a new state st+1 (this transition depends on current state st and action at), and updates Q.”, wherein the examiner interprets the current state st to be the same as the value corresponding to the at least one first parameter, because they are both directed to a current network state or parameter value used by a learning model. The examiner interprets the selected action at to be the same as the information associated with the operation corresponding to the at least one first parameter, because they are both directed to an operation selected for the network based on the current parameter state. The examiner further interprets observing a reward rt and updating Q based on the reward to be the same as using an entirety of a reward value identified based on an entirety of the value corresponding to the at least one first parameter, because they are both directed to using the complete observed reward value generated from the current state/action condition to update learning information.)
Katti, Devitt, Bendre, Peng, Shi, and the instant application are analogous art because they are all directed to machine-learning-based operation of radio access networks using network parameter values, operation information, and reward-based learning information to optimize network control. It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method claim 1 disclosed by Katti, Devitt, and Bendre to include the Q-learning update and reward-based state/action learning disclosed by Peng and Shi. One would be motivated to do so to effectively generate complete reward-based learning information from RAN state and operation data for improving model-based RAN resource decisions, as suggested by Shi ([Shi, page 4] “At each time t, the gNodeB selects an action at, observes a reward rt, and transitions from the current state st to a new state st+1…and updates Q.”).
Regarding claim 9, Katti, Devitt, Bendre, and Peng teach The method of claim 8, (see rejection of claim 8).
Katti, Devitt, Bendre, and Peng do not teach wherein the selecting of the part of the value corresponding to the at least one first parameter comprises: selecting the part of the value corresponding to the at least one first parameter, based on at least one operation among: selecting the part based on priority of each of the at least one first parameter, selecting the part based on a point in time at which each value corresponding to the at least one first parameter is obtained, or selecting the part in a random manner.
Shi teaches selecting the part of the value corresponding to the at least one first parameter, based on at least one operation among: selecting the part based on priority of each of the at least one first parameter, selecting the part based on a point in time at which each value corresponding to the at least one first parameter is obtained, or selecting the part in a random manner. ([Shi, page 5] “In optimization problem (18), weight wij assigns priority to request j of UE i.” and “Results indicate that if a UE’s weight is increased and it is larger than others, the number of served requests for this UE increases relative other UEs.”, wherein the examiner interprets assigning priority to a request using weight wij and increasing served requests based on higher weight to be the same as selecting the part based on priority of each of the at least one first parameter, because they are both directed to selecting network request or parameter information according to a priority value.), ([Shi, page 5] “FCFS algorithm: Available resources are allocated to network slice requests based on the arrival times of requests, i.e., at any given time, the oldest network slice request is answered first provided that the available resources are sufficient to grant this request.”, wherein the examiner interprets allocating resources based on arrival times and answering the oldest request first to be the same as selecting the part based on a point in time at which each value corresponding to the at least one first parameter is obtained, because they are both directed to selecting network request or parameter information according to when the information was obtained or arrived.), ([Shi, page 5] “Random algorithm: Available resources are allocated to uniformly randomly selected network slice requests.”, wherein the examiner interprets uniformly randomly selected network slice requests to be the same as selecting the part in a random manner, because they are both directed to selecting network request or parameter information using a random selection rule.)
Katti, Devitt, Bendre, Peng, Shi, and the instant application are analogous art because they are all directed to machine-learning-based RAN operation in which selected network-state values, request information, and reward-based learning are used to optimize RAN resource allocation and control decisions.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method claim 8 disclosed by Katti, Devitt, Bendre, and Peng to include the priority-based, time-based, and random selection operations disclosed by Shi. One would be motivated to do so to effectively select network-state or request information for learning-based RAN resource allocation under different network conditions, as suggested by Shi ([Shi, page 3] “assignments is to maximize the weighted number of supported requests or the total provided services, where weights represent priorities of these requests.”).
Claims 13 are rejected under 35 U.S.C. 103 as being unpatentable over Katti in view of Devitt in view of Bendre in view of Shi.
Regarding claim 13, Katti, Devitt, and Bendre teaches The method of claim 1 (see rejection of claim 1).
Katti, Devitt, and Bendre do not teach wherein the experience information comprises: a value corresponding to the at least one first parameter at a first point in time, a first operation performed by the RAN at the first point in time, a value corresponding to the at least one first parameter at a second point in time after the first point in time according to a result of performing the first operation, and a reward value at the first point in time.
Shi teaches wherein the experience information comprises: a value corresponding to the at least one first parameter at a first point in time, a first operation performed by the RAN at the first point in time, a value corresponding to the at least one first parameter at a second point in time after the first point in time according to a result of performing the first operation, and a reward value at the first point in time. [Shi, page 4, sec 3.A], “Starting Q as a random matrix and using the weighted average of the old value and the new information, Q-learning performs the value iteration update for Q as follows:
PNG
media_image1.png
62
463
media_image1.png
Greyscale
where α is the learning rate (0 < α ≤ 1) and γ is the discount factor (0 ≤ γ ≤ 1) for rewards over time. In (19), max_a Q(st+1, a) refers to the estimate of the optimal future value of Q” ,wherein the examiner interprets “using the weighted average of the old value and the new information” to be the same as “a value corresponding to the at least one first parameter at a first point in time”, as both describe incorporating past parameter values as a start into an iterative learning process. The examiner further interprets “Q-learning performs the value iteration update for Q” to be the same as “a first operation performed by the RAN at the first point in time”, as both describe executing an operation that updates system behavior. The examiner also interprets “max_a Q(st+1, a) refers to the estimate of the optimal future value of Q” to be the same as “a value corresponding to the at least one first parameter at a second point in time after the first point in time according to a result of performing the first operation”, as both describe computing a future parameter value based on the results of a prior operation. Finally, the examiner interprets “γ is the discount factor (0 ≤ γ ≤ 1) for rewards over time” to be the same as “a reward value at the first point in time”, as both describe assigning a reward value based on past actions and their anticipated impact over time.)
Katti, Devitt, Bendre, Shi, and the instant application are analogous art because they are all directed to machine-learning-based RAN operation in which network state values, RAN operations, subsequent network state values, and reward information are used as experience information for learning and optimizing RAN control decisions.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the method claim 1 disclosed by Katti, Devitt, and Bendre to include the state-action-reward-next-state experience information disclosed by Shi. One would be motivated to do so to effectively generate reinforcement-learning experience information over time for improving RAN resource-allocation decisions, as suggested by Shi ([Shi, page 4, Sec. III.A] “At each time t, the gNodeB selects an action at, observes a reward rt, and transitions from the current state st to a new state st+1…and updates Q.”).
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DEVAN KAPOOR whose telephone number is (703)756-1434. The examiner can normally be reached Monday - Friday: 9:00AM - 5:00 PM EST (times may vary).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Yi can be reached at (571) 270-7519. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DEVAN KAPOOR/Examiner, Art Unit 2126
/DAVID YI/Supervisory Patent Examiner, Art Unit 2126