DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This action is in response to the Application filed on 10/31/2023. Claims 1-9 are pending in the case.
Priority
The present application claims priority under 35 U.S.C. §119 to Chinese Patent Application No. 202311339042.3, filed October 16, 2023. Applicant’s claim for the benefit of a prior-filed application under 35 U.S.C. 119 and/or 35 U.S.C. 120 is acknowledged.
Information Disclosure Statement
4. As required by MPEP 609 (c), the Applicants’ submission of the Information Disclosure Statement(s) filed on 10/31/2023 and 11/26/2025 are acknowledged by the examiner and the cited references have been considered in the examination of the claims now pending.
Examiner Comments
6. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Claim Rejections - 35 USC § 103
7. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
8. Claims 1-9 are rejected under 35 U.S.C. 103 as being unpatentable over Chaitanya (Pub. No.: US 20240115320 A1, Pub. Date: 2024-04-11) , in view of Smith ( Pub. No.: US 20200384289 A1, Pub. Date2020-12-10) in view of Soyer ( Pub. No.: US 20240127060 A1, Pub. Date: 2024-04-18)
Regarding independent Claim 1,
Chaitanya a method for training a deep reinforcement learning model for generating a treatment plan (see Chaitanya: Fig.2, [0026], “automatic planning and guidance of liver tumor thermal ablation using AI (artificial intelligence) agents trained with deep reinforcement learning.”), wherein the deep reinforcement learning model is configured to include a:
plurality of actor network layers and a critic network layer (see Chaitanya: Fig.5, [0044], “Framework 500 comprises a 3D shared neural network 506 and two smaller dense layer networks: actor network 508 and critic network 512.”), [… ], wherein the method comprises:
performing a training process (see Chaitanya: Fig.4, [0042], “framework 400 comprises two networks: online network 406 parameterized by weights θ and target network 420 parameterized by weights. Target network 420 estimates an optimal Q-value function during a training stage. Online network 406 chooses a best action given the estimated Q-value function from target network 420 and eventually aims to reach the Q-value estimated by target network 420 by the end of the training stage.”), the training process including the following operations:
determining, based on the […] data of the objective target volume, current policy data of the plurality of actor network layers, and current policy data of the critic network layer, target data (see Chaitanya: Fig.5, [0044], “determining a continuous action using PPO, in accordance with one or more embodiments. Framework 500 comprises a 3D shared neural network 506 and two smaller dense layer networks: actor network 508 and critic network 512. 3D shared neural network 506 receives current state S.sub.t 504 observed in environment 502 as input and extracts latent features from current state S.sub.t 504 as output. The latent features represent the most relevant or important features of current state S.sub.t 504 in a latent space of smaller dimensionality. Actor network 508 and critic network 512 receive the output of 3D shared neural network 506 as input and respectively generate policy mean μ.sub.v 510 and value function VB(S.sub.t) 514 as output. Policy mean μ.sub.v is a multi-dimensional (with three values for the 3D coordinates of electrode skin endpoint P.sub.v) representing a mean of the optimal action policy to be applied to reach an optimal state.”); and
updating, based on the target data, the current policy data of the plurality of actor network layers and the current policy data of the critic network layer, so as to complete a current training for the deep reinforcement learning model (see Chaitanya: Fig.5, [0050], “the steps of determining the one or more actions (step 108) and defining the next state (step 110) are repeated for a plurality of iterations to iteratively update the one or more current positions of the respective ablation electrode using 1) the next state as the current state and 2) the one or more updated positions as the one or more current positions to determine one or more final positions of the respective ablation electrode for performing an ablation on the one or more tumors. The ablation may be, for example, a radiofrequency ablation, a microwave ablation, a laser ablation, a cryoablation ablation, or an ablation performed using any other suitable technique.”), and
iterating the training process until a count of training the deep reinforcement learning model reaches a preset count, so as to obtain the deep reinforcement learning model that has been trained (see Chaitanya: Fig.5, [0051] “Steps 108-110 are repeated until a stopping condition is reached. In one embodiment, the stopping condition is that the current state S.sub.t reaches a terminal or final state, which occurs when the value of the net cumulative reward r.sub.t reaches a predetermined threshold value indicating that all clinical constraints are satisfied. The clinical constraints in Table 1 are satisfied when the net cumulative reward r.sub.t is 2.12 with r.sub.d≥0.12 and r.sub.l>0. In another embodiment, the stopping condition may be a predetermined number of iterations. In one example, as shown in FIG. 2, workflow 200 is repeated using next state S.sub.t+1 212 as current state S.sub.t 208 and updated net cumulative reward r.sub.t+1 214 as current net cumulative reward r.sub.t 210.”)
wherein the target data comprises:
a plurality of action sets output by the plurality of actor network layers corresponding to the multiple dose distribution state data respectively (see Chaitanya: Fig.5, [0044], “A random continuous action value a.sub.v=(a.sub.xv, a.sub.yv, a.sub.zv) sampled from the Gaussian distribution N(μ.sub.v, Σ.sub.v) and applied to determine updated positions of electrode skin endpoint P.sub.1+1=P.sub.v(a.sub.v)=(x.sub.v+a.sub.x.sub.v, y.sub.v+a.sub.y.sub.v, z.sub.v+a.sub.Z.sub.v) in environment 502 and an updated net cumulative reward r.sub.t+1 518.”),
a plurality of predicted values output by the critic network layer corresponding to the plurality of action sets respectively (see Chaitanya: Fig.5, [0044], “Value function VB(S.sub.t) 514 is a multivariate Gaussian distribution N(μ.sub.v, Σ.sub.v), where mean value μ.sub.v is policy mean μ.sub.v 510 output from actor network 508 and Σ.sub.v is a fixed variance value. Value function VB(S.sub.t) 514 indicates how well the action is in relation to the considered action policy and how to adjust the action to improve the next action where the terminal state is not reached.”),
a plurality of actual rewards corresponding to the plurality of action sets respectively (see Chaitanya: Fig.1, [0059], “multi-agent reinforcement learning (MARL) was utilized for training the AI agents. In MARL, the AI agents collaborate to maximize the cumulative rewards by learning optimal policies used by the individual AI agents to select a set of actions to go from a current state to a terminal state. To foster collaboration between AI agents, a VDN (value decomposition networks) approach may be used in which DON-style agents select and execute actions independent, but receive a joint reward computed on the overall state.”),
Chaitanya does not teach the system wherein;
acquiring initial dose distribution state data of an objective target volume;
determining, based on the initial dose distribution state] data of the objective target volume;
the target data comprises a final dose distribution of the objective target volume,
the target data comprises multiple dose distribution state data of the objective target volume;
the target data comprises a predicted value of the objective target volume, and an actual reward of the objective target volume;
wherein each action set of the plurality of action sets comprises a target parameter combination composed of a plurality of different types of target parameters.
different actor network layers of the plurality of actor network layers are configured to output different types of target parameters included in the treatment plan;
However, Smith teaches the system wherein:
acquiring initial dose distribution state data of an objective target volume (see Smith: Fig.3, [0044], “a knowledge base 302 and a treatment planning tool set 310. The knowledge base 302 includes patient records 304 (e.g., radiation treatment plans), treatment types 306, and statistical models 308. The treatment planning tool set 310 in the example of FIG. 3 includes a current patient record 312, a treatment type 314, a medical image processing module 316, the optimizer model (module) 150, a dose distribution module 320, and a final radiation treatment plan 322.”)
determining, based on the initial dose distribution state data of the objective target volume (see Smith: Fig.5, [0054], “a prescribed dose to be delivered into and across the target is determined. Each portion of the target can be represented by at least one 3D element known as a voxel; a portion may include more than one voxel. A portion of a target or a voxel may also be referred to herein as a sub-volume; a sub-volume may include one or more portions or one or more voxels.”)
the target data comprises a final dose distribution of the objective target volume (see Smith: Fig.5, [0054], “a prescribed dose to be delivered into and across the target is determined. Each portion of the target can be represented by at least one 3D element known as a voxel; a portion may include more than one voxel. A portion of a target or a voxel may also be referred to herein as a sub-volume; a sub-volume may include one or more portions or one or more voxels.”)
the target data comprises multiple dose distribution state data of the objective target volume (see Smith: Fig.3, [0043], “The medical image processing module 316 provides automatic contouring and automatic segmentation of two-dimensional cross-sectional slides (e.g., from computed tomography or magnetic resonance imaging) to form a three-dimensional (3D) image using the medical images in the current patient record 312. Dose distribution maps are calculated by the dose distribution module 320, which may utilize the optimizer model 150.”)
the target data comprises a predicted value of the objective target volume, and an actual reward of the objective target volume (see Smith: Fig.3, [0042], “The treatment planning tool set 310 searches through the knowledge base 302 (through the patient records 304) for prior patient records that are similar to the current patient record 312. The statistical models 308 can be used to compare the predicted results for the current patient record 312 to a statistical patient. Using the current patient record 312, a selected treatment type 306, and selected statistical models 308, the tool set 310 generates a radiation treatment plan 322..”)
wherein each action set of the plurality of action sets comprises a target parameter combination composed of a plurality of different types of target parameters (see Smith: Fig.7, [0080], “Parameters include, for example, average field dose rate, local dose rate, spot dose rate, instantaneous dose rate, computed with active time or total time, or any other specific definition of biologically relevant dose rate, or time-depended flux pattern, as that information becomes available through pre-clinical research.”)
Because both Chaitanya and Smith are in the same/similar field of endeavor of applying reinforcement learning to optimize treatment plans, accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention to modify teaching of Chaitanya to include the system that comprises initial dose distribution state data of target that include a final dose distribution of the objective target volume, multiple dose distribution state data of the objective target volume, a predicted value of the objective target volume, and an actual reward of the objective target volume, and a target parameter combination composed of a plurality of different types of target parameters included in the treatment plan as taught by Smith. One would be motivated to make such a combination in order to improve developing efficient and effective treatment dosage plans to prevent mistakes and dependency on the medial professionals and provide clinicians and researchers valuable information that can further be correlated with biological parameters and patient outcome.. (see Smith [0007]
Chaitanya and Smith does not teach the system wherein:
different actor network layers of the plurality of actor network layers are configured to output different types of target parameters treatment plan.
However, Soyer teaches the system wherein:
different actor network layers of the plurality of actor network layers are configured to output different types of target parameters ( see Soyer: Fig.2, [0067], “The training system 200 is a distributed computing system which includes one or more learner computing units (e.g., 202-A, 202-B, . . . , 202-Y) and multiple actor computing units (e.g., 204-A, 204-B, . . . , 204-X)”)… [0075], the environments corresponding to different actor computing units may generate rewards that characterize the progress of the agent towards accomplishing different tasks.)( Examiner notes that the treatment plan parameters is taught by Smith reference above.)
It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention to modify the Actor- Critic treatment planning of Chaitanya as applied to the radiation treatment planning parameters of Smith to include the system that include different actor network layers of the plurality of actor network layers are configured to output different types of target parameters as taught by Soyer. One would be motivated to make such a combination in order to improve developing efficient and effective treatment dosage plans to prevent mistakes and dependency on the medial professionals and provide clinicians and researchers valuable information that can further be correlated with biological parameters and patient outcome.
Regarding Claim 2,
Chaitanya, Smith and Soyer teach all the limitations of Claim 1 Chaitanya, Smith and Soyer further teaches the system wherein:
the plurality of actor network layers comprise at least two of a size actor network layer of a target, a position actor network layer of a target, or a weight actor network layer of a target (see Chaitanya: Fig.5, [0044], “Value function VB(S.sub.t) 514 is a multivariate Gaussian distribution N(μ.sub.v, Σ.sub.v), where mean value μ.sub.v is policy mean μ.sub.v 510 output from actor network 508 and Σ.sub.v is a fixed variance value. Value function VB(S.sub.t) 514 indicates how well the action is in relation to the considered action policy and how to adjust the action to improve the next action where the terminal state is not reached.”),
Regarding Claim 3,
Chaitanya, Smith and Soyer teach all the limitations of Claim 2. Chaitanya, Smith and Soyer further teaches the system wherein:
the determining, based on […] data of the objective target volume, the current policy data of the plurality of actor network layers, and the current policy data of the critic network layer (see Chaitanya: Fig.4, [0044], “determining a continuous action using PPO, in accordance with one or more embodiments. Framework 500 comprises a 3D shared neural network 506 and two smaller dense layer networks: actor network 508 and critic network 512. 3D shared neural network 506 receives current state S.sub.t 504 observed in environment 502 as input and extracts latent features from current state S.sub.t 504 as output. The latent features represent the most relevant or important features of current state S.sub.t 504 in a latent space of smaller dimensionality. Actor network 508 and critic network 512 receive the output of 3D shared neural network 506 as input and respectively generate policy mean μ.sub.v 510 and value function VB(S.sub.t) 514 as output. Policy mean μ.sub.v is a multi-dimensional (with three values for the 3D coordinates of electrode skin endpoint P.sub.v) representing a mean of the optimal action policy to be applied to reach an optimal state.”);.”), the target data comprises:
determining, based on the […] data of the objective target volume and current policy data of each actor network layer of the plurality of actor network layers, an action set corresponding to the current [dose distribution] state data output by the plurality of actor network layers data (see Chaitanya: Fig.5, [0044], “Actor network 508 and critic network 512 receive the output of 3D shared neural network 506 as input and respectively generate policy mean μ.sub.v 510 and value function VB(S.sub.t) 514 as output. Policy mean μ.sub.v is a multi-dimensional (with three values for the 3D coordinates of electrode skin endpoint P.sub.v) representing a mean of the optimal action policy to be applied to reach an optimal state.”)
determining, based on [ … ] of the objective target volume and the current policy data of the critic network layer, a predicted value output by the critic network layer corresponding to the action set data (see Chaitanya: Fig.5, [0044], “The latent features represent the most relevant or important features of current state S.sub.t 504 in a latent space of smaller dimensionality. Actor network 508 and critic network 512 receive the output of 3D shared neural network 506 as input and respectively generate policy mean μ.sub.v 510 and value function VB(S.sub.t) 514 as output. Policy mean μ.sub.v is a multi-dimensional (with three values for the 3D coordinates of electrode skin endpoint P.sub.v) representing a mean of the optimal action policy to be applied to reach an optimal state.)
in response to that the dose distribution of the objective target volume does not meet a preset prescription dose, and a number of targets in the objective target volume is less than a preset maximum number of targets, updating the current dose distribution state data of the objective target volume based on the dose distribution of the objective target volume; or,
in response to that the dose distribution of the objective target volume meets the preset prescription dose, and/or the number of targets in the objective target volume is equal to the preset maximum number of targets (see Smith: Fig.7B, [0086], “In step 764, quality assurance is performed on the dose rate prescription. The dose rate prescription gets passed through to a QA step, where now dose delivered and dose rate delivered is verified before patient treatment.”)
determining the final dose distribution of the objective target volume (see Smith: Fig.5, [0054], “a prescribed dose to be delivered into and across the target is determined. Each portion of the target can be represented by at least one 3D element known as a voxel; a portion may include more than one voxel. A portion of a target or a voxel may also be referred to herein as a sub-volume; a sub-volume may include one or more portions or one or more voxels.”),
the multiple dose distribution state data of the objective target volume (see Smith: Fig.3, [0043], “The medical image processing module 316 provides automatic contouring and automatic segmentation of two-dimensional cross-sectional slides (e.g., from computed tomography or magnetic resonance imaging) to form a three-dimensional (3D) image using the medical images in the current patient record 312. Dose distribution maps are calculated by the dose distribution module 320, which may utilize the optimizer model 150.”),
the plurality of action sets output by the plurality of actor network layers corresponding to the multiple dose distribution state data respectively (see Chaitanya: Fig.5, [0044], “A random continuous action value a.sub.v=(a.sub.xv, a.sub.yv, a.sub.zv) sampled from the Gaussian distribution N(μ.sub.v, Σ.sub.v) and applied to determine updated positions of electrode skin endpoint P.sub.1+1= P.sub.v(a.sub.v)= (x.sub.v+a.sub.x.sub.v, y.sub.v+a.sub.y.sub.v, z.sub.v+a.sub.Z.sub.v) in environment 502 and an updated net cumulative reward r.sub.t+1 518.”),
the plurality of predicted values output by the critic network layer corresponding to the plurality of action sets respectively (see Chaitanya: Fig.5, [0044], “Value function VB(S.sub.t) 514 is a multivariate Gaussian distribution N(μ.sub.v, Σ.sub.v), where mean value μ.sub.v is policy mean μ.sub.v 510 output from actor network 508 and Σ.sub.v is a fixed variance value. Value function VB(S.sub.t) 514 indicates how well the action is in relation to the considered action policy and how to adjust the action to improve the next action where the terminal state is not reached.”),, and
the plurality of actual rewards corresponding to the plurality of action sets respectively data (see Chaitanya: Fig.1, [0059], “multi-agent reinforcement learning (MARL) was utilized for training the AI agents. In MARL, the AI agents collaborate to maximize the cumulative rewards by learning optimal policies used by the individual AI agents to select a set of actions to go from a current state to a terminal state. To foster collaboration between AI agents, a VDN (value decomposition networks) approach may be used in which DON-style agents select and execute actions independent, but receive a joint reward computed on the overall state.”), and
determining, based on the final dose distribution of the objective target volume and the plurality of predicted values corresponding to the plurality of action sets respectively, the predicted value of the objective target volume and the actual reward of the objective target volume data (see Chaitanya: Fig.1, [0059], “multi-agent reinforcement learning (MARL) was utilized for training the AI agents. In MARL, the AI agents collaborate to maximize the cumulative rewards by learning optimal policies used by the individual AI agents to select a set of actions to go from a current state to a terminal state. To foster collaboration between AI agents, a VDN (value decomposition networks) approach may be used in which DON-style agents select and execute actions independent, but receive a joint reward computed on the overall state.”),
Smith teaches the initial dose distribution state data, the current dose distribution state data and on the current dose distribution state data, final dose distribution (see Smith: Fig.5, [0054], “a prescribed dose to be delivered into and across the target is determined. Each portion of the target can be represented by at least one 3D element known as a voxel; a portion may include more than one voxel. A portion of a target or a voxel may also be referred to herein as a sub-volume; a sub-volume may include one or more portions or one or more voxels.”)
It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention to modify teaching of Chaitanya to include the system that comprises initial dose distribution state data of target that include a final dose distribution of the objective target volume, multiple dose distribution state data of the objective target volume, a predicted value of the objective target volume, and an actual reward of the objective target volume, and a target parameter combination composed of a plurality of different types of target parameters included in the treatment plan as taught by Smith. One would be motivated to make such a combination in order to improve developing efficient and effective treatment dosage plans to prevent mistakes and dependency on the medial professionals and provide clinicians and researchers valuable information that can further be correlated with biological parameters and patient outcome. (see Smith [0007]
Regarding Claim 4,
Chaitanya, Smith and Soyer teach all the limitations of Claim 3. Chaitanya, Smith and Soyer further teaches the system wherein:
when the plurality of actor network layers comprises a first actor network layer and a second actor network layer deployed from top to bottom, the determining based on [… ] state data of the objective target volume and the current policy data of each actor network layer of the plurality of actor network layers, the action set corresponding to the [,…] state data output by the plurality of actor network layers (see Chaitanya: Fig.4, [0044], “determining a continuous action using PPO, in accordance with one or more embodiments. Framework 500 comprises a 3D shared neural network 506 and two smaller dense layer networks: actor network 508 and critic network 512. 3D shared neural network 506 receives current state S.sub.t 504 observed in environment 502 as input and extracts latent features from current state S.sub.t 504 as output. The latent features represent the most relevant or important features of current state S.sub.t 504 in a latent space of smaller dimensionality. Actor network 508 and critic network 512 receive the output of 3D shared neural network 506 as input and respectively generate policy mean μ.sub.v 510 and value function VB(S.sub.t) 514 as output. Policy mean μ.sub.v is a multi-dimensional (with three values for the 3D coordinates of electrode skin endpoint P.sub.v) representing a mean of the optimal action policy to be applied to reach an optimal state.”); comprises:
determining, based on the […]state data and current policy data of the first actor network layer, a first action (see Chaitanya: Fig.5, [0050], “the steps of determining the one or more actions (step 108) and defining the next state (step 110) are repeated for a plurality of iterations to iteratively update the one or more current positions of the respective ablation electrode using 1) the next state as the current state and 2) the one or more updated positions as the one or more current positions to determine one or more final positions of the respective ablation electrode for performing an ablation on the one or more tumors. The ablation may be, for example, a radiofrequency ablation, a microwave ablation, a laser ablation, a cryoablation ablation, or an ablation performed using any other suitable technique.”) )and
determining, based on the […] the first action, and current policy data of the second actor network layer, a second action corresponding to the first action (see Chaitanya: Fig.5, [0044], “Value function VB(S.sub.t) 514 is a multivariate Gaussian distribution N(μ.sub.v, Σ.sub.v), where mean value μ.sub.v is policy mean μ.sub.v 510 output from actor network 508 and Σ.sub.v is a fixed variance value. Value function VB(S.sub.t) 514 indicates how well the action is in relation to the considered action policy and how to adjust the action to improve the next action where the terminal state is not reached.”)
Smith teaches the initial dose distribution state data, the current dose distribution state data and on the current dose distribution state data (see Smith: Fig.5, [0054], “a prescribed dose to be delivered into and across the target is determined. Each portion of the target can be represented by at least one 3D element known as a voxel; a portion may include more than one voxel. A portion of a target or a voxel may also be referred to herein as a sub-volume; a sub-volume may include one or more portions or one or more voxels.”)
It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention to modify teaching of Chaitanya to include the system that comprises initial dose distribution state data of target that include a final dose distribution of the objective target volume, multiple dose distribution state data of the objective target volume, a predicted value of the objective target volume, and an actual reward of the objective target volume, and a target parameter combination composed of a plurality of different types of target parameters included in the treatment plan as taught by Smith. One would be motivated to make such a combination in order to improve developing efficient and effective treatment dosage plans to prevent mistakes and dependency on the medial professionals and provide clinicians and researchers valuable information that can further be correlated with biological parameters and patient outcome. (see Smith [0007]
Regarding Claim 5,
Chaitanya, Smith and Soyer teach all the limitations of Claim 1. Chaitanya, Smith and Soyer further teaches the system wherein:
when the plurality of actor network layers comprises a first actor network layer, a second actor network layer, and a third actor network layer deployed from top to bottom, the determining, based on the [… ] state data of the objective target volume and the current policy data of each actor network layer of the plurality of actor network layers, the action set corresponding to the […] state data output by the plurality of actor network layers (see Chaitanya: Fig.4, [0044], “determining a continuous action using PPO, in accordance with one or more embodiments. Framework 500 comprises a 3D shared neural network 506 and two smaller dense layer networks: actor network 508 and critic network 512. 3D shared neural network 506 receives current state S.sub.t 504 observed in environment 502 as input and extracts latent features from current state S.sub.t 504 as output. The latent features represent the most relevant or important features of current state S.sub.t 504 in a latent space of smaller dimensionality. Actor network 508 and critic network 512 receive the output of 3D shared neural network 506 as input and respectively generate policy mean μ.sub.v 510 and value function VB(S.sub.t) 514 as output. Policy mean μ.sub.v is a multi-dimensional (with three values for the 3D coordinates of electrode skin endpoint P.sub.v) representing a mean of the optimal action policy to be applied to reach an optimal state.”); comprises:
determining, based on the […] state data and current policy data of the first actor network layer, a first action (see Chaitanya: Fig.5, [0044], “Value function VB(S.sub.t) 514 is a multivariate Gaussian distribution N(μ.sub.v, Σ.sub.v), where mean value μ.sub.v is policy mean μ.sub.v 510 output from actor network 508 and Σ.sub.v is a fixed variance value. Value function VB(S.sub.t) 514 indicates how well the action is in relation to the considered action policy and how to adjust the action to improve the next action where the terminal state is not reached.”),
determining, based on the […] state data, the first action, and current policy data of the second actor network layer, a second action corresponding to the first action (see Chaitanya: Fig.5, [0050], “the steps of determining the one or more actions (step 108) and defining the next state (step 110) are repeated for a plurality of iterations to iteratively update the one or more current positions of the respective ablation electrode using 1) the next state as the current state and 2) the one or more updated positions as the one or more current positions to determine one or more final positions of the respective ablation electrode for performing an ablation on the one or more tumors. The ablation may be, for example, a radiofrequency ablation, a microwave ablation, a laser ablation, a cryoablation ablation, or an ablation performed using any other suitable technique.”).”)and
determining, based on the […] state data, the first section action, the second section action, and current policy data of the third actor network layer, a third action corresponding to the first action (see Chaitanya: Fig.5, [0050], “the steps of determining the one or more actions (step 108) and defining the next state (step 110) are repeated for a plurality of iterations to iteratively update the one or more current positions of the respective ablation electrode using 1) the next state as the current state and 2) the one or more updated positions as the one or more current positions to determine one or more final positions of the respective ablation electrode for performing an ablation on the one or more tumors. The ablation may be, for example, a radiofrequency ablation, a microwave ablation, a laser ablation, a cryoablation ablation, or an ablation performed using any other suitable technique.”).”)
Smith teaches the initial dose distribution state data, the current dose distribution state data and on the current dose distribution state data (see Smith: Fig.5, [0054], “a prescribed dose to be delivered into and across the target is determined. Each portion of the target can be represented by at least one 3D element known as a voxel; a portion may include more than one voxel. A portion of a target or a voxel may also be referred to herein as a sub-volume; a sub-volume may include one or more portions or one or more voxels.”)
Because both Chaitanya and Smith are in the same/similar field of endeavor of applying reinforcement learning to optimize treatment plans, accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention to modify teaching of Chaitanya to include the system that comprises initial dose distribution state data of target that include a final dose distribution of the objective target volume, multiple dose distribution state data of the objective target volume, a predicted value of the objective target volume, and an actual reward of the objective target volume, and a target parameter combination composed of a plurality of different types of target parameters included in the treatment plan as taught by Smith. One would be motivated to make such a combination in order to improve developing efficient and effective treatment dosage plans to prevent mistakes and dependency on the medial professionals and provide clinicians and researchers valuable information that can further be correlated with biological parameters and patient outcome.. (see Smith [0007]
Regarding Claim 6,
Chaitanya, Smith and Soyer teach all the limitations of Claim 4. Chaitanya, Smith and Soyer further teaches the system wherein:
updating, based on the target data, the current policy data of the plurality of actor network layers and the current policy data of the critic network layer, so as to complete a current training for the deep reinforcement learning model (see Chaitanya: Fig.5, [0050], “the steps of determining the one or more actions (step 108) and defining the next state (step 110) are repeated for a plurality of iterations to iteratively update the one or more current positions of the respective ablation electrode using 1) the next state as the current state and 2) the one or more updated positions as the one or more current positions to determine one or more final positions of the respective ablation electrode for performing an ablation on the one or more tumors. The ablation may be, for example, a radiofrequency ablation, a microwave ablation, a laser ablation, a cryoablation ablation, or an ablation performed using any other suitable technique.”), comprises:
in response to that the [… ] of the objective target volume obtained from the current training meets a preset prescription dose, determining whether the actual reward of the objective target volume obtained from the current training is greater than a dynamic reward threshold (see Chaitanya: Fig.5, [0050], “the steps of determining the one or more actions (step 108) and defining the next state (step 110) are repeated for a plurality of iterations to iteratively update the one or more current positions of the respective ablation electrode using 1) the next state as the current state and 2) the one or more updated positions as the one or more current positions to determine one or more final positions of the respective ablation electrode for performing an ablation on the one or more tumors. The ablation may be, for example, a radiofrequency ablation, a microwave ablation, a laser ablation, a cryoablation ablation, or an ablation performed using any other suitable technique.”), wherein the dynamic reward threshold is an actual reward of the objective target volume corresponding to target data used for updating the current policy data of the plurality of actor network layers and the current policy data of the critic network layer previously (see Chaitanya: Fig.5, [0050], “the steps of determining the one or more actions (step 108) and defining the next state (step 110) are repeated for a plurality of iterations to iteratively update the one or more current positions of the respective ablation electrode using 1) the next state as the current state and 2) the one or more updated positions as the one or more current positions to determine one or more final positions of the respective ablation electrode for performing an ablation on the one or more tumors. The ablation may be, for example, a radiofrequency ablation, a microwave ablation, a laser ablation, a cryoablation ablation, or an ablation performed using any other suitable technique.”).
in response to that the actual reward of the objective target volume obtained from the current training is greater than the dynamic reward threshold, determining a loss value of the objective target volume corresponding to the current training based on the actual reward and the predicted value of the objective target volume obtained from the current training (see Chaitanya: Fig.1, [0059], “multi-agent reinforcement learning (MARL) was utilized for training the AI agents. In MARL, the AI agents collaborate to maximize the cumulative rewards by learning optimal policies used by the individual AI agents to select a set of actions to go from a current state to a terminal state. To foster collaboration between AI agents, a VDN (value decomposition networks) approach may be used in which DON-style agents select and execute actions independent, but receive a joint reward computed on the overall state.”),
in response to that the loss value of the objective target volume corresponding to the current training is less than a dynamic loss value, updating the current policy data of the plurality of actor network layers and the current policy data of the critic network layer based on the multiple dose distribution state data obtained from the current training, the plurality of action sets corresponding to the [] state data respectively, the plurality of predicted values corresponding to the plurality of action sets respectively, and the actual reward of the objective target volume (see Chaitanya: Fig.5, [0050], “the steps of determining the one or more actions (step 108) and defining the next state (step 110) are repeated for a plurality of iterations to iteratively update the one or more current positions of the respective ablation electrode using 1) the next state as the current state and 2) the one or more updated positions as the one or more current positions to determine one or more final positions of the respective ablation electrode for performing an ablation on the one or more tumors. The ablation may be, for example, a radiofrequency ablation, a microwave ablation, a laser ablation, a cryoablation ablation, or an ablation performed using any other suitable technique.”),
Smith teaches the initial dose distribution state data, the current dose distribution state data and on the current dose distribution state data (see Smith: Fig.5, [0054], “a prescribed dose to be delivered into and across the target is determined. Each portion of the target can be represented by at least one 3D element known as a voxel; a portion may include more than one voxel. A portion of a target or a voxel may also be referred to herein as a sub-volume; a sub-volume may include one or more portions or one or more voxels.”)
It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention to modify teaching of Chaitanya to include the system that comprises initial dose distribution state data of target that include a final dose distribution of the objective target volume, multiple dose distribution state data of the objective target volume, a predicted value of the objective target volume, and an actual reward of the objective target volume, and a target parameter combination composed of a plurality of different types of target parameters included in the treatment plan as taught by Smith. One would be motivated to make such a combination in order to improve developing efficient and effective treatment dosage plans to prevent mistakes and dependency on the medial professionals and provide clinicians and researchers valuable information that can further be correlated with biological parameters and patient outcome.. (see Smith [0007]
Regarding Claim 7,
Chaitanya, Smith and Soyer teach all the limitations of Claim 6. Chaitanya, Smith and Soyer further teaches the system wherein:
the updating the current policy data of the plurality of actor network layers and the current policy data of the critic network layer based on the multiple dose distribution state data obtained from the current training, the plurality of action sets corresponding to the multiple dose distribution state data respectively, the plurality of predicted values corresponding to the plurality of action sets respectively, and the actual reward of the objective target volume (see Chaitanya: Fig.5, [0050], “the steps of determining the one or more actions (step 108) and defining the next state (step 110) are repeated for a plurality of iterations to iteratively update the one or more current positions of the respective ablation electrode using 1) the next state as the current state and 2) the one or more updated positions as the one or more current positions to determine one or more final positions of the respective ablation electrode for performing an ablation on the one or more tumors. The ablation may be, for example, a radiofrequency ablation, a microwave ablation, a laser ablation, a cryoablation ablation, or an ablation performed using any other suitable technique.”) and comprises:
determining, based on the plurality of actual rewards corresponding to the plurality of action sets and the actual reward of the objective target volume, an actual cumulated reward value of the plurality of action sets, and updating the current policy data of the plurality of actor network layers based on the multiple dose distribution state data, the plurality of action sets corresponding to the multiple dose distribution state data respectively, and the actual cumulated reward value of the plurality of action sets (see Chaitanya: Fig.5, [0050], “the steps of determining the one or more actions (step 108) and defining the next state (step 110) are repeated for a plurality of iterations to iteratively update the one or more current positions of the respective ablation electrode using 1) the next state as the current state and 2) the one or more updated positions as the one or more current positions to determine one or more final positions of the respective ablation electrode for performing an ablation on the one or more tumors. The ablation may be, for example, a radiofrequency ablation, a microwave ablation, a laser ablation, a cryoablation ablation, or an ablation performed using any other suitable technique.”) and
updating the current policy data of the critic network layer based on the multiple dose distribution state data, the plurality of action sets corresponding to the [] state data respectively, and the plurality of predicted values corresponding to the plurality of action sets respectively (see Chaitanya: Fig.5, [0050], “the steps of determining the one or more actions (step 108) and defining the next state (step 110) are repeated for a plurality of iterations to iteratively update the one or more current positions of the respective ablation electrode using 1) the next state as the current state and 2) the one or more updated positions as the one or more current positions to determine one or more final positions of the respective ablation electrode for performing an ablation on the one or more tumors. The ablation may be, for example, a radiofrequency ablation, a microwave ablation, a laser ablation, a cryoablation ablation, or an ablation performed using any other suitable technique.”) and
Smith teaches the initial dose distribution state data, the current dose distribution state data and on the current dose distribution state data (see Smith: Fig.5, [0054], “a prescribed dose to be delivered into and across the target is determined. Each portion of the target can be represented by at least one 3D element known as a voxel; a portion may include more than one voxel. A portion of a target or a voxel may also be referred to herein as a sub-volume; a sub-volume may include one or more portions or one or more voxels.”)
It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the invention to modify teaching of Chaitanya to include the system that comprises initial dose distribution state data of target that include a final dose distribution of the objective target volume, multiple dose distribution state data of the objective target volume, a predicted value of the objective target volume, and an actual reward of the objective target volume, and a target parameter combination composed of a plurality of different types of target parameters included in the treatment plan as taught by Smith. One would be motivated to make such a combination in order to improve developing efficient and effective treatment dosage plans to prevent mistakes and dependency on the medial professionals and provide clinicians and researchers valuable information that can further be correlated with biological parameters and patient outcome.. (see Smith [0007])
Regarding independent Claim 8,
Claim 8 is a directed to a method claim and has similar/same claim limitation as Claim 1 and is rejected under the same rationale.
Regarding Claim independent 9,
Claim 9 is a non-transitory computer readable storage medium and has similar/same claim limitation as Claim 1 and is rejected under the same rationale.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
PGPUB
NUMBER:
INVENTOR-INFORMATION:
TITLE / DESCRIPTION
US 12370378 B2
Lofman; Fredrik
Title: Method Of Generating A Radiotherapy Treatment Plan For A Patient, A Computer Program Product, And A Computer System Comprising A Machine Learning System
Description: The present invention relates to a method of producing a radiotherapy treatment plan and a method of training a machine learning system, and to a computer program product and a computer system
US 20170177812 A1
SJÕLUND; JENS Ola
Title: SYSTEMS AND METHODS FOR OPTIMIZING TREATMENT PLANNING
Description: This disclosure relates generally to radiation therapy or radiotherapy. More specifically, this disclosure relates to systems and methods for developing a statistically optimal radiation therapy treatment plan to be used during radiotherapy.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ZELALEM W SHALU whose telephone number is (571)272-3003. The examiner can normally be reached M- F 0800am- 0500pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached at (571) 272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Zelalem Shalu/Examiner, Art Unit 2145
/CESAR B PAULA/Supervisory Patent Examiner, Art Unit 2145