Machine learning modeling to predict heuristic parameters for radiation therapy treatment planning
Methods and systems for configuring a plan optimizer model for radiotherapy treatment is presented herein in which a processor iteratively trains a machine learning model configured to predict a heuristic parameter, wherein with each iteration, an agent of the machine learning model identifies a test heuristic parameter; transmits the test heuristic parameter to the plan optimizer model configured to receive one or more radiotherapy treatment attributes and predict a treatment plan; and identifies a reward for the test heuristic parameter based on execution performance value of the plan optimizer model, wherein the processor iteratively trains a policy of the machine learning model until the policy satisfies an accuracy threshold based on maximizing the reward.
1 . A method of configuring a plan optimizer model for radiotherapy treatment using a machine learning model, the method comprising:
iteratively training, by a processor, the machine learning model implemented as a reinforcement learning agent configured to predict a heuristic parameter corresponding to a conjugate gradient mixing ratio, wherein with each iteration, an agent of the machine learning model:
identifies a test heuristic parameter;
transmits the test heuristic parameter to the plan optimizer model configured to receive one or more radiotherapy treatment attributes and predict a treatment plan generated based on the conjugate gradient mixing ratio; and
identifies a reward for the test heuristic parameter based on execution performance value of the plan optimizer model, the execution performance value comprising at least one measure of treatment plan quality and at least one indicator of operational performance of the plan optimizer model,
wherein the processor iteratively trains a policy of the machine learning model until the policy satisfies an accuracy threshold based on maximizing the reward; and
configuring, by the processor, the plan optimizer model based on a predicted test heuristic parameter predicted by the machine learning model.
2 . The method of claim 1 , wherein a category of the test heuristic parameter further corresponds to an initial step length in line search or a number of leaf tip mutation trials.
3 . The method of claim 1 , wherein the reward is based on whether the plan optimizer model converges upon a predicted treatment plan.
4 . The method of claim 1 , wherein the reward is based on an execution time of the plan optimizer model.
5 . The method of claim 1 , wherein the heuristic parameter corresponds to the test heuristic parameter having a maximum reward.
6 . The method of claim 1 , wherein the treatment plan comprises at least one radiotherapy machine attribute.
7 . The method of claim 1 , wherein the reward is based on a number of iterations for the plan optimizer model.
8 . The method of claim 1 , wherein the plan optimizer model is a machine learning model.
9 . The method of claim 1 , wherein the test heuristic parameter is within a defined range of values.
10 . A server comprising a processor and a non-transitory computer-readable medium containing instructions for configuring a plan optimizer model for radiotherapy treatment using a machine learning model, that when the instructions are executed by the processor, the instructions cause the processor to perform operations comprising:
iteratively training the machine learning model implemented as a reinforcement learning agent configured to predict a heuristic parameter, wherein with each iteration, an agent of the machine learning model:
identifies a test heuristic parameter;
transmits the test heuristic parameter to the plan optimizer model configured to receive one or more radiotherapy treatment attributes and predict a treatment plan generated based on a conjugate gradient mixing ratio;
identifies a reward for the test heuristic parameter based on execution performance value of the plan optimizer model, the execution performance value comprising at least one measure of treatment plan quality and at least one indicator of operational performance of the plan optimizer model,
wherein the processor iteratively trains a policy of the machine learning model until the policy satisfies an accuracy threshold based on maximizing the reward; and
configures the plan optimizer model based on a predicted test heuristic parameter predicted by the machine learning model.
11 . The server of claim 10 , wherein a category of the test heuristic parameter further corresponds to an initial step length in line search or a number of leaf tip mutation trials.
12 . The server of claim 10 , wherein the reward is based on whether the plan optimizer model converges upon a predicted treatment plan.
13 . The server of claim 10 , wherein the reward is based on an execution time of the plan optimizer model.
14 . The server of claim 10 , wherein the heuristic parameter corresponds to the test heuristic parameter having a maximum reward.
15 . The server of claim 10 , wherein the treatment plan comprises at least one radiotherapy machine attribute.
16 . The server of claim 10 , wherein the reward is based on a number of iterations for the plan optimizer model.
17 . The server of claim 10 , wherein the plan optimizer model is a machine learning model.
18 . The server of claim 10 , wherein the test heuristic parameter is within a defined range of values.
19 . A system for configuring a plan optimizer model for radiotherapy treatment using a machine learning model, the system comprising:
the plan optimizer model configured to receive one or more radiotherapy treatment attributes and predict a treatment plan; and
a server in communication with the plan optimizer model, the server comprising a processor and a non-transitory computer readable medium containing instructions that, when executed by the processor, cause the processor to:
iteratively train the machine learning model implemented as a reinforcement learning agent configured to predict a heuristic parameter corresponding to a conjugate gradient mixing ratio, wherein with each iteration, an agent of the machine learning model:
identifies a test heuristic parameter;
transmits the test heuristic parameter to the plan optimizer model; and
identifies a reward for the test heuristic parameter based on execution performance value of the plan optimizer model, the execution performance value comprising at least one measure of treatment plan quality and at least one indicator of operational performance of the plan optimizer model,
wherein the server iteratively trains a policy of the machine learning model until the policy satisfies an accuracy threshold based on maximizing the reward; and
configure the plan optimizer model based on a predicted test heuristic parameter predicted by the machine learning model.
20 . The system of claim 19 , wherein a category of the test heuristic parameter further corresponds to an initial step length in line search or a number of leaf tip mutation trials.