Quantum deep reinforcement learning for clinical decision support in oncology: application to adaptive radiotherapy | Scientific Reports – Nature.com
Quantum deep reinforcement learning
Quantum deep reinforcement learning is a novel action value-based decision-making framework derived from QRL23 and deep q-learning10 framework. Like conventional RL9,31, our qDRL based CDSS framework is comprised of 5 main elements: clinical AI agent, ARTE, radiation dose decision-making policy, reward, and q-value function. Here, the AI agent is a clinical decision-maker that learns to make dose decisions for achieving clinically desirable outcomes within the ARTE. The learning takes place by the agent-environment interaction, which can be sequentially ordered as: the AI decides on a dose and executes it, and in response, a patient (part of the ARTE) transits from one state to the next. Each transition provides the AI with feedback for its decision in terms of RT outcome and associated reward value. The goal of RL is for the AI to learn a decision-making policy that maximizes the reward in the long run, defined in terms of a specified q-value function that assigns a value to every state-dose-decision pair obtained from the accumulation of rewards over time (returns).
Assuming Markovs property (i.e., an environments response at time (t + 1) depends only on the state and dose-decision at time (t)), the qDRL task can be mathematically described as a 5-tuple ((S, left| D rightrangle , TF, P, R)), where (S) is a finite set of patients states, (left| D rightrangle) is a superimposed quantum state representing the finite set of eigen-dose decision, (TF:S times D to S^{prime }) is the transition function that maps patients state (s_{t}) and eigen-dose (left| d rightrangle_{t}) to the next state (s_{t + 1}), (P_{LC|RP2} :S^{prime } to left[ {0,1} right]) is the RT outcome estimator that assigns probability values (p_{LC}) and (p_{RP2}) to the state (s_{t + 1}), and (R:left[ {0,1} right] times left[ {0,1} right] to {mathbb{R}}) is the reward function that assigns a reward (r_{t + 1}) to the state-decision pair (left( {s_{t} ,left| d rightrangle_{t} } right)) based on the outcome probability estimates.
Eigen-dose (left| d rightrangle) is a physically performable decision that is selected via quantum methods from the superimposed quantum state (left| D rightrangle) which simultaneously represents all possible eigen-doses at once. In simple words, (left| D rightrangle) is the collection of all possible dose options and (left| d rightrangle) is one of those options which is selected after a decision is made. Selecting dose decision (left| d rightrangle) is carried out in two steps: (1) amplifying the optimal eigen-dose (left| d rightrangle^{*}) from the superimposed state (left| D rightrangle) (i.e., (left| D rightrangle^{prime } = widehat{Amp}_{{left| d rightrangle^{*} }} left| D rightrangle)) and (2) measuring the amplified state (i.e., (left| d rightrangle = widehat{Measure}(left| {D^{prime } } rightrangle )).
The optimal eigen-dose (left| d rightrangle^{*}) is obtained from deep Q-net, which is the AIs memory. Deep Q-net, (DQN:S to {mathbb{R}}^{d}), is a neural network that takes patients state as input and then outputs q-value for each eigen-dose ((left{ {q_{left| d rightrangle } } right})). The optimal dose is then selected following greedy policy where the dose with the maximum q-value is selected (i.e., (left| d rightrangle^{*} = begin{array}{*{20}c} {argmax} \ {left| {d^{prime } } rightrangle } \ end{array} { q_{left| d rightrangle } })). We have applied a double Q-learning 32 algorithm in training the deep Q-net. The schematic of a training cycle is presented in Fig.2 and additional technical details are presented in the Supplementary Material.
We initially employed Grovers amplification procedure33,34 for the decision selection mechanism. While Grovers procedure works on a quantum simulator, it fails to correctly work in a quantum computer. The quantum circuit depth of Grovers procedure (for 4 or higher qubits) is much greater than the coherence length of the current quantum processor35. Whenever the quantum circuit length exceeds the coherence length, quantum state becomes significantly affected by the system noise and loses vital information. Therefore, we designed a quantum controller circuit that is shorter than the coherence length and is suitable for the task of decision selection. The merit of our design is its fixed length; since its length is fixed for any number of qubits, it is suitable for higher qubit systems, as much as permitted by the circuit width. Technical details regarding its implementation in quantum processor is presented in the Supplementary Materials.
An example of a controller circuit is given in Fig.5. Controller circuits use twice the number of qubits (n), which can be divided into control and main. Optimal eigen-states obtained from deep Q-net are created in the control by selecting the appropriate pre-control gates. Then the control is entangled with the qubits from the main via controlled NOT (CNOT) gates. CNOT gates are connected between a control qubit from the control and a target qubit from the main. CNOT gates flip the target qubit from (left| 1 rightrangle) to (left| 0 rightrangle) only when the control is in (left| 1 rightrangle) state and does not perform any operation otherwise. Because all the main qubits are prepared in (left| 0 rightrangle) state, we introduced the reverse gates (n X-gates in parallel) to flip them to (left| 1 rightrangle). X-gates flip (left| 0 rightrangle) to (left| 1 rightrangle), and vice-versa. The CNOT flips all the qubits whose controls are in (left| 1 rightrangle) state, creating a state that is element-wise opposite to the marked state. Finally, another set of reverse gates is applied to the main before making a measurement.
Quantum controller circuit for a 5 qubit (32 bit) system. (a) Quantum controller circuit for the selection of the state (left| {10101} rightrangle). The probability distribution corresponding to (b) failed Grovers amplification procedure for one iteration run in the 5-qubit IBMQ Santiago quantum processor and (c) successful quantum controller selection run in the 15-qubit IBMQ Melbourne quantum processor.
Another advantage of the controller circuit is controlled uncertainty level. The controller circuit has additional degrees of freedom that can control the level of uncertainty that might be needed to model a highly dubious clinical situation. By replacing the CNOT gate by a more general (CU3left( {theta ,phi ,lambda } right)) gate, we can control the level of additional stochasticity with the rotation angles (theta), (phi), and (lambda), which corresponds to the angles in the Bloch sphere. The angles can either be fixed or, for additional control, changed with training episode.
The patients state in the ARTE is defined by 5 biological features: cytokine (IP10), PET imaging feature (GLSZM-ZSV), radiation doses (Tumor gEUD and lung gEUD), and genetics (cxcr1- Rs2234671). Their descriptions are presented in Table 2. These 5 variables were selected from a multi-objective Bayesian Network study13, which considered over 297 various biological features and found the best features for predicting the joint LC and RP2 RT outcomes.
The training data analyzed in this study are obtained from the University of Michigan study UMCC 2007.123 (NCI clinical trial NCT01190527) and the validation data analyzed in this study are obtained from the RTOG-0617 study (NCI clinical trial NCT00533949). Both trials were conducted in accordance with relevant guidelines and regulations and informed consent was obtained from all subjects and/or legal guardians. Details on training and validation datasets, and necessary model imputation carried out to accommodate the differences in the datasets are presented in the Supplementary Materials.
Deep Neural Networks (DNN) were applied as transition functions for IP10 and GLSZM-ZSV features. They were trained with a longitudinal (time-series) dataset, with the pre-irradiation patient state and corresponding radiation dose as input features and post-irradiation state as output. For lung and tumor gEUD, we utilized prior knowledge and applied a monotonic relationship for the transition function since we know that gEUD should increase with increasing radiation dose. We assumed that the change in gEUD is proportional to the dose fractionation and tissue radiosensitivity,
$$frac{{gleft( {t_{n} } right) - gleft( {t_{n - 1} } right)}}{{t_{n} - t_{n - 1} }} propto d_{n} left( {1 + frac{{d_{n} }}{{frac{alpha }{beta }}}} right).$$
(1)
Here (gleft( {t_{n} } right)) is the gEUD at time point (t_{n}), (d_{n}) is the radiation dose fractionation given during the nth time period, and (alpha /beta) ratio is the radiosensitivity parameter which differs between tissue type. Note that we first applied constrained training42 to maintain monotonicity with DNN model, however the gEUD over time trend was flatter than anticipated, thus we opted for a process-driven approach in the final implementation. The technical details on the NNs and its training are presented in the Supplementary Material.
DNN classifiers were applied as the RT outcome estimator for LC and RP2 treatment outcomes. They were trained with post irradiation patient states as input and binary LC and RP2 outcomes as its labels.
RT outcome estimator must also satisfy a monotone condition between increasing radiation dose and increasing probability of local control as well as probability of radiation induced pneumonitis. To maintain this monotonic relationship, we used a generic logistic function,
$$p_{LC|RP2} = frac{1}{{1 + exp left( {frac{{gleft( {t_{6} } right) - mu }}{T}} right)}},$$
(2)
where (gleft( {t_{6} } right)) is the gEUD at week 6, and (mu) and (T) are two patient-specific parameters that are learned from training the DNN. Here, (mu) and (T) are the outputs of two neural networks that are fed into the logistic function and tuned one after the other, leaving the other fixed. The training details are presented in the Supplementary Materials.
The task of the agent is to determine the optimal dose that maximizes (p_{LC}) while minimizing (p_{RP2}). Accordingly, we built a reward function on the base function (P^{ + } = P_{LC} left( {1 - P_{RP2} } right)) as shown in Fig.6. The algebraic form is as follows,
$$R = left{ {begin{array}{*{20}l} {P^{ + } + 10 } hfill & { {text{if}} 70% < p_{Lc} < 100% ;{text{and}}; 0% < p_{RP2} < 17.2% } hfill \ {P^{ + } + 5} hfill & {{text{if}} 50% < p_{Lc} < 70% ;{text{and}}; 17.2% < p_{RP2} < 50% } hfill \ {P^{ + } - 1} hfill & {{text{if}} 0% < p_{Lc} < 50% ;{text{and}}; 50 < p_{RP2} < 100% } hfill \ end{array} } right.$$
(3)
Reward function for reinforcement learning. Contour plot of reward function as a function of the probability of local control (PLC) and radiation induced pneumonitis of grade 2 or higher (PRP2). Area enclosed by the blue line corresponds to the clinically desirable outcome, i.e., (P_{LC} > 70{%}) and ({P_{RP2}} <17.2{%}). Similarly, the area enclosed by the green lines corresponds to the computationally desirable outcome, i.e., (P_{LC} > 50{%}) and ({P_{RP2}} <50{%}). Along with (P_{LC} times (1-P_{RP2})) the AI agent receives+10 reward for achieving clinically desirable outcome,+5 for achieving computationally desirable outcome, and -1 when unable to achieve a desirable outcome.
Here the AI agent receives additional 10 points for achieving clinically desirable outcome (i.e., (p_{LC} > 70% quad {text{and}} quad p_{RP2} < 17.2%)), 5 points for achieving computationally desirable outcome (i.e., (p_{LC} > 50% quad {text{and}} quad p_{RP2} < 50%)), and -1 point for failing to achieve a desirable outcome altogether. The negative point motivates the AI agent to search for the optimal dose as soon as possible.
To compensate for low number of data points we employed WGAN-GP43, which learns the underlying data distribution and generates more data points. We generated 4000 additional data points for training qDRL models. Having a larger training dataset helps the reinforcement learning algorithm in accurately representing the state space. The training details are presented in the Supplementary Material.
See the rest here:
Quantum deep reinforcement learning for clinical decision support in oncology: application to adaptive radiotherapy | Scientific Reports - Nature.com
- A Quantum Physicist Told His Team 'I Think AGI Is Here.' Here's What Convinced Him - Forbes - September 11th, 2026 [September 11th, 2026]
- What Does It Take for a Quantum Computer to Actually Be Quantum? - Quantinuum - September 11th, 2026 [September 11th, 2026]
- These 3 Quantum Stocks Are on the Rise as the US Government Takes Minority Stakes - Yahoo Finance - September 11th, 2026 [September 11th, 2026]
- Switzerland Is Getting Its First Dedicated IBM Quantum Computer - Tomorrow's World Today - September 11th, 2026 [September 11th, 2026]
- Beat the odds: Master and execute Monte Carlo Integration for finance on real quantum computers - Q-CTRL - September 11th, 2026 [September 11th, 2026]
- Why organizations need to act on quantum-safe cybersecurity - The World Economic Forum - September 11th, 2026 [September 11th, 2026]
- Could a Diamond Become the Heart of a Quantum Computer? - Fujitsu Global - September 11th, 2026 [September 11th, 2026]
- IBM, Lockheed Martin team up to bring Switzerland its first 120-qubit quantum computer - Interesting Engineering - September 11th, 2026 [September 11th, 2026]
- For banks, quantum computing is both a threat to security and a chance to build resilience - The World Economic Forum - September 11th, 2026 [September 11th, 2026]
- IBM, Lockheed Martin Announce Swiss Quantum Innovation Hub at ETH Zurich, Anchored by Switzerland's First IBM Quantum Computer - PR Newswire - September 11th, 2026 [September 11th, 2026]
- Fujitsu And Yaqumo Begin Testing on Neutral-Atom Quantum Computer Hardware - The Quantum Insider - September 11th, 2026 [September 11th, 2026]
- Quantum Computing Researchers Cut Estimated Cost of Attacking Bitcoin and Ethereum Encryption in Half - northeasttimes.com - September 11th, 2026 [September 11th, 2026]
- Cleveland Clinic, RIKEN and IBM Team Advance to Finals for 2026 ACM Gordon Bell Prize - IBM Newsroom - September 11th, 2026 [September 11th, 2026]
- Researchers Cut the Estimated Quantum Cost of a Key Step in Attacking Bitcoin and Ethereum by More Than Half - unchainedcrypto.com - September 11th, 2026 [September 11th, 2026]
- BWT Alpine and SEALSQ Explore Quantum-Classical Simulation for Formula One - The Quantum Insider - September 11th, 2026 [September 11th, 2026]
- 3 Quantum Stocks to Buy If You Already Have IonQ - Yahoo Finance - September 11th, 2026 [September 11th, 2026]
- China is poised to lead in research on quantum technologies. Democracies must act - The Strategist | ASPI's analysis and commentary site - September 11th, 2026 [September 11th, 2026]
- The commercialization of quantum computers is three years away. Although it is still in the stage of.. - - September 11th, 2026 [September 11th, 2026]
- Not IonQ, Not Rigetti Computing. This Quantum Computing Stock Could Be September's Biggest Winner. - The Globe and Mail - September 11th, 2026 [September 11th, 2026]
- Semiconductor and Quantum Developments Highlighted at Industry Conferences - IndexBox - September 11th, 2026 [September 11th, 2026]
- Quobly and Absolut System Partner on Cryogenic Infrastructure for Spin-Qubit Quantum Computers - The Quantum Insider - September 11th, 2026 [September 11th, 2026]
- Washington Bought Into Quantum Computing Crypto Should Be Paying Attention - blockhead.co - September 11th, 2026 [September 11th, 2026]
- When Seconds Replace Minutes: Quantum AI and the Fragility of Deterrence - Australian Institute of International Affairs - September 11th, 2026 [September 11th, 2026]
- Quantum computing will cause mayhem. Its an opportunity for investors - The Telegraph - September 11th, 2026 [September 11th, 2026]
- NEC has quietly quit quantum computing hardware development, report claims company says it will continue to evaluate practical applications and... - September 8th, 2026 [September 8th, 2026]
- U.S. Gets Minority Stakes in Some Quantum Computing Companies in $300 Million CHIPS Act Funding Deal - wsj.com - September 8th, 2026 [September 8th, 2026]
- Why quantum computing stocks are soaring today while the rest of the market falls - fastcompany.com - September 8th, 2026 [September 8th, 2026]
- 3 Top Quantum Computing Stocks to Buy in September - Yahoo Finance - September 8th, 2026 [September 8th, 2026]
- We Just Published the First Full-Stack Blueprint for Breaking 256-Bit Elliptic-Curve Signatures Heres What That Means - IonQ - September 8th, 2026 [September 8th, 2026]
- Ethereum Foundation Gives Itself Until 2029 to Guard Against Quantum Attacks - Northeast Times - September 8th, 2026 [September 8th, 2026]
- PsiQuantum Finalizes $100 Million Award with the U.S. Department of Commerce - The Quantum Insider - September 8th, 2026 [September 8th, 2026]
- Quantum Computing Stocks: IonQ Hosts Investor Day With SkyWater Deal In Focus - Investor's Business Daily - September 8th, 2026 [September 8th, 2026]
- Fujitsu Prototypes Diamond Quantum Computer Expandable via Optical Interconnects, Targeting 1,000 Logical Qubits by FY2035 - finance.biggo.com - September 8th, 2026 [September 8th, 2026]
- D-Wave Quantum Secures Up To $100 Million In CHIPS And Science Act Funding - pulse2.com - September 8th, 2026 [September 8th, 2026]
- Quantinuum Finalizes $100 Million CHIPS R&D Award with U.S. Department of Commerce to Advance Trapped-Ion Quantum Computer Manufacturing in the US -... - September 8th, 2026 [September 8th, 2026]
- Rigetti Signs Definitive Agreement for $100M with U.S. Government for Quantum Computing R&D - The Quantum Insider - September 8th, 2026 [September 8th, 2026]
- Biggest stock movers Tuesday: ZIM, NVS, quantum computing stocks, and more - Seeking Alpha - September 8th, 2026 [September 8th, 2026]
- The U.S. government agreed to award $100 million for quantum computing research with Rigetti - Stock Titan - September 8th, 2026 [September 8th, 2026]
- $100 million US push brings 300mm chips into the race for trapped-ion quantum hardware - Interesting Engineering - September 8th, 2026 [September 8th, 2026]
- Why Rigetti Computing Popped Today - The Motley Fool - September 8th, 2026 [September 8th, 2026]
- Is the Quantum Computing Threat to Bitcoin Overblown? These New Developments Suggest That's the Case. - The Globe and Mail - September 8th, 2026 [September 8th, 2026]
- Why D-Wave Quantum Stock Popped Today - Yahoo Finance - September 8th, 2026 [September 8th, 2026]
- IonQ and Congruity360 Partner to Enhance Quantum-Safe Protection - Business Wire - September 8th, 2026 [September 8th, 2026]
- Why Rigetti Computing Popped Today - Yahoo Finance - September 8th, 2026 [September 8th, 2026]
- The U.S.-China Quantum Race: How the State and Private Sector Are Being Mobilized for Quantum - The National Bureau of Asian Research (NBR) - September 4th, 2026 [September 4th, 2026]
- Quobly and TNO Partner on Silicon Spin Qubit Development - The Quantum Insider - September 4th, 2026 [September 4th, 2026]
- Nuclear reactor simulations advance with quantum particle transport breakthrough - Interesting Engineering - September 4th, 2026 [September 4th, 2026]
- Scientek and Classiq Partner to Accelerate Quantum Software Adoption in Taiwan - The Quantum Insider - September 4th, 2026 [September 4th, 2026]
- IBMs Nighthawk r2 Quantum Processor Targets a 25-Fold Increase in Circuit Speed - The Quantum Insider - September 4th, 2026 [September 4th, 2026]
- From MIT to IBM, expediting AI and quantum deployment - MIT News - September 4th, 2026 [September 4th, 2026]
- The Drawing on the Blackboard - The Times of Israel - September 4th, 2026 [September 4th, 2026]
- Quantum Explodes on the International Stage - afcea.org - September 4th, 2026 [September 4th, 2026]
- IonQ quantum computer Tempo goes online at KISTI in April 2027, Han River supercomputer 6 to launch this year - DongA Science - September 4th, 2026 [September 4th, 2026]
- Novo Nordisk's Foundation Is Building a Quantum Chip Factory in Copenhagen - Startup Fortune - September 4th, 2026 [September 4th, 2026]
- China Reaches 98% Efficiency With a Single Quantum Router on Origin Wukong - mediasat.info - September 4th, 2026 [September 4th, 2026]
- QC Ware and IonQ Demonstrate Hybrid Quantum Chemistry Workflow for Drug Discovery - The Quantum Insider - September 4th, 2026 [September 4th, 2026]
- Quantinuum to explore energy-industry uses for quantum tech with Saudi oil giant - BizWest - September 4th, 2026 [September 4th, 2026]
- IBM Stock Is Under Pressure. Is Its Quantum Business Reason Enough to Buy? - Barron's - September 4th, 2026 [September 4th, 2026]
- Old soup factory takes on new high-tech life as quantum computing firm gets $195M federal boost - CBC - August 29th, 2026 [August 29th, 2026]
- Quantum Computings Insiders Have Sold Nearly $863 Million More Than Theyve Bought. Should Investors Worry? - Yahoo Finance - August 29th, 2026 [August 29th, 2026]
- This Quantum Stock Is Surging 40% in Its Trading Debut. Meet Pasqal. - Barron's - August 29th, 2026 [August 29th, 2026]
- IBM Completes Acquisition of HRL Laboratories to Accelerate the Future of Quantum - IBM Newsroom - August 29th, 2026 [August 29th, 2026]
- Guest Post: Quantum Readiness Starts With the Infrastructure Were Building Today - The Quantum Insider - August 29th, 2026 [August 29th, 2026]
- Michaela Eichinger (Quantum Machines): Why classical compute and HPC integration will define useful quantum - The Quantum Insider - August 29th, 2026 [August 29th, 2026]
- Three Law Firms Hold Nearly Two-Thirds of Recorded Quantum Legal Work - The Quantum Insider - August 29th, 2026 [August 29th, 2026]
- From Roadmaps to Reality: How SoftBank Corp and Quantinuum Are Structuring the Path to Quantum Value - Quantinuum - August 29th, 2026 [August 29th, 2026]
- Quantum Computing Is Closer Than Insurers Think - Insurance Innovation Reporter - August 29th, 2026 [August 29th, 2026]
- Quantum Progress: From The Lab To The Real World - Forbes - August 29th, 2026 [August 29th, 2026]
- Ready before the hardware is: How agencies can get ahead of the quantum curve - Federal News Network - August 29th, 2026 [August 29th, 2026]
- Practical, self-correcting quantum computers could soon be developed in US with new funding - Interesting Engineering - August 29th, 2026 [August 29th, 2026]
- Stop Waiting for Q-Day: The Quantum Clock Is Ticking - Clearance Jobs - August 29th, 2026 [August 29th, 2026]
- Illinois-led regional quantum hub NSF HQAN renewed to pursue industry-ready computing and workforce development - The Grainger College of Engineering - August 29th, 2026 [August 29th, 2026]
- Scientists Are Building Quantum Computers That Fix Their Own Mistakes - Gadget Review - August 29th, 2026 [August 29th, 2026]
- Scientists Are Building Quantum Computers That Fix Their Own Mistakes - Yahoo Tech - August 29th, 2026 [August 29th, 2026]
- The Quantum Computer Revolution Is Tantalizingly Close - Advisor Perspectives - August 29th, 2026 [August 29th, 2026]
- 1 Quantum Computing Stock That Looks Like a Screaming Buy Right Now - Yahoo Finance - August 23rd, 2026 [August 23rd, 2026]
- Want to muck around with a real quantum computer? Now you can - The Conversation - August 23rd, 2026 [August 23rd, 2026]
- How Fujitsu is bridging quantum readiness and enterprise needs - cio.com - August 23rd, 2026 [August 23rd, 2026]
- New EPB CEO talks about the broadband companys work with quantum - Fierce Network - August 23rd, 2026 [August 23rd, 2026]
- D-Wave Quantum expands footprint for planned HQ - The Business Journals - August 23rd, 2026 [August 23rd, 2026]