Reinforcement Learning for Personalized Insulin Dosing: A Comparative Study of A2C, SAC and PPO on Real-World Clinical Data

Article Sidebar

Main Article Content

Chinatu M. Anyanwu
Nkiru C. Ogbonna
Mary Ofuru Kama
Stephen Uche Udeh
Ogechi Gift Onyedi

Personalized insulin dosing for Type 1 diabetes mellitus (T1DM) remains challenging because of complex glucose-insulin dynamics and substantial patient variability. Reinforcement learning (RL) has emerged as a promising approach for adaptive insulin management, yet the reliability of learned policies depends heavily on reward design and evaluation strategy. This study compares three actor–critic RL algorithms: Soft Actor-Critic (SAC), Advantage Actor-Critic (A2C), and Proximal Policy Optimization (PPO) for personalized insulin dosing using real-world continuous glucose monitoring, insulin delivery, basal insulin, and meal intake data from the OhioT1DM dataset. A custom Gymnasium-based environment was developed, and all algorithms were trained under identical conditions for 100,000 timesteps. Performance was evaluated using cumulative reward together with clinically relevant measures, including Time in Range (TIR) and insulin dosing behaviour. Although A2C and PPO achieved higher cumulative rewards than SAC, both converged to near-zero insulin dosing policies that exploited the reward formulation rather than learning clinically meaningful glucose regulation. In contrast, SAC maintained adaptive dosing behaviour, achieving a TIR of 72.71% with an average insulin dose of 1.769 U/step. These findings show that higher cumulative reward does not necessarily correspond to better clinical decision-making in open-loop reinforcement learning environments. The study highlights the importance of behaviour-focused evaluation alongside conventional reward metrics and provides practical insights for developing safer and more reliable reinforcement learning systems for personalized diabetes management.

Reinforcement Learning for Personalized Insulin Dosing: A Comparative Study of A2C, SAC and PPO on Real-World Clinical Data. (2026). International Journal of Latest Technology in Engineering Management & Applied Science, 15(6), 3700-3710. https://doi.org/10.51583/IJLTEMAS.2026.150600273

Downloads

References

Bolland, A., Lambrechts, G., & Ernst, D. (2024). Off-policy maximum entropy rl with future state and action visitation measures. arXiv preprint arXiv:2412.06655.

Dénes-Fazakas, L., Szilágyi, L., Kovács, L., De Gaetano, A., & Eigner, G. (2024). Reinforcement learning: a paradigm shift in personalized blood glucose management for diabetes. Biomedicines, 12(9), 2143.

Elsayed, N. A., Aleppo, G., Bannuru, R. R., Bruemmer, D., Collins, B. S., Ekhlaspour, L., & American Diabetes Association Professional Practice Committee. (2024). 16. Diabetes Care in the Hospital: Standards of Care in Diabetes—2024. Diabetes Care, 47.

Haarnoja, T., Zhou, A., Abbeel, P., & Levine, S. (2018). Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. Proceedings of the 35th International Conference on Machine Learning. https://doi.org/10.48550/arXiv.1801.01290

Lei, J., Sun, X., Li, Y., Li, K., Zhang, S., Zeng, H., & Zhang, Y. (2024, August). An Improved Adaptive Glucose Control Approach for Type 1 Diabetes with Temporal Dependence. In 2024 IEEE 9th International Conference on Computational Intelligence and Applications (ICCIA) (pp. 209-214). IEEE.

Manas, S., Pillai, G. N. & Gupta, M. K. (2023). Improved Soft Actor-Critic: Reducing Bias and Estimation Error for Fast Learning. IEEE International Student’s Conference on Electrical, Electronics and Computer Science (SCEECS), 1 - 9,2023,doi:10.1109/SCEECS57921.

Milton T. & Lieck R. (2024). Fully-Automated Patient-Agnostic Diabetes Management with Deep Reinforcement Learning. IEEE International Conference on Bioinformatics and Biomedicine (BIBM), 1085-1091.

Mnih, V. , Adria, P. B. , M. Mehdi, G. Alex, H. Tim, P. L. Timothy, S. David & K. Koray, (2016). Asynchronous Methods for Deep Reinforcement Learning. Proceedings of the 33rd International Conference on Machine Learning, New York. NY USA. JLMR. W & CP, 48, doi:10.48550/arXiv.1602.01783.

Parveen, A. (2021). A Personalized Deep Learning Approach for Blood Glucose Prediction in People with T1DM (Master's thesis, Stevens Institute of Technology).

Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal Policy Optimization Algorithms. arXiv preprint arXiv:1707.06347.

Singh, R., & Raj R. R. (2023). Optimizing Glycemic Control in Type 1 Diabetic Patients using a Deep Learning-Based Artificial Pancreas with a Secure Glucagon and Insulin Delivery System. bioRxiv, 12. doi: https://doi.org/10.1101/2023.12.07.566476.

Tuomas, H., Aurick, Z. Pieter, A. & Sergey, L. (2018). Soft Actor-Critic: Off - policy Maximum Entropy Deep reinforcement Learning with a Stochastic Actor. International Journal of Research and Innovation in Social Sciences, doi:10.48550/arXiv.1801.01290, https://www.researchgate.net/publication/322306636_Soft_Actor-Critic_Off-policy_Maximum_Entropy_Deep_Reinforcement_Learning_with_a_Stochastic_Actor

Zhao, X., Ding, S., An, Y., & Jia, W. (2019). Applications of asynchronous deep reinforcement learning based on dynamic updating weights: X. Zhao et al. Applied Intelligence, 49(2), 581-591.

Zheng, M., Zhang, J., Zhan, C., Ren, X., & Lü, S. (2025). Proximal policy optimization with reward-based prioritization. Expert Systems with Applications, 283, 127659.

Article Details

How to Cite

Reinforcement Learning for Personalized Insulin Dosing: A Comparative Study of A2C, SAC and PPO on Real-World Clinical Data. (2026). International Journal of Latest Technology in Engineering Management & Applied Science, 15(6), 3700-3710. https://doi.org/10.51583/IJLTEMAS.2026.150600273