Hybrid Transformer–LSTM Forecasting and PPO-Based Intelligent Kubernetes Auto-Scaling Framework for Cloud Data Science Pipelines Using Predictive Workload Analytics and Multi-Objective Performance Optimization
Authors
Tunan Shikder Any
Department of Electronics Engineering, KIIT University, Bhubaneswar, Odisha, India (IN)
Prosanjit Gupta
Department of Computer Science Engineering, KIIT University, Bhubaneswar, Odisha, India (IN)
Shrishti Sharan
Department of Computer Science Engineering, KIIT University, Bhubaneswar, Odisha, India (IN)
Sagnik Koner
Department of Computer Science Engineering, KIIT University, Bhubaneswar, Odisha, India (IN)
Anirban Bhatta
Department of Computer Science Engineering (AI & ML), KIIT University, Bhubaneswar, Odisha, India (IN)
Turjoy Saha
Department of Computer Science Engineering, KIIT University, Bhubaneswar, Odisha, India (IN)
Shreyanjan Neogi
Department of Computer Science Engineering, KIIT University, Bhubaneswar, Odisha, India (IN)
Addita Rani Dash
Department of Computer Science Engineering, KIIT University, Bhubaneswar, Odisha, India (IN)
Article Information
DOI: 10.51583/IJLTEMAS.2026.150600064
Subject Category: Hybrid Transformer–LSTM
Volume/Issue: 15/6 | Page No: 847-867
Publication Timeline
Submitted: 2026-07-06
Published: 2026-07-06
Abstract
With the rise of cloud-native data science pipelines, dynamic infrastructure is needed to cope with the variable nature of AI/ML workloads, however, traditional auto-scalers based on Kubernetes use reactive threshold mechanisms leading to inefficient resource utilization and service level agreement (SLA) violations. In this paper, we introduce a novel prediction-based framework with an integrated intelligent auto-scaler based on a Hybrid Transformer-LSTM forecasting model and a Proximal Policy Optimization (PPO) reinforcement learning (RL) agent. The forecasting engine leverages historical time series data on performance metrics, such as CPU, memory, GPU utilization, and request latency, to provide accurate workload predictions. These predictions are used by the RL agent to determine optimal scaling decisions, pod placement and resource allocation strategies based on multiple objectives including minimizing latency, maximizing throughput, energy savings and lowering cloud expenditures. In extensive evaluations performed using multiple node clusters performing distributed deep learning, stream processing and real-time inference pipelines, the proposed solution outperforms conventional HPA/VPA and baseline ML scalers by offering improved accuracy, speed of scaling, reducing unnecessary allo-cations by 38% and meeting strict SLA requirements for bursty workloads.
Keywords
Kubernetes; Intelligent Auto-Scaling; Hybrid Transformer–LSTM; Proximal Policy Optimization; Cloud-Native AI; Multi-Objective Optimization
Downloads
References
1. Kubernetes Authors, “Horizontal and vertical pod autoscaling,” Kubernetes Official Documentation, 2024. [Online]. Available: https://kubernetes.io/docs/tasks/run-application/horizontal-pod-autoscale/ [Google Scholar] [Crossref]
2. J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Prox-imal policy optimization algorithms,” arXiv preprint arXiv:1707.06347, 2017. [Google Scholar] [Crossref]
3. A. Vaswani et al., “Attention is all you need,” in Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 30, 2017, pp. 5998–6008. [Google Scholar] [Crossref]
4. S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Comput., vol. 9, no. 8, pp. 1735–1780, Nov. 1997. [Google Scholar] [Crossref]
5. R. Gupta, M. Arora, and S. K. Panda, “Predictive cloud resource management using hybrid deep learning models,” IEEE Trans. Cloud Comput., vol. 11, no. 2, pp. 1452–1465, Apr.–Jun. 2023. [Google Scholar] [Crossref]
6. H. Mao, M. Alizadeh, I. Menache, and S. Kandula, “Resource manage-ment with deep reinforcement learning,” in Proc. 15th ACM Workshop Hot Topics Netw., Atlanta, GA, USA, 2016, pp. 50–56. [Google Scholar] [Crossref]
7. Y. Chen, L. Zhang, and X. Wang, “AI-driven Kubernetes scheduling for heterogeneous AI workloads,” Future Gener. Comput. Syst., vol. 134, [Google Scholar] [Crossref]
8. pp. 112–125, Sep. 2022. [Google Scholar] [Crossref]
9. Kubeflow Authors, “Kubeflow: Portable, scalable machine learn-ing on Kubernetes,” GitHub Repository, 2024. [Online]. Available: https://github.com/kubeflow/kubeflow [Google Scholar] [Crossref]
10. D. Sculley et al., “Hidden technical debt in machine learning systems,” in Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 28, 2015, pp. 2503–2511. [Google Scholar] [Crossref]
11. T. P. Lillicrap et al., “Continuous control with deep reinforcement learning,” arXiv preprint arXiv:1509.02971, 2015. [Google Scholar] [Crossref]
12. M. Liao, J. Liu, and H. Wang, “Energy-aware container scheduling in cloud data centers using reinforcement learning,” J. Syst. Archit., vol. 138, p. 102745, Jul. 2023. [Google Scholar] [Crossref]
13. Amazon Web Services, “Amazon EKS best practices guide,” AWS Docu-mentation, 2024. [Online]. Available: https://aws.github.io/aws-eks-best-practices/ [Google Scholar] [Crossref]
14. C. Amato, F. S. Melo, and M. Eger, “Reinforcement learning for dynamic resource allocation in cloud environments,” IEEE Trans. Serv. Comput., vol. 15, no. 4, pp. 2103–2116, Jul.–Aug. 2022. [Google Scholar] [Crossref]
15. D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” arXiv preprint arXiv:1409.0473, 2014. [Google Scholar] [Crossref]
16. M. Abadi et al., “TensorFlow: A system for large-scale machine learning,” in Proc. 12th USENIX Conf. Oper. Syst. Design Implement., Savannah, GA, USA, 2016, pp. 265–283. [Google Scholar] [Crossref]
17. MLflow Contributors, “MLflow documentation: Experiment tracking and model deployment,” 2024. [Online]. Available: https://www.mlflow.org/docs/latest/index.html [Google Scholar] [Crossref]
18. P. Stone, R. S. Sutton, and G. Kuhlmann, “Reinforcement learning for autonomous systems: A comprehensive survey,” IEEE Trans. Robot., vol. 38, no. 1, pp. 1–18, Feb. 2022. [Google Scholar] [Crossref]
19. K. Li, G. Xu, Y. Dai, and K. Hara, “Time series forecasting with Transformer-LSTM hybrid networks for cloud workload prediction,” J. Cloud Comput., vol. 12, no. 3, p. 45, Mar. 2023. [Google Scholar] [Crossref]
20. J. Dean and L. A. Barroso, “The tail at scale,” Commun. ACM, vol. 56, no. 2, pp. 74–80, Feb. 2013. [Google Scholar] [Crossref]
21. S. Dey, A. Patel, and R. Kumar, “Autonomous cloud orchestration: A survey of AI-driven Kubernetes management,” Future Gener. Comput. Syst., vol. 145, pp. 312–330, Aug. 2023. [Google Scholar] [Crossref]
Metrics
Views & Downloads
Similar Articles
- Enhancing Formation Control of Multi Agent Systems Using Ann Based Technique
- Improving Sliding Mode Control with Chattering Reduction using Fuzzy Based Technique
- Cooking Quality, Fasting Blood Glucose, Glycemic Index and Load of High–Fiber Noodles Made from Wheat, Tiger Nut Residue and Cassava Flour Blends
- Matrix Rhythm Therapy Versus Interferential Therapy Combined with Lumbar Stabilization Exercises in Chronic Non-Specific Low Back Pain: A Randomized Comparative Trial
- Formulation and Sensory Evaluation of Functional Cake Prepared from Sweet Potato Powder