00
Days
00
Hrs
00
Min
00
Sec
Submit Your Paper

Hybrid Transformer–LSTM Forecasting and PPO-Based Intelligent Kubernetes Auto-Scaling Framework for Cloud Data Science Pipelines Using Predictive Workload Analytics and Multi-Objective Performance Optimization

Authors

Tunan Shikder Any

Department of Electronics Engineering, KIIT University, Bhubaneswar, Odisha, India (IN)

Prosanjit Gupta

Department of Computer Science Engineering, KIIT University, Bhubaneswar, Odisha, India (IN)

Shrishti Sharan

Department of Computer Science Engineering, KIIT University, Bhubaneswar, Odisha, India (IN)

Sagnik Koner

Department of Computer Science Engineering, KIIT University, Bhubaneswar, Odisha, India (IN)

Anirban Bhatta

Department of Computer Science Engineering (AI & ML), KIIT University, Bhubaneswar, Odisha, India (IN)

Turjoy Saha

Department of Computer Science Engineering, KIIT University, Bhubaneswar, Odisha, India (IN)

Shreyanjan Neogi

Department of Computer Science Engineering, KIIT University, Bhubaneswar, Odisha, India (IN)

Addita Rani Dash

Department of Computer Science Engineering, KIIT University, Bhubaneswar, Odisha, India (IN)

Article Information

DOI: 10.51583/IJLTEMAS.2026.150600064

Subject Category: Hybrid Transformer–LSTM

Volume/Issue: 15/6 | Page No: 847-867

Publication Timeline

Submitted: 2026-07-06

Published: 2026-07-06

Abstract

With the rise of cloud-native data science pipelines, dynamic infrastructure is needed to cope with the variable nature of AI/ML workloads, however, traditional auto-scalers based on Kubernetes use reactive threshold mechanisms leading to inefficient resource utilization and service level agreement (SLA) violations. In this paper, we introduce a novel prediction-based framework with an integrated intelligent auto-scaler based on a Hybrid Transformer-LSTM forecasting model and a Proximal Policy Optimization (PPO) reinforcement learning (RL) agent. The forecasting engine leverages historical time series data on performance metrics, such as CPU, memory, GPU utilization, and request latency, to provide accurate workload predictions. These predictions are used by the RL agent to determine optimal scaling decisions, pod placement and resource allocation strategies based on multiple objectives including minimizing latency, maximizing throughput, energy savings and lowering cloud expenditures. In extensive evaluations performed using multiple node clusters performing distributed deep learning, stream processing and real-time inference pipelines, the proposed solution outperforms conventional HPA/VPA and baseline ML scalers by offering improved accuracy, speed of scaling, reducing unnecessary allo-cations by 38% and meeting strict SLA requirements for bursty workloads.

Keywords

Kubernetes; Intelligent Auto-Scaling; Hybrid Transformer–LSTM; Proximal Policy Optimization; Cloud-Native AI; Multi-Objective Optimization

Downloads

References

1. Kubernetes Authors, “Horizontal and vertical pod autoscaling,” Kubernetes Official Documentation, 2024. [Online]. Available: https://kubernetes.io/docs/tasks/run-application/horizontal-pod-autoscale/ [Google Scholar] [Crossref]

2. J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Prox-imal policy optimization algorithms,” arXiv preprint arXiv:1707.06347, 2017. [Google Scholar] [Crossref]

3. A. Vaswani et al., “Attention is all you need,” in Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 30, 2017, pp. 5998–6008. [Google Scholar] [Crossref]

4. S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Comput., vol. 9, no. 8, pp. 1735–1780, Nov. 1997. [Google Scholar] [Crossref]

5. R. Gupta, M. Arora, and S. K. Panda, “Predictive cloud resource management using hybrid deep learning models,” IEEE Trans. Cloud Comput., vol. 11, no. 2, pp. 1452–1465, Apr.–Jun. 2023. [Google Scholar] [Crossref]

6. H. Mao, M. Alizadeh, I. Menache, and S. Kandula, “Resource manage-ment with deep reinforcement learning,” in Proc. 15th ACM Workshop Hot Topics Netw., Atlanta, GA, USA, 2016, pp. 50–56. [Google Scholar] [Crossref]

7. Y. Chen, L. Zhang, and X. Wang, “AI-driven Kubernetes scheduling for heterogeneous AI workloads,” Future Gener. Comput. Syst., vol. 134, [Google Scholar] [Crossref]

8. pp. 112–125, Sep. 2022. [Google Scholar] [Crossref]

9. Kubeflow Authors, “Kubeflow: Portable, scalable machine learn-ing on Kubernetes,” GitHub Repository, 2024. [Online]. Available: https://github.com/kubeflow/kubeflow [Google Scholar] [Crossref]

10. D. Sculley et al., “Hidden technical debt in machine learning systems,” in Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 28, 2015, pp. 2503–2511. [Google Scholar] [Crossref]

11. T. P. Lillicrap et al., “Continuous control with deep reinforcement learning,” arXiv preprint arXiv:1509.02971, 2015. [Google Scholar] [Crossref]

12. M. Liao, J. Liu, and H. Wang, “Energy-aware container scheduling in cloud data centers using reinforcement learning,” J. Syst. Archit., vol. 138, p. 102745, Jul. 2023. [Google Scholar] [Crossref]

13. Amazon Web Services, “Amazon EKS best practices guide,” AWS Docu-mentation, 2024. [Online]. Available: https://aws.github.io/aws-eks-best-practices/ [Google Scholar] [Crossref]

14. C. Amato, F. S. Melo, and M. Eger, “Reinforcement learning for dynamic resource allocation in cloud environments,” IEEE Trans. Serv. Comput., vol. 15, no. 4, pp. 2103–2116, Jul.–Aug. 2022. [Google Scholar] [Crossref]

15. D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” arXiv preprint arXiv:1409.0473, 2014. [Google Scholar] [Crossref]

16. M. Abadi et al., “TensorFlow: A system for large-scale machine learning,” in Proc. 12th USENIX Conf. Oper. Syst. Design Implement., Savannah, GA, USA, 2016, pp. 265–283. [Google Scholar] [Crossref]

17. MLflow Contributors, “MLflow documentation: Experiment tracking and model deployment,” 2024. [Online]. Available: https://www.mlflow.org/docs/latest/index.html [Google Scholar] [Crossref]

18. P. Stone, R. S. Sutton, and G. Kuhlmann, “Reinforcement learning for autonomous systems: A comprehensive survey,” IEEE Trans. Robot., vol. 38, no. 1, pp. 1–18, Feb. 2022. [Google Scholar] [Crossref]

19. K. Li, G. Xu, Y. Dai, and K. Hara, “Time series forecasting with Transformer-LSTM hybrid networks for cloud workload prediction,” J. Cloud Comput., vol. 12, no. 3, p. 45, Mar. 2023. [Google Scholar] [Crossref]

20. J. Dean and L. A. Barroso, “The tail at scale,” Commun. ACM, vol. 56, no. 2, pp. 74–80, Feb. 2013. [Google Scholar] [Crossref]

21. S. Dey, A. Patel, and R. Kumar, “Autonomous cloud orchestration: A survey of AI-driven Kubernetes management,” Future Gener. Comput. Syst., vol. 145, pp. 312–330, Aug. 2023. [Google Scholar] [Crossref]

Metrics

Views & Downloads

Similar Articles

© 2026 IJLTEMAS · RSIS International. All rights reserved. ISSN 2278-2540.