00
Days
00
Hrs
00
Min
00
Sec
Submit Your Paper

HyperNova++: A Novel Adaptive Activation Function for High-Accuracy Neural Learning on Nonlinear Synthetic Decision Manifolds

Authors

Sourish Dey

KIIT University, Bhubaneswar, Odisha, India (IN)

Sunil Kumar Sawant

KIIT University, Bhubaneswar, Odisha, India (IN)

Arunima Dutta

KIIT University, Bhubaneswar, Odisha, India (IN)

Abhradeep Hazra

KIIT University, Bhubaneswar, Odisha, India (IN)

Article Information

DOI: 10.51583/IJLTEMAS.2025.1412000109

Subject Category: Artificial Intelligence and Machine Learning

Volume/Issue: 14/12 | Page No: 1228-1252

Publication Timeline

Submitted: 2026-01-11

Published: 2026-01-10

Abstract

Activation functions are at the heart of how deep neural networks perform non-linear transformations. The use of an activation function allows a neural network to approximate highly complex functions, train using a gradient-based optimization technique and generalize to new data. However, existing activation functions, such as ReLU, GELU, and Swish, have limitations that restrict their use in practice. Specifically, they can saturate gradients during training due to their inherent structure, cause vanishing gradients on deeply stacked architectures, and are inefficient at learning periodic dependency relationships while performing poorly at modeling highly heterogeneous non-linear interactions. These limitations are of particular importance for scientific, financial, and engineering use cases where data represent polynomial, periodic, saturating, and exponential shapes on the same data manifold.


This paper introduces HyperNova++, a smooth, adaptive, parameterized activation function that unifies bounded saturation, periodic oscillation, and unbounded growth into a single learnable formula. HyperNova++ is architectured and designed to overcome the expressive constraints of existing activations which enables dynamic, data-driven modulation of curvature, frequency, and growth behavior using three trainable parameters (α,β,γ). These above mentioned parameters respectively govern contributions from the hyperbolic tangent (tanh) for bounded saturation, sine (sin) for periodic oscillations, and Softplus (log(1+ex)) for getting a smooth monotonic growth all thorughout. The resulting function obtained ensures non-vanishing gradients, smooth transitions, and controlled Lipschitz continuity, along with maintaining computational efficiency comparable to contemporary activations and other counterparts.


After doing a rigorous, large-scale evaluation on a meticulously crafted synthetic dataset with a known ground-truth decision boundary that stimulates reall life linear, polynomial, and periodic interactions. This controlled environment enables precise, unbiased comparisons against various functions including ReLU, GELU, and Swish under identical architectural, optimization, and hyperparameter settings. HyperNova++ achieves statistically significant superior performance compared to all , exceeding 99% accuracy (0.9903) compared to 98.34% for ReLU, 98.08% for GELU, and 97.60% for Swish,  while also attaining the highest F1-score (0.9906) and ROC-AUC (0.9997). Gradient analyses obtained confirm stable, non-vanishing gradients and accelerated convergence.


We supplement empirical results obtained during testing with comprehensive theoretical analysis, thus establishing HyperNova++’s universal approximation guarantee, Lipschitz properties, gradient bounds, and optimization landscape characteristics. Practical implementation guidelines, computational complexity dissections, and prospective applications in scientific machine learning, time-series analysis, and multimodal inference are being discussed. Collectively, this work positions HyperNova++ as a potent, versatile activation function for advanced deep learning architectures confronting intricate nonlinear manifolds in upcoming future.

Keywords

Activation Function, Deep Learning, Hyper-Nova++, Neural Networks, Nonlinear Modeling Synthetic Dataset, ROC-AUC Curve, Optimization, Adaptive Activation, Mixed Nonlinearities, Universal Approximation, Lipschitz Continuity

Downloads

References

1. G. Cybenko, “Approximation by superpositions of a sigmoidal function,” Mathematics of Control, Signals and Systems, vol. 2, no. 4, pp. 303–314, 1989. [Google Scholar] [Crossref]

2. I. Goodfellow, D. Warde-Farley, M. Mirza, A. Courville, and Y. Bengio, “Maxout networks,” in International Conference on Machine Learning, 2013, pp. 1319–1327. [Google Scholar] [Crossref]

3. K. He, X. Zhang, S. Ren, and J. Sun, “Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification,” in [Google Scholar] [Crossref]

4. Proceedings of the IEEE International Conference on Computer Vision, 2015, pp. 1026–1034. [Google Scholar] [Crossref]

5. D. Hendrycks and K. Gimpel, “Gaussian error linear units (GELUs),” arXiv preprint arXiv:1606.08415, 2016. [Google Scholar] [Crossref]

6. K. Hornik, “Approximation capabilities of multilayer feedforward networks,” Neural Networks, vol. 4, no. 2, pp. 251–257, 1991. [Google Scholar] [Crossref]

7. S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in International Conference on Machine Learning, 2015, pp. 448–456. [Google Scholar] [Crossref]

8. G. Klambauer, T. Unterthiner, A. Mayr, and S. Hochreiter, “Selfnormalizing neural networks,” in Advances in Neural Information Processing Systems, 2017, pp. 971–980. [Google Scholar] [Crossref]

9. M. Leshno, V. Y. Lin, A. Pinkus, and S. Schocken, “Multilayer feedforward networks with a nonpolynomial activation function can approximate any function,” Neural Networks, vol. 6, no. 6, pp. 861–867, 1993. [Google Scholar] [Crossref]

10. I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” arXiv preprint arXiv:1711.05101, 2017. [Google Scholar] [Crossref]

11. Z. Lu, H. Pu, F. Wang, Z. Hu, and L. Wang, “The expressive power of neural networks: A view from the width,” in Advances in Neural Information Processing Systems, 2017, pp. 6231–6239. [Google Scholar] [Crossref]

12. A. L. Maas, A. Y. Hannun, and A. Y. Ng, “Rectifier nonlinearities improve neural network acoustic models,” in Proceedings of the 30th International Conference on Machine Learning, vol. 30, no. 1, 2013, p.3. [Google Scholar] [Crossref]

13. G. F. Montufar, R. Pascanu, K. Cho, and Y. Bengio, “On the number of linear regions of deep neural networks,” in Advances in Neural Information Processing Systems, 2014, pp. 2924–2932. [Google Scholar] [Crossref]

14. P. Ramachandran, B. Zoph, and Q. V. Le, “Searching for activation functions,” arXiv preprint arXiv:1710.05941, 2017. [Google Scholar] [Crossref]

15. V. Sitzmann, J. Martel, A. Bergman, D. Lindell, and G. Wetzstein, “Implicit neural representations with periodic activation functions,” in Advances in Neural Information Processing Systems, 2020, pp. 7462–7473. [Google Scholar] [Crossref]

16. M. Tancik, P. P. Srinivasan, B. Mildenhall, S. Fridovich-Keil, N. Raghavan, U. Singhal, R. Ramamoorthi, J. T. Barron, and R. Ng, “Fourier features let networks learn high frequency functions in low dimensional domains,” in Advances in Neural Information Processing Systems, 2020, pp. 7537–7547. [Google Scholar] [Crossref]

17. B. Xu, N. Wang, T. Chen, and M. Li, “Empirical evaluation of rectified activations in convolutional network,” arXiv preprint arXiv:1505.00853, 2015. [Google Scholar] [Crossref]

18. D. Yarotsky, “Error bounds for approximations with deep ReLU networks,” Neural Networks, vol. 94, pp. 103–114, 2017. [Google Scholar] [Crossref]

19. D.-A. Clevert, T. Unterthiner, and S. Hochreiter, “Fast and accurate deep network learning by exponential linear units (ELUs),” arXiv preprint arXiv:1511.07289, 2015. [Google Scholar] [Crossref]

20. L. B. Godfrey and M. S. Gashler, “Adaptive blending units: Trainable activation functions for deep neural networks,” Neurocomputing, vol. 398, pp. 1–8, 2020. [Google Scholar] [Crossref]

Metrics

Views & Downloads

Similar Articles

© 2026 IJLTEMAS · RSIS International. All rights reserved. ISSN 2278-2540.