00
Days
00
Hrs
00
Min
00
Sec
Submit Your Paper

"Impact-X: A Causally-Grounded Interpretable Multimodal Deep Learning Framework for Transparent Early Disease Detection Using Imaging, Clinical, and Genomic Data."

Authors

Tunan Shikder Any

Department of Electronics Engineering KIIT University, Bhubaneswar, Odisha, India (IN)

Ananya Manna

Department of Computer Science Engineering KIIT University, Bhubaneswar, Odisha, India (IN)

MD SARWAR ISLAM

Department of Computer Science Engineering KIIT University, Bhubaneswar, Odisha, India (IN)

Anu Priya Yaduvanshi

Department of Computer Science Engineering KIIT University, Bhubaneswar, Odisha, India (IN)

Shreyanjan Neogi

Department of Computer Science Engineering KIIT University, Bhubaneswar, Odisha, India (IN)

Addita Rani Dash

Department of Computer Science Engineering KIIT University, Bhubaneswar, Odisha, India (IN)

Turjoy Saha

Department of Computer Science Engineering KIIT University, Bhubaneswar, Odisha, India (IN)

Article Information

DOI: 10.51583/IJLTEMAS.2026.150600149

Subject Category: Deep Learning

Volume/Issue: 15/6 | Page No: 2064-2086

Publication Timeline

Submitted: 2026-07-17

Published: 2026-07-17

Abstract

Recent developments in multimodal deep learning have brought great progress to early disease detection; yet, wide-scale implementation of such models in clinics is hindered by the inherently inscrutable reasoning of existing methods. Current frameworks often employ post-hoc explanations that are not cross-modal consistent and are unable to disambiguate between causality and correlation, compromising both clinician trust and patient safety. In order to resolve these key issues, we introduce IMPACT-X, a novel Causally-Grounded Interpretable Multimodal Deep Learning Framework. IMPACT-X fuses mul-tiple heterogeneous modalities—medical imaging with Vision Transformers, medical records with Tabular Transformers, and genetic sequences with Graph Neural Networks—into a single and interpretable model.


Our framework includes a novel Causal Multimodal Fusion Layer (CMFL) which leverages cross-modal attention alignment in order to align the representation in a dynamic manner. Fur-thermore, an SCM module with DAG learning capabilities helps identify latent confounders and ensures the causally-consistent nature of the predictions. An uncertainty-aware decision-making layer estimates epistemic uncertainty through Monte Carlo Dropout in order to produce confidence scores. A unique cross-modal interpretability alignment loss function ensures coherent explanations across multiple modalities. The experimental results show that IMPACT-X achieves an SOTA performance with AUC-ROC score of 0.94, beating the best black-box baseline by 5.2%. Quantitative evaluation shows that IMPACT-X is 40% better in terms of faithfulness than traditional attention mechanism-based explanation approaches. A qualitative study with practicing medical professionals shows the benefits of causality-grounded predictions by increasing the level of physician trust in the system output. With its combination of high prediction accuracy and causal interpretability, IMPACT-X can pave the way for the development of a regulatory compliant and interpretable paradigm of medical AI that can safely be implemented in clinics, while enabling more accurate personalized medicine practices.Index Terms—Multimodal Deep Learning; Causal Inference; Interpretability; Early Disease Detection; Clinical Decision Sup-port; Genomic Integration

Keywords

Causally-Grounded, Interpretable, Deep Learning

Downloads

References

1. Esteva et al., “A guide to deep learning in healthcare,” Nature Medicine, vol. 25, no. 1, pp. 24–29, 2019. [Google Scholar] [Crossref]

2. Dosovitskiy et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” in Int. Conf. Learning Representations (ICLR), 2021. [Google Scholar] [Crossref]

3. K. Huang, J. Altosaar, and R. Ranganath, “ClinicalBERT: Modeling clin-ical notes and predicting hospital readmission,” Journal of Biomedical Informatics, vol. 112, p. 103609, 2020. [Google Scholar] [Crossref]

4. M. Zitnik, M. Agrawal, and J. Leskovec, “Modeling polypharmacy side effects with graph convolutional networks,” Bioinformatics, vol. 34, no. 13, pp. i457–i466, 2018. [Google Scholar] [Crossref]

5. J. Pearl and D. Mackenzie, The Book of Why: The New Science of Cause and Effect. Basic Books, 2018. [Google Scholar] [Crossref]

6. X. Zheng, B. Aragam, P. K. Ravikumar, and E. P. Xing, “DAGs with NO TEARS: Continuous optimization for structure learning,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 31, pp. 9472–9483, 2018. [Google Scholar] [Crossref]

7. S. M. Lundberg and S. I. Lee, “A unified approach to interpreting model predictions,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 30, pp. 4765–4774, 2017. [Google Scholar] [Crossref]

8. M. T. Ribeiro, S. Singh, and C. Guestrin, “‘Why should I trust you?’: Explaining the predictions of any classifier,” in Proc. 22nd ACM SIGKDD, pp. 1135–1144, 2016. [Google Scholar] [Crossref]

9. R. R. Selvaraju et al., “Grad-CAM: Visual explanations from deep networks via gradient-based localization,” in Proc. IEEE ICCV, pp. 618–626, 2017. [Google Scholar] [Crossref]

10. Y. H. H. Tsai et al., “Multimodal transformer for unaligned multimodal language sequences,” in Proc. 57th Annual Meeting of the ACL, pp. 6558–6569, 2019. [Google Scholar] [Crossref]

11. Y. Gal and Z. Ghahramani, “Dropout as a Bayesian approximation: Representing model uncertainty in deep learning,” in Proc. 33rd ICML, vol. 48, pp. 1050–1059, 2016. [Google Scholar] [Crossref]

12. N. Rieke et al., “The future of digital health with federated learning,” NPJ Digital Medicine, vol. 3, no. 1, p. 119, 2020. [Google Scholar] [Crossref]

13. E. J. Topol, “High-performance medicine: the convergence of human and artificial intelligence,” Nature Medicine, vol. 25, no. 1, pp. 44–56, 2019. [Google Scholar] [Crossref]

14. G. Litjens et al., “A survey on deep learning in medical image analysis,” Medical Image Analysis, vol. 42, pp. 60–88, 2017. [Google Scholar] [Crossref]

15. Shickel, P. J. Tighe, A. Bihorac, and P. Rashidi, “Deep EHR: A survey of recent advances in deep learning techniques for EHR analysis,” IEEE Journal of Biomedical and Health Informatics, vol. 22, no. 5, pp. 1589–1604, 2018. [Google Scholar] [Crossref]

16. J. Zhou et al., “Graph neural networks: A review of methods and applications,” AI Open, vol. 1, pp. 57–81, 2021. [Google Scholar] [Crossref]

17. J. Peters, D. Janzing, and B. Scho¨lkopf, Elements of Causal Inference: Foundations and Learning Algorithms. MIT Press, 2017. [Google Scholar] [Crossref]

18. K. Yu, S. Budhathoki, and B. Scho¨lkopf, “Causal discovery and infer-ence: concepts and recent methodological advances,” Applied Informat-ics, vol. 3, no. 1, pp. 1–28, 2021. [Google Scholar] [Crossref]

19. W. Samek, T. Wiegand, and K. R. Mu¨ller, “Explainable artificial intelligence: Understanding, visualizing and interpreting deep learning models,” arXiv preprint arXiv:1708.08296, 2017. [Google Scholar] [Crossref]

20. B. Arrieta et al., “Explainable Artificial Intelligence (XAI): Con-cepts, taxonomies, opportunities and challenges toward responsible AI,” Information Fusion, vol. 58, pp. 82–115, 2020. [Google Scholar] [Crossref]

21. J. Vickers and E. B. Elkin, “Decision curve analysis: a novel method for evaluating prediction models,” Medical Decision Making, vol. 26, no. 6, pp. 565–574, 2006. [Google Scholar] [Crossref]

22. R. J. Chen et al., “Whole slide images are 2D point clouds: Context-aware survival prediction using patch-based graph convolutional net-works,” in MICCAI, pp. 339–349, 2021. [Google Scholar] [Crossref]

23. N. K. Tomar et al., “MMIT: Multi-modal medical image transformer for computer-aided diagnosis,” IEEE Journal of Biomedical and Health Informatics, vol. 27, no. 4, pp. 1895–1906, 2023. [Google Scholar] [Crossref]

24. Y. Yang et al., “Causal inference in healthcare: A review of methods and applications,” Journal of Biomedical Informatics, vol. 134, p. 104201, 2022. [Google Scholar] [Crossref]

25. M. Chen, S. Radhakrishnan, and F. Doshi-Velez, “Learning causal representations for robust domain adaptation,” in Proc. CLeaR, vol. 172, [Google Scholar] [Crossref]

26. pp. 156–182, 2022. [Google Scholar] [Crossref]

27. T. N. Kipf and M. Welling, “Semi-supervised classification with graph convolutional networks,” in Int. Conf. Learning Representations (ICLR), 2017. [Google Scholar] [Crossref]

28. Vaswani et al., “Attention is all you need,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 30, pp. 5998–6008, 2017. [Google Scholar] [Crossref]

29. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in Int. Conf. Learning Representations (ICLR), 2019. [Google Scholar] [Crossref]

30. J. Kelly et al., “Key challenges for delivering clinical impact with artificial intelligence,” BMC Medicine, vol. 17, no. 1, p. 195, 2019. [Google Scholar] [Crossref]

31. Amann et al., “Explainability for artificial intelligence in healthcare: a multidisciplinary perspective,” BMC Medical Informatics and Decision Making, vol. 20, no. 1, p. 310, 2020. [Google Scholar] [Crossref]

Metrics

Views & Downloads

Similar Articles

© 2026 IJLTEMAS · RSIS International. All rights reserved. ISSN 2278-2540.