00
Days
00
Hrs
00
Min
00
Sec
Submit Your Paper

Bridging Legal Language Barriers Using Explainable AI: Outcome Prediction and Multilingual Knowledge based answer retrieval for Indian Law

Authors

Manish Thirunavu D

School of Computer Science Engineering , Vellore Institute of Technology, Chennai (IN)

Article Information

DOI: 10.51583/IJLTEMAS.2026.150600109

Subject Category: Knowledge

Volume/Issue: 15/6 | Page No: 1564-1582

Publication Timeline

Submitted: 2026-07-16

Published: 2026-07-16

Abstract

This study presents an integrated legal AI platform that combines interpretable case outcome prediction with multilingual, retrieval-grounded legal question answering to improve access to Indian law. The work is motivated by the difficulty ordinary citizens face in understanding legal language, the scarcity of trustworthy guidance, and the need for tools that work across India’s major languages. To address this, the authors built two connected components: a prediction module for Supreme Court case outcomes and a question-answering module based on statutory retrieval and generation. For prediction, they compiled 26,688 Indian Supreme Court judgments from 1950 to 2024 and represented each case using TF-IDF text features, case-type encodings, and temporal metadata, then trained an interpretable logistic regression model. For legal QA, they indexed 21 Indian legal acts with sentence-transformer embeddings and FAISS, and used a locally hosted Mistral model to generate simplified answers grounded in retrieved legal passages. The system was designed for English, Hindi, and Tamil, with translation, speech input, speech output, and interactive visualizations to make legal information more accessible. The prediction model achieved 91.3% accuracy and 0.919 ROC-AUC, while confidence calibration showed a strong correlation between predicted and actual accuracy. In the QA module, the system reached 78.4% precision@5, 86% answer correctness, and only 7% hallucination, a substantial improvement over baseline generative approaches. User evaluation with 35 participants reported 4.05/5 overall satisfaction, with multilingual support and explainability among the most valued features. Overall, the study concludes that transparent machine learning, retrieval-augmented generation, and multilingual interfaces can work together to build a practical and trustworthy legal assistance system for Indian users.

Keywords

Explainable artificial intelligence, Indian Supreme Court, Legal judgment prediction, Logistic regression

Downloads

References

1. Chalkidis, I., Androutsopoulos, I., & Aletras, N. (2019). Neural legal judgment prediction in English. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 4317–4323. [Google Scholar] [Crossref]

2. Goel, R., et al. (2022). LexGLUE: A benchmark dataset for legal language understanding in English. Proceedings of the Neural Information Processing Systems Datasets and Benchmarks Track. https://arxiv.org/abs/2110.00976 [Google Scholar] [Crossref]

3. Kalamkar, P., Bhattacharya, P., Ghosh, K., & Dey, P. (2022). ILDC for CJPE: Indian Legal Documents Corpus for Court Judgment Prediction and Explanation. Findings of the Association for Computational Linguistics (ACL Findings). https://arxiv.org/abs/2105.13562 [Google Scholar] [Crossref]

4. Lewis, P., Perez, E., Piktus, A., et al. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459–9474. [Google Scholar] [Crossref]

5. Karpukhin, V., Oğuz, B., Min, S., et al. (2020). Dense passage retrieval for open-domain question answering. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, 6769–6781. [Google Scholar] [Crossref]

6. Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using Siamese BERT-networks. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, 3982–3992. [Google Scholar] [Crossref]

7. Johnson, J., Douze, M., & Jégou, H. (2019). Billion-scale similarity search with GPUs. IEEE Transactions on Big Data, 7(3), 535–547. [Google Scholar] [Crossref]

8. Rudin, C. (2019). Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence, 1(5), 206–215. [Google Scholar] [Crossref]

9. Miller, T. (2019). Explanation in artificial intelligence: Insights from the social sciences. Artificial Intelligence, 267, 1–38. [Google Scholar] [Crossref]

10. Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of NAACL-HLT 2019, 4171–4186. [Google Scholar] [Crossref]

11. Zheng, L., Guha, N., Anderson, B., et al. (2024). LegalBench-RAG: A benchmark for retrieval-augmented generation in the legal domain. https://arxiv.org/abs/2408.10343 [Google Scholar] [Crossref]

12. Guha, N., Nyarko, J., Ho, D. E., et al. (2023). LegalBench: A collaboratively built benchmark for measuring legal reasoning in large language models. Advances in Neural Information Processing Systems. https://arxiv.org/abs/2308.11462 [Google Scholar] [Crossref]

13. Bhattacharya, P., Ghosh, K., Pal, A., Mehta, P., Bhattacharya, A., & Majumder, P. (2019). A comparative study of summarization algorithms applied to legal case judgments. Lecture Notes in Computer Science, 11882, 413–428. [Google Scholar] [Crossref]

14. Zhong, H., Guo, Z., Tu, C., Xiao, C., Liu, Z., & Sun, M. (2020). How does NLP benefit legal system: A summary of legal artificial intelligence. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 5218–5230. [Google Scholar] [Crossref]

15. Bommasani, R., Hudson, D. A., Adeli, E., et al. (2021). On the opportunities and risks of foundation models. Stanford Center for Research on Foundation Models. https://arxiv.org/abs/2108.07258 [Google Scholar] [Crossref]

16. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30, 5998–6008. [Google Scholar] [Crossref]

Metrics

Views & Downloads

Similar Articles

© 2026 IJLTEMAS · RSIS International. All rights reserved. ISSN 2278-2540.