eSummarizer AI Service- Document Summarization Model Using BART
Authors
Rongdeep Pathak
Department of Computer Applications, Jorhat Engineering College, India (IN)
Mriganka Mohan Bora
Department of Computer Applications, Assam Engineering College, India (IN)
Nelson R Varte
Department of Computer Applications, Assam Engineering College, India (IN)
Article Information
DOI: 10.51583/IJLTEMAS.2025.1411000093
Subject Category: MACHINE LEARNING
Volume/Issue: 14/11 | Page No: 965-970
Publication Timeline
Submitted: 2025-12-19
Published: 2025-12-19
Abstract
The exponential growth of governmental documentation in India has created an urgent need for efficient automated summarization systems. This research presents eSummarizer AI Service, a custom Transformer-based machine learning model designed specifically for summarizing Indian government documents such as policy papers, circulars, legislative texts, and departmental reports. The study employs advanced Natural Language Processing (NLP) techniques, including BART (Bidirectional and Auto-Regressive Transformer), reinforcement learning, and a custom dataset curated from official government portals. Evaluation using ROUGE metrics demonstrates significant improvements over existing baseline models, achieving high coherence, contextual relevance, and factual consistency. The proposed system has practical applications for policymakers, researchers, and citizens, enhancing the accessibility and comprehension of complex governmental information.
Keywords
NLP, ROUGE, BART, Streamlit
Downloads
References
1. Babar S, Tech-Cse M, Rit, ”Text Summarization: An Overview”.2013 [Google Scholar] [Crossref]
2. Christian H, Agus M, Suhartono D,”Single Document Automatic Text Summarization using Term Frequency-Inverse Document Frequency (TF-IDF)”, ComTech: Computer, Mathematics and EngineeringApplications 2016. [Google Scholar] [Crossref]
3. Nomoto T ”Bayesian Learning in Text Summarization Models” ,2005 [Google Scholar] [Crossref]
4. Graves A ,”Generating Sequences With Recurrent Neural Net works”, CoRR abs/1308.0850:2013 [Google Scholar] [Crossref]
5. Nallapati R, Xiang B, Zhou B , ”Sequence-to-Sequence RNNs for Text Summarization”, CoRR abs/1602.06023: ,2016 [Google Scholar] [Crossref]
6. Hochreiter S, Schmidhuber J , ”Long Short-Term Memory”. Neural Comput 9:1735–1780, https://doi.org/10.1162/neco.1997.9.8.1735, 1997 [Google Scholar] [Crossref]
7. Shi T, Keneshloo Y, Ramakrishnan N, Reddy CK , ”Neural Abstractive Text Summarization with Sequence-to-Sequence Models”, CoRR abs/1812.02303:, 2018 [Google Scholar] [Crossref]
8. Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser L, Polosukhin I , ”Attention is All you Need”. ArXiv abs/1706.03762:,2017 [Google Scholar] [Crossref]
9. Devlin J, Chang M-W, Lee K, Toutanova K BERT: ”Pre-training of Deep Bidirectional Transformers for Language Understanding”, 2019,CoRR abs/1810.04805 [Google Scholar] [Crossref]
10. Zhang J, Zhao Y, Saleh M, Liu PJ , ”PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization”, 2019, CoRR abs/1912.08777 [Google Scholar] [Crossref]
11. Dong L, Yang N, Wang W, Wei F, Liu X, Wang Y, Gao J, Zhou M, Hon H-W, ”Unified Language ModelPre-training for Natural Language Understanding and Generation”, 2019 CoRRabs/1905.03197 [Google Scholar] [Crossref]
12. Radford A , ”Improving Language Understanding by Generative Pre-Training”,2018. [Google Scholar] [Crossref]
13. Wolf T, Debut L, Sanh V, Chaumond J, Delangue C, Moi A, Cistac P, Rault T, Louf R, Funtowicz M, ”HuggingFace’s Transformers: State-of-the-art Natural Language Processing”, 2019 CoRR abs/1910.03771. [Google Scholar] [Crossref]
14. Lin, C. Y. .” ROUGE: A Package for Automatic Evaluation of Summaries.” pages 74–81, 2004 Association for Computational Linguistics. [Google Scholar] [Crossref]
15. Balaji N, Karthik Pai B H, Bhaskar Bhat B, Praveen Barmavatu, ”Data Visualization in Splunk and Tableau: A Case Study Demonstration”, in Journal of Physics: Conference Series, 2021. [Google Scholar] [Crossref]
16. Lhoest Q, del Moral AV, Jernite Y, Thakur A, von Platen P, Patil S, Chaumond J, Drame M, Plu J, Tunstall L, Davison J. “Datasets: A community library for natural language” arXiv arXiv:2109.02846. 2021 [Google Scholar] [Crossref]
17. M.Lewis, Y Lin, N Goyal . “BART: Denoising Sequence-to-Sequence Pretraining for Natural Language Generation.” ACL, 2020. [Google Scholar] [Crossref]
Metrics
Views & Downloads
Similar Articles
- "Towards GDP to GEP-Centric Model: A Proposed GEP Index Framework and Its Application in Haryana"
- Physicochemical Analysis of Petroleum Products in Selected Depots in Calabar Metropolis Cross River State and The Effects on Motor Engine
- Impact of Employee Welfare Measures on Workforce Well-Being in SBI: Insights from the Public Banking Sector
- Problem of Small-Scale Farmers in Agricutlure Sector in Tirunelveli Taluk
- Globalization, Economic Development, and Ecological Footprint in Tunisia: A QARDL Approach