Transformer Models for Hausa Personal Name Spelling Correction: A Comparative Evaluation of Byt5, Flan-T5, and AfroLlama
Authors
Department of Computing Sciences, Admiralty University of Nigeria, Ibusa, Delta State, Nigeria/Department of Computer Science, Ebonyi State University Abakaliki, Ebonyi State, Nigeria (Nigeria)
Department of Computer Science, Ebonyi State University Abakaliki, Ebonyi State, Nigeria (Nigeria)
Article Information
DOI: 10.51583/IJLTEMAS.2026.150800074
Subject Category: Education
Volume/Issue: 15/8 | Page No: 1020-1039
Publication Timeline
Submitted: 2026-08-28
Accepted: 2026-09-02
Published: 2026-09-15
Abstract
Misspelling of personal names presents a significant challenge to data integrity, record linkage, identity verification, and information retrieval, particularly in low-resource linguistic environments. Hausa personal names exhibit substantial spelling variation arising from phonetic, orthographic, transliteration, and typographical differences, yet limited research has investigated Transformer-based approaches for their automatic correction. This study investigates the effectiveness of Transformer-based language models for automatically correcting misspelt Hausa personal names. A corpus comprising 148,710 spelling instances derived from 2,061 Hausa personal names was used, including 19,472 real human spelling attempts and synthetically generated variants. Three models, ByT5-small, Flan-T5-small, and AfroLlama, were evaluated under a three-phase cumulative curriculum-learning framework in which progressively larger edit-distance ranges were introduced during training. To reduce the possibility of canonical-name-level data leakage, the 2,061 canonical names were partitioned into mutually exclusive training, validation, and test sets containing 1,690, 185, and 186 canonical names, respectively. Model performance was evaluated using exact-match accuracy, character-level precision, recall and F1, SacreBLEU, and character error rate (CER), with McNemar's test used to assess differences in paired prediction outcomes. Fine-tuned ByT5-small achieved the strongest overall performance, correctly restoring 3,328 of 13,558 test instances (24.55%), with character precision of 92.45%, recall of 80.68%, F1 of 86.16%, SacreBLEU of 69.7286, and CER of 0.2048. Fine-tuned Flan-T5-small achieved 11.85% exact-match accuracy and 83.14% character F1. In contrast, AfroLlama declined from 6.84% to 4.46% exact-match accuracy following fine-tuning, accompanied by a reduction in character F1 from 74.93% to 69.70%. McNemar's tests indicated statistically significant differences between each pair of evaluated models (p < 0.001). The findings demonstrate that, under the experimental configuration examined, ByT5-small provided the strongest correction performance among the evaluated models. However, the results should not be interpreted as establishing the superiority of an entire model class. The study also highlights the need for human-versus-synthetic evaluation, curriculum ablation, architecture-specific optimization, and broader baseline comparisons in future research.
Keywords
ByT5; Flan-T5; Low-resource; Spelling correction; Transformer models.
Downloads
References
1. Aars, C., Adams, L., Tian, X., Wang, Z., Wismer, C., Wu, J., Rivas, P., Sooksatra, K., & Fendt, M. (2024). Efficacy of ByT5 in Multilingual Translation of Biblical Texts for Underrepresented Languages. In arXiv (Cornell University). Cornell University. https://doi.org/10.48550/arxiv.2405.13350 [Google Scholar] [Crossref]
2. Aliakbarzadeh, A., Flek, L., & Karimi, A. (2025). Exploring Robustness of Multilingual LLMs on Real-World Noisy Data. In Qeios. https://doi.org/10.32388/3x6cxv [Google Scholar] [Crossref]
3. Aliero, A. A., Bashir, S. A., Aliyu, H. O., Tafida, A. G., & Hussaini, M. (2025). Dual-BERT Adversarial Model for Text Normalization in Hausa User-Generated Contents. In Research Square. https://doi.org/10.21203/rs.3.rs-7446019/v1 [Google Scholar] [Crossref]
4. Al-Rfooh, B., Abandah, G. A., & Al‐Rfou, R. (2023). Fine-Tashkeel: Finetuning Byte-Level Models for Accurate Arabic Text Diacritization. In arXiv (Cornell University). Cornell University. https://doi.org/10.48550/arxiv.2303.14588 [Google Scholar] [Crossref]
5. Bernard, E., & Ajah, A. I. (2025). Sautex: A Language-Specific Phonetic Matching Algorithm for Resolving Spelling Variations in Hausa Personal Names. Journal of Computer, Software and Program., 2(2), 25–33. https://doi.org/10.69739/jcsp.v2i2.1141 [Google Scholar] [Crossref]
6. Christen, P. (2006). A Comparison of Personal Name Matching: Techniques and Practical Issues. The Austrilian National University. 290–294. https://doi.org/10.1109/icdmw.2006.2 [Google Scholar] [Crossref]
7. Feher, D., Vulić, I., & Minixhofer, B. (2024). Retrofitting Large Language Models with Dynamic Tokenization. In arXiv (Cornell University). Cornell University. https://doi.org/10.48550/arxiv.2411.18553 [Google Scholar] [Crossref]
8. Ismail, K., Abdou, S., Farouk, M., & Salem, A. (2025). Transformers to the rescue: alleviating data scarcity in arabic grammatical error correction with pre-trained models. Neural Computing and Applications, 37(18), 13011–13038. https://doi.org/10.1007/s00521-025-11145-1 [Google Scholar] [Crossref]
9. Khaleel, M. R., & Abandah, G. A. (2025). Efficient Stochastic Error Injection for Optimizing Large Language Models in Arabic Spelling Correction. 2025 International Conference on New Trends in Computing Sciences (ICTCS), Amman, Jordan . 505–510. https://doi.org/10.1109/ictcs65341.2025.10989319 [Google Scholar] [Crossref]
10. Kuparinen, O., Miletić, A., & Scherrer, Y. (2023). Dialect-to-Standard Normalization: A Large-Scale Multilingual Evaluation. Findings of the Association for Computational Linguistics: EMNLP 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.923 [Google Scholar] [Crossref]
11. Lauc, D., Rutherford, A., & Wongwarawipatr, W. (2024). AyutthayaAlpha: A Thai-Latin Script Transliteration Transformer. In arXiv (Cornell University). Cornell University. https://doi.org/10.48550/arxiv.2412.03877 [Google Scholar] [Crossref]
12. Lefrandt, M., Santoso, E. B., Gunawan, A. A. S., & Tedjasulaksana, J. J. (2025). Contextual Spelling Corrector for Indonesian Text Preprocessing: A Comparative Analysis of Large Language Models. 2025 IEEE International Conference on Industry 4.0, Artificial Intelligence, and Communications Technology (IAICT), Bali, Indonesia. 290–296. https://doi.org/10.1109/iaict65714.2025.11100636 [Google Scholar] [Crossref]
13. Li, H., Li, J., Jiang, W., Zhang, Z., Chen, M., Wang, S., & Xiao, J. (2021). PHMOSpell: Phonological and Morphological Knowledge Guided Chinese Spelling Check. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 5958–5967. https://doi.org/10.18653/v1/2021.acl-long.464 [Google Scholar] [Crossref]
14. Limisiewicz, T., Blevins, T., Gonen, H., Ahia, O., & Zettlemoyer, L. (2024). MYTE: Morphology-Driven Byte Encoding for Better and Fairer Multilingual Language Modeling. In arXiv (Cornell University). Cornell University. https://doi.org/10.48550/arxiv.2403.10691 [Google Scholar] [Crossref]
15. Lutgen, A.-M., Plum, A., Purschke, C., & Plank, B. (2024). Neural Text Normalization for Luxembourgish using Real-Life Variation Data. In arXiv (Cornell University). Cornell University. https://doi.org/10.48550/arxiv.2412.09383 [Google Scholar] [Crossref]
16. Okewunmi, P., James, F., & Fajemila, O. E. (2025). Evaluating Robustness of LLMs to Typographical Noise in Yorùbá QA. Proceedings of the Sixth Workshop on African Natural Language Processing (AfricaNLP 2025. 195–202. https://doi.org/10.18653/v1/2025.africanlp-1.29 [Google Scholar] [Crossref]
17. Owodunni, A. T., Ahia, O., & Kumar, S. (2025). FLEXITOKENS: Flexible Tokenization for Evolving Language Models. In ArXiv.org. https://doi.org/10.48550/arxiv.2507.12720 [Google Scholar] [Crossref]
18. Phonchai, T., Siripong, S., Patterson, N., & Campbell, O. (2025). Large Language Models for Zero-Shot Multicultural Name Recognition. arXiv (Cornell University). http://arxiv.org/abs/2507.04149 [Google Scholar] [Crossref]
19. S, B. R., Suri, G., Dewangan, V., & Sonavane, R. (2024). When Every Token Counts: Optimal Segmentation for Low-Resource Language Models. In arXiv (Cornell University). Cornell University. https://doi.org/10.48550/arxiv.2412.06926 [Google Scholar] [Crossref]
20. Sani, S. A., Muhammad, S. H., & Jarvis, D. (2025). Investigating the Impact of Language-Adaptive Fine-Tuning on Sentiment Analysis in Hausa Language Using AfriBERTa. In ArXiv.org. https://doi.org/10.48550/arxiv.2501.11023 [Google Scholar] [Crossref]
21. Sperduti, G., & Moreo, A. (2026). Misspellings in natural language processing: A survey of recent literature. Natural Language Processing., 32(2), 113–159. https://doi.org/10.1017/nlp.2026.10020 [Google Scholar] [Crossref]
22. Wali, A. M., & Nisioi, S. (2025). Automatic Correction of Writing Anomalies in Hausa Texts. In ArXiv.org. https://doi.org/10.48550/arxiv.2506.03820 [Google Scholar] [Crossref]
23. Wu, S., Tan, X., Wang, Z., Wang, R., Li, X., & Sun, M. (2024). Beyond Language Models: Byte Models are Digital World Simulators. In arXiv (Cornell University). Cornell University. https://doi.org/10.48550/arxiv.2402.19155 [Google Scholar] [Crossref]
24. Xue, L., Barua, A., Constant, N., Al‐Rfou, R., Narang, S., Kale, M., Roberts, A. P., & Raffel, C. (2022). ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models. Transactions of the Association for Computational Linguistics, 10, 291–306. https://doi.org/10.1162/tacl_a_00461 [Google Scholar] [Crossref]
25. Xue, L., Barua, A., Constant, N., Al‐Rfou, R., Narang, S., Kale, M., Roberts, A., & Raffel, C. (2021). ByT5: Towards a token-free future with pre-trained byte-to-byte models. In arXiv (Cornell University). Cornell University. https://doi.org/10.48550/arxiv.2105.13626 [Google Scholar] [Crossref]
26. Zanga, A. I., Abdulrahman, S. M., Ado, A., Bichi, A. A., Jibril, L. A., Umar, A. M., Adamu, A., Muhammad, S. H., & Abubakar, B. S. (2025). HausaMovieReview: A Benchmark Dataset for Sentiment Analysis in Low-Resource African Language. In ArXiv.org. https://doi.org/10.48550/arxiv.2509.16256 [Google Scholar] [Crossref]
Metrics
Views & Downloads
Similar Articles
- A Study to Assess the Impact of a Nurse-Led Educational Intervention on Knowledge Regarding Oral Health Among Primigravida in Selected Hospitals of Navi Mumbai
- Attitude towards Mathematics and Science in Relation to STEM Career Aspirations among Senior Secondary School Students
- AI-Driven Personalized Learning in Educational Systems: A Framework for Adaptive Learning and Decision Support
- Institutionalizing Indigenous Peoples Education in Philippine Basic Education: Development of the Integrated Institutionalization Framework for Indigenous Peoples Education (IIF-IPEd)
- Toward an Integrated Theory of Enterprise Risk Management: A Multi-Theoretical Conceptual Framework