Measuring Level of Agreement in Students' Academic Performance Comparing Kappa and Scott Pi Agreement Measures.
Authors
Onyenekwe Chukwuenyem Enoch
Department of Statistics, Nnamdi Azikiwe University, PMB 5025, Awka, Nigeria. (NG)
Adejumo Olusola Adebowale
Department of Statistics, University of Ilorin, PMB 1515, Ilorin, Nigeria. (NG)
Article Information
DOI: 10.51583/IJLTEMAS.2025.1407000023
Subject Category: Applied Science
Volume/Issue: 14/7 | Page No: 215-220
Publication Timeline
Submitted: 2025-08-01
Published: 2025-08-01
Abstract
Abstract: Kappa and Scott Pi statistics are agreement measures that are used to measure the level of agreement between two raters. The academic performance of students used is the Cumulative Grade Point Average (CGPA) of year one and the final year. Year One is classified as rater one and that of the final year is classified as rater two. The Analysis was done to see the consistency in the students’ performance, if the level of agreement is high, i.e. if there is agreement between their first-year result and that of the final year. A comparison of the Kappa Estimates and that of the Scot Pi was also done to determine the most efficient measure between the two. Kappa Agreement outperforms the Scot pi measures in terms of precision and efficiency. Also, there is no agreement between the Year One Performance and that of the Final Year
Keywords
Kappa, Scott Pi, Proportion, Contingency table
Downloads
References
1. Adejumo A.O (2010). A data based approach of imputing missing ratings in raters agreement measurement for two raters. ICASTOR Journal of mathematical Sciences. Vol 4,No 149-67 [Google Scholar] [Crossref]
2. Adejumo A.O, Heumann C, Toutenburg H.(2004) A review of agreement measure as a subset of association measure between raters, sonderforschungsbereich [Google Scholar] [Crossref]
3. Albert, J. H., Chib, S. (1993). Bayesian analysis of binary and polychotomous response data. J. Amer. Statist. Assoc. 88:669–679. [Google Scholar] [Crossref]
4. Alberto M. Marchevsky, Ann E. Walts, Birgit I. Lissenberg-White, Erk Thunnissen (2020) Pathologists should probably forget about kappa. Percent agreement, diagnostic specificity and related metrics provide more clinically applicable measures of interobserver variability. Annals of Diagnostic Pathology Volume 47, August 2020, 151561 [Google Scholar] [Crossref]
5. Banerjee, M., Capozzoli, M., McSweeney, L., Sinha, D. (1999). Beyond kappa: a review of interrater agreement measures. Can. J. Statist. 27:3–23. [Google Scholar] [Crossref]
6. Barlow, W. (1996). Measurement of interrater agreement with adjustment for covariates. Biometrics 52:695–702. [Google Scholar] [Crossref]
7. Basu, S., Nanerjee, M., Sen, A. (2000). Bayesian inference for kappa from single and multiple studies. Biometrics 56:577–582. [Google Scholar] [Crossref]
8. Bloch, D. A., Kraemer, H. C. (1989). 2 × 2 kappa coefficients: measures of agreement or association. Biometrics 45:269–287. [Google Scholar] [Crossref]
9. Bond, M. E., Higgins, J. J. (2001). A note on ‘a comparison of Bayes and maximum likelihood estimation of the intraclass correlation coefficient’. Commun. Statist. Theor. Meth. 30:371–380. [Google Scholar] [Crossref]
10. Chib, S., Greenberg, E. (1998). Analysis of multivariate probit models. Biometrika 85: 347–361. [Google Scholar] [Crossref]
11. Cohen, J. (1960). A coefficient of agreement for nominal scales. Educat. Psychol. Measure. 20:37–46. Cohen, J. (1968). Weighted kappa: nominal scale agreement with provision for scaled or partial credit. Psychol. Bull. 70:213–220. [Google Scholar] [Crossref]
12. D. Chicco, M. J. Matthijs, G. Jurman. (2021). MCC versus Cohen’s Kappa and Brier score. .DOI 10.1109/ACCESS.2021.3084050, IEEE Access [Google Scholar] [Crossref]
13. Drasgow, F. (1988). Polychoric and polyserial correlations. In: Kotz, L., Johnson, N. L., eds. Encyclopedia of Statistical Sciences. Vol. 7. New York: Wiley, pp. 69–74. [Google Scholar] [Crossref]
14. Farzan Madadizadeh ,Hesam Ghafari and Sajjad Bahariniya (2023) Kappa Statistics: A Method of Measuring Agreement in Dental Examinations. The Open Public Health Journal [Google Scholar] [Crossref]
15. Fleiss, J. L. (1971). Measuring nominal scale agreement among many raters. Psychol. Bull. 76:378–382. Hale, C. A., Fleiss, J. L. (1993). Interval estimation under two study designs for kappa with binary classifications. Biometrics 49:523–534. [Google Scholar] [Crossref]
16. Jonas Moss (2024). Measures of Agreement with Multiple Raters: Fréchet Variances and Inference. Psychometrika , Volume 89 , Issue 2 , June 2024 , pp. 517 - 541 [Google Scholar] [Crossref]
17. DOI: https://doi.org/10.1007/s11336-023-09945-2 [Google Scholar] [Crossref]
18. Kraemer, H. C., Periyakoil, V. S., Noda, A. (2002). Kappa coefficients in medical research. Statist. Med. 21:2109–2129. [Google Scholar] [Crossref]
19. Lee, J. J., Tu, Z. N. (1994). A better confidence interval for kappa on measuring agreement between two raters with binary outcomes. J. Computat. Graph. Statist. 3:301–321. [Google Scholar] [Crossref]
20. Lipsitz, S. R., Laird, N. M., Brennan, T. A. (1994). Simple moment estimates of the kappa coefficient and its variance. Appl. Statist. 43:309–323. [Google Scholar] [Crossref]
21. Liu, C. (2001). Bayesian analysis of multivariate probit model: discussion of “the art of dataaugmentation” by Van Dyk and Meng. J. Computat. Graph. Statist. 10:75–81. [Google Scholar] [Crossref]
22. Nam, J. (2000). Interval estimation of the kappa coefficient with binary classification and an equal marginal probability model. Biometrics 56:583–585. [Google Scholar] [Crossref]
23. Nandram, B., Chen, M.-H. (1994). Accelerating Gibbs sampler convergence in the generalized linear models via a reparameterization. J. Statist. Computat. Simul. 81: 27–40. [Google Scholar] [Crossref]
24. Ochi, Y., Prentice, R. L. (1984). Likelihood inference in a correlated probit regression model. Biometrika 71:531–543. [Google Scholar] [Crossref]
25. Pan, W., Wall, M. M. (2002). Small-sample adjustments in using the sandwich variance estimator in generalized estimating equations. Statist. Med. 21:1429–1441. [Google Scholar] [Crossref]
26. Pearson, K. (1900). Mathematical contributions to the theory of evolution. vii. on the correlation of characters not quantitatively measurable. [Google Scholar] [Crossref]
27. Zee, M. D., Mariani, L., Barisoni, L., Mahajan, S., & Gillespie, M. (2024). A novel agreement statistic using data on uncertainty in ratings. Journal of the Royal Statistical Society: Series C (Applied Statistics), 72(5), 1293–1314. https://doi.org/10.1093/jrsssc/qnad055 [Google Scholar] [Crossref]
Metrics
Views & Downloads
Similar Articles
- Advanced Digital Communication Strategies: A Comprehensive Analysis
- Developing an Explainable AI System for Digital Forensics: Enhancing Trust and Transparency in Flagging Events for Legal Evidence
- The Global Impact of Government Censorship on Women’s Access to Information: A Literature Review
- "Reimagining Higher Education Workspaces: A Review on The Transformative Role of Digital Technology Adoption"
- Improved CSP Efficiency: Innovations and Challenges in Thermal Energy Storage Systems