00
Days
00
Hrs
00
Min
00
Sec
Submit Your Paper

Measuring Level of Agreement in Students' Academic Performance Comparing Kappa and Scott Pi Agreement Measures.

Authors

Onyenekwe Chukwuenyem Enoch

Department of Statistics, Nnamdi Azikiwe University, PMB 5025, Awka, Nigeria. (NG)

Adejumo Olusola Adebowale

Department of Statistics, University of Ilorin, PMB 1515, Ilorin, Nigeria. (NG)

Article Information

DOI: 10.51583/IJLTEMAS.2025.1407000023

Subject Category: Applied Science

Volume/Issue: 14/7 | Page No: 215-220

Publication Timeline

Submitted: 2025-08-01

Published: 2025-08-01

Abstract

Abstract: Kappa and Scott Pi statistics are agreement measures that are used to measure the level of agreement between two raters. The academic performance of students used is the Cumulative Grade Point Average (CGPA) of year one and the final year.  Year One is classified as rater one and that of the final year is classified as rater two. The Analysis was done to see the consistency in the students’ performance, if the level of agreement is high, i.e. if there is agreement between their first-year result and that of the final year.  A comparison of the Kappa Estimates and that of the Scot Pi was also done to determine the most efficient measure between the two. Kappa Agreement outperforms the Scot pi measures in terms of precision and efficiency. Also, there is no agreement between the Year One Performance and that of the Final Year

Keywords

Kappa, Scott Pi, Proportion, Contingency table

Downloads

References

1. Adejumo A.O (2010). A data based approach of imputing missing ratings in raters agreement measurement for two raters. ICASTOR Journal of mathematical Sciences. Vol 4,No 149-67 [Google Scholar] [Crossref]

2. Adejumo A.O, Heumann C, Toutenburg H.(2004) A review of agreement measure as a subset of association measure between raters, sonderforschungsbereich [Google Scholar] [Crossref]

3. Albert, J. H., Chib, S. (1993). Bayesian analysis of binary and polychotomous response data. J. Amer. Statist. Assoc. 88:669–679. [Google Scholar] [Crossref]

4. Alberto M. Marchevsky, Ann E. Walts, Birgit I. Lissenberg-White, Erk Thunnissen (2020) Pathologists should probably forget about kappa. Percent agreement, diagnostic specificity and related metrics provide more clinically applicable measures of interobserver variability. Annals of Diagnostic Pathology Volume 47, August 2020, 151561 [Google Scholar] [Crossref]

5. Banerjee, M., Capozzoli, M., McSweeney, L., Sinha, D. (1999). Beyond kappa: a review of interrater agreement measures. Can. J. Statist. 27:3–23. [Google Scholar] [Crossref]

6. Barlow, W. (1996). Measurement of interrater agreement with adjustment for covariates. Biometrics 52:695–702. [Google Scholar] [Crossref]

7. Basu, S., Nanerjee, M., Sen, A. (2000). Bayesian inference for kappa from single and multiple studies. Biometrics 56:577–582. [Google Scholar] [Crossref]

8. Bloch, D. A., Kraemer, H. C. (1989). 2 × 2 kappa coefficients: measures of agreement or association. Biometrics 45:269–287. [Google Scholar] [Crossref]

9. Bond, M. E., Higgins, J. J. (2001). A note on ‘a comparison of Bayes and maximum likelihood estimation of the intraclass correlation coefficient’. Commun. Statist. Theor. Meth. 30:371–380. [Google Scholar] [Crossref]

10. Chib, S., Greenberg, E. (1998). Analysis of multivariate probit models. Biometrika 85: 347–361. [Google Scholar] [Crossref]

11. Cohen, J. (1960). A coefficient of agreement for nominal scales. Educat. Psychol. Measure. 20:37–46. Cohen, J. (1968). Weighted kappa: nominal scale agreement with provision for scaled or partial credit. Psychol. Bull. 70:213–220. [Google Scholar] [Crossref]

12. D. Chicco, M. J. Matthijs, G. Jurman. (2021). MCC versus Cohen’s Kappa and Brier score. .DOI 10.1109/ACCESS.2021.3084050, IEEE Access [Google Scholar] [Crossref]

13. Drasgow, F. (1988). Polychoric and polyserial correlations. In: Kotz, L., Johnson, N. L., eds. Encyclopedia of Statistical Sciences. Vol. 7. New York: Wiley, pp. 69–74. [Google Scholar] [Crossref]

14. Farzan Madadizadeh ,Hesam Ghafari and Sajjad Bahariniya (2023) Kappa Statistics: A Method of Measuring Agreement in Dental Examinations. The Open Public Health Journal [Google Scholar] [Crossref]

15. Fleiss, J. L. (1971). Measuring nominal scale agreement among many raters. Psychol. Bull. 76:378–382. Hale, C. A., Fleiss, J. L. (1993). Interval estimation under two study designs for kappa with binary classifications. Biometrics 49:523–534. [Google Scholar] [Crossref]

16. Jonas Moss (2024). Measures of Agreement with Multiple Raters: Fréchet Variances and Inference. Psychometrika , Volume 89 , Issue 2 , June 2024 , pp. 517 - 541 [Google Scholar] [Crossref]

17. DOI: https://doi.org/10.1007/s11336-023-09945-2 [Google Scholar] [Crossref]

18. Kraemer, H. C., Periyakoil, V. S., Noda, A. (2002). Kappa coefficients in medical research. Statist. Med. 21:2109–2129. [Google Scholar] [Crossref]

19. Lee, J. J., Tu, Z. N. (1994). A better confidence interval for kappa on measuring agreement between two raters with binary outcomes. J. Computat. Graph. Statist. 3:301–321. [Google Scholar] [Crossref]

20. Lipsitz, S. R., Laird, N. M., Brennan, T. A. (1994). Simple moment estimates of the kappa coefficient and its variance. Appl. Statist. 43:309–323. [Google Scholar] [Crossref]

21. Liu, C. (2001). Bayesian analysis of multivariate probit model: discussion of “the art of dataaugmentation” by Van Dyk and Meng. J. Computat. Graph. Statist. 10:75–81. [Google Scholar] [Crossref]

22. Nam, J. (2000). Interval estimation of the kappa coefficient with binary classification and an equal marginal probability model. Biometrics 56:583–585. [Google Scholar] [Crossref]

23. Nandram, B., Chen, M.-H. (1994). Accelerating Gibbs sampler convergence in the generalized linear models via a reparameterization. J. Statist. Computat. Simul. 81: 27–40. [Google Scholar] [Crossref]

24. Ochi, Y., Prentice, R. L. (1984). Likelihood inference in a correlated probit regression model. Biometrika 71:531–543. [Google Scholar] [Crossref]

25. Pan, W., Wall, M. M. (2002). Small-sample adjustments in using the sandwich variance estimator in generalized estimating equations. Statist. Med. 21:1429–1441. [Google Scholar] [Crossref]

26. Pearson, K. (1900). Mathematical contributions to the theory of evolution. vii. on the correlation of characters not quantitatively measurable. [Google Scholar] [Crossref]

27. Zee, M. D., Mariani, L., Barisoni, L., Mahajan, S., & Gillespie, M. (2024). A novel agreement statistic using data on uncertainty in ratings. Journal of the Royal Statistical Society: Series C (Applied Statistics), 72(5), 1293–1314. https://doi.org/10.1093/jrsssc/qnad055 [Google Scholar] [Crossref]

Metrics

Views & Downloads

Similar Articles

© 2026 IJLTEMAS · RSIS International. All rights reserved. ISSN 2278-2540.