SUN’IY INTELLEKT ASOSIDAGI AVTOMATLASHTIRILGAN BAHOLASH TIZIMLARI: ISHONCHLILIK, XOLISLIK VA ALGORITMIK NOXOLISLIK MUAMMOLARI
DOI:
https://doi.org/10.65164/rpw8as94Ключевые слова:
avtomatlashtirilgan baholash tizimlari, AES, sun'iy intellekt ta'limda, algoritmik noxolislik, ishonchlilik, xolislik, mashinali o'qitish, katta til modellari, XAI, psixometriya.Аннотация
Ushbu maqolada sun'iy intellekt (SI) asosidagi avtomatlashtirilgan baholash tizimlari (ABS/AES — Automated Essay/Assessment Scoring) ning ta'lim jarayonidagi qo'llanilishi, ularning ishonchlilik (reliability) va xolislik (fairness) ko'rsatkichlari, shuningdek algoritmik noxolislik (algorithmic bias) muammolari kompleks tarzda tahlil qilinadi. Katta til modellari (LLM) va mashinali o'qitish algoritmlariga asoslangan zamonaviy AES tizimlarining me'moriy yechimlari, psixometrik asoslari va ta'lim natijalariga ta'siri o'rganiladi. Algoritmik noxolislikning manbalari — demografik tafovut, o'quv ma'lumotlaridagi tarafkashlik va modelning cheklangan umumlashtirish qobiliyati — tahlil etilib, ularni minimallashtirishning texnik va metodologik yo'llari tavsiya etiladi. O'zbekiston oliy ta'lim tizimi uchun integratsion model taklif qilinadi.
Библиографические ссылки
1. Attali, Y., & Burstein, J. (2006). Automated essay scoring with e-rater® V.2. Journal of Technology, Learning, and Assessment, 4(3), 1–30. https://ejournals.bc.edu/index.php/jtla/article/view/1650
2. Barocas, S., Hardt, M., & Narayanan, A. (2023). Fairness and Machine Learning: Limitations and Opportunities. MIT Press. https://fairmlbook.org
3. Bridgeman, B., Trapani, C., & Attali, Y. (2012). Comparison of human and machine scoring of essays: Differences by gender, ethnicity, and country. Applied Measurement in Education, 25(1), 27–40. https://doi.org/10.1080/08957347.2012.635502
4. Chou, T., Murugadoss, K., Feng, J., & Bunt, A. (2023). Fairness in Automated Essay Scoring: A Comparative Analysis of Bias Mitigation Strategies. Proceedings of the 14th International Conference on Educational Data Mining (EDM 2023), 45–54.
5. Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37–46. https://doi.org/10.1177/001316446002000104
6. Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of NAACL-HLT 2019, 4171–4186. https://doi.org/10.18653/v1/N19-1423
7. Dikli, S. (2006). An overview of automated scoring of essays. Journal of Technology, Learning, and Assessment, 5(1). https://ejournals.bc.edu/index.php/jtla/article/view/1640
8. European Commission. (2024). AI Act: Regulation (EU) 2024/1689 of the European Parliament and of the Council on Artificial Intelligence. Official Journal of the European Union.
9. Ke, Z., & Ng, V. (2019). Automated essay scoring: A survey of the state of the art. Proceedings of the 28th International Joint Conference on Artificial Intelligence (IJCAI-2019), 6300–6308. https://doi.org/10.24963/ijcai.2019/879
10. Kendjayeva, D. X. (2023). Sun'iy intellektning iqtisodiy jarayonlar samaradorligini oshirishdagi roli. Toshkent amaliy fanlar universiteti ilmiy-amaliy anjumani materiallari to'plami. Toshkent: TATU nashriyoti, 142–148.
11. Loukina, A., Madnani, N., & Zechner, K. (2019). The many dimensions of algorithmic fairness in educational applications. Proceedings of the 14th Workshop on Innovative Use of NLP for Building Educational Applications, 1–10. https://doi.org/10.18653/v1/W19-4401
12. Lundberg, S. M., & Lee, S. I. (2017). A unified approach to interpreting model predictions. Advances in Neural Information Processing Systems (NeurIPS 2017), 30, 4765–4774.
13. Mansurov, B., & Mansurov, A. (2021). UzBERT: The first pre-trained language model for Uzbek. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP). https://aclanthology.org/2021.emnlp-main.655
14. Mayfield, E., Madaio, M., Prabhumoye, S., Gerritsen, D., McLaughlin, B., Dixon-Román, E., & Black, A. (2020). Equity beyond bias in language technologies for education. Proceedings of the 15th Workshop on Innovative Use of NLP for Building Educational Applications, 444–460. https://doi.org/10.18653/v1/2020.bea-1.34
15. O'zbekiston Respublikasi Prezidentining 2020-yil 5-oktyabrdagi "Raqamli O'zbekiston — 2030" strategiyasini tasdiqlash to'g'risidagi PF-6079-sonli Farmoni. https://lex.uz/docs/5031048
16. Page, E. B. (1966). The imminence of grading essays by computer. Phi Delta Kappan, 47(5), 238–243.
17. Perelman, L. (2012). Construct validity, length, score, and time in holistic large-scale writing assessments: The case against automated essay scoring (AES). International Journal of Learning, 18(10), 4-52.
695
18. Ramesh, D., & Sanampudi, S. K. (2022). An automated essay scoring systems: A systematic literature review. Artificial Intelligence Review, 55(3), 2495–2527. https://doi.org/10.1007/s10462-021-10068-2
19. UNESCO. (2022). K-12 AI curricula: A mapping of government-endorsed AI curricula. Paris: UNESCO. https://unesdoc.unesco.org/ark:/48223/pf0000380602
20. Wen, S., & Bastian, M. (2023). Gender bias in AI-based essay scoring: Evidence from PISA writing assessment data. Computers & Education: Artificial Intelligence, 4, 100120. https://doi.org/10.1016/j.caeai.2023.100120
21. Zhang, B. H., Lemoine, B., & Mitchell, M. (2018). Mitigating unwanted biases with adversarial learning. Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society, 335–340. https://doi.org/10.1145/3278721.3278779