O‘ZBEK TILIDAGI MATNLARNI VEKTORLASHTIRISH USULLARINING SEMANTIK O‘XSHASHLIKNI ANIQLASHDAGI SAMARADORLIGI
DOI:
https://doi.org/10.65164/2pypxe85Kalit so‘zlar:
matn vektorlashtirish, semantik o‘xshashlik, TF-IDF, Word2Vec, FastText, BERT Transformer modellar, tabiiy tilni qayta ishlash, o‘zbek tili korpusi, benchmark, lemmatizatsiya.Abstrak
Mazkur maqolada o‘zbek tilidagi matnlarni sonli ifodalash (vektorlashtirish) usullarining samaradorligi qiyosiy tahlil qilindi. TF-IDF, Word2Vec, FastText hamda BERT arxitekturasi asosidagi Transformer modellari o‘zaro taqqoslandi. Eksperimental tadqiqotlar 193000 dan ortiq hujjatdan iborat maxsus matn korpusi hamda 4200 ta belgilangan (annotatsiyalangan) matn juftliklaridan tashkil topgan baholash to‘plami asosida amalga oshirildi. UzBERT modeli F1 = 0,92 ko‘rsatkichiga erishib, tadqiqotda ko‘rib chiqilgan boshqa usullarga nisbatan yaxshi natija qayd etdi.
Havolalar
1. Manning C.D., Raghavan P., Schütze H. Introduction to Information Retrieval. -
Cambridge: Cambridge University Press, 2008. - 482 p.
2. Turney P.D., Pantel P. From frequency to meaning: Vector space models of semantics //
Journal of Artificial Intelligence Research. - 2010. - Vol. 37. - P. 141-188.
3. Joshi P., Santy S., Budhiraja A., Bali K., Choudhury M. The state and fate of linguistic
diversity in the NLP world // Proceedings of the 58th Annual Meeting of the Association
for Computational Linguistics (ACL). - 2020. - P. 6282-6293.
4. Tantuğ A.C., Adalı E., Oflazer K. A machine translation system from Uzbek to Turkish //
Proceedings of the 11th Conference of the European Association for Machine Translation
(EAMT). - Oslo, Norway, 2006. - P. 161-168.
5. Joulin A., Grave E., Bojanowski P., Mikolov T. Bag of tricks for efficient text classification
// Proceedings of the 15th Conference of the European Chapter of the Association for
Computational Linguistics (EACL). - Valencia, Spain, 2017. - P. 427-431.
6. Devlin J., Chang M.W., Lee K., Toutanova K. BERT: Pre-training of deep bidirectional
transformers for language understanding // Proceedings of the 2019 Conference of the
North American Chapter of the Association for Computational Linguistics (NAACL-HLT).
- Minneapolis, USA, 2019. - P. 4171-4186.
7. Conneau A., Khandelwal K., Goyal N., et al. Unsupervised cross-lingual representation
learning at scale // Proceedings of the 58th Annual Meeting of the Association for
Computational Linguistics (ACL). - 2020. - P. 8440-8451.
8. Mansurov B., Muminov B., Sobirov O., Kuriyozov E. Morphological analysis for Uzbek
using Hunspell // Proceedings of the ACL Student Research Workshop. - 2020. - P. 36-43.
9. Uzbek NLP. UzBERT: BERT model for Uzbek language [Elektron resurs]. - 2021. - URL:
https://github.com/UlugbekSobirov/uzbert (murojaat sanasi: 10.03.2026).
10. Allaberganova N.M. UzSemCorpus benchmark to‘plami. - Toshkent: Toshkent axborot
texnologiyalari universiteti repozitoriysi, 2024. - 1 elektron resurs.
11. Vajjala S., Majumder B., Gupta A., Surana H. Practical Natural Language Processing. -
Sebastopol, CA: O’Reilly Media, 2022. - 592 p.