O‘ZBEK TILIDAGI MATNLARNI VEKTORLASHTIRISH USULLARINING SEMANTIK O‘XSHASHLIKNI ANIQLASHDAGI SAMARADORLIGI

Авторы

  • Allaberganova Nasiba Toshkent amaliy fanlar universiteti, o‘qituvchisi, Автор

DOI:

https://doi.org/10.65164/2pypxe85

Ключевые слова:

matn vektorlashtirish, semantik o‘xshashlik, TF-IDF, Word2Vec, FastText, BERT Transformer modellar, tabiiy tilni qayta ishlash, o‘zbek tili korpusi, benchmark, lemmatizatsiya.

Аннотация

Mazkur maqolada o‘zbek tilidagi matnlarni sonli ifodalash (vektorlashtirish) usullarining samaradorligi qiyosiy tahlil qilindi. TF-IDF, Word2Vec, FastText hamda BERT arxitekturasi asosidagi Transformer modellari o‘zaro taqqoslandi. Eksperimental tadqiqotlar 193000 dan ortiq hujjatdan iborat maxsus matn korpusi hamda 4200 ta belgilangan (annotatsiyalangan) matn juftliklaridan tashkil topgan baholash to‘plami asosida amalga oshirildi. UzBERT modeli F1 = 0,92 ko‘rsatkichiga erishib, tadqiqotda ko‘rib chiqilgan boshqa usullarga nisbatan yaxshi natija qayd etdi. 

Библиографические ссылки

1. Manning C.D., Raghavan P., Schütze H. Introduction to Information Retrieval. -

Cambridge: Cambridge University Press, 2008. - 482 p.

2. Turney P.D., Pantel P. From frequency to meaning: Vector space models of semantics //

Journal of Artificial Intelligence Research. - 2010. - Vol. 37. - P. 141-188.

3. Joshi P., Santy S., Budhiraja A., Bali K., Choudhury M. The state and fate of linguistic

diversity in the NLP world // Proceedings of the 58th Annual Meeting of the Association

for Computational Linguistics (ACL). - 2020. - P. 6282-6293.

4. Tantuğ A.C., Adalı E., Oflazer K. A machine translation system from Uzbek to Turkish //

Proceedings of the 11th Conference of the European Association for Machine Translation

(EAMT). - Oslo, Norway, 2006. - P. 161-168.

5. Joulin A., Grave E., Bojanowski P., Mikolov T. Bag of tricks for efficient text classification

// Proceedings of the 15th Conference of the European Chapter of the Association for

Computational Linguistics (EACL). - Valencia, Spain, 2017. - P. 427-431.

6. Devlin J., Chang M.W., Lee K., Toutanova K. BERT: Pre-training of deep bidirectional

transformers for language understanding // Proceedings of the 2019 Conference of the

North American Chapter of the Association for Computational Linguistics (NAACL-HLT).

- Minneapolis, USA, 2019. - P. 4171-4186.

7. Conneau A., Khandelwal K., Goyal N., et al. Unsupervised cross-lingual representation

learning at scale // Proceedings of the 58th Annual Meeting of the Association for

Computational Linguistics (ACL). - 2020. - P. 8440-8451.

8. Mansurov B., Muminov B., Sobirov O., Kuriyozov E. Morphological analysis for Uzbek

using Hunspell // Proceedings of the ACL Student Research Workshop. - 2020. - P. 36-43.

9. Uzbek NLP. UzBERT: BERT model for Uzbek language [Elektron resurs]. - 2021. - URL:

https://github.com/UlugbekSobirov/uzbert (murojaat sanasi: 10.03.2026).

10. Allaberganova N.M. UzSemCorpus benchmark to‘plami. - Toshkent: Toshkent axborot

texnologiyalari universiteti repozitoriysi, 2024. - 1 elektron resurs.

11. Vajjala S., Majumder B., Gupta A., Surana H. Practical Natural Language Processing. -

Sebastopol, CA: O’Reilly Media, 2022. - 592 p.

Опубликован

2026-05-15