THE IMPORTANCE OF THE PARALLEL CORPUS AS A LINGUISTIC BASE
Keywords:
computational linguistics; parallel corpus; artificial intelligence; natural language processing (NLP); Uzbek language corpus; machine translation; GPT; BERT; corpus linguistics; linguistic database; transformative model; language modeling; translation technologies.Abstract
This article is devoted to the analysis of research conducted in the field of computational linguistics worldwide and in Uzbekistan during the period from 2020 to 2025. The study covers theoretical frameworks of computational linguistics, artificial intelligence–based language models (GPT, BERT, LLaMA, and others), as well as the role of parallel corpora as a linguistic foundation. An analysis of international experience demonstrates that over the past five years, the priority areas in computational linguistics have included the training of linguistic models based on neural transformer architectures, the development of multilingual parallel corpora, and the improvement of machine translation systems.
In Uzbekistan, the field of computational linguistics is observed to be transitioning from a predominantly applied stage to a more scientific and analytical phase through projects such as the Uzbek language corpus (uzbekcorpus.uz) and Paratranslator. The research employs systematic analysis of scientific literature, comparative linguistic and corpus-based methods, as well as the analytical capabilities of artificial intelligence models.
References
Devlin, J., Chang, M.-W., Lee, K., Toutanova, K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding // Proceedings of NAACL-HLT. – 2019. – P. 4171–4186. – URL: https://aclanthology.org/N19-1423/
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P. Language Models Are Few-Shot Learners // arXiv preprint arXiv:2005.14165. – 2020. – URL: https://arxiv.org/abs/2005.14165
Kenning, M. M. What Are Parallel and Comparable Corpora and How Can We Use Them? // Routledge Handbook of Translation Studies. – 2010. – P. 487–498. – DOI: 10.4324/9780203856949-42
Čermák, F., Rosen, A. The Case of InterCorp: Corpora of Parallel Texts // International Journal of Corpus Linguistics. – 2012. – Vol. 17(3). – P. 411–427. – DOI: 10.1075/ijcl.17.3.05cer
Zanettin, F. Parallel Corpora in Translation Studies: Issues in Corpus Design and Analysis // In: The Routledge Handbook of Translation Studies. – 2017. – P. 89–105. – DOI: 10.4324/9781315759951-8
Doval, I., Nieto, M. T. S. Parallel Corpora in Focus: Learning Corpora, Teaching and Research Applications // University of Santiago de Compostela Press. – 2019. – URL: https://minerva.usc.es/bitstreams/731b612a-e792-48d9-941c-43e29ea2cc10/download
Lefer, M. A. Parallel Corpora // In: Empirical Translation Studies: New Methodological and Theoretical Perspectives. – Springer, 2021. – P. 233–252. – DOI: 10.1007/978-3-030-46216-1_12
Abdurakhmonova, N., Shamsiyeva, G. Context-Based Multilingual Translation Technology: on the Example of the Paratranslator Platform. In: Proceedings of the 10th International Conference on Computer Science and Engineering (IEEE UBMK’25), Istanbul, Türkiye, 2025, pp. 1800–1804.