THE IMPORTANCE OF THE PARALLEL CORPUS AS A LINGUISTIC BASE

Authors

  • Shamsiyeva Gulshoda Asliddin qizi

Keywords:

computational linguistics; parallel corpus; artificial intelligence; natural language processing (NLP); Uzbek language corpus; machine translation; GPT; BERT; corpus linguistics; linguistic database; transformative model; language modeling; translation technologies.

Abstract

This article is devoted to the analysis of research conducted in the field of computational linguistics worldwide and in Uzbekistan during the period from 2020 to 2025. The study covers theoretical frameworks of computational linguistics, artificial intelligence–based language models (GPT, BERT, LLaMA, and others), as well as the role of parallel corpora as a linguistic foundation. An analysis of international experience demonstrates that over the past five years, the priority areas in computational linguistics have included the training of linguistic models based on neural transformer architectures, the development of multilingual parallel corpora, and the improvement of machine translation systems.

In Uzbekistan, the field of computational linguistics is observed to be transitioning from a predominantly applied stage to a more scientific and analytical phase through projects such as the Uzbek language corpus (uzbekcorpus.uz) and Paratranslator. The research employs systematic analysis of scientific literature, comparative linguistic and corpus-based methods, as well as the analytical capabilities of artificial intelligence models.

References

Devlin, J., Chang, M.-W., Lee, K., Toutanova, K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding // Proceedings of NAACL-HLT. – 2019. – P. 4171–4186. – URL: https://aclanthology.org/N19-1423/

Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P. Language Models Are Few-Shot Learners // arXiv preprint arXiv:2005.14165. – 2020. – URL: https://arxiv.org/abs/2005.14165

Kenning, M. M. What Are Parallel and Comparable Corpora and How Can We Use Them? // Routledge Handbook of Translation Studies. – 2010. – P. 487–498. – DOI: 10.4324/9780203856949-42

Čermák, F., Rosen, A. The Case of InterCorp: Corpora of Parallel Texts // International Journal of Corpus Linguistics. – 2012. – Vol. 17(3). – P. 411–427. – DOI: 10.1075/ijcl.17.3.05cer

Zanettin, F. Parallel Corpora in Translation Studies: Issues in Corpus Design and Analysis // In: The Routledge Handbook of Translation Studies. – 2017. – P. 89–105. – DOI: 10.4324/9781315759951-8

Doval, I., Nieto, M. T. S. Parallel Corpora in Focus: Learning Corpora, Teaching and Research Applications // University of Santiago de Compostela Press. – 2019. – URL: https://minerva.usc.es/bitstreams/731b612a-e792-48d9-941c-43e29ea2cc10/download

Lefer, M. A. Parallel Corpora // In: Empirical Translation Studies: New Methodological and Theoretical Perspectives. – Springer, 2021. – P. 233–252. – DOI: 10.1007/978-3-030-46216-1_12

Abdurakhmonova, N., Shamsiyeva, G. Context-Based Multilingual Translation Technology: on the Example of the Paratranslator Platform. In: Proceedings of the 10th International Conference on Computer Science and Engineering (IEEE UBMK’25), Istanbul, Türkiye, 2025, pp. 1800–1804.

Downloads

Published

2026-04-02