ALTERNATIVE APPROACHES TO NLP MODEL SCALE-UP: AN ANALYSIS OF APPROACHES TO OPTIMIZING DATA AND COMPUTATION VOLUME WHEN TRAINING LARGE-SCALE LANGUAGE MODELS

Abstract

This paper focuses on overcoming the systemic limitations of the large-scale language model (LLM) scaling paradigm, which are related to data exhaustion and exponential growth in computational costs. This enables the development of more efficient approaches to building NLP models without sacrificing their performance. The goal of this study is to compare the performance of a standard transformer architecture (nanoGPT) and a model using semantic embeddings (nanoSonar) for language modeling tasks under resource constraints. Working with conceptual embeddings allows us to identify deeper linguistic patterns and reduce the amount of required training data, significantly improving modeling efficiency. The study utilized the TinyStories dataset, which includes short narratives with a clear structure. Before implementing the models, the data was preprocessed: for nanoGPT, tokenization was performed using the BPE method, and for nanoSonar, text was converted into semantic embeddings using a pretrained Sonar model. The models were evaluated using the loss and perplexity metrics. The results showed that the nanoSonar model provides significantly lower perplexity (6.609 versus 39.151 for nanoGPT) and demonstrates more robust training dynamics at later stages. This paper presents an analysis of modern approaches to scaling optimization (MoE, distillation, PEFT) and promising architectures (LRM, SSM, RWKV), and provides practical recommendations for applying models operating in the space of semantic embeddings to domain-specific problems and systems with limited computational resources. The results of this study can be useful in developing efficient language models that combine high generation quality with a cost-effective architecture.

##article.references##

1. Kaplan J. et al. Scaling laws for neural language models, arxiv preprint arxiv:2001.08361, 2020.

2. Tihanyi N. et al. Dynamic intelligence assessment: Benchmarking llms on the road to agi with a focus on model confidence, 2024 IEEE International Conference on Big Data (bigdata). IEEE, 2024,

pp. 3313-3321.

3. Villalobos P. et al. Will we run out of data? Limits of LLM scaling based on human-generated data, arxiv preprint arxiv:2211.04325, 2022.

4. Hooker S. The hardware lottery, Communications of the ACM, 2021, Vol. 64, No. 12, pp. 58-65.

5. Barrault L. et al. Large concept models: Language modeling in a sentence representation space, arxiv preprint arxiv:2412.08821, 2024.

6. Shojaee P. et al. The illusion of thinking: Understanding the strengths and limitations of reasoning mod-els via the lens of problem complexity, arxiv preprint arxiv:2506.06941, 2025.

7. Zhou Y. et al. Divergences between language models and human brains, Advances in neural information processing systems, 2024, Vol. 37, pp. 137999-138031.

8. Pagliardini M. et al. Faster causal attention over large sequences through sparse flash attention, arxiv preprint arxiv:2306.01160, 2023.

9. Du N. et al. Glam: Efficient scaling of language models with mixture-of-experts, International confer-ence on machine learning. PMLR, 2022, pp. 5547-5569.

10. Gou J. et al. Knowledge distillation: A survey, International journal of computer vision, 2021,

Vol. 129, No. 6, pp. 1789-1819.

11. Mcdonald D., Papadopoulos R., Benningfield L. Reducing llm hallucination using knowledge distilla-tion: A case study with mistral large and mmlu benchmark, Authorea Preprints, 2024.

12. Han Z. et al. Parameter-efficient fine-tuning for large models: A comprehensive survey, arxiv preprint arxiv:2403.14608, 2024.

13. Wang T. et al. Dataset distillation, arxiv preprint arxiv:1811.10959, 2018.

14. Cazenavette G. et al. Dataset distillation by matching training trajectories, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 4750-4759.

15. Aguilar G. et al. Knowledge distillation from internal representations, Proceedings of the AAAI confer-ence on artificial intelligence, 2020, Vol. 34, No. 05, pp. 7350-7357.

16. Beltagy I., Peters M.E., Cohan A. Longformer: The long-document transformer, arxiv preprint arxiv:2004.05150, 2020.

17. Xu F. et al. Towards large reasoning models: A survey of reinforced reasoning with large language models, arxiv preprint arxiv:2501.09686, 2025.

18. Duquenne P.A., Schwenk H., Sagot B. SONAR: sentence-level multimodal and language-agnostic repre-sentations, arxiv preprint arxiv:2308.11466, 2023.

19. Hamilton J.D. State-space models, Handbook of econometrics, 1994, Vol. 4, pp. 3039-3080.

20. Peng B. et al. Rwkv: Reinventing rnns for the transformer era, arxiv preprint arxiv:2305.13048, 2023.

21. Ghosh-Dastidar S., Adeli H. Spiking neural networks, International journal of neural systems, 2009, Vol. 19, No. 04, pp. 295-308.

22. Dragunov N. et al. SONAR-LLM: Autoregressive Transformer that Thinks in Sentence Embeddings and Speaks in Tokens, arxiv preprint arxiv:2508.05305, 2025.

23. Eldan R., Li Y. Tinystories: How small can language models be and still speak coherent english?, arxiv preprint arxiv:2305.07759, 2023.

24. Kovalev V.V. Algoritm predvaritel'noy obrabotki izobrazheniy dlya snizheniya veroyatnosti pereobucheniya svertochnykh neyronnykh setey na neyronnom uskoritele [An algorithm for image pre-processing to reduce the probability of overfitting of convolutional neural networks on a neural accelera-tor], Izvestiya YuFU. Tekhnicheskie nauki [Izvestiya SFedU. Engineering Sciences], 2024, No. 5 (241), pp. 29-37.

25. Kovalev V.V., Sergeev N.E. Rasshirenie priznakovogo prostranstva v zadache poiska i raspoznavaniya malorazmernykh ob"ektov na izobrazheniyakh [Expansion of the feature space in the problem of search-ing and recognizing small-sized objects in images] Izvestiya YuFU. Tekhnicheskie nauki [Izvestiya SFedU. Engineering Sciences], 2024, No. 1 (237), pp. 267-276.

26. Kovalev V.V., Sergeev N.E. Realizatsiya svertochnykh neyronnykh setey na vstraivaemykh ustroystvakh s ogranichennym vychislitel'nym resursom [Implementation of convolutional neural networks on em-bedded devices with limited computing resources], Izvestiya YuFU. Tekhnicheskie nauki [Izvestiya SFedU. Engineering Sciences], 2021, No. 6 (223), pp. 64-72.

Скачивания

##article.published##:

2026-07-07

##article.issue##:

##article.section##:

SECTION III. MACHINE LEARNING AND DATA PROCESSING

DOI:

Keywords:

Large language models, scaling, computational optimization, Transformer architecture, semantic embeddings, NLP, machine learning, data efficiency

##submission.сitation##:

Ralko К.I. , Sergeev N. Е. ALTERNATIVE APPROACHES TO NLP MODEL SCALE-UP: AN ANALYSIS OF APPROACHES TO OPTIMIZING DATA AND COMPUTATION VOLUME WHEN TRAINING LARGE-SCALE LANGUAGE MODELS. IZVESTIYA SFedU. ENGINEERING SCIENCES. – 2026. - № 3. – ##article.page##. 152-172.