Skip to main content Skip to main navigation menu Skip to site footer
##common.pageHeaderLogo.altText##
Izvestiya SFedU
Engineering sciences
  • Current
  • Previous issues
    • Archive
    • Issues 1995 – 2019
  • Editorial Board
  • About journal
    • Officially
    • The main tasks
    • Main sections
    • Specialties of the Higher Attestation Commission of the Russian Federation
    • Editor-in-Chief
ISSN 1999-9429 print
ISSN 2311-3103 online
  • Login
  1. Home /
  2. Search

Search

Advanced filters
Published After
Published Before

Search Results

Found one item.
  • RECOGNITION OF EMOTIONAL STATES IN RUSSIAN SPEECH USING MFCC FUNCTIONS AND THE BLSTM MODEL FOR THE DUSHA DATASET

    P.G. Bukina , А.А. Merinov , S.S. Kharchenko , Е.Y. Kostyuchenko
    240-248
    2025-12-30
    Abstract ▼

    This paper investigates the task of automatic emotion recognition from speech signals using contemporary deep learning techniques. The relevance of this study arises from the increasing demand for intelligent systems capable of assessing human emotional states, with potential applications in medicine, psychology, information systems, and personnel management. The primary objective is to develop an efficient neural network model for emotion recognition in Russian speech that outperforms existing state-of-the-art architectures. The experiments were conducted using the open-source Russian-language dataset Dusha, which contains 300,000 audio recordings. A total of 183,055 samples from the Crowd subset, annotated with four emotional categories—joy, sadness, anger, and neutral state—were used for training. Mel-frequency cepstral coefficients (MFCCs) were extracted as input features (20 coefficients with a
    20 ms window and 10 ms overlap), followed by normalization. The baseline architecture employed a bidirectional long short-term memory network (BLSTM), capable of modeling both past and future temporal dependencies. To improve generalization and mitigate overfitting, the model was enhanced with convolutional layers (CNN), MaxPooling layers, and regularization mechanisms including Dropout and Batch Normalization. The resulting hybrid CNN–BLSTM architecture achieved 62.9% accuracy on the test set, exceeding the baseline performance (56.2%) by 6.7%. The results were further compared with state-of-the-art architectures such as MobileNetV2, HuBERT, and WavLM. The analysis highlights future directions for improving model performance through structural optimization, class balancing, and incorporation of additional acoustic features.

1 - 1 of 1 items

links

For authors
  • Submit article
  • Author Guidelines
  • Editorial Policy
  • Reviewing
  • Ethics of scientific publications
  • Open access policy
  • Supporting documents
Language
  • English
  • русский

journal

* not an advertisement

index

Индексация журнала
* not an advertisement
Information
  • For Readers
  • For Authors
  • For Librarians
Address: 347900, Taganrog, Chekhov St., 22, A-211 Phone: +7 (8634) 37-19-80 E-mail: iborodyanskiy@sfedu.ru
Publication is free
More information about the publishing system, Platform and Workflow by OJS/PKP.
logo Developed by RDCenter