A Multimodal Speech Emotion Recognition Framework for Malayalam through Audio-Text Fusion
DOI:
https://doi.org/10.68337/cpsm.v1.i1.2026-012Keywords:
Audio-text fusion, Malayalam language, multimodal learning, speech emotion recognition, wav2vec 2.0Abstract
Speech emotion recognition (SER) is used in many domains, such as translation, intelligent assistants, healthcare monitoring, large language models, and human-computer interaction. Emotion recognition in Malayalam, however, remains challenging because of the language's rich morphological structure. This work introduces a multimodal SER model for Malayalam that uses both speech and text information. Because few Malayalam emotion databases are available, a dataset was constructed from publicly available audiovisual content. The method uses a wav2vec 2.0 model to extract audio features, while the text features are obtained with a pretrained language model. The unified model then fuses the audio and text features for emotion prediction. Label consistency in the constructed dataset was further examined with an embedding-based analysis. On a dataset of 1,040 samples evenly distributed across four emotions, the proposed model achieved an accuracy of 78.85% and a macro-averaged F1-score of 0.79 on the test set, outperforming an audio-only baseline.
References
[1] Samuel Kakuba, Alwin Poulose, and Dong Seog Han, "Deep learning approaches for bimodal speech emotion recognition: Advancements, challenges, and a multi-learning model," IEEE Access, vol. 11, pp. 113769-113789, 2023, doi: 10.1109/ACCESS.2023.3325037.
[2] Sepideh Kalateh, Luis A. Estrada-Jimenez, Sanaz Nikghadam-Hojjati, and Jose Barata, "A systematic review on multimodal emotion recognition: Building blocks, current state, applications, and challenges," IEEE Access, vol. 12, pp. 103976-104019, 2024, doi: 10.1109/ACCESS.2024.3430850.
[3] Omkumar Chandraumakantham, N. Gowtham, Mohammed Zakariah, and Abdulaziz Almazyad, "Multimodal emotion recognition using feature fusion: An LLM-based approach," IEEE Access, vol. 12, pp. 108052-108071, 2024, doi: 10.1109/ACCESS.2024.3425953.
[4] Jing Zhao and Wei-Qiang Zhang, "Improving automatic speech recognition performance for low-resource languages with self-supervised models," IEEE J. Sel. Topics Signal Process., vol. 16, no. 6, pp. 1227-1241, 2022, doi: 10.1109/JSTSP.2022.3184480.
[5] Susmitha Vekkot and Deepa Gupta, "Fusion of spectral and prosody modelling for multilingual speech emotion conversion," Knowl.-Based Syst., vol. 242, Art. no. 108360, 2022, doi: 10.1016/j.knosys.2022.108360.
[6] Michael Neumann and Ngoc Thang Vu, "Cross-lingual and multilingual speech emotion recognition on English and French," in Proc. IEEE Int. Conf. Acoust., Speech Signal Process. (ICASSP), Calgary, AB, Canada, 2018, pp. 5769-5773, doi: 10.1109/ICASSP.2018.8462162.
[7] Konlakorn Wongpatikaseree, Sattaya Singkul, Narit Hnoohom, and Sumeth Yuenyong, "Real-time end-to-end speech emotion recognition with cross-domain adaptation," Big Data Cogn. Comput., vol. 6, no. 3, Art. no. 79, 2022, doi: 10.3390/bdcc6030079.
[8] Mamyr Altaibek, Altanbek Zulkhazhav, Banu Yergesh, Gulmira Bekmanova, and Tileukhan Aibol, "A multimodal framework for speech emotion recognition in low-resource languages," J. Artif. Intell. Technol., vol. 5, pp. 354-364, 2025, doi: 10.37965/jait.2025.0781.
[9] Athira Chandran, D. Pravena, and D. Govind, "Development of speech emotion recognition system using deep belief networks in Malayalam language," in Proc. Int. Conf. Adv. Comput., Commun. Informat. (ICACCI), Udupi, India, 2017, pp. 676-680, doi: 10.1109/ICACCI.2017.8125919.
[10] Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, "wav2vec 2.0: A framework for self-supervised learning of speech representations," in Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 33, 2020, pp. 12449-12460. [Online]. Available: https://proceedings.neurips.cc/paper/2020/hash/92d1e1eb1cd6f9fba3227870bb6d7f07-Abstract.html
[11] Simran Khanuja, Diksha Bansal, Sarvesh Mehtani, Savya Khosla, Atreyee Dey, Balaji Gopalan, Dilip Kumar Margam, Pooja Aggarwal, Rajiv Teja Nagipogu, Shachi Dave, Shruti Gupta, Subhash Chandra Bose Gali, Vish Subramanian, and Partha Talukdar, "MuRIL: Multilingual representations for Indian languages," arXiv:2103.10730, 2021, doi: 10.48550/arXiv.2103.10730.
[12] Leland McInnes, John Healy, and James Melville, "UMAP: Uniform manifold approximation and projection for dimension reduction," arXiv:1802.03426, 2018, doi: 10.48550/arXiv.1802.03426.
Downloads
Published
Versions
- 2026-10-01 (2)
- 2026-09-30 (1)
Data Availability Statement
The data supporting the findings of this study are available from the corresponding author upon reasonable request. The audio clips were drawn from commercial Malayalam films; the authors did not state whether copyright clearance was obtained for their use.Conference Proceedings Volume
Section
Categories
License
Copyright (c) 2026 The Authors

This work is licensed under a Creative Commons Attribution 4.0 International License.
Articles published in Conference Proceedings in Science and Management are made available under the Creative Commons Attribution 4.0 International License (CC BY 4.0). Authors retain copyright in their work and grant the journal first publication rights. Readers are free to share and adapt the published material for any purpose, including commercial use, provided appropriate credit is given to the author, a link to the license is provided, and any changes are indicated.
