A Multimodal Speech Emotion Recognition Framework for Malayalam through Audio-Text Fusion

Authors

  • Athira Raj Department of Computer Science and Engineering, College of Engineering Trivandrum, Thiruvananthapuram, India Author
  • Christy James Jose Department of Electronics and Communication Engineering, College of Engineering Trivandrum, Thiruvananthapuram, India Author
  • Biju K. S. Department of Electronics and Communication Engineering, College of Engineering Trivandrum, Thiruvananthapuram, India Author

DOI:

https://doi.org/10.68337/cpsm.v1.i1.2026-012

Keywords:

Audio-text fusion, Malayalam language, multimodal learning, speech emotion recognition, wav2vec 2.0

Abstract

Speech emotion recognition (SER) is used in many domains, such as translation, intelligent assistants, healthcare monitoring, large language models, and human-computer interaction. Emotion recognition in Malayalam, however, remains challenging because of the language's rich morphological structure. This work introduces a multimodal SER model for Malayalam that uses both speech and text information. Because few Malayalam emotion databases are available, a dataset was constructed from publicly available audiovisual content. The method uses a wav2vec 2.0 model to extract audio features, while the text features are obtained with a pretrained language model. The unified model then fuses the audio and text features for emotion prediction. Label consistency in the constructed dataset was further examined with an embedding-based analysis. On a dataset of 1,040 samples evenly distributed across four emotions, the proposed model achieved an accuracy of 78.85% and a macro-averaged F1-score of 0.79 on the test set, outperforming an audio-only baseline.

References

[1] Samuel Kakuba, Alwin Poulose, and Dong Seog Han, "Deep learning approaches for bimodal speech emotion recognition: Advancements, challenges, and a multi-learning model," IEEE Access, vol. 11, pp. 113769-113789, 2023, doi: 10.1109/ACCESS.2023.3325037.

[2] Sepideh Kalateh, Luis A. Estrada-Jimenez, Sanaz Nikghadam-Hojjati, and Jose Barata, "A systematic review on multimodal emotion recognition: Building blocks, current state, applications, and challenges," IEEE Access, vol. 12, pp. 103976-104019, 2024, doi: 10.1109/ACCESS.2024.3430850.

[3] Omkumar Chandraumakantham, N. Gowtham, Mohammed Zakariah, and Abdulaziz Almazyad, "Multimodal emotion recognition using feature fusion: An LLM-based approach," IEEE Access, vol. 12, pp. 108052-108071, 2024, doi: 10.1109/ACCESS.2024.3425953.

[4] Jing Zhao and Wei-Qiang Zhang, "Improving automatic speech recognition performance for low-resource languages with self-supervised models," IEEE J. Sel. Topics Signal Process., vol. 16, no. 6, pp. 1227-1241, 2022, doi: 10.1109/JSTSP.2022.3184480.

[5] Susmitha Vekkot and Deepa Gupta, "Fusion of spectral and prosody modelling for multilingual speech emotion conversion," Knowl.-Based Syst., vol. 242, Art. no. 108360, 2022, doi: 10.1016/j.knosys.2022.108360.

[6] Michael Neumann and Ngoc Thang Vu, "Cross-lingual and multilingual speech emotion recognition on English and French," in Proc. IEEE Int. Conf. Acoust., Speech Signal Process. (ICASSP), Calgary, AB, Canada, 2018, pp. 5769-5773, doi: 10.1109/ICASSP.2018.8462162.

[7] Konlakorn Wongpatikaseree, Sattaya Singkul, Narit Hnoohom, and Sumeth Yuenyong, "Real-time end-to-end speech emotion recognition with cross-domain adaptation," Big Data Cogn. Comput., vol. 6, no. 3, Art. no. 79, 2022, doi: 10.3390/bdcc6030079.

[8] Mamyr Altaibek, Altanbek Zulkhazhav, Banu Yergesh, Gulmira Bekmanova, and Tileukhan Aibol, "A multimodal framework for speech emotion recognition in low-resource languages," J. Artif. Intell. Technol., vol. 5, pp. 354-364, 2025, doi: 10.37965/jait.2025.0781.

[9] Athira Chandran, D. Pravena, and D. Govind, "Development of speech emotion recognition system using deep belief networks in Malayalam language," in Proc. Int. Conf. Adv. Comput., Commun. Informat. (ICACCI), Udupi, India, 2017, pp. 676-680, doi: 10.1109/ICACCI.2017.8125919.

[10] Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, "wav2vec 2.0: A framework for self-supervised learning of speech representations," in Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 33, 2020, pp. 12449-12460. [Online]. Available: https://proceedings.neurips.cc/paper/2020/hash/92d1e1eb1cd6f9fba3227870bb6d7f07-Abstract.html

[11] Simran Khanuja, Diksha Bansal, Sarvesh Mehtani, Savya Khosla, Atreyee Dey, Balaji Gopalan, Dilip Kumar Margam, Pooja Aggarwal, Rajiv Teja Nagipogu, Shachi Dave, Shruti Gupta, Subhash Chandra Bose Gali, Vish Subramanian, and Partha Talukdar, "MuRIL: Multilingual representations for Indian languages," arXiv:2103.10730, 2021, doi: 10.48550/arXiv.2103.10730.

[12] Leland McInnes, John Healy, and James Melville, "UMAP: Uniform manifold approximation and projection for dimension reduction," arXiv:1802.03426, 2018, doi: 10.48550/arXiv.1802.03426.

Downloads

Published

2026-09-30 — Updated on 2026-10-01

Versions

Data Availability Statement

The data supporting the findings of this study are available from the corresponding author upon reasonable request. The audio clips were drawn from commercial Malayalam films; the authors did not state whether copyright clearance was obtained for their use.

How to Cite

A Multimodal Speech Emotion Recognition Framework for Malayalam through Audio-Text Fusion. (2026). Conference Proceedings in Science and Management, 1(1), 49-52. https://doi.org/10.68337/cpsm.v1.i1.2026-012 (Original work published 2026)