Multimodal Sentiment and Emotion Detection in Social Media: Methods, Modalities, and Application Perspectives
DOI:
https://doi.org/10.68337/cpsm.v1.i1.2026-011Keywords:
Multimodal sentiment, emotion detection, social media analytics, deep learning, multimodal fusionAbstract
The growing popularity of social media has produced large volumes of multimodal user-generated content in the form of text, images, and speech, creating new opportunities to study how users express sentiment and emotion. This paper presents a narrative review of multimodal sentiment and emotion detection methods designed for social media data. The review covers representation methods for the textual, visual, and speech modalities, together with the machine learning and deep learning models used to analyze each modality and to fuse multimodal representations. The literature is analyzed in terms of feature extraction methods, fusion methods, and application contexts such as mental health monitoring, disaster management, education, and healthcare. The review shows that several weaknesses persist: text-based models still dominate, heterogeneous modalities are poorly integrated, standardized multimodal datasets are lacking, and deep multimodal systems are difficult to interpret. To address these concerns, the paper consolidates methodological knowledge across modalities and frames a unified multimodal analysis perspective that focuses on combining complementary features and on application-specific issues. The review offers a reference framework to guide future research and development in social media sentiment and emotion analysis by consolidating approaches, identifying research gaps, and clarifying the practical requirements of emotion-aware multimodal systems.
References
[1] Louis-Philippe Morency, Rada Mihalcea, and Payal Doshi, "Towards multimodal sentiment analysis: Harvesting opinions from the web," in Proc. 13th Int. Conf. Multimodal Interfaces (ICMI), 2011, pp. 169-176, doi: 10.1145/2070481.2070509.
[2] Soujanya Poria, Erik Cambria, Amir Hussain, and Guang-Bin Huang, "Towards an intelligent framework for multimodal affective data analysis," Neural Netw., vol. 63, pp. 104-116, 2015, doi: 10.1016/j.neunet.2014.10.005.
[3] Harika Abburi, Rajendra Prasath, Manish Shrivastava, and Suryakanth V. Gangashetty, "Multimodal sentiment analysis using deep neural networks," in Mining Intelligence and Knowledge Exploration (MIKE 2016), Lecture Notes in Computer Science, vol. 10089, 2017, pp. 58-65, doi: 10.1007/978-3-319-58130-9_6.
[4] Amir Zadeh, Minghai Chen, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency, "Tensor fusion network for multimodal sentiment analysis," in Proc. 2017 Conf. Empirical Methods in Natural Language Processing (EMNLP), 2017, pp. 1103-1114, doi: 10.18653/v1/D17-1115.
[5] Amir Zadeh, Paul Pu Liang, Navonil Mazumder, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency, "Memory fusion network for multi-view sequential learning," in Proc. AAAI Conf. Artif. Intell., vol. 32, no. 1, 2018, pp. 5634-5641, doi: 10.1609/aaai.v32i1.12021.
[6] Devamanyu Hazarika, Soujanya Poria, Amir Zadeh, Erik Cambria, Louis-Philippe Morency, and Roger Zimmermann, "Conversational memory network for emotion recognition in dyadic dialogue videos," in Proc. 2018 Conf. North American Chapter Assoc. Comput. Linguistics: Human Language Technologies (NAACL-HLT), vol. 1, 2018, pp. 2122-2132, doi: 10.18653/v1/N18-1193.
[7] Devamanyu Hazarika, Roger Zimmermann, and Soujanya Poria, "MISA: Modality-invariant and -specific representations for multimodal sentiment analysis," in Proc. 28th ACM Int. Conf. Multimedia, 2020, pp. 1122-1131, doi: 10.1145/3394171.3413678.
[8] Soujanya Poria, Erik Cambria, Devamanyu Hazarika, Navonil Mazumder, Amir Zadeh, and Louis-Philippe Morency, "Multi-level multiple attentions for contextual multimodal sentiment analysis," in Proc. 2017 IEEE Int. Conf. Data Mining (ICDM), 2017, pp. 1033-1038, doi: 10.1109/ICDM.2017.134.
[9] Navonil Majumder, Devamanyu Hazarika, Alexander Gelbukh, Erik Cambria, and Soujanya Poria, "Multimodal sentiment analysis using hierarchical fusion with context modeling," Knowl.-Based Syst., vol. 161, pp. 124-133, 2018, doi: 10.1016/j.knosys.2018.07.041.
[10] Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J. Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov, "Multimodal transformer for unaligned multimodal language sequences," in Proc. 57th Annu. Meeting Assoc. Comput. Linguistics (ACL), 2019, pp. 6558-6569, doi: 10.18653/v1/P19-1656.
[11] Yao-Hung Hubert Tsai, Martin Ma, Muqiao Yang, Ruslan Salakhutdinov, and Louis-Philippe Morency, "Multimodal routing: Improving local and global interpretability of multimodal language analysis," in Proc. 2020 Conf. Empirical Methods in Natural Language Processing (EMNLP), 2020, pp. 1823-1833, doi: 10.18653/v1/2020.emnlp-main.143.
[12] Wenmeng Yu, Hua Xu, Ziqi Yuan, and Jiele Wu, "Learning modality-specific representations with self-supervised multi-task learning for multimodal sentiment analysis," in Proc. AAAI Conf. Artif. Intell., vol. 35, no. 12, 2021, pp. 10790-10797, doi: 10.1609/aaai.v35i12.17289.
[13] Paul Pu Liang, Zihao Deng, Martin Q. Ma, James Y. Zou, Louis-Philippe Morency, and Ruslan Salakhutdinov, "Factorized contrastive learning: Going beyond multi-view redundancy," in Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 32971-32998. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2023/hash/6818dcc65fdf3cbd4b05770fb957803e-Abstract-Conference.html
[14] Mohammad Soleymani, Jeroen Lichtenauer, Thierry Pun, and Maja Pantic, "A multimodal database for affect recognition and implicit tagging," IEEE Trans. Affect. Comput., vol. 3, no. 1, pp. 42-55, 2012, doi: 10.1109/T-AFFC.2011.25.
[15] Swadha Gupta, Parteek Kumar, and Raj Kumar Tekchandani, "Facial emotion recognition based real-time learner engagement detection system in online learning context using deep learning models," Multimed. Tools Appl., vol. 82, no. 8, pp. 11365-11394, 2023, doi: 10.1007/s11042-022-13558-9.
[16] Trisha Mittal, Uttaran Bhattacharya, Rohan Chandra, Aniket Bera, and Dinesh Manocha, "M3ER: Multiplicative multimodal emotion recognition using facial, textual, and speech cues," in Proc. AAAI Conf. Artif. Intell., vol. 34, no. 2, 2020, pp. 1359-1367, doi: 10.1609/aaai.v34i02.5492.
[17] Amir Hossein Yazdavar, Mohammad Saeid Mahdavinejad, Goonmeet Bajaj, William Romine, Amit Sheth, Amir Hassan Monadjemi, Krishnaprasad Thirunarayan, John M. Meddar, Annie Myers, Jyotishman Pathak, and Pascal Hitzler, "Multimodal mental health analysis in social media," PLoS One, vol. 15, no. 4, Art. no. e0226248, 2020, doi: 10.1371/journal.pone.0226248.
[18] Rishit Jain, Revant Singh Rai, Sajal Jain, Ruchir Ahluwalia, and Jyoti Gupta, "Real time sentiment analysis of natural language using multimedia input," Multimed. Tools Appl., vol. 82, no. 26, pp. 41021-41036, 2023, doi: 10.1007/s11042-023-15213-3.
[19] Lydia Bryan-Smith, Jake Godsall, Franky George, Kelly Egode, Nina Dethlefs, and Dan Parsons, "Real-time social media sentiment analysis for rapid impact assessment of floods," Comput. Geosci., vol. 178, Art. no. 105405, 2023, doi: 10.1016/j.cageo.2023.105405.
[20] Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency, "Multimodal machine learning: A survey and taxonomy," IEEE Trans. Pattern Anal. Mach. Intell., vol. 41, no. 2, pp. 423-443, 2019, doi: 10.1109/TPAMI.2018.2798607.
[21] Songning Lai, Xifeng Hu, Haoxuan Xu, Zhaoxia Ren, and Zhi Liu, "Multimodal sentiment analysis: A survey," Displays, vol. 80, Art. no. 102563, 2023, doi: 10.1016/j.displa.2023.102563.
[22] Alaa Alslaity and Rita Orji, "Machine learning techniques for emotion detection and sentiment analysis: Current state, challenges, and future directions," Behav. Inf. Technol., vol. 43, no. 1, pp. 139-164, 2024, doi: 10.1080/0144929X.2022.2156387.
Downloads
Published
Versions
- 2026-10-01 (2)
- 2026-09-30 (1)
Data Availability Statement
No new data were generated or analyzed in this study; this narrative review draws only on previously published literature.Conference Proceedings Volume
Section
Categories
License
Copyright (c) 2026 The Authors

This work is licensed under a Creative Commons Attribution 4.0 International License.
Articles published in Conference Proceedings in Science and Management are made available under the Creative Commons Attribution 4.0 International License (CC BY 4.0). Authors retain copyright in their work and grant the journal first publication rights. Readers are free to share and adapt the published material for any purpose, including commercial use, provided appropriate credit is given to the author, a link to the license is provided, and any changes are indicated.
