Early Fake News Detection Using Linguistic Features: A Comparative Study of Machine Learning and Transformer Models

Authors

  • Priya Asha A. Department of Computer Science, Faculty of Engineering and Technology, SRM Institute of Science and Technology, Chennai, India Author
  • Gayathri R. Department of Computer Science, Faculty of Engineering and Technology, SRM Institute of Science and Technology, Chennai, India Author

DOI:

https://doi.org/10.68337/cpsm.v1.i1.2026-014

Keywords:

Fake news detection, machine learning, BERT, linguistic features, early detection, natural language processing

Abstract

Misinformation and disinformation spread rapidly on social media, so reliable mechanisms are needed to classify news as fake or real. Although deep learning models have become increasingly popular, traditional machine learning methods remain relevant because they are easier to interpret and can be very efficient. This paper proposes a lightweight framework for fake news detection that uses linguistic and sentiment-based features of the textual content. The framework was evaluated in detail, and the results show that the Bidirectional Encoder Representations from Transformers (BERT) model outperformed the other models in accuracy, while traditional models with engineered features performed adequately at a much lower computational cost. The results demonstrate the effectiveness of simple yet powerful methods for detecting fake news from the text of complete news articles.

References

[1] Kai Shu, Amy Sliva, Suhang Wang, Jiliang Tang, and Huan Liu, "Fake news detection on social media: A data mining perspective," ACM SIGKDD Explor. Newsl., vol. 19, no. 1, pp. 22-36, 2017, doi: 10.1145/3137597.3137600.

[2] A. B. Athira, S. D. Madhu Kumar, and Anu Mary Chacko, "A systematic survey on explainable AI applied to fake news detection," Eng. Appl. Artif. Intell., vol. 122, Art. no. 106087, 2023, doi: 10.1016/j.engappai.2023.106087.

[3] Bo Hu, Zhendong Mao, and Yongdong Zhang, "An overview of fake news detection: From a new perspective," Fundam. Res., vol. 5, no. 1, pp. 332-346, 2025, doi: 10.1016/j.fmre.2024.01.017.

[4] Carmela Comito, Luciano Caroprese, and Ester Zumpano, "Multimodal fake news detection on social media: A survey of deep learning techniques," Soc. Netw. Anal. Min., vol. 13, no. 1, Art. no. 101, 2023, doi: 10.1007/s13278-023-01104-w.

[5] Chanchal Kumar, Mani Bansal, Mohd Anas Khan, Vinay Kaushik, Md. Arquam, and Abdulatif Alabdultif, "Graph-augmented transformer ensemble framework for robust and scalable fake news detection in social media ecosystems," Sci. Rep., vol. 16, Art. no. 2001, 2026, doi: 10.1038/s41598-025-31653-3.

[6] Noureddine Seddari, Abdelouahid Derhab, Mohamed Belaoued, Waleed Halboob, Jalal Al-Muhtadi, and Abdelghani Bouras, "A hybrid linguistic and knowledge-based analysis approach for fake news detection on social media," IEEE Access, vol. 10, pp. 62097-62109, 2022, doi: 10.1109/ACCESS.2022.3181184.

[7] Shaina Raza and Chen Ding, "Fake news detection based on news content and social contexts: A transformer-based approach," Int. J. Data Sci. Anal., vol. 13, no. 4, pp. 335-362, 2022, doi: 10.1007/s41060-021-00302-z.

[8] William Yang Wang, "'Liar, liar pants on fire': A new benchmark dataset for fake news detection," in Proc. 55th Annu. Meeting Assoc. Comput. Linguistics (ACL), vol. 2, 2017, pp. 422-426, doi: 10.18653/v1/P17-2067.

[9] Natali Ruchansky, Sungyong Seo, and Yan Liu, "CSI: A hybrid deep model for fake news detection," in Proc. 2017 ACM Conf. Inf. Knowl. Manag. (CIKM), 2017, pp. 797-806, doi: 10.1145/3132847.3132877.

[10] Jayanti Rout, Minati Mishra, and Manob Jyoti Saikia, "Towards reliable fake news detection: Enhanced attention-based transformer model," J. Cybersecur. Priv., vol. 5, no. 3, Art. no. 43, 2025, doi: 10.3390/jcp5030043.

[11] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, "BERT: Pre-training of deep bidirectional transformers for language understanding," in Proc. 2019 Conf. North American Chapter Assoc. Comput. Linguistics: Human Language Technologies (NAACL-HLT), vol. 1, 2019, pp. 4171-4186, doi: 10.18653/v1/N19-1423.

[12] Aman Malik, Dayal Kumar Behera, Jhalak Hota, and Amulya Ratna Swain, "Ensemble graph neural networks for fake news detection using user engagement and text features," Results Eng., vol. 24, Art. no. 103081, 2024, doi: 10.1016/j.rineng.2024.103081.

[13] Litian Zhang, Xiaoming Zhang, Ziyi Zhou, Xi Zhang, Philip S. Yu, and Chaozhuo Li, "Knowledge-aware multimodal pre-training for fake news detection," Inf. Fusion, vol. 114, Art. no. 102715, 2025, doi: 10.1016/j.inffus.2024.102715.

Downloads

Published

2026-09-30 — Updated on 2026-10-01

Versions

Data Availability Statement

This study used the publicly available ISOT Fake News dataset (ISOT Research Lab, University of Victoria).

How to Cite

Early Fake News Detection Using Linguistic Features: A Comparative Study of Machine Learning and Transformer Models. (2026). Conference Proceedings in Science and Management, 1(1), 57-60. https://doi.org/10.68337/cpsm.v1.i1.2026-014 (Original work published 2026)