Researchers publish the FakeDiverse dataset in Scientific Reports, consolidating ten public datasets to improve AI-based fake news detection. While transformer models BERT and DeBERTa achieve high accuracies of 98% and 99% on this corpus, cross-dataset evaluations reveal significant challenges in generalizing to unseen data distributions.
FakeDiverse dataset creation
- ▪Researchers consolidated news articles from ten publicly available datasets into a single corpus named FakeDiverse to expose AI models to a broader spectrum of linguistic patterns
- ▪The FakeDiverse dataset is split in an 80:20 ratio for training and testing AI models
Transformer model performance
- ▪The DeBERTa model reached 99% accuracy when trained and evaluated on the FakeDiverse dataset
Traditional ML limitations
- ▪Earlier deep learning models, including CNNs and LSTMs, provide modest improvements but continue to struggle with long text and shifting topical contexts
- ▪Traditional machine learning approaches such as Support Vector Machines and Naïve Bayes rely on shallow linguistic cues and often fail to recognize sarcasm, implicit tone, or evolving writing styles
Cross-dataset generalization challenges
- ▪Cross-dataset evaluation shows that BERT and DeBERTa models face challenges in generalizing to unseen data distributions
- ▪The study's findings highlight the need for enhanced generalization strategies and domain adaptation in AI-based fake news detection
Story comments
Loading comments…