←BackN°01
01 · Project

Benchmark Validity in Fake News Classification

A 98 percent accurate model that was measuring the wrong thing.

Applied AIResearchModel Evaluation
Overview

A classifier can report excellent accuracy and still fail to measure the task it claims to solve. This study audits a widely used fake news benchmark and shows how contamination and source cues can create confidence without validity.

Role
Researcher
Timeline
2026
Stack
Python, pandas, scikit-learn, Model Evaluation
Status
Research Finding
Research audit

The setup

I trained a text classifier on a widely used fake news dataset. On the standard split, the model reached 98 percent reported accuracy, a result that looked strong enough to trust at first glance.

The problem

The score was not clean evidence of fake news detection. The benchmark contained patterns that let the model identify the dataset's construction instead of evaluating the underlying claim.

Finding 01

Duplicate contamination

Duplicate and near-duplicate records crossed the train and test split. The model was evaluated on text it had effectively already seen, so part of the reported score measured recall of the benchmark rather than performance on new examples. Replace with audited duplicate count: [NUMBER NEEDED].

Finding 02

Token dominance

A single source token, such as a wire service name, carried enough class signal to do much of the classification work. When one source marker dominates the decision, the model is learning where an article came from rather than whether its content is reliable. Replace with token contribution measure: [NUMBER NEEDED].

Finding 03

The validity gap

Once contaminated duplicates and source shortcuts are controlled, the distance between the reported score and useful real world performance becomes visible. Replace with corrected accuracy and external test result: [NUMBERS NEEDED].

What it means

Headline accuracy on a public benchmark is a starting point, not proof. A corrected evaluation removes duplicates before splitting, groups related records so they cannot cross folds, checks influential tokens for source leakage, and tests the final model on genuinely unseen sources. That process measures whether the model can generalize, not whether it can recognize the benchmark.

Accuracy comparisonPlaceholder values
Next Project · 02View→

Malware and Web Attack Detection

Teaching a model to tell the difference between normal traffic and an attack, and understanding why it gets it wrong.