Email Phishing Detection
A classifier tested for accuracy, generalization, and real world performance.
A phishing email classifier trained on 82,486 emails and stress-tested on unseen data to make sure the accuracy held up beyond the training set. Deployed as a Streamlit app.
- Role
- Builder
- Timeline
- 2026
- Stack
- Python, scikit-learn, pandas, HuggingFace, SMOTE, Streamlit
- Status
- Deployed (AI 311, UTK)
Phishing emails slip past filters. A model that scores well on a training set but fails on real-world emails is not useful. The goal was a classifier that actually generalizes.
Build
story.
Trained and evaluated three models: Logistic Regression, SVM (LinearSVC), and Random Forest on 82,486 emails (48% legitimate, 52% phishing) sourced from Kaggle.

Extended the dataset in Week 6 with 18,634 unseen emails to test generalization. Retrained on the combined set.

Used SMOTE for class balancing and deployed via Streamlit.

98.37% accuracy on combined dataset. 97.31% generalization on unseen emails. Real-world stress-test improved from 70% to 80% after retraining.
98.37%
Combined accuracy
97.31%
Generalization, unseen
70 → 80%
Real-world stress test
PhotoChain
Every image has a creator. PhotoChain makes sure the world knows who.