←BackN°02
02 · Project

Malware and Web Attack Detection

Teaching a model to tell the difference between normal traffic and an attack, and understanding why it gets it wrong.

Machine LearningSecurityClassification
Overview

A classifier built on the CICIDS 2017 network traffic dataset to separate malicious flows, including brute force, XSS, and SQL injection, from benign traffic. The work focused on feature engineering, class imbalance, and understanding what the model was actually learning.

Role
Builder
Timeline
2026
Stack
Python, scikit-learn, pandas, Random Forest, Logistic Regression, GridSearchCV
Status
Built (AI 311, UTK)
The Problem

Network traffic logs contain millions of data points, and malicious behavior hides inside patterns that look almost normal. The challenge was building a classifier that could detect web attacks, including brute force, XSS, and SQL injection, without overfitting to specific IPs or memorizing the training data.

Process

Build
story.

Worked with the Thursday network traffic dataset containing four classes: benign traffic, brute force, XSS, and SQL injection. Cleaned the data by removing identifier columns like IP addresses to prevent overfitting, handled divide-by-zero errors from zero-duration connections, and normalized features across wildly different scales.

Feature importance

Trained and compared Logistic Regression and Random Forest models. Random Forest won, catching 50% of attacks versus 30% for Logistic Regression, because it builds 100 decision trees and combines them rather than drawing a single boundary between safe and malicious.

Used hyperparameter tuning via GridSearchCV and analyzed feature importance to understand what the model was actually learning.

Results

Random Forest outperformed Logistic Regression across the board. Top features driving detection: Avg Bwd Segment Size, Destination Port, and Bwd Packet Length Mean, all signals about how the server responds during an attack. The tuned model achieved 0 false positives but 5 false negatives, surfacing a real cybersecurity trade-off: a cautious model that never cries wolf but misses half the actual attacks.

50%

Attack detection rate

0

False positives

RF > LR

Random Forest vs Logistic Regression

Next Project · 03View→

Email Phishing Detection

A classifier tested for accuracy, generalization, and real world performance.