Benchmarks

The numbers behind Heron

Six models were trained and compared across two modalities - text and logo images on a balanced dataset of 76,346 emails. The dual-tower fusion model tops them all at 99.45% accuracy, past the best single modality.

Model comparison

Traditional baselines (KNN, Logistic Regression) against the custom CNN specialists, a ResNet18 transfer-learning baseline, and the fused model that combines them.

ModelAccuracyNotes
KNN (text features)
81.71%
Baseline
Logistic Regression
80.00%
Baseline
Custom Text CNN
98.96%
Phase 1 · text specialist
Custom Image CNN
76.30%
Phase 1 · image specialist
ResNet18 (transfer learning)
97.43%
Comparison baseline
Dual-Tower Fusion
99.45%
Final model

Fusion model metrics

Precision, recall and F1 are reported on the phishing class, the costliest to miss, evaluated on the held-out validation split.

Accuracy99.45%
AUC-ROC0.999
Precision (phishing)99.5%
Recall (phishing)99.4%
F1-Score99.4%