Benchmarks
The numbers behind Heron
Six models were trained and compared across two modalities - text and logo images on a balanced dataset of 76,346 emails. The dual-tower fusion model tops them all at 99.45% accuracy, past the best single modality.
Model comparison
Traditional baselines (KNN, Logistic Regression) against the custom CNN specialists, a ResNet18 transfer-learning baseline, and the fused model that combines them.
| Model | Accuracy | Notes |
|---|---|---|
| KNN (text features) | 81.71% | Baseline |
| Logistic Regression | 80.00% | Baseline |
| Custom Text CNN | 98.96% | Phase 1 · text specialist |
| Custom Image CNN | 76.30% | Phase 1 · image specialist |
| ResNet18 (transfer learning) | 97.43% | Comparison baseline |
| Dual-Tower Fusion | 99.45% | Final model |
Fusion model metrics
Precision, recall and F1 are reported on the phishing class, the costliest to miss, evaluated on the held-out validation split.
Accuracy99.45%
AUC-ROC0.999
Precision (phishing)99.5%
Recall (phishing)99.4%
F1-Score99.4%