Robustness, Cost, and Attack-Surface Concentration in Phishing Detection
Julian Allagan, Mohamed Elbakary, Zohreh Safari, et al.
This paper reveals a critical weakness in phishing detection systems: while machine learning models achieve near-perfect accuracy in testing, they remain vulnerable to attackers who slightly modify website features to evade detection. The researchers show that across different model types (logistic regression, random forests, etc.), attackers need only 2 feature changes on average to bypass detection, and these attacks concentrate on just three cheap-to-modify features. The key finding is that robustness depends on the underlying feature economics (how easy features are to change) rather than the sophistication of the AI model itself.
adversarial robustnesssecurityphishing detectionevasion attacks