Catherine Olsson · Anthropic · 2019-03-26
Olsson takes apart one of ML safety's favorite talking points — "adversarial stickers will crash autonomous cars" — and asks whether it actually survives a real threat model. It mostly doesn't. Obscured signs, weather, and ordinary failure modes dominate the risk surface. So why study small-perturbation adversarial examples at all? Because they're a clean research testbed, not because they're the threat.
The essay is a lesson in research hygiene: don't let a vivid demo do the work of a serious argument. Useful for any learner trying to navigate the gap between what makes a good paper and what actually keeps deployed systems from breaking.
> The problem is worse than adversarial stickers.