Adversarial deception and data poisoning as active defense
Instead of only detecting or blocking AI-driven attacks, defenders can actively mislead attacker tooling by feeding it deceptive, corrupted, or ambiguous data. The goal is to raise attacker cost, break automation reliability, and force manual intervention.
Core ideas
- Adversarial examples: inputs crafted to confuse ML models while appearing normal to humans. Example: a CAPTCHA image that vision models misclassify.
- Data poisoning: deliberately polluting the data an attacker uses to train or refine their models, causing them to learn wrong patterns over time.
- Behavioral deception: serving inconsistent, fake, or misleading responses to suspected bots so automation cannot trust the environment.
Why it can work
Automation depends on reliability and low cost. If the environment is unpredictable, the attacker must:
- Build more robust and expensive tooling.
- Handle more false positives manually.
- Re-train or re-tune models constantly.
This raises the attacker's cost without requiring perfect detection.
Risks and limits
- False positives on legitimate users.
- Legal and ethical issues, especially around entrapment or harming third parties.
- Arms-race dynamics: attackers adapt once they detect the deception.
- Trust erosion: if your own systems or partners rely on consistent data, deception can backfire.
Connections
- Similar to honeypots but targets the attacker's ML pipeline rather than just wasting their time.
- Complements economic deterrence in security: the goal is not perfection, but raising cost enough to repel opportunistic attackers.
- Related to adversarial machine learning, which studies how ML systems fail under manipulated inputs.
Questions to ask before using it
- Who might be harmed if my deception misfires on a legitimate user?
- What signal tells me I am interacting with automation rather than a human?
- Is this deception legal and proportionate in my jurisdiction and organization?