Key takeaways
- Bots can be beneficial, neutral, or abusive.
- User-agent text is easy to imitate and needs verification.
- Behavioural anomalies need context from page purpose and audience.
- Response should match the harm and confidence level.
Define the unwanted behaviour
Automation can index public content, monitor availability, test systems, scrape data, abuse forms, probe inventory, or generate artificial interactions. Detection should begin with the undesirable effect, not with an assumption that every non-human visit is harmful.
Verified search crawlers should be handled differently from deceptive clients that only claim a familiar user agent.
Combine signal families
Protocol consistency, browser capabilities, device stability, request velocity, network ownership, navigation sequence, interaction timing, challenge outcomes, and conversion behaviour can contribute evidence. Each signal has legitimate edge cases.
- Verify known crawlers using current published methods.
- Baseline by route and expected user journey.
- Avoid permanent labels from one short observation window.
- Record data coverage and challenge accessibility.
Choose proportionate controls
Monitoring, rate limiting, form validation, step-up challenges, campaign action, and blocking impose different costs. Apply the least disruptive control that addresses the observed harm, then measure false positives and reversals.
Limitations
What this guide does not claim
Detection methods evolve and adversaries adapt. Accessibility tools, privacy software, shared devices, and unusual legitimate users can resemble automation, so no single technique is conclusive.
Evidence
Primary sources
- Automated Threats to Web ApplicationsOWASP Foundation
- About invalid trafficGoogle Ads Help
- AI Risk Management FrameworkNIST
Read how we source, review, update, and correct content in our editorial standards.
