← Back to News

Fake Crawlers, Real Threat: Catching an Attacker Impersonating Apple, OpenAI and Perplexity

14 September 2026

The Risk

Many organisations allowlist "known good" crawlers by User-Agent string alone — Googlebot, Applebot, GPTBot, PerplexityBot — on the assumption that a recognisable name in a request header is a trustworthy signal. It isn't. A User-Agent is just a string of text a client sends about itself, and nothing stops an attacker's own script from claiming to be any crawler it likes. Defenses built on that assumption have a blind spot exactly where it matters most.

The Threat

One of our honeypot sensors recently caught this blind spot being actively exploited. An attacker was systematically requesting exposed files — .env, .git/config, AWS credential files — the kind of reconnaissance that precedes a real credential-theft or account-takeover attempt. On its own, that's unremarkable; sensors like ours see this constantly. What stood out was the User-Agent on the requests: Apple's Applebot. Then OpenAI's GPTBot. Then ChatGPT-User. Then PerplexityBot. All four, from the same source IPs, and none of them real.

We checked the source IPs against each crawler's own published address ranges. Neither IP appeared in any of them. One had no reverse DNS record at all; the other resolved to a Google Cloud VM — consistent with a single operator rotating through four trusted-crawler identities specifically to slip past defenses that allowlist "known good bots" by User-Agent string alone. It didn't work here, because our detection reads what's actually being requested, not who the request claims to be.

The numbers behind it, from the last 7 days alone across our sensor network:

  • 55,000+ connection attempts logged
  • 599 requests specifically hunting for exposed secrets — .env files, .git configs, cloud credential paths
  • 296 of those requests disguised as a trusted AI or search crawler — 98 as GPTBot, 80 as ChatGPT-User, 68 as Applebot, 50 as PerplexityBot, all traced back to just 2 source IPs
  • 7 confirmed live WannaCry samples captured in the same window, still circulating almost a decade after the outbreak that made it infamous, among 20 total malware samples our sensors pulled in (one further identified as "ZombieBoy"; the rest remain unclassified)

A second, quieter finding sits behind the volume numbers. The single largest source block hitting one of our sensors that week — accounting for roughly 71% of its traffic — wasn't the compromised home routers or IoT devices that "botnet" usually brings to mind. WHOIS traced it to a commercial Bulgarian/US VPS provider (GBTCloud/QuickPacket). Rented, professionally hosted infrastructure, not hijacked consumer hardware. It's a useful reminder that a meaningful share of today's background internet attack traffic is running on infrastructure its operators paid for outright, not devices they broke into.

The Fix

  • Don't allowlist by User-Agent string alone. Where a "known good crawler" carve-out genuinely matters, verify the source IP against that crawler's own published address ranges — the identity claim in a header is free for anyone to make.
  • Detect on intent, not identity. A request hunting for .env or .git/config is suspicious regardless of what it claims to be. Behaviour-based detection catches spoofed-identity attacks that allowlist-based defenses miss entirely.
  • Treat "trusted bot" traffic touching sensitive paths as an anomaly, not an exception. Legitimate crawlers have no reason to request credential files.
  • Don't assume attack infrastructure is always compromised infrastructure. Rented commercial VPS capacity increasingly does the same job as a hijacked botnet, and it's traceable in ways a botnet isn't — WHOIS and hosting-provider abuse channels are underused tools here.
  • Low-and-slow reconnaissance like this rarely reaches a SOC dashboard on its own. It looks like ordinary internet background noise until it's correlated across sensors and traced back to source — exactly the gap dedicated sensor coverage is built to close.
⚙