The NeshatScan crawler
What NeshatScan fetches, how it identifies itself, how often it visits and how to block it: it obeys robots.txt rules for User-agent: NeshatScan.
NeshatScan is the fetcher behind the free scan on this site. It is not a crawler in the usual sense. It visits a site when a person asks for that site to be scanned, and in four other cases, each of which traces back to a person: once, seven days after someone emails themselves a report of that site; weekly, for a domain subscribed to Monitoring; when a Fix Pack for the domain is bought or its owner presses “Re-scan to verify”; and for our own research into which businesses in a field are reachable by answer engines, at most once a day per site. It fetches at most five URLs per visit.
How it identifies itself
Every request carries this user agent: NeshatScan/1.0 (+https://neshat.ai/scan/bot/)
What it fetches
/robots.txt- The single page a visitor typed (or the home page)
- The sitemap declared in robots.txt, or
/sitemap.xml /llms.txt
It reads at most two megabytes of a page and one megabyte of a sitemap, follows at most five redirects, and gives up on any request that takes longer than a few seconds. It sends no cookies and no credentials. Results for a URL are cached for fifteen minutes, and a single site is never scanned more than a handful of times in ten minutes, however many people ask.
What it never fetches
Anything that is not a public website: private networks, cloud metadata addresses, hostnames that resolve to internal ranges, non-standard ports, and anything reached through a redirect to one of those. Those requests are refused before a connection is opened.
How to opt out
Add a group for its token to your robots.txt and it will refuse to scan your site, showing the visitor that you asked it not to:
User-agent: NeshatScan Disallow: /
Where requests come from
The scan runs on Vercel’s serverless platform in the United States, so requests do not come from a fixed IP address. Identify it by the user agent.
Crawlers it reports on
The scan checks access for these tokens, each linked to the vendor documentation that describes it:
OAI-SearchBot— OpenAI · docs (checked 2026-09-22)Claude-SearchBot— Anthropic · docs (checked 2026-09-22)PerplexityBot— Perplexity · docs (checked 2026-09-22)ChatGPT-User— OpenAI · docs (checked 2026-09-22)Claude-User— Anthropic · docs (checked 2026-09-22)Perplexity-User— Perplexity · docs (checked 2026-09-22)Googlebot— Google · docs (checked 2026-09-22)Bingbot— Microsoft · docs (checked 2026-09-22)Applebot— Apple · docs (checked 2026-09-22)GPTBot— OpenAI · docs (checked 2026-09-22)ClaudeBot— Anthropic · docs (checked 2026-09-22)CCBot— Common Crawl · docs (checked 2026-09-22)Bytespider— ByteDance · docs (checked 2026-09-22)Meta-ExternalAgent— Meta · docs (checked 2026-09-22)Google-Extended— Google · docs (checked 2026-09-22)Applebot-Extended— Apple · docs (checked 2026-09-22)
Questions about a specific request? Write to me with the time and your domain, or run the scan on this site to see exactly what it looks at.
