AI crawler
An automated agent operated by an AI company to fetch web content for model training, search indexing, or live retrieval when answering a…
Perplexity's crawler for building its search index, alongside a separate user-triggered agent that fetches pages in response to specific requests.
Perplexity documents both agents and publishes IP ranges. The service has attracted scrutiny over whether user-triggered fetches should honour robots.txt, since Perplexity, like other vendors, treats a fetch made on behalf of a specific user differently from indexing. Publishers who want to appear in Perplexity's cited sources need PerplexityBot allowed; those objecting to the model entirely block both and enforce at the CDN, since robots.txt is voluntary.
A recipe site allows PerplexityBot and observes it as the third-largest AI crawler in its logs after Googlebot and GPTBot.
An automated agent operated by an AI company to fetch web content for model training, search indexing, or live retrieval when answering a…
A system that responds to a question with a synthesised answer and citations rather than a ranked list of documents — for example…
OpenAI's crawler used to collect publicly available web content that may be used to improve future models; controllable via robots.txt.
Anthropic's family of web crawlers and fetchers, including a general crawler and agents that retrieve pages for search results and for…
The proportion of AI answer citations across a tracked prompt set that point to a given domain, used as the AI-era analogue of ranking…