Model training opt-out
Mechanisms by which a publisher signals that its content should not be used to train AI models, principally robots.txt tokens and…
Commercial agreements in which publishers license their content to AI companies for training, retrieval or attribution, as an alternative to blocking crawlers.
Since 2023 a number of large publishers have signed licensing deals with model providers, and infrastructure vendors have introduced marketplace and pay-per-crawl mechanisms so smaller sites can charge for automated access. The strategic question for any publisher is whether its content is more valuable as licensed input or as freely retrievable material that earns citations and referral traffic. The answer differs sharply between reference publishers, news organisations and commercial sites.
A trade publisher licenses its archive for training while keeping current articles open to retrieval agents for citation value.
Mechanisms by which a publisher signals that its content should not be used to train AI models, principally robots.txt tokens and…
An automated agent operated by an AI company to fetch web content for model training, search indexing, or live retrieval when answering a…
OpenAI's crawler used to collect publicly available web content that may be used to improve future models; controllable via robots.txt.
A search where the user's need is satisfied on the results page itself — by a snippet, knowledge panel, calculator or AI Overview — and no…
Visits arriving from AI assistants and answer engines, identifiable in analytics by referrer hosts such as chatgpt.com, perplexity.ai…