Reference library

Dataset

Open datasets you can query yourself, from crawl corpora to real-user performance data.

7 references · 1 canonical · 6 free · Advanced / Intermediate

TitleAuthor / publisherYearLevelCost
Common CrawlCanon
Free, openly licensed petabyte-scale archive of web crawl data with monthly releases of WARC, WAT and WET files plus columnar indexes. Heavily used for link graph research and LLM pretraining.
commoncrawl.org
Common Crawl FoundationAdvancedFree
ai.robots.txt Community Blocklist
Community-maintained open-source list of AI-related crawlers with ready-made robots.txt, .htaccess, nginx, Caddy, HAProxy and JSON outputs for blocking or auditing them.
github.com
ai-robots-txt projectIntermediateFree
Chrome UX Report on BigQuery
The public BigQuery dataset containing monthly Core Web Vitals distributions for millions of origins, queryable for competitive performance benchmarking at scale.
console.cloud.google.com
Google Chrome teamAdvancedFreemium
ClueWeb22 Dataset
Large research web corpus derived from commercial search engine crawls, distributed under licence for academic information retrieval and web search research.
lemurproject.org
Lemur Project / Carnegie Mellon UniversityAdvancedFree
HTTP Archive Reports
Longitudinal trend reports on how the web is built and performs, backed by a public BigQuery dataset of monthly crawls of millions of URLs.
httparchive.org
HTTP ArchiveAdvancedFree
MS MARCO
Large-scale dataset of real anonymised Bing queries with human-generated answers and passage relevance labels. The standard benchmark for passage ranking and dense retrieval research.
microsoft.github.io
MicrosoftAdvancedFree
Wikimedia Downloads (Wikipedia Dumps)
Complete periodic exports of Wikipedia and sister projects, including page content, pagelinks and pageview data. A standard corpus for entity, knowledge graph and query understanding work.
dumps.wikimedia.org
Wikimedia FoundationAdvancedFree

Other formats