Rand Fishkin
Clickstream-based analysis finding that for every 1,000 Google searches only about 360 clicks reach the open web in the US and about 374 in the EU, with the remainder ending in no click or a click to a Google-owned property.
Reference library
188 references — the official documentation that is authoritative by definition, the RFCs and specifications underneath it, the books that are still worth the shelf space, the courses that teach rather than sell, and the studies people keep citing. 38 are marked as the canon: the ones to read before anything else.
If you read nothing else, read these. Primary sources first — the engines' own documentation beats every summary of it.
Rand Fishkin
Clickstream-based analysis finding that for every 1,000 Google searches only about 360 clicks reach the open web in the US and about 374 in the EU, with the remainder ending in no click or a click to a Google-owned property.
Google Search Central
Google's documentation on how AI Overviews and AI Mode source content, the eligibility requirements, and the preview controls site owners can use to limit inclusion.
Jon M. Kleinberg
Kleinberg's paper introducing the hubs-and-authorities model of link analysis, developed contemporaneously with PageRank and still the reference point for topical authority arguments.
Jacob Devlin, Ming-Wei Chang, Kenton Lee, Kristina Toutanova
The model Google deployed in Search in 2019 to better interpret query context and prepositions. The technical basis for the shift away from keyword matching toward passage-level meaning.
Microsoft Bing
Bing's published expectations for content, technical implementation and link acquisition, plus the practices Bing treats as abuse. The main non-Google engine policy document in English.
Common Crawl Foundation
Free, openly licensed petabyte-scale archive of web crawl data with monthly releases of WARC, WAT and WET files plus columnar indexes. Heavily used for link graph research and LLM pretraining.
Google Search Central
Google's self-assessment questions on content quality, expertise and originality. Includes Google's public explanation of E-E-A-T and of who, how and why in content creation.
Anthropic
Anthropic's official help article describing ClaudeBot, Claude-User and Claude-SearchBot, their robots.txt tokens and the crawling practices Anthropic follows.
US Federal Trade Commission
The FTC's practical guidance on disclosing material connections in endorsements, affiliate content, sponsored posts and reviews. The governing US rulebook for affiliate and influencer content in search results.
Pranjal Aggarwal et al.
The paper that named generative engine optimization, proposing a benchmark and testing which content modifications increase a source's visibility inside generative search responses.
Authoritative list of Google's crawlers, fetchers and user-triggered agents, including Googlebot, Google-Extended (AI training and grounding opt-out), GoogleOther and their robots.txt tokens and IP ranges.
Google Search Central
Google's official channel for announcing algorithm updates, documentation changes, new structured data types and crawler behaviour changes. The primary-source record of Google Search announcements since 2005.
Rand Fishkin
Clickstream-based analysis finding that for every 1,000 Google searches only about 360 clicks reach the open web in the US and about 374 in the EU, with the remainder ending in no click or a click to a Google-owned property.
Ahrefs
Analysis of referral traffic from AI assistants and answer engines across thousands of sites, quantifying how much measurable traffic AI surfaces currently send.
Ahrefs
Large-scale study of pages in the Ahrefs index showing the extreme skew of organic traffic distribution and the correlation between referring domains and any traffic at all.
Apple
Apple's support article covering Applebot and Applebot-Extended, the user agents behind Siri, Spotlight Suggestions and Apple Intelligence training data, and the robots.txt controls for each.
Advanced Web Ranking
Regularly refreshed CTR curves segmented by position, device, country, query length and SERP feature presence, drawn from a large Search Console sample.
Ahrefs
Ahrefs' free training hub teaching how to apply its crawler, backlink index and keyword data to practical digital marketing workflows.
Ahrefs
Free structured beginner video course covering how search engines work, keyword research, on-page SEO, link building and technical fundamentals.
Ahrefs
Continuously updated compilation of sourced SEO statistics drawn from Ahrefs' own studies and third-party research, useful as a citation index rather than as primary research.
Ahrefs
Free guide to link acquisition tactics, prospecting, outreach and link evaluation, including which link types Ahrefs' own studies find correlate with rankings.
Ahrefs
Ahrefs' free long-form SEO guide covering keyword research, on-page, link building, technical SEO and measurement, with worked examples from its own data.
Google Search Central
Google's documentation on how AI Overviews and AI Mode source content, the eligibility requirements, and the preview controls site owners can use to limit inclusion.
Ahrefs
Ahrefs study comparing average CTR for top-ranking pages on queries with and without an AI Overview, one of the most-cited quantifications of AI Overview click loss.
ai-robots-txt project
Community-maintained open-source list of AI-related crawlers with ready-made robots.txt, .htaccess, nginx, Caddy, HAProxy and JSON outputs for blocking or auditing them.
Spawning AI
Spawning's proposed ai.txt convention, a robots.txt-style file letting site owners state opt-in and opt-out preferences for AI training use of their media and text.
Rand Fishkin
The post that made the leaked Google Content Warehouse API documentation public, and contrasted the attribute names found in it with Google's prior public statements.
Matt Cutts
Google's April 2012 announcement of the webspam algorithm update later named Penguin, which targeted link schemes and keyword stuffing and reshaped the link building industry.
Advertising Standards Authority (UK)
The UK non-broadcast advertising codes governing marketing communications, including online claims, affiliate disclosure and influencer marketing. Enforced by the ASA.
Ashish Vaswani et al.
Introduced the Transformer architecture, which underpins BERT, MUM, T5 and the large language models that now generate AI search answers.
Jon M. Kleinberg
Kleinberg's paper introducing the hubs-and-authorities model of link analysis, developed contemporaneously with PageRank and still the reference point for topical authority arguments.
Backlinko
Free reference hub of SEO guides, definitions and data-backed studies, known for large-sample analyses of search results and CTR.
Baidu
Baidu's webmaster platform documentation and announcements, covering submission APIs, crawling behaviour and the algorithm updates Baidu publishes for Chinese-language sites.
Jacob Devlin, Ming-Wei Chang, Kenton Lee, Kristina Toutanova
The model Google deployed in Search in 2019 to better interpret query context and prepositions. The technical basis for the shift away from keyword matching toward passage-level meaning.
Microsoft Bing
Bing's published expectations for content, technical implementation and link acquisition, plus the practices Bing treats as abuse. The main non-Google engine policy document in English.
Microsoft Bing
Product documentation for Bing Webmaster Tools, including site verification, URL submission, crawl control, the Site Explorer and the IndexNow integration.
Google Search Central
How to use the noindex rule via meta tag or X-Robots-Tag header, and why disallowing a URL in robots.txt prevents the rule from being seen.
Brave
Brave's help page on its independent search index crawler. Notes that the crawler does not advertise a differentiated user agent, respects Googlebot blocks, and that delisting requires a robots noindex directive rather than robots.txt.
BrightEdge
Enterprise-panel research on channel share, SERP feature prevalence and AI Overview behaviour, drawn from BrightEdge's Data Cube of customer sites.
BrightLocal
The longest-running annual consumer survey on how people use online reviews when choosing local businesses, covering review volume, recency, star-rating thresholds and response expectations.
Donald Miller
Messaging framework applying narrative structure to marketing copy, commonly used to shape landing page and homepage messaging alongside search-driven content.
Common Crawl
Common Crawl's page describing CCBot, the crawler that builds the open web corpus widely used to pretrain large language models, and how to allow or block it in robots.txt.
Google Chrome team
Documentation for the CrUX dataset of real-user Core Web Vitals from opted-in Chrome users, covering the API, BigQuery tables, the CrUX History API and dashboard tooling.
Google Chrome team
The public BigQuery dataset containing monthly Core Web Vitals distributions for millions of origins, queryable for competitive performance benchmarking at scale.
Cloudflare
Cloudflare's documentation on bot classification, verified bots and the controls network operators use to allow, challenge or block AI crawlers at the edge.
Lemur Project / Carnegie Mellon University
Large research web corpus derived from commercial search engine crawls, distributed under licence for academic information retrieval and web search research.
Common Crawl Foundation
Free, openly licensed petabyte-scale archive of web crawl data with monthly releases of WARC, WAT and WET files plus columnar indexes. Heavily used for link graph research and LLM pretraining.
Conductor
Free enterprise-oriented SEO and answer engine optimisation learning library covering organic strategy, content operations and measurement.
Andy Crestodina
Illustrated content marketing handbook, currently in its 7th edition, split into a theory section on traffic and conversion and a practical lab section on content creation and promotion.
Content Marketing Institute
Annual B2B and B2C content marketing benchmarks, budgets and trends surveys, the long-standing reference for content team staffing and investment norms.
Google's explanation of crawl capacity limit and crawl demand, and the practical levers large sites have to influence how much Googlebot crawls.
Google Search Central
Requirements for Google News sitemaps, including the news-specific XML namespace, the two-day article window and publication metadata.
Google Search Central
Google's self-assessment questions on content quality, expertise and originality. Includes Google's public explanation of E-E-A-T and of who, how and why in content creation.
Google Chrome team
Definition of CLS, the visual stability Core Web Vital, covering session windows, expected versus unexpected shifts and the layout shift score calculation.
CXL
Paid practitioner-taught training in conversion optimisation, growth marketing, analytics and technical marketing, with minidegree programmes that include search and content tracks.
Google Search Central
Google's structured diagnostic for Search traffic declines, distinguishing technical issues, manual actions, algorithmic changes, seasonality and reporting glitches.
Vladimir Karpukhin et al.
Showed that learned dense embeddings can outperform BM25 for first-stage retrieval. The technical basis for vector search and for retrieval inside AI answer engines.
Pew Research Center
Independent behavioural study of a representative US panel finding users are less likely to click any result when an AI summary is present. Valuable as non-vendor evidence on AI Overview click behaviour.
Google Search Central
Google's advice for site owners evaluating SEO providers, including due-diligence questions and warning signs of unsafe practices.
Anthropic
Anthropic's official help article describing ClaudeBot, Claude-User and Claude-SearchBot, their robots.txt tokens and the crawling practices Anthropic follows.
Google's official help page mapping the former Analytics Academy curriculum onto the current Skillshop learning paths for beginner, intermediate and advanced GA4 users.
Sowmya Subramanian
Google's announcement of the page experience signal and the introduction of Core Web Vitals as measurable components of it.
Ann Handley
Practical writing handbook for marketers covering process, style, structure and content types. First published 2014; a revised second edition was later released.
Danny Sullivan, Gary Illyes
Introduced rel=sponsored and rel=ugc and reclassified nofollow as a hint rather than a directive, changing how paid and user-generated links are marked up.
Robby Stein
Google's announcement of AI Mode, a conversational, query-fan-out search experience, alongside the wider rollout of AI Overviews.
Colin Raffel et al.
The text-to-text framework Google cites as the architecture underlying MUM. Relevant to understanding multitask, multimodal and multilingual query handling in modern search.
US Federal Trade Commission
The FTC's practical guidance on disclosing material connections in endorsements, affiliate content, sponsored posts and reviews. The governing US rulebook for affiliate and influencer content in search results.
gdpr.eu
Reference resource covering the GDPR text and practical compliance requirements. Directly relevant to consent management, analytics measurement and personalisation in search marketing.
Pranjal Aggarwal et al.
The paper that named generative engine optimization, proposing a benchmark and testing which content modifications increase a source's visibility inside generative search responses.
The Google Analytics learning catalogue on Skillshop, including the Google Analytics Certification exam and beginner through advanced GA4 courses. Free to take; the certification expires and must be renewed.
Authoritative list of Google's crawlers, fetchers and user-triggered agents, including Googlebot, Google-Extended (AI training and grounding opt-out), GoogleOther and their robots.txt tokens and IP ranges.
Google Search Central
Hub of Google's technical documentation on how Googlebot discovers, crawls and indexes pages, covering robots.txt, sitemaps, canonicalization, redirects, mobile sites and JavaScript.
Liz Reid
Google's announcement of the general rollout of AI Overviews in the United States along with multi-step reasoning and planning features in Search.
Google Search Central
Google's official channel for announcing algorithm updates, documentation changes, new structured data types and crawler behaviour changes. The primary-source record of Google Search announcements since 2005.
Google Search Central
Google's official video channel with SEO office hours, Lightning Talks, the Search Console training series and JavaScript SEO explainers from Google engineers and search advocates.
Official product help for Search Console, covering property types, the Performance report data model, Index Coverage states, URL Inspection and manual actions.
Google Search Central
The successor to the Google Webmaster Guidelines. Defines the technical requirements, spam policies and key best practices a site must meet to be eligible for Google Search. Continuously updated.
Google's authoritative log of confirmed ranking updates with official start and completion timestamps, used to align traffic changes with named algorithm rollouts.
Google's official free training and certification platform covering Google Ads, Google Analytics, Google Marketing Platform and related products. The certifications are the only Google-issued marketing credentials.
Detail page for Google's common crawlers, listing user-agent strings, robots.txt tokens and reverse-DNS verification for Googlebot variants.
The rules governing Google Business Profile listings: eligibility, naming, categories, service areas and prohibited practices. The foundational policy document for local SEO.
Google's overview of generative AI in Search, describing the shift from a ranked list of links toward synthesised answers with supporting sources.
Maps each HTTP status code, network error and DNS failure to the behaviour Google's crawlers and indexing systems apply.
Google Search Central
Explains how Google consolidates duplicate URLs, which canonicalization signals it uses and how to troubleshoot unexpected canonical selection.
Google Search Central
Task-based guide to Search Console for common roles, linking to property verification, performance reporting, indexing reports and URL inspection.
Google Search Central
Google's guidance on snippet generation, meta description quality and the directives that limit snippet length or suppress snippets entirely.
WHATWG
The living HTML specification section defining head, title, base, link, meta and the standard metadata names, which is the normative basis for title tags and meta descriptions.
WHATWG
The living HTML specification section defining the a and link elements, the href attribute and every standard link relation including nofollow, sponsored, ugc, canonical and alternate.
HTTP Archive
Longitudinal trend reports on how the web is built and performs, backed by a public BigQuery dataset of monthly crawls of millions of URLs.
Zineb Ait Bahajji, Gary Illyes
Google's confirmation that HTTPS is a lightweight ranking signal, which triggered industry-wide migration to TLS.
HubSpot
Free HubSpot Academy courses and certifications on SEO, topic clusters and content strategy, now extended with answer engine optimisation material.
IAB Tech Lab
The technical standards library for digital advertising, including ads.txt, sellers.json, OpenRTB and transparency frameworks relevant to publisher monetisation alongside organic search.
IETF
IETF working group chartered to standardise a vocabulary and attachment mechanism for expressing preferences about automated processing of web content by AI systems. The formal successor track to ad hoc conventions such as ai.txt.
IndexNow
Open protocol, backed by Microsoft Bing and Yandex, that lets sites push URL change notifications to participating search engines instead of waiting to be recrawled.
Google Search Central
Explains how Google generates the title link shown in results, which sources it draws on and why the title element is sometimes rewritten.
Louis Rosenfeld, Peter Morville, Jorge Arango
The standard text on organising, labelling, navigating and searching information environments. Directly relevant to site structure, taxonomy design and internal linking. Multiple editions since 1998.
Stefan Büttcher, Charles L. A. Clarke, Gordon V. Cormack
Engineering-focused IR textbook emphasising index construction, compression, query evaluation efficiency and rigorous evaluation methodology. ISBN 9780262026512 (hardcover), 9780262528870 (paperback).
Google Chrome team
Definition of INP, the responsiveness metric that replaced First Input Delay as a Core Web Vital, including input delay, processing time and presentation delay components.
Google Search Central
Google's primer on structured data formats (JSON-LD, microdata, RDFa), placement, testing and the relationship between markup and rich result eligibility.
Google Search Central
Reference for Product, Offer, AggregateRating and merchant listing markup used for product rich results and Google Shopping surfaces.
Christopher D. Manning, Prabhakar Raghavan, Hinrich Schütze
The standard graduate IR textbook, covering inverted indexes, tolerant retrieval, vector space models, probabilistic ranking, link analysis and evaluation. The full text is available free online from Stanford NLP. ISBN 9780521865715.
W3C
The W3C Recommendation defining JSON-LD, the serialization format Google recommends for structured data. Specifies @context, @type, @id, framing and expansion.
JSON-LD Community
Community home for JSON-LD with the playground, tooling list and specification links. Useful for debugging structured data before deploying it.
Google Chrome team
Definition and measurement methodology for LCP, including which elements qualify, the good/needs-improvement/poor thresholds and the sub-part breakdown for optimisation.
Google Chrome team
Structured learning path on web.dev collecting the Core Web Vitals metric definitions, measurement tooling and optimisation techniques into a single course.
Google Chrome team
Free course covering web performance fundamentals from HTML and CSS delivery through image, font and JavaScript optimisation, with measurement in lab and field.
Aleyda Solis
Free curated SEO learning roadmap organising vetted third-party guides, videos, tools and courses by topic and experience level.
Google Chrome team
Official documentation for Lighthouse, the open-source lab auditing tool that scores performance, accessibility, best practices and SEO, and explains how each audit is computed.
Subscription video library of SEO courses spanning foundations, keyword strategy, technical SEO, local SEO and ecommerce, with completion certificates.
Google Search Central
Google's canonical reference for hreflang annotations across HTML, HTTP headers and sitemaps, including x-default and common implementation errors.
Rand Fishkin
The Moz founder's account of building and running a venture-backed SEO software company, covering fundraising, metrics, pricing and content-led growth. ISBN 9780735213326.
Lumar
Free technical SEO and website intelligence learning resources from Lumar (formerly Deepcrawl), covering crawling, site health, log analysis and enterprise SEO governance.
Mozilla
Practical reference for every HTTP status code with browser and crawler implications. The most commonly consulted working reference when auditing redirect and error handling.
Mozilla
Reference for the HTML link element and its rel values, including canonical, alternate, preload and preconnect, with attribute-level browser support notes.
Google Search Central
Complete list of meta tags and HTML attributes Google Search recognises, including google-site-verification, notranslate, nositelinkssearchbox and data-nosnippet.
Meta
Meta's developer documentation for its crawlers, including facebookexternalhit for link previews and meta-externalagent, and the robots.txt directives that control them.
Ricardo Baeza-Yates, Berthier Ribeiro-Neto
Comprehensive IR reference covering retrieval models, web search, crawling, indexing, user interfaces and evaluation. Second edition companion site hosts chapter slides and errata.
Moz
Moz's on-demand training platform with courses and certificates in SEO essentials, keyword research, technical SEO, local SEO and link building. Mix of free and paid tracks.
Moz
Moz's long-running ranking factors resource, historically built from practitioner surveys and correlation analysis. The reference point for how the industry has framed ranking factors since 2005.
Moz
Moz's free chaptered guide to link building strategy, tactics, outreach and measurement, including guidance on manual penalties and disavowal.
Moz
Moz's explainer and survey-derived summary of the signals that influence local pack and localised organic rankings.
Microsoft
Large-scale dataset of real anonymised Bing queries with human-generated answers and passage relevance labels. The standard benchmark for passage ranking and dense retrieval research.
Pandu Nayak
Google's introduction of the Multitask Unified Model, built on the T5 text-to-text architecture, trained across 75 languages and able to understand and generate language across modalities.
Naver
Naver's official webmaster guide for the dominant South Korean search engine, covering site registration, collection rules, markup and content policy.
Rand Fishkin
Clickstream analysis quantifying Google search volume growth against ChatGPT usage, widely cited as a counterweight to predictions of rapid search volume decline.
Google Search Quality team
The post announcing the addition of Experience to Expertise, Authoritativeness and Trustworthiness, and clarifying that Trust is the central member of the E-E-A-T family.
OpenAI
OpenAI's official documentation of GPTBot (model training), OAI-SearchBot (search index) and ChatGPT-User (user-initiated fetches), with robots.txt tokens and published IP ranges.
Google's public tool combining CrUX field data with a Lighthouse lab run for a single URL or origin. The standard reference point when discussing Core Web Vitals with stakeholders.
Rodrigo Nogueira, Kyunghyun Cho
Demonstrated large gains from applying BERT as a re-ranker over retrieved passages on MS MARCO. Foundational to passage ranking as deployed in modern search systems.
Perplexity AI
Perplexity's documentation of PerplexityBot and Perplexity-User, explaining which agent indexes content for its answer engine and which fetches pages on behalf of a user request.
Eli Schwartz
Argues that durable organic growth comes from building products and pages that serve genuine user demand rather than from following formulaic SEO checklists. Written for in-house practitioners and executives setting SEO strategy.
Google Search Central
Covers permanent and temporary server-side redirects, meta refresh, JavaScript redirects and how each is interpreted for canonicalization and site moves.
Google Search Central
Decision guide for temporary URL removals, permanent removals, outdated content removal and refresh requests in Search Console.
Patrick Lewis et al.
Introduced RAG, the retrieve-then-generate pattern behind AI Overviews, ChatGPT search and Perplexity. Essential for understanding why source retrievability drives AI citation.
T. Berners-Lee, R. Fielding, L. Masinter
The specification governing URL structure, percent-encoding, relative reference resolution and normalization. The basis for correct canonical URL and parameter handling.
A. Barth
The cookie specification. Relevant to SEO through cookie-based personalization, consent gating and crawler behaviour, since Googlebot does not persist cookies between requests.
M. Nottingham
Defines the Link HTTP header and the registry of link relation types, including canonical, alternate, prev and next. Obsoletes RFC 5988.
R. Fielding, M. Nottingham, J. Reschke
The current authoritative definition of HTTP methods, status codes, and header fields. The primary reference for correct use of 301, 302, 304, 404, 410 and 503 in SEO work.
R. Fielding, M. Nottingham, J. Reschke
Defines Cache-Control, ETag, Last-Modified and conditional requests. Underpins crawl efficiency, CDN behaviour and correct handling of If-Modified-Since by crawlers.
M. Koster, G. Illyes, H. Zeller, L. Sassman
The IETF Proposed Standard that finally formalised robots.txt after 25 years of de facto use, defining syntax, matching rules, caching behaviour and error handling.
Google's official validator for structured data eligibility, showing the rendered HTML Googlebot sees and which rich result types a page qualifies for.
Google Search Central
Specification of the robots meta tag and X-Robots-Tag HTTP header directives Google supports, including noindex, nofollow, nosnippet, max-snippet, max-image-preview and noai-adjacent controls.
Google Search Central
Google's reference for what robots.txt can and cannot do, syntax supported by Googlebot, and the interaction between disallow rules and indexing.
Google Search Central
The April 2015 announcement of mobile-friendliness as a ranking signal on mobile search, the update the trade press called Mobilegeddon.
Schema.org
The schema.org-hosted validator for generic structured data correctness, independent of whether a type is eligible for a Google rich result.
Schema.org
Introductory walkthrough of schema.org markup covering item types, properties, nesting and the differences between microdata, RDFa and JSON-LD encodings.
Schema.org (Google, Microsoft, Yahoo, Yandex)
The shared structured data vocabulary maintained by the major search engines and a W3C community group. The authoritative type and property reference behind almost all rich results.
Search Engine Journal
Annual practitioner survey covering budgets, team structures, tooling, priorities and the perceived impact of algorithm and AI-search changes.
Search Engine Journal
Search Engine Journal's free introductory SEO guide and topic hub, updated as the discipline changes and linked to its ebook library.
Search Engine Land
Search Engine Land's evergreen explainer and guide hub defining SEO, its disciplines and how the trade press categorises ranking factors and search updates.
Coursera
University-affiliated multi-course specialization on Coursera covering SEO fundamentals, on-page and off-page optimisation and an applied capstone project. Certificate available on paid enrolment.
W. Bruce Croft, Donald Metzler, Trevor Strohman
University textbook on how search engines are actually built: crawling, indexing, query processing, ranking models and evaluation. Companion site hosts slides and the Galago toolkit. ISBN 9780136072249.
The PDF handbook Google gives to its external search quality raters. Defines Page Quality, Needs Met, E-E-A-T, Your Money or Your Life topics and lowest-quality signals. The closest public statement of what Google considers a good result.
Mike King
The most detailed technical analysis of the leaked Google Content Warehouse API documentation, covering NavBoost, click signals, site authority, twiddlers and document scoring modules.
Semrush
Free course library and exam-based certifications covering SEO, PPC, content marketing, competitive research and the Semrush toolset itself.
Brian Dean
Free multi-lesson video course on keyword research, on-page optimisation, link building and content strategy, taught by the founder of Backlinko.
Google Search Central
Google's specialty documentation for online stores, covering URL structure for facets, pagination, product data, out-of-stock handling and merchant listings.
Bill Slawski
The archive of Bill Slawski's patent analyses, which for two decades was the primary public source translating Google and Bing patents into SEO implications. Slawski died in 2022; the archive remains a reference body of work.
John Jantsch, Phil Singleton
Introductory book positioning SEO within a broader marketing system for small businesses and agencies. ISBN 9780692769447.
Google Search Central
Google's own introductory guide to search engine optimization, covering how Google finds and understands pages, how to influence appearance in results and which practices are unnecessary.
Google Search Central
Requirements for influencing the site name displayed above results, including WebSite structured data and other supported signals.
sitemaps.org
The jointly supported sitemap specification from Google, Yahoo and Microsoft. Defines the XML schema, the 50,000 URL and 50 MB limits, sitemap index files and robots.txt autodiscovery.
Google Search Central
Google's enumeration of behaviours that can trigger ranking demotion or removal from Search, including cloaking, doorways, scaled content abuse, site reputation abuse, expired domain abuse and link spam.
Google Search Central
The search gallery listing every structured data type that can produce a rich result in Google Search, with required and recommended properties for each.
Jeremy Howard
Proposal for a markdown file at /llms.txt that gives large language models a curated, inference-time map of a site's most useful content. A community convention rather than a ratified standard; adoption by major AI providers remains limited.
HTTP Archive
The full 2024 edition of HTTP Archive's annual state of the web report, spanning performance, markup, media, security, CMS, accessibility and SEO chapters.
Sergey Brin, Lawrence Page
The paper that introduced Google, describing PageRank, anchor text as a ranking signal, the crawling and indexing architecture and the repository design. The single most important document in SEO's intellectual history.
Eric Enge, Stephan Spencer, Jessie Stricchiola, Rand Fishkin
The most comprehensive general SEO textbook, first published by O'Reilly in 2009 (ISBN 9780596518868) and revised across multiple editions. Covers search engine mechanics, strategy, keyword research, content, links, measurement and organisational issues.
Moz
One of the longest-running free introductory SEO guides on the web, covering crawling and indexing, keyword research, on-page and technical optimisation, link building and measurement.
ogp.me
Specification for the og: meta properties that control how URLs are rendered when shared on social platforms and in many chat and AI interfaces.
HTTP Archive
Annual, peer-reviewed analysis of SEO implementation across millions of real websites, measuring adoption of canonical tags, structured data, hreflang, robots directives and rendering patterns.
Marcus Sheridan
Content marketing framework built on answering customers' real buying questions, including pricing, comparisons, problems and reviews. Widely used as the editorial basis for bottom-of-funnel SEO content programmes.
Search Engine Land / Semrush
Paid SEO community and training programme combining an on-demand course library with a private practitioner community. Now operated as a Search Engine Land community.
GOV.UK
The UK regulator whose digital markets work covers search engine competition, self-preferencing and online choice architecture, including designations under the digital markets competition regime.
Google Search Central
Describes Google's three-stage crawl, render and index pipeline for JavaScript applications and the practices that make client-rendered content indexable.
Google Search Central
Google's current position on page experience signals, including Core Web Vitals, HTTPS and intrusive interstitials, and how they relate to ranking.
Pandu Nayak
Google's announcement of BERT in Search, described at the time as the biggest leap forward in five years, affecting roughly one in ten English queries.
Lawrence Page
The original PageRank patent, assigned to Stanford and licensed exclusively to Google. Describes ranking nodes in a linked database by recursively weighting inbound links by the rank of the linking node.
Google Inc.
The historical data patent, describing document inception dates, content update rates, link growth rates, anchor text change and domain registration signals as scoring inputs. Filed December 2003, published March 2008.
Google Inc.
Part of Anna Patterson's phrase-based indexing family. Identifies meaningful phrases and uses co-occurrence — which phrases predict the presence of other phrases — to rank and personalise results. Filed July 2004, published August 2009.
Google Inc.
Known in the industry as the reasonable surfer patent. Describes weighting links by the probability a user would click them, based on features such as position, font, anchor text and link context. Filed June 2004, issued May 2010.
Google Inc.
Describes using aggregated usage statistics, such as how often documents are selected and how long users engage with them, as inputs to document retrieval and scoring.
Google Inc.
Frequently cited in analyses of site-level quality scoring and of query-and-click based ranking modifiers. Filed September 2012, issued March 2014.
Yandex
Yandex's robots.txt reference, including directives such as Clean-param and Host that are specific to the Yandex robot.
Brian Dean
Correlation study across millions of results examining backlinks, content length, page speed, HTTPS, schema and other factors against ranking position. Correlational, not causal.
Brian Dean
Widely cited organic click-through-rate study by search position, including the effect of title length, question titles and power words on CTR.
Semrush
Large-sample analysis of AI Overview appearance by query type, length, intent and citation source, one of the primary datasets on which AI Overview optimisation advice is based.
Semrush
Semrush research quantifying how AI search surfaces change organic traffic patterns and the relative value of visits arriving from AI assistants.
W3C
The WCAG version referenced by most current accessibility legislation, including the European Accessibility Act and many public-sector procurement rules.
W3C
W3C Recommendation defining testable accessibility success criteria at levels A, AA and AAA. Overlaps heavily with SEO through semantic markup, alt text, heading structure and link text.
Google Chrome team
Canonical definition of the Web Vitals initiative and the Core Web Vitals set, with the current thresholds and the field-versus-lab measurement model.
Reiichiro Nakano et al.
OpenAI's paper on training a model to browse the web and cite sources when answering. An early blueprint for the browsing and citation behaviour of current AI answer engines.
Google Search Central
Google's guidance on when XML sitemaps help, the supported formats, sitemap index files and submission methods.
DuckDuckGo
DuckDuckGo's explanation of its result sources, including its own DuckDuckBot crawler, Bing, and other partners. Clarifies why there is no separate DuckDuckGo submission process.
Moz
Moz's long-running weekly video series, begun by Rand Fishkin, explaining SEO concepts at a whiteboard. Collectively one of the most-cited free educational bodies of work in the industry.
Whitespark
Annual expert survey ranking the relative importance of local pack and local organic ranking factors, including Google Business Profile signals, reviews, citations and links.
SISTRIX
SISTRIX's large-scale CTR analysis showing how SERP features, query intent and result layout distort the classic position-to-CTR curve.
Wikimedia Foundation
Complete periodic exports of Wikipedia and sister projects, including page content, pagelinks and pageview data. A standard corpus for entity, knowledge graph and query understanding work.
Yandex
Yandex's official English-language help for webmasters, covering site indexing, Yandex Webmaster tools, quality requirements and the behaviour of Yandex's robot.
Yoast
Training platform focused on WordPress-centric SEO, with courses on SEO copywriting, technical SEO, structured data and site structure. Some courses are free, others bundled with Yoast SEO Premium.