Independent auditor of media and advertising data, verifying circulation, digital traffic and ad-supply data for publishers and buyers. Its audits are one of the few third-party verifications of publisher traffic claims, relevant when evaluating link and placement partners.
Organizations
Standards Body
Standards bodies define the protocols SEO is built on — robots exclusion, sitemaps, structured data vocabularies, HTTP semantics and accessibility.
19 organizations · Global and Europe
Coalition maintaining an open technical standard, Content Credentials, that records the origin and edit history of digital media. Provenance metadata is increasingly used by search and AI answer engines as a trust and authenticity signal for images and video.
Metadata standards organisation that arose from a workshop series in the 1990s and now operates as a project of ASIS&T. It maintains the Dublin Core Metadata Element Set, standardised as ISO 15836, which remains a foundational vocabulary for describing web resources and is reused inside Schema.org and library-sector SEO.
The single authoritative voice on advertising self-regulation in Europe, coordinating national self-regulatory organisations to keep advertising legal, decent, honest and truthful. Its cross-border complaint mechanism and guidance, including on generative-AI disclosure, apply to online and search advertising across European markets.
Non-profit standards consortium established in 2014 that develops the technical specifications underpinning digital advertising, including OpenRTB, ads.txt and privacy signalling. Its newer agentic and AI workstreams are defining how AI agents identify themselves to publishers, which overlaps directly with crawler management.
IETF working group chartered to standardise how site owners express preferences about the use of their content by AI systems, including a common vocabulary and a means of attaching those preferences to web content. It is the venue where the successor to ad-hoc AI crawler directives in robots.txt is being negotiated.
International standards body whose committees produce specifications relevant to web content, including ISO 15836 (Dublin Core metadata), ISO 639 language codes and ISO 3166 country codes. Those code lists are the exact values required in hreflang annotations and locale targeting, making ISO an unavoidable dependency of international SEO.
Global standards body for the news media that maintains photo and video metadata standards, the Media Topics subject vocabulary, NewsML-G2, ninjs and RightsML. Its metadata and rights vocabularies are the mechanism publishers use to declare provenance and usage rights to search engines and AI crawlers.
Open standards development organisation established in 1986 that produces the RFC series governing internet protocols. It standardised the Robots Exclusion Protocol as RFC 9309 in 2022 and hosts working groups on HTTP, TLS and machine-readable content preferences, all of which set the rules crawlers must follow.
Proposal published by Jeremy Howard on 3 September 2024 for a /llms.txt markdown file that gives language models a concise, curated map of a site's content, plus clean markdown versions of pages. It is the most widely discussed community attempt at a robots.txt equivalent for AI answer engines, though it has no formal standards-body backing.
US body that sets minimum standards for audience measurement and accredits measurement services through independent audits. Its accreditation defines what counts as a valid impression or viewable ad, which anchors how digital and search advertising performance is reported.
Open governance body formed in November 2015 under the Linux Foundation by ten founding companies including Google, Microsoft and SmartBear, now with more than 40 members. It maintains the OpenAPI Specification, the machine-readable API description format that increasingly determines whether an organisation's data is consumable by AI agents rather than only by human search users.
Open content licensing standard that lets publishers declare machine-readable licensing terms for AI use of their content, building on the robots.txt and RSS lineage. It is a direct successor to crawler-blocking approaches, moving from disallow rules to explicit licensing terms attached to crawlable content.
Collaborative community activity started around 2011 by Google, Microsoft, Yahoo and Yandex to create and maintain shared schemas for structured data on the internet, now coordinated through a W3C Community Group chaired by Dan Brickley with a steering group chaired by R.V. Guha. Its vocabulary is the basis for virtually all rich result eligibility in search.
Cross-vendor home of the Sitemap protocol, currently at version 0.90 and released under a Creative Commons Attribution-ShareAlike licence with support from Google, Yahoo and Microsoft. It is the only jointly maintained specification telling search engines which URLs a site wants crawled, and it has been effectively frozen since 2020.
Cross-industry programme that certifies advertising supply-chain participants against fraud, malware and transparency standards, and which absorbed the JICWEBS digital trading standards. Its certifications underpin the traffic-quality and brand-safety assurances used across paid search and display ecosystems.
Non-profit that maintains the Unicode Standard, CLDR locale data and emoji encoding. Its work determines how non-Latin scripts, normalisation and internationalised domain names behave, which directly affects international SEO, hreflang implementation and multilingual URL handling.
Community founded in 2004 by people from Apple, Mozilla and Opera that maintains the HTML, URL, Fetch and DOM Living Standards, governed since 2017 by a steering group of browser engine vendors. Because the HTML Standard defines document semantics and the URL Standard defines how addresses are parsed and normalised, its decisions directly govern canonicalisation, link handling and markup validity.
Public-interest non-profit founded by Tim Berners-Lee in 1994 that develops open web standards, with over 330 member organisations and more than 14,700 developer participants. Its specifications for HTML semantics, accessibility, internationalisation and structured data vocabularies define the substrate that crawlers parse and that SEO recommendations are written against.