Index / 4 entries
AI Crawler Reference
Four official-source crawler references that separate search discovery, model training, user-triggered requests, and Gemini product-use controls.
Inspect the Atlas dossier01
What this index contains
Four official-source crawler references that separate search discovery, model training, user-triggered requests, and Gemini product-use controls.
The collection is organized around ai crawler. Every linked record has its own search intent, source set, checked date, remediation sequence, validation protocol, false-positive notes, and claim boundary. Routes are published only after those fields pass the same content contract used by the sitemap and static renderer.
The index is deliberately finite. It does not create pages for word-order changes, unsupported diagnostic codes, cities, industries, or other template slots that lack independent evidence and a distinct reader task.
02
How to use the collection
Start with the page that matches the observed state or decision you need to make. Collect the named evidence before applying a fix, then keep the validation record beside the original observation so later reviewers can distinguish a resolved signal from an unmeasured outcome.
Use hubs for discovery and leaves for action. The hub explains collection boundaries; each leaf owns the technical details and links to related pages, a relevant checklist, the Atlas dossier, and the broader research archive.
- Observed evidence describes what the named crawl or provider returned.
- Derived interpretation explains a bounded conclusion from that evidence.
- Measurement gaps remain visible instead of being converted into success or failure claims.
- Publish-ready language never promises ranking, traffic, citations, or rich-result selection.
03
Coverage and provenance
This hub currently owns 4 leaf records from the pinned Project Delta and Atlas public knowledge snapshots. Content dates reflect source review; sitemap lastmod values change only when a record changes.
Atlas-backed records pin commit 669d951, knowledge bundle atlas.seo.knowledge.v0, and schema registry 2026.06.23. Internal fixture names, runtime owners, provider payloads, credentials, and client data are excluded from the public snapshot.
04 / Canonical routes
Choose the evidence or decision you need.
Each row is one canonical task, with no query-variation or location fan-out.
- 01 / AI crawler reference
OAI-SearchBot: Search Visibility and Access
OAI-SearchBot robots.txt guidance covering its ChatGPT search role, independent access control, published IP ranges, verification, and claim limits.
Reviewed 2026-07-20 - 02 / AI crawler reference
GPTBot: Training Control, Not Search Inclusion
GPTBot robots.txt guidance explaining OpenAI model-training controls, its separation from OAI-SearchBot, verification steps, and safe publisher claims.
Reviewed 2026-07-20 - 03 / AI crawler reference
ChatGPT-User: User-Triggered Page Requests
ChatGPT-User robots.txt guidance covering user-triggered requests, why it is not an automatic search crawler, access verification, and policy limits.
Reviewed 2026-07-20 - 04 / AI crawler reference
Google-Extended: Gemini Use and Grounding Control
Google-Extended robots.txt guidance covering Gemini training and grounding controls, its lack of a separate request agent, and Search boundaries.
Reviewed 2026-07-20