Loading...
Loading...
05Module
How GPTBot, ClaudeBot, PerplexityBot, and other AI crawlers discover and index your content. What they look for and how to ensure access.
Available3 lessons19 min
AI crawlers do three different jobs, and confusing them is how sites make themselves invisible. Training crawlers collect data for future models, live retrieval crawlers fetch your page because someone asked a question right now, and search crawlers index for an AI product. Blocking the first is a defensible business decision. Blocking the second removes you from answers being generated today.
This module is the practical floor. Nothing else in the course matters if a crawler cannot reach your page or cannot read it once it arrives, and both failures are more common and more mundane than the strategy conversation suggests.
It separates the three jobs AI crawlers do, which is the distinction that decides your policy: blocking a training crawler is a defensible business decision, and blocking a live retrieval crawler removes you from answers being generated today. Most robots.txt files that block AI agents were written without that distinction in mind.
It then treats access as the commercial decision it is, and closes on the technical work that makes a page usable once it is fetched: substance in the initial HTML, honest heading structure, and provenance attached to anything you want repeated.
If you read one lesson: Optimizing for AI Discovery, and run the curl check in its first section before anything else on this site.