In short: A default WordPress site already passes most of the extraction test. The page arrives as HTML, the heading structure is real, and the author and date are in the template rather than assembled by script.
That is why WordPress sites usually fail for different reasons than JavaScript applications do. The problem is rarely the platform. It is a builder, a cache, or a setting somebody ticked years ago.
So the first hour is not implementation. It is confirming that the page a crawler receives is the page you think you published.
The Baseline You Start From
WordPress renders pages on the server. Whatever a theme produces arrives in the initial HTML response, which is the single property that decides whether retrieval can read a page at all.
Four other things come with it and each is work you do not have to do. URLs are real and stable. Headings come from the content rather than from styling choices. An author and a published date exist as fields and are usually printed by the theme. And archives, feeds and sitemaps are generated without anybody configuring them.
Compared with a single-page application, that is most of the technical pass already closed. It also means the remaining failures are less visible, because nothing about the site looks broken.
Confirm It Rather Than Assume It
The check is one command per page and it is the same one from Optimizing for AI Discovery:
curl -s https://example.com/your-page | grep -i "the sentence you want cited"
If the sentence is not there, it does not matter what the page looks like in a browser. Run it on your five most important pages before touching anything else.
The Three Things That Break The Baseline
A page builder that defers content. Builders produce the page from their own data, and most output it server-side. Some defer parts of it: tabs, accordions and anything that fetches when a section becomes visible. If the markup contains the text and CSS hides it, you are fine. If the text arrives on interaction, it is not there for a crawler.
A cache or CDN answering differently. Caching plugins and edge networks can serve a variant by user agent, by device, or from a stale entry generated before your last edit. The curl check tests exactly this, which is why it is worth running from outside your network rather than from a logged-in browser tab.
A theme that builds headings from styling. A heading that is a styled paragraph is not a heading to a parser. This is common in builder-made pages, where the visible hierarchy and the markup hierarchy are set independently.
View source, not the inspector. The browser inspector shows the page after scripts have run, which is the view a crawler may never get. On a WordPress site the two usually agree, and the cases where they disagree are precisely the ones you are looking for.
What Not To Install
There is no plugin that makes a site visible to AI systems, and the ones marketed that way generally do three things: add an llms.txt file, inject schema, and report a score.
The first is cheap and unconfirmed by any major platform. The second is a record of consistency rather than a source of it. The third is a number computed by a method you cannot inspect. None of them is the work, and The Mistakes Vendors Sell You covers why each one is attractive.
What is worth having is whatever you already use to control titles, canonicals and the head. That is ordinary SEO tooling and it was already doing most of this job.
Plugin names and menu paths change, so this course does not print them. Anything specific enough to be useful for a year is specific enough to be wrong in six months. Check the plugin's own documentation for where a setting lives, and use this module for what to look for rather than where to click.
The First Hour On A WordPress Site
- Fetch your five most important pages with curl and confirm the key sentence is in the response.
- Open the same pages in view source and check the heading levels are real and in order.
- Find any tab, accordion or "load more" block on those pages and confirm its text is in the source.
- Check that the author and date printed on the page are also in the markup rather than only in the design.
- Write down what you found before changing anything, because you will want to know which of these was already true.
Key Takeaways
- WordPress renders on the server, so a default site usually passes extraction. The failures come from builders, caches and old settings rather than from the platform.
- Confirm with curl rather than assuming. The browser inspector shows the page after scripts, which a crawler may never see.
- Content that arrives on interaction is not in the response. Markup that is present and hidden by CSS is fine.
- A cache or CDN can answer differently than your editor shows, which is exactly what the fetch test catches.
- No plugin makes a site visible to AI systems. The tooling worth having is what already controls titles, canonicals and the head.
Check yourself
Before you move on
Not scored, not recorded, and not part of the certificate. Both answers are settled by a sentence in this lesson, and the reasoning appears whichever option you pick.
- 01
Why do WordPress sites usually fail for different reasons than JavaScript applications?
- 02
An accordion on a WordPress page shows its text when clicked. When is that a problem?