In short: The second group of mistakes comes from doing a real skill harder in a place where more of it stops helping. They are made by capable people, which is what makes them hard to see.
Publishing more, writing for a phrase, treating markup as the work, and blocking access to protect content are all defensible instincts carried from search into a system that rewards something adjacent.
The correction is rarely to do less. It is to point the same effort at the step that is actually failing.
More Content, On The Same Question
Volume was a reasonable strategy when the surface had ten slots and a long tail beneath them. It is close to neutral where the answer names three sources, and it is negative when it dilutes.
Retrieval finds more material than fits, so most candidates are cut before merit is assessed. A comprehensive page can lose to a narrower one purely because the narrower one's answer is easier to locate and lift out intact. Every padded paragraph makes the useful passage harder to isolate.
The correction is not to publish less. It is to spend the same hours on making the existing answer findable and liftable, which is the subject of Optimizing for AI Discovery.
Writing For The Phrase Instead Of The Category
A deep habit, and it misses because the system does not search your phrase.
The reader asks one thing; the assistant decomposes it and searches in its own words, often several times, and one of those searches usually contains no brand at all. Optimising for the exact phrasing of the original question therefore matters far less than whether something already places you in the category that gets searched.
That work is mostly off your own site: being described as part of your category by other people, in places these systems read. It is slower than editing a page and it is the part that decides whether the page is ever fetched.
Schema Markup As The Work
Marking up entities is worth doing and it is a record of consistency rather than a source of it.
Marking up four different names for one company does not consolidate anything. It documents the inconsistency in a machine-readable format. No major AI system has confirmed schema as a direct input to answer selection, and it will not rescue a page that fails at extraction.
The order that makes it useful: pick one canonical name, use it in every position a machine reads, state your category in the words it is already known by, then mark up what you have just made consistent. Done in that order, markup earns its keep. Done first, it is a tidy description of a mess.
Blocking Everything To Protect The Content
The instinct is sound and the execution is usually one line too broad.
A team decides it does not want its content training future models, blocks every agent whose name contains a vendor string, and removes itself from live citations at the same time. Those are different decisions with different costs. Training collection, live retrieval and search indexing are three jobs, and one robots.txt line frequently makes all three at once.
The asymmetry is the part worth holding onto: blocking acts quickly and reverses slowly. Access restored takes as long as the crawler takes to return, and anything already absent from training data stays absent until a future model is trained. Managing AI Crawler Access is the business decision underneath it.
Auditing Everything Before Fixing Anything
Thoroughness misapplied, and it is the most expensive of the five in wall-clock time.
An extraction pass across four hundred pages is a project. Across the twenty that carry your commercial questions it is a week, and those twenty are the ones a category question would retrieve. The same applies to the passes themselves: if reachability fails, evidence about extraction is evidence about a hypothetical.
The tell that this is happening: a quarter has passed, the audit is comprehensive, and nothing has shipped. The audit was never the deliverable. The first failing step with its evidence, and a dated fix, was.
Key Takeaways
- Volume is roughly neutral and sometimes negative here, because selection cuts most candidates before merit and padding hides the useful passage.
- The system searches its own rewritten query, so being placed in the category by other people matters more than matching the original phrasing.
- Schema records consistency rather than creating it. Marking up four names for one company documents the problem.
- Blocking a vendor's whole name catches live retrieval agents alongside training crawlers, and blocking reverses far more slowly than it acts.
- Audit the twenty pages that carry the commercial questions, and stop at the first failing pass rather than completing the set.
Check yourself
Before you move on
Not scored, not recorded, and not part of the certificate. Both answers are settled by a sentence in this lesson, and the reasoning appears whichever option you pick.
- 01
Why does publishing more content on the same question help less here than it did in search?
- 02
A team marks up Organization schema for a company that uses four different names across its site. What has that achieved?