In short: Squarespace does not let you edit robots.txt. What it offers instead is one checkbox that adds twenty-three AI crawlers to the file at once, including GPTBot, ClaudeBot, Google-Extended, Amazonbot and Applebot-Extended.
It is off by default, deliberately, and Squarespace says why: leaving it off preserves the traffic those systems can send.
And the platform is straight about the limit that matters. There is currently no universal way to be excluded from training while still being featured by the same company's assistant, so on this platform the decision really is one decision.
What The Checkbox Does
Squarespace's own documentation puts it in Settings under Crawlers, and lists the agents it adds: AI2Bot, Amazonbot, anthropic-ai, Applebot-Extended, ClaudeBot, Google-Extended, GPTBot, Meta-ExternalAgent, TikTokSpider and the rest, twenty-three in total.
Three caveats come with it, all stated by the platform rather than discovered by anybody: a request in robots.txt is only a request, it does nothing about content already collected, and a crawler that ignores the file will ignore it.
The Limit Worth Understanding
Elsewhere this course argues for a precise policy: name the agent whose job you object to, leave the one that answers live questions alone. That advice depends on the operator publishing separate agents for separate jobs, and Meet the AI Crawlers is where that separation is set out.
Squarespace's page names the boundary of it plainly: there is currently no universal way to ask to be left out of training while still appearing in the same company's assistant. Where an operator has not separated the two, the choice collapses into one.
That is worth reading as a correction rather than as a contradiction. The precise policy is available with some operators and not with all of them, and a platform saying so is being straighter than one that implies otherwise.
So the question is the commercial one, undiluted. Is your content the product, or is it how people find out you exist. On a platform with one lever there is no configuration that splits the difference, and the four questions in Managing AI Crawler Access are the whole of the decision.
Off By Default, And Why That Matters
The box starts unchecked and Squarespace says that is deliberate, to preserve the traffic these systems can send.
Two things follow. If you have never touched it, you are open, and nothing about your absence from answers is explained by this setting. And if you inherit a site where it is on, somebody decided that, possibly years ago and possibly without recording why.
Check which state you are in by reading the file rather than the screen:
curl -s https://example.com/robots.txt
Recording A Decision The Platform Cannot Record
The weakness of a checkbox is that it holds no reasoning. A robots.txt you write by hand can carry a dated comment; one generated from a setting cannot.
So write it somewhere else. One paragraph, dated, saying who decided, what the alternative was, and when it should be revisited. Put it wherever your team keeps decisions rather than in a person's memory, because the cost of getting this wrong is asymmetric: blocking acts quickly and reversing takes as long as the crawler takes to return.
The setting does not reach content already collected. Squarespace says so directly. Anything already in a trained model stays there until a future model is trained, which is the part of this decision that cannot be undone and the reason to make it slowly.
Key Takeaways
- Squarespace has no robots.txt editor. One checkbox adds twenty-three AI crawlers, including GPTBot, ClaudeBot and Google-Extended.
- It is off by default, deliberately, so an untouched site is open and this setting explains nothing about absence.
- There is currently no universal way to be excluded from training while still appearing in the same company's assistant, so here the choice is genuinely one choice.
- A checkbox holds no reasoning. Record the decision, the date and the review point somewhere your team keeps decisions.
- Nothing here touches content already collected, which is the part that cannot be reversed.
Check yourself
Before you move on
Not scored, not recorded, and not part of the certificate. Both answers are settled by a sentence in this lesson, and the reasoning appears whichever option you pick.
- 01
Squarespace offers one checkbox that adds twenty-three AI crawlers to robots.txt. What does that cost you compared with editing the file yourself?
- 02
Squarespace notes there is no universal way to be left out of training while still appearing in the same company's assistant. How does that sit with the advice to name specific agents?