In short: A published Webflow site exists at two addresses: your domain and a .webflow.io subdomain. Both serve the same pages, and unless you turn the staging one off, both can be crawled.
Webflow has a setting for it. Turning staging indexing off publishes a robots.txt on the subdomain that blocks crawlers there, without touching your own domain. It needs a paid plan.
The wider lesson is the one underneath it: blocking a URL in robots.txt means a crawler never fetches it, which means it never reads the noindex tag on the page either.
Why Two Addresses Matter Here
The staging subdomain is genuinely useful: it gives you a working URL before a domain is connected and a place to check a build.
It is also a complete second copy of your content on a host you do not control the reputation of. Two consequences follow for retrieval.
The first is duplication. The same pages exist twice, so a system deciding which to draw on has two candidates that say the same thing, and nothing on either says which is the real one.
The second is worse and quieter: a citation can point at the staging address. An answer that names yoursite.webflow.io has credited you at a URL you would not print on anything, and it does nothing for the domain you are actually building.
The Setting
Webflow's own control is in the project's SEO settings, under indexing: turn staging indexing off, then save and publish. Doing so publishes a separate robots.txt on the .webflow.io subdomain that disallows crawlers, and leaves the robots.txt on your custom domain alone.
It is gated behind a paid site plan or workspace, which is worth knowing before you promise anybody it is a five minute job. Webflow's help centre carries the current path and the plan requirement, and both are more likely to change than the reason behind them.
Check the result rather than the checkbox:
curl -s https://yoursite.webflow.io/robots.txt
curl -s https://example.com/robots.txt
The two should not say the same thing. If they do, the setting has not taken effect, whatever the settings screen shows.
The Trap Worth Learning Beyond Webflow
There is a common instinct when something should not appear: block it in robots.txt and add a noindex tag, on the belief that two mechanisms are safer than one.
They work against each other. A noindex tag lives inside the page, so a crawler has to fetch the page to read it. Disallow the URL and the fetch never happens, the tag is never seen, and the instruction you meant to give is never delivered.
Pick one mechanism per URL. To keep something out of results while it stays reachable, use noindex and leave the URL crawlable so the tag can be read. To keep a crawler away from a path entirely, use robots.txt and accept that anything already known about it stays known. Doing both is how a page ends up neither blocked nor removed.
Which Address Are You Building
The practical rule is to decide which domain your work is accruing to, then make the other one silent.
Every link you place, every mention you earn and every citation you gather attaches to a host. Split across two, half the work builds an address that exists to be temporary. That is not a small loss on a site whose whole strategy is being named.
If the staging site has already been crawled, turning the setting on stops future crawling and does not remove what has been collected. Expect a lag, check what the subdomain serves now, and treat anything already in a model's training data as permanent, which is the asymmetry Managing AI Crawler Access describes.
Key Takeaways
- A Webflow site is published to both your domain and a .webflow.io subdomain, and both can be crawled unless you turn staging indexing off.
- The setting publishes a robots.txt on the subdomain only, leaves your domain untouched, and requires a paid plan.
- A citation pointing at the staging address credits a URL you would never print, and builds nothing for your domain.
- Blocking a URL in robots.txt prevents a crawler from ever reading a noindex tag on that page. Use one mechanism per URL, not both.
- Fetch both robots.txt files and compare them rather than trusting the settings screen.
Check yourself
Before you move on
Not scored, not recorded, and not part of the certificate. Both answers are settled by a sentence in this lesson, and the reasoning appears whichever option you pick.
- 01
Why does an answer that cites your .webflow.io address cost you something?
- 02
Why is blocking a URL in robots.txt and adding a noindex tag to the same page self-defeating?