What robots.txt Does on a Live Website
robots.txt is a file that tells crawlers which URLs they may fetch. Google says it manages crawl. It does not reliably keep an HTML page out of search results. A disallowed URL can still be indexed without a snippet if something else links to it [1]. To hide a page, use noindex or a password [1].
What should you not block?
CSS and JavaScript the page needs to render [1]. If the crawler cannot fetch those files, it cannot see the page the way a visitor does. The same mistake applies to an image folder you hoped to "save crawl" on. Google indexes images from an img src, and it cannot index a file it is not allowed to fetch [2].
Why does a disallow break noindex?
noindex works only when Google crawls the page and reads the rule, either in a meta tag or an X-Robots-Tag header. If robots.txt blocks that URL, Google never sees the noindex, and the URL can still appear [3]. A noindex line inside robots.txt is not supported [3]. Allow the crawl, then say noindex, if the goal is to drop the URL from results.
How long does a listed URL take to disappear?
It can take months after Google sees the noindex [3]. URL Inspection is how you ask for a recrawl. The request is not a delete button [3]. A disallow added on top of a page that is already indexed does not speed that up. It can prevent the noindex from being seen at all [3].
What about a staging copy?
Google's hosting-move guide says the temporary hostname should be noindex, and you must remove noindex and robots blocks before the real DNS cut or the live site stays invisible [4]. A demo left crawlable is also one of the ways Google says duplicate URLs appear [5]. Pick one: the copy is noindex and fetchable, or it is gone. A disallow is the option that does neither job cleanly [1].
Does the sitemap override the file?
No. A sitemap helps discovery and does not guarantee that a URL will be crawled or indexed [6]. Listing a disallowed URL does not grant a fetch. Listing a URL you want indexed, while robots.txt blocks the CSS it needs, still leaves the crawler half blind [1]. The sitemap and the robots file have to agree.
What if the blocked URL is a 404 anyway?
Then the robots rule is not the interesting part. A 4xx other than 429 means the content is ignored, and indexed URLs that keep erroring are removed over time [7]. Do not disallow a deleted page and also leave it in the sitemap. Remove it from the sitemap. Let the 404 be the answer [6] [7].
Who edits the file on a small site?
Whoever deploys the site. A one-line disallow of / copied from a tutorial will block the whole public site. Read the file after every launch. Google's introduction is short on purpose: allow what the pages need, and do not use the file as a privacy tool [1].
- Allow CSS and JS — the page is not readable without them.
- Do not disallow a page you want hidden — use noindex or a password, and let the crawler see the noindex.
- Do not put noindex in robots.txt — it is not supported.
- Staging — noindex while it is a copy, then remove the block on the copy that goes live.
- Sitemap — do not list URLs you have disallowed.
Where does this sit next to the rest of the work?
It is a launch check inside website development, on the checklist beside the form and the phone number. Maintenance is when a later deploy quietly blocks /assets. Security is a different file: a password and an update, not a disallow.
What is the file not for?
It is not a lock on a private page, and it is not a way to keep a URL out of Google [1]. A password does the first job. noindex, on a page the crawler can fetch, does the second [3]. robots.txt only answers whether a crawler may request the URL [1]. Use it for that, and for nothing that needs to stay secret.
What should you read this week?
The live robots.txt, out loud. If it disallows a folder the homepage needs, remove that line [1]. If it disallows a page you hoped to hide, switch that page to noindex and allow the fetch [3]. If the file is empty and the site is public, you may not need a clever rule at all [1].
Sources & references
- Google Search Central: Introduction to robots.txt.
- Google Search Central: Google image SEO best practices.
- Google Search Central: Block search indexing with noindex.
- Google Search Central: Site moves without URL changes.
- Google Search Central: URL canonicalization.
- Google Search Central: Learn about sitemaps.
- Google Search Central: HTTP status codes and network errors.
Crawl rules are on the linked Google pages. Read the current robots.txt introduction before you change a live file.
If a page you wanted public is blocked, or a page you wanted hidden is still in results, we can read the robots file against the live URLs.
Talk to us arrow_right_alt