Robots.txt: The File That Guides Crawlers

Robots.txt is a text file that tells search engine crawlers which parts of a site they may access. Learn the simple syntax and the mistakes that block whole sites.

Quick Definition

Robots.txt is a plain-text file at the root of a site that tells search engine crawlers which areas they may and may not access.

What Robots.txt Controls

Robots.txt controls crawling, not indexing. A rule here decides whether crawlers spend time on a URL, while separate directives like noindex decide whether it shows in results.

That distinction matters: blocking a page here usually keeps it out of the index, but a blocked page can still be indexed if it is linked to from elsewhere.

The Simple Syntax

The file is made of rules: a User-agent line names the crawler, and Disallow or Allow lines name the paths it may or may not visit.

A Sitemap line points crawlers at the XML sitemap, a useful convention even though it is not part of the classic protocol.

Common Mistakes

  • Disallowing CSS or JavaScript files, which can block the rendering of every page on the site.
  • Using Disallow instead of noindex to hide private pages, which leaks the intent while still allowing the pages to be found through links.
  • Shipping a rule meant for staging to the live site, quietly blocking entire sections.

Getting Started

  • Place the file at the root of the domain, like https://example.com/robots.txt.
  • Allow crawlers the resources they need to render pages, like CSS and images.
  • Add a Sitemap line pointing at your XML sitemap.
  • Test every rule before it goes live, and check the file again after site changes.
Example in Practice

The site: an agency adds "Disallow: /" to keep staging hidden and copies the file to production.

The reading: the live site vanishes from search overnight because crawlers are barred from the whole domain.

The fix: the file is corrected to allow the site and disallow only the admin and thank-you paths.

Why it works: one file had the power to remove the entire site, and reading it carefully brought it back.

💡

Quick Tip

Run the file through a robots tester after every edit; a single bad rule can erase weeks of indexing work.

Frequently Asked Questions

Robots.txt is a text file at the root of a site that tells search engine crawlers which parts of the site they may access.
Not directly. It stops crawling, which usually keeps a page out of the index, but a linked page can still be indexed from elsewhere.
Use a User-agent line followed by Allow or Disallow lines that name the paths crawlers may or may not visit.
At the root of the domain, like https://example.com/robots.txt.

Robots.txt, Bottom Line

Robots.txt is a small file with an oversized influence on how much of your site gets seen.

Write it to guide crawling, keep it honest, and always test it before it goes live.