← Glossary

What is robots.txt? Definition & Examples

Definition

robots.txt is a plain-text file located at the root of a domain (e.g., example.com/robots.txt) that tells web crawlers which URLs they are allowed or disallowed from requesting. It is part of the Robots Exclusion Protocol and is one of the first files a well-behaved crawler reads when visiting a site.

Common directives

  • User-agent: — names the crawler the rule applies to (* = all).
  • Disallow: — blocks a path from being crawled.
  • Allow: — explicitly permits a path inside a disallowed directory.
  • Sitemap: — declares the location of one or more XML sitemaps.
  • Crawl-delay: — requested pause between requests (honored by Bing, ignored by Googlebot).

Important caveats

  • robots.txt blocks crawling, not indexing — a disallowed URL can still appear in search results if other sites link to it. Use noindex to keep pages out of the index.
  • Malicious crawlers may ignore robots.txt entirely.
  • Blocking CSS or JavaScript files can break Google's ability to render and rank your pages.

robots.txt is a powerful but blunt instrument. Combine it with canonical tags, noindex directives and a clean XML sitemap for fine-grained crawl and index control.

Crawls your whole site, then keeps watching

Find broken links on your website

Detect dead links, missing images, and redirect loops before they hurt your SEO. Free, no signup required.

Check for broken links