Lesson 11 / 34
Robots.txt
Control crawler access with robots.txt directives, understand what it cannot do and reference your sitemap.
Robots.txt basics
robots.txt is a file at your site root (/robots.txt) that tells crawlers which paths to crawl. It is a request, not security: well-behaved bots follow it. Disallow stops crawling, not indexing; to keep a page out of search results, use noindex.
Robots.txt syntax
Use User-Agent to target specific crawlers, Disallow to block paths, and Allow to make exceptions. Always allow crawlers to access CSS, JS, and image files.
# All bots
User-agent: *
Allow: /
Disallow: /admin/
Disallow: /private/
Disallow: /temp/
# Block a specific bot
User-agent: BadBot
Disallow: /
# Google bot
User-agent: Googlebot
Allow: /admin/
# Always allow these paths
Allow: /assets/
Allow: /images/
# Sitemap location
Sitemap: https://example.com/sitemap.xml