पाठ 11 / 34
Robots.txt
robots.txt directives se crawler access control karein, jaanein ki yeh kya nahi kar sakta aur apna sitemap reference karein.
Robots.txt मूल बातें
robots.txt aapki site ke root (/robots.txt) par ek file hai jo crawlers ko batati hai ki kaun se paths crawl karne hain. Yeh ek request hai, security nahi: well-behaved bots ise follow karte hain. Disallow crawling rokta hai, indexing nahi; page ko search results se bahar rakhne ke liye noindex use karein.
Robots.txt सिंटैक्स
Specific crawlers ko target karne ke liye User-Agent, paths block karne ke liye Disallow aur exceptions ke liye Allow use karein. Crawlers ko CSS, JS aur image files access karne dein.
# All bots
User-agent: *
Allow: /
Disallow: /admin/
Disallow: /private/
Disallow: /temp/
# Block a specific bot
User-agent: BadBot
Disallow: /
# Google bot
User-agent: Googlebot
Allow: /admin/
# Always allow these paths
Allow: /assets/
Allow: /images/
# Sitemap location
Sitemap: https://example.com/sitemap.xml