SkillByAIइंटरैक्टिव संस्करण खोलें →

पाठ 11 / 34

Robots.txt

robots.txt directives se crawler access control karein, jaanein ki yeh kya nahi kar sakta aur apna sitemap reference karein.

Robots.txt मूल बातें

robots.txt aapki site ke root (/robots.txt) par ek file hai jo crawlers ko batati hai ki kaun se paths crawl karne hain. Yeh ek request hai, security nahi: well-behaved bots ise follow karte hain. Disallow crawling rokta hai, indexing nahi; page ko search results se bahar rakhne ke liye noindex use karein.

Robots.txt सिंटैक्स

Specific crawlers ko target karne ke liye User-Agent, paths block karne ke liye Disallow aur exceptions ke liye Allow use karein. Crawlers ko CSS, JS aur image files access karne dein.

# All bots
User-agent: *
Allow: /
Disallow: /admin/
Disallow: /private/
Disallow: /temp/

# Block a specific bot
User-agent: BadBot
Disallow: /

# Google bot
User-agent: Googlebot
Allow: /admin/

# Always allow these paths
Allow: /assets/
Allow: /images/

# Sitemap location
Sitemap: https://example.com/sitemap.xml