Short definition
A file that tells search engine bots which parts of a site they may crawl.
Beginner In plain English
robots.txt is a small text file at the root of a website with rules for bots, like "please do not go into this folder".
Professional How pros use it
Disallow controls crawling, not indexing; a blocked URL can still be indexed if linked. Never block CSS or JS needed for rendering. To keep a page out of the index, allow crawling and use noindex.
Examples
- Disallow: /cart/
- Sitemap: https://example.com/sitemap.xml
My notes
Keep private notes next to anything you read or practice.
Log in to take notes