Robots.txt Generator — Build, Check and Download robots.txt
Build a robots.txt with presets for allow-all, blocking AI crawlers, staging and WordPress, with warnings for the mistakes that de-index sites.
User-agent: * Allow: / Sitemap: https://www.example.com/sitemap.xml
Upload it to the root of your domain so it is served at https://yourdomain.com/robots.txt. Subfolders are ignored.
Runs entirely in your browser. Nothing you enter is sent to us or stored. robots.txt is one line item in a technical SEO audit. See what a full audit covers.
Found a problem, or something that could be better? We read every message and we fix things quickly.
This free robots.txt generator builds the file crawlers read before anything else on your site. Choose which user agents the rules apply to, list allowed and disallowed paths, add your sitemap, and download. Presets cover the situations people actually search for: allowing everything, blocking AI training crawlers such as GPTBot and ClaudeBot while keeping Google, locking down a staging site, and the sensible WordPress and e-commerce defaults.
robots.txt is small and easy to get catastrophically wrong: a single “Disallow: /” under “User-agent: *” removes a site from search, and it ships to production from staging more often than anyone admits. The generator warns about that and the other common errors, and the notes below explain what the file can and cannot do, including the fact that it controls crawling, not indexing.
How to use
- Pick a preset that is closest to your situation, or start from “Allow everything”.
- For each rule group, list the user agents it applies to (* means all), then the paths to allow or disallow. Paths start with / and are relative to the site root.
- Add the absolute URL of your sitemap; it is the one line that helps indexing rather than restricting it.
- Read any warnings shown above the output. Fix them unless you are deliberately blocking crawlers on a staging site.
- Download robots.txt and upload it to the root of your domain, so it is served at yourdomain.com/robots.txt. Verify in Google Search Console's robots.txt report.
Frequently asked questions
Does robots.txt remove pages from Google?
No. It stops crawling, not indexing. A page blocked in robots.txt can still appear in results (with no snippet) if other pages link to it. To keep a page out of the index, allow crawling and add a noindex meta tag or header; Google has to read the page to see the instruction. Use robots.txt to save crawl budget and keep bots out of admin, search-result and faceted URLs.
How do I block AI crawlers without affecting search?
Use the “Block AI training crawlers” preset. It leaves “User-agent: *” allowed and adds a separate group disallowing / for GPTBot, ClaudeBot, CCBot, Google-Extended and the others that honour robots.txt. Google-Extended controls Gemini training only; Googlebot and Search are unaffected. Note that robots.txt is voluntary: well-behaved crawlers obey it, scrapers do not.
What is the difference between Allow and Disallow?
Disallow blocks paths; Allow re-opens a narrower path inside a blocked one, for example disallowing /wp-admin/ while allowing /wp-admin/admin-ajax.php. When rules conflict, Google applies the most specific (longest) matching path. An empty Disallow: means allow everything.
Is Crawl-delay supported?
By Bing and Yandex, yes; Google ignores it and has since 2019. If Googlebot is crawling too hard, use the crawl rate setting in Search Console or return 503/429 temporarily. This tool warns when you set it so you do not expect an effect that will not come.
Can I use wildcards in paths?
Google and Bing support * (any sequence) and $ (end of URL): Disallow: /*?sort= blocks any URL containing ?sort=, and Disallow: /*.pdf$ blocks PDFs. Older or smaller crawlers may treat these literally, so keep wildcard rules for the big engines and keep the rest simple.
Where does robots.txt go?
At the root of each host, over the same protocol: https://www.example.com/robots.txt. A file in a subfolder is ignored, and example.com and www.example.com each need their own (usually identical) copy. It must be plain text, UTF-8, and under 500 KB.
A robots.txt that does the right things
For most public sites the whole file is: User-agent: * with no disallows, plus a Sitemap line. Add disallows only for URL spaces that are genuinely infinite or useless to search: internal search results, cart and checkout, admin, parameter permutations from filters and sorting. Do not block CSS, JavaScript or images; Google renders pages and needs them, and blocking them can hurt rankings.
The mistakes that take sites out of search
• Shipping the staging file (Disallow: /) to production. Check robots.txt as part of every deploy.
• Blocking a directory that contains the CSS or JS the site needs to render.
• Using robots.txt to hide pages instead of noindex, so they linger in results as bare URLs.
• A sitemap line pointing at a URL that 404s or redirects.
• Rules with paths that do not start with /, which are silently ignored.
Related free tools
- Meta Tag Generator — Fill in a page's title, description and social card details; get a complete, correctly escaped <head> block with live length counters.
- Schema Markup Generator — Fill a form, get valid JSON-LD for Organization, LocalBusiness, Person, Article, Product, FAQ, Event or Breadcrumbs. Copy it into <head>.
- Open Graph Preview — Enter a URL and see the exact card Facebook, LinkedIn, X and Slack will show, every tag found, and what is missing.