About this tool
A robots.txt file is a short set of instructions at the top of your website (example.com/robots.txt) that tells crawlers, the robots search engines and AI companies send out to read the web, which parts of your site they may visit. One wrong line can hide a whole site from Google, so this tool helps you write it, check it and prove what it does before it goes live.
The builder starts from a preset and lets you add rules for any robot. Two ticks handle the crawlers people ask about most: the AI training bots such as GPTBot, ClaudeBot and Google-Extended, and, separately, the bots that fetch pages for AI search answers.
The tester answers the question that matters: can this robot open this address? It follows Google’s published rules, longest match wins and Allow wins a tie, with * and $ wildcards, and shows the exact line that decided. Paste any existing robots.txt to have it checked for quiet mistakes.
The third tab builds the sitemap.xml that robots.txt points to: paste addresses or open a CSV, set last-modified dates, change frequency and priority per page or for all, and download it. Over 50,000 addresses are split into several files with a sitemap index. A checker reads any pasted sitemap and lists every problem by line.
How to use Robots.txt Generator
- In “Build a robots.txt”, press a preset such as “WordPress”, and tick “Block AI training crawlers” if you don’t want your pages used to train AI models.
- Under “Rules for each robot”, pick a robot and use “Add rule” to choose Disallow or Allow and type a path. Type your site address to fill in the sitemap line.
- Press “Test it” to open “Test & check a file”, choose a robot and type addresses to see Allowed or Blocked.
- Press “Copy” or “Download” and upload the file to the top folder of your website.
- For a sitemap, open “Sitemap.xml builder”, paste addresses or press “Open file” for a CSV or TXT list, and adjust dates, frequency and priority.
- Press “Download” (“Download all (.zip)” for a split sitemap), and paste any sitemap into “Check a sitemap” to test it.
With “Disallow: /admin/” and “Allow: /admin/help/”, the path /admin/help/faq is Allowed, because the Allow rule (12 characters) is longer than the Disallow (7), while /admin/settings is Blocked.
Features
- Five presets (allow all, block all, WordPress, online shop, staging site) with one-step undo.
- One-tick blocking for AI training crawlers (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot, meta-externalagent, Bytespider) and, separately, AI search robots (OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot).
- Rules for any robot, from over twenty common ones or any name you type, plus Crawl-delay per robot.
- Sitemap lines filled in from your site address, plus any extra sitemaps.
- A URL tester that follows Google’s rules (most specific group, longest match, Allow on ties, * and $ wildcards), for many addresses at once, showing the deciding line.
- A checker that spots misspelt directives, missing colons, rules before any User-agent, relative sitemap addresses, retired directives like Noindex and Host, blocked CSS and JavaScript, and files over Google’s 500 KiB limit.
- Sitemap builder: addresses, paths or CSV columns (url, lastmod, changefreq, priority), duplicates and # parts removed, & escaped, a table to edit each page, and automatic splitting into 50,000-address files plus an index.
- Sitemap checker: broken XML with the line, the wrong namespace, relative or overlong addresses, bad dates, priorities and frequencies, mixed sites and the 50,000 / 50 MB limits.
Tips and good to know
- robots.txt controls crawling, not indexing. A blocked page can still appear in Google as a bare link if other sites point to it. To keep a page out of results, let Google crawl it and add a noindex robots meta tag.
- Never use robots.txt to hide private things: the file is public, so it tells everyone where to look.
- Paths are case-sensitive: Disallow: /Photos does not block /photos.
- End a folder rule with a slash. Disallow: /blog blocks /blog-tips too, while /blog/ blocks only the folder.
- List only the pages you want found, at their final address: no redirects, no noindex pages, and an honest lastmod. Google ignores changefreq and priority.
Frequently asked questions
Is anything uploaded or fetched from my website?
No. Files are built and tested entirely in your browser, and the tool never contacts your site. To check a live robots.txt or sitemap, open it yourself and paste it in.
Is it free? Are there limits?
It is free with no sign-up. Add as many robots and rules as you like, test up to 200 addresses at a time, and build sitemaps of hundreds of thousands of addresses.
Does it work on a phone or offline?
Yes. It works in Safari, Chrome and Firefox on phones and computers, and keeps working without a connection once the page has loaded.
Will blocking AI crawlers hurt my Google ranking?
Blocking the training crawlers does not: Google-Extended only controls AI training, and Googlebot still crawls for search. Blocking the AI search robots may keep you out of answers in ChatGPT, Claude or Perplexity.
Do all robots obey robots.txt?
Search engines and the large AI companies say they do. It is a request, not a lock: to truly stop a robot, block it at your server or firewall.
Why does Google ignore Crawl-delay?
Google sets its own crawl speed from how quickly your server responds. Bing and Yandex do read Crawl-delay, as seconds between visits.
What happens when Allow and Disallow both match?
Google uses the rule with the longest path, as the most specific. If both are the same length, Allow wins. The tester shows which rule won.
Page last reviewed
