Robots.txt Generator
Build a robots.txt file from rule groups and presets, decide which AI crawlers may read your site, and test any address against the rules before you upload the file.
- Generator + tester
- AI crawler switches
- Checks the rules
- Runs in your browser
Presets
A preset replaces the rule groups below and keeps your sitemaps. Edit anything afterwards — Undo brings back what you had.
Rules for crawlers
There are no groups. Without rules every crawler may fetch every page.
A crawler follows only the group that names it most precisely, otherwise the * group. Inside a group the longest matching path wins; when an Allow and a Disallow rule are equally long, Allow wins. * stands for any characters, $ marks the end of the address.
AI crawlers
Training crawlers copy content into model training data; search and user-triggered fetchers read a page live and usually link back to it. Each switch adds or removes one group. Blocking Google-Extended does not change your ranking in Google Search. Reputable AI companies respect these rules, but robots.txt cannot force anyone.
Sitemaps and file header
Use the full address including https://. Sitemap lines apply to all crawlers, wherever they stand in the file.
Optional. Each line becomes a comment that starts with #.
Import an existing robots.txt
robots.txt
Nothing to show yet.
Save the file as robots.txt and upload it to the root of the site, so that it opens at https://your-domain.com/robots.txt. Each subdomain needs its own file. A blocked page can still appear in search results without a description — use noindex on the page itself to keep it out.
Checks
Test addresses against these rules
Full addresses and paths such as /shop/?sort=price both work. Only the path and the query string are compared.
Type a crawler name or paste a complete User-Agent header.
Enter one or more addresses to see whether the chosen crawler may fetch them.
| Address | Result | Deciding rule |
|---|
How to create a robots.txt file
The file is rebuilt after every change.
- 1
Start from a preset or your own file
Pick a preset such as WordPress, WooCommerce or “Allow everything”, or paste an existing robots.txt to load it into the editor.
- 2
Edit the groups
Add user-agents, Allow and Disallow paths and your sitemap addresses. Two switches add or remove the groups for AI crawlers.
- 3
Test, then download
Enter a few addresses to see whether a crawler may fetch them, read the checks, and save the file as robots.txt for the root of your site.
A robots.txt editor that explains itself
Matching follows RFC 9309 and the behavior Google documents.
Groups and rules
Each group names one or more user-agents and lists its Allow and Disallow paths. Groups and rules can be reordered, commented and removed.
Presets
Allow everything, block everything, WordPress, WooCommerce, a Shopify-like store, Joomla, Drupal, a staging site and a block list for AI training bots. Undo restores what you had.
AI crawler switches
One switch for training crawlers such as GPTBot, ClaudeBot, Google-Extended and CCBot, another for AI search and user-triggered fetchers such as OAI-SearchBot and PerplexityBot.
Built-in tester
Paste addresses or paths and pick a crawler. For each one you see Allowed or Blocked, the group that was used and the line of the deciding rule. The result exports as CSV.
Import with repair
Paste or open an existing file. Misspelled directives are recognized, and invalid lines, orphaned rules and unsupported <code>noindex</code> lines are reported.
Checks while you type
Warns when the whole site or Googlebot is blocked, when a rule may hide CSS or JavaScript, when a path is malformed and when the file exceeds 500 KiB.
Private by design
Everything runs in your browser. What you type, paste or open is not sent to a server.
Free, no sign-up
No account, no limits, no watermark — on a phone, tablet or computer.
What robots.txt does — and what it cannot do
A robots.txt file is a plain text file at the root of a host, for example https://example.com/robots.txt. It tells crawlers which paths they may request. The format is standardized as the Robots Exclusion Protocol in RFC 9309: groups start with one or more User-agent lines, followed by Allow and Disallow rules. Each protocol, subdomain and port needs its own file, and Sitemap lines apply to every crawler, wherever they stand.
A crawler follows only one group: the one that names it most precisely, otherwise the * group. Inside that group the rule with the longest matching path decides; if an Allow and a Disallow rule are equally long, Allow wins. * stands for any sequence of characters and $ marks the end of the address, so Disallow: /*.pdf$ blocks every PDF. Paths are case-sensitive. Crawl-delay is not part of the standard: Bing honors it, Google ignores it.
The most common mistake is to use robots.txt to keep a page out of search results. It only stops crawling: a blocked address can still be indexed, without a description, when other pages link to it. To keep a page out, leave it crawlable and add a robots noindex meta tag or an X-Robots-Tag header. Other frequent errors are a leftover Disallow: / from a staging site, blocked CSS or JavaScript that Google needs to render pages, and secrets listed in a file that anyone can read.
Common robots.txt rules
| Rule | Effect | Note |
|---|---|---|
| <code>Disallow:</code> (empty) | Allows everything | Same result as having no file |
| <code>Disallow: /</code> | Blocks the whole site | Only for sites that must stay out of search |
| <code>Disallow: /cart/</code> | Blocks a folder and everything below it | Without the last slash it also matches /cart-help |
| <code>Disallow: /*?sort=</code> | Blocks addresses that contain this parameter | <code>*</code> matches any characters |
| <code>Allow: /wp-admin/admin-ajax.php</code> | Re-opens one file inside a blocked folder | The longer rule wins |
| <code>Sitemap: https://example.com/sitemap.xml</code> | Points crawlers to your sitemap | Must be a full address |
Tips
- List your sitemap in the file — build one with the Sitemap Generator.
- After uploading, confirm that
/robots.txtanswers with status 200 using the HTTP Status Checker. A server error on this file can stop crawling of the whole site. - To keep a page out of search results, write a robots meta tag with the Meta Tag Generator instead of blocking the page here.
- Robots.txt is a request, not access control. Protect private folders with a password or an IP rule from the .htaccess Generator.
Frequently asked questions
Can’t find your answer? Contact us — we reply quickly.
Where do I put the robots.txt file?
In the root folder of the site, so that it opens at https://your-domain.com/robots.txt. A file in a sub-folder is ignored. Each subdomain, and http and https separately, uses its own file.
Does Disallow remove a page from Google?
No. Disallow stops crawling, not indexing. If other pages link to the address it can still appear in results without a description. Use a noindex robots meta tag or an X-Robots-Tag header — and do not block the page, otherwise the crawler never sees that instruction.
How do I block AI bots such as GPTBot or ClaudeBot?
Add a group that names the crawler and contains Disallow: /. The two AI switches of this generator write such groups for the known training crawlers and for AI search fetchers. Reputable operators follow these rules, but robots.txt cannot force a crawler to obey.
Does blocking Google-Extended affect my Google ranking?
No. Google-Extended is a control token for the use of your content in Google’s AI models. It does not influence crawling by Googlebot or your position in Google Search.
Do I need a robots.txt file at all?
Not necessarily. Without a file, crawlers assume that everything may be fetched. A file is useful to keep crawlers out of search result pages, carts and admin areas, to address individual crawlers and to announce your sitemap.
Is my data sent to a server?
No. The tool runs entirely in your browser. Your input is not uploaded, logged or stored on our servers.
More free SEO and website tools
Generate tags and files, check status codes, redirects, DNS and certificates — free, without an account.