Robots.txt Generator

Build a robots.txt file from rule groups and presets, decide which AI crawlers may read your site, and test any address against the rules before you upload the file.

  • Generator + tester
  • AI crawler switches
  • Checks the rules
  • Runs in your browser

Presets

A preset replaces the rule groups below and keeps your sitemaps. Edit anything afterwards — Undo brings back what you had.

Rules for crawlers

A crawler follows only the group that names it most precisely, otherwise the * group. Inside a group the longest matching path wins; when an Allow and a Disallow rule are equally long, Allow wins. * stands for any characters, $ marks the end of the address.

AI crawlers

Training crawlers copy content into model training data; search and user-triggered fetchers read a page live and usually link back to it. Each switch adds or removes one group. Blocking Google-Extended does not change your ranking in Google Search. Reputable AI companies respect these rules, but robots.txt cannot force anyone.

Sitemaps and file header

Sitemap addresses

Use the full address including https://. Sitemap lines apply to all crawlers, wherever they stand in the file.

Optional. Each line becomes a comment that starts with #.

Import an existing robots.txt

    robots.txt

    Save the file as robots.txt and upload it to the root of the site, so that it opens at https://your-domain.com/robots.txt. Each subdomain needs its own file. A blocked page can still appear in search results without a description — use noindex on the page itself to keep it out.

    Checks

      Test addresses against these rules

      Full addresses and paths such as /shop/?sort=price both work. Only the path and the query string are compared.

      Type a crawler name or paste a complete User-Agent header.

      Enter one or more addresses to see whether the chosen crawler may fetch them.

      How it works

      How to create a robots.txt file

      The file is rebuilt after every change.

      1. 1

        Start from a preset or your own file

        Pick a preset such as WordPress, WooCommerce or “Allow everything”, or paste an existing robots.txt to load it into the editor.

      2. 2

        Edit the groups

        Add user-agents, Allow and Disallow paths and your sitemap addresses. Two switches add or remove the groups for AI crawlers.

      3. 3

        Test, then download

        Enter a few addresses to see whether a crawler may fetch them, read the checks, and save the file as robots.txt for the root of your site.

      Why ToolCMB

      A robots.txt editor that explains itself

      Matching follows RFC 9309 and the behavior Google documents.

      Groups and rules

      Each group names one or more user-agents and lists its Allow and Disallow paths. Groups and rules can be reordered, commented and removed.

      Presets

      Allow everything, block everything, WordPress, WooCommerce, a Shopify-like store, Joomla, Drupal, a staging site and a block list for AI training bots. Undo restores what you had.

      AI crawler switches

      One switch for training crawlers such as GPTBot, ClaudeBot, Google-Extended and CCBot, another for AI search and user-triggered fetchers such as OAI-SearchBot and PerplexityBot.

      Built-in tester

      Paste addresses or paths and pick a crawler. For each one you see Allowed or Blocked, the group that was used and the line of the deciding rule. The result exports as CSV.

      Import with repair

      Paste or open an existing file. Misspelled directives are recognized, and invalid lines, orphaned rules and unsupported <code>noindex</code> lines are reported.

      Checks while you type

      Warns when the whole site or Googlebot is blocked, when a rule may hide CSS or JavaScript, when a path is malformed and when the file exceeds 500 KiB.

      Private by design

      Everything runs in your browser. What you type, paste or open is not sent to a server.

      Free, no sign-up

      No account, no limits, no watermark — on a phone, tablet or computer.

      What robots.txt does — and what it cannot do

      A robots.txt file is a plain text file at the root of a host, for example https://example.com/robots.txt. It tells crawlers which paths they may request. The format is standardized as the Robots Exclusion Protocol in RFC 9309: groups start with one or more User-agent lines, followed by Allow and Disallow rules. Each protocol, subdomain and port needs its own file, and Sitemap lines apply to every crawler, wherever they stand.

      A crawler follows only one group: the one that names it most precisely, otherwise the * group. Inside that group the rule with the longest matching path decides; if an Allow and a Disallow rule are equally long, Allow wins. * stands for any sequence of characters and $ marks the end of the address, so Disallow: /*.pdf$ blocks every PDF. Paths are case-sensitive. Crawl-delay is not part of the standard: Bing honors it, Google ignores it.

      The most common mistake is to use robots.txt to keep a page out of search results. It only stops crawling: a blocked address can still be indexed, without a description, when other pages link to it. To keep a page out, leave it crawlable and add a robots noindex meta tag or an X-Robots-Tag header. Other frequent errors are a leftover Disallow: / from a staging site, blocked CSS or JavaScript that Google needs to render pages, and secrets listed in a file that anyone can read.

      Common robots.txt rules

      RuleEffectNote
      <code>Disallow:</code> (empty)Allows everythingSame result as having no file
      <code>Disallow: /</code>Blocks the whole siteOnly for sites that must stay out of search
      <code>Disallow: /cart/</code>Blocks a folder and everything below itWithout the last slash it also matches /cart-help
      <code>Disallow: /*?sort=</code>Blocks addresses that contain this parameter<code>*</code> matches any characters
      <code>Allow: /wp-admin/admin-ajax.php</code>Re-opens one file inside a blocked folderThe longer rule wins
      <code>Sitemap: https://example.com/sitemap.xml</code>Points crawlers to your sitemapMust be a full address

      Tips

      • List your sitemap in the file — build one with the Sitemap Generator.
      • After uploading, confirm that /robots.txt answers with status 200 using the HTTP Status Checker. A server error on this file can stop crawling of the whole site.
      • To keep a page out of search results, write a robots meta tag with the Meta Tag Generator instead of blocking the page here.
      • Robots.txt is a request, not access control. Protect private folders with a password or an IP rule from the .htaccess Generator.
      FAQ

      Frequently asked questions

      Can’t find your answer? Contact us — we reply quickly.

      Where do I put the robots.txt file?

      In the root folder of the site, so that it opens at https://your-domain.com/robots.txt. A file in a sub-folder is ignored. Each subdomain, and http and https separately, uses its own file.

      Does Disallow remove a page from Google?

      No. Disallow stops crawling, not indexing. If other pages link to the address it can still appear in results without a description. Use a noindex robots meta tag or an X-Robots-Tag header — and do not block the page, otherwise the crawler never sees that instruction.

      How do I block AI bots such as GPTBot or ClaudeBot?

      Add a group that names the crawler and contains Disallow: /. The two AI switches of this generator write such groups for the known training crawlers and for AI search fetchers. Reputable operators follow these rules, but robots.txt cannot force a crawler to obey.

      Does blocking Google-Extended affect my Google ranking?

      No. Google-Extended is a control token for the use of your content in Google’s AI models. It does not influence crawling by Googlebot or your position in Google Search.

      Do I need a robots.txt file at all?

      Not necessarily. Without a file, crawlers assume that everything may be fetched. A file is useful to keep crawlers out of search result pages, carts and admin areas, to address individual crawlers and to announce your sitemap.

      Is my data sent to a server?

      No. The tool runs entirely in your browser. Your input is not uploaded, logged or stored on our servers.

      More free SEO and website tools

      Generate tags and files, check status codes, redirects, DNS and certificates — free, without an account.

      Browse all tools