XML Sitemap Generator
Paste a list of addresses and the sitemap is built in your browser. Or let the optional crawl mode collect the pages of a public website — that part runs through the ToolCMB server, which reads the site page by page.
- XML + sitemap index
- From a list: in the browser
- Crawl: via our server
- ZIP and .gz download
A sitemap from a list of addresses is built entirely in your browser — nothing is sent. Only the optional “Crawl a website” mode uses the ToolCMB server: it reads the public pages of the site you name, one by one, and returns the links it finds. We do not store the addresses or the result.
Paste a list, a spreadsheet column or an existing sitemap.xml — or open a .txt, .csv or .xml file. A date after the address (separated by a tab or comma) is used as “last modified”.
The crawl follows links on the same domain, starting at this address.
Parts of the path, separated by commas. * stands for anything.
Crawl log
| URL | Result |
|---|
Options
Google uses “last modified” when it is accurate, and ignores change frequency and priority. Leave both at “None” unless another search engine or tool needs them.
At most 50,000. Larger lists are split.
Needed for the index when the list is split. Empty = the site root.
Result
Upload the file to the root of your site, add the Sitemap line to robots.txt and submit the address in Google Search Console and Bing Webmaster Tools.
How to create an XML sitemap
Two sources, one result.
- 1
Paste a list — or crawl the site
A list of addresses, a spreadsheet column or an old sitemap.xml is processed in your browser. For “Crawl a website”, enter a start address: our server then fetches the public pages one by one and returns the links it finds.
- 2
Choose the options
Decide how “last modified” is filled, whether duplicates and tracking parameters are removed, and whether you want XML or a plain text file.
- 3
Download and submit
Save sitemap.xml (or a ZIP when the list was split), upload it to your site, copy the Sitemap line for robots.txt and submit the address to the search engines.
A valid sitemap, however you collect the addresses
Follows the sitemaps.org protocol and its limits.
From a list
Accepts one address per line, a table with dates, or an existing XML sitemap, pasted or opened as a .txt, .csv or .xml file. A date after the address becomes “last modified”.
Crawl mode
Follows links on the same domain from your start address, up to 50, 100, 250 or 500 pages, with exclude patterns, a stop button and a log of every page.
Only indexable pages
The crawl can respect robots.txt and leave out pages marked noindex or whose canonical points elsewhere, so the sitemap lists what should be indexed.
Splitting and index
Lists longer than your limit — at most 50,000 addresses or 50 MB per file — are split into several files plus a sitemap index, downloaded as one ZIP.
Lastmod, changefreq, priority
Keep given dates, fill in today where one is missing or set one date for all. Change frequency and priority are available, with a note that Google ignores them.
Checks
Reports invalid lines, mixed hosts, mixed http and https, addresses longer than 2,048 characters and duplicates before you download.
Honest about what is sent
The check runs on our server because a browser is not allowed to do it. Only the address you enter is sent; it is not stored.
Free, no sign-up
No account, no limits, no watermark — on a phone, tablet or computer.
What an XML sitemap is and what belongs in it
An XML sitemap is a file that lists the addresses of a website you want search engines to know about, each in a <loc> element, optionally with a <lastmod> date. It helps crawlers discover pages faster — especially on new sites, large sites and sites whose pages are poorly linked internally. A sitemap is a suggestion, not a guarantee: it does not force indexing and does not improve rankings by itself. Small, well-linked sites are usually crawled completely without one.
The sitemaps.org protocol sets firm limits: one file may hold at most 50,000 addresses and 50 MB uncompressed. Larger sites use several files and a sitemap index that lists them; you submit only the index. Addresses must be absolute, belong to the host the sitemap is stored on, and special characters such as & must be escaped — the generator does that for you. Google uses lastmod when it proves to be accurate and ignores changefreq and priority.
A good sitemap contains only canonical, indexable pages that answer with status 200. Common mistakes are listing redirects, error pages, pages with noindex or addresses blocked in robots.txt; mixing http and https or www and non-www; and stamping every page with today’s date on each export, which teaches search engines to distrust the dates. Do not list parameter variants of the same page. And keep the file current: a sitemap that is generated once and never updated loses its value quickly.
Sitemap elements and how they are used
| Element | Meaning | Used by Google |
|---|---|---|
| <code><loc></code> | Full address of the page (required) | Yes |
| <code><lastmod></code> | Date of the last significant change | Yes, when it is consistently accurate |
| <code><changefreq></code> | How often the page is expected to change | No |
| <code><priority></code> | Importance from 0.0 to 1.0 within the site | No |
| <code><image:image></code> | Images that belong to the page | Yes, for image search |
Tips
- Announce the sitemap with a
Sitemap:line in robots.txt — the Robots.txt Generator writes it for you. - Spot-check listed addresses with the HTTP Status Checker: every one should answer 200, and none should redirect (see the Redirect Checker).
- List only canonical addresses. The Canonical URL Generator cleans a whole list and marks duplicates.
- The crawl reads at most 500 pages per run. For a larger site, crawl sections separately by starting in a sub-folder, or export the address list from your CMS and paste it.
Frequently asked questions
Can’t find your answer? Contact us — we reply quickly.
Where do I upload the sitemap?
Normally to the root of the site, so that it opens at https://your-domain.com/sitemap.xml. A sitemap may only list addresses in its own folder and below, which is why the root is the safe place. Then add the address to robots.txt and submit it in Google Search Console and Bing Webmaster Tools.
How many URLs can a sitemap contain?
Up to 50,000 addresses and 50 MB uncompressed per file. Beyond that, split the list into several files and reference them in a sitemap index. This generator does the splitting and writes the index.
Does a sitemap improve my rankings?
Not directly. It helps search engines find and re-crawl your pages, which matters most for new, large or weakly linked sites. Whether a page is indexed and how it ranks depends on the page itself.
Should I set priority and changefreq?
For Google it makes no difference: both values are ignored. Leave them out unless another search engine or tool you use asks for them. An accurate lastmod date is the one optional value worth maintaining.
What is sent to your server?
Only the address or domain name you enter. A browser is not allowed to make this kind of request itself, so our server makes it for you and returns the result. We do not store the address or the result, and the content of the checked pages is never passed on.
What are the limits of the crawl mode?
One run reads at most 500 pages of a public website, follows only links on the same domain, and pauses when the number of requests per minute is reached. Private and internal network addresses are refused. The crawl reads the HTML the server delivers, so links created only by JavaScript are not found. The list mode has none of these limits apart from the protocol’s 50,000 addresses per file.
More free SEO and website tools
Generate tags and files, check status codes, redirects, DNS and certificates — free, without an account.