How to use this tool
- Choose Sitemap URL and enter the complete XML address, or switch to Paste XML or Upload XML or GZIP.
- Select whether to follow child sitemaps, then select Extract URLs. Review the status and any failed files before relying on the total.
- All extracted URLs appear in one text box. Optionally filter by section or text, then select extra fields such as Last modified or Source sitemap. Copy all or download CSV or TXT uses the complete displayed list.
Understanding the output
Unique URLs counts distinct decoded loc values. Total URL entries includes repeated occurrences across the successfully read documents. Duplicate entries is the difference. Sections group those unique URLs by hostname and one, two, or three path segments. They describe URL structure, not categories inferred from page content. Partial extraction means some documents or entries were not read; the displayed counts are not a total for the entire website.
Compare an exported URL inventory with a previous snapshot using List Compare to identify added and removed addresses.
Inspect a specific address with URL Parser when query parameters or URL components need closer review.
Five entries, four unique pages
The Try example input contains the homepage, two /blog/ pages, and the same /products/desk/ page twice. It produces 5 total entries, 4 unique URLs, 1 duplicate entry, and 3 sections at the first path level. Filtering for /blog/ leaves 2 URLs; a CSV download contains those two rows plus its header.
Turn a sitemap into a URL inventory
A sitemap is a publisher-provided list of URLs intended for discovery. Extracting that list is useful before a content review, a migration, or a comparison with another export. The tool reads the direct loc value of each url record. Image locations, video resources, and alternate-language links are not counted as additional page URLs. It does not visit each extracted page or guess which URLs are canonical.
A sitemap index lists other sitemap documents rather than pages. When following is enabled, the extractor reads those children, combines their entries, and avoids fetching the same document repeatedly. The file report distinguishes page counts in a urlset from child-document counts in a sitemapindex. A failed child remains visible instead of being silently excluded from a claim of complete coverage.
Inspect categories through URL paths
Path grouping helps compare sections such as /blog/, /products/, and /guides/. The first level is usually the clearest overview. Increase the depth to separate /blog/news/ from /blog/tutorials/. Root URLs appear under the hostname with a slash. Hostnames remain separate, so a shop subdomain is not combined with the main site.
These are structural groups, not an automatic classification of subject matter. A site with flat URLs may place each page in its own group. Query strings do not create new sections, but they remain part of the URL and therefore affect duplicate detection. The section summary uses all unique URLs; text filters narrow the output text box and exports.
Choose the fields to copy or download
The default text output and TXT download contain one URL per line, with no header. Select Last modified, Source sitemap, Section, Change frequency, or Priority to add columns. With extra fields selected, the text box and TXT use tab-separated columns with a header; CSV uses the same selected columns with a header. Missing metadata is blank. Dates and other metadata are taken from the sitemap, not verified against the page.
Duplicate detection compares exact decoded URL strings after surrounding whitespace is trimmed. It deliberately keeps differences in trailing slashes, path case, query order, and fragments. Those differences may matter on the source website. When duplicates are removed, the first occurrence supplies the metadata. Leave Show unique URLs unchecked when you need every occurrence and its source.
Copy all copies the entire displayed list, including selected metadata. Filters apply to the text box and both downloads. Clear filters restores all matching sections and paths. In text output with metadata, backslashes, tabs and line breaks inside values are escaped as text to keep one record per line. CSV retains those characters using quoted cells. Both formats prefix formula-like metadata with an apostrophe for spreadsheet safety.
Sitemap presence is not Google indexing
A URL in an XML sitemap is a discovery hint, not proof that Google has crawled or indexed it. This extractor does not access Search Console, test robots rules, check response codes for listed pages, or submit URLs to search engines. Use Search Console’s indexing reports for the verified property when you need indexing evidence.
The sitemap protocol permits larger documents than this interactive tool. Here each document is limited to 5 MB compressed input and 5 MB after decompression, with 30 MB total XML, 100 documents, 100,000 URL entries, and five child levels per extraction. These limits keep browser work bounded. For larger sites, extract individual child sitemaps separately.
Method and supported input
Well-formed UTF-8 XML is parsed as a urlset or sitemapindex. Standard XML entities are decoded, while DOCTYPE and custom entity declarations are rejected. Direct page loc values must be absolute HTTP(S) URLs without credentials and no longer than 2,048 characters. Invalid or missing loc entries are counted separately.
URL loading uses a restricted server fetcher with a 15-second request timeout, up to three redirects, and public HTTP(S) hostnames on standard ports. Files and pasted XML are parsed on the device. Gzip files are decompressed before parsing. Downloads are generated from the current filtered result set.
- The supplied sitemap is the intended inventory; unpublished or omitted website pages cannot be discovered from it.
- Optional metadata is reported as source text. No dates, priorities, content categories, or indexing status are inferred.
Limitations
- HTML sitemaps, robots.txt, RSS/Atom feeds, plain URL lists, password-protected URLs, private hosts, and nonstandard ports are not accepted.
- Sites can block automated requests. A failed fetch can be retried using a downloaded XML file or pasted XML.
- An empty sitemap can legitimately contain zero URLs. Invalid loc values, failed children, and extraction limits are reported separately.
- Requests are rate-limited; wait a minute if URL loading is temporarily limited. Metadata output protects leading spreadsheet formula characters.
Common questions
Can I extract URLs from a sitemap index?
Yes. Leave Follow child sitemaps in indexes selected. The tool reads nested index references within its document, depth, and size limits. Uncheck it to inspect the index without requesting its children; page totals will then be marked partial.
Are my uploaded files sent to a server?
Uploaded files and pasted XML are read on your device. If child following is enabled, the referenced child links are sent to our server for retrieval. URLs entered in Sitemap URL are also fetched through our server. The application does not persist the fetched XML or results.
Why are the unique URL and total entry counts different?
The same URL can occur more than once in one sitemap or across several child sitemaps. Total entries includes those repetitions. Unique URLs keeps the first occurrence of each exact decoded URL string.
Does extracting a sitemap improve Google indexing?
Extraction helps you inspect and export a sitemap’s contents. It does not submit the sitemap, request indexing, or prove that any listed page is indexed.
Can I download only one category?
Choose a section in Filter by section or open URLs by section and select its count. Copy all, CSV, and TXT include every matching URL and your selected fields. Clear filters returns to the full list.