Skip to content
Free tool · Technical SEO

Sitemap Validator

Check an XML sitemap for the problems that make crawlers skip URLs: duplicates, other hosts, http and bad dates.

No sign-up. We read the public file for you. Nothing is stored.

For example:

We look at /sitemap.xml, then whatever robots.txt declares, then /sitemap_index.xml.

In the output

What you get

  • How many URLs the sitemap lists
  • URLs crawlers will drop, and why
  • Duplicates and http:// links
  • Invalid, future or stamped lastmod dates
  • That every <loc> is an absolute URL on the sitemap’s own host, since relative and off-host URLs are dropped.
  • Duplicate URLs, http:// URLs on an https site, and URLs with #fragments, which crawlers collapse or redirect.
  • That lastmod values are valid W3C dates, none are in the future, and they are not one date stamped on every URL.
  • The 50,000-URL limit per file, and whether the file is a sitemap index rather than a list of pages.
Three steps

How to use the Sitemap Validator

1

Enter your domain

We look at /sitemap.xml, then whatever your robots.txt declares, then /sitemap_index.xml. An index is followed into its first sitemaps.

2

Read the checks

Errors are URLs crawlers will not use. Warnings are URLs that waste crawling or dates that undermine trust.

3

Fix the generator, not the file

Sitemaps are almost always generated. Fix the template or plugin that writes them, or the same problems return on the next build.

Why it matters

Why a sitemap that loads can still be wrong

Invalid URLs are dropped quietly.

A relative path, a URL on another host or a fragment is skipped without an error. The sitemap loads; those pages are simply missing from it.

lastmod has to be believable.

Google uses lastmod only when it matches reality. A build that stamps today’s date on every URL teaches Google to ignore your dates altogether.

http links start with a redirect.

Every http:// URL in a sitemap sends crawlers to a redirect before they reach the page.

Limits are hard limits.

One file may list 50,000 URLs. Past that, split it and list the parts in a sitemap index.

After you run it

What to do with the result

A sitemap is a list of the pages you want crawled, and it only helps if crawlers trust it. The failures are rarely visible: a sitemap that loads can still list URLs crawlers drop, or dates they have learned to ignore. Site Audit analyses the sitemap against the pages it actually finds, so you see both the URLs the sitemap lists that are broken and the pages it forgot.

Site Audit

Frequently asked questions

No. Google has said it ignores both. The only optional field it uses is lastmod, and only when the dates are consistently accurate. Leaving changefreq and priority in does no harm; it just does not help.

See where you actually stand.

A free tool checks one thing once. Veritas watches all of it, continuously, across every engine that decides whether you get cited.

GOOGLE · CHATGPT · PERPLEXITY · GEMINI · AI OVERVIEWS