lintlab Sitemap Doctor
Serveur MCP (via Apify) qui vérifie et valide les sitemaps XML : les trouve via robots.txt, suit les index de sitemaps et les fichiers gzip, extrait chaque URL avec lastmod/changefreq/priority, compare avec une exécution précédente, et vérifie optionnellement le statut HTTP. Les sitemaps cassés sont signalés, jamais ignorés.
Serveur MCP hébergé
npx add-mcp 'https://mcp.apify.com?tools=lintlab/sitemap-doctor'S’installe dans Claude Code, Codex, Cursor et plus
Documentation
XML Sitemap Checker, Validator & URL Extractor (lintlab/sitemap-doctor) Actor
Check and validate XML sitemaps: find them via robots.txt, follow sitemap indexes and gzip files, extract every URL with lastmod/changefreq/priority, diff against a previous run, and optionally check HTTP status. Broken sitemaps are reported, never skipped.
- URL: https://apify.com/lintlab/sitemap-doctor.md
- Developed by: lintlab (community)
- Categories: SEO tools, Developer tools
- Stats: 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- User rating: No ratings yet
Pricing
from $0.24 / 1,000 url extracteds
This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events. Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.
Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event
What's an Apify Actor?
An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes. In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours, and optionally produces a well-defined JSON output, datasets with results, or files in key-value store. In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.
How to integrate an Actor?
If asked about integration, you help developers integrate Actors into their projects. You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.
For examples already wired to this Actor's own input schema, see the API section below.
Each client library has reference documentation the quickstart does not restate: JavaScript/TypeScript (npm install apify-client) and Python (pip install apify-client).
README
XML Sitemap Checker, Validator & URL Extractor
Check a public XML sitemap and extract all of its URLs in one run. It works as a sitemap crawler and sitemap extractor: enter a domain or a sitemap URL, and the Actor finds sitemap files through robots.txt, follows sitemap indexes and gzip files, reports validation errors, and returns the sitemap URLs as one record per URL with its sitemap metadata. Optional HTTP status checks flag URLs that don't resolve, and a diff mode compares a run with a previous one. Parse and HTTP errors stay visible in the SUMMARY record instead of a sitemap being silently dropped.
Try a ready-made example: Extract URLs from a sitemap. Open it, click Start, then copy it and swap in your own domain.
Real output from a run on sitemaps.org (2026-10-01): 8 of the 84 URL records, with lastmod, date validity, HTTP status and issues.
What it does
- Accepts site roots,
.xmlsitemap URLs, and.xml.gzsitemap URLs. - For a site root, reads
Sitemap:lines from/robots.txt; if none exist, tries/sitemap.xmland/sitemap_index.xml. - Recursively follows sitemap indexes and handles gzip files and namespace-prefixed XML.
- Preserves
lastmod,changefreq, andpriority, including invalid values for diagnosis. - Reports malformed XML, non-2xx sitemap responses, invalid child entries, and sitemap protocol limits in
SUMMARY. - Emits duplicate occurrences with
duplicateOfset to the first sitemap where the URL appeared. Duplicate occurrences are not charged as extracted URLs. - Can check URL status with HEAD and a GET fallback for 405/501. Checks honor robots.txt and run at no more than five concurrent requests per host.
- Can mark URLs
added,unchanged, orremovedrelative to a previous Actor dataset. Removed records are not charged. - Rejects private, loopback, link-local, and other non-public IP ranges before every request and redirect hop.
Input
Only public sitemap and robots.txt files are read. Status checks honor robots.txt, and private or internal network addresses are always rejected.
{
"startUrls": [
{ "url": "https://www.sitemaps.org/sitemap.xml" }
],
"maxUrls": 10000,
"maxSitemapFiles": 200,
"checkStatus": false,
"statusSampleSize": 0,
"sameHostOnly": true
}
statusSampleSize: 0 means all unique URLs when checkStatus is enabled. With sameHostOnly: true, cross-host sitemap files referenced by an index are not followed. Cross-host page URLs are still emitted and flagged so they can be fixed.
previousDatasetId is optional. On the Apify platform it is opened from cloud storage. Its non-removed url values form the comparison baseline.
Dataset output
Each sitemap occurrence produces one record. This is an actual item from the local end-to-end run against https://www.sitemaps.org/sitemap.xml on 2026-09-25:
{
"url": "https://www.sitemaps.org/",
"sitemap": "https://www.sitemaps.org/sitemap.xml",
"lastmod": "2016-11-21",
"changefreq": null,
"priority": null,
"lastmodValid": true,
"hostMatches": true,
"duplicateOf": null,
"status": null,
"finalUrl": null,
"redirects": null,
"skippedByRobots": false,
"diff": null,
"issues": []
}
Stable issue IDs are:
DUPLICATE_URL, MULTIPLE_LOC, NON_XML_CONTENT_TYPE, INVALID_LASTMOD, FUTURE_LASTMOD, CROSS_HOST, NON_200, REDIRECTED, ROBOTS_DISALLOWED, INVALID_PRIORITY, INVALID_CHANGEFREQ, and URL_TOO_LONG. MULTIPLE_LOC uses the first location and is counted under structuralIssues on an index sitemap's SUMMARY entry. NON_XML_CONTENT_TYPE is also a structuralIssues entry, on any sitemap file that parses as XML but is served with an HTML or plain-text Content-Type (it records the type it saw); the file's URLs are still extracted.
The SUMMARY key-value-store record contains every attempted sitemap file with discovery method, type, gzip flag, uncompressed byte count, URL count, HTTP status, parse error, and protocol-limit violations. It also includes aggregate counts, billing counts, truncation state, and the most frequent issues.
The same end-to-end run parsed one sitemap file and emitted 84 unique URL records in 1.81 seconds:
{
"totals": {
"sitemapFilesAttempted": 1,
"sitemapFilesParsed": 1,
"urlRecords": 84,
"uniqueUrls": 84,
"removedUrls": 0,
"duplicates": 0,
"invalidLastmod": 0,
"non200": 0,
"crossHost": 0,
"statusChecksRequested": 0,
"statusResponses": 0
},
"topIssues": [],
"truncated": false,
"eventChargeLimitReached": false,
"billing": {
"sitemapFileParsed": 1,
"urlExtracted": 84,
"urlStatusChecked": 0
}
}
Pricing
Pay-per-event pricing (you pay only for work that succeeded):
$0.001 per sitemap fileforsitemap-file-parsed, charged once after a file is fetched and parsed successfully.$0.0003 per URLforurl-extracted, charged after the first current record for a unique URL is pushed. Duplicate occurrences and removed diff records are not charged.$0.0005 per status checkforurl-status-checked, charged after the URL record is pushed only when the check completed with an HTTP response. Network failures and robots skips are not charged.
The Actor stops adding work when an Apify event charge limit is reported.
Limits and behavior
maxUrls: default 10,000; maximum 200,000 total dataset records, including removed diff records.maxSitemapFiles: default 200.- Sitemap protocol violations are reported above 50,000 entries or 50 MB uncompressed.
- A 55 MB hard response/decompression cap, 20-second request timeout, 10-redirect cap, and SSRF checks bound network work.
- Status checks are limited to five simultaneous requests per origin.
- No proxies, authentication, login flows, or CAPTCHA handling are used.
Local development
Requires Node.js 20 or newer.
npm install
npm test
npm start
For a local Actor run, place INPUT.json in the default local key-value store and set APIFY_LOCAL_STORAGE_DIR. Newer Crawlee releases also recognize CRAWLEE_STORAGE_DIR for the same location.
More lintlab tools
- Broken Link Checker & Technical SEO Audit: page-level SEO findings and broken internal links
- Website Screenshot & Visual Regression Diff: full-page screenshots with pixel diffs
- PDF to Markdown & Text Extractor: page-aware Markdown for RAG and AI agents
Built by lintlab: tools that check web pages, sites and documents. Charged only for completed results. Support: open an issue on this Actor's Issues tab here on Apify.
Actor input Schema
startUrls (type: array):
Required list of HTTP(S) site roots or sitemap URLs. Site roots are inspected, then robots.txt Sitemap directives are used; /sitemap.xml and /sitemap_index.xml are fallbacks. Each successfully parsed sitemap file incurs the sitemap-file event price.
maxUrls (type: integer):
Maximum dataset URL records for the whole run, including duplicate occurrences and removed diff records. Default 10000; range 1–200000. Unique current URLs incur the URL extraction event price after they are pushed.
maxSitemapFiles (type: integer):
Maximum number of sitemap files fetched across the run. Default 200; each file that is fetched and parsed successfully incurs one sitemap-file event charge.
checkStatus (type: boolean):
When true, send HEAD requests (GET fallback for HTTP 405/501), at most 5 concurrently per host, while honoring robots.txt. Every URL check completed with an HTTP response incurs one status-check event charge. Default false.
statusSampleSize (type: integer):
Number of unique extracted URLs to status-check. Use 0 to check all when checkStatus is true. Default 0; unit is URLs.
previousDatasetId (type: string):
Optional Apify dataset ID from an earlier Sitemap Doctor run. Current URLs are marked added or unchanged, and URLs missing now are emitted as removed records. Removed records are not charged.
sameHostOnly (type: boolean):
When true (default), sitemap indexes do not lead to sitemap files on another hostname. Cross-host page URLs are still emitted and flagged so an agent can diagnose them.
Actor input object example
{
"startUrls": [
{
"url": "https://www.sitemaps.org/"
}
],
"maxUrls": 10000,
"maxSitemapFiles": 200,
"checkStatus": false,
"statusSampleSize": 0,
"sameHostOnly": true
}
Actor output Schema
items (type: string):
No description
summary (type: string):
No description
API
You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.
JavaScript example
import { ApifyClient } from 'apify-client';
// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
token: '<YOUR_API_TOKEN>',
});
// Prepare Actor input
const input = {
"startUrls": [
{
"url": "https://www.sitemaps.org/"
}
]
};
// Run the Actor and wait for it to finish
const run = await client.actor("lintlab/sitemap-doctor").call(input);
// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
console.dir(item);
});
// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs
Python example
from apify_client import ApifyClient
# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")
# Prepare the Actor input
run_input = { "startUrls": [{ "url": "https://www.sitemaps.org/" }] }
# Run the Actor and wait for it to finish
run = client.actor("lintlab/sitemap-doctor").call(run_input=run_input)
# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
print(item)
# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start
CLI example
echo '{
"startUrls": [
{
"url": "https://www.sitemaps.org/"
}
]
}' |
apify call lintlab/sitemap-doctor --silent --output-dataset
MCP server setup
{
"mcpServers": {
"apify": {
"type": "http",
"url": "https://mcp.apify.com/?tools=fetch-actor-details,lintlab/sitemap-doctor"
}
}
}
The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an Authorization: Bearer <APIFY_API_TOKEN> header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).
OpenAPI specification
Download the OpenAPI definition: https://api.apify.com/v2/actors/not7lGNg3TYpOYv0g/builds/MvppTE0IfCN5OVRCW/openapi.json