Watchtower
Job alerts for AI agents. Describe a role once in plain language and get only the new matching postings from 1,500+ tech company and startup job boards (Greenhouse, Lever, Ashby, Workday, Apple, Google, Amazon, Microsoft, and more), with salary and experience filters and webhooks. Free, no sign-up.
Hosted MCP Server
npx add-mcp 'https://watchtower.lat/mcp'Installs into Claude Code, Codex, Cursor and more
Documentation
Watchtower
Tech job monitoring for AI agents. Say what you are looking for once and get only the new postings that match, from the job boards of tech companies and startups, as structured JSON.
Agents are bad at waiting for a job to be posted. A session lasts minutes, careers pages are heavy to re-read, and "has anyone posted an iOS role in Austin yet?" turns into re-running the same search every day. With Watchtower the agent creates a watch once, in plain language:
{ "query": "iOS jobs in Austin making at least 150k a year with a maximum of 6 years of experience" }
Watchtower checks the job boards of tech companies and startups on a schedule and remembers which jobs were open. get_changes then returns only the new postings that match, and a webhook can wake the agent when one arrives. An agent can also watch one company's board by URL and get JOB_ADDED, JOB_REMOVED and JOB_UPDATED events.
- No URL needed. A watch created from a
querycovers every board Watchtower monitors: a built-in directory of tech company and startup boards plus every board anyone has watched by URL. - Pay and experience. Jobs carry
salaryandexperience_yearswhen the posting states them, and watches filter onmin_salaryandmax_experience_years. - Greenhouse, Lever, Ashby, Workable, SmartRecruiters, Recruitee, Workday, iCIMS, Oracle Recruiting, Eightfold and SuccessFactors are read through each platform's own endpoints: no bot walls. So are Apple (
jobs.apple.com), Google (careers.google.com), Amazon (amazon.jobs) and Microsoft (careers.microsoft.com), which run their own careers sites. Any other careers page works if it publishes schema.orgJobPostingmarkup. - Filters on the watch, so
get_changesonly returns what matters:keywords,all_keywords,exclude_keywords,locations,seniority,remote_only,min_salaryandmax_experience_years. Every job carries derivedremoteandseniorityfields. - Many companies in one call: pass
urlsinstead ofurl. - MCP (Streamable HTTP) and REST, backed by the same service layer.
- Free and anonymous: a client gets a token and up to 50 watches (a search across all boards is one watch).
- TypeScript, Node.js 22, Fastify 5, PostgreSQL, the official MCP TypeScript SDK. No LLM or third-party API keys required.
Use the hosted service
Watchtower runs at watchtower.lat, free, with no sign-up. MCP endpoint: https://watchtower.lat/mcp.
- Claude Code:
claude mcp add --transport http watchtower https://watchtower.lat/mcp, or install the plugin, which adds a skill that tells Claude when to use it:/plugin marketplace add connorlagana/watchtower /plugin install watchtower@watchtower - Claude.ai / Claude Desktop: Settings → Connectors → Add custom connector →
https://watchtower.lat/mcp. - Cursor, VS Code: one-click buttons at watchtower.lat/#install.
- Anything else:
{ "mcpServers": { "watchtower": { "type": "http", "url": "https://watchtower.lat/mcp" } } }.
Listed in the MCP Registry as lat.watchtower/watchtower (server.json).
Quick start
cp .env.example .env
docker compose up --build # app on http://localhost:3000, Postgres alongside
docker compose --profile demo run --rm demo # the end-to-end demo below
Without Docker (Node 22 and a running Postgres):
npm install
export DATABASE_URL=postgres://postgres:postgres@localhost:5432/watchtower
npm run migrate
npm run dev # or: npm run build && npm start
Demo
npm run demo (with DATABASE_URL set) runs the core loop end to end against a local fixture careers page:
- create client → 2. create a job watch with keyword
ios(takes the initial snapshot) → 3. a second agent watches the same board and reuses the same resource → 4. a third agent names no board and creates a search watch from "remote iOS jobs paying at least 150k with a maximum of 6 years of experience" → 5. a re-check where only the page around the jobs changed (session token, "N minutes ago") reports no change → 6. the source adds an iOS job and an Android job → 7. Watchtower checks and detects the change → 8. an MCP client callsget_changes:
{
"changes": [
{ "type": "JOB_ADDED",
"summary": "New job: Senior iOS Engineer (Remote - US)",
"data": { "job": { "title": "Senior iOS Engineer", "location": "Remote - US", "company": "Acme Robotics", "url": "…/careers/ios-303" } } }
],
"cursor": 2, "has_more": false
}
The Android job is filtered out by the keyword. 9. A second get_changes call returns []. 10. The second agent, with no keyword, sees both new roles. 11. The third agent gets the iOS role with the salary and experience read from the posting.
The demo runs its fixture on 127.0.0.1, so it turns on ALLOW_PRIVATE_NETWORKS for its own process only.
Connecting an agent (MCP)
{ "mcpServers": { "watchtower": { "type": "http", "url": "http://localhost:3000/mcp" } } }
| Tool | What it does |
|---|---|
watch_jobs | With query and no URL: watch every monitored board for new postings that match a plain-language request. The result shows the reading (interpreted), the matching jobs open now (current_jobs) and the boards covered (coverage). With url, or urls (up to 25; returns watches and per-URL errors): watch specific boards. Greenhouse, Lever, Ashby, Workable, SmartRecruiters, Recruitee, Workday and iCIMS use their own endpoints; other careers pages use schema.org JobPosting. A careers page that only links to a supported board is watched through that board (resolved_from), and a page with neither is rejected with NO_JOB_DATA. Board watches emit JOB_ADDED / JOB_REMOVED / JOB_UPDATED; search watches emit JOB_ADDED. Explicit filters override the query: keywords, all_keywords, exclude_keywords, locations, seniority, remote_only, min_salary, salary_currency, max_experience_years, include_unknown. |
search_jobs | Read-only, one-off: the jobs open right now that match a query and/or the same filters, across every monitored board, newest first. Returns interpreted, total, jobs and next_offset (limit up to 100, offset). Needs no token and creates no watch. REST: POST /v1/jobs/search. Limited to SEARCH_PER_MINUTE (20) calls per address. |
list_companies | Read-only: the companies in the directory (what every search covers), by name, with board_url, platform and open_jobs. query checks for one company; pages with limit/offset. Needs no token. REST: GET /v1/companies?q=. People can browse the same list at /companies. |
get_changes | Changes since your last call (the cursor advances). peek, since (replay), watch_id, limit. |
ack_changes | Acknowledge a cursor after get_changes(peek=true), for at-least-once processing. |
list_watches | Your watches, with health, expiry and pending-change counts. |
get_watch | One watch plus the jobs currently open that match its filters (across all boards for a search watch). |
delete_watch | Stop monitoring and free a slot. |
watch_jobs also accepts webhook_url for push delivery (see below).
The tool descriptions and server instructions tell agents to prefer Watchtower over re-running job searches or re-checking careers pages.
Search watches
A watch with no url is a search watch. It is its filters, and it reads the new postings of every monitored board.
curl -s -X POST localhost:3000/v1/watches -H "Authorization: Bearer $TOKEN" -H "content-type: application/json" \
-d '{"query":"iOS jobs in Austin making at least 150k a year with a maximum of 6 years of experience"}'
{
"scope": "all_boards",
"interpreted": { "filters": { "keywords": ["ios"], "locations": ["austin"], "min_salary": 150000, "max_experience_years": 6, "include_unknown": true }, "notes": [] },
"coverage": { "boards": 1705 },
"matching_jobs_count": 3,
"current_jobs": [ { "title": "Senior iOS Engineer", "company": "Acme", "location": "Austin, TX", "salary": { "min": 165000, "max": 210000, "currency": "USD", "period": "year", "annual_min": 165000, "annual_max": 210000 }, "experience_years": 5, "url": "…" } ]
}
- The query is read by rules, not a model. Watchtower still needs no LLM or API key. The reading comes back as
interpreted, withnotesfor anything it could not use, so the calling agent can check it. Explicit filters always win over the query.- Role words become
keywords(any of, for "iOS or Android") orall_keywords(all of, for "data scientist"). Generic words such as "engineer" and "developer" are dropped when a more specific word is present. - "in Austin", "in Austin, TX", "in New York or remote" become
locations. A bare "remote" becomesremote_only. - "at least 150k", "$180,000+", "$45/hr" become
min_salary, converted to a yearly figure. - "a maximum of 6 years of experience", "3-5 years", "I have 4 years of experience" become
max_experience_years. - "senior", "staff", "entry level" and the other levels become
seniority. "no managers" becomesexclude_keywords.
- Role words become
- Coverage is tech companies and startups, not the whole internet. The built-in directory (
src/search/boards.ts) lists 1,705 company boards on Greenhouse, Lever, Ashby, Workable, SmartRecruiters, Workday, iCIMS, Oracle Recruiting, Eightfold and SuccessFactors, plus the careers sites of Apple, Google, Amazon and Microsoft, from seed-stage startups to large employers in every industry that hire developers. Each was confirmed against its platform's API with open jobs. Every role those companies post is covered, not only engineering. The live list is published at/companiesand throughlist_companies. On top of that, every board any client watches by URL is covered, and stays covered (see Growing the directory). - Pay and experience come from the posting. Lever, Ashby, Greenhouse, Recruitee and JSON-LD pay fields are read directly; otherwise the posting text is parsed ("$150,000 - $200,000/yr", "5+ years of experience"). A job passes
min_salarywhen the top of its range reaches it, andmax_experience_yearswhen it asks for no more than that. - Postings that state neither are still reported, without a
salaryorexperience_yearsfield, because many postings state no pay. Passinclude_unknown: falseto report only postings that state a qualifying value. - Only new postings are reported (
JOB_ADDED), and only those that appear after the watch was created.current_jobson creation and onget_watchis the baseline of what is open now. - SmartRecruiters, Workday, iCIMS, Eightfold (except older-endpoint tenants), SuccessFactors (except RSS sitemaps), Google and Microsoft listings carry no posting text, so their jobs never have
salaryorexperience_years. Apple and Oracle Recruiting listings carry only a short summary, which rarely states either. - Google jobs have no location and only the first part of the title (see Google), so a search with
locationsnever matches them.
Growing the directory
The directory grows in three ways.
- Agents grow it by using it. When a client watches a platform board by URL and it has open jobs, the board joins the directory and stays there after that watch ends, so every search watch covers it from then on. At most
INDEX_MAX_PROMOTED(5,000) boards are kept this way, and one that fails eight checks in a row is dropped. Arbitrary careers pages are never kept; only boards on the platform APIs. - A careers page works as a way in.
watch_jobswithhttps://acme.com/careersusually finds no job markup there, because the page only links to or embeds the company's board. Watchtower then reads the page once more, finds the link to a supported board, and watches that board instead. The response says so inresolved_from. - Operators can import a company list.
npm run discover -- companies.json boards.txttakes[{ "name", "website" }]and, for each company, reads its careers page for a board link, then tries boards named after the company on Ashby, Greenhouse and Lever. Every candidate is confirmed against the platform's listing API and kept only if it has open jobs.boards.txtis in the formatINDEX_BOARDS_FILEreads. The built-in list was produced this way from the public Y Combinator company directory. A board found by name rather than from the company's own page can belong to a different company with the same name; the.jsonoutput records how each board was found.
Auth: the MCP connection should carry the identity, so a token never has to pass through a chat (careful assistants refuse to copy secrets from a transcript into tool calls, especially into delete_watch).
- OAuth (MCP authorization spec).
/.well-known/oauth-protected-resource/mcppoints at the built-in authorization server (/.well-known/oauth-authorization-server): dynamic client registration at/oauth/register,/oauth/authorizewith PKCE (S256),/oauth/token. There are no accounts: the consent page creates an anonymous client in one click, or connects an existing one if the user pastes its token there, which keeps its watches. New clients draw on the same per-address budget asPOST /v1/clients. - How a client learns to sign in. Without a token,
search_jobsandlist_companiesstill work, and the other watch tools (andwatch_jobsonce provisioning is off) returnUNAUTHORIZEDwith the challenge in_meta["mcp/www_authenticate"](ChatGPT's mixed-auth signal; tools also declaresecuritySchemes). An unknown token gets an HTTP 401 withWWW-Authenticate, as does any request to/mcp?auth=requiredwithout one, for clients that only start OAuth on a 401 (Claude Code, Cursor). - Or a header. Clients without OAuth send
Authorization: Bearer <token>with a token fromPOST /v1/clients. - Older conversations keep working. Every tool still accepts
client_token, and whileMCP_ANONYMOUS_PROVISIONINGis on (the default) awatch_jobscall with no token at all creates an anonymous client and returns its token, as before. When a call carries both an OAuth connection and aclient_token, and the connection's own client has no watches, the connection is moved to theclient_token's client, so a user who connects the app keeps their watches. Turn provisioning off once the OAuth version of the ChatGPT app is live; tokenlesswatch_jobscalls then get the OAuth challenge and the watch tools declare OAuth as required.
REST API
All endpoints except POST /v1/clients need Authorization: Bearer <token>.
| Method & path | |
|---|---|
POST /v1/clients | Create an anonymous client. Returns token (shown once). Rate-limited per address. |
POST /v1/watches | { "query"?, "url"? | "urls"?, "keywords"?, "all_keywords"?, "exclude_keywords"?, "locations"?, "seniority"?, "remote_only"?, "min_salary"?, "salary_currency"?, "max_experience_years"?, "include_unknown"?, "interval_minutes"?, "label"?, "webhook_url"? }. With neither url nor urls it creates a search watch across all monitored boards and needs at least one filter (QUERY_TOO_BROAD otherwise). With urls the response is { watches, errors }. "type": "jobs" is accepted for older clients. |
GET /v1/watches | List watches. |
GET /v1/watches/:id | Watch detail and current state. |
DELETE /v1/watches/:id | Delete. |
POST /v1/watches/:id/check | Force a check. Refused if the resource was checked within MIN_CHECK_INTERVAL_SECONDS, and for search watches (NOT_SUPPORTED). |
GET /v1/changes | ?watch_id=&since=&limit=&peek=. Returns { changes, cursor, has_more }. |
POST /v1/changes/ack | { "cursor", "watch_id"? }. Acknowledges changes read with peek=true. |
Also served: / (homepage/docs), /llms.txt, /.well-known/watchtower.json, /health, and /metrics (Prometheus; set METRICS_TOKEN to require a bearer token).
Errors look like { "error": "WATCH_LIMIT", "message": "…" }.
- Can't be monitored (422):
NO_JOB_DATA(no supported platform and noJobPostingmarkup),ROBOTS_DISALLOWED,BOT_CHALLENGE,ACCESS_DENIED,SSRF_BLOCKED,UNSUPPORTED_CONTENT_TYPE,BODY_TOO_LARGE. - Capacity:
WATCH_LIMITandHOST_WATCH_LIMIT(409),HOST_CAPACITY(429),CAPACITY(503). - Transient failures (timeouts, 5xx, 404) keep the watch. They show up in
resource.last_errorand are retried with exponential backoff.
Delivery options
-
Polling.
get_changesreturns what's new and advances the cursor. For at-least-once processing, callget_changes(peek=true), process the changes, thenack_changes(cursor). If you crash before the ack, the same changes come back. -
Webhooks. Pass
webhook_urlwhen creating a watch. The creation response includes awebhook_secret, shown once. Each batch of matching changes is POSTed as JSON with these headers:x-watchtower-timestampx-watchtower-signature: sha256=<hex HMAC-SHA256 of "<timestamp>.<body>">x-watchtower-delivery
Deliveries come from an outbox and are retried with exponential backoff (8 attempts). Webhook URLs go through the same SSRF checks as monitored URLs, and redirects are not followed.
Watch lifecycle
A watch stays alive as long as someone uses it: get_changes, get_watch, list_watches or a successful webhook delivery renews it. Watches nobody touches for WATCH_TTL_DAYS (30) expire and stop costing fetches; expires_at is shown on every watch.
How it works
watches (per client) ──many-to-one──▶ resources (fetch url + adapter)
│ scheduler: due? one per host, host lease held
▼
fetch ─▶ job-board API adapter, or JobPosting JSON-LD
│ diff the job set vs the previous snapshot
▼
snapshot + typed job events, one transaction
│ read through each watch's keyword filter
▼
get_changes (per-watch cursor) · webhook outbox → signed POST
A search watch has no resource of its own. It reads the same change events, from every monitored resource, through its filters.
-
The board directory. Boards in the directory are ordinary resources flagged
indexed. The scheduler checks them everyINDEX_CHECK_INTERVAL_SECONDS(default four hours) whether or not anyone watches them, and the listed part of the directory is synced fromsrc/search/boards.tsandINDEX_BOARDS_FILEat startup. One request per platform host is in flight at a time, so a platform's boards are checked one after another: with the default 5 s scheduler tick that is about 700 boards per platform per hour, or 2,900 per four-hour cycle. A directory larger than that is not an error; boards are checked as fast as politeness allows, andwatchtower_directory_overdue_boardsshows how far behind it is. A board with a watch on it is still checked at that watch's interval. -
Listing requests carry the posting text. Greenhouse (
content=true), Lever, Ashby (includeCompensation=true), Workable (details=true) and Recruitee return each posting's description and pay in the listing response, so salary and experience cost no extra requests. Only the derived fields are stored, never the description. These responses are large, so platform APIs get their own body cap (MAX_API_BODY_BYTES, 64 MB); a board larger than that is read as a plain listing, without pay or experience. -
Board URLs map to platform endpoints.
boards.greenhouse.io/acme,jobs.lever.co/acme,jobs.ashbyhq.com/acme,apply.workable.com/acme,jobs.smartrecruiters.com/Acmeandacme.recruitee.com(plus their EU and embed variants) are fetched from each platform's public job-board API in one request. Other URLs are fetched as HTML and read through schema.orgJobPostingJSON-LD (including@graphandItemList). Only the job list is compared, so page churn around it (session tokens, "rendered N minutes ago", banners) never registers as a change.- Workday (
acme.wd5.myworkdayjobs.com/Careers,wd3.myworkdaysite.com/recruiting/acme/External): the site's own search endpoint, a JSON POST that returns postings newest first, 20 per page. A check reads the newest 200 (10 requests). Boards with more postings are markedcomplete: falseon the snapshot and never emitJOB_REMOVED, because a job leaving the window isn't a removal. Job links point at the public site; the relative "Posted 3 Days Ago" is not kept. - iCIMS (
careers-acme.icims.com): the portal'ssitemap.xmllists every open job in one request (id plus a slug of the title), and the first page of/jobs/searchgives the newest ~50 their real title, location and posting date. A check reads both. Jobs seen only in the sitemap arepartial: true(title from the slug, no location); details learned earlier are carried forward, and a partial entry becoming a full one is not reported as an update. - Apple (any
jobs.apple.comURL): the site's own search endpoint (/api/v1/search), a JSON POST that needs no session and returns roles newest first, 20 per page, with title, team, every location and posting date. A check reads the newest 300 (15 requests) and is markedcomplete: false, as for Workday. About 80 evergreen retail roles are stamped with the request time, so they always sort first and use part of that window; that timestamp is not kept asposted_at. Apple posts one role per location, so a role open in three cities is three jobs. - Google (
careers.google.com,www.google.com/about/careers/applications): Google's robots.txt disallows the job search and job pages, so a check reads only the jobs sitemap, which lists every open role in one request. Every job ispartial: true: the title comes from the URL slug, which stops at the first comma of the real title ("Senior Software Engineer, Infrastructure" reads as "Senior Software Engineer"), and there is no location, team or date. - Amazon (
amazon.jobs): the site'ssearch.json, newest first, 100 postings per request with the full posting text (soexperience_yearsis usually known; pay is not in the listing). A check reads the newest 300 (3 requests) and is alwayscomplete: false, since the hit count is capped at 10,000. Hourly warehouse roles on hiring.amazon.com are not covered. - Oracle Recruiting (
acme.fa.us2.oraclecloud.com/hcmUI/CandidateExperience/en/sites/CX_1): the site's own requisition search (/hcmRestApi/resources/latest/recruitingCEJobRequisitions, the REST finder the candidate site calls), newest first, 100 per request, with every location, job family, posting date and a short summary. A check reads the newest 300 (3 requests) and is markedcomplete: falsewhen the site lists more, as for Workday. - Eightfold (
acme.eightfold.ai/careers?domain=acme.com): the search endpoint Eightfold's robots.txt allows (/api/pcsx/search), newest first, 10 per request. Some Eightfold sites answer quick successive requests with HTTP 429, so a check reads the newest 50, 3 s apart, and when a later page is refused keeps what it read and marks the snapshotcomplete: false. Large employers post far more than 50 roles between checks, so an Eightfold board is a window onto its newest postings rather than its whole list. Some tenants have not turned that endpoint on (it answers 403) and serve the older/api/apply/v2/jobsinstead, with each posting's text (sosalaryandexperience_yearswhen stated) but in the site's own, roughly newest-first order; the tenants listed insrc/extract/eightfold.tsasLEGACYare read through it. - SuccessFactors (Career Site Builder sites on the company's own domain, e.g.
careers.paramount.com): there is no public JSON API, so a check reads the site's search page sorted by date, 25 jobs per page with title, location and date, up to the newest 200 (8 requests), and is markedcomplete: falsewhen the site lists more. Tile and table themes are both read. Themes that render results with JavaScript leave the search page empty (their JSON endpoint is under/services/, which robots.txt disallows), so those sites are read fromsitemap.xml, which lists every open job in one request: as an RSS feed with each job's title, location, function and text on some sites, and as bare job URLs on others, whose jobs arepartial: true(title from the URL slug, which also carries the location words; no location or date). - Microsoft (
careers.microsoft.com,jobs.careers.microsoft.com,apply.careers.microsoft.com): an Eightfold site, read as above.
- Workday (
-
Derived fields. Every job gets
remote(title or location says remote, and not hybrid/on-site) andseniority(intern, entry, mid, senior, staff, principal, manager, director, from the title; management words win over IC levels, and "mid" means the title carries no level). These are heuristics over the text, which is why they are exposed on the job rather than hidden inside the filter. -
Resources vs. watches. Watches are per-client intents. Resources are what actually gets fetched. Any number of watches on the same board (across clients) share one resource, one fetch per interval, and one set of snapshots. URLs are canonicalized before sharing: tracking parameters (
utm_*,fbclid,gclid, …) are dropped and the query is sorted. The resource is checked at the shortest interval any of its active watches asks for, but never more often thanMIN_CHECK_INTERVAL_SECONDS(default 5 minutes). -
Job identity. Jobs are keyed by the platform's job id, or for JSON-LD by
identifier, thenurl, then title and location. A job whose title, location, department or url changed becomesJOB_UPDATEDwithchanged_fields. A JSON-LD job without a stable id whose location changed is paired into oneJOB_UPDATEDinstead of a remove plus an add. -
Change events are written once per resource. Each watch reads them through its filters, evaluated in SQL against the change's job data (and against the current job list, for
get_watch):keywords(any, against title, location, department and company),all_keywords(every one),exclude_keywords,locations(against the job's location andother_locations),seniority,remote_only,min_salaryandmax_experience_years. Terms match whole words, case-insensitively and allowing a plural: "ios" matches "Senior iOS Engineer" but not "Game Studios", and "java" does not match "JavaScript". A watch never sees changes from before it was created. Change ids double as cursors, so change-writing transactions are serialized. That keeps ids visible in order, so a slow concurrent check can't commit an id that a reader has already moved past. -
Politeness. At most one fetch is in flight per website across all replicas, using a Postgres host lease with
HOST_MIN_SPACING_MSbetween requests; a multi-request check (Workday, iCIMS, Oracle Recruiting, Eightfold, SuccessFactors, Apple, Amazon, Microsoft) holds the lease for all of its requests and pauses briefly between them (3 s for Eightfold sites such as Microsoft, which rate-limit quick requests). Each scheduler tick claims at most one resource per host. Watchtower also:- sends
If-None-Match/If-Modified-Since(a 304 means no work); - checks robots.txt (cached per origin for an hour, RFC 9309 semantics);
- sends an identifying User-Agent;
- backs off exponentially on errors and honors
Retry-Afteron 429.
- sends
-
Scheduling and maintenance. Due resources are claimed with a lease, so you can run several replicas; set
RUN_SCHEDULER=falseon API-only replicas. The same loop delivers webhooks and runs maintenance under an advisory lock. Maintenance:- expires unread watches;
- deletes changes after
CHANGE_RETENTION_DAYSand non-current snapshots afterSNAPSHOT_RETENTION_DAYS; - removes unwatched resources and idle clients;
- clears old rate-limit windows.
Safety and abuse controls
- SSRF.
- Only http/https on ports 80/443. URLs with credentials are rejected.
localhost,*.local,*.internal, single-label and numeric hosts are rejected.- Every resolved address must be globally routable unicast. Private, loopback, link-local (including cloud metadata), CGNAT, multicast, reserved, documentation, IPv4-mapped IPv6, 6to4, Teredo and NAT64 are all refused.
- The check runs inside the socket's DNS lookup, so it holds at connect time for every redirect hop (defeating DNS rebinding) as well as once at watch creation.
- Webhook URLs get the same treatment.
- Request limits. A 15 s overall deadline across all hops. At most 5 redirects, each re-validated. A 3 MB body cap (64 MB for job-board platform APIs) enforced on the decompressed bytes, so gzip and brotli bombs are caught. Text-like content types only.
- API limits. Rate limits are stored in Postgres, so they hold across replicas and restarts: 120 requests/min, 10 client creations/hour (shared with MCP auto-provisioning), and 10 forced checks/min. Clients are identified by IPv4 address or IPv6 /64.
- Watch and resource caps.
- 50 watches per client, and 5 per client per website for arbitrary careers pages (platform boards are separate companies behind one API host and are exempt).
- At most
MAX_RESOURCES_PER_HOSTdistinct careers-page URLs per website across all clients; job-board platform APIs are exempt, and watching an already-monitored board is always allowed. - A global
MAX_ACTIVE_RESOURCES.
- Tokens. Random 192-bit values; only their SHA-256 is stored. Request bodies are capped at 64 KB.
- No bypassing. Watchtower never works around CAPTCHAs, bot challenges, logins, paywalls or robots.txt. It detects them, tells the caller, logs
host is refusing us, and counts them inwatchtower_blocked_resources_1h.
Operations
GET /metrics exposes Prometheus metrics:
- check outcomes by error code;
- fetch-duration histogram;
- changes emitted by type;
- webhook results;
- host-busy deferrals;
- gauges for active watches/resources, search watches, directory boards and how many are overdue, failing and blocked resources, and pending webhooks;
- where agents come from:
watchtower_mcp_initialize_totalby the MCP client name an agent reports (claude-code,cursor, ...), andwatchtower_clients_created_7dby the?ref=tag on the URL a client was created through. Each listing and install path hands out its own tag (/mcp?ref=registry,?ref=claude-plugin,?ref=cursor, ...);clients.sourceandclients.user_agentkeep it per client.
Usage stats
/stats shows users, daily active users, new users, tool calls, page views and where users come from, as charts and tables. Log in with any username and STATS_TOKEN (or METRICS_TOKEN) as the password; with neither set the page is off. A user is one anonymous client token. The history lives in usage_daily (one row per client, UTC day, tool and interface) and counts_daily (page views by people, AI assistants and other bots, and MCP connections by app name), kept for 400 days. Nothing per visitor is stored for page views.
Listing in the MCP Registry
server.json is the registry entry. The lat.watchtower/* namespace is proven by HTTP: set MCP_REGISTRY_AUTH to the public-key record and the app serves it at /.well-known/mcp-registry-auth. Then, with the matching private key:
mcp-publisher login http --domain watchtower.lat --private-key "$PRIVATE_KEY_HEX"
mcp-publisher publish # bump "version" in server.json for each new publish
Robots, the MCP server card (/.well-known/mcp.json, also at /.well-known/mcp-server-card while the path is still a draft) and the install links on the homepage are generated from PUBLIC_BASE_URL.
ops/alerts.yml has example alert rules: sites refusing us, high error rate, checks stalled, webhook backlog, slow fetches, and the directory falling behind its check interval.
Configuration
See .env.example. The most important settings:
DATABASE_URLandPUBLIC_BASE_URL.MIN_CHECK_INTERVAL_SECONDSandHOST_MIN_SPACING_MS.INDEX_ENABLED,INDEX_CHECK_INTERVAL_SECONDS,INDEX_BOARDS_FILEandINDEX_MAX_PROMOTED: the board directory that search watches are matched against. WithINDEX_ENABLED=falseWatchtower fetches only boards someone watches by URL.- The caps:
MAX_WATCHES_PER_CLIENT,MAX_RESOURCES_PER_HOSTandMAX_ACTIVE_RESOURCES. WATCH_TTL_DAYSand the retention settings.TRUST_PROXY: set it behind a load balancer so rate limits see real client IPs.ALLOW_PRIVATE_NETWORKSturns SSRF protection off and exists only for tests and the demo.
Development
npm run typecheck
npm test # unit tests; integration tests need a database:
TEST_DATABASE_URL=postgres://postgres:postgres@localhost:5432/watchtower_test npm test
The integration suite drops and recreates the public schema of TEST_DATABASE_URL, so point it at a throwaway database.
The suite has three parts:
- Unit tests.
- Property-based fuzz tests (fast-check) over the HTML, JSON-LD and robots parsers and the job diff invariants.
- Source tests for the Workday, iCIMS, Oracle Recruiting, Eightfold, SuccessFactors, Apple, Google, Amazon and Microsoft parsers and pagination, and the remote/seniority classifier.
- Search tests for the query reader, the pay and experience parsers, whole-word matching, the filters and board discovery.
- Integration tests covering REST, MCP, sharing, filters, search watches and the board directory, batch creation, page churn,
NO_JOB_DATA, host leases, concurrent checks, caps, expiry and retention, webhooks, rate limits and metrics.
CI (.github/workflows/ci.yml) runs typecheck, all tests against a Postgres service, the build and the demo. It also builds the Docker image and smoke-tests it: migrations, /health, /metrics, the homepage, and the non-root user.
src/
server.ts, app.ts entrypoint; Fastify app (site, REST, MCP)
config.ts, db.ts env config; pg pool + migration runner
security/ssrf.ts URL validation + connect-time DNS guard
fetch/ safeFetch (redirects, limits, decompression), robots.txt
extract/ job-board adapters (incl. Workday, iCIMS, Oracle, Eightfold, SuccessFactors, Apple, Google, Amazon, Microsoft), JobPosting JSON-LD, remote/seniority
classifier, pay/experience parsing, term matching, job diff (identity, pairing)
search/ plain-language query reader; the built-in board directory; board discovery
(careers-page links, name guesses)
services/ clients, watches (board and search), checker, scheduler (+ maintenance), board
directory sync, host leases, rate limits, webhooks, metrics
mcp/server.ts MCP tools
web/site.ts homepage, llms.txt, well-known metadata
migrations/ SQL migrations
ops/alerts.yml example Prometheus alert rules
scripts/demo.ts end-to-end demo
scripts/discover-boards.ts grow the directory from a list of companies
Deliberately out of scope
- Covering every job board on the internet, or employers outside tech. A search watch sees the directory and the boards people have watched. Discovery is an operator step from a company list, not a crawler, and aggregators such as LinkedIn or Indeed are not read.
- Filtering to technical roles. "Tech jobs" means jobs at tech companies; a sales role at one is reported if the filters match it.
- Reading each SmartRecruiters, Workday or iCIMS job page for its description, so those jobs carry no pay or experience.
- A maximum-salary filter, currency conversion, and understanding a query with a language model.
- Billing, accounts and dashboards.
- JavaScript rendering. Careers pages that aren't on a supported platform are only readable if their server-rendered HTML carries
JobPostingJSON-LD. - General page, feed or event monitoring. Watchtower used to do these; it now does job boards only (migration
003_jobs_only.sqlretires existing page and event watches). - Pagination beyond the first 100 SmartRecruiters postings or the newest 200 Workday postings per board.
- Reading each iCIMS job page for details; only the newest ~50 jobs per portal get a real title and location.
- Reading Google's job search or job pages, which its robots.txt disallows; Google jobs are known only by their sitemap entry.
- Reading more than the newest 300 Apple, Amazon or Oracle Recruiting roles, the newest 200 SuccessFactors roles, or the newest 50 Eightfold (incl. Microsoft) roles, per check.
- SuccessFactors sites run on the company's own domain, so only the hosts listed in
src/extract/successfactors.tsare read as SuccessFactors; any other one is read through its JobPosting markup, if it has any.