Cheapest-LLM Router
MCP server that routes any prompt to the cheapest reachable free/cheap LLM (Kimi K2.6, Qwen, DeepSeek, Cloudflare Workers AI, Groq, Gemini...). Reuses Free & Cheap Tokens model channels. Pay-Per-Event: no monthly fee, pay only per tool call.
文件
Cheapest-LLM Router (CLR) - pick the cheapest LLM per prompt (neeenja/cheapest-llm-router) Actor
MCP server that routes any prompt to the cheapest reachable free/cheap LLM (Kimi K2.6, Qwen, DeepSeek, Cloudflare Workers AI, Groq, Gemini...). Reuses Free & Cheap Tokens model channels. Pay-Per-Event: no monthly fee, pay only per tool call.
- URL: https://apify.com/neeenja/cheapest-llm-router.md
- Developed by: Neeen Ja (community)
- Categories: MCP servers, AI, Developer tools
- Stats: 1 total users, 0 monthly users, 0.0% runs succeeded, 0 bookmarks
- User rating: No ratings yet
Pricing
from $0.50 / 1,000 mcp cache_route call - in-process cache demos
This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.
Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event
What's an Apify Actor?
An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes. In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours, and optionally produces a well-defined JSON output, datasets with results, or files in key-value store. In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.
How to integrate an Actor?
If asked about integration, you help developers integrate Actors into their projects. You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.
For examples already wired to this Actor's own input schema, see the API section below.
Each client library has reference documentation the quickstart does not restate: JavaScript/TypeScript (npm install apify-client) and Python (pip install apify-client).
README
Cheapest-LLM Router (CLR)
Route any prompt to the cheapest reachable free/cheap LLM — automatically. Reuses the free-model channels of Free & Cheap Tokens (FACT) (Kimi K2.6 · Qwen · DeepSeek · Cloudflare Workers AI · Groq · Gemini …). A zero-dependency MCP server (runs over stdio, no
npm install) with 4 tools:route·cost_compare·list_models·cache_route.
| It lives at | Link |
|---|---|
| MCP endpoint (hosted) | https://neeenja--cheapest-llm-router.apify.actor/mcp |
| Apify Store | https://apify.com/neeenja/cheapest-llm-router |
| Source | https://github.com/PanStories/cheapest-llm-router |
English
Cheapest-LLM Router is an MCP server that, given a prompt, automatically picks the cheapest model that can actually serve it — free models first, then the lowest-cost paid fallback — and returns the cost in both USD and CNY plus a fallback chain and plain-English reasoning. It does not call any model; it only does the routing math, so it costs nothing to run and nothing to call except the tiny Apify Pay-Per-Event fee when hosted.
Why this exists
- Zero new learning — it reuses the curated model list and free channels already built for FACT.
- Zero marginal cost — routing is pure local computation; no paid API is ever called.
- 100% brand synergy — it complements FACT's "free/cheap tokens" positioning and cross-sells to the same users.
- Monetization — advanced routing / caching / the cost-comparison report layer is the hosted paid tier.
- Cold-start revenue estimate — ¥0–¥500/month via FACT-user conversion.
What you get (MCP tools)
| Tool | What it does |
|---|---|
route | prompt → cheapest reachable model + USD/CNY cost estimate + fallback chain + reasoning |
cost_compare | ranks every reachable model by cost and reports the max savings vs the most expensive |
list_models | filtered registry listing (capability / region / free-only) |
cache_route | same as route, but demonstrates the in-process cache (cached: true on repeat hits) |
route input
{
"prompt": "Analyze the core risks of this earnings report",
"max_output_tokens": 512,
"required_capabilities": ["chinese", "reasoning"],
"region": "CN",
"priority": "cost",
"include_paid": true
}
route output (excerpt)
{
"ok": true,
"chosen": { "name": "Qwen3-Plus (via Aliyun Bailian)", "free": true, "cost_usd": 0 },
"fallback_chain": [ "Kimi K2 (Moonshot direct trial)", "DeepSeek-V3 (via SiliconFlow)" ],
"estimated_cost_usd": 0,
"estimated_cost_cny": 0,
"reasoning": "Routed by lowest cost. Estimated 18 input + 512 output tokens. Cheapest reachable is Qwen3-Plus ... free. No token cost."
}
Quick start
1. Run locally (stdio, zero install)
git clone https://github.com/PanStories/cheapest-llm-router.git
cd cheapest-llm-router
node src/server.mjs # plain node — no npm install needed
2. Verify (no dependencies)
node --test tests/ # routing core unit tests
node scripts/mcp-smoke.mjs # end-to-end MCP protocol smoke test
node scripts/http-e2e-test.mjs # Streamable HTTP transport test (needs deps)
node scripts/showcase.mjs # generate showcase.html from real routing output
3. Host remotely (Apify Pay-Per-Event)
npm install # pulls express + MCP SDK + apify (hosted variant only)
apify login && apify push # requires Apify KYC first
The hosted MCP endpoint is then reachable at https://neeenja--cheapest-llm-router.apify.actor/mcp (Bearer-token auth).
Connect it to your client
stdio (mcp.json) — for local desktop clients:
{
"mcpServers": {
"cheapest-llm-router": {
"command": "node",
"args": ["/abs/path/cheapest-llm-router/src/server.mjs"]
}
}
}
hosted (remote, via mcp-remote) — Apify's gateway requires a per-request Bearer token:
{
"mcpServers": {
"cheapest-llm-router": {
"command": "npx",
"args": [
"mcp-remote",
"https://neeenja--cheapest-llm-router.apify.actor/mcp",
"--header", "Authorization: Bearer <YOUR_APIFY_TOKEN>"
]
}
}
}
Pricing (Pay-Per-Event)
No subscription, no monthly fee. You pay only when a tool actually runs:
| Event | Price (USD) |
|---|---|
initialize / tools/list / report-issue | free |
route | $0.0005 |
cost_compare | $0.001 |
list_models / cache_route | $0.0005 |
Free events are never billed, so agents can connect and discover tools at zero cost. Apify also grants ~$5/month of free platform credits (≈ thousands of calls).
How routing works
- Token estimate — CJK ≈ 1.6 tokens/char, other text ≈ 0.25 tokens/char (no external tokenizer).
- Reachability filter —
required_capabilities ⊆ model capabilitiesandregionmatch (CN= mainland-accessible without a VPN). - Cost — free = $0; paid =
(in×pin + out×pout) / 1e6. - Ranking —
cost(default: free first, then by free-quota then latency; paid by cost) ·latency·quality. - Fallback chain — the next 3 reachable models.
⚠️ Prices are indicative (USD per 1M tokens) and drift with providers; verify on each provider's pricing page before production.
What it does NOT do (honest boundaries)
- It does not send real inference requests — it only selects a route.
- It does not cache/store your prompts (a persistent per-account cache is a hosted paid feature).
- It does not guarantee free quotas are live (provider-controlled; may return 429).
- Prices change; always trust the provider's official page.
Development
| Script | Purpose |
|---|---|
node src/server.mjs | stdio MCP server (zero dependency) |
node src/http.mjs | HTTP MCP server (PORT=3000) for local/dev |
node --test tests/ | routing core unit tests |
node scripts/mcp-smoke.mjs | stdio MCP protocol smoke test |
node scripts/http-e2e-test.mjs | HTTP transport e2e (PORT=3100) |
node scripts/showcase.mjs | build showcase.html from real output |
Project layout:
cheapest-llm-router/
├── data/models.json curated model registry (reuses FACT's verified list)
├── src/core/router.mjs routing engine (zero-dependency, single source of truth)
├── src/server.mjs zero-dependency stdio MCP server (default, verified)
├── src/http.mjs self-managed HTTP MCP server (Apify Standby)
├── src/handler.mjs MCP server factory (shared by both transports)
├── src/billing.mjs Apify Pay-Per-Event billing (hosted only)
├── tests/router.test.mjs unit tests (node --test)
├── scripts/ smoke / http-e2e / showcase scripts
└── .actor/ Apify actor.json + Dockerfile + schemas
License
MIT — fork, self-host, self-modify freely.
简体中文
Cheapest-LLM Router 是一个 MCP 服务器:给定一条 prompt,它会自动选出当下最便宜且能真正服务它的模型——优先免费模型,其次最低成本的付费兜底——并以美元和人民币双币种返回成本、兜底链与推理说明。它不会真正调用任何模型,只做路由计算,因此本地运行零成本,托管后除极低的 Apify 按事件计费外也无其他开销。
为什么做这个
- 零新学习——直接复用为 FACT 策展的模型清单与免费通道。
- 零边际成本——路由是纯本地计算,从不调用任何付费 API。
- 100% 品牌协同——与 FACT「免费/廉价 token」定位天然互补,可向同一批用户交叉转化。
- 变现点——高级路由 / 缓存 / 成本对比报表层即托管的付费能力。
- 冷启动月收入预估——¥0–¥500(靠 FACT 用户转化)。
你得到什么(MCP 工具)
| 工具 | 作用 |
|---|---|
route | 给定 prompt → 最便宜可达模型 + 美元/人民币成本估算 + 兜底链 + 推理说明 |
cost_compare | 把所有可达模型按成本排序,并给出相比最贵模型的最大节省额 |
list_models | 按 capability / region / free-only 过滤的模型清单 |
cache_route | 同 route,但演示进程内缓存(重复命中返回 cached: true) |
route 入参
{
"prompt": "分析这份财报的核心风险",
"max_output_tokens": 512,
"required_capabilities": ["chinese", "reasoning"],
"region": "CN",
"priority": "cost",
"include_paid": true
}
route 出参(节选)
{
"ok": true,
"chosen": { "name": "Qwen3-Plus (via Aliyun Bailian)", "free": true, "cost_usd": 0 },
"fallback_chain": [ "Kimi K2 (Moonshot direct trial)", "DeepSeek-V3 (via SiliconFlow)" ],
"estimated_cost_usd": 0,
"estimated_cost_cny": 0,
"reasoning": "Routed by lowest cost. Estimated 18 input + 512 output tokens. Cheapest reachable is Qwen3-Plus ... free. No token cost."
}
快速开始
1. 本地跑(stdio,零安装)
git clone https://github.com/PanStories/cheapest-llm-router.git
cd cheapest-llm-router
node src/server.mjs # 直接跑,无需 npm install
2. 验证(无需依赖)
node --test tests/ # 路由核心单元测试
node scripts/mcp-smoke.mjs # MCP 协议端到端冒烟测试
node scripts/http-e2e-test.mjs # Streamable HTTP 传输测试(需装依赖)
node scripts/showcase.mjs # 用真实路由输出生成 showcase.html
3. 远程托管(Apify 按事件计费)
npm install # 仅托管变体需要:express + MCP SDK + apify
apify login && apify push # 需先完成 Apify KYC
托管后的 MCP 端点:https://neeenja--cheapest-llm-router.apify.actor/mcp(Bearer token 鉴权)。
接入你的客户端
stdio(mcp.json)——本地桌面客户端:
{
"mcpServers": {
"cheapest-llm-router": {
"command": "node",
"args": ["/绝对路径/cheapest-llm-router/src/server.mjs"]
}
}
}
托管(远程,借助 mcp-remote)——Apify 网关要求每次请求带 Bearer token:
{
"mcpServers": {
"cheapest-llm-router": {
"command": "npx",
"args": [
"mcp-remote",
"https://neeenja--cheapest-llm-router.apify.actor/mcp",
"--header", "Authorization: Bearer <你的_APIFY_TOKEN>"
]
}
}
}
定价(按事件计费)
无订阅、无月费,只在工具真正运行时付费:
| 事件 | 价格(美元) |
|---|---|
initialize / tools/list / report-issue | 免费 |
route | $0.0005 |
cost_compare | $0.001 |
list_models / cache_route | $0.0005 |
免费事件永不计费,Agent 可零成本连接与发现工具。Apify 另送约 $5/月的免费平台额度(≈ 数千次调用)。
路由逻辑
- Token 估算:CJK 字符 ≈ 1.6 token,其他文本 ≈ 0.25 token(无外部 tokenizer)。
- 可达性过滤:
required_capabilities ⊆ 模型能力且region匹配(CN= 大陆免 VPN 可达)。 - 成本计算:免费模型 = $0;付费 =
(in×pin + out×pout) / 1e6。 - 排序:
cost(默认:免费优先 → 免费内按免费额度再按延迟;付费按成本升序)·latency·quality。 - 兜底链:次优 3 个可达模型。
⚠️ 价格为指示性(美元/百万 token),随厂商变动;生产前请以各厂商定价页为准。
本工具不做什么(诚实边界)
- 不替你发起真实推理请求——只选路由。
- 不缓存/存储你的 prompt(hosted 持久缓存为付费能力)。
- 不保证免费额度实时可用(额度由厂商控制,可能 429)。
- 价格随厂商变动,请以官方为准。
开发
| 脚本 | 用途 |
|---|---|
node src/server.mjs | stdio MCP 服务器(零依赖) |
node src/http.mjs | HTTP MCP 服务器(PORT=3000),本地/开发用 |
node --test tests/ | 路由核心单元测试 |
node scripts/mcp-smoke.mjs | stdio MCP 协议冒烟测试 |
node scripts/http-e2e-test.mjs | HTTP 传输端到端(PORT=3100) |
node scripts/showcase.mjs | 用真实输出生成 showcase.html |
项目结构:
cheapest-llm-router/
├── data/models.json 策展模型注册表(复用 FACT 已核验清单)
├── src/core/router.mjs 路由引擎(零依赖,唯一事实源)
├── src/server.mjs 零依赖 stdio MCP 服务器(默认、已验证)
├── src/http.mjs 自托管 HTTP MCP 服务器(Apify Standby)
├── src/handler.mjs MCP 服务器工厂(两个传输共用)
├── src/billing.mjs Apify 按事件计费(仅托管)
├── tests/router.test.mjs 单元测试(node --test)
├── scripts/ 冒烟 / HTTP e2e / 展示脚本
└── .actor/ Apify actor.json + Dockerfile + schema
许可证
MIT——可 fork、自部署、自托管、自修改。
繁體中文
Cheapest-LLM Router 是一個 MCP 伺服器:給定一條 prompt,它會自動選出當下最便宜且能真正服務它的模型——優先免費模型,其次最低成本的付費兜底——並以美元與人民幣雙幣種回傳成本、兜底鏈與推理說明。它不會真正呼叫任何模型,只做路由計算,因此本地執行零成本,託管後除極低的 Apify 按事件計費外也無其他開銷。
為什麼做這個
- 零新學習——直接複用為 FACT 策展的模型清單與免費通道。
- 零邊際成本——路由是純本地計算,從不呼叫任何付費 API。
- 100% 品牌協同——與 FACT「免費/廉價 token」定位天然互補,可向同一批用戶交叉轉化。
- 變現點——進階路由 / 快取 / 成本對比報表層即託管的付費能力。
- 冷啟動月收入預估——¥0–¥500(靠 FACT 用戶轉化)。
你得到什麼(MCP 工具)
| 工具 | 作用 |
|---|---|
route | 給定 prompt → 最便宜可達模型 + 美元/人民幣成本估算 + 兜底鏈 + 推理說明 |
cost_compare | 把所有可達模型按成本排序,並給出相比最貴模型的最大節省額 |
list_models | 按 capability / region / free-only 過濾的模型清單 |
cache_route | 同 route,但示範程序內快取(重複命中回傳 cached: true) |
route 入參
{
"prompt": "分析這份財報的核心風險",
"max_output_tokens": 512,
"required_capabilities": ["chinese", "reasoning"],
"region": "CN",
"priority": "cost",
"include_paid": true
}
route 出參(節選)
{
"ok": true,
"chosen": { "name": "Qwen3-Plus (via Aliyun Bailian)", "free": true, "cost_usd": 0 },
"fallback_chain": [ "Kimi K2 (Moonshot direct trial)", "DeepSeek-V3 (via SiliconFlow)" ],
"estimated_cost_usd": 0,
"estimated_cost_cny": 0,
"reasoning": "Routed by lowest cost. Estimated 18 input + 512 output tokens. Cheapest reachable is Qwen3-Plus ... free. No token cost."
}
快速開始
1. 本地執行(stdio,零安裝)
git clone https://github.com/PanStories/cheapest-llm-router.git
cd cheapest-llm-router
node src/server.mjs # 直接執行,無需 npm install
2. 驗證(無需依賴)
node --test tests/ # 路由核心單元測試
node scripts/mcp-smoke.mjs # MCP 協定端到端冒煙測試
node scripts/http-e2e-test.mjs # Streamable HTTP 傳輸測試(需裝依賴)
node scripts/showcase.mjs # 用真實路由輸出生成 showcase.html
3. 遠端託管(Apify 按事件計費)
npm install # 僅託管變體需要:express + MCP SDK + apify
apify login && apify push # 需先完成 Apify KYC
託管後的 MCP 端點:https://neeenja--cheapest-llm-router.apify.actor/mcp(Bearer token 鑑權)。
接入你的客戶端
stdio(mcp.json)——本地桌面客戶端:
{
"mcpServers": {
"cheapest-llm-router": {
"command": "node",
"args": ["/絕對路徑/cheapest-llm-router/src/server.mjs"]
}
}
}
託管(遠端,借助 mcp-remote)——Apify 閘道要求每次請求帶 Bearer token:
{
"mcpServers": {
"cheapest-llm-router": {
"command": "npx",
"args": [
"mcp-remote",
"https://neeenja--cheapest-llm-router.apify.actor/mcp",
"--header", "Authorization: Bearer <你的_APIFY_TOKEN>"
]
}
}
}
定價(按事件計費)
無訂閱、無月費,只在工具真正執行時付費:
| 事件 | 價格(美元) |
|---|---|
initialize / tools/list / report-issue | 免費 |
route | $0.0005 |
cost_compare | $0.001 |
list_models / cache_route | $0.0005 |
免費事件永不计費,Agent 可零成本連線與發現工具。Apify 另送約 $5/月的免費平台額度(≈ 數千次呼叫)。
路由邏輯
- Token 估算:CJK 字元 ≈ 1.6 token,其他文字 ≈ 0.25 token(無外部 tokenizer)。
- 可達性過濾:
required_capabilities ⊆ 模型能力且region匹配(CN= 大陸免 VPN 可達)。 - 成本計算:免費模型 = $0;付費 =
(in×pin + out×pout) / 1e6。 - 排序:
cost(預設:免費優先 → 免費內按免費額度再按延遲;付費按成本升序)·latency·quality。 - 兜底鏈:次優 3 個可達模型。
⚠️ 價格為指示性(美元/百萬 token),隨廠商變動;生產前請以各廠商定價頁為準。
本工具不做什么(誠實邊界)
- 不替你發起真實推理請求——只選路由。
- 不快取/儲存你的 prompt(託管持久快取為付費能力)。
- 不保證免費額度即時可用(額度由廠商控制,可能 429)。
- 價格隨廠商變動,請以官方為準。
開發
| 腳本 | 用途 |
|---|---|
node src/server.mjs | stdio MCP 伺服器(零依賴) |
node src/http.mjs | HTTP MCP 伺服器(PORT=3000),本地/開發用 |
node --test tests/ | 路由核心單元測試 |
node scripts/mcp-smoke.mjs | stdio MCP 協定冒煙測試 |
node scripts/http-e2e-test.mjs | HTTP 傳輸端到端(PORT=3100) |
node scripts/showcase.mjs | 用真實輸出生成 showcase.html |
專案結構:
cheapest-llm-router/
├── data/models.json 策展模型註冊表(複用 FACT 已核驗清單)
├── src/core/router.mjs 路由引擎(零依賴,唯一事實源)
├── src/server.mjs 零依賴 stdio MCP 伺服器(預設、已驗證)
├── src/http.mjs 自託管 HTTP MCP 伺服器(Apify Standby)
├── src/handler.mjs MCP 伺服器工廠(兩個傳輸共用)
├── src/billing.mjs Apify 按事件計費(僅託管)
├── tests/router.test.mjs 單元測試(node --test)
├── scripts/ 冒煙 / HTTP e2e / 展示腳本
└── .actor/ Apify actor.json + Dockerfile + schema
授權
MIT——可 fork、自部署、自託管、自修改。
Changelog
This Actor's version history is a separate document: https://apify.com/neeenja/cheapest-llm-router/changelog.md
Actor input Schema
Actor input object example
{}
Actor output Schema
mcpEndpoint (type: string):
Streamable HTTP MCP endpoint of this Actor run. Connect an MCP client here (Authorization: Bearer <APIFY_TOKEN>). POST JSON-RPC messages: initialize, tools/list, tools/call.
API
You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.
JavaScript example
import { ApifyClient } from 'apify-client';
// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
token: '<YOUR_API_TOKEN>',
});
// Prepare Actor input
const input = {};
// Run the Actor and wait for it to finish
const run = await client.actor("neeenja/cheapest-llm-router").call(input);
// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
console.dir(item);
});
// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs
Python example
from apify_client import ApifyClient
# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")
# Prepare the Actor input
run_input = {}
# Run the Actor and wait for it to finish
run = client.actor("neeenja/cheapest-llm-router").call(run_input=run_input)
# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
print(item)
# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start
CLI example
echo '{}' |
apify call neeenja/cheapest-llm-router --silent --output-dataset
MCP server setup
{
"mcpServers": {
"apify": {
"type": "http",
"url": "https://mcp.apify.com/?tools=fetch-actor-details,neeenja/cheapest-llm-router"
}
}
}
The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an Authorization: Bearer <APIFY_API_TOKEN> header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).
OpenAPI specification
Download the OpenAPI definition: https://api.apify.com/v2/actors/vYER8NboapafgX0U7/builds/SjbNgGetMTfOprnys/openapi.json