LLM Broker
一個相容於OpenAI的金鑰適用於所有模型:提供實測基準分數與即時價格,每個請求都會路由到符合您條件的最便宜方案。
託管 MCP 伺服器
npx add-mcp 'https://api.llm-broker.net/api/v1/broker/mcp'可安裝到 Claude Code、Codex、Cursor 等客戶端
文件
GLM 5.3 InferenceNet 0.079 / 4.40 Kimi K3 Wafer 0.90 / 15.00 Kimi K3 InferenceNet 0.80 / 13.50 ▼ GLM 5.3 Flash Wafer 0.100 / 0.50 Qwen3.8 27b ModelRun 0.70 / 4.70 ▼ DeepSeek V4 Pro Baidu 0.18 / 0.36 ▼ GLM 5.3 Wafer 0.051 / 6.00 ▲ GLM 5.3 Wafer 0.19 / 6.00 ▲ GLM 5.3 Flash OpenInference 0.032 / 0.083 ▼ GLM 5.3 Flash Inceptron 0.096 / 0.21 ▼ DeepSeek V4 Flash 0731 OpenInference 0.004 / 0.011 DeepSeek V4 Flash 0731 Relace 0.005 / 1.28 DeepSeek V4 Pro Relace 0.18 / 4.20 DeepSeek V4.1 Flash Wafer 0.070 / 1.20 DeepSeek V4.1 Flash Morph 0.072 / 0.60 Kimi K3 Morph 1.32 / 14.00 Kimi K3 Decart 2.55 / 12.75 ▲ DeepSeek V4 Flash 0731 Inceptron 0.006 / 0.013 GLM 5.2 Wafer 0.19 / 8.00 GLM 5.2 Wafer 0.060 / 6.00 ▼ DeepSeek V4 Pro 0813 Wafer 0.22 / 4.00 ▼ Kimi K3 Relace 2.00 / 14.00 ▲ GLM 5.3 Flash InferenceNet 0.070 / 0.50 DeepSeek V4.1 Flash InferenceNet 0.070 / 0.60 GLM 5.3 Flash Relace 0.040 / 0.50 DeepSeek V4.1 Flash Relace 0.016 / 0.60 ▼ GLM 5.3 Io Net 0.95 / 4.40 DeepSeek V4 Pro StreamLake 0.21 / 0.42 ▼ DeepSeek V4 Pro 0813 Ionstream 0.37 / 2.93 DeepSeek V4 Pro 0813 DeepSeek 0.66 / 1.98 ▼ DeepSeek V4.1 Flash Ionstream 0.100 / 1.10 ▲ DeepSeek V4.1 Flash Alibaba 0.15 / 0.60 ▼ DeepSeek V4.1 Flash Baidu 0.15 / 0.60 ▼ DeepSeek V4.1 Flash DeepSeek 0.15 / 0.60 ▼ DeepSeek V4.1 Flash DeepSeek 0.15 / 0.60 ▼ DeepSeek V4.1 Flash StreamLake 0.15 / 0.59 ▼ DeepSeek V4.1 Flash OpenInference 0.020 / 1.00 Mimo V2.6 Flash Io Net 0.29 / 0.45 DeepSeek V4.1 Flash AtlasCloud 0.30 / 1.20 ▲ DeepSeek V4.1 Flash DekaLLM 0.12 / 1.20 ▼ DeepSeek V4 Pro Reka 1.05 / 10.50 ▼ DeepSeek V4.1 Flash DigitalOcean 0.23 / 0.90 ▲ Claude Sonnet 5.5 Claude Platform on AWS 2.00 / 10.00 Claude Sonnet 5.5 Google 2.00 / 10.00 Claude Sonnet 5.5 Azure 2.20 / 11.00 Claude Sonnet 5.5 Amazon Bedrock 2.00 / 10.00 Claude Sonnet 5.5 Azure 2.00 / 10.00 Claude Sonnet 5.5 Google 2.20 / 11.00 Claude Sonnet 5.5 Google 2.20 / 11.00 Claude Sonnet 5.5 Anthropic 2.00 / 10.00 DeepSeek V4.1 Flash Novita 0.19 / 0.78 ▼ Mimo V2.6 Flash Darkbloom 0.100 / 0.28 DeepSeek V4.1 Flash Makora 0.27 / 1.15 ▲ Mimo V2.6 Pro GMICloud 0.44 / 0.87 ▲ Mimo V2.6 Flash GMICloud 0.14 / 0.28 ▲ DeepSeek V4.1 Flash Venice 0.30 / 1.20 ▼ DeepSeek V4.1 Flash Decart 0.090 / 0.18 GPT 6 Luna OpenAI 0.051 / 0.25 DeepSeek V4 Pro 0813 Baidu 0.14 / 0.43 Claude Haiku 5.5 Anthropic 0.100 / 0.50 DeepSeek V4 Flash Vision Exp DeepInfra 0.22 / 0.65 Mimo V2.6 Pro DeepInfra 0.43 / 0.87 Qwen3.8 27b Near AI 0.040 / 1.35 GLM 5.2 Decart 0.28 / 1.72 Gemini 3.7 Flash Google 0.38 / 1.87 GLM 5.3 Inceptron 0.60 / 2.20 Qwen3.7 Max Novita 1.25 / 3.75 Gemini 3.5 Flash Google 0.75 / 4.50 GPT 5.6 Sol OpenAI 1.00 / 5.00 GPT 6.1 Sol OpenAI 1.00 / 5.00 GPT 6 Sol OpenAI 1.00 / 5.00 Qwen3.8 2.4t A95b Novita 2.00 / 6.00 Grok 4.6 xAI 2.00 / 6.00 Grok 4.7 xAI 2.00 / 6.00 Claude Sonnet 5 Google 2.00 / 10.00 GLM 5.3 InferenceNet 0.079 / 4.40 Kimi K3 Wafer 0.90 / 15.00 Kimi K3 InferenceNet 0.80 / 13.50 ▼ GLM 5.3 Flash Wafer 0.100 / 0.50 Qwen3.8 27b ModelRun 0.70 / 4.70 ▼ DeepSeek V4 Pro Baidu 0.18 / 0.36 ▼ GLM 5.3 Wafer 0.051 / 6.00 ▲ GLM 5.3 Wafer 0.19 / 6.00 ▲ GLM 5.3 Flash OpenInference 0.032 / 0.083 ▼ GLM 5.3 Flash Inceptron 0.096 / 0.21 ▼ DeepSeek V4 Flash 0731 OpenInference 0.004 / 0.011 DeepSeek V4 Flash 0731 Relace 0.005 / 1.28 DeepSeek V4 Pro Relace 0.18 / 4.20 DeepSeek V4.1 Flash Wafer 0.070 / 1.20 DeepSeek V4.1 Flash Morph 0.072 / 0.60 Kimi K3 Morph 1.32 / 14.00 Kimi K3 Decart 2.55 / 12.75 ▲ DeepSeek V4 Flash 0731 Inceptron 0.006 / 0.013 GLM 5.2 Wafer 0.19 / 8.00 GLM 5.2 Wafer 0.060 / 6.00 ▼ DeepSeek V4 Pro 0813 Wafer 0.22 / 4.00 ▼ Kimi K3 Relace 2.00 / 14.00 ▲ GLM 5.3 Flash InferenceNet 0.070 / 0.50 DeepSeek V4.1 Flash InferenceNet 0.070 / 0.60 GLM 5.3 Flash Relace 0.040 / 0.50 DeepSeek V4.1 Flash Relace 0.016 / 0.60 ▼ GLM 5.3 Io Net 0.95 / 4.40 DeepSeek V4 Pro StreamLake 0.21 / 0.42 ▼ DeepSeek V4 Pro 0813 Ionstream 0.37 / 2.93 DeepSeek V4 Pro 0813 DeepSeek 0.66 / 1.98 ▼ DeepSeek V4.1 Flash Ionstream 0.100 / 1.10 ▲ DeepSeek V4.1 Flash Alibaba 0.15 / 0.60 ▼ DeepSeek V4.1 Flash Baidu 0.15 / 0.60 ▼ DeepSeek V4.1 Flash DeepSeek 0.15 / 0.60 ▼ DeepSeek V4.1 Flash DeepSeek 0.15 / 0.60 ▼ DeepSeek V4.1 Flash StreamLake 0.15 / 0.59 ▼ DeepSeek V4.1 Flash OpenInference 0.020 / 1.00 Mimo V2.6 Flash Io Net 0.29 / 0.45 DeepSeek V4.1 Flash AtlasCloud 0.30 / 1.20 ▲ DeepSeek V4.1 Flash DekaLLM 0.12 / 1.20 ▼ DeepSeek V4 Pro Reka 1.05 / 10.50 ▼ DeepSeek V4.1 Flash DigitalOcean 0.23 / 0.90 ▲ Claude Sonnet 5.5 Claude Platform on AWS 2.00 / 10.00 Claude Sonnet 5.5 Google 2.00 / 10.00 Claude Sonnet 5.5 Azure 2.20 / 11.00 Claude Sonnet 5.5 Amazon Bedrock 2.00 / 10.00 Claude Sonnet 5.5 Azure 2.00 / 10.00 Claude Sonnet 5.5 Google 2.20 / 11.00 Claude Sonnet 5.5 Google 2.20 / 11.00 Claude Sonnet 5.5 Anthropic 2.00 / 10.00 DeepSeek V4.1 Flash Novita 0.19 / 0.78 ▼ Mimo V2.6 Flash Darkbloom 0.100 / 0.28 DeepSeek V4.1 Flash Makora 0.27 / 1.15 ▲ Mimo V2.6 Pro GMICloud 0.44 / 0.87 ▲ Mimo V2.6 Flash GMICloud 0.14 / 0.28 ▲ DeepSeek V4.1 Flash Venice 0.30 / 1.20 ▼ DeepSeek V4.1 Flash Decart 0.090 / 0.18 GPT 6 Luna OpenAI 0.051 / 0.25 DeepSeek V4 Pro 0813 Baidu 0.14 / 0.43 Claude Haiku 5.5 Anthropic 0.100 / 0.50 DeepSeek V4 Flash Vision Exp DeepInfra 0.22 / 0.65 Mimo V2.6 Pro DeepInfra 0.43 / 0.87 Qwen3.8 27b Near AI 0.040 / 1.35 GLM 5.2 Decart 0.28 / 1.72 Gemini 3.7 Flash Google 0.38 / 1.87 GLM 5.3 Inceptron 0.60 / 2.20 Qwen3.7 Max Novita 1.25 / 3.75 Gemini 3.5 Flash Google 0.75 / 4.50 GPT 5.6 Sol OpenAI 1.00 / 5.00 GPT 6.1 Sol OpenAI 1.00 / 5.00 GPT 6 Sol OpenAI 1.00 / 5.00 Qwen3.8 2.4t A95b Novita 2.00 / 6.00 Grok 4.6 xAI 2.00 / 6.00 Grok 4.7 xAI 2.00 / 6.00 Claude Sonnet 5 Google 2.00 / 10.00
The best performance at the best price, automatically.
Authorization: Bearer ast_try · try without an account: model elliot, 20 requests a day for free
Tell us once what you need: score or model, model family, region, speed, capabilities, price. Your API key gets exactly that — today and with every new model generation, without changing a line. LLMs, images, speech and transcription with one key; prices compared live, performance measured by us, account and billing via API too. You pay the provider's price plus 5% — no subscription, no base fee.
33 models compared
56 platforms
0.011 USD per 1M output, from
You set the criteria. We keep them current.
Decide once what matters — your API key gets exactly that, re-decided on every request. All fields are part of the same OpenAI-compatible API.
Score or model
The best measured model per use case, a level (solid, strong, top), your own minimum score — say 0.80 in coding — or a fixed model name. Otherwise we take the cheapest one that meets it.
category + level | min_score · model
Model family
For example claude, qwen3 or deepseek/deepseek-v4. When a new generation ships, it is included as soon as it is on the market — with min_score, as soon as we have measured it.
family
Region
Switzerland, EU, USA or Asia: only providers whose execution region we have recorded (eu includes ch).
region
Speed
Rank by measured throughput when response time matters.
sort: "speed"
Capabilities
Tools, structured output, reasoning, image input, minimum context — every chosen provider must support it.
capabilities · min_context
Price
Cheapest first, best score per franc, or a hard maximum price per million tokens.
sort: "price" | "value" · max_price
Integrate once. Never chase models again.
If you build a platform or a tool, today you commit to a model — and tomorrow you switch when a better or cheaper one appears. Here you commit to your requirements, not to a model. The interface stays the same; we keep re-making the choice behind it.
- New model generations arrive without code changes.
- If a provider fails, the next one that meets your rules takes over.
- Market price cuts reach you automatically.
- If nothing meets your rules, we say so clearly (409) — no silent fallback.
{"model": "james-direct",
"routing": {"family": "deepseek", "category": "coding", "min_score": 0.8,
"sort": "value", "region": "eu", "capabilities": ["tools"]},
"messages": [{"role": "user", "content": "…"}]}
Market prices
Platform list prices in USD per 1 million tokens (daily rate, without our margin), checked every 15 minutes. Bars make price and speed comparable; index = Artificial Analysis via OpenRouter (pre-selection only — we commit only to what we measure ourselves). Region: as recorded by us — otherwise "global" (provider-dependent).
Performance map
Each dot is a model: cheap on the left, strong at the top. The glowing line is the efficiency frontier — no other model is both cheaper and better. Bubble size = context window.
↖ strong & cheap 0 20 40 60 0.01 0.1 1 10 100 USD per 1M output tokens (log scale) → Reasoning index ↑ GPT 6.1 Sol GPT 5.6 Sol Mimo V2.6 Pro GLM 5.3 GLM 5.3 Flash GLM 5.3 Flash DeepSeek V4.1 Flash DeepSeek V4 Flash 0731 DeepSeek V4 Flash 0731 DeepSeek V4 Flash 0731 Claude Opus 5.5 Claude Sonnet 5.5 Claude Fable 5.1 Claude Opus 5 Claude Haiku 5.5 Qwen3.8 27b
Efficiency frontier anthropicdeepseekopenaiqwenz-aigooglex-aixiaomi
Price range per platform
The same model costs different amounts on each platform. The glowing dot is the cheapest — that's where we route.
0.010.1110100
GPT 6 Astra
×2.2 · −55 %
GPT 6.1 Sol
×2.2 · −55 %
GPT 6 Sol
×2.2 · −55 %
GPT 5.6 Sol
×4.4 · −77 %
GLM 5.3
×2.7 · −63 %
Grok 4.6
×1.1 · −9 %
Kimi K3
×1.4 · −26 %
GLM 5.3 Flash
×19.3 · −95 %
Gemini 3.7 Flash
×2.0 · −50 %
DeepSeek V4.1 Flash
×6.6 · −85 %
Claude Sonnet 5
×1.5 · −33 %
GPT 6 Luna
×2.2 · −54 %
Price movement
Cheapest output price (USD per 1M tokens) across all platforms — candles per interval, red = more expensive, teal = cheaper.
Candle = 12 h · since 29.09.26 01:32 · 611 price changes
0.180 ▼ 69.9 %
0.0739 0.237 0.401 0.564 29.09.26 00:00 04.10.26 00:00 09.10.26 12:00 0.180
Value for money
Index points per USD of output price · price movement over the last 30 days.
1. D DeepSeek V4 Flash 0731 deepseek 3167.4 P/USD ↓ 95 %
2. Z GLM 5.3 Flash z-ai 503.5 P/USD ↓ 83 %
3. D DeepSeek V4.1 Flash deepseek 218.8 P/USD ↑ 2 %
4. O GPT 6 Luna openai 152.2 P/USD new
5. X Mimo V2.6 Flash xiaomi 135.2 P/USD ↑ 25 %
6. A Claude Haiku 5.5 anthropic 86.7 P/USD new
7. D DeepSeek V4 Pro 0813 deepseek 84.3 P/USD ↓ 84 %
8. D DeepSeek V4 Pro deepseek 83.9 P/USD ↓ 86 %
9. D DeepSeek V4 Flash Vision Exp deepseek 53.8 P/USD ↓ 51 %
10. X Mimo V2.6 Pro xiaomi 53.2 P/USD new
| Model | Reasoning | Code | Price in/out | Speed | Region | Trend | Platform | Context |
|---|---|---|---|---|---|---|---|---|
| Claude Opus 5.5 anthropic | 58 | HumanEval+ 93 % | 4.00 20.00 | 187 tok/s Uptime 100 % | global | Anthropic · 6 ▾ - Anthropic: 4.00 / 20.00 - Google: 4.00 / 20.00 - Amazon Bedrock: 4.00 / 20.00 - Azure: 4.00 / 20.00 - DeepInfra: 4.00 / 20.00 - Claude Platform on AWS: 4.00 / 20.00 | 1 M | |
| Claude Sonnet 5.5 anthropic | 56 | HumanEval+ 94 % | 2.00 10.00 | 113 tok/s Uptime 100 % | global | Google · 6 ▾ - Google: 2.00 / 10.00 - Amazon Bedrock: 2.00 / 10.00 - Anthropic: 2.00 / 10.00 - Azure: 2.00 / 10.00 - DeepInfra: 2.00 / 10.00 - Claude Platform on AWS: 2.00 / 10.00 | 1 M | |
| Claude Fable 5.1 anthropic | 53 | 82 | 10.00 50.00 | 72 tok/s Uptime 100 % | global | new | Amazon Bedrock · 4 ▾ - Amazon Bedrock: 10.00 / 50.00 - Google: 10.00 / 50.00 - Azure: 10.00 / 50.00 - Anthropic: 10.00 / 50.00 | 1 M |
| GPT 6 Astra openai | 53 | 77 | 5.00 25.00 | 81 tok/s Uptime 100 % | global | new | OpenAI · 3 ▾ - OpenAI: 5.00 / 25.00 - Azure: 10.00 / 50.00 - Amazon Bedrock: 11.00 / 55.00 | 1 M |
| GPT 6.1 Sol openai | 52 | HumanEval+ 95 % | 1.00 5.00 | 113 tok/s Uptime 100 % | global | OpenAI · 3 ▾ - OpenAI: 1.00 / 5.00 - Azure: 2.00 / 10.00 - Amazon Bedrock: 2.20 / 11.00 | 1 M | |
| Claude Opus 5 anthropic | 51 | 78 | 5.00 25.00 | 187 tok/s Uptime 100 % | global | new | Claude Platform on AWS · 6 ▾ - Claude Platform on AWS: 5.00 / 25.00 - Azure: 5.00 / 25.00 - Google: 5.00 / 25.00 - Anthropic: 5.00 / 25.00 - Amazon Bedrock: 5.00 / 25.00 - DeepInfra: 5.00 / 25.00 | 1 M |
| Claude Fable 5 anthropic | 50 | 77 | 10.00 50.00 | 60 tok/s Uptime 100 % | global | Claude Platform on AWS · 6 ▾ - Claude Platform on AWS: 10.00 / 50.00 - Azure: 10.00 / 50.00 - Google: 10.00 / 50.00 - Anthropic: 10.00 / 50.00 - Amazon Bedrock: 10.00 / 50.00 - DeepInfra: 10.00 / 50.00 | 1 M | |
| GPT 6 Sol openai | 48 | 1.00 5.00 | 92 tok/s Uptime 100 % | global | new | OpenAI · 3 ▾ - OpenAI: 1.00 / 5.00 - Azure: 2.00 / 10.00 - Amazon Bedrock: 2.20 / 11.00 | 1 M | |
| GPT 5.6 Sol openai | 47 | 77 | 1.00 5.00 | 84 tok/s Uptime 100 % | global | new | OpenAI · 3 ▾ - OpenAI: 1.00 / 5.00 - Azure: 4.00 / 20.00 - Amazon Bedrock: 4.40 / 22.00 | 1 M |
| Grok 4.7 x-ai | 46 | 2.00 6.00 | 78 tok/s Uptime 100 % | global | new | xAI · 1 ▾ - xAI: 2.00 / 6.00 | 500 k | |
| Mimo V2.6 Pro xiaomi | 46 | 0.43 0.87 | 40 tok/s Uptime 100 % | global | DeepInfra · 4 ▾ - DeepInfra: 0.43 / 0.87 - Novita: 0.44 / 0.87 - GMICloud: 0.44 / 0.87 - Xiaomi: 0.44 / 0.87 | 1 M | ||
| GLM 5.3 z-ai | 45 | 75 | 0.60 2.20 | 164 tok/s Uptime 100 % | global | Inceptron · 33 ▾ - Inceptron: 0.60 / 2.20 - Novita: 0.70 / 2.20 - DeepInfra: 0.56 / 2.50 - Phala: 0.84 / 2.64 - Decart: 0.84 / 2.65 - Sail Research: 0.20 / 3.40 - DigitalOcean: 0.91 / 2.86 - Morph: 0.21 / 3.74 - GMICloud: 0.98 / 3.08 - Reka: 0.14 / 4.20 - Relace: 0.031 / 4.40 - InferenceNet: 0.079 / 4.40 - Makora: 0.14 / 4.40 - AkashML: 0.19 / 4.40 - SiliconFlow: 1.12 / 3.52 - Alibaba: 1.19 / 3.74 - Friendli: 1.26 / 3.96 - Io Net: 0.95 / 4.40 - Mistral: 1.40 / 4.40 - Baidu: 1.40 / 4.40 - Crusoe: 1.40 / 4.40 - PrimeIntellect: 1.40 / 4.40 - Modal: 1.40 / 4.40 - Fireworks: 1.40 / 4.40 - Cloudflare: 1.40 / 4.40 - BaseTen: 1.40 / 4.40 - Nebius: 1.40 / 4.40 - AtlasCloud: 1.40 / 4.40 - Z.AI: 1.40 / 4.40 - Venice: 1.40 / 4.40 - Together: 1.40 / 4.40 - Parasail: 1.40 / 4.40 - Wafer: 0.051 / 6.00 | 1 M | |
| Grok 4.6 x-ai | 44 | 77 | 2.00 6.00 | 69 tok/s Uptime 100 % | global | new | xAI · 2 ▾ - xAI: 2.00 / 6.00 - Amazon Bedrock: 2.20 / 6.60 | 500 k |
| Kimi K3 moonshotai | 44 | 76 | 1.53 12.75 | 91 tok/s Uptime 100 % | global | Makora · 21 ▾ - Makora: 1.53 / 12.75 - InferenceNet: 0.80 / 13.50 - Sail Research: 0.84 / 13.50 - AkashML: 1.20 / 14.00 - Decart: 2.55 / 12.75 - Phala: 2.55 / 12.75 - Morph: 1.32 / 14.00 - DigitalOcean: 2.55 / 12.95 - Parasail: 2.60 / 13.00 - Wafer: 0.90 / 15.00 - Relace: 2.00 / 14.00 - Together: 2.70 / 13.50 - DeepInfra: 2.85 / 14.25 - Novita: 3.00 / 15.00 - Chutes: 3.00 / 15.00 - Fireworks: 3.00 / 15.00 - Modal: 3.00 / 15.00 - BaseTen: 3.00 / 15.00 - Moonshot AI: 3.00 / 15.00 - Amazon Bedrock: 3.00 / 15.00 - Alibaba: 3.45 / 17.25 | 1 M | |
| Claude Haiku 5.5 anthropic | 43 | 0.100 0.50 | 215 tok/s Uptime 100 % | global | new | Anthropic · 5 ▾ - Anthropic: 0.100 / 0.50 - Amazon Bedrock: 0.100 / 0.50 - Claude Platform on AWS: 0.100 / 0.50 - Google: 0.100 / 0.50 - Azure: 0.100 / 0.50 | 1 M | |
| GLM 5.3 Flash z-ai | 42 | 72 | 0.032 0.083 | 134 tok/s Uptime 100 % | global | OpenInference · 29 ▾ - OpenInference: 0.032 / 0.083 - Inceptron: 0.096 / 0.21 - DeepInfra: 0.075 / 0.25 - StreamLake: 0.084 / 0.28 - Novita: 0.084 / 0.28 - GMICloud: 0.090 / 0.30 - Decart: 0.093 / 0.31 - Near AI: 0.10 / 0.35 - Phala: 0.11 / 0.38 - Relace: 0.040 / 0.50 - InferenceNet: 0.070 / 0.50 - Wafer: 0.100 / 0.50 - Sail Research: 0.045 / 0.60 - Modal: 0.15 / 0.50 - Z.AI: 0.15 / 0.50 - Fireworks: 0.15 / 0.50 - AtlasCloud: 0.15 / 0.50 - BaseTen: 0.15 / 0.50 - Crusoe: 0.15 / 0.50 - CoreWeave: 0.15 / 0.50 - DigitalOcean: 0.15 / 0.50 - Friendli: 0.15 / 0.50 - SiliconFlow: 0.15 / 0.50 - Together: 0.15 / 0.50 - Parasail: 0.15 / 0.50 - Venice: 0.15 / 0.50 - Morph: 0.13 / 0.70 - DekaLLM: 0.100 / 1.00 - Reka: 0.060 / 1.60 | 1 M | |
| Claude Opus 4.8 anthropic | 42 | 74 | 5.00 25.00 | 158 tok/s Uptime 100 % | global | DeepInfra · 6 ▾ - DeepInfra: 5.00 / 25.00 - Azure: 5.00 / 25.00 - Claude Platform on AWS: 5.00 / 25.00 - Google: 5.00 / 25.00 - Anthropic: 5.00 / 25.00 - Amazon Bedrock: 5.00 / 25.00 | 1 M | |
| Claude Opus 4.7 anthropic | 41 | 74 | 5.00 25.00 | 81 tok/s Uptime 100 % | global | DeepInfra · 6 ▾ - DeepInfra: 5.00 / 25.00 - Claude Platform on AWS: 5.00 / 25.00 - Google: 5.00 / 25.00 - Amazon Bedrock: 5.00 / 25.00 - Anthropic: 5.00 / 25.00 - Azure: 5.00 / 25.00 | 1 M | |
| Qwen3.8 2.4t A95b qwen | 40 | 72 | 2.00 6.00 | 108 tok/s Uptime 100 % | global | Novita · 7 ▾ - Novita: 2.00 / 6.00 - Alibaba: 2.00 / 6.00 - SiliconFlow: 2.00 / 6.00 - Venice: 2.00 / 6.00 - Modal: 2.00 / 6.00 - DeepInfra: 2.00 / 6.00 - Together: 2.00 / 6.00 | 1 M | |
| Gemini 3.7 Flash google | 40 | 72 | 0.38 1.87 | 98 tok/s Uptime 100 % | global | Google · 3 ▾ - Google: 0.38 / 1.87 - Google AI Studio: 0.38 / 1.87 - DeepInfra: 0.75 / 3.75 | 1 M | |
| DeepSeek V4.1 Flash deepseek | 40 | HumanEval+ 89 % | 0.090 0.18 | 241 tok/s Uptime 100 % | global | Decart · 29 ▾ - Decart: 0.090 / 0.18 - Sail Research: 0.079 / 0.40 - DeepInfra: 0.14 / 0.42 - Relace: 0.016 / 0.60 - InferenceNet: 0.070 / 0.60 - Morph: 0.072 / 0.60 - StreamLake: 0.15 / 0.59 - Baidu: 0.15 / 0.60 - Alibaba: 0.15 / 0.60 - DeepSeek: 0.15 / 0.60 - CoreWeave: 0.20 / 0.65 - GMICloud: 0.18 / 0.72 - Novita: 0.19 / 0.78 - OpenInference: 0.020 / 1.00 - Phala: 0.21 / 0.84 - DigitalOcean: 0.23 / 0.90 - Ionstream: 0.100 / 1.10 - Wafer: 0.070 / 1.20 - DekaLLM: 0.12 / 1.20 - Makora: 0.27 / 1.15 - Crusoe: 0.29 / 1.20 - Modal: 0.30 / 1.20 - AtlasCloud: 0.30 / 1.20 - Venice: 0.30 / 1.20 - BaseTen: 0.30 / 1.20 - SiliconFlow: 0.30 / 1.20 - Parasail: 0.30 / 1.20 - Fireworks: 0.30 / 1.20 - Together: 0.30 / 1.20 | 1 M | |
| Claude Sonnet 5 anthropic | 38 | 72 | 2.00 10.00 | 94 tok/s Uptime 100 % | global | Google · 6 ▾ - Google: 2.00 / 10.00 - Claude Platform on AWS: 2.00 / 10.00 - Azure: 2.00 / 10.00 - Amazon Bedrock: 2.00 / 10.00 - Anthropic: 2.00 / 10.00 - DeepInfra: 3.00 / 15.00 | 1 M | |
| GPT 6 Luna openai | 38 | 0.051 0.25 | 97 tok/s Uptime 100 % | global | new | OpenAI · 3 ▾ - OpenAI: 0.051 / 0.25 - Azure: 0.100 / 0.50 - Amazon Bedrock: 0.11 / 0.55 | 1 M | |
| Mimo V2.6 Flash xiaomi | 38 | 0.100 0.28 | 149 tok/s Uptime 100 % | global | Darkbloom · 8 ▾ - Darkbloom: 0.100 / 0.28 - Makora: 0.13 / 0.28 - DeepInfra: 0.14 / 0.28 - GMICloud: 0.14 / 0.28 - Novita: 0.14 / 0.28 - Xiaomi: 0.14 / 0.28 - Venice: 0.17 / 0.35 - Io Net: 0.29 / 0.45 | 1 M | ||
| DeepSeek V4 Pro 0813 deepseek | 36 | 69 | 0.14 0.43 | 118 tok/s Uptime 100 % | global | Baidu · 20 ▾ - Baidu: 0.14 / 0.43 - Alibaba: 0.58 / 1.74 - StreamLake: 0.66 / 1.98 - DeepSeek: 0.66 / 1.98 - Ionstream: 0.37 / 2.93 - Phala: 0.96 / 2.88 - DeepInfra: 1.30 / 2.60 - Novita: 0.99 / 2.97 - Wafer: 0.22 / 4.00 - GMICloud: 1.06 / 3.17 - Sail Research: 0.40 / 4.30 - CoreWeave: 1.31 / 3.96 - Parasail: 1.32 / 3.96 - AtlasCloud: 1.32 / 3.96 - Together: 1.32 / 3.96 - DigitalOcean: 1.32 / 3.96 - SiliconFlow: 1.32 / 3.96 - Cloudflare: 1.32 / 3.96 - Venice: 1.65 / 4.95 - Relace: 0.13 / 8.00 | 1 M | |
| DeepSeek V4 Flash Vision Exp deepseek | 35 | 65 | 0.22 0.65 | 98 tok/s Uptime 100 % | global | DeepInfra · 3 ▾ - DeepInfra: 0.22 / 0.65 - GMICloud: 0.44 / 1.32 - SiliconFlow: 0.44 / 1.32 | 1 M | |
| DeepSeek V4 Flash 0731 deepseek | 34 | 69 | 0.004 0.011 | 152 tok/s Uptime 100 % | global | OpenInference · 24 ▾ - OpenInference: 0.004 / 0.011 - Inceptron: 0.006 / 0.013 - DeepInfra: 0.060 / 0.18 - StreamLake: 0.088 / 0.26 - DigitalOcean: 0.12 / 0.24 - Wafer: 0.13 / 0.23 - BaseTen: 0.13 / 0.26 - Venice: 0.13 / 0.26 - Sail Research: 0.100 / 0.30 - CoreWeave: 0.13 / 0.28 - Cohere: 0.14 / 0.28 - Together: 0.14 / 0.28 - Parasail: 0.14 / 0.28 - Reka: 0.020 / 0.53 - Alibaba: 0.18 / 0.53 - Mancer 2: 0.20 / 0.60 - SiliconFlow: 0.22 / 0.66 - GMICloud: 0.29 / 0.86 - Phala: 0.31 / 0.92 - Relace: 0.005 / 1.28 - Novita: 0.41 / 1.23 - Baidu: 0.44 / 1.32 - AtlasCloud: 0.44 / 1.32 - Cloudflare: 0.44 / 1.32 | 1 M | |
| Qwen3.8 27b qwen | 34 | 68 | 0.040 1.35 | 495 tok/s Uptime 100 % | global | Near AI · 19 ▾ - Near AI: 0.040 / 1.35 - Reka: 0.14 / 1.40 - DeepInfra: 0.15 / 1.87 - Phala: 0.15 / 1.87 - AkashML: 0.23 / 1.98 - Darkbloom: 0.051 / 2.20 - Wafer: 0.045 / 2.30 - Ionstream: 0.089 / 2.35 - Parasail: 0.24 / 2.20 - Chutes: 0.24 / 2.20 - Cerebras: 0.99 / 1.49 - Mancer 2: 0.20 / 2.50 - Alibaba: 0.42 / 2.55 - DekaLLM: 0.049 / 3.00 - CoreWeave: 0.40 / 3.00 - Novita: 0.42 / 3.00 - Cloudflare: 0.45 / 3.20 - Venice: 0.45 / 3.20 - ModelRun: 0.70 / 4.70 | 1 M | |
| GLM 5.2 z-ai | 34 | 69 | 0.28 1.72 | 274 tok/s Uptime 100 % | global | Decart · 26 ▾ - Decart: 0.28 / 1.72 - StreamLake: 0.56 / 1.75 - DeepInfra: 0.56 / 1.80 - Novita: 0.65 / 2.04 - DigitalOcean: 0.70 / 2.20 - CoreWeave: 0.76 / 2.42 - Inceptron: 1.25 / 2.06 - Baidu: 0.76 / 2.66 - AtlasCloud: 0.94 / 2.95 - Alibaba: 0.97 / 3.04 - Morph: 0.28 / 3.74 - Phala: 1.26 / 3.00 - Relace: 0.057 / 4.30 - InferenceNet: 0.18 / 4.40 - SiliconFlow: 1.19 / 3.74 - Cloudflare: 1.18 / 4.40 - Nebius: 1.40 / 4.40 - Mistral: 1.40 / 4.40 - BaseTen: 1.40 / 4.40 - Together: 1.40 / 4.40 - Z.AI: 1.40 / 4.40 - Friendli: 1.40 / 4.40 - Venice: 1.40 / 4.40 - GMICloud: 1.40 / 4.40 - Parasail: 1.40 / 4.40 - Wafer: 0.060 / 6.00 | 1 M | |
| Gemini 3.5 Flash google | 34 | 0.75 4.50 | 157 tok/s Uptime 100 % | global | Google · 3 ▾ - Google: 0.75 / 4.50 - Google AI Studio: 0.75 / 4.50 - DeepInfra: 1.50 / 9.00 | 1 M | ||
| DeepSeek V4 Pro deepseek | 30 | 59 | 0.18 0.36 | 108 tok/s Uptime 100 % | global | Baidu · 15 ▾ - Baidu: 0.18 / 0.36 - StreamLake: 0.21 / 0.42 - GMICloud: 0.96 / 1.91 - DigitalOcean: 1.04 / 2.09 - Cloudflare: 1.15 / 2.55 - DeepInfra: 1.30 / 2.60 - Parasail: 0.45 / 3.48 - Alibaba: 1.42 / 2.83 - Relace: 0.18 / 4.20 - SiliconFlow: 1.50 / 3.13 - Novita: 1.60 / 3.20 - Venice: 1.65 / 3.30 - AtlasCloud: 1.68 / 3.38 - Azure: 1.91 / 3.83 - Reka: 1.05 / 10.50 | 1 M | |
| Claude Sonnet 4.6 anthropic | 30 | 63 | 3.00 15.00 | 65 tok/s Uptime 100 % | global | Google · 6 ▾ - Google: 3.00 / 15.00 - Anthropic: 3.00 / 15.00 - Amazon Bedrock: 3.00 / 15.00 - Claude Platform on AWS: 3.00 / 15.00 - DeepInfra: 3.00 / 15.00 - Azure: 3.00 / 15.00 | 1 M | |
| Qwen3.7 Max qwen | 30 | 66 | 1.25 3.75 | 104 tok/s Uptime 100 % | global | Novita · 3 ▾ - Novita: 1.25 / 3.75 - Alibaba: 1.48 / 4.42 - DeepInfra: 2.50 / 7.50 | 1 M |
Image & speech — same key
Not just LLMs: image generation, speech and transcription from the market, with the same OpenAI-compatible API and the same billing. You name a market id, we pick the cheapest provider. Prices in USD, including our margin, live from /api/v1/broker/models?type=….

Image generation
From quick drafts to print quality: FLUX, Qwen-Image, Seedream and more — or iris, our own image model.
iris0.0078 USD per imagestabilityai/sdxl-turbo0.000211 USD per imageblack-forest-labs/FLUX-1-schnell0.000525 USD per imagePrunaAI/p-image0.0053 USD per imageblack-forest-labs/FLUX-1-dev0.0095 USD per imageblack-forest-labs/FLUX-2-dev0.0105 USD per image
curl https://llm-broker.net/api/v1/images/generations \
-H "Authorization: Bearer $BROKER_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"black-forest-labs/FLUX-1-schnell","prompt":"a lighthouse at dawn","size":"1024x1024"}'

Text to speech
Text to speech in many voices and languages — Kokoro, Qwen3-TTS, Chatterbox and more — or quill with cloned voices.
quill0.12 USD per 1,000 charactershexgrad/Kokoro-82M0.001 USD per 1,000 charactersResembleAI/chatterbox-multilingual0.0011 USD per 1,000 charactersResembleAI/chatterbox-turbo0.0011 USD per 1,000 charactersAudio8/Audio8-TTS-Preview-0.6b0.0053 USD per 1,000 characterssesame/csm-1b0.0074 USD per 1,000 characters
curl https://llm-broker.net/api/v1/audio/speech \
-H "Authorization: Bearer $BROKER_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"hexgrad/Kokoro-82M","input":"Hello from LLM Broker","voice":"af_bella"}' -o hello.mp3

Transcription
Speech to text: Whisper, Qwen3-ASR, Voxtral and more — billed by the duration the provider measures.
quill0.012 USD per minutenvidia/Nemotron-3.5-ASR-Streaming-Multilingual-0.6b0.000209 USD per minuteopenai/whisper-large-v3-turbo0.000209 USD per minuteQwen/Qwen3-ASR-0.6B0.000209 USD per minuteopenai/whisper-large-v30.000473 USD per minuteQwen/Qwen3-ASR-1.7B0.000473 USD per minute
curl https://llm-broker.net/api/v1/audio/transcriptions \
-H "Authorization: Bearer $BROKER_API_KEY" \
-F model=openai/whisper-large-v3-turbo -F [email protected]
One prompt — your AI assistant sets everything up
Paste the prompt into Claude Code, Cursor, Codex, OpenCode or any other coding assistant. It reads our docs, opens the account via API, stores the key safely, wires up the API and checks it with a smoke test. You only confirm the payment.

- llms.txt — entry point for agents https://llm-broker.net/llms.txt
- OpenAPI (machine-readable) https://llm-broker.net/api/openapi
- API docs in the browser https://llm-broker.net/api/docs
- Wiki: API https://astrioncore.wiki/api
- Wiki: account & billing via API https://astrioncore.wiki/api#konto
- Market list (no key needed) https://llm-broker.net/api/v1/broker/models
Claude Code: paste the prompt into the session · Cursor/Windsurf: into the agent chat · OpenCode: opencode run "…" · Codex: codex "…"
Set up LLM Broker in this project (https://llm-broker.net) — one OpenAI-compatible API for LLMs, images, speech
and transcription with a single key. Work on your own; only ask me when a payment must be confirmed or a secret
must be entered.
1. Read the descriptions first, don't guess:
- https://llm-broker.net/llms.txt (entry point for agents)
- https://llm-broker.net/api/openapi (all endpoints and schemas)
- https://astrioncore.wiki/api and https://astrioncore.wiki/api#konto (account and billing via API)
2. Key: if BROKER_API_KEY (starts with ast_sk_) is not set, open an account via the API:
POST https://llm-broker.net/api/v1/broker/accounts
{"billing": "topup", "amount": 25, "currency": "USD"} (every request picks its level via "routing")
Show me the checkout_url (the card is entered once at Stripe), then poll status_url until it returns 200.
The api_key is shown EXACTLY ONCE: store it right away in .env as BROKER_API_KEY, add .env to .gitignore,
never log or commit it. base_url is in the same response.
3. Integrate: the OpenAI SDK of the project's language with base_url and BROKER_API_KEY. Models are listed at
GET /api/v1/broker/models?type=chat|image|speech|transcription (no key needed).
- LLM: POST /api/v1/chat/completions, model "james-direct" with "routing" (category, level|min_score, sort,
region, family, capabilities, max_price) or a market id such as "deepseek/deepseek-v4.1-flash".
- Images: POST /api/v1/images/generations, speech: POST /api/v1/audio/speech,
transcription: POST /api/v1/audio/transcriptions — each with a market id of the matching type.
4. Add a small client layer (configuration from the environment, no hard-coded model names — prefer routing),
write a smoke test with one short chat call and run it.
5. Billing: GET /api/v1/broker/billing shows balance and spend; top up or change the monthly limit with
PATCH /api/v1/broker/billing, the card with POST|DELETE /api/v1/broker/card.
Errors: 402 insufficient_balance (top up), 409 no_matching_offer (relax routing),
404 model_not_found (check the market list), 429 rate_limited (retry later).
Finish with a summary of what is set up and where the key is stored.
Richte in diesem Projekt den LLM Broker ein (https://llm-broker.net) — eine OpenAI-kompatible API für LLM, Bilder,
Sprachausgabe und Transkription mit einem Schlüssel. Arbeite selbstständig; frag mich nur, wenn eine Zahlung
bestätigt oder ein Geheimnis eingegeben werden muss.
1. Lies zuerst die Beschreibungen, nicht raten:
- https://llm-broker.net/llms.txt (Einstieg für Agenten)
- https://llm-broker.net/api/openapi (alle Endpunkte und Schemata)
- https://astrioncore.wiki/api und https://astrioncore.wiki/api#konto (Konto und Bezahlung per API)
2. Schlüssel: Ist BROKER_API_KEY (beginnt mit ast_sk_) nicht gesetzt, eröffne ein Konto per API:
POST https://llm-broker.net/api/v1/broker/accounts
{"billing": "topup", "amount": 25, "currency": "USD"} (Leistung wählt jede Anfrage per "routing")
Zeig mir die checkout_url (die Karte wird einmal bei Stripe eingegeben) und frag danach status_url ab,
bis 200 kommt. Der api_key erscheint GENAU EINMAL: sofort in .env als BROKER_API_KEY speichern,
.env in .gitignore, nie ins Log oder in einen Commit. base_url steht in derselben Antwort.
3. Einbinden: OpenAI-SDK der Projektsprache mit base_url und BROKER_API_KEY. Modelle findest du unter
GET /api/v1/broker/models?type=chat|image|speech|transcription (ohne Schlüssel lesbar).
- LLM: POST /api/v1/chat/completions, model "james-direct" mit "routing" (category, level|min_score, sort,
region, family, capabilities, max_price) oder direkt eine Markt-ID wie "deepseek/deepseek-v4.1-flash".
- Bilder: POST /api/v1/images/generations, Sprache: POST /api/v1/audio/speech,
Transkription: POST /api/v1/audio/transcriptions — jeweils mit einer Markt-ID der passenden type.
4. Lege eine kleine Client-Schicht an (Konfiguration aus der Umgebung, keine festen Modellnamen im Code —
lieber routing), schreibe einen Rauchtest mit einem kurzen Chat-Aufruf und führe ihn aus.
5. Abrechnung: GET /api/v1/broker/billing zeigt Guthaben und Verbrauch; Aufladen oder Monatsrahmen ändern
mit PATCH /api/v1/broker/billing, Karte mit POST|DELETE /api/v1/broker/card.
Fehler: 402 insufficient_balance (aufladen), 409 no_matching_offer (routing lockern),
404 model_not_found (Marktliste prüfen), 429 rate_limited (später erneut).
Fasse am Ende zusammen, was eingerichtet ist und wo der Schlüssel liegt.
Performance per request
On every request you decide how strong the model has to be — via routing in the same OpenAI-compatible API. A level is a minimum score in OUR measurements per use case, best the highest-measured model; set your own threshold with "min_score": 0.85. The price is today's customer price (in the selected currency per 1M tokens in/out, including our margin) of the cheapest model that meets the level — works for taylor, james-direct and every market id.
Coding measured with HumanEval+ "category": "coding"
Solid ≥ 70 %
0.04 / 0.19
22 models qualify
"level": "solid"
Strong ≥ 80 %
0.04 / 0.19
20 models qualify
"level": "strong"
Top ≥ 88 %
0.06 / 0.19
17 models qualify
"level": "top"
Best top score
1.11 / 5.54
1 model qualifies
"level": "best"
Reasoning & knowledge measured with MMLU-Pro + GPQA Diamond "category": "reasoning"
Solid ≥ 55 %
0.04 / 0.19
17 models qualify
"level": "solid"
Strong ≥ 65 %
0.06 / 0.19
16 models qualify
"level": "strong"
Top ≥ 75 %
0.06 / 0.19
12 models qualify
"level": "top"
Best top score
1.11 / 5.54
1 model qualifies
"level": "best"
Math measured with MATH "category": "math"
Solid ≥ 60 %
0.04 / 0.19
18 models qualify
"level": "solid"
Strong ≥ 75 %
0.04 / 0.19
18 models qualify
"level": "strong"
Top ≥ 85 %
0.06 / 0.19
13 models qualify
"level": "top"
Best top score
4.20 / 21.00
1 model qualifies
"level": "best"
Instruction following measured with IFEval "category": "instruction_following"
Solid ≥ 75 %
0.04 / 0.19
20 models qualify
"level": "solid"
Strong ≥ 82 %
0.06 / 0.19
16 models qualify
"level": "strong"
Top ≥ 88 %
0.16 / 0.63
10 models qualify
"level": "top"
Best top score
2.22 / 6.65
1 model qualifies
"level": "best"
Get your key
Add a card once and you're set. Only Stripe sees the card number — we never do.
Sign up and pay — all via API
A program or agent opens the account, fetches the key and manages billing on its own. Money before key: no free credit, the key is created only after payment.
1 · ACCOUNT
Open an account
Choose billing — a level is optional, every request picks it via routing. The response contains checkout_url and status_url.
curl -X POST https://llm-broker.net/api/v1/broker/accounts \
-H "Content-Type: application/json" \
-d '{"billing":"topup",
"amount":25}'
2 · CARD
Pay once
The card is entered once at Stripe — by you or by an agent with a browser. We never see the card number.
3 · KEY
Collect the key
202 until payment, then api_key (ast_sk_…) and base_url — exactly once.
curl https://llm-broker.net/api/v1/broker/accounts/$SESSION_ID
4 · BILLING
Self-service
Read balance and usage, change top-up or monthly limit, replace or remove the card.
GET|PATCH /api/v1/broker/billing
POST|DELETE /api/v1/broker/card
After that everything runs without humans: top-ups with the stored card, monthly billing, cancellation. Only the very first card entry still needs a browser at Stripe today. Details: Wiki: account & billing via API.
Which model do I get with taylor?
Without a level, the most capable one; with a level (per request or as the key's default), the cheapest of all models that reach it in our measurements. The table above shows the market; which model answers a single request is decided by the price at that moment. Want the strongest instead of the cheapest: "routing": {"category": "…", "level": "best"} — works for taylor, james-direct and every market id.
What does "measured" mean?
We test every model ourselves on open benchmarks (samples, identical settings) and re-measure regularly — all without thinking mode so the numbers are comparable; reasoning models often do better on hard tasks with thinking switched on. You control that per request with reasoning_effort (e.g. "none"). Third-party indices are used for pre-selection only.
And if no model meets the level?
Then the API answers with level_unavailable and charges nothing — never a silent fallback to a weaker model.
Are there image and speech models too?
Yes: image generation, speech and transcription from the market with the same key — market ids under /api/v1/broker/models?type=image|speech|transcription, plus our own models iris and quill.
How does billing work?
Pay as you go at the daily price — tokens, images, characters or minutes — by credit card only: automatic top-up or monthly in arrears. No subscription, no free credit.