llm-advisor-mcp

Comparación en tiempo real de modelos LLM/VLM con benchmarks, precios y recomendaciones personalizadas de 5 fuentes de datos. No se requiere clave API.

Documentación

llm-advisor-mcp

npm version npm downloads CI License: MIT Node.js >= 18 TypeScript

Glama MCP server

Inglés | 日本語

Dale a tu asistente de IA conocimiento en tiempo real sobre LLM/VLM. Precios, benchmarks y recomendaciones — actualizados cada hora, no en cada ciclo de entrenamiento.

Los LLM tienen límites de conocimiento. Pregúntale a Claude "¿cuál es el mejor modelo de codificación ahora mismo?" y no podrá responder con datos actuales. Este servidor MCP lo soluciona alimentando inteligencia de modelos en vivo directamente en la ventana de contexto de tu asistente de IA.

  • Cero configuración — Sin claves API, sin registro. Un solo comando para instalar.
  • Bajo consumo de tokens — Tablas Markdown compactas (~300 tokens), no JSON crudo (~3,000 tokens). Tu ventana de contexto importa.
  • 5 fuentes de benchmarks — SWE-bench, LM Arena Elo, OpenCompass VLM, Aider Polyglot y precios de OpenRouter fusionados en una vista unificada.

Casos de uso

  • "¿Cuál es el mejor modelo de codificación ahora mismo?" — list_top_models con categoría coding
  • "Compara Claude vs GPT vs Gemini" — compare_models con tabla comparativa lado a lado
  • "Encuentra un modelo barato con contexto de 1M" — recommend_model con restricciones de presupuesto
  • "¿Qué benchmarks tiene el modelo X?" — get_model_info con rangos percentiles

Inicio rápido

Claude Code

claude mcp add llm-advisor -- npx -y llm-advisor-mcp

Claude Code (Windows)

claude mcp add llm-advisor -- cmd /c npx -y llm-advisor-mcp

Claude Desktop / Cursor / Windsurf

Añade a tu archivo de configuración MCP:

{
  "mcpServers": {
    "llm-advisor": {
      "command": "npx",
      "args": ["-y", "llm-advisor-mcp"]
    }
  }
}

Eso es todo. Sin claves API, sin archivos .env.

Clientes compatibles

ClienteCompatibleMétodo de instalación
Claude CodeSíclaude mcp add
Claude DesktopSíConfiguración JSON
CursorSíConfiguración JSON
WindsurfSíConfiguración JSON
Cualquier cliente MCPSíTransporte stdio

Herramientas

get_model_info

Especificaciones detalladas para un modelo específico: precios, benchmarks, rangos percentiles, capacidades y un ejemplo de código API listo para usar.

Parámetros

NombreTipoRequeridoPredeterminadoDescripción
modelstringSí—ID del modelo o nombre parcial (p. ej. "claude-sonnet-4", "gpt-5")
include_api_examplebooleanNotrueIncluir un fragmento de código listo para usar
api_formatenumNoopenai_sdkopenai_sdk, curl o python_requests

Ejemplo de salida

## anthropic/claude-sonnet-4

**Provider**: anthropic | **Modality**: text+image→text | **Released**: 2025-06-25

### Pricing
| Metric | Value |
|--------|-------|
| Input | $3.00 /1M tok |
| Output | $15.00 /1M tok |
| Cache Read | $0.30 /1M tok |
| Context | 200K |
| Max Output | 64K |

### Benchmarks
| Benchmark | Score |
|-----------|-------|
| SWE-bench Verified | 76.8% |
| Aider Polyglot | 72.1% |
| Arena Elo | 1467 |
| MMMU | 76.0% |

### Percentile Ranks
| Category | Percentile |
|----------|------------|
| Coding | P96 |
| General | P95 |
| Vision | P90 |

**Capabilities**: Tools, Reasoning, Vision

### API Example (openai_sdk)
```python
from openai import OpenAI
client = OpenAI(
    base_url="https://openrouter.ai/api/v1",
    api_key="<OPENROUTER_API_KEY>",
)
response = client.chat.completions.create(
    model="anthropic/claude-sonnet-4",
    messages=[{"role": "user", "content": "Hello"}],
)

---

### `list_top_models`

Top-ranked models for a category. Includes release dates for freshness awareness.

**Parameters**

| Name | Type | Required | Default | Description |
|------|------|----------|---------|-------------|
| `category` | enum | Yes | — | `coding`, `math`, `vision`, `general`, `cost-effective`, `open-source`, `speed`, `context-window`, `reasoning` |
| `limit` | number | No | `10` | Number of results (1-20) |
| `min_context` | number | No | — | Minimum context window in tokens |
| `min_release_date` | string | No | — | `YYYY-MM-DD`. Excludes models released before this date |

**Example output**

Top 5: codificación

#ModeloPuntuación claveEntrada $/1MSalida $/1MContextoLanzamiento
1openai/o3-proSWE 79.5%$20.00$80.00200K2025-06-10
2anthropic/claude-sonnet-4SWE 76.8%$3.00$15.00200K2025-06-25
3google/gemini-2.5-proSWE 75.2%$1.25$10.001M2025-03-25
4openai/o4-miniSWE 73.6%$1.10$4.40200K2025-04-16
5anthropic/claude-opus-4SWE 72.5%$15.00$75.00200K2025-05-22

---

### `compare_models`

Side-by-side comparison for 2-5 models. Best values are **bolded** automatically. Includes a `Released` row so you can spot outdated models at a glance.

**Parameters**

| Name | Type | Required | Default | Description |
|------|------|----------|---------|-------------|
| `models` | string[] | Yes | — | 2-5 model IDs or partial names |

**Example output**

Comparación de modelos (3 modelos)

anthropic/claude-sonnet-4openai/gpt-4.1google/gemini-2.5-pro
Entrada $/1M$3.00$2.00$1.25
Salida $/1M$15.00$8.00$5.00
Contexto200K1M1M
Salida máxima64K32K65K
SWE-bench76.8%55.0%75.2%
Aider Polyglot72.1%65.3%71.8%
Arena Elo146714921445
VisiónSíSíSí
HerramientasSíSíSí
RazonamientoSíNoSí
Código abiertoNoNoNo
Lanzamiento2025-06-252025-04-142025-03-25

---

### `recommend_model`

Personalized top-3 recommendations. Scores combine weighted benchmarks, pricing, capability bonuses, and a freshness bonus (+3 points for models released within 3 months, +1 within 6 months).

**Parameters**

| Name | Type | Required | Default | Description |
|------|------|----------|---------|-------------|
| `use_case` | enum | Yes | — | `coding`, `math`, `general`, `vision`, `creative`, `reasoning`, `cost-effective` |
| `max_input_price` | number | No | — | Max input price (USD/1M tokens) |
| `max_output_price` | number | No | — | Max output price (USD/1M tokens) |
| `min_context` | number | No | — | Minimum context window in tokens |
| `require_vision` | boolean | No | — | Require image input support |
| `require_tools` | boolean | No | — | Require tool/function calling support |
| `require_open_source` | boolean | No | — | Require open-source license |
| `min_release_date` | string | No | — | `YYYY-MM-DD`. Excludes older models |

**Example output**

Recomendado para: codificación

1. anthropic/claude-sonnet-4 (puntuación: 78)

Entrada: $3.00/1M | Salida: $15.00/1M | Contexto: 200K | Lanzamiento: 2025-06-25 Benchmarks: SWE-bench: 76.8%, Aider: 72.1%, Arena: 1467 Fortalezas: razonamiento, herramientas, visión

2. google/gemini-2.5-flash (puntuación: 74)

Entrada: $0.15/1M | Salida: $0.60/1M | Contexto: 1M | Lanzamiento: 2025-05-20 Benchmarks: SWE-bench: 62.9%, Arena: 1445 Fortalezas: herramientas, visión, contexto de 1M+

3. openai/o4-mini (puntuación: 71)

Entrada: $1.10/1M | Salida: $4.40/1M | Contexto: 200K | Lanzamiento: 2025-04-16 Benchmarks: SWE-bench: 73.6%, Arena: 1430 Fortalezas: razonamiento, herramientas


---

## Data Sources

All data is fetched in real time from free, public APIs. No authentication required.

| Source | Data | Models | Cache TTL |
|--------|------|--------|-----------|
| [OpenRouter](https://openrouter.ai/api/v1/models) | Pricing, context lengths, modalities, release dates | 300+ | 1 hour |
| [SWE-bench](https://github.com/SWE-bench/swe-bench.github.io) | Coding benchmark (Verified leaderboard) | 30+ | 6 hours |
| [LM Arena](https://lmarena.ai) | Human preference Elo ratings | 314+ | 6 hours |
| [OpenCompass VLM](https://opencompass.org.cn) | Vision benchmarks: MMMU, MMBench, OCRBench, AI2D, MathVista | 284+ | 6 hours |
| [Aider Polyglot](https://aider.chat/docs/leaderboards/) | Multi-language coding pass rate | 63+ | 6 hours |

---

## Context Cost

MCP tool definitions and responses consume your LLM's context window. This server is designed to be lean:

| Component | Tokens |
|-----------|--------|
| All 4 tool definitions | ~1,000 |
| Typical tool response | ~250-400 |

For comparison, most MCP servers that return raw JSON consume 3,000-10,000 tokens per response. Every response from llm-advisor-mcp is pre-formatted Markdown, keeping context costs roughly 10x lower.

---

## Architecture

┌──────────────────────────────────────────────┐ │ Cliente MCP (Claude, etc.) │ └──────────┬───────────────────────────────────┘ │ stdio (JSON-RPC) ┌──────────▼───────────────────────────────────┐ │ servidor llm-advisor-mcp │ │ │ │ ┌─────────┐ ┌───────────┐ ┌────────────┐ │ │ │ Herramientas │ Registro │ │ Caché │ │ │ │ (4 herramientas)│──│ (unificado) │──│ (en memoria)│ │ │ └─────────┘ └───────────┘ └────────────┘ │ │ │ │ │ ┌────────────┼────────────┐ │ │ ▼ ▼ ▼ │ │ ┌──────────┐ ┌──────────┐ ┌──────────┐ │ │ │Normalizador│ │Percentil│ │ Obtenedores │ │ │ │(mapa slug)│ │ (5 cats) │ │(5 fuentes│ │ │ └──────────┘ └──────────┘ └──────────┘ │ └──────────────────────────────────────────────┘ │ │ │ OpenRouter SWE-bench Arena / VLM / Aider


- **TypeScript + ESM** — Single entry point, `tsup` build
- **In-memory cache** — TTL-based (1h pricing, 6h benchmarks), stale-while-revalidate
- **Cross-source normalization** — Maps inconsistent model names (e.g. `Claude 3.5 Sonnet` vs `anthropic/claude-3.5-sonnet`) to canonical IDs
- **Percentile computation** — Ranks across 5 categories (coding, math, general, vision, cost efficiency)
- **Freshness scoring** — Recommendation algorithm gives a bonus to recently released models (+3 for <=3mo, +1 for <=6mo)
- **Zero runtime deps** beyond `@modelcontextprotocol/sdk` and `zod`

---

## Roadmap

| Version | Status | Highlights |
|---------|--------|------------|
| v0.1 | Done | `get_model_info` + `list_top_models` via OpenRouter |
| v0.2 | Done | `compare_models` + `recommend_model` + SWE-bench + Arena Elo |
| v0.3 | Done | VLM benchmarks (MMMU, MMBench, OCRBench, AI2D, MathVista) + Aider Polyglot + percentile ranks + 43 tests |
| v0.4 | **Current** | Release date display, date-based filtering, freshness scoring in recommendations + 51 tests |
| v1.0 | Planned | Community contributions, weekly static data snapshots via GitHub Actions |

---

## Development

```bash
git clone https://github.com/Daichi-Kudo/llm-advisor-mcp.git
cd llm-advisor-mcp
npm install
npm run build       # Build with tsup
npm run dev         # Run with tsx (hot reload)
npm test            # Run 51 unit tests (vitest)
npm run test:watch  # Watch mode

Estructura del proyecto

src/
  index.ts              # Server entry point
  types.ts              # Shared type definitions
  tools/
    model-info.ts       # get_model_info tool
    list-top.ts         # list_top_models tool
    compare.ts          # compare_models tool
    recommend.ts        # recommend_model tool
    formatters.ts       # Markdown output formatters
  data/
    registry.ts         # Unified model registry
    cache.ts            # In-memory TTL cache
    normalizer.ts       # Cross-source name normalization
    percentiles.ts      # Percentile rank computation
    fetchers/
      openrouter.ts     # OpenRouter API
      swe-bench.ts      # SWE-bench leaderboard
      arena.ts          # LM Arena Elo ratings
      vlm-leaderboard.ts # OpenCompass VLM benchmarks
      aider.ts          # Aider Polyglot scores
    static/
      api-examples.ts   # API code snippet templates

Contribuciones

  1. Haz un fork del repositorio
  2. Crea una rama de funcionalidad
  3. Añade pruebas para la nueva funcionalidad
  4. Ejecuta npm test para verificar que las 51 pruebas pasan
  5. Envía una solicitud de extracción (pull request)

Licencia

MIT — creado por Cognisant LLC