Multi-LLM-MCP

An MCP server that lets any IDE agent delegate coding tasks to any LLM: cloud APIs or local models, through a single unified interface.

Documentation

multi-llm-mcp

An MCP server that lets any IDE agent delegate coding tasks to any LLM — cloud APIs or local models — through a single unified interface.

Instead of being locked into one model, your coding agent can call NVIDIA NIM, OpenRouter, Groq, DeepSeek, or a local Ollama model for a second opinion, code review, or specialized task.


Features

  • 5 providers out of the box — NVIDIA NIM, OpenRouter, Groq, DeepSeek, Ollama
  • Local model support — Use Ollama for fully offline, private coding assistance
  • Fine-grained control — Set temperature, max_tokens, and system_prompt per call
  • Secure by design — API keys stay in environment variables, never in code
  • Connection pooling — Clients are cached for fast, efficient API calls
  • MCP standard — Works with any MCP-compatible IDE (Claude Desktop, VS Code, Cursor, Windsurf, etc.)

Supported Providers

ProviderTypeModels
OllamaLocalLlama 3, CodeLlama, Mistral, Gemma, etc.
NVIDIA NIMCloudLlama 3.1 405B, Mixtral, Code Llama, etc.
OpenRouterCloudClaude, GPT-4, Gemini, 200+ models
GroqCloudLlama 3, Mixtral, Gemma (ultra-fast inference)
DeepSeekCloudDeepSeek Coder, DeepSeek Chat

Quick Start

1. Clone and install

git clone https://github.com/arjunkr303/multi-llm-mcp.git cd multi-llm-mcp

python -m venv venv source venv/bin/activate # On Windows: venv\Scripts\activate

pip install -r requirements.txt

2. Configure API keys

cp .env.example .env

Edit .env and add your API keys. You only need keys for the providers you want to use. Ollama requires no API key.

NVIDIA_API_KEY=your_nvidia_key_here OPENROUTER_API_KEY=your_openrouter_key_here GROQ_API_KEY=your_groq_key_here DEEPSEEK_API_KEY=your_deepseek_key_here

3. Connect to your IDE

Add this to your MCP configuration:

Claude Desktop (claude_desktop_config.json):

{ "mcpServers": { "multi-llm-gateway": { "command": "/path/to/multi-llm-mcp/venv/bin/python", "args": ["/path/to/multi-llm-mcp/server.py"] } } }

VS Code / Cursor (.vscode/mcp.json or IDE MCP settings):

Replace /path/to/multi-llm-mcp with the actual path where you cloned the repo.

4. Use it

Once connected, your IDE agent has access to the ask_llm tool. You can invoke it by prompting your agent naturally. The agent will call the MCP server behind the scenes.

Tool: ask_llm

Delegates a coding question, code review, or rewrite task to an external LLM.

Parameters:

ParameterDefaultDescription
prompt(required)The question, code snippet, or task
system_prompt"You are a helpful coding assistant."Role context for the model
provider"ollama"Which provider to use
model"llama3"Model name for that provider
temperature0.00.0 = deterministic, 1.0 = creative
max_tokens4096Maximum response length

Example Prompts

Use these prompts in your IDE chat to trigger the MCP tool:

Rewrite code using a cloud model:

Use NVIDIA's Llama 3.1 to rewrite the database query in db.py to use async/await.

Get a code review from a different LLM:

Ask Groq's Llama 3 to review server.py for performance issues and suggest improvements.

Refactor with a local model (Ollama):

Use Ollama to refactor the authentication logic in auth.py into smaller, testable functions.

Generate tests using DeepSeek:

Ask DeepSeek Coder to write unit tests for the parse_config function in utils.py.

Get a second opinion on architecture:

Use OpenRouter's Claude to evaluate whether the current project structure follows
best practices and suggest a better layout.

Using with Ollama (Local Models)

For fully private, offline coding assistance:

Install Ollama: https://ollama.com/download

ollama pull llama3 ollama pull codellama

Then use provider: "ollama" with any pulled model name. No API key needed.

Security

  • API keys are loaded from environment variables only
  • .env is gitignored and never committed
  • No secrets are hardcoded in source code
  • All API communication happens server-side only

License

MIT — see LICENSE for details.