vibops-mcp

GPU infrastructure control plane + per-agent LLM FinOps. 74 tools. Deploy, scale, track cost per agent, enforce budgets. MIT license.

Documentation

vibops-mcp

License: MIT Python 3.11+ MCP Tools Tests

The MCP server for VibOps — The AI Infrastructure Engine. From code to GPU in one conversation.

The problem

Getting an AI app from code to production on GPUs requires stitching together 9+ tools — git, Docker, CI/CD, Helm, kubectl, GPU monitoring, cost management, compliance, alerting. Each with its own API, dashboard, and cost model. No single interface spans the full pipeline.

The solution

vibops-mcp connects your AI assistant to VibOps — the engine that clones, builds, deploys, scales, monitors, fixes, and bills your apps and agents on any GPU, any cluster, any cloud. One pip install, 117 tools, one conversation.

  • Ship — clone repos, build containers, deploy models, run Helm/kubectl, trigger pipelines, submit Slurm jobs
  • Operate — scale deployments, manage VMs (Proxmox/XO/vSphere/HPE VME), detect and remediate GPU anomalies
  • Observe — GPU utilisation, workload breakdown, MTTR, cost estimates, live K8s deployments
  • Govern — AI Act compliance, SOC 2/RGPD reports, immutable audit chain, policy management
  • FinOps — per-agent cost tracking, budget enforcement, chargeback, spend trends, waste analysis

Every operation goes through your VibOps instance and is recorded in the immutable audit log.

Installation

pip install git+https://github.com/VibOpsai/vibops-mcp.git

Configuration

You need two environment variables:

VariableDescription
VIBOPS_URLBase URL of your VibOps instance, e.g. https://vibops.example.com
VIBOPS_TOKENAPI token — create one in VibOps → Settings → API Tokens

Claude Desktop

Add to ~/.config/claude/claude_desktop_config.json (macOS: ~/Library/Application Support/Claude/claude_desktop_config.json):

{
  "mcpServers": {
    "vibops": {
      "command": "vibops-mcp",
      "env": {
        "VIBOPS_URL": "https://vibops.example.com",
        "VIBOPS_TOKEN": "your-token-here"
      }
    }
  }
}

Cursor

Add to .cursor/mcp.json in your project root, or to the global config:

{
  "mcpServers": {
    "vibops": {
      "command": "vibops-mcp",
      "env": {
        "VIBOPS_URL": "https://vibops.example.com",
        "VIBOPS_TOKEN": "your-token-here"
      }
    }
  }
}

Claude Code (CLI)

claude mcp add vibops vibops-mcp \
  -e VIBOPS_URL=https://vibops.example.com \
  -e VIBOPS_TOKEN=your-token-here

Available tools

Observation (16 tools — read-only)

ToolDescription
list_clustersList clusters and GPU utilisation
list_kubectl_contextsList available kubectl contexts
get_cluster_deploymentsLive K8s deployment status for a cluster
get_cluster_rateGet configured GPU cost rate for a cluster
list_jobsList recent jobs with optional filters
get_jobGet job details and result
get_job_metricsJob success rate, latency P50/P95/P99, error breakdown
get_gpu_metricsHourly GPU utilisation time-series
get_workload_breakdownJob count by workload type
get_mttrMean Time To Resolve GPU alerts
get_cost_estimateEstimated GPU spend
list_gatewaysList registered gateways and status
list_alertsList GPU alerts (open or resolved)
list_secretsList secrets (names only, never values)
list_providersList configured AI/GPU cloud providers
list_pipelinesList automation pipelines

Actions (18 tools — write)

ToolDescription
scale_deploymentScale a K8s deployment replica count
deploy_modelDeploy an AI model onto a GPU cluster
helm_upgradeRun helm upgrade --install (destructive — requires confirmed=True)
helm_uninstallUninstall a Helm release (destructive — requires confirmed=True)
run_kubectlRun an arbitrary kubectl command
git_cloneClone a git repository
create_secretStore an encrypted secret
trigger_pipelineManually trigger an automation pipeline
slurm_get_cluster_infoGet Slurm cluster info and partition details
slurm_list_jobsList Slurm jobs with optional filters
slurm_get_job_statusGet status of a specific Slurm job
slurm_get_job_outputRetrieve stdout/stderr of a completed Slurm job
slurm_submit_jobSubmit a new Slurm job
slurm_cancel_jobCancel a running or pending Slurm job
registry_list_reposList container registry repositories
registry_list_tagsList tags for a container image
registry_check_imageCheck image details (size, layers, created date)
registry_delete_tagDelete a stale image tag (requires confirmed=True)

HPE VME / Morpheus (13 tools)

ToolDescription
vme_list_instancesList all VMs managed by HPE VME
vme_get_instanceGet detailed VM status and configuration
vme_list_serversList physical hosts with CPU/memory
vme_list_cloudsList configured clouds/zones (KVM, VMware)
vme_start_instanceStart a stopped instance
vme_stop_instanceStop a running instance
vme_restart_instanceRestart an instance
vme_create_snapshotCreate a snapshot (pre-migration)
vme_list_snapshotsList snapshots for an instance
vme_convert_imageConvert disk format (vmdk → qcow2) — VMware migration
vme_list_virtual_imagesList available images and templates
vme_detect_vm_wasteDetect stopped/idle VMs with recommendations
vme_get_activityRecent audit log from Morpheus

Hypervisors — Proxmox, vSphere, Xen Orchestra (30 tools)

Documented last and listed here in full: these were registered and served but absent from this README, so a reader counting the sections found 87 of the 117 the first line promises. They now add up.

Proxmox VE (5 tools)

ToolDescription
proxmox_list_vmsList VMs across the cluster with status and sizing
proxmox_start_vmStart a stopped VM
proxmox_stop_vmStop a running VM (destructive — requires confirmed=True)
proxmox_migrate_vmMigrate a VM to another node (destructive — requires confirmed=True)
proxmox_create_snapshotSnapshot a VM before a risky change

VMware vSphere (9 tools)

ToolDescription
vsphere_list_vmsList VMs with power state, host and resources
vsphere_get_vmRead one VM: power state, host, vCPU, RAM, NICs
vsphere_get_vm_metricsCPU, memory and disk usage for one VM
vsphere_list_hostsList ESXi hosts with capacity
vsphere_start_vmPower a VM on
vsphere_stop_vmShut a VM down (destructive — requires confirmed=True)
vsphere_restart_vmRestart a VM (destructive — requires confirmed=True)
vsphere_migrate_vmvMotion a VM to another host (destructive — requires confirmed=True)
vsphere_create_snapshotSnapshot a VM

Xen Orchestra / XCP-ng (15 tools)

ToolDescription
xo_list_vmsList VMs across all pools
xo_start_vmStart a VM
xo_stop_vmCleanly shut a VM down (destructive — requires confirmed=True)
xo_migrate_vmMigrate a VM to another host (destructive — requires confirmed=True)
xo_snapshot_vmSnapshot a VM
xo_list_srsList storage repositories — capacity planning before a migration
xo_list_tasksRunning XO tasks
xo_list_backupsList configured backup jobs
xo_run_backupRun a backup job now
xo_restore_backupRestore a VM from a backup (destructive — requires confirmed=True)
xo_rolling_pool_updatePatch every host in the pool, one at a time (destructive — requires confirmed=True)
xo_v2v_list_esxiList reachable ESXi hosts, for a VMware migration
xo_v2v_list_vmware_vmsList the VMs on an ESXi host
xo_v2v_migrateMigrate a VMware VM onto XCP-ng (destructive — requires confirmed=True)
xo_v2v_statusFollow a V2V migration in progress

Cross-hypervisor

ToolDescription
get_vm_usageVM usage and cost across every declared hypervisor

Configuration (3 tools)

ToolDescription
set_cluster_rateSet GPU cost rate for a cluster (admin only)
register_gatewayRegister a new gateway (returns one-time token)
delete_gatewayRevoke a gateway

Agent Infrastructure Control Plane (12 tools)

The missing layer between your AI agents and your GPU fleet. Works with any framework (n8n, LangChain, CrewAI, Dify) — just point to the VibOps LLM Proxy.

ToolDescription
FinOps per agent
get_agent_usageGPU cost per agent — tokens, requests, cost, GPU-hours. "Which agent costs the most?"
get_agent_usage_detailDrill-down on one agent — daily breakdown, model distribution, cost trend
get_agent_budgetCurrent budget + MTD spend for an agent
set_agent_budgetSet monthly spend limit — soft alert at 80%, hard block at 100% (HTTP 429)
Model access control
get_agent_model_rulesList model access rules — which agent can use which LLM
update_agent_model_ruleCreate a rule: glob patterns, deny-first. "RH agents → Mistral only"
Identity lifecycle
list_agent_identitiesList machine identities for agents
create_agent_identityCreate a new machine identity (key shown once)
rotate_agent_identityRotate the key for an existing identity
revoke_agent_identityRevoke an identity immediately
Dependency graph
get_agent_dependency_graphFull org-wide graph: agent→model, agent→connector, agent→sub-agent
get_agent_dependenciesDependencies for one agent — impact analysis before migration

Governance & Compliance (21 tools)

ToolDescription
list_anomaliesList GPU anomalies with optional cluster/status filter
get_open_anomaliesGet all currently open anomalies
resolve_anomalyMark an anomaly as resolved
list_compliance_controlsList compliance controls (filter by framework)
get_compliance_scoreGet the compliance score for a framework (or all)
update_compliance_controlUpdate status, notes, or evidence URL for a control
list_compliance_reportsList generated compliance reports
generate_compliance_reportGenerate a SOC 2, RGPD, or HIPAA report asynchronously
get_compliance_reportPoll/retrieve a generated compliance report
list_audit_logsQuery the immutable audit log with filters
verify_audit_chainVerify HMAC-SHA256 integrity of the full audit chain
get_policyGet the current organisation policy
update_policyReplace the organisation policy (immediate effect)
list_eval_rubricsList LLM-as-judge evaluation rubrics
evaluate_jobTrigger LLM-as-judge evaluation for a job
get_job_evaluationsRetrieve evaluation results for a job
get_ldap_configGet LDAP / Active Directory configuration
update_ldap_configConfigure or enable/disable LDAP integration
get_siem_configGet SIEM push export configuration
update_siem_configSet Splunk/Datadog SIEM destination
push_to_siemExport audit events to configured SIEM

GPU FinOps (4 tools)

ToolDescription
get_budgetGet current GPU budget and consumed spend
get_chargebackGet chargeback breakdown by tenant for a given month
get_spend_trendGet daily GPU spend trend (default: last 30 days)
get_waste_analysisIdentify idle GPU resources and cost optimisation opportunities

LLM Inference Proxy

VibOps includes a transparent OpenAI-compatible proxy (port 8004) that sits between your AI agents and LLM inference servers (vLLM, Ollama, TGI). Every inference request is logged with agent attribution for FinOps.

Your agents point to the proxy instead of the LLM directly:

# Before
OPENAI_BASE_URL=http://vllm:8000/v1

# After
OPENAI_BASE_URL=http://vibops-proxy:8004/v1

Add a X-VibOps-Agent-Id header to attribute costs per agent:

curl -X POST http://vibops-proxy:8004/v1/chat/completions \
  -H "X-VibOps-Agent-Id: pricing-agent-v2" \
  -H "X-VibOps-Team: supply-chain" \
  -d '{"model": "mistral:7b", "messages": [...]}'

The proxy captures: agent ID, team, model, tokens, latency, GPU cost — visible in the console FinOps dashboard and queryable via get_agent_usage.

Example prompts

"Clone my repo and deploy it on the GPU cluster."
"Deploy llama3:8b on vibops-dev with 2 replicas."
"Scale the inference deployment to 4 replicas on prod-cluster."
"What's our GPU utilisation trend over the last 7 days?"
"Show me the cost breakdown per cluster this week."
"Which clusters have open critical GPU alerts?"
"Are there any open GPU anomalies right now?"
"Scan my infrastructure and show discovered services."
"What's our AI Act compliance score and which controls are non-compliant?"
"Generate a SOC 2 report for Q1 2026."
"Verify the audit chain hasn't been tampered with."
"Which agent costs the most in GPU this month?"
"Show me the inference cost breakdown for the pricing agent."
"Which agents depend on the claude-opus-4-6 model?"
"Create a machine identity for the pricing-agent with a 1-year expiry."
"Show me the spend trend for the last 7 days and flag any waste."

Contributing

See CONTRIBUTING.md. All contributions require a DCO sign-off (git commit -s).

License

MIT — free to use, modify, and distribute. See LICENSE.

Built on FastMCP and VibOps — The AI Infrastructure Engine.