GraphMem
ग्राफ-आधारित मेमोरी प्रबंधन के लिए एक MCP सर्वर, जो AI को ज्ञान इकाइयों और उनके संबंधों को बनाने, पुनर्प्राप्त करने और प्रबंधित करने में सक्षम बनाता है।
दस्तावेज़
GraphMem: Graph-Based Memory MCP Server
GraphMem is a Ruby on Rails application implementing a Model Context Protocol (MCP) server for graph-based memory management. It enables AI assistants and other clients to create, retrieve, search, and manage knowledge entities and their relationships through a standardized interface.
Overview
GraphMem provides persistent, structured storage for knowledge entities, their relationships, and observations. It's designed as an MCP server that enables AI assistants to maintain memory across sessions, build domain-specific knowledge graphs, and effectively reference past interactions.
Single-User Design
GraphMem is designed as a single-user, local-network server: one graph for one implicit owner, shared by a handful of that owner's agents. There is no per-user authorization model, deliberately -- the shared graph is the point.
Access is controlled by network reach plus an optional shared bearer token. By default the MCP endpoint accepts unauthenticated requests only from loopback and private (RFC1918) ranges, and it can never be simultaneously unauthenticated and reachable from a public address. See docs/mcp_access_control.md for configuration, and set GRAPH_MEM_MCP_TOKEN before exposing the server beyond a trusted LAN.
Per-agent project context: Several MCP clients can connect concurrently, provided each sends a distinct X-MCP-Client value -- context is stored per client id, so two agents sharing one value will overwrite each other's scope. GraphMem detects that case and adds a warning to set_context and get_context responses. Each agent identifies itself in its MCP configuration:
{
"mcpServers": {
"graph_mem": {
"url": "http://localhost:3030/mcp",
"headers": { "X-MCP-Client": "Agent-1" }
}
}
}
The example above assumes using the container setup, which is hardcoded to port 3030. (APP_PORT is currently hardcoded in Dockerfile and docker-compose.yml).
Use /mcp for the default 2025-03-26 Streamable HTTP profile. Use
/mcp/readonly for context and read tools only, or /mcp/maintenance for the
full catalog including maintenance operations. The legacy 2024-11-05 SSE
endpoint remains available at /mcp/sse and uses the default profile.
The server also publishes orient, recall, and persist MCP prompts. Every
successful tool call includes the GraphMem version, a concise next_move, and
a context banner when the client has not selected a project.
The 12 default tools advertise outputSchema and return matching
structuredContent plus mirrored JSON text for backward-compatible clients.
Context set via set_context is stored per client_id in the database and survives server restarts. Agents without the header share the "default" client bucket (backward-compatible single-agent behavior), so give every agent its own value once you run more than one.
The header is a cooperative scope key, not a credential: an agent can claim any client id, and the shared token grants identical full read/write access to everyone holding it. Multi-tenant isolation remains out of scope -- if two people need separate memories, run two instances against two databases.
Key capabilities:
- Vector semantic search via MariaDB 11.8 native VECTOR support + Ollama embeddings
- Project context scoping -- per-agent via
X-MCP-Client, persisted across restarts - Entity type canonicalization to prevent graph fragmentation
- Auto-deduplication on entity creation
- Hybrid search combining text tokenization with vector similarity (with context boosting)
- Docker Compose deployment with auto-start support
- LAN-wide embedding architecture using a centralized Ollama host
Technology Stack
- Ruby: 3.4.1+
- Rails: 8.1.2+
- MCP Implementation: fast-mcp gem with a custom
GraphMem::McpStreamableHttpTransportthat adds 2025-03-26 Streamable HTTP support while keeping the 2024-11-05 SSE transport - Database: MariaDB 11.8+ (VECTOR support required)
- Embeddings: Ollama with nomic-embed-text (768 dimensions)
Features
Rules & Skills
GraphMem ships agent-facing rules and a vendor-neutral skill:
Using GraphMem as an MCP toolset (your agent talks to a running server):
docs/rules/graph_mem_mcp_rules.md— always-on rules to copy into your agent's rule set (per-agent install targets:docs/rules/README.md)docs/rules/general_coding_rules.md— optional, project-agnostic coding rulesskills/graph-mem-mcp-toolset/SKILL.md— vendor-neutral operational skill (works in Cursor, Claude Code, Devin, ...; the.cursor/skills/copy is a thin adapter)
Developing GraphMem itself (editing this repo): see AGENTS.md,
docs/development.md, and the repo-scoped rules under .cursor/rules/ and
.devin/rules/.
MCP Tools
GraphMem exposes the following MCP tools:
Entity Management
get_entities-- Retrieve one or more entities with observations and all/internal relation projectionssearch-- Ranked summary search, projected subgraph search, or paginated catalog listing
Graph Mutation
graph_write-- Atomically create entities, observations, and relationsgraph_edit-- Atomically update entity metadata or observation content/lifecyclegraph_delete-- Atomically delete, obsolete, or merge graph records
Observations use an active, obsolete, or superseded lifecycle. Normal entity reads, graph traversal, relationship discovery, and search expose active observations only. Use get_entities(include_obsolete: true) or include_obsolete=true on observation REST/resource listings to inspect retained history. Explicit duplicate-cleanup maintenance operations still hard-delete redundant active rows.
Relationship Management
traverse_graph-- Bounded multi-hop traversal or direct endpoint/type relation queryfind_shortest_path-- Shortest path (by hop count) between two entities
The prior read and mutation names remain callable compatibility aliases but are
hidden from tools/list while telemetry measures remaining use.
Context and Workflow
set_context-- Scope subsequent operations to a project (perX-MCP-Client)get_context-- Check the active project contextclear_context-- Remove project scopingsuggest_merges-- Find potential duplicate entities via vector similaritydream_state_status-- Report background graph compaction state (running/paused/cursor)get_maintenance_reports-- Read maintenance/compaction reports, including thecompaction_reviewqueue
Utility
get_version-- Server versionget_current_time-- Server time in ISO 8601
Dream-State Compaction
A background dream-state job (DreamStateCompactionJob, scheduled via Solid Queue recurring.yml) periodically compacts the knowledge graph:
- Orphan phase -- attaches high-confidence orphan nodes to matching
Projectroots - Tree-walk phase -- deduplicates identical observations and auto-merges very similar entities (cosine distance < 0.10)
- Review queue -- lower-confidence merges and orphan matches are written to
maintenance_reports(compaction_review), readable viaget_maintenance_reports
The job is cooperatively pausable: mutating MCP tools request a pause when compaction is running, so live tool traffic takes priority. Use dream_state_status to inspect cursor position and stats, and get_maintenance_reports to review and action the queued suggestions. Paused runs resume on the next scheduled trigger.
MCP Resources
memory_entities-- Query entities with filtering, sorting, and relation inclusionmemory_observations-- Access observations with advanced filteringmemory_relations-- Query relationships with bidirectional entity inclusionmemory_graph-- Graph traversals starting from any entity
REST API
Full REST API at /api/v1 for direct integration. Swagger docs available at /api-docs.
Graph Visualization
Interactive Cytoscape.js-based graph visualization at the server root (/), with contextual menus, drag-and-drop operations, and data management features.
Quick Start with Docker
Prerequisites: Ollama must be running on the host with at least one embedding model pulled:
# Install Ollama (https://ollama.com) then pull the default embedding model
ollama pull nomic-embed-text
The
appcontainer usesnetwork_mode: host, so it shares the host's network stack.localhost:11434reaches Ollama with no bridge/firewall configuration needed. The app binds directly to host port 3030.
# Clone and enter the project
git clone https://github.com/steveoro/graph_mem.git
cd graph_mem
# Copy example config and set your master key
cp .env.example .env
# Edit .env: set RAILS_MASTER_KEY (from config/master.key) and DB_PASSWORD
# Start the stack (MariaDB 11.8 + Rails in production mode)
docker compose up -d
# Seed canonical entity types
docker compose exec app bin/rails db:seed
# Verify Ollama connectivity
docker compose exec app bin/rails embeddings:check
# Backfill embeddings
docker compose exec app bin/rails embeddings:backfill
# Update the container after a repository pull
docker compose down && docker compose up -d --build
The app is available at http://localhost:3030. Swagger API docs at http://localhost:3030/api-docs.
The app port (3030) is hardcoded on Dockerfile and docker-compose.yml because the service relies on host networking to access the embedding service by ollama. This allows a simpler container setup on different machines without resorting to iptables or firewall mangling.
This containerized app is a single-user server with no authentication layer -- it is designed to run locally on a machine and/or be accessible only through a trusted LAN. Do not expose this service to the public internet.
LAN sharing
To allow a local Ubuntu server running graph_mem in a container with ollama running as a service for embedding processing, remember to allow incoming trafic if you're using ufw (assuming your local LAN is set on 192.168.0.0/24):
sudo ufw allow from 192.168.0.0/24 to any port 3030 proto tcp comment "GraphMem from LAN"
sudo ufw allow from 192.168.0.0/24 to any port 11434 proto tcp comment "Ollama from LAN"
This way, the graph_mem UI will be accessible on http://<graph_mem_server_ip>:3030/ while the MCP server will be at http://<graph_mem_server_ip>:3030/mcp/sse.
Access control on a shared LAN
The MCP endpoint accepts unauthenticated requests from private ranges by default, so every device on the LAN has full read/write access to the graph. Once the network is not fully trusted, set a shared token on the server:
# in .env on the GraphMem host
GRAPH_MEM_MCP_TOKEN=$(openssl rand -hex 32)
GRAPH_MEM_MCP_ALLOWED_IPS=192.168.0.0/24
and add it to every client config alongside the X-MCP-Client header:
"headers": {
"Authorization": "Bearer <GRAPH_MEM_MCP_TOKEN>",
"X-MCP-Client": "cursor-1"
}
Clients may be updated before the server: while no token is configured the header is ignored, so you can roll it out without downtime. Full reference in docs/mcp_access_control.md.
Native Development Setup
For local development, run the app natively with MariaDB on localhost.
-
Prerequisites:
- Ruby 3.4.1+ (via RVM or rbenv)
- MariaDB 11.8+ (for vector search)
- Ollama with an embedding model
-
Install dependencies:
bundle install -
Database setup:
cp config/database.example.yml config/database.yml # Edit config/database.yml with your MariaDB credentials bin/rails db:prepare bin/rails db:seed -
Pull an embedding model:
ollama pull nomic-embed-text -
Backfill embeddings:
bin/rails embeddings:backfill -
Start the development server:
bin/dev
Setting Up the MCP Client
Cursor
Edit your Cursor's mcp.json:
Option A -- Streamable HTTP transport (recommended for modern clients):
{
"mcpServers": {
"graph_mem": {
"url": "http://localhost:3030/mcp",
"headers": { "X-MCP-Client": "Agent-1" }
}
}
}
Option B -- Legacy SSE transport (Docker / older clients):
{
"mcpServers": {
"graph_mem": {
"url": "http://localhost:3030/mcp/sse"
}
}
}
Option C -- stdio transport (native development / real-time changes applied):
{
"mcpServers": {
"graph_mem": {
"command": "/bin/bash",
"args": ["/absolute/path/to/graph_mem/bin/mcp_graph_mem_runner.sh"],
"env": { "RAILS_ENV": "development" }
}
}
}
Option D -- stdio via Docker:
{
"mcpServers": {
"graph_mem": {
"command": "/bin/bash",
"args": ["/absolute/path/to/graph_mem/bin/docker-mcp"]
}
}
}
Windsurf
Edit your Windsurf's mcp_config.json using the same approach as Cursor's. Usually the SSE container doesn't yield any issue.
Antigravity / Devin (Streamable HTTP)
GraphMem now supports the 2025-03-26 Streamable HTTP transport. Point these clients at the base MCP URL:
{
"mcpServers": {
"graph_mem": {
"url": "http://localhost:3030/mcp",
"headers": { "X-MCP-Client": "Agent-1" }
}
}
}
The legacy SSE endpoint (/mcp/sse) remains available for older clients.
Assuming graph_mem-app-1 is the running container name, edit mcp_servers.json:
{
"mcpServers": {
"graph_mem": {
"command": "/usr/bin/docker",
"args": [
"exec",
"-i",
"graph_mem-app-1",
"bash",
"-c",
"'bin/bundle exec ruby bin/mcp_stdio_runner.rb'"
// For development, change env below and replace the command with bash, first arg "-c" and second arg:
// "cd /home/steve/Projects/graph_mem && exec /usr/share/rvm/wrappers/ruby-3.4.1@graph_mem/bundle exec ruby bin/mcp_stdio_runner.rb"
],
"env": {
"RAILS_ENV": "production"
}
}
}
}
Claude Code
Typically both stdio and SSE should work as all 3 options highlighted above should.
Edit your Claude's mcp_servers.json with the one of your choosing.
LAN Access (Multiple Machines)
See Architecture for details.
Mind that the current version of GraphMem is specifically designed for single-user usage only, meaning 1 AI user per installation: currently there's no session storage per conversation, so multiple AI agents connecting and using the same running GraphMem instance may overwrite each other's work (one could reset the current context of another preventing the "context valve" to work as expected, thus leading to context bloat).
But nothing prevents you to share the same generated memory graph among different workstations, provided GraphMem will be used by a single AI user at a time. Or, more realistically, deploy GraphMem on a performant server and access it your usual workstation.
So, when running GraphMem on a local server and accessing it from other machines on the LAN:
1. Expose Ollama on the runner host
By default Ollama only listens on 127.0.0.1. Create a systemd drop-in override
(survives Ollama package upgrades):
sudo mkdir -p /etc/systemd/system/ollama.service.d
echo '[Service]
Environment="OLLAMA_HOST=0.0.0.0"' | sudo tee /etc/systemd/system/ollama.service.d/override.conf
sudo systemctl daemon-reload
sudo systemctl restart ollama
Verify it's bound to all interfaces:
ss -tlnp | grep 11434
# Should show *:11434 instead of 127.0.0.1:11434
2. Configure OLLAMA_URL
On the host running GraphMem, OLLAMA_URL=http://localhost:11434 (the default) works
because the app container uses host networking.
If Ollama runs on an another different machine, set in .env:
OLLAMA_URL=http://<ollama-host-ip>:11434
3. Connect MCP clients from other LAN machines
Use the Streamable HTTP endpoint for modern clients:
{ "url": "http://<workstation-ip>:3030/mcp" }
The legacy SSE endpoint is still available:
{ "url": "http://<workstation-ip>:3030/mcp/sse" }
Embedding Management
Operator dashboard
Sign in at /operator/login, then open Embeddings from the home dashboard or go to /operator/embeddings. The page shows coverage, index status, resolved configuration (with source badges: AppSettings / ENV / Default), connection test, backfill/regenerate jobs, and add/drop ANN index actions.
Embedding service settings live under System Settings → Embeddings (/operator/settings?tab=embeddings). Configuration resolves in this order:
| Priority | Source |
|---|---|
| 1 | AppSettings (operator UI) — blank string or 0 dims defers to ENV |
| 2 | Environment variables (OLLAMA_URL, EMBEDDING_MODEL, EMBEDDING_PROVIDER, EMBEDDING_DIMS) |
| 3 | Built-in defaults (http://localhost:11434, nomic-embed-text, ollama, 768) |
ENV variables remain the right choice for Docker and deployment manifests; the UI overrides them when values are set. Enable scheduled backfill in settings to run EmbeddingScheduledBackfillJob daily (see config/recurring.yml).
See docs/operator/embeddings.md for the recommended operator workflow.
Testing Ollama Connectivity
Before backfilling or regenerating embeddings, verify that the app can reach your Ollama instance:
# Docker
docker compose exec app bin/rails embeddings:check
# Native
bin/rails embeddings:check
This sends a single test embedding through EmbeddingService using the resolved configuration (AppSettings → ENV → defaults). It reports the resolved config, response latency, and vector dimensions — the exact same code path used by backfill and regenerate.
For a lower-level check, curl is available inside the production container (host networking
means localhost reaches Ollama directly):
# Verify Ollama is reachable and list available models
docker compose exec app curl -sf http://localhost:11434/api/tags
# Test a raw embedding request
docker compose exec app curl -sf http://localhost:11434/api/embed \
-d '{"model":"nomic-embed-text","input":"hello"}'
Rake Tasks
| Task | Description |
|---|---|
embeddings:check | Smoke-test Ollama connectivity and config |
embeddings:backfill | Generate embeddings for records missing them |
embeddings:regenerate | Recompute all embeddings in-place (e.g. after switching models) |
embeddings:add_indexes | Add VECTOR INDEX (HNSW, cosine) after all rows are populated |
embeddings:drop_indexes | Remove indexes and revert columns to nullable |
Switching Embedding Models
To change the model (e.g. from nomic-embed-text to a different one):
- Pull the new model on the Ollama host:
ollama pull <model-name> - Update the model (and dimensions if different) in System Settings → Embeddings or via
EMBEDDING_MODEL/EMBEDDING_DIMSin.env - Verify connectivity:
bin/rails embeddings:checkor the operator Test connection button - Recompute all vectors:
bin/rails embeddings:regenerateor the operator Regenerate all action
Database Backup & Restore
Backups are managed through System Settings (/operator/settings, session login) and rake tasks. Dumps are timestamped, environment-scoped, and retained according to backup_keep_max:
# Dump current database to <backup_folder>/<YYYYMMDDHHMM>_<env>.sql.bz2
bin/rails db:dump
# List backups for the current environment
bin/rails db:list_backups
# Restore from newest backup, or a specific file
bin/rails db:restore
FILE=202601011200_production.sql.bz2 bin/rails db:restore
Scheduled backups run via Solid Queue (DatabaseBackupJob) when Enable scheduled backups is on in System Settings. Production schedule: 1pm and 5pm GMT (config/recurring.yml). Use the Jobs dashboard at /operator/jobs (same operator session) to inspect queue status.
Sign in at /operator/login. Default operator credentials: operator / changeme (override with OPERATOR_USERNAME / OPERATOR_PASSWORD or Rails credentials under operator:).
Environment Variables
| Variable | Default | Description |
|---|---|---|
OLLAMA_URL | http://localhost:11434 | Ollama API base URL (overridden by AppSettings embedding_url when set) |
EMBEDDING_MODEL | nomic-embed-text | Ollama model name for embeddings |
EMBEDDING_PROVIDER | ollama | ollama or openai_compatible |
EMBEDDING_DIMS | 768 | Vector dimensions (must match model) |
DB_PASSWORD | my_password | MariaDB root password |
DB_NAME | graph_mem | Database name |
DB_PORT | 3307 | Host port for MariaDB (Docker) |
RAILS_MASTER_KEY | -- | Rails credentials key (required for Docker) |
DATABASE_URL | -- | Full database URL (overrides individual DB settings) |
DB_BACKUP_HOST_PATH | ./db/backup | full path to DB backup(s) folder (default is invalid: docker-compose won't expand special characters) |
OPERATOR_USERNAME | operator | Operator login username for the web dashboard |
OPERATOR_PASSWORD | changeme | Operator login password (change in production) |
Documentation
- Operator embeddings guide
- MCP Tools Reference
- App Settings Reference
- Architecture
- Development Guide
- Troubleshooting
- Memory Entity Resource
- Memory Observation Resource
- Memory Relation Resource
- Memory Graph Resource
Contributing
Pull requests are welcome. For major changes, please open an issue first to discuss what you would like to change. Please make sure to update tests as appropriate. Pull requests without proper test cases won't be accepted.
License
The project is available as open source under the terms of the LGPL-3.0 License.