LocaLLama MCP Server

by Heratiki

42 572 downloads Not rated yet

About

An MCP Server that works with Roo Code/Cline.Bot/Claude Desktop to optimize costs by intelligently routing coding tasks between local LLMs free APIs and paid APIs.

Explore

- Node.js 22+
- npm
- At least one of: Ollama, LM Studio, llama.cpp server, or an OpenRouter API key

| Variable | Default | Description |
|---|---|---|
| LM_STUDIO_ENDPOINT | — | LM Studio API base URL |
| OLLAMA_ENDPOINT | — | Ollama API base URL |
| LLAMA_CPP_ENDPOINT | — | llama-server URL; leave unset to disable provider |
| DEFAULT_LOCAL_MODEL | — | Model name used when offloading to local provider |
| TOKEN_THRESHOLD | 1500 | Token count above which local offload is considered |
| COST_THRESHOLD | 0.02 | USD cost above which local offload is preferred |
| QUALITY_THRESHOLD | 0.7 | Quality score below which paid API is always used |
| RELIABLE_BENCHMARK_COUNT | 3 | Benchmark runs required before empirical scores are treated as fully reliable |
| MIN_VALIDATOR_SCORE | 0.6 | Minimum validation score required before a model is eligible for external validation |
| VALIDATION_RETRY_BUDGET | 1 | Validation retry attempts allowed after an initial failed validation |
| PROVIDER_MAX_CONCURRENT_LOCAL | 1 | Shared local execution slot count |
| PROVIDER_MAX_CONCURRENT_REMOTE | 5 | Per-remote-provider slot count |
| OPENROUTER_API_KEY | — | Enables OpenRouter provider and related tools |
| OPENROUTER_FREE_ONLY | false | Restrict OpenRouter to free-tier models only |
| EXPECT_LOCAL_PROVIDER_DOWN | — | Set true in test-operational.mjs to assert no local suggestion |

Build the server, then point your MCP client at node dist/index.js:

{
  "mcpServers": {
    "locallama": {
      "command": "node",
      "args": ["/path/to/locallama-mcp/dist/index.js"],
      "env": {
        "LM_STUDIO_ENDPOINT": "http://localhost:1234/v1",
        "OLLAMA_ENDPOINT": "http://localhost:11434/api",
        "DEFAULT_LOCAL_MODEL": "qwen2.5-coder-3b-instruct",
        "TOKEN_THRESHOLD": "1500",
        "COST_THRESHOLD": "0.02",
        "QUALITY_THRESHOLD": "0.07",
        "OPENROUTER_API_KEY": "your_openrouter_api_key_here"
      }
    }
  }
}

Claude Code users can place this in .mcp.json (project-scoped) or ~/.claude/settings.json (global).

npm run benchmark
npm run benchmark:comprehensive

Results are stored in benchmark-results/ as JSON and Markdown summaries.

route_task

`task`, `context_length`, `expected_output_length?`, `complexity?`, `priority?`, `preemptive?`

get_task_status

`task_id`

cancel_task

`task_id`

cancel_job

`job_id`

preemptive_route_task

`task`, `context_length`, `expected_output_length?`, `complexity?`, `priority?`

get_cost_estimate

`context_length`, `expected_output_length?`, `model?`

benchmark_task

`task_id`, `task`, `context_length`, `expected_output_length?`, `complexity?`, `local_model?`, `paid_model?`, `runs_per_task?`

benchmark_tasks

`tasks[]`, `runs_per_task?`, `parallel?`, `max_parallel_tasks?`

benchmark_model

`model_id`, `provider_id?`, `task_categories?`

retriv_init

`directories[]`, `exclude_patterns?`, `chunk_size?`, `force_reindex?`, `bm25_options?`

retriv_search

`query`, `limit?`

reload_config

—

check_for_updates

—

update_server

—

get_free_models

—

clear_openrouter_tracking

—

benchmark_free_models

`tasks[]`, `runs_per_task?`, `parallel?`, `max_parallel_tasks?`

set_model_prompting_strategy

`model_id`, `system_prompt`, `user_prompt`, `use_chat`, `assistant_prompt?`, `success_rate?`, `quality_score?`

| Tool | Inputs | Description |
|---|---|---|
| route_task | task, context_length, expected_output_length?, complexity?, priority?, preemptive? | Queue a task asynchronously. Returns task_id immediately. Poll get_task_status for results. |
| get_task_status | task_id | Poll a non-blocking route_task submission. Returns status, progress, and inline result when complete. |
| cancel_task | task_id | Cancel all queued or in-progress jobs for a task. |
| cancel_job | job_id | Cancel a single background job. |
| preemptive_route_task | task, context_length, expected_output_length?, complexity?, priority? | Heuristic routing check with no LLM calls. Returns model/provider recommendation without executing the task. |
| get_cost_estimate | context_length, expected_output_length?, model? | Estimate USD cost before calling route_task. Local and free-tier models return 0. |
| benchmark_task | task_id, task, context_length, expected_output_length?, complexity?, local_model?, paid_model?, runs_per_task? | Benchmark one task across local vs paid models. |
| benchmark_tasks | tasks[], runs_per_task?, parallel?, max_parallel_tasks? | Benchmark multiple tasks in one call. |
| benchmark_model | model_id, provider_id?, task_categories? | Run built-in benchmark suites against a specific model. Persists results to benchmarks.db and updates ModelRegistry capability scores. |
| retriv_init | directories[], exclude_patterns?, chunk_size?, force_reindex?, bm25_options? | Index code with the native BM25 engine (no Python required). |
| retriv_search | query, limit? | Search indexed code using native BM25. |
| reload_config | — | Reload .env at runtime. Atomic: invalid config is rejected. |
| check_for_updates | — | Check whether the server is up to date with the latest GitHub commit. |
| update_server | — | Pull latest changes from GitHub, run npm install and npm run build. Restart the server manually after. |

| Tool | Inputs | Description |
|---|---|---|
| get_free_models | — | List free models available from OpenRouter. |
| clear_openrouter_tracking | — | Clear cached model list and force a fresh fetch. |
| benchmark_free_models | tasks[], runs_per_task?, parallel?, max_parallel_tasks? | Benchmark free OpenRouter models. Results written to benchmarks.db. |
| set_model_prompting_strategy | model_id, system_prompt, user_prompt, use_chat, assistant_prompt?, success_rate?, quality_score? | Set a custom prompting strategy for an OpenRouter model. |

retriv_init { "directories": ["/path/to/repo"], "force_reindex": true }
retriv_search { "query": "pagination logic" }
```

Status: experimental
Latest release
License: ISC

Local-first, provider-neutral Model Context Protocol server for coding-agent workflows. Routes tasks across local models (Ollama, LM Studio, llama.cpp), free OpenRouter models, and paid frontier models using cost, latency, context capacity, and benchmark history.

Node.js: >=22

> ⚠️ Early / experimental — not yet a stable release. This project is under active, rapid development and has not been fully verified end-to-end. MCP tool signatures, configuration, and behavior may change between releases without notice.
>
> Version numbers follow SemVer mechanically (they're derived from Conventional Commit messages, not hand-picked), so a 1.x number signals only "a public surface exists" — it is not a promise of stability or completeness. If you depend on this server, pin to an exact version.
>
> - Tagged releases on main are the relatively safer builds.
> - The testing channel publishes bleeding-edge pre-releases (x.y.z-testing.n) for trying unproven changes early.

Overview

LocalLama MCP reduces token costs without sacrificing quality. Tasks are queued asynchronously — route_task returns a task_id immediately; callers poll get_task_status for results. The decision engine chooses local → free → paid based on measured provider capabilities and configurable thresholds.

Supported MCP clients: Codex, Claude Code, Claw Code, Cursor, GitHub Copilot Agent mode, and any generic MCP stdio client.

Requirements

- Node.js 22+
- npm
- At least one of: Ollama, LM Studio, llama.cpp server, or an OpenRouter API key

Installation

git clone https://github.com/Heratiki/locallama-mcp.git
cd locallama-mcp
npm install
npm run build

Configuration

Copy .env.example to .env and edit with your values. The server resolves .env from its own root directory (or LOCALLAMA_ROOT_DIR when set), not from the MCP host's CWD.

```env

No reviews yet — be the first

Sign in to leave a review

Use Google, GitHub, or an email account so ratings stay tied to real people.

Email sign in

No reviews posted yet.