Llama Hot Swap

by oussama-kh

237 downloads Not rated yet
GitHub

About

MCP server for hot-swapping llama.cpp models in Claude Code - launchctl (macOS) + systemd (Linux)

Explore

- Hot-swap llama.cpp models without context loss
- Supports macOS (launchctl) and Linux (systemd)
- Mapped mode and directory mode for model discovery
- Create new model service configs from within Claude Code
- Works with any llama.cpp‑compatible model (GGUF)

Setting up with Highlight

This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:

  1. Download and install Highlight from highlightai.com/download
  2. Navigate to the plugins tab and select "Add Custom Plugin"
  3. Configure the plugin with the settings below
    Plugin Name Llama Hot Swap
    Command (node, npx, python, etc.)

    Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.

  4. Enable "Start Automatically" if you want the plugin to start when Highlight launches

From the repository


uvx mcp-llama-swap

pip install mcp-llama-swap

Add to ~/.claude.json:

{
  "mcpServers": {
    "llama-swap": {
      "command": "uvx",
      "args": ["mcp-llama-swap"],
      "env": {
        "LLAMA_SWAP_CONFIG": "/path/to/config.json"
      }
    }
  }
}

Create config.json (macOS):

{
  "plists_dir": "~/.llama-plists",
  "health_url": "http://localhost:8000/health",
  "health_timeout": 30,
  "models": {
    "planner": "qwen35-thinking.plist",
    "coder": "qwen3-coder.plist",
    "fast": "glm-flash.plist"
  }
}

Or on Linux:

{
  "services_dir": "~/.llama-services",
  "health_url": "http://localhost:8000/health",
  "health_timeout": 30,
  "models": {
    "planner": "llama-server-planner.service",
    "coder": "llama-server-coder.service"
  }
}
pip install mcp-llama-swap
pip install litellm

Create litellm_config.yaml:

model_list:
  - model_name: ""
    litellm_params:
      model: "openai/"
      api_base: "http://localhost:8000/v1"
      api_key: "sk-none"

litellm_settings:
drop_params: true
request_timeout: 300

Start it:

litellm --config litellm_config.yaml --port 4000

On macOS, you can use the included ai.litellm.proxy.plist.template to run it as a persistent launchd service (see setup.sh).

Copy config.example.json (macOS) or config.example.linux.json (Linux) and edit with your model aliases and service filenames.

You can create service configs manually, or use the create_model_config MCP tool inside Claude Code:

You: create a model config named "coder" for /path/to/model.gguf with 8192 context

This generates the appropriate launchd plist (macOS) or systemd unit file (Linux) in your services directory.

If you prefer a one-shot setup on macOS, clone this repo and run:

git clone https://github.com/oussama-kh/mcp-llama-swap.git ~/mcp-llama-swap
cd ~/mcp-llama-swap
chmod +x setup.sh
./setup.sh

The script creates a virtual environment, installs dependencies, configures the LiteLLM launchd service, and prints the exact config to add.

config.json fields:

| Field | Default | Description |
|-------|---------|-------------|
| services_dir | ~/.llama-plists (macOS) / ~/.llama-services (Linux) | Directory containing model service configs |
| plists_dir | — | macOS alias for services_dir (backwards compatible) |
| units_dir | — | Linux alias for services_dir |
| health_url | http://localhost:8000/health | llama-server health endpoint |
| health_timeout | 30 | Seconds to wait for health check after loading |
| models | {} | Alias-to-filename map. Empty = directory mode |
| platform | auto | Service manager: auto, launchctl, or systemd |
| launchctl_mode | legacy | macOS only: legacy (load/unload) or modern (bootstrap/bootout) |

Override config path via the LLAMA_SWAP_CONFIG environment variable.

pip install -e ".[test]"

list_models

Lists all configured models with load status and current mode

get_current_model

Returns the alias of the currently loaded model

swap_model

Unloads current model, loads the specified one, waits for health check

create_model_config

Generates a new launchd plist (macOS) or systemd unit (Linux) for a model

| Tool | Description |
|------|-------------|
| list_models | Lists all configured models with load status and current mode |
| get_current_model | Returns the alias of the currently loaded model |
| swap_model | Unloads current model, loads the specified one, waits for health check |
| create_model_config | Generates a new launchd plist (macOS) or systemd unit (Linux) for a model |

Claude Desktop / Cursor

Paste into your MCP client config file to install this server.

{
    "mcpServers": {
        "llama hot swap": {
            "mcp-llama-swap": {
                "command": "uvx",
                "args": [
                    "mcp-llama-swap"
                ]
            }
        }
    }
}

McpServers

{
    "mcp-llama-swap": {
        "command": "uvx",
        "args": [
            "mcp-llama-swap"
        ]
    }
}

mcp-llama-swap

PyPI version
License
Python 3.10+

Hot-swap llama.cpp models inside a running Claude Code session. No context loss. One command.

> Plan with a reasoning model. Implement with a coding model. Same session, same context, zero manual overhead.

Supports macOS (launchctl) and Linux (systemd).

<!-- TODO: Replace with actual recording
demo
-->

Why

Running local LLMs means choosing between a strong reasoning model and a fast coding model. You can't load both on a single machine. Manually swapping models kills your conversation context and flow.

mcp-llama-swap solves this by giving Claude Code a tool to swap the model behind llama-server via your system's service manager (launchctl on macOS, systemd on Linux), while preserving the full conversation history client-side.

Quick Start

Install

```bash

No reviews yet — be the first

Sign in to leave a review

Use Google, GitHub, or an email account so ratings stay tied to real people.

Email sign in

No reviews posted yet.