Llama Hot Swap
About
MCP server for hot-swapping llama.cpp models in Claude Code - launchctl (macOS) + systemd (Linux)
Explore
- Hot-swap llama.cpp models without context loss
- Supports macOS (launchctl) and Linux (systemd)
- Mapped mode and directory mode for model discovery
- Create new model service configs from within Claude Code
- Works with any llama.cpp‑compatible model (GGUF)
Setting up with Highlight
This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:
- Download and install Highlight from highlightai.com/download
- Navigate to the plugins tab and select "Add Custom Plugin"
-
Configure the plugin with the settings below
Plugin Name
Llama Hot SwapCommand (node, npx, python, etc.)Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.
- Enable "Start Automatically" if you want the plugin to start when Highlight launches
From the repository
uvx mcp-llama-swap
pip install mcp-llama-swap
Add to ~/.claude.json:
{
"mcpServers": {
"llama-swap": {
"command": "uvx",
"args": ["mcp-llama-swap"],
"env": {
"LLAMA_SWAP_CONFIG": "/path/to/config.json"
}
}
}
}
Create config.json (macOS):
{
"plists_dir": "~/.llama-plists",
"health_url": "http://localhost:8000/health",
"health_timeout": 30,
"models": {
"planner": "qwen35-thinking.plist",
"coder": "qwen3-coder.plist",
"fast": "glm-flash.plist"
}
}
Or on Linux:
{
"services_dir": "~/.llama-services",
"health_url": "http://localhost:8000/health",
"health_timeout": 30,
"models": {
"planner": "llama-server-planner.service",
"coder": "llama-server-coder.service"
}
}
pip install mcp-llama-swap
pip install litellm
Create litellm_config.yaml:
model_list:
- model_name: ""
litellm_params:
model: "openai/"
api_base: "http://localhost:8000/v1"
api_key: "sk-none"
litellm_settings:
drop_params: true
request_timeout: 300
Start it:
litellm --config litellm_config.yaml --port 4000
On macOS, you can use the included ai.litellm.proxy.plist.template to run it as a persistent launchd service (see setup.sh).
Copy config.example.json (macOS) or config.example.linux.json (Linux) and edit with your model aliases and service filenames.
You can create service configs manually, or use the create_model_config MCP tool inside Claude Code:
You: create a model config named "coder" for /path/to/model.gguf with 8192 context
This generates the appropriate launchd plist (macOS) or systemd unit file (Linux) in your services directory.
If you prefer a one-shot setup on macOS, clone this repo and run:
git clone https://github.com/oussama-kh/mcp-llama-swap.git ~/mcp-llama-swap
cd ~/mcp-llama-swap
chmod +x setup.sh
./setup.sh
The script creates a virtual environment, installs dependencies, configures the LiteLLM launchd service, and prints the exact config to add.
config.json fields:
| Field | Default | Description |
|-------|---------|-------------|
| services_dir | ~/.llama-plists (macOS) / ~/.llama-services (Linux) | Directory containing model service configs |
| plists_dir | — | macOS alias for services_dir (backwards compatible) |
| units_dir | — | Linux alias for services_dir |
| health_url | http://localhost:8000/health | llama-server health endpoint |
| health_timeout | 30 | Seconds to wait for health check after loading |
| models | {} | Alias-to-filename map. Empty = directory mode |
| platform | auto | Service manager: auto, launchctl, or systemd |
| launchctl_mode | legacy | macOS only: legacy (load/unload) or modern (bootstrap/bootout) |
Override config path via the LLAMA_SWAP_CONFIG environment variable.
pip install -e ".[test]"
list_models
Lists all configured models with load status and current mode
get_current_model
Returns the alias of the currently loaded model
swap_model
Unloads current model, loads the specified one, waits for health check
create_model_config
Generates a new launchd plist (macOS) or systemd unit (Linux) for a model
| Tool | Description |
|------|-------------|
| list_models | Lists all configured models with load status and current mode |
| get_current_model | Returns the alias of the currently loaded model |
| swap_model | Unloads current model, loads the specified one, waits for health check |
| create_model_config | Generates a new launchd plist (macOS) or systemd unit (Linux) for a model |
Claude Desktop / Cursor
Paste into your MCP client config file to install this server.
{
"mcpServers": {
"llama hot swap": {
"mcp-llama-swap": {
"command": "uvx",
"args": [
"mcp-llama-swap"
]
}
}
}
}
McpServers
{
"mcp-llama-swap": {
"command": "uvx",
"args": [
"mcp-llama-swap"
]
}
}
mcp-llama-swap
Hot-swap llama.cpp models inside a running Claude Code session. No context loss. One command.
> Plan with a reasoning model. Implement with a coding model. Same session, same context, zero manual overhead.
Supports macOS (launchctl) and Linux (systemd).
<!-- TODO: Replace with actual recording

-->
Why
Running local LLMs means choosing between a strong reasoning model and a fast coding model. You can't load both on a single machine. Manually swapping models kills your conversation context and flow.
mcp-llama-swap solves this by giving Claude Code a tool to swap the model behind llama-server via your system's service manager (launchctl on macOS, systemd on Linux), while preserving the full conversation history client-side.
Quick Start
Install
```bash
Sign in to leave a review
Use Google, GitHub, or an email account so ratings stay tied to real people.
No reviews posted yet.



