llm-cli-gateway
About
Unified MCP server providing access to Claude Code, Codex, and Gemini CLIs through a single gateway. Features multi-LLM orchestration, persistent session management, async job execution with polling, approval gates, retry with circuit breakers, and token optimization. Install: npx -y llm-cli-gateway
Explore
Setting up with Highlight
This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:
- Download and install Highlight from highlightai.com/download
- Navigate to the plugins tab and select "Add Custom Plugin"
-
Configure the plugin with the settings below
Plugin Name
llm-cli-gatewayCommand (node, npx, python, etc.)Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.
- Enable "Start Automatically" if you want the plugin to start when Highlight launches
From the repository
$Version = '<version>' $Base = "https://github.com/verivus-oss/llm-cli-gateway/releases/download/v$Version" $InstallDir = Join-Path (Join-Path $env:LOCALAPPDATA 'Programs') 'llm-cli-gateway' $ExeName = "llm-cli-gateway-$Version-windows-amd64.exe" $BundleName = "llm-cli-gateway-bundle-$Version-windows-amd64.tar.gz" $Exe = Join-Path $InstallDir 'llm-cli-gateway.exe' $Checksums = Join-Path $InstallDir 'SHA256SUMS' $ChecksumBundle = Join-Path $InstallDir 'SHA256SUMS.sigstore.json' New-Item -ItemType Directory -Force $InstallDir | Out-Null Invoke-WebRequest -UseBasicParsing "$Base/$ExeName" -OutFile $Exe Invoke-WebRequest -UseBasicParsing "$Base/SHA256SUMS" -OutFile $Checksums Invoke-WebRequest -UseBasicParsing "$Base/SHA256SUMS.sigstore.json" -OutFile $ChecksumBundle cosign verify-blob $Checksums --bundle $ChecksumBundle --certificate-identity "https://github.com/verivus-oss/llm-cli-gateway/.github/workflows/release-installer.yml@refs/tags/v$Version" --certificate-oidc-issuer "https://token.actions.githubusercontent.com" if ($LASTEXITCODE -ne 0) { throw "Sigstore verification failed for SHA256SUMS" } function Get-ReleaseSha256($Name) { $line = Select-String -Path $Checksums -Pattern "^](https://github.com/verivus-oss/llm-cli-gateway/blob/HEAD/setup/providers/)[a-fA-F0-9]{64}\s+$([regex]::Escape($Name))$" | Select-Object -First 1 if (-not $line) { throw "No SHA256SUMS entry found for $Name" } return (($line.Line -split "\s+")[0]).ToLowerInvariant() } if ((Get-FileHash $Exe -Algorithm SHA256).Hash.ToLowerInvariant() -ne (Get-ReleaseSha256 $ExeName)) { throw "Checksum mismatch for $ExeName" } $env:RVWR_GATEWAY_BUNDLE_URL = "$Base/$BundleName" $env:RVWR_GATEWAY_BUNDLE_SHA256 = Get-ReleaseSha256 $BundleName & $Exe setup & $Exe stop & $Exe install-bundle & $Exe start & $Exe status & $Exe doctor
The Windows installer keeps a stablellm-cli-gateway.execommand in%LOCALAPPDATA%\Programs\llm-cli-gatewayand adds that directory to the user PATH. Do not script against release-versioned exe names after install.
# After downloading the binary that matches your OS/arch from a release: cosign verify-blob SHA256SUMS --bundle SHA256SUMS.sigstore.json \ --certificate-identity "https://github.com/verivus-oss/llm-cli-gateway/.github/workflows/release-installer.yml@refs/tags/v<version>" \ --certificate-oidc-issuer "https://token.actions.githubusercontent.com" sha256sum --check SHA256SUMS # verify before run (or shasum -a 256 --check on macOS) chmod +x llm-cli-gateway-<ver>-<os>-<arch> ./llm-cli-gateway-<ver>-<os>-<arch> setup ./llm-cli-gateway-<ver>-<os>-<arch> install-bundle # uses the platform bundle URL/SHA256 ./llm-cli-gateway-<ver>-<os>-<arch> start ./llm-cli-gateway-<ver>-<os>-<arch> doctor # Upgrade: replace the binary, set the new bundle env vars, run upgrade. ./llm-cli-gateway-<new>-<os>-<arch> upgrade # Uninstall: dry-run first, then run with --yes. ./llm-cli-gateway-<ver>-<os>-<arch> uninstall ./llm-cli-gateway-<ver>-<os>-<arch> uninstall --yes
LLM_GATEWAY_AUTH_TOKEN=$(openssl rand -hex 32) \ docker compose -f docker/personal.compose.yml up -d docker compose -f docker/personal.compose.yml run --rm doctor
- Multi-LLM Orchestration: Unified interface for Claude Code, Codex, Gemini, Grok, Mistral (Vibe), Devin, and Cursor Agent CLIs
- Session Management: Track gateway session metadata and provider-specific continuity with persistent storage
- Gateway-owned worktrees: Run supported sync or async provider requests inside a managed git worktree with the local file-backed session manager. Same-session reuse requires same-host durable ownership plus a matching live Git registration and gateway branch; manager-level named path collisions fail closed. PostgreSQL-backed sessions reject worktrees before creation because a different database-connected host cannot safely own filesystem cleanup. Grok, Devin, and Mistral require an explicit provider-nativesessionId; fresh,createNewSession, andresumeLatest-only worktree requests fail closed. A worktree requires a registered workspace selected explicitly, through caller-owned session metadata, or by the configured default; it never inherits process cwd or combines withworkingDir,addDir, orincludeDirs. Materialization suppresses repository, system, and global Git hooks, configured clean, smudge, and process checkout filters, sparse checkout, and lazy object fetching. Filter-dependent content such as Git LFS remains in its repository representation instead of executing host commands. Session deletion and TTL eviction hide a durably owned worktree session while cleanup runs. If Git removal fails, the file-backed store retains a durable cleanup-pending tombstone, blocks reuse, and retries cleanup when that store is registered on the owning host. The record is finalized only after verified Git removal.
- Token Optimization: Automatic 44% reduction on prompts, 37% on responses (opt-in)
- Correlation ID Tracking: Full request tracing across all LLM interactions
- Cross-Tool Collaboration: LLMs can use each other via MCP (validated through dogfooding)
- SQLite Flight Recorder: Every request/response logged to~/.llm-cli-gateway/logs.dbwith correlation IDs, token usage, duration, retry counts, and circuit breaker state. Browse withDatasette:datasette ~/.llm-cli-gateway/logs.db
- Structured Metadata: Tool responses include machine-readablestructuredContent(model, cli, correlationId, sessionId, durationMs, token counts)
- Cache observability resources:cache-state://global,cache-state://session/{id}, andcache-state://prefix/{hash}MCP resources return aggregate cache hit/miss/savings — tokens and hashes only, no prompt text.session_getincludes acacheStateblock when the session has prior requests.
- Provider capability inventory:provider_tool_capabilitiesandprovider-tools://catalogexpose the gateway request fields, supported/degraded provider controls, local skill/tool discovery, and safe config-surface hints for Claude Code, Codex CLI, Gemini/Antigravity, Grok CLI/API, Mistral Vibe, Cognition Devin, and Cursor Agent.doctor --jsonincludes a compactprovider_capabilitiessummary for setup assistants.
Every_requestand_request_asynctool exceptdevin_request/devin_request_asyncandcursor_request/cursor_request_asyncaccepts an optionalpromptPartsfield that structures the prompt for better cache hit rates (the Devin and Cursor headless paths take a plainpromptonly). The gateway concatenates the parts in canonical order (system → tools → context → task) so that the stable prefix bytes precede the volatile task tail unchanged across calls, letting each provider's automatic prompt-caching land on the same content hash each time.
{ "promptParts": { "system": "You are a helpful code reviewer.", "tools": "You have access to Read, Grep, Bash.", "context": "<long stable context block — file dumps, etc.>", "task": "Review the changes in src/foo.ts for security issues." } }
promptandpromptPartsare mutually exclusive — pass exactly one.
Per-CLI capability matrix (prefix discipline is automatic viapromptPartsfor all providers except Devin and Cursor, which have nopromptPartssurface; explicit levers are provider-specific):
claude_request({ promptParts: { system: "You are a helpful code reviewer.", context: "<long stable file dump>", task: "Review the diff.", cacheControl: { system: true, context: true }, // task is never marked }, outputFormat: "stream-json", });
Gateway emits thestream-jsonstdin path withcache_control: {type:"ephemeral", ttl:"1h"}on marked blocks only.
grok_request({ promptParts: { system: "...", context: "...", task: "..." }, compactionMode: "segments", compactionDetail: "balanced", });
Emits--compaction-mode segments --compaction-detail balanced.
Seedocs/personal-mcp/PROVIDER_CACHE_SURFACES.mdfor full surfaces, telemetry differences (e.g. Grok-pvs ACP), exact stream-json payload shapes, and cross-LLM review notes.
Opt-in flags (all default off) live under[cache_awareness]in~/.llm-cli-gateway/config.toml.
- Retry Logic: Exponential backoff with circuit breaker for transient failures
- Atomic File Writes: Process-specific temp files with fsync for data integrity
- Host-protection backpressure: bounded HTTP session lifecycle (max sessions + idle reaper), global and per-provider job-execution limits with a bounded FIFO queue, and a configurable per-job output cap (default 50MB). SeeHost-protection limits.
- NVM Path Caching: Eliminates I/O overhead on every request
- Long-Running Jobs: Non-time-bound async execution via_request_async+ polling tools
- Comprehensive Testing: 1,700+ tests covering unit, integration, and regression scenarios with real CLI execution
- Input Validation: Zod schemas prevent injection attacks
- No Secret Leakage: Generic session descriptions only (file permissions 0o600)
- No ReDoS: Bounded regex patterns prevent catastrophic backtracking
- Type Safety: Strict TypeScript with comprehensive error handling
- Supply-chain hardening: a dedicated.github/workflows/security.ymlruns actionlint, zizmor, shellcheck, typos, osv-scanner, gitleaks, and lychee on every push and PR (seeSECURITY.mdfor the threat model)
Every provider is reachable through the same request, session, job, and validation machinery, but the underlying CLIs differ in what they natively expose. The table records what actually shipped per provider; discover the live surface at runtime withprovider_tool_capabilities,list_models, and theprovider-acp://<provider>/provider-tools://<provider>resources.
- Native ACPis reported honestly.grok,mistral,devin, andcursorexpose a native ACP entrypoint, soprovider-acp://<provider>carries the negotiatedinitializecapability set and the derived session-method availability, and the sync_requestacceptstransport: "acp"(fails closed unless[acp]and the provider'sruntime_enabledgate are set). ACP routing is sync-only: the_request_asyncvariants always run the CLI transport and do not accepttransport: "acp"(nor Devin'sagentType); async ACP parity is a later phase. ACP workspace selection is gateway-owned: an explicit ACPworkspacemust be a registered alias. A fresh remote ACP request uses that alias or[workspaces].default; a remote resume is fixed to its recorded canonical alias and cwd, and a different or unbound workspace is rejected. Local ACP may omitworkspace; each unscoped process then gets a fresh private0o700neutral directory that is removed after the process exits, never a shared predictable temp path.claude,codex, andgeminihave no native ACP entrypoint at their target CLI versions; theirprovider-acp://records reportnative: falsewith no methods and no adapter-as-native masquerade, and they expose notransport: "acp"selector.
- Managed approval is Claude-only today.approvalStrategy:"mcp_managed"is executable only by the Claude CLI adapter, which launches Claude with a request-scoped generated MCP configuration and--strict-mcp-config. It permits only provisioned, gateway-owned MCP definitions and rejects dynamicnpx, ambient-PATH, and Codex-config overrides. Codex, Gemini, Grok, Mistral, Devin, and Cursor rejectmcp_managedbefore launching a provider because their current adapters cannot isolate ambient MCP configuration. For those adapters, useapprovalStrategy:"legacy";approvalPolicyhas no effect.
- ACP has its own permission bridge.approvalStrategy:"mcp_managed"and anyapprovalPolicyare rejected whentransport:"acp"is selected. Use the Claude CLI transport for managed approval. ACP host services fail closed: reads are unavailable by default, and write or terminal callbacks need[acp]host-service configuration plus a one-time ApprovalManager decision. ACP never maps a raw CLI bypass input to a standing permission grant.
- Resourcesare generated from the provider registry for every CLI provider:models://<provider>,sessions://<provider>,provider-acp://<provider>,provider-tools://<provider>, andprovider-subcommands://<provider>.
- Model discoveryis live and account-aware: the discovery listed above reachesmodels://<provider>andlist_models, degrading to static registry facts when a live probe is unavailable (a resource read never spawns a CLI).
- Admin surfacesare discovery-driven and output-redacted.provider_admin_listandprovider_admin_runare read-only for every provider. State-mutating admin operations are exposed only throughprovider_admin_mutate, gated behind[admin] allow_mutating_cli_admin_ops, the remotecli:adminscope, an approval gate, and an audit record. Mutating ACP session operations are likewise gated behind[acp] allow_mutating_session_ops.
- Validationcommands work across every provider:review_changes,validate_with_models,second_opinion,compare_answers,red_team_review,consensus_check,ask_model, andsynthesize_validation, with canonically hashed immutable receipts viavalidation_receiptand thevalidation-receipt://{validationId}resource.review_changescaptures a complete, hashed Git artifact and starts repository-bound read-only reviewers.
Node.js >= 24.4.0is required (engines.nodeinpackage.json). The gateway uses Node's built-innode:sqlitemodule for persistence — there is no native binding to compile and no install scripts run. The 24.4 floor is whereallowBareNamedParametersdefaults totrue, which the persistence layer relies on.
Before using this gateway, you need to install the CLI tools you want to use:
# Installation instructions for Claude Code # Visit: https://docs.anthropic.com/claude-code npm install -g @anthropic-ai/claude-code
npm install -g @openai/codex codex login
The Gemini provider runs through Google Antigravity CLI (agy).
curl -fsSL https://antigravity.google/cli/install.sh | bash # Docs: https://antigravity.google/docs
curl -fsSL https://x.ai/cli/install.sh | bash grok login # OAuth flow; for headless auth, set XAI_API_KEY # Docs: https://docs.x.ai/build/overview
# Pick one — the gateway's cli_upgrade auto-detects which one you used. curl -LsSf https://mistral.ai/vibe/install.sh | bash pip install mistral-vibe uv tool install mistral-vibe brew install mistral-vibe vibe --setup # Complete the API-key setup locally. Do not paste the key into a chat. # Current Vibe defaults session logging to enabled. If an older config disabled it, # edit ~/.vibe/config.toml and set: # [session_logging] # enabled = true
- Model selection is via theVIBE_ACTIVE_MODELenvironment variable— Vibe has no--modelflag. The gateway discovers~/.vibe/config.toml/VIBE_MODELS, injectsVIBE_ACTIVE_MODELonly when a model is explicitly requested or Vibe config needs recovery, and retries once after a model-not-found failure with refreshed discovery.
- permissionModeis the Vibe--agentname.Builtins aredefault | plan | accept-edits | auto-approve; Vibe also accepts install-gated builtins (e.g.lean) and custom agents from~/.vibe/agents. Requests pass the selected name through for Vibe to validate.mcp_managedis not available for Vibe.
- Tool controls use Vibe's native flags.The gateway emits one--enabled-tools <tool>flag perallowedToolsentry and one--disabled-tools <tool>flag perdisallowedToolsentry. Vibe applies disabled tools after enabled-tool filtering.
- Usage telemetry is best-effort.Vibe does not emit token or cost data in programmatic stdout. When the gateway knows Vibe's native session UUID, it reads~/.vibe/logs/session/session_<...>/meta.jsonfor usage and cost. A missing, malformed, or not-yet-known session log leaves those fields empty.
- No self-update:cli_upgrade --cli mistraldetects whether you used pip / uv / brew and dispatches the matching upgrade command. Runningvibe updateis not a thing.
{ "mcpServers": { "llm-gateway": { "command": "npx", "args": ["-y", "llm-cli-gateway"] } } }
git clone https://github.com/verivus-oss/llm-cli-gateway.git cd llm-cli-gateway npm install npm run build
For clients that already support local stdio MCP servers, add a configuration like:
{ "mcpServers": { "llm-cli-gateway": { "command": "node", "args": ["/path/to/llm-cli-gateway/dist/index.js"] } } }
Stdio is the recommended path for unrestricted machine-local development access. HTTP MCP, including localhost HTTP and tunneled HTTPS, is treated as remote-capable for provider execution: provider tools must resolve a registered workspace alias, a session workspace, or[workspaces].defaultbefore spawning a CLI. Remote clients should pass relativeworkingDir,addDir, and include-directory values inside the selected workspace, and may resume only gateway-tracked sessions they own. Raw native provider session IDs are local-only. Disabling auth or using a no-auth connector path is not a filesystem bypass.
For a local CLI request with no resolvedworkingDir, registeredworkspace, or gateway-managedworktree, the child runs in a fresh private0o700temporary directory that is removed after the process exits. It never inherits the gateway repository cwd or its provider-native instruction context. The gateway canonicalizes the temp root and rejects or relocates it when any ancestor contains.git,AGENTS.md,AGENTS.override.md,Agents.md,AGENT.md,CLAUDE.md,Claude.md,CLAUDE.local.md,.claude/CLAUDE.md,.claude/rules/,.cursor/rules/,.cursorrules,GEMINI.md, or.vibe/config.toml, including through a symlinked or customTMPDIRbeneath that context. The list covers entries a provider discovers by walking up from its cwd. User-scope configuration such as~/.claude/settings.jsonis deliberately absent: it loads on every invocation regardless of cwd, so relocating the workspace would not isolate it. Provider-nativeresumeLatestoperations that use a cwd-scoped latest-session pointer therefore require an explicitworkingDir,workspace, or configured default workspace and fail closed when none is available. Use explicit target selection whenever several repositories are active at once.
CLI request schemas accept prompts up to 100,000 characters, but operating systems also impose byte limits on individual argv elements. Codex new and resume requests stream the exact prompt over stdin.codex_fork_sessionremains argv-bound and rejects an oversized UTF-8 prompt before spawn as non-retryableinput_too_large. Other providers whose current CLI contracts require an argv prompt use the same admission rule. Every other caller-controlled argv value is checked on its final encoded form too, including serialized agent/schema JSON, joined tool lists, instruction overrides, paths, model names, and native session IDs. The final spawn boundary checks every argv element plus the aggregate resolved command line against a conservative platform-specific byte budget and a 2,048-element cap. The aggregate byte budget excludes the environment but reserves headroom for it; on Windows, pre-resolution admission assumes the smaller npm.cmd/.batwrapper limit until command resolution proves a native executable. Native session and resume flags on non-Kit requests are included before workspace, session, provider-artifact handoff, or durable job side effects. Claude Kit projects its eventual argv before materializing its compiled context artifact or allocating a durable Kit session. An embedded NUL byte in the command or any argv element is rejected before spawn as non-retryableinvalid_input. Caller-facing results, long-lived job memory, durable job args, and async flight rows use a fixed invalid-argv marker; the optional duplicate durable payload is suppressed. None retains the rejected vector or Node's value-echoing native error. NativeE2BIG, including an environment-driven failure, is normalized without retaining the nativespawnargs. The gateway never truncates instructions or other values to make them fit. For stdin-backed requests, a clean provider exit is accepted only after the complete payload write callback succeeds. A closed or still-pending pipe becomes a fixed, non-sensitive incomplete-delivery failure; timeout, cancellation, and provider nonzero exits remain authoritative.
This generic stdio example is not provider-support verification for the Personal MCP Appliance. Client-specific setup guides for ChatGPT, Claude web, Claude Desktop, Codex, Gemini CLI, Gemini web, and Grok remain gated by the provider-support matrix indocs/personal-mcp/PRODUCT_CONTRACT.md.
The personal-appliance surface exposes simplified validation tools for non-developer clients. These tools start provider CLI jobs through the durable async job manager and return normalized provider status plus raw job references.
- validate_with_models: ask two or more providers to independently validate a question.
- review_changes: capture one complete Git review artifact, fence repository content as untrusted data, and start read-only independent reviewers. SeeRepository change review.
- second_opinion: ask one provider to review an answer.
- red_team_review: challenge a plan, answer, or document for risks and failure modes.
- consensus_check: check whether providers agree with a claim.
- ask_model: ask one provider through the simplified surface.
- synthesize_validation: run an explicit judge model after provider results have been collected. General validation requires the caller's question and terminal normalized results. Areview_changesrun instead reloads its exact owned durable results fromvalidationId; caller-supplied question/results are ignored.
- list_available_models: list the models each provider CLI exposes through the simplified surface.
- job_statusandjob_result: poll and collect validation job outputs.
- validation_receipt: retrieve the canonically hashed immutable receipt of a terminal cross-LLM validation run byvalidationId(returnsminted | pending | verification_failed | expired_unminted | not_found, own-or-not-found).verification_failedmeans a stored receipt exists but disagrees with its durable run, which is a defect to investigate;expired_unmintedonly ever means absence.format: "markdown"renders a human-readable report;includeRawResponsesinlines complete provider answer text when the linked job still exposes identity-verified output. Registered only when the attached job store provides the durable validation-run store capability (sqliteandpostgres).
The same receipt is also exposed as thevalidation-receipt://{validationId}MCP resource (same durable gate and own-or-not-found owner scoping).
The validation report preserves per-provider disagreement. Optional judge synthesis is explicit about which provider produced the judge job.
review_changesaccepts an absolute localworkingDiror a registeredworkspace, then resolvesscope: "auto" | "uncommitted" | "branch" | "commit". It can take an explicit Gitbase, literal repository-relativepaths,stance: "standard" | "adversarial", reviewermodels, an optionaljudgeModel,trustCursorWorkspace, and fail-closed artifact/prompt byte ceilings.
Cursor refuses to review a directory it does not trust, so a cursor seat is granted--trustonly when the reviewed directory is a registered[[workspaces.repos]]path whoseprovidersinclude cursor (a gateway worktree beneath one counts, and the nearest enclosing registration decides). On any other directory the cursor seat is skipped with an actionable reason, and the rest of the roster still runs.trustCursorWorkspace: trueaccepts the grant for one durable review run. The consent applies to Cursor seats in that run's reviewer roster and to its planned Cursor judge whensynthesize_validationlaunches it later. It is bound to the run owner, repository, and planned judge, so another principal or review run cannot replay it. That is a real decision rather than a formality: a trusted folder is also where cursor loads project rules andAGENTS.md, so the repository under review gains some influence over its own reviewer, which the fenced review prompt otherwise forbids. The artifact keeps committed, staged, unstaged, and regular non-ignored untracked file evidence separate. It forces tracked diffs to remain readable even when in-tree attributes mark them as non-diffable. Thereview-evidence.v2artifact exposescommittedPatch,stagedPatch, andunstagedPatchindependently; each segment carries its sorted path inventory, encoding, exact byte length, SHA-256 identity, and content. This prevents an index change and its worktree-only reversal from canceling out. The artifact is collision-fenced, byte-counted, SHA-256 identified, race-checked, and never truncated. Inautomode, a diverged branch is reviewed from its merge base with working-tree evidence included. Otherwise, a dirty tree selects uncommitted changes, while a clean tree falls back to the last commit (HEAD^..HEAD) without working-tree evidence. Unsafe untracked file types or a repository mutation during capture cause a refusal.
The tool starts asynchronous provider jobs and returns avalidationId, exact artifact and prompt identities, file inventory, and onerawJobReferenceper reviewer. Poll those references with validationjob_statusand collect them with validationjob_result, not the similarly namedllm_job_tools. If a judge was requested, wait for every reviewer to become terminal, then callsynthesize_validationwith thevalidationIdand the sameworkingDirorworkspaceselector. Continue collecting results for progress and human visibility, but do not pass them as review evidence: for areview_changesrun, the gateway ignores caller-suppliedquestionandproviderResults, reloads the exact owned durable linked terminal jobs, and reconstructs requested but unavailable seats as skipped. General validation synthesis still requires a caller-supplied question and terminal normalized results.
The review surface is registered only with durable SQLite or PostgreSQL job and validation-run storage. Each CLI review job retains the exact fenced prompt in its expiry-boundpayload_json; its persisted argv contains only a hash marker. The non-expiring flight recorder does not receive repository-review prompts. Configured HTTP/API reviewer seats require explicitallowApiUpload:truebecause the complete artifact leaves the local CLI boundary. Remote HTTP/OAuth workspace reviews reject API reviewer uploads even with that flag. Treat the durable job store as sensitive until the configured job retention expires. WhenjudgeModelis an HTTP/API provider,review_changesbinds that explicit consent, the judge provider, the resolved repository, and the caller identity to the durablevalidationId. The latersynthesize_validationcall must provide that id and the same repository selector. The stored judge, repository, owner, and upload consent are authoritative. The gateway atomically claims the planned judge once, so concurrent or repeated synthesis cannot start a second judge. A follow-up argument cannot grant or override upload consent.
Execute a Claude Code request with optional session management.
- prompt(string, optional): The prompt to send (1-100,000 chars). Exactly one ofpromptorpromptPartsis required (mutually exclusive)
- model(string, optional): Model name or alias (uselist_modelsfor available values; supportslatest)
- outputFormat(string, optional): Output format (text|json|stream-json), default:stream-json— the gateway parses NDJSON usage events for token/cost observability; override totextonly when you want unparsed stdout
- sessionId(string, optional): Specific session ID to use. Undermcp_managed, native continuation is a high-risk input because it can inherit an unverified provider posture; it requires approval andLLM_GATEWAY_APPROVAL_ALLOW_BYPASS=1, but does not select a full-permission profile.
- continueSession(boolean, optional): Continue the active session. It has the same managed-approval requirement assessionId. Because Claude--continueselects by cwd, it requiresworkingDiror a registered workspace selected explicitly, through caller-owned session metadata, or by the configured default. That workspace may optionally supply a gateway worktree. The request fails closed when no selection supplies a stable cwd.
- createNewSession(boolean, optional): Always create a new session
- forkSession(boolean, optional): Fork the resumed session instead of appending to it. Undermcp_managed, it is a high-risk native-fork input that requires approval andLLM_GATEWAY_APPROVAL_ALLOW_BYPASS=1, but stays bounded.
- allowedTools(string[], optional): Restrict Claude tools to this allow-list. A non-empty allow-list is a high-risk managed input because it can change the tool posture.
- disallowedTools(string[], optional): Explicitly deny listed Claude tools
- permissionMode(string, optional): Claude permission mode (default|acceptEdits|plan|auto|dontAsk|bypassPermissions); preferred overdangerouslySkipPermissions.bypassPermissionsis a direct full-permission request undermcp_managed.
- dangerouslySkipPermissions(boolean, optional): Deprecated, maps topermissionMode: "bypassPermissions";permissionModewins when both are set. It is a direct full-permission request undermcp_managed.
- agent(string, optional): Named sub-agent to run as. A non-empty value is a high-risk managed input because it can change tool and permission posture.
- agents(string, optional): Inline agent definitions JSON. A non-empty value is a high-risk managed input for the same reason.
- systemPrompt/appendSystemPrompt(string, optional): Replace or extend the system prompt. A non-empty value is a high-risk managed input.
- systemPromptFile/appendSystemPromptFile(string, optional): Replace or extend the system prompt from a file. A non-empty file path is a high-risk managed input.
- safeMode(boolean, optional): Start Claude with local customizations disabled, includingCLAUDE.md, skills, plugins, hooks, MCP, commands, and agents.trueis a high-risk managed input.
- bare(boolean, optional): Start Claude in minimal mode, skipping local customization discovery.trueis a high-risk managed input.
- debugFile(string, optional): Write Claude debug output to a file. A non-empty path is a high-risk managed input.
- maxBudgetUsd(number, optional): Budget cap in USD for the request
- maxTurns(integer, optional): Agent-loop turn cap
- effort(string, optional): Reasoning effort (low|medium|high|xhigh|max)
- fallbackModel(string, optional): Auto-fallback model when the default is overloaded
- jsonSchema(string, optional): JSON Schema literal constraining structured output
- addDir(string[], optional): Additional workspace directories. A non-empty value is a high-risk managed input.
- noSessionPersistence(boolean, optional): Ephemeral session (not persisted to disk)
- settingSources/settings/tools(optional): Setting sources to load, settings JSON path/literal, built-in tool restriction. Non-empty setting sources, settings, or tool selections are high-risk managed inputs.
- pluginDir/pluginUrl(string[], optional): Load Claude plugins from local directories or URLs. Non-empty values are high-risk managed inputs.
- excludeDynamicSystemPromptSections(boolean, optional): Trim dynamic system prompt sections
- approvalStrategy(string, optional):"legacy"(default) or"mcp_managed". Managed mode usesacceptEditsby default and forcesstrictMcpConfig:true, so Claude uses only the gateway-generated MCP configuration. A direct full-permission request requires all of an explicit caller request, an approval-manager approval, andLLM_GATEWAY_APPROVAL_ALLOW_BYPASS=1. Other high-risk inputs require the approval and operator setting too, but remain bounded and do not themselves select full permission.
- approvalPolicy(string, optional):"strict","balanced", or"permissive"
- mcpServers(string[], optional): Names of MCP servers to expose to Claude (default: none). Legacy requests resolve names from the local registry or Codex MCP config; unknown names are reported as unavailable. Undermcp_managed, Claude uses only the generated configuration and only registry entries explicitly provisioned as gateway-owned local commands are eligible. Dynamicnpxlaunchers, ambient-PATH commands, and Codex-config overrides are rejected. Configure and deploy the managed entries in the gateway environment.
- strictMcpConfig(boolean, optional): In legacy mode this defaults tofalse; settrueto require only the generated MCP config and fail if requested servers are unavailable. Undermcp_managed, the gateway forces it totrueand a caller-suppliedfalsecannot weaken that boundary.
- optimizePrompt(boolean, optional): Optimize prompt for token efficiency (44% reduction), default: false
- optimizeResponse(boolean, optional): Optimize response for token efficiency (37% reduction), default: false
- correlationId(string, optional): Request trace ID (auto-generated if omitted)
- idleTimeoutMs(integer, optional): Kill a stuck process after output inactivity; 30,000 to 3,600,000 ms. Idle enforcement applies only whenoutputFormatisstream-json; it is ignored for text/json, which produce no output until the run completes
- worktree(boolean|object, optional): Run inside a gateway-owned git worktree (slice λ). A worktree requires a registered workspace selected explicitly, through caller-owned session metadata, or by the configured default; it never inherits process cwd or combines withworkingDir,addDir, orincludeDirs. Materialization suppresses repository, system, and global Git hooks and configured clean, smudge, and process checkout filters, so filter-dependent content such as Git LFS remains in its repository representation instead of executing host commands. Requesting a worktree is a high-risk managed input that requires approval andLLM_GATEWAY_APPROVAL_ALLOW_BYPASS=1, but remains bounded.
- promptParts(object, optional): Cache-aware structured prompt{ system?, tools?, context?, task }; mutually exclusive withprompt
- forceRefresh(boolean, optional): Bypass dedup and force a fresh CLI run, default: false
Workspace boundary: stdio callers may use machine-local paths directly. HTTP/tunnel callers must passworkspaceor rely on a configured default/session workspace; path fields are then validated relative to that workspace.[workspaces].allow_unregistered_working_diris an inert legacy key: it is still accepted so old configs keep loading, but nothing reads it at either value, and setting it now logs a warning at startup. It never allowed arbitrary HTTP working directories or additional directories.
- approval: Approval decision record whenapprovalStrategy="mcp_managed"
- mcpServers: Requested/enabled/missing MCP servers for this call
{ "prompt": "Write a Python function to calculate fibonacci numbers", "model": "sonnet", "continueSession": true, "optimizePrompt": true, "optimizeResponse": true }
Execute a Codex request with optional session tracking.
- prompt(string, optional): The prompt to send (1-100,000 chars). Exactly one ofpromptorpromptPartsis required (mutually exclusive)
- model(string, optional): Model name or alias (uselist_modelsfor available values; supportslatest, recommended:gpt-5.5)
- fullAuto(boolean, optional): Deprecated — expands to--sandbox workspace-writeonly (current Codex no longer accepts approval-policy flags); prefersandboxMode
- sandboxMode(string, optional): Codex sandbox (read-only|workspace-write|danger-full-access).
- dangerouslyBypassApprovalsAndSandbox(boolean, optional): Request Codex's full approvals-and-sandbox bypass.
- dangerouslyBypassHookTrust(boolean, optional): Request Codex hook-trust bypass.
- approvalStrategy(string, optional):"legacy"is the only executable strategy."mcp_managed"is rejected before Codex launches because the adapter cannot isolate ambient MCP configuration.
- approvalPolicy(string, optional): Has no effect for Codex becausemcp_managedis unavailable.
- mcpServers(string[], optional): Metadata only. It does not configure or isolate Codex MCP servers.
- sessionId(string, optional): Session identifier for tracking.
- resumeLatest(boolean, optional): Resume a previous Codex session (codex exec resume --last). Do not rely on which session--lastselects or on the resumed working directory (#258):--lastis cwd-filtered upstream and the child is still spawned with the gateway-resolved cwd. Verify the target, or start a fresh session when it must be certain. Ignored ifsessionIdis set.
- createNewSession(boolean, optional): Always create a new session
- forceRefresh(boolean, optional): Bypass dedup and force a fresh CLI run, default: false
- outputFormat(string, optional):text(default) orjson(--jsonJSONL events for token usage extraction)
- outputSchema(string|object, optional): Codex--output-schema, path or inline JSON Schema.
- workingDir(string, optional): Working root for this session (-C/--cd; new sessions only). Personal Agent Config Kit mode requires an absolute path.
- addDir(string[], optional): Additional writable workspace directories (one--add-dirper entry; new sessions only).
- ephemeral(boolean, optional): Codex--ephemeral(no session persistence)
- images(string[], optional): Image attachments (one-i <path>per entry).
- profile(string, optional): Codex--profile <name>(new sessions only; ignored with a logged warning on resume).
- configOverrides(object, optional): Codex-c key=valueoverrides. Local callers only; remote HTTP/OAuth requests are rejected.
- enable/disable(string[], optional): Codex--enable/--disablefeature overrides. They are-c features.*equivalents and are also local-only.
- ignoreRules/ignoreUserConfig(boolean, optional): Codex--ignore-rules/--ignore-user-config.
- outputLastMessage(string, optional): Codex--output-last-message <path>.
- oss(boolean, optional): Codex--oss, selecting the open-source provider.
- localProvider(string, optional): Codex--local-provider <name>.
- worktree(boolean|object, optional): Run inside a gateway-owned git worktree (slice λ). A worktree requires a registered workspace selected explicitly, through caller-owned session metadata, or by the configured default; it never inherits process cwd or combines withworkingDir,addDir, orincludeDirs. Materialization suppresses repository, system, and global Git hooks and configured clean, smudge, and process checkout filters, so filter-dependent content such as Git LFS remains in its repository representation instead of executing host commands.
- promptParts(object, optional): Cache-aware structured prompt{ system?, tools?, context?, task }; mutually exclusive withprompt
- optimizePrompt(boolean, optional): Optimize prompt for token efficiency, default: false
- optimizeResponse(boolean, optional): Optimize response for token efficiency, default: false
- correlationId(string, optional): Request trace ID (auto-generated if omitted)
- idleTimeoutMs(integer, optional): Kill a stuck Codex process after output inactivity; 30,000 to 3,600,000 ms
Claude Desktop / Cursor
Paste into your MCP client config file to install this server.
{
"mcpServers": {
"llm-cli-gateway": {
"server": {
"command": "npx",
"args": [
"-y",
"llm-cli-gateway"
]
}
}
}
}
McpServers
{
"server": {
"command": "npx",
"args": [
"-y",
"llm-cli-gateway"
]
}
}
Transport
"stdio"
Package
"llm-cli-gateway"
Registry
"npm"
Unified MCP server providing access to Claude Code, Codex, and Gemini CLIs through a single gateway. Features multi-LLM orchestration, persistent session management, async job execution with polling, approval gates, retry with circuit breakers, and token optimization. Install: npx -y llm-cli-gateway
"Without consultation, plans are frustrated, but with many counselors they succeed."— Proverbs 15:22 (LSB)
Secure local control plane for AI coding agents.
llm-cli-gatewaylets supported MCP clients operate Claude Code, Codex, Gemini/Antigravity, Grok Build, Mistral Vibe, Cognition Devin, Cursor Agent, and configured HTTP API providers through one user-owned gateway while preserving native CLI sessions, local credentials, durable async jobs, validation receipts, and review workflows.
Why developers try it:use the client you are already in to delegate work to local coding agents, scope remote execution to registered workspaces, gate risky actions, survive disconnects, and collect auditable review evidence without turning those agents into a generic chat proxy.
Current signals:CI and security workflows pass onmain, OpenSSF Scorecard is published, OpenSSF Best Practices is passing, releases use Sigstore signing, and the package is MIT licensed.
Or use directly withnpxfrom an MCP client:
{ "mcpServers": { "llm-gateway": { "command": "npx", "args": ["-y", "llm-cli-gateway"] } } }
llm-cli-gatewayis a single-user MCP control plane for operating AI coding agents from supported local or remote clients. It is more than a thin CLI wrapper:
- Runs registered provider CLIs and configured HTTP API providers through consistent sync and async MCP tools.
- Persists long-running jobs, supports restart-safe result collection, deduplication, cancellation, and sync-to-async deferral.
- Tracks sessions, real CLI resume paths, structured response metadata, and cache telemetry.
- Supports cache-awarepromptParts, including explicit Claudecache_controlwhen opted in.
- Can run supported provider requests inside gateway-managed git worktrees for isolated multi-agent review and implementation loops when using the local file-backed session manager. PostgreSQL-backed sessions reject this filesystem-local feature before creation. Grok, Devin, and Mistral require an explicit provider-nativesessionIdfor a gateway worktree; fresh,createNewSession, andresumeLatest-only worktree requests fail closed because they cannot durably reselect it. Materialization suppresses repository, system, and global Git hooks, configured clean, smudge, and process checkout filters, sparse checkout, and lazy object fetching. Filter-dependent content such as Git LFS remains in its repository representation instead of executing host commands.
- Ships personal-appliance setup surfaces: HTTP transport with bearer-token auth,doctor --json, setup UI artifacts, provider setup snippets, Docker fallback, and checked release bundles.
- Remote web connectors use MCP OAuth discovery and authorization-code setup with static client or shared-secret gates. Client secrets are generated locally, stored only as hashes, and printed only by explicit copy-once commands.
- Provider CLI requests can select registered workspaces by alias viaworkspace; every HTTP/tunnel request must use a registered alias, session workspace, or[workspaces].defaultbefore provider execution. Local unrestricted filesystem access is the stdio transport.
The repo ships agent-ready workflow skills under.agents/skillsfor async orchestration, session continuity, multi-LLM review, implement-review-fix loops, retrospective evidence walks, secure approval-gated dispatch, and Personal Agent Config Kit operations. Nine caller-facing skills are bundled in the published npm package:async-job-orchestration,multi-llm-review,session-workflow,secure-orchestration,implement-review-fix,retrospective-walk,public-demo-session,least-cost-routing, andpersonal-agent-config-kit. Machine-readable DAG-TOML plans live underdocs/plansandsetup/install-plan.dag.tomlfor workflows that need deterministic sequencing and verification gates.
Skill packs can be updated outside the core npm release by placing skill directories in local, operator-controlled paths. The gateway loads bundled skills first, then[skills].paths, thenLLM_GATEWAY_SKILLS_PATH, then~/.llm-cli-gateway/skillswhen it exists; later roots override earlier skills by name. Each skill is a directory containingSKILL.md. A root may also carryskill-pack.jsonto pin expectedSKILL.mdhashes:
[skills] paths = ["/opt/llm-cli-gateway/skills"]
export LLM_GATEWAY_SKILLS_PATH="/opt/team-skill-pack:/opt/incident-skill-pack"
{ "name": "team-pack", "version": "1.0.0", "skills": [ { "name": "incident-retrospective", "sha256": "<sha256 of incident-retrospective/SKILL.md>" } ] }
The loader is intentionally local-only: it never fetches remote Markdown at startup. To update a pack, install or replace files through your package manager or deployment system, then restart the gateway so the advertisedskills://...resources refresh.
The next documentation focus is provider-specific skill and DAG-TOML pairs for each outbound CLI and API-provider family: Claude, Codex, Gemini/Antigravity, Grok, Mistral Vibe, Devin, Cursor Agent, OpenAI-compatible endpoints, Anthropic Messages, and xAI Responses. The implementation plan is tracked indocs/plans/provider-workflow-assets.dag.toml, with each provider asset expected to cover install/login checks or token-env checks, session behavior, approval modes, cache/telemetry surfaces, failure modes, and a smoke-test gate.
- CI runs build, lint, format, tests, package checks, and npm audit.
- Security CI runs actionlint, zizmor, shellcheck, typos, osv-scanner, gitleaks, and lychee.
- GitHub release installer artifacts are checksummed and signed with Sigstore keyless signing.
- npm releases use a generated prod-only shrinkwrap and release security audit; GitHub Actions Trusted Publishing exchanges the job's OIDC identity for short-lived npm publish credentials.
- The npm package intentionally ships a generated, prod-onlynpm-shrinkwrap.jsonso registry installs resolve the audited release tree. Release gates regenerate it frompackage-lock.json, compare for parity, and run a registry-fidelity consumer install before publishing.
- Socket behavioural alerts are documented insocket.ymland under "Security Considerations" below.shellAccessandshrinkwrapare reviewed package capabilities/configuration for this CLI appliance, not hidden install behaviour.
The personal-appliance contract keeps that surface intentionally narrow: one trusted user runs the gateway on a machine or volume they own, connects one MCP endpoint, and lets supported clients operate local coding agents through workspace-scoped, approval-gated, auditable requests.
The product contract is documented in[docs/personal-mcp/PRODUCT_CONTRACT.md. It defines the single-user scope, security posture, target support matrix, and provider-support verification gates. Public setup guides must not claim ChatGPT, Claude web, Claude Desktop, Codex, Gemini CLI, Gemini web, or Grok inbound support until the corresponding provider/client path has been verified.
This project does not provide hosted multi-tenant credential custody. Provider credentials stay on the user's machine or user-owned deployment volume.
…
Sign in to leave a review
Use Google, GitHub, or an email account so ratings stay tied to real people.
No reviews posted yet.



