ownvoice
About
MCP server wrapping the ownvoice CLI for voice/identity checks.
Explore
Setting up with Highlight
This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:
- Download and install Highlight from highlightai.com/download
- Navigate to the plugins tab and select "Add Custom Plugin"
-
Configure the plugin with the settings below
Plugin Name
ownvoiceCommand (node, npx, python, etc.)Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.
- Enable "Start Automatically" if you want the plugin to start when Highlight launches
From the repository
OwnVoice's own training run is real GPU time, honestly labeled and not hidden behind a fake progress bar, the same category norm Unsloth uses. What OwnVoice compresses to under two minutes is everythingbeforethat: confirming your environment actually works.
pocket-tts is a genuinely good, MIT-licensed, CPU-capable local text-to-speech model from Kyutai. Its own maintainers have been clear that fine-tuning code isn't coming any time soon: on](https://github.com/resemble-ai/Resemblyzer)issue #30, maintainer @vvolhejn wrote "We are not planning to release fine-tuning code for our TTS and STT models in the near future," and 18 people reacted to that thread asking for exactly this. OwnVoice is a small, standalone CLI that fills that specific gap: point it at a handful of your own voice recordings, and it trains a LoRA adapter you keep and run yourself.
It is not a hosted service, it has no billing, and it does not track usage. It is a training script, an inference script, and a scoring script, wired together behind three CLI commands.
pocket-tts already ships zero-shot voice cloning out of the box: pass a.wavfile to--voice(or callget_state_for_audio_prompt()from Python) and it clones that voice with no training step at all. If that is all you need, use pocket-tts directly, it is simpler and faster.
OwnVoice exists for a narrower case: baking a voice permanently into trained weights, so generation no longer depends on distributing or re-processing a reference audio clip at runtime, with (based on the training objective, not yet independently benchmarked at scale) more consistent output across many generations than a single-clip zero-shot embedding tends to produce. That is the specific gap the 18 reactors on issue #30 were describing, and it is the only thing OwnVoice adds on top of what pocket-tts already does well.
This tool clones a voice from audio you have the right to use. Do not clone someone else's voice, or a public figure's voice, without their explicit consent. OwnVoice ships no bulk-generation or auto-scaling feature in this version, keeping the blast radius of any single misuse case small.
This is a young, early-stage release.ownvoice check, the CLI argument parsing, voice-clip validation, the similarity scoring math, and the adapter/manifest save and load path are implemented and covered by the test suite (pytest). LoRA injection was verified structurally against pocket-tts's real source and then confirmed for real:ownvoice checkwas run against pocket-tts's actual downloaded weights, on CPU, and PEFT'starget_modules="all-linear"injection genuinely succeeded. The full training and generation path has since been verified end to end for real too: a real 2-epoch LoRA training run against loaded pocket-tts weights produced a finite, non-NaN flow-matching loss, and the resulting adapter produced a real, non-silent generated.wavfile viaownvoice infer. That validation surfaced two real gaps in the naive approach and fixed them: (1) pocket-tts's published, inference-only PyPI package does not actually expose a way to compute the training loss throughFlowLMModel.forward()despite its own docstring claiming otherwise, so OwnVoice computes the flow-matching loss directly fromflow_lm's real submodules instead; (2) swappingbase_model.flow_lmto the PEFT-wrapped model before callinggenerate_audio()breaks pocket-tts's internal KV-cache state lookup -- no swap is needed at all, since PEFT's LoRA injection already mutatesbase_model.flow_lmin place. One real, external limitation to know about: the publicly downloadable pocket-tts weights (kyutai/pocket-tts-without-voice-cloning) refuse a raw reference-clip path/URL outright; OwnVoice works around this by pre-loading and resampling the clip itself, but voice-cloning fidelity from that checkpoint is a known limitation of the base model, not an OwnVoice bug -- for kyutai's best-quality cloning weights, request gated access athuggingface.co/kyutai/pocket-tts. Runownvoice checkyourself and read the source before trusting any of it further, that is the right amount of skepticism for a project this early.
What is OwnVoice, and why not just use pocket-tts by itself?OwnVoice trains a LoRA adapter forpocket-ttsand saves it to your own disk asadapter.safetensorsplusmetadata.json. It exists because pocket-tts's own maintainers have said fine-tuning code is not on their near-term roadmap (seeissue #30). Once you have a trained adapter, you never need OwnVoice again to use it:ownvoice inferjust loads the adapter back onto the base model.
How is this different from pocket-tts's own built-in--voice <wav>zero-shot cloning?pocket-tts already clones a voice from a single reference clip with no training step,--voice <wav>at the CLI orget_state_for_audio_prompt()in Python. OwnVoice trades that speed for a permanently trained adapter, so generation no longer depends on carrying around a reference clip at runtime, with (based on the training objective, not yet independently benchmarked at scale) more consistent output across repeated generations than a single-clip zero-shot embedding tends to give. If zero-shot is enough for your use case, use pocket-tts directly, it is simpler and faster.
What do I need to install it, and does it run on Apple Silicon or a CPU-only machine?Python 3.11 or newer, thenpip install ownvoice-cli.ownvoice checkandownvoice inferneed no GPU at all and run fine on Apple Silicon or a CPU-only machine, matching pocket-tts's own CPU-capable design.ownvoice trainruns on CPU too, it just takes longer per epoch; install the CUDA build of PyTorch first if you have an NVIDIA GPU and want training to go faster.
How does OwnVoice compare to kokoro-tts and Unsloth?kokoro-ttsgets you synthesizing speech in under 2 minutes but has no fine-tuning step at all.Unslothgets a training run started in under a minute but is a general LLM fine-tuning framework, not TTS-specific. OwnVoice is narrower than either: one base model (pocket-tts only), one job (a voice adapter), plus a freeownvoice checkstep that confirms PEFT's LoRA injection actually works against your environment before you spend anything on a GPU, a check neither of those tools has an equivalent of.
My training run finished but printed "BELOW THRESHOLD", is that a bug?No. It is a labeled outcome, not a crash,ownvoice trainexits0either way. Below the 0.75 cosine-similarity bar, the adapter is still saved to disk and OwnVoice tells you plainly to try more or cleaner voice clips, more epochs, or a higher--lora-rank, then re-run. Only two things actually fail the command with a non-zero exit: no usable clips to load, or a caught PEFT-injection failure.
Can I use OwnVoice, and the adapters it produces, commercially?OwnVoice's own code is MIT (seeLICENSE). pocket-tts's code package is MIT too, but the model weights OwnVoice actually downloads and trains against,kyutai/pocket-tts-without-voice-cloningand the gatedkyutai/pocket-tts, are licensed CC-BY-4.0, not MIT. CC-BY-4.0 permits commercial use but requires attribution to Kyutai. Since any adapter you train is derived from those weights, check that attribution requirement before shipping a commercial product built on it.
Whose voice can I actually clone with this?Only your own, or someone else's with their explicit, checked consent, never a public figure's voice without it. SeeConsent and Misuseabove. OwnVoice ships no bulk-generation or auto-scaling feature in this version, which keeps the blast radius of any single misuse case small.
pip install "ownvoice-cli[mcp]"
Add it to an MCP client's config (for example, Claude Desktop'sclaude_desktop_config.json):
{ "mcpServers": { "ownvoice": { "command": "ownvoice-mcp" } } }
The server exposes a single tool,run(args: list[str]) -> dict, that shells out to the realownvoiceCLI with the given argv and returns its result as structured JSON, so a caller gets the exact same behavior the human-facing CLI has, including--jsonmode. Example call:run(args=["check", "--json"])returns{"result": {"success": true, "message": "...", "module_tree": null}}. A non-zero exit, a launch failure, or a subprocess timeout is always returned as{"error": "..."}rather than raised.
Issues and PRs welcome, MIT licensed throughout. If you want to help close the actual gap this project targets, the most useful contribution is upstream: a lightweight LoRA-adapter training script contributed back tokyutai-labs/pocket-ttsitself, discussed onissue #30.
This is a web browser that enables your coding agent, such as Claude Code, to visit websites on your behalf and assist you in identifying bugs or creating UI test cases.
Create crafted UI components inspired by the best 21st.dev design engineers.
Bring agent evaluations, observability, and synthetic test set generation directly into your IDE for free with Galileo's new MCP server
An MCP server to help AI assistants to answer questions and generate AccelByte Extend SDK code more effectively .
MCP server for AI Diagram Maker — generate beautiful software engineering diagrams directly inside Cursor, Claude Desktop, Claude Code, or any MCP-compatible AI agent
ALAPI MCP Tools,Call hundreds of API interfaces via MCP
AI-powered SVG animation generator that transforms static files into animated SVG components using the Allyson platform
MCP server that gives AI assistants on-demand access to 1,500+ amCharts docs, ~300 code examples, and 1000+ class API references.
APIMatic MCP Server is used to validate OpenAPI specifications using APIMatic. The server processes OpenAPI files and returns validation summaries by leveraging APIMatic’s API.
One shared context layer for AI agents and humans — live API specs, DB schemas, and versioned contracts across repos so every agent and teammate works from the same source of truth.
Build and deploy full-stack Next.js apps with 98 tools for React, AWS, and MongoDB
OwnVoice's own training run is real GPU time, honestly labeled and not hidden behind a fake progress bar, the same category norm Unsloth uses. What OwnVoice compresses to under two minutes is everythingbeforethat: confirming your environment actually works.
pocket-tts is a genuinely good, MIT-licensed, CPU-capable local text-to-speech model from Kyutai. Its own maintainers have been clear that fine-tuning code isn't coming any time soon: on](https://github.com/resemble-ai/Resemblyzer)issue #30, maintainer @vvolhejn wrote "We are not planning to release fine-tuning code for our TTS and STT models in the near future," and 18 people reacted to that thread asking for exactly this. OwnVoice is a small, standalone CLI that fills that specific gap: point it at a handful of your own voice recordings, and it trains a LoRA adapter you keep and run yourself.
It is not a hosted service, it has no billing, and it does not track usage. It is a training script, an inference script, and a scoring script, wired together behind three CLI commands.
pocket-tts already ships zero-shot voice cloning out of the box: pass a.wavfile to--voice(or callget_state_for_audio_prompt()from Python) and it clones that voice with no training step at all. If that is all you need, use pocket-tts directly, it is simpler and faster.
OwnVoice exists for a narrower case: baking a voice permanently into trained weights, so generation no longer depends on distributing or re-processing a reference audio clip at runtime, with (based on the training objective, not yet independently benchmarked at scale) more consistent output across many generations than a single-clip zero-shot embedding tends to produce. That is the specific gap the 18 reactors on issue #30 were describing, and it is the only thing OwnVoice adds on top of what pocket-tts already does well.
This tool clones a voice from audio you have the right to use. Do not clone someone else's voice, or a public figure's voice, without their explicit consent. OwnVoice ships no bulk-generation or auto-scaling feature in this version, keeping the blast radius of any single misuse case small.
This is a young, early-stage release.ownvoice check, the CLI argument parsing, voice-clip validation, the similarity scoring math, and the adapter/manifest save and load path are implemented and covered by the test suite (pytest). LoRA injection was verified structurally against pocket-tts's real source and then confirmed for real:ownvoice checkwas run against pocket-tts's actual downloaded weights, on CPU, and PEFT'starget_modules="all-linear"injection genuinely succeeded. The full training and generation path has since been verified end to end for real too: a real 2-epoch LoRA training run against loaded pocket-tts weights produced a finite, non-NaN flow-matching loss, and the resulting adapter produced a real, non-silent generated.wavfile viaownvoice infer. That validation surfaced two real gaps in the naive approach and fixed them: (1) pocket-tts's published, inference-only PyPI package does not actually expose a way to compute the training loss throughFlowLMModel.forward()despite its own docstring claiming otherwise, so OwnVoice computes the flow-matching loss directly fromflow_lm's real submodules instead; (2) swappingbase_model.flow_lmto the PEFT-wrapped model before callinggenerate_audio()breaks pocket-tts's internal KV-cache state lookup -- no swap is needed at all, since PEFT's LoRA injection already mutatesbase_model.flow_lmin place. One real, external limitation to know about: the publicly downloadable pocket-tts weights (kyutai/pocket-tts-without-voice-cloning) refuse a raw reference-clip path/URL outright; OwnVoice works around this by pre-loading and resampling the clip itself, but voice-cloning fidelity from that checkpoint is a known limitation of the base model, not an OwnVoice bug -- for kyutai's best-quality cloning weights, request gated access athuggingface.co/kyutai/pocket-tts. Runownvoice checkyourself and read the source before trusting any of it further, that is the right amount of skepticism for a project this early.
What is OwnVoice, and why not just use pocket-tts by itself?OwnVoice trains a LoRA adapter forpocket-ttsand saves it to your own disk asadapter.safetensorsplusmetadata.json. It exists because pocket-tts's own maintainers have said fine-tuning code is not on their near-term roadmap (seeissue #30). Once you have a trained adapter, you never need OwnVoice again to use it:ownvoice inferjust loads the adapter back onto the base model.
How is this different from pocket-tts's own built-in--voice <wav>zero-shot cloning?pocket-tts already clones a voice from a single reference clip with no training step,--voice <wav>at the CLI orget_state_for_audio_prompt()in Python. OwnVoice trades that speed for a permanently trained adapter, so generation no longer depends on carrying around a reference clip at runtime, with (based on the training objective, not yet independently benchmarked at scale) more consistent output across repeated generations than a single-clip zero-shot embedding tends to give. If zero-shot is enough for your use case, use pocket-tts directly, it is simpler and faster.
What do I need to install it, and does it run on Apple Silicon or a CPU-only machine?Python 3.11 or newer, thenpip install ownvoice-cli.ownvoice checkandownvoice inferneed no GPU at all and run fine on Apple Silicon or a CPU-only machine, matching pocket-tts's own CPU-capable design.ownvoice trainruns on CPU too, it just takes longer per epoch; install the CUDA build of PyTorch first if you have an NVIDIA GPU and want training to go faster.
How does OwnVoice compare to kokoro-tts and Unsloth?kokoro-ttsgets you synthesizing speech in under 2 minutes but has no fine-tuning step at all.Unslothgets a training run started in under a minute but is a general LLM fine-tuning framework, not TTS-specific. OwnVoice is narrower than either: one base model (pocket-tts only), one job (a voice adapter), plus a freeownvoice checkstep that confirms PEFT's LoRA injection actually works against your environment before you spend anything on a GPU, a check neither of those tools has an equivalent of.
My training run finished but printed "BELOW THRESHOLD", is that a bug?No. It is a labeled outcome, not a crash,ownvoice trainexits0either way. Below the 0.75 cosine-similarity bar, the adapter is still saved to disk and OwnVoice tells you plainly to try more or cleaner voice clips, more epochs, or a higher--lora-rank, then re-run. Only two things actually fail the command with a non-zero exit: no usable clips to load, or a caught PEFT-injection failure.
Can I use OwnVoice, and the adapters it produces, commercially?OwnVoice's own code is MIT (seeLICENSE). pocket-tts's code package is MIT too, but the model weights OwnVoice actually downloads and trains against,kyutai/pocket-tts-without-voice-cloningand the gatedkyutai/pocket-tts, are licensed CC-BY-4.0, not MIT. CC-BY-4.0 permits commercial use but requires attribution to Kyutai. Since any adapter you train is derived from those weights, check that attribution requirement before shipping a commercial product built on it.
Whose voice can I actually clone with this?Only your own, or someone else's with their explicit, checked consent, never a public figure's voice without it. SeeConsent and Misuseabove. OwnVoice ships no bulk-generation or auto-scaling feature in this version, which keeps the blast radius of any single misuse case small.
pip install "ownvoice-cli[mcp]"
Add it to an MCP client's config (for example, Claude Desktop'sclaude_desktop_config.json):
{ "mcpServers": { "ownvoice": { "command": "ownvoice-mcp" } } }
The server exposes a single tool,run(args: list[str]) -> dict, that shells out to the realownvoiceCLI with the given argv and returns its result as structured JSON, so a caller gets the exact same behavior the human-facing CLI has, including--jsonmode. Example call:run(args=["check", "--json"])returns{"result": {"success": true, "message": "...", "module_tree": null}}. A non-zero exit, a launch failure, or a subprocess timeout is always returned as{"error": "..."}rather than raised.
Issues and PRs welcome, MIT licensed throughout. If you want to help close the actual gap this project targets, the most useful contribution is upstream: a lightweight LoRA-adapter training script contributed back tokyutai-labs/pocket-ttsitself, discussed onissue #30.
This is a web browser that enables your coding agent, such as Claude Code, to visit websites on your behalf and assist you in identifying bugs or creating UI test cases.
Create crafted UI components inspired by the best 21st.dev design engineers.
Bring agent evaluations, observability, and synthetic test set generation directly into your IDE for free with Galileo's new MCP server
An MCP server to help AI assistants to answer questions and generate AccelByte Extend SDK code more effectively .
MCP server for AI Diagram Maker — generate beautiful software engineering diagrams directly inside Cursor, Claude Desktop, Claude Code, or any MCP-compatible AI agent
ALAPI MCP Tools,Call hundreds of API interfaces via MCP
AI-powered SVG animation generator that transforms static files into animated SVG components using the Allyson platform
MCP server that gives AI assistants on-demand access to 1,500+ amCharts docs, ~300 code examples, and 1000+ class API references.
APIMatic MCP Server is used to validate OpenAPI specifications using APIMatic. The server processes OpenAPI files and returns validation summaries by leveraging APIMatic’s API.
One shared context layer for AI agents and humans — live API specs, DB schemas, and versioned contracts across repos so every agent and teammate works from the same source of truth.
Build and deploy full-stack Next.js apps with 98 tools for React, AWS, and MongoDB
Claude Desktop / Cursor
Paste into your MCP client config file to install this server.
{
"mcpServers": {
"ownvoice": {
"server": {
"command": "uvx",
"args": [
"ownvoice-cli"
]
}
}
}
}
McpServers
{
"server": {
"command": "uvx",
"args": [
"ownvoice-cli"
]
}
}
Transport
"stdio"
Package
"ownvoice-cli"
Registry
"pypi"
1.ownvoice check, the free Day-0 validation
Before recording anything or renting a GPU, confirm that PEFT's LoRA injection actually works against pocket-tts's real model structure. This is entirely free: CPU only, no training, no GPU.
$ ownvoice check ](https://pytorch.org/get-started/locally/)[ownvoice check] PASS: PEFT LoRA injection succeeded against pocket-tts's flow_lm module (target_modules="all-linear").
If it fails, OwnVoice prints the model's real module tree instead of a raw stack trace, so you can see exactly what did not match and report it precisely:
$ ownvoice check [ownvoice check] FAIL: PEFT LoRA injection failed against pocket-tts's flow_lm module structure: <error detail>. Please post an honest blocker (this error plus the module tree above) as a comment on https://github.com/kyutai-labs/pocket-tts/issues/30 rather than working around it silently, that issue is exactly where this gap needs to be visible. Module tree (for debugging / for the issue #30 blocker post): <root>: FlowLMModel input_linear: Linear transformer: StreamingTransformer transformer.layers.0.self_attn.in_proj: Linear transformer.layers.0.self_attn.out_proj: Linear ...
Record 5 to 10 minutes of clean audio of the voice you want to train (your own voice, with your own consent, seeConsent and misuse), split into a few.wavclips in one directory, then point OwnVoice at it:
$ ownvoice train --voice-clips ./my-voice-clips [ownvoice train] USABLE ADAPTER Usable adapter (similarity 0.812 >= 0.75). Try it now: ownvoice infer --adapter ownvoice-adapter/adapter.safetensors --text "This is my own voice, trained with OwnVoice."
Only--voice-clipsis required. Every other flag has a sensible default (see the fullCLI Referencebelow).
A run that finishes but does not clear the similarity bar still exits0. It is a labeled result with a concrete next step, not a crash:
$ ownvoice train --voice-clips ./my-voice-clips [ownvoice train] BELOW THRESHOLD Below threshold (similarity 0.612 < 0.75). The adapter was still saved, try more/cleaner voice clips, more epochs, or a higher --lora-rank, then re-run. You can still listen to it: ownvoice infer --adapter ownvoice-adapter/adapter.safetensors --text "This is my own voice, trained with OwnVoice."
Only a data-loading problem (no usable clips) or a caught PEFT-injection failure exits non-zero. A finished run always writesadapter.safetensorsandmetadata.json(training config, similarity score, a timestamp) to the output directory: two files you keep, with no server round-trip needed to use them again.
$ ownvoice infer --adapter ownvoice-adapter/adapter.safetensors --text "Hello, this is my own voice." [ownvoice infer] Wrote ownvoice-output.wav
Every subcommand also supports--jsonfor a structured, machine-parseable output mode, useful if a script or an agent is callingownvoiceprogrammatically instead of a person reading the terminal:
$ ownvoice check --json {"success": true, "message": "PEFT LoRA injection succeeded against pocket-tts's flow_lm module (target_modules=\"all-linear\").", "module_tree": null}
Reference below is taken directly from each subcommand's real--helpoutput (ownvoice-cli0.1.2 on PyPI).
Free, CPU-only compatibility check: load pocket-tts and dry-run the LoRA injection. No GPU and no training required.
Train a LoRA voice adapter from a directory of.wavvoice clips. Only--voice-clipsis required.
Generate speech in the trained voice from a saved adapter, and save it to a.wavfile.
- A free compatibility check before you spend anything on a GPU.ownvoice checkloads pocket-tts and dry-runs PEFT's LoRA injection against its realflow_lmmodule tree, CPU only, no training. On failure it prints the actual module tree instead of a stack trace, so a real blocker is reportable instead of silent.
- An objective usable/not-usable signal, not a guess.Every training run resamples the generated test utterance to 16kHz mono and scores it against your reference clips with[Resemblyzercosine similarity.0.75or higher is labeledUSABLE ADAPTER; anything lower isBELOW THRESHOLD, a labeled outcome and not a crash, exit code0either way.
- Structured output on every subcommand.check,train, andinferall accept--json, returning one machine-parseable object instead of colored terminal text: confirmed directly,ownvoice check --jsonreturns{"success": true, "message": "...", "module_tree": null}.
- Two files you keep, no server round-trip.A finished training run writesadapter.safetensors(the trained weights, a few megabytes at the default--lora-rank 8) andmetadata.json(the full training config, similarity score, per-epoch loss, and a timestamp) to disk. Load them back any time later withownvoice infer, no network call required.
- One base model, on purpose.OwnVoice wraps pocket-tts only. There is no abstraction layer for a second base model, matching the codebase's own single-target-by-design architecture note: the LoRA injection path (target_modules="all-linear"against pocket-tts's realflow_lmlayers) stays exact instead of generic.
voice clips (wav) | v data.py --validate format/duration--> clean clip set | v train.py --PEFT LoRA (target_modules="all-linear")--> adapter.safetensors + metadata.json | v infer.py --generate test utterance--> synthesized audio | v score.py --resample to 16kHz mono--> Resemblyzer cosine similarity | v CLI report (>= 0.75 = usable adapter, below triggers a labeled next-step message)
ownvoice/data.pyloads and validates the voice-clip directory.ownvoice/train.pyloads pocket-tts's frozen base model, injects a LoRA adapter into itsflow_lmtransformer with PEFT (target_modules="all-linear"), runs the training loop, and saves the adapter plus a manifest.ownvoice/infer.pyloads a saved adapter back onto the base model and generates speech.ownvoice/score.pyresamples audio to 16kHz mono withtorchaudio.transforms.Resampleand scores speaker similarity with Resemblyzer.
OwnVoice is intentionally single-model: it wraps pocket-tts only, with no abstraction layer for a second base model, since none is in scope.
Sign in to leave a review
Use Google, GitHub, or an email account so ratings stay tied to real people.
No reviews posted yet.



