VocalRemover MCP
About
Audio tools MCP server for professional audio editing and AI augmentation
Details
- Author
- Unknown
- Categories
- Productivity, Other, Media
Jump to
Ask your assistant to edit audio in plain language — for example, "remove the vocals from this track."
- Remove vocals for karaoke— Get an instrumental backing track viaseparate_audiowith the instrumental task.
- Split songs into stems— Useseparate_audiowith the stems task to get vocals, drums, bass, guitar, and piano.
- Convert audio formats— Convert files between MP3, WAV, FLAC, OGG, M4A, AAC, and OPUS.
- Trim and probe audio— Trim a clip to an exact start time and length, or check duration, bitrate, and format.
- Extract audio from video— Pull the audio track from MP4, MKV, MOV, or AVI files.
- Clean up voice recordings— Denoise and de-reverb recordings to remove background noise and room echo.
Audio MCP Server for Claude, Cursor, ChatGPT & any AI agent
Remove vocals, split stems, clean up recordings and more — right inside Claude, Cursor or ChatGPT. Just say what you want and get the finished file back.
Set it up in under a minute — then let your assistant do the work.
Works with Claude· Cursor· Windsurf· Cline· Zed· ChatGPT · 22 tools · 18 free
Once connected, type what you want — your assistant calls the right tool and hands back a download link.
Remove the vocals from this song and give me the karaoke track.
Called separate_audio · task: instrumental
Done — vocals removed with our HQ Fusion model. Here is your karaoke track:
Now split the original into stems — vocals, drums, bass, guitar and piano.
Split into 5 stems with Multi-stem HQ. Here are your downloads: vocals.mp3 drums.mp3 bass.mp3 guitar.mp3 piano.mp3
Convert this interview to MP3 and clean up the background noise. De-reverb this voice recording. Make an 8D version of this song and trim it to 30 seconds. Extract the audio from this video as a WAV.
"Take these 12 practice recordings, convert each to MP3, then remove the vocals" runs end-to-end in a single agent session — convert, separate, poll, return 12 links. No app-switching, no glue code.
- Convert between MP3, WAV, FLAC, OGG, M4A, AAC and OPUS.
- Trim a clip to an exact start time and length.
- Extract the audio track from any video file.
- Probe a file for its duration, bitrate and format.
- Remove vocals or grab the instrumental for karaoke.
- Split a song into stems: vocals, drums, bass, guitar, piano.
- Clean up voice recordings — remove noise and room echo.
- Track progress and grab results without leaving the chat.
Many audio MCP servers ship older, open-source separation models. We run our own models — the same ones powering vocalremover.com — with a restoration pass on top.
Our flagship separation for cleaner vocals and instrumentals on complex mixes.
Splits a track into vocals, drums, bass, guitar and piano stems.
Restores clarity and detail to processed or low-quality audio.
- 1 Open connector settings In ChatGPT: Settings → Apps → enable Developer mode → Create app, then paste the server URL into the Connection field. In Claude.ai: Settings → Connectors → Add custom connector.
- Drop this address into the connector URL field:
https://vocalremover.com/mcp/audio
For Claude Desktop, Cursor, Windsurf, Cline, Zed and Claude Code — add the server with a Bearer API token.
Create your VocalRemover account, then copy your API token from the API page and paste it into the config below.
Add to claude_desktop_config.json, then restart Claude.
{ "mcpServers": { "vocalremover": { "command": "npx", "args": [ "-y", "mcp-remote", "https://vocalremover.com/mcp/audio", "--header", "Authorization:${AUTH}" ], "env": { "AUTH": "Bearer YOUR_API_TOKEN" } } } }
Native remote MCP — add to the client's mcp.json.
{ "mcpServers": { "vocalremover": { "url": "https://vocalremover.com/mcp/audio", "headers": { "Authorization": "Bearer YOUR_API_TOKEN" } } } }
One command — native remote MCP, no bridge.
claude mcp add --transport http \ vocalremover https://vocalremover.com/mcp/audio \ -H "Authorization: Bearer YOUR_API_TOKEN"
Your audio is fetched from the URL you pass, processed, and deleted — nothing is stored. Your access is scoped to your account and can be revoked any time.
Need details?Read the full documentation
Convert, trim, extract and clean up audio as much as you like — completely free.
Studio-grade vocal removal and stem splitting, right inside your assistant.
Turn any song into an instrumental on command.
De-reverb and denoise voice recordings in one step.
Pull clean stems for remixing and sampling.
Batch-process whole folders inside a single agent run.
An MCP (Model Context Protocol) server lets AI assistants like Claude or Cursor use audio tools as native capabilities. You simply type "remove the vocals from this track" and the assistant does it — no app switching, no manual uploads.
You can start for free. Lots of audio tools — format conversion, trimming, extracting audio from video and more — are free to use, and you can try the AI tools (vocal removal, stem separation, denoise, de-reverb) for yourself.
Claude Desktop, Claude Code, Cursor, Cline, Windsurf, Zed and the Anthropic Messages API connect with your Bearer token directly. The same processing is also available through our REST API.
Yes. It is built on the open Model Context Protocol, so it works in any compatible client — a standard connector, not a proprietary plugin.
Your audio is processed and then deleted immediately — we do not store it, whether you pass a link or upload it directly. Your access is scoped to your account and can be revoked any time.
Our AI separation runs on our own state-of-the-art models — HQ Fusion, Vocals HQ and Multi-stem HQ — with a Vocal Restoration pass for extra clarity. Many other audio MCP servers ship older open-source models.
What audio and video formats are supported?
MP3, WAV, FLAC, M4A, OGG, AAC and OPUS audio, plus common video formats (MP4, MKV, MOV, AVI) for audio extraction. Pass files as a public link or upload them directly (base64), up to 100 MB each.
Can I use it in automated agent pipelines?
Yes — that is a first-class use case. MCP is built for agent orchestration, so you can chain audio tasks across many files in a single agent run, for example "separate vocals from every track in this folder".
Can I make karaoke tracks with this MCP server?
Yes. Ask your assistant to "remove the vocals" or "give me the instrumental" and it returns a karaoke-ready backing track. You can do this for a single song or batch a whole folder in one agent run.
Can it clean up podcasts and voice recordings?
Yes. The denoise and de-reverb tools remove room sound and background noise from voice recordings, which is ideal for podcasters, voiceover artists and transcription pipelines.
Connect the server to your AI assistant and let Claude, Cursor or ChatGPT edit audio for you.
The World's First AI Music MCP Beyond images and video, your agent can now generate music.
MCP server for Audacity 3.x with 131 tools — effects, cleanup, mastering, format conversion, transcription.
Enables AI assistants to interact with DaVinci Resolve Studio for advanced control over video editing, color grading, and audio.
Transcribe and summarize video content from links using various transcription services.
Generate upload-ready clips from podcast.
AI-powered music production in REAPER via the Model Context Protocol — 150 tools for composition, mixing, mastering, and audio analysis.
: MCP server for AI media generation (imagesflux, videosveo3.1, music suno v5, with deterministic cost control using reserve-burn-refund billing
Create AI music videos and audio-reactive visuals from songs through MCP.
Create and publish unlimited podcast shows and episodes with ELEMENT.FM
Turn any language model into a multimodal powerhouse that can generate images, music, videos and more on the fly. Rostro's tools are designed to be used by language models from the ground up, expanding capabilities with minimal context bloat.
Sign in to leave a review
Use Google, GitHub, or an email account so ratings stay tied to real people.
No reviews posted yet.





