MCP Server Whisper

by arcaputo3

56 443 downloads Not rated yet MIT

About

Advanced audio transcription and processing using OpenAI's Whisper and GPT-4o models.

Details

License
MIT

Explore

- Advanced file searching with regex and metadata filters
- Multi‑model transcription (Whisper, GPT‑4o transcribe)
- Interactive audio chat with GPT‑4o audio models
- Text‑to‑speech with customizable voices and speed
- Automatic compression of files over 25 MB
- Enhanced transcription with specialized templates

Setting up with Highlight

This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:

  1. Download and install Highlight from highlightai.com/download
  2. Navigate to the plugins tab and select "Add Custom Plugin"
  3. Configure the plugin with the settings below
    Plugin Name MCP Server Whisper
    Command (node, npx, python, etc.)

    Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.

  4. Enable "Start Automatically" if you want the plugin to start when Highlight launches

From the repository


Create a .env file based on the provided .env.example:

bash
cp .env.example .env

Edit .env with your actual values:


OPENAI_API_KEY=your_openai_api_key
AUDIO_FILES_PATH=/path/to/your/audio/files

Note: Environment variables must be available at runtime. For local development with Claude, use a tool like dotenv-cli to load them (see Usage section below).

<details>
<summary>Basic Audio Transcription</summary>


Claude, please transcribe my latest audio file with detailed insights.

Claude will automatically:
1. Find the latest audio file using get_latest_audio
2. Determine the appropriate transcription method
3. Process the file with transcribe_with_enhancement using the "detailed" template
4. Return the enhanced transcription
</details>

<details>
<summary>Advanced Audio File Search and Filtering</summary>


Claude, list all my audio files that are longer than 5 minutes and were created after January 1st, 2024, sorted by size.

Claude will:
1. Convert the date to a timestamp
2. Use list_audio_files with appropriate filters:
- min_duration_seconds: 300 (5 minutes)
- min_modified_time: <timestamp for Jan 1, 2024>
- sort_by: "size"
3. Return a sorted list of matching audio files with comprehensive metadata
</details>

<details>
<summary>Batch Processing Multiple Files</summary>


Claude, find all MP3 files with "interview" in the filename and create professional transcripts for each one.

Claude will:
1. Search for files using list_audio_files with pattern and format filters
2. Make multiple parallel transcribe_with_enhancement tool calls (MCP handles parallelism natively)
3. Each call uses enhancement_type: "professional" and returns a typed TranscriptionResult
4. Return all transcriptions with full metadata in a well-formatted output
</details>

<details>
<summary>Generating Text-to-Speech Audio</summary>


Claude, create audio with this script: "Welcome to our podcast! Today we'll be discussing artificial intelligence trends in 2025." Use the shimmer voice.
``

Claude will:
1. Use the
create_audio tool with:
-
text_prompt containing the script
-
voice: "shimmer"
-
model: "gpt-4o-mini-tts" (default high-quality model)
-
instructions: "Speak in an enthusiastic, podcast host style" (optional)
-
speed: 1.0 (default, can be adjusted)
2. Generate the audio file and save it to the configured audio directory
3. Provide the path to the generated audio file
</details>

For production use with Claude Desktop (as opposed to local development), add this to your claude_desktop_config.json`:

Claude Desktop / Cursor

Paste into your MCP client config file to install this server.

{
    "mcpServers": {
        "mcp server whisper": {
            "mcp-server-whisper": {
                "command": "uv",
                "args": [
                    "sync"
                ]
            }
        }
    }
}

McpServers

{
    "mcp-server-whisper": {
        "command": "uv",
        "args": [
            "sync"
        ]
    }
}

<div align="center">

A Model Context Protocol (MCP) server for advanced audio transcription and processing using OpenAI's Whisper and GPT-4o models.

PyPI version
License: MIT
Python 3.10+
CI Status
Built with uv

</div>

> [!WARNING]
> This project has moved. Active development has migrated to TJC-LP/sanzaru. This repository is no longer maintained and will be archived. Please update your dependencies and issues to the new repo.

Overview

MCP Server Whisper provides a standardized way to process audio files through OpenAI's latest transcription and speech services. By implementing the Model Context Protocol, it enables AI assistants like Claude to seamlessly interact with audio processing capabilities.

Key features:
- 🔍 Advanced file searching with regex patterns, file metadata filtering, and sorting capabilities
- ⚡ MCP-native parallel processing - call multiple tools simultaneously
- 🔄 Format conversion between supported audio types
- 📦 Automatic compression for oversized files
- 🎯 Multi-model transcription with support for all OpenAI audio models
- 🗣️ Interactive audio chat with GPT-4o audio models
- ✏️ Enhanced transcription with specialized prompts and timestamp support
- 🎙️ Text-to-speech generation with customizable voices, instructions, and speed
- 📊 Comprehensive metadata including duration, file size, and format support
- 🚀 High-performance caching for repeated operations
- 🔒 Type-safe responses with Pydantic models for all tool outputs

> Note: This project is unofficial and not affiliated with, endorsed by, or sponsored by OpenAI. It provides a Model Context Protocol interface to OpenAI's publicly available APIs.

Installation

```bash

No reviews yet — be the first

Sign in to leave a review

Use Google, GitHub, or an email account so ratings stay tied to real people.

Email sign in

No reviews posted yet.

Videos about MCP Server Whisper

Relevant YouTube tutorials, setups, and demos