Parquet MCP Server

by DeepSpringAI

201 downloads Not rated yet
GitHub

About

# parquet_mcp_server [![smithery badge](https://smithery.ai/badge/@DeepSpringAI/parquet_mcp_server)](https://smithery.ai/server/@DeepSpringAI/parquet_mcp_server) A powerful MCP (Model Control Protocol) server that provides tools for manipulating and analyzing Parquet files. This server is designed to work with Claude…

Explore

- Generate text embeddings from Parquet columns
- Extract Parquet file schema and metadata
- Convert Parquet to DuckDB databases
- Convert Parquet to PostgreSQL tables with pgvector
- Process markdown files into structured chunks

Setting up with Highlight

This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:

  1. Download and install Highlight from highlightai.com/download
  2. Navigate to the plugins tab and select "Add Custom Plugin"
  3. Configure the plugin with the settings below
    Plugin Name Parquet MCP Server
    Command (node, npx, python, etc.)

    Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.

  4. Enable "Start Automatically" if you want the plugin to start when Highlight launches

From the repository

To install Parquet MCP Server for Claude Desktop automatically via Smithery:

npx -y @smithery/cli install @DeepSpringAI/parquet_mcp_server --client claude
uv venv
.venv\Scripts\activate  # On Windows
source .venv/bin/activate  # On macOS/Linux
uv pip install -e .

Create a .env file with the following variables:

EMBEDDING_URL=  # URL for the embedding service
OLLAMA_URL=    # URL for Ollama server
EMBEDDING_MODEL=nomic-embed-text  # Model to use for generating embeddings

POSTGRES_DB=your_database_name
POSTGRES_USER=your_username
POSTGRES_PASSWORD=your_password
POSTGRES_HOST=localhost
POSTGRES_PORT=5432

Add this to your Claude Desktop configuration file (claude_desktop_config.json):

{
  "mcpServers": {
    "parquet-mcp-server": {
      "command": "uv",
      "args": [
        "--directory",
        "/home/${USER}/workspace/parquet_mcp_server/src/parquet_mcp_server",
        "run",
        "main.py"
      ]
    }
  }
}

input_path

Path to input Parquet file

output_path

Path to save the output

column_name

Column containing text to embed

embedding_column

Name for the new embedding column

batch_size

Number of texts to process in each batch (for better performance)

file_path

Path to the Parquet file to analyze

parquet_path

Path to the input Parquet file

output_dir

Directory to save the DuckDB database (defaults to same directory as input file)

table_name

Name of the PostgreSQL table to create or append to

The server provides five main tools:

1. Embed Parquet: Adds embeddings to a specific column in a Parquet file
- Required parameters:
- input_path: Path to input Parquet file
- output_path: Path to save the output
- column_name: Column containing text to embed
- embedding_column: Name for the new embedding column
- batch_size: Number of texts to process in each batch (for better performance)

2. Parquet Information: Get details about a Parquet file
- Required parameters:
- file_path: Path to the Parquet file to analyze

3. Convert to DuckDB: Convert a Parquet file to a DuckDB database
- Required parameters:
- parquet_path: Path to the input Parquet file
- Optional parameters:
- output_dir: Directory to save the DuckDB database (defaults to same directory as input file)

4. Convert to PostgreSQL: Convert a Parquet file to a PostgreSQL table with pgvector support
- Required parameters:
- parquet_path: Path to the input Parquet file
- table_name: Name of the PostgreSQL table to create or append to

5. Process Markdown: Convert markdown files into structured chunks with metadata
- Required parameters:
- file_path: Path to the markdown file to process
- output_path: Path to save the output parquet file
- Features:
- Preserves document structure and links
- Extracts section headers and metadata
- Memory-optimized for large files
- Configurable chunk size and overlap

Claude Desktop / Cursor

Paste into your MCP client config file to install this server.

{
    "mcpServers": {
        "parquet mcp server": {
            "parquet_mcp_server": {
                "command": "npx",
                "args": [
                    "-y",
                    "@smithery/cli",
                    "install",
                    "@DeepSpringAI/parquet_mcp_server",
                    "--client",
                    "claude"
                ]
            }
        }
    }
}

McpServers

{
    "parquet_mcp_server": {
        "command": "npx",
        "args": [
            "-y",
            "@smithery/cli",
            "install",
            "@DeepSpringAI/parquet_mcp_server",
            "--client",
            "claude"
        ]
    }
}
smithery badge

A powerful MCP (Model Control Protocol) server that provides tools for manipulating and analyzing Parquet files. This server is designed to work with Claude Desktop and offers five main functionalities:

1. Text Embedding Generation: Convert text columns in Parquet files into vector embeddings using Ollama models
2. Parquet File Analysis: Extract detailed information about Parquet files including schema, row count, and file size
3. DuckDB Integration: Convert Parquet files to DuckDB databases for efficient querying and analysis
4. PostgreSQL Integration: Convert Parquet files to PostgreSQL tables with pgvector support for vector similarity search
5. Markdown Processing: Convert markdown files into chunked text with metadata, preserving document structure and links

This server is particularly useful for:
- Data scientists working with large Parquet datasets
- Applications requiring vector embeddings for text data
- Projects needing to analyze or convert Parquet files
- Workflows that benefit from DuckDB's fast querying capabilities
- Applications requiring vector similarity search with PostgreSQL and pgvector

Installation

Installing via Smithery

To install Parquet MCP Server for Claude Desktop automatically via Smithery:

npx -y @smithery/cli install @DeepSpringAI/parquet_mcp_server --client claude

Clone this repository

git clone ...
cd parquet_mcp_server

Create and activate virtual environment

uv venv
.venv\Scripts\activate  # On Windows
source .venv/bin/activate  # On macOS/Linux

Install the package

uv pip install -e .

Environment

Create a .env file with the following variables:

```bash
EMBEDDING_URL= # URL for the embedding service
OLLAMA_URL= # URL for Ollama server
EMBEDDING_MODEL=nomic-embed-text # Model to use for generating embeddings

No reviews yet — be the first

Sign in to leave a review

Use Google, GitHub, or an email account so ratings stay tied to real people.

Email sign in

No reviews posted yet.