Pandas MCP Server
About
# Pandas-MCP Server [](https://opensource.org/licenses/MIT) [](https://www.python.org/downloads/) [![Code Style…
Details
- License
- MIT
Explore
- Complete Value Distribution: Returns all unique values with their exact counts
- Sorted by Frequency: Values are sorted in descending order of occurrence
- Data Type Analysis: Identifies the underlying data type (object, int64, etc.)
- Quality Metrics: Provides null count and total values for data quality assessment
- Multi-column Support: Can analyze multiple columns in a single request
MCP Tool Usage:
{
"tool": "interpret_column_data",
"args": {
"file_path": "/path/to/sales_data.csv",
"column_names": ["Region", "Status"]
}
}
Excel File Usage (with optional sheet selection):
{
"tool": "interpret_column_data",
"args": {
"file_path": "/path/to/sales_data.xlsx",
"column_names": ["Region", "Status"],
"sheet_name": "Q3_Sales" // Optional: sheet name or index (default: 0)
}
}
Setting up with Highlight
This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:
- Download and install Highlight from highlightai.com/download
- Navigate to the plugins tab and select "Add Custom Plugin"
-
Configure the plugin with the settings below
Plugin Name
Pandas MCP ServerCommand (node, npx, python, etc.)Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.
- Enable "Start Automatically" if you want the plugin to start when Highlight launches
From the repository
- Python 3.10+
- pip package manager
- Git (for cloning the repository)
pip install -r requirements.txt
The server supports extensive configuration through environment variables. Copy the example configuration file:
cp .env.example .env
Edit the .env file to customize settings such as:
- Log levels and file locations
- File size limits
- Feature flags (enable/disable chart generation, code execution)
- Memory monitoring settings
- Security blacklist extensions
For detailed configuration options, see CONFIGURATION.md.
bash
{
"mcpServers": {
"pandas-server": {
"type": "stdio",
"command": "uvx",
"args": ["--from", "/path/to/pandas-mcp-server", "pandas-mcp-server"]
}
}
}
Path Format by Operating System:
- Windows: Use forward slashes (recommended) or escaped backslashes
- C:/Project/pandas-mcp-server (recommended)
- C:\\Project\\pandas-mcp-server (escaped backslashes)
- macOS/Linux: Use forward slashes
- /Users/username/projects/pandas-mcp-server
- /home/username/projects/pandas-mcp-server
Note: Replace /path/to/pandas-mcp-server with the absolute path to your pandas-mcp-server directory. The full path is required because Claude Desktop executes commands from its own working directory.
- Windows: %APPDATA%\Claude\claude_desktop_config.json
- macOS: ~/Library/Application Support/Claude/claude_desktop_config.json
- Linux: ~/.config/Claude/claude_desktop_config.json
- MAX_FILE_SIZE: 100MB file size limit
- BLACKLIST: Security restrictions for code execution
- CHARTS_DIR: Directory for generated charts
- Logging: Comprehensive logging with rotation
```bash
- Report Issues: Submit an issue on GitHub
- Discussions: Join our GitHub Discussions
The server exposes four main tools for LLM integration:
Extract comprehensive metadata from Excel and CSV files including:
- File type, size, encoding, and structure
- Column names, data types, and sample values
- Statistical summaries (null counts, unique values, min/max/mean)
- Data quality warnings and suggested operations
- Memory-optimized processing for large files
Purpose: Provides LLM with a high-level understanding of the data structure and characteristics, serving as the foundation for data analysis.
MCP Tool Usage:
{
"tool": "read_metadata_tool",
"args": {
"file_path": "/path/to/sales_data.xlsx"
}
}
Execute pandas operations with:
- Security filtering against malicious code
- Memory optimization for large datasets
- Comprehensive error handling and debugging
- Support for DataFrame, Series, and dictionary results
Purpose: Leverages insights from both read_metadata_tool and interpret_column_data to execute precise data analysis operations, particularly valuable when processing multiple CSV files with consistent value patterns.
Generate interactive charts with Chart.js:
- Bar charts - For categorical comparisons
- Line charts - For trend analysis
- Pie charts - For proportional data
- Interactive HTML templates with customization controls
Claude Desktop / Cursor
Paste into your MCP client config file to install this server.
{
"mcpServers": {
"pandas mcp server": {
"pandas-mcp-server": {
"command": "python",
"args": [
"cli.py"
]
}
}
}
}
McpServers
{
"pandas-mcp-server": {
"command": "python",
"args": [
"cli.py"
]
}
}
<div align="center">
🚀 Powerful tool for AI-powered data analysis - Through MCP protocol, enables LLMs to safely and efficiently execute pandas code and generate visualizations
If you find this project helpful, please consider giving it a ⭐️ star!
</div>
---
A comprehensive Model Context Protocol (MCP) server that enables LLMs to execute pandas code through a standardized workflow for data analysis and visualization.
✨ Key Features
- 🔒 Secure Execution Environment - Sandboxed code execution prevents malicious operations and protects system security
- 📊 Intelligent Data Analysis - Automatically extracts file metadata, understands data structure, and provides intelligent analysis suggestions
- 🎨 Interactive Visualizations - One-click generation of various interactive charts with real-time parameter adjustment
- 🧠 Memory Optimization - Intelligent memory management supports large file processing with automatic data type optimization
- 🔧 Easy Integration - Simple configuration for seamless integration with AI assistants like Claude Desktop
- 📝 CLI Support - Provides command-line interface for convenient testing and development
🎯 MCP Server Overview
The Pandas-MCP Server is designed as a Model Context Protocol (MCP) server that provides LLMs with powerful data processing capabilities. MCP is a standardized protocol that allows AI models to interact with external tools and services in a secure, structured way.
🛠️ Installation
Prerequisites
- Python 3.10+ - pip package manager - Git (for cloning the repository)Step 1: Clone the Repository
git clone https://github.com/marlonluo2018/pandas-mcp-server.git
cd pandas-mcp-server
Step 2: Install Dependencies
pip install -r requirements.txt
Step 3: Configure Environment Variables (Optional)
The server supports extensive configuration through environment variables. Copy the example configuration file:cp .env.example .env
Edit the .env file to customize settings such as:
- Log levels and file locations
- File size limits
- Feature flags (enable/disable chart generation, code execution)
- Memory monitoring settings
- Security blacklist extensions
For detailed configuration options, see CONFIGURATION.md.
Step 4: Verify Installation
# Test the CLI interface
python cli.py
Or test the MCP server directly
python server.py
Dependencies
- pandas>=2.0.0 - Data manipulation and analysis - fastmcp>=1.0.0 - MCP server framework - chardet>=5.0.0 - Character encoding detection - psutil - System monitoring for memory optimizationClaude Desktop Configuration
Using uvx (Recommended)
uvx is a fast Python package installer and runner that makes it easy to run Python tools without manual environment setup.
Install uvx
# Using pip
pip install uv
Or using pipx (recommended for isolation)
pipx install uv
uvx Advantages
- No manual installation: Automatically downloads and runs the package - Isolated environments: Each run uses a clean virtual environment - Fast: Uses uv's fast dependency resolver - Version pinning: Easy to run specific versionsAdd this configuration to your Claude Desktop settings:
{
"mcpServers": {
"pandas-server": {
"type": "stdio",
"command": "uvx",
"args": ["--from", "/path/to/pandas-mcp-server", "pandas-mcp-server"]
}
}
}
Path Format by Operating System:
- Windows: Use forward slashes (recommended) or escaped backslashes
- C:/Project/pandas-mcp-server (recommended)
- C:\\Project\\pandas-mcp-server (escaped backslashes)
- macOS/Linux: Use forward slashes
- /Users/username/projects/pandas-mcp-server
- /home/username/projects/pandas-mcp-server
Note: Replace /path/to/pandas-mcp-server with the absolute path to your pandas-mcp-server directory. The full path is required because Claude Desktop executes commands from its own working directory.
Using Python (Traditional)
Add this configuration to your Claude Desktop settings:{
"mcpServers": {
"pandas-server": {
"type": "stdio",
"command": "python",
"args": ["/path/to/your/pandas-mcp-server/server.py"]
}
}
}
Configuration File Location
- Windows:%APPDATA%\Claude\claude_desktop_config.json
- macOS: ~/Library/Application Support/Claude/claude_desktop_config.json
- Linux: ~/.config/Claude/claude_desktop_config.json
Verification
After configuration, restart Claude Desktop. The server should appear in the MCP tools list with four available tools: -read_metadata_tool - File analysis
- interpret_column_data - Column value interpretation
- run_pandas_code_tool - Code execution
- generate_chartjs_tool - Chart generation
🔄 Workflow
The pandas MCP server follows a structured workflow for data analysis and visualization:
Step 1: Read File Metadata
LLM callsread_metadata_tool to understand the file structure:
- Extract file type, size, encoding, and column information
- Get data types, sample values, and statistical summaries
- Receive data quality warnings and suggested operations
- Understand the dataset structure before processing
Step 2: Interpret Column Values (Optional)
LLM callsinterpret_column_data to understand specific columns:
- Extract all unique values from important columns
- Identify patterns in categorical data
- Understand the meaning behind codes or abbreviations
- Key Purpose: Complement metadata by providing deep understanding of column values, which helps LLM generate more accurate and effective pandas code in the next step, especially when working with multiple CSV files
When to Use interpret_column_data:
- High Value: Categorical fields with limited unique values (Region, Status, Category)
- High Value: Code fields that need interpretation (StatusCode "A", "B", "C")
- High Value: Fields with abbreviations or cryptic values
- Low Value: ID fields (usually unique values with no patterns)
- Low Value: Email fields (typically unique identifiers)
- Low Value: Numeric percentage fields (already self-explanatory)
- Conditional: Time fields (useful for non-standard formats or categorical time)
Step 3: Execute Pandas Operations
LLM callsrun_pandas_code_tool based on metadata and column analysis:
- Formulate pandas operations using the understood file structure
- Execute data processing, filtering, aggregation, or analysis
- Receive results in DataFrame, Series, or dictionary format
- Get optimized output with memory management
Step 4: Generate Visualizations
LLM callsgenerate_chartjs_tool to create interactive charts:
- Transform processed data into Chart.js compatible format
- Generate interactive HTML charts with customization controls
- Create bar, line, or pie charts based on data characteristics
- Output responsive visualizations for analysis presentation
How interpret_column_data Complements read_metadata
The interpret_column_data function is designed to complement the read_metadata_tool by providing deeper insights into column values:
…
Sign in to leave a review
Use Google, GitHub, or an email account so ratings stay tied to real people.
No reviews posted yet.



