Dataproc MCP Server

by dipseth

10 195 downloads Not rated yet MIT
GitHub

About

An MCP server for managing Google Cloud Dataproc operations and big data workflows, with seamless integration for VS Code.

Details

License
MIT

Explore

- ✅ Complete Tool Access - All 22 MCP tools available in Claude.ai
- ✅ HTTPS Tunneling - Cloudflare tunnel for secure external access
- ✅ OAuth Authentication - GitHub OAuth for secure authentication
- ✅ Trusted Certificates - No browser warnings or connection issues
- ✅ WebSocket Support - Full WebSocket compatibility with Claude.ai
- ✅ Production Ready - Tested and verified working solution

- 22 Production-Ready MCP Tools - Complete Dataproc management suite
- 🧠 Knowledge Base Semantic Search - Natural language queries with optional Qdrant integration
- 🚀 Response Optimization - 60-96% token reduction with Qdrant storage
- 🔄 Generic Type Conversion System - Automatic, type-safe data transformations
- 60-80% Parameter Reduction - Intelligent default injection
- Multi-Environment Support - Dev/staging/production configurations
- Service Account Impersonation - Enterprise authentication
- Real-time Job Monitoring - Comprehensive status tracking

- 🧠 Semantic Search: Natural language queries with Qdrant integration
- ⚡ Smart Defaults: 60-80% parameter reduction through intelligent injection
- 📊 Response Optimization: 96% token reduction with full data preservation
- 🔄 Async Support: Non-blocking job submission and monitoring
- 🏷️ Profile System: 8 production-ready cluster templates
- 📈 Analytics: Comprehensive insights and performance tracking

Setting up with Highlight

This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:

  1. Download and install Highlight from highlightai.com/download
  2. Navigate to the plugins tab and select "Add Custom Plugin"
  3. Configure the plugin with the settings below
    Plugin Name Dataproc MCP Server
    Command (node, npx, python, etc.)

    Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.

  4. Enable "Start Automatically" if you want the plugin to start when Highlight launches

From the repository

{
  "mcpServers": {
    "dataproc": {
      "command": "npx",
      "args": ["@dipseth/dataproc-mcp-server@latest"],
      "env": {
        "LOG_LEVEL": "info",
        "DATAPROC_CONFIG_PATH": "/path/to/your/config.json"
      }
    }
  }
}

npm install -g @dipseth/dataproc-mcp-server

1. Install the package:

bash
npm install -g @dipseth/dataproc-mcp-server@latest

2. Run the setup:
bash
dataproc-mcp --setup

3. Configure authentication:
bash

nano config/server.json


4. Start the server:
bash
dataproc-mcp

1. Setup GitHub OAuth (5 minutes)
2. Generate SSL certificates: npm run ssl:generate
3. Start services (2 terminals as shown above)
4. Connect Claude.ai to your tunnel URL

> 📖 Complete Guide: See docs/claude-ai-integration.md for detailed setup instructions, troubleshooting, and advanced features.

> 📖 Certificate Setup: See docs/trusted-certificates.md for SSL certificate configuration.

| Tool | Description | Smart Defaults | Key Features |
|------|-------------|----------------|--------------|
| list_profiles | List available cluster profiles | ✅ Category filtering | 8 production profiles |
| get_profile | Get detailed profile configuration | ✅ Profile ID only | Template access |
| query_cluster_data | Query stored cluster data | ✅ Natural language | Semantic search |

The server supports a project-based configuration format:

yaml


1. Service Account Impersonation (Recommended)
2. Direct Service Account Key
3. Application Default Credentials
4. Hybrid Authentication with fallbacks

bash

npm install

start_dataproc_cluster

Create and start new clusters

create_cluster_from_yaml

Create from YAML configuration

create_cluster_from_profile

Create using predefined profiles

list_clusters

List all clusters with filtering

list_tracked_clusters

List MCP-created clusters

get_cluster

Get detailed cluster information

delete_cluster

Delete existing clusters

get_zeppelin_url

Get Zeppelin notebook URL

submit_hive_query

Submit Hive queries to clusters

submit_dataproc_job

Submit Spark/PySpark/Presto jobs

cancel_dataproc_job

Cancel running or pending jobs

get_job_status

Get job execution status

get_job_results

Get job outputs and results

get_query_status

Get Hive query status

get_query_results

Get Hive query results

list_profiles

List available cluster profiles

get_profile

Get detailed profile configuration

query_cluster_data

Query stored cluster data

check_active_jobs

Quick status of all active jobs

get_cluster_insights

Comprehensive cluster analytics

get_job_analytics

Job performance analytics

query_knowledge

Query comprehensive knowledge base

> 🔄 Enhanced with Generic Type Conversion: All tools now benefit from automatic, type-safe data transformations with intelligent compression and field mapping.

| Tool | Description | Smart Defaults | Key Features |
|------|-------------|----------------|--------------|
| start_dataproc_cluster | Create and start new clusters | ✅ 80% fewer params | Profile-based, auto-config |
| create_cluster_from_yaml | Create from YAML configuration | ✅ Project/region injection | Template-driven setup |
| create_cluster_from_profile | Create using predefined profiles | ✅ 85% fewer params | 8 built-in profiles |
| list_clusters | List all clusters with filtering | ✅ No params needed | Semantic queries, pagination |
| list_tracked_clusters | List MCP-created clusters | ✅ Profile filtering | Creation tracking |
| get_cluster | Get detailed cluster information | ✅ 75% fewer params | Semantic data extraction |
| delete_cluster | Delete existing clusters | ✅ Project/region defaults | Safe deletion |
| get_zeppelin_url | Get Zeppelin notebook URL | ✅ Auto-discovery | Web interface access |

| Tool | Description | Smart Defaults | Key Features |
|------|-------------|----------------|--------------|
| submit_hive_query | Submit Hive queries to clusters | ✅ 70% fewer params | Async support, timeouts |
| submit_dataproc_job | Submit Spark/PySpark/Presto jobs | ✅ 75% fewer params | Multi-engine support, Local file staging |
| cancel_dataproc_job | Cancel running or pending jobs | ✅ JobID only needed | Emergency cancellation, cost control |
| get_job_status | Get job execution status | ✅ JobID only needed | Real-time monitoring |
| get_job_results | Get job outputs and results | ✅ Auto-pagination | Result formatting |
| get_query_status | Get Hive query status | ✅ Minimal params | Query tracking |
| get_query_results | Get Hive query results | ✅ Smart pagination | Enhanced async support |

| Tool | Description | Smart Defaults | Key Features |
|------|-------------|----------------|--------------|
| list_profiles | List available cluster profiles | ✅ Category filtering | 8 production profiles |
| get_profile | Get detailed profile configuration | ✅ Profile ID only | Template access |
| query_cluster_data | Query stored cluster data | ✅ Natural language | Semantic search |

| Tool | Description | Smart Defaults | Key Features |
|------|-------------|----------------|--------------|
| check_active_jobs | Quick status of all active jobs | ✅ No params needed | Multi-project view |
| get_cluster_insights | Comprehensive cluster analytics | ✅ Auto-discovery | Machine types, components |
| get_job_analytics | Job performance analytics | ✅ Success rates | Error patterns, metrics |
| query_knowledge | Query comprehensive knowledge base | ✅ Natural language | Clusters, jobs, errors |

Claude Desktop / Cursor

Paste into your MCP client config file to install this server.

{
    "mcpServers": {
        "dataproc mcp server": {
            "dataproc-mcp": {
                "command": "npx",
                "args": [
                    "@dipseth/dataproc-mcp-server@latest"
                ]
            }
        }
    }
}

McpServers

{
    "dataproc-mcp": {
        "command": "npx",
        "args": [
            "@dipseth/dataproc-mcp-server@latest"
        ]
    }
}

npm version
npm downloads
Build Status
Release Status
Coverage Status
License: MIT
Node.js Version
TypeScript
MCP Compatible
semantic-release

A production-ready Model Context Protocol (MCP) server for Google Cloud Dataproc operations with intelligent parameter injection, enterprise-grade security, and comprehensive tooling. Designed for seamless integration with Roo (VS Code).

🚀 Quick Start

Recommended: Roo (VS Code) Integration

Add this to your Roo MCP settings:

{
  "mcpServers": {
    "dataproc": {
      "command": "npx",
      "args": ["@dipseth/dataproc-mcp-server@latest"],
      "env": {
        "LOG_LEVEL": "info"
      }
    }
  }
}

With Custom Config File

{
  "mcpServers": {
    "dataproc": {
      "command": "npx",
      "args": ["@dipseth/dataproc-mcp-server@latest"],
      "env": {
        "LOG_LEVEL": "info",
        "DATAPROC_CONFIG_PATH": "/path/to/your/config.json"
      }
    }
  }
}

Alternative: Global Installation

```bash

No reviews yet — be the first

Sign in to leave a review

Use Google, GitHub, or an email account so ratings stay tied to real people.

Email sign in

No reviews posted yet.