Data Dictionary MCP

by jonahkeegan

276 downloads
Not rated
GitHub

About

A Model Context Protocol (MCP) server that coordinates AI agents to transform database tables into Wikipedia-style data dictionaries.

Details

Author
jonahkeegan
Downloads
276
Categories
Search, AI

- Multi-Format Support: JSON, CSV, and Plain Text files
- AI-Powered Analysis: generate field descriptions and relationships
- MCP Integration: coordinate AI agents via the protocol
- Schema Extraction: unify schemas from various formats
- Wikipedia-Style Output: familiar, accessible presentation format

Setting up with Highlight

This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:

  1. Download and install Highlight from highlightai.com/download
  2. Navigate to the plugins tab and select "Add Custom Plugin"
  3. Configure the plugin with the settings below
    Plugin Name Data Dictionary MCP
    Command (node, npx, python, etc.)

    Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.

  4. Enable "Start Automatically" if you want the plugin to start when Highlight launches

From the repository

Clone the repository, create a Python 3.9+ virtual environment, install dependencies from requirements.txt, then run python src/main.py.

Claude Desktop / Cursor

Paste into your MCP client config file to install this server.

{
    "mcpServers": {
        "data dictionary mcp": {
            "data-dictionary-mcp": {
                "command": "python",
                "args": [
                    "-m",
                    "venv",
                    "venv"
                ]
            }
        }
    }
}

McpServers

{
    "data-dictionary-mcp": {
        "command": "python",
        "args": [
            "-m",
            "venv",
            "venv"
        ]
    }
}

Data Dictionary MCP

A Model Context Protocol (MCP) server that coordinates AI agents to transform database tables into Wikipedia-style data dictionaries.

Overview

The Data Dictionary MCP project automates the conversion of various database formats into comprehensive, human-readable data dictionaries using AI-powered analysis and description. It leverages the Model Context Protocol (MCP) to coordinate AI agents for analyzing, describing, and verifying database structures.

Features

- Multi-Format Support: Process JSON, CSV, and Plain Text files (with more formats planned)
- AI-Powered Analysis: Generate field descriptions and identify relationships
- MCP Integration: Coordinate AI agents using the Model Context Protocol
- Schema Extraction: Extract database schemas from various formats into a unified representation
- Wikipedia-Style Output: Present data dictionaries in a familiar, accessible format

Project Status

This project is in active development. See the Project Roadmap for details.

Getting Started

Prerequisites

- Python 3.9+
- Git
- pip or poetry for dependency management

Installation

1. Clone the repository:

   git clone https://github.com/jonahkeegan/data-dictionary-mcp.git
cd data-dictionary-mcp

2. Create a virtual environment:

   python -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate

3. Install dependencies:

   pip install -r requirements.txt

4. Run the application:

   python src/main.py

Project Structure

data-dictionary-mcp/
├── docs/                  # Documentation
├── src/                   # Source code
│   ├── mcp/               # MCP server components
│   ├── analyzers/         # Format analyzers
│   ├── agents/            # Agent coordination
│   └── dictionary/        # Dictionary generation
├── tests/                 # Test suite
├── memory-bank/           # Cline memory bank
├── .gitignore
├── .clinerules            # Cline rules
├── README.md
└── requirements.txt

Project Roadmap

Milestone 1: MCP Server Foundation and Format Analyzers

- Implement MCP server with basic tool definitions - Develop format analyzers for JSON, CSV, and Plain Text - Create schema extraction system - Implement unit tests for core components

Milestone 2: AI Agent Coordination and Field Description

- Implement agent coordination system - Develop field description generation - Create task distribution and result aggregation - Add integration tests

Milestone 3: Content Verification and Publishing

- Implement content validation - Develop Wikipedia-style formatting - Create export capabilities - Add end-to-end tests

Milestone 4: User Interface and Deployment

- Develop web interface - Implement search capabilities - Add user feedback system - Create deployment infrastructure

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

License

This project is open source and available under the MIT License.

No reviews yet — be the first

Sign in to leave a review

Use Google, GitHub, or an email account so ratings stay tied to real people.

Email sign in

No reviews posted yet.