conKurrence

by alligatorc0der

Not rated yet
GitHub

About

AI evaluation toolkit — measure inter-rater agreement (Fleiss' κ, Kendall's W) across multiple LLM providers

Explore

Setting up with Highlight

This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:

  1. Download and install Highlight from highlightai.com/download
  2. Navigate to the plugins tab and select "Add Custom Plugin"
  3. Configure the plugin with the settings below
    Plugin Name conKurrence
    Command (node, npx, python, etc.)

    Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.

  4. Enable "Start Automatically" if you want the plugin to start when Highlight launches

Claude Desktop / Cursor

Paste into your MCP client config file to install this server.

{
    "mcpServers": {
        "conkurrence": {
            "server": {
                "command": "npx",
                "args": [
                    "-y",
                    "conkurrence"
                ]
            }
        }
    }
}

McpServers

{
    "server": {
        "command": "npx",
        "args": [
            "-y",
            "conkurrence"
        ]
    }
}

Transport

"stdio"

Package

"conkurrence"

Registry

"npm"

One command. Find out if your AI agrees with itself.

ConKurrence is a statistically validated consensus measurement toolkit for AI evaluation pipelines. It uses multiple AI models as independent raters, measures inter-rater reliability with Fleiss' kappa and bootstrap confidence intervals, and routes contested items to human experts.

Use ConKurrence as an MCP server in Claude Desktop or any MCP-compatible client:

Add to yourclaude_desktop_config.json:

{ "mcpServers": { "conkurrence": { "command": "npx", "args": ["-y", "conkurrence", "mcp"] } } }
/plugin marketplace add AlligatorC0der/conkurrence

- Multi-model evaluation— Run your schema against Bedrock, OpenAI, and Gemini models simultaneously
- Statistical rigor— Fleiss' kappa with bootstrap confidence intervals, Kendall's W for validity
- Self-consistency mode— No API keys needed; uses the host model via MCP Sampling
- Schema suggestion— AI-powered schema design from your data
- Trend tracking— Compare runs over time, detect agreement degradation
- Cost estimation— Know the cost before running

- Homepage:conkurrence.com
- npm:
npmjs.com/package/conkurrence
- Terms of Service:
app.conkurrence.com/terms
- Privacy Policy:
app.conkurrence.com/privacy

This is a web browser that enables your coding agent, such as Claude Code, to visit websites on your behalf and assist you in identifying bugs or creating UI test cases.

next-devtools-mcp is a MCP server that provides Next.js development tools and utilities for AI coding assistants like Claude and Cursor.

Word search, crossword, and sudoku generator MCP server with printable PDF worksheets, themed word banks, and verifiable LLM evals. Local-first, from the makers of puzzletide.com.

A demonstration server for ActionKit, providing access to Slack actions via Claude Desktop.

MCP server that lets Claude Code agents delegate tasks to agents in other project directories, with parallel dispatch, sessions, and async jobs.

Statistical regression testing for LLM agents: p-value, effect size, and CI on behavior change.

A Python MCP package that gives your LLM agents complete file system and shell capabilities — production-ready, sandboxed, and wired to any LLM in minutes.

Integrates with Google AI Studio/Gemini API for PDF to Markdown conversion and content generation.

Anchor Browser (https://anchorbrowser.io) is secure infrastructure for computer-use agents — stealth cloud browsers, authentication, captcha bypass, and a hosted MCP server for Cursor, Claude, and Windsurf.

No reviews yet — be the first

Sign in to leave a review

Use Google, GitHub, or an email account so ratings stay tied to real people.

Email sign in

No reviews posted yet.