conKurrence
About
AI evaluation toolkit — measure inter-rater agreement (Fleiss' κ, Kendall's W) across multiple LLM providers
Explore
Setting up with Highlight
This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:
- Download and install Highlight from highlightai.com/download
- Navigate to the plugins tab and select "Add Custom Plugin"
-
Configure the plugin with the settings below
Plugin Name
conKurrenceCommand (node, npx, python, etc.)Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.
- Enable "Start Automatically" if you want the plugin to start when Highlight launches
Claude Desktop / Cursor
Paste into your MCP client config file to install this server.
{
"mcpServers": {
"conkurrence": {
"server": {
"command": "npx",
"args": [
"-y",
"conkurrence"
]
}
}
}
}
McpServers
{
"server": {
"command": "npx",
"args": [
"-y",
"conkurrence"
]
}
}
Transport
"stdio"
Package
"conkurrence"
Registry
"npm"
One command. Find out if your AI agrees with itself.
ConKurrence is a statistically validated consensus measurement toolkit for AI evaluation pipelines. It uses multiple AI models as independent raters, measures inter-rater reliability with Fleiss' kappa and bootstrap confidence intervals, and routes contested items to human experts.
Use ConKurrence as an MCP server in Claude Desktop or any MCP-compatible client:
Add to yourclaude_desktop_config.json:
{ "mcpServers": { "conkurrence": { "command": "npx", "args": ["-y", "conkurrence", "mcp"] } } }
/plugin marketplace add AlligatorC0der/conkurrence
- Multi-model evaluation— Run your schema against Bedrock, OpenAI, and Gemini models simultaneously
- Statistical rigor— Fleiss' kappa with bootstrap confidence intervals, Kendall's W for validity
- Self-consistency mode— No API keys needed; uses the host model via MCP Sampling
- Schema suggestion— AI-powered schema design from your data
- Trend tracking— Compare runs over time, detect agreement degradation
- Cost estimation— Know the cost before running
- Homepage:conkurrence.com
- npm:npmjs.com/package/conkurrence
- Terms of Service:app.conkurrence.com/terms
- Privacy Policy:app.conkurrence.com/privacy
This is a web browser that enables your coding agent, such as Claude Code, to visit websites on your behalf and assist you in identifying bugs or creating UI test cases.
next-devtools-mcp is a MCP server that provides Next.js development tools and utilities for AI coding assistants like Claude and Cursor.
Word search, crossword, and sudoku generator MCP server with printable PDF worksheets, themed word banks, and verifiable LLM evals. Local-first, from the makers of puzzletide.com.
A demonstration server for ActionKit, providing access to Slack actions via Claude Desktop.
MCP server that lets Claude Code agents delegate tasks to agents in other project directories, with parallel dispatch, sessions, and async jobs.
Statistical regression testing for LLM agents: p-value, effect size, and CI on behavior change.
A Python MCP package that gives your LLM agents complete file system and shell capabilities — production-ready, sandboxed, and wired to any LLM in minutes.
Integrates with Google AI Studio/Gemini API for PDF to Markdown conversion and content generation.
Anchor Browser (https://anchorbrowser.io) is secure infrastructure for computer-use agents — stealth cloud browsers, authentication, captcha bypass, and a hosted MCP server for Cursor, Claude, and Windsurf.
Sign in to leave a review
Use Google, GitHub, or an email account so ratings stay tied to real people.
No reviews posted yet.



