Forge
About
GPU kernel optimization - 32 swarm agents turn PyTorch into fast CUDA/Triton kernels on real datacenter GPUs with up to 14x speedup
Explore
Setting up with Highlight
This MCP is not yet compatible with Highlight’s one-click setup. However, you can still use it with Highlight by following these steps:
- Download and install Highlight from highlightai.com/download
- Navigate to the plugins tab and select "Add Custom Plugin"
-
Configure the plugin with the settings below
Plugin Name
ForgeCommand (node, npx, python, etc.)Please refer to the README for specific instructions on how to obtain API keys or other required environment variables.
- Enable "Start Automatically" if you want the plugin to start when Highlight launches
forge_auth
Authenticate with the RightNow Forge GPU kernel optimization service. Opens the user's browser to sign in via the RightNow dashboard. Required before using any other Forge tools. If valid tokens already exist, verifies them without opening the browser.
forge_credits
Check your current Forge credit balance. Each single-GPU optimization costs 1 credit. Credits are only charged for successful optimizations (speedup >= 1.1x). Credits are lifetime (not monthly) and can be purchased on the dashboard.
forge_optimize
Submit PyTorch code to Forge for GPU kernel optimization. Forge generates optimized Triton or CUDA kernels, evaluates them on real H100/A100 GPUs, and returns the best result with speedup metrics over torch.compile(max-autotune). USE THIS TOOL WHEN YOU SEE: - PyTorch custom operations (torch.autograd.Function, custom forward/backward) - Manual CUDA kernels that could be faster - Performance-critical tensor operations (attention, convolution, normalization, softmax) - Code with comments like "slow", "bottleneck", "optimize", "performance" - torch.compile() targets or triton.jit kernels - Any nn.Module with significant compute in forward() - Matrix multiplication, reduction, or scan operations - Custom loss functions with reduction operations - Fused operation opportunities (e.g., LayerNorm + activation) The tool streams real-time progress and blocks until optimization completes (1-10 minutes). Cost: 1 credit per optimization (only charged if speedup >= 1.1x). Requires authentication via forge_auth first.
forge_generate
Generate an optimized GPU kernel from scratch based on a specification. Use this when you need to create a new high-performance kernel without existing PyTorch code. Forge will generate a PyTorch baseline first, then optimize it into Triton or CUDA. Requires authentication via forge_auth first. Cost: 1 credit per generation.
forge_status
Check the current status of a running or completed Forge optimization job.
forge_cancel
Cancel a running Forge optimization job. Credits may be refunded if less than 20% of iterations completed.
forge_sessions
List past Forge optimization sessions with their results. Useful for checking what has already been optimized.
- forge_auth: Authenticate with the RightNow Forge GPU kernel optimization service.
Opens the user's browser to sign in via the RightNow dashboard. Required before using any other Forge tools.
If valid tokens already exist, verifies them without opening the browser.
- forge_credits: Check your current Forge credit balance. Each single-GPU optimization costs 1 credit.
Credits are only charged for successful optimizations (speedup >= 1.1x).
Credits are lifetime (not monthly) and can be purchased on the dashboard.
- forge_optimize: Submit PyTorch code to Forge for GPU kernel optimization.
Forge generates optimized Triton or CUDA kernels, evaluates them on real H100/A100 GPUs,
and returns the best result with speedup metrics over torch.compile(max-autotune).
USE THIS TOOL WHEN YOU SEE:
- PyTorch custom operations (torch.autograd.Function, custom forward/backward)
- Manual CUDA kernels that could be faster
- Performance-critical tensor operations (attention, convolution, normalization, softmax)
- Code with comments like "slow", "bottleneck", "optimize", "performance"
- torch.compile() targets or triton.jit kernels
- Any nn.Module with significant compute in forward()
- Matrix multiplication, reduction, or scan operations
- Custom loss functions with reduction operations
- Fused operation opportunities (e.g., LayerNorm + activation)
The tool streams real-time progress and blocks until optimization completes (1-10 minutes).
Cost: 1 credit per optimization (only charged if speedup >= 1.1x).
Requires authentication via forge_auth first.
- forge_generate: Generate an optimized GPU kernel from scratch based on a specification.
Use this when you need to create a new high-performance kernel without existing PyTorch code.
Forge will generate a PyTorch baseline first, then optimize it into Triton or CUDA.
Requires authentication via forge_auth first. Cost: 1 credit per generation.
- forge_status: Check the current status of a running or completed Forge optimization job.
- forge_cancel: Cancel a running Forge optimization job. Credits may be refunded if less than 20% of iterations completed.
- forge_sessions: List past Forge optimization sessions with their results. Useful for checking what has already been optimized.
Claude Desktop / Cursor
Paste into your MCP client config file to install this server.
{
"mcpServers": {
"forge": {
"server": {
"command": "npx",
"args": [
"-y",
"@rightnow/forge-mcp-server"
]
}
}
}
}
McpServers
{
"server": {
"command": "npx",
"args": [
"-y",
"@rightnow/forge-mcp-server"
]
}
}
Transport
"stdio"
Package
"@rightnow/forge-mcp-server"
Registry
"npm"
This is a web browser that enables your coding agent, such as Claude Code, to visit websites on your behalf and assist you in identifying bugs or creating UI test cases.
Bring agent evaluations, observability, and synthetic test set generation directly into your IDE for free with Galileo's new MCP server
An MCP server to help AI assistants to answer questions and generate AccelByte Extend SDK code more effectively .
MCP server for AI Diagram Maker — generate beautiful software engineering diagrams directly inside Cursor, Claude Desktop, Claude Code, or any MCP-compatible AI agent
Official MCP server for Buildable AI-powered development platform. Enables AI assistants to manage tasks, track progress, get project context, and collaborate with humans on software projects.
What Shopify did for ecommerce, Chipp does for AI agents. Build, deploy, and monetize AI agents for your business — no engineering team required.
CodeVF MCP lets AI hand off problems to real engineers instantly, so your workflows don’t stall when models hit their limits.
Make your AI agent speak every language on the planet, using Lingo.dev Localization Engine.
Client implementation for Mastra, providing seamless integration with MCP-compatible AI models and tools.
Instead of direct calling MCP tools, mcpcode server transforms MCP tool calls into TypeScript programs, enabling smarter, lower-latency orchestration by LLMs.
Sign in to leave a review
Use Google, GitHub, or an email account so ratings stay tied to real people.
No reviews posted yet.



