Agentic Wallet Guardian

by rudimentall1

Not rated
GitHub

About

Self-hosted decision engine that sits between AI agents and blockchain execution, returning ALLOW/WARN/BLOCK before any transaction is signed.

Details

Author
rudimentall1
Categories
Finance

Setup

Install Agentic Wallet Guardian in your MCP client (Claude Desktop, Cursor, Windsurf, and others).

Repository: https://github.com/rudimentall1/agentic-wallet-guardian-v3

Follow the installation instructions in the repository README, then restart your MCP client.

A self-hosted decision engine that sits between an AI agent and blockchain execution.Agents submit a proposed action, Guardian returns an explainable ALLOW / WARN / BLOCK before anything gets signed or broadcast.

POST /decision -> ALLOW / WARN / BLOCK (with a reasoned explanation)

It runs on your own infrastructure, using your own policy rules and your own reputation data - seeWhy self-hostedfor why that matters and how this differs from calling a hosted security API directly.

There are good hosted alternatives for agent-transaction security (GoPlus's AgentGuard, Blockaid, Chainalysis/TRM for compliance). If you just want a risk score and don't care who sees the query, calling one of those directly is less work than running this. Guardian exists for the cases where that tradeoff doesn't work for you:

- Nothing about which wallets, contracts, or amounts your agents touch leaves your infrastructure.Threat-intel and contract allow/deny checks are local JSON files you populate yourself (seedata/threat_lists/README.md), not a lookup call to a third party. A hosted API inherently sees every address and amount you ask it about.
- Your policy rules live in your code, not a vendor's dashboard.Spending caps, reputation gates, and which action types require confirmation are plain Python inguardian/policy/, reviewable and changeable without waiting on anyone else's product roadmap.
- No per-call fees or rate limits imposed by someone else- only the ones you configure for your own users (GUARDIAN_RATE_LIMIT_PER_MINUTE).
- No vendor lock-in.Every external data source (RPC endpoint, Blockscout instance, DexScreener) is swappable behind a small provider interface - see
Architecture.

The honest tradeoff going the other way: you also take on running it, keeping your local threat lists current, and you don't get a hosted vendor's chain coverage or dedicated threat-research team for free. This is the right choice for teams that specifically need data sovereignty or deep policy customization - not a strict upgrade over every hosted option.

AI Agent | v Action Intent { agent_id, wallet, chain, action_type, target, amount, metadata } | v ┌───────────────────────────────┐ │ Guardian Decision Engine │ ├───────────────────────────────┤ │ 1. Hard Rules │ <- chain support, sanity checks │ 2. Wallet Intelligence │ <- mock | real RPC (web3.py) │ 3. Token Intelligence │ <- mock | real DexScreener | real GoPlus │ 4. Contract Intelligence │ <- local lists, then mock | real Blockscout | real GoPlus │ 5. Simulation │ <- mock | real eth_call dry-run (see below) │ 6. Threat Intelligence │ <- local JSON allow/deny lists │ 7. Anomaly Detection │ <- vs. this agent's own history (see below) │ 8. Policy Engine │ <- spending caps, reputation gates │ 9. Risk Fusion │ <- signals -> single 0-100 score │ 10. Reputation Adjustment │ │ 11. Explanation │ <- evidence -> human-readable reasons └───────────────────────────────┘ | v ALLOW / WARN / BLOCK | v Blockchain Execution

Every data source in steps 2-4 is a small provider interface with a mock implementation (zero config, zero network calls) and a real one, selected per-source by environment variable - see.env.example. Switching from demo mode to a real deployment is a config change, not a code change.

guardian/ config.py GuardianConfig - the one place that reads os.environ core/ ActionIntent, Signal, Decision, EvaluationContext (zero external dependencies - no pydantic/FastAPI) decision/ DecisionEngine (orchestrator), RiskFusionEngine, hard rules reasoning/ explanation + confidence builders intelligence/ wallet/ analyzer.py + providers.py (mock | RpcWalletDataProvider) token/ analyzer.py + providers.py (mock | DexScreenerTokenDataProvider | GoPlusTokenDataProvider) contract/ analyzer.py + providers.py (mock | BlockscoutContractDataProvider | GoPlusContractDataProvider) simulation/ pre-execution dry-run (mock | real eth_call) + tx_builder.py (real calldata for transfer/approve) goplus_client.py shared GoPlus Token Security API client (used by both contract + token) threat/ blocklist.py (local AddressList) + intelligence.py policy/ PolicyEngine + policy templates (spending caps, reputation gates) reputation/ AgentReputation (score derived from decision history) memory/ storage.py (protocol) + InMemoryStorage + sqlite_storage.py api/ main.py FastAPI app: /decision, /health, /capabilities, /agents/{id}/history, /demo/{scenario} security.py API-key auth dependency + rate-limit middleware schemas.py pydantic request/response models (API boundary only) mcp_server.py MCP stdio server - same DecisionEngine, no HTTP required data/threat_lists/ local, operator-maintained allow/deny lists (empty by default - see its README) scripts/ refresh_ofac_list.py fetch OFAC's public SDN list into the local threat list tests/ 101 tests covering the engine, policy, reputation, and every provider

guardian/is intentionally dependency-free (standard library only, except where a real provider needshttpxorweb3), so the decision core can be unit-tested, embedded in another service, or ported to a different web framework without dragging FastAPI along. Onlyapi/touches pydantic/FastAPI.

This is real, runnable, tested decision infrastructure with real (not mock) data sources available for every signal source - but "available" isn't the same as "flip a switch and trust it blindly." Specifics:

- Wallet (RPC provider):is_contractandtx_count(nonce-based) are reliable with any JSON-RPC endpoint. Walletagerequires an archive-capable node and is off by default (GUARDIAN_RPC_ESTIMATE_AGE=false) - most free public RPC endpoints don't serve historical state, so this fails closed to "unknown" rather than guessing.
- Contract (Blockscout provider):real verification-status lookups against a public Blockscout instance. Their exact response schema and rate limits can change - this is written to degrade to "unknown" on any unexpected response, never to fabricate an answer, but hasn't been load- tested against production traffic.
- Token (DexScreener provider):real liquidity data, but matching a bare ticker symbol to an on-chain pair is inherently ambiguous (many unrelated tokens share a symbol, and scammers deliberately mint look-alikes). The provider picks the highest-liquidity pair on the requested chain and reports its own match confidence rather than presenting a guess as certain - for anything where that ambiguity matters, match by contract address instead of symbol.
- Contract + Token (GoPlus provider):real contract-security (owner-can-drain, mintable, self-destruct, hidden owner) and trading-security (honeypot, buy/sell tax, blacklist, pausable transfers, holder concentration) from GoPlus's Token Security API - meaningfully more signal types than Blockscout/DexScreener give individually, since GoPlus's own static analysis covers both in one call. Two real limits: it only has data for contracts it's actually analyzed (mostly token contracts, not generic dApp/router contracts), andGoPlusTokenDataProviderneeds a contract
address- a bare symbol like "PEPE" can't be resolved and is honestly reported as unverifiable rather than guessed at.
- Sanctioned-address list is real, populated data: 103 addresses (100 EVM + 3 Solana) from OFAC's SDN list, via
0xB10C/ofac-sanctioned-digital-currency-addresses- verified end-to-end (a known-sanctioned address correctly triggersBLOCKthrough the full pipeline) and verified to correctly reflect delistings, not just additions (Tornado Cash's addresses, removed from the SDN list in March 2025, are correctly absent). Re-runscripts/refresh_ofac_list.pyperiodically - sanctions change in both directions.
- malicious_contracts.json/verified_contracts.jsonstill ship empty on purpose(seedata/threat_lists/README.md) - there's no single authoritative source for "malicious contract" the way OFAC's list is authoritative for sanctions, so populating these is a judgment call for whoever operates this instance, not something to seed by default with unverified entries.
- Simulation is real, but conditional.RpcSimulationProvider(GUARDIAN_SIMULATION_PROVIDER=rpc) genuinely dry-runs a transaction viaeth_call/eth_estimateGasagainst current chain state - a revert comes back with its actual reason, not a guess, and ERC-20approve()amounts are decoded from real calldata instead of inferred. This activates when the caller supplies raw calldata viaintent.metadata
["data"], OR - new - whenGUARDIAN_TX_BUILDER=rpcis also set and the intent is a plaintransferorapprove(see next bullet). Aswapintent with no transaction built yet still has nothing to dry-run - Guardian reports that honestly (simulation_not_attempted) rather than guessing.
- Transaction building closes part of that gap, deliberately not all of it.RpcTransactionBuilder(GUARDIAN_TX_BUILDER=rpc) turns a semantictransfer/approveintent into real calldata - it fetches the token's actualdecimals()via RPC rather than assuming 18 (a wrong assumption there would scale the amount by orders of magnitude), and deliberately has no hardcoded token-address registry: a bare symbol like "USDC" is refused rather than guessed at, since a wrong address here wouldn't just be a bad risk signal, it'd be an artifact that could end up in a real transaction.swapis built against Uniswap V2 Router02 only (one immutable, well-known contract - function selectors computed locally viaWeb3.keccak, not copied from memory) - realgetAmountsOut()on-chain quote, caller-suppliedmax_slippage_bpsrequired (never a default, same reasoning as decimals above).bridgeis a genuinely open-ended L2/bridge routing problem in general - dozens of protocols, wildly different trust models - but this module handles one well-scoped slice of it: L1 -> L2 deposits through a destination chain's own official OP Stack bridge (currently: Base and Optimism -depositETHTo/depositERC20ToonL1StandardBridge, both addresses independently cross-checked - Base against Etherscan's label plus basehub.org, Optimism against the official ethereum-optimism/superchain-registry plus a second independent dev-tool config - before being hardcoded). L2 -> L1 withdrawals are NOT built - that's a genuinely different, much slower proof/challenge-window flow, not a variant of the deposit call. Bridging to anywhere else, or via any non-canonical bridge, returnsNonerather than guessing.
- Storage:InMemoryStorage(default, zero setup),SQLiteStorage(GUARDIAN_STORAGE_BACKEND=sqlite- persists across restarts, no external infra), orPostgresStorage(GUARDIAN_STORAGE_BACKEND=postgres+GUARDIAN_POSTGRES_DSN- the fit for multiple replicas behind a load balancer, where SQLite's single-writer model becomes the bottleneck;pip install -r requirements-postgres.txt). Tested against a real local Postgres instance, not mocked - seetests/test_postgres_storage.py. Redis is still open if you specifically want it; the two-methodMemoryBackendinterface is small enough to implement against anything.
- API auth/rate-limitingare intentionally minimal - built for one self-hosted instance behind your own network boundary, not a multi-tenant gateway. Put a real API gateway in front if you need that.
- Not security-audited.The policy engine and risk fusion logic have not been reviewed by anyone outside this repo. TreatBLOCKas a strong signal, not a guarantee, until that's happened.

Everything downstream of aSignal- fusion, policy, reputation, explanation, the API - doesnotneed to change as any of the above gets hardened further. That boundary is the actual design contract here.

Zero-config demo mode (mock providers, in-memory storage, no auth):

pip install -r requirements.txt uvicorn api.main:app --reload
curl http://localhost:8000/demo/safe curl http://localhost:8000/demo/unknown curl http://localhost:8000/demo/malicious
curl -X POST http://localhost:8000/decision \ -H "Content-Type: application/json" \ -d '{ "agent_id": "trading-agent-001", "wallet": "0x742d35Cc6634C0532925a3b844Bc454e4438f44e", "chain": "ethereum", "action_type": "swap", "from_token": "ETH", "to_token": "USDC", "amount": 5 }'

Going from demo to a real self-hosted deployment

At minimum for a real deployment: setGUARDIAN_API_KEY(auth is off by default),GUARDIAN_STORAGE_BACKEND=sqlite(persistence), and whicheverGUARDIAN__PROVIDERvariables you want pointed at real data instead of mock - see the comments in.env.examplefor every option, andRpcWalletDataProvider/BlockscoutContractDataProvider/DexScreenerTokenDataProvider/GoPlusContractDataProvider/GoPlusTokenDataProvider's docstrings for what each one actually gives you.

For agent frameworks that speak MCP (LangChain, CrewAI, Claude Desktop, etc.),mcp_server.pyexposes the same decision engine as two tools (evaluate_action,get_agent_history) over stdio - installrequirements-mcp.txtalongsiderequirements.txt(they resolve into one environment; see the comment at the top ofrequirements-mcp.txt) and point your MCP client atpython mcp_server.py.

pip install -r requirements.txt -r requirements-chain.txt pytest -q

requirements-chain.txt(web3) is only needed for the RPC-provider tests; the rest of the suite runs with justrequirements.txt. Theguardian/core has no external dependencies beyond that, so it's also runnable with:

PYTHONPATH=. python3 -m unittest discover -s tests -v

CI (.github/workflows/ci.yml) runs the full suite on every push/PR against Python 3.11 and 3.12.

Every decision this service returns — from a single policy check up through the full pipeline — is signed as anOAA (Open Agent Attestation)token: an Ed25519-signed JWT wrapping the decision, the action, and the reason.

Anyone holding the public key can verify a decision offline, without calling back to whatever instance of Guardian issued it — useful for an auditor, a downstream service, or just a record you want to trust later without trusting the server that produced it.

python examples/example_oaa_attestation.py python examples/example_full_pipeline.py # capability -> intent -> engine -> OAA

The reference OAA implementation is ~150 lines (oaa.py/attestation.pyupstream) and is shared, unmodified, across this project andagent-guardrail— same signing format, same verification path, no per-project fork.

Using Guardian in front of MetaMask Agent Wallet

MetaMask Agent Wallet's Guard Mode / Beast Mode apply the same static spend limits and allowlists to every agent. Guardian is a second, independent check in front of it: doesthisspecific action look right forthisagent, right now — before themmCLI is ever invoked.

skills/guardian-check/is a standardAgent Skill— the same open format MetaMask itself uses formm(npx skills add MetaMask/agent-skills). Install it alongside MetaMask's own skill in any Agent-Skills-compatible runtime (Claude Code, Cursor, Codex, OpenClaw), and the agent will call a running Guardian instance for an ALLOW/WARN/BLOCK decision before running anymmcommand that moves funds —mm send,mm swap,mm bridge,mm perps,mm predict trade,mm earn,mm aave,mm pay.

Guardian never holds keys and never executes anything —mmremains the only thing that signs or broadcasts. This is a decision gate the agent is instructed to consult first, not a modification to MetaMask's own pipeline (there's no public hook for that today).

uvicorn api.main:app --reload # run Guardian locally export GUARDIAN_API_URL="http://localhost:8000" python skills/guardian-check/scripts/check.py \ --agent-id my-agent --wallet 0x... --chain ethereum \ --action-type transfer --target 0x... --amount 50

- Replace the mock wallet/token/contract analyzers with real data sources.Done - seeHonesty about the current statefor what "real" does and doesn't cover yet per source.
- Wire up real pre-execution simulation.Done fortransfer/approveend to end (GUARDIAN_SIMULATION_PROVIDER=rpc+GUARDIAN_TX_BUILDER=rpc- see
Honesty about the current state).swapneeds real DEX routing.Done against Uniswap V2 Router02 - real on-chaingetAmountsOut()quote, explicit caller-suppliedmax_slippage_bps(never defaulted), no calldata built without a real quote. Fixed a real bug found while building this: simulation was dry-running againstintent.target(the recipient/spender encoded
insideERC-20 calldata) instead of the actual contract being called (from_token) - meaning transfer/approve simulation silently "succeeded" against any EOA recipient regardless of whether the real call would have reverted.BuiltTransactionnow carries an explicitto; seetests/test_tx_builder.pyfor the regression tests that would have caught it.bridgestill open.Done for L1->L2 deposits to Base and Optimism via the official OP StackL1StandardBridge(depositETHTo/depositERC20To) - other destinations, other bridge protocols, and L2->L1 withdrawals all remain open; seeHonesty about the current state.
- Populate threat-intel / sanctions feeds; stop shipping empty sets.Done for sanctions (sanctioned_addresses.json- 103 real OFAC SDN addresses, refreshable viascripts/refresh_ofac_list.py).malicious_contracts.json/verified_contracts.jsonremain empty by design - no single authoritative source exists to seed them the way OFAC's list does for sanctions.
- SwapInMemoryStoragefor a persistent backend.SQLiteStorageis available;a Postgres/Redis backend is still open for multi-replica deployments.PostgresStoragedone - tested against a real local Postgres instance (tests/test_postgres_storage.py), same two-methodMemoryBackendinterface as the other backends. Redis remains open if specifically wanted.
- Add an MCP server wrapper.Done (mcp_server.py). A packaged Python/TypeScript SDK on top of the REST API is still open.
- Publish an OpenAPI spec and a hosted demo endpoint.
- Get the policy engine and risk fusion reviewed/audited before anyone relies on aBLOCKfrom this service in production - it's a security tool, so it needs the same scrutiny it applies to others.
- Add per-agent capability limits (delegation scoping).Done -guardian/policy/capabilities.py, and wired intoDecisionEngine.evaluate()via an optionalcapability_registryconstructor argument (previously it was a standalone module you had to call yourself outside the normal pipeline - seeexamples/example_capability_limits.py, which now runs throughDecisionEnginedirectly). Opt-in: pass no registry (the default) and nothing changes; an operator can grant a specific agent a scoped capability (allowed action types, allowed chains, per-action and daily spending caps, an expiry) with zero private-key material involved. Agents with no grant are unaffected. Real key management (session keys, account abstraction) remains deliberately out of scope - a categorically higher-stakes problem.
- Verify declared intent against decoded simulation results.Partially done -guardian/decision/intent_verification.pycatches the case where an agent declares one amount but the actual calldata it was handed encodes a meaningfully different (but still finite) one, andDecisionEngine.evaluate()now actually calls it (it didn't before - the module and its example script existed, but nothing in the real decision pipeline invoked it). It's still not a working guardrail on its own, though: comparing atomic units needs the token'sdecimals(), and no decimals provider exists yet, so the engine currently calls this withtoken_decimals=None- everyapprovewith a successful, finite-amount simulation gets an honest "cannot verify without decimals" WARN instead of either a false BLOCK or a silent skip. Seeexamples/example_intent_verification.pyfor the check actually blocking a real mismatch once decimals are supplied.
- Flag actions that deviate from an agent's own historical pattern.Done -guardian/intelligence/anomaly/analyzer.py. Distinct from reputation (a single trust score) and policy (static, operator-set limits): this compares the current intent against
this specific agent's*own recorded history - new action type, new chain, or an amount that's a statistical outlier versus what this agent has done before, even if it's within policy limits and the agent's reputation is fine. Honestly reports "insufficient history" rather than guessing a baseline from fewer than 5 prior data points - seetests/test_anomaly_detection.py.
- Sit in front of a real agent wallet, not just accept intents from a generic API caller.Done for MetaMask Agent Wallet -skills/guardian-check/is a standard Agent Skill an agent installs alongside MetaMask's ownmmskill; the agent calls it before running any fund-movingmmcommand and only proceeds on ALLOW. Tested end-to-end against a liveuvicorninstance (ALLOW/WARN/BLOCK/config-error all exercised for real, not just asserted) - see the "Using Guardian in front of MetaMask Agent Wallet" section above. No public hook exists (yet) to run inside MetaMask's own pipeline; this works at the agent-orchestration layer instead.

Same author, same principle applied elsewhere:

- agent-guardrail- a generic policy firewall for AI agent tool calls (not blockchain-specific). Published on PyPI, MIT, 46 tests.
-
x402-attest- cryptographically signed (Ed25519), independently verifiable attestations for agent-to-agent payment policy decisions. Early proof of concept.
-
open-agent-attestation- vendor-neutral open spec (JWT+EdDSA) for signing agent policy decisions, verifiable by anyone. x402-attest above uses a custom format; this is the generalized version. Draft v0.1.

Bridge Town is an MCP-native, git-versioned financial modeling platform for FP&A teams and finance leaders. AI agents use Bridge Town tools to create projects, write Python model files, run models in isolated cloud sandboxes, query data, write outputs to Google Sheets, create dashboards, branch scenarios, and collaborate with teammates.

The Capital.com MCP Server lets your AI assistant talk to your trading account directly. Market data, position checks, trade previews – all in plain language, without leaving your AI tool.

Coinrule Agentic Trading MCP enables investors to create, backtest, execute, and manage trading agents through natural language across stocks, crypto and ETFs

Invest with Claude and other AI assistants

Australian Consumer Data Right Product Data

Remote MCP server for historical crypto & prediction-market data: search ~500K instruments, live market stats (OHLC, turnover, spreads, depth, slippage) and tick-data purchase. Keyless for catalog & stats; optional OAuth for account tools. Endpoint: https://cryptostruct.com/mcp

Cross-border debt collection from your AI assistant: check cases, get pricing, submit new cases.

Read-only MCP server for your Evibe investment portfolio + live market data (holdings, performance, dividends, benchmarks, screeners). Works with Claude & ChatGPT.

Financial and quantitative modeling engine for AI agents. Typed, named, deterministic.

No reviews yet — be the first

Sign in to leave a review

Use Google, GitHub, or an email account so ratings stay tied to real people.

Email sign in

No reviews posted yet.