Honey Agent Skill by GreenPT
About
Open-source GreenPT skill that cuts coding-agent output 29% across mixed tasks and up to 70% in focused review workflows.
Explore
The chat edition,](https://github.com/Green-PT/honey-for-devs/blob/HEAD/bench/METHODOLOGY.md)skills/honey-chat/SKILL.md(~500 tokens), is the terse-prose core with the agent-harness levers removed — nothing in it needs tools. Two ways to use it:
- Project custom instructions or a Style (recommended):paste the file in. Instructions become part of the system prompt, so Honey applies toevery message in every conversation— always on, no triggering needed. The prefix is prompt-cached, and the ~500 input tokens are repaid many times over by the halved output.
- Uploaded Skill (paid plans):zip thehoney-chat/folder and upload it as a Skill. Cheaper at rest (only the description stays in context) but loads only when Claude judges it relevant — for an always-on writing style, Project instructions are the better default.
On the API, use the file as (part of) yoursystemprompt. Pin intensity by appending one line:Default to honey ultraorDefault to honey lite.
In a terminal it asks which agents you use, whether to wire the CO₂ badge, drop per-repo rule files, and your default mode — then sets up exactly that. The wizard prompts on/dev/tty, so it works throughcurl | bash. CI/pipes and--yesfall back to auto-detect.
curl -fsSL https://raw.githubusercontent.com/Green-PT/honey-for-devs/main/install.sh | bash
irm https://raw.githubusercontent.com/Green-PT/honey-for-devs/main/install.ps1 | iex
Windows (irm | iex) runs non-interactive; clone and runnode bin/install.jsfor the wizard. Addbash -s -- --yesto skip prompts. Requires Node.js on your PATH. Safe to re-run; skips tools you don't have.
All of these are also handled automatically by the one-line installer. SeeINSTALL.mdfor manual steps, flags, and uninstall.
When Honey is active, the statusline also shows a liveCO₂ estimatefor the session and theCO₂/$ savedvs a no-Honey baseline:
🍯 honey:full · 🌿 44g CO₂ (saved ~26g · $0.18)
(Illustrative — a ~2k-output-token Opus session.) The estimate is a faithful port ofEcoLogitsv0.8.2 (verified to match the package exactly).Model params come from EcoLogits' own registry(hooks/eco-models.json, exported byscripts/build-eco-models.py) — matched by exact id, falling back to a per-family alias for frontier models too new for the registry.Grid switches per provider— Anthropic on AWS Trainium (~500 gCO₂/kWh), OpenAI on Azure (~400), Google on GCP (~330). Aliases, grids, and per-mode savings live inhooks/eco-config.json.
The badge itself rendersonly in Claude Code(it reads Claude Code's transcript, where every model is a Claude model). The provider switching matters forscripts/eco_report.py, which runs against any transcript — Codex/Gemini CLIs would each need their own statusline hook to show a live badge there.
Params arespeculative— Anthropic discloses none. EcoLogits' raw coefficient is asingle-stream (batch-size-1) upper bound— it gives one request the whole GPU set for the full generation (for Opus, ~1.9 tok/s, ~30× slower than reality), which alone is ~1.4 kg per 1M output tokens. Production serves many requests concurrently, so the badge divides that ceiling by an effective batch concurrency (serving_concurrency, default 32 — calibrated so modeled throughput matches real ~50–70 tok/s serving) to show realisticservedimpact.eco_report.pyprints both the served figure and the single-stream ceiling. Treat these as a range, not a meter reading.
For the full breakdown (usage + embodied + primary energy) run the real package:
pip install ecologits python scripts/eco_report.py # newest session, or --transcript PATH
honey-usage(bin/usage.js, inspired bytokscale) reads the session data your coding agents already write to disk and reportsactual token usage— tokens, approximate USD, and served CO₂ — per app and model. Zero dependencies, no network, nothing leaves your machine.
Apps without data are skipped; adding another is a small scanner returning{app, model, ts, input, output, cacheRead, cacheWrite, cost}records.
honey-usage # table by app + model, totals row honey-usage --json # same aggregation as JSON honey-usage --daily --since 2026-08-01 # per-day breakdown, date-filtered honey-usage --client codex,opencode --today # scope by app and local day
APP MODEL INPUT OUTPUT CACHE-R CACHE-W USD CO2 claude claude-opus-5 85,540 4,296,591 1,814,741,458 44,864,917 $1295.62 94.45kg ...
- Dedup— Claude Code repeats assistant records across retries and continuations; each(message.id, requestId)counts once, globally.
- Cache-aware cost— rates frombench/pricing.json(cache writes/reads billed as multipliers on the input rate; unknown models fall back to_default, so treat $ as approximate). Codex'scached_input_tokensare split out ofinput_tokensand priced as cache reads; OpenCode rows use the app's own recorded cost.
- CO₂— the same served EcoLogits estimate as the badge (hooks/eco.js), from output tokens; the badge's caveats apply.
- Savings are ledger-gated— the default report has no "saved" column: it shows what was actually spent, and app logs don't record whether Honey was active.honey-usage --savingsclaims savingsonlyfor sessions the SessionStart hook logged to$CLAUDE_CONFIG_DIR/.honey-usage-ledger.jsonl(Claude Code, since Honey was installed — history before that is never claimed), and only for models with a committed bench stamp (hooks/eco-config.jsonsavings_provenance). Everything else is footnoted, not estimated. The figures stay modeled counterfactuals (est. modeled from bench/results/… — not measured), same basis as the badge.
The skill is authoredonceinskills/honey/SKILL.md. Every per-platform rule file (andAGENTS.md) is generated from it:
node scripts/build-rules.js # regenerate all rule files node scripts/build-rules.js --check # CI: fail if any copy drifted
The OpenClaw (.openclaw/skills/) and Hermes (.hermes/skills/) skill packages are generated the same way fromskills/; rerunnode scripts/build-openclaw-skills.js/node scripts/build-hermes-skills.jsafter changing a skill.tests/openclaw-skills.test.jsandtests/hermes-skills.test.jsfail if a committed copy is stale.
The carbon-estimation data and coefficients inhooks/eco-models.jsonandhooks/eco.jsare derived fromEcoLogitsand remain under theMPL-2.0. SeeNOTICEfor details.
This is a web browser that enables your coding agent, such as Claude Code, to visit websites on your behalf and assist you in identifying bugs or creating UI test cases.
Create crafted UI components inspired by the best 21st.dev design engineers.
Bring agent evaluations, observability, and synthetic test set generation directly into your IDE for free with Galileo's new MCP server
An MCP server to help AI assistants to answer questions and generate AccelByte Extend SDK code more effectively .
MCP server for AI Diagram Maker — generate beautiful software engineering diagrams directly inside Cursor, Claude Desktop, Claude Code, or any MCP-compatible AI agent
ALAPI MCP Tools,Call hundreds of API interfaces via MCP
ESON is lossless, for handoffs where every row matters.CCR(Compress-Cache-Retrieve) is the lossy-but-recoverable lever for the opposite case: a long uniform array you must read but mostly skim — logs, scan results, event streams. It keeps an informative sample (endpoints, anomalies/change-points, head/tail), caches the dropped rows locally, and leaves a<<ccr:HASH N_rows_offloaded>>sentinel. Nothing is lost —retrieverestores the original by hash on demand.
some-tool | eson crush # → sampled view + sentinel; originals cached in .honey-ccr/ eson retrieve <hash> # → the full original array, verbatim
Validated on a 90-row log (opus-4.8 + gpt-5.5):−82% tokens, crushed-only96%answer accuracy,100%with retrieve — and the lone crushed miss was a refusal, not a hallucination. Benches:npm run bench:ccr(tokens) andnpm run bench:ccr:comprehension(quality). Thehoney-ccrskill tells the agent when to reach for it.
Known limitation (upstream):Claude Code builds affected by[anthropics/claude-code#68951(a regression present since ~2.1.121, still open) ignore a PostToolUse hook'supdatedToolOutputfor the built-in Bash tool. On those versions the entry-time hook runs and stashes the original, but the model still receives the raw uncompressed output — honey warns once at session start when it detects an affected version. Piping explicitly (some-tool | eson crush) is unaffected: compression happens before the output leaves the tool. Separately, the hooks needNode >= 14on the PATH Claude Code spawns them with — desktop-app sessions inherit the launchd PATH, not your shell profile, so a stale/usr/local/bin/nodeis common; the hook now warns instead of failing silently.
Open-source GreenPT skill that cuts coding-agent output 29% across mixed tasks and up to 70% in focused review workflows.
Write less code and say less about it.Honey (I Shrunk the AI) byGreenPTis a cross-tool coding skill that cuts AI coding-agent token usage and LLM API costs — making agents emit less codeandless prose without losing correctness. It works withClaude (claude.ai and the API), Claude Code, Cursor, GitHub Copilot, Codex, Gemini CLI, Windsurf, Cline, OpenClaw, oh-my-pi, Kiro, Kilo Code, and Hermes Agent. Three independent levers, applied reflexively:
- Less code— YAGNI first. Walk a ladder (does it need to exist? → stdlib → language native → existing dependency → one line → minimum block) and stop at the first rung that works. The cheapest line is the one you never write.
- Less prose— drop the wind-up, the hedging, the narration of code that already speaks for itself. Answer first.
- Denser agent-to-agent handoffs— when the reader is another agent, not a human, hand it the most token-efficient format it parses losslessly (compact / columnar JSON, orESON). Cuts handoff size ~in half at zero loss of recovery. Fires only here — never as a user-facing answer.
Honey combines whatPonytail(minimal code) andCaveman(terse prose) do separately, then goes further:
- Auto-intensity—lite/full/ultrachosen reflexively from the request, with no deliberation tax (it never spends reasoning tokens decidinghowto comply — that would defeat the purpose on reasoning models).
- Safety carve-outs— input validation, error handling, auth, secrets, migrations, deletes, and anything you explicitly asked for arenevercompressed. Lazy ≠ broken.
- A skill family, not one prompt— an always-on core plus on-demand satellites (review, eco, gain, compress) and ahiveof read-only subagents that return compressed handoffs. SeeSkills & subagents.
Volume is cost. In agentic coding sessions, the volume of generated code and prose is what runs up the bill — and most of it is waste.
This repo ships areproducible benchmark(bench/) so you don't have to take the numbers on faith: 23 tasks across three kinds of work — baseline vsCavemanvsPonytailvs Honey — same model, same prompts, only the skill changes. Correctness is objective (unit tests, structural / accessibility checks, and lossless round-trip recovery for agent handoffs); quality is scored by a4-model cross-family judge panel(median of Opus 4.8 + Sonnet 4.6
- Haiku 4.5 + GPT-5.5) under aneutral rubricthat says nothing about length, so a terse skill gets no thumb on the scale. The figures below are the committed results (Claude Opus 4.8, 3 runs each) — runcd bench && npm run benchto reproduce.
Every number is apaired per-task deltavs baseline — runs collapse by median, tasks pair up, and the figure is the median of those paired deltas with a two-sided Wilcoxonp. Not a ratio of arm totals: that is dominated by whichever task happens to be longest, and it is how token-saving tools end up publishing numbers nobody can reproduce. Endpoints and the run ladder are pre-registered inbench/METHODOLOGY.md.
OnClaude Opus 5(23 tasks × 3 runs, 207 cells, zero refusals or truncation —full-opus5-lean):
Honey is the only arm with no failing cell — the no-skill baseline fails four. And the cut islarger on the newer model, not smaller: −71% LOC on Opus 5 against −39% on Opus 4.8. That runs against the 2026 prompting guidance that newer models need less instruction, which we tested directly and rejected — seeMETHODOLOGY.md.
A single blended number hides the story, because the levers fire differently per task type. Honey on Opus 4.8, where the full competitor set was run —Δ LOCmeasures Lever 1 directly,Δ outputmeasures the tokens (codeandthe prose around it):
Against the competitors on the whole suite (judge win/loss/tie by exact sign test):
- Code— the deepest cut (−39%) at 100% unit-test pass. Ponytail's mandatory self-checkinflatestrivial code (+60% on Opus, +92% on GPT-5.5).
- User-facing— the carve-out keeps Honey from compressing polish: the output delta here is astatistical tie, and Honey holds the only 100% accessibility pass while Ponytail drops to 81% on the structural/a11y checklist.
- Agent-to-agent— under adversarial relay queries (ordinal, nested, absence, cross-field count) Honey is theonly variant that stays 100% losslesswhile roughly halving handoff size; Caveman and Ponytail compress harderandlose recovery (67% / 50%). Its biggest, cleanest win — on 2 tasks, so no p-value.
- Quality is a tie overall(p=0.648) — fewer tokens at no measurable quality cost, not higher quality. But the whole-suite tie is two opposing effects cancelling: on Opus, Honeywins user-facing 6/0/1 (p=0.031)andloses the code judge 2/11/1 (p=0.022)— on tasks where every variant passes 100% of the unit tests, so that is a stylistic penalty for terseness, not a correctness one. Neither effect replicates on GPT-5.5 (p=0.375 / p=1.000), so treat the code-judge dip as suggestive, not established. Caveman's judgemeanalso ties baseline exactly — but paired, it loses 16 of 23 tasks (p=0.004). Means hide that; sign tests don't.
- The dollar saving is unproven at this sample size.−21% on Opus is p=0.104 — not significant on 23 tasks. Output volume is down; the bill is not yet a claim.
The output cut holds on GPT-5.5 (−20%, p=0.004; full two-provider table inbench/README.md), but therecost comes out +14% (ns)because no prompt caching engaged in that arm, so every task paid the skill prompt fresh. Honey is the only variant with no test regressions across all three tiers on Opus.
End-to-end agentic measurement (Cline harness)
npm run benchmakesone API callper task — clean for isolating the output lever, but it never exercises an agent loop, tool schemas, or multi-turn context growth, where a real agent's token bill actually lives.bench/src/cline-bench.js(npm run bench:cline) runs each taskthroughtheClineCLI headless, so the measured tokens are end-to-end agentic — harness prompt and every loop iteration included. Honey is injected as a Clinerule, recommended as the per-turn-cheapskills/honey/cline-rule.md(the operational core; the fullSKILL.mdre-sent every turn inflates input). Seebench/README.md.
ESON — Efficient Structured Object Notation
Honey includesESON, a zero-dependency, schema-first format for agent handoffs. Repeated record keys are emitted once; declared row counts catch truncated messages; JSON-compatible cells preserve types. ESON is developed in its own repo —Green-PT/honey-eson: the normative spec, JS + Python reference implementations, conformance vectors, the canonical LLM primer, the Honey Wire Profile, and negotiation. Honey vendors the codec in[eso/.
…
Sign in to leave a review
Use Google, GitHub, or an email account so ratings stay tied to real people.
No reviews posted yet.



