ponytail Skill Review: 6,637 Bytes of Always-On YAGNI, and a Disputed Cost Claim

The README of DietrichGebert/ponytail leads with "~54% less code (up to 94%) · ~20% cheaper · ~27% faster · 100% safe." The thing doing that work is one 6,637-byte markdown file. The repo around it has 149,188 stars, 8,020 forks, and install paths for 20 agents. On August 4 a user posted a controlled A/B on Sonnet 5 that found ponytail no cheaper than no skill at all, and 6.9% more expensive than a one-sentence prompt. Those two claims don't fit together. The gap between them is the most useful thing to understand before you install it. I read the skill, the hooks that inject it, the issue tracker, and two Hacker News threads. I didn't run it through our API.

TL;DR

  • ponytail's core skills/ponytail/SKILL.md is 6,637 bytes at v4.10.0 (commit e3ba2aa, 2026-09-14). The ruleset its hooks inject at the default full level is 5,252 bytes, about 1,313 estimated tokens at bytes ÷ 4.
  • DietrichGebert/ponytail had 149,188 GitHub stars, 8,020 forks, and 319 open issues on 2026-10-01. It's MIT-licensed and ships 6 skills, or 12 SKILL.md files counting the OpenClaw mirrors.
  • The Claude Code and Codex plugin injects that ruleset at every session start, resume, clear, and compact, and into every subagent by default. It behaves as resident context, not an on-demand skill.
  • The README reports −54% lines of code and −20% cost versus no skill on Haiku 4.5 (12 tasks, n=4). A user's A/B in issue #685 on Sonnet 5 found +0.7% cost versus no skill and +6.9% versus a one-sentence prompt, while still cutting lines of code by 20%.
  • The sharpest complaint, in issue #660: a user removed ponytail after two weeks because the agent "silently decided not to implement functional parts of features I had explicitly requested."

What the skill actually says

The core is a seven-rung "ladder": does this need to exist, is it already in this codebase, does the stdlib do it, does a native platform feature cover it, does an installed dependency solve it, can it be one line, and only then the minimum code. The agent stops at the first rung that holds. Around that sit rules against unrequested abstractions, an output format ("Code first. Then at most three short lines"), three intensity levels (lite, full, ultra), and a "When NOT to be lazy" list that protects trust-boundary validation, data-loss error handling, security, and accessibility.

"ACTIVE EVERY RESPONSE. No drift back to over-building."

"Shortest working diff wins — but only once you understand the problem."

"Lazy code without its check is unfinished."

The history explains the hedges. The "never lazy about understanding the problem" section and the root-cause bug-fix paragraph came in PR #253 (merged 2026-06-22), the same day issue #245 was filed under the title "Dangerously lazy." The maintainer's note on that fix: telling the model in prose to "trace the flow end to end" "did nothing in testing (0/3)." What worked was an operational rule, grep every caller and fix the shared function once. The file has had 13 commits and hasn't changed since 2026-07-10.

Five smaller skills ship alongside it. The most useful is ponytail-review, an over-engineering-only review that emits one line per finding with tags like delete:, stdlib:, and yagni:, and ends with net: -N lines possible.

What it costs you in context

All of this is arithmetic from file size (bytes ÷ 4), not a measured bill:

PathBytesEstimated tokens (÷4)When it loads
skills/ponytail/SKILL.md6,637~1,659Skill invocation; 937 B of it is frontmatter
Injected ruleset, full (hook output, frontmatter stripped)5,252~1,313Every session start/resume/clear/compact, and every subagent
Injected ruleset, lite / ultra5,225 / 5,290~1,306 / ~1,322Same, at the other levels
skills/ponytail-review/SKILL.md2,383~596On demand
skills/ponytail-help, -gain, -debt, -audit2,796 / 1,973 / 1,703 / 1,652~699 / ~493 / ~426 / ~413On demand
AGENTS.md (instruction-only fallback)2,593~648Always on, for harnesses that auto-load it
.openclaw/skills/ponytail/SKILL.md5,957~1,489OpenClaw copy

I got the injected sizes by reproducing the hook's mode filter in my own Python rather than running the repo's Node code. The hook scripts themselves (30,528 bytes across six .js files) run as processes. They aren't read into context.

Size isn't the real issue. Residency is. The issue #685 author argues that resident text gets re-read on every API call inside every turn, so the cost scales with session length. They measured the SessionStart payload at 11,158 characters, about 2.8k tokens, which is more than double my 1,313 estimate. I didn't reproduce their run, so I can't say where the extra bytes come from. Hook framing and the one-time statusline note (below) are candidates. Read my figure as a floor.

The hooks, briefly

The skill directory itself is one markdown file, read in full. The plugin adds Node lifecycle hooks, which I read without executing. They write a mode flag file, read settings.json to check for a statusline, and print the ruleset. A grep of hooks/*.js found no URLs, fetch, child_process, or spawn calls; the only matches were the word "spawn" in two code comments in ponytail-subagent.js. Two things to know. On the first session without a statusline, the hook tells the agent to "Proactively offer to set this up for the user on first interaction," which means an edit to your settings.json (it fires once). And open issue #824 says hook commands "still interpolate plugin roots into shell text."

What the community reports

SourceDateWhat was reportedLink
GitHub issue #1262026-06-16Colin Eberhardt found the original benchmark's baseline had no harness-style system prompt. Adding "Provide just one example for any given task, and no commentary or usage examples." cut baseline average LOC on Haiku from 108 to 16, against 8.25 for ponytail. The maintainer replied "You're right on both counts" and rebuilt the benchmark as agentic Claude Code runs.issue 126
GitHub issue #6852026-08-04RoxsLee's A/B (Claude Code 2.1.220, Sonnet 5, 5 tasks, n=8): cost +0.7% versus no skill (95% CI −6.5% to +9.8%) and +6.9% versus a one-sentence "leave a runnable check" prompt (p < 1e-4). They also note −20% LOC, −12% output tokens, and the only arm whose safety pass rate never dropped. These are their nominal list-price figures. Open, no replies after 58 days.issue 685
GitHub issue #2452026-06-22jayf0x reported that on Sonnet 4.6 the skill made the model skip the decision process under ambiguity and patch the wrong layer. The model's own post-mortem, which they quoted: "That's ponytail applied to the wrong layer." Fixed the same day by PR #253.issue 245
GitHub issue #6602026-07-31 / 08-15faangbait showed shortest-diff pressure leaking policy into callers on a config-loader task. sebthom removed ponytail after two weeks, calling it "a senior Perl developer obsessed with saving keystrokes." Fix PR #836 is open and unmerged.issue 660
GitHub issues #120 / #8872026-06-16 / 09-18xaccefy said the agent stamped ponytail: markers "every time ponytail is active" (17 comments; narrowed by PR #577 on 07-10). viktorianer reports there's still no opt-out for codebases that ban comments, and the rule is copied into 11 other files.issue 887
Hacker News, two threads2026-06-14 / 09-07The launch thread (98 points) ranged from "Is this the new leftpad?" to kamphey's "incredibly useful to have these basic 5 heuristics." In September, jghn, a user of "a couple of months": "it is hit or miss... There's no sense of nuance of context applied."HN 48527946 · HN 49592706

Dissent runs both ways. The #685 author writes that the point is "not 'ponytail is bad'" and credits it with less code and steady safety. The maintainer has been quick to respond elsewhere: #126 conceded and rebuilt within two days, #245 fixed the same day. One more open report, #745, says the always-on "Read fully, then be lazy" rule wedged agents mid-read on very large markdown files. It's one report with no replies, so weigh it accordingly.

How to install it

These commands come from the README at v4.10.0:

# Claude Code (send as two separate prompts)
/plugin marketplace add DietrichGebert/ponytail
/plugin install ponytail@ponytail

# Codex (then open /hooks and trust its two lifecycle hooks)
codex plugin marketplace add DietrichGebert/ponytail
codex plugin add ponytail@ponytail

# OpenClaw
clawhub install ponytail
clawhub install ponytail-review

# Gemini CLI
gemini extensions install https://github.com/DietrichGebert/ponytail

The plugins need node on the non-interactive shell's PATH. For Cursor rules, Windsurf, Cline, Kiro, and others, the README says to copy the matching rules file or AGENTS.md. To tune it: PONYTAIL_DEFAULT_MODE (off/lite/full/ultra) and PONYTAIL_SUBAGENT_MATCHER, a regex that limits which subagent types get the ruleset.

Where it earns its keep

New code where agents over-build. The maintainer's per-task numbers show the date picker dropping from 404 to 23 lines and the color picker from 287 to 23, because the agent reaches for a native <input>. The never-simplify list and the one-runnable-check rule are real counterweights, not decoration. The ponytail-debt ledger makes deliberate shortcuts greppable, which even #887's critic calls "one of the better ideas in this plugin." And ponytail-review is the cheapest way to get the value: on demand, about 596 estimated tokens, explicitly scoped to over-engineering only.

Where it doesn't

The objective function is my main criticism. "Shortest working diff wins" is still the tiebreaker, and the qualifiers around it ("only once you understand the problem," "anything explicitly requested") are prose. The maintainer's own finding in #245 was that a prose instruction to trace the flow did nothing until it became an operational rule. So I'd treat the remaining prose guards as hopes until someone tests them. #660's silently dropped features are what it looks like when they fail.

Second, it's always on by design, subagents included. A skill whose first rung is "Does this need to exist at all?" loads itself into read-only search agents unless you set the matcher. The README's own default: "Unset means inject into every subagent."

Third, the cost headline is model-specific. "~20% cheaper" comes from Haiku 4.5, n=4, one repo, and the writeup's limitations section says "One model. Haiku 4.5 only." The README admits cost "can go the other way (on GPT-5.5 it does)." #685 found no saving on Sonnet 5. Read it as "cheaper on Haiku, on this benchmark."

Verdict: if your agent over-builds new code, try ponytail-review or the 2,593-byte AGENTS.md first. Run the full plugin in lite, or scope subagents, for maintenance work in a mature codebase. If your project instructions already say "one solution, no commentary, reuse existing helpers," #126 and #685 suggest a one-line prompt gets you much of the reduction more cheaply. ponytail still wrote less code in both comparisons.

FAQ

What does the ponytail skill do?

It makes the agent climb a seven-rung ladder (skip it, reuse it, stdlib, native, installed dependency, one line, minimum code) and stop at the first rung that works. Validation, data-loss handling, security, and accessibility are off limits for simplification.

How much context does ponytail use?

SKILL.md is 6,637 bytes (~1,659 estimated tokens). The hook injects a 5,252-byte ruleset (~1,313 estimated tokens) per session and per subagent at the default level. File-size arithmetic, not a measured run.

Does ponytail make my agent cheaper?

On Haiku 4.5 the maintainer reports −20% cost. A user's Sonnet 5 A/B in #685 found +0.7% versus no skill and +6.9% versus a one-sentence prompt. It depends on the model.

How do I turn ponytail off or limit it?

"stop ponytail," "normal mode," or /ponytail off. Use /ponytail lite, PONYTAIL_DEFAULT_MODE, and PONYTAIL_SUBAGENT_MATCHER to scope it.

Skill by Dietrich Gebert and contributors. License: MIT, per the repo LICENSE and the skill's frontmatter. Repo: github.com/DietrichGebert/ponytail. Source read at commit e3ba2aa (v4.10.0). Review date: 2026-10-01. Community reception shifts over time, and #660, #685, and #887 were still open when this was written, so check the linked issues for their current state.