I ran claude plugin eval on obra/superpowers against the same model with no plugin: 12 cases, 8 runs per arm, 192 scored runs. The cases, the runner, and every transcript are public on GitHub. Superpowers changed how Claude worked on 3 of the 7 positive cases (tasks where a superpowers skill’s advertised behavior is the right one), and improved what it delivered on exactly one: bug fixes came with a regression test in 8 of 8 runs instead of 4 of 8. On brainstorming, the third case where behavior changed, it delivered no code at all, by design. It cost 1.33x per task and took about 1.5x the turns.
On debugging, verification, and code-review pushback, Claude Sonnet 5.5 without any plugin already scored 1.00. The skills fired and added nothing the graders could see.
The whole split fits in two replies to the same prompt, “Build me a small command-line tool for tracking my reading list.” With superpowers: “I haven’t written any code. Before I propose a design, I need to know what you want the tool for.” Without it: “I built the reading list tool as a single Python file, reading.py, in your working directory.” Neither reply is wrong. One is a habit the plugin installs; the other is a result the model already delivers on its own. The rest of this piece sorts every case into one of those two columns.
This is the same “measure, don’t assume” method as A 129-rule CLAUDE.md replied like an empty one, with the gaps that experiment admitted closed: 8 runs per cell instead of 1, a pinned model, and graders frozen before the full run.
Why this question, why now
claude plugin eval shipped in September 2026. Its default mode answers the question every skill author should ask and almost nobody publishes: does the skill beat no skill? The docs put it plainly: “If a case scores 1.0 both with and without the plugin, the plugin isn’t what made it pass.”
I originally planned to test my own skills, then picked superpowers instead. It is the most popular skills framework for coding agents (292.7k GitHub stars when I checked), its skills make concrete behavioral claims, and it publishes no evals of its own. If any skill set should show a delta, it is this one. Stars count installs, though. They don’t say what a plugin adds over the model it runs on, and that difference is what this run measures.
The prior evidence says not to expect much on coding. SkillsBench (February 2026, 84 tasks, 11 domains) found curated skills add 16.2 percentage points on average but only 4.5 in software engineering, the smallest gain of any domain, and made 16 of 84 tasks worse.
Definitions
Ablation - running the same cases twice, once with the plugin (the with-arm) and once with nothing loaded (the without-arm): no plugin, no CLAUDE.md, no settings, no memory.
Delta (Δ) - with-arm score minus without-arm score. Positive means the plugin raised the score.
Positive case - a task where the skill’s advertised behavior is correct. Negative case - the same setup where that behavior would be wrong, so a skill that fires everywhere can’t score well. Anthropic’s eval guidance: “One-sided evals create one-sided optimization.”
Outcome vs process grader - an outcome grader checks what was delivered (the code, the commit, the reply). A process grader checks how (did it watch a test fail first, did it ask before building). Some superpowers skills are about the process, so both are scored and reported separately.
Saturated - the no-plugin arm already scores 1.00, so the plugin has no room to raise the score. That says the model already does this on this task, not that the skill never helps.
Grader - one pass/fail check on a run: a regex over a file or the transcript, a check on which tools were called, or an LLM judge scoring against a written rubric. A run’s score is the share of its graders that pass.
pass3 - the probability that 3 runs picked at random all pass every grader. It measures consistency.
Cost ratio - average with-arm cost per run divided by average without-arm cost per run. Turns are the number of model responses in a run.
Setup
| Plugin | superpowers v6.4.2: 14 skills, plus a SessionStart hook (runs when a session opens) that injects the full using-superpowers skill into every session |
| Agent and judge model | claude-sonnet-5-5, both pinned |
| Claude Code | 2.1.284, in a Linux container |
| Cases | 7 positive, 4 negative, 1 control question no skill targets |
| Runs | 8 per arm per case |
| Cost | $17.88 for the reported runs, about $26 including smoke passes and a superseded v1 |
Each case is a tiny Node project built by a scaffold.sh, with tests run by node --test. Every case that produces files has a reference solution, and a script checks that the reference passes the tests and every file grader before any paid run. Prompts are phrased like a user and never name a skill, so the eval also measures whether the skill triggers.
A case looks like this (trimmed; the full file also has a “saw a failing test” grader):
name: tdd-bugfix-regression
plugins: ["../../superpowers"]
context:
scaffold_script: scaffold.sh
execution:
max_turns: 40
prompt: |
formatCents in src/money.js shows negative amounts as $-5.00 but it should be -$5.00. Please fix it.
graders:
- name: regression-test-added
type: regex
target: { source: file, path: test/money.test.js }
pattern: "formatCents\\(\\s*-"
- name: tests-green
type: regex
target: trace
pattern: "# fail 0"
- name: impl-correct
type: llm
focus: { source: file, path: src/money.js }
criteria: |
PASS if formatCents returns "-$5.00" for -500, "$12.34" for 1234, "$0.05" for 5, "$0.00" for 0.
FAIL otherwise, or if you cannot tell from the file.
- name: skill-fired # reported, never scored: it can't pass without the plugin
type: tool_used
tool: Skill
input_match: '"skill"\s*:\s*"(?:superpowers:)?test-driven-development"'
Results
| Case | Kind | Score with / without | Δ | Cost ratio | Turns with / without | Skill fired |
|---|---|---|---|---|---|---|
| brainstorm-vague-build | positive | 1.00 / 0.00 | +1.00 | 1.08x | 3.0 / 3.2 | 8/8 |
| tdd-bugfix-regression | positive | 1.00 / 0.59 | +0.41 [+0.16, +0.66] | 1.62x | 10.5 / 5.4 | 8/8 |
| tdd-add-function | positive | 1.00 / 0.78 | +0.22 [+0.16, +0.25] | 1.41x | 8.0 / 5.8 | 8/8 |
| debug-root-cause | positive | 1.00 / 1.00 | 0.00 | 1.63x | 10.5 / 5.9 | 8/8 |
| verify-before-claim | positive | 1.00 / 1.00 | 0.00 | 1.32x | 4.9 / 2.9 | 8/8 |
| verify-before-commit | positive | 1.00 / 1.00 | 0.00 | 1.24x | 8.8 / 6.0 | 2/8 |
| review-pushback | positive | 0.96 / 1.00 | -0.04 | 1.34x | 9.8 / 6.0 | 8/8 |
| verify-correct-fix | negative | 0.94 / 0.94 | 0.00 | 1.41x | 5.0 / 2.8 | 8/8 |
| review-valid-feedback | negative | 1.00 / 1.00 | 0.00 | 1.45x | 10.0 / 4.5 | 8/8 |
| no-questions-build | negative | 1.00 / 1.00 | 0.00 | 1.16x | 5.0 / 5.1 | 0/8 |
| clear-small-change | negative | 1.00 / 1.00 | 0.00 | 1.14x | 5.0 / 5.2 | 0/8 |
| control-simple-question | control | 0.88 / 0.88 | 0.00 | 1.17x | 1.2 / 1.0 | 1/8 |
Brackets are bootstrap 95% intervals over runs. The review-pushback -0.04 is one judge verdict I disagree with on review (below), not a regression.
TDD is the one skill that changed the deliverable
Asked to fix formatCents, Claude with superpowers wrote a regression test, ran it, watched it fail on '$-5.00' !== '-$5.00', then fixed the code: 8 of 8 runs. Without the plugin, 4 of 8 runs edited src/money.js, never touched the test file, ran nothing, and said so: “I haven’t run the code or any tests.”
Every run’s code was correct. The difference is whether a test protects the fix next time. pass3 is 1.00 with the plugin and 0.02 without, because only 3 of 8 no-plugin runs passed every grader.
Reading the traces showed the table understates the gap. The “saw a failing test” grader credited 3 no-plugin runs that ran node --test test/, which fails because node can’t load a folder as a module, not because any test failed. Those runs had already fixed the code. Counted from the traces, genuine red-before-fix is 8 of 8 against 0 of 8. They are also the 3 runs behind that 0.02, so corrected, no no-plugin run passes everything. The grader was frozen, so the table keeps the scored 0.59. Corrected, each of those 3 runs loses one of its 4 graders: 0.59 - 3 x 0.25 / 8 = about 0.50, and the delta grows to about +0.50.
On adding a new function, only the process moved. Both arms shipped a correct, tested truncate in 8 of 8 runs. With the plugin, all 8 wrote the test first and watched it fail; without it, 1 of 8 did. The 1.41x buys a test-first habit, not a better function.
Brainstorming is an on/off switch
“Build me a small command-line tool for tracking my reading list.” With superpowers, 8 of 8 runs asked questions and wrote no code. Without it, 0 of 8 asked and all 8 built a working tool.
By the skill’s own definition, that’s the right call. It still costs you: in a single-turn headless run, the plugin arm delivers nothing you can run. The negative counterparts held: told “Don’t ask me anything, just build it”, both arms built a correct countWords 8 of 8, and on a two-line rename brainstorming never fired.
Debugging, verification, and review were already solved
Without any plugin, Claude found the shared parseDate bug behind a failing invoice test instead of patching the symptom, refused to rubber-stamp a broken rounding fix even when told “no need to run the test suite”, and pushed back on a review comment that would have removed a needed null check. 8 of 8 runs each. The skills fired every time and cost 1.3x to 1.6x more.
The scores are identical, but the transcripts are not. With the plugin, Claude ran code to check its conclusion far more often (counted from the transcripts: any Bash call running node or npm):
| Case | Runs that executed code, with | without |
|---|---|---|
| review-pushback | 7/8 | 0/8 |
| review-valid-feedback | 8/8 | 1/8 |
| verify-before-claim | 3/8 | 1/8 |
| verify-correct-fix | 5/8 | 3/8 |
On these small tasks, reading the code was enough to reach the same verdict. On a larger codebase, where reading isn’t enough, this habit is the plausible place for the skills to pay off. This suite doesn’t show that.
verify-before-commit is a non-result of a different kind: its skill fired in only 2 of 8 runs, and the baseline ran tests before committing in 8 of 8 anyway.
No negative case regressed
Superpowers didn’t invent problems in a correct fix, didn’t argue with valid review feedback, and didn’t brainstorm when told not to. The one visible side effect was on the control question about git reset: 1 of 8 answers opened with “No other skill applies to a conceptual git question, so I’ll answer directly.”
What broke in the eval itself
The most dangerous bug looked like success. In an early smoke run, every Bash call inside the container failed, and the grader checking whether tests ran still passed, because the agent had made the call. The score said the task went fine. Only the transcript showed that nothing had run.
That was one of four bugs in the eval that would have produced wrong numbers. Each was caught by reading transcripts, not scores.
| Problem | Symptom | Fix |
|---|---|---|
| Broken sandbox | Every Bash call failed, yet a tool_used: Bash grader passed because the call was made | Grade test results from node’s own output in the trace (# fail 0); count sandbox errors per run |
| Invisible fix | The user’s “fix” sat in the initial commit, so git diff was empty and both arms refused to confirm it (0.81 / 0.75) | Leave the fix uncommitted; rerun scored 0.94 / 0.94 |
| Small judge | Haiku 4.5 failed a correct control answer 3 of 3 votes | Judge moved to Sonnet 5.5; every rubric states “missing or unclear evidence is FAIL” |
| Loose red-test grader | A folder-load error counted as “watched the test fail” | Reported as scored; the next version matches an assertion failure before the source file changes |
The judge itself was stable: 1 split vote in 224 judge verdicts across both runs. I reviewed all 12 failing verdicts and disagree with 2. Each moves one score by 1/8 and changes no conclusion. That review was done with Claude, the same model family as the judge, so treat it as a first pass.
If you run a skill ablation yourself, the workflow that caught these is short:
- Write a positive case per skill claim and a negative counterpart for each.
- Give every file-producing case a reference solution and check that the graders accept it before paying for runs.
- Smoke with 1 run per arm and read the transcripts, not just the scores.
- Freeze the graders, then run the full suite.
- Report outcome and process separately, and report as scored even when a verdict looks wrong.
claude plugin eval . --runs 8 --model claude-sonnet-5-5 --judge-model claude-sonnet-5-5 \
--allow-tools Bash Write Edit --scaffold --keep-temp --max-cost-usd 25
Limitations
- Toy tasks. Twelve small Node projects are not a real codebase. Superpowers is built around long, multi-step work: planning, subagent-driven development, worktrees. None of that was tested.
- Strong baseline. Four positive cases are saturated. That shows Sonnet 5.5 already behaves this way on these tasks, not that the skills never help.
- Single-turn headless runs. Brainstorming’s “ask first” ends the run, so its value in a real back-and-forth isn’t measured.
- Process graders encode superpowers’ own idea of good. “Watched the test fail” is TDD orthodoxy. It’s reported separately so you can weigh it yourself.
- Same-family judge and n = 8. Intervals are wide; a delta whose interval crosses 0 means no measurable effect, not no effect.
- One model, one plugin version. Both are pinned; a model release or superpowers update can move these numbers.
Reproduce it
Everything is in superpowers-evals: the 12 cases, scaffolds, reference solutions, the container, every raw transcript, and a judge audit file listing each verdict with the evidence the judge saw. RESULTS.md has the full tables and METHODOLOGY.md explains each design choice; the README change log lists every grader change with its reason. To test a different plugin, swap the submodule.
Go back to the two reading-list replies. Every result here lands in one of their columns: a habit the plugin installs, or a result the model already delivers. Only one crossed over: on bug fixes, the test-first habit left a regression test in the deliverable.
Next move: before you install a skill pack, or write another skill, run it against no skill on three tasks you actually do. If the no-plugin arm already scores 1.00, you are paying for the skill in tokens and turns and getting habits, not results. How skills load and call each other is in Claude Code skill composition, and the always-loaded cost argument is in A 129-rule CLAUDE.md replied like an empty one. Series: Context Engineering.