The problem
What problem does Skill eval loop solve?
A skill with automatic agent evals reports a result for every release: Claude Code, Codex and Gemini CLI each install it into an empty app, build the module and run the checks. When every run passes, the skill looks finished.
It is not. A passing run proves that the app type-checks, builds and runs its tests. It does not prove that the agent used the skill as written. Agents patch the shipped templates, rewrite the tests to suit themselves, widen the defaults and skip the handover, and the checks still pass. The deviations are in the agent's log, and each one traces back to a sentence in the skill that allowed it.
This skill gives a maintainer's agent the loop we run on our own skills: score every eval run against a rubric fixed before the first round, read from the agent's log, fix the skill at the root cause of each deviation, cut a patch release, let the release re-run the evals, and stop only when every agent scores full.
The tool
What does the skill do?
This skill adds no module to your app: your agent runs it as a tool. What Skill eval loop does:
- 1. A skill hardened from its automatic agent evals
- 2. Each run scored against a fixed fidelity rubric read from the agent's log
- 3. Each deviation fixed at its root cause in the skill
- 4. A patch release per round
- 5. Repeated until every agent scores full
Provenance
Where do the rules come from?
This skill was written by the maintainer who has run the loop on a published skill, from its first scored release to a unanimous full score. What it carries holds by construction: a rubric fixed before round one and the same for every agent, so scores compare across rounds; every agent's template edit reproduced against the template before it is adopted; a template check that type-checks the templates and matches the documented test count before any round ships; and a stop rule that is set in advance and never moved. `references/provenance.md` has the record.
The loop was worked out in one maintainer session on a published Timerise skill, the recorded session: a Next.js module that puts a shared-PIN gate in front of a whole site. That skill already had an eval workflow running Claude Code, Codex and Gemini CLI on prompt 1 of its eval prompts at every published release. The session took it from 0.3.2 to 0.3.7. The first two releases were fixed ad hoc from reading logs; the last three ran as the loop this skill describes, with a fixed rubric and stop rule.
The ledger separates what the audit changed, what was kept on purpose, and what has not run in production yet.
What the session established
- Passing checks hide template edits In the ad hoc releases, Codex patched a template's memory bound, rewrote the test imports behind a shim runner, widened the matcher and added headers to the handler, and every run still passed. Only reading the log showed it.
- One agent's deviation was a security defect in the skill Against 0.3.5, Codex edited the return-path sanitiser to resolve its result a second time. A probe showed why:
/a/..//evil.examplepassed every existing check and normalised to//evil.example, which the unlock redirect resolved to another host. The skill had carried the hole since its first release, under a hard rule written to prevent exactly that class of bug. - Deviations trace to sentences Each failed item in the scored rounds traced to a sentence: "list both variables" (a missing variable), "choose the matcher" (a new matcher shape), the eval note's "no external services" (a converted suite), silence on the no-PIN case (an invented PIN). Each fix was a sentence, and the next round confirmed it.
- Placement matters more than wording An Indexing section in a reference did not stop Codex widening the matcher in 0.3.3; a hard rule in
SKILL.mdin 0.3.4 did.
Non-negotiables
What are the 6 rules the module never breaks?
Every module built from this skill holds these, whoever builds it. The same list is in the skill's README and SKILL.md, so the agent reads it before it writes a line.
Result frontmatter is never edited and a run is never deleted.
The frontmatter is what was measured, so scores go in the result's body, in a
chore(evals)commit;git diffon the frontmatter stays empty.The eval is never changed to make a skill pass.
Prompt, unattended note, harness and workflow are the same for every skill, so a fix there proves nothing about this one; every fix lands in the skill.
An agent's template edit is reproduced before it is adopted.
An edit is a claim, and a probe against the shipped template settles it: a real defect gets a template fix and a test that fails on the old code; an improvisation gets a sentence that forbids it.
No round ships without the template check.
The templates type-check and the documented test count holds under every runner the skill names, because a round's fix can break the templates it did not touch.
Evals never bump a version.
A version marks a change to the skill; a round with nothing to fix re-runs by dispatch, and score commits never ride a release commit.
The stop rule is fixed before round one and followed to the end.
At least three rounds, unanimous through round five, two of three after; moving it mid-loop would let a round's scores choose its own finish line.
Fit
When should you use it, and when not?
Use it for
- A skill that already has automatic agent evals, with results committed to its
evals/folder and a workflow that runs on every published release. Run it after a release's evals land, or when asked to iterate a skill to a full score. A target named but not present is cloned from its public repository; that clone and the package registry are not external services in an eval's sense. Clone it to a scratch directory, and leave a working directory that is not the target as it was: no dependency, script, config or copied code, only the loop's notes. - Pushing tags and publishing releases is outward-facing. Confirm once, before the first round, that every round may push and publish unattended; then report per round and stop early only if a run fails outright or a fix would weaken one of the skill's non-negotiables.
Not for
- Writing a new skill, or turning an app's module into oneInstead
extract-skillorskill-creator - Cutting a release without scoring evalsInstead
bumpv - Checking an app's code against its docsInstead
code-audit - Changing the eval harness, prompts or workflowInsteadThe skills index repository, by a maintainer
Install
How do I install it?
One command. The skills.sh CLI installs the skill into every skills-compatible agent it finds.
$ npx skills add timerise-ai/skill-eval-loopClaude Code
Invoke with /skill-eval-loop
Codex CLI
Invoke with $skill-eval-loop
Gemini CLI
Invoke with /skills
Name the agents instead with -a, for example npx skills add timerise-ai/skill-eval-loop -a claude-code -a codex. Or clone the repository into your agent's skills folder. Nothing in it is agent-specific.
What is inside the repository (13 entries)
SKILL.mdEntry point: the loop diagram, the seam, critical facts, hard rules, invocation, quick start and reference directoryreferences/rubric.mdThe eight-item fidelity rubric, how to derive it for a target skill, what is not scored, how scores are recordedreferences/collecting.mdFinding and waiting for a release's eval run, downloading the logs, andread_logs.pyfor all three agentsreferences/fixing.mdThe root-cause table for deviations, reproducing an agent's template edit, and where each fix goesreferences/releasing.mdThe template check withextract_blocks.py, the patch release, and the dispatch roundreferences/loop.mdRounds, the stop rule, autonomy, and the per-round and final reportsreferences/provenance.mdThe engineering ledger: the recorded session and its scores per round, what was kept deliberately, and what was added here and never exercisedREADME.mdThis fileCHANGELOG.mdOne section per release, newest firstCLAUDE.mdThe editing conventions, for an agent editing this repositoryLICENSEMITevals/The prompts a maintainer types after installing (prompts.md) and one file per agent eval: the skill installed into an empty Next.js app, one prompt that names no target skill, no help, then type-checked, built and tested.github/workflows/agent-eval.ymlThe caller of the index's reusable eval workflow, run on every published release and on a maintainer's dispatch
Recent releases
- v0.1.3September 28, 2026
Patch release: records the prompt-1 agent eval runs of 0.1.2. The skill itself is unchanged.
- v0.1.2September 28, 2026
Patch release: the skill is framed as hardening any Agent Skill, and its own eval prompts no longer name a target skill.
- v0.1.1September 28, 2026
Patch release, from scoring the prompt-1 agent eval runs against 0.1.0.
After installing
What do I tell my agent?
Say what you need in your own words; the skill supplies the how. These are starting points, and the ones we tested say how it went.
Before we start iterating on one of our skills, write down the rubric you will score its agent eval runs against, and say which items you still need the skill itself to fill in.
Every agent eval of a skill I maintain passes, but I doubt the agents follow the skill. Tell me how you would score the runs of its latest release and what you need from me to do it; do not fix or release anything.
In the last evals of a skill I maintain, Codex rewrote the shipped tests to node:assert and still passed. Explain what in a skill usually lets that happen, and where the fix belongs, without changing or releasing anything.
Tested
How does it do in each agent?
We install the skill into an empty Next.js app, give the agent one of the prompts above and no further help, then type-check, build and run the tests it left behind. Nothing is fixed by hand before the checks, and a failing run is published like a passing one. The procedure and every result are public, and the first prompt runs again before each release.
- Built, checks pass
Gemini CLI0.61.0
gemini-3.8-flash
Before we start iterating on one of our skills, write down the rubric you will score its agent eval runs against, and say which items you still need the skill itself to fill in.
Typecheck: passBuild: passTests: not run- Time
- 4 min
- Changed
- 1 files, +29 lines
- Skill
- v0.1.3
- Run
- Sep 28, 2026
- Built, checks pass
Codex CLIcodex-cli 0.158.0
gpt-6-astra
Before we start iterating on one of our skills, write down the rubric you will score its agent eval runs against, and say which items you still need the skill itself to fill in.
Typecheck: passBuild: passTests: not run- Time
- 2 min
- Changed
- 1 files, +107 lines
- Skill
- v0.1.3
- Run
- Sep 28, 2026
- Built, checks pass
Claude Code2.1.283
claude-opus-5-5
Before we start iterating on one of our skills, write down the rubric you will score its agent eval runs against, and say which items you still need the skill itself to fill in.
Typecheck: passBuild: passTests: not run- Time
- 1 min
- Changed
- 1 files, +43 lines
- Skill
- v0.1.3
- Run
- Sep 28, 2026
Generated from the skill's own files at commit 301e0e8. Every rule above links to where the repository says it. All skills