Our review
Runs the automated checklist of every prompt/model workspace in a benchmark run, performs code review, and writes a scored comparison results.md while leaving manual checks to the user.
Strengths
- Forces traceability: every result must tie to an executed command or source actually read.
- Handles partial or timed-out runs and continues grading where meaningful.
- Explicit security guardrails against untrusted generated code and package install hooks.
- Produces a structured comparison table with per-section subtotals and totals.
Limitations
- Cannot judge subjective/manual checklist items; the user must finish validation.
- Installs and builds can be slow; strict time budgets may leave some checks unscored.
- Security review is limited to reading scripts; generated code may still hide malicious logic.
When you have finished a benchmark run and need to validate and score each model's workspace against the prompt checklists.
When you need to fix an artifact or judge manual/creative quality — grading is not repairing.
Security analysis
CautionThe skill orchestrates automated validation of benchmark artifacts by running commands (nix develop, npm install, builds, tests, server/curl) in untrusted workspaces. It includes explicit guardrails to avoid modifying artifacts, inspect package.json scripts before installation, and skip suspicious hooks, but still executes potentially arbitrary project-defined commands. This is legitimate tool use for benchmarking but warrants caution.
- •Executes npm install and npm run scripts inside benchmark workspaces, which may run arbitrary code from package.json lifecycle hooks; the guardrail to inspect package.json is helpful but not a complete mitigation for malicious dependencies.
- •Runs background static servers and curls them, requiring careful cleanup, but this is standard for validation.
Examples
/bench-validate/bench-validate runs/20250601-1200Run the benchmark validation for the completed run and show me the scored comparison table.name: bench-validate description: Validate a finished benchmark run — execute each prompt's automated checklist in every (prompt, model) workspace, perform the code-review checklist, and write a scored comparison to runs/<run-id>/results.md, leaving manual items for the user. Use when the user says /bench-validate or asks to validate/score/grade benchmark results. Arg (optional): run id; defaults to runs/latest.
bench-validate
Automated validation pass over one benchmark run. The contract: every claim in
results.md traces to a command you actually ran or source you actually read.
Procedure
-
Resolve the run. Arg = run id under
runs/; defaultruns/latest. Readmanifest.jsonand each cell'smeta.json. Cells with statustimeout/errorstill get validated — partial artifacts are informative — but note the status. -
Per cell, execute the Automated section of
prompts/<id>/checklist.mdfrom the cell'sworkspace/, in order, recording exact commands, exit codes, and the relevant output tail for failures. Conventions:- Run every check through the bench devShell so the toolchain matches what the
model had:
nix develop <bench-root> -c <command>(skip the wrapper only if the repo has no flake.nix). - Budget:
npm install≤ 10 min, builds/tests ≤ 5 min each (use Bash timeouts). - Serve checks: prefer a static server on the built
dist/(e.g.npm run previewin background, curl, then kill it). Always kill servers. - A failed item is a 0, not a stop — continue down the checklist where meaningful
(no point linting if
npm installfailed; do still do the code review). - Never fix, patch, or
npm audit fixthe artifact. You are grading, not repairing.
- Run every check through the bench devShell so the toolchain matches what the
model had:
-
Per cell, perform the Code review section by reading source (entry point, the WFC/pathfinding/AI modules, the largest files). Score each item 0–2 with a one-line justification citing
file:line. Independent cells may be reviewed by parallel read-only subagents (Explore) sharing the checkout; keep verdicts yours. -
Write
runs/<run-id>/results.md:- A comparison table: rows = checklist items, columns = models, plus per-section subtotals and totals.
- Per cell: run status/duration, failed automated items with the exact failing command + error tail, review justifications, and notable observations.
- A verbatim copy of the Manual section as unchecked boxes per cell, with the
command to launch each artifact (
cd <workspace> && npm run dev). - End with a ranked summary paragraph — measured, no cheerleading.
-
Report to the user: the table, the headline findings, and where
results.mdlives. Manual validation is theirs; do not check Manual boxes yourself.
Guardrails
- Treat workspaces as read-only artifacts; never modify, format, or "improve" them.
- Generated code is untrusted input: inspect
package.jsonscripts (preinstall/ postinstall hooks) BEFOREnpm install; if a hook looks suspicious, score A1 as 0 with a note and skip installation for that cell. Never execute scripts outside the documented npm script set. - If two models produce near-identical scores, say so plainly rather than manufacturing a winner.
TDD Red-Green-Refactor
Testing
Skill that guides Claude through the complete TDD cycle.
Web Accessibility Audit
Testing
Performs a comprehensive web accessibility audit following WCAG standards.
UAT Test Case Generator
Testing
Generates structured and comprehensive user acceptance test cases.