Benchmark Validation

VerifiedCaution

Executes automated checklists and code review to validate a benchmark run, producing a scored comparison report.

Sby Skills Guide Bot
TestingAdvanced
1508/18/2026
Claude Code
#benchmark-validation#automated-checklist#code-review#results-reporting

Recommended for

Our review

Runs the automated checklist of every prompt/model workspace in a benchmark run, performs code review, and writes a scored comparison results.md while leaving manual checks to the user.

Strengths

  • Forces traceability: every result must tie to an executed command or source actually read.
  • Handles partial or timed-out runs and continues grading where meaningful.
  • Explicit security guardrails against untrusted generated code and package install hooks.
  • Produces a structured comparison table with per-section subtotals and totals.

Limitations

  • Cannot judge subjective/manual checklist items; the user must finish validation.
  • Installs and builds can be slow; strict time budgets may leave some checks unscored.
  • Security review is limited to reading scripts; generated code may still hide malicious logic.
When to use it

When you have finished a benchmark run and need to validate and score each model's workspace against the prompt checklists.

When not to use it

When you need to fix an artifact or judge manual/creative quality — grading is not repairing.

Security analysis

Caution
Quality score88/100

The skill orchestrates automated validation of benchmark artifacts by running commands (nix develop, npm install, builds, tests, server/curl) in untrusted workspaces. It includes explicit guardrails to avoid modifying artifacts, inspect package.json scripts before installation, and skip suspicious hooks, but still executes potentially arbitrary project-defined commands. This is legitimate tool use for benchmarking but warrants caution.

Findings
  • Executes npm install and npm run scripts inside benchmark workspaces, which may run arbitrary code from package.json lifecycle hooks; the guardrail to inspect package.json is helpful but not a complete mitigation for malicious dependencies.
  • Runs background static servers and curls them, requiring careful cleanup, but this is standard for validation.

Examples

Validate default latest run
/bench-validate
Validate a specific run
/bench-validate runs/20250601-1200
Validate and show comparison
Run the benchmark validation for the completed run and show me the scored comparison table.

name: bench-validate description: Validate a finished benchmark run — execute each prompt's automated checklist in every (prompt, model) workspace, perform the code-review checklist, and write a scored comparison to runs/<run-id>/results.md, leaving manual items for the user. Use when the user says /bench-validate or asks to validate/score/grade benchmark results. Arg (optional): run id; defaults to runs/latest.

bench-validate

Automated validation pass over one benchmark run. The contract: every claim in results.md traces to a command you actually ran or source you actually read.

Procedure

  1. Resolve the run. Arg = run id under runs/; default runs/latest. Read manifest.json and each cell's meta.json. Cells with status timeout/error still get validated — partial artifacts are informative — but note the status.

  2. Per cell, execute the Automated section of prompts/<id>/checklist.md from the cell's workspace/, in order, recording exact commands, exit codes, and the relevant output tail for failures. Conventions:

    • Run every check through the bench devShell so the toolchain matches what the model had: nix develop <bench-root> -c <command> (skip the wrapper only if the repo has no flake.nix).
    • Budget: npm install ≤ 10 min, builds/tests ≤ 5 min each (use Bash timeouts).
    • Serve checks: prefer a static server on the built dist/ (e.g. npm run preview in background, curl, then kill it). Always kill servers.
    • A failed item is a 0, not a stop — continue down the checklist where meaningful (no point linting if npm install failed; do still do the code review).
    • Never fix, patch, or npm audit fix the artifact. You are grading, not repairing.
  3. Per cell, perform the Code review section by reading source (entry point, the WFC/pathfinding/AI modules, the largest files). Score each item 0–2 with a one-line justification citing file:line. Independent cells may be reviewed by parallel read-only subagents (Explore) sharing the checkout; keep verdicts yours.

  4. Write runs/<run-id>/results.md:

    • A comparison table: rows = checklist items, columns = models, plus per-section subtotals and totals.
    • Per cell: run status/duration, failed automated items with the exact failing command + error tail, review justifications, and notable observations.
    • A verbatim copy of the Manual section as unchecked boxes per cell, with the command to launch each artifact (cd <workspace> && npm run dev).
    • End with a ranked summary paragraph — measured, no cheerleading.
  5. Report to the user: the table, the headline findings, and where results.md lives. Manual validation is theirs; do not check Manual boxes yourself.

Guardrails

  • Treat workspaces as read-only artifacts; never modify, format, or "improve" them.
  • Generated code is untrusted input: inspect package.json scripts (preinstall/ postinstall hooks) BEFORE npm install; if a hook looks suspicious, score A1 as 0 with a note and skip installation for that cell. Never execute scripts outside the documented npm script set.
  • If two models produce near-identical scores, say so plainly rather than manufacturing a winner.
Related skills