Notre avis
Exécute la checklist automatisée de chaque espace de travail (prompt/modèle) d'un run de benchmark, effectue la revue de code et produit un comparatif scoré dans results.md, en laissant les vérifications manuelles à l'utilisateur.
Points forts
- Impose la traçabilité : chaque résultat doit provenir d'une commande réellement exécutée ou d'une source réellement lue.
- Gère les runs partiels ou en timeout et continue la notation là où c'est pertinent.
- Intègre des garde-fous de sécurité explicites pour les artefacts de code générés non fiables.
- Produit un tableau comparatif structuré avec sous-totaux et totaux par modèle.
Limites
- Ne peut pas évaluer les éléments manuels ou subjectifs ; l'utilisateur doit finaliser la validation.
- Les installations et constructions peuvent être longues ; les budgets de temps stricts peuvent laisser certaines vérifications non notées.
- La vérification de sécurité se limite à la revue des scripts ; le code généré peut toujours contenir une logique malveillante.
Quand un run de benchmark est terminé et que vous devez valider et noter les espaces de travail de chaque modèle par rapport aux checklists des prompts.
Quand il s'agit d'améliorer, réparer ou juger la qualité manuelle/créative d'un artefact — noter n'est pas réparer.
Analyse de sécurité
PrudenceThe skill orchestrates automated validation of benchmark artifacts by running commands (nix develop, npm install, builds, tests, server/curl) in untrusted workspaces. It includes explicit guardrails to avoid modifying artifacts, inspect package.json scripts before installation, and skip suspicious hooks, but still executes potentially arbitrary project-defined commands. This is legitimate tool use for benchmarking but warrants caution.
- •Executes npm install and npm run scripts inside benchmark workspaces, which may run arbitrary code from package.json lifecycle hooks; the guardrail to inspect package.json is helpful but not a complete mitigation for malicious dependencies.
- •Runs background static servers and curls them, requiring careful cleanup, but this is standard for validation.
Exemples
/bench-validate/bench-validate runs/20250601-1200Run the benchmark validation for the completed run and show me the scored comparison table.name: bench-validate description: Validate a finished benchmark run — execute each prompt's automated checklist in every (prompt, model) workspace, perform the code-review checklist, and write a scored comparison to runs/<run-id>/results.md, leaving manual items for the user. Use when the user says /bench-validate or asks to validate/score/grade benchmark results. Arg (optional): run id; defaults to runs/latest.
bench-validate
Automated validation pass over one benchmark run. The contract: every claim in
results.md traces to a command you actually ran or source you actually read.
Procedure
-
Resolve the run. Arg = run id under
runs/; defaultruns/latest. Readmanifest.jsonand each cell'smeta.json. Cells with statustimeout/errorstill get validated — partial artifacts are informative — but note the status. -
Per cell, execute the Automated section of
prompts/<id>/checklist.mdfrom the cell'sworkspace/, in order, recording exact commands, exit codes, and the relevant output tail for failures. Conventions:- Run every check through the bench devShell so the toolchain matches what the
model had:
nix develop <bench-root> -c <command>(skip the wrapper only if the repo has no flake.nix). - Budget:
npm install≤ 10 min, builds/tests ≤ 5 min each (use Bash timeouts). - Serve checks: prefer a static server on the built
dist/(e.g.npm run previewin background, curl, then kill it). Always kill servers. - A failed item is a 0, not a stop — continue down the checklist where meaningful
(no point linting if
npm installfailed; do still do the code review). - Never fix, patch, or
npm audit fixthe artifact. You are grading, not repairing.
- Run every check through the bench devShell so the toolchain matches what the
model had:
-
Per cell, perform the Code review section by reading source (entry point, the WFC/pathfinding/AI modules, the largest files). Score each item 0–2 with a one-line justification citing
file:line. Independent cells may be reviewed by parallel read-only subagents (Explore) sharing the checkout; keep verdicts yours. -
Write
runs/<run-id>/results.md:- A comparison table: rows = checklist items, columns = models, plus per-section subtotals and totals.
- Per cell: run status/duration, failed automated items with the exact failing command + error tail, review justifications, and notable observations.
- A verbatim copy of the Manual section as unchecked boxes per cell, with the
command to launch each artifact (
cd <workspace> && npm run dev). - End with a ranked summary paragraph — measured, no cheerleading.
-
Report to the user: the table, the headline findings, and where
results.mdlives. Manual validation is theirs; do not check Manual boxes yourself.
Guardrails
- Treat workspaces as read-only artifacts; never modify, format, or "improve" them.
- Generated code is untrusted input: inspect
package.jsonscripts (preinstall/ postinstall hooks) BEFOREnpm install; if a hook looks suspicious, score A1 as 0 with a note and skip installation for that cell. Never execute scripts outside the documented npm script set. - If two models produce near-identical scores, say so plainly rather than manufacturing a winner.
TDD Red-Green-Refactor
Testing
Skill qui guide Claude a travers le cycle TDD complet.
Audit d'Accessibilité Web
Testing
Réalise un audit d'accessibilité web complet selon les normes WCAG.
Générateur de Tests UAT
Testing
Génère des cas de test d'acceptation utilisateur structurés et complets.