Find the perfect skill
Unit Test Generator
Testing
Automatically generate comprehensive unit tests with edge cases, mocks, and proper assertions for any codebase.
AI Agent Evaluation Systems
Testing
Design and implement comprehensive evaluation systems for AI agents, covering grader types (code-based, model-based, human), benchmarks (SWE-bench, Terminal-Bench), and production integration. Use this skill when building evals for coding, conversational, research, or computer-use agents to measure and improve agent performance reproducibly.
Web QA testing and bug fixing
Testing
Systematically test a web app, fix bugs found, and verify fixes with atomic commits. Produces before/after health reports and ship-readiness summary.
Evidence-First Proof Kit for Deterministic Changes
Testing
Skill for assembling proof kits with explicit safety justification, using artifacts and reports over narrative claims. Applies strict determinism rules and negative test suites for risky operations.
Gatling Scenario Creator
Testing
Automated assistance for creating Gatling performance test scenarios. Provides step-by-step guidance and best practices.
Verifying cv-builder changes end-to-end
Testing
How to build, run and drive cv-builder to verify a change end-to-end, including dev server, Playwright recipe, data hygiene, and cleanup.
UI Visual Validator
Testing
Rigorous visual validation expert specializing in UI testing, design system compliance, and accessibility verification.
Playwright Automation for JSF Apps
Testing
Expert in Playwright automation with Python for JSF/RichFaces applications. Uses suffix selectors, AJAX waits, and explicit viewport.
Fallout 4 Compatibility Audit
Testing
Per-game audit of Fallout 4 compatibility for ByroRedux, covering BA2, half-float verts, BGSM materials, and M49 precombines (CSG).
Slop-Guard QA Audit
Testing
Black-box QA audit of slop-guard across MCP, CLI, docs, and workflows. Files genuine GitHub issues.
Naive UX Check
Testing
Runs a naive-user UX comprehension check by surfacing the next unrun prompt (kid, guardian, or admin persona) and logging the response to a dated report. Use before releases for a fresh user perspective.
Promptfoo - LLM Evaluation Framework
Testing
CLI tool for testing and comparing LLM outputs. Create evaluation configs, test cases, and custom assertions for validating model behavior.