evaluation workflow · DietrichGebert/ponytail
Agentic Claude Code benchmark
Runs isolated headless Claude Code sessions against seeded tasks and measures correctness, safety, source size, completeness, cost, duration, and turns.
Location
Repository path
benchmarks/agentic/README.mdInvocation
Use the documented `run.py` self-test and task workflows from `benchmarks/agentic/README.md`.
Setup
Installation / activation
Explicit evaluation run with Claude Code, Python 3, and the pinned template checkout.
Keywords
Repository context
DietrichGebert/ponytail
MIT-licensed Claude Code plugin and portable Agent Skill that applies a YAGNI-first decision ladder, persistent mode controls, lifecycle-hook injection, review and audit commands, and an optional read-only MCP interface to steer coding agents toward the smallest correct implementation without removing stated safety, validation, accessibility, or data-loss guards.