evaluation workflow · DietrichGebert/ponytail

Agentic Claude Code benchmark

Runs isolated headless Claude Code sessions against seeded tasks and measures correctness, safety, source size, completeness, cost, duration, and turns.

Type
evaluation workflow
Repository
DietrichGebert/ponytail
Readiness
Usable for Claude Code with documented installation and tests, with caveats around Node-based hook activation, model adherence, and unverified benchmark and security claims.
Keywords
6

Location

Repository path

benchmarks/agentic/README.md

Invocation

Use the documented `run.py` self-test and task workflows from `benchmarks/agentic/README.md`.

Setup

Installation / activation

Explicit evaluation run with Claude Code, Python 3, and the pinned template checkout.

Keywords

Claude-Codeagentic-benchmarksafetycompletenessLOCevaluation

Repository context

DietrichGebert/ponytail

MIT-licensed Claude Code plugin and portable Agent Skill that applies a YAGNI-first decision ladder, persistent mode controls, lifecycle-hook injection, review and audit commands, and an optional read-only MCP interface to steer coding agents toward the smallest correct implementation without removing stated safety, validation, accessibility, or data-loss guards.

software engineeringcode generationcode reviewrefactoringdependency selectiontechnical debtprompt engineeringagent productivity

Open full repository research