benchmark workflow · DietrichGebert/ponytail
Single-shot benchmark
Compares no-skill, Caveman, and Ponytail outputs across everyday coding tasks, recording LOC, correctness, cost, and latency.
Location
Repository path
benchmarks/README.mdInvocation
Use the documented Promptfoo workflow in `benchmarks/README.md`; local-model runs use `benchmarks/benchmark-local.py`.
Setup
Installation / activation
Explicit benchmark run with the documented prerequisites.
Keywords
benchmarkPromptfooClaudeOllamaLOCcorrectnesscostlatency
Repository context
DietrichGebert/ponytail
MIT-licensed Claude Code plugin and portable Agent Skill that applies a YAGNI-first decision ladder, persistent mode controls, lifecycle-hook injection, review and audit commands, and an optional read-only MCP interface to steer coding agents toward the smallest correct implementation without removing stated safety, validation, accessibility, or data-loss guards.
software engineeringcode generationcode reviewrefactoringdependency selectiontechnical debtprompt engineeringagent productivity