benchmark workflow · DietrichGebert/ponytail

Single-shot benchmark

Compares no-skill, Caveman, and Ponytail outputs across everyday coding tasks, recording LOC, correctness, cost, and latency.

Type
benchmark workflow
Repository
DietrichGebert/ponytail
Readiness
Usable for Claude Code with documented installation and tests, with caveats around Node-based hook activation, model adherence, and unverified benchmark and security claims.
Keywords
8

Location

Repository path

benchmarks/README.md

Invocation

Use the documented Promptfoo workflow in `benchmarks/README.md`; local-model runs use `benchmarks/benchmark-local.py`.

Setup

Installation / activation

Explicit benchmark run with the documented prerequisites.

Keywords

benchmarkPromptfooClaudeOllamaLOCcorrectnesscostlatency

Repository context

DietrichGebert/ponytail

MIT-licensed Claude Code plugin and portable Agent Skill that applies a YAGNI-first decision ladder, persistent mode controls, lifecycle-hook injection, review and audit commands, and an optional read-only MCP interface to steer coding agents toward the smallest correct implementation without removing stated safety, validation, accessibility, or data-loss guards.

software engineeringcode generationcode reviewrefactoringdependency selectiontechnical debtprompt engineeringagent productivity

Open full repository research