skill · moonlight-lupin/agent-skills
model-compare
Runs blind multi-model A/B comparisons in simple, tool, coding, review, and embedding-benchmark modes.
Location
Repository path
mlops/model-compare/SKILL.mdInvocation
Load/install the skill and request a blind comparison or embedding benchmark.
Setup
Installation / activation
Evaluating multiple models while reducing presentation and identity bias.
Keywords
model evaluationA/B testingblind comparisonembeddingsbenchmark
Repository context
moonlight-lupin/agent-skills
MIT-licensed collection of 29 self-contained Agent Skills built for Hermes Agent, plus a Hermes-only BM25 skill-retrieval plugin; Claude Code can install the SKILL.md packages, but workflows may require mapping Hermes tool names and the plugin itself is explicitly incompatible with Claude Code.
Research and fact-checkingCreative media generationProductivity and document automationAgent operationsMLOps and model evaluationWeb scrapingDevOpsRetrieval-augmented generationCompliance-support research