skill · moonlight-lupin/agent-skills

model-compare

Runs blind multi-model A/B comparisons in simple, tool, coding, review, and embedding-benchmark modes.

Type
skill
Repository
moonlight-lupin/agent-skills
Readiness
Usable for Claude Code with adaptation; individual dependencies, credentials, and host-tool mappings vary by skill.
Keywords
5

Location

Repository path

mlops/model-compare/SKILL.md

Invocation

Load/install the skill and request a blind comparison or embedding benchmark.

Setup

Installation / activation

Evaluating multiple models while reducing presentation and identity bias.

Keywords

model evaluationA/B testingblind comparisonembeddingsbenchmark

Repository context

moonlight-lupin/agent-skills

MIT-licensed collection of 29 self-contained Agent Skills built for Hermes Agent, plus a Hermes-only BM25 skill-retrieval plugin; Claude Code can install the SKILL.md packages, but workflows may require mapping Hermes tool names and the plugin itself is explicitly incompatible with Claude Code.

Research and fact-checkingCreative media generationProductivity and document automationAgent operationsMLOps and model evaluationWeb scrapingDevOpsRetrieval-augmented generationCompliance-support research

Open full repository research