skill category · Orchestra-Research/AI-Research-SKILLs
Evaluation
Benchmarking workflows for lm-evaluation-harness, BigCode Evaluation Harness, and NeMo Evaluator.
Location
Repository path
11-evaluation/Invocation
Select Evaluation in the interactive installer, then request a benchmark workflow.
Setup
Installation / activation
LLM, code-model, safety, or multimodal evaluation tasks.
Keywords
evaluationbenchmarkinglm-evalBigCodeNeMo Evaluator
Repository context
Orchestra-Research/AI-Research-SKILLs
MIT-licensed, cross-agent library of 98 stated SKILL.md research and engineering playbooks in 23 categories; for Claude Code it installs as individual or category skills and adds an autoresearch layer that routes literature, ideation, experimentation, analysis, artifact, and paper-writing work across domain skills.
AI researchMachine learning engineeringLarge language modelsModel trainingModel evaluationMechanistic interpretabilityInference and servingMLOps and infrastructureRetrieval-augmented generationMultimodal AIAI safety and alignmentScientific communication