skill category · Orchestra-Research/AI-Research-SKILLs

Evaluation

Benchmarking workflows for lm-evaluation-harness, BigCode Evaluation Harness, and NeMo Evaluator.

Type
skill category
Repository
Orchestra-Research/AI-Research-SKILLs
Readiness
Usable for guided research workflows, but framework guidance, autonomous claims, and demo results require project-specific validation
Keywords
5

Location

Repository path

11-evaluation/

Invocation

Select Evaluation in the interactive installer, then request a benchmark workflow.

Setup

Installation / activation

LLM, code-model, safety, or multimodal evaluation tasks.

Keywords

evaluationbenchmarkinglm-evalBigCodeNeMo Evaluator

Repository context

Orchestra-Research/AI-Research-SKILLs

MIT-licensed, cross-agent library of 98 stated SKILL.md research and engineering playbooks in 23 categories; for Claude Code it installs as individual or category skills and adds an autoresearch layer that routes literature, ideation, experimentation, analysis, artifact, and paper-writing work across domain skills.

AI researchMachine learning engineeringLarge language modelsModel trainingModel evaluationMechanistic interpretabilityInference and servingMLOps and infrastructureRetrieval-augmented generationMultimodal AIAI safety and alignmentScientific communication

Open full repository research