AI research Agent Skills library · Active, versioned repository with installer, marketplace packaging, demos, and stated CI maintenance

Orchestra-Research/AI-Research-SKILLs

MIT-licensed, cross-agent library of 98 stated SKILL.md research and engineering playbooks in 23 categories; for Claude Code it installs as individual or category skills and adds an autoresearch layer that routes literature, ideation, experimentation, analysis, artifact, and paper-writing work across domain skills.

AI researchMachine learning engineeringLarge language modelsModel trainingModel evaluationMechanistic interpretabilityInference and servingMLOps and infrastructureRetrieval-augmented generationMultimodal AIAI safety and alignmentScientific communication
Routing score
80.0
Readiness
Usable for guided research workflows, but framework guidance, autonomous claims, and demo results require project-specific validation
License
MIT
Maintenance
active
Components
20
Revision
0

Selection

Select when

  • You want Claude Code guidance spanning both research reasoning and ML engineering.
  • You need several framework-specific skills from one installable collection.
  • You want category-level or individual skill selection.
  • You want an autoresearch workflow that routes across specialized skills.
  • You need guidance for fine-tuning, RL post-training, distributed training, evaluation, or serving.
  • You need mechanistic-interpretability workflows alongside training guidance.
  • You want paper-writing, LaTeX-template, and academic-plotting support.
  • You want project-local skills that can be version controlled.
  • You use multiple SKILL.md-compatible coding agents.
  • You want research outputs organized into auditable agent-native artifacts.

Boundaries

Avoid when

  • You need independently verified scientific conclusions rather than agent guidance.
  • You need a turnkey compute environment with dependencies and credentials already configured.
  • You want a small domain-neutral Claude Code extension.
  • You cannot review generated code, experimental design, citations, or claims.
  • You require offline-only operation for skills that reference remote APIs or services.
  • You need guaranteed compatibility with a specific framework or host version.
  • You require an independently audited installer or skill corpus.
  • You do not want an autonomous agent writing files or running experiments.
  • You need deterministic reproduction of the showcased autonomous demos.
  • You require consistent inventory details with no documentation drift.

Strengths

Capabilities

Provides 98 stated skills across 23 AI-research categories.Installs all skills, bundles, categories, or individual skills through an interactive npm installer.Installs skills globally through canonical storage and agent-directory symlinks, with a Windows copy fallback.Supports per-project copied installations with a tracking file for version-controlled skill sets.Offers category plugins through a Claude Code marketplace alternative.Places Claude Code skills under global or project-local .claude/skills directories.Uses SKILL.md entry points with metadata, triggers, patterns, examples, and progressively disclosed references.Provides an autoresearch skill that routes work through a two-loop research workflow and other domain skills.Covers architecture, tokenization, fine-tuning, post-training, distributed training, optimization, evaluation, serving, and infrastructure.Covers mechanistic interpretability through TransformerLens, SAELens, pyvene, and nnsight guidance.Covers agents, RAG, prompt engineering, multimodal systems, observability, MLOps, and safety.Provides research brainstorming and creative-thinking workflows.Provides ML paper-writing guidance, venue templates, citation verification, and academic plotting workflows.Includes optional scripts, templates, examples, and assets inside skill folders.Adds three Agent-Native Research Artifact skills for compilation, session recording, and rigor review.Documents demos for autonomous research, model evaluation, quantization, multilingual embeddings, and scientific plotting.Supports installation for Claude Code and eight other stated agent locations.States that its installer can list, update, and selectively uninstall installed skills.Reports a CI inventory drift guard and marketplace packaging checks in repository release notes.Documents a reusable skill directory format and contribution path for adding or improving skills.

Risk profile

Risks and limitations

  • Uncertainty: only selected first-party files were supplied, so the 98 SKILL.md files, installer source, marketplace manifest, CI workflows, optional scripts, and most references were not directly inspected.
  • Uncertainty: metadata and release notes are dated in 2026 relative to this analysis context, so activity, version, adoption, and recency claims cannot be independently reconciled.
  • The root README contains stale 87-skill statistics despite stating that inventory was reconciled to 98 skills, indicating residual documentation drift.
  • The package category examples disagree with the root inventory on model-architecture and fine-tuning counts and named inclusions.
  • Repository claims of production-ready, battle-tested, autonomous, comprehensive, or research-grade quality were not independently validated.
  • Demo conclusions and performance numbers are repository-reported, sometimes use synthetic data, and do not establish general scientific reliability.
  • Autonomous literature review, experimentation, causal interpretation, citation handling, and paper writing still require expert human oversight.
  • Many skills depend on third-party frameworks, GPUs, cloud services, API keys, datasets, system packages, or paid model access.
  • Optional helper scripts can execute code or transmit inputs to external services, and their implementations were not supplied for inspection.
  • No minimum Claude Code version, host compatibility test matrix, telemetry disclosure, threat model, or independent security audit was supplied.
  • The installer downloads skills and creates symlinks or copies in multiple agent directories, which changes local configuration and expands the review surface.
  • Individual referenced libraries and bundled conference materials may have licenses, versions, deadlines, or requirements distinct from the repository's MIT license.

Evidence: release notes state that installer dependencies were pinned to patched versions and that npm audit reported zero vulnerabilities; this was not verified from manifests or CI output. The installer downloads content and writes symlinks or copies into agent directories, while optional skill scripts may execute code or use external APIs. Users should inspect installed skills, scripts, templates, dependency versions, and data flows before use; no independent audit or threat model was supplied.

Component inventory

20 documented components

Claude Code skill suite

AI Research Skills Library

MIT-licensed library claiming 98 SKILL.md research and engineering skills across 23 categories; extends Claude Code through contextual skills and orchestration.

README.md
installer CLI

AI Research Skills npm CLI

Interactive cross-agent installer supporting all, quickstart, category, individual, global, and project-local skill sets.

packages/ai-research-skills/
plugin marketplace

Claude Code Marketplace

Alternative Claude Code distribution exposing the skill library as category plugins.

.claude-plugin/marketplace.json
orchestration skill

Autoresearch

Two-loop autonomous workflow for literature survey, ideation, experiments, synthesis, and paper writing with domain-skill routing.

0-autoresearch-skill/
skill category

Research Ideation

Research Brainstorming and Creative Thinking skills for structured, high-impact idea generation.

21-research-ideation/
skill category

ML Paper Writing

ML Paper Writing and Academic Plotting workflows with conference templates, citation verification, diagrams, and publication charts.

20-ml-paper-writing/
skill category

Agent-Native Research Artifact

ARA Compiler, Research Manager, and Rigor Reviewer create, record, and assess auditable research artifacts.

22-agent-native-research-artifact/
skill category

Model Architecture

Guidance for LitGPT, Mamba, NanoGPT, RWKV, and TorchTitan model architecture workflows.

01-model-architecture/
skill categories

Tokenization and Data Processing

Workflows for HuggingFace Tokenizers, SentencePiece, Ray Data, and NeMo Curator.

02-tokenization/; 05-data-processing/
skill category / Claude plugin

Fine-Tuning

Fine-tuning guidance for Axolotl, LLaMA-Factory, PEFT, and Unsloth.

03-fine-tuning/
skill category / Claude plugin

Post-Training

Eight post-training skills covering TRL, GRPO, OpenRLHF, SimPO, verl, slime, miles, and torchforge.

06-post-training/
skill categories

Distributed Training and Infrastructure

Training-scale and compute workflows spanning DeepSpeed, FSDP2, Accelerate, Megatron-Core, Lightning, Ray Train, Modal, Lambda Labs, and SkyPilot.

08-distributed-training/; 09-infrastructure/
skill categories / Claude plugins

Optimization and Inference Serving

Optimization and serving guidance for Flash Attention, quantization formats and methods, vLLM, TensorRT-LLM, llama.cpp, and SGLang.

10-optimization/; 12-inference-serving/
skill category

Evaluation

Benchmarking workflows for lm-evaluation-harness, BigCode Evaluation Harness, and NeMo Evaluator.

11-evaluation/
skill category

Safety and Alignment

Guidance for Constitutional AI, LlamaGuard, NeMo Guardrails, and Prompt Guard.

07-safety-alignment/
skill categories

Agents and RAG

Agent-framework and retrieval guidance for LangChain, LlamaIndex, CrewAI, AutoGPT, Chroma, FAISS, Pinecone, Qdrant, and Sentence Transformers.

14-agents/; 15-rag/
skill category

Prompt Engineering

Structured and constrained generation workflows using DSPy, Instructor, Guidance, and Outlines.

16-prompt-engineering/
skill categories

MLOps and Observability

Experiment tracking and LLM observability guidance for W&B, MLflow, TensorBoard, LangSmith, and Phoenix.

13-mlops/; 17-observability/
skill category

Multimodal

Seven workflows covering CLIP, Whisper, LLaVA, Stable Diffusion, Segment Anything, BLIP-2, and AudioCraft.

18-multimodal/
skill categories

Mechanistic Interpretability and Emerging Techniques

Interpretability workflows for TransformerLens, SAELens, pyvene, and nnsight, plus MoE, merging, long context, speculative decoding, distillation, and pruning.

04-mechanistic-interpretability/; 19-emerging-techniques/

Technical profile

Requirements and configuration

Repository
Orchestra-Research/AI-Research-SKILLs
License
MIT for the repository; referenced projects may use other licenses.
Language
GitHub metadata identifies TeX as the primary language.
Skill Count
98 stated skills.
Category Count
23 stated categories.
Skill Format
SKILL.md with metadata, triggers, examples, references, and optional scripts or assets.
Claude Paths
Global ~/.claude/skills and project-local .claude/skills are documented.
Canonical Storage
Global skills are stated to live under ~/.orchestra/skills with agent-directory symlinks.
Local Tracking
Project-local installations are tracked by .orchestra-skills.json.
Package
@orchestra-research/ai-research-skills on npm is documented.
Plugin Packaging
Claude Code marketplace category plugins are documented.
Orchestration
Autoresearch uses stated inner optimization and outer synthesis loops.
Updates
The installer documents list, update, and uninstall operations.
Repository Status
Public, unarchived repository with MIT metadata.
Evidence Scope
README files and selected references were supplied; core skill and installer implementations were mostly absent.

Classification

How it enters the stack

Context InjectionOrchestrationSkill Authoring

Evidence: the installer or marketplace places SKILL.md packages in Claude Code skill locations, and autoresearch routes work across domain skills. Inference: skill_authoring applies because the repository documents a reusable skill structure and contributions. Basic install is short, but research execution often needs substantial dependencies, compute, and credentials.

Claude Extension · medium setup effort · high confidence · automated

Evidence and risk

Primary sources

first_party_fileREADME.mdhttps://github.com/Orchestra-Research/AI-Research-SKILLs/blob/main/README.md
first_party_filedemos/README.mdhttps://github.com/Orchestra-Research/AI-Research-SKILLs/blob/main/demos/README.md

Routing context

Conflicts, complements, and synergies

complements

gh_d4vinci_scrapling

Research playbooks can use web-derived evidence, while Scrapling can provide authorized extraction and RAG-ready Markdown; no direct integration is evidenced.

low confidence
alternative_to

gh_dzhng_skills

Inference: both install composable skills and orchestrate multi-stage work, but this repository targets general software delivery and visual verification while Orchestra-Research focuses on AI research and engineering.

medium confidence
similar_to

gh_google_mantis

Catalog inference: both route work across reusable domain skills, but Mantis focuses on vulnerability review and includes a dedicated sandboxed reference harness.

low confidence
overlaps_with

gh_moonlight_lupin_agent_skills

Both provide MIT-licensed, installable research-oriented Agent Skills for Claude Code, but this repository is Hermes-first and broader in productivity and creative operations.

medium confidence
similar_to

gh_mengto_skills

Inference: both are cross-agent SKILL.md libraries with optional supporting files, but this repository specializes in AI research while MengTo/Skills emphasizes design, frontend, games, and media.

medium confidence
alternative_to

gh_voltagent_awesome_claude_code_subagents

Inference: both provide broad specialist routing for Claude Code, but this repository installs skills and an autoresearch router rather than a catalog of Markdown subagents.

medium confidence
similar_to

gh_owl_listener_designer_skills

Inference: both offer category-installable Claude Code skill suites with workflow orchestration, but their research and design domains differ substantially.

low confidence