AI Research Skills Library
MIT-licensed library claiming 98 SKILL.md research and engineering skills across 23 categories; extends Claude Code through contextual skills and orchestration.
README.mdAI research Agent Skills library · Active, versioned repository with installer, marketplace packaging, demos, and stated CI maintenance
MIT-licensed, cross-agent library of 98 stated SKILL.md research and engineering playbooks in 23 categories; for Claude Code it installs as individual or category skills and adds an autoresearch layer that routes literature, ideation, experimentation, analysis, artifact, and paper-writing work across domain skills.
Selection
Boundaries
Strengths
Risk profile
Evidence: release notes state that installer dependencies were pinned to patched versions and that npm audit reported zero vulnerabilities; this was not verified from manifests or CI output. The installer downloads content and writes symlinks or copies into agent directories, while optional skill scripts may execute code or use external APIs. Users should inspect installed skills, scripts, templates, dependency versions, and data flows before use; no independent audit or threat model was supplied.
Component inventory
MIT-licensed library claiming 98 SKILL.md research and engineering skills across 23 categories; extends Claude Code through contextual skills and orchestration.
README.mdInteractive cross-agent installer supporting all, quickstart, category, individual, global, and project-local skill sets.
packages/ai-research-skills/Alternative Claude Code distribution exposing the skill library as category plugins.
.claude-plugin/marketplace.jsonTwo-loop autonomous workflow for literature survey, ideation, experiments, synthesis, and paper writing with domain-skill routing.
0-autoresearch-skill/Research Brainstorming and Creative Thinking skills for structured, high-impact idea generation.
21-research-ideation/ML Paper Writing and Academic Plotting workflows with conference templates, citation verification, diagrams, and publication charts.
20-ml-paper-writing/ARA Compiler, Research Manager, and Rigor Reviewer create, record, and assess auditable research artifacts.
22-agent-native-research-artifact/Guidance for LitGPT, Mamba, NanoGPT, RWKV, and TorchTitan model architecture workflows.
01-model-architecture/Workflows for HuggingFace Tokenizers, SentencePiece, Ray Data, and NeMo Curator.
02-tokenization/; 05-data-processing/Fine-tuning guidance for Axolotl, LLaMA-Factory, PEFT, and Unsloth.
03-fine-tuning/Eight post-training skills covering TRL, GRPO, OpenRLHF, SimPO, verl, slime, miles, and torchforge.
06-post-training/Training-scale and compute workflows spanning DeepSpeed, FSDP2, Accelerate, Megatron-Core, Lightning, Ray Train, Modal, Lambda Labs, and SkyPilot.
08-distributed-training/; 09-infrastructure/Optimization and serving guidance for Flash Attention, quantization formats and methods, vLLM, TensorRT-LLM, llama.cpp, and SGLang.
10-optimization/; 12-inference-serving/Benchmarking workflows for lm-evaluation-harness, BigCode Evaluation Harness, and NeMo Evaluator.
11-evaluation/Guidance for Constitutional AI, LlamaGuard, NeMo Guardrails, and Prompt Guard.
07-safety-alignment/Agent-framework and retrieval guidance for LangChain, LlamaIndex, CrewAI, AutoGPT, Chroma, FAISS, Pinecone, Qdrant, and Sentence Transformers.
14-agents/; 15-rag/Structured and constrained generation workflows using DSPy, Instructor, Guidance, and Outlines.
16-prompt-engineering/Experiment tracking and LLM observability guidance for W&B, MLflow, TensorBoard, LangSmith, and Phoenix.
13-mlops/; 17-observability/Seven workflows covering CLIP, Whisper, LLaVA, Stable Diffusion, Segment Anything, BLIP-2, and AudioCraft.
18-multimodal/Interpretability workflows for TransformerLens, SAELens, pyvene, and nnsight, plus MoE, merging, long context, speculative decoding, distillation, and pruning.
04-mechanistic-interpretability/; 19-emerging-techniques/Technical profile
Classification
Evidence: the installer or marketplace places SKILL.md packages in Claude Code skill locations, and autoresearch routes work across domain skills. Inference: skill_authoring applies because the repository documents a reusable skill structure and contributions. Basic install is short, but research execution often needs substantial dependencies, compute, and credentials.
Claude Extension · medium setup effort · high confidence · automatedEvidence and risk
Routing context
Research playbooks can use web-derived evidence, while Scrapling can provide authorized extraction and RAG-ready Markdown; no direct integration is evidenced.
low confidenceInference: both install composable skills and orchestrate multi-stage work, but this repository targets general software delivery and visual verification while Orchestra-Research focuses on AI research and engineering.
medium confidenceCatalog inference: both route work across reusable domain skills, but Mantis focuses on vulnerability review and includes a dedicated sandboxed reference harness.
low confidenceBoth provide MIT-licensed, installable research-oriented Agent Skills for Claude Code, but this repository is Hermes-first and broader in productivity and creative operations.
medium confidenceInference: both are cross-agent SKILL.md libraries with optional supporting files, but this repository specializes in AI research while MengTo/Skills emphasizes design, frontend, games, and media.
medium confidenceInference: both provide broad specialist routing for Claude Code, but this repository installs skills and an autoresearch router rather than a catalog of Markdown subagents.
medium confidenceInference: both offer category-installable Claude Code skill suites with workflow orchestration, but their research and design domains differ substantially.
low confidence