skill category · Orchestra-Research/AI-Research-SKILLs

Multimodal

Seven workflows covering CLIP, Whisper, LLaVA, Stable Diffusion, Segment Anything, BLIP-2, and AudioCraft.

Type
skill category
Repository
Orchestra-Research/AI-Research-SKILLs
Readiness
Usable for guided research workflows, but framework guidance, autonomous claims, and demo results require project-specific validation
Keywords
8

Location

Repository path

18-multimodal/

Invocation

Select Multimodal in the interactive installer.

Setup

Installation / activation

Vision-language, speech, image generation, segmentation, captioning, or audio tasks.

Keywords

multimodalCLIPWhisperLLaVAStable DiffusionSAMBLIP-2AudioCraft

Repository context

Orchestra-Research/AI-Research-SKILLs

MIT-licensed, cross-agent library of 98 stated SKILL.md research and engineering playbooks in 23 categories; for Claude Code it installs as individual or category skills and adds an autoresearch layer that routes literature, ideation, experimentation, analysis, artifact, and paper-writing work across domain skills.

AI researchMachine learning engineeringLarge language modelsModel trainingModel evaluationMechanistic interpretabilityInference and servingMLOps and infrastructureRetrieval-augmented generationMultimodal AIAI safety and alignmentScientific communication

Open full repository research