OptionalCreativeVersion 1.0.0

Kanban Video Orchestrator: Multi-Agent Video Production with Hermes

Plan and run multi-agent video production pipelines.

Written by Neura Market from the official Hermes Agent documentation for Kanban Video Orchestrator. Commands, paths, and version numbers are reproduced from the source unchanged.

Read the official documentation

The Kanban Video Orchestrator is a meta-pipeline for Hermes Agent that turns a video brief into a multi-agent production. You describe what you want, and the orchestrator designs a team of specialized agent profiles, sets up a kanban board, and hands off execution to a director profile that decomposes the work and routes tasks to the right renderers. This is for anyone who needs to produce a video that requires multiple distinct skills, a narrative short with animation and music, a product teaser that mixes live footage with motion graphics, or a music video that needs beat analysis and ASCII art, and wants the coordination handled by agents rather than a single monolithic script.

What it does

The orchestrator does not render a single frame. It is a workflow engine that scopes the request, designs a team, generates a setup script, and monitors execution. The actual rendering happens inside the kanban once it is running, via whichever existing skills and tools fit the scenes: ascii-video, manim-video, p5js, comfyui, touchdesigner-mcp, blender-mcp, songwriting-and-ai-music, heartmula, external APIs, or plain Python with PIL and ffmpeg.

Before you start

This is an optional skill. Install it on demand into your Hermes Agent instance. It runs on Linux, macOS, and Windows. You need Hermes Agent installed and configured, with the kanban toolset available. The skill itself is a set of reference files and scripts that guide the agent through the pipeline; it does not add new binaries.

You will need API keys for any external services your video requires, TTS, image generation, image-to-video, and so on. These go in ${HERMES_HOME:-~/.hermes}/.env or the user's secret store. The setup script's check_key helper aborts cleanly if a required key is missing.

Workflow

The pipeline follows six stages:

DISCOVER  →  BRIEF  →  TEAM DESIGN  →  SETUP  →  EXECUTE  →  MONITOR

Step 1, Discover (ask the right questions)

The discovery process is adaptive: ask only what is actually needed. Always start with three questions to identify the broad shape:

  • What is the video? (one-sentence brief)
  • How long? (5-30s teaser / 30-90s short / 90s-3min explainer / 3-10min film / longer)
  • What aspect ratio + target platform? (1:1 / 9:16 / 16:9; X, IG, YouTube, internal, etc.)

From the answer, classify the style category. The style determines which follow-up questions to ask. Do not ask all questions at once. Ask 2-4 at a time, listen, then proceed. Make reasonable assumptions whenever the user implies an answer.

For complete intake patterns and per-style question banks, see references/intake.md.

Step 2, Brief

Once enough is known, produce a structured brief.md using the template in assets/brief.md.tmpl. Stages:

  1. Concept, the one-sentence pitch + emotional north star
  2. Scope, duration, aspect, platform, deadline
  3. Style, visual references, brand constraints, tone
  4. Scenes, beat-by-beat breakdown (durations, content, target tool)
  5. Audio, narration / music / SFX / silent (per scene if needed)
  6. Deliverables, file format, resolution, optional alternates (vertical cut, GIF, etc.)

Show the brief to the user for confirmation before designing the team. The brief is the contract, every downstream task references it.

Step 3, Team design

Pick role archetypes from the library that fit this video. Compose, don't clone. Most videos need 4-7 profiles. The director is always present; the rest are picked by what the brief actually requires.

For the role library and per-style team compositions, see references/role-archetypes.md.

For mapping role → which Hermes skills + toolsets it loads, see references/tool-matrix.md.

Step 4, Setup

Generate a setup script (setup.sh) and run it. The script:

  1. Creates the project workspace (~/projects/video-pipeline//)
  2. Copies any provided assets into taste/, audio/, assets/
  3. Creates each Hermes profile via hermes profile create --clone
  4. Writes per-profile SOUL.md (personality + role definition)
  5. Configures profile YAML (toolsets, always_load skills, cwd)
  6. Writes brief.md, TEAM.md, and taste/ content
  7. Fires the initial hermes kanban create task assigned to the director

Use scripts/bootstrap_pipeline.py to generate setup.sh from a brief + team-design JSON. See references/kanban-setup.md for the setup script structure, profile config patterns, and the critical "shared workspace" rule.

Step 5, Execute

Run setup.sh. Then provide the user with monitoring commands:

hermes kanban watch --tenant <project-tenant>     # live events
hermes kanban list  --tenant <project-tenant>     # board snapshot
hermes dashboard                                   # visual board UI

The director profile takes over from here, decomposing the work and routing tasks to specialist profiles via the kanban toolset.

Step 6, Monitor and intervene

Stay engaged, the kanban runs autonomously but a stuck task or bad output needs human (or AI) judgment.

Monitoring patterns: poll kanban list periodically, inspect any RUNNING task that exceeds its expected duration with kanban show , and check heartbeats. When a worker's output fails review, the standard interventions are:

  1. Comment on the worker's task with specific feedback (kanban_comment)
  2. Create a re-run task with the original as parent
  3. Adjust the brief's scope and let the director re-decompose

For diagnostic patterns, intervention recipes, and the "task is stuck" playbook, see references/monitoring.md.

Reference: worked examples

Six concrete pipelines covering very different video styles, narrative film, product/marketing, music video, math/algorithm explainer, ASCII video, real-time installation, showing how the same workflow yields very different teams and task graphs. See references/examples.md.

Critical rules

  1. Discovery before action. Never start generating a brief or team without asking at least the three baseline questions. A bad brief cascades through the entire pipeline.
  2. Match the team to the video. Don't reuse the same 4-profile setup for every job. A music video that doesn't have a beat-analysis profile will misfire. A narrative film that doesn't have a writer profile will produce incoherent scenes. See references/role-archetypes.md.
  3. One workspace per project. All profiles for a given video share the same dir: workspace. Tasks pass artifacts via shared filesystem and structured handoffs. Every kanban_create call passes workspace_kind="dir" + workspace_path="".
  4. Tenant every project. Use a project-specific tenant (--tenant ). Keeps the dashboard scoped and prevents cross-pollination with other ongoing kanbans.
  5. Respect existing skills. When a scene fits an existing skill, the relevant renderer should load that skill via --skill on its task or always_load in its profile. Do not re-derive what a skill already provides.
  6. The director never executes. Even with the full kanban + terminal + file toolset, the director's SOUL.md rules forbid it from executing work itself. It decomposes and routes only, every concrete task becomes a hermes kanban create call to a specialist profile. The kanban orchestration guidance auto-injected into every kanban worker's system prompt spells this out further.
  7. Don't over-decompose. A 30-second product video does NOT need 20 tasks. Aim for the smallest task graph that still parallelizes well and exposes the right human-review gates.
  8. Verify API keys BEFORE firing. External APIs (TTS, image-gen, image-to-video) need keys in ${HERMES_HOME:-~/.hermes}/.env or the user's secret store. A worker that hits a missing-key error wastes a task slot. The setup script's check_key helper aborts cleanly if a required key is missing.

File map

SKILL.md                            ← this file (workflow + rules)
references/
  intake.md                         ← discovery question banks per style
  role-archetypes.md                ← role library (writer, designer, animator, …)
  tool-matrix.md                    ← skill + toolset mapping per role
  kanban-setup.md                   ← setup script structure & profile config
  monitoring.md                     ← watch + intervene patterns
  examples.md                       ← six worked pipelines
assets/
  brief.md.tmpl                     ← brief skeleton
  setup.sh.tmpl                     ← setup script skeleton
  soul.md.tmpl                      ← profile personality skeleton
scripts/
  bootstrap_pipeline.py             ← generate setup.sh from brief + team JSON
  monitor.py                        ← polling + intervention helpers

When not to use it

  • The video is one continuous procedural project that needs no specialists. Just write the code directly.
  • The user wants a quick one-shot conversion (e.g. "convert this mp4 to a GIF"), use ffmpeg directly.
  • The output is a static image, GIF, or audio-only artifact, use the matching specific skill (ascii-art, gifs, meme-generation, songwriting-and-ai-music).
  • The work fits a single existing skill cleanly (e.g. a pure ASCII video, just use ascii-video).

Limits and gotchas

The orchestrator is a meta-pipeline, not a renderer. It depends on the availability and correctness of the underlying skills and tools. If a required skill is not installed or an API key is missing, the pipeline will stall. The director profile is autonomous but not infallible, a stuck task or bad output requires human or AI judgment to correct. The kanban board is scoped to a project tenant; cross-pollination with other kanbans is prevented by using unique tenants.

What pairs with this

This skill is designed to work alongside the rendering skills it orchestrates: ascii-video, manim-video, p5js, comfyui, touchdesigner-mcp, blender-mcp, songwriting-and-ai-music, heartmula, and others in the Creative category. The related skills list also includes ascii-art, pixel-art, gif-search, meme-generation, youtube-content, claude-design, excalidraw, architecture-diagram, concept-diagrams, baoyu-comic, baoyu-infographic, humanizer, and songsee.

Skills the docs pair this with

More Creative skills