Narrating over a Video using Multimodal AI

This workflow extracts frames from a video, uses a multimodal LLM like GPT-4o to generate a narration script, creates voiceover audio via TTS, and uploads it to Google Drive.

This n8n workflow automates the process of adding AI-generated narration to videos. It starts by downloading the video via HTTP, then uses a Python code node with OpenCV to extract key frames. These frames are batched using a Loop node and fed into a multimodal LLM (e.g., GPT-4o) to generate partial narration scripts based on visual content. The partial scripts are combined into a full script, which is converted to high-quality voiceover audio using OpenAI's TTS API. Finally, the audio file is u
Platform
n8n
Category
Development & IT
Price
$19.99
Creator
Jana Hedlund

How to import this workflow into n8n

  1. 1Purchase or download the workflow to get the n8n workflow JSON file.
  2. 2In your n8n instance, open Workflows and choose "Import from File" (or paste the JSON with Ctrl+V on the canvas).
  3. 3Open each node marked with a credential warning and connect your own accounts and API keys.
  4. 4Run the workflow once manually to verify the data flow, then toggle it to Active.

Related Development & IT workflows

More from Jana Hedlund

Need this deployed? We'll set it up for you.

Our automation experts deploy this workflow in your stack, connect your accounts, and verify it works — or build a custom solution from scratch.