# Narrating over a video using multimodal AI

## Overview
This n8n template takes a video and extracts frames from it which are used with a multimodal LLM to generate a script. The script is then passed to the same multimodal LLM to generate a voiceover clip.

This template was inspired by [Processing and narrating a video with GPT’s visual capabilities and the TTS API](https://cookbook.openai.com/examples/gpt_with_vision_for_video_understanding).

## Workflow Steps
### 1. Download Video
In this demonstration, we’ll download a stock video from pixabay using the HTTP Request node. Feel free to use other sources but ensure they are in a format supported by OpenCV ([See docs](https://docs.opencv.org/3.4/dd/d43/tutorial_py_video_display.html)).

### 2. Split Video into Frames
We need to think of videos as a sum of 2 parts; a visual track and an audio track. The visual track is technically just a collection of images displayed one after the other and are typically referred to as frames. Here, we use the Python Code node to extract the frames from the video using OpenCV, a computer vision library. For performance reasons, we’ll also capture only a max of 90 frames from the video but ensure they are evenly distributed across the video. This step takes about 1-2 mins to complete on a 3mb video.

### 🚨 PERFORMANCE WARNING!
Using large videos or capturing a large number of frames is really memory intensive and could crash your n8n instance. Be sure you have sufficient memory and to optimise the video beforehand!

### 3. Use Vision AI to Narrate on Batches of Frames
To keep within token limits of our LLM, we’ll need to send our frames in sequential batches to represent chunks of our original video. We’ll use the loop node to create batches of 15 frames - this is because of our max of 90 frames, this fits perfectly for a total of 6 loops. A wait node is used to stay within service rate limits. This is useful for new users who are still on lower tiers. If you do not have such restrictions, feel free to remove this wait node!

### 4. Generate Voice Over Clip Using TTS
Finally, with our generated script parts, we can combine them into one and use OpenAI’s Audio generation capabilities to generate a voice over from the full script.

**The video used in this demonstration is** © [Coverr-Free-Footage](https://pixabay.com/users/coverr-free-footage-1281706/) via [Pixabay](https://pixabay.com/videos/india-street-busy-rickshaw-people-3175/)

### Sample Output
[Listen to the finished product here](https://drive.google.com/file/d/1-XCoii0leGB2MffBMPpCZoxboVyeyeIX/view?usp=sharing).

## Requirements
- OpenAI for LLM
- Ideally, a mid-range (16GB RAM) machine for acceptable performance!

## Customising this workflow
- For larger videos, consider splitting into smaller clips for better performance.
- Use a multimodal LLM which supports fully video such as Google's Gemini.
