Training
FreeTrain and fine-tune LTX-2 for audio-video generation.
About Training
LTX-2 Trainer is an open-source package by Lightricks for training and fine-tuning the LTX-2 audio-video generation model. It supports LoRA training and full fine-tuning with a flexible conditioning framework that covers text-to-video, text-to-audio, image-to-video, video extension, audio extension, video inpainting, audio inpainting, video outpainting, IC-LoRA for video, audio, and joint audio-video references, audio-to-video, and video-to-audio. The trainer includes a quick start guide, dataset preparation instructions, configuration reference, and troubleshooting guide. It requires an LTX-2 model checkpoint, a Gemma text encoder, Linux with CUDA 13+, and an Nvidia GPU with 80GB+ VRAM for standard config, with a low VRAM mode available for GPUs like the RTX 5090 with 32GB VRAM.
Key Features
Pros & Cons
- Open-source and free to use
- Supports multiple training modes (LoRA, full fine-tuning)
- Flexible conditioning for various input-output combinations
- Low VRAM configuration allows use on consumer GPUs like RTX 5090
- Active community on Discord and GitHub for support and collaboration
- Agent-assisted training automates data probing, mode selection, and monitoring
- Standard configuration requires Nvidia GPU with 80GB+ VRAM
- Only supports Linux with CUDA 13+ (no Windows or macOS native support)
- Requires separate model checkpoints (LTX-2 and Gemma text encoder) to be downloaded
- Limited to the LTX-2 model architecture; not a general-purpose trainer