Kandinsky 6.0 Video Adds Native Audio
Kandinsky 6.0 Video brings 3B Lite and 29B Pro models to fal for five-second text- and image-to-video generation with synchronized 44 kHz audio, lip-sync, and built-in upscaling. The models and tooling are released under MIT, positioning Kandinsky as an unusually open alternative for audio-video generation.
Native audio is the meaningful upgrade here—not another silent video endpoint. Kandinsky’s open release and Lite/Pro split make it practical for both rapid prototyping and higher-fidelity production, though five-second outputs and a 29B Pro model still impose real infrastructure costs.
- –Lite targets fast, lower-cost iteration, while Pro prioritizes cinematic quality and speech performance.
- –Synchronized audio and lip-sync reduce the need for separate sound-generation and post-production pipelines.
- –Built-in Full HD upscaling, plus standalone VSR tools, supports a cleaner path from draft generation to delivery.
- –MIT-licensed weights, code, and Diffusers integration make self-hosting and fine-tuning viable beyond fal. [Kandinsky Lab](https://www.kandinskylab.ai/)
- –Developers should benchmark latency, cost, and temporal consistency against leading closed video models before committing to production.
DISCOVERED
1h ago
2026-10-06
PUBLISHED
1h ago
2026-10-06
RELEVANCE
AUTHOR
fal