Kandinsky 6.0 Video: The Open AI Video Model That Makes Its Own Sound, Under a License You Can Actually Sell
A free model that writes the dialogue track while it draws the shot. What it makes, what the MIT license lets you do with it, and why a five-second clip can still cost you twenty minutes.
The hardest part of open AI video with sound this year has not been the sound. It has been the small print.
Every free model that can draw a person and make them talk has arrived with a clause attached. One is free until your company earns ten million dollars a year. Another excludes whole regions of the world, including the United States, the UK and the EU, from its license entirely. You could download them, run them, and admire the results. Whether you could put the results in front of a paying client was a separate question, and the answer usually involved reading a legal document at midnight.
Kandinsky 6.0 Video, which its makers open-sourced on October 6, comes with a short license that anyone who has shipped open-source software will recognise: MIT. The team behind it, Kandinsky Lab, says it plainly in their paper's summary: they "release the code, model checkpoints, and diffusers integration under the MIT license." In plain terms: the code, the trained model files, and the plug-in for Hugging Face's popular video toolkit are all free to use commercially.
That one line is the most interesting thing about the release, and almost nobody will lead with it.
What it actually makes
Kandinsky 6.0 Video makes five-second clips with sound baked in. You type a description, or hand it a starting picture plus a description, and it gives back a video with a matching audio track at the same sample rate as most music releases (44 kHz): speech, room sound, effects. The makers say it handles lip-sync, so a character you describe as saying a line should move their mouth to the words.
It comes in two sizes. Lite is the small one, around 3 billion internal settings, which is what lets it fit on a home graphics card. Pro is roughly ten times bigger at 29 billion and needs more memory and patience. Each size also has a fast version, labelled "distill" in its file names, that finishes in 10 passes instead of 50. The clips come out at a modest size (about 480 by 864 pixels), and a separate upscaling step takes them to full HD, 1920 by 1080.
The way it works explains why the sound lines up. The model has two halves, one that draws frames and one that writes audio, and the two keep checking each other the whole time a clip is being made. The team trained the audio half first on sound alone, then trained both halves together on real video with real soundtracks. So the sound is not a dub laid over a silent clip afterwards. It grows alongside the picture, which is the likeliest way to get a mouth and a word onto the same frame.
Why the license is the story
Here is the comparison that matters if you make things for money.
The LTX-2 license, which covers Lightricks' popular open video models, is generous. It says the company "claims no rights in the Output you generate," and it only asks for a paid license from businesses earning $10 million or more a year. For most freelancers and small studios that is effectively free. Still, it is a clause, and a client's legal team will ask about it.
MiniMax H3's community license, which covers one of this season's most-discussed open video models with sound, carries a territory clause that, as I reported in September, excludes the EU, the UK, the Republic of Korea and the USA. If you work from any of those places, that model is something you can study and not much else.
MIT has no revenue line and no map. You can use it, change it, sell what you make with it, and build it into a product. The main obligation is keeping the copyright notice with the software if you pass the software on. Your rendered clip carries no such obligation.
I would not call that a small thing. Licensing has been the hidden tax on open creative AI all year, the reason a great free model sits unused in a studio while the team keeps paying for a hosted one. A model you can hand straight to a client changes which tool gets opened on a Monday morning.
The price is time
So what is the catch? The clock.
The Kandinsky README publishes a timing table for the full 50-pass models, measured after the model has warmed up and leaving out loading time. On an RTX 4090, a Lite clip takes about 437 seconds, a little over seven minutes. On a mid-range RTX 5060 Ti, the same clip takes about 1,310 seconds, roughly 22 minutes. Pro at full HD on a 4090 runs about 1,247 seconds, just under 21 minutes. A data-center H100 brings Lite down to around four minutes.
Twenty-two minutes for five seconds of video is a render, not a sketch. You will not sit and iterate on prompts at that pace.
The fast 10-pass versions should be much quicker, but the README gives no timings for them, and I would rather tell you that than guess. The honest plan is to time one run on your own card before you promise anybody a schedule.
This is the trade the whole open side is offering right now. You stop paying per clip, and you start paying in waiting.
Put this into practice
Here is how I would spend a first afternoon with it.
1. Start in the browser. Kandinsky Lab runs a free demo of the fast Pro version on Hugging Face, under the name Kandinsky-6.0-Pro-distill-5s. It runs on free shared hardware, so expect a queue, but it costs nothing and needs no install. Use it to answer one question before you download anything: does the mouth match the words?
2. Write for five seconds. Five seconds holds one line of dialogue, one action and one sound, and not much more. A prompt like "a woman at a rainy bus stop looks up from her phone and says, 'You're late again,' traffic hissing past" gives the model a single beat to land. A paragraph describing a whole scene gives it nothing to prioritise. Treat each clip as a single shot from a storyboard, not a scene.
3. Use a starting image for continuity. The model also accepts a picture to begin from. If you need the same character across several clips, design them once as a still, then start every clip from that still. That will hold a face far better than describing it again in words each time.
4. Install it locally when you have a keeper. The project uses two small command-line helpers. After cloning the repository from GitHub (kandinskylab/kandinsky-6), the README's quick start is three commands: just setup to install everything for your card, just download pro-distill to fetch the fast Pro version, and just generate "your prompt" to render. Each result lands in its own dated folder along with the expanded prompt the model actually used, which is worth reading because it shows you how your words were interpreted. If you live in ComfyUI, install the kandinsky6 and kandinsky6-sr nodes through ComfyUI Manager instead.
5. Render overnight, draft by day. Queue the clips that passed the browser test, let the machine work while you sleep, and upscale only the ones you will use. That rhythm turns the slow part into something you never watch.
Where it falls short
Five seconds is the ceiling. There is no long-shot mode. Anything longer is an edit of several clips, and joining clips with matched sound is your job.
It needs an NVIDIA card. The README requires NVIDIA hardware and a recent version of Python. Mac and AMD users are on the browser demo or a rented machine.
The quality claims are the makers' own. The paper says Pro "remains competitive with leading audio-video generation models, particularly in speech quality," and the paper reports a 47% drop in Pro's speech errors after a round of extra training. Those numbers come from the team's own tests. Nobody outside has checked them yet, a day after release.
The fast version has no published speed. Covered above, and worth repeating because it is the number that decides whether this fits your week.
Language support is not spelled out. The examples are in English. The README I read does not say which languages the speech handles well, so test yours before you plan a project around it.
The license covers the model, not your prompt. MIT lets you sell what the model makes. It says nothing about a real person's face, a real actor's voice or a trademark you describe into the scene. Those rules are the same as they have always been.
It is very new. The code repository had 138 stars on GitHub when I checked, one day in. Expect rough installs and quick changes.
The question worth asking
For two years, the useful question about any AI video tool has been "how good does it look?" Kandinsky 6.0 Video will be compared on that question, and it may well lose some of those comparisons to the paid models.
I think the better question for a working creative is different: "Can I hand this to a client without a phone call to a lawyer?" Among the open models that make picture and sound together, this one gives the plainest yes I have seen: no threshold to track, no map to check.
Try one line of dialogue in the free demo today. If the mouth lands on the word, you have found a tool you can own outright, and the only thing it asks of you is patience.
Medium metadata
Title: Kandinsky 6.0 Video: The Open AI Video Model That Makes Its Own Sound, Under a License You Can Actually Sell
Subtitle: A free model that writes the dialogue track while it draws the shot. What it makes, what the MIT license lets you do with it, and why a five-second clip can still cost you twenty minutes.
Tags: AI Video, Filmmaking, Open Source, Video Production, Generative AI
Estimated read time: 8 minutes