
KittenTTS 2 Brings Local Voice Cloning
KittenTTS 2 is a 1.7B speech model that runs locally, clones voices from 5–30 seconds of audio, and includes 47 built-in voices with expression controls. Its packed weights are about 947 MiB, but they use the Stellon Labs Community License rather than Apache 2.0. [Model card](https://huggingface.co/KittenML/kitten-tts-2)
KittenTTS 2 makes local voice cloning substantially more practical, moving a capability usually associated with hosted APIs onto developer-controlled hardware. The important caveat is licensing: the code is Apache-2.0, but the weights carry commercial thresholds and attribution obligations.
- –CPU support and a llama.cpp fork reduce deployment friction, although 1.7B parameters make it much heavier than KittenTTS’s original 15M–80M models. [GitHub README](https://github.com/KittenML/KittenTTS)
- –In-context cloning eliminates fine-tuning, enabling private narrators, accessibility tools, game characters, and local voice agents.
- –Built-in voices, multilingual support, and inline emotion controls make it more useful than basic waveform-generation libraries.
- –Organizations exceeding $1 million in annual revenue or cumulative funding need a separate license, and distributions require “Powered by Stellon Labs” attribution. [License](https://huggingface.co/KittenML/kitten-tts-2/blob/main/LICENSE.md)
DISCOVERED
1h ago
2026-10-07
PUBLISHED
1h ago
2026-10-07
RELEVANCE
AUTHOR
Better Stack