MOSS-TTS is a new open-source speech and sound generation model family that unifies voice cloning, sound effects, and real-time text-to-speech in one framework.
MOSS-TTS is a unified, open-source speech and sound generation model family capable of handling voice cloning, sound effects generation, and low-latency streaming text-to-speech. Designed for high-fidelity and expressive audio generation, it supports multi-speaker dialogue and long-form speech synthesis, while also offering zero-shot voice design from text descriptions. The suite contains multiple specialized models, including lightweight variants optimized for CPU environments, providing a flexible and license-friendly alternative to commercial speech APIs.
Unifying voice cloning, dialogue generation, and sound effects under a single open-source ecosystem lowers the barrier for developers building complex audio applications.
- –Unified Architecture: Combining diverse audio generation tasks in one framework simplifies developer integration and reduces pipeline overhead.
- –Local and Efficient Execution: The inclusion of CPU-optimized lightweight models allows for offline, cost-effective deployments on edge devices.
- –Advanced Voice Customization: Support for instruction-driven voice generation from text prompts bypasses the need for high-quality audio references.
DISCOVERED
51d ago
2026-06-10
PUBLISHED
51d ago
2026-06-10
RELEVANCE
AUTHOR
GithubProjects