Iris-3B Goes Open, Skips VAE
Sperid Labs released Iris-3B, a fully open 3B-parameter pixel-space diffusion transformer that generates images without a VAE and fine-tunes for monocular depth and 4× restoration. The model matches Qwen-Image on OneIG while shipping weights, training code, and a demo under Apache 2.0. [GitHub](https://github.com/speridlabs/iris-3b) [Paper](https://huggingface.co/papers/2610.09450)
Iris-3B is a meaningful scaling result for pixel-space generation, but its own evaluations temper the VAE-free hype: the representation is viable, not yet clearly superior.
- –The 256→512→1024 training curriculum shows pixel-space diffusion can scale to 3B parameters and competitive text-to-image quality.
- –The same backbone supports generation, depth estimation, and restoration, offering an unusually broad open research platform.
- –Depth performance roughly matches latent FLUX.2 Klein, while 4× restoration does not beat the latent baseline.
- –Developers should expect research-grade tradeoffs, including patch-grid artifacts and potentially slower inference at high resolution. [AI Weekly](https://aiweekly.co/alerts/speridlabs-scales-iris-3b-to-3b-finds-no-downstream-edge) [Community discussion](https://www.reddit.com/r/StableDiffusion/comments/1x17iqb/iris3b_speridlabs_research/)
DISCOVERED
1h ago
2026-10-11
PUBLISHED
1h ago
2026-10-11
RELEVANCE
AUTHOR
AI Search