NVIDIA PixelUMM releases pixel-space multimodal model
PixelUMM is an encoder-free NVIDIA model that processes raw image patches and video tubelets directly, supporting image and video understanding plus generation through a shared Qwen3-8B-based Transformer. The checkpoint is available for non-commercial research and evaluation. [Project page](https://nv-tlabs.github.io/PixelUMM/) [Model card](https://huggingface.co/nvidia/PixelUMM)
PixelUMM is a compelling architecture experiment, but its raw-pixel simplicity comes with substantial compute and quality tradeoffs.
- –Removes both the vision encoder and VAE, using a shared pixel-space interface for understanding and generation.
- –Handles text-to-image, image-conditioned text, video understanding, and video generation in one checkpoint.
- –Uses 16×16 image patches, four-frame video tubelets, and separate understanding/generation experts with shared attention.
- –Developers should expect research-grade limitations, including prompt-following issues, temporal inconsistencies, patch artifacts, demanding NVIDIA GPU requirements, and non-commercial model weights. [Technical details](https://github.com/nv-tlabs/PixelUMM)
DISCOVERED
1h ago
2026-10-04
PUBLISHED
1h ago
2026-10-04
RELEVANCE
AUTHOR
AI Search