Google DeepMind unveils an encoder-free multimodal architecture in Gemma 4 12B that ingests raw images and audio directly.
Google DeepMind's Gemma 4 family introduces an innovative 12B variant featuring a unified, encoder-free multimodal architecture. By streaming raw image patches and 40ms audio chunks directly into the main transformer backbone, this approach completely removes the need for separate vision and audio encoders, significantly simplifying multimodal model pipelines and reducing processing overhead.
Eliminating visual and audio pre-encoders is a massive architectural shift towards true natively unified multimodal models.
- –Streamlines inference pipelines by removing auxiliary feature extraction models.
- –Directly ingests 40ms audio chunks and image patches into the main transformer backbone.
- –Demonstrates how open-weight model architectures continue pushing efficiency and integration boundaries.
DISCOVERED
46d ago
2026-08-07
PUBLISHED
46d ago
2026-08-07
RELEVANCE
AUTHOR
Two Minute Papers