Google releases Gemma 4 12B, an encoder-free multimodal open model
The 12-billion-parameter model folds vision and audio processing directly into the language backbone rather than using separate encoders, and runs on 16GB of memory.
Open weights & ecosystem