ModelsOpen Source 🇺🇸 27.07.2026 12:03

Google DeepMind Unveils Gemma 4 12B: A Unified Multimodal Model Without Encoders

Google/DeepMindGoogle/DeepMind Google DeepMindGoogle DeepMind
Google DeepMind has announced Gemma 4 12B, a compact multimodal model capable of processing text, images, and audio without separate encoders. The model runs on standard laptops with 16 GB of RAM, outperforms many solutions in performance, and is available under the Apache 2.0 license.
Google DeepMind has unveiled Gemma 4 12B, a new medium-sized model that integrates text, image, and audio processing capabilities without using separate encoders. Unlike traditional multimodal models that rely on dedicated encoders to translate images and audio, Gemma 4 12B uses a lightweight embedding module for vision and projects raw audio signals directly into the text token space. The model achieves performance close to the larger Gemma 4 26B MoE model while consuming less than half the memory and can run locally on consumer laptops with 16 GB of RAM. Gemma 4 12B supports Multi-Token Prediction (MTP) to reduce latency and is released under the Apache 2.0 license. Model weights are available on Hugging Face and Kaggle, and developers can use popular tools such as Hugging Face Transformers, llama.cpp, MLX, SGLang, and vLLM. An official Gemma Skills library has also been released to support agents working with the model.
Source: Google DeepMind — original
Our earlier posts on this topic ↓
Fresh news