Right now, we’re introducing Gemma 4 12B, our newest mannequin designed to deliver agentic multimodal intelligence on to laptops. Bridging the hole between our edge-friendly E4B and our extra superior 26B Combination of Specialists (MoE), Gemma 4 12B packages highly effective capabilities inside a decreased reminiscence footprint. Additionally it is our first mid-sized mannequin to function native audio inputs.
Because of the developer group, Gemma 4 fashions have now crossed 150 million downloads. You’ve constructed all the things from wearable robotic arms for bodily help to enterprise-grade AI safety. We’re excited to see what you construct with this newest addition.
Right here’s an outline of what makes Gemma 4 12B distinctive:
Novel unified structure: No multimodal encoders. The imaginative and prescient and audio inputs circulation immediately into the LLM spine.Superior reasoning: Benchmark efficiency nearing our 26B mannequin, unlocking highly effective multi-step reasoning and agentic workflows.Laptop computer prepared: Sufficiently small to run regionally with simply 16GB of VRAM or unified reminiscence.Open and accessible: Launched beneath an Apache 2.0 license with assist throughout the developer ecosystem.Drafter-ready: Gemma 4 12B comes outfitted with Multi-Token Prediction (MTP) drafters to scale back latency.
Collectively, these options deliver superior multimodal capabilities to on a regular basis {hardware} with out sacrificing velocity or reasoning. Let’s now take a better take a look at how Gemma 4 12B achieves this.
Run state-of-the-art brokers regionally
Gemma 4 12B delivers efficiency nearing our bigger 26B MoE mannequin on customary benchmarks, however at lower than half the entire reminiscence footprint. Sufficiently small to run regionally on client laptops with 16GB of RAM, it unlocks highly effective multimodal and agentic experiences proper in your machine.

