DiffusionGemma (by way of) Final Might Google briefly launched an experimental Gemini Diffusion mannequin. I attempted the preview on the time and recorded it operating at 857 tokens/second. It was an thrilling mannequin, however Google made no additional bulletins about it.
That analysis has returned in the absolute best method: as a brand new open weight (Apache 2 licensed) Gemma mannequin, google/diffusiongemma-26B-A4B-it.
NVIDIA are presently internet hosting the mannequin without cost on their NIM cloud API. I used that API to generate this pelican, which took 4.4s (in accordance with time uv run generate.py) to return 2,409 tokens – so at the least 500 tokens/second.

