An open model drafts text in blocks, not tokens
A DeepMind technical report posted on 31 July 2026 introduces DiffusionGemma, an open-weight model made by adapting a Gemma 4 mixture-of-experts model so it denoises blocks of 256 tokens instead of emitting one token at a time. The authors report about 20 tokens per forward pass and about 1,500 tokens a second on one H100 graphics processor, using under a tenth of the original model's training-token budget for the conversion. The model keeps a thinking mode, image input, and long context, and it can still generate left to right with some loss of score. They also note occasional repetition and a remaining gap versus the autoregressive model it started from.
The record ends here. Everything below this line is invented.
What if a sentence could change its beginning after the end existed?
Part one
The Rewritten First Word
Hana wrote captions for instruments that could not wait. The new drafter did not type. It held a block of 256 tokens in a fog and let every word look at every other word until the fog decided.
She used it because the first word and the last number had to agree. The old models committed to a sign before they had done the arithmetic, then apologized. This one sometimes printed the right sign on the first pass, the calculation having leaked backward through the block.
A pressure alarm needed a caption before the valve moved. Hana accepted a block that began CLEAR and ended with a number inside the limit. The valve controller took CLEAR.
On the drafter's screen the block had not locked. A later denoising step, chasing a unit it had misread, revised the opening word.
The caption now began HOLD. The valve had already started to open.
What’s real
- DiffusionGemma denoises blocks of 256 tokens
- The report gives about 1,500 tokens a second on one H100
- Conversion used under a tenth of the original training tokens
- The authors report a quality gap and rare repetition
What’s invented
- Hana's valve that obeys the first caption
- A block that revises CLEAR into HOLD after acceptance
- A pressure alarm on her shift
- A controller that cannot unread a word
Part two: The Word the Valve Heard
Part two isn’t on sale yet. Check back soon.
Filed under technology, text diffusion, language models, inference.