Recirculating a model's states lifts its accuracy
A DeepMind preprint posted on 18 August 2026 describes recirculation, an inference-time change that lets an off-the-shelf transformer reuse its own hidden states, adding recurrence without retraining the weights. The authors say generation latency barely changes, while the prefill step becomes serial. On the Gemma 3 family, an adaptive version lowered perplexity across a suite of datasets and raised accuracy on GSM8K by 21 percent relative to the untouched baseline, with gains on other tasks as well. The method is a research modification, not a new trained model.
The record ends here. Everything below this line is invented.
What if a model could rethink a line without taking longer to speak?
Part one
The Second Pass
Jules filed load tickets for a crane yard that wanted the number before the lift. The model on his screen was the same one as last year. Someone had turned on a second pass through its own states. It did not pause. It simply arrived surer.
He liked the surety on the word problems in the qualification set. A grade-school item it used to miss, it now cleared, the hidden tally having looped once before the sentence formed. The vendor called that free.
A real ticket was not a word problem. Jules sent a clearance for four tonnes, the figure on the screen, and heard the printer in the yard take the slip.
The second pass finished a moment later. The figure on his screen was two.
The printer kept going. Jules stood up with the first slip still warm in the tray and the corrected number only on the glass.
What’s real
- Recirculation reuses a model's states without new training
- The authors report little added latency while speaking
- GSM8K accuracy rose 21 percent relative on Gemma 3
- Prefill becomes serial
What’s invented
- Jules's crane-yard printer and four-tonne slip
- A second pass that finishes after the send
- A rule that the first number is the one that prints
- A lift already waiting on the paper
Part two: The Slip in the Tray
Part two isn’t on sale yet. Check back soon.
Filed under technology, inference, language models, reasoning.