Google DeepMind has released a new paper describing an inference-time method called "recirculation" that feeds deep activation information back into earlier layers. According to the report, the approach cuts perplexity without retraining and can deliver results comparable to full fine-tuning. Tests were run on the Gemma3 model family, including 1B, 4B, and 12B parameter versions. The base variant reduced perplexity by 8.5%, while an adaptive version reached a 23% reduction. The paper says the trade-off is higher processing cost during the prefilling stage, but no added latency during generation. The technique has also been independently reproduced on models including Llama 3.2 1B. The item was cited by Techub, with CryptoBriefing named as the source in the brief.
Google DeepMind has published a paper introducing an inference-time technique called "recirculation," according to a Techub News brief citing CryptoBriefing. The method feeds activation information from deeper layers back into shallower ones and, according to the report, reduces model perplexity by 23% without retraining. Its performance was described as comparable to full fine-tuning.
The paper tested the approach on the Gemma3 model family, covering 1B, 4B, and 12B parameter versions. In those tests, the base version reduced perplexity by 8.5%, while the adaptive version reached 23%.
The report said the method raises processing costs during the prefilling stage, but adds no latency during generation.
The technique has also been independently reproduced on models including Llama 3.2 1B, the brief said.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.