EvolveScaler tests whether LLMs can keep up with changing information, with the best GPT-5.5 score at 59.3%

EvolveScaler tests whether LLMs can keep up with changing information, with the best GPT-5.5 score at 59.3%

N
News Editor
2026-09-15 09:11:00
Teams including Tencent Hunyuan have released EvolveScaler, a benchmark built to test whether large language models can reason over information that keeps changing over time. The setup injects updates, reversals, additions, and invalidations into long passages, then asks models to answer based on the latest state rather than retrieve a ready-made sentence from the text. One example starts with 40 days of game records and then asks a counterfactual question about what would happen if a small boss had not been defeated on day 7, forcing the model to recompute later equipment, health, and battle outcomes. The benchmark includes 117 task categories and 159 question types, with the longest sequences containing about 1,200 events. It evaluates 14 model configurations, including GPT-5.5, Gemini 3.1 Pro, DeepSeek V4 Preview Pro, GLM-5.2, and Qwen3.5 Plus. In the hardest setting, the median score was only 11.3%, while the top-performing GPT-5.5-xhigh reached 59.3%. The team also said the dataset can be used for training: after adding 6,000 such samples, an internal A3B model improved on all eight external tests, gaining an average of 5.25 points.

Teams including Tencent Hunyuan have released EvolveScaler, a benchmark designed to test whether large language models can keep up with information that keeps changing. The framework repeatedly inserts modifications, withdrawals, additions, and invalidations into long-form text, then asks models to answer according to the latest state.

How the benchmark works

One example begins with 40 days of game logs and then asks: 「If the small boss had not been defeated on day 7, could the player still win in the end?」 That single change would alter later equipment, health, and battle outcomes as well.

According to the description, the model must recompute everything from day 7 onward across the following dozens of days. The answer is not explicitly present in the original text.

117 task categories and 159 question types

The team built 117 task categories and 159 question types into the benchmark, with the longest cases spanning about 1,200 events. The evaluation covered 14 model configurations, including GPT-5.5, Gemini 3.1 Pro, DeepSeek V4 Preview Pro, GLM-5.2, and Qwen3.5 Plus.

Median score in the hardest setting was 11.3%

In the hardest tier, the median score was only 11.3%. Even the best-performing model, GPT-5.5-xhigh, scored 59.3%.

The dataset can also be used for training

The team said the dataset is not limited to evaluation. It can also be used to train models. After adding 6,000 samples of this type, an internal A3B model improved on all eight external tests, with an average gain of 5.25 points.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
8100

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.