Tencent has introduced T1, a terminal-focused Agent model trained from Qwen3.5-122B-A10B for complex Linux command-line tasks. The model can call tools for more than 300 rounds within a single task, targeting longer and more demanding terminal workflows. On Terminal-Bench 2.1, T1 scored 64.0%, up from the base model’s 43.8%, a gain of 20.2 points. Tencent said most of that improvement came from reinforcement learning. Supervised fine-tuning, or SFT, lifted the score to 49.4%, and reinforcement learning added another 14.6 points after that. The team also prepared about 15,000 terminal tasks, each paired with automated tests. During training, the Agent could still receive rewards for completing only part of a task’s requirements rather than the full assignment. Tencent also described a training-inference mismatch in long tasks: the same content may be tokenized differently between execution and later training, while a mixture-of-experts model may route the sequence to a different set of experts. To address that, T1 records the generated tokens and the experts selected at the time of execution, then replays them during training. Tencent said this reduced the mismatch by about one-third.
Tencent has introduced T1, a terminal Agent model built by training Qwen3.5-122B-A10B for complex tasks in Linux terminal environments. The model is designed for long task chains and can make more than 300 consecutive tool calls within a single assignment.
On Terminal-Bench 2.1, T1 scored 64.0%, compared with 43.8% for the base model, for a total improvement of 20.2 points. Tencent said reinforcement learning accounted for most of that increase. SFT raised the score to 49.4%, and reinforcement learning added another 14.6 points afterward.
The team prepared about 15,000 terminal tasks, with each task paired with automated tests. During training, the Agent did not need to finish an entire task to receive rewards. It could earn partial rewards by completing some of the requirements.
Tencent also pointed to a problem that appears in long tasks. When the Agent is carrying out work and later training on that same experience, the same sentence can be segmented into different tokens, and a Mixture-of-Experts, or MoE, model may route it to a different group of experts. T1 records both the tokens generated at the time and the experts selected, then replays that path during training. According to the disclosed information, that cut the training-inference mismatch by about one-third.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.