MarsBit’s featured article, titled “Behind AI Scorecards Is a Chinese ‘Question Setter,’” focuses on Chinese scholar Chen Wenhu and his team’s work on AI model evaluation. The report centers on three benchmarks developed by the team: MMLU-Pro, MMMU and MMMU-Pro. These benchmarks are presented as tools for assessing foundation models and for examining how AI systems perform under more demanding evaluation settings.
According to the article, the team’s contribution is not limited to extending earlier testing formats. Instead, the work restructures several core parts of AI evaluation, including the difficulty of questions, the design of answer options and the setup of multimodal tasks. Through these changes, MMLU-Pro, MMMU and MMMU-Pro are described as addressing the problem of inaccurate results in traditional evaluation methods, while helping model scores better reflect performance across different tasks.
The article also introduces Chen Wenhu’s academic background, his laboratory work and the team’s key contributions to the evaluation system for foundation models. MarsBit describes these benchmarks as having become part of the industry’s common standards, while bringing more attention to the people and design logic behind the “scorecards” used to judge AI model capability.

