Tencent Hunyuan, together with Tsinghua University and Peking University, has introduced IWC-Bench, a benchmark designed to evaluate how usable AI-generated webpages are in real-world use. Instead of relying mainly on code inspection or screenshots, the benchmark has an agent open the webpage, click buttons, enter text, switch pages, and then verify which functions actually run. The system scores three areas: visual quality, ease of use, and whether the requested user functions were delivered. According to the release, IWC-Bench includes 369 real user requests and 5,088 acceptance criteria, while 17 models generated a total of 6,273 web applications for testing. In a separate comparison covering 197 manually reviewed cases, IWC-Bench matched human choices in 168 cases, for an agreement rate of 85.3%. The results also suggest that a page looking polished does not necessarily mean it is usable. GPT-5.6-Sol ranked first in visual appeal, but only sixth in usability. For individual webpages, the correlation between appearance and usability was 0.36.
Tencent Hunyuan, working with Tsinghua University and Peking University, has released IWC-Bench, a benchmark built to measure how AI-generated webpages perform in actual use.
Many existing webpage-generation evaluations focus on code or screenshots. That can miss problems that only show up during real interaction. A page may look polished, and the code may include the intended functions, yet buttons may not respond, forms may fail to submit, or switching pages may trigger errors.
How IWC-Bench evaluates usability
IWC-Bench is designed more like a human acceptance test. It has an agent open the webpage, click buttons, enter content, move between pages, and then check which functions actually work in practice.
The benchmark scores three things separately: how the page looks, how smooth the interaction is, and whether the functions requested by the user were actually delivered.
Benchmark size and results
According to the release, IWC-Bench includes 369 real user requests and 5,088 acceptance criteria. Across the test set, 17 models generated 6,273 web applications.
In another comparison based on 197 manually reviewed cases, IWC-Bench agreed with human selections in 168 cases, giving it an agreement rate of 85.3%.
Looking good is not the same as working well
The results indicate that a webpage that looks good is not necessarily one that works well. GPT-5.6-Sol ranked highest on visual appeal, but placed sixth on usability.
At the level of individual webpages, appearance was only weakly linked to usability, with a correlation coefficient of 0.36.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.