USTC2026-08-06 08:25:09USTC study tests whether AI can run a real chemistry labA new University of Science and Technology of China study pushed AI agents beyond text-based planning and into a real machine-catalysis lab. The setup includes 45 modular automated workstations covering synthesis, characterization, and catalytic testing, wrapped as machine-readable skills that agents can call under real equipment constraints. Researchers evaluated 48 configurations across six agent frameworks and nine large language models, running 4,608 trials on 32 expert-defined research tasks. Only 151 workflows, or 3.3%, could execute without manual repair. The best result came from Claude Code with Claude Opus 4.7, at 28.1%, followed by Codex with GPT 5.5 at 19.8%. In a five-round closed loop, Codex/GPT 5.5 could adjust formulations and operating conditions, but it did not redesign analysis methods or fix persistent omissions such as missing electrode binders and assay-specific color reagents. The paper separates three different abilities often lumped together in the “AI scientist” debate: writing an experimental plan, producing a workflow that can actually run in a physical lab, and changing the broader research strategy after seeing experimental results.1900