BackAIRA-dojo

AIRA-dojo

Meta FAIR
2026-09-06 20:36:51

Meta FAIR releases RPM to rank AI experiments before costly GPU runs

Meta FAIR, working with research teams from the University of Oxford and University College London, has introduced an AI Research Preference Model, or RPM, designed to sort candidate experiment plans before running expensive GPU-heavy tests. The idea is to let an AI research agent spend compute on the most promising options instead of trying every possible setup. The model comes in two versions: an inference-only variant and an agent-driven variant. Both use a frozen pretrained large model, Qwen3.6-27B. The team also open-sourced the related AIRA-dojo framework and the AIRS-Bench benchmark. On AIRS-Bench, RPM improved the average normalized score from 0.684 under random selection to 0.711 for the inference-only version and 0.729 for the agent-driven version. In efficiency terms, both variants reached the baseline model’s 24-hour performance level in about 15 hours, translating to roughly 1.5x to 1.6x speedup. The research also set new state-of-the-art results on the WinoGrande and SVAMP benchmarks, according to the summary published by Techub citing MarkTechPost.

20
Meta FAIR releases RPM to rank AI experiments before costly GPU runs