Meta FAIR, working with research teams from the University of Oxford and University College London, has introduced an AI Research Preference Model, or RPM, designed to sort candidate experiment plans before running expensive GPU-heavy tests. The idea is to let an AI research agent spend compute on the most promising options instead of trying every possible setup.
The model comes in two versions: an inference-only variant and an agent-driven variant. Both use a frozen pretrained large model, Qwen3.6-27B. The team also open-sourced the related AIRA-dojo framework and the AIRS-Bench benchmark.
On AIRS-Bench, RPM improved the average normalized score from 0.684 under random selection to 0.711 for the inference-only version and 0.729 for the agent-driven version. In efficiency terms, both variants reached the baseline model’s 24-hour performance level in about 15 hours, translating to roughly 1.5x to 1.6x speedup. The research also set new state-of-the-art results on the WinoGrande and SVAMP benchmarks, according to the summary published by Techub citing MarkTechPost.
Meta FAIR and research teams from the University of Oxford and University College London have released an AI Research Preference Model, or RPM, built to rank candidate experiment plans before costly GPU runs begin. The goal is to let AI research agents run only the options with the strongest promise instead of spending compute across a wider set of trials.
RPM includes two variants: an inference-only version and an agent-driven version. The model uses a frozen pretrained large model, Qwen3.6-27B. The related AIRA-dojo framework and AIRS-Bench benchmark have also been open-sourced.
AIRS-Bench results
In tests on AIRS-Bench, RPM raised the average normalized score from 0.684 under random selection to 0.711 for the inference-only setup and 0.729 for the agent-driven setup.
On efficiency, both RPM variants reached the baseline model’s 24-hour performance level in about 15 hours, equal to roughly 1.5x to 1.6x acceleration.
Other benchmark gains
The research also posted new state-of-the-art results on the WinoGrande and SVAMP benchmarks.
The item was published by Techub and attributed to MarkTechPost.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.