Meta FAIR, the University of Oxford, and University College London introduce AI Research Preference Models (RPMs) to rank unexecuted machine learning experiments before consuming GPU hours. Operating within the AIRA-dojo evolutionary tree search, RPMs evaluate 15 parallel candidates using frozen pretrained LLMs. The inference-only and agentic variants achieve 1.5x to 1.6x training speedups on AIRS-Bench and set new state-of-the-art scores of 94.1% on WinoGrande and 95.7% on SVAMP.
AI Research Preference Models
- ▪AI Research Preference Models rank unexecuted candidates and pick one to execute rather than forecasting absolute scores or execution outcomes.
- ▪AI Research Preference Models utilize frozen pretrained large language models with no fine-tuning.
- ▪Meta FAIR, the University of Oxford, and University College London introduced AI Research Preference Models to rank unexecuted machine learning experiment candidates.
AIRA-dojo evolutionary tree search
- ▪The AIRA-dojo scaffold is an evolutionary tree search that uses greedy parent selection, Draft, Improve, and Debug operators.
- ▪In the AIRA-dojo scaffold, the AI Research Preference Model generates 15 unexecuted candidates in parallel and compares them pairwise in a knockout tournament.
Inference-only RPM variant
- ▪The prompt for the inference-only AI Research Preference Model was optimized using MIPROv2 from DSPy, achieving an offline accuracy of 57.7% to 59.0%.
- ▪The inference-only AI Research Preference Model variant acts as a large language model judge over candidate plans, code, and search history.
Agentic RPM variant
- ▪The agentic AI Research Preference Model variant runs small-scale pilot experiments capped at 30 with a 60-second threshold.
- ▪The agentic AI Research Preference Model variant utilizes a sandbox environment cloning the agent's environment, including a single NVIDIA H200 GPU.
AIRS-Bench performance results
- ▪The inference-only AI Research Preference Model reached the baseline's final score in 14.88 hours, representing a 1.61 times speedup.
- ▪On the AIRS-Bench benchmark, the average normalized score rose from 0.684 to 0.711 and 0.729 when using AI Research Preference Models.
Story comments
Loading comments…