Language Models Learn to Rank Research Ideas Before Testing
Researchers trained LMs to forecast which of two competing research ideas will perform better on a benchmark, without running experiments.
A fine-tuned 8B model reached 77.1% accuracy on 11,488 idea pairs from PapersWithCode, beating GPT-4's 61.1%.
Solves a bottleneck: as AI generates hundreds of hypotheses, humans need a fast filter before costly empirical work.