Paper
Pretraining Validation Loss Predicts Reinforcement Learning Gains: A Study on LLM Pipeline for Chess
The study finds that pretraining validation loss can predict pass@1 accuracy after RL with high precision (correlation |ρ| improves from 0.93 to 0.99), and the reward gain per 10x RL compute is positively correlated with the logarithm of pretraining tokens (r=+0.84). For a 680M parameter model, the optimal RL compute ratio is about 28%. Mechanistically, RL does not uniformly enhance reasoning but instead generates correct actions for hard problems while amplifying already preferred incorrect actions by several times, thus improving only pass@1 but not pass@16.
Read the original (opens in a new tab)
News stream data aggregated by AI HOT