From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space
arXiv:2604.14142v1 Announce Type: cross
Abstract: While reinforcement learning with verifiable rewards (RLVR) significantly enhances LLM reasoning by optimizing the conditional distribution P(y|x), its potential is fundamentally bounded by the base mo…