cs.AI, cs.CL, cs.LG

From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space

arXiv:2604.14142v1 Announce Type: cross
Abstract: While reinforcement learning with verifiable rewards (RLVR) significantly enhances LLM reasoning by optimizing the conditional distribution P(y|x), its potential is fundamentally bounded by the base mo…