Reinforcement learning trains a policy to choose actions that maximize expected cumulative reward. The learner interacts with an environment or generated tasks, observes outcomes and updates its behavior according to the reward signal.
For reasoning models, verifiable tasks can provide scalable rewards when outputs can be checked automatically. The quality of the learned behavior depends on whether the reward accurately represents the intended goal and resists exploitation.
