Designing Reward Functions in Open R1: Code Execution Verification and Reasoning Process Evaluation
Architecture of the Reward System
The Open R1 project, a fully open-source reproduction of DeepSeek-R1, relies heavily on a sophisticated reinforcement learning pipeline. The core of its performance lies in the precise construction of its reward mechanisms. The project implements a multi-faceted scoring system within its training stack—specific ...
Posted on Wed, 05 Aug 2026 16:53:56 +0000 by hamishrock
Implementing RLHF from Scratch with PPO and RLOO
Reinforcement Learning from Human Feeedback (RLHF) aligns language models with human preferences through a multi-stage process. Modern large language model pipelines commonly incorporate RLHF, often combining techniques like Direct Preference Optimization (DPO) followed by Proximal Policy Optimization (PPO). This guide focuses on implementing t ...
Posted on Mon, 03 Aug 2026 16:49:31 +0000 by alsal
Reinforcement Learning: Q-values and V-values, Monte Carlo and Temporal Difference Methods
Table of Contents
Introduction
Uncertainty in Reinforcement Learning
Q-values and V-values
a. V-value
b. Q-value
Monte Carlo and Temporal Difference Methods
a. Monte Carlo Method (MC)
i. MC Estimation Algorithm
ii. Relationship between G and V
iii. V-value Estimation Formula
iv. V-value and Policy Correlation
v. Characteristics of MC
b. Tempor ...
Posted on Mon, 18 May 2026 16:25:07 +0000 by v4g