Notes on the Foundations of Reinforcement Learning and LLM Post-Training (Part 2)
Part 2 covers RL in the LLM/VLM setting, RLHF, the PPO training pipeline, and GRPO.
Part 2 covers RL in the LLM/VLM setting, RLHF, the PPO training pipeline, and GRPO.
Part 1 covers RL fundamentals, policy gradients, REINFORCE, actor-critic methods, TRPO, and PPO.