Reinforcement Learning

Notes on the Foundations of Reinforcement Learning and LLM Post-Training (Part 2)

Part 2 covers RL in the LLM/VLM setting, RLHF, the PPO training pipeline, and GRPO.

avatar
Chi Phan
•

Notes on the Foundations of Reinforcement Learning and LLM Post-Training (Part 1)

Part 1 covers RL fundamentals, policy gradients, REINFORCE, actor-critic methods, TRPO, and PPO.

avatar
Chi Phan
•