Intro

Nathan Labenz introduces Kyle Corbitt, frames the episode as an RL master class, and previews topics including GRPO, reward hacking, environments, and distillation.

Play episode from 00:00

Transcript

The AI-powered Podcast Player

Save insights by tapping your headphones, chat with episodes, discover the best highlights - and more!

Get the app