
technologyMay 1, 20261:46:35pending
The RL Fine-Tuning Playbook: CoreWeave's Kyle Corbitt on GRPO, Rubrics, Environments, Reward Hacking
About this episode
Kyle Corbitt, founder of OpenPipe, breaks down reinforcement learning and custom fine-tuning for modern AI models. He explains how RL differs from supervised fine-tuning, why GRPO and LLM-as-judge post-training matter, and how these techniques can improve performance, latency, and cost on open source models. The conversation also covers reward hacking, evaluation design, LoRA adapters, and how Chinese labs are using distillation to fast-follow frontier models.
Sponsors:
Sequence:
Sequence handles the full revenue workflow for complex pricing, from quoting and metering to invoicing, revenue recognition, and collections. Book a public demo at https://sequencehq.com and use code Cognizant in the source field to save 20% off year one
AvePoint:
AvePoint is building the control layer for AI agents so you can securely govern, audit, and recover every action at scale. Design trusted agentic outcomes from day one at https://avpt.co/tcr
VCX:
VCX, by Fundrise, is the public ticker for private tech, giving everyday investors access to high-growth private companies in AI, space, defense tech, and more. Learn how to invest at https://getvcx.com
Claude:
Claude by Anthropic is an AI collaborator that understands your workflow and helps you tackle research, writing, coding, and organization with deep context. Get started with Claude and explore Claude Pro at https://claude.ai/tcr
Get every episode summarized
Each time "The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
No transcript yet
This episode has not been transcribed. Request it and it moves to the front of the queue.
More episodes
More from "The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

Nathan Goes to China #3: US-China Relations, the Art of the AI Deal & the Road t...
"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis
Sep 10, 20263:17:07completed

AI:AM Highlights: Welcome to the AGI Era
"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis
Sep 5, 20262:20:44completed

Write, Change, Recall, Forget: MongoDB's Pete Johnson on How Retrieval Drives Ag...
"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis
Sep 1, 20261:36:43pending

AI:AM Highlights: Recursive Self-Improvement, Rushed and Vibe-Coded?
"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis
Aug 28, 20262:11:56pending