logo
episode-header-image
Apr 2025
51m 45s

Teaching LLMs to Self-Reflect with Reinf...

Sam Charrington
About this episode
Today, we're joined by Maohao Shen, PhD student at MIT to discuss his paper, “Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search.” We dig into how Satori leverages reinforcement learning to improve language model reasoning—enabling model self-reflection, self-correction, and exploration of alternative ... Show More
Up next
Jul 27
Why Models Are AI’s Next Training Dataset with Damian Borth - #772
For more than a decade, AI has advanced by training ever-larger models on ever-larger datasets. But as high-quality training data becomes harder to find and pretraining grows increasingly expensive, researchers are looking for new ways to keep foundation models improving. In this ... Show More
47 m
Jul 8
How AI Learns to Smell with Alex Wiltschko - #771
In this episode, Alex Wiltschko, founder and CEO of Osmo, joins the show to discuss his goal of giving computers a sense of smell and what it takes to build olfactory intelligence. We explore the science behind smell, from the hundreds of olfactory receptors in the human nose to ... Show More
59m 55s
Jun 16
Why AI Agents Break the GenAI Security Model with Devvret Rishi - #770
In this episode, Sam talks with Dev Rishi, GM of AI at Rubrik, about what happens when agents move beyond answering questions and start taking action across tools, systems, and business processes. We explore why the enterprise playbook of static guardrails plus human approval sta ... Show More
56m 18s
Recommended Episodes
Aug 2023
Cuttlefish Model Tuning
<p>Hongyi Wang, a Senior Researcher at the Machine Learning Department at Carnegie Mellon University, joins us. His research is in the intersection of systems and machine learning. He discussed his research paper, Cuttlefish: Low-Rank Model Training without All the Tuning, on tod ... Show More
27m 8s
Jul 2023
#130 Mathew Lodge: The Future of Large Language Models in AI
<p data-pm-slice="1 1 []">Welcome to episode #130 of Eye on AI with <a href="https://www.linkedin.com/in/mathew/">Mathew Lodge</a>. In this episode, we explore the world of reinforcement learning and code generation. Mathew Lodge, the CEO of Diffblue, shares insights into how rei ... Show More
49m 44s
Mar 2021
Goodhart's Law in Reinforcement Learning
tail spinning
37m 11s
Sep 2023
Unlocking Language Models: The Power of Prompt Engineering (Ep. 238)
<p>Join me on an enlightening journey through the world of prompt engineering. Explore the multifaceted skills and strategies involved in harnessing the potential of large language models for various applications. From enhancing safety measures to augmenting models with domain kn ... Show More
28m 8s
Sep 2023
The Defeat of the Winograd Schema Challenge
Our guest today is Vid Kocijan, a Machine Learning Engineer at Kumo AI. Vid has a Ph.D. in Computer Science at the University of Oxford. His research focused on common sense reasoning, pre-training in LLMs, pretraining in knowledge-based completion, and how these pre-trainings im ... Show More
31m 3s
Feb 2025
LLMs and Graphs Synergy
<p>In this episode, Garima Agrawal, a senior researcher and AI consultant, brings her years of experience in data science and artificial intelligence. Listeners will learn about the evolving role of knowledge graphs in augmenting large language models (LLMs) for domain-specific t ... Show More
34m 47s
Jul 2023
Computable AGI
<p>On today's show, we are joined by Michael Timothy Bennett, a Ph.D. student at the Australian National University. Michael's research is centered around Artificial General Intelligence (AGI), specifically the mathematical formalism of AGIs. He joins us to discuss findings from ... Show More
36m 13s
Jul 2020
GPT-3, Limits of Deep Learning, Deepfakes in the Real World
Stanford AI Lab PhDs Andrey Kurenkov and Sharon Zhou discuss this week's major AI news stories. Check out all the stories discussed here and more at www.skynettoday.com Theme: Deliberate Thought Kevin MacLeod (incompetech.com)See Privacy Policy at https://art19.com/privacy and Ca ... Show More
33m 52s