logo
episode-header-image
Apr 2025
51m 45s

Teaching LLMs to Self-Reflect with Reinf...

Sam Charrington
About this episode
Today, we're joined by Maohao Shen, PhD student at MIT to discuss his paper, “Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search.” We dig into how Satori leverages reinforcement learning to improve language model reasoning—enabling model self-reflection, self-correction, and exploration of alternative ... Show More
Up next
Oct 6
Why Jev Is Changing How We Build With AI with Diogo Almeida - #779
In this episode, Diogo Almeida, co-founder and CEO of TypeSafe, joins us to discuss Jev, TypeSafe’s recently released model for bringing fast, reliable intelligence directly into software. We explore the idea of “machine-native intelligence” and why Diogo believes models optimize ... Show More
1h 30m
Sep 29
From Math Olympiads to Navier-Stokes: How Fast Is AI Progressing? with Greg Burnham - #778
AI systems have gone from struggling with grade-school math to helping solve research problems that have resisted mathematicians for decades, including Navier-Stokes. In this episode, Greg Burnham, who leads AI capabilities research at Epoch AI, joins us to examine what that prog ... Show More
1h 7m
Sep 17
From Voice Agents to AI Avatars with Alexander Smola - #777
Voice AI has gotten remarkably good, but natural conversation remains a high bar. Small delays, awkward interruptions, or the wrong tone can quickly break the illusion—and adding vision and visual presence only raises the stakes. In this episode, Alex Smola, co-founder and CEO of ... Show More
1h 4m
Recommended Episodes
Aug 2023
Cuttlefish Model Tuning
<p>Hongyi Wang, a Senior Researcher at the Machine Learning Department at Carnegie Mellon University, joins us. His research is in the intersection of systems and machine learning. He discussed his research paper, Cuttlefish: Low-Rank Model Training without All the Tuning, on tod ... Show More
27m 8s
Jul 2023
#130 Mathew Lodge: The Future of Large Language Models in AI
<p data-pm-slice="1 1 []">Welcome to episode #130 of Eye on AI with <a href="https://www.linkedin.com/in/mathew/">Mathew Lodge</a>. In this episode, we explore the world of reinforcement learning and code generation. Mathew Lodge, the CEO of Diffblue, shares insights into how rei ... Show More
49m 44s
Dec 2024
Nature of Intelligence, Ep. 6: AI’s changing seasons
Guest: Melanie Mitchell, Resident Professor, Santa Fe InstituteHosts: Abha Eli PhobooProducer: Katherine MoncurePodcast theme music by: Mitch MignanoFollow us on:Twitter • YouTube • Facebook • Instagram • LinkedIn • BlueskyMore info:Tutorial: Fundamentals of Machine LearningLectu ... Show More
44m 1s
Mar 2021
Goodhart's Law in Reinforcement Learning
tail spinning
37m 11s
Sep 2023
Unlocking Language Models: The Power of Prompt Engineering (Ep. 238)
<p>Join me on an enlightening journey through the world of prompt engineering. Explore the multifaceted skills and strategies involved in harnessing the potential of large language models for various applications. From enhancing safety measures to augmenting models with domain kn ... Show More
28m 8s
Sep 2023
The Defeat of the Winograd Schema Challenge
Our guest today is Vid Kocijan, a Machine Learning Engineer at Kumo AI. Vid has a Ph.D. in Computer Science at the University of Oxford. His research focused on common sense reasoning, pre-training in LLMs, pretraining in knowledge-based completion, and how these pre-trainings im ... Show More
31m 3s
Feb 2025
LLMs and Graphs Synergy
<p>In this episode, Garima Agrawal, a senior researcher and AI consultant, brings her years of experience in data science and artificial intelligence. Listeners will learn about the evolving role of knowledge graphs in augmenting large language models (LLMs) for domain-specific t ... Show More
34m 47s
Jul 2023
Computable AGI
<p>On today's show, we are joined by Michael Timothy Bennett, a Ph.D. student at the Australian National University. Michael's research is centered around Artificial General Intelligence (AGI), specifically the mathematical formalism of AGIs. He joins us to discuss findings from ... Show More
36m 13s
Jul 2020
GPT-3, Limits of Deep Learning, Deepfakes in the Real World
Stanford AI Lab PhDs Andrey Kurenkov and Sharon Zhou discuss this week's major AI news stories. Check out all the stories discussed here and more at www.skynettoday.com Theme: Deliberate Thought Kevin MacLeod (incompetech.com)See Privacy Policy at https://art19.com/privacy and Ca ... Show More
33m 52s