logo
episode-header-image
Jul 2025
1h 28m

903: LLM Benchmarks Are Lying to You (An...

Jon Krohn
About this episode
Has AI benchmarking reached its limit, and what do we have to fill this gap? Sinan Ozdemir speaks to Jon Krohn about the lack of transparency in training data and the necessity of human-led quality assurance to detect AI hallucinations, when and why to be skeptical of AI benchmarks, and the future of benchmarking agentic and multimodal models. Additional ... Show More
Up next
Sep 15
1027: Building an Always-On AI Agent for Busy Parents, with Dr. Dilani Kahawala
In Episode #1027, Dr. Dilani Kahawala (Co-Founder and CEO of Anna) joins Jon Krohn to explain what it takes to build an always-on AI assistant that busy parents will trust with their inboxes. Anna watches the email, school apps, WhatsApp messages and calendars flowing into a fami ... Show More
1h 2m
Sep 11
1026: OpenAI’s GPT-6 Astra
In Episode #1026, Jon Krohn breaks down GPT-6 Astra, OpenAI’s new flagship that its president has floated as a possible marker of AGI. Jon covers what the model is, what it costs, its state-of-the-art results across computer use, coding, abstract reasoning and science and the saf ... Show More
19m 37s
Sep 8
1025: Word Gravity: How Transformers Bend Space, with Dr. Luis Serrano
In Episode #1025, Dr. Luis Serrano (Founder of Serrano Academy) joins Jon Krohn to explain the paper he co-authored on the curved spacetime of transformer architectures, in which attention stops being a lookup table and becomes something closer to gravity: words bend the space ar ... Show More
1h 10m
Recommended Episodes
Nov 2024
The Future of AI: Predictions and Realities
In this episode, Jaeden Schafer discusses the current challenges and developments in the AI industry, particularly focusing on the limitations faced by major players like OpenAI and Anthropic. The conversation explores the anticipated improvements in AI models, the predictions fo ... Show More
19m 52s
Jan 2024
Why AI Should Be Taught to Know Its Limits
One of AI’s biggest, unsolved problems is what the advanced algorithms should do when they confront a situation they don’t have an answer for. For programs like Chat GPT, that could mean providing a confidently wrong answer, what’s often called a “hallucination”; for others, as w ... Show More
14m 58s
Jul 2022
Why Artificial Intelligence Projects Fail ?
Podcast with Gautam Siwach and Jin Vanstee ! Speaker - Elpida Tzortzatos is an IBM Fellow and CTO AI on IBM zSystems. In this Podcast we listen to Elpida's thoughts about driving Artificial Intelligence strategies, associated risks, and Values.We will learn how to drive industry- ... Show More
10m 50s
Mar 2023
#312 — The Trouble with AI
<p dir="ltr">Sam Harris speaks with Stuart Russell and Gary Marcus about recent developments in artificial intelligence and the long-term risks of producing artificial general intelligence (AGI). They discuss the limitations of Deep Learning, the surprising power of narrow AI, Ch ... Show More
1h 26m
Oct 2024
What Big Tech Isn’t Telling You About AI (Ep. 267)
<p>Are AI giants really building trustworthy systems? A groundbreaking transparency report by Stanford, MIT, and Princeton says no. In this episode, we expose the shocking lack of transparency in AI development and how it impacts bias, safety, and trust in the technology. We’ll b ... Show More
19m 15s
Jan 2024
Careers, Skills, and the Evolution of AI (Ep. 248)
<p>!!WARNING!!</p> <p>Due to some technical issues the volume is not always constant during the show. I sincerely apologise for any inconvenience Francesco </p> <p> </p> <p> </p> <p>In this episode, I speak with Richie Cotton, Data Evangelist at DataCamp, as he delves into the d ... Show More
32m 27s
Feb 2025
Industry Roundup #3: The Rise of Reasoning LLMs, OpenAI Operator, Project Stargate, and Gemini’s Struggle for Recognition
Welcome to DataFramed Industry Roundups! In this series of episodes, Adel & Richie sit down to discuss the latest and greatest in data & AI. In this episode, we discuss the rise of reasoning LLMs like DeepSeek R1 and the competition shaping the AI space, OpenAI’s Operator and the ... Show More
29m 46s
Oct 2023
Hinton and Bengio on Managing AI Risks
A group of scientists, academics and researchers has released a new framework on managing AI risks. NLW explores whether we're moving to more specific policy proposals. Today's Sponsors: Listen to the chart-topping podcast 'web3 with a16z crypto' wherever you get your podcasts ... Show More
17m 2s