About this episode
Aug 27
Specialized AI for Data Engineers: Inside Astronomer’s Otto
Summary In this episode Yetunde Dada discusses Otto, Astronomer’s AI agent for Airflow, and the broader challenge of making agentic tooling actually useful for data engineers. She explored why generic coding assistants often fall short in data workflows, how Otto adds the missing ... Show More
45m 30s
Aug 2
Why Multi-Agent Systems Need Shared State, Graph Semantics, and Governance
Summary In this episode Ragnor Comerford talks about OmniGraph, a lakehouse-native graph storage layer designed around the needs of agentic systems. He explores how graphs are primarily a semantic model for representing the world, rather than just a specialized engine for travers ... Show More
1h 2m
Jul 6
Building the Context Flywheel for AI Data Agents
Summary In this episode Prukalpa Sankar, co-founder of Atlan, talks about what it takes to build a “context flywheel” for AI agents in data-intensive organizations. She explained why model intelligence alone isn’t enough to make AI useful in production, and how real performance d ... Show More
1 h
Aug 2024
Snowflake's Baris Gultekin on Unlocking the Value of Data With Large Language Models - Ep. 231
Snowflake is using AI to help enterprises transform data into insights and applications. In this episode of NVIDIA’s AI Podcast, host Noah Kravitz and Baris Gultekin, head of AI at Snowflake, discuss how the company’s AI Data Cloud platform enables customers to access and manage ... Show More
32m 10s
Dec 2022
Update Your Model's View Of The World In Real Time With Streaming Machine Learning Using River
Preamble
This is a cross-over episode from our new show The Machine Learning Podcast, the show about going from idea to production with machine learning.
Summary
The majority of machine learning projects that you read about or work on are built around batch processes. The model i ... Show More
1h 16m
May 2020
Episode 102 - Complex flavors - complex systems
We’ve all been there, a project starts off simple, but quickly becomes more complex. In this episode, we are joined by Sarah Drasner to talk with us about how to deal with complex codebases and systems. Guests: Sarah Drasner - @sarah_edo Panelists: Ryan Burgess - @burgessdryan Je ... Show More
59m 10s
May 2024
D2C242: Data Engineering and its Streams, Rivers, and Lakes
Keith Gregory teaches us about data engineering in a way DevOps folks (and hydrologists) can understand. He explains that the role of a data engineer is to create pipelines to transport data from metaphorical rivers and make it usable for data analysts. Keith walks us through the ... Show More
47m 37s
May 2022
From Kubernetes to PaaS - now what? (Ship It! #51)
Today we talk to Mark Ericksen about all the things that we could be doing on the new platform - this is a follow-up to episode 50.
Mark specialises in Elixir, he hosts the Thinking Elixir podcast, and he also helps make Fly.io the best place to run Phoenix apps, such as chang ... Show More
58m 10s
Jan 2024
Fivetran COO Unravels Enterprise Data Movement
Automating the collection of dispersed, divergent and disjointed pools of enterprise data on to a single repository to drive analytics and build applications remains complex. In this edition of Bloomberg Intelligence’s Tech Disruptors podcast, Fivetran cofounder and COO Taylor Br ... Show More
40m 15s
Jan 2024
SingleStore CEO on High-Speed Database Currents
Enterprise data architecture is highly complex, databases deeply fragmented and demand for high-speed information flows continues to grow. In this edition of the Tech Disruptors podcast, SingleStore CEO Raj Verma joins Sunil Rajgopal, Bloomberg Intelligence senior software analys ... Show More
47m 26s
Dec 2024
#491: DuckDB and Python: Ducks and Snakes living together
See the full show notes for this episode on the website at <a href="https://talkpython.fm/491">talkpython.fm/491</a>
1h 2m
Nov 2023
#153 Ed Anuff: Unpacking AI's Role in Data Management
<p>This episode is sponsored by Celonis ,the global leader in process mining. AI has landed and enterprises are adapting. To give customers slick experiences and teams the technology to deliver. The road is long, but you're closer than you think. Your business processes run throu ... Show More
59m 5s
Summary
Building applications on top of unbounded event streams is a complex endeavor, requiring careful integration of multiple disparate systems that were engineered in isolation. The ksqlDB project was created to address this state of affairs by building a unified layer on top of the Kafka ecosystem for stream processing. Developers can work with the SQL constructs that they are familiar with while automatically getting the durability and reliability that Kafka offers. In this episode Michael Drogalis, product manager for ksqlDB at Confluent, explains how the system is implemented, how you can use it for building your own stream processing applications, and how it fits into the lifecycle of your data infrastructure. If you have been struggling with building services on low level streaming interfaces then give this episode a listen and try it out for yourself.
Announcements
- Hello and welcome to the Data Engineering Podcast, the show about modern data management
- When you’re ready to build your next pipeline, or want to test out the projects you hear about on the show, you’ll need somewhere to deploy it, so check out our friends at Linode. With 200Gbit private networking, scalable shared block storage, a 40Gbit public network, fast object storage, and a brand new managed Kubernetes platform, you’ve got everything you need to run a fast, reliable, and bullet-proof data platform. And for your machine learning workloads, they’ve got dedicated CPU and GPU instances. Go to dataengineeringpodcast.com/linode today to get a $20 credit and launch a new server in under a minute. And don’t forget to thank them for their continued support of this show!
- Are you spending too much time maintaining your data pipeline? Snowplow empowers your business with a real-time event data pipeline running in your own cloud account without the hassle of maintenance. Snowplow takes care of everything from installing your pipeline in a couple of hours to upgrading and autoscaling so you can focus on your exciting data projects. Your team will get the most complete, accurate and ready-to-use behavioral web and mobile data, delivered into your data warehouse, data lake and real-time streams. Go to dataengineeringpodcast.com/snowplow today to find out why more than 600,000 websites run Snowplow. Set up a demo and mention you’re a listener for a special offer!
- You listen to this show to learn and stay up to date with what’s happening in databases, streaming platforms, big data, and everything else you need to know about modern data management. For even more opportunities to meet, listen, and learn from your peers you don’t want to miss out on this year’s conference season. We have partnered with organizations such as O’Reilly Media, Corinium Global Intelligence, ODSC, and Data Council. Upcoming events include the Software Architecture Conference in NYC, Strata Data in San Jose, and PyCon US in Pittsburgh. Go to dataengineeringpodcast.com/conferences to learn more about these and other events, and take advantage of our partner discounts to save money when you register today.
- Your host is Tobias Macey and today I’m interviewing Michael Drogalis about ksqlDB, the open source streaming database layer for Kafka
Interview
- Introduction
- How did you get involved in the area of data management?
- Can you start by describing what ksqlDB is?
- What are some of the use cases that it is designed for?
- How do the capabilities and design of ksqlDB compare to other solutions for querying streaming data with SQL such as Pulsar SQL, PipelineDB, or Materialize?
- What was the motivation for building a unified project for providing a database interface on the data stored in Kafka?
- How is ksqlDB architected?
- If you were to rebuild the entire platform and its components from scratch today, what would you do differently?
- What is the workflow for an analyst or engineer to design and build an application on top of ksqlDB?
- What dialect of SQL is supported?
- What kinds of extensions or built in functions have been added to aid in the creation of streaming queries?
- How are table schemas defined and enforced?
- How do you handle schema migrations on active streams?
- Typically a database is considered a long term storage location for data, whereas Kafka is a streaming layer with a bounded amount of durable storage. What is a typical lifecycle of information in ksqlDB?
- Can you talk through an example architecture that might incorporate ksqlDB including the source systems, applications that might interact with the data in transit, and any destinations sytems for long term persistence?
- What are some of the less obvious features of ksqlDB or capabilities that you think should be more widely publicized?
- What are some of the edge cases or potential pitfalls that users should be aware of as they are designing their streaming applications?
- What is involved in deploying and maintaining an installation of ksqlDB?
- What are some of the operational characteristics of the system that should be considered while planning an installation such as scaling factors, high availability, or potential bottlenecks in the architecture?
- When is ksqlDB the wrong choice?
- What are some of the most interesting/unexpected/innovative projects that you have seen built with ksqlDB?
- What are some of the most interesting/unexpected/challenging lessons that you have learned while working on ksqlDB?
- What is in store for the future of the project?
Contact Info
Parting Question
- From your perspective, what is the biggest gap in the tooling or technology for data management today?
Links
The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
Support Data Engineering Podcast