About this episode
Jul 6
Building the Context Flywheel for AI Data Agents
Summary In this episode Prukalpa Sankar, co-founder of Atlan, talks about what it takes to build a “context flywheel” for AI agents in data-intensive organizations. She explained why model intelligence alone isn’t enough to make AI useful in production, and how real performance d ... Show More
1 h
Jun 18
Holding Kafka Right: Product-Friendly Streaming with TypeStream
Summary In this episode Jevin Maltais talks about the practical realities of building reliable, product-focused streaming systems with Kafka. Jevin shares lessons from roles at Zapier, Humi, and Clio, where real-time synchronization, customer data unification, and document sync a ... Show More
49m 51s
Jun 8
Text to Data Products: Kaarvi’s End-to-End AI for Ingestion, Quality, and Dashboards
Summary In this episode Shravan Gunda, founder and CEO of Kaarvi AI, talks about building an AI-native, agent-driven data platform designed to eliminate the janitorial work that consumes most data teams. He explores Kaarvi’s multi-agent architecture that runs queries across seven ... Show More
52m 52s
Aug 2020
Building the world's most popular data science platform
Everyone working in data science and AI knows about Anaconda and has probably “conda” installed something. But how did Anaconda get started and what are they working on now? Peter Wang, CEO of Anaconda and creator of PyData and popular packages like Bokeh and DataShader, joins us ... Show More
59m 13s
Aug 2024
Snowflake's Baris Gultekin on Unlocking the Value of Data With Large Language Models - Ep. 231
Snowflake is using AI to help enterprises transform data into insights and applications. In this episode of NVIDIA’s AI Podcast, host Noah Kravitz and Baris Gultekin, head of AI at Snowflake, discuss how the company’s AI Data Cloud platform enables customers to access and manage ... Show More
32m 10s
Oct 2023
#628: Data on EKS
Organizations use their data to make better decisions and build innovative experiences for their customers. With the exponential growth in data, and the rapid pace of innovation in machine learning (ML), there is a growing need to build modern data applications that are agile and ... Show More
20m 56s
Jun 2024
780: Cloud Storage: Bandwidth, Storage and BIG ZIPS
Today, Scott and Wes dive into cloud storage solutions—why you might need them, how they use them, and what you need to know about the big players, fees, and more.
Show Notes 00:00 Welcome to Syntax!
01:14 Brought to you by Sentry.io.
02:05 Why you might need a cloud stora ... Show More
29m 2s
May 2025
Strategies for Decentralized Cloud Data Storage
Murphy John discusses strategies for secure, scalable, and efficient cloud data storage. Murphy is the Chief Operating Officer of StorX Network, a decentralized cloud storage platform with 117,000 and a network of more than 2,500 storage nodes. Host, Kevin Craine Want to be a gue ... Show More
24m 52s
Nov 2021
AI Today Podcast: AI Education Series: Managing Data for AI
This podcast episode provides a snippet of Cognilytica’s AI and ML education from our Cognilytica Education Subscription. Data is at the heart of AI. It should be no surprise then that proper data management is crucial for AI projects. This podcast is an excerpt from our Cognilyt ... Show More
24m 25s
Feb 2024
Virtual Gigantic Storage – The Business of Charging
<p><strong>In today’s digital age, the demand for virtual storage has skyrocketed. With the accumulation of vast amounts of personal and professional data, individuals and organisations are constantly seeking reliable and secure solutions to store their files.</strong></p>
<p><st ... Show More
6m 56s
Aug 2022
Introduction and Kickoff of In the Clouds with AWS
This episode launches our podcast In the Clouds with AWS. Our future sessions we will cover topics, blogs, and support articles about AWS"Amazon Web Services provides a highly reliable, scalable, low-cost infrastructure platform in the cloud that powers hundreds of thousands of b ... Show More
9m 25s
Jan 2022
Making the last database you’ll ever need (Founders Talk #85)
This week Adam is joined by Sam Lambert, CEO of PlanetScale. Now that PlanetScale is in general availability, Adam had to get Sam on the show to talk about the behind the scenes of building this database platform, how this is the last database you’ll ever need and what that means ... Show More
1h 25m
Summary
Object storage is quickly becoming the unifying layer for data intensive applications and analytics. Modern, cloud oriented data warehouses and data lakes both rely on the durability and ease of use that it provides. S3 from Amazon has quickly become the de-facto API for interacting with this service, so the team at MinIO have built a production grade, easy to manage storage engine that replicates that interface. In this episode Anand Babu Periasamy shares the origin story for the MinIO platform, the myriad use cases that it supports, and the challenges that they have faced in replicating the functionality of S3. He also explains the technical implementation, innovative design, and broad vision for the project.
Announcements
- Hello and welcome to the Data Engineering Podcast, the show about modern data management
- When you’re ready to build your next pipeline, or want to test out the projects you hear about on the show, you’ll need somewhere to deploy it, so check out our friends at Linode. With 200Gbit private networking, scalable shared block storage, and a 40Gbit public network, you’ve got everything you need to run a fast, reliable, and bullet-proof data platform. If you need global distribution, they’ve got that covered too with world-wide datacenters including new ones in Toronto and Mumbai. And for your machine learning workloads, they just announced dedicated CPU instances. Go to dataengineeringpodcast.com/linode today to get a $20 credit and launch a new server in under a minute. And don’t forget to thank them for their continued support of this show!
- You listen to this show to learn and stay up to date with what’s happening in databases, streaming platforms, big data, and everything else you need to know about modern data management.For even more opportunities to meet, listen, and learn from your peers you don’t want to miss out on this year’s conference season. We have partnered with organizations such as O’Reilly Media, Dataversity, Corinium Global Intelligence, and Data Council. Upcoming events include the O’Reilly AI conference, the Strata Data conference, the combined events of the Data Architecture Summit and Graphorum, and Data Council in Barcelona. Go to dataengineeringpodcast.com/conferences to learn more about these and other events, and take advantage of our partner discounts to save money when you register today.
- Your host is Tobias Macey and today I’m interviewing Anand Babu Periasamy about MinIO, the neutral, open source, enterprise grade object storage system.
Interview
- Introduction
- How did you get involved in the area of data management?
- Can you explain what MinIO is and its origin story?
- What are some of the main use cases that MinIO enables?
- How does MinIO compare to other object storage options and what benefits does it provide over other open source platforms?
- Your marketing focuses on the utility of MinIO for ML and AI workloads. What benefits does object storage provide as compared to distributed file systems? (e.g. HDFS, GlusterFS, Ceph)
- What are some of the challenges that you face in terms of maintaining compatibility with the S3 interface?
- What are the constraints and opportunities that are provided by adhering to that API?
- Can you describe how MinIO is implemented and the overall system design?
- How has that design evolved since you first began working on it?
- What assumptions did you have at the outset and how have they been challenged or updated?
- What are the axes for scaling that MinIO provides and how does it handle clustering?
- Where does it fall on the axes of availability and consistency in the CAP theorem?
- One of the useful features that you provide is efficient erasure coding, as well as protection against data corruption. How much overhead do those capabilties incur, in terms of computational efficiency and, in a clustered scenario, storage volume?
- For someone who is interested in running MinIO, what is involved in deploying and maintaining an installation of it?
- What are the cases where it makes sense to use MinIO in place of a cloud-native object store such as S3 or Google Cloud Storage?
- How do you approach project governance and sustainability?
- What are some of the most interesting/innovative/unexpected ways that you have seen MinIO used?
- What do you have planned for the future of MinIO?
Contact Info
Parting Question
- From your perspective, what is the biggest gap in the tooling or technology for data management today?
Closing Announcements
- Thank you for listening! Don’t forget to check out our other show, Podcast.__init__ to learn about the Python language, its community, and the innovative ways it is being used.
- Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
- If you’ve learned something or tried out a project from the show then tell us about it! Email hosts@dataengineeringpodcast.com) with your story.
- To help other people find the show please leave a review on iTunes and tell your friends and co-workers
- Join the community in the new Zulip chat workspace at dataengineeringpodcast.com/chat
Links
The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
Support Data Engineering Podcast