About this episode
Aug 2
Why Multi-Agent Systems Need Shared State, Graph Semantics, and Governance
Summary In this episode Ragnor Comerford talks about OmniGraph, a lakehouse-native graph storage layer designed around the needs of agentic systems. He explores how graphs are primarily a semantic model for representing the world, rather than just a specialized engine for travers ... Show More
1h 2m
Jul 6
Building the Context Flywheel for AI Data Agents
Summary In this episode Prukalpa Sankar, co-founder of Atlan, talks about what it takes to build a “context flywheel” for AI agents in data-intensive organizations. She explained why model intelligence alone isn’t enough to make AI useful in production, and how real performance d ... Show More
1 h
Jun 18
Holding Kafka Right: Product-Friendly Streaming with TypeStream
Summary In this episode Jevin Maltais talks about the practical realities of building reliable, product-focused streaming systems with Kafka. Jevin shares lessons from roles at Zapier, Humi, and Clio, where real-time synchronization, customer data unification, and document sync a ... Show More
49m 51s
Sep 2024
Data for Dummies: A Crash Course for Non-Technical PMs (with Mo Hallaba)
<p>In today's data-driven landscape, organizations often find themselves drowning in a sea of data, yet struggling to glean actionable insights from it. Many companies are eager to label themselves as data-centric, but the reality is that not everyone is equally adept at int ... Show More
23m 23s
Nov 2021
Time Plus Data Equals Efficiency with Paul Dix, the Founder and CTO of InfluxData and the Creator of InfluxDB
<p>If the topic of databases is brought up to certain people, their eyes may gloss over. But if that happened, that would be because they just don’t know the awesome power of databases. Data can be valuable but only if it is contextualized, and time is an extremely relevant aspec ... Show More
36m 4s
Nov 2021
AI Today Podcast: AI Education Series: Managing Data for AI
This podcast episode provides a snippet of Cognilytica’s AI and ML education from our Cognilytica Education Subscription. Data is at the heart of AI. It should be no surprise then that proper data management is crucial for AI projects. This podcast is an excerpt from our Cognilyt ... Show More
24m 25s
May 2019
263: Communicating Data
In this episode of the SuperDataScience Podcast, I chat with Eoin Murray, the founder of Kyso.io, a platform where you can blog about your data science projects using tools such as Jupyter notebooks. You will learn what the platform means for data scientists and how you can use i ... Show More
1h 1m
Oct 2025
Too much data, not enough insight
Companies have access to more of their data than ever before. But using it to inform real-time decision-making remains a challenge. Chip Kleinheksel talks to Skyworks Solutions’ CIO, Satya Jayadev, and Deloitte's Som Suresh, about how leading companies are turning data into reven ... Show More
31m 37s
May 2022
How to Link Data to Business Outcomes
<p><span style="font-weight: 400;">Every business likes to claim that it is "data-driven" or at least "data-informed," but too often, that's not the way things actually work. Data is relegated to an IT function, siloed and, in some cases, boils down to simply producing more repor ... Show More
17m 32s
Jan 2015
[MINI] Data Provenance
This episode introduces a high level discussion on the topic of Data Provenance, with more MINI episodes to follow to get into specific topics. Thanks to listener Sara L who wrote in to point out the Data Skeptic Podcast has focused alot about using data to be skeptical, but not ... Show More
10m 56s
Mar 2023
51: How to Solve Any Data Analytics Error You Face...
<p>Let's face it; data analytics can sometimes be a real pain in the spreadsheet.</p>
<p>You spend hours analysing data, writing code, and crafting reports, only to have one mystery bug stop your work from progressing.</p>
<p>I'll show you my 10-step guide for solving ANY ... Show More
18m 58s
Summary
Data is useless if it isn’t being used, and you can’t use it if you don’t know where it is. Data catalogs were the first solution to this problem, but they are only helpful if you know what you are looking for. In this episode Shinji Kim discusses the challenges of data discovery and how to collect and preserve additional context about each piece of information so that you can find what you need when you don’t even know what you’re looking for yet.
Announcements
- Hello and welcome to the Data Engineering Podcast, the show about modern data management
- When you’re ready to build your next pipeline, or want to test out the projects you hear about on the show, you’ll need somewhere to deploy it, so check out our friends at Linode. With their new managed database service you can launch a production ready MySQL, Postgres, or MongoDB cluster in minutes, with automated backups, 40 Gbps connections from your application hosts, and high throughput SSDs. Go to dataengineeringpodcast.com/linode today and get a $100 credit to launch a database, create a Kubernetes cluster, or take advantage of all of their other services. And don’t forget to thank them for their continued support of this show!
- Data stacks are becoming more and more complex. This brings infinite possibilities for data pipelines to break and a host of other issues, severely deteriorating the quality of the data and causing teams to lose trust. Sifflet solves this problem by acting as an overseeing layer to the data stack – observing data and ensuring it’s reliable from ingestion all the way to consumption. Whether the data is in transit or at rest, Sifflet can detect data quality anomalies, assess business impact, identify the root cause, and alert data teams’ on their preferred channels. All thanks to 50+ quality checks, extensive column-level lineage, and 20+ connectors across the Data Stack. In addition, data discovery is made easy through Sifflet’s information-rich data catalog with a powerful search engine and real-time health statuses. Listeners of the podcast will get $2000 to use as platform credits when signing up to use Sifflet. Sifflet also offers a 2-week free trial. Find out more at dataengineeringpodcast.com/sifflet today!
- The biggest challenge with modern data systems is understanding what data you have, where it is located, and who is using it. Select Star’s data discovery platform solves that out of the box, with an automated catalog that includes lineage from where the data originated, all the way to which dashboards rely on it and who is viewing them every day. Just connect it to your database/data warehouse/data lakehouse/whatever you’re using and let them do the rest. Go to dataengineeringpodcast.com/selectstar today to double the length of your free trial and get a swag package when you convert to a paid plan.
- Data teams are increasingly under pressure to deliver. According to a recent survey by Ascend.io, 95% in fact reported being at or over capacity. With 72% of data experts reporting demands on their team going up faster than they can hire, it’s no surprise they are increasingly turning to automation. In fact, while only 3.5% report having current investments in automation, 85% of data teams plan on investing in automation in the next 12 months. 85%!!! That’s where our friends at Ascend.io come in. The Ascend Data Automation Cloud provides a unified platform for data ingestion, transformation, orchestration, and observability. Ascend users love its declarative pipelines, powerful SDK, elegant UI, and extensible plug-in architecture, as well as its support for Python, SQL, Scala, and Java. Ascend automates workloads on Snowflake, Databricks, BigQuery, and open source Spark, and can be deployed in AWS, Azure, or GCP. Go to dataengineeringpodcast.com/ascend and sign up for a free trial. If you’re a data engineering podcast listener, you get credits worth $5,000 when you become a customer.
- Your host is Tobias Macey and today I’m interviewing Shinji Kim about data discovery and what is required to build and maintain useful context for your information assets
Interview
- Introduction
- How did you get involved in the area of data management?
- Can you share your definition of "data discovery" and the technical/social/process components that are required to make it viable?
- What are the differences between "data discovery" and the capabilities of a "data catalog" and how do they overlap?
- discovery of assets outside the bounds of the warehouse
- capturing and codifying tribal knowledge
- creating a useful structure/framework for capturing data context and operationalizing it
- What are the most interesting, innovative, or unexpected ways that you have seen data discovery implemented?
- What are the most interesting, unexpected, or challenging lessons that you have learned while working on data discovery at SelectStar?
- When might a data discovery effort be more work than is required?
- What do you have planned for the future of SelectStar?
Contact Info
Parting Question
- From your perspective, what is the biggest gap in the tooling or technology for data management today?
Closing Announcements
- Thank you for listening! Don’t forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The Machine Learning Podcast helps you go from idea to production with machine learning.
- Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
- If you’ve learned something or tried out a project from the show then tell us about it! Email hosts@dataengineeringpodcast.com) with your story.
- To help other people find the show please leave a review on Apple Podcasts and tell your friends and co-workers
Links
The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
Support Data Engineering Podcast