About this episode
Aug 2
Why Multi-Agent Systems Need Shared State, Graph Semantics, and Governance
Summary In this episode Ragnor Comerford talks about OmniGraph, a lakehouse-native graph storage layer designed around the needs of agentic systems. He explores how graphs are primarily a semantic model for representing the world, rather than just a specialized engine for travers ... Show More
1h 2m
Jul 6
Building the Context Flywheel for AI Data Agents
Summary In this episode Prukalpa Sankar, co-founder of Atlan, talks about what it takes to build a “context flywheel” for AI agents in data-intensive organizations. She explained why model intelligence alone isn’t enough to make AI useful in production, and how real performance d ... Show More
1 h
Jun 18
Holding Kafka Right: Product-Friendly Streaming with TypeStream
Summary In this episode Jevin Maltais talks about the practical realities of building reliable, product-focused streaming systems with Kafka. Jevin shares lessons from roles at Zapier, Humi, and Clio, where real-time synchronization, customer data unification, and document sync a ... Show More
49m 51s
Oct 2024
825: Data Contracts: The Key to Data Quality, with Chad Sanderson
Data contracts are redefining data quality and governance, and Chad Sanderson, CEO of Gable.ai, joins host Jon Krohn to explain how they can transform your data strategy. He breaks down what data contracts are, how they shift data quality checks closer to production, and why they ... Show More
1h 2m
Jun 2023
#593: AWS Glue Data Quality
Hundreds of thousands of customers build data lakes everyday and these data lakes can quickly become data swamps if they dont pay attention to the data quality. Setting up data quality can be a time-consuming, tedious process. Shiv Narayanan, Product Manager for AWS Glue chats wi ... Show More
19m 7s
Feb 2025
How Can GenAI Make Analytics More Accessible to Product Teams? (with Mario Ciabarra)
<p>Whether you prefer the term data-driven, or data-informed, or data-dazzled, it doesn't matter—today's tech cannot survive without high quality data sets AND the tools to use them effectively. But we also can't afford to think about data as the responsibility of ... Show More
27m 46s
Mar 2024
How Data and Analytics Can Bring AI and Humans Together
As organizations implement AI in decision making and how work gets done, trustworthy data is essential to adding value every step of the way. Inadequate data governance or unclear AI ambitions leaves enterprises at risk of falling behind. In this keynote address from Gartner Data ... Show More
26m 4s
Jan 2015
[MINI] Data Provenance
This episode introduces a high level discussion on the topic of Data Provenance, with more MINI episodes to follow to get into specific topics. Thanks to listener Sara L who wrote in to point out the Data Skeptic Podcast has focused alot about using data to be skeptical, but not ... Show More
10m 56s
Oct 2022
Data for All
People are starting to wake up to the fact that they have control and ownership over their data, and governments are moving quickly to legislate these rights. John K. Thompson has written a new book on the topic that is a must read! We talk about the new book in this episode alon ... Show More
49m 27s
Aug 2024
Only as good as the data
You might have heard that “AI is only as good as the data.” What does that mean and what data are we talking about? Chris and Daniel dig into that topic in the episode exploring the categories of data that you might encounter working in AI (for training, testing, fine-tuning, ben ... Show More
45m 41s
Apr 2024
BDTP. Data-powered Growth with Arpit Choudhury
Today we have another episode of Better Done Than Perfect. Listen in as we talk to Arpit Choudhury, founder of Databeats. You'll learn what makes a robust data foundation, how teams can work together to collect data, why less data is better, and more.Please head over to the episo ... Show More
40m 26s
Sep 2016
Best Practices for Data Quality in Case Management
On this episode of the META Podcast, we discuss data quality in case management programs for resettled refugees. Hear common challenges and tips for improvement from Abigail Clarke-Sayer and Prospero Herrera of the IRC as well as Jennifer Malloy of the Washington State Office of ... Show More
21m 9s
May 2022
How to Use Data in the PM Role by LinkedIn Product Leader
<p>LinkedIn Product Leader <a href='https://productschool.com/product-leaders/Ming-Chang?utm_campaign=202205-mkt_podcast_episode-ming_chang&utm_medium=social&utm_source=buzzsprout&utm_term=b2c_current_pm'>Ming Chang</a> believes that one of the best PM superpowers is ... Show More
32m 2s
Summary
The most important gauge of success for a data platform is the level of trust in the accuracy of the information that it provides. In order to build and maintain that trust it is necessary to invest in defining, monitoring, and enforcing data quality metrics. In this episode Michael Harper advocates for proactive data quality and starting with the source, rather than being reactive and having to work backwards from when a problem is found.
Announcements
- Hello and welcome to the Data Engineering Podcast, the show about modern data management
- When you’re ready to build your next pipeline, or want to test out the projects you hear about on the show, you’ll need somewhere to deploy it, so check out our friends at Linode. With their managed Kubernetes platform it’s now even easier to deploy and scale your workflows, or try out the latest Helm charts from tools like Pulsar and Pachyderm. With simple pricing, fast networking, object storage, and worldwide data centers, you’ve got everything you need to run a bulletproof data platform. Go to dataengineeringpodcast.com/linode today and get a $100 credit to try out a Kubernetes cluster of your own. And don’t forget to thank them for their continued support of this show!
- Atlan is a collaborative workspace for data-driven teams, like Github for engineering or Figma for design teams. By acting as a virtual hub for data assets ranging from tables and dashboards to SQL snippets & code, Atlan enables teams to create a single source of truth for all their data assets, and collaborate across the modern data stack through deep integrations with tools like Snowflake, Slack, Looker and more. Go to dataengineeringpodcast.com/atlan today and sign up for a free trial. If you’re a data engineering podcast listener, you get credits worth $3000 on an annual subscription
- Modern Data teams are dealing with a lot of complexity in their data pipelines and analytical code. Monitoring data quality, tracing incidents, and testing changes can be daunting and often takes hours to days. Datafold helps Data teams gain visibility and confidence in the quality of their analytical data through data profiling, column-level lineage and intelligent anomaly detection. Datafold also helps automate regression testing of ETL code with its Data Diff feature that instantly shows how a change in ETL or BI code affects the produced data, both on a statistical level and down to individual rows and values. Datafold integrates with all major data warehouses as well as frameworks such as Airflow & dbt and seamlessly plugs into CI workflows. Go to dataengineeringpodcast.com/datafold today to start a 30-day trial of Datafold.
- Your host is Tobias Macey and today I’m interviewing Michael Harper about definitions of data quality and where to define and enforce it in the data platform
Interview
- Introduction
- How did you get involved in the area of data management?
- What is your definition for the term "data quality" and what are the implied goals that it embodies?
- What are some ways that different stakeholders and participants in the data lifecycle might disagree about the definitions and manifestations of data quality?
- The market for "data quality tools" has been growing and gaining attention recently. How would you categorize the different approaches taken by open source and commercial options in the ecosystem?
- What are the tradeoffs that you see in each approach? (e.g. data warehouse as a chokepoint vs quality checks on extract)
- What are the difficulties that engineers and stakeholders encounter when identifying and defining information that is necessary to identify issues in their workflows?
- Can you describe some examples of adding data quality checks to the beginning stages of a data workflow and the kinds of issues that can be identified?
- What are some ways that quality and observability metrics can be aggregated across multiple pipeline stages to identify more complex issues?
- In application observability the metrics across multiple processes are often associated with a given service. What is the equivalent concept in data platform observabiliity?
- In your work at Databand what are some of the ways that your ideas and assumptions around data quality have been challenged or changed?
- What are the most interesting, innovative, or unexpected ways that you have seen Databand used?
- What are the most interesting, unexpected, or challenging lessons that you have learned while working at Databand?
- When is Databand the wrong choice?
- What do you have planned for the future of Databand?
Contact Info
Parting Question
- From your perspective, what is the biggest gap in the tooling or technology for data management today?
Links
The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
Support Data Engineering Podcast