logo
episode-header-image
Jan 2015
10m 56s

[MINI] Data Provenance

Kyle Polich
About this episode

This episode introduces a high level discussion on the topic of Data Provenance, with more MINI episodes to follow to get into specific topics. Thanks to listener Sara L who wrote in to point out the Data Skeptic Podcast has focused alot about using data to be skeptical, but not necessarily being skeptical of data.

Data Provenance is the concept of knowing the full origin of your dataset. Where did it come from? Who collected it? How as it collected? Does it combine independent sources or one singular source? What are the error bounds on the way it was measured? These are just some of the questions one should ask to understand their data. After all, if the antecedent of an argument is built on dubious grounds, the consequent of the argument is equally dubious.

For a more technical discussion than what we get into in this mini epiosode, I recommend A Survey of Data Provenance Techniques by authors Simmhan, Plale, and Gannon.

Up next
Sep 25
The Lived Informatics Model
The data we collect about ourselves can tell us a lot—but only if the technology collecting it actually fits into our lives. Daniel Epstein explores personal informatics, from fitness trackers and food journals to baby tracking and AI, and explains why abandoning a tracking tool ... Show More
34m 13s
Sep 9
Recommender Systems Today and Tomorrow
In the final episode of our Recommender Systems season, we explore the growing questions of trust, manipulation, privacy, fairness, sustainability, and user control. From fake reviews and shilling attacks to explainable recommendations and user-selected algorithms, we look at wha ... Show More
22m 46s
Sep 1
Recommender Systems Optimization Goals
In part two of the Data Skeptic Recommender Systems season finale, Kyle asks a deceptively difficult question: what should recommender systems actually optimize for? Drawing on conversations from across the season, the episode explores engagement, filter bubbles, popularity bias, ... Show More
31m 15s
Recommended Episodes
Jun 2019
Data Trusts and Citation Trends
<p>In episode eleven of season five, we dig in to just what a data trust actually is, take a look at <a href="http://maithraraghu.com/blog/2019/Citation_Statistics_of_Machine_Learning_Papers/" target="_blank">citation trends </a>and other places <a href="http://proceedings.mlr.pr ... Show More
54m 15s
Aug 2022
Collecting And Retaining Contextual Metadata For Powerful And Effective Data Discovery
<div class="wp-block-jetpack-markdown"><h2>Summary</h2> <p>Data is useless if it isn&#8217;t being used, and you can&#8217;t use it if you don&#8217;t know where it is. Data catalogs were the first solution to this problem, but they are only helpful if you know what you are look ... Show More
53m 24s
Apr 2017
041: An Inspiring Journey from a Totally Different Background to Data Science
In this episode of the SuperDataScience Podcast, I chat with Aspiring Data Scientist Nicholas Cepeda. You will be able to discover the many different pathways in Data Science, learn R from a SQL background and hear about SAS Enterprise Miner. If you enjoyed this episode, check o ... Show More
49m 18s
Nov 2021
Data Quality Starts At The Source
<div class="wp-block-jetpack-markdown"><h2>Summary</h2> <p>The most important gauge of success for a data platform is the level of trust in the accuracy of the information that it provides. In order to build and maintain that trust it is necessary to invest in defining, monitori ... Show More
58m 55s
Aug 2022
An Exploration Of The Expectations, Ecosystem, and Realities Of Real-Time Data Applications
<div class="wp-block-jetpack-markdown"><h2>Summary</h2> <p>Data has permeated every aspect of our lives and the products that we interact with. As a result, end users and customers have come to expect interactions and updates with services and analytics to be fast and up to date ... Show More
1h 6m
Nov 2017
103: Why this Is the Golden Age of Data Science & How to Get In
In this episode of the SuperDataScience Podcast, I chat with Data Science Freelance Consultant, Emanuele Carbone. You will hear about creating a massive value with Tableau and Excel, learn how to get established as a data science consultant and learn how Emanuele went from novice ... Show More
52m 10s
Aug 2024
Only as good as the data
You might have heard that “AI is only as good as the data.” What does that mean and what data are we talking about? Chris and Daniel dig into that topic in the episode exploring the categories of data that you might encounter working in AI (for training, testing, fine-tuning, ben ... Show More
45m 41s
May 2019
263: Communicating Data
In this episode of the SuperDataScience Podcast, I chat with Eoin Murray, the founder of Kyso.io, a platform where you can blog about your data science projects using tools such as Jupyter notebooks. You will learn what the platform means for data scientists and how you can use i ... Show More
1h 1m