logo
episode-header-image
Jan 2015
10m 56s

[MINI] Data Provenance

Kyle Polich
About this episode

This episode introduces a high level discussion on the topic of Data Provenance, with more MINI episodes to follow to get into specific topics. Thanks to listener Sara L who wrote in to point out the Data Skeptic Podcast has focused alot about using data to be skeptical, but not necessarily being skeptical of data.

Data Provenance is the concept of knowing the full origin of your dataset. Where did it come from? Who collected it? How as it collected? Does it combine independent sources or one singular source? What are the error bounds on the way it was measured? These are just some of the questions one should ask to understand their data. After all, if the antecedent of an argument is built on dubious grounds, the consequent of the argument is equally dubious.

For a more technical discussion than what we get into in this mini epiosode, I recommend A Survey of Data Provenance Techniques by authors Simmhan, Plale, and Gannon.

Up next
Jul 27
Social Choice for Fair Recommendations
Recommender systems influence nearly every aspect of our digital lives—but what does it mean for those systems to be fair? Robin Burke joins Data Skeptic to discuss the history of recommender systems, the limitations of optimizing purely for accuracy, and how ideas from social ch ... Show More
42m 56s
Jul 2
News Recommendations
News recommendation algorithms influence far more than what stories we click—they can shape our understanding of the world. In this episode, Kyle Polich speaks with Andreea Iana about responsible AI, filter bubbles, multilingual news recommendation, and her open-source NewsRecLib ... Show More
46m 6s
Jun 23
Give Users the Wheel
What if you could simply tell a recommendation system what you want instead of relying on likes, dislikes, and watch history? Kyle Polich talks with Fuyuan Lyu about the DPR framework, which combines large language models and traditional recommender systems to give users direct c ... Show More
35m 28s
Recommended Episodes
Jun 2019
Data Trusts and Citation Trends
<p>In episode eleven of season five, we dig in to just what a data trust actually is, take a look at <a href="http://maithraraghu.com/blog/2019/Citation_Statistics_of_Machine_Learning_Papers/" target="_blank">citation trends </a>and other places <a href="http://proceedings.mlr.pr ... Show More
54m 15s
Aug 2022
Collecting And Retaining Contextual Metadata For Powerful And Effective Data Discovery
<div class="wp-block-jetpack-markdown"><h2>Summary</h2> <p>Data is useless if it isn&#8217;t being used, and you can&#8217;t use it if you don&#8217;t know where it is. Data catalogs were the first solution to this problem, but they are only helpful if you know what you are look ... Show More
53m 24s
Apr 2017
041: An Inspiring Journey from a Totally Different Background to Data Science
In this episode of the SuperDataScience Podcast, I chat with Aspiring Data Scientist Nicholas Cepeda. You will be able to discover the many different pathways in Data Science, learn R from a SQL background and hear about SAS Enterprise Miner. If you enjoyed this episode, check o ... Show More
49m 18s
Nov 2021
Data Quality Starts At The Source
<div class="wp-block-jetpack-markdown"><h2>Summary</h2> <p>The most important gauge of success for a data platform is the level of trust in the accuracy of the information that it provides. In order to build and maintain that trust it is necessary to invest in defining, monitori ... Show More
58m 55s
Aug 2022
An Exploration Of The Expectations, Ecosystem, and Realities Of Real-Time Data Applications
<div class="wp-block-jetpack-markdown"><h2>Summary</h2> <p>Data has permeated every aspect of our lives and the products that we interact with. As a result, end users and customers have come to expect interactions and updates with services and analytics to be fast and up to date ... Show More
1h 6m
Nov 2017
103: Why this Is the Golden Age of Data Science & How to Get In
In this episode of the SuperDataScience Podcast, I chat with Data Science Freelance Consultant, Emanuele Carbone. You will hear about creating a massive value with Tableau and Excel, learn how to get established as a data science consultant and learn how Emanuele went from novice ... Show More
52m 10s
Aug 2024
Only as good as the data
You might have heard that “AI is only as good as the data.” What does that mean and what data are we talking about? Chris and Daniel dig into that topic in the episode exploring the categories of data that you might encounter working in AI (for training, testing, fine-tuning, ben ... Show More
45m 41s
May 2019
263: Communicating Data
In this episode of the SuperDataScience Podcast, I chat with Eoin Murray, the founder of Kyso.io, a platform where you can blog about your data science projects using tools such as Jupyter notebooks. You will learn what the platform means for data scientists and how you can use i ... Show More
1h 1m