Time Travel and Provenance for Machine Learning Pipelines

Authors

Alexandru A. Ormenisan, Moritz Meister, Fabio Buso, Robin Andersson, Seif Haridi, Jim Dowling

Abstract

Machine learning pipelines have become the defacto paradigm for productionizing machine learning applications as they clearly abstract the processing steps involved in trans-forming raw data into engineered features that are then used to train models. In this paper, we use a bottom-up method for capturing provenance information regarding the processing steps and artifacts produced in ML pipelines. Our approach is based on replacing traditional intrusive hooks in application code (to capture ML pipeline events) with standardized change-data-capture support in the systems involved in ML pipelines: the distributed file system, feature store, resource manager, and applications themselves. In particular, we lever- age data versioning and time-travel capabilities in our feature store to show how provenance can enable model reproducibil- ity and debugging.

Download Paper

Hopsworks RonDB Maggy HopsFS

Research White Papers Newsroom Use Cases

Events Career Contact Us About Us

Time Travel and Provenance for Machine Learning Pipelines

Authors

Abstract

Products

Resources

Company