August 12, 2016

Towards Optimal Cardinality Estimation of Unions and Intersections with Sketches

ACM Conference on Knowledge Discovery and Data Mining

By: Daniel Ting

Abstract

Estimating the cardinality of unions and intersections of sets is a problem of interest in OLAP. Large data applications often require the use of approximate methods based on small sketches of the data. We give new estimators for the cardinality of unions and intersection and show they approximate an optimal estimation procedure. These estimators enable the improved accuracy of the streaming MinCount sketch to be exploited in distributed settings. Both theoretical and empirical results demonstrate substantial improvements over existing methods.