Our new X account is live! Follow @wizwand_team for updates
WorkDL logo mark

Nonparametric semi-supervised learning of class proportions

About

The problem of developing binary classifiers from positive and unlabeled data is often encountered in machine learning. A common requirement in this setting is to approximate posterior probabilities of positive and negative classes for a previously unseen data point. This problem can be decomposed into two steps: (i) the development of accurate predictors that discriminate between positive and unlabeled data, and (ii) the accurate estimation of the prior probabilities of positive and negative examples. In this work we primarily focus on the latter subproblem. We study nonparametric class prior estimation and formulate this problem as an estimation of mixing proportions in two-component mixture models, given a sample from one of the components and another sample from the mixture itself. We show that estimation of mixing proportions is generally ill-defined and propose a canonical form to obtain identifiability while maintaining the flexibility to model any distribution. We use insights from this theory to elucidate the optimization surface of the class priors and propose an algorithm for estimating them. To address the problems of high-dimensional density estimation, we provide practical transformations to low-dimensional spaces that preserve class priors. Finally, we demonstrate the efficacy of our method on univariate and multivariate data.

Shantanu Jain, Martha White, Michael W. Trosset, Predrag Radivojac• 2016

Related benchmarks

TaskDatasetResultRank
Mixture Proportion EstimationBinarized CIFAR
Absolute Estimation Error0.09
17
Mixture Proportion EstimationCIFAR Dog vs Cat
Abs. Estimation Error0.17
12
Mixture Proportion EstimationBinarized MNIST
Absolute Estimation Error (%)9
7
Mixture Proportion EstimationMNIST 17
Abs Estimation Error7.5
7
Mixture Proportion EstimationIMDB
Absolute Error0.07
5
Showing 5 of 5 rows

Other info

Follow for update