Feature Selection using Stochastic Gates

About

Feature selection problems have been extensively studied for linear estimation, for instance, Lasso, but less emphasis has been placed on feature selection for non-linear functions. In this study, we propose a method for feature selection in high-dimensional non-linear function estimation problems. The new procedure is based on minimizing the $\ell_0$ norm of the vector of indicator variables that represent if a feature is selected or not. Our approach relies on the continuous relaxation of Bernoulli distributions, which allows our model to learn the parameters of the approximate Bernoulli distributions via gradient descent. This general framework simultaneously minimizes a loss function while selecting relevant features. Furthermore, we provide an information-theoretic justification of incorporating Bernoulli distribution into our approach and demonstrate the potential of the approach on synthetic and real-life applications.

Yutaro Yamada, Ofir Lindenbaum, Sahand Negahban, Yuval Kluger• 2018

Related benchmarks

Task	Dataset	Result
Classification	Lung	ACC93.3	96
Classification	Adult	Accuracy56.52	86
Classification	TOX_171	Accuracy87.95	78
Classification	GLI_85	Accuracy82.48	78
Classification	Colon	Accuracy79.55	78
Classification	ALLAML	Accuracy86.08	72
Classification	SMK_CAN_187	Accuracy57.25	72
Classification	HE	Accuracy38.9	66
Classification	HDLSS Datasets Summary	Average Rank9.17	66
Classification	GE	Accuracy60.9	65

Showing 10 of 85 rows

...

Other info

Follow for update

@wizwand_team Discord