Standard Gaussian Process is All You Need for High-Dimensional Bayesian Optimization

About

A long-standing belief holds that Bayesian Optimization (BO) with standard Gaussian processes (GP) -- referred to as standard BO -- underperforms in high-dimensional optimization problems. While this belief seems plausible, it lacks both robust empirical evidence and theoretical justification. To address this gap, we present a systematic investigation. First, through a comprehensive evaluation across twelve benchmarks, we found that while the popular Square Exponential (SE) kernel often leads to poor performance, using Mat\'ern kernels enables standard BO to consistently achieve top-tier results, frequently surpassing methods specifically designed for high-dimensional optimization. Second, our theoretical analysis reveals that the SE kernel's failure primarily stems from improper initialization of the length-scale parameters, which are commonly used in practice but can cause gradient vanishing in training. We provide a probabilistic bound to characterize this issue, showing that Mat\'ern kernels are less susceptible and can robustly handle much higher dimensions. Third, we propose a simple robust initialization strategy that dramatically improves the performance of the SE kernel, bringing it close to state-of-the-art methods, without requiring additional priors or regularization. We prove another probabilistic bound that demonstrates how the gradient vanishing issue can be effectively mitigated with our method. Our findings advocate for a re-evaluation of standard BO's potential in high-dimensional settings.

Zhitong Xu, Haitao Wang, Jeff M Phillips, Shandian Zhe• 2024

Related benchmarks

Task	Dataset	Result
High-dimensional optimization	MSLR	Convergence Value-8.7492	21
High-dimensional optimization	LIMO	Convergence Value-4.6146	20
High-dimensional optimization	Lasso-Hard	Convergence Value49.3089	20
Function Optimization	Michalewicz D=1000	Convergence Value-7.712	19
Function Optimization	Rosenbrock D=1000	Convergence Value6.44e+5	19
Function Optimization	Sphere D=1000	Final Value152.3	19
Function Optimization	Levy D=1000	Convergence Value184.6	19
Function Optimization	Dixon D=1000	Convergence Value1.21e+6	19
Function Optimization	Griewank D=1000	Convergence Value (Statistic)88.753	19

Showing 9 of 9 rows

Other info

Follow for update

@wizwand_team Discord