Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Bernstein-Schur Kernels: Random Features by Sketched Modulation and Radial Randomization

About

Bernstein--Schur kernels are products of a finite-feature kernel and a completely monotone shift-invariant kernel: nonstationary kernels falling between the shift-invariant and dot-product templates random features exploit, so neither Bochner sampling nor polynomial sketching applies to the full kernel directly. We give one random-feature construction for the whole class that randomizes both factors: it sketches the finite modulation and samples the radial factor's one-dimensional Bernstein--Widder scale before applying Gaussian random Fourier features, giving feature dimension $Dm$, free of the $O(d^2)$ size of the exact modulation feature. With the modulation kept exact (the $m\to\infty$ limit), we prove unbiasedness, an exact variance, and a matrix-Bernstein operator-norm bound controlled by the top kernel and modulation eigenvalues and an intrinsic dimension rather than the crude $N\max_{ij}$ route. Whitening this argument at the ridge makes the effective dimension $d_{\mathrm{eff}}(\lambda)$ the \emph{exact} intrinsic dimension of the matrix variance, so $O((1+\|P\|_{\mathrm{op}}/\lambda)\log(d_{\mathrm{eff}}/\delta))$ radial draws preserve the kernel-ridge solution; tilting the draw by a closed-form whitened leverage improves this to the effective-dimension count $O((1+d_{\mathrm{eff}})\log(d_{\mathrm{eff}}/\delta))$. Conditioning on the sketch carries every guarantee to the deployed doubly-randomized estimator up to one additive sketch term, and all hold for the whole class with the modulation Gram in place of the polynomial one. The flagship instance is the biased $yat$-kernel $k_{yat,b}(w,x)=(w^\top x+b)^2/(\|w-x\|^2+\varepsilon)$, whose family span contains the inverse-multiquadric kernel by finite differences in $b$.

Taha Bouhsine• 2026

Related benchmarks

TaskDatasetResultRank
ClassificationHiggs (test)
AUC74.9
21
Gram matrix approximationOff-sphere bounded ball (||x|| ∈ [0.25, 1], N=1000)
Relative Frobenius Gram Error0.00e+0
20
Kernel ApproximationUnit sphere
Samples Required200
12
Digit ClassificationSpherical MNIST (test)
Accuracy97.9
9
RegressionSynthetic Coupled Target off-sphere (held-out split)
RMSE0.034
6
ClassificationDigits off-sphere (held-out split)
Accuracy97.9
6
Classificationsphere-normalized Digits d=64 (3 splits)
Accuracy98.6
5
Gram matrix approximationUnit sphere d=5 N=1000, b=1, epsilon=1 (train test)
Relative Frobenius Error0.045
5
Gram matrix approximationUnit sphere d=10 N=1000, b=1, epsilon=1 (train test)
Relative Frobenius Error6.3
5
Gram matrix approximationUnit sphere d=20 N=1000, b=1, epsilon=1 (train test)
Relative Frobenius Error0.073
5
Showing 10 of 12 rows

Other info

Follow for update