Research & Papers

Infinite vs. finite width: New theorem proves learnability equivalence in Bayesian neural networks

Infinite-width limits don't give neural networks unfair generalization advantages, new proof shows

Deep Dive

A new paper by Dmitry Vaintrob and Kaarel Hänni tackles a foundational question in deep learning theory: Do infinite-width neural networks have fundamentally different inductive biases than their finite-width counterparts? While infinite-width limits (like neural tangent kernels and mean-field models) are widely used for theoretical analysis, the authors show that at the critical mean-field (feature-learning) scaling, the difference is at most polynomial. Their main result is a width-robust learnability theorem for Bayesian neural networks with fixed depth: a family of Boolean-cube targets is learnable from polynomially many samples at infinite width if and only if its reduced entropy — the intensive prior cost of representing a target function to a given accuracy — is polynomially bounded.

Crucially, this equivalence holds both ways: if a target is learnable with a finite-width network, it is also learnable in the infinite-width limit, and vice versa. The proof introduces a subsampling method that selects polynomially many hidden neurons from the infinite-width posterior. These neurons split into an "active" component (capturing data-dependent low-dimensional statistics) and a "lazy" component (resampling entropy-dominated directions from the prior). This constructive argument shows that infinite-width mean-field limits provide a clean analytical description without introducing spurious width-dependent generalization power, bridging a key gap between theory and practice in neural network learning.

Key Points
  • Proves equivalence: infinite-width learnability iff polynomial-width learnability for mean-field Bayesian neural networks on Boolean-cube targets.
  • Key condition is "reduced entropy" — a measure of prior cost for representing target functions to a given MSE.
  • Novel subsampling technique extracts polynomially many active and lazy neurons from infinite-width posterior to preserve learned function on all inputs.

Why It Matters

Validates infinite-width theoretical analyses as faithful proxies for finite neural networks, strengthening foundations of deep learning theory.

📬 Get the top 10 AI stories daily