De-Concentration Phenomena in High Dimensions — Epoche C1
What concentration actually asserts Draw a point at random from the standard Gaussian distribution in ten thousand dimensions and measure its distance from the origin: the answer will be within about one unit of $100$, virtually every time, and it will be within about one unit of $\sqrt{d}$ in $d$ dimensions no matter how large $d$ becomes. That stubbornness is the concentration of measure phenomenon, and this essay is about where it fails and where it only appears to. The distinction matters because the two examples most often produced as failures — the length of a random vector and the maximum of its coordinates — are in fact the phenomenon's flagship successes, while the genuine failures have a different and identifiable cause. The precise statement is worth having in front of us, since everything below is a comparison against it. Call a function $f:\mathbb{R}^d \to \mathbb{R}$ $L$-Lipschitz if moving the input by a distance $r$ never moves the output by more than $Lr$; equivalently, where $f$ is differentiable, its gradient has length at most $L$. For $X$ a standard Gaussian vector — coordinates independent, each with mean zero and variance one — and $f$ any $L$-Lipschitz function, $$P\big(|f(X) - E f(X)| \ge t\big) \le 2\,e^{-t^2/(2L^2)}.$$ Read the right-hand side carefully: the dimension $d$ does not appear. This is the content of the theorem, presented in Ledoux (2001) and in Vershynin (2018). A Lipschitz function of a high-dimensional Gaussian has fluctuations of size $L$, an amount fixed by the function's own steepness and not by the size of the space it lives in. Anything called de-concentration must therefore be a failure of one of the theorem's two conditions — Gaussian coordinates, or a Lipschitz constant that does not grow — and identifying which condition breaks is the whole diagnosis. The length of a random vector: concentration, not its failure The first example usually offered as evidence of de-concentration is the Euclidean norm, and it needs to be corrected rather than defended, because the calculation that is supposed to establish de-concentration is done on the wrong quantity. Let $X = (X_1, \dots, X_d)$ with each $X_i$ an independent standard Gaussian. Then $\|X\|^2 = \sum_i X_i^2$ has a chi-squared distribution with $d$ degrees of freedom. Its mean is $d$ because $E[X_i^2] = 1$ for each of the $d$ terms. Its variance is $2d$ because $\operatorname{Var}(X_i^2) = E[X_i^4] - (E[X_i^2])^2 = 3 - 1 = 2$, the fourth moment of a standard Gaussian being $3$, and because independent terms add their variances. So the squared norm is $d \pm \sqrt{2d}$: at $d = 100$, a mean of $100$ and a standard deviation of about $14.1$. The compressed version of this argument stops here and reports the growing absolute fluctuation as evidence of weak concentration. But the geometrically meaningful quantity is the norm, not its square, and taking the square root changes the answer entirely. If $Y$ has mean $\mu$ and small relative fluctuation, then $\sqrt{Y}$ has standard deviation approximately $\operatorname{sd}(Y)/(2\sqrt{\mu})$, since the derivative of the square root at $\mu$ is $1/(2\sqrt{\mu})$. Here that gives $$\operatorname{sd}(\|X\|) \approx \frac{\sqrt{2d}}{2\sqrt{d}} = \frac{1}{\sqrt{2}} \approx 0.707,$$ independent of $d$. At $d = 100$ the norm sits at about $10 \pm 0.71$; at $d = 10^4$ it sits at about $100 \pm 0.71$. The absolute spread does not grow at all, and the relative spread falls from seven per cent to seven parts in a thousand. This is not a coincidence of the Gaussian case: the map $x \mapsto \|x\|$ is $1$-Lipschitz by the triangle inequality, so the theorem above applies with $L=1$ and predicts exactly this dimension-free fluctuation. The thin-shell picture — that in high dimensions almost all Gaussian mass lies in a thin annulus at radius $\sqrt{d}$ — is the standard illustration of concentration, and it should not be filed under de-concentration. The maximum coordinate: a moving centre is not a spreading distribution The second example needs the same correction, and here the compressed claim is not merely misfiled but false as stated. Let $M_d = \max_{i \le d} X_i$ for the same independent standard Gaussians. Its typical size follows from a count. Since the Gaussian tail satisfies $P(Z \ge t) \approx e^{-t^2/2}/(t\sqrt{2\pi})$ for large $t$, the expected number of coordinates exceeding $t$ is about $d\,e^{-t^2/2}$ up to the polynomial factor; setting that expected number to one and solving for $t$ gives $$M_d \approx \sqrt{2\ln d},$$ which is the value the essay's original text reports, and it is right. The exponent $2$ inside the root comes directly from the $t^2/2$ in the Gaussian exponent, and the logarithm from inverting an exponential tail — which is why the maximum grows so slowly: going from a thousand coordinates to a million raises it only from about $3.7$ to about $5.3$. The claim that follows in the original — that the distribution of $M_d$ spreads out as $d$ increases, rather than tightening — is incorrect, and the reason is instructive. The function $x \mapsto \max_i x_i$ is also $1$-Lipschitz, since changing the input vector by a distance $r$ cannot change any coordinate by more than $r$ and so cannot change their maximum by more than $r$. The theorem therefore applies unchanged, and the fluctuation of $M_d$ about its mean is bounded by a constant for every $d$. In fact the extreme-value limit is sharper still: $\sqrt{2\ln d}\,(M_d - b_d)$ converges to a Gumbel distribution for a suitable centring $b_d$, and since the Gumbel law has standard deviation $\pi/\sqrt{6} \approx 1.28$, the standard deviation of $M_d$ itself is about $1.28/\sqrt{2\ln d}$ — roughly $0.42$ at $d = 100$ and $0.24$ at $d = 10^6$. The maximum concentrates better as the dimension grows. What is true is that $M_d$ does not converge to a fixed number: its centre drifts upwards as $\sqrt{2\ln d}$. That is a statement about the location of the distribution, not its width, and conf